Summary
AI agents, like genies, interpret instructions literally, potentially leading to disastrous outcomes, as demonstrated by a recent Hugging Face hack attributed to an OpenAI GPT model. Experts Bruce Schneier and Barath Raghavan argue for new measurement methods to ensure AI aligns with human intent.