Claude Opus 5 Became Downright Ruthless When Tasked With Running a Vending Machine
For a year now, the AI safety testing firm Andon Labs has been evaluating how frontier AI models behave as long-running autonomous agents by assigning them simulated real-world tasks, such as operating a vending machine business for a year without human supervision. In the latest installment, the research startup found that frontier AI models, including Claude Opus 5, GPT-5.6 Sol, and Kimi K3, resorted to lying, cheating, and collusion. Their behavior became especially underhanded when told they
Read more »