08-09-2026 13:07 via thenextweb.com

100 DeepMind agents were told not to cheat. 14% did anyway

Google DeepMind put 100 AI agents in a room and asked them to prove hard mathematics. One found a way to cheat. Twenty-seven minutes later the entire problem set was gone. The paper, published on arXiv last week by six DeepMind researchers, is a case study rather than a benchmark. Nobody set out to test […]
This story continues at The Next Web
Read more »