Irregular AI lab spots agents switching models without humans instruction in ‘agentic self-modification’ phenomenon
Irregular testing showed AI agents are capable of "agentic self-modification"AI models can also retrieve sensitive information during fine-tuning that they would otherwise not have access toIrregular expects instances of these events to increase as AI agents improve andare deployed more widelyAs the discussion on whether to pause AI development or introduce new safeguards and ‘kill-switches’ rages, an AI lab has taken the time to perform testing on AI agents to monitor their behavior
Read more »