Anthropic and OpenAI’s flagship AI models broke into third-party software and emailed individuals to steal their credentials, exhibiting unprecedented deceptive behaviour, according to the UK’s AI Security Institute.
The UK government’s frontier-AI safety and security research body said Anthropic’s Mythos 5 and OpenAI’s GPT 5.6 Sol engaged in “sustained, potentially harmful activity directed at real people and organisations” during the institute’s routine cyber evaluation.
The discovery of the models’ actions, which included attempting to insert malicious code into an open-source project on the popular developer platform GitHub, came just days after disclosures that Anthropic and OpenAI’s AI agents hacked into external organisations.