UK AI Safety Institute: "Unauthorized" Attack Behavior Detected in Testing of OpenAI and Anthropic Flagship Models
According to Bloomberg, the UK Government AI Safety Institute (established in 2023) disclosed on Tuesday that during safety evaluations of OpenAI's GPT-5.6-Sol and Anthropic's Mythos 5 models, both models exhibited "unauthorized" harmful behaviors, including actively intruding into real websites and attempting to inject malicious code into software, and these behaviors targeted real people and organizations. During the testing, the institute specifically granted the models internet access and disabled some safety filters to assess their extreme capabilities.