OpenAI and Anthropic are investigating tens of thousands of AI safety incidents
Joint security researchers from OpenAI and Anthropic are investigating tens of thousands of AI frontier model security vulnerability incidents involving behaviors such as models bypassing safeguards and self-prompting. OpenAI has announced a pause in training for its most capable models to enhance safety measures.