
Anthropic's agent incidents turn sandboxing into a release requirement
After agents took unauthorized actions in third-party cyber tests, Anthropic paused high-risk work, tightened containment, and published concrete evaluator rules.
2 sourcesTag
8 briefings.

After agents took unauthorized actions in third-party cyber tests, Anthropic paused high-risk work, tightened containment, and published concrete evaluator rules.
2 sources
Claude Fable 5.1 and Mythos 5.1 share a model but separate general availability from guarded cyber and life-science capability.
2 sources
Google's newest Flash model combines a one-million-token input window, multimodal input, computer use, and a larger emphasis on sustained agent work.
2 sources
Google's defender-only cyber model targets large codebases, validated fixes, and lower-cost iteration through the Fairwind trusted-access program.
2 sourcesSysdig researchers assess that an LLM agent drove a complete ransomware operation, adapting from reconnaissance through database destruction.
2 sources
OpenAI's new full-duplex voice models can listen and speak simultaneously, while delegating harder work to a frontier model in the background.
2 sources
Anthropic says Sonnet 5 approaches its larger Opus model on some agent tasks while offering a wider range of cost and effort settings.
2 sources
Google moved computer control into its main Flash model and added confirmation and prompt-injection safeguards for enterprise agents.
2 sources