Signal archive

Every briefing.
Newest signal first.

Concise reporting on models, agents, policy, security, and the business of AI.

OpenAI ExploitBench chart reports Astra at 100% capability coverage and GPT-5.6 Sol at 78.5%, plotted against output tokens.
Research chart© OpenAIEditorial excerptOriginal storyOpenAI’s ExploitBench results without production safeguards, from Astra’s safety report. Company-reported evaluation; not an independent test or the unrestricted capability of the public release.

openai · models

OpenAI starts GPT-6 Astra rollout with critical cyber capability gated

OpenAI's September 5 Astra launch combines a new computer-use flagship with restricted exploit generation and a staged expansion through Daybreak.

4 sources
Google’s HLE-Verified chart compares Gemini 3.8 Flash at 54.9% with earlier Flash and competing models.
Research chart© Google DeepMindEditorial excerptOriginal storyGoogle’s HLE-Verified comparison for Gemini 3.8 Flash, from its model page. Vendor-reported reasoning scores do not measure reliability on a complete agent workflow.

google · gemini

Gemini 3.8 Flash moves Google's workhorse model toward long-running agents

Google's newest Flash model combines a one-million-token input window, multimodal input, computer use, and a larger emphasis on sustained agent work.

2 sources
OpenAI’s ChatGPT Voice launch photo shows three women sharing a phone, with the labels The all-new ChatGPT Voice and Powered by GPT-Live-1.
Source photo© OpenAIEditorial excerptOriginal storyOpenAI’s campaign photo for the GPT-Live-powered ChatGPT Voice launch. Promotional imagery from the announcement; not an independent test of conversational performance.

openai · voice

GPT-Live turns voice AI from walkie-talkie into conversation

OpenAI's new full-duplex voice models can listen and speak simultaneously, while delegating harder work to a frontier model in the background.

2 sources
OpenRouter stacked weekly token-volume chart shows Chinese-author models overtaking American-author models in early June 2026.
Research chart© OpenRouterEditorial excerptOriginal storyOpenRouter’s weekly token volume by model-author country through June 14, 2026. This measures OpenRouter traffic, not the whole AI market or user nationality.

open-weights · glm-5.2

Chinese open-weight models are winning more AI workloads

OpenRouter data shows Chinese models overtook US models in token share in early June, driven by capable, lower-cost open weights.

4 sources
Still from Google’s demo labeled UI controlled by Gemini 3.5 Flash, showing the Gemini app being explored through its mobile environment.
Product screenshot© GoogleEditorial excerptOriginal storyA still from Google’s Gemini 3.5 Flash computer-use demo. Google says the model explored the app across 73 turns; the demo shortens the sequence.

google · gemini

Gemini 3.5 Flash gets computer use as a built-in tool

Google moved computer control into its main Flash model and added confirmation and prompt-injection safeguards for enterprise agents.

2 sources
OpenAI Codex Security screenshot shows a SQL-injection finding, a proposed patch and the Apply patch locally action.
Product screenshot© OpenAIEditorial excerptOriginal storyOpenAI’s Daybreak announcement shows Codex Security moving from a finding to a proposed patch for local review. This is the publisher’s demo, not a BigBrainZ test.

openai · security

OpenAI's Daybreak shifts cyber AI from finding bugs to landing patches

Daybreak combines Codex Security, GPT-5.5-Cyber, and an open-source maintainer program around the harder half of security work: fixing what models find.

2 sources