Tag

Gemini

3 briefings.

Google’s HLE-Verified chart compares Gemini 3.8 Flash at 54.9% with earlier Flash and competing models.
Research chart© Google DeepMindEditorial excerptOriginal storyGoogle’s HLE-Verified comparison for Gemini 3.8 Flash, from its model page. Vendor-reported reasoning scores do not measure reliability on a complete agent workflow.

google · gemini

Gemini 3.8 Flash moves Google's workhorse model toward long-running agents

Google's newest Flash model combines a one-million-token input window, multimodal input, computer use, and a larger emphasis on sustained agent work.

2 sources
Still from Google’s demo labeled UI controlled by Gemini 3.5 Flash, showing the Gemini app being explored through its mobile environment.
Product screenshot© GoogleEditorial excerptOriginal storyA still from Google’s Gemini 3.5 Flash computer-use demo. Google says the model explored the app across 73 turns; the demo shortens the sequence.

google · gemini

Gemini 3.5 Flash gets computer use as a built-in tool

Google moved computer control into its main Flash model and added confirmation and prompt-injection safeguards for enterprise agents.

2 sources