Google I/O 2026: Gemini 3.5 Flash, multimodal video, and agent infra as standard
At Google I/O 2026, Google shipped Gemini 3.5 Flash — which outperforms the previous Pro model on agentic evals at $1.50/M — plus multimodal Omni and an API that hands you remote Linux per call. Agent infra is no longer a differentiator.
Summary
At Google I/O on May 19, Google launched four products at once: Gemini 3.5 Flash, the multimodal Omni model, Spark agents, and the Managed Agents API. Flash is priced at $1.50/M input and $9/M output, supports 1M token context, and scored 76.2% on Terminal-Bench 2.1 — outperforming Gemini 3.1 Pro on agentic benchmarks. Omni takes image, audio, video, and text as simultaneous input. The Managed Agents API hands you a remote Linux environment per API call.
In practice
Gemini 3.5 Flash becomes the default option for cost-sensitive workloads needing long context — it beats the previous Pro model on agents at a fraction of the price. The Managed Agents API, alongside equivalent launches from Cursor and Anthropic in the same week, signals that the three largest providers are no longer just selling models: they're selling agent runtimes. Teams still running their own orchestration glue are paying maintenance on a problem the market has already solved.
Context
I/O happened in a week dominated by DeepSeek's price cut. Google's answer wasn't price — it was product. Flash 3.5 isn't a low-cost model with reduced capabilities: it competes at the top on agentic evals at a price most builders can absorb. Omni complements Veo 3 (Google's video generation model) with multimodal understanding. Midjourney launched its V1 video model the same week, at roughly 25× cheaper per second than direct rivals.
Why it matters
- Gemini 3.5 Flash outperforms Gemini 3.1 Pro on agentic evals at $1.50/M — new cost-tier reference point
- Managed Agents API brings remote Linux to a single API call; Cursor and Anthropic took the same step in the same week
- Agent infra is now a baseline expectation across all three major providers, not a differentiator
- Omni pairs multimodal understanding with the generation ecosystem (Veo 3) in an integrated stack