Google held its I/O 2026 developer keynote on May 19, 2026, with CEO Sundar Pichai declaring that the company is "firmly in our agentic Gemini era." The phrase set the tone for a keynote built less around individual chatbot features and more around AI systems that complete multi-step tasks with limited supervision. Pichai's own numbers illustrate how fast the underlying products have grown: the Gemini app now has 900 million monthly active users, more than double the 400 million Google reported a year earlier, while AI Mode in Google Search has separately passed 1 billion monthly active users roughly a year after launch.
A split-chip strategy: TPU 8t and TPU 8i
The most structurally significant announcement wasn't a model at all β it was a chip strategy Google had first unveiled at Cloud Next in April 2026 and revisited at I/O. For the first time, Google's eighth-generation Tensor Processing Units split into two distinct designs rather than one general-purpose chip. TPU 8t is built for large-scale training and scales up to 9,600 chips in a single "superpod," delivering nearly three times the raw compute of the previous generation, according to Google's own technical documentation. TPU 8i is the inference-focused counterpart, connecting up to 1,152 chips per pod through a new network layout called Boardfly and packing three times the on-chip memory (SRAM) of its predecessor β prioritizing the low latency that matters when an AI agent, not a human, is the one waiting on a response. Google says both chips deliver up to twice the performance-per-watt of the previous generation and will reach cloud customers "soon," without committing to a firm general-availability date.
Gemini 3.5 Flash: cheaper, faster, and built for "long-horizon" tasks
Gemini 3.5 Flash, the first release in the 3.5 series, is now generally available in the Gemini app, Google Search, Google AI Studio, Android Studio, and via the Gemini API. Google says it outperforms the earlier Gemini 3.1 Pro on a set of agentic and coding benchmarks β including Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%) β while running at roughly four times the output-token speed of comparable frontier models and at less than half the typical cost. A heavier sibling, Gemini 3.5 Pro, was still in internal testing at I/O and was expected within weeks. Alongside the models, Google cut the price of its top-tier AI Ultra subscription from $250 to $200 a month and introduced a new $100-a-month tier beneath it, signaling a push to widen the paying user base rather than protect margins on existing subscribers.
New agents: Google Flow, Daily Brief, Spark, and Omni
Beyond the models, Google unveiled several consumer-facing "out-of-the-box" agents. Google Flow, previously a generative-video creation tool, gained an agent that can plan and reason through multi-step creative projects β including letting users "vibe code" custom tools for effects like hand-drawn animation, under the user's ongoing direction rather than fully autonomously. Daily Brief is a new Gemini app feature that synthesizes a user's inbox, calendar, and tasks into a short morning digest, and β Google says β actively suggests next steps rather than only summarizing. Google also introduced Gemini Spark, a "24/7" personal agent built on Gemini 3.5 Flash that keeps working in the background β Google says even while a user's phone or laptop is turned off β and is designed to check in with the user before taking major actions on their behalf; it's rolling out to trusted testers first, with a beta planned for U.S. AI Ultra subscribers the following week. Alongside it, Google unveiled Gemini Omni, a new multimodal model the company frames as a step toward deeper "world understanding": it can take text, image, audio, or video as a reference and turn it into a new output, starting with video generation and editing.
Our take
The headline-grabbing chip specs matter less to most readers than what they enable: Google is betting that owning both the training silicon (TPU 8t) and the inference silicon (TPU 8i) lets it offer agentic AI more cheaply than rivals that lease GPU capacity from third parties. That's the real subtext behind Gemini 3.5 Flash's aggressive pricing and the AI Ultra discount β Google can absorb thinner margins on individual queries because it controls the full stack. Whether Daily Brief and Spark actually save people time, versus adding another app-and-permissions layer to manage, is the open question that benchmarks can't answer; that will only be settled by how these agents perform once ordinary users β not keynote demos β start relying on them.
What to watch next
Gemini 3.5 Pro's release, which Google said at I/O was coming "next month," will show whether Google can pair Flash's speed with genuinely deeper reasoning. It's also worth watching whether Gemini Spark's early third-party integrations β Google named Canva, Dropbox, and Instacart as launch partners via the Model Context Protocol β expand into the much larger roster the company has promised for later in the summer, and whether the TPU 8t/8i rollout measurably narrows Google Cloud's price gap with Nvidia-based competitors.
π‘ The Nexalytics Take
Google is shifting gears from simple chatbots to full "agents." By unveiling the lightning-fast Gemini 3.5 Flash and the specialized TPU 8t chip, Google is building an ecosystem where AI doesn't just answer questionsβit actively manages your daily workflow. The real story is the vertical integration: Google's control over both training and inference silicon is what lets it undercut rivals on price while pushing agents like Spark and Daily Brief into daily use.