BREAKING
πŸ”₯ Microsoft cuts 3,200 Xbox jobs, divests four studios in historic reset  |  Anthropic in early talks with Samsung to build custom 2nm AI chip  |  SpaceX joins Nasdaq-100 on July 7, unlocking wave of passive money  |  Qualcomm's Dragonfly C1000 lands Meta data center deal for 2028  |  Samsung rolls out ChatGPT Enterprise & Codex to workers worldwide  |  Meta bets ~$900M on Cred, Kunal Shah to lead WhatsApp globally  |  OpenAI latest model GPT-5.5  |  Starlink hits 10-Gigabit speeds in global beta  |  Nvidia's 'Rubin' GPUs Promise 4x Efficiency Jump  |  Generative UI frameworks end static web design  |  Hackers manipulate chatbot to steal 20,000 Instagram accounts  |  McDonald's tests Google-backed AI drive-thrus
Google I/O 2026: The Agentic Gemini Era and TPU 8t AI
June 3, 2026 5 min read

Google I/O 2026: The Agentic Gemini Era and TPU 8t

N
Nexalytics Tech Editorial Team Reporting & analysis by our staff
⚑ Short on time? Jump to The Nexalytics Take for a quick summary.

Google held its I/O 2026 developer keynote on May 19, 2026, with CEO Sundar Pichai declaring that the company is "firmly in our agentic Gemini era." The phrase set the tone for a keynote built less around individual chatbot features and more around AI systems that complete multi-step tasks with limited supervision. Pichai's own numbers illustrate how fast the underlying products have grown: the Gemini app now has 900 million monthly active users, more than double the 400 million Google reported a year earlier, while AI Mode in Google Search has separately passed 1 billion monthly active users roughly a year after launch.

A split-chip strategy: TPU 8t and TPU 8i

The most structurally significant announcement wasn't a model at all β€” it was a chip strategy Google had first unveiled at Cloud Next in April 2026 and revisited at I/O. For the first time, Google's eighth-generation Tensor Processing Units split into two distinct designs rather than one general-purpose chip. TPU 8t is built for large-scale training and scales up to 9,600 chips in a single "superpod," delivering nearly three times the raw compute of the previous generation, according to Google's own technical documentation. TPU 8i is the inference-focused counterpart, connecting up to 1,152 chips per pod through a new network layout called Boardfly and packing three times the on-chip memory (SRAM) of its predecessor β€” prioritizing the low latency that matters when an AI agent, not a human, is the one waiting on a response. Google says both chips deliver up to twice the performance-per-watt of the previous generation and will reach cloud customers "soon," without committing to a firm general-availability date.

Gemini 3.5 Flash: cheaper, faster, and built for "long-horizon" tasks

Gemini 3.5 Flash, the first release in the 3.5 series, is now generally available in the Gemini app, Google Search, Google AI Studio, Android Studio, and via the Gemini API. Google says it outperforms the earlier Gemini 3.1 Pro on a set of agentic and coding benchmarks β€” including Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo), and MCP Atlas (83.6%) β€” while running at roughly four times the output-token speed of comparable frontier models and at less than half the typical cost. A heavier sibling, Gemini 3.5 Pro, was still in internal testing at I/O and was expected within weeks. Alongside the models, Google cut the price of its top-tier AI Ultra subscription from $250 to $200 a month and introduced a new $100-a-month tier beneath it, signaling a push to widen the paying user base rather than protect margins on existing subscribers.

New agents: Google Flow, Daily Brief, Spark, and Omni

Beyond the models, Google unveiled several consumer-facing "out-of-the-box" agents. Google Flow, previously a generative-video creation tool, gained an agent that can plan and reason through multi-step creative projects β€” including letting users "vibe code" custom tools for effects like hand-drawn animation, under the user's ongoing direction rather than fully autonomously. Daily Brief is a new Gemini app feature that synthesizes a user's inbox, calendar, and tasks into a short morning digest, and β€” Google says β€” actively suggests next steps rather than only summarizing. Google also introduced Gemini Spark, a "24/7" personal agent built on Gemini 3.5 Flash that keeps working in the background β€” Google says even while a user's phone or laptop is turned off β€” and is designed to check in with the user before taking major actions on their behalf; it's rolling out to trusted testers first, with a beta planned for U.S. AI Ultra subscribers the following week. Alongside it, Google unveiled Gemini Omni, a new multimodal model the company frames as a step toward deeper "world understanding": it can take text, image, audio, or video as a reference and turn it into a new output, starting with video generation and editing.

Our take

The headline-grabbing chip specs matter less to most readers than what they enable: Google is betting that owning both the training silicon (TPU 8t) and the inference silicon (TPU 8i) lets it offer agentic AI more cheaply than rivals that lease GPU capacity from third parties. That's the real subtext behind Gemini 3.5 Flash's aggressive pricing and the AI Ultra discount β€” Google can absorb thinner margins on individual queries because it controls the full stack. Whether Daily Brief and Spark actually save people time, versus adding another app-and-permissions layer to manage, is the open question that benchmarks can't answer; that will only be settled by how these agents perform once ordinary users β€” not keynote demos β€” start relying on them.

What to watch next

Gemini 3.5 Pro's release, which Google said at I/O was coming "next month," will show whether Google can pair Flash's speed with genuinely deeper reasoning. It's also worth watching whether Gemini Spark's early third-party integrations β€” Google named Canva, Dropbox, and Instacart as launch partners via the Model Context Protocol β€” expand into the much larger roster the company has promised for later in the summer, and whether the TPU 8t/8i rollout measurably narrows Google Cloud's price gap with Nvidia-based competitors.

πŸ’‘ The Nexalytics Take

Google is shifting gears from simple chatbots to full "agents." By unveiling the lightning-fast Gemini 3.5 Flash and the specialized TPU 8t chip, Google is building an ecosystem where AI doesn't just answer questionsβ€”it actively manages your daily workflow. The real story is the vertical integration: Google's control over both training and inference silicon is what lets it undercut rivals on price while pushing agents like Spark and Daily Brief into daily use.

Share: