Command Palette

Search for a command to run...

Google

Google Cloud's AI model-serving platform — spanning the Gemini API, AI Studio, and the Gemini Enterprise Agent Platform (formerly Vertex AI) — is Alphabet's full-stack bet to be the enterprise AI infrastructure layer of the agentic era, combining custom TPU silicon, frontier Gemini models, and a 200+ model catalog into a single pay-as-you-go cloud platform.

Intelligence over time

Google Cloud AI

  Executive Briefing

Google Cloud AI — operating today through three product surfaces: the Gemini API / Google AI Studio, the Gemini Enterprise Agent Platform (rebranded from Vertex AI at Google Cloud Next in April 2026), and the Model Garden — is Alphabet's end-to-end attempt to own the AI infrastructure stack from custom silicon to enterprise workflow. Launched in its current unified form as Vertex AI on 18 May 2021, the platform combines Google's proprietary Tensor Processing Units (TPUs), the frontier Gemini model family from Google DeepMind, and a curated catalog of more than 200 models — including third-party offerings from Anthropic, Meta, and Mistral AI — into a single consumption-based cloud service.1

The AI unit sits at the intersection of three Alphabet organizations: Google Cloud, led by CEO Thomas Kurian (who joined from Oracle in 2019 and is the principal commercial steward of the platform); Google DeepMind, led by Nobel laureate CEO Demis Hassabis (who co-founded DeepMind in London in 2010 and merged it with Google Brain in April 2023 to form the research engine behind all Gemini models); and Alphabet itself, overseen by CEO Sundar Pichai. Day-to-day AI product direction is further shaped by Jeff Dean (Chief Scientist, Google DeepMind, and co-founder of Google Brain in 2011), Koray Kavukcuoglu (CTO, Google DeepMind), and Sissie Hsiao (VP, Gemini AI Experiences).2

Financially, Google Cloud is the fastest-growing of the three major hyperscalers: Q4 2025 revenue reached $17.7 billion (up 48% year-over-year), with a cloud order backlog reported at $240 billion (more than doubled year-over-year).3 Planned capital expenditure for 2026 is reported at $175–185 billion, nearly double the estimated $91–93 billion deployed in 2025, signaling an infrastructure build-out at a scale unmatched in the company's history.4 Google has stated that it reduced the per-query cost of serving Gemini by approximately 78% during 2025 through TPU optimizations and data-center efficiency improvements — a figure that, if sustained, changes the unit economics of frontier inference at enterprise scale.

The April 2026 rename from Vertex AI to Gemini Enterprise Agent Platform is more than a marketing refresh: it signals a strategic pivot from ML operations infrastructure (training pipelines, model registries, batch serving) to end-to-end agentic workflow orchestration as the platform's primary value proposition. With the Agent2Agent (A2A) protocol in production at 150 organizations, a generally available Agent Development Kit (ADK v1.0), managed MCP servers via Apigee, and a Workspace Studio no-code agent builder embedded into Gmail and Google Docs, Google is betting that vertical integration — custom chips to frontier models to consumer distribution — gives it a structural moat that neither pure cloud providers nor pure AI labs can easily replicate.

  At a Glance

ItemDetail
Platform GA (Vertex AI)18 May 2021
TypeHyperscaler — AI model-serving and agentic platform
HeadquartersMountain View, CA (Alphabet / Google); London, UK (Google DeepMind); Sunnyvale, CA (Google Cloud)
StatusActive
LeadershipThomas Kurian (CEO, Google Cloud), Demis Hassabis (CEO, Google DeepMind), Sundar Pichai (CEO, Alphabet)
Parent / ownershipAlphabet Inc. (NASDAQ: GOOGL / GOOG)
SpecialtiesCustom TPU silicon, Gemini model family, 200+ model catalog, agentic AI platform, Workspace AI
Google Cloud Q4 2025 revenue$17.7B (+48% YoY, reported)3
Cloud backlog (end Q4 2025)$240B (reported, more than doubled YoY)3
2025 capex~$91–93B (reported/estimated)4
2026 planned capex$175–185B (reported)4

  Origins and Founding

Google's cloud AI platform has roots in two founding acts separated by a year and an ocean. In 2010, Demis Hassabis, Shane Legg, and Mustafa Suleyman co-founded DeepMind in London as an independent AI safety and research lab. In 2011, Jeff Dean, Greg Corrado, and Andrew Ng founded Google Brain as an internal research group at Google, pioneering large-scale neural network training on Google's compute infrastructure and publishing landmark work including the 2012 "Google cat" unsupervised learning paper. Google acquired DeepMind in January 2014 for a reported ~$500 million, keeping it as a semi-autonomous unit in London.5

The original product thesis for a managed ML platform emerged from Google's frustration with the friction its own ML teams experienced moving from research to production. Engineers building products on TensorFlow needed to stitch together separate data pipelines, training clusters, serving infrastructure, and monitoring dashboards. Cloud ML Engine, launched in 2017, offered managed training and serving. Cloud AutoML (2018) extended no-code model training to domain experts without ML backgrounds. The core insight was that the entire pipeline — data preparation, training, serving, and monitoring — belonged on a unified managed surface, not as a collection of separate services.1

On 18 May 2021, Google announced Vertex AI at Google I/O, consolidating AutoML and Cloud ML Engine into a single MLOps platform and establishing the product architecture still in use today. In April 2023, the merger of Google Brain and DeepMind into Google DeepMind under Hassabis's leadership created the unified research organization responsible for the Gemini model family — the models that have since become the commercial engine of the platform.5

  History and Timeline

    2010–2020: Research foundations and the TPU era

Google's two AI research arms — Brain and DeepMind — operated largely in parallel through the 2010s, producing transformative research (the transformer architecture paper "Attention Is All You Need" emerged from Google Brain researchers in 2017) while Google's cloud product team assembled the first generation of managed ML services. The TPU program began in 2015 with the first-generation chip deployed in Google data centers; by 2018, TPU v3 was available to Cloud customers, establishing custom silicon as a Google Cloud differentiator long before it was a competitive necessity.

    2021–2022: Vertex AI unification

The May 2021 Vertex AI launch was immediately followed by rapid catalog expansion. By mid-2022 the platform supported custom training on GPUs and TPUs, AutoML for tabular and vision tasks, and a nascent model registry. The generative AI wave had not yet arrived in the cloud product, but the infrastructure foundation was laid.

    2023: Generative AI arrives on Vertex

June 2023 brought generative AI support to general availability on Vertex AI, including PaLM 2, Imagen, and Codey (code generation), with roughly 60 models in the Model Garden. By August 2023 the catalog had expanded to include third-party open models: Claude 2 and Llama 2 were added, establishing the multi-model philosophy that now defines the platform. Vertex AI Extensions launched to support tool use and API integrations. The April 2023 merger of Google Brain and DeepMind into Google DeepMind under Demis Hassabis created the unified research organization that would produce all subsequent Gemini models.5

In December 2023, Google announced Gemini 1.0 in three sizes (Ultra, Pro, Nano), marking the transition from PaLM 2 to Gemini as the platform's flagship model family. Vertex AI Agent Builder debuted in April 2024 as a no-code conversational agent tool, and Gemini 1.5 expanded context windows dramatically.

    2024–2025: Gemini 2.x and the reasoning turn

Gemini 2.0 Flash launched in December 2024 with native tool integration, and Gemini 2.5 Pro followed in March 2025 as an experimental reasoning model. At Google I/O 2025 (May 20), Google announced the seventh-generation TPU — Ironwood (TPU v7) — as the first purpose-built inference TPU, alongside Veo 3 (video generation), Imagen 4, and the Gemma 3n open model family. Gemini 2.5 Flash reached general availability in June 2025. In October 2025, Anthropic announced an expansion of its Google Cloud TPU commitment to up to one million chips — a figure referencing chip-equivalents over a multi-year period, with exact terms not disclosed.6

Gemini 3 Pro previewed in November 2025, followed by Gemini 3 Flash GA in December 2025 and Gemini 3.1 Pro in February 2026. Alphabet's Q4 2025 earnings, reported February 4 2026, recorded Google Cloud revenue of $17.7B (+48% YoY) and a $240B backlog.3

    2026: The agentic platform era

Google Cloud Next 2026 (April 22, Las Vegas) was the platform's most significant product event to date. The centerpiece was the rename of Vertex AI to the Gemini Enterprise Agent Platform, reflecting the strategic shift toward agentic AI orchestration. Google simultaneously announced TPU 8t (training, co-designed with Broadcom) and TPU 8i (inference, co-designed with MediaTek) as bifurcated silicon for distinct workloads; the A2A protocol v1.0 in production at 150 organizations; Workspace Studio (no-code agent builder for Gmail, Docs, and Sheets); and a $750 million partner fund targeting SI-led enterprise deployments across Accenture, Deloitte, KPMG, and others.78

Google I/O 2026 (May 19) added Gemini 3.5 Flash, Gemini Omni Flash, and Antigravity 2.0. On June 9, Gemini 3.5 Live Translate brought near-real-time speech-to-speech translation in 70+ languages.

  What They Offer — Products and Platform

Google Cloud AI operates through three primary product surfaces, each targeting a distinct segment of the developer and enterprise market.

Gemini API / Google AI Studio is the developer-first surface, offering a free tier (up to 500 RPM on Flash models via AI Studio) and pay-as-you-go billing. AI Studio provides a prompt engineering UI, model comparison tools, and direct API key generation. This surface prioritizes low-friction prototyping and is where individual developers and startups typically begin.1

Gemini Enterprise Agent Platform (formerly Vertex AI) is the enterprise-grade cloud service, adding IAM and role-based access control, SLAs, data-privacy guarantees (no training on customer data by default), VPC Service Controls for network isolation, and support for fine-tuning, batch inference, and full MLOps pipelines. Enterprise features include: Vertex AI Pipelines (ML workflow orchestration), Vertex AI Feature Store, Model Registry and Monitoring, Agent Builder / Agent Engine (managed runtime for agentic workloads), and Vertex AI Studio (prompt engineering with enterprise guardrails). An Express Mode provides limited-quota access without billing enablement for evaluation.1

Model Garden is the curated catalog of 200+ models, combining Google's own Gemini and Gemma families with third-party models from Anthropic (Claude 3 family and Claude Fable 5), Meta (Llama), Mistral AI, and a broad range of open-source models. Model Garden enables side-by-side evaluation, one-click deployment, and consistent billing across the full catalog — a deliberate multi-model strategy that reduces customer lock-in risk while capturing more of the inference spend regardless of which model wins a given use case.

Beyond these three surfaces, the platform encompasses: Agent Development Kit (ADK v1.0), generally available in four languages; managed MCP servers via Apigee; the Agent2Agent (A2A) protocol v1.0 for cross-platform agent communication (now a Linux Foundation project); Workspace Studio, a no-code agent builder embedded in Gmail, Docs, and Sheets for building agents that interact with productivity workflows; and Imagen 4 and Veo 3 for image and video generation respectively. The business model is consumption-based — pay-per-token for model APIs, per-node-hour for training, per data point for AutoML forecasting — with enterprise committed-use discounts available.

  Technology and Infrastructure

Google's infrastructure advantage is built on a deliberate strategy of vertical integration: designing its own silicon, training its own frontier models on that silicon, and serving those models through its own cloud — capturing margin and performance improvements at every layer of the stack.

    Custom Silicon: The TPU program

The Tensor Processing Unit (TPU) program began in 2015 and has produced seven generations of custom AI accelerators. Each generation has targeted either training or inference, with the seventh and eighth generations representing the most explicit specialization to date.

Ironwood (TPU v7), announced at Google Cloud Next 2025, was Google's first TPU designed explicitly for high-volume, low-latency inference, delivering 4.6 petaFLOPS per chip. It established the inference-optimized silicon path that TPU 8i continues.9

At Google Cloud Next 2026, Google announced a bifurcated eighth-generation architecture reflecting diverging requirements for training and inference workloads:9

ChipPurposeKey specsCo-designed with
TPU 8tTraining121 FP4 ExaFLOPs per 9,600-chip superpod; 2 PB shared HBM; ~2.8x training price-performance vs. Ironwood; 134,000-chip single data center via Virgo NetworkBroadcom
TPU 8iInference / RL384 MB on-chip SRAM (3x increase); 288 GB HBM; Boardfly topology (56% reduction in network diameter); 80% better inference price-performance vs. IronwoodMediaTek

The Virgo Network underpinning TPU 8t superpods delivers a 4x bandwidth improvement over prior generations, enabling near-linear scaling efficiency across massive clusters. Management reported approximately a 78% reduction in per-Gemini-query serving cost during 2025 through TPU optimizations and data-center efficiency gains.9

    GPU and CPU infrastructure

For customers requiring NVIDIA silicon, Google offers A5X bare-metal instances powered by NVIDIA Vera Rubin NVL72, scaling to 80,000 GPUs per data center and 960,000 GPUs multi-site via the Falcon networking fabric co-engineered with NVIDIA. The Google Axion N4A provides custom ARM-based CPU VMs. Storage is served by Google Cloud Managed Lustre at 10 TB/s throughput and Rapid Buckets with sub-millisecond latency and 20 million operations per second.9

    Software stack and inference optimizations

The software stack runs JAX natively (the primary framework for Gemini model training), with PyTorch / TorchTPU in preview. Inference serving supports vLLM and llm-d (a CNCF sandbox project). Google Kubernetes Engine (GKE) includes an AI-powered Inference Gateway that Google claims delivers 70%+ reduction in time-to-first-token (TTFT) latency. The full hardware and software environment is branded the AI Hypercomputer, representing the integration of TPUs, GPUs, Axion CPUs, Virgo networking, and the software stack into a heterogeneous compute environment.9

    Capital expenditure

2025 capital expenditure for Alphabet has been estimated at approximately $91–93 billion (approximately 60% servers, 40% data centers and networking), with $175–185 billion planned for 2026 — nearly double the prior year.4 These figures encompass all of Alphabet's infrastructure, not Google Cloud AI alone, but the company has publicly attributed the acceleration to AI infrastructure demand.

  Model Catalog and Performance

The Gemini Enterprise Agent Platform's Model Garden offers more than 200 models spanning first-party, third-party commercial, and open-source families. Google's own Gemini and Gemma families anchor the catalog; third-party integrations address enterprise requirements for model diversity and vendor risk management.

    Google first-party models

The Gemini family spans four current generations (2.5, 3, 3.1, and 3.5) with three tiers at each generation: Pro (highest capability), Flash (balanced price-performance), and Flash-Lite or equivalent (lowest cost). Key models as of mid-2026:

ModelPositioningNotable benchmarks
Gemini 2.5 ProFlagship reasoning (prior gen)
Gemini 2.5 FlashFast multimodal; 1M token context207 t/s (Google API); 165.7 t/s (Vertex)
Gemini 2.5 Flash-LiteLowest cost in family392.8 t/s; 0.29s TTFT
Gemini 3.1 ProCurrent flagship reasoning94.3% GPQA Diamond; 77.1% ARC-AGI-2; top-1 on 12/18 tracked benchmarks (reported)10
Gemini 3.5 FlashFast agentic / coding76.2% Terminal-bench 2.1; 55.1% SWE-Bench Pro (third-party reported)10
Gemini Omni FlashAnnounced Google I/O 2026

The Gemma family (open models, Apache 2.0 license) supports on-premises and sovereign cloud deployment, with Gemma 4 announced at Google Cloud Next 2026. Imagen 4 handles text-to-image generation and Veo 3 video generation.

    Third-party models via Model Garden

Model Garden gives enterprise customers access to models from Anthropic (Claude 3 family and Claude Fable 5, optimized for coding and autonomous work), Meta (multiple Llama versions), Mistral AI, and a broad set of open-source models — all with consistent billing, IAM, and data-privacy guarantees through the Gemini Enterprise Agent Platform.1

  Pricing and Performance Position

Google Cloud AI competes across the full price-performance spectrum, from the lowest-cost inference available from a frontier lab to high-capability reasoning models positioned against OpenAI and Anthropic.

ModelInput ($/M tokens)Output ($/M tokens)SpeedNotes
Gemini 2.5 Flash-Lite$0.10$0.40392.8 t/s; 0.29s TTFTCheapest in family
Gemini 2.5 Flash$0.30$2.50207 t/s (Google API); 1M token contextCache hits: ~$0.03/M (~90% discount)
Gemini 2.5 Pro$1.25 (under 200K tokens)$10$2.50/M input over 200K tokens
Gemini 3.1 Pro$2 (under 200K tokens)$12$4/$18 over 200K tokens
Gemini 3.5 Flash$1.50$9Pricing reported; verify current rates

Free tier: AI Studio provides access to Flash and Flash-Lite models (up to 500 RPM on Flash) without billing enablement. Pro models require billing activation. Consumer subscriptions include Google AI Pro at $19.99/month (Gemini 3.1 Pro access with 1M token context) and Google AI Plus at $7.99/month.10

Context caching delivers approximately a 90% discount on cache-hit tokens, making long-context and repeated-document use cases substantially more economical. Management's reported 78% reduction in per-query serving costs during 2025 suggests continued price compression for future model generations.9

  People and Leadership

The Gemini Enterprise Agent Platform sits at the intersection of three leadership structures within Alphabet, each with distinct accountability.

Thomas Kurian, CEO of Google Cloud, is the primary commercial leader of the platform. Kurian joined Google in 2019 after nearly 25 years at Oracle, where he served as President of Product Development. His enterprise sales orientation has been credited with accelerating Google Cloud's penetration of Fortune 500 customers and structuring the large committed-use deals that underpin the $240B backlog.

Demis Hassabis, CEO of Google DeepMind and co-founder of the original DeepMind (London, 2010), leads the research organization that produces all Gemini models. Hassabis was awarded the Nobel Prize in Chemistry in 2024 for AlphaFold's contributions to protein structure prediction — a recognition that lends Google DeepMind unusual scientific credibility among enterprise research customers.2 He assumed leadership of the merged Google Brain / DeepMind organization in April 2023.

Jeff Dean, Chief Scientist at Google DeepMind, co-founded Google Brain in 2011 and led the original TPU program. His research contributions include the MapReduce programming model, the Bigtable distributed storage system, and foundational work on deep learning at scale. Dean provides scientific continuity between Google's research history and its current Gemini-era direction.

Koray Kavukcuoglu serves as CTO of Google DeepMind, leading the engineering teams responsible for integrating Gemini models with Google Cloud's serving infrastructure. Sissie Hsiao (VP, Gemini AI Experiences) leads consumer-facing Gemini products and the integration of AI capabilities into Google Assistant and Workspace. Sundar Pichai (CEO, Alphabet) provides board-level oversight and has consistently positioned AI as Alphabet's primary strategic priority across public communications.

  Position within Parent Org

Google Cloud AI is a business unit of Alphabet Inc. (NASDAQ: GOOGL / GOOG), a publicly traded company with an approximate market capitalization exceeding $2 trillion as of mid-2026. There is no separately capitalized AI entity; investment flows from Alphabet's balance sheet and operating cash flow.

Google Cloud as a segment reported $17.7 billion in revenue in Q4 2025 (+48% year-over-year), with operating income of approximately $5.3 billion (~30% margin).3 Full-year 2025 marked the first year Alphabet's total revenue surpassed $400 billion. Google Cloud's order backlog grew 55% sequentially and more than doubled year-over-year to $240 billion at end of Q4 2025, making it the fastest-growing of the three major hyperscalers in Q4 2025 (~50% YoY vs. Azure's ~31% and AWS's ~19% in comparable periods, per analyst estimates).3

Google does not disclose Vertex AI / Gemini API revenue separately from overall Google Cloud revenue, so the AI model-serving contribution to these figures is estimated, not reported. The business model is consumption-based: pay-per-token for model APIs, per-node-hour for dedicated training compute, per data point for AutoML forecasting. Enterprise committed-use discounts and the $300 credit / 90-day free trial serve as acquisition mechanisms. Two consumer subscription tiers — Google AI Pro ($19.99/month) and Google AI Plus ($7.99/month) — extend the commercial model to individual users, with Gemini paid users reported as growing approximately 40% quarter-over-quarter as of April 2026.7

A potentially significant business model extension is TPU-as-a-service for external AI labs: Anthropic has committed to access up to one million TPU chips (chip-equivalents over a multi-year period; exact terms not disclosed) for training and serving Claude models.6 Meta has been reported in advanced discussions to lease TPUs beginning 2026 and potentially purchase TPU systems outright from 2027 — a deal described as multibillion-dollar but not confirmed closed as of the research date.3 If these arrangements mature into a hardware-licensing revenue stream, they would represent a meaningful structural expansion of Google Cloud's business model beyond traditional cloud consumption.

  Customers and Partnerships

Enterprise deployments of the Gemini Enterprise Agent Platform span manufacturing, natural resources, professional services, and the AI lab sector. Danfoss, the Danish industrial manufacturer, deployed a natural-language agentic workflow to automate 80% of transactional email order-processing decisions, reducing response times from an estimated 42 hours to near real-time. Suzano, the Brazilian pulp and paper company, built a natural-language-to-SQL agent that reduced query times by approximately 95% for 50,000 employees.7

In the AI lab segment, Anthropic's expanded TPU commitment (up to one million chips, announced October 2025) makes Google Cloud the primary infrastructure provider for Claude model training and serving.6 Meta's reported TPU discussions, if concluded, would establish Google as a major compute provider to a direct competitor — an unusual arrangement reflecting the scarcity of TPU-scale compute alternatives.

Strategic systems integrator partnerships include Accenture (approximately 45% of joint client projects advanced from generative AI proof-of-concept to production, per reported figures), Deloitte, and KPMG. The $750 million Google partner fund announced at Cloud Next 2026 targets these SI-led enterprise deployments.8 ISV integrations via Workspace Studio and the A2A protocol extend the ecosystem to Salesforce, ServiceNow, Workday, Box, Asana, and Jira.

The A2A protocol v1.0, now governed under the Linux Foundation, reached production deployment at 150 organizations by April 2026.7 This open governance model is a deliberate strategy to establish A2A as an industry standard rather than a proprietary lock-in mechanism — mirroring the approach taken with Android and Kubernetes in prior technology cycles.

  Competitive Position

Google Cloud holds third position in global cloud market share behind AWS and Microsoft Azure, a standing that has persisted for years, though the growth rate gap has narrowed substantially. Google Cloud's approximately 50% year-over-year growth in Q4 2025 outpaced both Azure (estimated ~31% YoY) and AWS (estimated ~19% YoY) in comparable periods, a trend driven primarily by AI workload demand.3

Google's differentiation claims rest on five pillars. First, vertical integration: custom TPU silicon designed in-house, frontier Gemini models from Google DeepMind, unified cloud platform, and a three-billion-user Workspace distribution channel for agent deployment — a combination no rival can replicate in full. Second, TPU price-performance: management reports 78% reduction in per-query serving costs during 2025, and TPU 8i claims 80% better inference price-performance over the prior generation, though these figures are based on Google's own benchmarking.9 Third, model breadth: Model Garden's 200+ models include first-party (Gemini, Gemma), third-party (Claude, Mistral), and open-source (Llama) options, enabling enterprises to consolidate inference spend on a single platform without betting on a single model provider. Fourth, agent infrastructure: A2A protocol at production scale, ADK v1.0, and managed MCP servers represent a more mature agentic runtime than most competitors have shipped at equivalent scale. Fifth, consumer distribution: Workspace with billions of users provides a built-in go-to-market for Workspace Studio agents that enterprise software competitors cannot match.

The key competitive risks are correspondingly structural. Microsoft Azure AI Foundry and its OpenAI partnership have the deepest enterprise software penetration through Microsoft 365 and Azure Active Directory identity. AWS Bedrock has the broadest third-party model catalog and the largest absolute hyperscaler cloud customer base. OpenAI Operator and Anthropic MCP ecosystem are building developer loyalty that is partly independent of any cloud provider's stack. And the fastest-improving open-weight models — from DeepSeek, Meta, and the Gemma family itself — exert commoditization pressure on paid model APIs across all providers.

  Outlook and Roadmap

Google's stated strategic direction centers on owning "the full stack from chip to inbox" in the agentic era — a phrase that encapsulates the vertical integration bet described throughout this dossier. Several near-term directions are visible from public announcements and financial signals, though specific model release timelines and product roadmap details beyond what has been announced should be treated as speculative.

The TPU 8t and 8i silicon announced at Cloud Next 2026 will roll out through 2026–2027, providing the infrastructure base for both Gemini model training and high-volume inference.9 Expansion of the A2A protocol's governance under the Linux Foundation is likely to accelerate third-party adoption, potentially establishing A2A as a cross-cloud standard for agent communication — a strategic outcome that would benefit Google whether customers build agents on Google Cloud or elsewhere. The Workspace Studio no-code agent builder targets the large population of enterprise users who interact with AI through Gmail, Docs, and Sheets rather than through developer APIs, a segment that neither pure AI labs nor pure infrastructure providers currently reach at scale.

The Gemini 3.x model line — currently at 3.1 Pro and 3.5 Flash — will continue to evolve, with the open-model Gemma 4 family (Apache 2.0) addressing sovereign cloud, on-premises, and regulated-industry use cases that cannot route data through public APIs. The $240 billion cloud backlog and $175–185 billion in planned 2026 capex represent the financial commitment underpinning these directions.34

The potential maturation of TPU-as-a-service into a standalone revenue stream — via Anthropic's multi-year chip commitment and Meta's reported lease discussions — would represent a meaningful expansion of the business model, transforming TPUs from a cost advantage into a product. Whether this develops into a systematic external-leasing business analogous to AWS's original "excess capacity" thesis will be one of the more consequential strategic questions for Google Cloud AI over the next 24 months.


  References

  1. Google Cloud Vertex AI / Gemini Enterprise Agent Platform — Product Overview
  2. Gemini Enterprise Agent Platform — Wikipedia
  3. Google Cloud Next 2026 — AI Agents and Agentic Era (The Next Web)
  4. Google Cloud Next 2026 — AI Infrastructure Announcement
  5. Gemini 2.5 Flash — Artificial Analysis
  6. Alphabet Q4 FY2025 Earnings — CNBC
  7. Google Cloud $750M Partner Fund — The Next Web
  8. Anthropic Expands Google Cloud TPU Usage — Google Cloud Press Corner
  9. Google Consolidates AI Research into Google DeepMind — TechCrunch
  10. Google AI Timeline — Script by AI
  11. Google 2025 Capex and 2026 Data Center Spend — Data Center Dynamics
  12. Alphabet Q4 FY2025 Highlights — Futurum Group
  13. TPU 8t and 8i Architecture Analysis — Hyperframe Research

  References

  1. Google Cloud Vertex AI / Gemini Enterprise Agent Platform product documentation — cloud.google.com/vertex-ai. 2 3 4 5

  2. Demis Hassabis biography and Nobel Prize — Wikipedia: Gemini Enterprise Agent Platform. 2

  3. Alphabet Q4 2025 earnings — CNBC; Futurum Group. Growth rate comparisons with Azure and AWS are analyst estimates, not directly comparable official figures. 2 3 4 5 6 7 8 9

  4. 2025 and 2026 capex figures are reported estimates — Data Center Dynamics. 2 3 4 5

  5. Google DeepMind merger — TechCrunch, April 2023. 2 3

  6. Anthropic TPU commitment — Google Cloud Press Corner, Oct 2025. The "1 million TPU chip" figure may refer to chip-equivalents over a multi-year period; exact terms not publicly disclosed. 2 3

  7. Google Cloud Next 2026 announcements — The Next Web. 2 3 4

  8. Google Cloud $750M partner fund — The Next Web. 2

  9. Google Cloud AI infrastructure at Next 2026; TPU 8t/8i architecture — cloud.google.com blog; Hyperframe Research. 2 3 4 5 6 7 8

  10. Pricing and benchmark data sourced from third-party analysis — Artificial Analysis. Benchmark scores for Gemini 3.5 Flash are third-party reported; official Google benchmarks may differ. 2 3