Command Palette

Search for a command to run...

Cerebras

Wafer-scale AI silicon company delivering 15–30x faster LLM inference than GPU-based solutions via its CS-3 system and WSE-3 chip — a single 46,225 mm² silicon wafer with 4 trillion transistors — plus a public inference cloud API.

Cerebras

  Executive Briefing

Cerebras Systems is the only company in the world that manufactures a processor occupying an entire silicon wafer, and has turned that singular engineering bet into the fastest publicly available LLM inference service. Founded in March 2016 in Sunnyvale, California, by a team of five co-founders — Andrew Feldman (CEO), Sean Lie (CTO), Michael James (Chief Software Architect), Jean-Philippe Fricker (Chief System Architect), and Gary Lauterbach (original CTO, now retired) — Cerebras was built on a single conviction: that the real bottleneck in AI compute was not raw processing power but the cost, latency, and energy of moving data between thousands of discrete chips strung together across PCIe and NVLink fabrics. The answer, the founders argued, was to build a processor that fills an entire 300mm silicon wafer, eliminating inter-chip communication overhead entirely.1

The result is the Wafer Scale Engine (WSE), now in its third generation. The WSE-3, manufactured by TSMC on a 5nm process, contains 4 trillion transistors, 900,000 AI-optimized compute cores, and 44 GB of on-chip SRAM delivering 21 petabytes per second of memory bandwidth — figures that dwarf any conventional GPU die by orders of magnitude.2 The WSE-3 is housed in the CS-3, a 15U water-cooled rack system, and clusters of CS-3 systems form the Condor Galaxy supercomputer network that powers the Cerebras Inference Cloud — a public API reporting speeds of up to 2,500 tokens per second per user on models such as Llama 4 Maverick, against GPU baselines that typically deliver 100–200 tokens per second.3

From the outside, Cerebras looks like an inference-API company competing on speed. From the inside, it is an integrated semiconductor designer, system builder, and cloud operator — a vertically integrated stack from the silicon up. Its business spans three channels: hardware sales and leases (CS-3 systems to enterprises and research labs), cloud inference API access (pay-per-token and subscription tiers), and long-term cluster supply agreements with hyperscalers and sovereign AI programs. The most significant of these is a reported $10 billion-plus compute supply agreement signed with OpenAI in January 2026, described as the largest high-speed AI inference deployment in the world.4

Cerebras listed on the Nasdaq under the ticker CBRS on 14 May 2026, raising $5.55 billion in the largest U.S. technology IPO since Uber in 2019. Shares opened at $350 against a $185 IPO price, closing the first day at $311, a 68% gain, and briefly trading as high as $386 intraday.5 The IPO valued the company at roughly $95 billion at its first-day close. The path to the public markets was not straightforward: Cerebras withdrew an earlier IPO attempt in October 2025 after regulators scrutinized the company's heavy revenue concentration in UAE-linked entities — G42 and MBZUAI together accounted for an estimated 86% of 2025 revenue, a concentration that generated both CFIUS and SEC concerns.6 The OpenAI and a separate Amazon Web Services binding term sheet signed in March 2026 mark the beginning of a customer diversification push that the company's long-term trajectory depends on.

  At a Glance

ItemDetail
FoundedMarch 2016 (Sunnyvale, CA)
TypeInference hardware manufacturer and cloud operator
HeadquartersSunnyvale, CA, USA
StatusActive (Nasdaq: CBRS, IPO May 2026)
LeadershipAndrew Feldman (CEO), Sean Lie (CTO), Dhiraj Mallick (COO), Robert Komin (CFO)
OwnershipIndependent (public as of May 2026); AMD is a strategic investor
SpecialtiesWafer-scale AI silicon, ultra-high-throughput LLM inference, on-premises AI supercomputers
Key hardwareWSE-3 chip, CS-3 system, MemoryX external memory, Condor Galaxy clusters
Reported pre-IPO funding~$2.91 billion across multiple rounds7
IPO raise$5.55 billion at $185/share (May 2026)5
Post-IPO market cap (day 1 close)~$95 billion
Last private valuation~$23 billion (Series H, February 2026)8

  Origins & Founding

The five Cerebras co-founders share a common origin: they all worked together at SeaMicro, a microserver startup founded in 2007 and acquired by AMD in 2012 for a reported $334–357 million.1 SeaMicro's thesis was that conventional high-density servers wasted power and space on interconnect; it packaged hundreds of low-power Intel Atom chips onto a single board using a custom high-bandwidth fabric. Cerebras applied an analogous insight one level deeper: if the bottleneck for server compute was the overhead of connecting many chips, the bottleneck for AI compute was the overhead of connecting many GPU dies.

Andrew Feldman had been the CEO of SeaMicro and became CEO of Cerebras. Gary Lauterbach served as original CTO and guided early silicon architecture before retiring. Sean Lie (Chief Hardware Architect, later CTO) led chip design. Michael James (Chief Software Architect) built the software stack, and Jean-Philippe Fricker (Chief System Architect) owned system integration. Cerebras was incorporated in March 2016, with some sources attributing an informal founding to late 2015; the March 2016 date is more consistently cited for the legal entity.1

The founding thesis was architectural, not merely incremental. GPU-based AI systems spend enormous energy and introduce substantial latency routing activations across PCIe buses and NVLink fabrics between hundreds or thousands of physically separate chips. A single processor filling an entire 300mm silicon wafer would host all its compute cores and memory on a single substrate, connected by on-wafer metal interconnects rather than chip-to-chip links. The tradeoff was radical: manufacturing yields on a wafer-scale device are dramatically more sensitive to defects than on a small chip, requiring entirely new approaches to wafer design, defect tolerance, and thermal management. Cerebras spent its first three years proving that these challenges were solvable before announcing its first product publicly.

  History & Timeline

    2016–2019: Stealth development and the WSE-1

Cerebras operated in stealth for its first three years, raising a $27M Series A in May 2016 from Benchmark, Foundation Capital, and Eclipse Ventures, followed by further rounds as the team built out the WSE-1.7 The WSE-1 was unveiled at the Hot Chips conference in August 2019 alongside the CS-1 compute system — the first commercially offered wafer-scale AI computer. The WSE-1 was manufactured by TSMC on a 16nm process, contained 1.2 trillion transistors and 400,000 cores, and occupied 46,225 mm² of silicon, making it approximately 56 times larger than a contemporary Nvidia V100 GPU die.2 The announcement was a striking proof-of-concept: Cerebras had actually shipped a product based on an architecture that many in the semiconductor industry had considered impractical. A $270M Series E at a $2.4B valuation in November 2019 confirmed that investors believed the bet could scale.7

    2020–2022: WSE-2, Andromeda, and the first customers

The WSE-2, announced in April 2021 on TSMC 7nm, doubled core count to 850,000 and transistor count to 2.6 trillion, while improving memory bandwidth and power efficiency. The CS-2 system housing the WSE-2 became the workhorse for early customers including national laboratories and research institutions. Cerebras demonstrated in June 2022 that a single CS-2 could train a 20-billion-parameter model without model parallelism — a feat that would require a rack of high-end GPUs to achieve. In November 2022, the company unveiled Andromeda, a 16-CS-2 cluster delivering one exaflop of AI compute in 500 kW — demonstrating that multiple CS-2 systems could be tightly coupled into a supercomputer. The WSE was inducted into the Computer History Museum in August 2022, an unusual recognition for a product less than four years old.1

    2023–2024: Condor Galaxy, WSE-3, and the inference cloud

The partnership with G42 — the Abu Dhabi AI conglomerate — proved transformative and ultimately problematic in equal measure. Announced in July 2023, the agreement saw G42 commit to operating Condor Galaxy data centers globally, with Cerebras supplying the CS-3 hardware. Condor Galaxy 1 (CG-1), a 64-node CS-2 cluster delivering 4 exaFLOPS, launched in Santa Clara, California, in Colovore's facility in July 2023. Condor Galaxy 3 (CG-3) broke ground in Dallas, Texas, in September 2023, targeting 8 exaFLOPS using CS-3 systems.

The WSE-3 was unveiled in March 2024, moving to TSMC 5nm: 4 trillion transistors, 900,000 cores, 44 GB on-chip SRAM with 21 PB/s of memory bandwidth, and 125 peak petaflops of AI compute per system.2 Cerebras also launched the Cerebras Inference Cloud in August 2024 — a public API offering open-weight LLM inference at reported speeds that exceeded any publicly available GPU-based service, including a free developer tier providing 1 million tokens per day without a credit card. In October 2024, the company claimed to have tripled inference performance relative to its August baseline.

G42 and its affiliated entity MBZUAI (Mohamed bin Zayed University of Artificial Intelligence) together represented approximately 85% of 2024 revenue, a concentration that would define Cerebras' regulatory challenges.6

    2025: The blocked IPO, Series G, and competitive shock

Cerebras filed a confidential Draft Registration Statement (DRS) with the SEC in September 2024, targeting a public listing. The filing was not immediately approved; by October 2025 the company formally withdrew the IPO application after the SEC and the Committee on Foreign Investment in the United States (CFIUS) raised concerns about the company's revenue dependence on G42, which was itself under a separate CFIUS investigation related to its ties to Chinese technology firms.6 The withdrawal was a significant setback.

Despite it, Cerebras closed a $1.1 billion Series G in September 2025 at an $8.1 billion valuation, led by Fidelity Management and Research and Atreides Management, with participation from Tiger Global, Valor Equity Partners, 1789 Capital, Altimeter, Alpha Wave Global, and existing investor Benchmark.9 December 2025 brought a major competitive development: Nvidia announced a reported $20 billion licensing or acquisition agreement with Groq — Cerebras' closest throughput competitor — signaling that the incumbent GPU leader was directly targeting the inference-speed market that Cerebras had pioneered.

    2026: OpenAI deal, AWS, Series H, and the IPO

January 2026 saw the announcement that opened the customer diversification path: OpenAI agreed to a reported $10 billion-plus compute supply agreement with Cerebras, covering 750 MW of inference capacity through 2028, described as the largest high-speed AI inference deployment in the world.4 OpenAI CEO Sam Altman is also a personal investor in Cerebras. In February 2026, Cerebras closed a $1 billion Series H at approximately $23 billion post-money valuation, with new investors including AMD, Coatue, and Tiger Global joining existing backers.8

A binding term sheet with Amazon Web Services for Cerebras inference capacity was signed in March 2026 — the final agreement was pending as of the IPO filing, but the term sheet was described as establishing pricing, exclusivity terms, and minimum commitments. Cerebras refiled its S-1/A in April 2026, and the IPO priced at $185 per share on 13 May 2026, above the indicated range. Trading began on Nasdaq on 14 May 2026; shares opened at $350, peaked intraday at $386, and closed at $311, a 68% gain, as the company raised $5.55 billion — the largest U.S. technology IPO since Uber's 2019 debut.5

  What They Offer — Products & Platform

Cerebras' commercial offering is organized into three interconnected products that share the same underlying hardware stack.

The CS-3 AI Supercomputer is the company's primary hardware product: a 15U rack unit housing the WSE-3 wafer, proprietary water cooling, and supporting electronics, consuming approximately 23 kW. The CS-3 can be deployed standalone or in clusters, and interfaces with the MemoryX external memory system that provides 1.5 TB, 12 TB, or 1.2 PB of off-chip storage accessible via a proprietary high-bandwidth interconnect. MemoryX enables a weight-streaming architecture in which model weights too large to fit in the WSE-3's 44 GB on-chip SRAM are streamed in at inference time from MemoryX, allowing a single CS-3 to serve models of hundreds of billions of parameters — or, for training, models reported up to 24 trillion parameters. Multiple CS-3 systems can be interconnected with a proprietary fabric, with clusters called Condor Galaxy installations ranging from tens to hundreds of CS-3 nodes.3

The Cerebras Inference Cloud is a public API service offering OpenAI-compatible endpoints for LLM inference. It runs on Condor Galaxy clusters operated by Cerebras in its own data centers and by partner G42 across six global sites. The API offers three tiers: a free developer tier providing 1 million tokens per day without a credit card; a Developer pay-per-use tier billed per token; and Enterprise agreements with reserved capacity, higher rate limits, compliance options, and SLA guarantees. Fine-tuning and model customization services are also offered. The API's primary differentiator is throughput: Cerebras publicly reports token-generation rates of 1,800–2,500 tokens per second per user on major open-weight models, compared with GPU-based inference services that typically deliver 100–400 tokens per second.3

On-premises CS-3 deployments form the third channel, serving enterprises and research labs requiring dedicated capacity, data residency, or air-gapped operation. Customers purchase or lease CS-3 systems and receive access to the Cerebras Software Platform (CSP), which provides PyTorch and JAX compatibility, model compilation tooling, and profiling utilities for deploying on the WSE-3.

  Technology & Infrastructure

    The Wafer Scale Engine (WSE-3)

The central innovation is the decision to use an entire TSMC 300mm silicon wafer as a single processor die rather than dicing it into hundreds of individual chips. The WSE-3, introduced in March 2024 and built on TSMC's 5nm process node, occupies an active area of approximately 46,225 mm² — roughly 57 to 58 times larger than the die area of a Nvidia H100 GPU.2 It contains 4 trillion transistors, 900,000 AI-optimized compute cores, and 44 GB of on-chip SRAM delivering 21 petabytes per second of memory bandwidth. By comparison, an Nvidia H100 SXM offers approximately 3.35 TB/s of HBM3 memory bandwidth — about 1/6,000th the on-chip bandwidth of the WSE-3.

The on-wafer fabric delivers 214 petabits per second of aggregate interconnect bandwidth between cores, with single-clock-cycle core-to-core latency and no off-chip routing for intra-wafer communication. This eliminates the primary bottleneck in transformer inference: the need to repeatedly load weight tensors from slow, external DRAM through narrow memory buses during each attention and feed-forward computation. For large language models, where memory bandwidth (not raw compute) is typically the binding constraint during auto-regressive token generation, this architecture produces the throughput advantages Cerebras claims.2

Manufacturing a wafer-scale device poses yield challenges that conventional chip production does not face. A single dust particle or process defect can destroy a traditional chip; at wafer scale, such defects are statistically unavoidable. Cerebras addresses this through a redundancy and defect-tolerance design: the WSE-3 includes spare cores and on-wafer routing that bypasses known-bad regions identified during post-fabrication testing. Thermal management is similarly challenging — the WSE-3 requires a custom water-cooling system for each unit, designed around the heat distribution profile of a single large silicon plate rather than a standard chip package.2

Prior WSE generations: the WSE-1 (2019, TSMC 16nm, 1.2 trillion transistors, 400,000 cores) and WSE-2 (2021, TSMC 7nm, 2.6 trillion transistors, 850,000 cores) each roughly doubled key metrics generation-over-generation.

    Weight-Streaming Architecture and MemoryX

The WSE-3's 44 GB of on-chip SRAM is fast but finite; a 70B-parameter model in bf16 requires approximately 140 GB of weight storage, and a 405B model requires roughly 810 GB. Cerebras addresses this via the weight-streaming architecture: model weights reside in MemoryX external memory (1.5 TB to 1.2 PB per system) and are streamed onto the WSE-3 compute fabric at inference time. Because the WSE-3's aggregate internal bandwidth far exceeds what any off-chip interface can deliver, compute cores process incoming weights at full speed — the weight streaming rate becomes the pacing factor, but the architecture allows very large models to run on a single WSE-3 rather than requiring model parallelism across many GPUs. Training on the WSE-3 uses a related distributed pipeline where a single physical CS-3 can handle models that would require tensor parallelism across dozens of GPU nodes.3

    Condor Galaxy Network and Data Centers

Cerebras operates and enables a global network of Condor Galaxy supercomputer clusters:

  • Condor Galaxy 1 (CG-1): 64 CS-2 nodes, 4 exaFLOPS, Santa Clara, CA (Colovore facility)
  • Condor Galaxy 3 (CG-3): 64 CS-3 systems, 8 exaFLOPS, Dallas, TX
  • Oklahoma City: Cerebras-owned facility
  • Montreal: Cerebras-owned facility
  • Six G42-operated facilities: globally distributed, exact locations not fully public

By end of 2025, the Condor Galaxy network collectively targeted over 40 million tokens per second of total inference throughput.10

    Software Stack

The Cerebras Software Platform (CSP) exposes PyTorch and JAX interfaces over the WSE-3 hardware. It includes a compiler that maps computational graphs to the WSE-3's core layout, a runtime for weight streaming and MemoryX management, and profiling tools. The inference API is largely OpenAI-compatible — applications using the OpenAI Python SDK can typically switch providers by changing the base URL and API key. Cerebras also offers fine-tuning as a managed service for select models.

  Model Catalog & Performance

Cerebras' inference catalog focuses on leading open-weight models where throughput advantage is most pronounced. The company serves models from Meta AI, DeepSeek, Zhipu AI, and others, with the catalog expanding as software support matures:

ModelReported speedNotes
Llama 4 Maverick (~400B parameters)~2,500 tokens/sec per userReported 2x faster than Nvidia DGX B2003
Llama 3.3 70B~1,800–2,220 tokens/secSee pricing data in frontmatter
Llama 3.1 8B~1,800–2,047 tokens/secReported 20x faster than GPU baselines3
Llama 3.1 70B~1,200 tokens/secSee pricing data in frontmatter
Llama 3.1 405BNot publicly specifiedAvailable via weight-streaming
DeepSeek R1 Distill Llama 70B1,500+ tokens/secReported 57x faster than GPU solutions3
Llama 4 ScoutNot publicly specifiedAvailable on the inference API
GLM-4.7Not publicly specifiedZhipu AI model

The catalog is narrower than GPU-based inference providers, reflecting the software investment required to map each new model architecture to the WSE-3 compiler. Cerebras has stated intent to expand coverage as toolchain support broadens. The company also deploys select Qwen-family models and has signaled ongoing work to add models from additional open-weight families.

  Pricing & Performance Position

Cerebras' pricing strategy pairs a deliberately accessible free tier with competitive per-token rates designed to undercut GPU-based inference providers on cost while offering superior speed. Public pricing as of early-to-mid 2025 placed Llama 3.1 8B at $0.10 per million tokens and Llama 3.1 70B at $0.60 per million tokens, with the broader range across models at approximately $0.10–$1.50 per million tokens.3 As of mid-2026, the company has moved toward subscription and reserved-capacity tiers for larger customers, and specific per-token public rates for newer models may vary from historical figures — treat current pricing as indicative only.

The free developer tier (1 million tokens per day, no credit card required) positions Cerebras as the fastest zero-cost way to evaluate leading open-weight models, and serves as a customer acquisition funnel. Enterprise agreements offer reserved capacity, higher rate limits, data-residency options, compliance documentation, and uptime SLAs. Reserved-capacity batch contracts are available for high-volume workloads.

On throughput, Cerebras' reported benchmarks — if taken at face value — represent order-of-magnitude improvements over GPU-based alternatives for per-user token generation rates. Independent benchmarks broadly confirm that Cerebras delivers the fastest single-request token throughput available on a public API for the models it supports. The tradeoff is model breadth: GPU-based providers offer far larger catalogs covering hundreds of models, including proprietary frontier models such as GPT-5, Claude, and Gemini, which Cerebras does not serve.

  People & Leadership

The founding team has remained largely intact through the decade since incorporation — an unusual degree of stability for a deep-tech startup. Andrew Feldman (CEO) remains the public face of the company and its primary spokesperson, a veteran of the SeaMicro acquisition who brings both hardware and business-development expertise. Sean Lie (CTO and Chief Hardware Architect) leads silicon development and has guided the WSE across three generations. Michael James (Chief Software Architect) owns the compiler and runtime stack that makes the WSE programmable. Jean-Philippe Fricker (Chief System Architect) oversees system integration — the thermal, mechanical, and electrical engineering that turns a silicon wafer into a data-center-deployable product.

Gary Lauterbach, the fifth co-founder who served as original CTO, has since retired from the company. On the executive side, Robert Komin serves as CFO and Dhiraj Mallick as COO. Both were recruited as the company scaled its commercial operations ahead of the IPO. The IPO S-1 filings estimated the net worth of Andrew Feldman at approximately $3.4 billion at the initial offering price, reflecting his stake as co-founder and CEO.5

Headcount details have not been publicly disclosed in detail, but Cerebras operates across engineering, sales, and operations functions consistent with a late-stage private company that has raised nearly $3 billion pre-IPO.

  Funding, Ownership & Business

Cerebras has raised a reported $2.91 billion in private capital across multiple rounds before its May 2026 IPO:7

RoundDateAmountKey investors / notes
Series AMay 2016$27MBenchmark, Foundation Capital, Eclipse Ventures
Series DNov 2018$88MUnicorn milestone
Series ENov 2019$270M$2.4B valuation
Series FNov 2021$250M$4B+ valuation
Series F-1Jul–Sep 2024$85MBridge round
Series GSep 2025$1.1B$8.1B valuation; led by Fidelity, Atreides; Tiger Global, Valor, Altimeter, Alpha Wave, Benchmark9
Series HJan–Feb 2026$1.0B~$23B valuation; led by Tiger Global; new investors AMD, Coatue8
IPO (CBRS)May 14, 2026$5.55B$185/share; 30M Class A shares; Nasdaq5

The business model combines four revenue streams: hardware sales and leases (CS-3 systems to enterprises, research labs, and government programs); cloud inference API (pay-per-token and subscription tiers for Cerebras Inference Cloud); Condor Galaxy cluster supply agreements (long-term capacity contracts with partners such as G42 and MBZUAI); and compute supply agreements with hyperscalers (the OpenAI $10B+ contract and the AWS term sheet). The G42 and MBZUAI relationships represented approximately 85% of 2024 revenue and an estimated 86% of 2025 revenue — a concentration that generated both the CFIUS regulatory challenge that blocked the first IPO and the strategic imperative behind the 2026 OpenAI and AWS deals.6

AMD is both a strategic investor and the acquirer of SeaMicro, the company where all five Cerebras founders previously worked — an alignment that reflects AMD's interest in alternative AI accelerator ecosystems. Sam Altman (OpenAI CEO) is a personal investor. Nvidia acquired Groq in December 2025 in a reported $20 billion deal, making the GPU giant a direct competitor in the inference-speed market.11

Reported financial figures vary across sources and should be treated with appropriate uncertainty: one source cites a 2025 GAAP net income figure of $87M while others cite $237.8M; the discrepancy may reflect different reporting periods or the inclusion of one-time items.

  Customers & Partnerships

Cerebras' most significant commercial relationships as of mid-2026 are:

OpenAI is the company's largest reported customer by contract value. The reported $10 billion-plus compute supply agreement announced in January 2026 covers 750 MW of Cerebras inference capacity through 2028 and is described as the largest high-speed AI inference deployment in the world.4 OpenAI and Cerebras have a longstanding technical relationship dating to 2017, and OpenAI CEO Sam Altman is a personal investor in Cerebras. Some sources have cited the deal value as $20 billion; the exact contract value is unconfirmed — treat as reported.

G42 (Alpha Technology Investments, Abu Dhabi) committed $1.43 billion in Cerebras hardware purchases in May 2024 and operates six global data centers running Condor Galaxy clusters. G42 represented approximately 24% of 2025 revenue (down from around 60% in 2024 as MBZUAI grew).6 The G42 relationship is the subject of ongoing CFIUS review; the ultimate resolution of that review was not disclosed in the IPO filing.

MBZUAI (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi) represented approximately 62% of 2025 revenue, making it the single largest customer.6 MBZUAI's purchases are linked to UAE sovereign AI infrastructure development and are connected to the G42 network.

Amazon Web Services signed a binding term sheet for Cerebras inference capacity in March 2026, establishing pricing, exclusivity terms, and minimum commitments. Final contractual agreements were pending as of the IPO filing and should be treated as unfinalized.5

Beyond named enterprise and government contracts, Cerebras serves a broad base of developers, researchers, and startups through the public inference API.

  Competitive Position

Cerebras occupies a structurally unique position in AI infrastructure: it is the only wafer-scale processor vendor in commercial production, which means its primary differentiator — eliminating inter-chip communication overhead — is not available to any competitor operating on conventional die architectures. This provides a durable technical moat as long as no other company successfully brings wafer-scale silicon to market at comparable yields and performance.

The competitive landscape has nonetheless shifted materially since the WSE-3 launch:

Nvidia remains the dominant AI chip vendor, and its December 2025 reported acquisition of Groq signals a direct move into the inference-throughput market where Cerebras competes. Groq's Language Processing Unit (LPU) architecture targets single-query latency; the combination of Groq's throughput-optimized chips with Nvidia's Blackwell GPU line and distribution network is the primary competitive threat Cerebras faces as of mid-2026.11

SambaNova (Reconfigurable Dataflow Unit / RDU) serves a similar inference-hardware market, primarily targeting enterprise on-premises deployments, but has not established the same public inference API presence as Cerebras.

GPU-based inference providers — including Together AI, Fireworks AI, Groq (pre-acquisition), DeepInfra, and the hyperscalers' managed inference services — offer far broader model catalogs and the ability to serve proprietary frontier models. Their throughput on large open-weight models is substantially lower than Cerebras, but they benefit from the Nvidia CUDA ecosystem's breadth and maturity.

Hyperscalers (AWS, Azure, Google Cloud, Oracle, CoreWeave) cannot yet match Cerebras' throughput on autoregressive LLM generation using GPU clusters, but they can bundle inference with storage, databases, networking, and enterprise compliance in ways a specialized inference provider cannot. The AWS term sheet may partially resolve this tension by making Cerebras capacity available through AWS's distribution.

Cerebras' key vulnerabilities are model catalog breadth, CFIUS overhang from G42 exposure, and the risk that Nvidia's acquisition of Groq, combined with Blackwell-generation GPU improvements, narrows the throughput gap sufficiently to remove the speed argument for customers who value ecosystem breadth.

  Notable Events

Several episodes have defined Cerebras' public narrative beyond its chip launches:

The October 2025 IPO withdrawal stands as the most significant setback. The company had filed confidentially in September 2024 and been preparing for a public listing through 2025, but withdrew after the SEC and CFIUS raised concerns about the concentration of revenue in G42 and MBZUAI, entities with ties to Abu Dhabi and, in G42's case, prior technology partnerships with Chinese firms under CFIUS scrutiny.6 The withdrawal was publicly disclosed in October 2025; the subsequent Series G closed the same month, providing a capital bridge.

The Nvidia-Groq announcement (December 2025) — a reported $20 billion deal — marked the first major competitive escalation directly targeting Cerebras' throughput positioning. Groq had been Cerebras' closest public-API throughput competitor, and its integration into Nvidia's portfolio signals that the incumbent GPU leader views inference speed as a strategic priority.11

The Computer History Museum recognition (August 2022) inducted the WSE into its permanent collection, an acknowledgment that a product still in active commercial development had already altered the trajectory of semiconductor design.1

The May 2026 IPO (CBRS) raised $5.55 billion with a 68% first-day gain, the largest U.S. technology IPO since Uber's 2019 listing. The second trading day saw approximately a 10% decline, and by early June 2026 the stock had traded as low as approximately $196.73 — a reminder that first-day pop and long-term valuation are distinct phenomena. The IPO S-1 filings' disclosure that 86% of revenue came from UAE entities remained a subject of investor and press commentary through the post-IPO period.56

  Outlook & Roadmap

Cerebras enters its public-company era with roughly $5.55 billion in fresh capital, a high-profile flagship customer in OpenAI, and a competitive moat built on a decade of wafer-scale manufacturing expertise that no rival has yet replicated. The strategic priorities are customer diversification, next-generation silicon, and software ecosystem expansion.

On silicon, the WSE-4 is expected on the product roadmap, likely targeting TSMC's 3nm process node with further increases to core count, memory bandwidth, and compute density relative to the WSE-3. No confirmed specifications or release date have been publicly announced — analyst estimates project a 2026–2027 timeframe, but this should be treated as speculative.12

On customers and revenue, the OpenAI supply agreement through 2028 provides meaningful revenue visibility. The AWS term sheet, once finalized, would open access to AWS's enterprise customer base and could represent the most significant customer acquisition channel Cerebras has ever had. Morgan Stanley reportedly projected $8 billion in Cerebras revenue by 2028, though such projections are inherently uncertain.12

On software and model catalog, the company is investing in compiler improvements to accelerate the onboarding of new model architectures onto the WSE-3. The current catalog's concentration on Llama-family and a small number of other open-weight models reflects the compilation work required for each new architecture; broader catalog coverage would expand the addressable use-case set and reduce the competitive vulnerability to GPU-based providers' greater breadth.

The CFIUS overhang related to the G42 relationship remains unresolved as of the IPO filing, and an adverse CFIUS determination could require Cerebras to unwind or restrict aspects of that commercial relationship. The OpenAI and AWS deals are in part a hedge against this risk, but the G42/MBZUAI revenue base cannot be replaced quickly if that relationship were materially constrained.

The competitive risk from Nvidia-Groq is the company's most significant medium-term strategic threat. Cerebras must continue to extend its throughput advantage through successive WSE generations while simultaneously building the software ecosystem — model breadth, developer tooling, enterprise integrations — that converts speed advantage into stickiness.


  References

  1. Cerebras — Wikipedia
  2. WSE-3 announcement — Cerebras press release
  3. Cerebras inference cloud performance — Cerebras
  4. OpenAI $10B+ Cerebras compute deal — Dataconomy, Jan 2026
  5. Cerebras IPO — CNBC, May 2026
  6. Cerebras IPO and UAE revenue concentration — TechTimes, May 2026
  7. Cerebras funding history — Contrary Research
  8. Series H press release — Cerebras, Feb 2026
  9. Series G press release — Cerebras, Sep 2025
  10. Cerebras six data centers — Data Center Frontier
  11. Competitive landscape — Sacra research
  12. Cerebras S-1 filing — SEC EDGAR

  References

  1. Founding history, SeaMicro background, Computer History Museum — Wikipedia: Cerebras. 2 3 4 5

  2. WSE-3 specifications — Cerebras WSE-3 press release; technical analysis in IEEE Spectrum. 2 3 4 5 6

  3. Cerebras Inference Cloud throughput figures and product descriptions — cerebras.ai; these are company-reported claims. 2 3 4 5 6 7 8

  4. OpenAI compute supply agreement — Dataconomy, January 2026. Deal value is reported as $10B+; some sources cite $20B — exact figure unconfirmed. 2 3

  5. IPO pricing, first-day trading, and capital raise — CNBC, May 14, 2026; SEC S-1/A filing. 2 3 4 5 6 7

  6. UAE revenue concentration, CFIUS concerns, IPO withdrawal — TechTimes, May 2026; revenue share figures are as reported in IPO filings. 2 3 4 5 6 7 8

  7. Funding round history — Contrary Research; Sacra. 2 3 4

  8. Series H — Cerebras press release. Post-money valuation ~$23B is as reported. 2 3

  9. Series G — Cerebras press release. $8.1B valuation as reported. 2

  10. Condor Galaxy network and data centers — Data Center Frontier.

  11. Nvidia-Groq deal — reported December 2025; $20B figure is as cited in market reporting and should be treated as reported/unconfirmed. 2 3

  12. WSE-4 roadmap and revenue projections are analyst estimates and should be treated as speculative; no confirmed specifications or timeline from Cerebras as of June 2026. 2