Command Palette

Search for a command to run...

Replicate

Serverless model hosting marketplace giving developers one-line API access to 50,000+ open-source ML models — acquired by Cloudflare in late 2025 to power its global edge inference platform.

Replicate

  Executive Briefing

Replicate is a serverless inference API and open-source model hosting marketplace that lets any developer run machine learning models with a single API call — no GPU infrastructure, no CUDA configuration, no ML ops expertise required. Founded in October 2019 by Ben Firshman and Andreas Jansson and headquartered in San Francisco, the company built what many in the industry call the "Heroku of AI models": a platform that abstracts away every layer of GPU complexity and exposes a catalog of more than 50,000 community-contributed open-source models through a clean REST API in Python, JavaScript, or any HTTP client.1

The platform rests on Cog, an open-source containerization tool that Firshman and Jansson built as Replicate's foundational primitive. Cog packages any ML model into a Docker-compatible container with the correct CUDA, cuDNN, Python, and framework versions resolved automatically — solving the notorious "works on my machine" problem that had made sharing ML research painful. On top of Cog, Replicate layered a cloud marketplace where developers can discover, run, fine-tune, and privately host models on a pay-per-prediction basis, with no idle charges for public serverless workloads.2

In November 2025, Cloudflare (NYSE: NET) announced the acquisition of Replicate in a move designed to build what Cloudflare's CEO Matthew Prince described as the most seamless AI cloud for developers.3 The acquisition, which closed around December 2025, brings Replicate's 50,000+ model catalog, Cog containerization technology, and model-serving expertise into Cloudflare's global network — with a stated roadmap to make inference available at Cloudflare's edge PoPs worldwide and to integrate model inference natively with Workers, R2, Vectorize, Durable Objects, and AI Gateway. Replicate retains its brand and API compatibility, and its co-founder team has continued in technical roles through the integration.

Replicate matters in the AI infrastructure landscape as the canonical example of the "model marketplace" archetype: a platform that chose breadth and simplicity over optimization depth, winning a community of millions of developers by making the question of "which GPU, which version of PyTorch, which driver" disappear entirely. The Cloudflare acquisition is a bet that this breadth, combined with global edge distribution and a full-stack developer platform, can compete with hyperscaler inference APIs on a dimension none of them can easily replicate — proximity to the developer's own application stack.

  At a Glance

ItemDetail
FoundedOctober 2019
TypeInference API / model hosting marketplace
HeadquartersSan Francisco, CA, USA
StatusAcquired (by Cloudflare, ~December 2025)
LeadershipBen Firshman (Co-Founder); Matthew Prince (Cloudflare CEO, integration sponsor)
Parent / ownershipCloudflare, Inc. (NYSE: NET)
Specialties50,000+ open-source models, serverless pay-per-prediction billing, Cog containerization, fine-tuning
Funding (pre-acquisition)~$57.8M raised; reported ~$350M post-Series B valuation
Reported ARR (2024)~$40M (estimated, per Latka)4
Paying customers (Dec 2023)30,000; 2 million registered users

  Origins & Founding

Replicate was founded in October 2019 by Ben Firshman and Andreas Jansson, who had first met in 2012 while working together on This Is My Jam, a social music platform. Firshman had gone on to become a product lead at Docker and is credited as the creator of Docker Compose, bringing deep intuitions about developer tooling and containerization to the new company. Jansson had spent years at Spotify building ML research tools and infrastructure, and holds a PhD in ML applied to music — giving the pair complementary expertise in both the infrastructure and research sides of machine learning.5

Their founding insight was that while AI research was accelerating rapidly, developers still lacked a standardized, reproducible way to share, run, and deploy ML models. The same problem that Docker had solved for general software — "it works on my machine" — remained acute in ML, where every model carried an implicit stack of CUDA versions, Python environments, and framework-specific dependencies. Firshman and Jansson built Cog as the foundational primitive: an open-source tool that packages any ML model into a standard Docker-compatible container with all hardware dependencies resolved, along with a prediction API interface. The cloud marketplace came second, layered on top of Cog as the distribution and monetization surface.5

Replicate participated in Y Combinator's Winter 2020 batch (W20), giving it access to YC's network and initial institutional validation before any significant external financing.6

  History & Timeline

    2019–2022: Building the Primitive and Weathering the Surge

Replicate's first years were spent building Cog and establishing the model marketplace with an initial catalog of community-contributed models. The platform operated quietly until August 2022, when Stability AI released Stable Diffusion as an open-weight model. The release sent an immediate, massive traffic spike through Replicate, which had quickly become the easiest way for developers to run image generation models without a local GPU. The company rebuilt significant portions of its infrastructure in weeks to cope with the demand — an operational test that validated both the market demand and the brittleness of the original stack at scale.5

In 2021, Replicate open-sourced Cog, making it the de-facto standard containerization tool for ML models in the open-source community and seeding the supply side of the marketplace by making it easy for researchers and developers to contribute their models.

    2023: Institutional Funding and Explosive Growth

February 2023 brought the announcement of a $12.5M Series A led by Andreessen Horowitz, with participation from Sequoia Capital, Y Combinator, and angels including Dylan Field (Figma) and Guillermo Rauch (Vercel). The round validated Replicate's position as a serious developer infrastructure company and funded continued growth of the GPU fleet and engineering team.7

July 2023 proved to be the company's biggest-ever growth week when Meta released Llama 2 as an open-weight model. Replicate was once again the path of least resistance for developers who wanted to run the model immediately, and the resulting traffic surge established the pattern: every major open-weight model release became a Replicate growth event. The platform's flywheel — more models attract more developers, more developers attract more model contributors — was accelerating.

On December 5, 2023, Replicate announced a $40M Series B led again by Andreessen Horowitz, with new participation from NVentures (NVIDIA's corporate venture arm), Heavybit, and continued backing from Sequoia and Y Combinator.8 The round valued the company at a reported ~$350M. At the time of the announcement, the company reported 2 million registered users and 30,000 paying customers. NVIDIA's participation through NVentures was notable as a strategic alignment between the platform's GPU demand and the chip vendor's interest in expanding consumption of its hardware through developer-friendly abstractions.

    2024: Scaling to 50,000 Models and $40M ARR

Through 2024, Replicate grew its model catalog past 50,000 community-contributed models and its ARR to an estimated $40M (reported via Latka; not audited).4 The platform extended its hardware fleet to include multi-GPU configurations (up to 8× H100) for large model workloads and expanded model coverage into video, music, and speech. Net revenue retention was estimated at approximately 150%, reflecting strong expansion within the developer customer base. Reported headcount as of 2024 was approximately 37 employees — an exceptionally lean team for the scale of infrastructure operated.4

    2025: Cloudflare Acquisition

On November 17, 2025, Cloudflare announced a definitive agreement to acquire Replicate.3 Terms of the acquisition were not disclosed. The announcement framed the deal as a merger of Replicate's model catalog and serverless inference expertise with Cloudflare's global edge network and developer platform — with a specific ambition to make AI inference available at every Cloudflare PoP and to unify inference, storage, compute, and observability into a single developer experience. The acquisition was reported to have closed around December 1, 2025. Ben Firshman's post-acquisition role has not been formally announced by either company as of the time of writing, though some sources report a transition to a role at Anthropic (unverified).9

  What They Offer — Products & Platform

Replicate's platform centers on three interlocking products. The first is the model marketplace: a catalog of 50,000+ open-source ML models contributed by the community and first-party hosted models from labs including Black Forest Labs, Meta, Stability AI, Minimax, ByteDance, and others. Any developer can run any model in the catalog with a single API call, without provisioning infrastructure, managing drivers, or understanding the underlying ML framework. API clients exist for Python and JavaScript, and the REST API works with any HTTP client.2

The second product is Cog itself, released and maintained as open-source. Cog allows any developer or researcher to package their own ML model into a standardized container, publish it to Replicate's marketplace (publicly or privately), and serve it through the platform's inference infrastructure. The combination of a containerization standard and a distribution marketplace makes Replicate the closest analogue to Docker Hub or npm for ML models.

The third tier is dedicated deployments: persistent, always-warm GPU instances for production workloads where cold-start latency is unacceptable. Dedicated deployments differ from the default serverless track in that they bill for setup, idle, and active time — trading cost efficiency for predictable latency. They are targeted at production applications with consistent traffic, while the serverless public-model track is optimized for experimentation and variable-traffic use cases.

Additional platform capabilities include fine-tuning support for customizing public models on proprietary data, private model hosting for organizations that need to run proprietary models without making them publicly accessible, automatic scale-to-zero for public models and scale-from-zero for traffic bursts, and integrated logging and monitoring. Enterprise plans add volume discounts, dedicated account management, and higher GPU quotas.

  Technology & Infrastructure

Replicate's infrastructure is built around a fleet of NVIDIA GPUs provisioned in configurations to match the compute requirements of different model sizes. The available hardware tiers include Nvidia T4, Nvidia L40S, Nvidia A100 (80GB), and Nvidia H100, with multi-GPU configurations available up to 8× H100, 8× A100, and 8× L40S for large frontier models. These configurations allow Replicate to host models across the full spectrum from lightweight image classifiers to 405B-parameter LLMs.2

The software layer is anchored by Cog, which handles containerization, GPU driver compatibility resolution, and the standardized prediction API interface. Every model on the platform runs in a Cog container, which means hardware environment consistency is a platform-level guarantee rather than a per-model responsibility. Cog is maintained as open-source (MIT-licensed) and contributes to the supply side of the marketplace by giving model authors a well-documented, reproducible path to publication.

A significant known weakness in the serverless public-model track is cold-start latency. Community benchmarks from the Latency.io Developer Report 2025 placed average cold-start time for non-warmed public Replicate models at approximately 18.4 seconds — materially higher than platforms with persistent warm model pools such as Baseten or Modal, and considerably higher than dedicated inference APIs.10 This is an intrinsic trade-off of the scale-to-zero model: cost efficiency for low-traffic models comes at the expense of first-request latency. Dedicated deployments address this for production workloads, but at higher cost.

Post-acquisition, Replicate's inference infrastructure is in the process of integration with Cloudflare's global network. The stated roadmap calls for migrating the 50,000+ model catalog to Cloudflare Workers AI and serving inference from Cloudflare's edge PoPs globally — a migration that, if completed, would address the latency disadvantage by placing warm model replicas closer to end users. The timeline for full migration was not confirmed as of mid-2026, and the extent of infrastructure transition completed by this date remains uncertain.3

  Model Catalog & Performance

Replicate's catalog breadth is its primary differentiator. With 50,000+ models spanning image generation, video generation, speech, music, text, code, upscaling, and specialized scientific domains, it is the largest publicly accessible open-source model marketplace by count. The catalog includes first-party hosted models from major labs alongside an enormous long tail of community contributions.

Featured and high-traffic models on the platform include:

ModelTypeNotes
FLUX 1.1 ProImage generationBlack Forest Labs flagship; top-tier output quality
FLUX SchnellImage generationFast variant; 4-step distilled
FLUX DevImage generationDevelopment-grade FLUX variant
Stable Diffusion XLImage generationStability AI; one of Replicate's historically highest-traffic models
Stable Diffusion 3Image generationStability AI latest generation
Llama 3.1 405B InstructLLMMeta; largest publicly available open-weight LLM
Llama 3 70B InstructLLMMeta; high-quality instruction-following
WhisperSpeech-to-textOpenAI open-source ASR; among most-run models on platform
ByteDance Seedream / SeedanceImage / video generationByteDance generative media models
Minimax Music 2.6Music generationHigh-fidelity audio generation
xAI Grok Imagine VideoVideo generationxAI video generation model
Alibaba HappyHorseVideo generationAlibaba generative video

Replicate also serves access to some closed-source models via API pass-through, including models from Anthropic and Google. The catalog quality is variable — as with any open marketplace, models range from production-grade first-party releases to experimental community submissions of uncertain provenance. Replicate does not apply systematic quality gating to community models, which is consistent with its open-marketplace philosophy and a deliberate trade-off against curation overhead.2

  Pricing & Performance Position

Replicate uses two billing tracks depending on the model. Most community models bill by hardware-per-second — you pay only for active inference time, with no charge for idle or cold-start periods. A second track bills on output units (per token, per image, or per video second) for select first-party models. There is no monthly subscription; all billing is consumption-based.2

Hardware-per-second rates (as of mid-2024 published pricing):

HardwarePer-second rateEffective hourly rate
CPU (small)$0.000025$0.09
CPU$0.000100$0.36
Nvidia T4$0.000225$0.81
Nvidia L40S$0.000975$3.51
Nvidia A100 (80GB)$0.001400$5.04
Nvidia H100$0.001525$5.49
8× H100$0.011200$40.32

Private and dedicated deployments add billing for setup and idle time in addition to active inference time. The serverless public-model track has no idle charges, making it cost-effective for low-volume experimentation but potentially expensive per-request for high-volume production workloads where a dedicated deployment or a specialist inference API would offer better unit economics.

The platform's principal performance disadvantage is cold-start latency. At an average of approximately 18.4 seconds for non-warmed public models per community benchmarks, Replicate is materially slower on first request than specialist inference platforms including Baseten (persistent warm pools), Modal (sub-second cold starts for many workloads), and fal.ai (optimized for image model cold starts).10 For output quality and throughput on warmed workloads, Replicate's performance is comparable to similar GPU configurations elsewhere — the hardware is standard and the optimization surface is limited by the generality of the platform.

  People & Leadership

Ben Firshman co-founded Replicate and served as CEO from 2019 through the Cloudflare acquisition in late 2025. His prior experience as a product lead at Docker and creator of Docker Compose gave him a specific and relevant mental model for what a containerization and distribution layer for ML models should look like. His post-acquisition role as of mid-2026 has not been formally confirmed; some sources indicate a transition to a role at Anthropic, but this is unverified.9

Andreas Jansson co-founded Replicate and contributed the ML research and infrastructure expertise to the founding team. His background building ML infrastructure at Spotify, combined with doctoral research in ML for music, oriented Replicate's technology toward the needs of practitioners who understand both the research and engineering sides of model development. Jansson's specific role post-acquisition has not been publicly detailed.

At the time of acquisition, Replicate reported approximately 37 employees — an intentionally lean headcount relative to the scale of infrastructure operated, consistent with a platform-centric business model where Cog's open-source adoption and community contributions substitute for a large professional services or model-development team.4

On the acquirer side, Matthew Prince (Cloudflare CEO and co-founder) has been the public sponsor of the integration, framing the acquisition as central to Cloudflare's strategy of building the most complete AI cloud for developers.

  Funding, Ownership & Business

Replicate raised approximately $57.8M across three disclosed rounds before its acquisition:

RoundDateAmountLead investor
Seed~2020–2022~$5.3MSequoia Capital, Y Combinator
Series AFebruary 2023~$12.5MAndreessen Horowitz (a16z)
Series BDecember 5, 2023$40MAndreessen Horowitz (a16z)

The Series B also included NVentures (NVIDIA's corporate venture arm), Heavybit, Sequoia Capital, and Y Combinator, and valued the company at a reported ~$350M.8 NVIDIA's participation through NVentures created a strategic alignment between Replicate's GPU consumption and NVIDIA's interest in expanding demand for its hardware through developer-accessible abstractions.

Replicate's revenue model is a usage-based markup on GPU costs, with enterprise contracts adding volume discounts and dedicated support. Reported revenue of approximately $5.3M for December 2024 and an estimated $40M ARR for 2024 (per Latka, unaudited) suggest strong growth through the year of acquisition.4 Net revenue retention was estimated at approximately 150%, indicating meaningful expansion revenue from existing customers.

The acquisition price was not disclosed by Cloudflare or Replicate. The transaction was announced November 17, 2025 and closed approximately December 1, 2025, absorbing Replicate into Cloudflare as a wholly-owned unit.3

  Customers & Partnerships

Replicate's paying customer base is weighted toward individual developers and early-stage startups, with growing enterprise adoption as of 2023–2024. As of the Series B announcement in December 2023, the platform reported 30,000 paying customers and 2 million registered users.8

Notable customers and design partners include Character.ai, Photo.ai (PhotoAI), Magnific (AI image enhancement), BuzzFeed, Unsplash, Labelbox, and HeadshotPro — a set that spans consumer generative AI applications, media, and developer tooling. These customers represent the middle tier of the market: organizations large enough to have consistent inference volume but preferring Replicate's simplicity and marketplace breadth over building their own ML ops stack.

Post-acquisition, Cloudflare's own developer ecosystem becomes a primary integration surface. The technical integration roadmap calls for deep interoperability between Replicate's model catalog and Cloudflare Workers, R2 object storage, Vectorize (vector database), Durable Objects (agent state), AI Gateway (observability and cost analytics), and WebRTC/WebSockets for real-time generative applications. This integration effectively makes every Cloudflare developer a potential Replicate customer, extending the addressable market from the existing AI-specialist developer community to the broader Cloudflare Workers ecosystem.3

NVentures (NVIDIA's venture arm) is both a financial backer from the Series B and a strategic partner, given Replicate's role in driving GPU consumption across a wide developer base. This relationship pre-dates the Cloudflare acquisition and continues to inform hardware prioritization.

  Competitive Position

Replicate's competitive identity is breadth and accessibility. No competing platform offers 50,000+ models through a single API; the next largest open-source model marketplace is Hugging Face's Inference API and Endpoints ecosystem, which has a larger community but more operational complexity. Replicate's primary edge is that a developer with no ML ops background can run any model in the catalog in under five minutes.

Its primary weaknesses are cold-start latency (~18.4 seconds average for non-warmed public models) and less operational control compared to platforms designed for production ML engineering. The competitive map:

  • Hugging Face Inference API / Endpoints — largest model community, more self-hosted and bring-your-own-compute options; more complex to operate at scale.
  • Together AI — LLM-focused, competitive per-token pricing, optimized for text inference; narrower model catalog.
  • Baseten — compliance-grade, persistent warm pools, lower cold-start latency; aimed at enterprise ML engineering teams.
  • Modal — arbitrary Python plus GPU, more flexible for custom workloads; requires more ML engineering sophistication.
  • fal.ai — faster cold starts for image models; narrower catalog.
  • RunPod — cheaper raw GPU access for teams comfortable managing their own containers.
  • AWS Bedrock / Google Vertex AI / Azure AI Foundry — hyperscaler managed inference with compliance, SLAs, and deep ecosystem integration; closed-model catalog bias, not open-weight marketplaces.

Post-acquisition, Replicate's long-term competitive position shifts toward the unique combination of model breadth, global edge distribution via Cloudflare's network, and native integration with a developer platform used by millions of existing Cloudflare customers. No pure-play inference API competitor can currently match that combination.3

  Notable Events

August 2022 — Stable Diffusion surge. The open-weight release of Stable Diffusion sent an immediate, massive traffic spike through Replicate, which had become the easiest path for developers to run image generation. The company rebuilt significant portions of its infrastructure within weeks to accommodate the load — an operational test that both validated market demand and revealed the infrastructure investment required to serve as a marketplace for viral model releases.5

July 2023 — Llama 2 launch. Meta's release of Llama 2 as an open-weight model produced Replicate's biggest-ever growth week. The pattern — every major open-weight release becoming a Replicate growth event — confirmed the flywheel model and the strategic value of the marketplace position.5

December 5, 2023 — Series B. The $40M Series B led by a16z with NVIDIA NVentures participation validated Replicate's position in the AI infrastructure market at a reported ~$350M valuation.8

November 17, 2025 — Cloudflare acquisition announced. Cloudflare announced a definitive agreement to acquire Replicate, framing the deal as the foundation of an edge-native, globally distributed AI inference cloud. Terms were not disclosed.3

  Outlook & Roadmap

The post-acquisition roadmap, as communicated by Cloudflare, centers on three integration pillars. First, migrating Replicate's 50,000+ model catalog into Cloudflare Workers AI, making those models accessible from any Cloudflare PoP and removing the centralized GPU fleet bottleneck that drives the current cold-start latency disadvantage. Second, bringing Replicate's fine-tuning capabilities into Workers AI — adding customization to the global inference layer. Third, achieving full-stack developer platform integration: running inference and persisting outputs in R2, triggering inference from Workers and Queues, using Durable Objects for agent state, building real-time generative UIs via WebRTC and WebSockets, and providing unified observability through Cloudflare AI Gateway.3

Replicate retains its brand and API compatibility through the integration; existing customers and integrations continue without disruption. The strategic goal is to make Cloudflare Workers the default end-to-end platform for AI application development — a full-stack, edge-native inference layer that rivals hyperscaler inference APIs on distribution while outcompeting them on developer simplicity and ecosystem integration.

Whether the full migration of 50,000+ models to Cloudflare's edge network will be completed during 2026, and what performance characteristics the integrated platform will achieve, remains to be seen. The technical challenge of distributing and warming that catalog across Cloudflare's network at inference-ready latency is substantial. What is clear is that the combination of Replicate's model breadth and Cog containerization technology with Cloudflare's global network and developer ecosystem represents a genuinely differentiated position in the AI inference market — one that no existing competitor was in a position to replicate at the time of the acquisition.


  References

  1. Replicate — Platform Overview
  2. Replicate — Documentation & Pricing
  3. Cloudflare Press Release — Cloudflare to Acquire Replicate
  4. Latka — Replicate revenue estimates
  5. Sequoia Capital — Replicate Spotlight
  6. Replicate — Y Combinator W20
  7. Replicate — Series A announcement
  8. Replicate — Series B announcement
  9. Ben Firshman post-acquisition role — unverified; LinkedIn (status uncertain as of 2026)
  10. Latency.io Developer Report 2025 — cold-start benchmark

  References

  1. Replicate platform overview — replicate.com.

  2. Replicate documentation, API reference, and pricing — replicate.com/pricing. 2 3 4 5

  3. Cloudflare acquisition announcement — Cloudflare press release, November 17 2025 and Cloudflare blog. 2 3 4 5 6 7 8

  4. Revenue and employee figures are estimated/reported by Latka and are not audited — getlatka.com/companies/replicate.com. 2 3 4 5

  5. Founding story, Stable Diffusion surge, and Llama 2 growth week — Sequoia Capital Replicate Spotlight. 2 3 4 5

  6. Y Combinator W20 batch participation — ycombinator.com/companies/replicate.

  7. Series A announcement — replicate.com/blog/series-a.

  8. Series B announcement — replicate.com/blog/series-b; also Maginative. 2 3 4

  9. Ben Firshman's post-acquisition role is unverified as of mid-2026; some LinkedIn sources indicate a transition to Anthropic but this has not been confirmed by either party and may reflect stale or inaccurate profile data. 2

  10. Cold-start latency benchmark — Latency.io Developer Report 2025; community-reported average of ~18.4 seconds for non-warmed public Replicate models. 2