Command Palette

Search for a command to run...

DeepInfra

Purpose-built AI inference cloud offering 190+ open-source models via serverless APIs and dedicated GPU clusters, competing on price and throughput for developers and enterprises building production AI workloads.

DeepInfra

  Executive Briefing

DeepInfra is a purpose-built AI inference cloud headquartered in Palo Alto, California, founded in September 2022 by Nikola Borisov (CEO), Georgios Papoutsis, and Yessenzhar Kanapin — three engineers who previously built and scaled backend infrastructure for imo, a messaging application that served 200 million monthly active users and processed billions of daily messages. That operational experience, including hands-on negotiation and operation of on-premises data centers across multiple countries, gave the founders an unusually concrete view of cloud compute economics and a conviction that hyperscaler pricing was dramatically inflated relative to self-owned infrastructure.1 All three bring competitive programming backgrounds that Felicis Ventures, their earliest institutional backer, described as international olympiad pedigree.2

The company operates what it positions as the inference cloud for the open-source AI era: a serverless pay-per-token API spanning more than 190 open-source models — covering text generation, vision, embeddings, speech recognition, text-to-image, text-to-video, text-to-speech, and multimodal categories — alongside private dedicated GPU deployments, on-demand bare-metal NVIDIA Blackwell instances, and long-term dedicated GPU clusters through its DeepCluster product. All APIs are OpenAI-compatible, and the platform carries SOC 2 and ISO 27001 certifications with zero data retention on the serverless tier. As of May 2026 the platform processes an estimated 5 trillion tokens per week, roughly 30 percent of which originates from agentic systems.3

DeepInfra raised a $107 million Series B in May 2026 co-led by 500 Global and angel investor Georges Harik (an early Google engineer), with NVIDIA, Samsung Next, Supermicro, and existing investors participating.3 The round followed reported revenue that tripled in the first half of 2026 and a 25-fold increase in token volume since the April 2025 Series A — metrics that reflect both the explosive growth in inference demand from reasoning and agentic models and DeepInfra's positioning as consistently one of the cheapest major providers across the 69-provider landscape tracked by Infrabase.4

The company's central thesis — that inference, not training, would become the dominant enterprise AI bottleneck — proved prescient as reasoning models like DeepSeek-R1 and multi-step agentic workflows drove step-change growth in compute requirements. DeepInfra's vertical integration from owned GPU hardware through the API layer, combined with early access to NVIDIA's Blackwell and forthcoming Vera Rubin architectures and integration with NVIDIA's Dynamo distributed inference software, positions it to compete on both price and throughput at a time when efficient token generation at scale is the defining axis of inference competition.

  At a Glance

ItemDetail
FoundedSeptember 2022
TypeInference API / GPU Cloud
HeadquartersPalo Alto, CA
StatusActive
LeadershipNikola Borisov (CEO), Georgios Papoutsis (Co-Founder), Yessenzhar Kanapin (Co-Founder)
Parent / ownershipIndependent
SpecialtiesLowest-cost open-source inference APIs, 190+ model catalog, owned GPU infrastructure, DeepCluster dedicated clusters
Disclosed funding~$133M total across Seed, Series A, and Series B
Notable backersA.Capital Ventures, Felicis Ventures, 500 Global, Georges Harik, NVIDIA, Samsung Next, Supermicro
Weekly token volume~5 trillion tokens (reported, May 2026)3

  Origins & Founding

DeepInfra was incorporated in September 2022, emerging from the infrastructure experience of its three co-founders at imo, a consumer messaging platform with over 200 million monthly active users. At imo, Borisov, Papoutsis, and Kanapin were responsible for building and operating a global backend that handled billions of messages daily — work that required directly negotiating hardware contracts and standing up on-premises data centers in markets where commercial cloud infrastructure was either unavailable or prohibitively expensive.2

That experience shaped the company's founding conviction: large-scale inference compute was structurally overpriced in the cloud, and a team willing to own and operate its own hardware could deliver meaningfully cheaper tokens without sacrificing reliability. The original product concept was broad — hosting a wide catalog of machine-learning models across modalities via a simple API abstraction. The launch of ChatGPT in late 2022 accelerated a pivot toward LLMs specifically, as it became clear that text generation at scale would become the dominant inference workload. The team bet that inference rather than training would be the long-term enterprise bottleneck, a thesis that has held as agentic multi-step workloads and reasoning models have multiplied per-request compute requirements.2

Nikola Borisov (CEO) studied at Northwestern University and worked at Microsoft and HalloApp before joining imo. Georgios Papoutsis holds a physics degree from Technische Universität Berlin and studied at TU München. Yessenzhar Kanapin brings a competitive programming background consistent with the team's described olympiad pedigree. The company's early investors at A.Capital Ventures and Felicis Ventures were attracted specifically by the operational depth the founders brought from imo's infrastructure build-out — a differentiated background relative to most inference API founders entering from pure software or research contexts.2

  History & Timeline

    2022–2023: Founding and seed

DeepInfra was incorporated in September 2022, initially targeting a broad ML model hosting catalog. The team spent early months building the infrastructure layer and model serving platform while iterating on the product thesis. In November 2023, the company closed an $8 million Seed round led by A.Capital Ventures and Felicis Ventures, providing capital to expand GPU capacity and accelerate the model catalog.5 By this point the pivot to LLM-focused inference was fully underway, with the platform beginning to attract developers seeking cheaper and more flexible alternatives to the major cloud providers for open-source model access.

    2024: Scale and market validation

During 2024, DeepInfra expanded its model catalog past 100 open-source models and began accumulating meaningful token volume. 500 Global made an initial investment in the company during this period, ahead of its later co-lead role in the Series B.6 The company continued to scale its owned GPU infrastructure and develop the dedicated deployment products that would eventually become its DeepCluster offering. The inference market was beginning to fragment as developers recognized that open-source models had become competitive with closed alternatives for many workloads, and DeepInfra's cost positioning attracted developer traction.

    2025: Series A and Blackwell

On April 22, 2025, DeepInfra announced an $18 million Series A co-led by Felicis Ventures and Georges Harik, with A.Capital Ventures also participating.7 The announcement cited 8,000-fold growth in processing volume since the seed stage — a figure that illustrates how rapidly the inference market had expanded as reasoning models and agentic workloads emerged. The company also received its first large shipment of NVIDIA Blackwell GPUs during this period, enabling it to offer bare-metal B200 instances and begin developing the next-generation cluster product.7

The arrival of reasoning models — particularly DeepSeek-R1 and similar "thinking" architectures — drove a step-change increase in compute per request as models generated long chains of thought before producing answers. For DeepInfra this represented both a demand surge and a validation of its infrastructure thesis: workloads that generate millions of tokens per session create exactly the kind of sustained, high-throughput compute demand that owned GPU infrastructure handles most efficiently. Agentic workflows began emerging as a distinct category, eventually accounting for roughly 30 percent of the platform's weekly token volume.3

    2026: Series B and DeepCluster

On May 4, 2026, DeepInfra announced a $107 million Series B co-led by 500 Global and Georges Harik, with additional participation from A.Capital Ventures, Crescent Cove, Felicis Ventures, NVIDIA, Peak6, Samsung Next, Supermicro, and Upper90.3 The round brings total disclosed funding to approximately $133 million. The company reported that token volume had grown 25-fold since the Series A and that revenue had tripled in the first half of 2026, though no absolute revenue figures were disclosed. The platform was processing approximately 5 trillion tokens per week across 190+ models at the time of announcement. NVIDIA's participation as a strategic investor reflects both the company's early Blackwell access and its integration with NVIDIA's Dynamo distributed inference software stack.

  What They Offer — Products & Platform

DeepInfra operates four distinct product layers, structured around a land-and-expand motion from self-serve API access through long-term infrastructure contracts.

Serverless shared inference is the entry point: a pay-per-token, OpenAI-compatible API covering 190+ open-source models across text generation, vision, embeddings, rerankers, automatic speech recognition, text-to-image, text-to-video, text-to-speech, text-to-music, OCR, and multimodal categories. The serverless tier carries zero data retention, SOC 2 Type II and ISO 27001 certification, and is accessible without minimum commitments. Pricing is denominated per million tokens for language models and per unit (image, character, or second) for other modalities.

Private dedicated deployments allow customers to deploy custom or fine-tuned models — including LoRA adapter variants — on dedicated GPU capacity billed per GPU-hour rather than per token. This tier is designed for organizations that need consistent capacity, lower per-token costs at volume, or privacy requirements that rule out shared infrastructure. Supported base models for LoRA fine-tuning span Llama, Qwen, Mistral, DeepSeek, and others. Deployment is accessible through the dashboard without requiring direct infrastructure management.

GPU Instances provide on-demand bare-metal NVIDIA Blackwell B200 SXM6 compute in 1x, 2x, 4x, and 8x GPU configurations at $2.79 per GPU-hour, with no minimum commitment and no egress fees.8 Each B200 carries 180 GB HBM3e memory and 7.7 TB/s bandwidth, enabling deployment of large models without the overhead of a managed inference layer. This tier targets teams that want raw compute access for custom serving, experimentation, or workloads that do not fit the serverless model catalog.

DeepCluster is the company's enterprise cluster product: long-term dedicated NVIDIA Blackwell B300 GPU clusters ranging from 256 to 5,000 GPUs, co-located in Tier 3 data centers and operated by DeepInfra under a 99.982% uptime SLA.9 Pricing is structured on 3-year terms at $2.99 per GPU-hour and 5-year terms at $1.98 per GPU-hour, compared to a reference public cloud rate of approximately $6.50 per GPU-hour — an advertised saving of 54 to 70 percent. DeepCluster targets enterprises with predictable, large-scale inference or training workloads willing to commit to longer contract terms in exchange for cost certainty and dedicated capacity.

Fine-tuning via LoRA adapters is supported as a platform-level capability rather than a separate product, accessible across the dedicated deployment tier.

  Technology & Infrastructure

DeepInfra's technology strategy is predicated on vertical integration: owning and operating GPU hardware rather than renting from hyperscalers allows the company to pass the margin differential to customers as lower token prices while maintaining direct control over the hardware configuration, networking, and software stack. The company operates infrastructure across eight U.S. data centers, with EU expansion planned in part to address requirements arising from the EU AI Act.3

The current GPU fleet centers on two NVIDIA Blackwell generations. The B200 SXM6 — deployed in on-demand GPU instances — delivers 180 GB HBM3e, 7.7 TB/s memory bandwidth, and 18 petaFLOPS FP4 throughput, making it well-suited for large-model inference at high batch sizes. The B300, deployed in DeepCluster, extends this with 288 GB HBM3e, 8 TB/s memory bandwidth, NVLink 5 at 1.8 TB/s, and InfiniBand at 800 Gbps per GPU — specifications optimized for distributed multi-node inference of the largest frontier models.9 DeepInfra is an early deployment partner for NVIDIA's next-generation Vera Rubin architecture, though deployment timelines have not been publicly confirmed.

On the software side, DeepInfra has integrated with NVIDIA Dynamo, a distributed inference software framework that the company reports enables up to 20x inference cost efficiency improvements through disaggregated prefill/decode and intelligent request routing across the GPU fleet.3 The platform is optimized for continuous, high-volume token generation rather than minimizing individual request latency — a deliberate trade-off that prioritizes the sustained throughput characteristics of agentic and batch workloads over the first-token speed that specialized hardware like Groq's LPUs optimize for.

Self-reported throughput benchmarks at P50 over 72-hour windows (as of mid-2026) illustrate the platform's performance profile on smaller models: Qwen3.5 0.8B at 403.5 tokens/second with a 0.37-second time-to-first-token; Qwen3.5 2B at 347.6 tokens/second and 0.36s TTFT; Qwen3.5 4B at 250 tokens/second and 0.45s TTFT; GLM-4.7-Flash at 74.6 tokens/second and 0.75s TTFT; Kimi K2 0905 at 77.7 tokens/second and 0.53s TTFT.3 The platform processes approximately 5 trillion tokens per week across all tiers as of May 2026. The engineering team numbers approximately 12, within a company of roughly 20+ people — a notably lean ratio relative to the infrastructure footprint.

  Model Catalog & Performance

DeepInfra's serverless API catalog spans 190+ open-source models, with the text generation tier representing the largest segment. The catalog is updated rapidly as new open-source releases appear, and the company has consistently been among the first inference providers to deploy major new models. Cross-modality coverage — including text-to-image, text-to-video, automatic speech recognition, and text-to-speech — differentiates the platform from inference providers focused narrowly on language models.

Key models available as of mid-2026 include:

Large language models and reasoning models:

  • Llama 4 Maverick — Meta's frontier mixture-of-experts model; $0.12 per million input tokens, $0.30 per million output tokens4
  • DeepSeek-R1 — the influential open reasoning model; $0.55 per million input, $2.19 per million output
  • DeepSeek-V4-Flash — $0.10 per million input, $0.20 per million output; 1M-token context window
  • DeepSeek-R1-0528 — updated reasoning model at $0.50 per million input, $2.15 per million output
  • NVIDIA Nemotron-Ultra-550B — $0.50 per million input, $2.50 per million output
  • Qwen3.5-397B-A17B — $0.45 per million input, $3.00 per million output; 262K context
  • THUDM GLM-5.1 — $1.05 per million input, $3.50 per million output
  • MiniMax-M2 — $0.25 per million input, $1.00 per million output
  • Kimi K2 0905 — available on the serverless tier
  • gpt-oss-120B — approximately $0.08 per million tokens blended

Small and efficient models:

  • Qwen3.5 0.8B, 2B, and 4B — starting from approximately $0.06 per million tokens
  • Llama 3.2 3B, Llama 3.1 8B, and other Llama variants
  • Mistral Small, Phi-4, and Gemma 3 families

Multimodal and specialized:

  • FLUX-2-klein-4b for text-to-image at $0.014 per image unit
  • Qwen3-TTS for text-to-speech at $20.00 per million characters
  • Whisper-family models for automatic speech recognition
  • Phi-4-multimodal and other vision-language models

The catalog's breadth is itself a competitive asset: developers who begin with one model family can access alternatives, embeddings, rerankers, and non-text modalities through the same API and billing relationship without switching providers.

  Pricing & Performance Position

DeepInfra's pricing strategy centers on being the lowest-cost major inference API for open-source frontier models, enabled by the cost structure of self-owned GPU infrastructure rather than rented hyperscaler compute. On Llama 4 Maverick input tokens, the platform charges $0.12 per million — reported as approximately 76% cheaper than Together AI on the same model.4 Small models including Llama 3.1 8B and Mistral 7B start at approximately $0.06 per million tokens. gpt-oss-120B blended pricing sits around $0.08 per million tokens. Independent comparisons from Infrabase.ai, which tracks 69 inference API providers, consistently rank DeepInfra among the cheapest across model categories.4

For dedicated infrastructure, DeepCluster B300 clusters are priced at $2.99 per GPU-hour on 3-year terms and $1.98 per GPU-hour on 5-year terms, against a reference public cloud GPU rate of approximately $6.50 per GPU-hour — an advertised cost reduction of 54 to 70 percent for organizations with predictable long-term workloads.9 On-demand B200 GPU instances at $2.79 per GPU-hour with no minimums and no egress fees provide a middle tier between serverless and long-term commitment.

The platform's throughput profile is competitive for sustained, high-volume generation — the workload type associated with agentic pipelines and batch inference — while it is not optimized for raw first-token latency at low batch sizes, where hardware-specialized providers like Groq and Cerebras hold an advantage. This is a deliberate positioning choice: DeepInfra targets price-sensitive developers and enterprises running continuous high-throughput workloads rather than interactive applications where milliseconds of TTFT are the primary constraint.

  People & Leadership

The company is led by its three co-founders, all of whom remain active. Nikola Borisov serves as CEO, bringing prior experience at Microsoft, HalloApp, and imo, and a computer science background from Northwestern University. Georgios Papoutsis co-founded the company after a career at imo, with an academic background in physics from Technische Universität Berlin and TU München. Yessenzhar Kanapin co-founded the company with an extensive competitive programming background alongside the imo infrastructure experience shared by all three founders.2

The company employs approximately 20+ people overall, with roughly 12 engineers — an exceptionally lean engineering team relative to the scale of infrastructure the company operates and the token volume it processes. This ratio reflects the automation-first approach inherited from the founders' imo experience, where small teams operated infrastructure serving hundreds of millions of users.

Advisory relationships reportedly include executives from WhatsApp and founders from Twitch, Vercel, and Weights and Biases, providing access to operational expertise in scaling consumer and developer-facing platforms.3

  Funding, Ownership & Business

DeepInfra has raised approximately $133 million across three disclosed rounds since its September 2022 founding, remaining independent with no parent organization.

The $8 million Seed round closed in November 2023, led by A.Capital Ventures and Felicis Ventures.5 The $18 million Series A was announced April 22, 2025, co-led by Felicis Ventures and Georges Harik, with A.Capital Ventures participating.7 Note that some databases record a $20.6 million round with a November 2024 date, which may reflect a two-tranche structure or a data error; the company's own announcement cites $18 million. The $107 million Series B was announced May 4, 2026, co-led by 500 Global and Georges Harik, with additional participation from A.Capital Ventures, Crescent Cove, Felicis Ventures, NVIDIA, Peak6, Samsung Next, Supermicro, and Upper90.3

The post-Series B valuation has been reported by Sacra at approximately $125 million, but this figure appears inconsistent with a $107 million round and should be treated as uncertain — likely a stale or pre-money figure, with the true post-money valuation undisclosed.10

The business model operates across three distinct pricing structures: per-token on the serverless tier, per-GPU-hour on dedicated deployments and on-demand instances, and long-term CapEx-style contracts on DeepCluster. This layered structure allows the company to serve developers at low entry costs while building toward larger, more predictable revenue from enterprise cluster contracts. Revenue tripled in the first half of 2026 according to the company's Series B announcement, though no absolute ARR or MRR figure has been disclosed.3

NVIDIA's participation in the Series B as a strategic investor is notable beyond the capital: it reflects DeepInfra's status as an early deployment partner for Blackwell hardware and a Dynamo integration partner, suggesting a degree of supply-chain alignment that may provide preferential access to future GPU generations including Vera Rubin.

  Customers & Partnerships

DeepInfra's publicly named customer base is limited, but the scale metrics — 5 trillion tokens per week, 190+ models, revenue tripling — suggest broad developer and enterprise adoption. The most prominently cited customer is Venice AI, whose CEO Jesse Proudman was quoted in the Series B press release: "DeepInfra gives us access to best-in-class models with the reliability and speed we need to ship."3 Agentic systems, including a platform identified as "OpenClaw," account for roughly 30 percent of the weekly token volume — indicating that the platform is genuinely embedded in production agentic workflows rather than serving primarily one-off API experimentation.3

The strategic partnership with NVIDIA spans supply-chain access (early Blackwell and Vera Rubin hardware), software integration (NVIDIA Dynamo), and the Series B investment. Samsung Next and Supermicro's participation in the Series B also suggest hardware and systems supply relationships beyond their investor roles. 500 Global began as an investor in 2024 before co-leading the Series B, indicating a multi-year conviction in the company's trajectory.6

The target customer segments are developers building AI-powered products on open-source models, scaleup companies needing reliable high-throughput inference at competitive prices, and enterprises willing to commit to dedicated GPU capacity through DeepCluster. The OpenAI-compatible API lowers the switching cost for teams already using the OpenAI ecosystem.

  Competitive Position

DeepInfra competes in the inference API market against a set of peers differentiated on price, speed, model breadth, and infrastructure ownership. Its primary direct competitors are Together AI, Fireworks AI, Baseten, and Replicate in the serverless open-source inference segment. Hardware-specialized alternatives — Groq (LPU-based, optimizing for lowest latency) and Cerebras (wafer-scale, similarly latency-focused) — serve customers for whom time-to-first-token is the primary constraint. Hyperscaler managed services including AWS Bedrock, Azure AI Foundry, and Google Vertex AI offer bundled compliance and ecosystem benefits at higher per-token costs. Distribution aggregators such as OpenRouter and Hugging Face Inference Endpoints create commoditization pressure by layering on top of providers including DeepInfra itself.

DeepInfra's reported 76% price advantage over Together AI on Llama 4 Maverick input tokens is the sharpest expression of its cost-leadership thesis.4 Against Groq and Cerebras, the company trades raw latency for breadth: 190+ models versus a much smaller curated catalog, and throughput optimization versus TTFT minimization. Against hyperscalers, the differentiation is cost and open-source focus — DeepInfra does not host proprietary models and does not impose the procurement overhead typical of enterprise cloud relationships.

In the dedicated cluster market, DeepCluster competes with CoreWeave and Lambda Labs. DeepInfra's differentiation here is its full-stack operation — the same team operating the serverless API operates the clusters, providing a unified platform and potentially tighter integration between inference software optimization and bare-metal hardware management.

The NVIDIA partnership provides a structural advantage on the supply side: early access to each GPU generation translates into a performance and cost lead for the period between a chip's launch and broad availability, which in the current GPU market can span months to years.

  Outlook & Roadmap

DeepInfra has articulated several concrete directions for the capital raised in the Series B. Global compute capacity expansion beyond the eight current U.S. data centers is a stated priority, with EU expansion explicitly mentioned as driven partly by the EU AI Act's data residency and compliance requirements.3 Enhanced developer tooling and support for next-generation open-source and agentic AI models — including NVIDIA Vera Rubin architecture when it becomes available — are also cited as investment areas.

The company's medium-term strategic bet remains on inference compute as the defining AI infrastructure bottleneck as agentic workloads proliferate. Multi-step agent frameworks that chain dozens or hundreds of model calls per user session drive token volumes that are orders of magnitude higher than single-turn chat interactions, and DeepInfra's throughput-optimized infrastructure is designed for exactly this workload profile. The 30% of volume already attributable to agentic systems as of mid-2026 suggests this transition is underway rather than speculative.3

DeepCluster's long-term cluster contracts represent an intentional move up-market into enterprise CapEx relationships, providing more predictable revenue than the per-token serverless business while extending the company's relationship with customers beyond API access into infrastructure partnership. NVIDIA Dynamo integration and continued early access to new GPU generations are expected to maintain DeepInfra's efficiency advantage as the hardware landscape evolves.

Specific timelines for Vera Rubin deployment, EU data center openings, or new product categories have not been publicly confirmed and should be treated as directional rather than committed roadmap items.


  References

  1. DeepInfra — About
  2. Felicis Ventures — Deep Dive: DeepInfra
  3. DeepInfra — Series B Announcement
  4. Infrabase.ai — AI Inference API Providers Compared
  5. Felicis Ventures — Investing in DeepInfra
  6. 500 Global — DeepInfra
  7. DeepInfra — $18M Milestone Blog Post
  8. DeepInfra — GPU Instances
  9. DeepInfra — DeepCluster
  10. Sacra — DeepInfra
  11. GlobeNewswire — DeepInfra Closes $107M Series B
  12. SiliconAngle — DeepInfra $107M Funding

  References

  1. DeepInfra company background and infrastructure thesis — deepinfra.com/about.

  2. Founders' background, imo experience, and competitive programming credentials — Felicis Ventures deep dive. 2 3 4 5

  3. Series B metrics, token volume, revenue growth, agentic workloads, and investor detail — DeepInfra Series B blog post and GlobeNewswire. 2 3 4 5 6 7 8 9 10 11 12 13 14 15

  4. Pricing comparisons vs. Together AI and Infrabase.ai ranking — Infrabase.ai. 2 3 4 5

  5. Seed round details — Felicis Ventures: Investing in DeepInfra. 2

  6. 500 Global's investment history with DeepInfra — 500.co. 2

  7. Series A announcement, 8,000x volume growth, and Blackwell GPU delivery — DeepInfra $18M milestone post. 2 3

  8. B200 GPU instance pricing and specifications — DeepInfra GPU Instances.

  9. DeepCluster specifications, B300 hardware, and pricing tiers — DeepInfra DeepCluster. 2 3

  10. Post-Series B valuation cited by Sacra is reported as approximately $125M and should be treated as uncertain — Sacra.