Novita
Executive Briefing
Novita AI is a San Francisco-based AI infrastructure company offering developer-focused, serverless inference across more than 200 open-source models — spanning large language models, image generation, video generation, text-to-speech, voice cloning, and embeddings — alongside a GPU cloud and, since April 2026, a secure agent sandbox built on Firecracker microVMs. Founded in late 2023 by Frank Lewis (CEO) and Junyu Huang (COO), the company entered a market already crowded with inference providers and carved a position around price leadership, multi-modal breadth, and day-0 availability for newly released open-weight models. Bootstrapped throughout its first two-plus years of operation, Novita reached a reported $1.1 million in annual recurring revenue with a team of roughly ten people — an unusual trajectory for a capital-intensive infrastructure play.1
Junyu Huang brings direct AI infrastructure experience from his time on the Scale AI engineering team, co-founded Toma.com (a voice AI startup backed by Y Combinator's W24 cohort and Andreessen Horowitz), and also worked at Wizard. That pedigree shaped Novita's founding thesis: as the 2023 open-weight model explosion — DeepSeek, Qwen, Llama, Mistral, and dozens of other releases — outpaced the ability of developers to self-host, a serverless, affordable inference layer with a broad catalog would capture significant demand that neither the hyperscalers nor the narrow speed-specialist providers were addressing cheaply. The company launched its first image-generation APIs on Product Hunt in December 2023, added LLM inference in April 2024, and has steadily expanded its surface area since.
By April 2026 Novita had earned two signal moments that raised its profile beyond cost-consciousness alone. First, it was named an official Hugging Face Inference Partner, unlocking a "Deploy on Novita" integration for more than five million Hugging Face developers and making it a day-0 launch partner for Google's Gemma 4.2 Second, independent benchmarker Artificial Analysis ranked Novita the #1 performing and most reliable inference layer across all providers on GPQA Diamond (scientific reasoning), scoring 79.0% and outperforming every other inference endpoint tested — a result that made clear the platform had moved past its early positioning as merely the cheapest option.3
Today Novita operates as what it calls an "AI and agent cloud platform," integrating three product pillars — a Model API marketplace, a GPU Cloud (on-demand instances through to bare metal clusters), and an Agent Sandbox for secure autonomous-agent execution. It serves a developer-first customer base including integrations with Vercel, OpenRouter, Quora, TiDB, Genspark, and Fish Audio, among others, and positions itself as the affordable, multi-modal alternative to more narrowly scoped inference providers.
At a Glance
Origins & Founding
Novita AI was founded in San Francisco in the second half of 2023 by Frank Lewis and Junyu Huang. The founding moment coincided with an inflection point in the open-weight model ecosystem: following Meta's release of Llama 2, the proliferation of Stable Diffusion variants, and the early Mistral and Qwen releases, there was suddenly a large catalog of capable open-source models that developers wanted to call via API rather than operate themselves. The major cloud providers were slow to aggregate these models or price them competitively. Narrow inference specialists like Groq focused on speed using custom silicon rather than breadth or cost. The gap Novita targeted was simple: the most cost-effective, broadest-catalog, developer-friendly API layer for the open-source AI ecosystem, with pricing designed to compete with self-hosting.
Huang's background on the Scale AI infrastructure team gave the founding pair credibility in production AI systems, while his co-founding of Toma.com (a voice AI company, YC W24, backed by a16z) illustrated a pattern of building fast in applied AI before Novita. The company chose to remain bootstrapped — unusual for a cloud infrastructure startup — which focused the team on revenue-positive unit economics from launch rather than growth-at-all-costs. The initial product was image generation: Stable Diffusion-based generation and manipulation APIs, launched on Product Hunt in December 2023. LLM inference followed in April 2024, and the catalog has expanded continuously since.
History & Timeline
2023 — Launch: Image APIs and first developer traction
Novita's first public product was the Image to Video - Motion API, launched on Product Hunt on 25 December 2023.4 Within the following three weeks (January 3–16, 2024) the team shipped Text to Video, AI Image Outpainting, and a Stable Diffusion Reimagine API — establishing a pattern of rapid successive launches across image-manipulation and generation capabilities. These products validated developer demand and generated early ARR from a media-generation audience before the company expanded into language models.
2024 — LLM expansion and catalog breadth
On 30 April 2024, Novita launched its LLM API on Product Hunt, marketing it as "the most reliable, cost-effective uncensored LLMs API."4 The launch extended the platform's reach from image/video generation into the much larger language model market. Through the remainder of 2024 the team expanded the LLM catalog rapidly, adding models from DeepSeek, Qwen, Llama, Mistral, and other open-weight families, while building out GPU cloud products (on-demand instances, serverless GPU, and bare metal) to complement the API offering. By the end of 2024 the company had established integrations with OpenRouter and Vercel AI Gateway, two high-traffic distribution channels for developer-facing inference products.
2025 — Hugging Face listing and ARR milestone
In February 2025 Novita was listed as a serverless inference provider on Hugging Face Hub,5 giving it visibility with the largest open-source AI developer community. Latka data (reported, self-estimated, unverified by official announcement) placed the company at approximately $1.1M ARR with a team of roughly ten employees as of mid-2025.1 This milestone was achieved without external venture capital, relying entirely on usage-based revenue.
2026 — Official partnerships, benchmark wins, and Agent Sandbox
April 2026 represented Novita's highest-profile month to date. On 14 April 2026 the company was named an official Hugging Face Inference Partner — a formal tier above mere listing — enabling a "Deploy on Novita" integration surfaced directly to the platform's 5 million-plus developers.2 Novita was simultaneously announced as the day-0 launch partner for Google's Gemma 4, meaning the model was available on the platform at the moment of its release.2
One week later, on 21 April 2026, Artificial Analysis published benchmarks placing Novita #1 among all inference providers for GPQA Diamond (scientific reasoning) at 79.0%, with 93.3% on AIME 2025 and a #5 ranking on IFBench (68.9%).3 These results were the first time a cost-positioned provider had topped quality rankings for a major benchmark suite, and Novita amplified them via PR Newswire as evidence of simultaneous price and quality leadership.3
On 28 April 2026 the company launched its Agent Sandbox — a secure microVM execution environment for autonomous agents, built on Firecracker technology, with sub-200ms startup times, ephemeral filesystem isolation, a dedicated kernel per task, stateful pause/resume, and per-second billing.6 The launch targeted autonomous agent builders and was positioned around security isolation for workloads like the OpenClaw and Hermes Agent frameworks.
What They Offer — Products & Platform
Novita organizes its offering across three integrated product lines, each addressing a different layer of AI infrastructure.
Model APIs form the core of the platform: serverless, token-priced access to 200+ models across six modalities — LLMs, image generation, video generation, text-to-speech, voice cloning, and embeddings. The LLM API layer is compatible with both OpenAI and Anthropic API formats, enabling zero-code migration for teams already using those SDKs. Key capabilities include day-0 availability for newly released open-weight models, batch inference at a 50% discount over standard pricing, prompt caching, structured outputs, and tool-calling support. The image and video APIs cover generation, inpainting, outpainting, upscaling, and format conversion through models such as Flux, Hunyuan Image 3, Kling v3.0, and the Wan series.
GPU Cloud spans three tiers: on-demand GPU Instances (H200, H100, A100, RTX 5090, RTX 4090, L40S, and RTX 3090 options), Serverless GPU with auto-scaling and no idle cost, and Bare Metal dedicated physical clusters for large-scale training or fine-tuning operations. Spot instances are available at up to 50% below on-demand rates. The fleet connects to a Hugging Face "Deploy on Novita" integration that surfaces Novita instances directly in the model hub workflow.
Agent Sandbox (launched April 2026) provides secure Firecracker microVM execution environments purpose-built for autonomous agents. Each microVM runs with a dedicated kernel, ephemeral filesystem isolation, and stateful pause/resume — properties that matter for long-running, multi-step agentic workloads that interact with external resources. Startup latency is targeted below 200ms; billing is per-second to match bursty, task-driven usage patterns. The product targets frameworks and platforms building autonomous systems that need compute isolation guarantees beyond standard container runtimes.
A Novita Startup Program provides credits and platform access to early-stage teams, functioning as a developer acquisition channel for the next generation of AI-native startups.
Technology & Infrastructure
Novita's infrastructure is built on NVIDIA GPU hardware interconnected with high-bandwidth networking, distributed across a global multi-region footprint. On the hardware side, the GPU fleet includes NVIDIA H200 SXM (141 GB HBM3e, 8× per bare metal node, 1.128 TB total VRAM per node), H100 SXM (80 GB HBM3, 8× per node, 640 GB total), A100 SXM, RTX 5090, RTX 4090, L40S, and RTX 3090 — a range covering the full price-performance spectrum from cost-sensitive single-GPU inference to large-scale multi-node training. Node interconnect uses NVLink 4th Gen (900 GB/s) and 400 Gb/s RDMA networking for GPU-to-GPU communication within clusters.7
On the geographic side, Novita operates at least 16 data center locations across six continents — including the United States (California), United Kingdom, Germany, Australia, Brazil, Japan, Singapore, and India — enabling regional placement for latency-sensitive deployments and data-residency requirements.7 Exact location counts are reported by third-party infrastructure trackers and should be treated as approximate.
The software stack includes an OpenAI-compatible API layer and an Anthropic-compatible API layer, meaning existing applications targeting either SDK can switch providers with a single endpoint change. The Agent Sandbox layer uses Firecracker microVMs, the same open-source virtual machine monitor that Amazon Web Services uses for AWS Lambda, providing hardware-level isolation between tenant workloads with a much lower overhead than full virtual machines. The platform holds SOC 2 certification.8
Performance targets published by Novita include a 200ms time-to-first-token (TTFT) for model APIs — as low as 50ms for some models in claimed configurations — and a 99.5% uptime SLA. Independent Artificial Analysis measurements as of April 2026 observed average LLM throughput of approximately 45–53 tokens per second across the provider's endpoint mix, with the fastest model (Qwen3 35B A3B) reaching 201 tokens per second, average TTFT of 0.95 seconds, and best TTFT of 0.73 seconds.3
Model Catalog & Performance
Novita's catalog spans more than 200 models at the time of writing, with 120+ LLM entries — making breadth one of the platform's primary differentiators versus providers with narrower, curated catalogs. The company prioritizes day-0 availability: new open-weight models are typically available on the platform within hours of their public release.
The LLM catalog includes major families from the leading open-weight labs:
- DeepSeek: DeepSeek V4 Pro, DeepSeek V4 Flash, DeepSeek R1 0528
- Qwen (Alibaba): Qwen3.7-Max, Qwen3 Coder 30B, Qwen3 30B A3B, Qwen3 VL 8B Instruct
- Meta: Llama 3.1 8B, Llama 4 Scout
- Google: Gemma 4 (day-0 launch partner)
- MiniMax: MiniMax-M3
- Moonshot: Kimi K2.6
- MiMo: MiMo-V2.5-Pro
- GLM (Zai-org) and Baidu ERNIE for Chinese-language and multilingual coverage
Image generation is served through Flux (the FOFR/Black Forest Labs family), Hunyuan Image 3, and numerous Stable Diffusion variants. Video generation covers Kling v3.0, Vidu Q3 Turbo, and the Wan series. Audio includes Fish Audio TTS and MiniMax speech-2.6-hd for text-to-speech and voice cloning.
The April 2026 Artificial Analysis benchmarks on the GPT-OSS 120B evaluation suite placed Novita's inference endpoint #1 among all providers on GPQA Diamond (79.0%, scientific reasoning), 93.3% on AIME 2025 (mathematics), and #5 on IFBench (68.9%, instruction following).3 These figures represent quality of served outputs under the Artificial Analysis testing methodology and are a function of both the underlying model weights and the provider's serving configuration.
Pricing & Performance Position
Novita positions itself as the price leader in the multi-modal inference API space, with LLM inference starting at $0.02 per million tokens for models such as Llama 3.1 8B — placing it at or near the floor of the inference market. The pricing range across its LLM catalog spans from $0.02/M tokens to $4.00/M tokens (approximately a 174× spread), with the higher end covering large frontier models such as DeepSeek V4 Pro ($1.60/M input, $3.20/M output) and Qwen3.7-Max ($1.25/M input, $3.75/M output).8
Batch inference is priced at a 50% discount relative to standard (synchronous) pricing, making high-volume offline workloads significantly cheaper. The platform also supports prompt caching for repeat context, further reducing effective costs for applications with stable system prompts or long shared contexts.
GPU cloud pricing is similarly positioned toward the affordable end of the market: H100 SXM at $2.59/hr (on-demand; one source cites $2.89/hr — likely reflecting rate changes or spot vs. on-demand variation), A100 SXM at $1.60/hr, RTX 5090 at $0.72/hr, RTX 4090 at $0.69/hr, L40S at $0.55/hr, and RTX 3090 at $0.21/hr, with spot instances available at up to 50% below those figures.8 Image generation starts at $0.001 per image for standard models; Flux models range from $0.018 to $0.072 per image. The Agent Sandbox is billed per second of microVM execution time.6
Novita claims pricing up to 50% lower than competing inference endpoints across its catalog, a figure that is consistent with public comparisons against Together AI and Fireworks AI for commodity LLM workloads.
People & Leadership
The company is led by its two co-founders. Frank Lewis serves as CEO; detailed public information about his background prior to Novita is limited. Junyu Huang serves as COO and is the more publicly documented of the two: he previously worked on AI infrastructure at Scale AI, co-founded Toma.com (a voice AI startup that was part of Y Combinator's Winter 2024 cohort and backed by Andreessen Horowitz), and worked at Wizard. Huang's combination of large-scale AI infrastructure experience and applied AI startup execution shaped the founding architecture of Novita.
The team is reported at approximately ten full-time employees as of mid-2025,1 which is notably small for a platform operating 200+ models across 16+ global data centers — suggesting significant reliance on managed compute providers, automated tooling, and lean operations rather than a large human workforce for day-to-day serving.
Funding, Ownership & Business
Novita AI is reported to be bootstrapped — operating without external venture capital as of mid-2026.1 This makes it unusual among AI infrastructure startups, virtually all of which have raised significant outside capital to fund GPU acquisition, networking, and data center costs. Latka, a software revenue data aggregator, reports the company at approximately $1.1M ARR (as of roughly mid-2025) and estimates a valuation of approximately $3.3M — figures that are self-reported or estimated by Latka's methodology and have not been confirmed via any announced funding round.1
The business model is usage-based throughout: LLM inference is priced per input and output token, image and video generation per asset produced, audio services per character or second, GPU instances per hour, and the Agent Sandbox per second of execution. This structure aligns Novita's revenue directly with developer usage growth and keeps customer acquisition cost low by removing upfront commitments. The Novita Startup Program — offering credits and access to early-stage companies — functions as a loss-leader for future paid consumption.
At the scale implied by the reported ARR and team size, Novita's unit economics depend heavily on its ability to aggregate demand across a very large model catalog and amortize GPU capacity across many concurrent tenants. The bootstrapped structure suggests the company has achieved sufficient gross margin to sustain operations without external capital, though it limits the speed at which it can expand hardware capacity relative to well-funded competitors.
Customers & Partnerships
Novita's most significant external relationship is with Hugging Face, formalized in April 2026 with official Inference Partner status.2 The partnership surfaces Novita as a serving option directly within the Hugging Face model hub — the primary discovery and deployment interface for the open-source AI community — and enables a "Deploy on Novita" workflow for more than five million registered developers. The simultaneous designation as a day-0 launch partner for Google's Gemma 4 demonstrated that Novita now participates in top-tier model launches rather than being a downstream aggregator.
Beyond Hugging Face, Novita is listed as an inference provider on OpenRouter (a widely used API aggregator allowing developers to route across providers), appears in the top-10 providers on the Vercel AI Gateway (which routes AI calls for applications deployed on the Vercel platform), and maintains integrations with TiDB (distributed database), Genspark (AI search), Quora, Fish Audio (TTS and voice cloning), Kilo Code (AI coding tool), beBee, Gizmo, and Wiz.ai.8 The Novita Startup Program provides credits to early-stage teams and serves as a structured on-ramp for companies that may become larger paying customers.
Competitive Position
Novita competes in the multi-provider inference API and GPU cloud market against Together AI, Fireworks AI, Replicate, DeepInfra, Groq, Cerebras, and RunPod, among others. Its competitive positioning rests on several distinct axes.
Price leadership is the most established differentiator: LLM inference from $0.02/M tokens is among the lowest available in the market, and GPU instance pricing undercuts most branded cloud providers. Multi-modal breadth is the second: most inference competitors focus primarily on LLMs, while Novita aggregates LLMs, image generation, video generation, text-to-speech, and agent sandbox infrastructure into a single billing relationship — reducing integration overhead for teams building multi-modal applications. API compatibility (both OpenAI and Anthropic format support) lowers switching costs for the large developer base already using one of those two SDK conventions. Day-0 model availability addresses a pain point for teams that need to evaluate new releases immediately rather than waiting for providers to onboard them.
Unlike Groq and Cerebras, which compete on raw generation speed using custom silicon, Novita does not claim speed leadership for typical workloads — its independently measured throughput (45–53 tok/s average, up to 201 tok/s for MoE models) is competitive but not fastest-in-class. Unlike Together AI, which has invested heavily in enterprise fine-tuning and training infrastructure, Novita's training offering remains secondary to its inference and GPU cloud products. The April 2026 Artificial Analysis benchmark win on GPQA Diamond introduced a quality narrative that augments the price story, though this reflects serving quality for a specific model at a point in time rather than a systematic architectural advantage.3
Outlook & Roadmap
Novita's strategic direction through 2026 is a deliberate pivot from pure inference API toward what the company calls an "AI and agent cloud platform." The Agent Sandbox launch in April 2026 represents the clearest signal of this direction: rather than competing only on who serves a given LLM most cheaply, Novita is building secure, isolated execution infrastructure for the autonomous agent use case — a higher-value tier of infrastructure that commands a different pricing model (per-second execution) and addresses a differentiated buyer (agent platform builders rather than API consumers).6
Within the inference API business, the Hugging Face partnership is the primary developer acquisition engine, providing passive discovery by millions of developers without direct sales cost. The Vercel AI Gateway integration extends reach into the application deployment layer. Batch inference and prompt caching features are aimed at serving large-scale, cost-sensitive customers who need to process millions of documents or queries offline.
The company's bootstrapped posture implies that near-term priorities are likely organic growth, catalog expansion, and margin improvement rather than a large infrastructure build-out or fundraising milestone. Whether the current model scales to the infrastructure investment required to remain competitive as GPU costs evolve, model sizes grow, and well-funded competitors expand, remains the central strategic question for Novita over the next 12–24 months. Specific roadmap commitments beyond the Agent Sandbox and Hugging Face integration are not publicly disclosed; forward-looking characterizations here should be treated as directional.8
References
- Novita AI — Official website
- Novita AI ranked best performing and reliable inference layer — PR Newswire, April 2026
- Novita AI joins Hugging Face as Official Inference Partner — PR Newswire, April 2026
- Novita AI launches Agent Sandbox — PR Newswire, April 2026
- Novita AI infrastructure overview — GetDeploying
- Novita AI provider benchmarks — Artificial Analysis
- Novita AI ARR and team data — Latka
- Novita AI product history — Product Hunt
- Novita AI on Hugging Face — Novita blog
References
-
ARR, valuation, and team-size figures are reported by Latka, a data aggregator; they are self-reported or estimated and have not been confirmed via official Novita announcements — Latka. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7
-
Hugging Face Official Inference Partner announcement — PR Newswire, April 2026. ↩ ↩2 ↩3 ↩4
-
Artificial Analysis benchmark results, April 2026 — PR Newswire and Artificial Analysis. ↩ ↩2 ↩3 ↩4 ↩5 ↩6
-
Product launch history — Product Hunt. ↩ ↩2
-
Hugging Face Hub listing — Novita blog. ↩
-
Agent Sandbox launch — PR Newswire, April 2026. ↩ ↩2 ↩3
-
Infrastructure data from GetDeploying; exact data center counts are reported by third-party trackers and may differ from current Novita figures. ↩ ↩2