Command Palette

Search for a command to run...

Fireworks

Enterprise inference and fine-tuning platform built by Meta PyTorch veterans, delivering high-speed production-grade serving of open-weight models via a serverless and dedicated GPU cloud across 400+ models and five modalities.

Fireworks AI

  Executive Briefing

Fireworks AI is an enterprise inference and fine-tuning platform headquartered in Redwood City, California, built on a founding thesis that the future of AI belongs not to a handful of closed model labs but to the thousands of enterprises that need efficient, controllable, and customizable production-grade serving of open-weight models. Founded in October 2022 by seven engineers — all veterans of elite AI infrastructure roles at Meta and Google — the company has grown from a boutique inference API into one of the fastest-scaling platforms in the AI infrastructure market, reportedly processing more than 10 trillion tokens per day across 18 regions as of late 2025.1

The company is led by CEO and co-founder Lin Qiao, who spent seven years at Meta as Senior Director of Engineering, building and later rebuilding the company's entire training and inference stack around PyTorch, overseeing more than 300 engineers. Her six co-founders — James Reed (PyTorch compiler), Dmytro Dzhulgakov (PyTorch core maintainer), Dmytro Ivchenko (PyTorch for ranking), Benny Yufei Chen (Meta ads infrastructure lead), Chenyu Zhao (Google Vertex AI lead), and Pawel Garbacki — form one of the most concentrated teams of production AI systems expertise at any inference company. Their shared background in building the frameworks that the entire industry runs on is the direct source of Fireworks's software differentiation: custom CUDA kernels (FireAttention), proprietary model sharding and quantization (FireOptimizer), and scheduling logic tuned for open-model workloads at scale.2

Commercially, Fireworks has grown at a pace that rivals the most aggressive SaaS trajectories the venture world has seen. Annualized revenue reached $280M at its Series C in October 2025, climbed to a reported $315M ARR by February 2026, and was reported by sources to Bloomberg at approximately $800M annualized as of May 2026 — representing year-over-year growth of more than 400%.34 The platform serves more than 10,000 enterprise customers including Samsung, Uber, DoorDash, Notion, Shopify, Cursor, Perplexity, and Sourcegraph, and has secured strategic partnerships with both AWS (via an AWS Generative AI Competency agreement) and Microsoft (via a multi-year Microsoft Foundry partnership launched in March 2026).56 As of June 2026, Bloomberg reported the company is in talks to raise a new round at a reported $15B valuation, though the round size, series designation, and close date remain unconfirmed.4

Fireworks positions itself as the "TSMC of AI factories" — production infrastructure beneath a broad and fast-moving model ecosystem — and its April 2026 "Own Your AI" and Fireworks Training Preview announcements signal a deliberate expansion up the stack from inference serving into full model training and fine-tuning.7

  At a Glance

ItemDetail
FoundedOctober 2022
TypeInference API platform
HeadquartersRedwood City, CA, USA
StatusActive
LeadershipLin Qiao (CEO), James Reed, Dmytro Dzhulgakov, Dmytro Ivchenko, Benny Yufei Chen, Chenyu Zhao, Pawel Garbacki (Co-Founders)
Parent / ownershipIndependent
SpecialtiesOpen-model serverless inference, custom CUDA kernels, LoRA and reinforcement fine-tuning, multi-cloud GPU scheduling, enterprise compliance
Models served400+ across text, vision, image, and audio (company figure; third-party catalogs vary)
Confirmed funding$327M+ across Seed, Series A ($25M), Series B ($52M), Series C ($250M)
Last known valuation$4B post-money (October 2025 Series C)1
Reported ARR~$800M annualized (May 2026, reported)4

  Origins & Founding

Fireworks AI was founded in October 2022 in Redwood City, California, by seven engineers who had worked together at Meta and Google. The founding team's origin was the PyTorch ecosystem: Lin Qiao led Meta's training and inference engineering org for seven years, overseeing the migration of Meta's production stack from Caffe2 to PyTorch and managing more than 300 engineers. Her co-founders occupied the engine room of that same stack — Reed on the PyTorch compiler, Dzhulgakov as PyTorch's core maintainer, Ivchenko building PyTorch-based ranking systems, Chen leading Meta ads infrastructure, and Zhao running Vertex AI at Google before joining as a co-founder.2

The founding thesis was deliberately contrarian to the model-lab race. Qiao and her co-founders argued that the real bottleneck in enterprise AI adoption was not model quality — open-weight models from Meta AI, Mistral, and others were rapidly closing the gap with closed frontier models — but rather the cost, latency, and operational complexity of running those models reliably at production scale. Their answer was to build software-level optimizations (custom kernels, quantization, scheduling) on top of commodity GPU cloud infrastructure, delivering substantially better price-performance than naively running open-source serving frameworks like vLLM.8

The early product was a serverless inference API for open-source models, requiring no GPU provisioning or model management from the developer. The founding team's depth in production ML systems gave them an unusual ability to extract hardware efficiency without sacrificing output quality, and that asymmetry became the company's durable commercial advantage.

  History & Timeline

    2022–2023: Formation and early product

Fireworks launched in late 2022 and spent most of 2023 building out the core serverless API, establishing its initial model catalog, and onboarding early enterprise customers. Seed funding details are not publicly disclosed. The company began attracting engineers and customers on the basis of its open-model serving speed and ease of use.

    2024: FireAttention, Series A and B, and FireFunction

January 2024 saw the public debut of FireAttention V1, a custom CUDA kernel optimizing multi-query attention for transformer inference, which Fireworks benchmarked as delivering significantly faster throughput versus vLLM on equivalent hardware.8 A Series A of $25M closed in March 2024, providing fuel for infrastructure expansion and engineering hiring. The more significant signal came in June 2024, when FireAttention V2 was published, delivering a reported 12x speedup for long-context inference versus prior approaches and a 4x improvement over vLLM on certain workloads — conditions and baselines for that claim are specified in the company's technical blog rather than a peer-reviewed venue.8

The Series B closed in July 2024 at $52M on a $552M post-money valuation, led by Sequoia Capital with participation from NVIDIA, AMD, Databricks Ventures, MongoDB Ventures, and Benchmark — a backer list that spans both the software and hardware layers of the AI stack.2 The same month, FireFunction-V2 launched, a purpose-built function-calling model benchmarked at GPT-4o parity on multi-tool accuracy at roughly 2.5x the speed and 10% of the cost. On-demand GPU deployment (private GPU instances with autoscaling) launched alongside serverless and reserved-capacity options, completing the platform's three original deployment modes.

In November 2024, Fireworks launched f1, its own compound reasoning model, and a smaller f1-mini, positioned as beating GPT-4o and Claude 3.5 Sonnet on select coding and mathematics benchmarks.9

    2025: Scale, Series C, and platform expansion

By mid-2025 the platform had reached 5 trillion tokens per day across 8 cloud providers and 18 regions via the newly GA Fireworks Virtual Cloud architecture.10 The company launched Fireworks RFT (Reinforcement Fine-Tuning), a managed service that trains models against custom evaluators using reinforcement learning — targeted at enterprises building specialized agents. Deployment Shapes, one-click pre-configured templates pairing a model to a hardware configuration, launched in October 2025.

The Series C closed on October 28, 2025 at $250M on a $4B post-money valuation, co-led by Lightspeed Venture Partners, Index Ventures, and Evantic Capital, with continued participation from Sequoia ($230M primary, $20M secondary).1 By this point Fireworks had surpassed 10,000 customer companies and reached 10+ trillion tokens per day in traffic volume.

    2026: Enterprise partnerships and revenue acceleration

March 11, 2026: Fireworks on Microsoft Foundry launched in public preview, a multi-year strategic partnership enabling Fireworks-served models — including DeepSeek V3.2, Kimi K2.5, and MiniMax M2.5 — via Azure-native endpoints.5 April 6, 2026: Fireworks Training Preview and the "Own Your AI" initiative announced, signaling the company's intent to move from inference serving into full model training.7 April 27, 2026: DeepSeek V4 Pro added to the platform, with third-party measurements indicating Fireworks delivers 5x faster throughput on that model than leading alternatives at the same price point. May 26, 2026: Serverless 2.0 launched, restructuring the serverless tier into Standard, Priority, and Fast service levels under a unified API.11 Bloomberg reported on May 27, 2026 that Fireworks is in talks to raise a new round at a reported $15B valuation, with Index Ventures set to co-lead; the round has not been confirmed closed as of June 2026.4

  What They Offer — Products & Platform

Fireworks operates an inference and fine-tuning platform structured around four deployment modes and an expanding set of model creation tools.

Serverless inference is the core product: pay-per-token access to 400+ models with no GPU provisioning, no cold starts, and multi-cloud failover. Following the May 2026 Serverless 2.0 launch, serverless is offered in three tiers — Standard, Priority, and Fast — selectable per request through a single unified API, letting developers trade cost against throughput and time-to-first-token without changing their integration.11 Cached input tokens receive a 50% discount; batch inference is priced at 50% of standard serverless rates.

On-demand GPU deployment provisions private GPU instances with autoscaling for workloads requiring dedicated capacity, model privacy, or consistent latency guarantees. Reserved capacity provides negotiated long-term GPU contracts for predictable high-volume workloads. Bring Your Own Cloud (BYOC) deploys the Fireworks inference engine inside the customer's own VPC, meeting data residency and sovereignty requirements while preserving the platform's optimization layer.

Fine-tuning is a first-class product rather than an afterthought. The platform supports LoRA SFT, LoRA DPO, full-parameter SFT, and full-parameter DPO fine-tuning, all fed into the same serverless serving infrastructure. The April 2026 Fireworks Training Preview ("Own Your AI") adds full model pre-training to the platform's scope. The most distinctive offering is Fireworks RFT (Reinforcement Fine-Tuning): a managed RL service that trains models against custom evaluators using reinforcement signals, intended for enterprises building specialized agents that need their own reward functions.12

The platform also includes multi-tiered prompt caching with reported 60–90% hit rates, a global scheduler with geographic locality and compliance constraints, and Deployment Shapes (one-click pre-configured model-plus-hardware templates introduced in October 2025) that eliminate the need for manual hardware selection.1

Enterprise compliance covers SOC 2, HIPAA, and GDPR, with zero data retention — no customer data is stored or used to train Fireworks's own models.

  Technology & Infrastructure

Fireworks does not own its GPU fleet. Instead, the company operates a multi-cloud procurement and scheduling layer — branded Fireworks Virtual Cloud — that spans 8 cloud providers and 18 global regions, with the inference engine deployed uniformly across all of them. Hardware available on the platform includes NVIDIA H100 80GB ($7/hr), H200 141GB ($7/hr), B200 180GB ($10/hr), and B300 288GB ($12/hr), plus AMD MI300X. The multi-cloud architecture provides geographic failover, compliance-region routing, and supply-chain resilience without the capital requirements of owning data centers.10

The proprietary software stack is where Fireworks's competitive moat lives:

FireAttention is the company's custom CUDA kernel library for transformer attention. V1 (January 2024) targeted multi-query attention for standard context lengths. V2 (June 2024) focused on long-context inference, with Fireworks reporting a 12x speedup over prior approaches and 4x improvement over vLLM on relevant workloads — note that these figures are from Fireworks's own benchmarks and the specific baseline conditions are detailed in their technical publications rather than independently verified third parties.8 FireAttention implements custom memory management, FP8 and FP16 mixed-precision kernels, and fused operator patterns that reduce memory bandwidth bottlenecks during multi-head and grouped-query attention.

FireOptimizer handles model sharding, quantization strategy selection, and deployment-shape optimization. Together with FireAttention, it enables Fireworks to serve models that would otherwise require more GPU memory, increasing density and reducing per-token cost.

Additional infrastructure capabilities include disaggregated prefill and decode scheduling (separating the compute-intensive prefill phase from memory-bandwidth-bound decode to maximize GPU utilization), semantic caching (caching inference results for semantically similar prompts, not just identical strings), and a global scheduler that routes requests based on geographic locality, service-level tier, compliance constraints, and real-time autoscaling needs.1

At the scale reached by mid-2025, the platform was processing 5 trillion tokens per day; by October 2025 this had grown to 10+ trillion tokens per day at over 100,000 requests per second.1 Gross margin is estimated at approximately 50%, with a stated target of 60% through improved GPU utilization — these are analyst/estimated figures and have not been formally disclosed by the company.3

  Model Catalog & Performance

Fireworks serves 400+ models across five modalities — text, vision, image generation, audio, and code — though third-party catalog counts (such as Artificial Analysis) have measured lower numbers depending on how variants and deprecated models are counted. The company is known for fast day-zero or near-zero launches of frontier open-weight models.

The catalog highlights below are cross-linked where model cards exist on this platform:

Text / reasoning

ModelNotes
DeepSeek V4 ProFrontier open reasoning model; Fireworks reported as 5x faster than leading alternatives
DeepSeek V3.2Available via Microsoft Foundry partnership as well as direct API
Llama 3.1 70B InstructMeta open-weight; high-throughput serverless
Llama 3.1 8B InstructLow-latency, cost-optimized tier
Qwen 3.6 PlusAlibaba open-weight; Mixture-of-Experts
Kimi K2Moonshot AI open-weight agent model
MiniMax M2.5 and M3MiniMax open-weight; available via Microsoft Foundry
GLM-5.1Zhipu AI open-weight
GPT-OSS 120BOpenAI open-weight model

Fireworks-produced models

ModelNotes
FireFunction-V2Purpose-built function calling; 92.1% multi-tool accuracy; reported GPT-4o parity at 2.5x speed, 10% cost
f1Compound reasoning model (November 2024); reported to beat GPT-4o and Claude 3.5 Sonnet on coding and math benchmarks9

Image and audio

ModelNotes
FLUX.1 DevBlack Forest Labs image generation model
Whisper V3OpenAI open-weight speech recognition

Fireworks is notable for serving a high proportion of Chinese open-weight models (DeepSeek, Kimi, MiniMax, GLM, Qwen) alongside the Meta Llama family, giving enterprise customers unusual breadth without vendor lock-in to any single model family.

  Pricing & Performance Position

Fireworks's pricing philosophy is usage-based and publicly listed, with explicit tiers for serverless service level, fine-tuning method, and model size. The figures below are from Fireworks's published pricing and may change; cached tokens receive a 50% discount across all tiers.11

Serverless inference (per million tokens, standard tier)

ModelInputOutput
GPT-OSS 120B / MiniMax M2.5$0.20$0.20
Llama 3.1 8B~$0.20~$0.20
Llama 3.1 70B~$0.90~$0.90
DeepSeek V4 Pro$1.74$3.48
GLM-5.1$4.40

Batch inference is priced at 50% of serverless rates. On-demand GPU pricing ranges from $7/hr (H100 or H200) to $12/hr (B300).

Fine-tuning (per million training tokens)

MethodModel sizePrice
LoRA SFT≤16B$0.50
Full-param SFTlargehigher
Full-param DPO>300B$40.00

Performance claims and third-party context

Fireworks publishes internal benchmarks claiming up to 40x faster inference and 8x cost reduction versus unnamed competitors; these should be treated as upper-bound marketing figures for specific model-hardware configurations. More reliable data points: the fastest model on the platform (GPT-OSS 120B high tier) reaches 664 tokens/second output throughput; MiniMax M2.5 achieves 0.72 seconds time-to-first-token; DeepSeek V3 exceeds 250 tokens/second on the latest hardware configurations. Artificial Analysis's third-party measurements as of 2026 indicate Fireworks delivers approximately 5x faster throughput on DeepSeek V4 Pro at comparable price to leading alternatives.13 The company reports 99.8% uptime and multi-cloud failover as primary enterprise reliability differentiators.

  People & Leadership

The founding team of seven is unusually cohesive: all came from either Meta or Google AI infrastructure roles, most worked together directly, and all remain with the company. This continuity is notable in a market where co-founder departures within the first three years are common.

Lin Qiao (CEO) is the company's primary public voice. Her background spans the full ML infrastructure stack — she was responsible for Meta's production ML systems through the pivotal transition from Caffe2 to PyTorch and the scaling of inference infrastructure to support hundreds of millions of daily users. She has articulated the "TSMC of AI factories" positioning as the company's long-term identity: deep production infrastructure serving a broad ecosystem rather than competing on model research.

James Reed and Dmytro Dzhulgakov bring PyTorch compiler and core maintainer expertise respectively — the kind of systems depth that directly enables the custom kernel work underlying FireAttention. Dmytro Ivchenko's background in ranking-system inference (where ultra-low latency and high throughput are table stakes) maps directly onto the demands of production LLM serving. Benny Yufei Chen's experience in Meta ads infrastructure — one of the world's highest-scale, most latency-sensitive ML serving environments — provides additional credibility in serving at internet scale. Chenyu Zhao brings the enterprise and cloud partnership perspective from her time leading Google Vertex AI.

Headcount is not publicly disclosed. The company has grown from a small founding team through multiple engineering hiring rounds funded by Series A through C capital, but no specific employee count has been confirmed.

  Funding, Ownership & Business

Fireworks is an independent, venture-backed company. Its funding history to date:

RoundDateAmountValuationKey investors
Seed2022–2023UndisclosedUndisclosed
Series AMarch 2024$25MUndisclosed
Series BJuly 2024$52M$552M postSequoia (lead), NVIDIA, AMD, Databricks Ventures, MongoDB Ventures, Benchmark
Series COctober 2025$250M$4B postLightspeed (co-lead), Index Ventures (co-lead), Evantic Capital (co-lead), Sequoia ($230M primary + $20M secondary)

Total confirmed funding exceeds $327M. A further round at a reported $15B valuation was reported by Bloomberg on May 27, 2026, with Index Ventures set to co-lead; this round has not been confirmed closed and the round size, series letter, and final investor composition remain unconfirmed as of June 2026.4

The business model is usage-based B2B SaaS: per-token serverless revenue, per-training-token fine-tuning fees, per-GPU-hour on-demand billing, and negotiated reserved contracts for large enterprise customers. Annualized revenue milestones: $280M ARR at Series C (October 2025), $315M ARR in February 2026 (reported 416% year-over-year growth), and approximately $800M annualized as of May 2026 (reported — not audited figures).34 Blended ARPU across the 10,000+ customer base is estimated at approximately $28K per company, though this is an analyst estimate from Sacra rather than a figure disclosed by Fireworks.3

The NVIDIA and AMD participation in the Series B is strategically notable: both chip manufacturers have an incentive to ensure their silicon is served efficiently and that the open-model ecosystem thrives, making Fireworks a natural infrastructure partner.

  Customers & Partnerships

Fireworks serves more than 10,000 enterprise companies as of October 2025. Named customers span consumer applications, developer tools, enterprise SaaS, and AI-native startups: Samsung, Uber, DoorDash, Notion, Shopify, Upwork, Cursor, Perplexity, Sourcegraph, Cresta, Liner, Quora, Superhuman, Sentient, Genspark, Vercel, Innovative Solutions, and Trilogy. Customer count grew approximately 10x between the Series B (July 2024, ~1,000 companies) and Series C (October 2025, 10,000+).1

AWS: Fireworks holds an AWS Generative AI Competency Partner designation and has signed an AWS Strategic Collaboration Agreement. The platform integrates with Amazon SageMaker AI and Amazon Bedrock AgentCore, enabling AWS customers to access Fireworks-served models within existing AWS workflows.6

Microsoft: A multi-year strategic partnership launched in March 2026 makes Fireworks models available via Microsoft Foundry (Azure AI Foundry) as first-class Azure endpoints. At launch, the partnership covers DeepSeek V3.2, Kimi K2.5, and MiniMax M2.5, with Fireworks providing the inference optimization layer behind the Azure-branded endpoints.5 The dual AWS and Microsoft partnership posture is unusual in the inference market and reduces enterprise lock-in concerns by enabling Fireworks models to be consumed from within existing cloud-spend commitments.

  Competitive Position

Fireworks competes in the open-model inference API market alongside Together AI (reported $150M+ ARR, $305M Series B), Baseten (reported $5B valuation), Replicate, Hyperbolic, and the major cloud providers' managed inference offerings (AWS Bedrock, Google Vertex AI, Azure AI Foundry). On the hardware-differentiation side, Groq and Cerebras attack inference latency from custom silicon (LPU and wafer-scale architectures respectively), delivering ultra-low latency for specific model sizes but with narrower catalogs and different cost profiles.

Fireworks's differentiation rests on five claims: first, software-level performance optimizations (FireAttention, FireOptimizer) that extract better throughput and latency from standard NVIDIA GPUs than generic serving frameworks; second, the broadest open-model catalog in the market (400+ models across five modalities, with fast day-zero launches of frontier open models from Meta, Alibaba, DeepSeek, Moonshot, MiniMax, and Zhipu); third, enterprise reliability and compliance (99.8% uptime SLA, multi-cloud failover, SOC 2, HIPAA, GDPR, zero data retention); fourth, the only major inference platform that pairs serving with a full fine-tuning-to-production pipeline including reinforcement fine-tuning; and fifth, embedded access through both AWS and Microsoft Azure, reducing the friction of enterprise procurement.1256

Groq and Cerebras can outperform Fireworks on raw throughput for certain model/chip combinations but do not serve the breadth of models or modalities that enterprise customers increasingly demand. Together AI is the closest catalog and pricing competitor; the two companies differ primarily on infrastructure approach (Together operates more owned infrastructure) and fine-tuning depth. Hyperscalers bundle inference with broader cloud services but historically lag on open-model performance optimization and day-zero model launches.

  Outlook & Roadmap

Plans articulated at the Series C in October 2025 called for scaling global compute infrastructure 3–4x over the following year, expanding the AI creation toolchain, and investing in tuning and inference alignment research.1 The trajectory since then has been consistent with that direction: the Microsoft Foundry partnership adds a major distribution channel; the AWS collaboration provides another; and the April 2026 Fireworks Training Preview signals the company's intention to compete not just on inference but on the full model development lifecycle — serving, fine-tuning, and now training under one platform.

The "Own Your AI" framing is the most significant strategic signal to date: it positions Fireworks as a counterweight to the closed-model platforms (OpenAI, Anthropic, Google), offering enterprises the ability to own their models' weights, train on proprietary data, and serve from their own cloud environment via BYOC — without sacrificing the performance optimization that Fireworks provides on shared infrastructure.

A pending funding round at a reported $15B valuation (Bloomberg, May 2026) would, if confirmed and closed at that scale, provide capital for substantial GPU procurement, international data-center expansion, and possibly acquisitions in the training or tooling space.4 Leadership has explicitly framed the company's ambition as becoming the production infrastructure layer — the "TSMC" — beneath the AI model ecosystem: a specialist platform that thousands of enterprises depend on, rather than a model lab competing to own the frontier.


  References

  1. Fireworks AI Series C announcement
  2. Fireworks AI Series B and compound AI announcement
  3. Sacra — Fireworks AI company profile
  4. Bloomberg — Fireworks AI in talks for funding at $15B valuation
  5. Fireworks on Microsoft Foundry — Fireworks blog
  6. Fireworks expands AWS alliance
  7. Fireworks Training Preview — "Own Your AI"
  8. FireAttention — serving open-source models 4x faster than vLLM
  9. Fireworks f1 compound AI system
  10. Fireworks Virtual Cloud
  11. Serverless 2.0 — Fireworks blog
  12. Fireworks RFT (Reinforcement Fine-Tuning)
  13. Artificial Analysis — Fireworks provider page

  References

  1. Fireworks Series C announcement, October 2025 — fireworks.ai/blog/series-c. 2 3 4 5 6 7 8 9

  2. Fireworks Series B announcement and team background — fireworks.ai/blog/fireworks-ai-series-b-compound-ai. 2 3 4

  3. Revenue and margin figures are reported/estimated; Sacra company profile — sacra.com/c/fireworks-ai/. 2 3 4

  4. $15B valuation round reported but unconfirmed as of June 2026 — Bloomberg, May 27 2026. 2 3 4 5 6 7

  5. Microsoft Foundry partnership launch, March 2026 — fireworks.ai/blog/fireworks-on-microsoft-foundry. 2 3 4

  6. AWS Strategic Collaboration Agreement — fireworks.ai/blog/fireworks-expands-aws-alliance. 2 3

  7. Fireworks Training Preview "Own Your AI" — fireworks.ai/blog. 2

  8. FireAttention technical blog — fireworks.ai/blog/fire-attention-serving-open-source-models-4x-faster-than-vllm-by-quantizing-with-no-tradeoffs. Performance claims are from Fireworks's own benchmarks; independent verification is limited. 2 3 4

  9. f1 compound reasoning model launch — fireworks.ai/blog/fireworks-compound-ai-system-f1. 2

  10. Fireworks Virtual Cloud GA — fireworks.ai/blog/virtual-cloud. 2

  11. Serverless 2.0 launch, May 2026 — fireworks.ai/blog/serverless-2. 2 3

  12. Fireworks RFT — fireworks.ai/blog/fireworks-rft.

  13. Artificial Analysis provider measurements — artificialanalysis.ai/providers/fireworks.