Command Palette

Search for a command to run...

Together.ai

AI Native Cloud platform offering high-performance serverless inference APIs across 200+ open-source models, dedicated GPU clusters, and fine-tuning infrastructure powered by proprietary kernel research including FlashAttention.

Together AI

  Executive Briefing

Together AI is a San Francisco–based AI infrastructure company positioning itself as the "AI Native Cloud" — a full-stack platform that sits between raw GPU-rental clouds and proprietary hyperscaler APIs. Founded in June 2022, it offers developers and enterprises a unified path from rapid prototyping (via serverless, pay-per-token inference across 200+ open-source models) to dedicated, large-scale training and deployment (via owned GPU clusters running NVIDIA Hopper and Blackwell hardware). The company's founding thesis is that the rise of open-source foundation models represents a generational technology shift, and that the teams best positioned to win will be those who can democratize access to the compute and software stack needed to train, fine-tune, and serve those models at frontier quality and speed.

The company was co-founded by five people whose combined credentials span industry, academia, and open-source research. Vipul Ved Prakash (CEO) built and sold two prior companies — Topsy to Apple and Cloudmark to Proofpoint — and leads the business. Ce Zhang (CTO) came from ETH Zurich, where he led research on distributed ML systems and data management. Chris Ré and Percy Liang are Stanford professors whose labs (Hazy Research and CRFM, respectively) have shaped the data-centric and evaluation sides of foundation model research. In 2023, Tri Dao — the creator of FlashAttention, the algorithm that became a standard component of virtually every modern transformer training stack — joined as co-founder and Chief Scientist. That hire both validated Together's research credibility and brought one of the field's most commercially valuable inference optimization techniques in-house.

Together AI has grown rapidly: from a seed round in late 2022 to reported annualized revenue approaching $1 billion by early 2026, a community of more than one million developers, and a $3.3 billion Series B valuation as of February 2025.12 The company was reportedly in discussions as of March 2026 to raise approximately $1 billion at a pre-money valuation of roughly $7.5 billion — though those talks had not been publicly confirmed as closed at the time of writing.3 The revenue mix is estimated at roughly 30–40% API inference fees and 60–70% GPU rental and cluster revenue, a split that reflects Together's deliberate bridging strategy between the two dominant AI infrastructure business models.4

What sets Together apart from pure-play inference competitors is its vertically integrated software stack. FlashAttention, FlashDecoding, ThunderKittens, ThunderAgent, ATLAS-2, and the newly announced together.compile are proprietary or in-house-led optimizations that compound into meaningful latency and throughput advantages over providers running off-the-shelf serving software. The company also contributes back to the open-source ecosystem — most prominently through the RedPajama training datasets — which has built developer trust and community in a way that purely commercial inference providers struggle to replicate.

  At a Glance

ItemDetail
FoundedJune 2022
TypeInference API + GPU Cloud
HeadquartersSan Francisco, CA, USA
StatusActive
LeadershipVipul Ved Prakash (CEO), Ce Zhang (CTO), Tri Dao (Chief Scientist)
Parent / ownershipIndependent
SpecialtiesOpen-model serverless inference, dedicated GPU clusters, proprietary inference kernels, fine-tuning
Total funding~$533.5M reported across 4 rounds1
Last known valuation$3.3B (Series B, February 2025)1
Developer community1M+ developers as of March 20262

  Origins & Founding

Together AI was incorporated in June 2022 by a founding team with unusually deep roots in both academic ML research and commercial infrastructure. The original four co-founders — Vipul Ved Prakash, Ce Zhang, Chris Ré, and Percy Liang — brought together complementary skill sets: serial entrepreneurship, distributed training research, data management for ML, and the scientific rigor of Stanford's Center for Research on Foundation Models (CRFM).

The founding thesis was precise: a growing compute moat was concentrating the ability to train and deploy large language models inside a small number of well-resourced corporations, and the proliferation of open-source foundation models (led by Meta's LLaMA release in early 2023) represented a generational opportunity to democratize that access. Rather than build proprietary models, Together would build the infrastructure layer beneath open models — making it as easy to run Llama or Mixtral at scale as it is to call an OpenAI endpoint, while also providing the GPU clusters needed to train and fine-tune frontier-scale systems.

In Summer 2023, Tri Dao joined as co-founder and Chief Scientist. Dao had invented FlashAttention — the IO-aware exact attention algorithm that dramatically reduced memory bandwidth requirements for transformer training and inference — while completing his PhD at Stanford. His decision to join Together rather than found an independent lab or join a hyperscaler was a significant signal about the company's research positioning, and his continued work on FlashAttention-2 and the subsequent FlashAttention-4 (announced March 2026) has given Together a durable, defensible software advantage over providers relying purely on off-the-shelf inference frameworks.5

The initial seed round of approximately $20M, led by Lux Capital in late 2022, funded the first iteration of the inference API platform and the early research agenda, including the RedPajama open training dataset project — a collaboration with ETH DS3Lab, Stanford CRFM, Hazy Research, and MILA that produced 1.2 trillion tokens replicating the LLaMA pre-training data.6

  History & Timeline

    2022–2023: Seed, research foundation, and platform launch

Together closed its seed round in late 2022 and moved quickly to release its first developer-facing inference API. In April 2023, the company published the RedPajama-V1 dataset — a 1.2 trillion token open corpus designed to replicate the data used to train LLaMA, enabling the community to build fully open reproduction of Meta's model.6 That contribution established Together's credibility as a genuine contributor to the open-source ecosystem, not merely a commercial consumer of it.

The Series A closed in May 2023. In October 2023, the team released RedPajama-Data-V2, expanding to 30 trillion tokens of multilingual web data — at the time one of the largest publicly available pre-training datasets in the world. The same month, Tri Dao formally joined as co-founder and Chief Scientist. In November 2023, Prosperity7 Ventures and NVIDIA made strategic investments in a Series A extension, pushing Together's cumulative funding above $100M and signaling NVIDIA's bet on Together as a key inference-layer partner.

    2024: Unicorn status, infrastructure deals, and enterprise pivot

Through 2024, Together expanded its infrastructure footprint and platform capabilities. In March 2024, Applied Digital signed a $75M contract to supply GPU clusters to Together AI, onboarding new capacity for the inference and cluster offerings.7 The company grew its developer base to 450,000+ and crossed the $1.25B valuation threshold, achieving unicorn status.

In September 2024, the Together Enterprise Platform launched with AWS Marketplace availability, marking a deliberate push into larger enterprise accounts. The codeless fine-tuning platform was expanded, and the catalog of serverless models grew to over 200. In November 2024, the company announced a partnership with Hypertec to co-deploy a cluster of 36,000 NVIDIA GB200 NVL72 GPUs — one of the largest single Blackwell GPU deployments announced at the time — across European data centers.1 On 12 December 2024, Together acquired CodeSandbox, a cloud-based code execution environment, and used it to launch a built-in Code Interpreter capability directly within the platform.

    2025–2026: Series B, Blackwell deployment, and the AI Native Cloud

February 2025 brought the $305M Series B at a $3.3B valuation, led by General Catalyst and co-led by Prosperity7 Ventures, with participation from Salesforce Ventures, Coatue, Kleiner Perkins, DAMAC Capital, and others.1 The same month, Together launched its GPU Cluster product running NVIDIA Blackwell B200 and GB200 hardware. Together Instant Clusters — self-service, API-first GPU infrastructure for multi-node workloads — reached general availability in September 2025.

On 6 March 2026, Together hosted its first "AI Native Conf" and announced a wave of new capabilities: FlashAttention-4 (claiming up to 4x throughput improvement at long sequence lengths versus FlashAttention-3, per company benchmarks),2 ThunderAgent (3.6x throughput improvement for multi-turn agentic workloads), a Reinforcement Learning API for distributed RL pipelines, ATLAS-2 (1.5x faster inference via real-time data adaptation), and together.compile (a compilation framework for custom model deployment). As of March 2026, Together was reported by The Information to be in talks to raise approximately $1B at a ~$7.5B pre-money valuation, with annualized revenue near $1B.3

  What They Offer — Products & Platform

Together AI's platform is organized around four major surface areas, unified by an OpenAI-compatible API and a single developer account.

Serverless Inference API is the core consumer-facing product: pay-per-token access to more than 200 open-source models spanning text, code, image, video, and audio modalities. The API is fully compatible with OpenAI client libraries, making migration straightforward. Pricing ranges from $0.05 per million tokens for budget text models to $9.00 per million tokens for the largest frontier open models. A Batch API provides asynchronous processing at up to 50% lower cost than synchronous serverless calls, suited to large-scale data transformation, evaluation, and labeling workloads.

Together Instant Clusters (GA September 2025) are self-service, API-first GPU infrastructure packages scaling from single-node 8-GPU configurations to large multi-node clusters with hundreds of interconnected GPUs. Orchestration options include managed Kubernetes and Slurm on Kubernetes; networking uses high-speed InfiniBand; storage options include Weka and VAST high-performance shared filesystems. Observability ships with pre-built Grafana dashboards. The product targets ML engineering teams that need reproducible, large-scale compute for training runs, fine-tuning experiments, and high-throughput inference jobs that exceed what serverless can deliver.

Fine-Tuning Platform supports models above 100B parameters, extended context lengths, Hugging Face Hub integration, and advanced optimization methods including DPO (Direct Preference Optimization) and SimPO. Leading open models including DeepSeek-R1, Qwen3-235B, and Llama 4 Maverick are available for fine-tuning.

Together Reasoning Clusters are purpose-built dedicated infrastructure for token-heavy, latency-sensitive workloads — particularly long-context and multi-step reasoning tasks where shared serverless infrastructure introduces too much variance. The Reinforcement Learning API (announced March 2026) provides distributed RL pipeline infrastructure for teams building post-training and RLHF workflows without managing their own RL scheduler.

The Code Interpreter capability, powered by the December 2024 CodeSandbox acquisition, enables sandboxed code execution natively within API calls — reducing the integration burden for agent and coding-assistant applications. The platform is SOC 2 Type II certified, with HIPAA and GDPR support available for regulated enterprise deployments.

  Technology & Infrastructure

Together AI's technical differentiation is grounded in proprietary inference software built atop owned and contracted GPU infrastructure. The two pillars — hardware fleet and software stack — are increasingly designed to compound rather than operate independently.

Hardware fleet. Together operates across 25+ cities globally. North American capacity is anchored by a 2GW+ portfolio with 600MW of near-term US capacity; European capacity exceeds 150MW across the UK, Spain, France, Portugal, and Iceland. Active GPU generations include the NVIDIA H100 SXM (80GB), H200 (141GB), B200, and GB200 NVL72. The company has secured 200MW of power specifically allocated to NVIDIA Blackwell deployments. Key infrastructure partnerships include: Applied Digital ($75M contract, onboarded March 2024)7; Hypertec (36,000 GB200 NVL72 GPU cluster across European facilities, announced November 2024)1; and 5C Group, deploying NVIDIA B200s in Maryland and GB200/GB300 systems in Memphis, Tennessee. Together AI is transitioning from leasing third-party GPU capacity to owning and operating its own data center infrastructure — a move toward the CoreWeave and Lambda Labs model, expected to improve gross margins from an estimated ~45% (per Sacra analyst estimates, not audited).4

Software stack. The inference optimization stack is where Together's research heritage is most visible:

  • FlashAttention / FlashAttention-2 / FlashAttention-4 — IO-aware exact attention algorithms originated by co-founder Tri Dao. FlashAttention-4 was announced in March 2026, with company-reported claims of up to 4x throughput improvement at long sequence lengths versus FA-3.2 These claims are from internal benchmarks and have not been independently verified at time of writing.
  • FlashDecoding — extends FlashAttention's parallelism to the autoregressive decoding phase, improving throughput on long-context generation.
  • ThunderKittens — a GPU kernel programming framework enabling rapid development of custom CUDA kernels.
  • ThunderAgent — agentic inference optimization reporting 3.6x throughput improvement for multi-turn agentic workloads.2
  • ATLAS-2 — real-time data adaptation layer claiming 1.5x faster inference by adapting model behavior to incoming data distributions.2
  • together.compile — a compilation and deployment framework for custom model configurations.
  • Together Kernel Collection — a curated set of optimized GPU kernels bundled into the cluster product.

Networking across multi-node cluster deployments uses high-speed InfiniBand. The combination of hardware and software is what enables Together's speed claims — reported 2–3x faster inference than hyperscaler competitors overall, with specific benchmark leaders including gpt-oss-20B at 0.45s TTFT and Kimi K2 at over 65% faster than the next-fastest provider, per company figures.

  Model Catalog & Performance

Together's catalog of 200+ open-source models is the broadest available from any single inference-API provider as of mid-2026. The catalog spans text, code, image, video, and audio modalities and is updated continuously as new open-weight models are released. Key models available at time of writing include:

ModelInput (per 1M tokens)Output (per 1M tokens)Notes
DeepSeek V3.1$0.60$1.70High-capability open mixture-of-experts
DeepSeek V4 Pro$2.10$4.40Latest DeepSeek generation
DeepSeek R1$3.00$7.00Reasoning model
GPT-OSS 20B$0.05$0.20Lowest-cost catalog entry; 0.45s TTFT
GPT-OSS 120B$0.15$0.60Compact high-throughput model; 0.54s TTFT
Llama 4 Maverick$0.27$0.85Meta's frontier open model
Qwen 3.6 Plus$0.50$3.00Alibaba Qwen family
Kimi K2.6$1.20$4.50Moonshot AI; claimed 65%+ faster vs. peers
GLM 5.1$1.40$4.40Zhipu AI
MiniMax M2.7$0.30$1.20Efficient mid-range model
MiniMax-M31M token context window
LFM2 24B A2B0.59s TTFT; lightweight fast model

The catalog also includes the full Llama 3.x family, Mixtral and Mistral variants, Qwen 2.5, and a range of image and audio models. Together claims category-leading speeds on several models: up to 2.75x faster on Qwen3 235B 2507 versus the next-fastest provider, and 2x faster on gpt-oss-20B, per company benchmarks as of early 2026.

  Pricing & Performance Position

Together's pricing strategy is deliberately tiered to serve both cost-sensitive developers and latency-sensitive enterprise users. At the bottom of the range, GPT-OSS 20B at $0.05 per million input tokens is among the lowest-priced inference endpoints for capable language models available from any provider. At the top, DeepSeek R1 at $7.00 per million output tokens reflects the premium for running the largest, most capable open reasoning models on optimized hardware.

The Batch API extends the value proposition for offline workloads, providing up to 50% discount versus synchronous serverless pricing for asynchronous jobs — making Together competitive with purpose-built batch inference services for high-volume data processing.

On throughput and latency, Together competes aggressively with Fireworks AI as its closest direct rival. Against hyperscaler inference services (AWS Bedrock, Azure AI Foundry, Google Vertex AI), the company claims 2–3x faster inference overall, attributing the advantage to its custom kernel stack rather than pure hardware differences. Against custom-silicon providers such as Groq and Cerebras, Together trades some raw tokens-per-second speed for dramatically broader model coverage and GPU cluster flexibility. Company-stated gross margins of approximately 45% (an analyst estimate from Sacra, not audited) are expected to improve as Together transitions from leasing to owning infrastructure.4

  People & Leadership

Together AI's leadership bench combines serial entrepreneurship with frontier ML research — an unusual combination that reflects both the product and research ambitions of the platform.

Vipul Ved Prakash (Co-founder & CEO) founded Topsy, a social media analytics company acquired by Apple in 2013, and Cloudmark, an email security company acquired by Proofpoint in 2017. He leads business strategy, fundraising, and go-to-market.

Ce Zhang (Co-founder & CTO) was a professor at ETH Zurich leading the DS3Lab, with research spanning data management for machine learning, efficient model training, and distributed systems. He oversees engineering and infrastructure architecture.

Tri Dao (Co-founder & Chief Scientist) invented FlashAttention while a PhD student at Stanford under Christopher Ré, and subsequently published FlashAttention-2 and FlashAttention-3. His continued work at Together on FlashAttention-4, ThunderKittens, ThunderAgent, and the Together Kernel Collection makes him one of the most directly commercially influential ML researchers in the inference optimization space.

Chris Ré (Co-founder) is a professor at Stanford and MacArthur Fellow whose Hazy Research lab has produced influential work on weak supervision, data programming, and the role of data quality in foundation model performance.

Percy Liang (Co-founder) is a Stanford professor and director of the Center for Research on Foundation Models (CRFM), the group behind the HELM benchmark suite and holistic evaluation standards for language models.

The executive team beyond the founders includes Charles Zedlewski (CPO), Kai Mak (CRO), Meicheng Shi (SVP Finance), Mahadev Konar (SVP Engineering Infrastructure), Albert Meixner (SVP Engineering), Dan Fu (VP Kernels), Max Ryabinin (VP Model Shaping), Arielle Fidel (VP Strategic Partnerships), Jon Fields (VP Sales), Prem Prakash (VP Marketing), and James Barker (VP Sales, EMEA). In February 2026, Alon Gavrielov joined as VP of Infrastructure Strategy, having previously served as VP Infrastructure at Cloudflare — a hire that signals Together's growing emphasis on enterprise infrastructure reliability and global network buildout.

  Funding, Ownership & Business

Together AI is an independent, venture-backed company with no corporate parent. Its reported funding history spans four rounds totaling approximately $533.5M:

  • Late 2022: ~$20M seed round led by Lux Capital.
  • May 2023: Series A (amount not separately disclosed).
  • November 2023: Series A extension led by Prosperity7 Ventures (the venture arm of Saudi Aramco) and NVIDIA, pushing cumulative funding above $100M and valuation to approximately $1.25B.
  • February 2025: $305M Series B led by General Catalyst and co-led by Prosperity7 Ventures, at a $3.3B valuation.1 Co-investors included Salesforce Ventures, Coatue Management, Kleiner Perkins, DAMAC Capital, March Capital, Emergence Capital, SE Ventures, Greycroft, Definition, Cadenza Ventures, and others.

As of March 2026, The Information reported Together was in discussions to raise approximately $1B at a pre-money valuation of $7.5B.3 That fundraise had not been publicly confirmed as closed at time of writing, and both the valuation and revenue figures ($1B ARR) cited in that reporting are from sources not officially confirmed by Together AI.

The business model operates across two streams. API inference fees (estimated 30–40% of revenue) are billed per token against the serverless catalog. GPU cluster and rental revenue (estimated 60–70%) comes from dedicated cluster contracts and reserved capacity deals. The $75M Applied Digital contract and the Hypertec partnership are examples of the supply-side infrastructure that supports this. Together has also signed at least 27 enterprise deals exceeding $1M in value, and one contract reported to exceed $1B in total value (counterparty and terms not publicly disclosed).2

NVIDIA's participation as a Series A strategic investor and its Preferred Partner designation for Together AI reflects the mutual interest: Together is one of NVIDIA's larger inference infrastructure customers and a showcase for Blackwell GPU deployments.

  Customers & Partnerships

Together AI's customer base spans from individual developers to large enterprises. The platform serves more than one million developers and thousands of enterprise customers as of March 2026.2 Publicly named customers include Cursor (AI coding assistant), Decagon (enterprise AI agents), and Cartesia (audio AI). Enterprise customers named in conjunction with the Series B announcement include Salesforce, Zoom, SK Telecom, and The Washington Post.1

Key technology and infrastructure partnerships include:

  • NVIDIA — Preferred Partner; strategic investor; Together is a reference customer for Hopper and Blackwell GPU deployments.
  • Hugging Face — Hub integration for model discovery and fine-tuning workflows.
  • MongoDB — Database integration partnership.
  • Black Forest Labs — Image generation model collaboration.
  • Applied Digital — $75M GPU cluster supply contract (onboarded March 2024).7
  • Hypertec — Co-deployment of 36,000 NVIDIA GB200 NVL72 GPUs across European data centers.1
  • 5C Group — NVIDIA B200 deployments in Maryland and GB200/GB300 systems in Tennessee.
  • AWS Marketplace — Platform available via AWS Marketplace as of September 2024.

Together AI's co-founder Chris Ré's Hazy Research lab and Percy Liang's CRFM maintain ongoing research relationships with the company, and the RedPajama dataset project involved collaboration with ETH DS3Lab, Stanford CRFM, Hazy Research, and MILA.

  Competitive Position

Together AI occupies a distinctive position in the AI infrastructure stack: it is neither a pure-play hyperscaler serving proprietary models, nor a raw GPU rental cloud, nor a custom-silicon inference specialist. Its dual-mode offering — serverless API plus dedicated GPU clusters — allows it to serve developers prototyping at $0.05/M tokens and enterprises running sustained multi-node training jobs under the same account and brand.

In the serverless inference API market, the closest competitor is Fireworks AI, which operates a similar open-model catalog with its own kernel optimization stack. Deepinfra, Replicate, and Hugging Face Inference Endpoints serve overlapping audiences, generally with smaller catalogs or less aggressive pricing. Anyscale and OctoAI (since acquired by Meta) addressed enterprise model serving with more managed-service positioning. Against Groq and Cerebras, which use custom silicon to achieve best-in-class raw tokens-per-second speeds, Together trades some peak throughput for model breadth and the ability to serve the full open-model catalog rather than a curated subset.

In GPU cloud, Together competes with CoreWeave, Lambda Labs, Nebius, and Parasail for dedicated cluster customers. The differentiator here is Together's inference software stack: a team buying GPU capacity from Together also gets FlashAttention-4, ThunderAgent, and together.compile, which are not available from a generic GPU rental provider.

The open-source research contribution strategy — RedPajama, FlashAttention, the Together Kernel Collection — is a deliberate moat-building mechanism that mirrors how companies like Databricks and Hugging Face used open-source to acquire community before converting to enterprise revenue. The combination of research credibility (from the academic co-founders) and enterprise execution (from Vipul Ved Prakash's background) is a rare pairing in the AI infrastructure space.

  Outlook & Roadmap

Together AI is expanding along three strategic axes simultaneously, all of which are reflected in recent hiring and capital deployment decisions.

Infrastructure ownership. The company is actively transitioning from leasing third-party GPU capacity to owning and operating its own data center infrastructure — the same shift CoreWeave and Lambda Labs have made before it. The 200MW of secured Blackwell capacity and the Hypertec partnership for European deployment represent the vanguard of this strategy. Owning infrastructure improves the margin profile (currently estimated ~45% gross margin per analyst estimates) and provides more control over SLA commitments to enterprise customers.

Geographic expansion. European data center presence is being built out through the Hypertec partnership across the UK, France, Italy, Portugal, and Spain, supplemented by additional planned locations in Iceland and Sweden. The EMEA VP of Sales hire (James Barker) signals that revenue follow-up in Europe is a parallel priority.

Product and capability expansion. The March 2026 AI Native Conf announcements — FlashAttention-4, ThunderAgent, the Reinforcement Learning API, ATLAS-2, and together.compile — indicate the company's product roadmap is moving up the stack toward agentic and RL workloads, not just faster token generation. Voice AI and data transformation are cited as additional roadmap directions. The CodeSandbox acquisition and resulting Code Interpreter capability suggest Together is targeting the agentic application developer as a distinct persona alongside the model-serving-focused ML engineer.

The reported $1B fundraise, if completed at or near the $7.5B valuation, would be among the largest single rounds ever raised by an inference infrastructure company.3 That capital would fund continued GPU fleet growth, data center buildout, and international expansion. FlashAttention-4 and together.compile are the current leading bets on proprietary inference optimization as a durable technical moat. Performance claims for these systems are from company sources and should be treated as preliminary pending independent benchmarking.


  References

  1. Together AI $305M Series B announcement — PR Newswire
  2. Together AI AI Native Conf announcements — PR Newswire
  3. Together AI in talks to raise $1B at $7.5B valuation — The Information
  4. Together AI financial analysis — Sacra
  5. Together AI About Us
  6. RedPajama blog post — Together AI
  7. Applied Digital GPU contract with Together AI — IR press release
  8. Together AI CodeSandbox acquisition — Together AI blog
  9. Together AI $1B fundraise report — Data Center Dynamics
  10. Together AI provider benchmarks — Artificial Analysis
  11. Together AI data center locations

  References

  1. $305M Series B at $3.3B valuation, investor list, and customer names from the February 2025 Series B press release — PR Newswire. 2 3 4 5 6 7 8 9

  2. AI Native Conf announcements, FlashAttention-4 and ThunderAgent performance claims, developer count, and enterprise deal metrics are from company PR as of March 2026 — PR Newswire. Performance figures have not been independently benchmarked at time of writing. 2 3 4 5 6 7 8

  3. Reported $1B raise at ~$7.5B valuation and ~$1B ARR figure are from anonymous sources cited by The Information (March 2026) and have not been officially confirmed by Together AI — The Information; see also Data Center Dynamics. 2 3 4

  4. Gross margin estimate (~45%) and revenue split (30–40% API / 60–70% GPU rental) are analyst estimates from Sacra and are not audited figures — Sacra. 2 3

  5. FlashAttention background and Tri Dao's role — Together AI About Us.

  6. RedPajama-V1 dataset release — Together AI blog. 2

  7. Applied Digital $75M contract — Applied Digital IR. 2 3