Command Palette

Search for a command to run...

kluster.ai

Decentralized, developer-first AI inference platform offering OpenAI-compatible serverless and batch inference across open-source LLMs — shut down in 2025–2026 after a pivot to AI code verification.

kluster.ai

  Executive Briefing

kluster.ai was a developer-first AI inference startup founded in 2024 and incubated at Moonsong Labs, a Web3 and AI venture studio rooted in the Polkadot ecosystem. Led by CEO and co-founder Julio Viera, COO Jacob Rettig, and CTO Ryan McConville, the company launched with the conviction that GPU compute was the primary bottleneck constraining AI development and that a decentralized network of GPU providers — coordinated via blockchain-inspired scheduling — could unlock access at costs far below those charged by hyperscalers. Its OpenAI-compatible REST API offered serverless real-time inference, asynchronous webhook-driven jobs, and a distinctive batch mode called "Adaptive Inference" that delivered predictable turnaround times with dynamic rate limits — a combination not commonly available from peers at the time.1

At its operational peak in early 2025, kluster.ai gained industry notice by becoming what it described as the first developer platform to host DeepSeek-R1 on 3 February 2025, pricing the model at $7 per million tokens and claiming savings of up to 95% against competing offerings at the time.2 The platform reported strong early traction: 63% month-on-month growth in real-time inference jobs, 140% growth in batch workloads, and a 36% increase in its developer community in February 2025 alone.3 The company also opened a fine-tuning beta and signed a partnership with decentralized GPU network Aethir in June 2025 to access enterprise NVIDIA H100 and H200 clusters across more than 20 global regions.4

The momentum was short-lived. On 17 July 2025 kluster.ai issued an end-of-life notice for its inference, fine-tuning, and dedicated deployment products, pivoting to a new product called Verify by kluster.ai — an IDE-integrated AI code verification tool for VS Code, Cursor, and Claude Code that detected bugs, security vulnerabilities, regressions, and intent mismatches in LLM-generated code. The company also published an LLM Hallucination Detection Leaderboard on Hugging Face during this period. That pivot did not sustain the business: all kluster.ai services were fully sunset on 9 June 2026, and the team subsequently joined MITO, an AI video-creation startup.5

kluster.ai operated for roughly two years with a reported team of seven people, an estimated $770K ARR at its operational peak (per Latka, unaudited), and reported funding of approximately $4.25M — making it one of the smaller entrants in an inference-API market dominated by well-capitalized players such as Together AI and Fireworks AI. Its arc is instructive for the competitive dynamics of the 2024–2025 inference commoditization wave: aggressive early pricing and a decentralized supply thesis were ultimately insufficient to sustain an independent business against better-funded alternatives.

  At a Glance

ItemDetail
Founded2024
TypeInference API (serverless + batch)
HeadquartersFort Lauderdale, FL, USA (also listed as Miami, FL)
StatusDefunct — services sunset 9 June 2026
LeadershipJulio Viera (CEO), Jacob Rettig (COO), Ryan McConville (CTO)
Parent / ownershipIndependent; incubated at Moonsong Labs
SpecialtiesDecentralized GPU inference, batch/async LLM APIs, OpenAI-compatible endpoints
Reported funding~$4.25M (reported by Tracxn; some sources list as bootstrapped/unfunded)6
Reported ARR~$770K (reported by Latka, unaudited, at peak)
Team at peak~7 people

  Origins & Founding

kluster.ai was founded in 2024 by Julio Viera and incubated through Moonsong Labs, a venture studio and engineering-services firm with roots in the Polkadot blockchain ecosystem.1 Moonsong Labs led the initial ideation, technical exploration, and team assembly. The core premise was that GPU compute scarcity was a structural tax on AI development, and that coordinating a decentralized network of heterogeneous GPU providers — using scheduling logic inspired by blockchain-era distributed-systems thinking — could drive down per-token costs dramatically.

Early technical concepts articulated by the founding team included Tensor Fragments (distributing model shards across multiple nodes), Selective Activation (sparse activation of model components to reduce per-token compute), and a Compute Scheduler that would orchestrate inference jobs across geographically distributed GPUs in real time. Whether these concepts were fully implemented in the shipping product or remained architectural aspirations is not fully documented in public sources. The API launched as fully OpenAI-SDK-compatible, suggesting that developer experience and ecosystem fit were prioritized alongside the novel infrastructure thesis.1

Seed investors included Life Extension Ventures and Protagonist, with Moonsong Labs serving as incubator; whether Moonsong took equity or provided only incubation services is not publicly confirmed. Total reported funding is approximately $4.25M per Tracxn, though some data sources classify the company as bootstrapped or unfunded, which may reflect an unannounced or informal seed arrangement.6

  History & Timeline

    2024 — Launch and early infrastructure

kluster.ai launched its decentralized inference platform in 2024, accepting consumer and professional GPUs (including RTX 4090 class hardware) as supply nodes alongside enterprise-grade chips. The initial model catalog centered on the Meta Llama 3.1 family, with the 8B and 405B parameter variants offered as custom-branded endpoints (klusterai/Meta-Llama-3.1-8B-Instruct-Turbo, etc.). The company introduced "Adaptive Inference," its batch-mode product offering dynamic rate limits and predictable turnaround times at reduced cost — positioning it against both traditional cloud inference and early serverless-API competitors.

New users received $5 in free credits on signup. The base API URL (https://api.kluster.ai/v1) was designed to be a drop-in replacement for the OpenAI SDK, lowering the friction of adoption for developers already building against OpenAI-compatible endpoints.

    Early 2025 — DeepSeek-R1, fine-tuning beta, and peak growth

On 3 February 2025, kluster.ai announced that it was the first developer platform to offer access to DeepSeek-R1, the open-weight reasoning model from Chinese lab DeepSeek that had attracted enormous industry attention.2 The model was priced at $7 per million tokens, with the company claiming savings of up to 95% relative to comparable reasoning-model offerings. A promotional campaign offered $100 in credits (approximately 167 million tokens) and a 1:1 credit match for deposits of $10 or more through 10 February 2025.2

The same month, kluster.ai opened a fine-tuning beta and reported striking month-on-month growth: 63% in real-time inference jobs, 140% in batch workloads, and 36% in developer community size.3 The company's catalog expanded to include Llama 3.3 70B, Llama 4 Maverick, Qwen3-235B-A22B, and Google Gemma-3. By mid-2025 the company reportedly had seven employees and an estimated $770K ARR (per Latka; this figure is self-reported and unaudited).

    June 2025 — Aethir partnership and infrastructure upgrade

On 17 June 2025, kluster.ai announced a partnership with Aethir, a decentralized GPU-as-a-Service platform, to access enterprise bare-metal NVIDIA H100 and H200 clusters across more than 20 global regions.4 The Aethir partnership addressed a key criticism of decentralized inference networks — variability in hardware quality and latency — by providing access to enterprise-grade, non-virtualized compute with no resource oversubscription and no hidden egress fees. The deal represented a hybrid approach: decentralized scheduling logic on top of a more conventional enterprise GPU supply layer.

    July 2025 — End-of-life and pivot to Verify

On 17 July 2025, kluster.ai issued a public end-of-life notice for its inference platform, fine-tuning service, and dedicated deployment product.5 The company simultaneously pivoted to Verify by kluster.ai, an IDE-integrated tool for detecting bugs, security vulnerabilities, regressions, and intent mismatches in AI-generated code, distributed as extensions for VS Code, Cursor, and Claude Code. A companion LLM Hallucination Detection Leaderboard was published on Hugging Face to support developer trust-evaluation workflows.

    2026 — Shutdown and team transition

All kluster.ai services were fully sunset on 9 June 2026. The kluster.ai team subsequently joined MITO, an AI video-creation startup, in what appears to have been an acqui-hire or voluntary team transition; formal acquisition terms were not publicly disclosed. The Verify product was also discontinued as part of this transition.

  What They Offer — Products & Platform

kluster.ai's core offering was an OpenAI-compatible REST API with three inference modes designed to serve different latency-cost trade-offs. Real-time inference targeted sub-second latency workloads where responsiveness was paramount. Asynchronous inference supported webhook callbacks for applications where immediate response was not required, allowing the platform to schedule jobs across its distributed GPU network with greater efficiency. Adaptive Inference (batch mode) was positioned for large-scale, cost-optimized LLM jobs where developers could tolerate variable turnaround in exchange for significantly lower per-token pricing and dynamic rate limits that scaled with demand.

Beyond the core inference API, the platform offered a fine-tuning service (launched in beta in early 2025), enabling developers to customize models on proprietary data, and dedicated deployment instances for teams requiring isolated compute. The company published comprehensive documentation and integration guides, with a notable partnership integration for Bespoke Labs' Curator library enabling batch inference workflows out of the box.7

In mid-2025 the company launched Verify by kluster.ai, repositioning toward developer tooling rather than infrastructure. Verify integrated into developer IDEs and applied AI analysis to flag potential issues in LLM-generated code before it shipped — targeting the growing segment of developers who relied on AI coding assistants and needed a verification layer. The product also included the LLM Hallucination Detection Leaderboard on Hugging Face, surfacing model reliability comparisons for developer reference.8

  Technology & Infrastructure

kluster.ai's infrastructure was built on a decentralized, distributed GPU network that accepted heterogeneous hardware — from consumer GPUs (RTX 4090 class) to professional data-center chips — as supply nodes. The architecture drew on coordination concepts from blockchain-era distributed systems, with Moonsong Labs' Polkadot background informing the distributed-scheduling philosophy. The software stack was OpenAI SDK-compatible end-to-end, with adaptive resource scheduling, dynamic rate limits, and data encryption cited as core features of the platform.1

Specific scale metrics — total node count, aggregate GPU capacity, geographic distribution of supply nodes — were not publicly disclosed at any point during the company's operation. The lack of transparency around infrastructure scale was a notable constraint on independent verification of the platform's claimed performance and redundancy characteristics.

The June 2025 partnership with Aethir marked a significant infrastructure evolution.4 Aethir's decentralized GPU-as-a-service network provided access to enterprise bare-metal NVIDIA H100 and H200 clusters across more than 20 global regions, with no virtualization overhead, no resource oversubscription, and no hidden egress fees — addressing the quality and latency variance that could affect consumer-grade distributed inference. This hybrid model (decentralized scheduling over enterprise bare-metal supply) represented a pragmatic response to the performance expectations of production API customers.

The original "Tensor Fragments" (model-shard distribution) and "Selective Activation" (sparse component activation) concepts articulated at founding were not confirmed as shipping features in public documentation; the production platform's architecture more closely resembled conventional distributed inference over heterogeneous GPUs, coordinated by the Compute Scheduler.

  Model Catalog & Performance

kluster.ai served a catalog of open-weight models across the Meta Llama, DeepSeek, Google Gemma, and Qwen families. Models were exposed under both their original identifiers and custom kluster.ai-branded endpoint names:

ModelEndpoint / IdentifierNotes
Meta Llama 3.1 8B Instructklusterai/Meta-Llama-3.1-8B-Instruct-TurboCore catalog; fast, low-cost
Meta Llama 3.3 70B Instructklusterai/Meta-Llama-3.3-70B-Instruct-TurboMid-tier capability/cost
Meta Llama 3.1 405B Instructklusterai/Meta-Llama-3.1-405B-Instruct-TurboLargest Llama served
DeepSeek-R1deepseek-ai/DeepSeek-R1First hosted Feb 3, 20252
Llama 4 Maverickmeta-llama/Llama-4-MaverickAdded in 2025
Qwen3 235B A22BQwen/Qwen3-235B-A22BLarge MoE model
Google Gemma-3google/Gemma-3Added in 2025

The platform claimed sub-second latency for real-time inference jobs. Specific throughput metrics (tokens per second) and independent benchmark results were not published. The full model catalog at the time of shutdown in July 2025 was not comprehensively documented in public sources.

  Pricing & Performance Position

kluster.ai's pricing strategy was explicit and aggressive: the company positioned itself as the lowest-cost route to open-weight model inference, with marketing claims of up to 95% savings relative to hyperscaler inference pricing (benchmarked against GPT-o1 at the time of the DeepSeek-R1 launch in February 2025).2 The flagship data point was DeepSeek-R1 at $7 per million tokens, which the company cited as dramatically undercutting competing access points for a comparable reasoning-capable model.

Llama model pricing was described as offering "up to 50% savings vs. competitors," but specific per-token input and output rates for the Llama family were not publicly disclosed in the company's documentation or press materials. New users received $5 in free credits on signup, lowering the barrier to first-use evaluation. A promotional period in February 2025 offered a 1:1 credit match for deposits of $10 or more, and a separate DeepSeek-R1 promotion provided $100 in credits (approximately 167 million tokens).2

The business model was consumption-based pay-per-token, with no disclosed minimum commitment for the self-serve tier. Dedicated deployment instances likely carried different pricing arrangements, but terms were not published. The free-tier entry point and aggressive promotional credits were consistent with a growth-over-margin strategy aimed at building developer adoption during the platform's early phase.

  People & Leadership

kluster.ai operated with a lean team of approximately seven people at its operational peak. The core leadership comprised:

  • Julio Viera — CEO and Co-founder. Led the company from founding through shutdown. Background in Web3 and AI infrastructure through his work at Moonsong Labs.
  • Jacob Rettig — COO. Responsible for operations and business execution.
  • Ryan McConville — CTO. Led technical architecture and engineering.
  • Anjin Stewart-Funai — Head of Marketing. Led developer community growth and communications.

Headcount details beyond the reported figure of seven are not publicly documented. The team's 2026 transition to MITO (AI video creation) suggests the core group remained intact through the wind-down period.

  Funding, Ownership & Business

kluster.ai's capitalization was modest relative to its ambitions and the competitive landscape it entered. Reported investors included Life Extension Ventures and Protagonist as seed-round backers, with Moonsong Labs serving as incubator and potentially as a source of initial capital, though whether Moonsong took equity or provided only incubation services is not publicly confirmed.6

Total reported funding is approximately $4.25M per Tracxn. Some data aggregators classify the company as bootstrapped or unfunded, which may reflect the absence of a formally announced seed round rather than a genuinely self-funded operation. Latka reported an estimated valuation of approximately $2.3M — a figure the platform derives algorithmically from revenue multiples and is not a confirmed, negotiated valuation.6 The company's reported $770K ARR (per Latka, unaudited) at its operational peak in early 2025 represents a consumption-based revenue run-rate consistent with a small but growing developer-focused API business.

The consumption-based business model — pay-per-token with free credits for new users — is standard for inference-API providers. At $770K ARR with seven employees and approximately $4.25M raised, kluster.ai was operating in a capital-efficient mode but well below the scale required to compete durably against peers like Together AI (which raised $300M+) or Fireworks AI (which raised $52M+).9

  Customers & Partnerships

kluster.ai's most documented customer integration was with Bespoke Labs, whose Curator library published official documentation for using kluster.ai as a batch inference backend — a notable endorsement that placed kluster.ai alongside better-known providers in a developer-facing integration guide.7 No other named enterprise customers were publicly disclosed during the company's operation.

On the supply side, the Aethir partnership (June 2025) was the platform's most significant business relationship, providing access to enterprise-grade GPU infrastructure at scale across 20+ global regions.4 Early community discussions also explored kluster.ai's position relative to OpenRouter (an inference aggregator), with developers comparing the two as alternative access points for open-weight models. No formal OpenRouter integration was announced.

  Competitive Position

kluster.ai entered one of the most contested markets in the 2024–2025 AI landscape: serverless inference APIs for open-weight LLMs. Its direct competitors included Together AI, Fireworks AI, Groq, Replicate, DeepInfra, and Hyperbolic, among others. The company's differentiation rested on four pillars: a decentralized GPU supply network that purported to reduce infrastructure costs structurally; aggressive per-token pricing (up to 95% cheaper than hyperscaler equivalents); the Adaptive Inference batch mode with dynamic rate limits; and early-mover positioning on high-demand models, most notably being the first platform to offer DeepSeek-R1 access.2

The central challenge was scale and capital. Well-funded competitors had raised orders of magnitude more capital and could sustain lower margins, absorb infrastructure costs, and invest in optimizations (custom kernels, speculative decoding, quantization, purpose-built silicon) that meaningfully differentiated performance. The decentralized GPU supply thesis, while potentially compelling as a long-run cost story, introduced quality and reliability variables that enterprise customers were unlikely to tolerate without extensive mitigation — which the Aethir partnership began to address only months before the platform's shutdown. At seven employees and $4.25M raised, kluster.ai lacked the organizational surface area to execute fast enough across pricing, performance, reliability, and go-to-market.9

  Notable Events

  • 3 February 2025 — kluster.ai announced as the first developer platform to host DeepSeek-R1, priced at $7 per million tokens with up to 95% claimed savings vs. competitors.2
  • February 2025 — Fine-tuning beta opened; platform reported 63% MoM growth in real-time inference jobs, 140% in batch workloads, 36% growth in developer community.
  • 17 June 2025 — Partnership with Aethir for enterprise H100/H200 GPU infrastructure across 20+ regions announced.4
  • 17 July 2025 — End-of-life notice issued for inference, fine-tuning, and dedicated deployment services.5
  • Mid-to-late 2025 — Pivot to Verify by kluster.ai (IDE-integrated AI code verification) and publication of LLM Hallucination Detection Leaderboard on Hugging Face.8
  • 9 June 2026 — All kluster.ai services fully sunset.
  • 2026 — kluster.ai team joins MITO (AI video creation startup).

  Outlook & Roadmap

kluster.ai did not survive as an independent inference provider. The company's trajectory — aggressive early pricing, rapid growth in early 2025, a hardware partnership to address infrastructure quality concerns, followed within weeks by an end-of-life announcement and pivot — reflects the structural difficulty facing small, undercapitalized inference APIs in a market where the cost of compute, the speed of model releases, and the pricing pressure from well-funded competitors all move faster than a seven-person team can respond.

The pivot to Verify represented a product-strategy bet that developer tooling layered on top of LLMs might offer more durable differentiation than inference infrastructure alone. That bet did not generate sufficient momentum to sustain the business, and the team's transition to MITO (AI video creation) in 2026 marked the end of both the inference platform and the Verify product.

No forward roadmap exists. The kluster.ai domain and services are inactive as of 9 June 2026.


  References

  1. Moonsong Labs — kluster.ai founding blog post
  2. PRWeb — kluster.ai first to host DeepSeek-R1 (Feb 3, 2025)
  3. AI Thority — kluster.ai DeepSeek-R1 announcement
  4. Aethir ecosystem blog — kluster.ai partnership
  5. kluster.ai — End-of-life announcement
  6. Tracxn — kluster.ai company profile
  7. Bespoke Labs Curator docs — kluster.ai batch inference
  8. Hugging Face — kluster.ai LLM Hallucination Detection Leaderboard
  9. Latka — kluster.ai revenue and team data

  References

  1. kluster.ai founding thesis and technical architecture — Moonsong Labs blog. 2 3 4

  2. DeepSeek-R1 hosting announcement and pricing — PRWeb, Feb 3 2025; AI Thority. 2 3 4 5 6 7 8

  3. February 2025 growth metrics — AI Thority. 2

  4. Aethir partnership (June 17, 2025) — Aethir ecosystem blog. 2 3 4 5

  5. End-of-life and shutdown — kluster.ai blog. 2 3

  6. Funding and valuation figures are reported estimates from Tracxn and Latka and have not been publicly confirmed by the company — Tracxn; Latka. 2 3 4

  7. Bespoke Labs integration — Bespoke Curator docs. 2

  8. Verify by kluster.ai and Hallucination Leaderboard — Hugging Face. 2

  9. Competitive context and funding comparisons are drawn from public reporting; Together AI and Fireworks AI funding figures are widely reported estimates. 2