Command Palette

Search for a command to run...

Simplismart

India-founded AI inference and MLOps platform (Verute Technologies Pvt Ltd) delivering software-optimized, model-agnostic inference for enterprises and cloud providers on NVIDIA GPU infrastructure, with a proprietary megakernel engine claimed to outperform vLLM and SGLang.

Simplismart

  Executive Briefing

Simplismart (legal entity: Verute Technologies Private Limited) is a Bengaluru-founded AI inference and MLOps platform that delivers software-optimized, model-agnostic inference for enterprises and cloud providers on NVIDIA GPU infrastructure. The company was co-founded in 2022 by Amritanshu Jain (CEO, ex-Oracle Cloud and Capillary Technologies) and Devansh Ghatak (CTO, ex-Google Search), two BITS Pilani alumni who started with an AutoML ambition and discovered, mid-build, that the inference engine they had assembled as a side capability was the genuinely competitive product. That pivot — executed with less than $1 million in seed capital — became the founding thesis of what is now one of India's most closely watched AI infrastructure startups.1

The company's core technical claim is a proprietary megakernel inference engine that fuses an entire LLM forward pass into a single GPU kernel, eliminating idle GPU time between the hundreds of sequential CUDA operations typically required. Benchmarked internally, the engine reportedly reaches approximately 78% GPU bandwidth utilization (up from under 50% with vLLM), claims 2.5× faster decoding throughput than vLLM and 1.5× faster than SGLang on LLaMA-1B, and drives LLaMA-1B single forward-pass latency to under 1 ms on H100 and 0.68 ms on B200.2 On top of this engine, Simplismart has built four branded product lines — SimpliLLM, SimpliScribe, SimpliDiffuse, and SimpliSpeak — serving text, speech, image, and voice workloads respectively.

Simplismart's commercial trajectory accelerated sharply around its October 2024 Series A: a $7 million round led by Accel, with Shastra VC, Titan Capital, and angel investor Akshay Kothari (co-founder of Notion) participating.3 By that point the company reported over 30 enterprise customers and approximately $1 million in annual recurring revenue. In February 2026 it extended its reach into a new distribution layer, launching a white-label advanced inference platform for NVIDIA Cloud Partners (NCPs) that integrates NVIDIA Inference Microservices (NIMs) and pre-built workflow templates.4 As of mid-2026, Simplismart is reportedly in advanced discussions for a further ~$20 million round that would be led by NVIDIA — a round that, if closed, would mark a significant strategic endorsement from the GPU incumbent at a reported post-money valuation of approximately $100 million.5

Simplismart operates in direct competition with inference API platforms such as Together AI and Fireworks AI, differentiating on software-level throughput optimization, enterprise deployment flexibility (serverless, dedicated, bring-your-own-cloud, fully on-premises / air-gapped Kubernetes and Slurm), and a cost structure anchored in its India-based engineering organization. Its security posture — ISO 27001 v2022, SOC 2 Type II, GDPR, and AICPA compliance — enables deployment in regulated enterprise and government contexts where data sovereignty is non-negotiable.

  At a Glance

ItemDetail
Founded2022 (as Verute Technologies Pvt Ltd)
TypeInference API / MLOps platform
HeadquartersBengaluru, India (San Francisco, CA, USA presence)
StatusActive
LeadershipAmritanshu Jain (CEO), Devansh Ghatak (CTO)
Parent / ownershipIndependent
SpecialtiesMegakernel LLM inference, serverless + dedicated GPU APIs, white-label NCP inference, on-prem/air-gapped deployments, multimodal API suite
Total funding raised~$8 million (reported, pre-NVIDIA round)6
Reported post-Series A valuation~$25 million3
Prospective Series B~$20 million at ~$100 million valuation (not yet confirmed as of June 2026)5
Security certificationsISO 27001 v2022, SOC 2 Type II, GDPR, AICPA

  Origins & Founding

Simplismart was founded in 2022 in Bengaluru by Amritanshu Jain and Devansh Ghatak, both graduates of the Birla Institute of Technology and Science (BITS) Pilani.1 Jain had previously worked on cloud infrastructure and machine-learning pipelines at Oracle Cloud and as an ML engineer at Capillary Technologies, while Ghatak brought search algorithm expertise developed at Google Search. Their shared experience with the practical friction of deploying machine-learning models in production — the gap between a research artifact and a reliable, low-latency production endpoint — seeded the initial idea for an AutoML and no-code ML platform.

The pivot to inference infrastructure came from a customer conversation. While demonstrating their AutoML platform to a prospective client, Jain and Ghatak realized that the inference engine they had built as a supporting component was substantially faster than anything the prospect had seen. The competitive advantage was not in the no-code interface they had set out to build — it was in the inference engine they had built to power it. The founders reframed the company accordingly: inference engine speed was the product, and everything else was built on top of it. This reorientation was accomplished before significant outside capital had been deployed, with the initial pivot funded by under $1 million in seed financing.7

Early development was validated through hackathon wins and an admission to the NVIDIA Inception Program, which provided early access to hardware and go-to-market support. The company launched four named product lines in 2023, secured early design-partner customers including healthcare platform Tata 1mg and AI video creation platform InVideo, and built toward its first meaningful financing event in late 2024.

  History & Timeline

    2022–2023: From AutoML to Inference Infrastructure

Simplismart launched as an AutoML and no-code ML platform before pivoting to AI inference middleware after recognizing that the internal inference engine was the company's genuine competitive differentiator. During this period the team raised less than $1 million in seed financing, won multiple hackathons, joined the NVIDIA Inception Program, and released the first versions of its four product lines — SimpliLLM, SimpliScribe, SimpliDiffuse, and SimpliSpeak. Early design-partner relationships were established with enterprise customers in healthcare, e-commerce, and media, validating the inference-as-a-service thesis.

    2024: Series A and Public Platform Launch

In October 2024, Simplismart announced a $7 million Series A round led by Accel, with participation from Shastra VC, Titan Capital, and angel investors including Akshay Kothari (co-founder of Notion).3 The round was accompanied by the public launch of the Simplismart platform, at which point the company reported over 30 enterprise customers and approximately $1 million ARR, with a stated target of $5 million ARR by Q1 2025 (attainment not publicly confirmed). During the same period, AWS published a case study documenting Simplismart's deployment of EC2 Auto Scaling warm pools that reduced GPU scale-up time from 5–6 minutes to under 70 seconds, alongside reported 40% infrastructure cost reductions for customers and 8× growth in GPU-hours deployed over three months.8

    2026: NVIDIA Cloud Partner Integration and Reported Series B

In February 2026, Simplismart announced the launch of an advanced AI inference platform for NVIDIA Cloud Partners (NCPs), enabling cloud providers to white-label Simplismart's managed inference endpoints with integrated NVIDIA Inference Microservices (NIMs) and pre-built workflow templates.4 This announcement represented Simplismart's entry into an infrastructure-distribution role — powering managed inference for other cloud providers rather than serving only direct enterprise customers. In May 2026, reports emerged that NVIDIA was in advanced discussions to lead a further ~$20 million funding round at a prospective post-money valuation of approximately $100 million.5 As of June 2026, this round has not been publicly confirmed as closed.

  What They Offer — Products & Platform

Simplismart's commercial offering is organized into four named product lines and a set of deployment modalities that span serverless to fully air-gapped environments.

SimpliLLM provides serverless and dedicated API endpoints for large language models and vision-language models. The API supports over 150 pre-deployed open-source models importable from 10+ cloud repositories including Hugging Face. Supported quantization formats include FP16, FP8, BF16, AWQ, and GPTQ; KV caching options include Paged, Static, and ShadowKV; and tensor parallelism from TP1 to TP8 is available for large-model serving.

SimpliScribe delivers speech-to-text APIs supporting over 100 languages, backed by Whisper family models. Simplismart claims 8% higher accuracy and 36% lower latency than alternatives, with time-to-first-audio under 250 ms. SimpliSpeak covers text-to-speech with a claimed industry-leading 250 ms time-to-first-audio chunk. SimpliDiffuse offers optimized text-to-image generation APIs for Stable Diffusion and Flux model families.

Across all product lines, Simplismart supports four deployment modalities:

  • Serverless API — pay-as-you-go consumption with no GPU reservation; suitable for variable workloads.
  • Dedicated GPU clusters — cloud-hosted reserved capacity for latency-sensitive or high-volume production deployments.
  • Bring Your Own Cloud (BYOC) — Simplismart's software stack deployed on the customer's own cloud account.
  • On-premises / air-gapped — Kubernetes or Slurm deployments for regulated industries requiring full data sovereignty.

A Terraform-like declarative language standardizes fine-tuning and deployment workflow configuration across all modes. The platform also functions as white-label inference infrastructure for NVIDIA Cloud Partners, offering managed endpoints with integrated NIMs and pre-built templates that NCPs can expose under their own brand.

  Technology & Infrastructure

Simplismart's technical foundation is a proprietary megakernel inference engine that fuses the entire forward pass of an LLM into a single GPU kernel. Conventional inference runtimes execute hundreds of sequential CUDA kernels per forward pass; the transitions between kernels introduce idle time during which GPU compute sits unused. Simplismart's megakernel eliminates these transitions, driving reported GPU bandwidth utilization from under 50% (vLLM baseline) to approximately 78%.2 The company has also written 28 custom CUDA kernels and implemented Flash Attention to further optimize memory bandwidth during attention computation.

Benchmarked on LLaMA-1B decoding throughput, Simplismart reports 2.5× faster performance than vLLM and 1.5× faster than SGLang. Single forward-pass latency for LLaMA-1B reaches under 1 ms on H100 and 0.68 ms on B200.2 For larger batch workloads, the company reports aggregate throughputs of approximately 11,000 tokens/second for Llama 2 7B on A100 and approximately 9,000 tokens/second for Mistral on A100. Individual token generation rates for Llama 3.1 8B reach 440–501 tokens per second in reported benchmarks.

The hardware fleet spans NVIDIA T4, L4, A10G, A100, H100, H200, and B200 GPUs. Cloud infrastructure runs primarily on AWS (EC2 P5 and P4d instances, EKS for orchestration, EC2 Auto Scaling warm pools for sub-70-second scale-up).8 The software stack layers multiple open-source inference runtimes — vLLM, Triton Inference Server, LMDeploy, and TensorRT — beneath the megakernel engine, which operates as a drop-in replacement for the decode phase. Planned roadmap extensions include automated kernel generation during model fine-tuning, which would allow the megakernel to be produced on-demand for newly fine-tuned model variants.

Auto-scaling response time is reported at under 500 ms, and Simplismart claims 99.99% uptime across its managed endpoints. Enterprise security certifications include ISO 27001 v2022, SOC 2 Type II, GDPR, and AICPA compliance.

  Model Catalog & Performance

Simplismart's catalog spans large language models, vision-language models, image-generation models, and speech models. The platform carries 150+ pre-deployed open-source models and supports import from Hugging Face and 10+ other cloud repositories. Key catalog entries and their reported serving performance are listed below.

Large Language Models

ModelNotes
Llama 3.1 8B440–501 tokens/second reported; primary megakernel benchmark model
Llama 3.1 70BServed on A100/H100 with tensor parallelism
Llama 3.1 405BMulti-GPU TP8 serving
DeepSeek-R1Available via serverless API
DeepSeek-V3Available via serverless API
Gemma 3 4BLightweight serving tier
Qwen2.5 72BAvailable; megakernel support on roadmap
Qwen2.5 7B InstructAvailable via serverless API
Phi-3 128KLong-context serving
MistralAvailable; megakernel support on roadmap

Image Generation

ModelNotes
Flux 1.1 ProVia SimpliDiffuse
Flux DevVia SimpliDiffuse
Flux.1 KontextVia SimpliDiffuse
SDXLVia SimpliDiffuse

Speech

ModelNotes
Whisper Large v2Via SimpliScribe; 100+ languages
Whisper Large v3Via SimpliScribe
Whisper v3 TurboVia SimpliScribe; lowest latency tier

  Pricing & Performance Position

Simplismart competes primarily on price-per-token and latency, targeting workloads where Together AI and Fireworks AI are the incumbent alternatives. Pricing below reflects the published serverless API rates as of June 2026; dedicated GPU and enterprise on-prem pricing is quoted separately.9

LLM Serverless API (per 1M tokens, input and output)

ModelPrice
Llama 3.1 8B$0.13
Llama 3.1 70B$0.74
Llama 3.1 405B$3.00
DeepSeek-R1$3.90
DeepSeek-V3$0.90
Gemma 3 4B$0.10
Qwen2.5 72B$1.08
Phi-3 128K$0.08

Image Generation (per image, 1024×1024)

ModelPrice
Flux 1.1 Pro$0.05
Flux Dev$0.03
Flux.1 Kontext$0.04
SDXL$0.028 (note: $0.28 listed on pricing page — may reflect per-batch or a listing error)9

Speech-to-Text (per audio minute)

ModelPrice
Whisper Large v2$0.0028
Whisper Large v3$0.0030
Whisper v3 Turbo$0.0018

Dedicated GPU Hourly Rates

GPUHourly
NVIDIA T4$1.20
NVIDIA L4$1.50
NVIDIA A10G$2.00
NVIDIA A100$3.00
NVIDIA H100$4.00
NVIDIA H200$5.20
NVIDIA B200By quote

Customer-reported outcomes illustrate the pricing thesis in practice. Mindtickle (sales readiness SaaS) reported a 97% reduction in image-generation infrastructure spend, from approximately $30,000 to under $1,000 per month, alongside a 50% reduction in inference latency after moving image generation workloads to Simplismart.7 Dashtoon (AI comics platform) reported cutting compute costs by 60%. The AWS case study documented customer infrastructure cost reductions of approximately 40% and 8× growth in GPU-hours deployed over three months.8

  People & Leadership

Simplismart is led by its two co-founders, who retain operational control across product, engineering, and commercial functions.

Amritanshu Jain is co-founder and CEO. Before Simplismart he worked as an ML engineer and cloud infrastructure specialist at Oracle Cloud and prior to that at Capillary Technologies, a retail technology company. He has publicly articulated the thesis that open-source inference at the production layer is the critical battleground for the next phase of enterprise AI adoption, and that no single inference configuration fits all workloads — a position that underpins Simplismart's emphasis on configurable latency-versus-cost tradeoffs.

Devansh Ghatak is co-founder and CTO. He spent time at Google Search working on search algorithms before co-founding Simplismart, bringing a systems and optimization orientation that shaped the megakernel engineering program and the company's broader focus on GPU bandwidth utilization as the primary efficiency lever.7

The company's headcount is not publicly disclosed, but its engineering team is based primarily in Bengaluru.

  Funding, Ownership & Business

Simplismart is an independent company operating as Verute Technologies Private Limited under Indian corporate law. Its capital history comprises two publicly acknowledged rounds and a third that had been widely reported but not confirmed closed as of the date of this dossier.

The seed round consisted of less than $1 million raised from undisclosed investors and angels, used to fund the pivot from AutoML to inference infrastructure and the early product launches.1 The Series A of $7 million, announced in October 2024, was led by Accel — one of India's most active venture franchises — with co-investment from Shastra VC, Titan Capital, and individual angels including Akshay Kothari (co-founder of Notion).3 The post-Series A valuation was reported at approximately $25 million. Total capital raised through the Series A is estimated at approximately $8–8.26 million across seed and Series A tranches, with variations in cited figures likely reflecting how earlier-stage investments are counted.6

In May 2026, multiple outlets including The Tech Portal and NewsBytesApp reported that NVIDIA was in advanced discussions to lead a further ~$20 million round at a prospective post-money valuation of approximately $100 million, with Accel expected to re-participate.5 Additional earlier-stage institutional backers cited in secondary sources include Google for Startups Accelerator, Dallas Venture Capital, ML Elevate, and Anicut Capital; these are not confirmed in primary press releases.

Simplismart's revenue model is multi-modal: pay-as-you-go consumption billing (tokens, images, audio minutes), hourly dedicated GPU cluster rental, enterprise on-premises software licensing, and white-label managed inference fees from NVIDIA Cloud Partners. The NCP white-label channel in particular represents a distribution wedge distinct from direct API competition — positioning Simplismart as infrastructure middleware for other cloud providers rather than solely a direct enterprise vendor.

  Customers & Partnerships

By the time of its October 2024 Series A, Simplismart reported serving over 30 enterprise customers, with a customer base spanning healthcare, media, sales technology, and e-commerce.3 Notable production deployments include:

  • Tata 1mg — healthcare and pharmacy platform (India); early design partner.
  • Mindtickle — sales readiness SaaS; reported 97% image-generation cost reduction and 50% latency improvement.7
  • InVideo — AI-powered video creation platform; leverages SimpliDiffuse and SimpliLLM.
  • Dashtoon — AI comics and manga platform; reported 60% compute cost reduction.7
  • Dubverse — AI dubbing and video localization.
  • Vodex — AI voice automation.

On the partnership side, Simplismart holds NVIDIA Inception Program membership and, from February 2026, operates as a white-label inference infrastructure provider for select NVIDIA Cloud Partners (NCPs) globally — enabling NCPs to expose Simplismart-powered endpoints under their own brand with integrated NIM catalogs and workflow templates.4 The AWS case study is the most detailed publicly available account of Simplismart's infrastructure optimization work, documenting sub-70-second auto-scaling via EC2 warm pools and customer cost outcomes.8

  Competitive Position

Simplismart frames its primary competition as Together AI and Fireworks AI, two US-based inference API providers. Its differentiation rests on four claims: higher raw throughput via the megakernel engine (2.5× vLLM, 1.5× SGLang in internal benchmarks); deployment flexibility from serverless to air-gapped on-prem; white-label NIM integration for cloud provider resellers; and a price point enabled in part by an India-based engineering cost structure.10

The broader competitive set includes Replicate, Modal, BentoML, Hyperbolic, and FriendliAI in the inference-API category, along with hyperscaler inference services such as AWS SageMaker, Azure AI Foundry, and Google Vertex AI at the enterprise end. Against hyperscalers, Simplismart differentiates on open-model focus, latency-versus-cost configurability, and data-sovereignty deployment options that managed cloud services cannot easily match. The prospective NVIDIA investment, if completed, would introduce a new dimension: deep integration with the dominant GPU ecosystem and the distribution reach of NVIDIA's cloud partner network.

  Outlook & Roadmap

Simplismart's near-term roadmap centers on extending the megakernel engine to additional model families — Mistral and Qwen are the next planned targets — and building automated kernel generation that produces megakernels during the fine-tuning process, making the optimization applicable to custom model variants without manual kernel engineering.2 The NVIDIA Cloud Partner channel is expected to expand through a growing NIM catalog and new workflow template libraries, positioning Simplismart increasingly as infrastructure middleware rather than solely a direct inference API vendor.

The prospective ~$20 million round reportedly led by NVIDIA would, if closed, fund R&D growth, commercial sales expansion, and deeper infrastructure partnerships.5 NVIDIA's strategic participation would signal potential integration of Simplismart's megakernel technology with the broader NIM ecosystem, though the terms of any such technical collaboration have not been disclosed. CEO Jain has described the long-term thesis as winning the production inference layer for open-source models globally — a market where, in his view, no single inference profile satisfies all enterprise workloads, making configurability and continuous optimization a durable differentiator.


  References

  1. GlobeNewswire — Simplismart $7M Series A announcement (Oct 2024)
  2. SiliconAngle — Simplismart raises $7M (Oct 2024)
  3. Simplismart blog — Advanced AI inference platform for cloud providers on NVIDIA infrastructure (Feb 2026)
  4. Simplismart — Pricing page
  5. Simplismart blog — Megakernel inference: unlocking blazing-fast responses
  6. AWS — Simplismart case study
  7. The Tech Portal — NVIDIA reportedly to lead $20M Simplismart round (May 2026)
  8. NewsBytesApp — NVIDIA to lead $20M round at ~$100M valuation
  9. Seed to Scale podcast — Building AI infrastructure from India for the world (Simplismart)
  10. Analytics India Magazine — Bengaluru startup makes fastest inference engine beating Together AI and Fireworks AI

  References

  1. Founding story, pivot, and seed conditions — GlobeNewswire Series A announcement. 2 3

  2. Megakernel engine technical claims — Simplismart blog; Analytics India Magazine. Internal benchmarks; independent verification pending. 2 3 4

  3. Series A details, investors, customer count, and ARR at launch — GlobeNewswire; SiliconAngle. 2 3 4 5

  4. NCP white-label platform launch — Simplismart blog, Feb 2026. 2 3

  5. NVIDIA Series B reports — The Tech Portal; NewsBytesApp. Round not confirmed closed as of June 2026. 2 3 4 5

  6. Total funding estimate from PitchBook/Crunchbase via secondary sources; exact seed tranches vary by source. 2

  7. Founder backgrounds, customer case studies, and early-stage narrative — Seed to Scale podcast. 2 3 4 5

  8. AWS case study — aws.amazon.com. Figures reflect reported customer outcomes. 2 3 4

  9. Pricing figures from Simplismart pricing page — simplismart.ai/pricing. SDXL price discrepancy noted; verify against current page. 2

  10. Competitive positioning claims — Analytics India Magazine.