Three GPU clouds, three completely different bets on what AI infrastructure actually is.
Runpod ($240M ARR, $1B valuation) is the GPU marketplace for teams who want price and breadth across the full AI lifecycle. Modal ($300M ARR, $4.65B valuation) is a Python-native serverless runtime increasingly betting its future on agentic AI sandboxes. Cerebrium (YC W22, ~$8.5M raised) is the compliance-first niche player built specifically for real-time voice and video AI. They share a pricing model but serve three different buyers.
Key takeaways
Winners by dimension
Side-by-side
| Runpod | Modal Labs | Cerebrium | |
|---|---|---|---|
| Free tier | No free tier; no minimum spend required | $30/month free compute credits on Starter plan | Hobby plan: free platform access, pay compute at per-second rates |
| Starting paid price | Pay-as-you-go from $0.27/hr (RTX A5000); no plan fee | $250/mo Team plan + compute; Starter is free | $100/mo Standard plan + compute at per-second rates |
| H100 effective hourly rate | $2.89/hr (PCIe, listed July 2026) | $3.95/hr (listed July 2026) | $0.000944/sec = ~$3.40/hr (listed July 2026) |
| Cold start performance | Sub-200ms for 48% of serverless requests (FlashBoot); 4.2s at P99 | Sub-second for most workloads; GPU snapshotting improves by 10x vs baseline | 2-4s typical; GPU memory snapshotting reduces by 71% |
| Pricing model | Per-second billing; no idle cost; no egress fees | Per-second (CPU cycle + GPU second); optional $250/mo plan fee | Per-second (GPU + CPU vCPU + memory GB); $100/mo platform fee on Standard |
| GPU catalog size | 30+ SKUs (RTX A5000 to B300/H200/B200); 31 global regions | ~10 GPU options (T4 to B300); multi-cloud (AWS, GCP, Oracle) | 12+ GPU types (T4 to B200, AMD MI300X, TPU v5e, AWS Trainium) |
| Compliance certifications | SOC 2 Type II, HIPAA, GDPR (all achieved by Feb 2026) | SOC 2 Type II; HIPAA on Enterprise plan only | SOC 2 Type II, ISO 27001, HIPAA, GDPR (all active as of July 2026) |
| Uptime SLA | 99.99% (Secure Cloud); no SLA on Community Cloud | Reportedly no published uptime SLA | Reportedly 99.999% with multi-region failover |
| Training / multi-node clusters | Up to 64 GPUs on-demand; 10,000+ on Reserved; InfiniBand/RoCE v2 networking | Up to 128 B200s; 3,200 Gbps InfiniBand for multi-node RL/training | Serverless inference focus; fine-tuning supported but not primary use case |
| Developer experience / SDK | Python, JS, Go SDKs; Docker-based; CLI + REST API; ~1-2 hr first deploy | Python decorator API; no Docker needed; 5-10 min first deploy | Bring-your-own Dockerfile; no code rewrites; CLI; minimal SDK lock-in |
| Open roles (growth signal) | 22 open roles (11 engineering, 6 sales, 2 product, 2 marketing) | Reportedly 27 open roles (heavy in ML engineering, enterprise sales, GTM) | Reportedly 2 open roles (1 engineering, 1 GTM/sales) |
| Agentic AI / Sandbox capability | Serverless endpoints support agentic workflows; no dedicated sandbox primitive | 1B+ Sandboxes launched; drives >1/3 of $300M ARR; sub-second scheduling | Not a focus; inference and voice AI are the primary use cases |
Who should pick whom
What we found
Who They're Actually Selling To
Runpod sells to the broadest audience: individual developers who want cheap GPUs, ML teams who need the full training-to-inference pipeline, and increasingly enterprises that need compliance. Modal sells to Python developers who want infrastructure to disappear, and to AI-native companies building agentic systems. Cerebrium sells to a specific niche: teams building real-time voice, video, and digital avatar products who need sub-500ms latency, ISO 27001, and a support team that actually picks up Slack. These aren't three versions of the same product. They're three different answers to what AI infrastructure should feel like.
Pricing: Where the Math Gets Interesting
All three use per-second billing, but the economics diverge fast. Runpod's H100 at $2.89/hr beats Modal's $3.95/hr by 37%, which matters enormously on long training runs. Per DeployBase's March 2026 analysis, RunPod is roughly 25x cheaper than Modal on a 10M-token batch inference job. Modal flips the math for bursty short workloads: its sub-second cold starts mean you're not paying 60-90 seconds of H100 time to wake up a container. Cerebrium's $100/mo Standard plan adds a fixed overhead that stings small teams but unlocks HIPAA and ISO 27001 that neither competitor offers at that price point.
Where the Real Moat Lives
Runpod's moat is its 1 million developers and 120% Net Dollar Retention. That community is a distribution flywheel: developers who start on $0.27/hr RTX A5000s organically expand to H100 clusters. Modal's moat is its custom Rust-based container runtime and GPU snapshotting, which competitors can't easily copy. The Sandboxes product, now driving over a third of $300M ARR (per Benzinga, May 2026), is becoming a platform bet: Modal wants to be the default execution environment for AI agents, not just a GPU rental shop. Cerebrium's moat is its compliance stack combined with white-glove support. A 12-person team can't outbuild Runpod or Modal on features, but it can out-support them on enterprise deals where a dedicated Slack channel and ML engineering services close contracts.
The Reliability Gap Nobody Talks About Enough
Runpod's Community Cloud explicitly carries no uptime warranty in its terms of service. A January 2026 AWS us-east-1 incident disrupted Runpod's control plane through a Vercel dependency, affecting pod provisioning and payment processing (per GMI Cloud analysis, May 2026). Practitioner sentiment on Reddit flags GPU availability shortages as a recurring issue in Community Cloud, though Secure Cloud and Serverless are notably more stable. Modal reportedly publishes no uptime SLA at all. Cerebrium's reported 99.999% SLA with multi-region failover is the standout here, making it the only one of the three that enterprise procurement teams can sign off on without a negotiation.
What the Hiring Signals Reveal
Modal's reportedly 27 open roles skew heavily toward ML engineering research and enterprise sales, which tells you the company is racing to win model-ownership workloads before hyperscalers catch up. Runpod's 22 roles include 6 in sales, which is new for a company that grew entirely on developer word-of-mouth. The sales motion is being built now, not later. Cerebrium's reportedly 2 open roles reflect a company at a decision point: stay lean and profitable serving a niche, or raise more capital and hire aggressively. With only $8.5M raised versus Runpod's $122M and Modal's $466M, Cerebrium is playing a different game entirely.
Sources & references
Every claim in this report was triangulated against 17 third-party sources (analyst reports, developer surveys, news coverage, and pricing pages). Sources are listed below in citation order.
- AI cloud startup Runpod hits $120M in ARR, and it started with a Reddit post(narrative, key_stat, side_by_side)
- RunPod: $100M Series A at $1B, Rejected $500M Buyouts(headline, key_stat, winners, narrative)
- Runpod raises $100M at $1B valuation, rejects $500M buyout offers(key_stat, headline, narrative)
- Modal's Series C: Raising $355M at a $4.65B valuation(key_stat, narrative, winners)
- Modal Labs revenue, valuation & funding | Sacra(narrative, side_by_side, winners)
- Modal Plan Pricing(side_by_side, narrative)
- Pay-Per-Second Pricing for Serverless AI | Cerebrium(side_by_side, narrative)
- Runpod Reliability Deep-Dive Report 2026 | Endplan(narrative, winners, side_by_side)
- 10 Best Modal Alternatives in 2026: Serverless GPU Without the Lock-In | Spheron Blog(side_by_side, narrative, winners)
- 10 Best RunPod Alternatives in 2026 (Compared) | Spheron Blog(side_by_side, narrative)
- NVIDIA H100 Pricing (July 2026): Cheapest Cloud GPU Rates | Thunder Compute(side_by_side, winners, key_stat)
- Modal vs RunPod: Python-First Serverless vs GPU Marketplace | DeployBase(narrative, side_by_side, personas)
- Best RunPod Alternatives for Scalable GPU Inference 2026 | GMI Cloud(narrative, winners)
- Pitch Deck Cerebrium Used to Nab $8.5 Million From Gradient Ventures | Business Insider(narrative, personas, side_by_side)
- Stanford CS336: Language Modeling from Scratch (Spring 2026)(winners, personas, narrative)
- Top Serverless GPU Clouds for 2026: Comparing Runpod, Modal, and More(side_by_side, narrative)
- General Catalyst, Redpoint Fuel Modal's AI Infrastructure Push With $355M Raise | Benzinga(key_stat, narrative, winners)
Want to share this?
Three GPU clouds. Three completely different businesses. Runpod turned down $500M to stay independent. Here's what actually separates them:
Want this kind of report on your competitors?
ClientCues runs deep AI scans + side-by-side comparisons every week.
Try ClientCues Free