Every AI product eventually outgrows calling someone else's API and needs to run a model of its own: a fine-tune, an open-weights LLM, a custom vision pipeline. That puts you in the market for an inference platform, where the prices are strange.
As of August 22, 2026, the same NVIDIA H100 costs $3.95/hr on Modal (billed per second), $5.49/hr on Replicate (per second), $5.49/hr on a Together AI dedicated endpoint (per hour), and $6.50/hr on Baseten (per minute), all read from the vendors' own pricing pages the day this was written. That 65% spread on identical silicon is four different bets about what you are really buying: raw seconds, managed replicas, a model catalog, or tokens. Meanwhile the money says this category is the center of gravity: Baseten announced a $1.5 billion Series F at a reported $13 billion valuation in June 2026, and Cloudflare bought Replicate outright.
Quick Comparison
| Platform | Billing unit | H100 price | Free tier | Best for |
|---|---|---|---|---|
| Modal | Per second | $0.001097/sec ($3.95/hr) | $30/mo credits | Custom Python workloads, scale-to-zero |
| Baseten | Per minute | $0.10833/min ($6.50/hr) | Credits on request | Production replicas, dedicated deployments |
| Replicate | Per second, or per token | $0.001525/sec ($5.49/hr) | Pay as you go | Running catalog models with one API call |
| Together AI | Per 1M tokens, or per GPU-hour | $5.49/hr dedicated, $3.99/hr clusters | Pay as you go | Open-weights LLMs behind an OpenAI-style API |
| vLLM | Your GPU bill | Apache 2.0, 89,671 stars | Free | Self-hosting when you already have GPUs |
The billing unit is the decision that compounds. Per-second billing rewards bursty traffic that scales to zero. Per-minute billing assumes replicas that stay warm. Per-token billing means you never think about GPUs at all, and someone else pockets the efficiency gains.
Modal

Modal sells compute the way a language runtime would: you decorate a Python function, name a GPU, and it runs. Billing is per second with no charge for idle, CPU is $0.0000131 per core-second and memory $0.00000222 per GiB-second, and the GPU list runs from a T4 at $0.000164/sec ($0.59/hr) through the H100 at $3.95/hr to a B300 at $7.10/hr. Starter is $0 with $30/month of free credits; Team is $250/month including $100/month of credits.
The whole deployment, verified against modal 1.5.4:
Pythonimport modal app = modal.App("llm-inference") image = modal.Image.debian_slim().pip_install("vllm") @app.function(image=image, gpu="H100", timeout=600) def generate(prompt: str) -> str: # load once per container, serve many requests ... # modal deploy app.py -> live endpoint, scales to zero between calls
What it does well. The developer experience is the best here if your model is custom Python rather than a catalog entry. Scale-to-zero is real, cold starts are engineered around aggressively, and per-second billing means a job that runs 40 seconds costs 40 seconds. It is also the cheapest H100 of the four.
What to watch. You are writing infrastructure as code, not clicking deploy on a model page. And Modal is a general compute platform (we used its sandboxes for the code interpreter guide), so inference-specific niceties like a model registry or canary rollouts are yours to build.
Baseten

Baseten is the production-replica specialist, and the market has noticed: a $300 million raise at a $5 billion valuation in January 2026 (per Bloomberg), then a $1.5 billion Series F announced June 2026 at a reported $13 billion, on reported annualized revenue around $600 million.
Pricing is per minute of deployment uptime: the H100 is $0.10833/min ($6.50/hr), an A100 $4.00/hr, a B200 $9.98/hr, and an H100 MIG slice $3.75/hr, with the same rates for training and volume discounts advertised. Models deploy through Truss, its open-source packaging format, with autoscaling, canary deploys, and a Model APIs tier where hosted open-weights models (currently the Kimi K3 family and DeepSeek V4) bill per million tokens instead.
What it does well. The operational layer between "model file" and "SLA" is the product: replica autoscaling, observability, and dedicated deployments that do not share queues. Teams running revenue-critical inference at steady volume are who the $6.50/hr H100 is priced for.
What to watch. That is a 65% premium over Modal for the same chip, per-minute billing quietly rounds up bursty workloads, and the free tier is "contact us for credits" rather than a standing allowance.
Replicate

Replicate is the model catalog: thousands of community and proprietary models behind one API, from FLUX image generation ($0.04 per output image for flux-1.1-pro) to hosted LLMs priced per token. Running one is genuinely one call, verified against replicate 1.0.7:
Pythonimport replicate output = replicate.run( "black-forest-labs/flux-schnell", input={"prompt": "a watercolor map of San Francisco"}, )
The 2026 headline is ownership: Cloudflare announced on November 17, 2025 that it is acquiring Replicate, with the deal expected to close within two months, folding the catalog into Workers AI. If your stack already leans Cloudflare, this is now the default answer; if you were counting on Replicate's independence, it is gone.
Billing splits in two. Public models bill per second of runtime (or per token for language models), so you pay only while your request processes. Private models run on dedicated hardware, and the pricing page is unusually honest about what that means: you pay for setup time, idle time, and active time. The H100 is $0.001525/sec ($5.49/hr), an A100 $5.04/hr, an L40S $3.51/hr, and multi-GPU configs stack linearly. Fast-booting fine-tunes are the exception that skips idle billing.
What to watch. That idle-time clause is the bill surprise: a private model with trickle traffic pays for every quiet hour. Cold starts on public models also vary with model popularity, since hot models stay warm on shared hardware.
Together AI

Together AI starts from the other end: tokens, not GPUs. Its serverless tier prices open-weights models per million tokens with an OpenAI-compatible API, and the current table reads like the frontier of open models: DeepSeek V4 Flash at $0.14 input / $0.28 output, DeepSeek V4 Pro at $1.32 / $3.96, Kimi K3 at $3.00 / $15.00, with cached-input discounts (Kimi K3 cached input drops to $0.30) and a batch-price toggle. The company has reportedly passed $1 billion in annualized revenue, with a $7.5 billion valuation reported after its $305 million Series B.
When serverless stops fitting, the same platform sells the ladder down: dedicated endpoints at $5.49 per H100-hour ($8.99 for B200), provisioned throughput units for reserved token capacity, and raw GPU clusters at $3.99 per H100-hour on demand, dropping to $3.19 on a 91-to-180-day reservation. The verified client shape (together 2.31.0):
Pythonfrom together import Together client = Together() # TOGETHER_API_KEY in env resp = client.chat.completions.create( model="deepseek-ai/DeepSeek-V4", messages=[{"role": "user", "content": "Summarize this diff"}], stream=False, )
What to watch. Serverless per-token pricing is the fastest start and the least control: model deprecations, shared-queue latency, and per-token margin are all Together's calls. The moment you need a custom model rather than a catalog one, you are back in per-GPU-hour land with everyone else.
The Self-Hosting Escape Hatch
Every platform above ultimately runs an open-source serving engine you can run yourself. vLLM (Apache 2.0, 89,671 stars, pushed today) is the default: continuous batching and paged attention turn one H100 into a respectable API server. SGLang (32,244 stars, also active today) frequently beats it on structured output and multi-turn workloads. Hugging Face's TGI (10,887 stars) still works but has visibly slowed, with no push since March 2026.
The honest math: self-hosting wins with steady, high utilization and someone to own CUDA drivers and 3 a.m. OOM pages. A $3.19/hr reserved H100 running vLLM at 70% utilization beats every managed per-token price; at 10% it loses to all of them.
Running the Numbers
One model, one always-on H100 replica, 730 hours: Modal $2,884, Together dedicated $4,008, Replicate private $4,008, Baseten $4,745. Give it two hours of real GPU work per day on a scale-to-zero platform instead, and Modal bills roughly $237/month while anything holding a warm replica still bills the always-on price. That is the category in two sentences: the sticker matters less than whether idle time is billed, and scale-to-zero is a 10x lever for bursty traffic.
For token-shaped work the inversion runs the other way: at light volume, Together serverless (or Baseten's Model APIs, or Replicate's per-token models) costs almost nothing and needs no capacity planning, while past a few billion tokens a month dedicated GPUs undercut per-token pricing.
How to Pick
Custom Python model, bursty or unpredictable traffic: Modal, for per-second billing, scale-to-zero, and the cheapest H100 here. Steady production traffic with an SLA and a team that wants deploys, canaries, and autoscaling handled: Baseten, priced like the premium option because it is one. You mostly want to call existing models (image, video, speech, open LLMs) without owning infrastructure: Replicate, with the Cloudflare acquisition as a plus if you are already there and a governance question if you are not. Open-weights LLMs behind an OpenAI-style API, from prototype tokens to reserved clusters: Together AI. GPUs you already own or reserved capacity plus engineering time: vLLM or SGLang.
Conclusion
Price your traffic shape, not the H100 sticker. The $3.95-to-$6.50 spread is real money, but the spread between billed-idle and scale-to-zero is bigger: it is why the same workload can cost $237 or $4,745 a month on platforms that look interchangeable in a feature table.
And the category is consolidating upward: Cloudflare bought Replicate, Baseten tripled to a reported $13 billion in five months, Together crossed a reported billion in annualized revenue. Every vendor sells a cheap on-ramp and an expensive destination and assumes you will migrate from one to the other. Re-run the math on your own bill before they turn out to be right.
Related DevToolLab Tools
- Latency Percentile Calculator - turn a list of inference response times into the p50/p95/p99 you actually compare platforms on.
- Uptime SLA Calculator - convert a platform's promised nines into allowed downtime per month before you sign.
- Download Time Calculator - estimate how long pulling 140 GB of model weights takes at your bandwidth, which is most of what a cold start is.
- JSON to Python Converter - turn a sample inference API response into typed Python for your client code.
Related Guides
- Best LLM Gateways and API Routers - the routing layer that sits in front of whatever platform you pick here
- Best AI Fine-Tuning Platforms - training the custom model these platforms then serve
- Local LLM VRAM Requirements - sizing the GPU before you rent or buy it
- Build an AI Code Interpreter with E2B and Modal - Modal's other identity as a sandbox platform, hands on
