AI inference providers explained: the top platforms in 2026 and how they compare
“Inference provider” is one word for five businesses. Which key to hold depends on what you generate. Prices from twelve vendor pages.
Pick a model and the next question arrives before the first one is settled: where does it run? The same open-weight text model is sold by at least three of the platforms below, at input prices that differ by nearly four to one. The picture model you just chose sits on three catalogs with three different price units. The word for all of them is “inference provider,” and it is one word for five different businesses.
This guide maps the five shapes, then walks the top platform in each, using numbers read from the vendors’ own pricing pages on September 10, 2026. By the end you should know which key, or keys, to hold. The short answer for most people who generate text, images, audio and video: one media host, one router, and a local runtime for everything that can stay on the machine.
A provider is where the model runs, and that is a separate choice
A model vendor trains the weights: OpenAI, Anthropic, Google, Meta, DeepSeek, ByteDance. An inference provider owns the GPUs those weights run on and sells the output, by the token, by the image or by the second. Sometimes vendor and provider are the same company. OpenAI hosts GPT-5.6 itself. Often they are not. Meta trains Llama and hosts none of it for you, while Together, Fireworks, Groq, DeepInfra and Amazon all do. ByteDance’s Seedream image model is on Runware, fal and Replicate, priced three ways.
A gateway, or router, is a third thing: a switchboard that holds keys to many providers and resells them behind one API. We covered how gateways work, what they quietly cost and who now owns them in the gateways and routers post. Here OpenRouter is treated as one provider shape among five, and that story is not re-argued.
The shapes matter because they price differently, host different things and fail differently. A media host bills per output and does not charge for a failed generation. A GPU host bills per token, with a cheaper rate for input it has seen before. A hyperscaler bills through a cloud account your finance team already pays. The hero above lays out the five. The rest of this post walks them in the order most readers meet them.
Media hosts bill by the picture or the second
If you generate images, video or audio, read the price unit before the price. A per-image price tells you what a picture costs before you make it. A per-second GPU price tells you only after you know how long the model ran, and for a diffusion model that varies with resolution, step count and whether the model was warm. Both units exist on these three platforms, sometimes on the same page.
- Hosts
- Image, video, audio, 3D, editing and upscaling, plus a side catalog of text models. The homepage counts 198 image model families, 114 video and 29 audio, and “400K+ models” once community checkpoints are included (runware.ai).
- Pricing
- Per output. The pricing page puts image generation at $0.0006 to $0.24 per image depending on model, resolution and quality. Serverless compute is also sold by the second: an H100 at $0.000767 per second, a B200 at $0.001386.
- Billing
- Pay as you go, no prepaid minimum. $2 in credits with a business email and no card. Only successful requests are charged.
- API
- REST for stateless calls, WebSockets for persistent low-latency sessions, and an OpenAI-compatible layer for tools that expect one.
- Strong where
- Breadth of media models behind one fixed price per output.
- Real limits
- Text models are a side catalog, not the main event. Nothing free beyond the $2.
The four illustrations in this post came from Runware’s API while it was being written: Seedream 5 Pro at 2560 by 1440, first result each, $0.39 for all four, 46 to 59 seconds per image. That is the whole point of per-output pricing. The number was known before the request went out.
- Hosts
- “Thousands of models contributed by our community,” alongside official models from OpenAI, Google, Anthropic and Black Forest Labs, across images, speech, music, video and language models (replicate.com).
- Pricing
- Two units on one pricing page. Community models bill by hardware time: a T4 at $0.000225 per second, an A100 80 GB at $0.0014, an H100 at $0.001525 ($5.49 per hour). Official models bill per output: FLUX Dev $0.025 per image, FLUX Pro $0.04, Ideogram v3 $0.09, Wan video $0.09 to $0.25 per second of output.
- Billing
- Prepaid credit or monthly postpaid. For public models “setup and idle time for the model is free,” and failed runs are not charged (billing docs).
- API
- Its own prediction API and SDKs. Deployments give you a private endpoint with always-on instances, and in exchange you pay for setup and idle time too.
- Strong where
- The long tail. If a researcher published it last week, it is probably here.
- Real limits
- Shared hardware on the long tail means cold boots. The fix, a deployment, changes the billing unit from per output to per hour.
- Hosts
- “1,000+ production ready image, video, audio and 3D models” (fal.ai).
- Pricing
- Per output, on the pricing page: Seedream V4 $0.03 per image, FLUX Kontext Pro $0.04, Qwen image $0.02 per megapixel. Video by the second: Wan 2.5 $0.05, Kling 2.5 Turbo Pro $0.07, Veo 3 $0.40. Raw GPUs from $1.89 per hour for an H100.
- Billing
- Pay only for what you use; no free-credit amount is stated on the pages we read.
- API
- Its own queue-based API and client libraries.
- Strong where
- Diffusion speed. fal says its inference engine is “up to 10x faster,” a vendor claim we did not measure.
- Real limits
- Media only. There is no text catalog worth holding a key for.
Open-weight GPU hosts turn text into a commodity
Open-weight text is the most competitive market in AI. The same DeepSeek V4 Flash checkpoint is on all four of these platforms, and its input price runs from $0.06 to $0.22 per million tokens depending on which one you ask. Every one of them speaks the OpenAI request shape, so switching is a base URL and a model name. That is why prices converge, and why the differences that remain are speed tiers, cache discounts and billing terms.
- Hosts
- Open text, vision, image, audio, video and embedding models, plus dedicated endpoints, GPU clusters and fine-tuning.
- Pricing
- Per million tokens on the pricing page: DeepSeek V4 Flash $0.14 in and $0.28 out, Llama 3.3 70B $1.04 each way, GLM-5.3 $1.40 and $4.40, Kimi K3 $3 and $15. Images $0.0006 to $0.134 each. An on-demand H100 is $3.99 per hour on promotion ($5.49 standard).
- Billing
- Pay as you go; the pricing page says “start for free” without stating a credit amount.
- API
- OpenAI-compatible at api.together.ai/v1 “across chat, completions, vision, image generation, text-to-speech, and embeddings” (docs).
- Strong where
- The full stack in one account, from serverless to reserved clusters to fine-tuning.
- Real limits
- Price. Its Llama 3.3 70B list price is ten times DeepInfra’s for the same weights.
- Hosts
- Open text and vision models, Whisper for audio, FLUX Kontext for images, embeddings, fine-tuning.
- Pricing
- Three serverless tiers per model on the pricing docs. Standard: DeepSeek V4 Flash $0.22 in and $0.66 out, DeepSeek V4 Pro $1.32 and $3.96 with cached input at $0.044, gpt-oss-120B $0.15 and $0.60, Kimi K3 $3 and $15. Priority costs about 25% more; the Fast tier for Kimi K3 is $4.50 and $22.50.
- Billing
- Postpaid per token, with $1 in free credits to start (pricing page).
- API
- “OpenAI and Anthropic compatible” (fireworks.ai).
- Strong where
- Choosing latency per request. The same model has a cheap lane and a fast lane.
- Real limits
- The fast lane costs 50% more, and on-demand GPUs rose to $8 per H100-hour on September 1.
- Hosts
- A short list of open text models and speech models on its own LPU hardware.
- Pricing
- The model table lists gpt-oss-120B at $0.15 in and $0.60 out at about 500 tokens per second, and gpt-oss-20B at $0.075 and $0.30 at about 1,000. Llama 3.1 8B is shown at 560 tokens per second with price marked “contact sales.”
- Billing
- A rate-limited free plan and a paid Developer plan with batch and flex processing (rate limits).
- API
- “Mostly compatible” with OpenAI’s client libraries at api.groq.com/openai/v1 (docs).
- Strong where
- Tokens per second. Nothing else on this list publishes speeds in its price table.
- Real limits
- Text and speech only, and a small catalog. On December 24, 2025 NVIDIA licensed Groq’s inference technology and the founder moved across; Groq says GroqCloud “will continue to operate without interruption” as an independent company (newsroom).
- Hosts
- Open text, embeddings, rerankers, image, speech, music and video models, and it now resells closed models from Anthropic and Google at list price.
- Pricing
- The lowest list prices in this group on the pricing page: Llama 3.1 8B $0.02 in and $0.04 out, Llama 3.3 70B $0.10 and $0.32, DeepSeek V4 Flash $0.06 and $0.18, DeepSeek V4 Pro $1.30 and $2.60. FLUX-2 Pro images at $0.015. A Flex lane at 0.8 times base price and a Priority lane at 1.5 times.
- Billing
- A card on file or a prepaid balance is required before any request runs. No free credits are stated.
- API
- OpenAI-compatible: “just change the base URL and model name” to api.deepinfra.com/v1/openai (docs).
- Strong where
- Price. If the model is open and you can tolerate the Flex lane, this is usually the floor.
- Real limits
- No free tier to try it on, and the closed models it resells cost the same as at source.
A router sells breadth for a 5.5% fee
OpenRouter owns no GPUs. It holds keys to the hosts above and to the first-party APIs, and resells all of them behind one OpenAI-shaped key. Stripe’s August 19, 2026 announcement of its acquisition put the count at “400+ models from more than 80 providers.”
- Hosts
- Text and vision through chat completions, plus dedicated endpoints for image generation, video generation, text-to-speech and transcription (multimodal docs).
- Pricing
- The provider’s own price with no inference markup. The fee is 5.5% (minimum $0.80) when you buy credits by card, 5% by crypto (FAQ). Bring your own provider key and the fee is 5% of what the same call would cost on OpenRouter, above a monthly allowance.
- Billing
- Prepaid credits. Free-tier models allow 50 requests per day, or 1,000 per day once you have bought $10 of credits.
- API
- Implements the OpenAI specification as a drop-in replacement. If a provider fails it falls back to the next; a :nitro suffix asks for the fastest throughput, :floor for the lowest price.
- Strong where
- One key, one bill, one request shape, and a provider list that changes weekly without touching your code.
- Real limits
- It can be no cheaper than the cheapest provider it holds plus the fee, it cannot fix a provider’s cold start, and it now belongs to a payments company.
First-party APIs sell exclusivity, hyperscalers sell procurement
Some models you cannot get anywhere else on launch day, and the first weeks of a frontier model are often the only weeks that matter for a product decision. That is the first-party business. The prices are public and the discounts are structural: cached input and batch processing, not negotiation.
- OpenAI. GPT-5.6 Terra at $2 in and $12 out per million tokens, GPT-5.6 Sol at $4 and $20, GPT-6 Astra at $10 and $50, with cached input at a tenth of the input price. Image output for gpt-image-2 is $30 per million output tokens. Batch runs at half price (pricing).
- Anthropic. Haiku 4.5 at $1 and $5, Sonnet 5 at $2 and $10, Opus 5 at $5 and $25, Fable 5.1 at $10 and $50. Cache reads cut input to $0.10 to $0.50 per million; batch is 50% off (pricing).
- Google. Gemini 3.8 Flash at $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.50. Gemini 2.5 Flash-Lite at $0.10 and $0.40. A free tier exists, with the note that content “may be used to improve our products.” Veo 3.1 video runs $0.05 to $0.60 per second, and Gemini 2.5 Flash Image is $0.039 per picture (pricing).
Hyperscaler gateways sell the same models to a different buyer. Amazon Bedrock’s model catalog lists eighteen providers, from Anthropic and OpenAI to DeepSeek, Qwen, Mistral, xAI and Z.AI, billed on demand per token, at half price in batch, or as provisioned throughput on one-month and six-month terms (pricing). Google’s Vertex AI, which the docs now file under its Gemini Enterprise Agent Platform, hosts Claude Fable 5.1, Opus 5 and Sonnet 5 beside Gemini with Standard, Priority and Flex pay-as-you-go lanes and provisioned throughput (docs). Microsoft Foundry claims “over 11,000” models including OpenAI, Anthropic, Meta, Google and xAI, billed per service (product page).
What you buy from a hyperscaler is not a model. It is one invoice, one identity system, a data-residency choice and compliance paperwork someone already signed. What you give up is the router’s breadth and the GPU hosts’ bargains. The price you see is the list price, and the model you want sometimes arrives there after it arrives at source.
The whole map on one table
Read the table by the column that matters to you. If you generate pictures, the “bills by” column decides. If you run one open text model at volume, “weights” and “free” narrow it to four rows and the price cards above settle it. If you need a signature from procurement, only the last row applies.
| Platform | Shape | Modalities | Weights | Bills by | OpenAI API | Free | Strong where |
|---|---|---|---|---|---|---|---|
| Runware | Media host | Image, video, audio, 3D, text | Open + closed | Per output; GPU per second | Compatible layer | $2, no card | Widest media catalog at fixed per-output prices |
| Replicate | Media host | Image, video, audio, text | Open + closed | Per second or per output | Own API | Select models free | The long tail; anyone can publish a model |
| fal | Media host | Image, video, audio, 3D | Open + closed | Per output; GPU per hour | Own API | Not stated | Diffusion speed, per-output video prices |
| Together | GPU host | Text, image, audio, video | Open | Per million tokens | Yes | Not stated | Serverless to clusters to fine-tuning |
| Fireworks | GPU host | Text, vision, image, audio | Open | Per million tokens, tiered | Yes, plus Anthropic | $1 | Standard / Priority / Fast per request |
| Groq | GPU host | Text, speech | Open | Per million tokens | Mostly | Free plan | Tokens per second |
| DeepInfra | GPU host | Text, image, audio, video | Open + resold closed | Per million tokens, ×0.8 to ×1.5 | Yes | None stated | Lowest list prices in the group |
| OpenRouter | Router | Text, image, video, speech | Open + closed | Provider price + 5.5% on credits | Yes | Free models, 50/day | 400+ models from 80+ providers, one key |
| OpenAI | First-party | Text, image, audio, video | Closed (plus gpt-oss) | Per million tokens | Is the standard | None stated | GPT-5.6 and GPT-6 at source |
| Anthropic | First-party | Text, vision | Closed | Per million tokens | Own SDK | None stated | Claude at source, deepest cache discount |
| Google Gemini API | First-party | Text, image, video, audio | Closed (plus Gemma) | Per million tokens | Compatible endpoint (beta) | Free tier | Cheapest frontier tier, free tier |
| Bedrock · Vertex · Foundry | Hyperscaler | All, by model | Open + closed | Per token; provisioned | Varies | Cloud credits | One invoice, one identity system |
The scenario list below is the same table read backwards, from the need to the key. It assumes you are a builder or a power user paying your own bill; a company with a cloud contract should start from the last line.
Two things the table cannot show. First, per-output prices and per-token prices are not comparable until you know your task, which is the argument of our post on what one AI image costs. Second, every number here moves. Google’s Gemini 3.8 Flash doubles on January 1, 2027; Fireworks raised on-demand GPU rates on September 1; Together’s H100 price is a promotion. Treat the cards as a snapshot with links, not a contract.
Three keys and a local runtime cover nearly everything
Here is how the map turns into a decision. Hold one media host key, because per-output pricing is the honest unit for pictures, video and sound. Hold one router key, because open text is a commodity and the router lets you follow the price without rewriting anything. Run a local runtime for everything small enough to stay on the machine, which is more than most people expect; the cloud API versus local runtime post draws that line. Add a first-party key only when a model is exclusive, and a hyperscaler only when procurement demands it.
That is the shape CSuite is built around. The desktop app ships Runware, Replicate and OpenRouter as cloud platforms next to local Ollama, Hugging Face and GGML runtimes. A model appears once in the catalog with the platforms able to run it listed beside it, so the same Seedream or DeepSeek entry runs under whichever key you hold. The keys never leave the machine: they sit in the app’s local database, encrypted through the operating system’s keyring, the Keychain on macOS, DPAPI on Windows and Secret Service on Linux.
The provider question does not go away after you answer it. Prices move, catalogs change, and a company that owned no GPUs last year is owned by a payments company this year. But it gets easier once you stop treating “provider” as one word. There are five businesses in it, and you only need to be a customer of two.


