Skip to content
CSuite
ListicleAI ModelsEconomicsSep 10, 202611 min read

AI inference providers explained: the top platforms in 2026 and how they compare

“Inference provider” is one word for five businesses. Which key to hold depends on what you generate. Prices from twelve vendor pages.

Where a model runs when you do not run it · September 2026
“Inference provider” is one word for five businesses. Pick by what you generate.
Media hosts
Image, video, audio range
Per image · per second
Open-weight GPU hosts
Price and speed on open text
Per million tokens
Router
Breadth behind one key
Provider price + 5.5%
First-party APIs
Frontier models, exclusive
Per million tokens
Hyperscaler gateways
Procurement and compliance
Cloud invoice, list price
Every price in this post was read from the vendor’s own pricing or docs page on September 10, 2026, and is linked where it appears. Platforms are compared on published facts only; nothing here ranks them by opinion.

Pick a model and the next question arrives before the first one is settled: where does it run? The same open-weight text model is sold by at least three of the platforms below, at input prices that differ by nearly four to one. The picture model you just chose sits on three catalogs with three different price units. The word for all of them is “inference provider,” and it is one word for five different businesses.

This guide maps the five shapes, then walks the top platform in each, using numbers read from the vendors’ own pricing pages on September 10, 2026. By the end you should know which key, or keys, to hold. The short answer for most people who generate text, images, audio and video: one media host, one router, and a local runtime for everything that can stay on the machine.

A provider is where the model runs, and that is a separate choice

A model vendor trains the weights: OpenAI, Anthropic, Google, Meta, DeepSeek, ByteDance. An inference provider owns the GPUs those weights run on and sells the output, by the token, by the image or by the second. Sometimes vendor and provider are the same company. OpenAI hosts GPT-5.6 itself. Often they are not. Meta trains Llama and hosts none of it for you, while Together, Fireworks, Groq, DeepInfra and Amazon all do. ByteDance’s Seedream image model is on Runware, fal and Replicate, priced three ways.

A gateway, or router, is a third thing: a switchboard that holds keys to many providers and resells them behind one API. We covered how gateways work, what they quietly cost and who now owns them in the gateways and routers post. Here OpenRouter is treated as one provider shape among five, and that story is not re-argued.

The shapes matter because they price differently, host different things and fail differently. A media host bills per output and does not charge for a failed generation. A GPU host bills per token, with a cheaper rate for input it has seen before. A hyperscaler bills through a cloud account your finance team already pays. The hero above lays out the five. The rest of this post walks them in the order most readers meet them.

Media hosts bill by the picture or the second

If you generate images, video or audio, read the price unit before the price. A per-image price tells you what a picture costs before you make it. A per-second GPU price tells you only after you know how long the model ran, and for a diffusion model that varies with resolution, step count and whether the model was warm. Both units exist on these three platforms, sometimes on the same page.

RunwareMedia host
Hosts
Image, video, audio, 3D, editing and upscaling, plus a side catalog of text models. The homepage counts 198 image model families, 114 video and 29 audio, and “400K+ models” once community checkpoints are included (runware.ai).
Pricing
Per output. The pricing page puts image generation at $0.0006 to $0.24 per image depending on model, resolution and quality. Serverless compute is also sold by the second: an H100 at $0.000767 per second, a B200 at $0.001386.
Billing
Pay as you go, no prepaid minimum. $2 in credits with a business email and no card. Only successful requests are charged.
API
REST for stateless calls, WebSockets for persistent low-latency sessions, and an OpenAI-compatible layer for tools that expect one.
Strong where
Breadth of media models behind one fixed price per output.
Real limits
Text models are a side catalog, not the main event. Nothing free beyond the $2.

The four illustrations in this post came from Runware’s API while it was being written: Seedream 5 Pro at 2560 by 1440, first result each, $0.39 for all four, 46 to 59 seconds per image. That is the whole point of per-output pricing. The number was known before the request went out.

ReplicateMedia host
Hosts
“Thousands of models contributed by our community,” alongside official models from OpenAI, Google, Anthropic and Black Forest Labs, across images, speech, music, video and language models (replicate.com).
Pricing
Two units on one pricing page. Community models bill by hardware time: a T4 at $0.000225 per second, an A100 80 GB at $0.0014, an H100 at $0.001525 ($5.49 per hour). Official models bill per output: FLUX Dev $0.025 per image, FLUX Pro $0.04, Ideogram v3 $0.09, Wan video $0.09 to $0.25 per second of output.
Billing
Prepaid credit or monthly postpaid. For public models “setup and idle time for the model is free,” and failed runs are not charged (billing docs).
API
Its own prediction API and SDKs. Deployments give you a private endpoint with always-on instances, and in exchange you pay for setup and idle time too.
Strong where
The long tail. If a researcher published it last week, it is probably here.
Real limits
Shared hardware on the long tail means cold boots. The fix, a deployment, changes the billing unit from per output to per hour.
falMedia host
Hosts
“1,000+ production ready image, video, audio and 3D models” (fal.ai).
Pricing
Per output, on the pricing page: Seedream V4 $0.03 per image, FLUX Kontext Pro $0.04, Qwen image $0.02 per megapixel. Video by the second: Wan 2.5 $0.05, Kling 2.5 Turbo Pro $0.07, Veo 3 $0.40. Raw GPUs from $1.89 per hour for an H100.
Billing
Pay only for what you use; no free-credit amount is stated on the pages we read.
API
Its own queue-based API and client libraries.
Strong where
Diffusion speed. fal says its inference engine is “up to 10x faster,” a vendor claim we did not measure.
Real limits
Media only. There is no text catalog worth holding a key for.
A media host sells finished outputs, not GPU time: one price per picture, per second of video, per minute of speech. Illustration generated with Seedream 5 Pro via Runware.

Open-weight GPU hosts turn text into a commodity

Open-weight text is the most competitive market in AI. The same DeepSeek V4 Flash checkpoint is on all four of these platforms, and its input price runs from $0.06 to $0.22 per million tokens depending on which one you ask. Every one of them speaks the OpenAI request shape, so switching is a base URL and a model name. That is why prices converge, and why the differences that remain are speed tiers, cache discounts and billing terms.

TogetherOpen-weight GPU host
Hosts
Open text, vision, image, audio, video and embedding models, plus dedicated endpoints, GPU clusters and fine-tuning.
Pricing
Per million tokens on the pricing page: DeepSeek V4 Flash $0.14 in and $0.28 out, Llama 3.3 70B $1.04 each way, GLM-5.3 $1.40 and $4.40, Kimi K3 $3 and $15. Images $0.0006 to $0.134 each. An on-demand H100 is $3.99 per hour on promotion ($5.49 standard).
Billing
Pay as you go; the pricing page says “start for free” without stating a credit amount.
API
OpenAI-compatible at api.together.ai/v1 “across chat, completions, vision, image generation, text-to-speech, and embeddings” (docs).
Strong where
The full stack in one account, from serverless to reserved clusters to fine-tuning.
Real limits
Price. Its Llama 3.3 70B list price is ten times DeepInfra’s for the same weights.
FireworksOpen-weight GPU host
Hosts
Open text and vision models, Whisper for audio, FLUX Kontext for images, embeddings, fine-tuning.
Pricing
Three serverless tiers per model on the pricing docs. Standard: DeepSeek V4 Flash $0.22 in and $0.66 out, DeepSeek V4 Pro $1.32 and $3.96 with cached input at $0.044, gpt-oss-120B $0.15 and $0.60, Kimi K3 $3 and $15. Priority costs about 25% more; the Fast tier for Kimi K3 is $4.50 and $22.50.
Billing
Postpaid per token, with $1 in free credits to start (pricing page).
API
“OpenAI and Anthropic compatible” (fireworks.ai).
Strong where
Choosing latency per request. The same model has a cheap lane and a fast lane.
Real limits
The fast lane costs 50% more, and on-demand GPUs rose to $8 per H100-hour on September 1.
GroqOpen-weight GPU host
Hosts
A short list of open text models and speech models on its own LPU hardware.
Pricing
The model table lists gpt-oss-120B at $0.15 in and $0.60 out at about 500 tokens per second, and gpt-oss-20B at $0.075 and $0.30 at about 1,000. Llama 3.1 8B is shown at 560 tokens per second with price marked “contact sales.”
Billing
A rate-limited free plan and a paid Developer plan with batch and flex processing (rate limits).
API
“Mostly compatible” with OpenAI’s client libraries at api.groq.com/openai/v1 (docs).
Strong where
Tokens per second. Nothing else on this list publishes speeds in its price table.
Real limits
Text and speech only, and a small catalog. On December 24, 2025 NVIDIA licensed Groq’s inference technology and the founder moved across; Groq says GroqCloud “will continue to operate without interruption” as an independent company (newsroom).
DeepInfraOpen-weight GPU host
Hosts
Open text, embeddings, rerankers, image, speech, music and video models, and it now resells closed models from Anthropic and Google at list price.
Pricing
The lowest list prices in this group on the pricing page: Llama 3.1 8B $0.02 in and $0.04 out, Llama 3.3 70B $0.10 and $0.32, DeepSeek V4 Flash $0.06 and $0.18, DeepSeek V4 Pro $1.30 and $2.60. FLUX-2 Pro images at $0.015. A Flex lane at 0.8 times base price and a Priority lane at 1.5 times.
Billing
A card on file or a prepaid balance is required before any request runs. No free credits are stated.
API
OpenAI-compatible: “just change the base URL and model name” to api.deepinfra.com/v1/openai (docs).
Strong where
Price. If the model is open and you can tolerate the Flex lane, this is usually the floor.
Real limits
No free tier to try it on, and the closed models it resells cost the same as at source.
Four companies, the same weights, four prices. When the model is open, the provider competes on the rack, not the model. Illustration generated with Seedream 5 Pro via Runware.

A router sells breadth for a 5.5% fee

OpenRouter owns no GPUs. It holds keys to the hosts above and to the first-party APIs, and resells all of them behind one OpenAI-shaped key. Stripe’s August 19, 2026 announcement of its acquisition put the count at “400+ models from more than 80 providers.”

OpenRouterRouter
Hosts
Text and vision through chat completions, plus dedicated endpoints for image generation, video generation, text-to-speech and transcription (multimodal docs).
Pricing
The provider’s own price with no inference markup. The fee is 5.5% (minimum $0.80) when you buy credits by card, 5% by crypto (FAQ). Bring your own provider key and the fee is 5% of what the same call would cost on OpenRouter, above a monthly allowance.
Billing
Prepaid credits. Free-tier models allow 50 requests per day, or 1,000 per day once you have bought $10 of credits.
API
Implements the OpenAI specification as a drop-in replacement. If a provider fails it falls back to the next; a :nitro suffix asks for the fastest throughput, :floor for the lowest price.
Strong where
One key, one bill, one request shape, and a provider list that changes weekly without touching your code.
Real limits
It can be no cheaper than the cheapest provider it holds plus the fee, it cannot fix a provider’s cold start, and it now belongs to a payments company.

First-party APIs sell exclusivity, hyperscalers sell procurement

Some models you cannot get anywhere else on launch day, and the first weeks of a frontier model are often the only weeks that matter for a product decision. That is the first-party business. The prices are public and the discounts are structural: cached input and batch processing, not negotiation.

  • OpenAI. GPT-5.6 Terra at $2 in and $12 out per million tokens, GPT-5.6 Sol at $4 and $20, GPT-6 Astra at $10 and $50, with cached input at a tenth of the input price. Image output for gpt-image-2 is $30 per million output tokens. Batch runs at half price (pricing).
  • Anthropic. Haiku 4.5 at $1 and $5, Sonnet 5 at $2 and $10, Opus 5 at $5 and $25, Fable 5.1 at $10 and $50. Cache reads cut input to $0.10 to $0.50 per million; batch is 50% off (pricing).
  • Google. Gemini 3.8 Flash at $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.50. Gemini 2.5 Flash-Lite at $0.10 and $0.40. A free tier exists, with the note that content “may be used to improve our products.” Veo 3.1 video runs $0.05 to $0.60 per second, and Gemini 2.5 Flash Image is $0.039 per picture (pricing).

Hyperscaler gateways sell the same models to a different buyer. Amazon Bedrock’s model catalog lists eighteen providers, from Anthropic and OpenAI to DeepSeek, Qwen, Mistral, xAI and Z.AI, billed on demand per token, at half price in batch, or as provisioned throughput on one-month and six-month terms (pricing). Google’s Vertex AI, which the docs now file under its Gemini Enterprise Agent Platform, hosts Claude Fable 5.1, Opus 5 and Sonnet 5 beside Gemini with Standard, Priority and Flex pay-as-you-go lanes and provisioned throughput (docs). Microsoft Foundry claims “over 11,000” models including OpenAI, Anthropic, Meta, Google and xAI, billed per service (product page).

What you buy from a hyperscaler is not a model. It is one invoice, one identity system, a data-residency choice and compliance paperwork someone already signed. What you give up is the router’s breadth and the GPU hosts’ bargains. The price you see is the list price, and the model you want sometimes arrives there after it arrives at source.

The hyperscaler sale happens in a room like this one. The model is the same; the invoice and the signatures are the product. Illustration generated with Seedream 5 Pro via Runware.

The whole map on one table

Read the table by the column that matters to you. If you generate pictures, the “bills by” column decides. If you run one open text model at volume, “weights” and “free” narrow it to four rows and the price cards above settle it. If you need a signature from procurement, only the last row applies.

Twelve platforms, seven columns
PlatformShapeModalitiesWeightsBills byOpenAI APIFreeStrong where
RunwareMedia hostImage, video, audio, 3D, textOpen + closedPer output; GPU per secondCompatible layer$2, no cardWidest media catalog at fixed per-output prices
ReplicateMedia hostImage, video, audio, textOpen + closedPer second or per outputOwn APISelect models freeThe long tail; anyone can publish a model
falMedia hostImage, video, audio, 3DOpen + closedPer output; GPU per hourOwn APINot statedDiffusion speed, per-output video prices
TogetherGPU hostText, image, audio, videoOpenPer million tokensYesNot statedServerless to clusters to fine-tuning
FireworksGPU hostText, vision, image, audioOpenPer million tokens, tieredYes, plus Anthropic$1Standard / Priority / Fast per request
GroqGPU hostText, speechOpenPer million tokensMostlyFree planTokens per second
DeepInfraGPU hostText, image, audio, videoOpen + resold closedPer million tokens, ×0.8 to ×1.5YesNone statedLowest list prices in the group
OpenRouterRouterText, image, video, speechOpen + closedProvider price + 5.5% on creditsYesFree models, 50/day400+ models from 80+ providers, one key
OpenAIFirst-partyText, image, audio, videoClosed (plus gpt-oss)Per million tokensIs the standardNone statedGPT-5.6 and GPT-6 at source
AnthropicFirst-partyText, visionClosedPer million tokensOwn SDKNone statedClaude at source, deepest cache discount
Google Gemini APIFirst-partyText, image, video, audioClosed (plus Gemma)Per million tokensCompatible endpoint (beta)Free tierCheapest frontier tier, free tier
Bedrock · Vertex · FoundryHyperscalerAll, by modelOpen + closedPer token; provisionedVariesCloud creditsOne invoice, one identity system
“OpenAI API” means the platform accepts the OpenAI request shape at a different base URL. “Free” is what the vendor states on its pricing or docs page; “not stated” means nothing was found there, not that no offer exists. Sources are linked in each platform’s card above.

The scenario list below is the same table read backwards, from the need to the key. It assumes you are a builder or a power user paying your own bill; a company with a cloud contract should start from the last line.

Which key to hold, by what you need
One key for everything text
OpenRouter
400+ models, OpenAI request shape, automatic fallback. Pay 5.5% on top of the provider’s own price.
Cheapest open text model
DeepInfra or Groq
DeepSeek V4 Flash at $0.06 in and $0.18 out per million on DeepInfra; gpt-oss-20B at $0.075 in and $0.30 out on Groq.
Fastest text
Groq
gpt-oss-120B at roughly 500 tokens per second, gpt-oss-20B at roughly 1,000, from Groq’s own model table.
Widest image, video, audio range
Runware, then fal or Replicate
Fixed per-output prices, every major media model, and no charge for a failed generation.
A frontier model nobody else hosts
The first-party API
GPT-6 Astra, Claude Fable 5.1 and Gemini 3.8 Flash arrive at source first; hyperscalers and routers follow.
Enterprise procurement
Bedrock, Vertex or Azure Foundry
One cloud invoice, existing identity and compliance, provisioned throughput. You pay list price for the convenience.

Two things the table cannot show. First, per-output prices and per-token prices are not comparable until you know your task, which is the argument of our post on what one AI image costs. Second, every number here moves. Google’s Gemini 3.8 Flash doubles on January 1, 2027; Fireworks raised on-demand GPU rates on September 1; Together’s H100 price is a promotion. Treat the cards as a snapshot with links, not a contract.

Three keys and a local runtime cover nearly everything

Here is how the map turns into a decision. Hold one media host key, because per-output pricing is the honest unit for pictures, video and sound. Hold one router key, because open text is a commodity and the router lets you follow the price without rewriting anything. Run a local runtime for everything small enough to stay on the machine, which is more than most people expect; the cloud API versus local runtime post draws that line. Add a first-party key only when a model is exclusive, and a hyperscaler only when procurement demands it.

That is the shape CSuite is built around. The desktop app ships Runware, Replicate and OpenRouter as cloud platforms next to local Ollama, Hugging Face and GGML runtimes. A model appears once in the catalog with the platforms able to run it listed beside it, so the same Seedream or DeepSeek entry runs under whichever key you hold. The keys never leave the machine: they sit in the app’s local database, encrypted through the operating system’s keyring, the Keychain on macOS, DPAPI on Windows and Secret Service on Linux.

Three keys on one ring: a media host, a router, and whatever runs locally for free. Illustration generated with Seedream 5 Pro via Runware.

The provider question does not go away after you answer it. Prices move, catalogs change, and a company that owned no GPUs last year is owned by a payments company this year. But it gets easier once you stop treating “provider” as one word. There are five businesses in it, and you only need to be a customer of two.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app