The right AI model, for each task, at the right price.
Pick a task and how you access the models: Routéo ranks the AI models on the market by fit, price and sovereignty, and estimates your monthly bill.
- 30+AI models tracked
- 12typical tasks, scored from 0 to 100
- every daylast price check ·
For which task?
Twelve typical uses, from code to real-time conversation. Each model gets a fit score from 0 to 100.
How do you access the models?
The access mode changes the price you pay, and what the ranking means. The other settings refine the selection.
The three models to keep
For the task and settings you chose.
Every model that passes your filters
The top three are highlighted. Prices in dollars per million tokens; monthly cost calculated on your volumes.
| # | Model | Fit | Input | Output | Average price | $ / month | Context | Speed | Worth knowing |
|---|
A hundredfold gap between the cheapest and the most expensive
Average price on a logarithmic scale; the height of each bar shows the fit for the task.
Choosing the model remains the first budget lever, before any optimisation of instructions or caching.
Every model, unfiltered
Click a header to sort. The provider's name links to its official pricing page; the sovereignty column also gives the weights licence.
The catalogue ignores the range, open-weights, sovereignty and minimum-fit filters. It keeps the input / output split and the OpenRouter fee option, to stay consistent with the ranking.
What the grid does not tell you
Six points that change a budget or a choice, before you commit.
The listed price is not the price you pay
Four gaps keep appearing between the price list and the bill:
- Reasoning tokens are billed as output. On a model that "thinks", this invisible part often exceeds the answer itself: it is the first cause of underestimated budgets.
- Context is billed again on every turn. An agent resends the whole history at each step: a ten-turn task costs far more than ten times the first call.
- Caching changes the rankings. Rereading context already sent costs 0.1 times the price at Anthropic and DeepSeek, 0.25 times at Google and xAI, $1 per million on GPT-6 Astra. Writing to the cache, however, costs 1.25 to 2 times the price: caching only pays off from the second reread, and it stops working if the start of your requests changes order.
- The
max_priceparameter protects you. Set on every production request, it refuses the call rather than let the bill run away.
The input / output split decides the ranking
The average price only makes sense against your real traffic, which varies enormously from one use to another. Four development-agent sessions, measured in-house:
- agent on MiniMax M3: 171,225 input tokens for 493 output, i.e. 0.3% output;
- task creation: 42,556 against 671, i.e. 1.6%;
- Claude Opus 5, heavy reasoning: 86,902 against 9,510, i.e. 9.9%, the highest observed;
- Claude Opus 5, review: 112,058 against 791, i.e. 0.7%.
An agent therefore consumes 90 to 99% input, not 75%. Conversely, writing a text from a three-line instruction reverses the ratio. The slider covers the whole range and defaults to the expected value for the chosen task. Untick "automatic" to test your own traffic: on almost-pure input use, models with cheap input move up.
One limit to know: on three of these sessions, 94 to 99.9% of the input was reread from the cache, billed at 10 to 25% of the full price depending on the provider. The average price shown is therefore a ceiling, not a forecast: on an agent that makes good use of caching, the real bill can fall to a quarter.
Via OpenRouter: what really decides
- No margin on tokens. OpenRouter charges 5.5% on card top-ups, 5% in crypto, and nothing with BYOK — your own provider key — under $25,000 of monthly usage. As soon as you have a provider key, switch to BYOK: you keep routing, fallbacks and statistics, without the fees.
- Default routing favours the cheapest: a host three times cheaper is nine times more likely to be chosen. An advantage on open-weight models, a risk on a sensitive model, as quantisation is not always documented. The
onlyandignoreparameters lock the choice. - Three routing modes: balanced,
:nitrofor throughput, and a mode geared to tool-call accuracy. For an agent that chains tools, this accuracy matters more than the token price. data_collection: "deny"excludes hosts that may train their models on your requests: set it by default as soon as data is sensitive.zdr: truefails on Anthropic models, which do not offer zero retention.:batchvariants cost half as much at Anthropic, OpenAI and Google: anything that can wait 24 hours should go there.
Three access modes, three readings of the table
The usage context chosen in your settings does not just change a fee calculation: it changes what the table means.
- Routed via OpenRouter. What counts: the effective price after routing, batch variants, automatic fallback when a host goes down. The $0.05 floor obtained on DeepSeek through third-party hosts is a real gain, provided you test quality: quantisation is not documented per host.
- Direct API from the provider. The right choice when your customers, your data or your sector require fewer intermediaries: one more American intermediary is hard to justify in a compliance process. The prices remain valid, as a basis for negotiating a direct contract. Two points are negotiated at the same time and do not appear in the table: rate limits, which drive sizing, and the guaranteed service level.
- Self-hosted weights. The per-token price disappears: the cost becomes a fixed server bill, independent of volume. The question is no longer "how much does a token cost?" but "which machine do I need?". The price columns give way to model size, active share, weights footprint and minimum hardware; the ranking follows fit alone, among open-weight models.
Self-hosting: two traps
An open-weight model is not free: its cost changes nature. Below a few hundred million tokens a month, a direct API is almost always cheaper than an amortised GPU server, and comes without the operations on-call. Self-hosting is justified by sovereignty, by a predictable bill or by a volume that truly fills the machine — rarely by price alone.
Second trap, specific to mixture-of-experts (MoE) architectures: the active share drives speed, the total drives memory. GLM-5.3 activates only 39 of its 744 billion parameters, which makes it fast, but all 744 must fit in memory: 756 GB in FP8, i.e. an eight-card server. The "active" column says nothing about hardware; the "parameters" column does.
Four hardware classes: workstation for a consumer card, one 80 GB card for the next tier, multi-GPU node for two to eight cards in one machine, multi-node cluster beyond. The gap is stark: Ministral 3 8B runs on a workstation, Kimi K3 needs 1.56 TB of weights and at least sixty-four accelerators. Yet both are "open weights".
Two licence points to know: Mistral Large 3 is under Apache 2.0 — 675 billion parameters, 41 billion active — and is the best self-hostable sovereign candidate in the table; GLM-5.3 is no longer under MIT, unlike its Flash variant: read Z.ai's own licence before any deployment.
Sovereignty: two questions, not a scale
The word covers two distinct requirements. Before choosing, be clear which one you are after — or ask your customer.
- Who operates the processing? The legal question, the one behind "trusted cloud" policies and the French SecNumCloud qualification: a provider established in the EU, a contract under European law, no exposure to the CLOUD Act. On this point, a model you serve yourself does even better than the provider's API: there is no third party left.
- Where do the weights come from? The industrial question. Mistral wins everywhere; a Qwen or a GLM loses, even when run in France. Some public buyers reject an origin wherever the model runs; others only look at where the data travels.
Only one choice answers both at once: Mistral open-weight models, served on French infrastructure. Mistral Small 4, Ministral 3 8B and Devstral 2 cover general use, edge devices and code, with no per-token cost.
The setting offers two levels. Strict keeps only providers established in the EU. Broad adds open weights you can deploy on your own infrastructure, accepting an origin outside the EU. Everything else is out of scope: Azure, Vertex and Bedrock offer European regions, which means data residency, not sovereignty — the operator remains subject to US law.
One last point: the licence decides what you are allowed to do with open weights. The catalogue shows it, and flags "to be checked" when it is not confirmed for the exact version. Settle it before any commitment. To go further: where your data goes, and how to require that it stays in France.
In production: three tiers, not one model
A serious agent does not run on a single model. The pattern that holds in production separates three tiers: a small model that sorts and prepares, a mid-range model that runs most of the steps, and a frontier model called only for hard decisions or after two failures. For 240 million input tokens and 90 million output tokens a month, the gap between the top and the bottom of the table exceeds a factor of sixty, for the same traffic.
This split works for all three access modes. Via OpenRouter, it is implemented with routing rules. Direct or self-hosted, it is an architecture decision: three providers to contract with, or a self-hosted model for the first two tiers and a direct API for the third. In that last pattern, Mistral Large 3, at $0.50 input and $1.50 output, remains the best sovereignty-price trade-off on the market.
Prices checked every day
An undated grid is a wrong grid: every model carries the date it was last checked.
The prices live in a public repository, data-routeo, in a single file. Routéo reloads it on every visit: when a price moves, the page is up to date without being republished. The file also holds the allowed values for each field, the definition of the sovereignty levels, the official pricing page of each provider, and a dated log of every change, with its source. Model data is written in French; the notes shown here are translated.
Every day, an automatic reconciliation
Every morning at 6 am (UTC), an automated job queries OpenRouter's public API — prices and contexts for all models —, reconciles these figures with the dataset, then strictly checks the result before writing anything: allowed values, numeric prices, locked fields unchanged, no model deleted, an entry dated today in the log. If a check fails, nothing is changed.
Why every day: in a single quarter, GPT-5.6 Luna dropped by 80%, Terra by 20%, a scheduled Sonnet 5 increase was cancelled, DeepSeek introduced off-peak hours, and two frontier models launched at $50 per million output tokens.
Every week, a human review
The automatic reconciliation covers prices and contexts, not the nuances: promotional prices, notes, licences, models missing from OpenRouter such as Mistral's. A weekly review checks these points on the providers' pages.
Known deadlines
- 16 October 2026: Gemini 2.5 Flash-Lite retired.
- 21 November 2026: end of the price guarantee on GPT-5.6 Sol. Not an expiry: it is the first date on which OpenAI reserves the right to change its price. No later price has been published.
- 31 December 2026: end of the launch price for the Gemini Flash family, which doubles on 1 January.
Benchmarks and pricing pages
Fit scores are an editorial rating, calibrated on these public benchmarks: they guide a first choice, they do not replace a test on your own tasks.
- openrouter.ai/models and OpenRouter documentation: fees, routing, caching, BYOK, limits
- platform.claude.com, Anthropic pricing pages
- developers.openai.com/api/docs/pricing
- ai.google.dev/gemini-api/docs/pricing and Google Cloud
- api-docs.deepseek.com: peak and off-peak hours
- docs.x.ai, platform.kimi.ai, z.ai, mistral.ai/pricing
- artificialanalysis.ai, Intelligence Index v4.3
- SWE-bench Verified, SWE-bench Pro, Terminal-Bench 4.0: code
- BFCL v4, MCP Atlas, Tau-Bench, Toolathlon: tool calling
- LongBench v2, MRCRv2, RULER: long context
- OSWorld 2.0, ScreenSpot-Pro: computer use
- MMLU-ProX, MGSM: multilingual
Saturated benchmarks such as SWE-bench Verified no longer separate the leading models: to decide between frontier models, prefer SWE-bench Pro and Terminal-Bench.
Choosing the model is the easy part.
The hard part is connecting AI to your tools, your rules and your teams. Synergetik puts AI agents into service that take over your time-consuming tasks and processes, with published prices and measured gains. Start with a one-hour interview, free of charge: within 72 hours, a written report of what you can stop doing by hand.