devtools.codes

Multi-model Cost Comparison

RUNS LOCALLY

Price one workload across several models at once, using dated official list prices. The table is ordered by cost and names no winner, because the cheapest model for a task depends on whether it is good enough at it — which is a judgement the numbers cannot make for you.

Your tool input is processed locally in your browser and is not intentionally uploaded to our servers. Advertising and analytics providers may still process normal page, device, cookie and network information.

How to use the Multi-model Cost Comparison

  1. Enter tokens in and out for a representative request, including the system prompt.
  2. Enter expected requests per month. The ratio of input to output matters more than either alone.
  3. Select at least two models. Each row carries the date its price was last checked.
  4. Read the multiple of the cheapest option, which shows the spread more usefully than absolute figures.

A worked example, and what it shows

The example prices a classification workload: a long prompt, a short answer, and high volume. That shape is chosen because it is where model choice changes the bill by an order of magnitude rather than a few per cent.

The spread across models is wide, and an advisory says what to do about it — try the smallest candidate first and measure how often it is wrong, rather than assuming it will be. No row is marked as the recommendation, and none is described as best.

Press Example in the workspace above to load it.

Common questions about the Multi-model Cost Comparison

Which model should I choose?

This tool will not tell you, deliberately. It shows what each costs for your workload; whether a cheaper model is good enough at your task is something only testing on your own data can answer. A model that fails more often can cost more once retries and fallbacks are counted.

Why is output priced separately from input?

Because providers price them separately, and output is usually several times more expensive per token. Blending them into one figure would hide which half actually drives your bill. For most workloads shortening responses saves considerably more than shortening prompts, which is not the intuition most teams start with.

How current are the prices in this table?

Each carries the date it was last checked, shown in its row. Providers change pricing without notice, so treat the figures as a planning baseline rather than a quotation, and check the provider’s own pricing page before committing to a number.

What does this exclude?

Free tiers, batch processing, prompt caching and negotiated rates, all of which reduce a real bill. Long-context pricing, tool use, media, retries and taxes, all of which increase it. It covers standard published rates for text input and output, which is the comparable part.

Does my input leave the browser?

No. Your tool input is processed locally in your browser and is not intentionally uploaded to our servers. Advertising and analytics providers may still process normal page, device, cookie and network information. The token counts you enter stay in the page, as does everything else here.