Cost per task for model configurations
Illustrative values entered in your browser. No real records.
What it shows
Compares the cost per task and the monthly cost of up to three configurations, through an API or with local inference, starting from the monthly volume and the tokens per task. It shows which configuration costs least at the entered volume, computes the break-even between a local and an API configuration, first against the cheapest one, and shows how the cost changes with volume.
The demonstration
Results computed in the page
| Configuration | Cost per task | Monthly cost | Capacity per unit | Units needed |
|---|
Local–API break-even (first against the cheapest configuration)
How it works
For an API configuration, the cost per task is tokens_in × price_in / 1,000,000 + tokens_out × price_out / 1,000,000, in EUR. The monthly cost is that cost multiplied by the monthly volume.
A local configuration has a fixed monthly cost per unit: amortisation, power and hosting. A task occupies a unit for tokens_in / prompt_throughput + tokens_out / generation_throughput seconds, with separate throughputs because reading the prompt is usually much faster than generating. The capacity of one unit is 2,592,000 seconds (30 days) × utilisation, divided by that time. If the volume exceeds the capacity, N = ⌈volume / capacity⌉ units are needed. The monthly cost is then N × the cost of one unit, so it rises in steps.
The break-even is the monthly volume at which the local cost equals the API cost: unit_cost / API_cost_per_task, rounded up to a whole task. It exists only if that volume fits in a single unit; otherwise the local configuration never becomes cheaper. After the first break-even, the page also shows the intervals, just after each step, in which the API becomes cheaper again.
The calculator holds no model names, vendor names or quality scores. Choosing a configuration also depends on the quality of results on the real task, on confidentiality and on control. Virtual Soft evaluates these criteria separately in consulting, comparing open models hosted locally with EU API services.
What is real and what is simulated
Real, computed in the page:
- the cost per task and the monthly cost of each configuration, recomputed on every change;
- the capacity per unit, the number of units needed and the load;
- the cheapest configuration at the entered volume and the break-even between each local configuration and each API configuration (first against the cheapest), with the intervals in which the advantage reverses;
- the sensitivity chart and its table;
- the downloadable technical sheet, with the formulas and the values entered.
Simulated:
- the pre-filled prices, hardware costs, throughputs and utilisation are illustrative values, not the prices or measurements of any particular provider;
- the model assumes constant average throughputs, separate for prompt processing and generation, and a fixed monthly cost per unit; it does not include volume discounts, caching, staff costs or traffic peaks.
Last verified: