Cost per task for model configurations

Functional — runs in your browser

Illustrative values entered in your browser. No real records.

What it shows

Compares the cost per task and the monthly cost of up to three configurations, through an API or with local inference, starting from the monthly volume and the tokens per task. It shows which configuration costs least at the entered volume, computes the break-even between a local and an API configuration, first against the cheapest one, and shows how the cost changes with volume.

The demonstration

Cost per task calculator All amounts are in EUR, excluding VAT. A month has 30 days.
Workload
Average per task: the text sent and the answer generated.
Configuration A — API
Type

illustrative values — enter the provider's current prices

Configuration B — API
Type

illustrative values — enter the provider's current prices

Configuration C — local
Type

illustrative values — enter your own costs and measurements

Amortisation, power and hosting, as one amount.
Aggregate throughputs, with requests running in parallel, measured on the real task. Reading the prompt is usually much faster than generating.

Results computed in the page

Capacity per unit is in tasks per month. API configurations have no capacity limit in the model.
ConfigurationCost per taskMonthly costCapacity per unitUnits needed

Local–API break-even (first against the cheapest configuration)

    Monthly cost in EUR for each configuration, on a logarithmic volume axis. API configurations rise linearly; local ones rise in steps, one unit at a time. The values are in the table below.
    The formulas, the values entered and the results, generated in your browser.
    Monthly cost (EUR) per configuration, at volumes chosen on a logarithmic scale

    How it works

    For an API configuration, the cost per task is tokens_in × price_in / 1,000,000 + tokens_out × price_out / 1,000,000, in EUR. The monthly cost is that cost multiplied by the monthly volume.

    A local configuration has a fixed monthly cost per unit: amortisation, power and hosting. A task occupies a unit for tokens_in / prompt_throughput + tokens_out / generation_throughput seconds, with separate throughputs because reading the prompt is usually much faster than generating. The capacity of one unit is 2,592,000 seconds (30 days) × utilisation, divided by that time. If the volume exceeds the capacity, N = ⌈volume / capacity⌉ units are needed. The monthly cost is then N × the cost of one unit, so it rises in steps.

    The break-even is the monthly volume at which the local cost equals the API cost: unit_cost / API_cost_per_task, rounded up to a whole task. It exists only if that volume fits in a single unit; otherwise the local configuration never becomes cheaper. After the first break-even, the page also shows the intervals, just after each step, in which the API becomes cheaper again.

    The calculator holds no model names, vendor names or quality scores. Choosing a configuration also depends on the quality of results on the real task, on confidentiality and on control. Virtual Soft evaluates these criteria separately in consulting, comparing open models hosted locally with EU API services.

    What is real and what is simulated

    Real, computed in the page:

    • the cost per task and the monthly cost of each configuration, recomputed on every change;
    • the capacity per unit, the number of units needed and the load;
    • the cheapest configuration at the entered volume and the break-even between each local configuration and each API configuration (first against the cheapest), with the intervals in which the advantage reverses;
    • the sensitivity chart and its table;
    • the downloadable technical sheet, with the formulas and the values entered.

    Simulated:

    • the pre-filled prices, hardware costs, throughputs and utilisation are illustrative values, not the prices or measurements of any particular provider;
    • the model assumes constant average throughputs, separate for prompt processing and generation, and a fixed monthly cost per unit; it does not include volume discounts, caching, staff costs or traffic peaks.

    Last verified: