All insights

Inference economics

How to Compare Input, Output, Cache, Tool, Media, and Other AI Inference Costs

The most useful comparison is a line-item cost ledger normalized to the same workload and reporting period. Record each billable component in its native unit, apply the relevant unit rate to expected usage, and calculate its extended cost. Then compare total cost per request, per successful task or business outcome, and per month—not just cost per million tokens.

The most useful comparison is a line-item cost ledger normalized to the same workload and reporting period. Record each billable component in its native unit, apply the relevant unit rate to expected usage, and calculate its extended cost. Then compare total cost per request, per successful task or business outcome, and per month—not just cost per million tokens.

The most useful breakdown: line items, workload assumptions, and extended cost

A single token price is rarely enough for production planning. It can obscure differences between input and output pricing, cache behavior, tool use, media processing, infrastructure commitments, and operational overhead. A useful comparison starts with an identical workload definition and keeps each cost driver separate until the final total.

Use one row for each distinct billing unit

Create a row for every component that may appear on a provider statement or contribute to a private deployment's operating cost. Do not force unlike units into an artificial token equivalent. Images may be billed per image, audio by duration, search by call, and private infrastructure by GPU or instance time. Retain those native units and convert them into workload-level cost only after expected usage has been calculated.

A reusable ledger can use the following structure:

ComponentNative billing unitUnit rateExpected monthly usageExtended monthly costCost typePrice source and access dateExclusions or notes
Input or uncached inputProvider-defined token unitCurrent documented rateWorkload estimateRate × usageVariablePrimary pricing source; date accessedModel, region, and tier
OutputProvider-defined token unitCurrent documented rateWorkload estimateRate × usageVariablePrimary pricing source; date accessedInclude expected output growth
Cache readsProvider-native cache unitCurrent documented rateCache-hit workload estimateRate × usageVariablePrimary pricing source; date accessedRecord eligibility rules
Cache writes or storageWrite unit, token-time, or storage unitCurrent documented rateWorkload estimateRate × usageVariablePrimary pricing source; date accessedInclude retention assumptions
Tools and retrievalCall, query, execution time, or pass-through unitCurrent documented rateCalls or runtime estimateRate × usageVariablePrimary pricing source; date accessedSeparate third-party fees
MediaImage, duration, frame, or other native unitCurrent documented rateMedia-volume estimateRate × usageVariablePrimary pricing source; date accessedRecord quality or format tier
Compute or capacityGPU-hour, instance-hour, or committed-capacity unitContract or infrastructure rateProvisioned or consumed capacityRate × usageFixed, committed, or variableContract or infrastructure sourceInclude utilization assumption
Commercial and operating costsContract-defined unitApplicable rateReporting-period estimateApplicable calculationFixed or variableContract, invoice, or tax sourceSupport, transfer, tax, or minimums

Providers may not charge every category. Keep the row visible and mark it not charged or not applicable rather than deleting it. This makes differences in billing scope explicit and helps prevent a missing line item from being mistaken for a zero-cost service.

Show unit rate, expected usage, and monthly extended cost separately

Separating rate from usage makes the comparison easier to audit and update. A low unit rate can still produce a high monthly cost if output is verbose, cache eligibility is limited, tools are invoked frequently, or traffic is routed to a higher-cost model. Conversely, a higher listed rate may not determine outcome cost if the workload needs fewer attempts to complete successfully.

Use this general calculation:

Total monthly cost = token costs + cache costs + tool costs + media costs + infrastructure costs + storage and transfer + commercial charges + applicable taxes

Include only the components that apply, but document exclusions. For each line item:

Extended cost = unit rate × expected usage in the provider's native billing unit

After summing the ledger, calculate decision-oriented metrics:

  • Cost per request: total applicable cost divided by total requests.
  • Cost per successful task: total applicable cost divided by tasks that satisfy the defined completion criteria.
  • Cost per business outcome: total applicable cost divided by completed outcomes such as resolved cases, processed documents, or accepted content units.
  • Monthly cost: variable usage plus allocated fixed and committed costs for the reporting period.

The denominator matters. If a workflow requires retries, human rework, or repeated tool calls, nominal request cost may understate the cost of producing a usable result. Define success before comparing services, and evaluate quality and task completion alongside price.

Keep non-applicable components visible but marked as not charged

A consistent taxonomy improves statement reconciliation. If one service includes a capability in another meter while another bills it separately, note that treatment instead of blending the charges without explanation. Useful status labels include:

  • Charged separately
  • Included in another rate
  • Not charged
  • Not used by this workload
  • Requires contractual confirmation

Record the currency, region, model or version, service tier, price source, access date, and any volume or commitment condition for every rate. This metadata is essential because pricing pages, model versions, regional terms, and negotiated arrangements can change.

Normalize all options to the same workload

Define a workload scenario before entering rates. At minimum, document request volume, average input and output size, expected cache-hit behavior, tool-use frequency, media volume, concurrency, latency needs, and service tier. Also identify whether the workload is latency-sensitive chat, batch enrichment, an agentic workflow, or another operating pattern.

An illustrative worksheet might describe:

  • A fixed monthly request forecast with separate normal and peak periods.
  • Average uncached input and generated output per request.
  • An expected share of eligible requests served from cache.
  • Tool and media use expressed as calls, runtime, files, images, or duration per request.
  • Required concurrency and whether work can be queued or batched.
  • Retry, failure, and successful-completion assumptions.

These assumptions should be identical across providers unless a technical difference requires a documented adjustment. If one model needs a different prompt, produces longer answers, or completes the task at a different rate, reflect that in the workload rather than pretending the requests are equivalent.

Separate every variable usage meter that can change the bill

The ledger should preserve the relationship between operational behavior and billed usage. This makes it possible to identify whether cost movement comes from traffic, context size, generated output, cache performance, agent behavior, media volume, or infrastructure utilization.

Input, output, uncached input, cache reads, and cache writes or storage

Keep input and output separate because they may have different rates and usage patterns. Output length is often especially sensitive to application behavior, response policies, and user demand. Track system instructions, retrieved context, conversation history, and repeated context where they contribute to billed input.

Cache accounting should also remain granular when applicable:

  • Uncached input: content processed without a qualifying cache reuse event.
  • Cached-input reads: reusable content served under the provider's cache rules.
  • Cache writes: content initially written or prepared for later reuse.
  • Cache storage or retention: time- or capacity-based charges, where present.

Do not assume every cache hit has the same economic treatment. Verify eligibility windows, minimum content requirements, write charges, retention rules, and model-specific conditions in current primary documentation. For planning, calculate cached and uncached volumes from the workload rather than applying a discount to all input.

Context growth deserves its own assumption. Multi-turn conversations, retrieval-augmented prompts, and agent histories can increase input over time. An average based only on the first request may therefore understate production usage.

Tool calls, execution time, search, retrieval, code execution, and pass-through fees

Agentic applications can incur costs beyond model tokens. Separate tool-related charges according to the unit actually billed, including calls, queries, runtime, documents retrieved, code-execution sessions, or third-party fees.

Tool frequency should be measured per attempted and successful task. A workflow that repeatedly searches, retries code, or calls an external API may have a materially different outcome cost from one that completes with fewer steps. Where third-party services are involved, distinguish the model provider's charge from pass-through or separately contracted charges.

Retries and failed requests should not be hidden in an average token figure. Track their input, output, tool, and media consumption, then include that usage in the total cost divided by successful tasks.

Preserve native units for image, audio, video, and embeddings

Media and multimodal services may use units that do not map cleanly to text tokens. Record image generation or analysis in the documented image unit, audio in duration or another provider-defined measure, and video in its applicable duration, frame, resolution, or generation unit. Embeddings should likewise retain their documented billing basis.

Convert each category into cost for the shared workload:

Media workload cost = native media quantity × applicable unit rate

Document the format, quality setting, duration, resolution, or processing mode when it affects the rate. This avoids comparing a low-resolution asynchronous job with a higher-service-tier real-time workload as though they were equivalent.

Add fixed, committed, and operational costs

Managed API access is commonly dominated by variable usage, while self-deployed model serving and a private inference control plane may include fixed or committed infrastructure. Compare both using the same reporting period, but do not collapse their economics into the same unit-rate column.

Private-inference analysis may include GPU or instance time, idle capacity, storage, data transfer, orchestration, observability, support, and minimum commitments. Allocate shared fixed costs to the workload using a documented method. For example:

Allocated infrastructure cost per task = workload share of fixed and committed cost ÷ successful tasks

Utilization is critical. Provisioned capacity that remains idle still contributes to cost, while high concurrency may change capacity requirements. Batch processing, latency targets, traffic peaks, and headroom policies should therefore be included in the scenario.

Test the assumptions that can change the result

A base case alone can create false precision. Run sensitivity analysis around the factors most likely to move effective cost:

VariableLower-case scenarioBase caseHigher-case scenarioWhat it tests
Output lengthShorter responsesExpected responsesLonger responsesExposure to generated-token growth
Cache-hit rateLimited reuseExpected reuseGreater eligible reuseSensitivity to cache behavior
Tool frequencyFewer callsExpected callsMore calls or retriesAgent and third-party cost exposure
Traffic volumeLower demandForecast demandPeak demandVariable spend and commitment fit
Infrastructure utilizationMore idle capacityPlanned utilizationHigher utilizationPrivate-serving unit economics

Routing and batching should also be modeled as operational policies rather than assumed benefits. Routing can change the mix of models used, while batching can change how work consumes serving capacity and interacts with latency requirements. Quantization and GPU scheduling may affect infrastructure planning, but the financial effect should be evaluated against measured workload behavior rather than treated as an automatic saving.

Connecting the ledger to Token Forge Cloud

Token Forge Cloud approaches inference economics at the serving layer rather than treating raw token price as the complete cost picture. Token Forge Cloud Private LLM Inference supports evaluation of serving controls including caching, model routing, batching, quantization, and GPU scheduling. Teams can map these controls to the ledger's measured drivers—for example, cache-eligible input, model mix, batchable demand, and infrastructure utilization—without assuming a predetermined financial result.

Different workloads require different serving policies. Latency-sensitive chat, batch enrichment, and agentic workflows should be modeled separately because their concurrency, output, tool use, and service requirements differ. This workload-level view also helps teams decide which traffic is predictable enough to evaluate for private deployment.

For organizations still validating demand, Token Forge Cloud Managed Model APIs provides an API-first route to model access and usage data. That usage history can help establish workload assumptions before comparing ongoing managed API consumption with private deployment. The comparison should retain the same success criteria and reporting period on both sides.

A practical reconciliation workflow

Use the ledger as a repeatable monthly process rather than a one-time pricing exercise:

  1. Define the workload and success measure. Specify what counts as a request, successful task, and business outcome.
  2. Capture current pricing metadata. Record the source, date, currency, region, model version, tier, commitments, and exclusions.
  3. Map every charge to a native meter. Keep tokens, calls, runtime, media, capacity, storage, transfer, and commercial charges separate.
  4. Reconcile forecast with actual usage. Investigate changes in context length, output, retries, cache behavior, routing, tools, traffic, and utilization.
  5. Recalculate effective unit economics. Report cost per request, successful task, outcome, and month.
  6. Run sensitivity cases. Test whether the preferred option changes under realistic demand and behavior changes.

This process produces a more useful comparison than an undated vendor price table. It shows not only what was charged, but which workload and serving decisions generated the cost.

Next step

Contact Token Forge Cloud to discuss API access, private deployment, and LLM inference cost control.

Contact us