Tool-call costs should be attributed to a single agent run or trace while preserving a separate cost event for every language-model invocation and external API call. The run total should be calculated as the sum of direct model charges, direct external API charges, and any separately allocated share of serving or infrastructure costs. This creates one understandable workflow total without obscuring how each component was consumed, priced, or assigned.
The short answer: roll up costs by agent run without merging the underlying charges
A reliable attribution method separates correlation from calculation. The agent run connects all related activity, but each child event retains its own usage unit, pricing rule, status, and provider information.
A practical rollup is:
> Agent run cost = direct model-call charges + direct external API charges + documented allocation of shared serving or infrastructure costs
This approach is more useful than converting every activity into a token-equivalent metric. Model providers may charge for input tokens, output tokens, cached tokens, images, audio, or other units. External tools may charge by request, record, transaction, search, minute, generated asset, or subscription tier. Preserve those native units and normalize only the resulting monetary amounts into a reporting currency.
Use the agent run or trace as the reporting parent
Every cost-producing event should point to one run or trace. That parent represents a user request, scheduled job, autonomous task, or another defined unit of work.
The reporting parent makes it possible to answer questions such as:
- What did this customer-support workflow cost from start to finish?
- Which model and tool calls contributed to the total?
- Did retries or fallback routing materially affect the run?
- Which tenant, team, project, or application owns the cost?
- Is the run complete, or are asynchronous child operations still pending?
A session can contain several runs, and a workflow can contain multiple steps. Avoid using the broader session as the only allocation object because it can make individual tasks difficult to compare. The run should remain the primary rollup unit even when session-level reporting is also required.
Keep every model invocation and external API call as a separate cost event
Each model invocation should record the model actually used, measured consumption, applicable pricing inputs, and execution result. Each external API call should independently record the service, operation, native usage quantity, and fee calculation.
For example, consider an agent that uses a planning model, calls a paid search API, and sends the returned results to a second model for synthesis:
| Event | Parent | Event type | Native usage | Direct cost | Status |
|---|---|---|---|---|---|
run-4821 | ā | Agent run | Rollup only | $0.046 | Complete |
model-01 | run-4821 | Planning model call | Input and output tokens | $0.008 | Succeeded |
tool-01 | run-4821 | External search API | One request | $0.020 | Succeeded |
model-02 | tool-01 | Synthesis model call | Input and output tokens | $0.018 | Succeeded |
The model call beneath tool-01 illustrates a tool that triggers additional inference. It still rolls up to the same run, but its immediate parent records why the invocation occurred. The $0.046 run value is a calculated total, not another charge to add to the ledger.
In production, the direct-cost calculation should use the price actually applicable to the event whenever that information is available. Public list prices may not reflect contracted rates, volume tiers, credits, commitments, minimum charges, or negotiated bundles.
Treat parent totals as rollups, not additional billable events
Double-counting commonly occurs when systems store both child charges and parent totals in the same ledger without distinguishing their roles. A safer design marks every record as either:
- A billable or allocable cost event, which contributes to a total
- A calculated rollup, which summarizes contributing events but is not added again
The same rule applies at higher levels. A workflow total can summarize several run totals, while a tenant total can summarize many workflows. These aggregates should not be treated as new consumption.
Direct costs should also remain distinct from shared costs. Direct model and API charges can often be associated with individual calls. Shared GPU capacity, orchestration services, observability, networking, and platform operations may require an allocation convention instead.
Possible shared-cost drivers include active compute time, reserved capacity, request volume, weighted model usage, or a fixed allocation by team. No allocation rule perfectly recreates directly metered consumption. Select a rule that is understandable, consistently applied, and documented so finance and engineering teams can interpret the result correctly.
Build a cost-event schema that preserves execution context
A cost ledger becomes useful when each event combines financial information with enough execution context to explain what happened. The schema does not need to depend on one tracing standard, but it should provide stable identifiers and explicit parent-child relationships across the agent runtime, model gateway, external tools, and billing pipeline.
Core identifiers: run, trace, session, agent, workflow step, and environment
The following fields provide a practical starting point. Teams can adapt the names to their existing telemetry and accounting systems.
| Field group | Recommended fields | Why they matter |
|---|---|---|
| Correlation | event_id, run_id, trace_id, session_id | Connects individual charges to a complete unit of work |
| Execution structure | parent_event_id, agent_id, workflow_id, workflow_step_id, tool_call_id, model_call_id | Represents nested calls and prevents ambiguous rollups |
| Ownership | tenant_id, team_id, project_id, application_id, cost_center | Supports operational reporting, showback, or chargeback |
| Runtime context | environment, region, timestamp_started, timestamp_completed | Distinguishes production, testing, and other execution contexts |
| Service identity | provider, service, operation, requested_model, selected_model | Identifies the resource that generated the charge |
| Usage | usage_unit, input_quantity, output_quantity, billable_quantity | Preserves tokens, requests, records, media, time, or other native units |
| Pricing | currency, unit_price, pricing_rule_id, pricing_version, effective_date | Reproduces the calculation using the applicable commercial terms |
| Outcome | status, error_type, retry_number, cache_status, timeout_flag | Explains why a call succeeded, failed, repeated, or avoided execution |
| Reconciliation | provider_account, invoice_period, invoice_reference, source_event_reference | Helps compare internal records with provider billing data |
Not every system will provide every field. The central design principle is to retain the most granular source event available and avoid replacing it with an early aggregate that cannot later be explained.
Cost ownership fields: tenant, team, project, application, and cost center
Technical identifiers explain execution; ownership dimensions explain responsibility. Capture ownership at request time where possible rather than attempting to reconstruct it after monthly invoices arrive.
A shared agent platform may need several simultaneous views. A request can belong to a tenant, originate from an application, be operated by one team, and map financially to another cost center. Storing these dimensions separately allows the same event ledger to support engineering analysis and finance reporting without changing the underlying direct charges.
These records can support showback, where teams see the costs associated with their usage, or chargeback, where costs are transferred to a budget or business unit. The attribution method should work for either approach, but the organization should define which costs are eligible, how shared expenses are allocated, and how disputes or late adjustments are handled.
Record model usage separately from external API fees
A model event and a tool event should not share one generic token_count field. Use an explicit unit type and retain the provider's pricing components.
For model calls, relevant quantities may include input, output, cached, image, audio, or other provider-defined consumption. For external APIs, the unit might be requests, search results, data records, transactions, generated assets, or elapsed processing time. Subscription fees and monthly minimums may need separate period-level records rather than being attached arbitrarily to one request.
Convert charges into a common reporting currency only after calculating them under their native pricing rules. If currency conversion is required, retain the original currency, original amount, exchange rate, and conversion date. This preserves the ability to reproduce the total later.
Handle retries, failures, cache hits, nested calls, and parallel branches explicitly
Operational edge cases can materially change run economics. Represent them as events rather than hiding them inside an unexplained total:
- Retries: Record each attempt separately and connect it to the original logical operation. If the provider charges for failed attempts, preserve that cost.
- Failures: Record measured usage and fees even when the workflow does not produce a usable answer. A failed run can still incur model or API charges.
- Timeouts: Distinguish an application timeout from provider completion. A timed-out client may still receive a billable provider response later.
- Cache hits: Record the hit as its own outcome. If the cache avoids a model call, do not create a fictional model charge; record any actual cache or serving cost under the applicable method.
- Nested calls: Connect child model or API calls to the tool, sub-agent, or workflow step that initiated them while retaining the top-level run identifier.
- Parallel execution: Give every branch a distinct event identifier. Sum all completed billable branches, including results that were generated but not ultimately selected.
- Fallbacks: Preserve both the failed or rejected first attempt and the successful fallback when either created a charge.
Idempotency controls are also important. If the billing pipeline receives the same usage event more than once, a stable source-event identifier should prevent duplicate ledger entries.
Attribute routed calls to the model actually selected
When an agent can use multiple language models, the requested model or routing policy is not enough for cost attribution. Record the model that actually served each call, along with its measured usage and applicable price version.
For example, a workflow might request a general capability class while the router selects different models for planning, extraction, and response generation. Reporting the whole run against the default model would distort model-level economics and make routing decisions difficult to evaluate.
Keep these concepts separate:
- Requested model or policy: What the application asked for
- Selected model: What the serving layer used
- Fallback sequence: Which alternatives were attempted
- Measured usage: What each selected model consumed
- Calculated charge: What that usage cost under the applicable pricing terms
This design allows teams to compare routing policies using actual run outcomes rather than assumptions based on a default model.
Account for asynchronous and multimodal operations
Some tool calls complete after the original agent process has returned. Video generation, batch enrichment, long-running analysis, and other asynchronous operations may begin in one trace and finish later. Keep the run financially open until expected child events reach a terminal state or the organization applies a documented close rule.
Useful event states include provisional, completed, failed, canceled, and adjusted. A provisional estimate should not silently become a final amount. When the provider later supplies final usage, append or update the record with a clear adjustment history.
Multimodal workflows require the same discipline around native units. Do not force image, audio, or video consumption into a token metric merely to simplify dashboards. Calculate each component according to its own pricing rule, then aggregate the monetary charges at the run level.
Preserve raw events and pricing versions for invoice reconciliation
Monthly provider totals rarely align perfectly with a naive sum of public list prices. Differences can come from pricing changes, contract terms, rounding, credits, volume tiers, taxes, minimum commitments, delayed events, or billing-period cutoffs.
Retain three layers of data:
- Raw usage data: The quantities and identifiers returned by the runtime or provider.
- Rated cost event: The pricing rule and version used to calculate the estimated or effective charge.
- Financial adjustment: Credits, commitments, invoice corrections, or allocated period-level fees applied later.
Reconciliation should compare like with like: provider account, service, model or operation, currency, and billing period. Do not overwrite raw events to force them to match an invoice. Record the variance and any adjustment separately so the history remains understandable.
Design reporting views around operational decisions
A single total-cost dashboard is rarely sufficient. The same event ledger should support several views without changing the underlying calculation:
- Per-run cost: The complete cost of answering one request or executing one task
- Per-tool cost: Direct API charges plus any child calls initiated by that tool
- Per-model cost: Actual usage and charges for each selected model
- Per-workflow cost: Aggregated cost across runs and workflow steps
- Tenant or team cost: Ownership-based showback or chargeback views
- Unit economics: Cost per completed task, supported user, processed document, generated asset, or other business output
- Failure cost: Spend associated with errors, timeouts, abandoned branches, and unsuccessful retries
Unit-economics denominators should represent completed business outcomes rather than raw requests alone. If one user action launches several runs, reporting only cost per API request may understate the true cost of delivering the outcome.
Connect attribution to serving-layer cost control
Attribution explains where money was spent; serving policy influences how inference resources are used. The two functions should exchange consistent identifiers, but they are not the same system.
Token Forge Cloud focuses on LLM inference cost control at the serving layer. Token Forge Cloud Private LLM Inference supports private deployment and serving-layer approaches involving caching, model routing, batching, quantization, and GPU scheduling. These controls can influence the cost profile that an attribution ledger measures:
- Caching can change whether a model invocation is required and which usage event should be recorded.
- Model routing can change the selected model and applicable pricing inputs.
- Batching can create a need to allocate shared execution costs across contributing requests.
- Quantization and GPU scheduling can affect private-serving resource economics and the method used to allocate infrastructure costs.
For API-first evaluation, Token Forge Cloud Managed Model APIs offers model access and usage data, with a path toward private deployment as workloads become more predictable. Teams should define their required run identifiers, external-tool records, pricing inputs, exports, and financial workflows as part of implementation planning rather than assuming usage data alone provides complete cost attribution.
The operational goal is a clean handoff: the serving layer produces accurate execution and usage context, while the cost ledger applies transparent pricing and allocation rules. That separation makes it easier to evolve routing or infrastructure policies without rewriting the financial history of earlier runs.
Next step
A sound agent-cost model starts with granular usage events, stable run correlation, versioned pricing, and a clear distinction between direct charges and shared infrastructure allocations. Once that foundation is established, teams can evaluate how API access and private inference policies affect workload economics.
Contact Token Forge Cloud to discuss API access, private deployment, and LLM inference cost control.