All insights

Inference economics

What Columns Should Be Included in a Downloadable Request-Level Billing Report?

A downloadable request-level billing report should minimally include a request ID, request timestamp, billing period, account or project identifier, model, API operation, measured usage, unit price, currency, calculated charge, processing status, and invoice or reconciliation identifiers. Keep usage, pricing, and charges in separate columns so finance and engineering teams can trace consumption, inspect the applied rate, allocate costs, and reconcile the result without confusing operational telemetry with billable usage.

A downloadable request-level billing report should minimally include a request ID, request timestamp, billing period, account or project identifier, model, API operation, measured usage, unit price, currency, calculated charge, processing status, and invoice or reconciliation identifiers. Keep usage, pricing, and charges in separate columns so finance and engineering teams can trace consumption, inspect the applied rate, allocate costs, and reconcile the result without confusing operational telemetry with billable usage.

The exact schema will vary by provider, contract, modality, and deployment architecture. Teams can adapt the recommendations below to their billing model, using essential billing fields alongside conditional allocation and diagnostic fields where appropriate.

The Minimum Viable Schema for Request-Level Billing

A useful report needs to answer five basic questions for every row:

  1. What request or billable event occurred?
  2. Who or what should own the cost?
  3. What usage was measured?
  4. Which rate and adjustments were applied?
  5. How does the row reconcile to a statement or invoice?

The following categories provide a practical starting point.

CategoryMinimum columnsConditional or optional columns
Request identity and timerequest_id, request_timestamp, billing_period_start, billing_period_end, processing_statusparent_request_id, batch_id
Cost ownershiporganization_id or account_id, project_id or workspace_idenvironment, cost_center, team_id, service_identity, pseudonymous user_id
Workloadmodel_requested, api_operationmodel_served, endpoint_or_deployment, region_or_location, request_type
Measured usageusage_quantity, unit_type, request_countinput_tokens, output_tokens, total_tokens, cached_input_tokens, modality-specific units
Pricingunit_price, currencypricing_tier, rate_card_version, pricing_effective_date
Chargesgross_charge, net_chargeinput_charge, output_charge, cache_charge_or_credit, other_usage_charge, discount_or_credit, adjustment_amount, tax_treatment
Reconciliationline_item_id, source_system, export_generated_at, schema_versioninvoice_or_statement_id, correction_indicator, idempotency_key
DiagnosticsNone by defaultroute_selected, provider_selected, serving_configuration, response_status, error_category, processing_duration_ms, infrastructure attribution

A minimum schema should remain small enough to understand but complete enough to reproduce the commercial logic at the level supported by the billing model. Additional columns are valuable only when their definitions are stable, their values are reliable, and their exposure is appropriate.

Request-level records can support several related workflows:

  • Showback: showing teams how their applications contribute to AI usage and spend.
  • Chargeback: allocating eligible charges to projects, departments, products, or cost centers.
  • Anomaly review: finding unexpected changes in request volume, model selection, or unit consumption.
  • Vendor reconciliation: comparing detailed usage and charges with an invoice or statement.
  • Optimization analysis: examining whether workload and serving choices correlate with cost patterns.

Identify Each Request, Billing Period, and Cost Owner

Every row needs a stable identity. Use request_id as the primary reference for an individual inference request or billable event. If retries, asynchronous jobs, agent steps, or batch submissions can produce multiple related events, add parent_request_id or batch_id without replacing the unique row-level identifier.

Time fields need precise and distinct meanings:

  • request_timestamp records when the request or metered event occurred.
  • billing_period_start and billing_period_end identify the accounting interval containing the charge.
  • pricing_effective_date identifies when the applicable rate became effective.
  • export_generated_at records when the downloadable file was produced.

Do not overload one timestamp to represent all four concepts. Define the time zone, use a consistent machine-readable format, and state whether period boundaries are inclusive or exclusive.

Cost ownership should reflect the organization’s operating model. An organization_id or account_id identifies the contractual owner, while project_id, workspace_id, environment, cost_center, and team_id enable more granular allocation. These fields may originate from an identity system, gateway, application tag, or deployment configuration, so the data dictionary should identify the authoritative source.

Where user- or service-level attribution is permitted, prefer a stable pseudonymous user_id or service_identity. Names, email addresses, API keys, prompts, responses, secrets, and unnecessary personal data should not appear in billing exports by default. Finance usually needs durable allocation keys—not sensitive request content.

processing_status should indicate whether the row is final, pending, failed, corrected, or otherwise subject to change. The allowed values must be documented so downstream systems do not treat an incomplete record as a settled charge.

Record the Model, Operation, and Measured Usage

LLM billing records should identify both the workload and the meter. At minimum, include model_requested and api_operation. When routing is involved, model_served may differ from the model originally requested, so retaining both fields can improve diagnosis when that information is available and appropriate to expose.

Useful workload dimensions include:

  • endpoint_or_deployment for the logical API endpoint, private deployment, or serving target.
  • region_or_location where location is operationally relevant and consistently defined.
  • request_type to distinguish patterns such as interactive chat, batch enrichment, embeddings, or other supported operations.
  • api_operation for the specific operation invoked.

Token-based workloads commonly benefit from separate input_tokens, output_tokens, and total_tokens columns. If cached input is measured, use a separately defined field such as cached_input_tokens rather than subtracting it implicitly from another value. The data dictionary should explain whether totals include cached units and how failed or partially completed requests are handled.

The schema should also accommodate workloads that are not priced in tokens. A general pair such as usage_quantity and unit_type can represent the applicable meter without forcing image, audio, time-based, request-based, or other services into a token-only structure. Do not assume a meter is billable merely because telemetry exists for it.

Include request_count even in a request-level file, typically with a value representing one billable event. This makes aggregation easier and supports architectures where a row may represent an adjusted, grouped, or non-request event. Define token-counting rules, modality units, and null behavior separately rather than relying on column names alone.

Token Forge Cloud Managed Model APIs provide an API-first route to model access and usage data for teams validating demand before considering private deployment. When evaluating any managed API, buyers should confirm which workload and usage dimensions are available, how each unit is defined, and whether reported data can be mapped consistently into their internal cost model.

Connect Usage to Rates, Credits, and the Final Charge

A strong billing report does not collapse measured consumption and money into one opaque value. It records the quantity, the rate applied to that quantity, and the resulting charge as distinct concepts.

Recommended pricing fields include:

  • usage_quantity and unit_type for measured consumption.
  • unit_price for the applicable price per defined unit.
  • pricing_tier or rate_card_version where multiple commercial schedules may apply.
  • currency using one documented currency convention.
  • pricing_effective_date where rates can change over time.

If the pricing model distinguishes input, output, cached, or other usage, preserve separate charge columns such as input_charge, output_charge, cache_charge_or_credit, and other_usage_charge. Otherwise, use a clearly defined aggregate charge instead of inventing unsupported components.

Adjustments should remain visible rather than being silently folded into the total. Depending on the contract, relevant fields may include discount_or_credit, adjustment_amount, gross_charge, tax_treatment, and net_charge. Tax may be applied above the request level, in which case tax_treatment should indicate that treatment rather than allocating a fabricated per-request amount.

Document the calculation order and precision policy. Finance teams need to know whether rounding occurs at the request, daily aggregate, statement line, or invoice level. Small differences can otherwise accumulate even when the underlying usage records agree. Separating these values makes calculation review easier, but it does not guarantee an exact match when invoice-level adjustments or different rounding stages apply.

Add Serving and Outcome Context Without Treating It as Billable by Default

Operational context helps teams explain why costs changed, but it should not be presented as a billing determinant unless documented commercial rules make it one.

Potential diagnostic columns include:

  • route_selected or provider_selected when a routing layer chooses among serving paths.
  • batch_id or another batch-attribution field for grouped processing.
  • cache_status or measured cache usage where definitions are reliable.
  • serving_configuration for a controlled configuration identifier rather than an unstructured settings dump.
  • response_status and error_category for successful, failed, rejected, or partial outcomes.
  • processing_duration_ms for cost and performance investigation.
  • GPU or infrastructure allocation only when it can be measured consistently and is safe to expose.

These columns can help distinguish, for example, a rise in request volume from a shift in model routing or serving configuration. They should not imply that latency, a cache hit, quantization, routing, batching, or GPU time affected the charge unless the applicable pricing model explicitly establishes that relationship.

Token Forge Cloud treats latency-sensitive chat, batch enrichment, and agentic workflows as different serving-policy problems. Token Forge Cloud Private LLM Inference focuses on private deployment and serving-layer optimization using concepts including caching, model routing, batching, quantization, and GPU scheduling. Request-level attribution can help teams investigate how those serving decisions relate to workload economics, provided each exported value has a stable definition. It should be used for analysis rather than as a promise that a particular technique will produce a specific saving or performance result.

Make the Export Reconciliable, Versioned, and Safe to Share

A downloadable report should remain interpretable after it leaves the billing interface. Add metadata that lets downstream systems determine where each record came from, whether it has changed, and how it relates to a statement.

Recommended reconciliation fields include:

  • invoice_or_statement_id to associate records with a commercial document when available.
  • line_item_id as a stable identifier for the detailed billing row.
  • source_system to identify the system that produced the record.
  • export_generated_at to distinguish export time from request time.
  • schema_version so importers can handle structural changes safely.
  • correction_indicator or adjustment_type to identify reversals, replacements, and corrections.
  • idempotency_key or deduplication_key to reduce double-counting risk during repeated imports.

Correction handling deserves an explicit policy. A corrected record might replace an earlier row, reverse it with an offsetting line, or introduce a new adjustment. Consumers need to know which approach is used and which identifier links the records.

The accompanying data dictionary should define:

  • Time zones and billing-period boundaries.
  • Currency codes, decimal precision, and rounding stages.
  • Token definitions and all non-token units.
  • Nullable fields and the difference between zero, unknown, and not applicable.
  • Allowed enum values for statuses, units, and adjustment types.
  • Gross-to-net calculation order.
  • Schema-version changes and deprecation handling.

Data minimization should guide export design. Use pseudonymous allocation identifiers and omit prompts, responses, API keys, authentication tokens, secrets, and unnecessary personal information. More telemetry does not automatically make a report audit-ready or compliant; access policy, definitions, change control, and operational governance still matter.

Recommended Downloadable Header and Data Dictionary

The following extensible header can be adapted to the applicable architecture, contract, and metering model. Conditional fields may be blank or omitted when they do not apply.

request_id,parent_request_id,batch_id,request_timestamp,billing_period_start,billing_period_end,processing_status,organization_id,project_id,workspace_id,environment,cost_center,team_id,service_identity,model_requested,model_served,endpoint_or_deployment,region_or_location,request_type,api_operation,input_tokens,output_tokens,total_tokens,cached_input_tokens,request_count,usage_quantity,unit_type,unit_price,pricing_tier,rate_card_version,pricing_effective_date,currency,input_charge,output_charge,cache_charge_or_credit,other_usage_charge,discount_or_credit,adjustment_amount,gross_charge,tax_treatment,net_charge,route_selected,provider_selected,serving_configuration,response_status,error_category,processing_duration_ms,invoice_or_statement_id,line_item_id,source_system,correction_indicator,idempotency_key,export_generated_at,schema_version

A compact data dictionary can group related fields while still documenting each column in the full implementation:

Field or field groupSuggested type or unitLevelDefinition and caveat
request_idStringMinimumStable identifier for one request or billable event; do not use sensitive request content.
Request and period timestampsISO 8601 datetimeMinimumDefine time zone and period boundaries; keep request, pricing, billing-period, and export times separate.
Account and project fieldsStringMinimumStable allocation identifiers tied to an authoritative source.
Cost center, team, and identity fieldsStringConditionalInclude only when needed and permitted; prefer pseudonymous values.
Model and operation fieldsString or documented enumMinimumIdentify the requested workload; populate served-model and deployment fields only when available.
Token fieldsInteger tokensConditionalDefine counting rules, cache treatment, failed requests, and total-token semantics.
usage_quantity and unit_typeDecimal plus enumMinimumExtensible meter pair for token and non-token pricing models.
Rate fieldsDecimal and stringMinimum or conditionalState the unit basis, currency, effective date, and applicable rate-card identifier.
Charge fieldsDecimal currency amountMinimum or conditionalPreserve gross, adjustments, and net values separately where relevant.
Serving diagnosticsString, enum, integer, or decimalOptionalUse for operational analysis; do not assume these values determine billing.
Reconciliation identifiersStringMinimum or conditionalLink rows to source records, statements, corrections, and repeated imports.
schema_versionStringMinimumIdentifies the field definitions and semantics used to generate the export.

Before implementing the export, test the schema against real workflows: aggregate it by project for showback, join it to a statement for reconciliation, re-import the same file to test deduplication, process a correction, and investigate a model or usage anomaly. Those exercises expose ambiguous units and identifiers faster than adding more columns without clear definitions.

For organizations moving from early API validation to controlled private inference, the same conceptual schema can provide continuity across deployment stages. The implementation details will differ, but stable workload, ownership, usage, and reconciliation dimensions make it easier to evaluate serving economics over time.

Contact Token Forge Cloud to discuss API access, private deployment, and LLM inference cost control.

Contact us