All insights

Inference economics

What a Sub-Agent Trace Should Contain for Parent-Task Attribution

A sub-agent execution should be represented as a child span—or a linked trace unit for asynchronous or shared work—with stable trace, parent-task, span, agent, run, delegation, and attempt identifiers. Its trace should separately record timing, usage and allocated cost, policy decisions, failures, lineage, serving choices, lifecycle events, and effects on the parent outcome.

A sub-agent execution should be represented as a child span—or a linked trace unit for asynchronous or shared work—with stable trace, parent-task, span, agent, run, delegation, and attempt identifiers. Its trace should separately record timing, usage and allocated cost, policy decisions, failures, lineage, serving choices, lifecycle events, and effects on the parent outcome.

The direct answer: identify every execution, decision, resource use, and outcome

An attributable trace must answer more than “what ran?” It should show why the sub-agent ran, which parent task authorized it, what resources it consumed, what decisions affected its behavior, whether it fulfilled its objective, and how its result changed the parent task.

Treat each sub-agent execution as an independently inspectable unit of work. Record structured attributes for facts that describe the execution, events for significant transitions within it, and status information for its final condition. Store sensitive inputs and outputs separately—or retain only controlled references, hashes, and redacted summaries—rather than copying unrestricted content into operational telemetry.

Minimum identifiers for an attributable sub-agent execution

At minimum, record:

  • Trace ID: Correlates work belonging to the same distributed execution.
  • Parent task ID: Identifies the business or application task to which cost and outcomes must ultimately be assigned.
  • Parent span ID: Expresses the immediate causal parent when the sub-agent was invoked directly.
  • Sub-agent span ID: Uniquely identifies this execution unit.
  • Workflow or run ID: Distinguishes one workflow execution from another, including repeated runs of the same task.
  • Agent identity and version: Identifies the agent definition, configuration, prompt template version, or release that ran.
  • Delegation ID: Connects the parent’s assignment event with the sub-agent’s accepted objective.
  • Attempt number: Separates the initial execution from retries, fallbacks, and resumed attempts.

Use stable identifiers rather than names alone. A label such as research-agent may describe a role, but it cannot distinguish versions, concurrent invocations, or repeated attempts.

Why trace IDs alone do not establish complete attribution

A shared trace ID establishes correlation, not necessarily causation or allocation. It does not reveal whether work was initiated by the parent, reused from a cache, shared by several tasks, retried after a failure, or performed speculatively and later discarded.

Complete reconstruction also depends on:

  • Explicit parent-child relationships or links
  • Delegation and attempt identifiers
  • Ordered lifecycle events
  • Usage and cost-allocation rules
  • Policy-decision records
  • Failure-propagation status
  • Outcome acceptance or rejection by the parent

Without these details, an observability system may find related activity while still being unable to explain who should bear the cost, which operation delayed completion, or why a result was excluded.

Separate timing fields by operational stage

A sub-agent trace needs timestamps and durations that distinguish waiting from active work. Useful fields include:

  • Delegated, enqueued, dequeued, started, and completed timestamps
  • Queue duration and scheduler delay
  • Active execution duration
  • Model request duration
  • Tool-call duration, separated by call when practical
  • Policy-evaluation duration
  • Handoff or network wait time
  • Timeout and cancellation timestamps
  • Parent-result acceptance timestamp
  • Critical-path contribution

The parent task’s latency is not the sum of all child-span durations. Parallel agents overlap, and background work may continue without blocking the parent. Reconstruct latency from sequencing, dependencies, and the critical path: the chain of blocking operations that determines when the parent can complete.

For example, suppose a parent task takes 900 ms. Child A runs for 500 ms, Child B runs in parallel for 400 ms, and one tool retry consumes 200 ms within Child A. Adding those child durations would overstate parent latency. The trace should instead show whether the retry extended the blocking path and whether either child completed after the parent already had enough information to proceed.

Record usage and cost with its calculation basis

Cost attribution should capture both consumption and the method used to value or allocate it. Recommended fields include:

  • Model, model version, provider, endpoint, or deployment pool
  • Input and output token counts when available
  • Cached-input, cache-hit, and cache-write status where applicable
  • Tool usage and externally billed operations
  • Compute, accelerator, memory, or infrastructure consumption where measurable
  • Pricing version, rate-card reference, or internal cost basis
  • Currency and estimated or reconciled cost
  • Allocation method for shared infrastructure or batched requests
  • Treatment of failed attempts, fallbacks, speculative work, and retries
  • Whether the amount is estimated, allocated, invoiced, or reconciled

Token counts alone are not an authoritative cost record when caching, batching, shared infrastructure, retries, or changing prices affect economics. Keep the raw usage facts separate from the valuation calculation so finance and engineering teams can recompute totals when the cost basis changes.

Capture policy decisions as structured records

A policy event should explain the decision without placing sensitive policy context into broadly propagated telemetry. Record:

  • Policy name and version
  • Decision outcome, such as allow, deny, redirect, require review, or apply restriction
  • Stable reason code
  • Evaluated principal, service role, or workload class
  • Requested resource and action
  • Applicable rule identifiers
  • Evaluation timestamp
  • Enforcement point
  • Safe reference to supporting evidence
  • Resulting change to routing, tool access, model access, or workflow behavior

A human-readable explanation can help investigation, but stable reason and rule codes are essential for aggregation. Evidence references should point to access-controlled records rather than embedding confidential content in the trace.

Make failures, retries, and propagation visible

A failure record should show where the execution failed and what happened next. Capture the final status, failed stage, sanitized error category, safe error code, retryability, attempt number, timeout or cancellation state, and whether the error propagated to the parent.

Also distinguish among materially different outcomes:

  • The sub-agent failed and the parent failed.
  • The sub-agent failed, but a retry succeeded.
  • The sub-agent failed, and a fallback agent or model completed the assignment.
  • The sub-agent result arrived after cancellation and was discarded.
  • The sub-agent completed, but the parent rejected its output.
  • The sub-agent produced a partial result that the parent used.

Avoid recording raw stack traces, credentials, prompts, tool arguments, or provider payloads by default. Use sanitized details plus a restricted diagnostic reference when deeper investigation is necessary.

Recommended sub-agent trace schema

The following is general design guidance, not a description of a Token Forge Cloud implementation.

Field groupRecommended contentsAttribution question answered
IdentityTrace ID, parent task ID, parent span ID, sub-agent span ID, workflow ID, agent identity and version, delegation ID, attemptWhich execution belongs to which task?
TimingQueue, start, model, tool, policy, handoff, completion, cancellation, acceptance, critical-path contributionWhat delayed the parent?
Usage and costModel or endpoint, tokens, cache state, tool and infrastructure usage, cost basis, currency, estimated cost, allocation methodWhat resources were consumed and how were they valued?
PolicyPolicy and version, decision, reason code, principal, resource, action, rule IDs, evidence referenceWhy was an action allowed, denied, or redirected?
FailureStatus, failed stage, error category, retryability, attempt, timeout, cancellation, propagationWhere did the failure occur and what did it affect?
LineageInput and output references, hashes, redacted summaries, artifact versionsWhich controlled inputs and outputs influenced the result?
Serving metadataRouting decision, cache result, batch reference, quantization choice, compute pool or scheduler referenceWhich serving decisions affected behavior or economics?
EventsDelegation, model and tool calls, policy checks, handoffs, retries, completion, cancellationIn what order did material actions occur?
OutcomeObjective status, result reference, quality or acceptance signal, parent action, contribution classificationDid the work achieve its objective, and was it used?

Only retain serving metadata when it helps explain cost, latency, behavior, or reproducibility. High-cardinality operational details without a defined use can raise storage costs and complicate governance without improving attribution.

Build a causal hierarchy across parent tasks, agents, runs, and attempts

A useful trace model separates the business task from the technical execution hierarchy. The parent task represents the outcome the organization cares about. A workflow run represents one attempt to produce that outcome. Spans and linked units then represent agent, model, tool, policy, and infrastructure activity inside or associated with that run.

Trace ID, parent span ID, sub-agent span ID, and parent task ID

Use parent-child spans for direct causality: the parent delegates work, the sub-agent accepts it, and the parent waits for or consumes the result. Preserve both the technical parent span and the durable parent task ID. Technical traces may be split, sampled, or continued across services, while the task identifier provides the stable business allocation key.

Do not overload one identifier to represent multiple concepts. Trace, task, workflow, span, delegation, and billing-allocation identifiers have different lifecycles and should remain independently queryable.

Workflow, agent version, attempt, and delegation identifiers

A workflow ID groups the execution graph for one run. Agent identity and version explain which behavior was deployed. The delegation ID joins the assignment with its response, while the attempt number distinguishes retries.

This separation is especially important when a retry uses a different model, endpoint, tool, routing policy, or agent version. The parent should be able to show the complete attempt history while identifying which attempt supplied the accepted result.

Links for asynchronous handoffs and shared work

A single parent is not always sufficient. Use links when work is asynchronous, initiated from a queue, shared across parent tasks, or involved in fan-out and fan-in processing. Examples include a background sub-agent whose result is consumed later, a cached computation reused by several tasks, and a batch containing requests from multiple workflows.

A link should preserve association without falsely claiming exclusive ownership. Shared work needs a separate allocation key and rule; assigning its full cost to every linked parent would double count it.

Sequence lifecycle events, not just start and end times

Events make the causal path reconstructable. Useful event types include:

  1. Delegation created and accepted
  2. Queue entry and execution start
  3. Policy check and enforcement action
  4. Routing or cache decision
  5. Model and tool calls
  6. Handoff or fan-out
  7. Failure, retry, or fallback
  8. Result produced and evaluated
  9. Result accepted, partially used, rejected, or superseded
  10. Completion or cancellation

Each event should carry a timestamp, event type, relevant operation ID, attempt, and safe references to associated records. Sequence numbers can help when clocks differ across services, but they should complement rather than replace timestamps.

Prevent double counting in parent-task totals

Define aggregation rules before presenting parent-task cost as a single number. The rules should specify:

  • Whether failed and abandoned attempts remain chargeable
  • How cache reads and cache creation are assigned
  • How batched compute is divided across requests
  • How shared model or tool calls are allocated
  • Whether speculative calls are assigned to the initiating parent
  • How infrastructure overhead is apportioned
  • Which source takes precedence when telemetry and billing records differ
  • How currency conversion and pricing-version changes are handled

Maintain separate measures for gross work, allocated work, and accepted-result work. Gross work shows everything consumed; allocated work applies the chosen accounting method; accepted-result work shows the activity that contributed to the final parent outcome. These views answer different operational and financial questions.

Control lineage and propagated context

Use trace context for causal correlation, but tightly restrict any metadata propagated across services. Propagated context can spread farther than expected and may appear in logs, queues, or third-party systems.

Do not use propagated metadata as a default store for prompts, responses, credentials, personal data, unrestricted tool payloads, or detailed policy evidence. Prefer opaque identifiers and access-controlled references. Apply data minimization, redaction, access restrictions, retention rules, encryption, and separation between operational metadata and sensitive content according to the organization’s architecture and risk model.

For multimodal workloads, lineage references should identify controlled artifacts—such as an image, audio segment, document, or generated asset—without embedding the asset itself in every span. Record the transformation or version relationship needed to understand which artifact informed which execution.

Buyer tests for an agent observability system

Before adopting a tracing or governance approach, ask the system to reconstruct a realistic parent task and verify whether it can:

  • Calculate total allocated cost without charging every parent the full value of shared, cached, or batched work
  • Reconcile estimated usage with an authoritative billing source when one exists
  • Identify the critical path rather than summing parallel spans
  • Explain which policy and rule version produced a decision
  • Locate the failed model, tool, policy, routing, or handoff stage
  • Show whether retries increased cost, changed routing, or delayed completion
  • Distinguish successful execution from parent acceptance of the result
  • Follow asynchronous handoffs across queues and services
  • Limit access to sensitive diagnostic and lineage records
  • Apply deletion and retention controls without destroying essential aggregate accounting records

The strongest evaluation is a replayable scenario containing parallel agents, a failed attempt, a fallback, cached or shared work, and a policy decision. If the system cannot explain both the causal graph and the allocation method, a polished dashboard alone does not establish parent-task attribution.

Connecting trace design to inference economics

Token Forge Cloud focuses on LLM inference economics and control at the serving layer. Token Forge Cloud Private LLM Inference addresses areas such as caching, routing, batching, quantization, and GPU scheduling, while Token Forge Cloud Managed Model APIs provides an API-first entry point for teams validating model demand before considering private deployment.

These serving decisions are relevant to trace design because they can affect usage, latency, allocation, and reproducibility. Teams evaluating an inference architecture should therefore determine which routing, cache, batch, quantization, and scheduling metadata must be exported into their own observability and accounting workflows. Token Forge Cloud also supports private deployment paths where models, prompts, and telemetry remain in the customer’s controlled environment; telemetry handling, retention, access, and integration details should be defined for the intended deployment.

Next Step

Contact Token Forge Cloud to discuss API access, private deployment, and LLM inference cost control.

Contact us