Customers should receive a stable job ID, an explicit lifecycle state, queue-entry and last-updated timestamps, meaningful queue context, work-based progress signals, reserved-credit details, available controls, and a final credit reconciliation. Most importantly, the interface should distinguish credits reserved for a job from credits ultimately consumed or charged. Queue position, progress, and ETA should be presented as qualified operational estimates—not promises—because capacity, priority, batching, routing, retries, and workload size can change during execution.
The minimum visibility customers should receive
A useful customer-facing status contract connects three views of the same asynchronous job:
- Operational state: What is happening to the job now?
- Execution progress: What measurable work has been completed?
- Financial state: What credits are reserved, consumed, released, or awaiting reconciliation?
Customers should not have to combine an ambiguous “processing” label with a separate billing page to understand whether a job is waiting, running, finalizing, or no longer able to proceed.
| Visibility area | Recommended customer-facing information |
|---|---|
| Job identity | Stable job ID, workload type, submission timestamp, and idempotency or request reference where applicable |
| Queue status | Current state, queue-entry time, priority class, qualified work-ahead indicator, capacity-constraint category, and last update |
| Execution progress | Current stage, completed and total work units when measurable, produced outputs, checkpoints, and progress timestamp |
| Credit status | Amount reserved, reservation time and status, estimated or maximum exposure where supportable, amount consumed, amount released, and final reconciliation |
| Controls | Cancellation availability, timeout information, retry state, and links or identifiers for retrieving results |
| Auditability | State-transition history, timestamps, reason codes, human-readable explanations, and initiating actor or system where appropriate |
The amount of detail can vary by workload. A batch enrichment job may report completed records out of a known total. A multimodal generation job might expose stages such as preprocessing, generation, encoding, and output validation. An agentic workflow may report completed steps, tool calls, checkpoints, or subtasks without knowing the final number of steps in advance.
Treat credit reservation and final charging as separate events
A credit reservation means an amount has been set aside for a job under the provider’s billing rules. It should not be displayed as completed consumption or a final charge.
The financial lifecycle should keep these concepts distinct:
- Reservation: Credits are set aside when the job is accepted or prepared for scheduling.
- Consumption: Billable work is recorded according to the applicable usage policy.
- Release: Some or all of the reserved amount becomes available again when it is no longer required.
- Reconciliation: The final record explains the relationship between the amount reserved, consumed, released, and charged.
Where the billing model supports them, customers should be able to see the reserved amount, reservation timestamp, reservation status, current consumed amount, released amount, and final reconciled amount. If the maximum possible exposure can be calculated, the interface should explain whether that figure is a limit, an estimate, or simply the current reservation.
Exact treatment following cancellation, timeout, retry, failure, or capacity loss depends on the provider’s documented billing policy. The status interface should link the operational outcome to the relevant financial record rather than asking users to infer what happened to their credits.
Keep operational status and financial status visible together
Operational and credit records do not need to use the same state names, but they should be correlated through the same stable job ID. For example, a job could be operationally running while its reservation remains active. After execution ends, the job might move to finalizing while usage is still being reconciled.
A useful status view might show:
- Job:
running - Current stage:
generating_outputs - Progress: 620 of 1,000 eligible items processed
- Credit reservation:
active - Reserved amount: visible in the applicable credit unit
- Consumed so far: visible when supported by the billing model
- Last updated: a precise timestamp
This prevents two common misunderstandings: that a reserved amount has already been fully spent, or that a completed computational task has already reached final billing reconciliation.
A job lifecycle that people and systems can interpret
A long-running job needs a state model that works for both automation and human diagnosis. State names may differ between implementations, but the model should preserve the distinction between request acceptance, credit reservation, queueing, scheduling, execution, finalization, and terminal outcomes.
Recommended states from acceptance through completion
A practical illustrative lifecycle is:
accepted → credits_reserved → queued → scheduled → running → finalizing → completed
The lifecycle should also support terminal or exceptional outcomes such as:
failedcanceledexpired
Each state should answer a specific question:
- Accepted: Did the service accept responsibility for tracking the request?
- Credits reserved: Was the required reservation established?
- Queued: Is the job waiting for eligible capacity or another prerequisite?
- Scheduled: Has the job been assigned to an execution window or resource class?
- Running: Has measurable execution begun?
- Finalizing: Is primary execution complete while outputs or billing records are being prepared?
- Completed: Are the result and terminal status available?
- Failed, canceled, or expired: Why did execution stop, and what happens next?
Implementations may combine short-lived states, but collapsing everything into pending, processing, and done usually leaves too much ambiguity. In particular, queued should not imply that execution has begun, and completed should not automatically imply that financial reconciliation is final unless the billing record confirms it.
Retries also need explicit treatment. Customers should be able to determine whether the original job is still active, whether an individual attempt failed, whether another attempt is pending, and whether the retry changes the credit reservation or expected exposure. Attempt IDs can supplement the stable job ID without forcing users to track an entirely new job for every retry.
Timestamps, reason codes, explanations, and audit history
Every state transition should have a timestamp. A current state without transition and freshness information cannot reveal whether the job is progressing normally or whether its status has become stale.
A robust record should include:
- A machine-readable state for automation
- A human-readable explanation for operators
- State-entry and last-updated timestamps
- A structured reason code for waits, failures, or interruptions
- Retry and attempt information where applicable
- A chronological transition history
- Result identifiers or links when outputs become available
- Correlated reservation and reconciliation references
Reason codes should be actionable without disclosing sensitive infrastructure information. Categories such as eligible_capacity_pending, dependency_pending, retry_scheduled, or output_finalization are generally more useful than “delayed.” Human-readable text can explain the next expected event and whether customer action is possible.
An illustrative status payload could look like this:
``json { "job_id": "job_…", "status": "running", "status_message": "Processing the current batch", "submitted_at": "…", "queue_entered_at": "…", "started_at": "…", "last_updated_at": "…", "progress": { "stage": "generate_outputs", "completed_units": 620, "total_units": 1000, "unit": "items" }, "credits": { "reservation_status": "active", "reserved_amount": "…", "consumed_amount": "…", "released_amount": "…", "reconciliation_status": "pending" }, "reason_code": null, "retry": { "attempt": 1, "next_attempt_at": null } } ``
This is an illustrative design pattern, not a Token Forge Cloud API schema.
Queue information that is useful without creating false certainty
Queue visibility should help customers decide whether to wait, cancel, reschedule dependent work, or investigate a constraint. It should not create a false impression that an exact position or completion time is fixed.
Useful queue information can include:
- Queue-entry timestamp
- Current queue or scheduling state
- Priority or service class
- A non-sensitive capacity-constraint category
- Work ahead, queue position, or a range when it remains meaningful
- An estimated start or completion window when supportable
- The timestamp and freshness of the estimate
An exact “position 14” may be less useful than it appears. The jobs ahead may require different models, hardware classes, batch sizes, routing policies, or execution times. Priority changes, retries, canceled work, and newly available capacity can also move a job forward or backward.
For that reason, queue position and ETA should carry clear qualifications. The interface should indicate when an estimate was calculated, what broad factors can change it, and whether confidence is high enough to support operational planning. If no credible estimate is available, an honest capacity state and next update time are preferable to a fabricated countdown.
Token Forge Cloud treats latency-sensitive chat, batch enrichment, and agentic workflows as different serving-policy problems. Token Forge Cloud Private LLM Inference applies workload-aware caching, routing, batching, quantization, and GPU scheduling. These serving-layer considerations illustrate why queue and progress semantics should reflect workload type rather than forcing every job into one reporting pattern.
Report measurable work instead of invented percentages
Progress should be tied to evidence of completed work. Depending on the job, that may mean:
- Stages completed out of a defined sequence
- Records, files, frames, pages, or segments processed
- Agent steps or subtasks completed
- Checkpoints reached
- Outputs produced and validated
- Batches completed out of a known total
A percentage is appropriate only when the numerator and denominator have a defensible meaning. If an agent can dynamically add steps, “60% complete” may be misleading. Reporting “4 steps completed; currently waiting for tool result” is more informative, even if a final percentage is unavailable.
Progress interfaces should also distinguish active work from periods spent waiting on dependencies, retry backoff, eligible capacity, or final output handling. This helps operations teams determine whether the job is computing, blocked, or simply awaiting its next scheduled action.
Provide status through suitable access patterns
Different customers and workloads need different delivery patterns. Common implementation options include:
- Polling endpoints for simple integrations and recovery after client downtime
- Webhooks for event-driven application updates
- Event streams for high-frequency or multi-stage workflows
- Dashboards for operators, support teams, finance users, and business stakeholders
Whichever patterns are used, they should refer to the same stable job ID and consistent state semantics. Event-driven delivery should not eliminate the ability to retrieve current authoritative state, because events can be delayed, duplicated, or processed out of order.
Implementation details should cover authentication, authorization, update cadence, webhook retry behavior, event ordering, rate limits, history retention, and recovery from missed updates. Available mechanisms can vary by Token Forge Cloud access path.
Make cancellation, timeout, retry, and failure outcomes explicit
A cancellation request and an effective cancellation are not necessarily the same event. The interface should show when cancellation was requested, whether execution could be stopped, how much work had completed, and the eventual terminal state.
Similar clarity is needed for timeout, failure, retry, expiration, and capacity loss. Customers should be able to see:
- What happened and when
- Whether execution stopped or another attempt is scheduled
- Which outputs, if any, remain available
- Whether the credit reservation remains active
- Whether reconciliation is pending or complete
- Where to find the final financial record
The interface should not imply a refund, release, forfeiture, or charge merely from the operational state. Those outcomes must follow the documented commercial and billing policy.
Protect tenant and infrastructure information
Queue transparency should remain tenant-safe. Customers need enough information to understand their own job, but they should not see other tenants’ identities, workload details, prompts, queue entries, or resource allocations.
Role-aware access is also important. An application operator may need technical failure details, while a finance user may need reservation and reconciliation records. The status model should limit each user or service account to appropriate job and billing information while retaining a useful audit trail.
Long-running AI job visibility checklist
A clear implementation should provide the following:
- Every accepted job receives a stable identifier.
- Reservation is distinct from consumption and final charging.
- Queue, execution, and financial states are visible together.
- State transitions include timestamps and freshness indicators.
- Progress is based on stages, work units, checkpoints, or outputs.
- Queue position and ETA are qualified rather than presented as guarantees.
- Cancellation, timeout, retry, failure, and expiration have explicit outcomes.
- Reason codes are machine-readable and explanations are useful to people.
- A final record reconciles reserved, consumed, and released credits.
- Status can be retrieved after missed notifications or client downtime.
- Access is role-aware and does not reveal other tenants’ activity.
- Different agent, multimodal, batch, and routed workloads can use appropriate progress semantics.
Token Forge Cloud Managed Model APIs offer an API-first path with managed model access and usage data for teams evaluating demand before private deployment. Token Forge Cloud Private LLM Inference provides a serving-layer control plane for private LLM deployments. For long-running workloads, the exact job-status interfaces, lifecycle fields, queue signals, cancellation semantics, and credit policies depend on the intended deployment and should be confirmed before implementation.
Next Step
Contact Token Forge Cloud to discuss API access, private deployment, and LLM inference cost control.