All insights

Inference economics

Designing Webhook Idempotency for Job Completion and Billing Settlement

Treat webhook delivery as at least once , record each event durably, and make job completion and billing settlement independently idempotent. Do not use one global processed flag: one operation may succeed while the other fails. Instead, assign stable keys, maintain operation-specific states, acknowledge the webhook after durable receipt, process downstream work with retries, and reconcile local records against the billing provider’s final status.

Treat webhook delivery as at least once, record each event durably, and make job completion and billing settlement independently idempotent. Do not use one global processed flag: one operation may succeed while the other fails. Instead, assign stable keys, maintain operation-specific states, acknowledge the webhook after durable receipt, process downstream work with retries, and reconcile local records against the billing provider’s final status.

The short answer: deduplicate delivery, but make each business effect independently idempotent

A webhook is a delivery mechanism, not a transaction spanning your job system, usage records, and billing provider. The endpoint may receive duplicate, delayed, concurrent, or out-of-order deliveries. A connection can also fail after the sender transmits an event but before it receives your response, causing a retry even though your system accepted the original request.

The design therefore needs two layers of protection:

  1. Delivery deduplication: recognize that the same provider event has already been received.
  2. Operation idempotency: ensure each logical business operation can be attempted repeatedly without creating an unintended second effect.

For an event that affects both job status and billing, these are distinct operations. A duplicate must not complete the job twice, and it must not create a second charge or settlement request. At the same time, a previous billing failure must not be hidden merely because job completion succeeded.

Why at-least-once delivery must be the operating assumption

Unless a provider contract explicitly says otherwise, design for the possibility that an event will arrive more than once and in an unexpected order. Even a provider that normally delivers quickly and sequentially can retry after timeouts or operational interruptions.

A resilient consumer should be able to receive the same event repeatedly and reach the same intended state. This requires durable database constraints and guarded transitions rather than an in-memory cache alone. A cache may reduce duplicate work, but eviction, restart, failover, or a delayed replay can make it insufficient for financially significant operations.

The deduplication retention period should reflect the provider’s documented retry and replay periods, internal replay procedures, financial record-retention needs, and the lifetime of the underlying business operation. A short arbitrary expiration can allow an old event to produce a new effect later.

Why an HTTP acknowledgment is not business completion

Return a successful HTTP acknowledgment after the request has been authenticated and validated and its receipt has been stored durably. The acknowledgment should mean, “we accepted responsibility for processing this event,” not, “every downstream workflow has finished.”

When the architecture permits, keep the synchronous webhook path short:

  1. Authenticate the sender and validate the event envelope.
  2. Resolve the stable event identity.
  3. Insert or find the durable receipt record.
  4. Create the required local work records or outbox entries.
  5. Commit the transaction.
  6. Return a success response.
  7. Complete job and billing work asynchronously.

If durable storage is unavailable, returning a success response can cause the provider to stop retrying even though the event has effectively been lost. Conversely, waiting synchronously for external billing settlement increases timeout risk and can produce more duplicate deliveries.

Why one global processed flag fails under partial success

Consider this sequence:

  • The webhook is received.
  • The job transition commits successfully.
  • The billing provider is unavailable.
  • The handler sets—or has already set—a single processed=true flag.

On retry, the global flag suppresses all processing, so billing never resumes. If the flag is not set, the retry may repeat both operations and risk duplicating whichever one already succeeded.

Use separate state or idempotency records for each consumer and operation. For example:

  • job_completion: pending, applied, failed
  • billing_settlement: pending, submitted, settled, failed
  • event_receipt: received, dispatching, complete, requires_reconciliation

This makes partial completion visible and allows one operation to retry without reopening the other.

Choose stable keys and model separate operation states

Delivery identity and business-operation identity are related but not interchangeable. A provider event ID identifies a delivered event. A billing idempotency key identifies the logical request sent to the billing provider. A job transition may use a job ID plus a target state or completion version.

Use a provider event ID or documented business-operation key

Prefer an immutable event ID supplied by the webhook provider. If no suitable event ID exists, derive a documented key from stable business fields rather than request time, worker attempt, or a random value generated on receipt.

An illustrative uniqueness scope is:

``text (provider, provider_event_id, consumer, operation) ``

Example operation records might use:

``text (acme-events, evt_123, job-worker, complete-job) (acme-events, evt_123, billing-worker, settle-usage) ``

This design lets the event receipt be deduplicated while preserving independent outcomes for the job and billing consumers. A database uniqueness constraint should arbitrate concurrent inserts. Application-level “check, then insert” logic without a constraint can race when two workers process the same delivery simultaneously.

For the external billing request, derive or persist a separate key representing the same logical financial operation, such as a settlement record ID. Reuse that key on every retry of that operation, subject to the billing provider’s documented key semantics and retention period. Generating a new key for each retry can make the provider treat every attempt as a new request.

Store the original payload or a payload hash and reject conflicting key reuse

The same idempotency key should not silently accept materially different instructions. Retain the original payload where privacy, security, and retention policies allow, or store a canonical hash of the fields that define the operation.

When an existing key is presented again:

  • If the payload or canonical hash matches, treat it as a duplicate or retry.
  • If it differs, reject or quarantine the event as conflicting key reuse.
  • Record enough diagnostic information to investigate without unnecessarily retaining sensitive payload data.

Canonicalization matters. Semantically equivalent JSON can differ in whitespace or field order, so hash a normalized representation or an explicit set of business fields rather than arbitrary raw bytes when appropriate.

A compact illustrative data model could include:

RecordImportant fieldsKey constraintPurpose
Event receiptProvider, event ID, payload hash, received timeProvider + event IDEstablish durable receipt and detect duplicates
Operation attemptEvent ID, consumer, operation, state, attempt countEvent ID + consumer + operationTrack job and billing independently
Job transitionJob ID, completion version, resulting stateJob ID + completion versionPrevent repeated completion effects
Billing operationSettlement ID, usage record ID, provider key, statusSettlement ID or logical charge keyPreserve the identity of the financial operation
Outbox messageAggregate ID, event type, payload, publish stateMessage IDPublish committed work reliably

The precise keys depend on whether one event can represent multiple jobs, usage periods, line items, or settlements. Define identity around the logical operation rather than assuming one webhook always maps to one database row.

Use durable transactions and a transactional outbox where they fit

When the webhook receipt, job transition, and outbox record share one transactional database, commit the related local changes atomically. This avoids a local dual-write failure in which the database commits but publishing downstream work fails—or a message is published before the database transaction later rolls back.

A typical flow is:

``text Webhook provider -> verification and validation -> durable event receipt -> operation-specific state records -> transactional outbox -> job worker -> billing worker -> reconciliation process ``

Inside one local database transaction, the handler can:

  • insert the event receipt if it is new;
  • create pending job and billing operation records;
  • apply an eligible local job transition if appropriate; and
  • insert outbox messages for downstream processing.

A separate publisher reads committed outbox rows and sends them to a queue. Publishing may still happen more than once, so downstream consumers remain idempotent. The outbox addresses the database-and-message dual write; it does not make consumer execution exactly once.

If records live in different databases or services, do not imply atomicity across them. Use durable messaging, idempotent consumers, explicit workflow states, and reconciliation. A distributed transaction may be possible in some environments, but it is not universally available or necessary.

Design retries, ordering, and concurrency explicitly

Retries should distinguish transient failures from terminal failures. Network timeouts, rate limits, and temporary service unavailability commonly justify retry with exponential backoff and jitter. Invalid requests, conflicting key reuse, or an ineligible state transition usually require investigation or correction rather than unlimited retries.

Set operational limits for:

  • maximum attempts or elapsed retry time;
  • exponential backoff and jitter;
  • terminal error classification;
  • dead-letter handling;
  • safe replay authorization;
  • observability for pending and partially failed operations; and
  • retention of idempotency and operation records.

Replay should call the same idempotent operation path as live processing. Avoid a recovery script that bypasses uniqueness constraints or invents a new billing key.

For concurrency, use a uniqueness constraint, conditional update, compare-and-swap version, or row lock appropriate to the database. The goal is to ensure that two workers cannot both move the same logical operation from pending to applied.

For ordering, attach a sequence number, aggregate version, or source timestamp if the provider supplies a meaningful ordering signal. Guard state transitions so an older event cannot move a completed job backward or overwrite a newer settlement status. If a prerequisite event is missing, defer processing or reconcile from the authoritative source rather than guessing.

Handle partial success with workflow states and compensation

A useful state model separates transport receipt from business outcomes. Illustrative states include:

  • received
  • job-completed
  • billing-pending
  • billing-submitted
  • billing-settled
  • partially-failed
  • reconciliation-required
  • reconciled

These states do not need to form one linear sequence. Job completion and billing settlement can progress on separate branches, with an aggregate workflow status derived from their individual results.

When the operations cannot participate in one transaction, use saga-style coordination. Each step commits locally and emits or records the next action. If a later step fails, the system either retries it or applies a defined compensation.

Compensation is business-specific. A billing failure might place an account or usage record into review rather than reverse a successfully completed compute job. If a charge must be voided or refunded, model that as a new auditable financial operation with its own stable idempotency key—not by deleting the original record or mutating history without traceability.

Failure matrix for the two-workflow design

ScenarioPersisted evidenceImmediate responseRetry behaviorReconciliation action
Duplicate webhook deliveryExisting receipt and matching payload hashReturn success if the original was durably acceptedResume only incomplete operationsConfirm both operation states are accounted for
Job succeeds; billing failsApplied job transition and failed or pending billing recordPreserve job completion; mark partial failureRetry billing with the same logical billing keyCompare usage, billing request, and provider status
Billing succeeds; local call times outBilling request key and uncertain local statusDo not create a new logical chargeQuery or retry using the same key under provider rulesImport provider result and close the local operation
Events arrive out of orderEvent versions, timestamps, or guarded state historyDefer or reject an invalid transitionRetry after prerequisites or authoritative lookupRebuild state from source records where necessary
Concurrent duplicate deliveryUnique-key conflict or conditional-update loserTreat one worker as owner; suppress duplicate effectLosing worker reads the committed stateVerify no operation remains pending without work

The most important ambiguous case is billing success followed by a local timeout. The caller does not know whether the provider committed the operation. Retrying with the same provider-recognized idempotency key, or querying by the persisted provider operation reference, is safer than issuing a new request with a new key.

Reconcile job, usage, billing, and provider records

Idempotency reduces duplicate effects, but reconciliation detects gaps that retries alone may not repair. Run a periodic process that compares:

  • completed job records;
  • immutable or append-oriented usage records;
  • local billing-operation records;
  • external billing or settlement status; and
  • unresolved event and outbox records.

Useful reconciliation questions include:

  • Does every billable completed job have the expected usage record?
  • Does each usage record map to one intended billing operation?
  • Are any billing operations stuck in submitted or unknown status?
  • Did the external provider settle an operation that remains pending locally?
  • Are there completed local operations with unpublished outbox messages?
  • Were any events quarantined because their IDs were reused with conflicting content?

Reconciliation should produce repairable cases, not merely alerts. Define whether each discrepancy triggers an automatic retry, an authoritative status query, a compensating operation, or manual review. Preserve an audit trail of the original event, attempted transitions, external references, and repair decision.

Exactly-once business effects are engineered, not assumed

Webhook transport should not be treated as an exactly-once execution guarantee. In practice, the desired business outcome is approximated through several reinforcing controls:

  • stable event and operation identities;
  • durable receipt before acknowledgment;
  • database uniqueness constraints;
  • guarded and versioned state transitions;
  • independently idempotent consumers;
  • consistent reuse of external idempotency keys;
  • transactional outbox publication where applicable;
  • controlled retries and dead-letter replay; and
  • routine reconciliation.

No single flag or queue setting replaces this combination. The system is reliable because duplicate and ambiguous outcomes are represented explicitly and can be resolved—not because duplicates are assumed never to occur.

Applying the pattern to enterprise inference operations

Asynchronous model jobs can produce completion and usage information that later influences reporting, cost allocation, invoicing, or an external billing workflow. Agentic and multimodal workloads make this especially important because one user request may create multiple model calls, tools, retries, or long-running tasks.

Keep the operational facts separate:

  • A job record answers whether the requested inference work completed.
  • A usage record captures the metered activity used for reporting or downstream accounting.
  • A billing operation represents an intended financial action.
  • A settlement record captures the external provider’s result.

Token Forge Cloud Managed Model APIs provide an API-first route to model access and usage data, with a path toward private deployment as workloads become more predictable. Token Forge Cloud Private LLM Inference focuses on private deployment and serving-layer controls such as caching, routing, batching, quantization, and GPU scheduling.

Teams using managed APIs or private inference can apply the architecture in this guide to their own surrounding job, usage, and billing systems. The exact implementation depends on the event source, database, queue, payment provider, privacy policy, and consistency requirements. These architectural patterns can guide the surrounding implementation; Token Forge Cloud does not supply the webhook settlement or ledger functionality described in this guide.

Conclusion

When one event influences both job completion and billing settlement, store the event once but track each business effect separately. Acknowledge only after durable receipt, use an outbox for reliable downstream publication where appropriate, reuse stable operation keys on retries, and make partial failure visible through explicit states. Complete the design with concurrency controls, dead-letter replay, compensation, and reconciliation across job, usage, billing, and provider records.

Contact Token Forge Cloud to discuss API access, private deployment, and LLM inference cost control.

Contact us