Evidence density per token is an evaluator-defined measure of how much validated support a research answer contains relative to its length. For Kimi K3 research tasks, enterprise teams can use it to identify concise, well-supported outputs, but it is not an official Kimi K3 metric and should not be treated as a standalone score for accuracy or research quality. A useful evaluation must define the task, evidence rules, tokenizer, comparison controls, and companion quality metrics before scoring begins.
What Evidence Density per Token Measures—and What It Does Not
Evidence density per token is best understood as a diagnostic efficiency metric. It asks whether an answer delivers supported information economically, rather than simply producing more citations or fewer words.
A high score may indicate that an output contains many supported claims within a relatively small token budget. It does not establish that the answer is complete, that its sources are authoritative, or that its conclusions are suitable for a business decision. A short answer can score well while omitting critical context, and a long answer can score poorly even when its additional detail is useful.
Define the research task and expected answer before scoring
Begin with a written task specification. Without one, reviewers may apply different standards to different outputs or reward answers that are short but fail to satisfy the original request.
For each research task, define:
- The research question: What decision or knowledge need should the answer address?
- The expected format: Should the model produce an executive summary, comparison, literature synthesis, risk analysis, or recommendation?
- The claim unit: Will reviewers assess sentences, propositions, table cells, or another consistently defined unit?
- What counts as evidence: Must support come from accessible primary sources, or are credible secondary sources acceptable?
- The accepted source window: Are recency requirements relevant to the topic?
- The required coverage: Which questions or decision factors must appear for the answer to be considered complete?
- The token convention: Which tokenizer and which portions of the output will be counted?
The unit of analysis also matters. A citation attached to a paragraph may support one claim but not every assertion in that paragraph. Reviewers should break compound statements into independently assessable claims when those statements contain several factual propositions.
Separate evidence efficiency from accuracy and overall research quality
Several concepts that appear similar should remain separate during evaluation:
- Citation presence asks whether the answer includes a citation or source link.
- Citation entailment asks whether the cited source actually supports the associated claim.
- Source authority considers the credibility and relevance of the source for that claim.
- Source diversity measures whether an answer relies on independent sources rather than repeatedly citing one source.
- Claim coverage measures how much of the answer's verifiable content has adequate support.
- Completeness asks whether the answer addresses the important parts of the research task.
- Factual accuracy considers whether claims are correct, including claims that may not require external citations.
- Decision usefulness considers whether the answer provides the context, qualifications, and structure required by its intended reader.
Evidence density combines only selected parts of this picture. It should not collapse all these dimensions into one number. In particular, more citations do not necessarily mean more evidence: repeated, irrelevant, inaccessible, or weakly connected citations can inflate a basic citation count without improving support.
Choose an Operational Formula for the Evaluation
There is no universal evidence-density-per-token formula for Kimi K3 research tasks. Each enterprise should adopt an operational definition suited to its tasks and document that definition so the evaluation can be reproduced.
Candidate formula: weighted supported claims divided by output tokens
One practical candidate is:
Evidence density per 1,000 tokens = (sum of weights for validated supported claims / counted output tokens) × 1,000
This is a proposed evaluator-defined formula, not a vendor or industry standard. Multiplying by 1,000 makes results easier to read; it does not change the underlying relationship.
Consider a hypothetical answer containing 12 reviewable claims. Nine claims pass the support test, and their documented quality weights total 7.5. If the answer contains 1,500 counted tokens, the result is:
(7.5 / 1,500) × 1,000 = 5.0 weighted supported claims per 1,000 tokens
This illustrative score has no universal good-or-bad threshold. It becomes meaningful only when compared with results from equivalent tasks evaluated under the same rules.
An unweighted version may be easier to audit during an initial pilot. Weighting can add useful distinctions, but it also adds reviewer judgment and can make results harder to reproduce.
Numerator options: supported claims, valid links, unique evidence units, and quality weights
The numerator determines what the metric rewards. Common options include:
- Supported claims: Count claims for which a reviewer confirms that the cited evidence entails the claim.
- Independently supported claims: Require corroboration from more than one independent source when the task warrants it.
- Valid source links: Count accessible, relevant links, although this is weaker than claim-level support because a valid link may not substantiate the surrounding text.
- Unique evidence units: Count distinct supporting findings rather than every repeated citation to the same source or fact.
- Quality-weighted evidence: Assign documented weights based on criteria such as source authority, directness, recency, or independence.
A weighting rubric should be defined before outputs are reviewed. For example, a team might distinguish direct support from partial support, but reviewers should not invent weights after seeing which configuration performs best.
Human review remains important for ambiguous cases. Automated citation extraction can locate links and markers, but entailment may depend on qualifiers, scope, dates, definitions, or contradictory passages within the source.
Denominator options: total output, answer body, or content excluding references
The denominator should match the purpose of the evaluation. Possible conventions include:
- Total generated tokens: Includes the answer, headings, citations, bibliography, and formatting.
- Answer-body tokens: Includes substantive prose and in-text citations but excludes the bibliography.
- Content-only tokens: Excludes references and selected formatting syntax.
None is inherently correct for every task. Total tokens align closely with complete output consumption, while answer-body tokens can reduce distortion from different bibliography formats. Whatever convention is selected should remain fixed across the comparison.
Tokenizer consistency is equally important. Token counts can vary with the tokenizer and with the handling of URLs, tables, code, multilingual text, citation markers, and bibliographies. Record the tokenizer version or counting method alongside every evaluation run.
Build a Repeatable Kimi K3 Evaluation Workflow
A reproducible evaluation separates generation, evidence validation, token counting, and reporting. The following workflow can be used for Kimi K3 without assuming any particular model architecture or citation behavior.
- Create a fixed task set. Include representative research scenarios and define the expected answer structure for each one.
- Write the scoring rubric in advance. Define claims, accepted sources, evidence units, entailment rules, weighting, exclusions, and treatment of partially supported claims.
- Control generation conditions. Hold prompts, settings, tool access, retrieval corpus, and source availability constant when comparing modes or configurations.
- Capture complete outputs. Retain the prompt, response, timestamps, configuration identifiers, tool results, citations, and relevant usage records.
- Extract factual claims and citations. Break compound claims into assessable units and map each citation to the claim it appears to support.
- Validate the sources. Check accessibility, relevance, publication context, date, and whether each source supports the associated claim.
- Score claims independently. Mark claims as supported, partially supported, unsupported, contradicted, or not requiring external support according to the predefined rubric.
- Count tokens consistently. Apply the same tokenizer and inclusion rules to every output.
- Calculate density and companion metrics. Keep raw counts so the final score remains auditable.
- Aggregate by task category. Report distributions and inspect outliers rather than relying only on a global average.
A small calibration round can improve reviewer consistency. Have multiple reviewers score the same sample, discuss disagreements, and refine the rubric before evaluating the complete dataset.
Compare Modes or Configurations Fairly
Evidence-density results are not comparable when generation and retrieval conditions change without being documented. If one configuration has browsing or a curated corpus and another does not, the test is measuring a broader system difference—not just model behavior.
| Evaluation control | Hold constant or document | Why it matters |
|---|---|---|
| Task set | Questions, task categories, and difficulty | Different tasks create different evidence opportunities |
| Prompt | Instructions, answer format, and citation requirements | Prompt wording can change length and sourcing behavior |
| Tool access | Search, browsing, retrieval, and external tools | Access affects which evidence is available |
| Retrieval corpus | Documents, versions, filters, and time window | Corpus quality affects support and source diversity |
| Generation settings | Relevant decoding and response controls | Settings can affect verbosity and output variability |
| Scoring rubric | Claim units, entailment rules, and weights | Rubric changes alter the numerator |
| Token method | Tokenizer and inclusion rules | Counting differences alter the denominator |
| Review process | Reviewer training and adjudication | Judgment differences can change support classifications |
Run repeated trials when output variability is material to the use case. Record exceptions rather than silently removing inconvenient results.
Report Evidence Density with Companion Metrics
A useful report shows what produced the score and what the density metric leaves out. Recommended companion measures include:
- Unsupported-claim rate: The share of reviewable claims lacking adequate support.
- Citation precision: The share of citations that validly support their associated claims.
- Claim coverage or citation recall: The share of claims requiring support that receive valid support.
- Source-quality score: A separately defined assessment of source authority, relevance, directness, and recency.
- Source diversity: The number and distribution of independent sources used.
- Completeness score: Coverage of required task elements.
- Latency: Time required to produce an answer under the tested conditions.
- Review effort: Human time needed to validate and accept the output.
- Cost per accepted answer: Total measured inference and review cost divided by answers that pass the acceptance rubric.
Report the per-task distribution, median, range, and sample size. Where the dataset supports it, include confidence intervals; otherwise, provide a clear uncertainty note. Break results down by task category because a score for document extraction may not transfer to open-ended market research or technical synthesis.
Outliers deserve individual review. An unusually high score may reveal excellent evidence efficiency, but it may also indicate citation padding, an unusually short response, or a task with many easy-to-support claims.
Recognize Common Failure Modes
| Failure mode | How it distorts the metric | Detection and corrective action |
|---|---|---|
| Citation padding | Raises citation counts without adding support | Review entailment at claim level and deduplicate repeated evidence |
| Fragmented citations | Makes one source appear to support several separate claims | Map each citation to the exact proposition it supports |
| Repeated sources | Inflates volume without improving independence | Count unique evidence units and report source diversity separately |
| Terse but incomplete answers | Produces a small denominator and an apparently strong score | Apply a completeness requirement before accepting the density result |
| Verbose unsupported synthesis | Expands tokens and introduces weakly supported claims | Track unsupported-claim rate and distinguish analysis from sourced facts |
| Inaccessible sources | Prevents reviewers from validating support | Record accessibility and exclude or separately classify unverifiable sources |
| Irrelevant sources | Creates citation presence without entailment | Require direct claim-to-source validation |
| Bibliography inflation | Changes the denominator based on formatting rather than substance | Use a fixed rule for including or excluding reference-list tokens |
These checks make gaming harder, but no scoring system eliminates judgment. Preserve reviewed examples so future evaluators can apply the rubric consistently.
Connect Evaluation Design to Enterprise Serving Operations
Serving conditions influence whether an evaluation is reproducible and economically meaningful, even though they do not establish evidence quality by themselves. Teams should record model routing, caching behavior, batching, quantization, GPU scheduling, retrieval availability, and other operational settings that could change output conditions, latency, throughput, or measured consumption.
Token Forge Cloud Managed Model APIs can provide an API-first route for teams validating model demand before committing to private serving capacity. Token Forge Cloud Private LLM Inference supports private deployment and serving-layer control through capabilities including caching, model routing, batching, quantization, and GPU scheduling.
These controls can help teams manage evaluation conditions and examine serving tradeoffs. They should not be interpreted as native evidence scoring or citation verification, and they do not by themselves improve factuality, citation entailment, or research completeness.
For workloads moving from experimentation to ongoing operation, evaluate the complete economics of an accepted answer rather than raw token consumption alone. A lower-cost response that requires extensive correction may be less economical than a more expensive response that passes review quickly. Conversely, a highly detailed answer may add token and review cost without adding decision-relevant evidence.
Enterprise Evaluation Checklist
Before operationalizing the metric, confirm that the evaluation design addresses:
- Data control: Where prompts, outputs, retrieved documents, citations, and reviewer annotations are processed and retained.
- Reproducibility: Whether prompts, settings, corpus versions, scoring rules, and token-counting methods can be reconstructed.
- Logging: Which generation, usage, routing, retrieval, and timing records are available for analysis.
- Source retention: Whether cited material or sufficient source metadata can be retained for later review.
- Review effort: How long claim extraction, source checking, and adjudication take.
- Deployment model: Whether managed model API access or private deployment better fits the workload and operating requirements.
- Latency and throughput: Whether evaluation conditions reflect the expected production workload.
- Cost: Inference, retrieval, storage, observability, and human-review costs per accepted answer.
- Governance: Who owns the rubric, approves changes, resolves disputed scores, and reviews drift over time.
Limitations
Evidence density per token rewards supported information relative to length, so it can create an incentive to shorten outputs. That incentive must be balanced with minimum requirements for completeness, clarity, uncertainty disclosure, and decision usefulness.
Scores also depend heavily on task design. Research questions differ in how many claims they require, how available authoritative sources are, and how much explanation a responsible answer needs. Avoid publishing universal thresholds until a task-specific baseline has been validated.
Finally, citation assessment is not fully objective. Reviewers may disagree about whether a source entails a claim or whether a source is sufficiently authoritative. Document uncertainty, use adjudication for consequential evaluations, and retain examples of borderline decisions.