Milestone 9 of 9

Write and audit the near-duplicate review card

State token and group rules, short-document policy, exact threshold, pair scope, evidence limits, sensitivity, replay, and limits on interpretation.

The review card explains a bounded lexical flagging rule. It must not turn a pair decision into a semantic, legal, or automatic deletion claim.

Goal

Write and audit project-root report.md as a concise near-duplicate review card whose claims point to saved evidence and respect every interpretation limit.

Inputs

Use the manifest, selected tokens and excluded summaries, groups and complete occurrences, all pair metrics, flagged pairs, shared evidence, excerpts, sensitivity summaries and additions/removals, figures and plot data, read-back checks, and replay agreement or mismatch.

Deliverables

The card must explain in plain language:

  • corpus, tokenizer, source identity, selected token kind, and both coordinates;
  • group size, exact tuple windows, overlap, unique sets, repeated occurrences, and short-document insufficient_groups policy;
  • pair eligibility, exhaustive scope, generated order, excluded reasons, and exact intersection and union counts;
  • rational threshold cross-products, equality inclusion, decimal display-only status, and flagged-pair ordering;
  • complete shared-group evidence, bounded excerpts, full-token ranges, and display truncation limits;
  • group-size versus threshold sensitivity, primary-setting separation, exact additions/removals, artifacts, figures, read-back, and replay; and
  • what the fixture cannot establish about meaning, deletion, accuracy, scalability, legal equivalence, or best parameters.

Include one small worked path from selected tokens through a group set, pair fraction and threshold products, shared occurrences, and an excerpt. Distinguish observations from interpretation and name any first mismatch honestly.

Checks

Audit every statement against a saved record or explicit rule. Reject wording that calls a flagged pair a confirmed duplicate, semantically equivalent, accurate, legally equivalent, scalable, or suitable for automatic removal. Check that excluded pairs retain reasons, repeated groups retain occurrences, fractions and cross-products control decisions, excerpts do not alter source content, and sensitivity does not choose a winner.

Check the card names primary and audit settings, complete versus rendered evidence, pair scope, deterministic ordering, plot data, artifact read-back, and replay outcome. A reader should be able to locate the evidence for every claim.

Workspace

Keep the card at the project root as report.md. Do not replace it with a similarity gauge, semantic-document icons, or a polished review interface that hides groups and evidence. Preserve machine-readable artifacts and figures.

Hints

HintStart with one pair
Build the worked path from a saved pair row and shared-occurrence record, then explain the general rule. Keep the report traceable.
HintName the fraction
State the intersection and union counts and the integer threshold products. A rounded decimal is display evidence, not the decision rule.
HintKeep the claim small
A flagged pair met this lexical rule on this corpus. That statement is useful; a claim about meaning or automatic action is not supported.

Review

Give the card to someone who has not opened the source code. Can they explain why one pair is eligible, reproduce its boundary decision, find all shared occurrences, distinguish complete from rendered evidence, and tell whether replay agreed?

How to check your work

Checks compare report.md with the evidence-limited reference card after your own audit. The supplied fixture names the lexical contract and its limits without calling flags confirmed duplicates or choosing a parameter.

LLM PrimerWrite and audit the near-duplicate review cardhttps://llmprimer.com/python/projects/find-near-duplicate-documents/write-and-audit-the-near-duplicate-review-card© 2026 LLM Primer