Milestone 9 of 9
Write and audit the near-duplicate review card
State token and group rules, short-document policy, exact threshold, pair scope, evidence limits, sensitivity, replay, and limits on interpretation.
The review card explains a bounded lexical flagging rule. It must not turn a pair decision into a semantic, legal, or automatic deletion claim.
Goal
Write and audit project-root report.md as a concise near-duplicate review card
whose claims point to saved evidence and respect every interpretation limit.
Inputs
Use the manifest, selected tokens and excluded summaries, groups and complete occurrences, all pair metrics, flagged pairs, shared evidence, excerpts, sensitivity summaries and additions/removals, figures and plot data, read-back checks, and replay agreement or mismatch.
Deliverables
The card must explain in plain language:
- corpus, tokenizer, source identity, selected token kind, and both coordinates;
- group size, exact tuple windows, overlap, unique sets, repeated occurrences,
and short-document
insufficient_groupspolicy; - pair eligibility, exhaustive scope, generated order, excluded reasons, and exact intersection and union counts;
- rational threshold cross-products, equality inclusion, decimal display-only status, and flagged-pair ordering;
- complete shared-group evidence, bounded excerpts, full-token ranges, and display truncation limits;
- group-size versus threshold sensitivity, primary-setting separation, exact additions/removals, artifacts, figures, read-back, and replay; and
- what the fixture cannot establish about meaning, deletion, accuracy, scalability, legal equivalence, or best parameters.
Include one small worked path from selected tokens through a group set, pair fraction and threshold products, shared occurrences, and an excerpt. Distinguish observations from interpretation and name any first mismatch honestly.
Checks
Audit every statement against a saved record or explicit rule. Reject wording that calls a flagged pair a confirmed duplicate, semantically equivalent, accurate, legally equivalent, scalable, or suitable for automatic removal. Check that excluded pairs retain reasons, repeated groups retain occurrences, fractions and cross-products control decisions, excerpts do not alter source content, and sensitivity does not choose a winner.
Check the card names primary and audit settings, complete versus rendered evidence, pair scope, deterministic ordering, plot data, artifact read-back, and replay outcome. A reader should be able to locate the evidence for every claim.
Workspace
Keep the card at the project root as report.md. Do not replace it with a
similarity gauge, semantic-document icons, or a polished review interface that
hides groups and evidence. Preserve machine-readable artifacts and figures.
Hints
HintStart with one pair
HintName the fraction
HintKeep the claim small
Review
Give the card to someone who has not opened the source code. Can they explain why one pair is eligible, reproduce its boundary decision, find all shared occurrences, distinguish complete from rendered evidence, and tell whether replay agreed?
How to check your work
Checks compare report.md with the evidence-limited reference card after your own
audit. The supplied fixture names the lexical contract and its limits without calling
flags confirmed duplicates or choosing a parameter.