AI evaluation · Framework review · 31 Aug 2026
A. Jacobs argues that fabrication is the wrong headline failure for AI language systems and erosion is the right one. He argues too that the operations which look safest — summarising, paraphrasing, simplifying — do the most damage. He is right on both counts, and the evaluation stack really cannot see it. But he calls this a measurement framework, and in practice it is a set of proposals with one demonstration behind it, whose strongest borrowed citation points at the wrong volume of Nature.
A transformation can keep every fact correct and still lose what the text was for. Jacobs’s split between accuracy and fidelity names that gap, and it survives contact with practice.
Semantic Fidelity is the degree to which meaning, intent and interpretive structure survive a transformation — retrieval, compression, summarisation, generation, reuse. His one-line version: “Accuracy ensures correctness. Semantic fidelity ensures meaning.” Put more usefully, accuracy asks whether the output is true and fidelity asks whether the point survived.
Both come apart constantly, and once you hold the distinction you stop being able to unsee it. For example: a summary of a bug report that lists every symptom correctly and drops the sentence explaining which symptom the reporter thought was weird. Nothing in it is false. Anyone acting on it investigates the wrong thing.
Jacobs turns this into four questions to put to any transformation:
His fourth question is the only one treating communication as having a destination, and no automated metric attempts it.
The operations that look safest — summarising, paraphrasing, simplifying — do the most damage, because compression drops meaning before it drops facts. A hollowed summary announces nothing, so nobody checks it.
Jacobs states it plainly in SFL-04: “The safer and more efficient a task appears, the more likely it is to degrade semantic fidelity.” He frames the stakes with the same inversion. “Meaning often collapses long before models hallucinate.”
Nothing about the mechanism is mysterious. Compression works by deciding what to drop. Easiest to drop are exactly the things carrying meaning rather than information — the qualification, the aside about why this mattered, the ordering that made an argument an argument. Whatever remains is shorter, correct, and load-bearing in a way nobody notices until something gets built on it.
“The safer and more efficient a task appears, the more likely it is to degrade semantic fidelity.”
Two things make this operationally nastier than fabrication. Fabricated citations announce themselves the moment somebody checks them, so that error has a natural discovery event. Hollowed summaries announce nothing, and they get forwarded. For instance the bug-report summary in §01 would pass every review a team actually runs on it. And summarisation feels like the low-risk task, so people delegate it without review. In most pipelines the highest-damage operation is therefore also the least supervised one.
Jacobs gives this a name worth stealing: meaning debt, the accumulated cost of locally convenient omissions. Each shortcut looks manageable at the time. Payment comes due when a later system depends on the compressed result and the omitted structure cannot re-enter.
Lexical, structural, contextual and relational decay turn a vague complaint about AI writing into four things you could go and count, and not one of them surfaces as a factual error.
Lexical decay narrows expressive range as precise or situated language gets replaced by statistically safe formulations. In practice that is “stakeholder” arriving where the draft said “the warehouse manager.” Drift across generations mutates meaning across recursive summarise-and-paraphrase cycles while the structure survives. Ground erosion collapses the tacit background that gave the explicit language its weight, so a sentence survives and the reason for writing it does not. Semantic noise drowns signal in fluent redundancy.
His demonstration is a good one. Run Martin Luther King’s “I Have a Dream” through ten rounds of summarisation. By round five, he reports, dream has shifted from prophecy to generic aspiration; by round ten the speech reads like a policy memo. Nothing false enters at any step, which is the whole point.
He is right that no standard metric measures meaning and wrong that the field is indifferent to it. The stronger version: every metric compares an output to a reference text, and fidelity is about the relationship between text and purpose, which the reference does not contain.
| Metric | What it measures | What Jacobs says it cannot see |
|---|---|---|
| Faithfulness | grounding in the source text | tonal or contextual loss |
| Adequacy | completeness of information | communicative intent |
| Semantic similarity | overlap between two texts | shifts in nuance and purpose |
| Accuracy | factual correctness | meaning degradation |
His conclusion: “Together, these metrics evaluate whether a response is correct. None determine whether it remains meaningful.”
He overstates by treating the field as indifferent to meaning. Summarisation research spent years moving past ROUGE precisely because overlap scores were known to be inadequate, and model-graded rubrics now routinely ask about faithfulness to intent. So researchers are chasing this already. Missing is a metric that survives the thing Jacobs actually points at.
Every metric on that list compares an output to a reference text. Semantic fidelity asks whether the output preserves the relationship between the text and its purpose — a thing absent from the reference, and often written down nowhere. Hence his fourth question, whether the recipient received the meaning the sender intended, has no automated analogue. Concretely, no metric can score that bug-report summary as bad, because the information about which symptom mattered lived in the reporter’s head.
Jacobs has a term for a benchmark failing to see past its own criteria. He calls it evaluation blindness — a measurement that can be “internally reliable and still measure too little.” Which is a better description of the situation than “nobody measures meaning.”
Figure · the accuracy camera
Ten rounds of recursive summarisation, plotted against three things a transformation can keep or lose. You are currently looking at them from where an accuracy metric stands.
Accuracy check
100%
Eleven rounds, one point. Every fact in the source is still present in round 10, so the check passes at every step.
Fidelitynot measured
None of it is built. No index, no curve, no dataset, no rubric, no number — one summarisation run described in two sentences of prose is the whole evidence base.
His proposals are sensible and specific. Frequency–specificity tests and anchor-correlation tests for lexical decay. Recursive summarisation chains scored for metaphor retention and tone preservation, aggregated into a Semantic Drift Index or a Fidelity Decay Curve. Retrieval precision measured before and after synthetic-text infusion for semantic noise. Any competent team could build these.
None of them are built. Nowhere in the repository is there an index, a curve, a dataset, a scoring rubric, or a single number. He describes the MLK run in two sentences of prose: which round the meaning shifted, and which round it read as a memo. No rubric, no rater, and no second text to compare against. Call that a demonstration, and a persuasive one. Nobody should call it a measurement.
Closing the gap would be unglamorous and small: one text, ten rounds, two raters, a published rubric, and the disagreement rate reported. In other words the MLK run again, done properly and written down. Until that exists, semantic fidelity is a well-specified hypothesis rather than a measurement framework, and citing it as the latter overstates it.
“LLMs do not hallucinate; they drift” is a true comparative claim pushed into a false absolute, and the reference list of the paper making it contains a citation to a volume of Nature that does not hold the paper.
Measuring Fidelity Decay opens its first section with: “LLMs do not hallucinate; they drift.” The supporting argument is that hallucination implies perceptual mis-seeing, and models do not see, they predict.
His etymological point is fine, and his conclusion does not follow from it. Fabrication is a real, distinct and well-documented failure: invented case citations that got lawyers sanctioned, APIs that never existed, plausible page numbers for real journals. None of that is erosion of meaning. Each is a confident assertion of something that was never true, and no amount of fidelity preservation prevents it.
Proof sits in his own reference list. Measuring Fidelity Decay cites the Shumailov model-collapse paper as Nature, volume 628, pages 555–560. Crossref puts that paper in Nature volume 631, pages 755–759. Both numbers are wrong, and wrong in the specific way a language model gets citations wrong — right journal, right year, plausible volume, plausible page range, all invented.
Jacobs argues that fabrication is the wrong thing to worry about, in a paper whose only borrowed evidence arrives through a fabricated citation.
I want to be careful about what this does and does not show. One error, touching none of the compression argument, and I have no way to know whether a model or a person produced it. Still, it shows the category Jacobs wants to retire doing real work. Catching this needs a fact check rather than a fidelity check, and nothing in his framework would have flagged it. Keep both gates from §01. His correct claim was always the comparative one: erosion is more common and less visible than fabrication, and gets a fraction of the attention.
The thirty-term lexicon is the most useful document in the repository, because every entry names the term it gets confused with. The distinction is what turns vocabulary into diagnosis.
Every entry carries a structural pattern and a distinction — the term it is most often confused with, and what separates them. Distinctions are what turn vocabulary into diagnosis, because naming a problem only helps if the name rules something out. Here is a sample, chosen for terms naming a failure you can walk into:
| Term | Distinguished from | What having the distinction lets you say |
|---|---|---|
| Retrieval drift | retrieval failure | Search returned something plausible and topically adjacent that quietly redirected the frame. Failure returns nothing; drift returns the wrong frame, confidently. |
| Interpretation failure | retrieval quality | The right context was assembled into the wrong meaning. “Presence is not preservation.” |
| Memory drift | context drift | Context drift happens inside one chain; memory drift persists across sessions, because each update reinterprets prior state rather than preserving it. |
| Instruction sedimentation | memory drift | Memory drift changes what is stored. Sedimentation is stored layers acquiring governing force — old prompts, inferred preferences and summaries hardening into constraints nobody chose. |
| Agent drift | task failure | Task failure is visible. Agent drift produces “successful-looking completion of the wrong objective.” |
| Procedural completion | proxy substitution | The workflow ran to the end and the underlying need is untouched. The procedure closes; the problem stays open. |
| Synthetic coherence | hallucination | Hallucination is one false statement. Synthetic coherence is a whole structure that feels valid because its parts fit each other rather than because they fit the world. |
| Corrective contact | feedback | Feedback is any returned signal. Corrective contact is a signal still connected to the underlying condition — the only kind that can interrupt a coherent wrong answer. |
Instruction sedimentation describes what a growing pile of always-loaded configuration does to a system. Past decisions acquire authority nobody granted them, and they stay invisible precisely because they are always present. Anyone running a global config file, a skills library and a memory store at once should sit with that sentence.
And procedural completion names the specific way a well-designed workflow lies to you — which means it reports done because the steps ran, not because the need was met.
Reel to input note to atomic to argument to weekly to monthly is six compression passes with no review step. Two of Jacobs’s preconditions for measuring fidelity are already met here; the gate itself is the thing missing.
Consider one chain. A Reel becomes an input note, that becomes an atomic, and the atomic feeds an argument note. Weekly reports summarise those, and monthly reports summarise the weeklies. Every hop is a compression pass performed by a model, which means the chain runs five hops deep by the time a claim reaches a monthly review.
Two of Jacobs’s preconditions for measuring fidelity are already met, which is the good news, and luck had nothing to do with it. Stan’s source/ directory keeps raw captures unfiltered, so an uncompressed original always exists to measure decay against. Jacobs calls the lack of one provenance failure, and git plus plain text closes it. Missing is the gate itself. Every hop gets read for whether the facts survived, and no hop gets read for whether the point did.
■ the risk concentrates here ● present and working ◖ partial · absent
| Hop in the chain | Raw keptcan re-derive | Provenancegit history | Accuracy readfacts checked | Fidelity readpoint checked |
|---|---|---|---|---|
| Reel → input note | ●present | ●present | ●present | ·absent |
| Input note → atomic | ●present | ●present | ●present | ·absent |
| Atomic → argument note | ●present | ●present | ◖partial | ·absent |
| Notes → weekly report | ◖partial | ●present | ◖partial | ·absent |
| Weekly → monthly review | ·absent | ●present | ·absent | ■the risk concentrates here |
Figure · the ungated chain
The same meaning space as the figure above — but nothing on this chain has ever been measured, so each hop is drawn as the volume the artefact could be anywhere inside, not as a point. Click any hop to insert the fidelity gate that is missing.
Three tests, none needing a tool: explain a result without reproducing its wording, walk three claims from the last monthly review back to their sources, and watch for the moment a name becomes a role.
One question is the best test in the whole corpus, borrowed from the cognition half of the project: can you explain and use the result without reproducing the system’s wording? Applied to a note, that means closing the file and saying what it argues in your own words. Any note that fails was copied into the vault rather than understood, and it will read as knowledge on every future retrieval.
Test two costs about ten minutes a month and targets the bottom row of the matrix. Take three claims from the most recent monthly review and walk each one back to a source note. Then ask whether that source note still means what the monthly says it means. Three claims need no rubric, which makes this a fidelity read of the highest-risk hop for almost nothing.
Test three is structural, and it follows from lexical decay: when a summarising step replaces a specific noun with a category, that is the signal. In practice, watch for the moment a name becomes a role — a person becomes “the stakeholder”, a tool becomes “the platform”. Jacobs is right that this is measurable, and you can notice it without measuring anything.
Jacobs’s strongest claim inverts normal practice: the task delegated without review is the one doing the damage. Fabrication announces itself; a hollowed summary gets forwarded.
Thirty terms, each with the distinction that separates it from its neighbour. Instruction sedimentation, procedural completion, agent drift and corrective contact name things worth having names for.
Jacobs proposes a Semantic Drift Index and a Fidelity Decay Curve, and builds neither. One text, ten rounds, two raters and a published rubric would change that; until it exists, describing the corpus as offering benchmarks overstates it.
“LLMs do not hallucinate; they drift” is the one line to reject, and the corpus refutes itself. Its citation to Nature gives a plausible, wrong volume and a plausible, wrong page range. Only a fact check catches that.
Raw capture and git provenance are already in place, so the baseline for measuring decay exists. Three claims walked back to source, once a month, covers the hop where the chain is deepest and the raw is furthest away.
Any framework about whether meaning survives transmission invites having its own references checked. I checked every external citation in Measuring Fidelity Decay carrying a resolvable identifier, on 31 August 2026.
| Citation as given | What resolving it returns | Status |
|---|---|---|
| Shumailov et al. (2024), model collapse, Nature 628, 555–560 | Real paper, wrong location. It is Nature 631, 755–759, DOI 10.1038/s41586-024-07566-y. There is also a 2025 author correction at Nature 640, E6. | wrong |
| Meta AI (2024), “Know when to stop: a study of semantic drift in text generation”, NAACL | Confirmed via Crossref: NAACL-HLT 2024, Volume 1 Long Papers. The one citation that directly supports the drift claim, and it holds. | verified |
| Arora et al. (2024), F-Fidelity, arXiv:2410.02970 | Identifier and title both correct — but the paper is F-Fidelity: A Robust Framework for Faithfulness Evaluation of Explainable AI. It evaluates XAI explanation methods, not meaning preservation in generated text. Right paper, adjacent problem. | off-target |
| Jacobs (2025), “The Meaning Equation”, figshare 10.6084/m9.figshare.30128110 | Resolves to the named preprint on figshare. Self-published, self-cited, and honest about being a preprint. | verified |
| Lakoff & Johnson (1980); Sperber & Wilson (1986) | Standard works, correctly named. Not checked to page level, and neither is load-bearing for any specific claim. | not checked |
One wrong citation in five hardly counts as a scandal, and I would not lead a page with it. Keeping it here has one reason. §06 argues the corpus is wrong to retire the fabrication category, and here is the corpus demonstrating why on its own reference list. Everything else in the audit holds, and the central argument depends on none of it.
Scope, stated plainly: this page judges the semantic-fidelity-project repository as of its 13 August 2026 state — 23 PDFs, all downloaded and read. I am not judging the Substack essays, the main Reality Drift library beyond what §07 borrows, or the author.