03
Preprints
All experimental work is released as open preprints with code, data-generation scripts, and evaluation harnesses included; figures are regenerated from source result files. The three papers below are the current record. Earlier consolidations they subsume are listed after them rather than removed, so the versioning is visible.
Preprint v1.0 · Aug 2026
Source Relations Under Recursive Transformation: Provenance Path Dependence and External Re-grounding in Causal Language Models
Factual preservation and provenance preservation are not the same problem.
A mechanism ladder narrows a structural source-boundary effect in Pythia-410M to source-token prediction in a first-occurrence regime, then tests recursive provenance in a frozen Qwen2.5-7B-Instruct protocol: 32 records, five conditions, ten recursive passes, 1,600 generations. Initial attribution state strongly constrains later attribution, and re-grounding against an immutable provenance ledger outperforms recursive self-relay by 15.1 points.
Christopher W. Sweeney · 11 August 2026 · Pythia-410M, Qwen2.5-7B-Instruct · Preregistered recursive provenance assay
DOI 10.5281/zenodo.21890267
PDF
Preprint v1 · Aug 2026
The Anchor Protects What It Names
Compact context anchors protect the source content they enumerate and not the content they omit — and refreshing them from working state removes the protection.
Forty technical passages carrying 708 curated terms, rewritten recursively ten times under five conditions, replicated across three independent GPU runs and two architectures. Naming a term raises its ten-generation survival by +0.143. A randomised, yoked assignment settles the direction of causation. A self-refreshed anchor is statistically equivalent to having no anchor at all on the very content it began by naming — a checklist that forgets its items.
Christopher W. Sweeney · August 2026 · Qwen2.5-7B-Instruct (×2), Mistral-7B-Instruct-v0.3 · 6,000 generations, 6,600 scored texts
DOI 10.5281/zenodo.21855188
PDF
Consolidated preprint v1 · Aug 2026
Three Jobs of a Source Tag
Semantic conditioning, attribution integrity and stability under recursive self-training depend on different properties of the tag — and only one of the three requires the source to be true.
Eleven experiments across 152 fine-tuned models separate three outcomes that are routinely conflated. The semantic job requires the tag to be about the text; substituting an author for a title collapses the conditioning benefit to the opaque-identifier floor. The structural job requires only the occupancy of a position. The attribution job is the only one that requires truth — and getting it wrong costs roughly twice what getting it right buys.
Christopher W. Sweeney · 6 August 2026 · GPT-2 124M/355M/774M, Qwen2.5-0.5B · Wikipedia and post-cutoff arXiv · Supersedes Zenodo 21782935, 21796258, 21810517
PDF
Superseded — retained for the record
Both papers below are subsumed by Three Jobs of a Source Tag, which reports the same experiments at greater depth and corrects two claims made here. They remain listed because the corrections are part of the record.
Earlier consolidation · Aug 2026
The Positional Function of Source Attribution in Language Models
Correct attribution lowers loss by 0.22 nats and slows recursive drift by a quarter — but a meaningless placeholder occupying the same position captures two-thirds of that protection.
Eight experiments across approximately seventy fine-tuned models. The conditioning effect is not memorisation: it is undiminished when all test titles are unseen during fine-tuning, it strengthens with model scale, it transfers across architectures, and it is largest on a post-cutoff corpus where pretraining familiarity is impossible. The decisive comparison is the one that fails — what the tag contributes structurally is the occupancy of a position, not the transmission of a source.
Christopher W. Sweeney · 6 August 2026 · ~70 fine-tuned models, a subset of the 152 · Superseded by Three Jobs of a Source Tag, which replicates the suffix arm and narrows this paper's trailing-tag claim
PDF
Earlier preprint · Aug 2026
Attribution Slows Model Collapse Whether or Not It Is True
Structure contributes 87% of the protection; accuracy protects only the attribution chain, and that effect compounds with recursion depth.
Every published mitigation for recursive collapse works by importing information from outside the loop. This paper tests an intervention that imports nothing: a source tag inside the model's own generated text. Thirty GPT-2 models trained across five recursive generations in three lineages. Meanwhile attribution does not collapse under recursion — it inflates: a corpus of 2,073 real sources becomes 6,700–14,044 distinct generated titles, 28.2% of generated sentences carrying a fabricated source by generation five.
Christopher W. Sweeney · 5 August 2026 · 30 GPT-2 fine-tunes · 3 conditions × 5 generations × 2 seeds · This experiment is reproduced and extended within Three Jobs of a Source Tag, which supersedes both companion records
Companion DOI 10.5281/zenodo.21782935
Companion DOI 10.5281/zenodo.21796258
PDF