Content Quality: Well-structured Analysis piece (Overview / What We Know / What We Don't Know / Analysis). Word count is 1,172, within the Analysis category range (800-2000). Neutral, non-sensational tone throughout; no AI self-reference.
Source Verification: Both cited sources are the arXiv abstract page (source-0.html.gz, sha256 verified 4eca130a...) and the arXiv full-text HTML page (source-1.html.gz, sha256 verified fc3a40e2...) for the same preprint, arXiv:2608.16630, 'The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks' by Bardia Mohammadi, Lars Klein, Aman Chadha, Akhil Arora, and Laurent Bindschaedler, submitted 17 Aug 2026. Both snapshots read in full (decompressed and text-extracted, not re-fetched). Title, author list, and submission date all confirmed verbatim against source-0. All 14 direct quotes in the article body were checked against the actual snapshot text: every quote appears verbatim somewhere across the two snapshots (confirmed via exact substring search after HTML-to-text extraction and whitespace normalization). Two quotes ('putting the same facts in the prompt lifts 299 of 300 matched trials...' and 'the tokens consumed over a run differ by 12.8x...') are hyperlinked in the article to the arXiv abstract page (source-0) but the exact sentence is only present verbatim in the full-text HTML page (source-1) -- the abstract itself only paraphrases these findings ('putting the facts in the prompt restores success' / 'differ more than tenfold in tokens consumed'). This is a same-paper cross-linking imprecision, not a factual error or wrong-outlet citation: both URLs are already in article.sources, the quoted content is 100% accurate and independently verified in the full-text snapshot, and a reader following either source link lands on the same paper. Logged as a minor concern below rather than a corrections-worthy issue since no fact is wrong or unsourced.
Factual Accuracy: All specific numbers checked against source-1 full text and confirmed accurate: 154 closed-book trials scoring 0/12 with a 2.4% Wilson upper bound; 299/300 matched trials reaching >=9/12 requirements after front-loading; 66/70 trials across seven models converging on 24/79 tests on the renamed-Pydantic migration; the 128,000-character context-distance test; 12.8x token-spend variance vs 1.8x peak-context variance; 293,882 to 3,752,134 token range with five vs seventy-nine tool calls; 0.46% reasoning-token share; the block-diagnostics table (Opus 100%, Fable 75%, Sonnet 25%, Haiku 12.5%, GPT-5/Codex CLI 0%, GLM-5.2/opencode 0%, pooled 37% of 46 trials); the 39-trial stale-standard-vs-code experiment; the SWE-bench slice (100 instances, 8 repositories, 400 attempted/397 scored, resolved rates 1.0% Sonnet+Aider to 39.4% GPT-5+Codex CLI); and the AUC ~=0.49 residency-metric result after correcting an earlier leaky ~0.92 version. The 'coherence debt' definition and the article's central quotes about fabrication-over-refusal both trace verbatim to the abstract. All five harnesses named in the article (Claude Code, Codex CLI, Aider, OpenHands, opencode) and all seven models/model families referenced are independently confirmed present in the full-text snapshot. No fabricated or hallucinated specifics found.
Overall Assessment: High-quality, well-sourced Analysis piece. Every specific number and every direct quote was independently verified against the decompressed source snapshots and confirmed accurate; the five harnesses and seven models are all corroborated in the full paper; the preprint/non-peer-reviewed status and small-sample caveats are surfaced honestly rather than glossed over. The only issue found -- two quotes linked to the abstract instead of the full-text page where they verbatim appear -- is a minor citation-precision nitpick, not a factual problem, and does not warrant a corrections record. Approved as submitted.