Content Quality: Well-structured Analysis piece (Overview / What We Know / What We Don't Know / Analysis), 1,329 words — within the Analysis category's 800-2000 word range. Technical detail (the two-phase Spec-Driven Test Generation framework, Contract Coverage metric, Greenfield Test Generation eval protocol) is explained clearly for a software-engineering-literate audience without oversimplifying the statistics.
Source Verification: Read both source snapshots from sources/2026-08/google-researchers-show-spec-driven-test-generation-cuts-missed-bugs-in-ai-coding-agents/ after verifying each snapshot's sha256 against manifest.json (both matched exactly: source-0.html.gz 4aab41b0...511b53 [abstract page], source-1.html.gz ad45a51f...2971cc [full HTML paper] — no tampering). source-0.html.gz (arxiv.org/abs/2608.17177) confirms: title 'Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation', all 10 authors (Michele Tufano, James McClure, José Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivančić, Livio Dalloro, Pat Rondon), submitted 17 Aug 2026, and the abstract's headline stats (9.8pp bug-detection improvement p=0.0352, 2.5pp branch-coverage improvement p=0.0034, 77.8%/56.7% LLM-judge superiority rates) verbatim. source-1.html.gz (arxiv.org/html/2608.17177v1, the full paper) independently confirms every specific and quote used in the article: all 10 authors listed with @google.com emails and 'Affiliation: Google, USA'; the 90-historical-bug-fixes dataset description ('90 historical bug-fixes (pairs of buggy and fixed code) from Google's Internal Issue Tracking System') verbatim; the C++/Java/Python/Go language span verbatim; the 'entirely blind to the bug, the commit message, and the fix diff' quote verbatim; the detection-rate table (baseline 53.4% [43.3%,63.3%] vs spec-driven 63.2% [53.3%,73.3%], +9.8%, McNemar's p=0.0352 at k=5) verbatim; the k=1 gap (36.9% vs 41.1%, +4.2%, not significant) verbatim, matching the article's characterization; branch coverage (46.4% vs 48.9%, +2.5%, Wilcoxon p=0.0034) verbatim; the 45-shared/12-unique-spec-driven/3-unique-baseline bug breakdown verbatim; the LLM-as-a-Judge results (77.8% vs baseline, 56.7% vs human-written, 'showing parity with human-level engineering rigor') verbatim; the Contract Coverage metric (61.1% at k=1, 78.9% at k=5, and the 54.9%-vs-19.4% predictive split at p=3.62×10⁻¹⁴) verbatim; the token-overhead figures (243.9M baseline vs 336.7M spec-driven tokens, 38.0% overhead) matching the paper's summary Table 4 (note: one prose sentence elsewhere in the paper says '336.6M' rather than '336.7M' — the article's figure matches the table, which is the paper's own canonical number, so this is the paper's internal rounding inconsistency, not an article error); the 'no explicit contracts to anchor their exploration... miss edge cases, hallucinate logic, and generate superficial tests' quote verbatim; the 'retroactively extracting and articulating a code component's underlying contract before authoring any test code' quote verbatim; the optional Human-in-the-Loop Specification Curation step (§3.2: developer can 'refine the inferred Pre_c and Post_c, as well as accept/reject/amend... before any code is generated') and the explicit confirmation that reported results bypass HITL ('we bypass the HITL phase and accept all agent-generated test suggestions') both verified — matching the article's claim that reported results reflect the fully autonomous version; and the Hoare logic / Eiffel / Design-by-Contract 'co-locating their logical assertions with their implementation and automatically checking these at runtime' quote verbatim in the Related Work section. No hallucinated quotes, no fabricated statistics, no misattribution found.
Factual Accuracy: Every number, name, date, and direct quote in the article traces to the decompressed full-text paper, independently confirmed, not inferred from the manifest or abstract alone. The internal link to the prior 'coherence debt' article (/article/2026-08/18-researchers-coin-coherence-debt-to-explain-why-coding-agents-fabricate-fixes-instead-of-asking-when-facts-go-missing) resolves to a real, existing article file using the correct singular /article/ path form.
Overall Assessment: Clean, rigorously sourced Analysis piece on a preprint (2 sources, both arxiv.org, both the primary paper). Every statistic, quote, and attribution independently verified verbatim against the full-text HTML snapshot, not just the abstract. No fabrication, no misattribution, no orphan sources, no duplicate coverage of the two specifically-checked related articles. The only automated finding is a false-positive stylistic lint on an accurately quoted source word, documented and dismissed above. APPROVE.