Content Quality: Well-structured News piece (626 words, within the 400-1200 word News range) using the standard Overview / What We Know / What We Don't Know / Analysis format. Claims are consistently attributed inline to the specific cited source (Hugging Face card, SCMP, or GitHub README) rather than left unattributed. Bulleted 'What We Know' section is specific and each bullet traces to one named source.
Source Verification: All 3 sources were read from the committed gzip snapshots in sources/2026-09/tencents-hy4-preview-open-weight-model-beats-qwen38-max-and-deepseek-v4-pro-on-deepswe-benchmark/ (not re-fetched). Decompressed content was rehashed and matched the manifest sha256 for all three files (source-0: ac3566bb..., source-1: 6098f3aa..., source-2: 8611f9e7...). No suspicious_patterns entries in the manifest for any of the 3 sources. (1) source-0.html.gz (SCMP, paywalled article, full text recovered from the page's embedded JSON-LD articleBody): confirms the DeepSWE benchmark sentence verbatim - 'On the DeepSWE benchmark, which tests software engineering capabilities, the Hy4 preview scored 64.3, surpassing Alibaba's Qwen-3.8 Max at 56.6 and DeepSeek-V4 Pro at 62.7' - appearing in the article's own subheading/deck AND body text, not an image. This directly resolves the flagged risk that the benchmark numbers might trace only to an image-only table: the final submission attributes this exact comparison sentence solely to SCMP, whose text (not an image) states it. Also confirms verbatim the Goldman Sachs quote attributed to Ronald Keung, the leaderboard placement (8th, up from 34th, behind Claude Fable 5, ahead of Qwen 3.8-Flash-Next), and 'full commercial iteration...launch later this year' (article renders this as 'later in 2026', an accurate paraphrase since the piece is dated Sept 2026). (2) source-1.html.gz (Hugging Face model card): confirms 770B total / 49B activated parameters, 78 layers, 256 routed + 1 shared expert, top-8 routed activation, 1M context length, 120,832 vocabulary size, Gated DSA with IndexCache cross-layer reuse, iHC (identity Hyper-Connections) residual pathway, Apache 2.0 license, availability on Hugging Face/ModelScope/GitCode/CNB, the 163-expert/203-task blind eval numbers (GLM 5.3 2.99 vs 2.92, 46.8% win rate; Kimi K3 2.99 vs 2.94, 51.2% win rate) verbatim, the CodeBuddy/WorkBuddy co-design claim, and the closing quote 'As with Hy3 preview, we would rather ship early and hear what breaks...' verbatim. Note: the HF card's own 'Benchmark' section is an image-only table (extracts to empty text between the 'Benchmark' and 'Appendix' headings) - the submission correctly does NOT cite Hugging Face for the comparative DeepSWE numbers, attributing that sentence to SCMP's text instead, so no claim in the final article rests on the unreadable image table. Hugging Face's own structured Evaluation Results widget separately lists Hy4's DeepSWE score as 64.3, independently corroborating the SCMP figure for Hy4 itself. (3) source-2.html.gz (GitHub README): mirrors the Hugging Face card content and independently confirms the Apache 2.0 license and architecture specifics cited from 'its GitHub repository.' Both internal cross-links (/article/2026-08/07-alibaba-unveils-qwen38-max... and /article/2026-04/25-deepseek-releases-v4...) resolve to existing published articles.
Factual Accuracy: Nearly every claim checked out verbatim against source text, including all quotes, the architecture specs, the benchmark comparison, and the internal-eval numbers. One factual error found: the 'What We Don't Know' paragraph states 'Neither Hugging Face nor SCMP's coverage discloses Hy3's own parameter count or context window, so the scale of the jump in raw model size between generations is not confirmed by any source used here.' This is incorrect - the cited SCMP article explicitly states 'The Hy4 preview marks a significant leap in scale, featuring 770 billion parameters - a 2.6-fold increase over Hy3's 295 billion parameters - and a context window that quadruples to 1 million tokens, up from 256,000.' SCMP is cited elsewhere in this same article for other facts, so the omission is not a matter of an unavailable source - the information was available in a source the bot had already fetched and read. This is a subordinate claim (not in the headline, summary, or lead) and is isolated to one paragraph; every other claim in the article is independently verified against source text, so this does not reflect a pattern of fabrication.
Overall Assessment: APPROVE_WITH_CORRECTIONS. All specifics used to support the headline, summary, and lead are independently verified against the actual source text (not images), including the previously flagged risk around the DeepSWE benchmark table - the final article correctly sources that comparison to SCMP's readable text rather than Hugging Face's image-only benchmark table. The single issue found is a subordinate claim in the 'What We Don't Know' section that inaccurately says information is unavailable when it is in fact present in a source cited elsewhere in the same piece; this is honestly and fully coverable in a single corrections note and does not undermine the article's central thesis or lead.