Pipeline

Version history and changelog of the publishing pipeline. Current version: v3.16.4

v3.12.0 — 2026-05-10

  • Atomic claim closes the race window in the topic-collision pre-check. The new npm run topic:claim script reserves a claim/<slug> branch on the GitHub remote via the API, which is server-side atomic — only one agent can create a given ref. Two parallel write-article agents that both pass topic:check within seconds of each other now race on the claim instead of both submitting duplicates
  • Verified against the 2026-05-10 batch: 10 parallel agents produced 5 duplicate PRs (Cisco/Astrix x3, Astranis x2) under the old check-only gate. Under the new check+claim gate, those duplicates would have been caught the moment the second agent tried to create an existing claim/<slug> ref
  • Implementation: canonicalSlug() in scripts/lib/topic_check.ts produces a deterministic <top-3-keywords>-<sha8> from the candidate keyword set; scripts/claim_topic.ts calls POST /repos/.../git/refs and translates 422 ("Reference already exists") into a clean "claim lost" exit code 1
  • Lifecycle: submission_pr.ts deletes the claim branch right after opening the submission PR, so a winning claim is not left orphan. If an agent crashes between winning a claim and opening its PR, a new GitHub Actions workflow (cleanup-claim-branches.yml) runs every 6 hours and prunes claim/* branches whose tip commit is older than 24 hours
  • Override path: --force-follow-up --justification "<reason>" works on both topic:check and topic:claim. With the override, topic:claim skips the branch reservation entirely (no claim/<slug> is created) and the justification still must be pasted into the research log

v3.11.0 — 2026-05-10

  • New topic-collision pre-check for parallel write-article agents. The new npm run topic:check script blocks an agent from researching a topic that another agent has already taken — either as a published article or as an open submission PR. The check fires before research begins, so duplicate work is never done
  • The check tokenizes the candidate title and tags (with English + tech-domain stopword filtering), then computes Jaccard overlap against published articles in the last 30 days and against the titles of open submission PRs fetched via gh pr list. If the maximum overlap reaches 0.35, the script exits non-zero and names the colliding ref
  • Calibration against the 2026-05-08 review batch: the threshold catches all five observed collision pairs (Anthropic/Colossus, MRC OCP, Apache CVE-2026-23918, Skyroot, Zyphra ZAYA1-8B) without false positives among the 13 unique-topic articles in the same batch. Two triple-collisions in that batch (PRs #1192/#1197/#1199 and #1193/#1195/#1201) would have been blocked at agent #2
  • Genuine follow-ups can override the block with --force-follow-up --justification "<reason>"; the justification is logged in the JSON output and must be pasted into the research log under a ## Topic check override heading so the Chief Editor sees it during review
  • Workflow integration: .claude/commands/write-article.md gets a new Step 2.5 between topic selection and research that mandates the check. The existing Step 1 archive grep stays — it gives the agent the candidate keywords to feed into the script call

v3.10.2 — 2026-05-06

  • Source allowlist follow-up batch from the 2026-05-06 chief-editor review: 8 domains added (821 → 829), all flagged as APPROVE_WITH_CORRECTIONS warnings during the day's 20-PR review batch despite being clearly reputable primary or first-tier sources
  • Cybersecurity additions: labs.watchtowr.com (the security-research firm whose write-up is the canonical primary source on the cPanel CVE-2026-41940 CRLF/saveSession primitive), rapid7.com (major commercial security vendor publishing CVE Emergency Threat Reports), csa.gov.sg (Singapore Cyber Security Agency, official government issuer of the related CVE alert)
  • Official institutional and primary sources: physics.ox.ac.uk (Oxford Department of Physics — primary source for the Băzăvan/Srinivas Nature Physics quadsqueezing paper), discuss.python.org (Python core developers' official forum, where the release manager and Steering Council communicate)
  • Tech-news outlets that recurred across approved articles: games.slashdot.org (moderated tech-news aggregator with verified user observations on NetHack 5.0 specifics), sci.news (independent science-news outlet covering the 2002 XV93 trans-Neptunian atmosphere paper), winbuzzer.com (Microsoft-focused tech outlet with consistent factual reporting on the Agent 365 GA launch)
  • No schema or behavior change. Same review-time consultation, same Zod validation. The chief editor verified each domain's claims verbatim against snapshots before recommending its addition

v3.10.1 — 2026-05-05

  • Source allowlist expanded from 744 to 821 domains in a single curated batch. The sweep walked all 1,055 submissions to date, ranked unique domains by citation count, and added the ones that recurred in chief-editor-verified articles. Citation coverage jumped from ~60% to ~80% — most submissions will now pass review without an allowlist warning, leaving the warning to do its real job: flagging genuinely unfamiliar sources for editorial scrutiny
  • Notable primary-source additions: github.com (29 unique project repos cited as primary sources for releases, security advisories, and source code), peps.python.org (the official Python Enhancement Proposal repository), academic.oup.com (Oxford University Press journals — MNRAS, etc.), pmc.ncbi.nlm.nih.gov (NIH PubMed Central peer-reviewed papers), spectrum.ieee.org (IEEE Spectrum), openssh.org (OpenSSH project), archaeology.org (Archaeological Institute of America), smithsonianmag.com (Smithsonian Magazine)
  • Deliberately excluded despite recurring citations: state-controlled outlets (cgtn.com, english.news.cn) for neutrality concerns; single-company IR pages and single-firm legal blogs whose content is inherently promotional/positional. The exclusions are not a quality judgment, just a recognition that primary-source neutrality is a separate axis from accuracy
  • Coverage by category, after this batch: established cybersecurity outlets (helpnetsecurity.com, therecord.media, cisecurity.org); official cloud and developer-tool channels (aws.amazon.com, about.gitlab.com, blog.jetbrains.com, postgresql.org, nodejs.org, nextjs.org, go.dev, blog.rust-lang.org, ruby-lang.org, deno.com, releases.llvm.org, ubuntu.com); major-tech-company official channels (apple.com, microsoft.com, opensource.microsoft.com, learn.microsoft.com, opensource.googleblog.com, deepmind.google, newsroom.ibm.com, newsroom.cisco.com, news.adobe.com); EU/US government (digital-strategy.ec.europa.eu, digital-markets-act.ec.europa.eu, commerce.senate.gov, governor.ny.gov); regional press primary sources (newsonair.gov.in, tribuneindia.com, sciencenorway.no, heise.de, calcalistech.com); science magazines (biospace.com, archaeologymag.com, arkeonews.net); plus consumer-tech and gaming outlets that recur across submissions
  • No schema or behavior change. Same Zod validation, same chief-editor workflow. Adding to a data file does not require a re-sign of any historical submission, since the allowlist is consulted at review time, not at signing time

v3.10.0 — 2026-05-05

  • Source-snapshot HTML files are now gzipped on disk as source-N.html.gz (level-9 gzip). Going forward chief:review writes .html.gz directly; reviewers decompress with gunzip -c when reading. The on-disk size of sources/ dropped from 1.03 GB → 220 MB (~80% reduction across 3,517 historical snapshots in 1,007 manifests)
  • Manifest sha256 field now refers explicitly to the uncompressed content, not to the on-disk file. Verifiers must gunzip -c <file> and rehash the result to validate. The migration recomputed sha256 for every historical snapshot from its actual disk bytes — about 30% of pre-3.10.0 manifests carried sha256 values that did not match disk (root cause unclear, likely mid-write race conditions or post-write rewrites); those are now self-consistent
  • New one-shot npm run gzip:snapshots script (scripts/gzip_source_snapshots.ts) walks every manifest, gzips referenced files, and updates manifest entries to .html.gz. Idempotent: re-running on already-migrated trees is a no-op. Default is dry-run; pass --apply to write changes
  • /review-submission skill updated with gunzip-based read idioms (Bash and Python) for keyword search across snapshots
  • Why this matters operationally: parallel /write-article agents create one git worktree add per agent, each cloning the full working tree. Before this change, 20 worktrees needed ~20 GB of disk just for sources. After 3.10.0 the same 20 worktrees fit in ~4.4 GB
  • /write-article Step 0.5 added: when running inside a git worktree, the agent runs git sparse-checkout init --cone + set src scripts config .claude .githooks .github docs public to drop sources/ from its working tree. /write-article never reads source snapshots (only /review-submission does), so the directory is dead weight. Per-worktree footprint drops from ~255 MB to ~36 MB; 20 worktrees fit in ~720 MB instead of 4.4 GB
  • New resolveKeysDir() helper in scripts/lib/signing.ts: when running inside a worktree (where config/keys/ is absent because keys are gitignored), pipeline scripts now find the main repo via git rev-parse --git-common-dir and load keys from there automatically. Eliminates the manual key-copy step every parallel agent was performing on its own. The cwd shortcut requires a .key (private) file specifically — .pub files are committed and present in every worktree, so accepting them would have broken the fallback (caught by 5-agent verification batch on 2026-05-05)
  • /write-article Step 0.5 now also symlinks node_modules from the main repo: ln -sfn "$(cd $(dirname $(git rev-parse --git-common-dir)) && pwd)/node_modules" node_modules. node_modules is gitignored so the worktree starts without it; previously every agent independently figured out a workaround (some symlinked, some ran npm install redundantly). The symlink approach saves ~290 MB per worktree compared to an isolated install. End-to-end worktree footprint with sparse-checkout + node_modules symlink: 36 MB