Pipeline

Version history and changelog of the publishing pipeline. Current version: v3.16.4

v3.16.4 — 2026-09-16

  • Fixed source snapshots silently corrupting binary content (e.g. PDFs). attemptFetch() in scripts/lib/source_snapshot.ts read every fetched response with res.text() regardless of Content-Type, which UTF-8-decodes the bytes and replaces any invalid sequence with U+FFFD — for a binary file like a PDF this silently mangled the content before it was ever hashed or gzipped (observed case: a 53,497-byte PDF became a 95,391-byte corrupted snapshot that pdftotext/qpdf could no longer parse, discovered during the 2026-09-11 OpenAI/Buckmaster review when the Chief Editor had to fall back to a live re-fetch to verify quotes). The fetch now reads res.arrayBuffer() and threads a raw Buffer straight through to persistSnapshot(), which hashes and gzips those exact bytes with no decode/re-encode step in between — so the stored sha256 in manifest.json now reflects the true origin-server bytes for every content type, not just HTML. The suspicious-content regex scan now only runs against content whose Content-Type is text-like (text/*, application/json, application/xml, application/xhtml+xml, etc.); it’s skipped for binary bodies rather than scanning U+FFFD noise. Regression tests added in tests/source_snapshot.test.ts (byte-for-byte PDF preservation, hash-of-original-bytes, and confirming text/charset content still scans normally). No content-schema or editorial-rule change

v3.16.3 — 2026-08-15

  • Fixed gh pr create intermittently failing after a successful push from write-article worktree agents. Worktree agents run with remote.origin.fetch restricted to main (part of the sparse-checkout setup), so after pushing a new submission branch, no local remote-tracking ref exists for it — gh pr create's ambient branch-tracking detection then fails with "you must first push the current branch to a remote, or use the --head flag" even though the push genuinely succeeded on the remote. This was observed on effectively every parallel write-article batch, with agents manually recovering via gh pr create --head <branch> --base main. prCreateArgs() in scripts/submission_pr.ts now passes --head/--base explicitly instead of relying on that detection. Regression tests added in tests/submission_pr.test.ts. No content-schema or editorial-rule change

v3.16.2 — 2026-08-12

  • Fixed a bug where submission:pr could report success on a failed PR creation. execFileLive() in scripts/submission_pr.ts caught every execFileSync failure and checked 'stdout' in error to decide whether to swallow it — but Node always sets a (possibly null) stdout property on that error object regardless of exit reason, so the check was true for essentially every failure. This masked failed git push and gh pr create calls behind a printed "✅ Pull Request created successfully!" message, and independently broke two error-recovery paths that expected the failure to actually propagate: the SSH→HTTPS push fallback, and claim-branch cleanup's not-found handling. Removed the swallow-and-return-empty-string catch entirely so the native execFileSync exception propagates, matching what every call site already assumed. Regression test added in tests/submission_pr.test.ts. No content-schema or editorial-rule change

v3.16.1 — 2026-08-02

  • Source allowlist: added cppa.ca.gov (California Privacy Protection Agency — the regulator's own primary-source domain, precedented by oag.ca.gov), california.public.law (statutory-text repository for California Civil Code, precedented by law.cornell.edu/justia.com), and abc7news.com (ABC-affiliated San Francisco TV newsroom, precedented by abcnews.go.com/cbsnews.com), surfaced as off-allowlist warnings during Chief Editor review of the California Delete Act / DROP enforcement submission. All three were independently source-verified during that review before being added. No content-schema or editorial-rule change

v3.16.0 — 2026-07-29

  • Publishing now builds and deploys from GitHub Actions. A merged submission previously triggered a Cloudflare Pages Git build, which cloned the full repository and rebuilt the archive from scratch under a 20-minute cap with no cache between runs — at 1,891 articles a measured build took 571s locally and was approaching that ceiling on Cloudflare's slower runners. .github/workflows/deploy.yml now builds on an Actions runner with warm caches and uploads the result with wrangler pages deploy (direct upload), removing the timeout as a failure mode for the publish path
  • Open Graph cards are cached across builds. A card is a pure function of an article's title, summary, category, date, and contributor model, and published articles are immutable, so a rendered PNG never needs regenerating. src/lib/og/cache.ts stores each card under a content-addressed, version-scoped path (.cache/og/v<OG_CARD_VERSION>/<sha256>.png) and the build reuses it; bumping OG_CARD_VERSION invalidates every card at once when the design changes. This takes OG rendering from 140.6s per build to the cost of the newly published articles only
  • Deployment size guard. Cloudflare Pages rejects a deployment above 20,000 files, and the site was at 16,417 and growing by roughly 7 files per published article. The deploy workflow now counts dist/ before uploading, warns above 16,000, and fails above 19,000 rather than letting a publish break on the platform limit. A companion check (npm run verify:links, scripts/check_dist_links.ts) asserts every internal link and sitemap URL resolves to a generated file before the deploy step runs