AI Crawlers Now Devour 20% of git.kernel.org's CPU, Kernel.org Administrator Reports
Kernel.org administrator Konstantin Ryabitsev says AI scrapers generate 98% of git.kernel.org's 6 million daily requests, burning more CPU than all legitimate traffic combined.
Overview
AI scrapers are consuming more CPU capacity on git.kernel.org, the Linux kernel’s git hosting infrastructure, than all legitimate human and tooling traffic combined, according to Konstantin Ryabitsev, the site’s administrator. In a blog post published August 29, titled “Creepy crawlies,” Ryabitsev laid out the first hard numbers behind a problem he says he has informally complained about for some time, writing that “we spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones.”
What We Know
- Git.kernel.org receives about 6 million daily requests “demanding to see random commits,” according to Ryabitsev’s post, a figure also quoted by LWN. Of those requests, 66% are “immediately batted away” by Anubis, a proof-of-work challenge system, while 33% now solve the challenge and reach the main site.
- With “a bunch of generous assumptions,” Ryabitsev estimates legitimate requests make up only about 2% of git.kernel.org’s traffic, a figure corroborated by LWN’s write-up of the same post.
- Across five geographically distributed nodes totaling 90 CPU cores, 14 to 16 cores are constantly occupied rendering git commits as HTML for scrapers — about 20% of the site’s total capacity on average, according to Ryabitsev. He notes the real load is spikier than a flat 20%, arriving in waves.
- Linux’s main repository, linux.git, holds about 1.48 million commits, and git.kernel.org hosts roughly 922 forks of it. Rather than using
git cloneto pull that history efficiently, Ryabitsev writes that scrapers instead request individually rendered HTML pages for commits, diffs, and patches one at a time — a method that multiplies the same underlying commit history into what he calls “several BILLION valid URLs” across all the forks. - Ryabitsev argues kernel.org’s appeal to AI crawlers comes from the purity of its data: because Linux kernel history predates the current wave of AI-generated content, it is, in his words, a source that is “easy to filter in order to guarantee pure unadulterated pre-AI content” for training large language models.
- Kernel.org’s defenses have escalated over time. Administrators first blocked scrapers identified by user-agent string, then moved to banning individual IP addresses and, eventually, entire autonomous system numbers, as bots began spoofing browser identities. That approach broke down once scraping shifted to “millions of random residential or mobile IPs, all pretending to be random modern browsers,” Ryabitsev writes, since each address might make only four or five requests before disappearing.
- About a year before the post, kernel.org deployed Anubis, a system that forces incoming requests to solve a computational proof-of-work puzzle before reaching the site. It was “immediately extremely effective,” per Ryabitsev, but bots adapted within months to solve the initial difficulty level; administrators raised the difficulty, and the bots eventually adapted to that level too.
- Despite the load, Ryabitsev says git.kernel.org is not currently overwhelmed and “will likely be snappy and responsive” for visitors. He notes that outages, when they happen, are usually caused by misconfigured continuous-integration systems performing repeated shallow clones of the stable kernel repository rather than by scraper traffic itself.
- Going forward, Ryabitsev says the project is “turning off features to reduce the number of crawlable URLs and to gate off actions that are expensive for us to run,” warning that some functionality will be lost, particularly for anonymous, unauthenticated access. He adds that the project will keep offering its full data for download to anyone who requests it, though with more steps required to get it.
What We Don’t Know
Ryabitsev’s post does not identify which companies or specific crawlers are responsible for the traffic, nor does it provide a breakdown of which AI labs or scraping services are driving the load. The post also does not specify an exact date for when Anubis was first deployed, describing it only as “about a year ago.” It is not clear from the post what specific features or anonymous-access capabilities will be disabled, or on what timeline.
Analysis
Ryabitsev frames the trend as largely intractable in the near term, writing that the outcome depends on factors outside kernel.org’s control — either “the AI bubble bursts” and demand for training data drops, or crawler operators “smarten up and stop consuming our data in the dumbest way possible.” In the meantime, his account illustrates a recurring tension in the arms race between open source infrastructure and AI data collection: the Anubis proof-of-work system was designed to make scraping expensive, but as commenters on the LWN discussion thread point out, that cost can be offloaded onto third parties when scrapers route requests through residential proxy networks running on devices like televisions, making the computational toll fall on unwitting device owners rather than the scraping operators themselves.