New SWE-Prime Method Trains AI Coding Agents on Just 10% of Data, Lifting Bug-Fix Accuracy Up to 24.2%
A new study finds that curating just 10% of AI coding-agent training trajectories outperforms training on the full dataset, with gains up to 24.2%.
Editor's Note ·
- Correction:
- The article states the paper's "Random-10%" baseline scored below the raw, un-fine-tuned model "in two of the six comparisons." The paper's results table (three base models x two benchmarks) shows Random-10% actually scored below the raw model in four of the six comparisons: GLM-4.7-Flash on both SWE-Bench Verified (39.6 vs. raw 40.4) and SWE-Bench Pro (10.81 vs. raw 11.22), and Qwen3-Coder-30B-A3B-Instruct on both SWE-Bench Verified (41.2 vs. raw 44.8) and SWE-Bench Pro (27.22 vs. raw 29.55). Only Qwen3-30B-A3B-Instruct-2507 scored above raw with Random-10% on both benchmarks. The article's broader point -- that Random-10% underperformed training on the full resolved dataset across all six comparisons -- is correct and unaffected by this correction.
Overview
A new study, “SWE-Prime: Fewer Trajectories, Better Performance”, posted to arXiv on August 27, 2026 by researchers affiliated with Sun Yat-sen University, Huawei Cloud Computing Technologies, and Chongqing University, proposes a way to train AI coding agents on a fraction of the usual data while producing agents that resolve more bugs. The paper argues that “task success alone does not guarantee high-quality supervision” for the training process most coding agents rely on, according to the paper.
What We Know
Most AI coding agents are improved through supervised fine-tuning: researchers run an agent against real GitHub issues, keep the trajectories — full step-by-step logs of everything the agent did — where the agent successfully resolved the issue, and retrain the model on those successful transcripts. The paper’s abstract describes this as the current norm, noting that “prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories,” according to the paper.
The researchers argue this approach has a blind spot: “successful trajectories may still contain ineffective, redundant, or risky behaviors that should not be used as supervision for SFT,” according to the full paper. In other words, an agent transcript that ends with a passing test can still contain wasted steps or bad habits worth filtering out before it’s used to train the next model.
To address that, the team built SWE-Prime, a two-stage data-selection pipeline. According to the paper, “Stage 1 performs trajectory-level selection by jointly considering three dimensions: process quality, result quality, and data representativeness,” while “Stage 2 performs segment-level selection within the trajectories retained by Stage 1 through three components: Segment Chunking, Segment Scoring, and Segment Filtering.”
The starting material was the “SWE-rebench OpenHands Trajectories” dataset released by Nebius, which the paper says “contains 67,074 trajectories generated by Qwen3-Coder-480B-A35B-Instruct with OpenHands.” Of those, “32,161 trajectories successfully resolve their corresponding issues,” according to the paper, and that resolved subset became SWE-Prime’s candidate pool for selection.
Testing SWE-Prime’s curated 10% subset against three open base models — GLM-4.7-Flash, Qwen3-30B-A3B-Instruct-2507, and Qwen3-Coder-30B-A3B-Instruct — on the SWE-Bench Verified and SWE-Bench Pro benchmarks, the paper reports that “training on the 10% trajectory subset selected by SWE-Prime outperforms training on the full resolved dataset, yielding relative performance gains of up to 12.2% and 24.2%, respectively,” according to the paper. The largest single gain came with GLM-4.7-Flash on SWE-Bench Verified, where fine-tuning on the entire resolved dataset moved the raw model’s score from 40.4 to only 41.4, while SWE-Prime’s 10% subset pushed it to 51.4.
The gains were not simply a function of using less data. The paper also reports results for a “Random-10%” baseline — fine-tuning on a random 10% slice of the same resolved trajectories instead of SWE-Prime’s curated selection — according to the full paper. That random baseline underperformed training on the full resolved dataset across every model and benchmark tested, and in two of the six comparisons it scored below the raw, un-fine-tuned model.
What We Don’t Know
The paper is an arXiv preprint that has not, as of publication, undergone formal peer review. It does not state whether the SWE-Prime selection code or the curated 10% training subset will be released publicly; no repository or artifact link appears anywhere in the paper’s text. It also contains no dedicated limitations section, so questions such as how the method performs on trajectory data generated by agent scaffolds other than OpenHands, or on base models larger than the roughly 30-billion-parameter class tested here, are not addressed in the material reviewed for this article.
Analysis
For teams building or fine-tuning their own coding agents, the practical hook is cost. Agent trajectory data is expensive to produce, since generating each one means actually running an agent against a real repository until it resolves an issue, then keeping only the successful runs. If a curated tenth of that data trains a stronger agent than the full dataset does, that points to a real reduction in the compute and data-curation budget needed to reach a given accuracy bar — provided the selection is done deliberately, since the paper’s own “Random-10%” comparison shows that simply throwing away 90% of the data at random does not produce the same benefit and can even underperform the unmodified base model. Whether SWE-Prime’s specific selection pipeline holds up on trajectory sets generated by agent scaffolds other than OpenHands, or on base models well outside the roughly 30-billion-parameter range tested, is a question that further work — or formal peer review of this preprint — will need to answer.