TrueFoundry Open-Sources TrueForge, an Agent Harness Benchmarked Up to 75% Cheaper Than Claude Managed Agents
TrueFoundry's MIT-licensed TrueForge harness matched Claude Managed Agents on a 14-task enterprise benchmark while costing up to 75% less per run.
Overview
TrueFoundry has open-sourced TrueForge, an MIT-licensed agent harness that the company says matched Anthropic’s Claude Managed Agents on accuracy across a 14-task enterprise benchmark while costing up to 75% less to run, according to VentureBeat. TrueFoundry detailed the benchmark methodology and results in a companion post on its engineering blog.
What Was Launched
TrueForge is a self-hosted “harness” — the runtime layer that sits between a language model and the tools, sandboxes and session state an agent needs to actually complete a task. TrueFoundry’s engineering blog describes it as “the harness we run in production,” adding that it is MIT-licensed and can be started locally with a single npx @truefoundry/trueforge command. According to VentureBeat, TrueFoundry is a San Francisco B2B machine-learning startup co-founded in 2021 by former Meta engineers, and the harness is published on GitHub under the same permissive license, meaning it can be forked, self-hosted or built into commercial products.
TrueForge’s architecture is built around what the company calls context engineering: VentureBeat reports that the harness delays loading MCP tool schemas until they’re needed, delegates isolated work to subagents, offloads oversized tool results into files instead of the active context window, and automatically compacts long-running conversations, with a default compaction threshold of 50,000 tokens. It also treats the execution sandbox as a tool rather than an always-on environment — TrueFoundry’s blog says TrueForge “only spins one up when the agent needs to run code, so one server runs many agents at once and non-code turns stay cheap.” Developers can run TrueForge locally with SQLite for personal use, or move the same harness into a shared production deployment using Docker Compose or Helm with Postgres and Redis, per VentureBeat; TrueFoundry explicitly warns the local, SQLite-backed configuration is meant only for a developer’s own machine, not as an internet-facing production service.
The Benchmark
TrueFoundry’s benchmark methodology post says the company tested TrueForge against “the two obvious alternatives on the same 14 enterprise tasks, with the same tools and one blind judge: the closed Claude Managed Agents and the open-source deepagents” — LangChain’s LangGraph-based agent harness. The benchmark used DevRev’s Enterprise-Bench, which the post describes as “14 cross-system tasks at varied difficulty that read like real B2B ops work,” requiring an agent to call MCP tools across a Salesforce-style CRM, a Jira-style project tracker, and a Drive-style document and call-transcript store, then join the results and pitch the answer at the right level of detail. Grading was blind: an LLM judge scored each answer against the task’s criteria without knowing which harness or model produced it, and a task counted as solved only if it met every criterion, with no partial credit.
When TrueFoundry ran TrueForge with the open-weight model GLM-5.2 against Claude Managed Agents running Claude Opus 4.8, both configurations solved roughly 11 of the 14 tasks, but TrueForge’s run cost about $2.90 against $11.80 for Claude Managed Agents — a difference of roughly 75%, according to both VentureBeat and TrueFoundry’s benchmark post. TrueFoundry’s post attributes the gap partly to the fact that “Claude Managed Agents can’t run this setup at all: it’s Anthropic-only, so the cheaper-model lever was never on the table.” The post prices the underlying tokens at “$5 / $25 per 1M in/out for Opus 4.8; $0.73 / $2.28 for GLM-5.2.”
Holding the model constant at Opus 4.8 across all three harnesses, TrueFoundry’s benchmark post reports TrueForge solved about 11 of 14 tasks for roughly $8.50 per run in 3.8 million tokens and 40 minutes; Claude Managed Agents solved about 11 of 14 for roughly $11.80 in 10 million tokens and 63 minutes; and deepagents solved about 10 of 14 for roughly $21 in 16.5 million tokens and 64 minutes. TrueFoundry says that works out to TrueForge being about 30% cheaper than Claude Managed Agents and about 2.5 times cheaper than deepagents on identical tasks and the identical model, attributing deepagents’ higher token count to the fact that it “re-reads its own accumulated context step after step.”
Why TrueFoundry Is Giving It Away
Speaking to VentureBeat, Anuraag Gutgutia, TrueFoundry’s co-founder and COO, said the release responds to enterprise demand: “We’ve had this ask from a bunch of customers. You have an ability where you bring in agents and MCPs — can we also get something where you can actually launch these managed agents? I think that is the need we are satisfying. It is not a replacement. People will use this alongside other harnesses, like the cloud-managed ones or the commercial-provider-managed ones, but this will serve as a way for people to use them in a vendor-neutral way and also at a lower cost.”
TrueFoundry already sells a commercial “AI Gateway” that centralizes model and MCP access, credentials, permissions, budgets and observability for enterprises, and Gutgutia framed TrueForge as sitting above that layer rather than replacing it: “There will be a set of companies that will use our harness as the way to launch managed agents,” he told VentureBeat, “but all that traffic should still be flowing through our gateway.” He also cautioned that open-sourcing the harness doesn’t automatically bring enterprise governance controls with it: “If you are using just the open source version of our agent harness, yes, you will need to put the right controls therein or in front of some other internal control system,” he said. VentureBeat reports TrueFoundry named NetApp as a beta user that contributed requirements during development, and Automattic as an early user.
The release follows Anthropic’s own April launch of Claude Managed Agents, a fully managed, Claude-only runtime that Anthropic bills at standard token rates plus a flat per-hour fee for active agent runtime.
What We Don’t Know
TrueFoundry’s benchmark is self-reported and was not independently reproduced by a third party in the sources reviewed for this article. The 11-of-14 and 10-of-14 task-solve figures are presented by TrueFoundry itself as approximate. It is also not yet clear how the harness performs on tasks outside the DevRev Enterprise-Bench suite, or how its cost advantage holds up at larger scale in production rather than in a benchmark run.