Skip to content
Insights
Above

TrueForge: The Hard Part of AI Agents Isn't the Model

TrueFoundry's TrueForge is an open-source agent runtime claiming up to 75% lower costs than Claude Managed Agents. What a harness actually does, what the benchmark really shows, and why it's good news for open models.

By Adam Maguire Wilson12 min read
On this page

Ask a room of engineers what makes an AI agent hard to build, and most will say the model. Ask a room of engineers who have actually shipped one, and you get a different list: sessions that survive a restart, tool credentials that don't leak into the sandbox, a human approval step that works at 2am, a context window that doesn't quietly cost you ten times what you budgeted. The model is the part you can swap in an afternoon. The machinery around it is the part that eats the quarter.

That machinery is what TrueFoundry open-sourced on 19 August. TrueForge is an agent harness, the runtime layer that runs the loop around a model: plan, call a tool, feed the result back, repeat, while managing sessions, sandboxes, approvals and context along the way. The launch pitch is a vendor-neutral alternative to Anthropic's Claude Managed Agents at roughly half the operating cost, and the benchmark behind that pitch has a Chinese open model doing the heavy lifting, which is the part I find most interesting. I've read the repo, the announcement post and the coverage. Here's what holds up.

Key Takeaways - TrueForge is an MIT-licensed agent harness from TrueFoundry: a core server running the agent loop, an HTTP API with a TypeScript SDK, and an embeddable chat UI. Any model provider, any MCP server, your own infrastructure. - The headline benchmark: on a 14-task enterprise eval, TrueForge matched Claude Managed Agents on task completion at about 30% lower cost on the same model, and about 75% lower when paired with GLM-5.2 instead of Opus 4.8. It's TrueFoundry's own benchmark, not independently validated. - The runtime is free; the business is the gateway underneath it. TrueFoundry sells the governance layer (budgets, rate limits, audit, observability) that every model and tool call routes through. - For open models, this is the more important story: a neutral harness means the cheapest model that finishes the task wins the workload, and right now that race increasingly involves Chinese labs. - Local mode is a single process with SQLite and no login. Treat it as a way to kick the tyres, nothing more.

What happened

On 19 August, TrueFoundry, a San Francisco AI infrastructure startup founded in 2021 by former Meta engineers, released TrueForge under the MIT licence. The launch coverage landed in VentureBeat and InfoWorld the same week, and the repo passed 1,800 stars within days. The key facts:

  • The harness runs the full agent execution loop: model calls, MCP tools, skills, sandboxed code execution, human approvals, context management and session state, streamed end to end.

  • It's exposed three ways: a bundled chat UI, an HTTP API with a TypeScript SDK (@truefoundry/trueforge-sdk), and an embeddable UI SDK (@truefoundry/trueforge-ui) for putting the interface inside your own product.

  • Models come from YAML catalogs: OpenAI, Anthropic, Gemini and twenty-odd others, plus any OpenAI-compatible endpoint. Skills are git-backed SKILL.md packs, the same convention Claude Code popularised.

  • It runs in two shapes: local mode (one process, SQLite, started with npx @truefoundry/trueforge) or hosted mode (Postgres plus Redis, via Docker Compose or Helm) for teams.

  • A hosted, pay-per-usage version is available from TrueFoundry for teams that don't want to run it themselves. NetApp and Automatiq are named as early users, alongside TrueFoundry's own internal assistant.

TrueForge is an open-source (MIT) agent harness released by TrueFoundry on 19 August 2026. It runs the agent loop (model calls, MCP tools, sandboxed execution, approvals, context management, session state) and exposes it through a chat UI, an HTTP API with a TypeScript SDK, and an embeddable UI, per TrueFoundry's announcement.

Why the harness is the layer worth fighting over

Building an agent demo is easy. Running fifty of them in production is not, and the gap between the two is entirely harness work. The model reasons, but it can't act: it won't open a file, call an API, remember what it said three turns ago, or stop itself before deleting the wrong table. Somebody's code has to do all of that, repeatedly, safely, at a cost you can predict.

Until recently you had two options. Build the harness yourself (months of plumbing, and you now maintain a runtime instead of a product), or rent one from a model provider, which is what Claude Managed Agents is: Anthropic runs the loop for you, billed as Claude tokens plus $0.08 per running session-hour, metered to the millisecond. The second option is convenient and structurally awkward. The moment the runtime, the tools and the billing all belong to one vendor, so does your roadmap.

This is also why the context-engineering features matter more than the model list. TrueForge ships subagents, deferred tool loading (tool definitions are loaded when needed, not stuffed into every prompt), large-result offloading, and context compaction, all the tricks that keep a long agent run from drowning in its own history. These aren't luxuries. In TrueFoundry's benchmark, its configuration burned 3.8 million tokens per run against Claude Managed Agents' 10 million on the same tasks and the same model. Even if you take that number with a pinch of salt, the direction is right: most agent cost is context waste, and context handling is harness work. It's the same reason I tell clients the interesting question in agentic architecture is rarely which model, it's what wraps it.

An agent harness handles the production concerns models don't: sessions, tool credentials, sandboxes, approvals and context growth. In TrueFoundry's benchmark, its harness used 3.8 million tokens per run versus 10 million for Claude Managed Agents on identical tasks, which it attributes to context compaction and deferred tool loading.

The benchmark, read carefully

The headline number everywhere is 75%, so let's take it apart, because it's really two numbers wearing a trenchcoat.

TrueFoundry tested on DevRev's Enterprise-Bench, a 14-task evaluation spanning simulated CRM, issue-tracking and document systems. Same tasks, same tools, same model:

  • TrueForge running Anthropic's Opus 4.8 completed about 11 of 14 tasks at an average of $8.50 per run.

  • Claude Managed Agents running the same Opus 4.8 completed about the same number of tasks at $11.80 per run.

  • TrueForge running GLM-5.2, Zhipu's open-weight model, again completed about the same number of tasks, at $2.90 per run.

So the honest decomposition: harness versus harness on the same model saves roughly 30%, mostly managed-service markup plus the token reduction above. The jump to 75% comes from also switching the model, from Opus 4.8 to GLM-5.2. That comparison changes two variables at once, as one sharp write-up noted, and it tells you more about how far open models have come than about TrueForge itself. Which is, in a way, TrueFoundry's actual argument: if the harness is neutral, you're free to take that saving.

Two more caveats. The benchmark was run by TrueFoundry and hasn't been independently validated across larger production workloads. And 14 tasks is a small sample: three tasks of noise is the difference between a good and a bad week. Treat the figures as reported rather than audited.

Bar chart of average cost per run on the DevRev Enterprise-Bench evaluation. TrueForge with Opus 4.8 averaged 8.50 dollars, Claude Managed Agents with Opus 4.8 averaged 11.80 dollars, and TrueForge with GLM-5.2 averaged 2.90 dollars, with all three completing roughly 11 of 14 tasks.

Same eval, three configurations. The 30% saving is harness versus harness; the 75% saving mostly comes from swapping the model. Vendor-reported figures, reproducible from the benchmark directory in the repo.

In TrueFoundry's self-run benchmark on DevRev's Enterprise-Bench, TrueForge matched Claude Managed Agents at roughly 11 of 14 tasks: $8.50 versus $11.80 per run on the same Opus 4.8 model (about 30% less), and $2.90 per run when paired with GLM-5.2 (about 75% less), per the announcement. The results have not been independently validated, and the 75% figure changes both the harness and the model at once.

The business model is not the licence

Give away the runtime, sell the layer beneath it. The open-source harness is complete and genuinely free, but TrueForge is designed to route every model call and MCP interaction through TrueFoundry's AI Gateway, and the gateway is where budgets, rate limits, guardrails, access policies and observability live. That's the commercial product, along with the hosted pay-per-usage tier for teams who'd rather not run Postgres themselves.

I don't say that as a criticism; it's the same bargain as a lot of good open-source infrastructure, and the Meta pedigree shows. The founders' pitch is that control and convenience aren't opposites if the platform is right, and giving away the harness is how you get your control layer underneath everyone's agents. Robert Rubin, Senior Director of Platform Engineering at NetApp, describes TrueFoundry as "the central platform in IT where agentic apps and agents are routed through", which is precisely the position the company wants to own. Worth knowing, as a buyer, which layer you're committing to. The harness you can fork. The gateway is a subscription.

The competitive set is getting crowded, which tells you the category is real: Anthropic's managed runtime at the proprietary end, LangChain's Deep Agents in framework land, DeepSeek's own MIT-licensed harness still in developer preview, and now TrueForge aiming at general-purpose production use. When three continents ship the same layer of software within a year, the layer is load-bearing.

TrueForge is free under MIT, but it's designed to route model and MCP traffic through TrueFoundry's commercial AI Gateway, where budget enforcement, rate limits and observability are sold. Early adopters named at launch include NetApp and Automatiq, per the launch coverage.

What it means for open models

Here's the angle I think most Western coverage underplayed. A managed runtime tied to one model family is a moat for that family. A neutral runtime is a leveller: every workload routed through TrueForge is a workload where GLM-5.2, Qwen, DeepSeek or Kimi can win on cost per completed task rather than on brand. The 75% number exists because a Chinese open-weight model finished enterprise tasks at a quarter of the Opus price. A year ago, "the cheap model finished about as many tasks" was not a sentence you could write with a straight face.

That's the pattern I've been watching from Hangzhou for a while now: the open-weight labs don't need to beat the frontier everywhere, they need to be close enough that the layer above stops caring which model is underneath. A neutral harness is exactly that layer, and it's now open source with a governance story attached. If you're weighing whether open Chinese models are ready for real workloads, that calculus is covered properly in the risk framework I use with clients. The short version: "the harness treats it as just another endpoint" removes one of the better excuses for not testing.

What to do now

If you build or run agents, three steps, in order.

  1. Today: run npx @truefoundry/trueforge on your machine and wire in one model and one MCP server. Local mode is a single process on SQLite, which makes it the cheapest way I know to build intuition for what a harness actually does. Keep it on localhost: there's no login by default, and the docs are blunt that local mode is not an internet-facing setup.

  2. This week: take one agent workload you already run, note its token spend and completion rate, then reproduce the benchmark's shape on your own tasks. The repo's benchmark/ directory is designed for exactly this. Your number is the only one that matters; vendor benchmarks, including this one, are marketing until you rerun them.

  3. This month: if you're paying managed-runtime prices, model the 30% harness-only saving against your actual bill, and separately test whether an open model clears your quality bar on the same tasks. Two independent decisions, two independent spreadsheets. If you're earlier in the journey and still choosing your first stack, start with build versus buy and the self-hosting trade-offs before committing to any runtime.

FAQ

What is an agent harness?

The runtime layer that turns a language model into a working agent. It runs the loop around the model (plan, call tools, feed results back, repeat) and handles everything the model can't: session persistence, tool credentials, sandboxed execution, human approvals, and keeping the context window under control on long tasks. TrueForge, Claude Managed Agents and LangChain's Deep Agents are all harnesses; the models underneath them are interchangeable to varying degrees.

Is TrueForge free?

The harness itself is MIT-licensed and free, including for commercial use. What you still pay for: the models, the sandbox provider (Daytona today), and your own infrastructure. TrueFoundry sells an optional governance layer (the AI Gateway, with budgets, rate limits, audit and observability) and a hosted pay-per-usage version for teams that don't want to self-host.

How does TrueForge compare to Claude Managed Agents?

Claude Managed Agents is fully managed by Anthropic, billed as Claude tokens plus $0.08 per running session-hour, and Claude-only. TrueForge is self-hosted (or hosted by TrueFoundry), model-neutral, and free to run. In TrueFoundry's own benchmark, TrueForge matched it on task completion at about 30% lower cost on the same model. The trade: you operate more yourself, and the benchmark is the vendor's, not an independent one.

Can TrueForge use Chinese open models like GLM or Qwen?

Yes. It supports any OpenAI-compatible endpoint, and the launch benchmark's headline figure came from pairing it with GLM-5.2. That pairing is where the 75% cost reduction comes from, which says as much about open-model pricing as it does about the harness.

The bottom line

TrueForge is a month old, its benchmark is its own, and nobody should be routing payroll through it yet. But the shape of the thing is right: the harness is becoming its own layer of the stack, that layer was always going to get an open-source neutral option, and the fact that its launch maths leans on a Chinese open model to make the headline number is the quiet tell about where agent economics are heading. Watch whether independent replications of the cost claims appear over the next quarter. That's the difference between a launch and a shift.

If you're trying to work out where a harness sits in your own agent stack, that's a conversation I have with clients regularly. Get in touch.

Sources

  • TrueFoundry, "TrueForge: Open-Source Alternative to Claude Managed Agents": https://www.truefoundry.com/blog/engineering/trueforge-open-source-agent-harness/ (published 2026-08-18, retrieved 2026-08-29)

  • TrueFoundry, TrueForge repository: https://github.com/truefoundry/trueforge (retrieved 2026-08-29)

  • TrueFoundry Docs, "What TrueFoundry Adds on Top of TrueForge": https://www.truefoundry.com/docs/agent-platform/agent-harness/what-truefoundry-adds (retrieved 2026-08-29)

  • VentureBeat, "TrueFoundry's open source AI agent harness TrueForge boasts 30%-75% cheaper task completion": https://venturebeat.com/orchestration/truefoundrys-open-source-ai-agent-harness-trueforge-boasts-30-75-cheaper-task-completion-than-claude-managed-agents (published 2026-08-19, retrieved 2026-08-29)

  • InfoWorld, "TrueFoundry debuts open-source AI agent harness, claiming up to 75% lower costs": https://www.infoworld.com/article/4211969/truefoundry-debuts-open-source-ai-agent-harness-claiming-up-to-75-lower-costs.html (published 2026-08-20, retrieved 2026-08-29)

  • Superpower Daily, "TrueFoundry Gives Away Its Agent Runtime to Sell the Layer Beneath It": https://www.superpowerdaily.com/posts/truefoundry-gives-away-its-agent-runtime-to-sell-the-layer-beneath-it (published 2026-08-20, retrieved 2026-08-29)

  • Business Wire, "TrueFoundry Launches TrueForge" (press release): https://finance.yahoo.com/technology/ai/articles/truefoundry-launches-trueforge-open-source-120000790.html (published 2026-08-19, retrieved 2026-08-29)

Keep reading

Agent Field Notes

Get the next issue.

Agent harnesses, runtimes, security and governance, explained for the people who have to operate them.

Facing a decision like this?

We run architecture reviews, governance assessments and version-pinned framework evaluations for teams making consequential agent decisions.

About the author

Adam Maguire Wilson

Founder and independent advisor on AI agent systems.

adam.mw