Back to feed

Not a Coding Tool but an Agent Runtime: Running Your Own Agent Layer with TrueForge

NeuralNine demonstrates TrueForge, TrueFoundry's MIT-licensed open-source agent harness, as a vendor-neutral runtime layer and alternative to Claude Managed Agents. The walkthrough covers the agent loop idea, a 14-task cost benchmark, local setup, test chats, a Python API stock assistant, a Daytona sandbox report agent, skills and local models.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — lcMf0W0cxig
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

NeuralNine frames TrueForge from the first minute as infrastructure, not as another coding assistant next to Pi or OpenCode. The product is an MIT-licensed harness from TrueFoundry that supplies the runtime layer for general-purpose agents. I find the framing useful because it sets the right expectation: this is the backbone of an agentic system, not a chat window.

The video then explains what a harness actually covers, and I think this is its best two minutes. A language model by itself only maps input tokens to output tokens, while an agent also needs tool calls, MCP connections, code sandboxes, approvals from a human and waiting states. TrueForge owns that surrounding loop: reasoning, acting, reading results and continuing until the job is done. Once you see the loop as the product, the comparison with Claude Managed Agents starts to make sense.

Before the demo, the presenter puts a disclosure on the table: the video is a paid collaboration with TrueForge, though the opinions are his own. I credit that openness, and it matters for everything that follows because the cost claims originate from the vendor. The invitation is to try everything shown for free and to star the repository, so at least the viewer can verify the basics without paying.

The benchmark comes from an X post and covers 14 tasks from DevRev's Enterprise-Bench, spanning CRM records, issue trackers and document stores. Three configurations are compared: Claude Managed Agents with Opus 4.8, TrueForge with the same Opus 4.8, and TrueForge with the open GLM-5.2 model. Reported completion is 11 of 14 in each comparable run, with per-task cost at 11.80 dollars managed, 8.50 dollars on TrueForge with the same model, and 2.90 dollars with the open model. That is roughly 30 percent cheaper on equal models and 75 percent cheaper with the open model.

TrueFoundry itself is worth a sentence of context. It is a San Francisco machine-learning infrastructure company founded in 2021 by former Meta engineers, and it already sells enterprises a paid gateway for model access, credentials and budgets. TrueForge sits above that gateway as the open execution loop, which explains the business logic: give away the harness, keep the control plane. Forbes describes the same split, with the 50 percent saving headline needing qualification.

Installation is deliberately boring, which I mean as praise. You copy an NPX command from the GitHub repository, run it once, and it downloads the package, creates a local SQLite database and serves a page on localhost. The browser view that opens is your control panel, and from there you attach whatever models you like. A local-first start like this lowers the trial barrier to nearly zero.

Model setup covers the big providers plus a self-hosted escape hatch. You can paste an OpenAI key, do the same for Anthropic or Gemini, or register a custom provider with a base URL pointing at Ollama. The video walks through key creation on the provider console and then shows the models appearing inside TrueForge. That vendor-neutral shelf, including local models, is the core promise of the whole project.

The chat view that greets you is explicitly a laboratory, not the product. You open a new chat, toggle connectors such as Linear, Notion or Exa web search, and experiment with prompts before anything is permanent. A visible save button marks the boundary between playing and shipping. I like this separation because it forces you to prove the setup works before you name it and expose it.

The first full example builds a stock research assistant. A throwaway test prompt asks for Nvidia financials, the agent fans out into tool calls and a sub-agent, and it visibly digs through 10-Q and 10-K style filings before summarizing. Then comes the key lesson: none of that chat history becomes the agent. The saved definition is a short system prompt plus the chosen model, connectors and skills, after which a bare MSFT input produces a fresh Microsoft rundown.

The payoff section moves the agent out of the browser and into Python. There is no Python SDK, only a TypeScript one, so the demo uses plain requests against the local API: open a session with the agent name, keep the returned session id, post a user message and read back a server-sent-events stream. The script filters the stream for model message deltas and prints them live, and a Tesla query returns streamed findings while the dashboard shows each agent step. A short program like this turns the harness into a genuine runtime layer for your own apps.

The second agent raises the stakes with a Daytona code sandbox. After creating a full-access API key and pasting it into the sandbox settings, the agent can execute Python instead of merely describing numbers. Asked for Meta financials as a PDF with Matplotlib plots, it searches the web, runs the code remotely and hands back a downloadable report with generated charts. Saving that as a second agent with charting baked into its instructions makes every later ticker, such as Nvidia, follow the same pipeline automatically.

Skills get their own compact demonstration through algorithmic art. Without any skill attached, a creative request yields a plain SVG or HTML file; with the dedicated skill enabled, the agent first reads the skill instructions and produces a visibly more professional piece with adjustable parameters. The contrast is the whole argument for skills in miniature: packaged know-how beats raw generation. GitHub-importable skills plus custom MCP servers are what turn the trivial web-search demo into something proprietary.

The closing stretch returns to vendor neutrality with local models. A custom provider entry aimed at the standard Ollama base URL, a manually entered model id from the Qwen family with around 27 billion parameters, and pasted context and output limits are enough to bring a home machine into the same workflow. The presenter runs it from his own workstation hardware and gets answers through the identical loop that earlier served the OpenAI model. Everything open, everything owned: that is the note the video ends on, next to a reminder that the project lives or dies by community stars.

Visualization: nodesdaily AI
SetupCompletedCost
TrueForge + GLM-5.211/14$2.90
TrueForge + Opus 4.811/14$8.50
Claude Managed + Opus 4.811/14$11.80

AI commentary

"I see TrueForge as infrastructure rather than yet another chatbot wrapper, and that distinction is exactly why I find this video worth covering."

AI assessment

The strongest objection is that a managed service exists precisely so teams do not babysit infrastructure, and for a small team the hosted option can still be the rational choice. Running your own loop means owning upgrades, credential rotation, sandbox hygiene and incident response, and those hours rarely appear in a benchmark chart. I would steelman the other side like this: the 30 percent saving buys you a second job as your own platform team.

What the demo does not test is the adversarial side of self-hosting. A Microsoft security write-up on self-hosted runtimes warns that downloaded skills and untrusted text inputs converge into one execution loop, so an unguarded deployment risks credential exposure, poisoned memory and host compromise. The video connects third-party connectors, custom MCP servers and downloaded skills with no discussion of scoping or isolation, and the sandbox segment covers code execution without covering credential boundaries. That gap matters more than any missing feature.

On provenance, two things need daylight. The video is paid for by TrueForge, which the presenter discloses openly, and the headline numbers come from the vendor's own blog post rather than an independent rerun. The per-task figures of 2.90, 8.50 and 11.80 dollars check out through VentureBeat's reporting, but exact reproducibility on DevRev's Enterprise-Bench would need a fresh run with pinned prompts and tool versions. I treat the direction as credible and the decimals as vendor-supplied.

My practical read is that TrueForge fits builders who already operate infrastructure and want model freedom, local execution and auditable loops. If you only need one assistant with web search, this is overkill and a managed service will serve you faster. I would start with the local install, wire one real internal tool through a scoped credential, and only then decide whether the savings survive contact with your own operations load.

Sources

9 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

trueforge · agent harness · open source · claude · truefoundry · daytona · ollama

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…

Not a Coding Tool but an Agent Runtime | Nodesdaily