Harness Router starts from a simple gap: you already have unified layers like OpenRouter that let you call every large language model from one place, but there has been no equivalent router for the harness layer that puts those models to work. The presenter makes this clear on screen: a model alone can write text or make images, but it cannot open your files, walk through them step by step and hand back a finished spreadsheet — that job belongs to the harness. Harness Router, as described on harnessrouter.ai 's open-source page, is exactly that router: an open, local-first hub where you pick Claude Code , Codex CLI , Gemini CLI or Hermes and run them from a single dashboard. The repository on GitHub at github.com/harnessrouter/harnessrouter keeps the whole setup auditable, and the app itself runs on your own machine.
A strong model is not enough on its own
The core claim is that the harness matters even when the model is fixed. In a study the team calls Harness Bench , 50 realistic spreadsheet tasks were given to four different harnesses wrapping the exact same model; the best harness solved 85% correctly while the weakest solved 77% and the winner took roughly half the time. The methodology and leaderboard live on harness-bench.ai and stress sandboxed, offline, reproducible runs — not cherry-picked prompts. To make the idea concrete: the model is the engine that predicts the next token, the harness is the chassis that adds file I/O, tool use, planning and memory. Anthropic 's own engineering essay on anthropic.com about harness design for long-running apps makes the same point with three patterns for reliability. Bringing that lens to Harness Router helps you see it as OpenRouter for harnesses, not just another launcher.
Setup is deliberately boring, which is a good sign. If you already have free Docker Desktop , Harness Router boots with a single command and then lives in your browser on localhost — the same local-agent pattern documented on docker.com in the Docker Agent getting-started guides. Because it runs locally, your data stays on disk and only leaves when you attach a key and make a model call. That also keeps cost and privacy sane: you bring your own keys, there is no mandatory subscription, and through openrouter.ai you can fan out to 300+ models behind one OpenAI-compatible endpoint and one bill. In the demo the presenter adds a Gemini key from Google AI Studio and an OpenRouter key in the dashboard's key vault, then freely swaps models without touching the app.
One command to a local control plane
The dashboard reads in four blocks. The first lists runnable tools — Claude Code , Codex, Pie and others — with live health and model tags. At the top you can create your own tool, which is how the presenter bakes a policy into a reusable harness. Next are Groups : one-click starter apps such as dashboards and slide decks. Then Extensions : you give your harnesses a browser or other powers. Finally key management. The customization the video shows is telling: a new harness called "expense checker" is built on top of Gemini CLI and the spending policy is written once into its instructions, so the team no longer pastes the policy every month. That pattern mirrors the model-agnostic call style described in openrouter.ai 's quickstart — the harness can change while the underlying model stays constant.
Then comes a real job. The presenter prepares a 34-row August expense report and hides seven violations on purpose: meals over $75 need manager approval, any item over $25 needs a receipt, and no duplicate claims are allowed. First run uses Gemini CLI with the Gemini 3 Flash model; files are attached, the policy is pasted, and after a couple of minutes the same spreadsheet returns with a new "verify" column and a short manager summary: seven rows need a second look, totaling $487.65. Exactly the seven hidden errors. Even nicer, the harness spots that two people claimed the same taxi ride on the same day and asks whether they shared the cab — contextual reasoning beyond a rule filter. Opening the step trace shows every file check that led there, a transparency habit that harness-bench.ai argues is essential for repeatable agent work.
A real workload: 34 rows, 7 errors and $487
The second run proves portability. Without any new setup the presenter flips to the Pie harness, keeps the same model, file and prompt, and gets the same seven rows — the harness switched, the result did not drift. Next the custom expense-checker harness is used: no pasted policy, just "check this month's report", and again the same seven rows and total appear. For a team that does this monthly, that is the payoff: the policy lives in one place, everyone presses the same button, and outcomes stay deterministic. The community edition docs on harnessrouter.ai and the examples on GitHub recommend exactly this: harden one harness, publish it, and bind your org to a single contract.
The last act tests two starter groups. The dashboards group is wired to a fake e-commerce database, Hermes is chosen as the harness and the OpenRouter key is used for calls. The prompt is plain English: "show revenue by month, top five products, revenue by region". In under two minutes the harness discovers the schema on its own, writes the queries and renders the charts; total revenue is about $1.48 million and the presenter checks each number against the database. The slides group then builds a five-slide leadership deck from the same data: title slide, top five products, two things leadership should watch, and every slide remains manually editable. Both artifacts are produced locally, finish in the browser and cost nothing beyond the model call. The closing line lands: Harness Router is free , open source , runs on your laptop, and only needs a key from a model provider.
Key moments
- OpenRouter analogy: we have a model router, no harness router yet
- What a harness is: opens files and returns a finished sheet
- Harness Bench: 85% vs 77% on 50 tasks, half the time
- One Docker command, browser localhost panel
- 34-row expense report with seven hidden violations
- Gemini CLI catches seven rows at $487.65 and flags shared taxi
- Switch to Pie, same model and file, same result
- Custom expense checker harness with policy baked in
- Dashboards group: Hermes builds $1.48M dashboard
- Slides group: five-slide leadership deck, editable
AI commentary
"Harness Router is the clearest sign that leverage is moving from picking a model to picking the harness that lets it work. Comparing that layer in one panel and baking policy into a reusable harness is a pragmatic win for teams."
AI assessment
The strongest pushback is abstraction cost. With new CLIs, flags and MCP servers shipping weekly, a single router cannot cover every edge on day one; the freshest streaming or tool-calling modes often land natively first and only later in the router. Anthropic 's essay on anthropic.com about long-running harness design also reminds you that state management and recovery are harness-specific: a uniform interface can smooth over important differences. Running everything under Docker is not universal either; on managed corporate laptops Docker Desktop policies or virtual-network restrictions can block the one-command story. So treat Harness Router as an excellent comparison bench, not as the sole pillar for a production pipeline before you have piloted it on your own repo.
What the video leaves out matters too. A 50-task Harness Bench spread of 85% versus 77% is striking but narrow — it zooms in on spreadsheets, and the ranking can shift on code generation, multi-step web research or long-context memory tasks. The exact model and version behind the score also changes outcomes; the presenter says Gemini 3 Flash , yet the full breakdowns on harness-bench.ai show different orderings per model. Security is the other open question: a browser extension and file access are powerful, but each harness carries its own permission model, and granting keys to many harnesses from one panel can widen blast radius if misconfigured. The video rightly stays on the demo, but those gaps deserve your own tests.
Speaker interest looks limited and transparently disclosed as personal opinion with a "views are my own" note, and there is no paid brand placement in the flow. Still, the examples lean toward the best case — a perfect 7-out-of-7 and a dashboard in under two minutes — and messier real schemas or a hallucinating model can break charts unless you keep the verification step the presenter does. The practical takeaway is to use Harness Router as a comparison rig : run the same task, same model, same files across two or three harnesses pulled from GitHub , try the install once via docker.com on a spare laptop, then publish a policy-infused harness for your team and meter cost centrally through openrouter.ai . The leverage is less about picking the model and more about picking the harness that lets it do real work.
Sources
7 links; 1 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — HarnessRouter: Run Claude Code, Codex & Gemini CLI From One Place
- @github.com GitHub — HarnessRouter/harnessrouter
- @harnessrouter.ai HarnessRouter — Community Edition open-source
- @harness-bench.ai Harness Bench — Measuring Harness Effects
- @openrouter.ai OpenRouter — Quickstart
- @anthropic.com Anthropic — Harness design for long-running apps
Also cited by: Agent Harness Explained: From Tiny Context Windows to Loops That Ship
- @docker.com Docker — Docker Agent installation
harness router · ai agent · docker · openrouter · harness bench