The video opens with a clear aspiration: not a model that never errs, but a system that never repeats the same mistake twice. Most interactions remain one-way, you send a prompt and receive an answer. The loop proposed here inverts that. A prompt generates a strategy, the strategy trades and produces an outcome, the outcome becomes training signal for a new prompt that rewrites the agent. Built from a single copy-paste prompt and described as free to run, the design is framed as a 24/7 loop rather than a one-off script.
The author, acting as an architect, anchors the design in four criteria: accuracy, reliability, a well-defined goal, and self-improvement. Without all four, even a strong model fails to become a consistent trading system. Manually scripting self-learning in tools like Claude is described as tedious and brittle. Against the buzz around fully autonomous frameworks such as OpenClaude, Hermes is introduced as a quieter, background option presented as more ambitious on the self-learning side.
Accuracy starts at the data source. The author says he tested every major model on trading tasks in recent weeks and the most striking result was inconsistency. Models expected to pull from the same feed produced different numbers, and the same news article led to different conclusions across agents. The remedy proposed is resilient API connectivity, consistent interpretation of news flow, and explicit rules that keep conclusions verifiable and objective, so both prices and narratives align before the agent reasons.
Reliability means continuous execution. The agent must stay on day and night, even when the local machine is off. In the video this is solved by hosting on Railway. A one-shot setup opens a CLI integration that syncs every strategy update to a live server. Railway is presented as staying in a generous free tier for a long time and comfortably carrying dozens of projects, so the agent remains active after the terminal closes and keeps its schedule intact.
The third criterion is a sharply defined goal. Nearly everyone building a strategy, the author argues, lacks a destination. What counts as success or failure, ten dollars a month or a million, which Sharpe, which maximum drawdown? A deliberately impossible anchor, like aiming for a million a month from ten dollars, illustrates why calibration matters. The more concrete the target, the more meaningful feedback becomes. Each result is then positioned as toward the goal or away from it, and the loop steers by that compass until the target is met.
Self-improvement is cast as a scientific method running on that compass. The agent gathers and organizes information, analyzes the outcome for proximity to the goal, forms a hypothesis about why the result occurred, and forms a second hypothesis about what to try next. Only one variable changes per iteration, because changing many at once hides the cause of any gain. Every improvement becomes the new baseline for the next series of tests, and learning accumulates rather than resets.
Hermes is positioned alongside a distribution channel. The heart of the free setup lives in a community called 01 Systems. From the top link in the description, the classroom section holds ready prompts among YouTube entries, and the package evolves with feedback so each download delivers the latest revision. Positive responses to a prior Markov-method strategy are shown as evidence that this community feedback loop is live and responsive.
The demo starts in the terminal. A new session opens with permissive flags and the one-shot prompt is pasted. Phase one is an environment check that detects Mac versus Windows and confirms Node.js and Claude Code. Phase two defines the strategy: how a file that scores every trade will encode success and failure, and which asset to trade. Solana, dollar, Ethereum and Bitcoin are listed as options, with three paths offered, create a basic strategy from scratch, co-build it with the onboarding agent, or point to an existing one. The author chooses the fourth path and points to his live strategy.
That strategy is identified as Wacko Alpha, described as a D Tau momentum and yield system that arrives with more than one and a half million data points collected over six to eight weeks. It trades real money and is tracked on a dashboard for a tenfold challenge from fifty thousand to five hundred thousand pounds. The system scans the filesystem and pulls targets into a confirmation screen: a thirty-day return target of 4.7, which maps to about forty seven percent per interval, a six-month tenfold perspective, a minimum Sharpe of one and a capped drawdown. Those thresholds travel with the deployment.
Phase three scaffolds state for Hermes, preparing folders and files for review. Phase four is skipped because the strategy is already live; otherwise API hooks and a TradingView execution flow would be provisioned through code generation. Railway login hits an interactive session guard, so the author splits the terminal, pastes the command externally, completes browser authentication and returns with success. The CLI integration then takes over, pushing every future change to the hosted service and keeping local and remote in sync, a pattern aligned with Railway's remote MCP offering.
The closing board is explicit: a ledger of twenty four winning and twenty two losing trades is converted into a Hermes-readable form, and a strategy document is populated with a twelve-position cap, slippage tolerance, gas reserve and scorer weights. Hermes installs in seconds and becomes callable as a command. The deployment is framed as running on Bittensor subnets, with a weekly Hermes review and thirty-minute plus daily rebalancing. The first cycle is read-only and produces a markdown review with no writes; going live awaits a mode flip. A second agent named Cornelius tunes learned parameters weekly on a calendar offset by three days from Hermes, and check-in commands become available a day after review.
Key moments
- Holy grail: an agent that does not repeat mistakes
Prompt creates strategy, strategy creates outcome, outcome feeds the next prompt.
- Four criteria: accuracy, reliability, goal, self-improvement
Without all four, a strong model does not become a consistent system.
- Accuracy test: same source, different numbers across models
- Defining the goal: what success and failure mean in numbers
Each result is positioned toward or away from the goal.
- Scientific method: one variable per iteration
Each improvement becomes the new baseline.
- Demo phases: environment check and strategy selection
- Wacko Alpha: 1.5M data points and the 10x dashboard
A 4.7 target per thirty days maps to the board.
- Railway setup: browser login and CLI sync
- Close: 24 wins 22 losses ledger and weekly loop
First Hermes cycle is read-only and review-only.
AI commentary
"In my view, this video does not reduce automation to a magic prompt; it ties it to a four-part architecture and a human-approved weekly loop."
AI assessment
Steelmanning the pitch, a single-prompt self-learning trader is compelling but remains unmeasured. The HKU Business School AI-Trader study and BurningTheta's analysis of the same work show large models delivering weak live returns and fragile risk control. Figures in the video such as twenty four wins against twenty two losses and one and a half million data points describe mechanics, not proof of edge; automation without verification can become marketing.
Methodology leaves gaps. Which asset, which period, which trading costs, what slippage and gas drag on net return, and which dataset is free of look-ahead? The TradingAgents framework quietly fixing a forward-looking data leak in v0.3.1 after passing one hundred thousand stars is a reminder that bright backtests deserve a second look. The video rightly stresses a one-variable scientific method, yet it does not show an independent validation step, cost accounting, or external audit before going live.
On incentives and verifiability, the beneficiaries are clear: the creator community, the Hermes framework, the Railway hosting layer and Bittensor subnets. Claims such as forty seven percent per thirty days, tenfold in six months, a Sharpe floor of one and a capped drawdown need independent replication. Early evidence of recursive self-improvement reported by Weco and MIT Technology Review's note that recursive gains may arrive more slowly both temper expectations; without lab confirmation those numbers stay as targets.
In my view the practical take is this: the setup fits builders who want to encode trading logic, keep memory across runs, and review weekly with a human gate. It does not fit those chasing quick profit or skipping cost, slippage and drawdown accounting. Running in paper mode, simulating costs thoroughly and piloting live at small size before flipping the mode aligns better with the weekly review and one-variable discipline the video itself advocates.
Sources
8 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube.com YouTube — Hermes Agent Video (Lewis Jackson)
- @hermesagents.net https://hermesagents.net/blog/hermes-self-improving-loop/
- @burningtheta.com https://www.burningtheta.com/article/ai-trading-agents-fail-live-market-benchmark
- @hku.hk https://www-1.hku.hk/press/news_detail_29241.html
- @railway.com https://blog.railway.com/p/agent-rails-remote-mcp-cli
- @notatechguy.com https://www.notatechguy.com/tradingagents-hits-100k-stars-with-data-leak-fix-in-v0-3-1/
- @weco.ai https://www.weco.ai/blog/first-evidence-of-recursive-self-improvement
- @technologyreview.com https://www.technologyreview.com/2026/08/18/1142188/ai-recursive-self-improvement-might-not-come-quickly/
hermes · trading agent · railway · bittensor · wacko alpha