Our position-sizing tool sat quiet for 30 trades before we let it touch anything. It watched real trades come in. It logged what it would have told us to do. It changed nothing.
That wasn’t caution for caution’s sake. That was the plan from day one.
This is post one in a ten-part series on what we’ve actually learned building tools and testing strategies with Claude, Anthropic’s AI. Glenn and I aren’t AI developers by trade. We’re traders who started using Claude to build things we needed and couldn’t buy off the shelf. Some of what we built worked. Some of it didn’t, until we fixed how we were building it. This series is the honest version of that process. What to do, what not to do, no polish added.
Today’s post covers the do’s. Specifically, the three things that kept an AI-built risk tool from ever putting our account in danger, even while it was still rough around the edges.
I (Reid) handle most of the AI and content systems at HTA, so I was the one sitting with Claude when we started building a position-sizing guardrail tool. The idea was simple. Feed it a trade log, have it track drawdown, and have it tell us the maximum size we should trade next.
Not the minimum. Never “size up.” Maximum, only.
Before a single line of code got written, we wrote one governing rule and made Claude repeat it back to us. This tool may only ever reduce position size. It may never increase it.
Every design decision after that got checked against that one sentence. New feature idea? Does it violate the rule. Edge case in the logic? Does it violate the rule. It sounds almost too simple to matter. It mattered more than anything else in the build.
Here’s the reasoning. A risk tool doesn’t earn trust by promising bigger wins. It earns trust by protecting you on your worst days. A tool that can only ever tell you to trade the same size or smaller can’t blow up your account, even if the code underneath it is buggy. Cap the downside of the tool itself, and you’ve capped the downside of every mistake you haven’t found yet.
That’s risk management applied to the risk management tool. If you’ve read our post on Trading Risk Management Strategy: The Psychology Edge, you know this is the same instinct we teach traders to apply to their own rules. Decide the boundary before emotion, urgency, or a good-looking result can talk you out of it.
Once the tool existed, the temptation was to turn it on and let it start capping our size in real time. We didn’t.
Instead, we ran it in shadow mode. It watched real trades and recorded what it would have said. Reduce to this size, hold at this size. It touched nothing, for more than 30 trades. No live influence at all.
Before we collected a single data point, we wrote down our pass/fail criteria. Would the tool’s caps have actually reduced our worst drawdown, and by how much? How much profit would we have given up on the trades where it capped us and we would have been fine anyway? We locked those questions in first, on purpose, so we couldn’t quietly redefine “success” once we saw how the data leaned.
That’s the part people skip. It’s easy to build something, glance at the output, and decide it’s good because it feels good. Deciding your criteria in advance is a discipline move as much as a technical one. It’s the same reason we push traders to journal their plan before the trade, not after. It’s Process Over Profits, applied to software instead of a setup.
Shadow mode is slow. It’s also the only way to know if a tool actually helps before you hand it the keys.
Here’s where it got humbling. We eventually audited a version of the tool Claude had built for us. On the surface, it looked professional. Tests passed. Documentation was clean and well organized. It read like something you’d trust.
Then we checked it against our real environment, and it fell apart. The tool was reading from a dead data file, not the live trade log we thought it was pointed at. Results weren’t reproducible. Run it twice, get two different answers. The config file referenced settings that didn’t even exist in the code.
None of that showed up in the tests, because the tests were checking whether the code ran, not whether it did the job. Passing tests is not the same as a working tool. It’s the same lesson we teach around backtesting. A strategy that passes on paper still has to survive contact with live conditions before you trust it with real size. We wrote about this exact trap in our TradeZella Backtesting Guide. A clean backtest report tells you the math works, not that the strategy will hold up when the market gets messy.
The fix wasn’t more tests. It was checking the tool against the thing it was actually supposed to do, in the actual environment it would run in.
Everything, honestly. REPs, which stands for Risk, Edge, and Psychology, is our core framework for a reason. It applies past the chart. Building an AI tool has the same failure points as trading a live account. Overconfidence in something that looks good. Skipping the step that would tell you the truth. Believing your own result before you’ve stress-tested it.
The governing rule protected us from ourselves as much as from the code. The shadow mode protected us from confirmation bias. The audit protected us from mistaking “looks finished” for “is finished.” Patience is key in all three, and none of them are exciting. That’s kind of the point.
We talked through more of this build on the Edge Up Podcast, Episode 077, “Using Claude and AI in Trading,” if you want the longer conversation on Spotify.
Amateurs trust a tool because it looks done. Professionals check it against reality first.
Everything we build at HTA starts with the same idea you just read: process over profits, risk before edge. If that’s the way you want to trade, here’s where to go next.
Start free. Take our free Trader Process Assessment. Thirteen questions, about three minutes, and you’ll know the single process bottleneck holding your trading back. Find your bottleneck →
Go deeper on the research. Our NQ Research Lab is a growing library of certified historical NQ futures studies, the honest results behind what actually holds up and what doesn’t. Explore the NQ Research Lab →
When you’re ready for the full system. Net Alpha Pro is our complete rules-based process for futures traders: the Risk, Edge & Psychology playbooks, the Trade Feedback Loop, monthly coaching with Glenn & Reid, and the full HTA indicator suite. $97/month, cancel anytime. See Net Alpha Pro →
No signals. No promises. Just the work, done right, at your own pace. Join the trading ohana when it’s your time.
Mahalo for reading and trade well! Glenn & Reid | Hawai’i Trading Academy