What is actually holding back your trading? Get your free Process Score in 3 minutes.

Trading Discipline Lesson: Write the Test Before You Build the Tool

Third time running the same test. Same trade, same broker export, same expected P&L number sitting in a text file since before we wrote a single line of code. I (Reid) had already watched this exact test fail twice. Once because the tool couldn't find its own data file. Once because the output changed every time we ran it. I hit run a third time and didn't move until the number came back.

We were building a position-sizing guardrail tool with Claude. Something to catch us before we oversized a trade or broke our own rules. Round one of review found 14 real problems. A dead data file. Output that wasn't reproducible. A config file referencing settings that didn't actually exist anywhere in the code.

Not small stuff. The kind of stuff that, if it ships, quietly wrecks your risk management the day you need it most.

Fourteen problems is a lot. It's also, it turns out, the least interesting number in this whole story.

Why Does the Same Test Have to Run Every Time?

Here's what actually mattered. Across all four review rounds, we never changed the acceptance test. Same file. Same real trade, pulled from an actual day of trading, not a made-up scenario. Same exact expected number, written down before Claude touched the build.

That sounds small. It isn't.

If you change the test each time, you can't tell whether the tool improved or the test got easier. That's not a coding problem. That's a trading problem wearing a coding costume.

It's the same move as a trader who backtests a strategy, doesn't like the drawdown, and quietly loosens the stop until the equity curve looks better. You didn't build a better edge. You built a friendlier test.

Round one's 14 problems dropped to 4 in round two, then 2 in round three, then zero in round four. That shrinking pattern only means something because the target never moved.

What Happens When You Test Against Real Trading Data?

Round two is where it got uncomfortable. We ran the tool against one real day of our own trading and found 4 new problems. A bug that double-counted losses. And a broker "no data" placeholder that the tool read as a massive profit.

On paper, that second one would've told us we had a great day. We didn't. We had a data gap dressed up as a win.

That's the exact reason we don't trust a clean equity curve until we've stress-tested it against something real. It's the whole idea behind a proper TradeZella Backtesting Guide. You don't get to trust the summary stat until you've traced it back to the actual trades underneath it. A tool that can't survive contact with one real trading day has no business anywhere near your position sizing.

Round three found 2 smaller problems. A bot logging under two slightly different names in two different messages, which split what should've been one trade record into two. Small bug. Same discipline applied to catch it.

How Do You Know a Suspicious Zero Is Actually Right?

Round four came back clean. Zero problems. And that's exactly the moment to get suspicious, not relieved.

On that pass, the tool output "size: 0" for one specific test trade. Zero contracts. We didn't shrug and move on. We changed one input, used a smaller stop distance on that same trade, and confirmed the output correctly flipped to "1." That's the only way to know the zero was real math doing its job, not a silent failure dressed up as caution.

Any time your tool spits out a suspicious constant, perturb one input and confirm the output moves the way the math says it should. A flat zero. A perfectly round number. An output that never changes no matter what you feed it. If it doesn't move, you don't have a tool. You have a coin that always lands on the same side.

This is Psychology, the third leg of our REPs framework, showing up in a spreadsheet. It's tempting to accept the answer you were hoping for. Discipline is checking it anyway.

Can You Grade Your Own Homework?

By the last round, we ended up fixing the final small bugs ourselves instead of sending them back to another Claude session. That's fine. But only under one condition we're strict about.

It's only okay because the acceptance test was written before the build, on real data, with an exact expected number locked in advance. And because someone else reruns that same test before anything touches live capital. A fresh review. A second pass.

If the builder and the checker become the same person, the test carries the independence. You don't. Never grade homework you wrote the answer key for after doing the homework.

We talked through this whole build in more depth on Edge Up Podcast, Episode 077, "Using Claude and AI in Trading," on Spotify. Worth a listen if you're using AI as a research or coding partner and want to hear where it actually broke.

This isn't just a software lesson. It's Process Over Profits in its purest form. Your trading edge needs the same locked-in acceptance test. A defined setup, a defined outcome, checked against real trades, not adjusted after the fact to make the numbers look better.

That's the same discipline behind Positive Expectancy: Finding Your Trading Edge. You don't get to move the goalposts once you've already seen where the ball landed.

Write your exit rules before you're in the trade. Write your acceptance test before you build the tool. Same discipline, different spreadsheet. Skip it, and you're not testing anything. You're just hoping in a nicer font.

Want to trade with more structure and less guessing?

Everything we build at HTA starts with the same idea you just read: process over profits, risk before edge. If that’s the way you want to trade, here’s where to go next.

Start free. Take our free Trader Process Assessment. Thirteen questions, about three minutes, and you’ll know the single process bottleneck holding your trading back. Find your bottleneck →

Go deeper on the research. Our NQ Research Lab is a growing library of certified historical NQ futures studies, the honest results behind what actually holds up and what doesn’t. Explore the NQ Research Lab →

When you’re ready for the full system. Net Alpha Pro is our complete rules-based process for futures traders: the Risk, Edge & Psychology playbooks, the Trade Feedback Loop, monthly coaching with Glenn & Reid, and the full HTA indicator suite. $97/month, cancel anytime. See Net Alpha Pro →

No signals. No promises. Just the work, done right, at your own pace. Join the trading ohana when it’s your time.

Mahalo for reading and trade well! Glenn & Reid | Hawai’i Trading Academy