What is actually holding back your trading? Get your free Process Score in 3 minutes.

Passing Tests ≠ A Working Tool: Our AI Trading Tool Audit Checklist

"If the tests pass, the tool works." That's the assumption almost everyone makes about AI-built software. We made it too. For about twenty minutes.

Earlier this year we had Claude build us a position-sizing guardrail tool. Something to catch us before we sized a trade too big. The first version came back looking sharp. Every test passed. The documentation read like a senior engineer wrote it on a good day. Our gut said ship it.

We didn't. And that decision is the whole point of this post.

What Was A "Perfect" Build Actually Hiding?

I (Reid) run point on our AI builds, so I was the one staring at that first version, ready to call it done. Then we did what we tell every student to do with a new strategy before it touches real money. We audited it instead of trusting it.

What we found wasn't a small bug. It was three of them, stacked underneath a shiny surface.

The tool was reading from a dead data file. A source that no longer existed in the pipeline it was supposedly checking. Run it twice with the exact same inputs, and you'd get two different answers. And buried in the config, it referenced settings that were never actually defined anywhere in the code.

None of that showed up in testing. All of it would have shown up the first time we trusted this tool with real risk.

Why Doesn't Passing Tests Mean The Tool Works?

Here's the distinction that matters. Tests prove the code runs. They don't prove the tool works.

A test suite checks that a function returns something when you feed it the inputs the test expects. It doesn't know your file paths are wrong. It doesn't know your data source went dead last month. It doesn't know that "works in isolation" and "works in your real environment, with your real data, today" are two completely different claims.

Only checking a tool against your actual environment proves it works. That's not a Claude problem, and it's not an AI problem. It's the same reason we don't trust a backtest until we've walked it forward against live conditions. Polish is not proof. A confident-sounding output is not the same as a correct one.

This is where Psychology shows up in a coding session and not just a trading session. Psychology is the third leg of our REPs framework, alongside Risk Management and Edge. The discipline to slow down and verify, instead of relaxing because something looks finished, is the same muscle whether you're staring at a chart or a codebase. Skipping that step because a tool feels trustworthy is the exact same mistake as skipping your process because a trade feels obvious.

We talked through this whole build on Edge Up Podcast Episode 077, "Using Claude and AI in Trading" (Spotify). Worth a listen if you want the full blow-by-blow, including the parts that didn't make it into a tidy blog post.

The 4-Point Audit We Run On Every Claude Build

We built a checklist out of that experience. It's simple enough to run on anything Claude (or any AI) hands you, and it takes longer to read than it does to actually do.

1. Ask a fresh Claude conversation to attack it. Not the same thread that built the tool. A new one, with no investment in the answer. Prompt it directly: "Try to find ways this tool could silently give a wrong answer. Assume the author was rushed." A fresh conversation isn't defending its own work, so it's a lot more honest about where things break.

2. Check every file path, column name, and data source by hand. Does it actually exist on your machine, spelled exactly the way the code spells it? This is the check that would have caught our dead data file in about ten seconds. Instead it took a full audit to surface, because nobody had opened the file and looked.

3. Run it twice with the same inputs. Compare the output. If you get two different answers from identical inputs, stop right there and find out why before you trust another number that tool gives you. Reproducibility isn't optional in a tool that's going to sit near your risk management.

4. Feed it garbage on purpose. An empty file. A missing column. Numbers that make no sense. A tool that shrugs and returns a quiet, plausible-looking answer to bad input will shrug the same way at your money. A tool worth trusting should complain loudly.

Run those four checks on a Claude build before it goes near anything real. The same way you'd want a new strategy proven out before you'd size it up. It's the practical, unglamorous work. Same as journaling every trade in TradeZella instead of trusting your memory of how the week went, or reading the "Trading Risk Management Strategy: The Psychology Edge" post instead of assuming discipline will show up on its own.

One More Trap: The Tool That Never Says Anything

There's a sneakier version of this same problem we ran into later. A monitoring tool can be built to always output "nothing." Empty, zero, all-clear. It passes every test you throw at it, because an empty result is still a valid result. It just never actually tells you anything.

The question to ask of any monitoring or guardrail tool isn't "does it run clean." It's this: under today's real conditions, will this ever produce a non-trivial signal? If the honest answer is no, the fix isn't tweaking what it decides. It's redesigning what it records in the first place.

This is why we treat "process over profits" as more than a slogan. It's the whole idea behind our "Process Over Profits" post. A tool that looks busy but never actually flags anything is a process failure wearing a working-tool costume.

Trust The Checklist, Not The Checkmark

A green test suite feels like permission to stop looking. It isn't. It's the start of the audit, not the end of it.

We didn't build a position-sizing guardrail so we could relax the moment it compiled. We built it because risk management is the whole game. A broken guardrail is worse than no guardrail. It gives you false confidence right when you need real numbers.

Run the four checks. Every time. On every tool, Claude-built or otherwise, before it gets anywhere near your account.

Want to trade with more structure and less guessing?

Everything we build at HTA starts with the same idea you just read: process over profits, risk before edge. If that’s the way you want to trade, here’s where to go next.

Start free. Take our free Trader Process Assessment. Thirteen questions, about three minutes, and you’ll know the single process bottleneck holding your trading back. Find your bottleneck →

Go deeper on the research. Our NQ Research Lab is a growing library of certified historical NQ futures studies, the honest results behind what actually holds up and what doesn’t. Explore the NQ Research Lab →

When you’re ready for the full system. Net Alpha Pro is our complete rules-based process for futures traders: the Risk, Edge & Psychology playbooks, the Trade Feedback Loop, monthly coaching with Glenn & Reid, and the full HTA indicator suite. $97/month, cancel anytime. See Net Alpha Pro →

No signals. No promises. Just the work, done right, at your own pace. Join the trading ohana when it’s your time.

Mahalo for reading and trade well! Glenn & Reid | Hawai’i Trading Academy