A trading statistic can be technically correct and still lead you to a bad decision. That is the problem with screenshots that say '72% win rate' and stop there. The number might be real, but without the sample, condition, distribution, and downside, you do not know what it actually means.
If the HTA Research Lab is going to be useful, traders need to know how to read NQ futures statistics without turning historical probabilities into predictions. Here is the framework we use.
Before you look at a result, define the event. 'Gap fill' sounds obvious until two traders use different closes, different opens, different sessions, and different thresholds. 'Trend day' is even worse. One person means close above open. Another means one-directional price action with shallow pullbacks. Those are different studies.
A clean research page should tell you exactly how the condition was measured. For NQ, session boundaries matter. So do the timeframe, whether data is RTH or Globex, and whether a threshold is measured in raw points, ATR, or a percentage of recent range.
Ten examples can teach you what a pattern looks like. They cannot tell you much about how stable the probability is. A larger sample does not automatically make a study good, but a tiny sample should make you cautious no matter how impressive the percentage looks.
This is why we prefer to show the eligible sample right next to the headline statistic. A 70% rate across 20 events and a 70% rate across 2,000 events should not create the same level of confidence.
Our existing backtesting process makes the same point from a strategy-development angle: define the rules, log enough repetitions to learn something, and avoid making a live-capital decision from a handful of pretty trades.
The mean is the arithmetic average. The median is the middle observation after results are sorted. NIST notes that different distributions can make different measures of location more useful. In trading, that matters because a few huge NQ sessions can drag an average far away from what a typical day looks like.
Imagine five retracements of 25, 28, 31, 34, and 182 points. The mean is 60 points. The median is 31. If you planned risk around the mean without seeing the distribution, you would be describing the outlier almost as much as the normal event.
That is why good NQ futures statistics should show more than one center point. When useful, we want the median, mean, and percentile range together.
Traders naturally ask, 'How far will it go?' Historical research cannot answer that for the next session. It can answer a better question: how far did comparable sessions go, and how were those outcomes distributed?
This is where research becomes risk management. A distribution gives you a range of plausible historical outcomes. It does not tell you which one arrives today.
Pooled averages are convenient, but they can hide the mechanism. Gap behavior may change with gap size. Session range may change with overnight volatility. A first-hour move may mean something different after a quiet Globex session than after a violent one.
One of the easiest ways to fool yourself is to discover a pattern in a pooled sample, then trade it without checking whether the result survives when the obvious conditioning variables are added.
Markets change. A result that looks excellent over 10 years can still be carried by one era. We care about whether the finding behaves similarly across earlier and later periods, and whether a holdout period tells the same basic story.
This is also why a failed out-of-sample result is useful. It saves you from promoting a historical coincidence into a trading rule.
This might be the most important sentence in the article: a statistically interesting market behavior is not automatically a trading edge.
A study can show that one state happens more often than another and still fail as an entry strategy once you add timing, stop placement, slippage, and costs. We have published the same principle in our positive expectancy work: the math has to survive the actual decision process, not just look good in a summary table.
Every public HTA study is designed to include an honest catch. Maybe the sample is smaller in one bucket. Maybe the result is stable for volatility but not direction. Maybe a popular claim simply fails. The caveat is not legal fine print. It is part of the result.
That standard matters in futures because leverage can punish overconfidence quickly. The CFTC advises traders to understand the market and its risks before acting on internet hype or unfamiliar information. We agree. Evidence should make you more precise, not more reckless.
That is the kind of statistical literacy we want the Research Lab to teach. For traders looking for Hawaii Trading Education, the goal should not be another guru telling you what NQ will do. It should be a better process for deciding what the evidence is strong enough to say.
Educational content only, not financial advice. Historical statistics are not forecasts or recommendations to trade.
Mahalo for reading and trade well! — Glenn & Reid | Hawai'i Trading Academy