7 Days in the Temple · Week 2
7 Days in the Temple — Issue 2: 36,300 Candidates, Zero Survivors
A full parameter sweep, a null control that embarrassed us, and the discovery that we had been counting our own trades wrong. The week the fees won.
By Ansel Grau · Reporting period: – · Updated:
Quick answer
In week 2 we swept 36,300 signal combinations. In-sample, 260 looked positive. Out-of-sample, zero survived. A permutation null control produced a better best result than the real sweep in 14 of 60 runs. The honest finding: what we found is indistinguishable from chance, and the binding constraint is cost per trade, not market direction.
Numbers of the week
Signals tested
signals_tested
Out-of-sample survivors
oos_survivors
Trades executed
trades_count
Win rate
win_rate_pct
All figures are aggregated and relative. They come from our lab’s weekly export — account balances, amounts and position sizes are structurally excluded from it. Raw data (JSON)
Week two. Eight trading days on the live path, five green, three red, 29 trades across seven markets on a single venue, relative performance at −0.13 percent. That is the surface. Below the surface, this was the week we ran the largest backtest sweep the Lab has attempted so far — and the week the sweep turned around and told us something we did not want to hear.
What we tested
The plan was simple to state and tedious to execute: take the signal family from issue 1, parameterize it properly, and grind through every combination the grid allows. That came to 36,300 labeled candidates. Each one got the same treatment — an in-sample window of 18 days, an out-of-sample window of 12 days, identical cost assumptions, identical data.
In parallel, we ran a null control, because a sweep without a null control is a machine for manufacturing confidence. We block-permuted the signals — random noise with the same temporal texture as the real thing — and ran the identical grid 60 times over.
And we rebuilt our own trade accounting from scratch, because something in the fill-level numbers had smelled wrong since the first issue. More on that below. It is not flattering.
What worked
Two things, and I want to be precise about what “worked” means here, because neither of them is a trading edge.
First: the directional calls held up. Six of the nine markets we traded were positive before costs. The win rate on the live path came in at 75.9 percent against a loss rate of 24.1 percent. If you only looked at direction, you would think the week was fine. It was not fine, and the reason it was not fine is the actual finding of this issue: the accounting eats what the direction earns. The binding constraint is the fee per trade, not the market view. We wrote about this mechanism at length in the fee truth, and this week the data confirmed it with unpleasant clarity.
Second: the quality gate did its job. Of everything that reached the final review stage, 8 candidates passed the formal quality checks. That number sounds like progress until you read the next section.
The graveyard
This is a long section this week. It has earned its length.
The sweep found fee avoidance, not edge. Across all 36,292 tradeable combinations, a single variable explains 98 percent of the spread in results — an r-squared of 0.981. That variable is trade frequency. The combinations that looked best were, almost without exception, the ones that traded least. We did not discover a strategy. We rediscovered the cost model, wearing 36,300 different costumes.
The in-sample winners were a mirage. 260 of 36,292 combinations were positive in-sample — 0.72 percent. After stripping out aliases, combinations that are labeled differently but produce effectively the same trades, 215 genuinely distinct candidates remained. Out-of-sample, not one survived as a defensible find. Zero. The word overfitting exists for exactly this shape of result.
The null control was the kill shot. The block-permuted random signals produced a better best result than the real sweep in 14 of 60 runs. The z-score of the real best candidate: 0.75. After multiple-testing correction across 24,560 effective tests, the p-value is 1.0. Not marginal. Not “needs more data.” The best thing our real sweep found is the kind of thing pure noise finds roughly a quarter of the time.
The grid lied about its own size. 36,300 labels produced only 24,560 distinguishable outcomes. Worse, the neighborhood argument — “the parameters next to the winner also look good, so it must be real” — collapsed on inspection. Neighboring candidates shared 93 percent of their trades. They are not independent confirmation. They are the same witness in a different hat.
We had been counting our own trades wrong. This one stings. What our records presented as 130 decisions was in fact 35. The count was keyed to fill IDs, not order IDs, and a single signal can shatter into up to 27 partial executions at the venue. Every statistic we had computed at fill level was quietly flattered by this. The correction is in, the historical numbers are re-stated in the data page, and the lesson is pinned to the wall: a metric is only as honest as its denominator.
The break-even math is structural, not unlucky. The current trade geometry requires an 88 percent hit rate just to break even after costs. That figure independently reproduces from pure market data at 83 to 86 percent — so it is not an artifact of our bookkeeping. We achieve 74 percent at signal level. The gap between what the geometry demands and what the signals deliver is not variance waiting to revert. It is a wall.
The numbers that remain
If you keep only two numbers from this issue, keep these.
Zero out-of-sample survivors from 36,300 candidates. Read it the right way around: this is not proof that no profitable strategy exists. It is proof that this grid, on this data, with these signals — one regime, 18 days in-sample, 12 days out-of-sample — contains nothing distinguishable from chance. The negative result is narrow. It is also, as far as it goes, solid, and a narrow solid result beats a wide soft one every week of the year.
An r-squared of 0.981 on trade frequency. When one variable explains 98 percent of a sweep’s outcomes and that variable is “how often you pay the fee,” the sweep has stopped searching for alpha and started ranking cost exposure. Any future sweep in this Lab gets checked against this failure mode first, before anyone gets excited about a leaderboard.
Market backdrop, for the record: the sentiment reading on 2026-07-19 sat in FEAR territory. Make of that what you will; we make nothing of a single reading.
Outlook
Three consequences for week 3, all of them boring and all of them earned.
The cost structure goes on trial before any new signal does. If the break-even hit rate demanded by the geometry sits 88 percent high while the signals deliver 74, the useful lever is the geometry — fewer, larger, cheaper decisions — not a fresh batch of candidates thrown at the same wall.
The null control becomes standard equipment. Sixty permutation runs cost compute and patience; they saved us from publishing 260 phantom winners. That trade we will take every time.
And the order-versus-fill accounting gets an automated check so it cannot regress silently. All raw counts, the corrected series, and the sweep artifacts are documented in the Lab, with prior issues in the weekly report archive.
A week that produces zero survivors and one corrected counting error is not a wasted week. It is a week in which the map got more honest. The territory was always going to look like this; we just could not see it before we tortured the numbers properly.
FAQ
- Does zero out-of-sample survivors mean no profitable strategy exists?
- No. It means that in this grid, on this data, with these signals, none was found — and that the best in-sample result cannot be distinguished from chance. That is a statement about our search, not about the market.
- Why did 260 in-sample winners all die out-of-sample?
- Because the sweep was mostly measuring one thing: how often a combination trades. A single variable explained 98 percent of the variance. After deduplication, 215 genuinely distinct candidates remained, and after multiple-testing correction across 24,560 effective tests, the corrected p-value was 1.0.
- What was the miscounted-trades error?
- Our count was keyed to fill IDs instead of order IDs. One signal can split into up to 27 partial executions, so what looked like 130 decisions was actually 35. Every fill-level statistic before the fix was flattered.
- What does the break-even hit-rate gap mean?
- The current trade geometry requires an 88 percent hit rate to break even after costs, a figure that independently reproduces at 83 to 86 percent from market data alone. We achieve 74 percent at signal level. That gap is structural — it will not close by waiting for better luck.
Past results are not an indicator of future performance. All figures are aggregated and relative, without base amounts. Not investment, tax or legal advice. Investing and trading carry substantial risk up to total loss. Do your own research and decide responsibly.