7 Days in the Temple · Week 1
7 Days in the Temple — Week 1: 4.57 million calculations, zero findings
We ran a grid of 36,300 strategy variants across nine markets. 260 of them looked good. Not one survived. This is the most honest report we could write.
By Ansel Grau · Reporting period: – · Updated:
Quick answer
In week 1 we evaluated 4.57 million strategy combinations. 0.72 percent were profitable in the training window — and not a single one survived the independent test window. Random signals beat our real sweep in 14 out of 60 runs. The finding of the week is not a finding but a law: 98 percent of the variance in results is explained by one variable alone — how often you trade.
Numbers of the week
Combinations tested in the grid
lab_weekly.json · signals_tested
Out-of-sample survivors
lab_weekly.json · oos_survivors
Trading decisions (signal level)
lab_weekly.json · trades_count
Win rate
lab_weekly.json · win_rate_pct
All figures are aggregated and relative. They come from our lab’s weekly export — account balances, amounts and position sizes are structurally excluded from it. Raw data (JSON)
A sweep is a machine that manufactures hope. You hand it a grid of parameters, it computes, and at the end it places a winner on the table. It always places a winner on the table. That is its design principle, and it is why I trust no number that comes from a maximum. If you want to know how a backtest lies systematically, that is the long version — and every running metric lives in our Lab data hub.
This week we ran the machine. And then we audited it.
What we tested
The grid: eleven profit targets × eleven stop levels × five holding periods × six trailing variants × ten entry signals. That makes 36,300 combinations per market. Across nine markets and two cost assumptions: 653,400 runs — in 18.3 seconds. Plus the control calculation we will come to shortly: another 3.92 million. Total: roughly 4.57 million evaluations.
Before the first computation we split the data: 432 hours to search in, 288 hours to test against. The test portion was written to disk, checksummed, and removed from scope — and we recomputed that checksum independently at the end. The sweep never saw the test data. This is not a formality. It is the only reason the rest of this report is worth anything.
What worked
The method. Nothing else — and that is not false modesty.
One honest side-finding that surprised us: six of nine markets were positive before costs. The directional decisions of our live path hold up. They are, in fact, better than anything the sweep could generate. The problem is not in the seeing. It is in the settling.
And something else worked, in the uncomfortable way: we had been counting our own trades wrong. What looked like 130 decisions were 35. The counting basis was an execution ID, not an order ID — and a single signal shatters into as many as 27 partial fills. Every statistic we had built on that basis was cosmetically flattered: many small partial fills lift the win rate without contributing anything. As of this week we count at signal level. The 35 in the metrics box above is that new, less comfortable number.
The graveyard
Now the part that sweep reports normally leave out.
260 of 36,292 tradeable combinations were profitable in the search window. That is 0.72 percent. After removing look-alikes, 215 genuinely distinct ones remain. The median across the entire grid: deep in the red. 99.3 percent of the grid loses.
In the test window: zero survivors. Not “few”. None.
The top 20 from the search window added up to a clear profit there. In the test window the entire portfolio flipped sign. Two of the twenty were still nominally positive — but they sat at ranks where 13 to 20 percent of the entire grid did better. That is the definition of noise, not of selection.
Then the robustness test, and here it gets interesting. The rule is: a find whose neighbours lose is worthless. The neighbours do not lose. Five out of six are positive too — a plateau, not a lonely peak. That sounds like good news and is the worst news in this report: the neighbours share 93 percent of their trades. They are not independent witnesses, they are the same witness wearing a different hat. A plateau built from copies of the same trade is just as overfitted as a lonely peak — it merely looks more respectable.
And finally the null control, which caught the sweep red-handed. We replaced the real signals with 60 sets of random signals and ran the same grid. Result: pure noise produces a “best find” — of course it does, it always does. In 14 of 60 runs chance beat our real sweep. The best real find sits 0.75 standard deviations above what dice deliver. After correcting for the number of effective tests: p = 1.0.
Not one find reaches significance. Not even when treated as though it were the only hypothesis ever tested.
The grid also lies about its own size, incidentally: 36,300 labels produce only 24,560 distinguishable results. And 19 of the 20 best combinations are the same signal in a light disguise.
The two numbers that remain
First: 0.72 percent to 0. Of 36,292 combinations, 260 looked good in training. In the independent test, none survived. The distance between those two numbers is everything you need to know about backtest optimisation.
Second: 98 percent. Across the whole grid, a single variable explains 98 percent of the variance in all results — not the signal, not the profit target, not the holding period: how often you trade. Factor that law out, and what remains of the apparent transfer between search and test window is a correlation residual of 0.19. Practically nothing.
The sweep did not find an edge. It found fee avoidance and dressed it up as a strategy. In this grid, “optimising” meant: trade less. Why cost per trade dominates everything for us, we dissected in The Fee Truth — and the narrated version is the fee story from the lab.
Outlook
We are changing nothing about the strategy — because we have nothing to change. Any parameter set we pulled from this distribution would be an extreme-value artefact. That would be precisely the error this report dissects, only at smaller scale.
Instead, three things:
- Count cleanly. From now on at order level, with a cost flag per side. This week showed that we did not reliably know basic facts about our own system. We fix that first.
- Verify the cost finding. There is a reasoned suspicion that our resting exit orders are being charged more than necessary. The suspicion is inferred, not measured. It gets tested before anyone derives anything from it.
- Wait until n is enough. 35 decisions is not a sample, it is an anecdote with decimal places. Three weeks, then we recompute.
The brief was “the big sweep”. What came out was 4.57 million evaluations and zero robust findings. That is not a failed experiment. An experiment that kills a hypothesis has worked — it just does not feel like it.
Next week: week 2. Until then, everything else from the laboratory lives in the The Lab section.
FAQ
- Does “zero survivors” mean no profitable strategy exists?
- No. It means: in this grid, on this data, with these signals, none was found — and what was found is indistinguishable from chance. That is a statement about our test, not about the market.
- Why publish a failure?
- Because a sweep that finds nothing reveals more about the method than one that finds something. Show only hits and you are showing a selection. We show the distribution.
- What is a null control?
- We replace the real signals with random ones and run exactly the same grid. If chance finds “best strategies” as good as ours, our best find was not a discovery but an extreme value.
Past results are not an indicator of future performance. All figures are aggregated and relative, without base amounts. Not investment, tax or legal advice. Investing and trading carry substantial risk up to total loss. Do your own research and decide responsibly.