We spent months building a trading strategy around an edge that did not exist. It was not a bad idea badly tested. The edge was manufactured by a defect in our own price data — two different versions of history mixed together, bar by bar, inside the same trading day. What makes this worth your time is not that we got it wrong. It is why nothing caught it: every integrity check we had said the data was healthy, and every one of them was right, because every single price in the file was a real price. They were just not all measured from the same place.
Imagine measuring a child’s height every month and writing it on a chart. In March someone measures from the floor. In April, without telling anyone, they measure from the top of a 3-centimetre doorstep. In May, back to the floor. June, the doorstep again.
Every measurement is real. Nobody lied, nobody fat-fingered a number, and if you check any single entry against the child standing there it is correct. But the chart is nonsense. The growth spurts and shrinkages it shows are the doorstep, not the child. And here is the cruel part: the more carefully you analyse the chart, the more confident you become in a pattern that is purely an artefact of where the tape was held.
That is exactly what happened to our price data, and futures markets make it unusually easy to do.
A futures contract expires. The one everybody trades today is replaced by next quarter’s, and the two trade at slightly different prices for perfectly ordinary reasons — storage, financing, time. So a chart of “the price of the S&P 500 future” going back ten years is not one contract. It is dozens, stitched together.
To stitch them without fake jumps at every changeover, data providers back-adjust: they shift all the older prices by whatever the gap was, so the seam disappears. It works well. But it means the historical prices in your file are not the prices anyone ever traded at — they have been shifted to make the line continuous. That is fine, and it is standard, as long as the whole file is shifted the same way.
Adjust the same history twice, at different times, with different contract gaps, and you get two versions of the past that are each internally perfect and about 2.4% apart. Put both in one table and you have the doorstep.
Here is the fingerprint. For every minute where we could compare our database against an independent vendor copy of the same contract, we took the ratio between them. If the data were clean, that ratio would be 1.0000 everywhere.
| Ratio, our data vs the vendor | Share of bars | Meaning |
|---|---|---|
| exactly 1.0000 | 42.3% | Correct version |
| about 1.024 | 28.5% | The other version, ~2.4% adrift |
| in between / other | 29.2% | Boundaries and smaller mixes |
Two clusters. That is the signature of two versions, not of noise — noise is smooth and centred, not lumpy at two specific values. And the mixing was not neatly separated by month or by year, which would have been survivable. Ninety-two of one hundred and five trading days contained both versions inside the same session.
We had not discovered an edge. We had discovered our own data flipping between two rulers, and built a trading strategy that detected it.
We are not careless about this. We run integrity checks on every series before it is used. They look for the ways data normally goes wrong, and they all passed. Look at why:
| The check | What it asks | Why it passed |
|---|---|---|
| Missing data | Are there gaps? | No gaps. Every minute present. |
| Stale data | Is the price stuck? | Prices moved normally all day. |
| Out of range | Any impossible values? | Every price plausible. They are real prices. |
| Bad bar shape | Is the high below the low? | Every bar internally consistent. |
| Discontinuity | Any sudden jumps? | 0.06% of bars — below any sane threshold. |
That last row is the one that stings. We had a jump detector, aimed at exactly this class of problem, and it saw almost nothing — because in percentage terms a 2.4% step between two adjacent minutes is unusual but not absurd. Markets do that. The detector was calibrated to catch a series that had been visibly broken in half, and this series had been quietly shuffled instead.
This is the part we find hardest to write, and the most useful to publish.
The one result we still believed in — a small, real effect, our best-evidenced idea — had been independently verified. We had a separate script, written later, whose entire job was to recompute it from scratch and check the answer. It reproduced the original figure exactly, to two decimal places. We recorded that as confirmation and moved on.
The verification script read the same database. It loaded the same contaminated numbers through a slightly different door and did the arithmetic again. Of course it agreed. It was never checking the data; it was checking our multiplication.
When we finally pointed it at a genuinely different source, it disagreed immediately — and the effect it had been “confirming” turned out to be 2.3 times smaller than we had been claiming.
| What we believed | What it was |
|---|---|
| The core strategy had a strong statistical edge | On clean data the same code has a negative expectation. It never worked. |
| We had a benchmark to measure live results against | The benchmark was built from the same bad data. We were grading against a fiction. |
| A promising effect had faded and was downgraded | The “fade” was the defect. On clean data the effect held. Downgrade withdrawn. |
| Our best idea was strong | Real, but 2.3x smaller than advertised. |
Note the third row, because it cuts against the self-flattering version of this story. Bad data does not only invent things that are not there. It also destroys things that are. We had thrown away a genuine finding because corrupt numbers made it look dead. Contamination is not a bias in one direction; it is a fog.
The one consolation, and it is a real one: our live paper-trading results had been quietly disagreeing with the backtest for months. We had noticed and assumed the live execution was flawed. It was not. The ledger of things that actually happened was right the whole time, and the model was wrong. When a clean record and a clever model disagree, the record is usually not the thing to fix.
Databases usually record when each row was written. Ours did, and it made the forensics almost anticlimactic. Grouping every price bar by the date it was inserted:
Nobody did anything obviously wrong. Each backfill was a sensible response to a real gap. The mistake was structural: the table had no concept of where a row came from. Two versions of history were allowed to sit side by side with nothing recording that they were measured differently, because it had never occurred to anyone that they could be.
The specifics here are futures data, but the shape is completely general. It happens whenever numbers from more than one source, vintage, or definition end up in the same column — which is most spreadsheets that have existed for more than a year.
Three questions catch most of it, and none needs any technical skill.
Prices adjusted two ways. Revenue before and after a policy change. Headcount including and excluding contractors. Temperatures from two instruments. The failure never announces itself, because both definitions produce believable numbers. If a column has been appended to, patched, or merged at any point, the answer is probably yes.
Recalculating, having a colleague review the formula, and running it again tomorrow all test the arithmetic. None of them can see a data defect. Only a genuinely separate source can — a different vendor, a different system, a physical count, a printed statement. If you cannot name the second source, you have not verified anything.
Our instinct was to admire the edge and then stress-test the strategy. The faster and more reliable move is the opposite: a surprisingly good result is a data question before it is a skill question. Genuine edges are usually small and awkward. Beautiful ones are usually a measurement artefact, and it costs ten minutes to rule that out before spending five months not ruling it out.
All figures from our own files, measured 1–11 September 2026. Share-of-bars figures are for the contract we examined in most depth; the pattern is present to differing degrees across the affected symbols. Nothing here is financial advice, and none of the affected conclusions was ever published as a subscriber signal.