The Lens · What We Got Wrong

Our Data Was Lying for Five Months. Every Alarm Said It Was Fine.

We spent months building a trading strategy around an edge that did not exist. It was not a bad idea badly tested. The edge was manufactured by a defect in our own price data — two different versions of history mixed together, bar by bar, inside the same trading day. What makes this worth your time is not that we got it wrong. It is why nothing caught it: every integrity check we had said the data was healthy, and every one of them was right, because every single price in the file was a real price. They were just not all measured from the same place.

Dragonfly Lens · 11 September 2026 · Defect found 1 September 2026; full extent measured 10–11 September. Every figure below is from our own files.

The short version

1. The growth chart with a moving floor

Imagine measuring a child’s height every month and writing it on a chart. In March someone measures from the floor. In April, without telling anyone, they measure from the top of a 3-centimetre doorstep. In May, back to the floor. June, the doorstep again.

Every measurement is real. Nobody lied, nobody fat-fingered a number, and if you check any single entry against the child standing there it is correct. But the chart is nonsense. The growth spurts and shrinkages it shows are the doorstep, not the child. And here is the cruel part: the more carefully you analyse the chart, the more confident you become in a pattern that is purely an artefact of where the tape was held.

That is exactly what happened to our price data, and futures markets make it unusually easy to do.

Why futures data needs a “floor” at all

A futures contract expires. The one everybody trades today is replaced by next quarter’s, and the two trade at slightly different prices for perfectly ordinary reasons — storage, financing, time. So a chart of “the price of the S&P 500 future” going back ten years is not one contract. It is dozens, stitched together.

To stitch them without fake jumps at every changeover, data providers back-adjust: they shift all the older prices by whatever the gap was, so the seam disappears. It works well. But it means the historical prices in your file are not the prices anyone ever traded at — they have been shifted to make the line continuous. That is fine, and it is standard, as long as the whole file is shifted the same way.

Adjust the same history twice, at different times, with different contract gaps, and you get two versions of the past that are each internally perfect and about 2.4% apart. Put both in one table and you have the doorstep.

2. What it actually looked like

Here is the fingerprint. For every minute where we could compare our database against an independent vendor copy of the same contract, we took the ratio between them. If the data were clean, that ratio would be 1.0000 everywhere.

Ratio, our data vs the vendorShare of barsMeaning
exactly 1.000042.3%Correct version
about 1.02428.5%The other version, ~2.4% adrift
in between / other29.2%Boundaries and smaller mixes

Two clusters. That is the signature of two versions, not of noise — noise is smooth and centred, not lumpy at two specific values. And the mixing was not neatly separated by month or by year, which would have been survivable. Ninety-two of one hundred and five trading days contained both versions inside the same session.

Why that specific pattern is so destructive. A 2.4% step is roughly sixty times the size of a typical overnight move in these markets. So any calculation that happened to span one of those switches did not get a slightly wrong answer. It got an enormous fake move, in whichever direction the switch went. Strategies that look for sharp moves and bet on them reversing — which is precisely what we had built — were being fed a steady diet of manufactured sharp moves that reliably “reversed” when the data flipped back.

We had not discovered an edge. We had discovered our own data flipping between two rulers, and built a trading strategy that detected it.

3. Why every alarm stayed silent

We are not careless about this. We run integrity checks on every series before it is used. They look for the ways data normally goes wrong, and they all passed. Look at why:

The checkWhat it asksWhy it passed
Missing dataAre there gaps?No gaps. Every minute present.
Stale dataIs the price stuck?Prices moved normally all day.
Out of rangeAny impossible values?Every price plausible. They are real prices.
Bad bar shapeIs the high below the low?Every bar internally consistent.
DiscontinuityAny sudden jumps?0.06% of bars — below any sane threshold.

That last row is the one that stings. We had a jump detector, aimed at exactly this class of problem, and it saw almost nothing — because in percentage terms a 2.4% step between two adjacent minutes is unusual but not absurd. Markets do that. The detector was calibrated to catch a series that had been visibly broken in half, and this series had been quietly shuffled instead.

The general form of the problem. Every one of those checks examines the file on its own terms — is it internally sensible? A mixed-vintage series is internally sensible. Each number is real, each bar is well-formed, the whole thing looks like a market. It is wrong only in relation to something outside itself. No amount of staring at the file can reveal that, which is why five months of looking did not.

4. The verification that verified nothing

This is the part we find hardest to write, and the most useful to publish.

The one result we still believed in — a small, real effect, our best-evidenced idea — had been independently verified. We had a separate script, written later, whose entire job was to recompute it from scratch and check the answer. It reproduced the original figure exactly, to two decimal places. We recorded that as confirmation and moved on.

The verification script read the same database. It loaded the same contaminated numbers through a slightly different door and did the arithmetic again. Of course it agreed. It was never checking the data; it was checking our multiplication.

When we finally pointed it at a genuinely different source, it disagreed immediately — and the effect it had been “confirming” turned out to be 2.3 times smaller than we had been claiming.

The rule we now hold ourselves to. Reproducing a number from the same source verifies the calculation, not the data. A second opinion has to come from a second source. Anything else is agreement with yourself, wearing a lab coat.

5. What it cost

What we believedWhat it was
The core strategy had a strong statistical edgeOn clean data the same code has a negative expectation. It never worked.
We had a benchmark to measure live results againstThe benchmark was built from the same bad data. We were grading against a fiction.
A promising effect had faded and was downgradedThe “fade” was the defect. On clean data the effect held. Downgrade withdrawn.
Our best idea was strongReal, but 2.3x smaller than advertised.

Note the third row, because it cuts against the self-flattering version of this story. Bad data does not only invent things that are not there. It also destroys things that are. We had thrown away a genuine finding because corrupt numbers made it look dead. Contamination is not a bias in one direction; it is a fog.

The one consolation, and it is a real one: our live paper-trading results had been quietly disagreeing with the backtest for months. We had noticed and assumed the live execution was flawed. It was not. The ledger of things that actually happened was right the whole time, and the model was wrong. When a clean record and a clever model disagree, the record is usually not the thing to fix.

6. Where it came from

Databases usually record when each row was written. Ours did, and it made the forensics almost anticlimactic. Grouping every price bar by the date it was inserted:

Nobody did anything obviously wrong. Each backfill was a sensible response to a real gap. The mistake was structural: the table had no concept of where a row came from. Two versions of history were allowed to sit side by side with nothing recording that they were measured differently, because it had never occurred to anyone that they could be.

7. Zoom all the way out: this is not a trading problem

The specifics here are futures data, but the shape is completely general. It happens whenever numbers from more than one source, vintage, or definition end up in the same column — which is most spreadsheets that have existed for more than a year.

Three questions catch most of it, and none needs any technical skill.

Question 1: Could this file contain two different definitions of the same thing?

Prices adjusted two ways. Revenue before and after a policy change. Headcount including and excluding contractors. Temperatures from two instruments. The failure never announces itself, because both definitions produce believable numbers. If a column has been appended to, patched, or merged at any point, the answer is probably yes.

Question 2: Am I checking this against something, or just re-checking myself?

Recalculating, having a colleague review the formula, and running it again tomorrow all test the arithmetic. None of them can see a data defect. Only a genuinely separate source can — a different vendor, a different system, a physical count, a printed statement. If you cannot name the second source, you have not verified anything.

Question 3: When a result looks unusually good, do I check the data before I check the idea?

Our instinct was to admire the edge and then stress-test the strategy. The faster and more reliable move is the opposite: a surprisingly good result is a data question before it is a skill question. Genuine edges are usually small and awkward. Beautiful ones are usually a measurement artefact, and it costs ten minutes to rule that out before spending five months not ruling it out.

What we changed

Why publish this. Nobody would have known. The defect is invisible from outside, the conclusions it poisoned were never sold to anyone, and quietly fixing it would have been easy and comfortable. We are writing it down because a research operation that only reports its wins is not producing evidence, it is producing advertising — and because the single most valuable thing we own is not a strategy. It is knowing which of our numbers we are allowed to believe.

All figures from our own files, measured 1–11 September 2026. Share-of-bars figures are for the contract we examined in most depth; the pattern is present to differing degrees across the affected symbols. Nothing here is financial advice, and none of the affected conclusions was ever published as a subscriber signal.