Mercury · evidence dossier

What survived a holdout that was locked before we looked.

A track record is only worth reading if the rules were fixed before the results were known. This page states the rules first, then the numbers, then the sleeves that did not make it, then the conditions under which the ones that did will be demoted.

← Back to the engine overview

Method, stated first

Three windows were fixed before any sleeve was scanned: the observation window, the discovery/holdout cutoff, and the timestamp at which open positions are marked. Nothing below was re-cut after the fact.

Cohort
Positions opened between 5 Aug 2026 and 19 Aug 2026, 07:00 UTC. Membership is decided by entry date, so fast closes cannot crowd the sample.
Holdout
Cutoff locked at 12 Aug 2026, 00:00 UTC before the first scan. Everything earlier is discovery; everything later is out of sample.
Metric
Return per position in percentage points of notional, size-weighted within the position. Closed positions use realised return; open positions are marked at the cutoff price.
Grain
Single-sleeve positions only. Positions carrying overlapping sleeve exposure are excluded from ranking, so no sleeve is credited with another's result.

Three tracks are three independently deployed instances of the engine, each on its own configuration pin and datastore. They share a code base, so they are correlated samples of the same engine under different parameterisation — replication across tracks rules out a configuration artifact, not a code-level one.

Discovery versus holdout

The only chart that matters here. A sleeve is interesting when the out-of-sample bar stands close to the in-sample bar; a tall discovery bar next to a short holdout bar is a warning, and is reported as such.

Discovery Holdout (out of sample) scale 0 → 1.2 pp
  • Sleeve A — short book, tactical

    Track II

    0.596
    0.592
  • Sleeve B — long book, regime

    Track I

    0.862
    0.718
  • Sleeve C — short book, structural

    Track III

    1.126
    0.336
  • Sleeve C — short book, structural

    Track I

    0.730
    0.400

Sleeve A holds its out-of-sample mean almost exactly. Sleeve B gives back roughly a sixth. Sleeve C decays hard on Track III and mildly on Track I — which is why it is presented as a two-track result rather than a Track III headline.

Promoted sleeves

Ranked on expectancy, sample size, interval width, holdout survival, and symbol breadth. Pockets with fewer than 20 positions, or with a failing holdout, are ineligible regardless of average.

SleeveTracknMean pp95% CISymbolsHHI
A — short book, tacticalII158+0.594[+0.34, +0.85]730.02
B — long book, regimeI84+0.751[+0.32, +1.21]730.02
C — short book, structuralIII78+0.782[+0.30, +1.33]540.02
C — short book, structuralI72+0.520[+0.06, +0.93]540.02

Strongest current position

Sleeve A on Track II. Out-of-sample mean is within 0.004 pp of in-sample — the closest thing to a non-result from the overfitting test that this dataset can produce. It clears on 158 positions across 73 symbols with a concentration index of 0.02, and the same sleeve prints positive on Track III (+0.367, n=115). On Track I it is still positive (+0.304) but its interval touches zero, so Track I is reported as unconfirmed rather than supporting.

Portfolio level, all sleeves

Every position in the cohort, promoted sleeves and rejected ones together. This is the number that describes the engine as deployed, and it is deliberately lower than the crowned sleeves — the difference is the cost of running research in production.

Track I

+0.35 pp

Positions
651
95% CI
[+0.22, +0.48]
Open, marked
16

Track II

+0.33 pp

Positions
856
95% CI
[+0.24, +0.42]
Open, marked
5

Track III

+0.39 pp

Positions
754
95% CI
[+0.28, +0.50]
Open, marked
24

All three intervals exclude zero. For completeness: grouped instead by close date, the same period settles 857, 1,095, and 931 positions for a summed +224, +282, and +284 pp respectively — a flattering cut that over-represents fast closes, which is exactly why it is not the ranking metric.

Where the money is actually made and lost

Each closed position records the path it exited through. Decomposing by exit path is how a sleeve's average is audited: it separates the mechanism that earns from the mechanism that bleeds, and it points at the next experiment rather than the next slide.

Sleeve A, Track II — both books

BookExit pathnMean
Short bookTrailing stop (profit harvest)122+0.95
Short bookThesis invalidation34-0.67
Long bookTrailing stop125+0.56
Long bookThesis invalidation42-0.69

Invalidation costs both books the same. The split between them is entirely in the exit geometry of the winners — which localises the edge to a design decision we control, not to a market accident.

Sleeve B, long book — Track I versus Track II

TrackExit pathnMean
ITarget reached15+2.70
ITrailing stop61+0.28
IITrailing stop84+0.80
IIThesis invalidation75-0.52

Same sleeve, two parameterisations. Track I lets a small tail of positions reach target and pays +2.70 pp on them; Track II invalidates roughly 45% of its closes and taxes the average away. The open work is to bring Track II's invalidation behaviour to Track I's, not to switch the sleeve off.

Cleared the bar, not promoted

These pockets passed the holdout and the interval test and were still held back. Publishing them is the point: a shortlist with no rejects behind it is a shortlist that was chosen after the fact.

PocketnMeanWhy it was held back
Sleeve D — long book, swing (Track III)25+0.85Sample too thin to carry a promotion: only 9 positions fall in the discovery half.
Sleeve B — short book (Track III)53+0.57Weaker than the long side of the same sleeve, and does not replicate on Track I, where its holdout fails.
Sleeve C — long book (Track II)48+0.47Real, but a smaller sample and the weaker sibling of a short book that replicates on two tracks.
Sleeve A — long book (Tracks II, III)170 / 184+0.25 / +0.22Positive and stable, but materially weaker than the short book on the same engine. Held, not crowned.

What this study refused to do

  • No sleeve was ranked on its full-period average alone; the holdout half had to stand by itself.
  • No holdout cutoff was chosen after seeing results — the date was fixed before the first scan.
  • No expectancy was computed on positions grouped by close date, which over-represents fast winners.
  • No sleeve was crowned because it beat a sibling; a contrast is mechanism evidence, never a ranking.
  • No positions were dropped for still being open — open risk is marked to market and carried into the mean.
  • No sleeve was retired on a single weak window without its written falsification condition being met.

How each result dies

Written before the next window opens, so a demotion is a scheduled outcome rather than an argument.

Sleeve A, Track II
On the next seven-day window, the holdout mean is at or below zero, or the confidence interval covers zero. Demote.
Sleeve B, Track I
The holdout mean turns non-positive on the next window, or the profit tail collapses onto a single symbol. Demote.
Sleeve C, Track I
The confidence interval covers zero on the next holdout at a sample of 40 or more. Demote.

Known limits of this cohort

  • Gross, not net. Fees, funding, and slippage are not deducted. At this per-position scale that is a material haircut and is modelled separately under a specific mandate.
  • Demo execution. Orders were placed on demo accounts against live venue data. Fill realism, not signal quality, is the open question this cohort cannot answer.
  • Fourteen days. One regime, one fortnight. Holdout survival is evidence against overfitting the window, not evidence of durability across regimes.
  • Correlated tracks. The three tracks share a code base. Cross-track replication rules out a parameterisation artifact; it does not rule out a shared implementation error.

Each of these is a scheduled piece of work, not a caveat added at the end. The next window extends the cohort, prices the fee drag explicitly, and puts the strongest sleeve on live capital at pilot size.

Next step

Position-level data, the full sleeve inventory including everything that failed, and the operating profile are available under a direct conversation — [email protected].

All trading involves risk of loss. Figures on this page are from demo-account execution against live venue data over a fourteen-day window, gross of fees and financing, and are not a representation of net client returns. Past results do not predict future performance.

kaido.team — one operator, a fleet of agents, under one flag.

the name