Appearance
Phase 2f: Strategy Search Campaign
Phase 2f ran the owner's brief — improve H1 with meaningful structural changes (no data farming), and search for a genuinely new strategy that beats H1's Sharpe, diversifies it (negative/zero correlation), or has verified expansion potential — as two agent-orchestrated search workflows plus one focused follow-up round. Eighteen pre-registered hypotheses went through the full ledger protocol; every promising result was adversarially audited (patch re-applied in a clean worktree, numbers reproduced, diff reviewed for look-ahead, the append-only run log checked for undisclosed variants).
How the search was run
- Pre-registration before data. Every candidate's rule, grid (≤ 9 combos, most ≤ 6), and fixed priors were written into its module docstring before the first real-data run. Grids were frozen after first OOS contact; one disclosed bug-fix mulligan allowed.
- One scoring path. Every run goes through
research.run_protocol— the ledger protocol (walk-forward grid re-fit per 504/126-bar window, stitched-OOS-only scoring, 1 tick/side slippage debited, moving-block bootstrap 90% Sharpe CI, same-span buy-and-hold, slippage and fill sensitivity) plus a Phase 2f addition: correlation to champion H1's own stitched OOS returns, the number the diversification criterion actually needs. Every invocation appends to an append-onlyruns.jsonlaudited after the fact. - Adversarial verification. Candidates claiming results were re-run from their exported patch by an independent auditor agent; every audited reproduction matched to four decimal places.
- Deep-history stress. Daily candidates that cleared their micro CIs were re-run on the ES/NQ e-mini parents (2012–2026 OOS, ~3× the micro span) — now possible offline via the committed
data/fallback/parquets (this environment's R2 token is read-only; see the storage-layer fallback added this phase).
Champion reference throughout: H1 OOS Sharpe 0.84 (MES) / 0.86 (MNQ), deep parents 0.74 (ES) / 0.58 (NQ), all CIs excluding zero.
Improve-H1 workflow — 8 structural candidates
| Candidate | Mechanism | MES / MNQ (ES / NQ) OOS Sharpe | Verdict |
|---|---|---|---|
i_stress_exit | VIX-latched exit: stress entries get prior-10-high target + 15-bar leash | 1.06 / 0.76 (0.77 / 0.60) | ⭐ finalist — note; registered, forward-validate on MES; not adopted (MNQ −0.10, best-of-10 selection risk) |
i_dip_provenance | Size 2× when the dip formed intraday vs overnight (hourly-derived, info H1 can't see) | 1.00 / 0.83 | ❌ shelved — MES active-stream alpha real (0.95 vs 0.84) but MNQ gain is pure leverage (active-stream 0.56 vs 0.86, maxDD −4.9%→−6.6%) |
i_volume_capitulation_size | 2× size on volume-spike capitulation days | 0.93 / 0.88 (0.71 / 0.51) | ❌ rejected — entire gain sits in the single final OOS window; ex-final-window it underperforms H1 on both micros; deep-history increment negative |
i_vol_target_exit | ATR-scaled profit target anchored at entry | 0.79 / 0.83 | ❌ CONFIRMED-negative — static vol targets are procyclically demanding (time-stops rose 26%→37%); H1's decaying prior-high exit is already implicitly vol-adaptive |
i_bear_short_fade | Add symmetric short side below the 200-SMA | 0.67 / 0.71 | ❌ shorts earned ~0 net at 1.2–1.4× the variance; the SMA lags bear→bull turns and fades genuine recoveries; shared grids let short noise poison long selection |
i_atr_dip_entry | Replace N-day-low breach with ATR-normalized depth trigger | 0.31 / 0.50 | ❌ selectivity starves the edge (35–48 trades); H1's edge is breadth of shallow dips |
i_bear_dip_long | Buy dips below the 200-SMA (faster vol-scaled exit) | 0.29 / 0.12 (0.24 / 0.40) | ❌ bear-regime dips have positive means but 1.6–1.8× the vol — risk-adjusted thinner everywhere |
i_ts_inversion_short | Short rallies only while VIX9D/VIX3M inverted | −0.19 / 0.10 (−0.22 / −0.11) | ❌ front-end vol clocks mark spike-stress (V-recovery days), not grind-bears; systematically sells the wrong day |
Improve verdict (round-2 judge, quoted): adopt no change; H1 stands. The below-200-SMA cell is closed on both sides; sizing overlays are attribution-trapped under walk-forward re-fit (quantity changes train-window Sharpe → parameter selection, so "sizing-only" deltas conflate mechanism with selection drift); exit engineering was the only axis that ever added value. Every same-book candidate landed MES-up/MNQ-flat-or-down — cross-symbol disagreement on a shared mechanism is the signature of noise around a real edge. Continuing would have been the small-improvement farming the brief prohibited.
New-strategy workflow — 8 mechanisms + 2 follow-ups
| Candidate | Mechanism | OOS Sharpe (corr to H1) | Verdict |
|---|---|---|---|
hourly_momentum (sleeve) | Long-only hourly Donchian — the untuned H5/H6 baseline promoted to a candidate | MES 0.47 (0.36) / MNQ 0.66, CI [0.14, 1.16] (0.28) | ⭐ finalist — MNQ 50/50 blend with H1: Sharpe 0.95 vs 0.86 H1-alone; registered as hourly_momentum; MES leg unverified (CI spans 0) |
n_vix_spike_reclaim | Buy first VIX down-tick after a ≥20% spike, ungated by trend | MES 0.51 CI [0.16, 1.05] (0.30) / MNQ 0.25 / ES 0.28 / NQ 0.22 | ❌ audit-CONFIRMED but mechanism inverted: all profit comes from above-SMA spikes (H1's own regime); below-SMA legs lose; blends are a wash |
n_xasset_tsmom | Donchian trend on MCL/QO/BZ (new markets) | MCL 0.45 (−0.28) / QO 0.22 (0.02) / BZ 0.02 (−0.08) | ❌ no CI clears zero; but the protocol's cold-start test windows structurally cannot express 100–200-day channels — slow TSMOM was not fairly falsified (see open questions) |
n_xasset_pullback | H1's exact machinery on MCL/QO/BZ | MCL 0.66 / QO 0.19 / BZ 0.06 | ❌ closes the twice-flagged generality question: H1's reversion does not transport off equity indices (QO/BZ flat over 12–14y; MCL's 0.66 is one hot half-year, and 4×-history Brent says the bet is empty) |
n_hourly_trend_breakout | Slower hourly channels (48–168h) + daily trend gate | 0.09 / −0.02 | ❌ bounds the momentum lead: the intraweek edge lives at ≤ ~1 CME day; slowing channels to cut fees kills the signal |
n_hourly_fast_symm | Fast symmetric hourly Donchian (long/short) + time stop | −0.32 / −0.23 | ❌ lost money even at zero slippage; hourly momentum is long-only-and-unstopped or nothing |
n_bear_rally_fade | Short/flat rally fade below the 200-SMA | −0.34 / 0.01 (0.00) | ❌ perfect complementarity (zero same-day overlap with H1 across 5.25y) — and no edge to put in the slot |
n_fomc_predrift | Long K days into FOMC decisions (daily bars) | 0.17 / 0.04 (ES 0.27 / NQ 0.27) | ❌ direction right (PF 1.36–1.47 over 16y of events), never significant — daily bars bundle the drift with the post-2pm reaction |
n_cot_flow | CFTC TFF positioning flow as directional signal | 0.10 / 0.08 (ES −0.02 / NQ −0.06) | ❌ COT on equity indexes now fully falsified (level as gate, flow as signal, micros + 14y parents) |
overnight_drift (final round) | Trend-gated overnight hold, 16:00 ET close → 10:00 ET close | MES 0.45 CI spans 0 (0.43) / MNQ 0.83, CI [0.33, 1.36] (0.46) | ⭐ finalist (MNQ) — the campaign's only CI-verified non-H1 edge: slippage-robust (0.87/0.83/0.78 at 0/1/2 ticks) after 2.5% fee drag; "always" won 25/26 train windows (the premium is unconditional). MES is fee-destroyed (gross 0.67 → net 0.45 → 0.22 at 2 ticks) — tick-to-notional geometry, not the anomaly, decides survival. 50/50 MNQ blend 0.99 vs 0.86, but the delta is not itself significant (P≈0.75) — paper-trade, never deploy the MES leg |
fomc_hourly (final round) | Long K days into FOMC, exit 13:00 ET strictly before the statement | MES 0.58 CI [−0.05, 1.18] (0.04) / MNQ 0.54 CI [−0.12, 1.12] (0.13) | ⭐ finalist (power-bound) — the most orthogonal stream found (corr-to-H1 ≈ 0), PF ~1.9, tiny drawdowns at 7% exposure, blends +0.10–0.14 on both symbols — but 42 events cannot exclude zero at Sharpe ~0.55. Isolating the pre-statement window roughly halved per-event noise vs the failed daily variant, vindicating its bounded conclusion. Accumulate events live/paper; ES/NQ parent hourly history would settle it |
What the campaign established
- H1 is a local optimum on its information set. Eight structural edits — new datapoints (hourly provenance, volume, VIX family), new sides, new entries, new exits — and none survives an adoption bar on both symbols. The two flagged finalists both need forward evidence, not more backtests.
- Vol information belongs in the exit. Three entry vetoes failed (2c); the stress-exit improved 3 of 4 spans without ever worsening risk. This is the campaign's one directional mechanism finding.
- The intraweek momentum sleeve is real on MNQ and is the portfolio's one verified diversifier. corr to H1 ≈ 0.3, own CI excludes zero, 50/50 blend +0.10 Sharpe. It was never tuned toward (it was the baseline), so selection bias is minimal — but the MES leg is unverified and fee drag (456–468 trades, ~1.3% of equity over the span) is the honest cost.
- The diversification slot exists but is hard to fill. A disjoint-regime sleeve is reachable (bear_rally_fade: zero overlap days) and finding low-correlation activity is easy; finding low-correlation edge is the entire problem. Every 50/50 blend except hourly-momentum-on-MNQ underperformed H1 alone.
- Closed doors (do not re-test without new data or a protocol change): below-200-SMA in either direction; COT on equity indexes; H1-style reversion on commodities; hourly reversion; hourly momentum outside the fast long-only band; daily-resolution event drift; static entry-anchored vol targets; sizing overlays under re-fit walk-forward.
Open questions the protocol itself raised
- Slow trend is untestable under cold-start windows. The 126-bar OOS windows give a 200-day channel zero live bars; commodity TSMOM at the literature's 6–12-month horizon was handicapped, not falsified. Warm-started (position-carrying) test windows are an owner-level protocol decision that would unlock the strongest remaining diversifier class (the QO sleeve made +5.3% through 2022 while H1 idled, directionally exactly the crisis-alpha shape).
- Portfolio capacity, not exit design, may be H1's real bottleneck: the stress-exit's MNQ loss came entirely from blocked re-entries during longer holds. A multi-cycle capacity study is a new hypothesis.
Infrastructure this phase leaves behind
research.run_protocol+ corr-to-champion scoring + append-only run log.- Storage-layer local fallback for read-only R2 tokens; committed deep-history parquets: ES/NQ (2010–), QO (2010–), BZ (2012–), MCL (2021–) under
data/fallback/. - MCL/QO/BZ
Marketentries with sourced constants (BZ margin flagged as estimate) and per-market Databento roll rules (roll_rule="c"for thin books whose volume ranking flip-flops). - New registered candidates:
pullback_stress_exit,hourly_momentum,overnight_drift,fomc_hourly; plus the cross-asset machinery modules (xasset_pullback.py,xasset_tsmom.py) for future commodity hypotheses.
The forward-testing stack (post-hoc, selection-biased — a hypothesis, not a claim)
The final judge's independent portfolio math from the OOS return CSVs (1,636 inner-joined days): the two-symbol H1 book scores 0.92; adding the FOMC sleeve on both symbols → ~1.03; adding the MNQ overnight sleeve → ~1.10; adding the MNQ hourly-momentum sleeve → ~1.14. No individual blend delta is itself statistically significant (block bootstrap: overnight-MNQ P(Δ>0)=0.75, FOMC-MES P=0.82), and the sleeves were chosen after seeing OOS — so this stack is the forward-testing agenda, with every increment owed its own live burn-in.
Recommended next steps (the final judge's ranked plan)
- Keep H1 as the sole live core on MES+MNQ (two-symbol book Sharpe 0.92); nothing in 18 candidates beat it or diversified it with significance. Build the live path (Phase 2e's recommendation), and run
pullback_stress_exitbeside it on MES as the cheap forward experiment. - Paper-trade the MNQ-only overnight sleeve (
overnight_drift) — the campaign's strongest non-H1 result. Blend point estimate 0.86 → 0.99, delta insignificant (P=0.75): let live paper data decide. Never deploy the MES leg (fee-destroyed). - Run the FOMC hourly sleeve live/paper on both symbols to accumulate events (~0.02%/yr fee cost, 7% exposure, orthogonal). If buying data, ES/NQ parent hourly history is the single highest-value purchase — ~14y of events would settle FOMC significance.
- Extend the pre-announcement-drift template to other scheduled macro events (CPI, NFP): same module pattern, public calendars, pre-registrable — the one untested hypothesis with a mechanism prior.
- Optional third sleeve: the untuned hourly-momentum Donchian on MNQ (+0.08–0.10 blend at corr 0.28), behind the two event/session sleeves.
- Owner call: warm-started test windows, then one slow-trend commodity candidate (QO/BZ/MCL, 100–200-day horizon) — the remaining untested mechanism class with a real diversification thesis.
- Do not fund further same-book improvement searches or any of the closed-door mechanisms above.