ID: S035
Slug: etf-momentum-rotation-12-1-v1
Run date: 2026-07-04
Failed at: Stage 1 (Quick-screen)
Outcome: FAILED — 4 of 5 gates fail.
Headline metric: strategy 5.8%/yr vs SPY 11.0%/yr → net excess −5.3%/yr (gate needs ≥ +2.0%)
Fail reason: Under honest methodology — survivorship-free universe, real friction, dividends, a point-in-time liquidity filter, and significance testing — the 12-1 ETF momentum rotation loses to simply holding SPY by ~5%/yr, with a worse drawdown. The cross-sectional signal has the right sign but is not significant. The edge the source study reported was an artifact of the defects this spec set out to correct.
What we tested
A confirmatory re-test of the 12-month-lookback / 1-month-hold cross-sectional ETF momentum rotation that a third-party course selected in-sample on a survivor-only 40-ETF list with zero commissions, no dividends, and no significance test. Rebuilt Mechaniq-standard: a survivorship-free point-in-time ETF universe (including delisted funds), a $25M median-dollar-volume liquidity filter, 20 bps round-trip friction, next-day adjusted-open execution, dividends, and block-bootstrap significance — judged against a SPY total-return benchmark over Dec 2005 → Mar 2026 on five pre-registered gates.
What we found
| Gate | Threshold | Result | |
|---|---|---|---|
| G1 — net excess return | ≥ +2.0%/yr vs SPY | −5.26%/yr | ✗ |
| G2 — bootstrap significance | excess 95% CI lower bound > 0 | −0.84% (LB) | ✗ |
| G3 — cross-sectional validity | top−bottom > 0, p < 0.05 | +0.62%/mo, p = 0.118 | ✗ |
| G4 — drawdown tail | ≤ 1.25× SPY max-DD (18%) | 23% | ✗ |
| G5 — CVaR-95 tail | ≤ 1.25× SPY CVaR (−13.2%) | −14.0% | ✓ |
Sample: 244 months, Dec 2005 → Mar 2026. Turnover ~30%/month. Leveraged/inverse funds excluded: 314. The pre-run forecast called this almost exactly — "much of the apparent alpha may be a beta-timing artifact that a SPY benchmark absorbs… the most likely failure points are G4/G5… the fully-invested, no-cash-filter design eats every benchmark drawdown plus rotation risk."
The universe question, answered
A frequent and fair worry: why thousands of ETFs, when the source used ~40? The answer is that the 6,049 is only the raw ingest pool (every US ETF ever classified, incl. ~1,900 later delisted) — needed to build the universe honestly. The tradeable count each month, after the $25M point-in-time liquidity filter, is far smaller: 32 ETFs at the first formation (Dec 2005) — right next to the source's 40, but survivorship-free — growing to a median of 152 and a max of 480 as the ETF market expanded (see the per-month chart). The source's hand-picked 40 survivors is exactly the bias this corrects: it silently excludes the funds momentum rotated into that later died, which flatters the backtest.
What we learned
The source's headline came from a stack of defects, each of which this spec removed:
- Survivorship bias — a survivor-only list hides the hot ETFs momentum buys that later delist. The survivorship-free universe here includes them, and it hurts.
- Zero friction — at ~30%/month turnover, real costs bite; but even gross, the strategy trails SPY.
- No significance test — the cross-sectional spread (G3) looks positive until you bootstrap it: p = 0.12, not real.
- Beta-timing artifact — 12-month momentum spent the sample concentrating into US large-cap growth/tech ETFs; against a SPY benchmark that absorbs it, there is no residual alpha, only extra drawdown.
What this doesn't tell us yet
The fail is robust across lookbacks — every 12-1-style window (4, 6, 9, 12 months) loses to SPY — so it is a stable result, not a knife-edge. What it does not test is a different object: a cash / absolute-momentum overlay that steps aside in downtrends would change the drawdown profile and is a separate, pre-registered S-NNN. This condemns fully-invested relative rotation over this sample, not every momentum expression.
What happens next
FAILED — cut at Stage 1, counts toward the swing pillar's C1 (cross-asset rotation) family (owner-assigned), reinforcing S031's finding: liquid-ETF rotation buys drawdown, not diversification, and loses to passive. A famous, heavily-marketed edge, failed transparently on honest gates — and a clean demonstration that the methodology corrections (survivorship, friction, PIT universe, significance) are the whole difference between the source's "success" and reality. No variant is queued; a cash/absolute-momentum overlay would change the object (a new S-NNN), and the bar is high given the base strategy loses to SPY across every lookback.
What we tested — the recipe
Slice & dice
For the specialist — methodology details (click to expand)
- Universe: survivorship-free point-in-time ETF set from a 6,049 raw ingest pool (incl. ~1,900 delisted); after the $25M PIT liquidity filter, 32 tradeable at the first formation (Dec 2005), rising to a median of 152 and a max of 480. Leveraged/inverse excluded (314).
- Execution & friction: next-day adjusted-open, 20 bps round-trip on turned-over notional (~30%/month turnover), dividends included.
- Significance: block bootstrap (mean block 6 months, 10,000 resamples); cross-sectional top−bottom +0.62%/month at p = 0.118 (not significant).
- Sample & baseline: 244 months, Dec 2005 → Mar 2026, SPY total return; counts toward swing-pillar family C1 (with S031).
Baseline SPY total return. Friction 20 bps round-trip on turned-over notional. Survivorship-free PIT ETF universe (Tiingo, incl. delisted). Block bootstrap: mean block 6 months, 10,000 resamples. English per repo convention. Ledger: swing-pillar-ledger.md (family C1).