← Lab
S048

S-048 — return-stacking-tbill-gated-v1 (T-bill base + market-neutral overlay + gated sector sleeve)

return-stacking-tbill-gated-v1
pre-registered 2026-07-13
Share: Twitter / X LinkedIn
Stacked equity curve — S048
236 months, including 2008 and 2022 · max drawdown 13.1% · 1.0x = start
1.0x7.2x200620082010201220142016201820202022202420262006-09
T-bill base + market-neutral book at 1.0x/1.0x + gated 9-sector sleeve at 50% of NAV, at the REAL historical 1-year T-bill rate (which averaged ~1.75%/yr over this window because it contains the zero-rate decade). Not a live track record.
-591.1%-257.0%+77.0%+411.1%+745.2% 350 Distribution of monthly returns — S048 monthly return of the stacked portfolio, % · dashed line = zero
gate
Distribution of monthly returns — S048

The idea and the mechanism

Sleeve Source track Native normalisation Risk driver
Overlay 1 — S-038 confirm.series()["C"] — monthly, vol-scaled to 10%/yr in-sample 10%/yr vol; owner sizes it 1.0× long / 1.0× short (cut from 1.5/1.5 for the buffer) equity cross-section / factor crowding (β≈−0.03)

AMENDMENT 2026-07-15 — 0DTE sleeve upgraded to the S-049 dual-fly. The 0DTE component is now the S-049 book (1 SPX + 5 QQQ) on the complete SPX (+missing-strikes backfill) ∩ QQQ data, not the SPX-only S-045 ×2. Effects: window extended to 2023-01→2026-04 (40 months, vs 21); S-038·0DTE correlation −0.03 (vs +0.21); the 0DTE is now a diversifier at ~25% risk-share (S-038 anchors ~78%), not co-dominant. Sizing is now N_BOOKS (units of the 1-SPX+5-QQQ book), default 1. Results in gates.md. The mentions of "S-045 ×2 / n_contracts / 2022-06→2025-12" below are the superseded original.

AMENDMENT 2026-09-23 — the 0DTE sleeve is replaced by a gated equity sleeve. The S-049 dual-fly FAILED the settle-bug audit, leaving overlay 2 empty. What fills it is the absolute-momentum gate, the one candidate this repo has measured on three asset classes (dual-momentum-gem/, sector-gate-overwrite/, swing-pillar.md §6a): in every case it roughly halves drawdown and costs return — it fixes path, not expectancy. That is precisely the trade a stacking sleeve should make.

Two inputs also had to be fixed before the study could answer anything. The window ran 2023-01→2026-04 (52 months) only because the dual-fly data started there; with that sleeve gone, tbill_dgs1.csv became the binding limit at 2022-01. The full DGS1 was pulled from FRED into tbill_dgs1_full.csv (the original is left untouched), opening the window to 236 months, 2006-09→2026-04, including 2008 and 2022 — i.e. the study can finally see the correction it is meant to survive.

Result at the declared 50% sizing (criteria as amended below):

ann vol Sharpe maxDD
base + S-038 +8.04% 8.04% 1.00 24.89%
+ gated 9 sectors +12.79% 9.40% 1.33 19.49%

Return up, drawdown down, corr(S-038, sleeve) = +0.02 over 236 months. (1a) PASS · (1b) PASS · (2) OK at 25% and 50%. A single gated SPY sleeve was tested alongside and is worse: it passes at 25% but fails (1b) at 50% and both at vol-matched, because one gate concentrates the on/off decision where nine sector gates spread it. Sizing above ~50% fails on drawdown — a real ceiling, not a preference.

On inheriting in-sample flattery — measured, and the earlier caveat was wrong. The concern was that 2006-2022 is S-038's in-sample period and would flatter the blend. It does not: S-038 scores Sharpe 0.79 / maxDD 27.67% over that window versus 1.25 / 7.85% on its sealed 2023+ holdout — the in-sample window is harder, because 2008 is in it. And criterion (1b) measures a difference (stack with the sleeve minus stack without), so any level bias sits in both terms and largely cancels. Independently: the sleeve's own stream is price-only with no S-038 input at all (Sharpe 0.94 / maxDD 12.66% over 236 months), and on S-038's clean holdout the sleeve still earns its place (base+S-038 1.76 Sharpe / 6.53% DD → 1.96 / 6.39% at 25%, 2.04 / 6.32% at 50%). Nothing here depends on the in-sample window.

What remains is ordinary and small: six configurations were reported (2 sleeves × 3 sizings), all of them including the three that fail, and the winner wins for a structural reason rather than a better number.

AMENDMENT 2026-09-23 (later the same day) — the table above is SUPERSEDED by the sized, margin-correct configuration. Two things changed after it was written: the buying-power charge was corrected from a per-leg ~10% to 14.99% of gross with zero netting, measured on a live account rather than looked up, and the account size was pinned to a real one. At the original 1.5/1.5 the stack no longer fits the buffer, so S-038 was cut to 1.0× long / 1.0× short; the owner then sized the sleeve at 50% of NAV, with the free-buffer target lowered from 50% to 25%. The binding numbers are now:

full window, 236 months ann vol Sharpe maxDD
base + S-038 (1.0×) +5.98% 5.38% 1.11 16.28%
+ gated 9 sectors at 50% NAV +10.65% 7.22% 1.44 13.12%

(1a) PASS · (1b) PASS · (2) 67% of stressed BP used, 33% free against the 25% target.

AND THE CAVEAT THAT GOES WITH IT, which is not a footnote. The backtest earns the real historical 1-year rate, which averaged only ~1.75%/yr across 2006-2026 because the window contains the ZIRP decade. The owner's actual fill is ~4%. That changes which sleeve sizes pass (1b), in the direction that matters: with a 1.75% base, base + S-038 draws down 16.28% and almost any sleeve appears to improve it; with a 4% base the same pair draws only 11.06%, and the sleeve must be genuinely small not to make it worse. The breakpoint sits between 35% and 40% of NAV:

sleeve backtest (1b) at 4% carry (1b) sealed holdout (1b)
25% PASS PASS PASS
35% PASS PASS PASS
40% PASS FAIL FAIL
50% (chosen) PASS FAIL FAIL

So the (1b) pass at 50% is an artifact of the interest-rate environment of 2009-2015, not a property of the sleeve, and must not be reported as a passed criterion. What 50% actually buys, at the carry the money will really sit in: Sharpe 1.54 → 1.76, max-drawdown 11.06% → 12.17%, return +8.39% → +13.16%. That is a deliberate exchange of 2.4pp of drawdown for 4.8pp of return — an owner risk decision, taken with the failure stated, not a criterion cleared.

A withdrawn claim. An earlier version of this amendment said "25% is the only sleeve size that passes (1b) on both windows." That was true only of the three pre-declared sizes (25%, 50%, vol-matched). A finer sweep shows 35% also passes on every reading and dominates 25% on every axis — more return (+9.25% vs +8.32%), higher Sharpe (1.43 vs 1.39) and lower drawdown (12.19% vs 12.27%). The line is withdrawn.

On choosing sizes after seeing results. 35%/40%/45% etc. were swept after the pre-declared 25/50/ vol-matched set, i.e. on the data. CLAUDE.md §2A permits this pre-freeze, but it is exploration, not evidence: the Sharpe plateau from 25% to 75% is flat (1.36–1.44), so the size should be picked on the buffer and the tolerable drawdown, not on the peak of a curve fitted in-sample.

You park all the money in one-year Treasury bills. They pay you about 4% a year and the broker still counts them as collateral, so you can run strategies on top of the same money instead of choosing between them. That is the whole idea of return stacking: the T-bills are the income, so no strategy has to manufacture it.

On top sit two engines that make money in unrelated ways. The first is the market-neutral book, which buys and short-sells roughly equal amounts and therefore does not care much whether the market rises or falls. The second is a simple trend rule across the nine US sector funds: hold a sector only while it has beaten cash over the past year, otherwise sit in cash for that slice. When markets fall apart it steps aside on its own.

Neither engine protects you by buying insurance, which is the point. Insurance bleeds money every month you do not need it. These two cost nothing to carry: one is roughly indifferent to market direction, the other simply gets out of the way. Together with the T-bill income they earned more, and fell less, than the two of them without the trend sleeve — over a window that includes both 2008 and 2022.

Why there is no "best mix", in plain words

The obvious next question is what is the ideal split between the two engines? It turns out the question has no useful answer here, for two reasons that are worth understanding rather than glossing over.

The usual measuring stick breaks. The standard way to rank a portfolio is Sharpe — return divided by how much it bounces around. But T-bills earn a return with essentially no bouncing, so by that measure they score near-perfectly. Ask a computer to find the mix with the highest Sharpe and it answers hold almost nothing but T-bills, every single time. It is like asking for the safest way to drive to work and being told not to go. The measure is not broken; the question is. So the split is a decision about how much risk you want to carry, not a sum with a right answer.

And the surface is flat anyway. Across sleeve sizes from 25% to 100% of the account, the quality score barely moves — 1.78, 1.79, 1.76, 1.70, 1.64, 1.57. The "best" one beats the runner-up by 0.01, which is noise. There is no peak to climb, only a long gentle slope downward as you add risk. Anyone reporting the top of that list as an optimum would be reporting a rounding error.

What can be answered. Two sensible, non-fitted questions do have answers. Are the two engines pulling their weight equally? Yes — they contribute 56% and 44% of the portfolio's risk, and they move independently of each other (correlation +0.02), which is about as close to a balanced pair as you get without engineering it. And is the money being spent efficiently? Not entirely: the market-neutral book returns about 0.14 per dollar of margin it consumes, the sector sleeve about 0.59 — roughly four times better, because the broker charges the market-neutral book for its long and short legs separately and gives no credit for the fact they offset. That is a real inefficiency. It is tolerated deliberately, because the market-neutral book is what holds up when the market falls: in the twelve worst months for the S&P, it lost 0.14% a month while the sector sleeve lost 1.64%. You are paying margin for protection, not for return.

What this is not: a promise, a frozen result, or anything running with real money. It is a construction study that says these pieces fit together on paper, at a size stated in advance.

The frozen gate

AMENDED 2026-09-23 on owner instruction. The spec is REGISTERED / PRE-REGISTRATION and explicitly NOT frozen, so amendment is permitted under CLAUDE.md §2A; a frozen gate would have required a new identifier instead. The original (1) was unpassable by construction and is superseded. It read: "combined Sharpe > max(single-sleeve Sharpe), AND combined max-drawdown ≤ the worst single sleeve's." The flaw: the T-bill base is riskless, so its Sharpe is degenerate — measured at 3.33 over the full 2006-2026 window and ~25 over the 2023+ holdout, purely because its volatility is near zero. No stack that contains its own collateral can ever exceed that, regardless of how good the sleeves are. The test was therefore rejecting every configuration for a reason that had nothing to do with stacking.

What we expect to find

Written before the sizing work, and reported below whether or not it held. We expected the gate to fix the path and not the expectancy — roughly halving drawdown while costing return, which is what this repo has measured on three separate asset classes. We expected the two engines to be close to uncorrelated, since one is a cross-sectional stock book and the other a sector-level trend rule. And we expected the binding constraint to be margin rather than returns: a long/short book is charged on its gross exposure, so the plumbing was always the more likely thing to fail. All three held. What we did not anticipate is that the interest-rate environment inside the backtest would decide which sleeve sizes pass the test — see the amendments.

Methodology appendix — gates, exact parameters, look-ahead audit — is visible to subscribers. See the plans →

← OlderSPX 0DTE at-the-money fly with tight scalp exits Newer →Spreading the 0DTE fly across SPX and QQQ

Mechaniq provides information, not investment advice. We do not execute trades. Past results are no guarantee for future performance. You are solely responsible for your trading decisions.

Lab · Methodology · Greeks Lab · About · FAQ · Glossary · Privacy · Terms · mechaniq.trade © 2026

24 Sep 2026, 19:45