S-056 — statarb-expression-neutrality-vol-v1 (better hedge estimate + book-level vol targeting on the frozen S-038 signal)
The idea and the mechanism
S-038's worst stretch was not a crash and was not, on the evidence so far, a market event. Between 2021-07 and 2023-02 it lost 21.2% while the S&P fell only 7.5%; across the full 34-month episode it gave back book losing money in a rising market is a failure of expression, not of direction.
Two expression weaknesses are identifiable in advance:
1. The hedge ratio lags. Each name is hedged against a β-weighted combination of the nine sector ETFs, with the βs estimated on a rolling window. A rolling estimate is by construction late to a change in a name's factor loading, and the error does not cancel across names — it accumulates into book-level beta. This is not speculation about the mechanism: its sibling D-005 measured exactly this failure, where the momentum signal ordering worked as predicted and realised beta came out at +0.99 against a frozen |β| ≤ 0.15 gate. Twice now the signal has held while neutrality broke.
2. Gross exposure is constant. The book carries the same size into a high-volatility regime as into a calm one, so its realised risk is whatever the market hands it. Volatility clustering is among the most robust facts in this data, which makes predicted book volatility a far more tractable target than return. There is a second payoff specific to this account: S-048's margin buffer is a hard constraint, so a book that de-levers itself before a stress month is worth money independently of its Sharpe.
This is not an alpha claim. It proposes no new edge and takes no trades S-038 does not already take. Like S-048 it is a construction / risk-management claim, and it is labelled as one: the honest success statement is "the same signals, better hedged and better sized, with less drawdown per unit of return".
The frozen gate
deliberately left open — a gate written after seeing the data is not a gate.
| field | value |
|---|---|
| Hedge estimator (arm 1) | [CALIBRATE] — one named specification, frozen before the run |
| Vol-target specification (arm 2) | [CALIBRATE] — target vol, estimator, cap on gross, floor |
| Stage-1 bar, arm 1 | [CALIBRATE] — stated on realised |β| and its dispersion, not on return |
| Stage-1 bar, arm 2 | [CALIBRATE] — stated on max-drawdown and stressed buying power, with return as a constraint that must not fall |
| Cost standard | net of S-038's own ≤4 bps/side realised execution standard — an arm that wins gross and loses net is a fail |
| Baseline | the frozen S-038, on identical signals and dates. Never equal weight, never unhedged |
| Frozen | [CALIBRATE] — date the owner says "freeze" |
Two rules that are not calibratable and hold whatever the thresholds turn out to be:
- Each arm is judged alone, on its own objective. A combined run that improves on net Sharpe while realised |β| gets worse is a failure of arm 1, not a success of the pair.
- The family's sealed holdout is spent (D-010, 5/5 touches, "forward-only henceforth"). Stage 2 here is a walk-forward on in-sample data with purging and embargo; it is not a holdout confirmation, and the write-up must not call it one.
What we expect to find
Written before the run, and reported either way.
We expect arm 1 to be the one that pays, because there is a measured failure to correct rather than a hoped-for improvement, and because beta is an estimation problem with abundant data and persistent structure. We expect the improvement to show up as lower realised |β| and a shorter drawdown, not as higher return — if return rises materially that is a warning sign, not a result, because better hedging should cost a little return in exchange for less of someone else's risk.
We expect arm 2 to reduce max-drawdown and to cost some return, which is the trade it is supposed to make; the test is whether the drawdown reduction is larger than the return given up, net of the turnover that re-scaling creates.
We expect the most likely failure mode to be that the rolling hedge is already good enough, so arm 1 reproduces it with more variance and no gain. And we hold this lab's base rate in view: seven of eight discovery studies NULL, D-001 at probability of backtest overfitting (PBO) 0.93, D-011 at universe PBO 0.486 after ten names cleared Deflated Sharpe Ratio (DSR) > 0.95 under FDR control.
Methodology appendix — gates, exact parameters, look-ahead audit — is visible to subscribers. See the plans →