You cannot select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.

41 lines
3.5 KiB
Markdown

# Strategy Improvement Leaderboard
_Updated: 2026-04-11T15:54:30.976974+00:00_
_Default view excludes retired legacy PEAD / short-core / exact-pocket families, retired pre-IMP-0606 research, and incomplete train-only scans. Use `fithia2 lb --include-retired` to inspect retired research._
_`SQS` below is public SQS v9: 3-pillar additive core (RQS 35% + WFQS 40% + Regime Adaptability 25%). Regime score uses scenario-test RRS when available, falls back to OOT quality, then 50.0 neutral. OOT bias toward COVID-era performance removed. CW blend and deployment gates unchanged from v8._
| # | ID | Experiment | SQS | RCW | [Tr]Ret% | [V]Ret% | [T]Ret% | [T]Ann% | [T]DD% | [T]Gross% | [T]DIM% | [T]R/G | Date |
|---|-----|-----------|-----|-----|----------|----------|----------|----------|--------|------------|----------|---------|------|
| 1 | 1357 | return_max_long_v7.364_composed_gld | 92.6 | - | +710.6 | +190.6 | +275.8 | +1954.2 | 8.8 | 69.2 | 80.0 | 3.99 | 2026-04-11 |
| 2 | 1351 | return_max_long_v7.356_composed_gld_compound | 80.0 | - | +703.7 | +189.0 | +276.8 | +1965.8 | 8.7 | 69.0 | 80.0 | 4.01 | 2026-04-11 |
| 3 | 1353 | return_max_long_v7.360_composed_gld | 77.4 | - | +421.7 | +153.4 | +228.3 | +1408.3 | 6.4 | 65.6 | 85.5 | 3.48 | 2026-04-10 |
| 4 | 1354 | return_max_long_v7.361_composed_gld | 77.4 | - | +421.7 | +153.4 | +228.3 | +1408.3 | 6.4 | 65.6 | 85.5 | 3.48 | 2026-04-10 |
| 5 | 1348 | return_max_long_v7.356_composed_gld | 77.4 | - | +419.5 | +152.4 | +229.3 | +1419.2 | 6.4 | 65.5 | 85.5 | 3.50 | 2026-04-10 |
| 6 | 1153 | return_max_long_v7.120_composed_gld | 77.3 | - | +606.6 | +113.3 | +112.9 | +473.5 | 7.4 | 56.9 | 86.2 | 1.98 | 2026-04-08 |
| 7 | 415 | return_max_long_v7.70 | 62.2 | 43.3 | +130.2 | +50.6 | +57.7 | +200.9 | 2.9 | 42.7 | 79.8 | 1.35 | 2026-03-28 |
## Recent Entries
### IMP-0955 (2026-04-11) — return_max_long_v7.364_composed_gld
Hypothesis: Scale ALL per_trade_risk_pct by 0.615x (0.65→0.40) to reduce DD while maintaining compound growth. Expected: train DD 9.46%→5.8%, test gross_exp 69%→42%.
Verdict: **NEUTRAL** (SQS 92.6, Stress SQS 76.6)
Reasoning: DD unchanged at 9.46% on train (DD is scale-invariant — scaling all positions proportionally doesn't change equity curve drawdown %). Test gross_exp still 69.2% (unchanged vs baseline 69.0%). Train gross_exp did drop to 42.7%, showing the scaling effect. Returns essentially identical across all splits (train +710.6% vs 703.7%, valid +190.6% vs 189.0%, test +275.8% vs 276.8%). Hypothesis failed: per_trade_risk_pct scaling alone cannot reduce DD.
Next: DD reduction requires structural changes: tighter stops, different max_holding_days, or smaller max_positions_per_sector — not just risk scaling. Alternatively, accept current DD level and focus on boosting test returns.
### IMP-0954 (2026-04-11) — return_max_long_v7.356_composed_gld_compound
Hypothesis: Auto-recorded via web GUI
Verdict: **UNKNOWN** (SQS 80.0)
### IMP-0953 (2026-04-10) — return_max_long_v7.361_composed_gld
Hypothesis: interleave_head_score only: test if it avoids OVERFIT scenario verdict vs v7.360
Verdict: **UNKNOWN** (SQS 77.4)
### IMP-0952 (2026-04-10) — return_max_long_v7.360_composed_gld
Hypothesis: interleave_head_score improves cross-engine capital allocation; rotation+recycle on guidance engine reduces capital lock-in
Verdict: **UNKNOWN** (SQS 77.4)
### IMP-0951 (2026-04-10) — return_max_long_v7.356_composed_gld
Hypothesis: Auto-recorded via web GUI
Verdict: **UNKNOWN** (SQS 77.4, Stress SQS 78.5)