You cannot select more than 25 topics
Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
43 lines
3.7 KiB
Markdown
43 lines
3.7 KiB
Markdown
# Strategy Improvement Leaderboard
|
|
_Updated: 2026-05-12T20:15:10.358786+00:00_
|
|
|
|
_Default view excludes retired legacy PEAD / short-core / exact-pocket families, retired pre-IMP-0606 research, and incomplete train-only scans. Use `fithia2 lb --include-retired` to inspect retired research._
|
|
|
|
_`SQS` below is public SQS v9: 3-pillar additive core (RQS 35% + WFQS 40% + Regime Adaptability 25%). Regime score uses scenario-test RRS when available, falls back to OOT quality, then 50.0 neutral. OOT bias toward COVID-era performance removed. CW blend and deployment gates unchanged from v8._
|
|
|
|
| # | ID | Experiment | SQS | RCW | [Tr]Ret% | [V]Ret% | [T]Ret% | [T]Ann% | [T]DD% | [T]Gross% | [T]DIM% | [T]R/G | Date |
|
|
|---|-----|-----------|-----|-----|----------|----------|----------|----------|--------|------------|----------|---------|------|
|
|
| 1 | 1357 | return_max_long_v7.364_composed_gld | 92.6 | - | +710.6 | +190.6 | +275.8 | +1954.2 | 8.8 | 69.2 | 80.0 | 3.99 | 2026-04-11 |
|
|
| 2 | 1348 | return_max_long_v7.356_composed_gld_compound | 80.0 | - | +703.7 | +189.0 | +276.8 | +1965.8 | 8.7 | 69.0 | 80.0 | 4.01 | 2026-04-11 |
|
|
| 3 | 1348 | return_max_long_v7.356_composed_gld | 77.4 | - | +419.5 | +152.4 | +229.3 | +1419.2 | 6.4 | 65.5 | 85.5 | 3.50 | 2026-04-10 |
|
|
| 4 | 1153 | return_max_long_v7.120_composed_gld | 77.3 | - | +606.6 | +113.3 | +112.9 | +473.5 | 7.4 | 56.9 | 86.2 | 1.98 | 2026-04-08 |
|
|
| 5 | 415 | return_max_long_v7.70 | 62.2 | 43.3 | +130.2 | +50.6 | +57.7 | +200.9 | 2.9 | 42.7 | 79.8 | 1.35 | 2026-03-28 |
|
|
|
|
## Recent Entries
|
|
### IMP-0959 (2026-05-12) — return_max_long_v9.3_v92_er30_xs10
|
|
Hypothesis: v9.2 PEAD + ER 30% + xsmom 10% = 3-way hybrid (Pareto-better in compound).
|
|
Verdict: **BETTER** (SQS None)
|
|
Reasoning: Compound 4y: ret=8319%/MDD=8.95%/Sharpe=3.65/SQS=85.9 vs v9.2 ret=8091%/MDD=9.78%/SQS=84.8. 3-split: train 1935%/SQS 86.8, valid 1326%/SQS 87.2, test 151%/SQS 74.1 (OOT softer).
|
|
Next: Sweep ER allocation 0-20% in fixed_capital mode to recover SQS; validate winners under compound.
|
|
|
|
### IMP-0955 (2026-04-11) — return_max_long_v7.364_composed_gld
|
|
Hypothesis: Scale ALL per_trade_risk_pct by 0.615x (0.65→0.40) to reduce DD while maintaining compound growth. Expected: train DD 9.46%→5.8%, test gross_exp 69%→42%.
|
|
Verdict: **NEUTRAL** (SQS 92.6, Stress SQS 76.6)
|
|
Reasoning: DD unchanged at 9.46% on train (DD is scale-invariant — scaling all positions proportionally doesn't change equity curve drawdown %). Test gross_exp still 69.2% (unchanged vs baseline 69.0%). Train gross_exp did drop to 42.7%, showing the scaling effect. Returns essentially identical across all splits (train +710.6% vs 703.7%, valid +190.6% vs 189.0%, test +275.8% vs 276.8%). Hypothesis failed: per_trade_risk_pct scaling alone cannot reduce DD.
|
|
Next: DD reduction requires structural changes: tighter stops, different max_holding_days, or smaller max_positions_per_sector — not just risk scaling. Alternatively, accept current DD level and focus on boosting test returns.
|
|
|
|
### IMP-0954 (2026-04-11) — return_max_long_v7.356_composed_gld_compound
|
|
Hypothesis: Auto-recorded via web GUI
|
|
Verdict: **UNKNOWN** (SQS 80.0)
|
|
|
|
### IMP-0951 (2026-04-10) — return_max_long_v7.356_composed_gld
|
|
Hypothesis: Auto-recorded via web GUI
|
|
Verdict: **UNKNOWN** (SQS 77.4, Stress SQS 78.5)
|
|
|
|
### IMP-0937 (2026-04-08) — return_max_long_v7.120_composed_gld
|
|
Hypothesis: Mega-cap material_contract engine ($100B+, 2-12% reaction day) captures AVGO-type strategic partnership events. Engine #9 reaction_max widened 6%->7%.
|
|
Verdict: **BETTER** (SQS 77.3, Stress SQS 74.8)
|
|
Reasoning: New engine fired 4 trades (COST+9%, DIS+11%, MS+6.5%, GOOGL breakeven). All orderly. Internal SQS 88.7, total return 4175% vs 3245% baseline. Mega-cap lane fills $100B+ gap.
|
|
Next: Compute v9 SQS and compare to 78.6. Promote if positive.
|
|
|