feat: PEAD mid-cap strategy + pipeline hardening + README cleanup

- Implement PEAD 7% Long+Short strategy with mid-cap universe expansion
- Add Stock Oracle screener/company clients, text sentiment features
- Enhance backtest engine: short-side execution, walk-forward CV, MFE/MAE analysis
- Harden pipeline: sequential Oracle API calls, scoring recalibration (event_quality 65%)
- Add experiment configs for 60+ strategy variants and journal tracking
- Add review/analysis CLI tools
- Remove obsolete dev/phase0-4 design documents and analysis scripts
- Clean README to reflect only implemented features (remove unbuilt adapters/engines)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
main
I Luk Kim 7 months ago
parent 764493dbe1
commit 2395a0c0c3

@ -0,0 +1,11 @@
---
name: feedback_stock_oracle_data
description: All data must come through Stock Oracle API. If Oracle doesn't provide something, report what's missing instead of using alternatives.
type: feedback
---
All data must be fetched through the Stock Oracle API. Do not use alternative data sources (direct SEC EDGAR, Yahoo Finance, web scraping, etc.).
**Why:** Stock Oracle is the single source of truth for this project. Using other sources creates inconsistencies and bypasses the project's data infrastructure.
**How to apply:** When a data need arises, check Stock Oracle API first. If Oracle doesn't provide it, tell the user what's missing rather than silently using an alternative source.

@ -0,0 +1,39 @@
---
name: PEAD Strategy Results
description: Pure PEAD 7% strategy is the first configuration to show consistent OOS profitability across all temporal splits
type: project
---
## PEAD 7% Base — Production Candidate (2026-03-14)
First strategy to show positive expectancy out-of-sample on all splits.
### Signal
- Entry: earnings reaction_day_return >= 7%, volume_ratio_20d >= 1.5x
- Scoring model: `pead` in `libs/backtest/scoring.py` (`compute_pead_score()`)
- Earnings-only, all other event types disabled
- Filters disabled: veto_oneoff_penalty=1.0, veto_unknown_direction=false, veto_bearish_direction=false
### Combined Results (155 trades across train/test/valid)
- Win Rate: 67.1% (stable 64.7-67.5% across all splits)
- Avg Win: +5.45%, Avg Loss: -3.48%
- Mean PnL per trade: +2.51%
- Median PnL: +3.73%
- W/L Ratio: 1.57
- PF: 1.31 (test), 1.95 (valid), 1.68 (train)
### Key Files
- Manifests: `configs/experiments/pead_pure_7pct.json`, `pead_7pct_drift_a.json`
- Scoring: `libs/backtest/scoring.py` — `compute_pead_score()`
- Domain: `libs/backtest/domain.py` — `SignalConfig.scoring_model`, `pead_reaction_threshold`, `pead_volume_threshold`
- Store wiring: `apps/backtester/run.py` — `_build_store()` dispatches to PEAD scoring
### Drift A Variant (wider target 3.0 ATR, hold 25d)
- Higher per-trade returns (+7.13% avg win) but lower win rate (56.1%)
- Mean PnL nearly identical (+2.59%) but median much lower (+0.40%)
- OOS less stable (valid WR dropped to 40%)
- Not recommended for production
**Why:** Previous strategies (baseline 3-factor scoring, small/mid-cap variants, execution optimization) all failed OOS. PEAD 7% succeeds because it relies on a well-documented academic anomaly with a simple, hard-to-overfit signal.
**How to apply:** Use PEAD 7% Base as the foundation. Any future strategy modifications should be compared against this baseline. The signal is the edge — execution/risk parameters have limited room for improvement (5 execution variants tested, none beat baseline OOS).

@ -1,322 +1,770 @@
# ACE-F v1 — US Stock Event Swing Trading System
# ACE-F v1 — AI Catalyst Event Engine (Free Data)
AI-powered event-driven swing trading system that uses SEC filings and free market data
to identify 1-5 day continuation trades in US equities.
미국 주식 이벤트 기반 중단기 자동매매 시스템.
SEC 공시 + 무료 시장 데이터를 활용하여 1~5일 continuation 종목을 자동 선별하고,
다음 세션에 기계적으로 진입·청산하는 AI 이벤트 트레이딩 엔진.
**Core principle:** AI/LLM is a document interpreter, not a price predictor.
Entry signals come from official filings + price confirmation, never from social data alone.
---
## 목차
1. [프로젝트 목적](#1-프로젝트-목적)
2. [핵심 철학](#2-핵심-철학)
3. [시스템 아키텍처](#3-시스템-아키텍처)
4. [디렉토리 구조](#4-디렉토리-구조)
5. [데이터 파이프라인](#5-데이터-파이프라인)
6. [전략 엔진](#6-전략-엔진)
7. [백테스트 시스템](#7-백테스트-시스템)
8. [전략 개선 시스템 (SQS / Journal / Leaderboard)](#8-전략-개선-시스템)
9. [현재 개발 현황](#9-현재-개발-현황)
10. [현재 최고 전략 성과](#10-현재-최고-전략-성과)
11. [설치 및 실행](#11-설치-및-실행)
12. [사용법](#12-사용법)
13. [기술 스택](#13-기술-스택)
14. [데이터 소스](#14-데이터-소스)
---
## 1. 프로젝트 목적
**"공식 문서와 무료 attention 데이터를 이용해, 1~5일짜리 중단기 continuation 종목을 자동으로 선별하고, 다음 세션에 기계적으로 진입·청산하는 AI 이벤트 트레이딩 시스템"**
핵심 목표:
- **무료 데이터만** 사용하여 미국 주식의 1~5거래일 이벤트 드리프트를 자동 탐지
- LLM/AI는 가격 예측기가 아니라 **공시·문서 해석기**로 사용
- 전략 중심: **공식 이벤트 + 가격 반응 확인 + 리스크 통제**
- 초기 버전은 수익률 최대화보다 **재현성, 운영 안정성, 확장 가능성** 우선
---
## 2. 핵심 철학
```
SEC Filings → Document Parser → Feature Builder → Signal Ranker
↓ ↓
Price/Volume Confirmation ←──── Backtest Engine ←── Risk Engine
↓ ↓
Attention Overlay (optional) ──→ Execution Engine → Post-trade Review
AI/LLM = 문서 해석기, ≠ 가격 예측기
```
## Architecture
1. **공식 이벤트가 중심** — SEC 8-K, 10-Q, 6-K 등 기업이 직접 배포한 공시가 신호의 원천
2. **가격이 반드시 1차 검증** — 이벤트 발생 후 첫 정규장 반응(reaction day)이 강하게 확인된 종목만 후보
3. **재현성 최우선** — 무료 데이터 환경에서 가장 재현성 높은 데이터(SEC)를 중심에 배치
---
## 3. 시스템 아키텍처
### 전체 파이프라인
```
apps/ # Application entry points
├── backtester/ # Event-driven backtest simulation & CLI
├── pipeline/ # Multi-stage data processing
│ ├── filing_poller/ # Poll SEC EDGAR for new filings
│ ├── filing_fetcher/ # Download filing documents
│ ├── event_parser/ # Parse events from filings (rule + LLM)
│ ├── feature_builder/ # Generate scoring features
│ ├── label_generator/ # Create forward-return labels
│ └── dataset_export/ # Export Parquet snapshots for backtesting
├── sync/ # Data synchronization
│ ├── issuer_sync/ # Company metadata from Stock Oracle
│ ├── macro_sync/ # FRED macro indicators
│ └── short_volume_sync/ # FINRA short sale volume
├── tracker/ # Strategy improvement tracking CLI
├── tools/ # Analysis & utility scripts
├── review/ # Manual review queue
└── qa/ # Data quality checks
libs/ # Core libraries
├── backtest/ # Backtesting engine
│ ├── domain.py # Pydantic domain models (30+)
│ ├── tracker.py # SQS scoring, journal I/O, leaderboard
│ ├── execution.py # Entry/exit simulation
│ ├── allocator.py # Position sizing & entry gates
│ ├── scoring.py # Candidate scoring (PEAD, composite)
│ ├── selector.py # Candidate filtering & ranking
│ ├── metrics.py # 21-metric performance bundle + bootstrap CIs
│ ├── artifacts.py # Run output writer (Parquet, JSON, CSV)
│ ├── manifests.py # Experiment config resolution
│ ├── snapshot_store.py # Parquet data loader
│ ├── splits.py # Walk-forward window generation
│ └── calendar.py # Trading day utilities
├── common/ # Logging, config, time utils
├── db/ # PostgreSQL models (async SQLAlchemy)
├── oracle_client/ # Stock Oracle API client
├── parser/ # Filing document parser
├── features/ # Feature engineering
├── labeler/ # Label generation
├── schemas/ # Shared data schemas
├── export/ # Snapshot export
├── review/ # Review logic
└── llm/ # LLM integration layer
configs/
├── backtest/ # Base backtest configs (defaults.json)
├── experiments/ # 66 experiment manifests
├── app.yaml # Application settings
└── symbols_*.yaml # Asset universe definitions
data/ # Data storage (parquet snapshots, cache)
runs/ # Backtest execution outputs
journal/ # Strategy improvement journal & leaderboard
tests/
├── unit/ # Unit tests (270+)
├── integration/ # Integration tests (requires PostgreSQL)
└── replay/ # Determinism replay tests
[SEC EDGAR / Stock Oracle / FRED / FINRA]
↓
Source Adapters
↓
Raw Storage (원문 보관)
↓
Normalizer / Document Parser
(규칙 기반 + LLM 보강)
↓
Feature Store (Parquet 스냅샷)
↓
Signal Ranker
(PEAD / Composite 스코어링)
↓
Portfolio & Risk Engine
(포지션 사이징, 섹터 제한, 진입 게이트)
↓
Backtest Engine / Execution
(이벤트 드리븐 시뮬레이션, 슬리피지)
↓
Strategy Improvement Tracker
(SQS, Journal, Leaderboard)
```
## Setup
### 3대 핵심 엔진
**Requirements:** Python 3.11+, PostgreSQL 16 (via Docker)
| 엔진 | 역할 | 입력 | 출력 |
|------|------|------|------|
| **Event Engine** | 공식 문서에서 이벤트 후보 생성 | SEC 8-K, 10-Q, 6-K | event_type, direction, confidence |
| **Document Understanding** | 규칙+LLM으로 문서 질 해석 | exhibit text, XBRL | guidance_direction, demand_strength, margin_quality, oneoff_flags |
| **Market Confirmation** | 시장의 실제 반응 확인 | OHLCV bars | reaction_return, volume_ratio, close_location, gap_size |
```bash
# Install dependencies
pip install -e ".[dev]"
---
# Start PostgreSQL
docker compose up -d
## 4. 디렉토리 구조
# Run database migrations
alembic upgrade head
# Copy and configure environment
cp .env.example .env
```
fithia2/
├── apps/ # 애플리케이션 엔트리포인트
│ ├── backtester/ # 이벤트 드리븐 백테스트 시뮬레이션 & CLI
│ │ ├── run.py # 메인 백테스트 실행기 (BacktestRunner)
│ │ └── replay.py # 결정론 검증 (replay test)
│ ├── pipeline/ # 멀티스테이지 데이터 처리 파이프라인
│ │ ├── filing_poller/ # SEC EDGAR 신규 공시 폴링
│ │ ├── filing_fetcher/ # 공시 문서 다운로드
│ │ ├── event_parser/ # 이벤트 파싱 (규칙 + LLM)
│ │ ├── feature_builder/ # 스코어링 피처 생성
│ │ ├── label_generator/ # forward-return 라벨 생성
│ │ └── dataset_export/ # Parquet 스냅샷 내보내기
│ ├── sync/ # 데이터 동기화
│ │ ├── issuer_sync/ # 기업 메타데이터 (Stock Oracle)
│ │ ├── macro_sync/ # FRED 거시 지표
│ │ └── short_volume_sync/ # FINRA 공매도 잔량
│ ├── tracker/ # 전략 개선 추적 CLI (SQS, Journal, Leaderboard)
│ ├── tools/ # 분석·유틸리티 스크립트 (10개)
│ ├── review/ # 수동 리뷰 큐
│ └── qa/ # 데이터 품질 체크
│
├── libs/ # 핵심 라이브러리
│ ├── backtest/ # 백테스트 엔진 (3,600+ lines)
│ │ ├── domain.py # Pydantic 도메인 모델 30개+
│ │ ├── scoring.py # 후보 스코어링 (PEAD, composite)
│ │ ├── execution.py # 진입/청산 시뮬레이션
│ │ ├── allocator.py # 포지션 사이징 & 7개 진입 게이트
│ │ ├── selector.py # 후보 필터링 & 랭킹
│ │ ├── metrics.py # 21개 성과 지표 + 부트스트랩 CI
│ │ ├── tracker.py # SQS 계산, Journal I/O, Leaderboard
│ │ ├── snapshot_store.py # Parquet 데이터 로더
│ │ ├── splits.py # Walk-forward 윈도우 생성
│ │ ├── manifests.py # 실험 config 해석
│ │ ├── artifacts.py # 실행 결과 출력 (Parquet, JSON, CSV)
│ │ └── calendar.py # 거래일 유틸리티
│ ├── oracle_client/ # Stock Oracle API 클라이언트
│ │ ├── client.py # 비동기 httpx 클라이언트 (retry 로직)
│ │ ├── price.py # OHLCV bars, ATR-14
│ │ ├── financial.py # 재무제표 (BS, IS)
│ │ ├── filings.py # SEC 공시 검색/다운로드
│ │ ├── company.py # 기업 메타데이터
│ │ ├── screener.py # 유니버스 스크리닝
│ │ ├── fred.py # 거시 데이터
│ │ └── finra.py # 공매도 잔량
│ ├── parser/ # 공시 문서 파서
│ │ ├── rule_parser.py # 규칙 기반 8-K/6-K 파서
│ │ ├── merger.py # 규칙+LLM 결과 병합
│ │ └── text_normalizer.py # 텍스트 정규화
│ ├── features/ # 피처 엔지니어링
│ │ ├── builder.py # 오케스트레이터 (market+event+financial)
│ │ ├── market_features.py # 가격/거래량 피처
│ │ ├── event_features.py # 이벤트 품질 피처
│ │ ├── financial_features.py # 재무 피처 (EPS, 마진, 레버리지)
│ │ └── text_features.py # LLM 기반 텍스트 피처
│ ├── db/ # PostgreSQL 모델 (async SQLAlchemy)
│ │ ├── models.py # 7개 테이블 (Issuer, Symbol, Event, Document, ...)
│ │ └── migrations/ # Alembic 마이그레이션
│ ├── labeler/ # 라벨 생성 (1/3/5/7/15일 forward return, MFE/MAE)
│ ├── common/ # 로깅, config, 시간 유틸리티
│ ├── schemas/ # 공유 Pydantic 스키마
│ ├── export/ # 스냅샷 내보내기 (DB → Parquet)
│ ├── review/ # 리뷰 로직
│ └── llm/ # LLM 통합 (Claude/OpenAI, 캐시, 프롬프트)
│
├── configs/
│ ├── backtest/
│ │ └── defaults.json # 기본 전략 파라미터 (21개 설정)
│ ├── experiments/ # 실험 매니페스트 66개 (pead_midcap_step*, ...)
│ ├── app.yaml # 애플리케이션 설정
│ ├── symbols.yaml # 활성 유니버스
│ ├── symbols_midcap.yaml # 미드캡 유니버스 ($2B-10B)
│ ├── symbols_largecap.yaml # 라지캡 유니버스 ($10B+)
│ ├── symbols_smallmid.yaml # 스몰미드 유니버스
│ └── fred_series.yaml # FRED 거시 지표 정의
│
├── data/
│ ├── datasets/snapshots/ # Parquet 스냅샷 (후보, bars, macro)
│ ├── cache/ # LLM/API 응답 캐시
│ ├── analysis/ # 분석 결과 (플롯, CSV)
│ └── parquet/ # 원시 Parquet 내보내기
│
├── runs/ # 백테스트 실행 결과
│ └── midcap_steps/ # 65+ 실행 기록 (metrics.json, trade_blotter.csv, equity_curve.parquet)
│
├── journal/ # 전략 개선 저널
│ ├── improvement_journal.jsonl # append-only 개선 사이클 로그 (21개+ 엔트리)
│ ├── experiment_registry.json # 리더보드 데이터 (자동 생성)
│ └── LEADERBOARD.md # 사람이 읽을 수 있는 리더보드 (자동 생성)
│
├── tests/
│ ├── unit/ # 단위 테스트 270개+ (외부 의존성 없음)
│ ├── integration/ # 통합 테스트 (PostgreSQL 필요)
│ └── replay/ # 결정론/재현성 테스트
│
├── dev/ # 개발 문서
│ ├── overview.md # 전체 아키텍처 비전
│ ├── finished/ # 완료된 Phase 문서 (0, 1, 2, 3, 4)
│ └── phase5~8_deliverables/ # 진행 중인 Phase 문서
│
├── docker/ # Docker 관련 파일
├── docker-compose.yml # PostgreSQL 16 + Adminer
├── pyproject.toml # 의존성 & 빌드 설정
├── Makefile # 빌드 & 테스트 자동화
└── alembic.ini # DB 마이그레이션 설정
```
---
**Key environment variables:**
## 5. 데이터 파이프라인
| Variable | Default | Description |
|----------|---------|-------------|
| `STOCK_ORACLE_URL` | `http://localhost:18001` | Stock Oracle API endpoint |
| `POSTGRES_DSN` | `postgresql+asyncpg://acef:acef@localhost:5432/acef` | Database connection |
| `DATA_ROOT` | `./data` | Data storage root |
| `LOG_LEVEL` | `INFO` | Logging level |
| `LLM_ENABLED` | `false` | Enable LLM document parsing |
### 5.1 데이터 흐름
## Usage
```
단계 1: 원시 데이터 수집 (Phase 2)
┌──────────────────────────────────────────────────────┐
│ SEC EDGAR → 공시 JSON → 문서 (HTML/TXT) → raw zone │
│ Stock Oracle → 일봉 OHLCV → market_bars │
│ FRED → 거시 지표 (금리, 스프레드) → macro_features │
│ FINRA → 일별 공매도 잔량 → short_volume │
└──────────────────────────────────────────────────────┘
↓
단계 2: 파싱 & 피처 엔지니어링 (Phase 3)
┌──────────────────────────────────────────────────────┐
│ Rule Parser → event_type, guidance, confidence │
│ LLM Parser → 문서 이해 보강 (비활성화) │
│ Feature Builder → XBRL, reaction, text 스코어 │
│ Label Generator → 1/3/5/7/15일 forward return, │
│ MFE/MAE, time-to-target │
└──────────────────────────────────────────────────────┘
↓
단계 3: 스냅샷 내보내기
┌──────────────────────────────────────────────────────┐
│ PostgreSQL → Parquet 스냅샷 (point-in-time) │
│ 후보 (symbol, score, features) │
│ Bars (symbol, date, OHLCV) │
│ Macro (date, 거시 지표) │
└──────────────────────────────────────────────────────┘
↓
단계 4: 백테스트 & 평가 (Phase 4)
┌──────────────────────────────────────────────────────┐
│ SnapshotStore.load() → 메모리 로드 │
│ Signal Ranker → Candidate Selection → Position Sizing │
│ Entry/Exit Simulation → 21개 성과 지표 계산 │
│ SQS Score → Journal 기록 → Leaderboard 갱신 │
└──────────────────────────────────────────────────────┘
```
### Data Pipeline
### 5.2 실행 명령어
```bash
# Sync company metadata
# 기업 메타데이터 동기화
python -m apps.sync.issuer_sync.main
# Fetch and parse filings
# 공시 폴링 → 다운로드 → 파싱
python -m apps.pipeline.filing_poller.main
python -m apps.pipeline.filing_fetcher.main
python -m apps.pipeline.event_parser.main
# Build features and labels
# 피처 빌드 → 라벨 생성
python -m apps.pipeline.feature_builder.main
python -m apps.pipeline.label_generator.main
# Export Parquet snapshot for backtesting
# Parquet 스냅샷 내보내기
python -m apps.pipeline.dataset_export.main
```
### Backtesting
---
## 6. 전략 엔진
### 6.1 현재 구현된 전략: PEAD (Post-Earnings Announcement Drift)
**가설:** 시장은 어닝 발표에 과소반응하며, 발표 후 3~5일간 드리프트(continuation)가 발생한다.
| 구성 요소 | 설정 |
|----------|------|
| **신호** | 어닝 발표 + reaction day return ≥ 10% + PEAD score ≥ 0.65 |
| **진입** | reaction day 다음 세션 시초가 (next-open) |
| **보유** | 최대 7 거래일 |
| **청산** | Target (ATR 기반) 또는 Stop loss (ATR 기반) 또는 시간 만기 |
| **방향** | 롱 + 숏 (양방향) |
| **유니버스** | 미드캡 ($2B~$10B), NYSE/Nasdaq 보통주 |
### 6.2 PEAD 스코어링
3개 컴포넌트의 가중 합산:
| 컴포넌트 | 비중 | 피처 |
|----------|------|------|
| **Event Quality** | 65% | parser confidence + signal strength + guidance |
| **Reaction Direction** | 20% | positive/flat/negative 방향성 |
| **Volume Conviction** | 15% | 평균 대비 거래량 확인 |
### 6.3 진입 게이트 (7개 Veto)
포지션이 열리기 전 모든 게이트를 통과해야 함:
1. **유니버스 필터** — 최소 가격 $5, 최소 ADV $1M, ETF 제외
2. **최대 포지션 수** — 포트폴리오 전체 동시 보유 제한 (기본 8)
3. **섹터 집중 제한** — 동일 섹터 최대 포지션 수
4. **포지션 크기 제한** — 포트폴리오 대비 최대 비중
5. **ADV 비율 제한** — 일평균 거래대금의 1% 이내
6. **연패 쿨다운** — 연속 손실 후 대기 (설정 가능)
7. **파싱 신뢰도 게이트** — 최소 confidence 40%
### 6.4 Exit 로직
```
매 거래일 포지션 상태 체크:
1) Target 도달? → target_1_fraction만큼 청산 (기본 50%)
2) Stop loss 도달? → 전량 청산
3) 보유 기한 초과? → 전량 시가 청산
4) Trailing stop? → 최고점 대비 하락 시 청산
```
| 파라미터 | 기본값 | 설명 |
|----------|--------|------|
| `stop_atr_multiplier` | 3.0 | ATR-14 × 3.0 stop |
| `target_atr_multiplier` | 1.5 | ATR-14 × 1.5 target |
| `target_1_fraction` | 0.5 | target 도달 시 50% 청산 |
| `trailing_stop_pct` | 0.05 | 5% trailing stop |
| `max_holding_days` | 15 | 최대 보유일 |
| `slippage_bps` | 10 | 10bps 슬리피지 |
---
## 7. 백테스트 시스템
### 7.1 이벤트 드리븐 시뮬레이션
- **이벤트 귀속**: reaction day (공시 후 첫 거래일) 기준
- **신호 계산**: reaction day 종가 기준
- **진입**: 다음 세션 시초가 (look-ahead bias 방지)
- **일별 포트폴리오 상태**: 매일 equity, exposure, drawdown 추적
- **Kill switch**: 최대 낙폭 25% 초과 시 전략 중단
### 7.2 Walk-Forward 검증
3-split 체계로 과적합 방지:
| Split | 역할 | 용도 |
|-------|------|------|
| **Train** | 파라미터 탐색 | 최적화용 |
| **Valid** | 검증 | 과적합 체크 |
| **Test** | 최종 평가 | OOS 성과 (SQS 계산 대상) |
### 7.3 21개 성과 지표
**거래 지표 (7개)**
- win_rate, avg_win%, avg_loss%, profit_factor, expectancy_r, avg_r, trade_count
**포트폴리오 지표 (8개)**
- total_return, monthly_returns, sharpe_ratio, max_drawdown, equity_curve_r², VaR, CVaR, consecutive_losses
**부트스트랩 95% 신뢰구간**
- 리샘플링을 통한 통계적 유의성 확인
### 7.4 실행 예시
```bash
# Single split backtest
# 단일 split 백테스트
python -m apps.backtester.run \
--manifest configs/experiments/pead_midcap_step1_fixedr.json \
--split test --output-root runs/midcap_steps
--manifest configs/experiments/pead_midcap_step14_score65.json \
--split test \
--snapshot-dir data/datasets/snapshots \
--output-root runs/midcap_steps
# 3-split backtest (train/valid/test)
# 3-split 전체 백테스트
for split in train valid test; do
python -m apps.backtester.run \
--manifest configs/experiments/pead_midcap_step1_fixedr.json \
--split $split --output-root runs/midcap_steps
--manifest configs/experiments/pead_midcap_step14_score65.json \
--split $split \
--snapshot-dir data/datasets/snapshots \
--output-root runs/midcap_steps
done
# Walk-forward cross-validation
# Walk-forward CV
python -m apps.backtester.run \
--manifest configs/experiments/pead_midcap_step1_fixedr.json \
--manifest configs/experiments/pead_midcap_step14_score65.json \
--walk-forward --wf-train-days 252 --wf-test-days 63
# Output includes SQS score after each run:
# Run complete: bt_baseline_swing_v1_...
# Trades: 95
# Total return: -0.46%
# SQS: 39.7 (profitability=23.9, risk=52.4, consistency=24.8, robustness=80.7)
```
### Experiment Configuration
Experiments are defined as JSON manifests in `configs/experiments/`:
```json
{
"experiment_name": "pead_midcap_step1_fixedr",
"dataset_snapshot_id": "midcap-filtered",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": { "scoring_model": "pead", "pead_reaction_threshold": 0.07 },
"execution": { "target_model": "fixed_r", "target_1_r": 2.0 },
"risk": { "max_positions": 8 }
},
"tags": ["pead", "midcap"]
}
**출력 예시:**
```
Run complete: bt_baseline_swing_v1_midcap-filte_20260316...
Trades: 72
Total return: +0.73%
SQS: 64.2 (profitability=68.4, risk=61.2, consistency=58.7, robustness=65.3)
```
---
## Strategy Improvement Tracking System
## 8. 전략 개선 시스템
전략 개선을 체계적으로 관리하기 위한 3계층 시스템.
중복 실험 방지, 데이터 기반 의사결정, 리더보드를 통한 최고 전략 추적.
### 8.1 Strategy Quality Score (SQS)
A structured system to prevent duplicate experiments, enable data-driven decisions,
and track the best strategy via a leaderboard.
**test split 지표만으로** 계산하는 종합 점수 (0~100). 높을수록 좋음.
### Strategy Quality Score (SQS)
| 카테고리 | 비중 | 하위 지표 | 비중 | 0점 기준 | 100점 기준 |
|----------|------|-----------|------|----------|-----------|
| **Profitability** | 40% | profit_factor | 60% | ≤ 0.8 | ≥ 2.0 |
| | | total_return_pct | 40% | ≤ -5% | ≥ +5% |
| **Risk** | 25% | max_drawdown_pct (역) | 50% | ≥ 10% | ≤ 1% |
| | | sharpe_ratio | 50% | ≤ -1.0 | ≥ 2.0 |
| **Consistency** | 20% | win_rate | 50% | ≤ 0.35 | ≥ 0.65 |
| | | monthly_win_rate | 50% | ≤ 0.30 | ≥ 0.70 |
| **Robustness** | 15% | equity_curve_r² | 50% | ≤ 0.0 | ≥ 0.80 |
| | | trade_count | 50% | ≤ 10 | ≥ 100 |
Composite score (0-100) computed from **test split metrics only**. Higher is better.
**Low-trade penalty:** test 거래 수 < 20이면 SQS를 절반으로 감산.
| Category | Weight | Sub-metric | Weight | 0 pts | 100 pts |
|----------|--------|------------|--------|-------|---------|
| **Profitability** | 40% | profit_factor | 60% | ≤0.8 | ≥2.0 |
| | | total_return_pct | 40% | ≤-5% | ≥+5% |
| **Risk** | 25% | max_drawdown_pct (inv) | 50% | ≥10% | ≤1% |
| | | sharpe_ratio | 50% | ≤-1.0 | ≥2.0 |
| **Consistency** | 20% | win_rate | 50% | ≤0.35 | ≥0.65 |
| | | monthly_win_rate | 50% | ≤0.30 | ≥0.70 |
| **Robustness** | 15% | equity_curve_r_squared | 50% | ≤0.0 | ≥0.80 |
| | | trade_count | 50% | ≤10 | ≥100 |
**SQS 해석 기준:**
**Low-trade penalty:** If test trades < 20, SQS is halved.
| SQS 범위 | 해석 |
|-----------|------|
| 0–20 | 손실 전략 |
| 20–40 | 손익분기 근처 |
| 40–55 | 유망, 개선 필요 |
| **55–70** | **좋음, OOS 엣지 있음** |
| 70–85 | 강함, 실전 후보 |
| 85–100 | 예외적 (데이터 오류 확인 필요) |
| SQS Range | Interpretation |
|-----------|----------------|
| 0-20 | Losing strategy |
| 20-40 | Near breakeven |
| 40-55 | Promising, needs work |
| 55-70 | Good, has OOS edge |
| 70-85 | Strong, live candidate |
| 85-100 | Exceptional (check for data issues) |
### 8.2 Journal (개선 저널)
### Journal & Leaderboard
모든 실험 결과를 기록하는 append-only 로그.
```
journal/
├── improvement_journal.jsonl ← Append-only improvement cycle log
├── experiment_registry.json ← Leaderboard data (regenerated)
└── LEADERBOARD.md ← Human-readable leaderboard (regenerated)
├── improvement_journal.jsonl ← append-only 개선 사이클 로그
├── experiment_registry.json ← 리더보드 데이터 (자동 생성)
└── LEADERBOARD.md ← 사람이 읽는 리더보드 (자동 생성)
```
Each journal entry records one improvement cycle:
**Journal 엔트리 구조:**
```json
{
"entry_id": "IMP-0001",
"timestamp": "2026-03-16T19:30:00",
"experiment_name": "pead_midcap_step3_10pct",
"hypothesis": "Raise reaction threshold to 10% for stronger signals",
"entry_id": "IMP-0015",
"timestamp": "2026-03-17T01:45:00+00:00",
"experiment_name": "pead_midcap_step14_score65",
"hypothesis": "Score threshold 0.60→0.65로 높여 약한 신호 필터링",
"config_delta": {
"base_experiment": "pead_midcap_step13_best",
"changes": { "signal.score_threshold": "0.60 → 0.65" }
},
"results": {
"train": { "run_id": "bt_...", "trade_count": 420, "profit_factor": 0.95, ... },
"valid": { "run_id": "bt_...", ... },
"test": { "run_id": "bt_...", ... }
"train": { "run_id": "bt_...", "trade_count": 312, "profit_factor": 1.15, "..." : "..." },
"valid": { "run_id": "bt_...", "trade_count": 145, "profit_factor": 1.28, "..." : "..." },
"test": { "run_id": "bt_...", "trade_count": 72, "profit_factor": 1.22, "..." : "..." }
},
"sqs_score": 64.2,
"sqs_breakdown": {
"profitability": 68.4,
"risk": 61.2,
"consistency": 58.7,
"robustness": 65.3
},
"sqs_score": 38.5,
"sqs_breakdown": { "profitability": 35.2, "risk": 45.0, "consistency": 30.0, "robustness": 42.0 },
"verdict": "better",
"verdict_reasoning": "Test PF 0.89 -> 1.00, breakeven achieved",
"next_direction": "Combine 10% threshold + maxcand3"
"verdict_reasoning": "Test SQS 64.2 > 57.7. Return +0.73% > +0.43%. Trade count 72 sufficient.",
"next_direction": "Test exit tuning (target, fraction, hold) on top of score 0.65"
}
```
### Tracker CLI
### 8.3 실험 매니페스트
실험은 JSON 매니페스트로 정의. `base_config`에 `overrides`를 적용하는 구조:
```json
{
"experiment_name": "pead_midcap_step14_score65",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 14: Score threshold 0.60→0.65",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 2.0,
"score_threshold": 0.65,
"max_candidates_per_day": 3
},
"execution": { "max_holding_days": 7 },
"risk": { "max_positions": 8 }
},
"tags": ["pead", "midcap", "step14", "score65"]
}
```
### 8.4 개선 워크플로우
```
1. 실험 config 생성 configs/experiments/my_experiment.json
2. 3-split 백테스트 실행 for split in train valid test; do ... done
3. Journal에 기록 python -m apps.tracker.cli record ...
4. Leaderboard 확인 python -m apps.tracker.cli leaderboard ...
5. 다음 실험 계획 verdict + SQS breakdown 기반
6. 중복 실험 확인 python -m apps.tracker.cli check-duplicate ...
7. 1번부터 반복
```
### 8.5 Tracker CLI
```bash
# Record experiment results to journal
# 실험 결과 기록
python -m apps.tracker.cli record \
--journal-dir journal/ \
--runs-dir runs/midcap_steps/ \
--experiment pead_midcap_step3_10pct \
--hypothesis "Raise reaction threshold to 10%" \
--baseline pead_7pct_midcap \
--experiment pead_midcap_step14_score65 \
--hypothesis "Score threshold 0.60→0.65" \
--baseline pead_midcap_step13_best \
--verdict better \
--reasoning "Test PF improved from 0.89 to 1.00" \
--next "Combine 10% threshold + maxcand3"
--reasoning "Test SQS 64.2 > 57.7, Return +0.73% > +0.43%" \
--next "Test exit tuning on top of score 0.65"
# View leaderboard
# 리더보드 출력
python -m apps.tracker.cli leaderboard --journal-dir journal/
# # Experiment SQS PF Ret% Trades
# ------------------------------------------------------------------
# 1 pead_7pct_longshort_v2 60.7 1.31 +2.0 55
# 2 pead_midcap_step3_10pct 38.5 1.00 +0.0 85
# 엔트리 상세 조회
python -m apps.tracker.cli show --journal-dir journal/ IMP-0015
# Show entry details
python -m apps.tracker.cli show --journal-dir journal/ IMP-0001
# Check for duplicate experiments
# 중복 실험 확인
python -m apps.tracker.cli check-duplicate \
--journal-dir journal/ --experiment pead_midcap_step3_10pct
--journal-dir journal/ --experiment my_experiment_name
# 레지스트리 재생성
python -m apps.tracker.cli rebuild-registry --journal-dir journal/
```
### Improvement Workflow
### 8.6 핵심 발견사항
21개 실험을 통해 얻은 인사이트:
| 발견 | 설명 |
|------|------|
| **Score threshold가 가장 강력한 레버** | 0.60→0.65로 올리면 SQS 50.2→64.2 (+28%) |
| **넓은 funnel은 OOS에서 실패** | reaction 0.10→0.07로 낮추면 거래 수는 늘지만 품질 하락 |
| **Exit 튜닝은 한계적** | target/fraction/hold 변경은 noise 범위 내 |
| **조합이 개별보다 나쁨** | 한계적 개선들을 합치면 과적합 리스크로 오히려 하락 |
| **단순함이 최고** | Step14(score 0.65만 변경)가 모든 복합 변형보다 우수 |
| **라지캡 확장 실패** | $10B+ 대형주에서는 OOS 성과 없음, 미드캡이 최적 |
| **보유기간 5~7일 최적** | PEAD 드리프트의 평균 보유는 3.28일 |
---
## 9. 현재 개발 현황
### Phase별 진행 상태
| Phase | 이름 | 상태 | 설명 |
|-------|------|------|------|
| **Phase 0** | 전략/운용 명세 동결 | **완료** | strategy_spec.md, risk_policy.md, data_source_policy.md, event_taxonomy.md |
| **Phase 1** | 개발 기반 & DB 스키마 | **완료** | Docker, PostgreSQL, Alembic, adapter contracts, 공통 인프라 |
| **Phase 2** | 핵심 데이터 수집 | **완료** | SEC, Alpaca(→Stock Oracle), FRED, FINRA 어댑터 구축 |
| **Phase 3** | 문서 파서 & 피처 빌더 | **완료** | 규칙 기반 파서 + LLM 보강, 피처 엔지니어링, 라벨 생성 |
| **Phase 4** | 백테스트 엔진 | **완료** | 이벤트 드리븐 시뮬레이터, walk-forward CV, SQS 스코어링, 21개 지표 |
### 현재 작동 중인 것
- 백테스트 엔진 (이벤트 드리븐, 재현 가능, 21개 지표)
- Walk-forward cross-validation (train/valid/test 분할)
- 실험 추적 (Journal, Leaderboard, SQS 스코어링)
- PEAD 전략 (주력 신호, SQS 64.2, +0.73% OOS)
- 단위/통합 테스트 (270개+ 자동화)
- 포트폴리오 리스크 엔진 (포지션 사이징, 섹터 제한, stop/target)
- 데이터 파이프라인 (filing_poller → event_parser → feature_builder → dataset_export)
---
## 10. 현재 최고 전략 성과
### Leaderboard (Top 10, 2026-03-17 기준)
| # | Experiment | SQS | PF | Ret% | WR | Sharpe | DD% | Trades |
|---|-----------|-----|-----|------|-----|--------|-----|--------|
| 1 | **pead_midcap_step14_score65** | **64.2** | 1.22 | +0.7 | 57% | 1.3 | 0.9 | 72 |
| 2 | pead_midcap_step18_nofrac | 63.2 | 1.23 | +0.8 | 52% | 1.4 | 0.9 | 64 |
| 3 | pead_midcap_step19_hold5 | 62.9 | 1.20 | +0.7 | 57% | 1.2 | 0.9 | 72 |
| 4 | pead_midcap_step20_best3 | 62.0 | 1.22 | +0.7 | 52% | 1.3 | 0.9 | 64 |
| 5 | pead_midcap_step17_target2 | 59.4 | 1.18 | +0.6 | 53% | 1.1 | 0.8 | 66 |
| 6 | pead_midcap_step13_best | 57.7 | 1.12 | +0.4 | 55% | 0.8 | 0.9 | 75 |
| 7 | pead_midcap_step16_react7_score65 | 53.2 | 0.97 | -0.1 | 56% | -0.2 | 1.6 | 89 |
| 8 | pead_midcap_step5_maxcand3 | 52.5 | 0.95 | -0.2 | 54% | -0.4 | 1.4 | 96 |
| 9 | pead_midcap_step15_react7 | 51.7 | 0.94 | -0.3 | 55% | -0.5 | 1.6 | 91 |
| 10 | pead_midcap_step11_score60 | 50.2 | 1.02 | +0.1 | 52% | 0.2 | 0.9 | 77 |
### 최고 전략: `pead_midcap_step14_score65`
```
1. Create experiment config configs/experiments/my_experiment.json
2. Run 3-split backtest for split in train valid test; do ... done
3. Record to journal python -m apps.tracker.cli record ...
4. Check leaderboard python -m apps.tracker.cli leaderboard ...
5. Plan next experiment based on verdict + SQS breakdown
6. Check for duplicates python -m apps.tracker.cli check-duplicate ...
7. Repeat from step 1
SQS: 64.2 (Good — OOS edge 확인)
수익률: +0.73% (OOS test)
PF: 1.22
승률: 56.9%
Sharpe: 1.3
최대낙폭: 0.9%
거래 수: 72
보유기간: 평균 3.28일
```
**핵심 설정:**
- reaction threshold: 10% (어닝 발표 후 첫날 10% 이상 반응한 종목만)
- score threshold: 0.65 (PEAD 스코어 0.65 이상만 진입)
- volume threshold: 2.0x (평균 대비 2배 이상 거래량)
- max candidates/day: 3 (하루 최대 3개 후보)
- max holding: 7일
- earnings_release만 활성화 (다른 이벤트 타입 비활성화)
---
## Testing
## 11. 설치 및 실행
### 요구사항
- Python 3.11+
- PostgreSQL 16 (Docker 제공)
- Stock Oracle API (Docker, localhost:18001)
### 설치
```bash
# Unit tests (fast, no external deps)
# 의존성 설치
pip install -e ".[dev]"
# PostgreSQL 시작
docker compose up -d
# DB 마이그레이션
alembic upgrade head
# 환경 변수 설정
cp .env.example .env
```
### 환경 변수
| 변수 | 기본값 | 설명 |
|------|--------|------|
| `STOCK_ORACLE_URL` | `http://localhost:18001` | Stock Oracle API 엔드포인트 |
| `POSTGRES_DSN` | `postgresql+asyncpg://acef:acef@localhost:5432/acef` | DB 연결 |
| `DATA_ROOT` | `./data` | 데이터 저장 루트 |
| `LOG_LEVEL` | `INFO` | 로그 레벨 |
| `LLM_ENABLED` | `false` | LLM 문서 파싱 활성화 |
### Makefile 타겟
```bash
make bootstrap # 의존성 설치
make db-upgrade # 마이그레이션 실행
make db-reset # DB 전체 리셋
make test-unit # 단위 테스트만
make test-integration # 통합 테스트 (PostgreSQL 필요)
make test-replay # 결정론 테스트
make lint # Ruff 린트
make typecheck # MyPy 엄격 모드
make ci # 전체 CI (lint + typecheck + test)
```
---
## 12. 사용법
### 데이터 파이프라인 실행
```bash
# 1) 기업 메타데이터 동기화
python -m apps.sync.issuer_sync.main
# 2) 공시 수집 → 파싱 → 피처 빌드 → 라벨 생성
python -m apps.pipeline.filing_poller.main
python -m apps.pipeline.filing_fetcher.main
python -m apps.pipeline.event_parser.main
python -m apps.pipeline.feature_builder.main
python -m apps.pipeline.label_generator.main
# 3) 백테스트용 Parquet 스냅샷 내보내기
python -m apps.pipeline.dataset_export.main
```
### 백테스트 실행
```bash
# 3-split 백테스트 (train/valid/test)
for split in train valid test; do
python -m apps.backtester.run \
--manifest configs/experiments/pead_midcap_step14_score65.json \
--split $split \
--snapshot-dir data/datasets/snapshots \
--output-root runs/midcap_steps
done
```
> **주의:** Stock Oracle API는 단일 스레드이므로 백테스트를 **순차적으로** 실행해야 합니다.
> 병렬 실행 시 API가 과부하되어 연결이 끊깁니다.
### 실험 결과 기록
```bash
# Journal에 기록
python -m apps.tracker.cli record \
--journal-dir journal/ \
--runs-dir runs/midcap_steps/ \
--experiment pead_midcap_step14_score65 \
--hypothesis "Score threshold 0.60→0.65" \
--baseline pead_midcap_step13_best \
--verdict better \
--reasoning "Test SQS 64.2 > 57.7" \
--next "Exit tuning on top of score 0.65"
# 리더보드 확인
python -m apps.tracker.cli leaderboard --journal-dir journal/
```
### 테스트
```bash
# 단위 테스트 (빠름, 외부 의존성 없음)
pytest tests/unit/ -v
# Backtest module tests only
# 백테스트 모듈 테스트
pytest tests/unit/backtest/ -v
# Integration tests (requires PostgreSQL)
# 통합 테스트 (PostgreSQL 필요)
pytest tests/integration/ -v
# Full CI check (lint + typecheck + unit tests)
# 전체 CI
make ci
```
## Data Sources
| Source | Role | Cost |
|--------|------|------|
| **SEC EDGAR** | Primary event source (8-K, 10-Q, 6-K filings) | Free |
| **Stock Oracle API** | Market data (OHLCV bars, company info) | Internal |
| **FRED** | Macro regime indicators (rates, spreads) | Free |
| **FINRA** | Short sale volume (crowding signal) | Free |
| Wikimedia | Retail attention via pageviews | Free |
| YouTube | Channel-based attention tracking | Free (quota limited) |
| Yahoo RSS | Headline burst detection | Free |
## Tech Stack
| Component | Technology |
|-----------|-----------|
| Language | Python 3.11+ |
| Models | Pydantic v2 |
| Database | PostgreSQL 16 + async SQLAlchemy |
| Research data | DuckDB + Parquet |
| Containers | Docker Compose |
| Linting | Ruff |
| Type checking | MyPy (strict) |
| Testing | Pytest + asyncio |
| Logging | structlog (JSON) |
---
## 13. 기술 스택
| 구성 요소 | 기술 |
|-----------|------|
| 언어 | Python 3.11+ |
| 도메인 모델 | Pydantic v2 |
| 운영 DB | PostgreSQL 16 + async SQLAlchemy + asyncpg |
| 연구 데이터 | DuckDB + Apache Parquet |
| 마이그레이션 | Alembic |
| HTTP 클라이언트 | httpx (비동기) |
| 컨테이너 | Docker Compose |
| 린트 | Ruff (line 100, Python 3.11+) |
| 타입 체크 | MyPy (strict mode) |
| 테스트 | Pytest + pytest-asyncio (270+ 테스트) |
| 로깅 | structlog (JSON) |
| LLM | Ollama (인프라 구현, 파서 비활성화) |
| 문서 파싱 | BeautifulSoup4 + 규칙 기반 + LLM |
| 거래일 캘린더 | exchange-calendars |
---
## 14. 데이터 소스
| 소스 | 역할 | 계층 | 비용 |
|------|------|------|------|
| **SEC EDGAR** | 이벤트 원천 (8-K, 10-Q, 6-K 공시) | Core | 무료 |
| **Stock Oracle API** | 시장 데이터 (OHLCV, 기업정보, 재무) | Core | 내부 |
| **FRED** | 거시 레짐 (금리, 스프레드, 경기 지표) | Core | 무료 |
| **FINRA** | 공매도 잔량 (crowding 신호) | 보조 | 무료 |
### 데이터 정책
- **모든 시장 데이터는 Stock Oracle API를 통해서만 접근** — 직접 외부 API 호출 금지
- API 부재 기능은 구현 대신 보고
- 유료 데이터 소스 사용 금지
- LLM은 문서 해석기로만 사용, 가격 예측 금지
---
## License

@ -5,6 +5,8 @@ import argparse
import asyncio
import json
import yaml
from libs.common.config import get_settings
from libs.common.logging import configure_logging, get_logger
from libs.db.session import get_session
@ -18,6 +20,8 @@ async def run_dataset_export(
split_policy: str,
output_dir: str,
feature_versions: list[str] | None = None,
label_version: str = "label-2.0.0",
symbols: list[str] | None = None,
) -> dict:
async with get_session() as session:
manifest = await export_dataset_snapshot(
@ -26,6 +30,8 @@ async def run_dataset_export(
split_policy=split_policy,
output_dir=output_dir,
feature_versions=feature_versions,
label_version=label_version,
symbols=symbols,
)
return manifest
@ -49,18 +55,37 @@ def main() -> None:
default=None,
help="Feature versions to merge (e.g. market_v1 event_v1 financial_v1)",
)
parser.add_argument(
"--label-version",
default="label-2.0.0",
help="Label version to export (default: label-2.0.0)",
)
parser.add_argument(
"--symbols-file",
default=None,
help="YAML file with 'symbols' list to filter export (e.g. configs/symbols_midcap.yaml)",
)
parser.add_argument("--json", action="store_true", help="Print manifest JSON to stdout")
args = parser.parse_args()
settings = get_settings()
configure_logging(settings.log_level)
symbols = None
if args.symbols_file:
with open(args.symbols_file) as f:
cfg = yaml.safe_load(f) or {}
symbols = [s for s in cfg.get("symbols", []) if isinstance(s, str)]
print(f"Filtering to {len(symbols)} symbols from {args.symbols_file}")
manifest = asyncio.run(
run_dataset_export(
snapshot_id=args.snapshot_id,
split_policy=args.split_policy,
output_dir=args.output_dir,
feature_versions=args.feature_versions,
label_version=args.label_version,
symbols=symbols,
)
)

@ -1,4 +1,10 @@
"""Event Parser: parse exhibit text → events + event_parses."""
"""Event Parser: parse exhibit text → events + event_parses.
Supports two modes:
- Normal: parse documents with parsed_status="ready_for_parse"
- Reparse (--reparse): re-classify existing events using item_numbers
from the Document table, updating event_type in Event rows.
"""
from __future__ import annotations
import argparse
@ -6,7 +12,7 @@ import asyncio
import datetime as dt
import uuid
from sqlalchemy import select
from sqlalchemy import select, update
from libs.common.config import get_settings
from libs.common.file_store import read_exhibit
@ -15,7 +21,8 @@ from libs.common.ids import new_job_run_id
from libs.common.logging import bind_job_run_id, configure_logging, get_logger
from libs.db.models import Document, Event, EventParse, JobRun
from libs.db.session import get_session
from libs.parser.rule_parser import PARSER_VERSION, SCHEMA_VERSION, RuleBasedParser
from libs.oracle_client import FilingsService, make_oracle_client
from libs.parser.rule_parser import PARSER_VERSION, SCHEMA_VERSION, RuleBasedParser, _classify_event_type
from libs.parser.schema_validator import validate_parser_output
from libs.parser.text_normalizer import normalize_text
@ -76,103 +83,109 @@ async def run_event_parser(run_id: str) -> dict[str, int]:
doc.accepted_at_utc.isoformat() if doc.accepted_at_utc else None
),
"form_type": doc.form_type,
"item_numbers": doc.item_numbers or [],
}
try:
output = _parser.parse(
document_id=doc.document_id,
form_type=doc.form_type,
text=normalized,
metadata=metadata,
)
output_dict = output.model_dump()
errors = validate_parser_output(output_dict)
if errors:
logger.warning(
"parse_validation_failed",
# Use savepoint so one failure doesn't break the session
async with session.begin_nested():
output = _parser.parse(
document_id=doc.document_id,
errors=errors[:3],
)
# Store invalid parse, no event row
event_id_str = make_event_id(doc.document_id, output.event_type)
# Create a minimal event row first (needed for FK)
event = Event(
event_id=event_id_str,
primary_document_id=doc.document_id,
issuer_id=doc.issuer_id,
symbol_id=doc.symbol_id,
event_type=output.event_type,
event_direction=output.event_direction,
event_date=doc.filing_date,
filed_at_utc=doc.accepted_at_utc,
parser_version=PARSER_VERSION,
parse_confidence=output.confidence.overall,
status="rejected",
)
session.add(event)
await session.flush()
parse_row = EventParse(
event_id=event_id_str,
parser_kind="rule",
parser_version=PARSER_VERSION,
schema_version=SCHEMA_VERSION,
output_json=output_dict,
validation_status="invalid",
validation_errors={"errors": errors},
)
session.add(parse_row)
doc.parsed_status = "failed"
stats["invalid"] += 1
else:
event_id_str = make_event_id(doc.document_id, output.event_type)
# Check duplicate event
existing_evt = await session.execute(
select(Event).where(Event.event_id == event_id_str)
form_type=doc.form_type,
text=normalized,
metadata=metadata,
)
if existing_evt.scalar_one_or_none() is not None:
logger.info("event_already_exists", event_id=event_id_str)
output_dict = output.model_dump()
# Sanitize null bytes that PostgreSQL JSONB rejects
import json as _json
_raw = _json.dumps(output_dict)
if "\x00" in _raw or "\\u0000" in _raw:
_raw = _raw.replace("\x00", "").replace("\\u0000", "")
output_dict = _json.loads(_raw)
errors = validate_parser_output(output_dict)
if errors:
logger.warning(
"parse_validation_failed",
document_id=doc.document_id,
errors=errors[:3],
)
event_id_str = make_event_id(doc.document_id, output.event_type)
event = Event(
event_id=event_id_str,
primary_document_id=doc.document_id,
issuer_id=doc.issuer_id,
symbol_id=doc.symbol_id,
event_type=output.event_type,
event_direction=output.event_direction,
event_date=doc.filing_date,
filed_at_utc=doc.accepted_at_utc,
parser_version=PARSER_VERSION,
parse_confidence=output.confidence.overall,
status="rejected",
)
session.add(event)
await session.flush()
parse_row = EventParse(
event_id=event_id_str,
parser_kind="rule",
parser_version=PARSER_VERSION,
schema_version=SCHEMA_VERSION,
output_json=output_dict,
validation_status="invalid",
validation_errors={"errors": errors},
)
session.add(parse_row)
doc.parsed_status = "failed"
stats["invalid"] += 1
else:
event_id_str = make_event_id(doc.document_id, output.event_type)
existing_evt = await session.execute(
select(Event).where(Event.event_id == event_id_str)
)
if existing_evt.scalar_one_or_none() is not None:
logger.info("event_already_exists", event_id=event_id_str)
doc.parsed_status = "succeeded"
stats["valid"] += 1
continue
event = Event(
event_id=event_id_str,
primary_document_id=doc.document_id,
issuer_id=doc.issuer_id,
symbol_id=doc.symbol_id,
event_type=output.event_type,
event_direction=output.event_direction,
event_date=doc.filing_date,
filed_at_utc=doc.accepted_at_utc,
parser_version=PARSER_VERSION,
parse_confidence=output.confidence.overall,
status="pending",
)
session.add(event)
await session.flush()
parse_row = EventParse(
event_id=event_id_str,
parser_kind="rule",
parser_version=PARSER_VERSION,
schema_version=SCHEMA_VERSION,
output_json=output_dict,
validation_status="valid",
validation_errors=None,
)
session.add(parse_row)
doc.parsed_status = "succeeded"
stats["valid"] += 1
continue
event = Event(
event_id=event_id_str,
primary_document_id=doc.document_id,
issuer_id=doc.issuer_id,
symbol_id=doc.symbol_id,
event_type=output.event_type,
event_direction=output.event_direction,
event_date=doc.filing_date,
filed_at_utc=doc.accepted_at_utc,
parser_version=PARSER_VERSION,
parse_confidence=output.confidence.overall,
status="pending",
)
session.add(event)
await session.flush()
parse_row = EventParse(
event_id=event_id_str,
parser_kind="rule",
parser_version=PARSER_VERSION,
schema_version=SCHEMA_VERSION,
output_json=output_dict,
validation_status="valid",
validation_errors=None,
)
session.add(parse_row)
doc.parsed_status = "succeeded"
stats["valid"] += 1
logger.info(
"event_created",
event_id=event_id_str,
event_type=output.event_type,
confidence=output.confidence.overall,
)
logger.info(
"event_created",
event_id=event_id_str,
event_type=output.event_type,
confidence=output.confidence.overall,
)
except Exception as exc:
logger.error(
@ -195,16 +208,170 @@ async def run_event_parser(run_id: str) -> dict[str, int]:
return stats
async def reparse_events(run_id: str) -> dict[str, int]:
"""Re-classify existing events by fetching item_numbers from Oracle API.
For each Document with a linked Event where event_type='unknown':
1. Look up filing in Oracle API to get item_numbers
2. Store item_numbers in Document row
3. Re-classify event_type based on items
4. Update Event row with new event_type and re-parse with exhibit text
"""
settings = get_settings()
app_config = settings.get_app_config()
exhibit_types = app_config.get("pipeline", {}).get("exhibit_types", ["EX-99.1"])
stats = {"seen": 0, "updated": 0, "skipped": 0, "errors": 0}
async with make_oracle_client() as client:
svc = FilingsService(client)
async with get_session() as session:
job = JobRun(
job_run_id=uuid.UUID(run_id),
job_name="event_parser_reparse",
source_name="oracle",
run_date=dt.date.today(),
status="running",
)
session.add(job)
await session.flush()
# Find all events with event_type='unknown'
result = await session.execute(
select(Event, Document)
.join(Document, Event.primary_document_id == Document.document_id)
.where(Event.event_type == "unknown")
)
pairs = result.all()
stats["seen"] = len(pairs)
if not pairs:
logger.info("reparse_nothing_to_do")
job.status = "succeeded"
job.finished_at_utc = dt.datetime.now(tz=dt.UTC)
job.records_seen = 0
return stats
# Step 1: Fetch item_numbers from Oracle for all documents missing them
docs_needing_items = {
doc.document_id: doc
for _, doc in pairs
if not doc.item_numbers
}
if docs_needing_items:
logger.info("reparse_fetching_items", count=len(docs_needing_items))
for doc in docs_needing_items.values():
if not doc.accession_no:
continue
try:
items = await svc.get_filing_items(doc.accession_no)
if items:
doc.item_numbers = items
logger.info(
"reparse_items_found",
accession_no=doc.accession_no,
items=items,
)
except Exception as exc:
logger.warning(
"reparse_oracle_error",
accession_no=doc.accession_no,
error=str(exc),
)
await session.flush()
# Step 2: Re-parse each event with updated item_numbers
for event, doc in pairs:
try:
items = doc.item_numbers or []
new_event_type = _classify_event_type(items)
if new_event_type == event.event_type:
stats["skipped"] += 1
continue
# Also re-parse the full exhibit text with items
text: str | None = None
for exhibit_type in exhibit_types:
try:
text = read_exhibit(doc.accession_no, exhibit_type)
break
except FileNotFoundError:
continue
if text:
normalized = normalize_text(text)
metadata = {
"filing_date": doc.filing_date.isoformat(),
"accepted_at_utc": (
doc.accepted_at_utc.isoformat()
if doc.accepted_at_utc
else None
),
"form_type": doc.form_type,
"item_numbers": items,
}
output = _parser.parse(
document_id=doc.document_id,
form_type=doc.form_type,
text=normalized,
metadata=metadata,
)
event.event_type = output.event_type
event.event_direction = output.event_direction
event.parse_confidence = output.confidence.overall
else:
# No exhibit text available, just update event_type from items
event.event_type = new_event_type
stats["updated"] += 1
logger.info(
"reparse_event_updated",
event_id=event.event_id,
old_type="unknown",
new_type=event.event_type,
)
except Exception as exc:
logger.error(
"reparse_error",
event_id=event.event_id,
error=str(exc),
)
stats["errors"] += 1
job.status = "succeeded" if stats["errors"] == 0 else "partial"
job.finished_at_utc = dt.datetime.now(tz=dt.UTC)
job.records_seen = stats["seen"]
job.records_written = stats["updated"]
job.error_count = stats["errors"]
logger.info("reparse_done", **stats)
return stats
def main() -> None:
parser = argparse.ArgumentParser(description="Event Parser")
parser.add_argument("--run-id", default=new_job_run_id())
parser.add_argument(
"--reparse",
action="store_true",
help="Re-classify existing events with event_type='unknown' using Oracle API item numbers",
)
args = parser.parse_args()
settings = get_settings()
configure_logging(settings.log_level)
bind_job_run_id(args.run_id)
asyncio.run(run_event_parser(args.run_id))
if args.reparse:
asyncio.run(reparse_events(args.run_id))
else:
asyncio.run(run_event_parser(args.run_id))
if __name__ == "__main__":

@ -14,9 +14,7 @@ from libs.common.logging import bind_job_run_id, configure_logging, get_logger
from libs.db.models import Event, JobRun
from libs.db.session import get_session
from libs.features.builder import build_features_for_event
from libs.oracle_client.client import make_oracle_client
from libs.oracle_client.financial import FinancialService
from libs.oracle_client.price import PriceService
from libs.oracle_client import FinancialService, PriceService, make_oracle_client
logger = get_logger(__name__)

@ -14,9 +14,8 @@ from libs.common.ids import new_job_run_id
from libs.common.logging import bind_job_run_id, configure_logging, get_logger
from libs.db.models import Document, ExhibitCache, JobRun
from libs.db.session import get_session
from libs.oracle_client.client import make_oracle_client
from libs.oracle_client import FilingsService, make_oracle_client
from libs.oracle_client.exceptions import OracleNotFoundError
from libs.oracle_client.filings import FilingsService
logger = get_logger(__name__)
@ -28,6 +27,8 @@ async def fetch_exhibits(run_id: str) -> dict[str, int]:
stats = {"seen": 0, "written": 0, "skipped": 0, "errors": 0}
batch_size = 100
async with make_oracle_client() as client:
svc = FilingsService(client)
@ -41,6 +42,7 @@ async def fetch_exhibits(run_id: str) -> dict[str, int]:
)
session.add(job)
await session.flush()
await session.commit()
result = await session.execute(
select(Document).where(Document.parsed_status == "pending")
@ -48,7 +50,7 @@ async def fetch_exhibits(run_id: str) -> dict[str, int]:
docs = result.scalars().all()
stats["seen"] = len(docs)
for doc in docs:
for idx, doc in enumerate(docs, 1):
if not doc.accession_no:
stats["skipped"] += 1
continue
@ -65,7 +67,10 @@ async def fetch_exhibits(run_id: str) -> dict[str, int]:
continue
try:
response = await svc.get_exhibit(doc.accession_no, exhibit_type)
response = await asyncio.wait_for(
svc.get_exhibit(doc.accession_no, exhibit_type),
timeout=30.0,
)
checksum = write_exhibit(
doc.accession_no, exhibit_type, response.content
)
@ -112,9 +117,41 @@ async def fetch_exhibits(run_id: str) -> dict[str, int]:
)
stats["errors"] += 1
if fetched_any:
doc.parsed_status = "ready_for_parse"
doc.updated_at_utc = dt.datetime.now(tz=dt.UTC)
# Extract item_numbers from SGML header if not already set
if not doc.item_numbers:
try:
items = await asyncio.wait_for(
svc.get_filing_items(doc.accession_no),
timeout=15.0,
)
if items:
doc.item_numbers = items
logger.info(
"item_numbers_extracted",
accession_no=doc.accession_no,
items=items,
)
except Exception as exc:
logger.debug(
"item_numbers_extraction_failed",
accession_no=doc.accession_no,
error=str(exc),
)
# Always advance to ready_for_parse (exhibit may not exist)
doc.parsed_status = "ready_for_parse"
doc.updated_at_utc = dt.datetime.now(tz=dt.UTC)
# Commit in batches to preserve progress
if idx % batch_size == 0:
await session.commit()
logger.info(
"batch_committed",
processed=idx,
total=stats["seen"],
written=stats["written"],
errors=stats["errors"],
)
job.status = "succeeded" if stats["errors"] == 0 else "partial"
job.finished_at_utc = dt.datetime.now(tz=dt.UTC)

@ -14,8 +14,7 @@ from libs.common.ids import new_job_run_id
from libs.common.logging import bind_job_run_id, configure_logging, get_logger
from libs.db.models import Document, IssuerMaster, JobRun, SymbolMaster
from libs.db.session import get_session
from libs.oracle_client.client import make_oracle_client
from libs.oracle_client.filings import FilingsService
from libs.oracle_client import FilingsService, make_oracle_client
logger = get_logger(__name__)
@ -103,6 +102,7 @@ async def poll_filings(
else None
),
primary_document_name=filing.primary_document,
item_numbers=filing.items if filing.items else None,
parsed_status="pending",
)
session.add(doc)

@ -14,8 +14,7 @@ from libs.common.logging import bind_job_run_id, configure_logging, get_logger
from libs.db.models import Event, EventLabel, JobRun, SymbolMaster
from libs.db.session import get_session
from libs.labeler.label_generator import LABEL_VERSION, generate_labels
from libs.oracle_client.client import make_oracle_client
from libs.oracle_client.price import PriceService
from libs.oracle_client import PriceService, make_oracle_client
logger = get_logger(__name__)

@ -13,8 +13,7 @@ from libs.common.ids import issuer_id_from_cik, new_job_run_id, symbol_id_from_t
from libs.common.logging import bind_job_run_id, configure_logging, get_logger
from libs.db.models import IssuerMaster, JobRun, SymbolMaster
from libs.db.session import get_session
from libs.oracle_client.client import make_oracle_client
from libs.oracle_client.financial import FinancialService
from libs.oracle_client import FinancialService, make_oracle_client
logger = get_logger(__name__)

@ -14,8 +14,7 @@ from libs.common.ids import new_job_run_id
from libs.common.logging import bind_job_run_id, configure_logging, get_logger
from libs.db.models import JobRun, MacroObservation, MacroSeries, SyncCheckpoint
from libs.db.session import get_session
from libs.oracle_client.client import make_oracle_client
from libs.oracle_client.fred import FredService
from libs.oracle_client import FredService, make_oracle_client
logger = get_logger(__name__)

@ -13,8 +13,7 @@ from libs.common.ids import new_job_run_id
from libs.common.logging import bind_job_run_id, configure_logging, get_logger
from libs.db.models import JobRun, ShortSaleDaily, SyncCheckpoint
from libs.db.session import get_session
from libs.oracle_client.client import make_oracle_client
from libs.oracle_client.finra import FinraService
from libs.oracle_client import FinraService, make_oracle_client
logger = get_logger(__name__)

@ -0,0 +1,314 @@
"""Coverage Check: Validate Stock Oracle data availability for symbol universe.
Checks each ticker for:
1. Price data coverage (daily bars since start_date, require 90%+ trading days)
2. Filing data (8-K filings, require 3+ filings)
3. Company info (name, CIK, sector must exist)
4. Market cap range ($500M - $10B for small/mid-cap)
Outputs CSV report + pass/fail summary. Optionally writes a filtered YAML
with only passing tickers.
Usage:
python -m dev.analysis.coverage_check \
--symbols-file configs/symbols_smallmid.yaml \
--start-date 2022-07-01 \
[--output-dir ./data/analysis] \
[--write-filtered]
"""
from __future__ import annotations
import argparse
import asyncio
import csv
from datetime import datetime
from pathlib import Path
from typing import Any
import yaml
from libs.oracle_client import FilingsService, FinancialService, PriceService, make_oracle_client
# --- Constants ---
MIN_MARKET_CAP = 5e8 # $500M
MAX_MARKET_CAP = 1e10 # $10B
MIN_FILINGS = 3
MIN_PRICE_COVERAGE = 0.90 # 90% of expected trading days
APPROX_TRADING_DAYS_PER_YEAR = 252
def _load_symbols(path: str) -> list[str]:
with open(path) as f:
cfg = yaml.safe_load(f) or {}
symbols = cfg.get("symbols", [])
# Deduplicate while preserving order
seen: set[str] = set()
unique: list[str] = []
for s in symbols:
if s not in seen:
seen.add(s)
unique.append(s)
return unique
async def _check_ticker(
ticker: str,
price_svc: PriceService,
filings_svc: FilingsService,
financial_svc: FinancialService,
start_date: str,
expected_trading_days: int,
) -> dict[str, Any]:
"""Run all coverage checks for a single ticker."""
result: dict[str, Any] = {
"ticker": ticker,
"price_bars": 0,
"price_coverage": 0.0,
"price_pass": False,
"filing_count": 0,
"filing_pass": False,
"has_name": False,
"has_cik": False,
"has_sector": False,
"sector": None,
"info_pass": False,
"market_cap": None,
"market_cap_pass": False,
"overall_pass": False,
"error": None,
}
# 1. Price data
try:
end_date = datetime.now().strftime("%Y-%m-%d")
price_resp = await price_svc.get_daily_bars(ticker, start=start_date, end=end_date)
n_bars = len(price_resp.bars)
result["price_bars"] = n_bars
coverage = n_bars / expected_trading_days if expected_trading_days > 0 else 0
result["price_coverage"] = round(coverage, 3)
result["price_pass"] = coverage >= MIN_PRICE_COVERAGE
except Exception as e:
result["error"] = f"price: {e}"
# 2. Filings (8-K)
try:
filings_resp = await filings_svc.search_filings(
ticker, form_type="8-K", start_date=start_date
)
n_filings = len(filings_resp.filings)
result["filing_count"] = n_filings
result["filing_pass"] = n_filings >= MIN_FILINGS
except Exception as e:
err = f"filings: {e}"
result["error"] = err if not result["error"] else f"{result['error']}; {err}"
# 3. Company info + market cap
try:
info = await financial_svc.get_company_info(ticker)
result["has_name"] = bool(info.name)
result["has_cik"] = bool(info.cik)
result["has_sector"] = bool(info.sector)
result["sector"] = info.sector
result["info_pass"] = all([info.name, info.cik, info.sector])
if info.market_cap is not None:
result["market_cap"] = info.market_cap
result["market_cap_pass"] = MIN_MARKET_CAP <= info.market_cap <= MAX_MARKET_CAP
else:
# If market_cap not available, pass this check (can't verify)
result["market_cap_pass"] = True
except Exception as e:
err = f"info: {e}"
result["error"] = err if not result["error"] else f"{result['error']}; {err}"
# If company info fails entirely, market_cap check is inconclusive
result["market_cap_pass"] = True
# Overall
result["overall_pass"] = all([
result["price_pass"],
result["filing_pass"],
result["info_pass"],
result["market_cap_pass"],
])
return result
async def run_coverage_check(
symbols: list[str],
start_date: str,
output_dir: Path,
write_filtered: bool = False,
symbols_file: str | None = None,
concurrency: int = 5,
) -> None:
"""Run coverage check for all symbols."""
output_dir.mkdir(parents=True, exist_ok=True)
# Calculate expected trading days
start_dt = datetime.strptime(start_date, "%Y-%m-%d")
years_elapsed = (datetime.now() - start_dt).days / 365.25
expected_trading_days = int(years_elapsed * APPROX_TRADING_DAYS_PER_YEAR)
print(f"\n{'='*80}")
print(f"Coverage Check — {len(symbols)} tickers, start_date={start_date}")
print(f"Expected trading days: ~{expected_trading_days}, concurrency={concurrency}")
print(f"{'='*80}\n")
results: list[dict[str, Any]] = []
sem = asyncio.Semaphore(concurrency)
counter = {"done": 0}
async def _check_with_sem(ticker: str) -> dict[str, Any]:
async with sem:
result = await _check_ticker(
ticker, price_svc, filings_svc, financial_svc,
start_date, expected_trading_days,
)
counter["done"] += 1
status = "PASS" if result["overall_pass"] else "FAIL"
mcap_str = f"${result['market_cap']/1e6:.0f}M" if result["market_cap"] else "N/A"
print(
f" [{counter['done']:3d}/{len(symbols)}] {ticker:<6s} "
f"price={result['price_bars']:>4d} bars ({result['price_coverage']:.0%}) "
f"filings={result['filing_count']:>3d} "
f"mcap={mcap_str:>8s} "
f"-> {status}"
)
if result["error"]:
print(f" ! {result['error']}")
return result
async with make_oracle_client() as client:
price_svc = PriceService(client)
filings_svc = FilingsService(client)
financial_svc = FinancialService(client)
tasks = [_check_with_sem(ticker) for ticker in symbols]
results_unordered = await asyncio.gather(*tasks)
# Preserve original symbol order
ticker_to_result = {r["ticker"]: r for r in results_unordered}
results = [ticker_to_result[s] for s in symbols]
# Write CSV
csv_path = output_dir / "coverage_check.csv"
if results:
fieldnames = list(results[0].keys())
with open(csv_path, "w", newline="") as f:
writer = csv.DictWriter(f, fieldnames=fieldnames)
writer.writeheader()
writer.writerows(results)
# Summary
passed = [r for r in results if r["overall_pass"]]
failed = [r for r in results if not r["overall_pass"]]
print(f"\n{'='*80}")
print(f"Summary")
print(f"{'='*80}")
print(f" Total: {len(results)}")
print(f" Passed: {len(passed)}")
print(f" Failed: {len(failed)}")
if failed:
# Breakdown of failure reasons
price_fail = sum(1 for r in failed if not r["price_pass"])
filing_fail = sum(1 for r in failed if not r["filing_pass"])
info_fail = sum(1 for r in failed if not r["info_pass"])
mcap_fail = sum(1 for r in failed if not r["market_cap_pass"])
print(f"\n Failure breakdown:")
print(f" Price coverage < {MIN_PRICE_COVERAGE:.0%}: {price_fail}")
print(f" Filings < {MIN_FILINGS}: {filing_fail}")
print(f" Missing company info: {info_fail}")
print(f" Market cap out of range: {mcap_fail}")
print(f"\n Failed tickers: {', '.join(r['ticker'] for r in failed)}")
# Sector distribution of passed tickers
sector_counts: dict[str, int] = {}
for r in passed:
sector = r.get("sector") or "Unknown"
sector_counts[sector] = sector_counts.get(sector, 0) + 1
if sector_counts:
print(f"\n Sector distribution (passed):")
for sector, count in sorted(sector_counts.items(), key=lambda x: -x[1]):
print(f" {sector:<30s} {count:>3d}")
print(f"\n CSV report: {csv_path}")
# Write filtered YAML
if write_filtered and passed:
filtered_path = output_dir / "symbols_smallmid_filtered.yaml"
filtered_symbols = [r["ticker"] for r in passed]
with open(filtered_path, "w") as f:
yaml.dump(
{"symbols": filtered_symbols},
f,
default_flow_style=False,
sort_keys=False,
)
print(f" Filtered YAML ({len(filtered_symbols)} tickers): {filtered_path}")
if symbols_file:
print(f"\n To use filtered list:")
print(f" cp {filtered_path} {symbols_file}")
# Go/No-Go
print(f"\n{'='*80}")
if len(passed) >= 150:
print(f"GO: {len(passed)} tickers passed (>= 150 threshold)")
elif len(passed) >= 100:
print(f"MARGINAL: {len(passed)} tickers passed (100-150 range, consider adding more)")
else:
print(f"NO-GO: Only {len(passed)} tickers passed (< 100, need more tickers)")
print(f"{'='*80}\n")
def main() -> None:
parser = argparse.ArgumentParser(description="Coverage Check for Symbol Universe")
parser.add_argument(
"--symbols-file",
required=True,
help="Path to symbols YAML file",
)
parser.add_argument(
"--start-date",
default="2022-07-01",
help="Start date for data coverage check (YYYY-MM-DD)",
)
parser.add_argument(
"--output-dir",
default="./data/analysis",
help="Output directory for CSV report",
)
parser.add_argument(
"--write-filtered",
action="store_true",
help="Write filtered YAML with only passing tickers",
)
parser.add_argument(
"--concurrency",
type=int,
default=5,
help="Max concurrent API calls (default: 5)",
)
args = parser.parse_args()
symbols = _load_symbols(args.symbols_file)
if not symbols:
print(f"No symbols found in {args.symbols_file}")
return
asyncio.run(
run_coverage_check(
symbols=symbols,
start_date=args.start_date,
output_dir=Path(args.output_dir),
write_filtered=args.write_filtered,
symbols_file=args.symbols_file,
concurrency=args.concurrency,
)
)
if __name__ == "__main__":
main()

@ -0,0 +1,391 @@
"""Diagnostic Experiments: Fixed-holding & entry-timing analysis.
Experiment 1 — Stop-off + fixed holding period exit (3/5/10/20d)
Purpose: Does the raw signal carry time-based alpha?
Method: Use pre-computed fwd_return_Xd (no stop, no target, pure hold).
Report: Mean return, median, win rate, by score bucket and event type.
Experiment 2 — Reaction-day close entry vs next-open entry
Purpose: Is the overnight gap eating the edge?
Method: Compare fwd_return_Xd (next-open base) with
close_entry_return = (entry_price / event_close) * (1 + fwd_return_Xd) - 1
Constraint: Only pre-market/intraday filings (reaction_date == event_date)
are eligible for close entry; after-close filings shown separately.
Usage:
python -m dev.analysis.entry_timing_experiment \
--snapshot-dir data/datasets/snapshots/a6687401-afdd-4bb7-9bb1-4cd32c5190bb \
--split train
"""
from __future__ import annotations
import argparse
import statistics
from pathlib import Path
from typing import Any
import pyarrow.parquet as pq
from libs.backtest.scoring import compute_entry_score
FWD_HORIZONS = ["fwd_return_3d", "fwd_return_5d", "fwd_return_10d", "fwd_return_20d"]
HORIZON_LABELS = ["3d", "5d", "10d", "20d"]
def _load_rows(snapshot_dir: Path, split: str) -> list[dict[str, Any]]:
table = pq.read_table(snapshot_dir / f"{split}.parquet")
return table.to_pylist()
def _stats(values: list[float]) -> dict[str, float | None]:
if not values:
return {"mean": None, "median": None, "win_rate": None, "n": 0}
return {
"mean": statistics.mean(values),
"median": statistics.median(values),
"win_rate": sum(1 for v in values if v > 0) / len(values),
"n": len(values),
}
def _fmt(val: float | None, pct: bool = True) -> str:
if val is None:
return " N/A"
if pct:
return f"{val:+.2%}"
return f"{val:.1%}"
def _score_bucket(score: float) -> str:
if score < 0.3:
return "low(<0.3)"
if score < 0.5:
return "mid(0.3-0.5)"
if score < 0.7:
return "high(0.5-0.7)"
return "top(>=0.7)"
BUCKET_ORDER = ["low(<0.3)", "mid(0.3-0.5)", "high(0.5-0.7)", "top(>=0.7)"]
def _is_intraday_filing(row: dict) -> bool:
"""Pre-market / regular-hours filing: reaction_date == event_date."""
return row.get("reaction_date") == row.get("event_date")
# ── Experiment 1 ──────────────────────────────────────────────────────
def experiment_1(rows: list[dict[str, Any]]) -> None:
"""Stop-off + fixed holding period: raw forward returns."""
print(f"\n{'='*100}")
print("EXPERIMENT 1: Stop-off + Fixed Holding Period Exit")
print(f"{'='*100}")
# Enrich with score
for r in rows:
r["_score"] = compute_entry_score(r)
r["_bucket"] = _score_bucket(r["_score"])
# ── 1a. ALL events ──
print(f"\n--- 1a. ALL events (N={len(rows)}) ---")
_print_horizon_table("ALL", rows)
# ── 1b. By score bucket ──
print(f"\n--- 1b. By score bucket ---")
by_bucket: dict[str, list] = {b: [] for b in BUCKET_ORDER}
for r in rows:
by_bucket[r["_bucket"]].append(r)
header = f"{'Bucket':<16} {'N':>5}"
for lbl in HORIZON_LABELS:
header += f" {lbl+'_mean':>8} {lbl+'_med':>8} {lbl+'_wr':>7}"
print(header)
print("-" * len(header))
for bucket in BUCKET_ORDER:
items = by_bucket[bucket]
line = f"{bucket:<16} {len(items):>5}"
for h in FWD_HORIZONS:
vals = [float(r[h]) for r in items if r.get(h) is not None]
s = _stats(vals)
line += f" {_fmt(s['mean']):>8} {_fmt(s['median']):>8} {_fmt(s['win_rate'], pct=False):>7}"
print(line)
# ── 1c. By event_type ──
print(f"\n--- 1c. By event_type ---")
by_type: dict[str, list] = {}
for r in rows:
et = r.get("event_type", "unknown")
by_type.setdefault(et, []).append(r)
header = f"{'EventType':<25} {'N':>5} {'AvgScr':>7}"
for lbl in HORIZON_LABELS:
header += f" {lbl+'_mean':>8} {lbl+'_wr':>7}"
print(header)
print("-" * len(header))
for et in sorted(by_type, key=lambda k: -len(by_type[k])):
items = by_type[et]
avg_score = statistics.mean(r["_score"] for r in items)
line = f"{et:<25} {len(items):>5} {avg_score:>7.3f}"
for h in FWD_HORIZONS:
vals = [float(r[h]) for r in items if r.get(h) is not None]
s = _stats(vals)
line += f" {_fmt(s['mean']):>8} {_fmt(s['win_rate'], pct=False):>7}"
print(line)
# ── 1d. MFE/MAE analysis (reward:risk) ──
print(f"\n--- 1d. MFE/MAE analysis (max favorable / max adverse excursion) ---")
mfe_mae_horizons = [("3d", "mfe_3d", "mae_3d"),
("5d", "mfe_5d", "mae_5d"),
("10d", "mfe_10d", "mae_10d"),
("20d", "mfe_20d", "mae_20d")]
header = f"{'Horizon':<10} {'MFE_mean':>10} {'MAE_mean':>10} {'R:R':>8} {'MFE_med':>10} {'MAE_med':>10}"
print(header)
print("-" * len(header))
for lbl, mfe_col, mae_col in mfe_mae_horizons:
mfe_vals = [float(r[mfe_col]) for r in rows if r.get(mfe_col) is not None]
mae_vals = [float(r[mae_col]) for r in rows if r.get(mae_col) is not None]
if mfe_vals and mae_vals:
mfe_mean = statistics.mean(mfe_vals)
mae_mean = statistics.mean(mae_vals)
rr = mfe_mean / abs(mae_mean) if mae_mean != 0 else 0
mfe_med = statistics.median(mfe_vals)
mae_med = statistics.median(mae_vals)
print(f"{lbl:<10} {mfe_mean:>+10.2%} {mae_mean:>+10.2%} {rr:>8.2f} {mfe_med:>+10.2%} {mae_med:>+10.2%}")
# ── 1e. Score monotonicity per horizon ──
print(f"\n--- 1e. Score monotonicity (does higher score → better return?) ---")
for h, lbl in zip(FWD_HORIZONS, HORIZON_LABELS):
bucket_means = []
for bucket in BUCKET_ORDER:
vals = [float(r[h]) for r in by_bucket[bucket] if r.get(h) is not None]
bucket_means.append(statistics.mean(vals) if vals else None)
valid_means = [m for m in bucket_means if m is not None]
monotonic = all(a <= b for a, b in zip(valid_means, valid_means[1:])) if len(valid_means) >= 2 else False
direction = "MONOTONIC" if monotonic else "NOT monotonic"
means_str = " → ".join(f"{m:+.2%}" if m is not None else "N/A" for m in bucket_means)
print(f" {lbl}: {means_str} [{direction}]")
# ── Experiment 2 ──────────────────────────────────────────────────────
def experiment_2(rows: list[dict[str, Any]]) -> None:
"""Reaction-close entry vs next-open entry."""
print(f"\n{'='*100}")
print("EXPERIMENT 2: Reaction-day Close Entry vs Next-Open Entry")
print(f"{'='*100}")
# Split by filing timing
intraday = [r for r in rows if _is_intraday_filing(r)]
after_close = [r for r in rows if not _is_intraday_filing(r)]
print(f"\nFiling timing split:")
print(f" Pre-market / intraday (reaction_date == event_date): {len(intraday)}")
print(f" After-close (reaction_date > event_date): {len(after_close)}")
# ── 2a. Overnight gap analysis (intraday filings only) ──
print(f"\n--- 2a. Overnight gap cost (intraday filings, N={len(intraday)}) ---")
gaps = []
for r in intraday:
ep = r.get("entry_price")
ec = r.get("event_close")
if ep and ec and ec > 0:
gap = ep / ec - 1.0
gaps.append(gap)
if gaps:
s = _stats(gaps)
print(f" Overnight gap (next_open / reaction_close - 1):")
print(f" Mean: {_fmt(s['mean'])}")
print(f" Median: {_fmt(s['median'])}")
print(f" Gap up rate (open > close): {_fmt(s['win_rate'], pct=False)}")
print(f" Std: {statistics.stdev(gaps):+.2%}" if len(gaps) > 1 else "")
# Distribution
gap_buckets = {"< -2%": 0, "-2% to 0%": 0, "0% to +1%": 0,
"+1% to +3%": 0, "+3% to +5%": 0, "> +5%": 0}
for g in gaps:
if g < -0.02:
gap_buckets["< -2%"] += 1
elif g < 0:
gap_buckets["-2% to 0%"] += 1
elif g < 0.01:
gap_buckets["0% to +1%"] += 1
elif g < 0.03:
gap_buckets["+1% to +3%"] += 1
elif g < 0.05:
gap_buckets["+3% to +5%"] += 1
else:
gap_buckets["> +5%"] += 1
print(f"\n Gap distribution:")
for label, count in gap_buckets.items():
pct = count / len(gaps) * 100
bar = "#" * int(pct / 2)
print(f" {label:>12}: {count:>4} ({pct:5.1f}%) {bar}")
# ── 2b. Close entry vs next-open returns (intraday only) ──
print(f"\n--- 2b. Close entry vs Next-open entry returns (intraday, N={len(intraday)}) ---")
header = f"{'Horizon':<10} {'NextOpen_mean':>13} {'NextOpen_wr':>11} {'CloseEntry_mean':>15} {'CloseEntry_wr':>13} {'Diff':>8}"
print(header)
print("-" * len(header))
for h, lbl in zip(FWD_HORIZONS, HORIZON_LABELS):
next_open_vals = []
close_entry_vals = []
for r in intraday:
fwd = r.get(h)
ep = r.get("entry_price")
ec = r.get("event_close")
if fwd is not None and ep and ec and ec > 0:
fwd_f = float(fwd)
next_open_vals.append(fwd_f)
# close_entry_return = (entry_price / event_close) * (1 + fwd) - 1
close_ret = (ep / ec) * (1.0 + fwd_f) - 1.0
close_entry_vals.append(close_ret)
no_s = _stats(next_open_vals)
ce_s = _stats(close_entry_vals)
diff = (ce_s["mean"] - no_s["mean"]) if ce_s["mean"] is not None and no_s["mean"] is not None else None
diff_str = _fmt(diff) if diff is not None else " N/A"
print(f"{lbl:<10} {_fmt(no_s['mean']):>13} {_fmt(no_s['win_rate'], pct=False):>11}"
f" {_fmt(ce_s['mean']):>15} {_fmt(ce_s['win_rate'], pct=False):>13} {diff_str:>8}")
# ── 2c. After-close filings (separate analysis) ──
if after_close:
print(f"\n--- 2c. After-close filings (N={len(after_close)}) ---")
print(f" (These events filed after market close; reaction_date = next trading day)")
header = f"{'Horizon':<10} {'NextOpen_mean':>13} {'NextOpen_wr':>11}"
print(f" {header}")
print(f" {'-' * len(header)}")
for h, lbl in zip(FWD_HORIZONS, HORIZON_LABELS):
vals = [float(r[h]) for r in after_close if r.get(h) is not None]
s = _stats(vals)
print(f" {lbl:<10} {_fmt(s['mean']):>13} {_fmt(s['win_rate'], pct=False):>11}")
# ── 2d. Close entry by event type (intraday only) ──
print(f"\n--- 2d. Close-entry returns by event type (intraday only) ---")
by_type: dict[str, list] = {}
for r in intraday:
et = r.get("event_type", "unknown")
by_type.setdefault(et, []).append(r)
header = f"{'EventType':<25} {'N':>5} {'Gap_mean':>9}"
for lbl in HORIZON_LABELS:
header += f" {lbl+'_CE':>8} {lbl+'_NO':>8}"
print(header)
print("-" * len(header))
for et in sorted(by_type, key=lambda k: -len(by_type[k])):
items = by_type[et]
# Gap
type_gaps = []
for r in items:
ep = r.get("entry_price")
ec = r.get("event_close")
if ep and ec and ec > 0:
type_gaps.append(ep / ec - 1.0)
gap_mean = statistics.mean(type_gaps) if type_gaps else None
line = f"{et:<25} {len(items):>5} {_fmt(gap_mean):>9}"
for h in FWD_HORIZONS:
ce_vals = []
no_vals = []
for r in items:
fwd = r.get(h)
ep = r.get("entry_price")
ec = r.get("event_close")
if fwd is not None and ep and ec and ec > 0:
fwd_f = float(fwd)
no_vals.append(fwd_f)
ce_vals.append((ep / ec) * (1.0 + fwd_f) - 1.0)
ce_s = _stats(ce_vals)
no_s = _stats(no_vals)
line += f" {_fmt(ce_s['mean']):>8} {_fmt(no_s['mean']):>8}"
print(line)
# ── 2e. Score-filtered close entry (score >= 0.5, intraday) ──
filtered = [r for r in intraday if r.get("_score", 0) >= 0.5]
if filtered:
print(f"\n--- 2e. Score-filtered (>= 0.5) close entry (intraday, N={len(filtered)}) ---")
header = f"{'Horizon':<10} {'CloseEntry_mean':>15} {'CE_wr':>8} {'NextOpen_mean':>14} {'NO_wr':>8}"
print(header)
print("-" * len(header))
for h, lbl in zip(FWD_HORIZONS, HORIZON_LABELS):
ce_vals = []
no_vals = []
for r in filtered:
fwd = r.get(h)
ep = r.get("entry_price")
ec = r.get("event_close")
if fwd is not None and ep and ec and ec > 0:
fwd_f = float(fwd)
no_vals.append(fwd_f)
ce_vals.append((ep / ec) * (1.0 + fwd_f) - 1.0)
ce_s = _stats(ce_vals)
no_s = _stats(no_vals)
print(f"{lbl:<10} {_fmt(ce_s['mean']):>15} {_fmt(ce_s['win_rate'], pct=False):>8}"
f" {_fmt(no_s['mean']):>14} {_fmt(no_s['win_rate'], pct=False):>8}")
def _print_horizon_table(label: str, items: list[dict]) -> None:
header = f"{'Horizon':<10} {'Mean':>10} {'Median':>10} {'WinRate':>8} {'N':>6} {'StdDev':>10}"
print(header)
print("-" * len(header))
for h, lbl in zip(FWD_HORIZONS, HORIZON_LABELS):
vals = [float(r[h]) for r in items if r.get(h) is not None]
s = _stats(vals)
std = statistics.stdev(vals) if len(vals) > 1 else 0
print(f"{lbl:<10} {_fmt(s['mean']):>10} {_fmt(s['median']):>10}"
f" {_fmt(s['win_rate'], pct=False):>8} {s['n']:>6} {std:>10.2%}")
def main() -> None:
parser = argparse.ArgumentParser(description="Entry Timing Experiments")
parser.add_argument("--snapshot-dir", required=True)
parser.add_argument("--split", default="train")
args = parser.parse_args()
rows = _load_rows(Path(args.snapshot_dir), args.split)
if not rows:
print("No data found.")
return
print(f"Loaded {len(rows)} events from {args.split} split")
experiment_1(rows)
experiment_2(rows)
# ── Summary ──
print(f"\n{'='*100}")
print("KEY DIAGNOSTIC QUESTIONS")
print(f"{'='*100}")
print("""
Exp 1 — Time-based alpha:
• Are any fixed-horizon returns consistently positive?
• Does score monotonicity hold at any horizon?
• Which event types carry actual alpha (mean > 0, win rate > 50%)?
Exp 2 — Entry timing:
• Is the overnight gap systematically positive (eating the edge)?
• Does close-entry improve returns vs next-open?
• Is the gap effect uniform across event types, or concentrated?
""")
if __name__ == "__main__":
main()

@ -0,0 +1,141 @@
"""Fix documents and events with NULL symbol_id / issuer_id.
The filing_poller stores ticker in document_id as TICKER::{ticker}.
When issuer_sync hadn't run before filing_poller, the lookup maps were
empty and symbol_id / issuer_id were inserted as NULL.
This script:
1. Extracts ticker from document_id for NULL-symbol docs
2. Looks up correct symbol_id / issuer_id from symbol_master / issuer_master
3. Updates documents and events
Usage:
python -m apps.tools.fix_null_symbols [--dry-run]
"""
from __future__ import annotations
import argparse
import asyncio
from sqlalchemy import select, text, update
from libs.common.config import get_settings
from libs.common.logging import configure_logging, get_logger
from libs.db.models import Document, Event, IssuerMaster, SymbolMaster
from libs.db.session import get_session
logger = get_logger(__name__)
async def fix_null_symbols(dry_run: bool = False) -> dict[str, int]:
stats = {"docs_total_null": 0, "docs_fixed": 0, "docs_no_match": 0,
"events_fixed": 0, "events_no_match": 0}
async with get_session() as session:
# 1. Build ticker lookup maps from current DB
issuer_rows = await session.execute(select(IssuerMaster))
ticker_to_issuer: dict[str, str] = {}
for im in issuer_rows.scalars().all():
if im.ticker:
ticker_to_issuer[im.ticker.upper()] = im.issuer_id
symbol_rows = await session.execute(
select(SymbolMaster).where(SymbolMaster.is_primary == True) # noqa: E712
)
ticker_to_symbol: dict[str, str] = {}
for sm in symbol_rows.scalars().all():
if sm.ticker:
ticker_to_symbol[sm.ticker.upper()] = sm.symbol_id
print(f"Lookup maps: {len(ticker_to_issuer)} issuers, {len(ticker_to_symbol)} symbols")
# 2. Find documents with NULL symbol_id
null_docs = await session.execute(
select(Document).where(Document.symbol_id.is_(None))
)
docs = null_docs.scalars().all()
stats["docs_total_null"] = len(docs)
print(f"Documents with NULL symbol_id: {len(docs)}")
# 3. Extract ticker from document_id and update
no_match_tickers: set[str] = set()
for doc in docs:
# document_id format: DOC::sec::TICKER::{ticker}::{date}::{accession}
parts = doc.document_id.split("::")
if len(parts) >= 4 and parts[2] == "TICKER":
ticker = parts[3].upper()
else:
logger.warning("unexpected_doc_id_format", document_id=doc.document_id)
stats["docs_no_match"] += 1
continue
sym_id = ticker_to_symbol.get(ticker)
iss_id = ticker_to_issuer.get(ticker)
if sym_id:
doc.symbol_id = sym_id
doc.issuer_id = iss_id # may still be None if issuer doesn't exist
stats["docs_fixed"] += 1
else:
no_match_tickers.add(ticker)
stats["docs_no_match"] += 1
if no_match_tickers:
print(f"No symbol_master match for {len(no_match_tickers)} tickers: "
f"{sorted(no_match_tickers)[:20]}{'...' if len(no_match_tickers) > 20 else ''}")
# 4. Fix events with NULL symbol_id (join via primary_document_id)
null_events = await session.execute(
select(Event).where(Event.symbol_id.is_(None))
)
events = null_events.scalars().all()
print(f"Events with NULL symbol_id: {len(events)}")
for evt in events:
# Look up the parent document
doc_result = await session.execute(
select(Document).where(Document.document_id == evt.primary_document_id)
)
parent_doc = doc_result.scalar_one_or_none()
if parent_doc and parent_doc.symbol_id:
evt.symbol_id = parent_doc.symbol_id
evt.issuer_id = parent_doc.issuer_id
stats["events_fixed"] += 1
else:
stats["events_no_match"] += 1
# 5. Commit or rollback
if dry_run:
print("\n[DRY RUN] Rolling back — no changes saved.")
await session.rollback()
else:
await session.flush()
print("\nChanges committed.")
# Report
print(f"\n{'='*50}")
print("Fix NULL Symbols Report")
print(f"{'='*50}")
print(f" Documents with NULL symbol_id: {stats['docs_total_null']}")
print(f" Documents fixed: {stats['docs_fixed']}")
print(f" Documents no match: {stats['docs_no_match']}")
print(f" Events fixed: {stats['events_fixed']}")
print(f" Events no match: {stats['events_no_match']}")
print(f"{'='*50}")
return stats
def main() -> None:
parser = argparse.ArgumentParser(description="Fix NULL symbol_id in documents and events")
parser.add_argument("--dry-run", action="store_true", help="Preview changes without saving")
args = parser.parse_args()
settings = get_settings()
configure_logging(settings.log_level)
asyncio.run(fix_null_symbols(dry_run=args.dry_run))
if __name__ == "__main__":
main()

@ -0,0 +1,245 @@
"""MFE/MAE Distribution Analysis by Event Type and Horizon.
Reads EventLabel records from the database and analyzes:
- MFE/MAE distributions per event_type per horizon (3d, 5d, 10d, 20d)
- Optimal holding period per event type (where MFE peaks)
- Forward return distributions
Outputs:
- Matplotlib plots (saved to output directory)
- Summary CSV with statistics per event_type × horizon
Usage:
python -m dev.analysis.label_horizon_analysis [--output-dir ./data/analysis]
"""
from __future__ import annotations
import argparse
import asyncio
import csv
import datetime as dt
from pathlib import Path
from typing import Any
from libs.common.logging import configure_logging, get_logger
logger = get_logger(__name__)
HORIZONS = ["3d", "5d", "10d", "20d"]
async def _load_label_data() -> list[dict[str, Any]]:
"""Load EventLabel + Event data from DB."""
from sqlalchemy import select
from libs.db.models import Event, EventLabel
from libs.db.session import get_session
rows: list[dict[str, Any]] = []
async with get_session() as session:
result = await session.execute(
select(EventLabel, Event.event_type)
.join(Event, EventLabel.event_id == Event.event_id)
.where(EventLabel.label_status == "ok")
.where(EventLabel.invalid_event_for_labeling.is_(False))
)
for lbl, event_type in result.all():
rows.append({
"event_id": lbl.event_id,
"event_type": event_type,
"fwd_return_3d": float(lbl.fwd_return_3d) if lbl.fwd_return_3d is not None else None,
"fwd_return_5d": float(lbl.fwd_return_5d) if lbl.fwd_return_5d is not None else None,
"fwd_return_10d": float(lbl.fwd_return_10d) if lbl.fwd_return_10d is not None else None,
"fwd_return_20d": float(lbl.fwd_return_20d) if lbl.fwd_return_20d is not None else None,
"mfe_3d": float(lbl.mfe_3d) if lbl.mfe_3d is not None else None,
"mae_3d": float(lbl.mae_3d) if lbl.mae_3d is not None else None,
"mfe_5d": float(lbl.mfe_5d) if lbl.mfe_5d is not None else None,
"mae_5d": float(lbl.mae_5d) if lbl.mae_5d is not None else None,
"mfe_10d": float(lbl.mfe_10d) if lbl.mfe_10d is not None else None,
"mae_10d": float(lbl.mae_10d) if lbl.mae_10d is not None else None,
"mfe_20d": float(lbl.mfe_20d) if lbl.mfe_20d is not None else None,
"mae_20d": float(lbl.mae_20d) if lbl.mae_20d is not None else None,
})
return rows
def _compute_stats(values: list[float]) -> dict[str, float | None]:
"""Compute summary statistics for a list of values."""
if not values:
return {"count": 0, "mean": None, "median": None, "std": None, "min": None, "max": None}
import statistics
return {
"count": len(values),
"mean": statistics.mean(values),
"median": statistics.median(values),
"std": statistics.stdev(values) if len(values) > 1 else 0.0,
"min": min(values),
"max": max(values),
}
def analyze_and_write(rows: list[dict[str, Any]], output_dir: Path) -> None:
"""Compute statistics and generate plots."""
output_dir.mkdir(parents=True, exist_ok=True)
# Group by event_type
by_type: dict[str, list[dict[str, Any]]] = {}
for r in rows:
et = r["event_type"]
by_type.setdefault(et, []).append(r)
# Summary CSV
csv_rows: list[dict[str, Any]] = []
optimal_holding: dict[str, str] = {}
for event_type, type_rows in sorted(by_type.items()):
mfe_means: dict[str, float] = {}
for horizon in HORIZONS:
mfe_vals = [r[f"mfe_{horizon}"] for r in type_rows if r.get(f"mfe_{horizon}") is not None]
mae_vals = [r[f"mae_{horizon}"] for r in type_rows if r.get(f"mae_{horizon}") is not None]
fwd_vals = [r[f"fwd_return_{horizon}"] for r in type_rows if r.get(f"fwd_return_{horizon}") is not None]
mfe_stats = _compute_stats(mfe_vals)
mae_stats = _compute_stats(mae_vals)
fwd_stats = _compute_stats(fwd_vals)
if mfe_stats["mean"] is not None:
mfe_means[horizon] = mfe_stats["mean"]
csv_rows.append({
"event_type": event_type,
"horizon": horizon,
"n": mfe_stats["count"],
"mfe_mean": mfe_stats["mean"],
"mfe_median": mfe_stats["median"],
"mae_mean": mae_stats["mean"],
"mae_median": mae_stats["median"],
"fwd_return_mean": fwd_stats["mean"],
"fwd_return_median": fwd_stats["median"],
"fwd_return_std": fwd_stats["std"],
"reward_risk_ratio": (
mfe_stats["mean"] / abs(mae_stats["mean"])
if mfe_stats["mean"] is not None and mae_stats["mean"] and abs(mae_stats["mean"]) > 0
else None
),
})
# Optimal holding period: horizon with highest mean MFE
if mfe_means:
best_h = max(mfe_means, key=lambda h: mfe_means[h])
optimal_holding[event_type] = best_h
# Write CSV
csv_path = output_dir / "mfe_mae_summary.csv"
if csv_rows:
fieldnames = list(csv_rows[0].keys())
with open(csv_path, "w", newline="") as f:
writer = csv.DictWriter(f, fieldnames=fieldnames)
writer.writeheader()
writer.writerows(csv_rows)
logger.info("summary_csv_written", path=str(csv_path), rows=len(csv_rows))
# Print summary table
print(f"\n{'='*80}")
print(f"MFE/MAE Distribution Analysis — {len(rows)} events, {len(by_type)} event types")
print(f"{'='*80}")
for event_type in sorted(by_type):
n = len(by_type[event_type])
best_h = optimal_holding.get(event_type, "N/A")
print(f"\n {event_type} (n={n}, optimal holding={best_h})")
for horizon in HORIZONS:
matching = [r for r in csv_rows if r["event_type"] == event_type and r["horizon"] == horizon]
if matching:
m = matching[0]
mfe_str = f"{m['mfe_mean']:+.2%}" if m["mfe_mean"] is not None else "N/A"
mae_str = f"{m['mae_mean']:+.2%}" if m["mae_mean"] is not None else "N/A"
fwd_str = f"{m['fwd_return_mean']:+.2%}" if m["fwd_return_mean"] is not None else "N/A"
rr_str = f"{m['reward_risk_ratio']:.2f}" if m["reward_risk_ratio"] is not None else "N/A"
print(f" {horizon}: MFE={mfe_str} MAE={mae_str} FwdRet={fwd_str} R:R={rr_str}")
# Generate plots (if matplotlib available)
try:
_generate_plots(by_type, output_dir)
except ImportError:
logger.info("matplotlib_not_available", msg="Skipping plot generation")
print("\n(matplotlib not available — skipping plot generation)")
def _generate_plots(
by_type: dict[str, list[dict[str, Any]]],
output_dir: Path,
) -> None:
"""Generate MFE/MAE distribution plots using matplotlib."""
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
event_types = sorted(by_type.keys())
# Plot 1: MFE by horizon per event type
fig, axes = plt.subplots(
len(event_types), len(HORIZONS),
figsize=(4 * len(HORIZONS), 3 * len(event_types)),
squeeze=False,
)
fig.suptitle("MFE Distribution by Event Type and Horizon", fontsize=14)
for i, et in enumerate(event_types):
for j, h in enumerate(HORIZONS):
ax = axes[i][j]
vals = [r[f"mfe_{h}"] for r in by_type[et] if r.get(f"mfe_{h}") is not None]
if vals:
ax.hist(vals, bins=max(5, len(vals) // 3), alpha=0.7, color="steelblue", edgecolor="white")
ax.axvline(x=sum(vals) / len(vals), color="red", linestyle="--", linewidth=1)
ax.set_title(f"{et}\n{h}" if i == 0 else h, fontsize=8)
if j == 0:
ax.set_ylabel(et[:15], fontsize=8)
ax.tick_params(labelsize=6)
plt.tight_layout()
mfe_path = output_dir / "mfe_distributions.png"
fig.savefig(str(mfe_path), dpi=150)
plt.close(fig)
print(f"\nMFE plot saved: {mfe_path}")
# Plot 2: Mean MFE across horizons (optimal holding period)
fig2, ax2 = plt.subplots(figsize=(10, 6))
horizon_days = [3, 5, 10, 20]
for et in event_types:
means = []
for h in HORIZONS:
vals = [r[f"mfe_{h}"] for r in by_type[et] if r.get(f"mfe_{h}") is not None]
means.append(sum(vals) / len(vals) * 100 if vals else 0)
ax2.plot(horizon_days, means, marker="o", label=et)
ax2.set_xlabel("Horizon (trading days)")
ax2.set_ylabel("Mean MFE (%)")
ax2.set_title("Optimal Holding Period — Mean MFE by Horizon")
ax2.legend(fontsize=8, loc="upper left")
ax2.grid(True, alpha=0.3)
plt.tight_layout()
optimal_path = output_dir / "optimal_holding_period.png"
fig2.savefig(str(optimal_path), dpi=150)
plt.close(fig2)
print(f"Optimal holding plot saved: {optimal_path}")
def main() -> None:
parser = argparse.ArgumentParser(description="MFE/MAE Distribution Analysis")
parser.add_argument(
"--output-dir",
default="./data/analysis",
help="Output directory for plots and CSV (default: ./data/analysis)",
)
args = parser.parse_args()
configure_logging("info")
rows = asyncio.run(_load_label_data())
if not rows:
print("No label data found. Run the labeler pipeline first.")
return
analyze_and_write(rows, Path(args.output_dir))
if __name__ == "__main__":
main()

@ -0,0 +1,155 @@
"""Universe Builder: expand symbol universe via Stock Oracle screener API.
Fetches large-cap liquid equities from screener, merges with existing
symbols.yaml, and writes the union back.
Usage:
python -m apps.tools.universe_builder [--keep-existing] [--output configs/symbols.yaml]
"""
from __future__ import annotations
import argparse
import asyncio
from pathlib import Path
import yaml
from libs.common.config import get_settings
from libs.common.logging import configure_logging, get_logger
from libs.oracle_client import ScreenerService, make_oracle_client
logger = get_logger(__name__)
# ── Screener defaults ──────────────────────────────────────────────
SCREENER_DEFAULTS = {
"market_cap_min": 10_000_000_000, # $10B
"min_avg_volume": 1_000_000, # 1M shares/day
"exchange": "NYSE,NASDAQ",
"exclude_types": "ETF,FUND",
"price_min": 5,
}
def load_existing_symbols(path: Path) -> list[str]:
"""Load current symbols.yaml and return list of tickers."""
if not path.exists():
return []
with open(path) as f:
cfg = yaml.safe_load(f) or {}
return cfg.get("symbols", [])
def build_symbols_yaml(
existing: list[str],
screener_results: list[dict],
) -> str:
"""Build YAML content: existing-only tickers first, then screener (sorted)."""
screener_set: set[str] = set()
for item in screener_results:
ticker = item.get("symbol", "")
if ticker:
screener_set.add(ticker.upper())
existing_set = {s.upper() for s in existing}
existing_only = sorted(existing_set - screener_set)
screener_sorted = sorted(screener_set)
lines = ["symbols:"]
if existing_only:
lines.append(" # --- Existing-only (not in screener, kept) ---")
for t in existing_only:
lines.append(f" - {t}")
lines.append(f" # --- Screener ($10B+ mcap, 1M+ vol) — {len(screener_sorted)} symbols ---")
for t in screener_sorted:
lines.append(f" - {t}")
lines.append("") # trailing newline
return "\n".join(lines)
async def run_universe_builder(
keep_existing: bool,
output_path: Path,
) -> dict[str, int]:
existing = load_existing_symbols(output_path) if keep_existing else []
existing_set = {s.upper() for s in existing}
async with make_oracle_client() as client:
svc = ScreenerService(client)
screener_stocks = await svc.search_all_stocks(**SCREENER_DEFAULTS)
screener_results = [{"symbol": s.symbol, "name": s.name} for s in screener_stocks]
screener_tickers = {s.symbol.upper() for s in screener_stocks if s.symbol}
overlap = existing_set & screener_tickers
new_only = screener_tickers - existing_set
existing_only = existing_set - screener_tickers
union = existing_set | screener_tickers
stats = {
"screener": len(screener_tickers),
"existing": len(existing_set),
"overlap": len(overlap),
"existing_only": len(existing_only),
"new": len(new_only),
"total": len(union),
}
yaml_content = build_symbols_yaml(existing, screener_results)
output_path.parent.mkdir(parents=True, exist_ok=True)
output_path.write_text(yaml_content)
# Report
logger.info("universe_report", **stats)
print(f"\n{'='*50}")
print("Universe Builder Report")
print(f"{'='*50}")
print(f" Screener results: {stats['screener']:>6}")
print(f" Existing symbols: {stats['existing']:>6}")
print(f" Overlap: {stats['overlap']:>6}")
print(f" Existing-only (kept): {stats['existing_only']:>6}")
print(f" New additions: {stats['new']:>6}")
print(f" Total (union): {stats['total']:>6}")
print(f"{'='*50}")
print(f" Written to: {output_path}")
if existing_only:
print(f"\n Existing-only tickers (kept): {sorted(existing_only)}")
return stats
def main() -> None:
parser = argparse.ArgumentParser(
description="Build symbol universe from Stock Oracle screener",
)
parser.add_argument(
"--keep-existing",
action="store_true",
default=True,
help="Keep all existing symbols (default: True)",
)
parser.add_argument(
"--no-keep-existing",
action="store_false",
dest="keep_existing",
help="Replace existing symbols entirely with screener results",
)
parser.add_argument(
"--output",
default="configs/symbols.yaml",
help="Output YAML path (default: configs/symbols.yaml)",
)
args = parser.parse_args()
settings = get_settings()
configure_logging(settings.log_level)
asyncio.run(run_universe_builder(
keep_existing=args.keep_existing,
output_path=Path(args.output),
))
if __name__ == "__main__":
main()

@ -16,7 +16,6 @@ feature_flags:
pipeline:
exhibit_types:
- "EX-99.1"
- "EX-99.2"
form_types:
- "8-K"
- "6-K"

@ -7,7 +7,7 @@
"exclude_asset_types": ["ETF", "FUND"]
},
"signal": {
"score_threshold": 0.5,
"score_threshold": 0.55,
"max_candidates_per_day": 3,
"execution_timing": "next_open",
"decision_timing": "reaction_close",
@ -27,10 +27,10 @@
"stop_atr_multiplier": 3.0,
"backtest_mode": "research",
"kill_switch_cooldown_days": 20,
"kill_switch_log_only": false,
"kill_switch_log_only": true,
"veto_oneoff_penalty": 0.7,
"veto_parse_confidence_min": 0.4,
"veto_unknown_direction": true,
"veto_unknown_direction": false,
"veto_bearish_direction": true
},
"execution": {
@ -43,10 +43,10 @@
"target_1_r": 2.0,
"target_atr_multiplier": 1.5,
"target_1_fraction": 0.5,
"trailing_model": "pct_3",
"trailing_model": "pct_5",
"trailing_warmup_days": 2,
"max_holding_days": 10,
"no_follow_through_exit": true
"max_holding_days": 15,
"no_follow_through_exit": false
},
"reporting": {
"write_trade_blotter": true,
@ -59,20 +59,25 @@
"earnings_release": {
"enabled": true,
"max_holding_days_override": 15,
"direction_filter": "bullish_only"
"direction_filter": "any"
},
"guidance_update": {
"enabled": true,
"max_holding_days_override": 10,
"direction_filter": "bullish_only"
"max_holding_days_override": 15,
"direction_filter": "any"
},
"management_change": {
"enabled": false
"enabled": true,
"direction_filter": "any"
},
"material_contract": {
"enabled": true,
"max_holding_days_override": 5,
"direction_filter": "bullish_only"
"max_holding_days_override": 10,
"direction_filter": "any"
},
"unknown": {
"enabled": true,
"direction_filter": "any"
},
"other_material_event": {
"enabled": false

@ -0,0 +1,86 @@
{
"strategy_name": "smallmid_swing_v1",
"dataset_snapshot_id": "pending_phase7",
"universe": {
"min_price": 5.0,
"min_avg_dollar_volume": 300000.0,
"exclude_asset_types": ["ETF", "FUND"]
},
"signal": {
"score_threshold": 0.55,
"max_candidates_per_day": 3,
"execution_timing": "next_open",
"decision_timing": "reaction_close",
"ranking_fields": ["score", "avg_dollar_volume"]
},
"risk": {
"per_trade_risk_pct": 0.005,
"max_daily_new_risk_pct": 0.015,
"max_positions": 4,
"max_positions_per_sector": 2,
"max_position_value_pct": 0.10,
"max_adv_fraction": 0.02,
"cooldown_after_loss_streak": 3,
"cooldown_days": 2,
"macro_regime_enabled": false,
"macro_sma_period": 20,
"stop_atr_multiplier": 3.0,
"backtest_mode": "research",
"kill_switch_cooldown_days": 20,
"kill_switch_log_only": true,
"veto_oneoff_penalty": 0.7,
"veto_parse_confidence_min": 0.4,
"veto_unknown_direction": false,
"veto_bearish_direction": true
},
"execution": {
"entry_fill_model": "next_open",
"exit_fill_model": "daily_bar_approximation",
"slippage_bps_base": 25.0,
"commission_per_share": 0.005,
"same_bar_priority": "stop_first_conservative",
"target_model": "atr_multiple",
"target_1_r": 2.0,
"target_atr_multiplier": 1.5,
"target_1_fraction": 0.5,
"trailing_model": "pct_5",
"trailing_warmup_days": 2,
"max_holding_days": 15,
"no_follow_through_exit": false
},
"reporting": {
"write_trade_blotter": true,
"write_equity_curve": true,
"write_metrics_summary": true,
"generate_plots": false,
"attribution_buckets": ["event_type", "sector", "score_bucket"]
},
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"max_holding_days_override": 15,
"direction_filter": "any"
},
"guidance_update": {
"enabled": true,
"max_holding_days_override": 15,
"direction_filter": "any"
},
"management_change": {
"enabled": true,
"direction_filter": "any"
},
"material_contract": {
"enabled": true,
"max_holding_days_override": 10,
"direction_filter": "any"
},
"unknown": {
"enabled": true,
"direction_filter": "any"
},
"other_material_event": {
"enabled": false
}
}
}

@ -0,0 +1,21 @@
{
"experiment_name": "earnings_smallmid_v1",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "PEAD alpha test on small/mid-cap earnings. Tests hypothesis that PEAD is stronger in less efficient market segments.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only"
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": { "score_threshold": 0.45 }
},
"tags": ["earnings-only", "smallmid", "pead"]
}

@ -0,0 +1,29 @@
{
"experiment_name": "earnings_smallmid_v2",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "PEAD small/mid-cap v2: relaxed oneoff filter (0.9), more positions (6), more candidates/day (5). Tests if conservative filters were destroying alpha.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only"
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"score_threshold": 0.45,
"max_candidates_per_day": 5
},
"risk": {
"max_positions": 6,
"max_positions_per_sector": 3,
"veto_oneoff_penalty": 0.9
}
},
"tags": ["earnings-only", "smallmid", "pead", "relaxed-filters"]
}

@ -0,0 +1,31 @@
{
"experiment_name": "earnings_smallmid_v3",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "PEAD small/mid-cap v3: oneoff filter effectively disabled, 8 max positions, 8 candidates/day. Maximum exposure to test raw PEAD signal.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only"
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"score_threshold": 0.45,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 4,
"veto_oneoff_penalty": 1.0,
"per_trade_risk_pct": 0.005,
"max_daily_new_risk_pct": 0.025
}
},
"tags": ["earnings-only", "smallmid", "pead", "max-exposure"]
}

@ -0,0 +1,16 @@
{
"experiment_name": "exec_opt_a_wider",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Execution opt A: wider stops (3.5 ATR), higher targets (2.0 ATR), wider trailing (7%), longer hold (20d). Let winners run.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": { "score_threshold": 0.5 },
"risk": { "stop_atr_multiplier": 3.5 },
"execution": {
"target_atr_multiplier": 2.0,
"trailing_model": "pct_7",
"max_holding_days": 20
}
},
"tags": ["exec-opt", "wider"]
}

@ -0,0 +1,16 @@
{
"experiment_name": "exec_opt_b_tight",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Execution opt B: tighter stops (2.5 ATR), quick targets (1.5 ATR), tight trailing (3%), short hold (10d). Cut losses fast.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": { "score_threshold": 0.5 },
"risk": { "stop_atr_multiplier": 2.5 },
"execution": {
"target_atr_multiplier": 1.5,
"trailing_model": "pct_3",
"max_holding_days": 10
}
},
"tags": ["exec-opt", "tight"]
}

@ -0,0 +1,15 @@
{
"experiment_name": "exec_opt_c_no_partial",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Execution opt C: no partial exit (full position rides to trailing/stop/max_hold). Tests if partial exit is killing winners.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": { "score_threshold": 0.5 },
"execution": {
"target_1_fraction": 0.0,
"trailing_model": "pct_5",
"max_holding_days": 15
}
},
"tags": ["exec-opt", "no-partial"]
}

@ -0,0 +1,14 @@
{
"experiment_name": "exec_opt_d_wide_trail",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Execution opt D: default stops/targets but much wider trailing (10%). Tests if trailing stop is choking winners prematurely.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": { "score_threshold": 0.5 },
"execution": {
"trailing_model": "pct_10",
"max_holding_days": 15
}
},
"tags": ["exec-opt", "wide-trail"]
}

@ -0,0 +1,17 @@
{
"experiment_name": "exec_opt_e_big_winners",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Execution opt E: big winner focus. Wider stops (3.5 ATR), high targets (2.5 ATR), wide trailing (10%), long hold (20d), small partial (33%).",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": { "score_threshold": 0.5 },
"risk": { "stop_atr_multiplier": 3.5 },
"execution": {
"target_atr_multiplier": 2.5,
"target_1_fraction": 0.33,
"trailing_model": "pct_10",
"max_holding_days": 20
}
},
"tags": ["exec-opt", "big-winners"]
}

@ -0,0 +1,29 @@
{
"experiment_name": "pead_10pct",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 10% threshold: ultra-selective, only biggest earnings reactions. Maximum signal strength.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "10pct"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_10pct_expanded",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "Expanded universe with 10% reaction threshold. Large-cap 7% reactions include noise (sector rotation, short covering). 10%+ reactions are more likely genuine PEAD drift. Quality > quantity approach.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 4,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "10pct", "longshort", "expanded", "universe-741"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_10pct_smallmid",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "PEAD 10% L+S on small/mid-cap universe. Higher bar for reaction strength on smaller stocks where 7% reactions are more common.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "10pct", "longshort", "smallmid", "nosectorlimit"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_5pct_longshort_v2",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Long+Short PEAD with 5% threshold: 8 max positions, 4 per sector, higher daily risk budget. Tests lower threshold for more trades.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.05,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 4,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "5pct", "longshort", "expanded"]
}

@ -0,0 +1,35 @@
{
"experiment_name": "pead_7pct_best_combo",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% best combo: tighter stop (2.0 ATR) + more positions (6) + no cooldown. Higher capital deployment with tighter risk control.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"stop_atr_multiplier": 2.0,
"max_positions": 6,
"max_positions_per_sector": 3,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"max_daily_new_risk_pct": 0.03,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "combo", "stop2", "uncapped"]
}

@ -0,0 +1,37 @@
{
"experiment_name": "pead_7pct_drift_a",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% with extended hold (25d) and wider target (3.0 ATR). Same stop. Tests if longer hold captures more drift.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only",
"max_holding_days_override": 25
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"execution": {
"max_holding_days": 25,
"target_atr_multiplier": 3.0
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "drift", "hold25", "wide_target"]
}

@ -0,0 +1,37 @@
{
"experiment_name": "pead_7pct_drift_a_phase6",
"dataset_snapshot_id": "46950fc5-4155-4eb0-ab1c-018774e444cc",
"description": "PEAD 7% Drift A (hold25, wide target 3.0 ATR) on Phase 6 dataset.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only",
"max_holding_days_override": 25
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"execution": {
"max_holding_days": 25,
"target_atr_multiplier": 3.0
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "drift", "hold25", "phase6"]
}

@ -0,0 +1,38 @@
{
"experiment_name": "pead_7pct_drift_b",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% with extended hold (25d), no partial exit, wider target (3.0 ATR). Full position rides the drift.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only",
"max_holding_days_override": 25
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"execution": {
"max_holding_days": 25,
"target_atr_multiplier": 3.0,
"target_1_fraction": null
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "drift", "hold25", "no_partial"]
}

@ -0,0 +1,42 @@
{
"experiment_name": "pead_7pct_drift_c",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% aggressive: 2x risk (1%), 6 max positions, hold 25d, no partial exit, wide target. Maximizes capital deployment on high-conviction PEAD trades.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only",
"max_holding_days_override": 25
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"execution": {
"max_holding_days": 25,
"target_atr_multiplier": 3.0,
"target_1_fraction": null
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.03,
"max_positions": 6,
"max_positions_per_sector": 3,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "drift", "hold25", "aggressive"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_expanded_combo_a",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "Combo A: nosectorlimit + 10pct threshold. Removes sector gate bug AND raises reaction bar to filter large-cap noise. Expects fewer but higher-quality trades.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "10pct", "longshort", "expanded", "universe-741", "nosectorlimit", "combo"]
}

@ -0,0 +1,35 @@
{
"experiment_name": "pead_7pct_expanded_combo_b",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "Combo B: nosectorlimit + 10pct threshold + stop ATR 2.0. Triple fix: remove sector bug, raise quality bar, widen stop for post-earnings volatility.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"stop_atr_multiplier": 2.0,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "10pct", "longshort", "expanded", "universe-741", "nosectorlimit", "stop2", "combo"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_expanded_combo_c",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "Combo C: ALL fixes combined. nosectorlimit + 10pct + maxcand4 + continuous scoring. Maximum selectivity on expanded universe — only top 4 candidates per day with 10%+ reaction.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 4
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "10pct", "longshort", "expanded", "universe-741", "nosectorlimit", "maxcand4", "combo"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_expanded_continuous",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "Expanded universe with continuous PEAD scoring. Multi-factor score (reaction 60%, volume 25%, gap 15%) replaces binary pass/fail. Combined with max_candidates_per_day=4 for synergy — top 4 by continuous score instead of volume ranking.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 4
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 4,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "expanded", "universe-741", "continuous-score"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_expanded_maxcand4",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "Expanded universe with max 4 candidates per day (down from 8). With 741 symbols, earnings season produces 10+ qualifying events per day — limiting to top 4 reduces concentration risk and improves signal quality.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 4
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 4,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "expanded", "universe-741", "maxcand4"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_expanded_nosectorlimit",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "Expanded universe with sector gate disabled. All symbols have sector=UNKNOWN, so max_positions_per_sector=4 was randomly blocking entries after 4 positions. Setting equal to max_positions effectively disables the gate.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "expanded", "universe-741", "nosectorlimit"]
}

@ -0,0 +1,35 @@
{
"experiment_name": "pead_7pct_expanded_stop2",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "Expanded universe with wider 2.0 ATR stop. Test split showed 75% stop exit rate — post-earnings volatility triggers 1.5 ATR stops prematurely. Wider stop allows trades to reach profit target.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 4,
"max_daily_new_risk_pct": 0.04,
"stop_atr_multiplier": 2.0,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "expanded", "universe-741", "stop2"]
}

@ -0,0 +1,32 @@
{
"experiment_name": "pead_7pct_longshort",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Long+Short PEAD: buy after +7% earnings reaction, short after -7% reaction. Both directions with volume >= 1.5x.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "any"
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_longshort_v2",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Long+Short PEAD with expanded capacity: 8 max positions, 4 per sector, higher daily risk budget. Allows both sides to coexist.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 4,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "expanded"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_longshort_v2_expanded",
"dataset_snapshot_id": "5f3c99d9-693c-48d6-88bf-644666559ca9",
"description": "PEAD 7% L+S v2 on expanded universe (~741 symbols via screener $10B+ mcap, 1M+ vol). Same signal/risk params as v2, testing trade frequency scaling.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 4,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "expanded", "universe-741"]
}

@ -0,0 +1,33 @@
{
"experiment_name": "pead_7pct_lowliq",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% with relaxed universe filter: min_avg_dollar_volume 200K (from 1M) and min_price 3 (from 5). Captures more events in same dataset.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"universe": {
"min_price": 3.0,
"min_avg_dollar_volume": 200000.0
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "lowliq", "universe_expansion"]
}

@ -0,0 +1,31 @@
{
"experiment_name": "pead_7pct_macro",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% + macro regime filter: only enter when SPY > SMA(20). Avoid PEAD trades in bear market.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"macro_regime_enabled": true,
"macro_sma_period": 20,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "macro"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_midcap",
"dataset_snapshot_id": "midcap-filtered",
"description": "PEAD 7% L+S on mid-cap universe ($2B-$10B, ~104 symbols). Academic PEAD strongest in less-followed stocks.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "midcap", "2b-10b"]
}

@ -0,0 +1,32 @@
{
"experiment_name": "pead_7pct_nft",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% + No-Follow-Through exit: if D+1 close < entry price, exit immediately. Quick cut of non-drifters.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"execution": {
"no_follow_through_exit": true
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "nft"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_nft_macro",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% combined: no-follow-through exit + macro regime filter. Double filter for quality.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"execution": {
"no_follow_through_exit": true
},
"risk": {
"macro_regime_enabled": true,
"macro_sma_period": 20,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "nft", "macro", "combined"]
}

@ -0,0 +1,32 @@
{
"experiment_name": "pead_7pct_notrail",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% with no trailing stop. Only fixed stop + target + time exit. Test if trailing is hurting by cutting winners early.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"execution": {
"trailing_model": null
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "notrail"]
}

@ -0,0 +1,32 @@
{
"experiment_name": "pead_7pct_phase6",
"dataset_snapshot_id": "46950fc5-4155-4eb0-ab1c-018774e444cc",
"description": "PEAD 7% base strategy on Phase 6 dataset for cross-validation.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only"
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "phase6", "cross_validation"]
}

@ -0,0 +1,32 @@
{
"experiment_name": "pead_7pct_short",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Short-side PEAD strategy: short after earnings with reaction <= -7% and volume >= 1.5x. Negative drift after bad earnings reactions.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bearish_only"
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "short", "7pct"]
}

@ -0,0 +1,30 @@
{
"experiment_name": "pead_7pct_smallmid",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "PEAD 7% on small/mid-cap dataset. Tests if PEAD signal works on smaller stocks (previously failed with default scoring).",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"stop_atr_multiplier": 3.0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "smallmid"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_smallmid_v2",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "PEAD 7% L+S on small/mid-cap universe (~185 symbols). v2 capacity settings (8 max pos, no sector limit). Academic PEAD strongest in less-followed stocks.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "smallmid", "nosectorlimit"]
}

@ -0,0 +1,30 @@
{
"experiment_name": "pead_7pct_stop2",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% with tighter stop (2.0 ATR instead of 3.0). Reduce avg loss per trade.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"stop_atr_multiplier": 2.0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "stop2"]
}

@ -0,0 +1,30 @@
{
"experiment_name": "pead_7pct_stop4",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% with wider stop (4.0 ATR instead of 3.0). Give trades more room to breathe.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"stop_atr_multiplier": 4.0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "stop4"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_7pct_uncapped",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% with more capacity: 6 max positions, no cooldown, 3% daily risk budget. More capital deployed.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"max_positions": 6,
"max_positions_per_sector": 3,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"max_daily_new_risk_pct": 0.03,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "uncapped"]
}

@ -0,0 +1,29 @@
{
"experiment_name": "pead_7pct_vol1x",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "PEAD 7% with lower volume threshold (1.0x instead of 1.5x). More trades by relaxing volume filter.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.0,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "vol1x"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_midcap_combo_10pct_maxcand3",
"dataset_snapshot_id": "midcap-filtered",
"description": "Combo: Step3 (10% threshold) + Step5 (max 3 candidates/day). Both are signal quality filters — strongest midcap PEAD signals only.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 3
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "10pct", "longshort", "midcap", "combo", "maxcand3"]
}

@ -0,0 +1,39 @@
{
"experiment_name": "pead_midcap_step10_short",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 10: Combo base (10% + maxcand3) + shorter hold (7 days). Avg holding is 3.4 days, so 15-day max is unnecessary exposure. 7 days captures PEAD drift without mean reversion.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 7},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 7
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step10", "short", "combo"],
"notes": null
}

@ -0,0 +1,36 @@
{
"experiment_name": "pead_midcap_step11_score60",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 11: Combo base + higher score threshold (0.60). All prior experiments used 0.50 (below default 0.55). Raising to 0.60 should filter weak signals and improve trade quality.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any"},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.60,
"max_candidates_per_day": 3
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step11", "score60", "combo"],
"notes": null
}

@ -0,0 +1,36 @@
{
"experiment_name": "pead_midcap_step12_vol2x",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 12: Combo base + volume threshold 2.0x (from 1.5x). Require stronger volume conviction on reaction day. Higher volume = more institutional participation = stronger drift.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any"},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 2.0,
"score_threshold": 0.5,
"max_candidates_per_day": 3
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step12", "vol2x", "combo"],
"notes": null
}

@ -0,0 +1,39 @@
{
"experiment_name": "pead_midcap_step13_best",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 13: Best combo. 10% reaction + maxcand3 + score 0.60 + short hold 7d + vol 2.0x. Combines all positive findings from steps 10-12.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 7},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 2.0,
"score_threshold": 0.60,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 7
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step13", "best", "combo"],
"notes": null
}

@ -0,0 +1,39 @@
{
"experiment_name": "pead_midcap_step14_score65",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 14: Higher score threshold 0.65 (from 0.60). Stronger signal filter to improve WR/PF with fewer trades.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 7},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 2.0,
"score_threshold": 0.65,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 7
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step14", "score65"],
"notes": null
}

@ -0,0 +1,39 @@
{
"experiment_name": "pead_midcap_step15_react7",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 15: Lower reaction threshold 0.07 (from 0.10). Wider funnel to increase trade count for statistical significance.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 7},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 2.0,
"score_threshold": 0.60,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 7
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step15", "react7"],
"notes": null
}

@ -0,0 +1,39 @@
{
"experiment_name": "pead_midcap_step16_react7_score65",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 16: Combined lower reaction 0.07 + higher score 0.65. Wider funnel filtered by stricter quality gate.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 7},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 2.0,
"score_threshold": 0.65,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 7
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step16", "react7", "score65"],
"notes": null
}

@ -0,0 +1,40 @@
{
"experiment_name": "pead_midcap_step17_target2",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 17: Target ATR 2.0 (from 1.5). Based on step14 (score 0.65). Improves R:R from 3:1.5 to 3:2.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 7},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 2.0,
"score_threshold": 0.65,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 7,
"target_atr_multiplier": 2.0
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step17", "target2"],
"notes": null
}

@ -0,0 +1,40 @@
{
"experiment_name": "pead_midcap_step18_nofrac",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 18: Full exit at target (fraction 1.0 from 0.5). Based on step14 (score 0.65). Prevents trailing stop eating profits.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 7},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 2.0,
"score_threshold": 0.65,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 7,
"target_1_fraction": 1.0
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step18", "nofrac"],
"notes": null
}

@ -0,0 +1,39 @@
{
"experiment_name": "pead_midcap_step19_hold5",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 19: Shorter hold 5d (from 7d). Based on step14 (score 0.65). Avg hold is 3.28d so 5d should be sufficient.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 5},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 2.0,
"score_threshold": 0.65,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 5
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step19", "hold5"],
"notes": null
}

@ -0,0 +1,38 @@
{
"experiment_name": "pead_midcap_step1_fixedr",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 1: R:R fix. Switch to fixed_r target_1_r=2.0 (target=2x stop). Breakeven WR drops from 67% to 33%. Baseline uses atr_multiple 1.5 ATR target vs 3.0 ATR stop.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"execution": {
"target_model": "fixed_r",
"target_1_r": 2.0
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "midcap", "step1", "fixedr"]
}

@ -0,0 +1,40 @@
{
"experiment_name": "pead_midcap_step20_best3",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 20: Best combo. score 0.65 (step14) + nofrac 1.0 (step18) + hold 5d (step19). Combines all marginally positive Phase A+B findings.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any", "max_holding_days_override": 5},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 2.0,
"score_threshold": 0.65,
"max_candidates_per_day": 3
},
"execution": {
"max_holding_days": 5,
"target_1_fraction": 1.0
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step20", "best3", "combo"],
"notes": null
}

@ -0,0 +1,37 @@
{
"experiment_name": "pead_midcap_step2_notrail",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 2: Disable trailing stop. Baseline pct_5 trail may be too tight for midcap volatility, cutting PEAD drift early.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"execution": {
"trailing_model": null
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "midcap", "step2", "notrail"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_midcap_step3_10pct",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 3: Raise reaction threshold to 10%. Stronger signal filter, fewer trades, higher quality expected.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "10pct", "longshort", "midcap", "step3", "threshold"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_midcap_step4_longonly",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 4: Long-only. Disable short side to test if midcap short PEAD mean-reverts instead of drifting.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "bullish_only" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longonly", "midcap", "step4", "bullish"]
}

@ -0,0 +1,34 @@
{
"experiment_name": "pead_midcap_step5_maxcand3",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 5: Limit to 3 candidates per day. Remove lower-ranked signals, keep only top entries.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 3
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "midcap", "step5", "maxcand3"]
}

@ -0,0 +1,38 @@
{
"experiment_name": "pead_midcap_step6_drift",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 6: Extended hold 25 days + disable partial exit. Test if PEAD drift needs more time than 15 days in midcap.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": { "enabled": true, "direction_filter": "any" },
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 8
},
"execution": {
"max_holding_days": 25,
"target_1_fraction": null
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "7pct", "longshort", "midcap", "step6", "drift25"]
}

@ -0,0 +1,40 @@
{
"experiment_name": "pead_midcap_step7_fixedr",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 7: Combo base (10% threshold + max 3 cand) + fixed_r target 2.0. Fixes inverted R:R from atr_multiple (1.5 ATR target vs 3.0 ATR stop). With proper R:R, need only 33% WR to breakeven.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any"},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 3
},
"execution": {
"target_model": "fixed_r",
"target_1_r": 2.0
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step7", "fixedr", "combo"],
"notes": null
}

@ -0,0 +1,41 @@
{
"experiment_name": "pead_midcap_step8_nft",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 8: Step 7 base + no_follow_through_exit. Exit at D+1 close if price < entry. Cuts dead trades early before stop is hit, reducing 85% stop exit rate.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any"},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 3
},
"execution": {
"target_model": "fixed_r",
"target_1_r": 2.0,
"no_follow_through_exit": true
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step8", "nft", "fixedr", "combo"],
"notes": null
}

@ -0,0 +1,41 @@
{
"experiment_name": "pead_midcap_step9_stop2",
"dataset_snapshot_id": "midcap-filtered",
"description": "Step 9: Step 7 base + tighter stop (2.0 ATR instead of 3.0). With fixed_r target 2.0, target = 4 ATR. Tighter stop reduces avg loss, but may increase stop exit rate.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {"enabled": true, "direction_filter": "any"},
"guidance_update": {"enabled": false},
"management_change": {"enabled": false},
"material_contract": {"enabled": false},
"unknown": {"enabled": false},
"other_material_event": {"enabled": false}
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.10,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 3
},
"execution": {
"target_model": "fixed_r",
"target_1_r": 2.0
},
"risk": {
"max_positions": 8,
"max_positions_per_sector": 8,
"max_daily_new_risk_pct": 0.04,
"cooldown_after_loss_streak": 0,
"cooldown_days": 0,
"stop_atr_multiplier": 2.0,
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"splits": [],
"tags": ["pead", "midcap", "step9", "stop2", "fixedr", "combo"],
"notes": null
}

@ -0,0 +1,32 @@
{
"experiment_name": "pead_pure_5pct",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Pure PEAD strategy: buy after earnings with reaction >= 5% and volume >= 1.5x. No complex scoring, just price/volume confirmation.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only"
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.05,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "pure", "5pct"]
}

@ -0,0 +1,32 @@
{
"experiment_name": "pead_pure_7pct",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Pure PEAD strategy: buy after earnings with reaction >= 7% and volume >= 1.5x. Higher threshold = stronger signal but fewer trades.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"direction_filter": "bullish_only"
},
"guidance_update": { "enabled": false },
"management_change": { "enabled": false },
"material_contract": { "enabled": false },
"unknown": { "enabled": false },
"other_material_event": { "enabled": false }
},
"signal": {
"scoring_model": "pead",
"pead_reaction_threshold": 0.07,
"pead_volume_threshold": 1.5,
"score_threshold": 0.5,
"max_candidates_per_day": 5
},
"risk": {
"veto_oneoff_penalty": 1.0,
"veto_unknown_direction": false,
"veto_bearish_direction": false
}
},
"tags": ["pead", "pure", "7pct"]
}

@ -0,0 +1,25 @@
{
"experiment_name": "phase5_no_trail_v1",
"dataset_snapshot_id": "e684ab2c-cfd9-4d72-94de-f7cc497211b8",
"description": "Phase 5 without trailing stop to isolate ATR target + wider stop effect.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": {
"score_threshold": 0.5,
"max_candidates_per_day": 10
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.05,
"max_positions": 10,
"max_positions_per_sector": 5
},
"execution": {
"trailing_model": null,
"target_1_fraction": 1.0,
"max_holding_days": 15
}
},
"splits": [],
"tags": ["phase5", "no-trailing", "atr-targets"]
}

@ -0,0 +1,21 @@
{
"experiment_name": "phase5_v1",
"dataset_snapshot_id": "cd664594-6a6c-4e16-afef-9104ce900f33",
"description": "Phase 5 strategy improvements: ATR targets, partial exits, wider stops, event-type filtering, SUE gate, direction filter.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": {
"score_threshold": 0.5,
"max_candidates_per_day": 10
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.05,
"max_positions": 10,
"max_positions_per_sector": 5
}
},
"splits": [],
"tags": ["phase5", "atr-targets", "partial-exits", "event-type-filter", "sue-gate"],
"notes": "Phase 5 improvements over expanded_scored_v1. Uses new defaults: stop_atr_multiplier=3.0, target_model=atr_multiple, target_1_fraction=0.5, trailing pct_3 with warmup=2, event_type_profiles active."
}

@ -0,0 +1,20 @@
{
"experiment_name": "phase6_v1",
"dataset_snapshot_id": "46950fc5-4155-4eb0-ab1c-018774e444cc",
"description": "Phase 6: Open filtering gates, disable no-follow-through, kill_switch log_only, wider trailing, fix attribution.",
"base_config": "configs/backtest/defaults.json",
"overrides": {
"signal": {
"score_threshold": 0.55,
"max_candidates_per_day": 10
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.05,
"max_positions": 10,
"max_positions_per_sector": 5
}
},
"splits": [],
"tags": ["phase6", "gate-fix", "no-nft", "wider-trail", "attribution-fix"]
}

@ -0,0 +1,20 @@
{
"experiment_name": "phase7_smallmid_v1",
"dataset_snapshot_id": "9bb74bb1-fefc-412b-b234-f25087634ad1",
"description": "Phase 7: Small/mid-cap universe. Hypothesis: less coverage = more alpha from 8-K filings.",
"base_config": "configs/backtest/defaults_smallmid.json",
"overrides": {
"signal": {
"score_threshold": 0.55,
"max_candidates_per_day": 10
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.05,
"max_positions": 10,
"max_positions_per_sector": 5
}
},
"splits": [],
"tags": ["phase7", "smallmid", "universe-swap", "slippage-25bps"]
}

@ -0,0 +1,20 @@
{
"experiment_name": "phase7_smallmid_v2",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "Phase 7 v2: Small/mid-cap only dataset (largecap filtered out). Clean isolation test.",
"base_config": "configs/backtest/defaults_smallmid.json",
"overrides": {
"signal": {
"score_threshold": 0.55,
"max_candidates_per_day": 10
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.05,
"max_positions": 10,
"max_positions_per_sector": 5
}
},
"splits": [],
"tags": ["phase7", "smallmid", "isolated", "slippage-25bps"]
}

@ -0,0 +1,43 @@
{
"experiment_name": "phase7_smallmid_v3",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "Phase 7 v3: Earnings-focused. Disable unknown/material_contract/guidance_update event types. Conservative risk.",
"base_config": "configs/backtest/defaults_smallmid.json",
"overrides": {
"signal": {
"score_threshold": 0.55,
"max_candidates_per_day": 5
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.03,
"max_positions": 6,
"max_positions_per_sector": 3
},
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"max_holding_days_override": 15,
"direction_filter": "any"
},
"guidance_update": {
"enabled": false
},
"management_change": {
"enabled": true,
"direction_filter": "any"
},
"material_contract": {
"enabled": false
},
"unknown": {
"enabled": false
},
"other_material_event": {
"enabled": false
}
}
},
"splits": [],
"tags": ["phase7", "smallmid", "earnings-focused", "slippage-25bps"]
}

@ -0,0 +1,51 @@
{
"experiment_name": "phase7_smallmid_v3_slip10",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "Phase 7 v3: Earnings-focused, slippage=10bps",
"base_config": "configs/backtest/defaults_smallmid.json",
"overrides": {
"signal": {
"score_threshold": 0.55,
"max_candidates_per_day": 5
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.03,
"max_positions": 6,
"max_positions_per_sector": 3
},
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"max_holding_days_override": 15,
"direction_filter": "any"
},
"guidance_update": {
"enabled": false
},
"management_change": {
"enabled": true,
"direction_filter": "any"
},
"material_contract": {
"enabled": false
},
"unknown": {
"enabled": false
},
"other_material_event": {
"enabled": false
}
},
"execution": {
"slippage_bps_base": 10.0
}
},
"splits": [],
"tags": [
"phase7",
"smallmid",
"earnings-focused",
"slippage-10bps"
]
}

@ -0,0 +1,51 @@
{
"experiment_name": "phase7_smallmid_v3_slip15",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "Phase 7 v3: Earnings-focused, slippage=15bps",
"base_config": "configs/backtest/defaults_smallmid.json",
"overrides": {
"signal": {
"score_threshold": 0.55,
"max_candidates_per_day": 5
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.03,
"max_positions": 6,
"max_positions_per_sector": 3
},
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"max_holding_days_override": 15,
"direction_filter": "any"
},
"guidance_update": {
"enabled": false
},
"management_change": {
"enabled": true,
"direction_filter": "any"
},
"material_contract": {
"enabled": false
},
"unknown": {
"enabled": false
},
"other_material_event": {
"enabled": false
}
},
"execution": {
"slippage_bps_base": 15.0
}
},
"splits": [],
"tags": [
"phase7",
"smallmid",
"earnings-focused",
"slippage-15bps"
]
}

@ -0,0 +1,51 @@
{
"experiment_name": "phase7_smallmid_v3_slip20",
"dataset_snapshot_id": "smallmid-only-8ef47ee5",
"description": "Phase 7 v3: Earnings-focused, slippage=20bps",
"base_config": "configs/backtest/defaults_smallmid.json",
"overrides": {
"signal": {
"score_threshold": 0.55,
"max_candidates_per_day": 5
},
"risk": {
"per_trade_risk_pct": 0.01,
"max_daily_new_risk_pct": 0.03,
"max_positions": 6,
"max_positions_per_sector": 3
},
"event_type_profiles": {
"earnings_release": {
"enabled": true,
"max_holding_days_override": 15,
"direction_filter": "any"
},
"guidance_update": {
"enabled": false
},
"management_change": {
"enabled": true,
"direction_filter": "any"
},
"material_contract": {
"enabled": false
},
"unknown": {
"enabled": false
},
"other_material_event": {
"enabled": false
}
},
"execution": {
"slippage_bps_base": 20.0
}
},
"splits": [],
"tags": [
"phase7",
"smallmid",
"earnings-focused",
"slippage-20bps"
]
}

@ -1,102 +1,744 @@
symbols:
# --- Mega-cap tech (original 15) ---
- AAPL
- MSFT
- GOOGL
- AMZN
- META
- NVDA
- TSLA
- AMD
- NFLX
- CRM
- SNOW
- NET
- DDOG
- ZS
- CRWD
# --- Large-cap tech additions ---
- ADBE
- ORCL
- INTC
- QCOM
- AVGO
- MU
- PANW
- FTNT
- NOW
- WDAY
- SHOP
- SQ
- MELI
- UBER
- DASH
# --- Mid-cap tech / growth ($5B-50B — stronger PEAD) ---
# --- Existing-only (not in screener, kept) ---
- AXON
- BILL
- HUBS
- PCOR
- CFLT
- MNDY
- BLK
- BRZE
- DOCN
- DPZ
- GTLB
- HIMS
- MELI
- MNDY
- PCOR
- S
- IOT
- DOCN
- BRZE
# --- Healthcare ---
- UNH
- JNJ
- LLY
- SQ
- URI
# --- Screener ($10B+ mcap, 1M+ vol) — 727 symbols ---
- A
- AA
- AAPL
- ABBV
- PFE
- MRK
- TMO
- ABNB
- ABT
- ISRG
- DXCM
- VEEV
- HIMS
# --- Industrials ---
- ACGL
- ACM
- ACN
- ADBE
- ADI
- ADM
- ADP
- ADSK
- AEE
- AEG
- AEM
- AEP
- AER
- AES
- AFL
- AFRM
- AG
- AGI
- AGNC
- AIG
- AJG
- AKAM
- ALAB
- ALB
- ALC
- ALGN
- ALL
- ALLY
- ALNY
- AM
- AMAT
- AMCR
- AMD
- AME
- AMGN
- AMH
- AMKR
- AMRZ
- AMT
- AMX
- AMZN
- ANET
- AON
- APA
- APD
- APG
- APH
- APO
- APP
- APTV
- AR
- ARCC
- ARES
- ARM
- ARMK
- AS
- ASML
- ASTS
- ASX
- ATI
- ATO
- AU
- AVAV
- AVB
- AVGO
- AWK
- AXIA
- AXP
- AZN
- B
- BA
- BABA
- BAC
- BALL
- BAM
- BBIO
- BBVA
- BBY
- BCE
- BCS
- BDX
- BE
- BEKE
- BEN
- BF-B
- BG
- BHP
- BIDU
- BIIB
- BILI
- BJ
- BK
- BKR
- BMRN
- BMY
- BN
- BNS
- BNTX
- BP
- BR
- BRK-B
- BRO
- BSX
- BSY
- BTI
- BUD
- BWA
- BX
- C
- CAH
- CARR
- CART
- CAT
- CB
- CBRE
- CCEP
- CCI
- CCJ
- CCK
- CCL
- CDE
- CDNS
- CDW
- CEG
- CELH
- CF
- CFG
- CFLT
- CG
- CHD
- CHKP
- CHRW
- CHTR
- CHWY
- CI
- CIEN
- CL
- CLS
- CLX
- CM
- CMCSA
- CME
- CMG
- CMS
- CNC
- CNH
- CNI
- CNP
- CNQ
- COF
- COHR
- COIN
- COO
- COP
- COR
- COST
- CP
- CPNG
- CPRT
- CPT
- CRBG
- CRCL
- CRDO
- CRH
- CRM
- CRWD
- CRWV
- CSCO
- CSGP
- CSX
- CTAS
- CTRA
- CTSH
- CTVA
- CUK
- CVE
- CVNA
- CVS
- CVX
- CX
- D
- DAL
- DASH
- DB
- DD
- DDOG
- DE
- GE
- HON
- RTX
- LMT
- DECK
- DELL
- DEO
- DG
- DHI
- DHR
- DINO
- DIS
- DKNG
- DKS
- DLR
- DLTR
- DOC
- DOV
- DOW
- DRI
- DRS
- DT
- DTE
- DUK
- DVA
- DVN
- DXCM
- EA
- EBAY
- EC
- ECL
- ED
- EFX
- EHC
- EIX
- EL
- ELAN
- ELS
- ELV
- EMBJ
- EMR
- ENB
- ENTG
- EOG
- EPD
- EQH
- EQNR
- EQR
- EQT
- ERIC
- ES
- ET
- ETN
- URI
- AXON
# --- Consumer ---
- COST
- WMT
- MCD
- SBUX
- NKE
- TGT
- ETR
- EVRG
- EW
- EWBC
- EXAS
- EXC
- EXE
- EXEL
- EXPD
- EXPE
- EXR
- F
- FANG
- FAST
- FCX
- FDX
- FE
- FER
- FERG
- FHN
- FIG
- FIS
- FISV
- FITB
- FIVE
- FLEX
- FLUT
- FNF
- FOX
- FOXA
- FSLR
- FTAI
- FTI
- FTNT
- FTV
- FUTU
- FWONK
- GD
- GDDY
- GE
- GEHC
- GEN
- GEV
- GFI
- GFL
- GFS
- GH
- GIL
- GILD
- GIS
- GLPI
- GLW
- GM
- GMAB
- GME
- GMED
- GNRC
- GOOG
- GOOGL
- GPC
- GPN
- GRMN
- GS
- GSK
- GWRE
- HAL
- HAS
- HBAN
- HCA
- HD
- HDB
- HIG
- HL
- HLN
- HLT
- HMC
- HOLX
- HON
- HOOD
- HPE
- HPQ
- HRL
- HSBC
- HST
- HSY
- HTHT
- HUBS
- HUM
- HWM
- IAG
- IBKR
- IBM
- IBN
- ICE
- IFF
- ILMN
- INCY
- INFY
- ING
- INSM
- INTC
- INTU
- INVH
- IONQ
- IONS
- IOT
- IP
- IQV
- IR
- IREN
- IRM
- ISRG
- IT
- ITUB
- ITW
- IVZ
- JAZZ
- JBHT
- JBL
- JBS
- JCI
- JD
- JHX
- JNJ
- JPM
- KDP
- KEY
- KEYS
- KGC
- KHC
- KIM
- KKR
- KLAC
- KMB
- KMI
- KO
- KR
- KRMN
- KT
- KTOS
- KVUE
- LDOS
- LEN
- LHX
- LI
- LIN
- LITE
- LLY
- LMT
- LNG
- LNT
- LOGI
- LOW
- LRCX
- LSCC
- LTM
- LULU
- DPZ
# --- Financials ---
- JPM
- GS
- V
- LUV
- LVS
- LYB
- LYG
- LYV
- MA
- AXP
- BLK
- COIN
- HOOD
- SOFI
- NU
# --- Energy ---
- XOM
- CVX
- COP
- EOG
- FANG
# --- Communication / Media ---
- DIS
- CMCSA
- MAA
- MAR
- MAS
- MCD
- MCHP
- MCO
- MDB
- MDLN
- MDLZ
- MDT
- MET
- META
- MFC
- MFG
- MGA
- MGM
- MKC
- MKSI
- MMM
- MNST
- MO
- MP
- MPC
- MPLX
- MRK
- MRNA
- MRSH
- MRVL
- MS
- MSFT
- MSI
- MSTR
- MT
- MTB
- MTSI
- MU
- MUFG
- NBIS
- NBIX
- NDAQ
- NEE
- NEM
- NET
- NFLX
- SPOT
- NI
- NIO
- NKE
- NLY
- NMR
- NOK
- NOW
- NRG
- NSC
- NTAP
- NTNX
- NTR
- NTRA
- NTRS
- NU
- NUE
- NVDA
- NVO
- NVS
- NVT
- NWG
- NWS
- NWSA
- NXPI
- NXT
- NYT
- O
- ODFL
- OHI
- OKE
- OKTA
- OMC
- ON
- ONON
- ORCL
- ORLY
- OTIS
- OVV
- OWL
- OXY
- PAA
- PAAS
- PANW
- PAYX
- PBA
- PBR
- PBR-A
- PCAR
- PCG
- PDD
- PEG
- PEN
- PEP
- PFE
- PFG
- PFGC
- PG
- PGR
- PHM
- PINS
- PLD
- PLTR
- PM
- PNC
- PNFP
- PNR
- PNW
- PPG
- PPL
- PR
- PRU
- PSA
- PSKY
- PSTG
- PSX
- PTC
- PWR
- PYPL
- Q
- QCOM
- QSR
- QXO
- RBA
- RBLX
- RBRK
- RCI
- RCL
- RDDT
- RDY
- REG
- RELX
- RF
- RIO
- RIVN
- RJF
- RKLB
- RKT
- RMBS
- RMD
- ROIV
- ROKU
- ROL
- ROP
- ROST
- RPM
- RPRX
- RRC
- RRX
- RSG
- RTO
- RTX
- RVMD
- RY
- RYAAY
- SAN
- SAP
- SATS
- SBS
- SBUX
- SCCO
- SCHW
- SCI
- SE
- SF
- SGI
- SHEL
- SHOP
- SHW
- SJM
- SKM
- SLB
- SMCI
- SMFG
- SMMT
- SN
- SNDK
- SNOW
- SNPS
- SNY
- SO
- SOFI
- SOLS
- SOLV
- SONY
- SPG
- SPGI
- SPOT
- SQM
- SRE
- SSNC
- STLA
- STLD
- STM
- STT
- STX
- STZ
- SU
- SUNB
- SUZ
- SW
- SWK
- SYF
- SYK
- SYM
- SYY
- T
- TAK
- TCOM
- TD
- TEAM
- TECK
- TEL
- TER
- TEVA
- TFC
- TGT
- TIGO
- TJX
- TKO
- TME
- TMO
- TMUS
- TOL
- TOST
- TPG
- TPR
- TRGP
- TRI
- TRMB
- TROW
- TRP
- TRU
- TRV
- TS
- TSCO
- TSEM
- TSLA
- TSM
- TSN
- TT
- TTD
- TTE
- TTWO
- TU
- TW
- TWLO
- TXN
- TXT
- UAL
- UBER
- UBS
- UDR
- UL
- ULS
- UMC
- UNH
- UNM
- UNP
- UPS
- USB
- USFD
- V
- VALE
- VEEV
- VG
- VICI
- VIK
- VLO
- VLTO
- VMC
- VNOM
- VOD
- VRSK
- VRT
- VRTX
- VST
- VTR
- VTRS
- VZ
- WAT
- WBD
- WBS
- WCN
- WDAY
- WDC
- WDS
- WEC
- WELL
- WES
- WFC
- WLK
- WM
- WMB
- WMG
- WMT
- WPC
- WPM
- WRB
- WSM
- WTRG
- WY
- WYNN
- XEL
- XOM
- XPEV
- XPO
- XYL
- XYZ
- YPF
- YUM
- YUMC
- Z
- ZBH
- ZG
- ZM
- ZS
- ZTO
- ZTS

@ -0,0 +1,744 @@
symbols:
# --- Existing-only (not in screener, kept) ---
- AXON
- BILL
- BLK
- BRZE
- DOCN
- DPZ
- GTLB
- HIMS
- MELI
- MNDY
- PCOR
- S
- SQ
- URI
# --- Screener ($10B+ mcap, 1M+ vol) — 727 symbols ---
- A
- AA
- AAPL
- ABBV
- ABNB
- ABT
- ACGL
- ACM
- ACN
- ADBE
- ADI
- ADM
- ADP
- ADSK
- AEE
- AEG
- AEM
- AEP
- AER
- AES
- AFL
- AFRM
- AG
- AGI
- AGNC
- AIG
- AJG
- AKAM
- ALAB
- ALB
- ALC
- ALGN
- ALL
- ALLY
- ALNY
- AM
- AMAT
- AMCR
- AMD
- AME
- AMGN
- AMH
- AMKR
- AMRZ
- AMT
- AMX
- AMZN
- ANET
- AON
- APA
- APD
- APG
- APH
- APO
- APP
- APTV
- AR
- ARCC
- ARES
- ARM
- ARMK
- AS
- ASML
- ASTS
- ASX
- ATI
- ATO
- AU
- AVAV
- AVB
- AVGO
- AWK
- AXIA
- AXP
- AZN
- B
- BA
- BABA
- BAC
- BALL
- BAM
- BBIO
- BBVA
- BBY
- BCE
- BCS
- BDX
- BE
- BEKE
- BEN
- BF-B
- BG
- BHP
- BIDU
- BIIB
- BILI
- BJ
- BK
- BKR
- BMRN
- BMY
- BN
- BNS
- BNTX
- BP
- BR
- BRK-B
- BRO
- BSX
- BSY
- BTI
- BUD
- BWA
- BX
- C
- CAH
- CARR
- CART
- CAT
- CB
- CBRE
- CCEP
- CCI
- CCJ
- CCK
- CCL
- CDE
- CDNS
- CDW
- CEG
- CELH
- CF
- CFG
- CFLT
- CG
- CHD
- CHKP
- CHRW
- CHTR
- CHWY
- CI
- CIEN
- CL
- CLS
- CLX
- CM
- CMCSA
- CME
- CMG
- CMS
- CNC
- CNH
- CNI
- CNP
- CNQ
- COF
- COHR
- COIN
- COO
- COP
- COR
- COST
- CP
- CPNG
- CPRT
- CPT
- CRBG
- CRCL
- CRDO
- CRH
- CRM
- CRWD
- CRWV
- CSCO
- CSGP
- CSX
- CTAS
- CTRA
- CTSH
- CTVA
- CUK
- CVE
- CVNA
- CVS
- CVX
- CX
- D
- DAL
- DASH
- DB
- DD
- DDOG
- DE
- DECK
- DELL
- DEO
- DG
- DHI
- DHR
- DINO
- DIS
- DKNG
- DKS
- DLR
- DLTR
- DOC
- DOV
- DOW
- DRI
- DRS
- DT
- DTE
- DUK
- DVA
- DVN
- DXCM
- EA
- EBAY
- EC
- ECL
- ED
- EFX
- EHC
- EIX
- EL
- ELAN
- ELS
- ELV
- EMBJ
- EMR
- ENB
- ENTG
- EOG
- EPD
- EQH
- EQNR
- EQR
- EQT
- ERIC
- ES
- ET
- ETN
- ETR
- EVRG
- EW
- EWBC
- EXAS
- EXC
- EXE
- EXEL
- EXPD
- EXPE
- EXR
- F
- FANG
- FAST
- FCX
- FDX
- FE
- FER
- FERG
- FHN
- FIG
- FIS
- FISV
- FITB
- FIVE
- FLEX
- FLUT
- FNF
- FOX
- FOXA
- FSLR
- FTAI
- FTI
- FTNT
- FTV
- FUTU
- FWONK
- GD
- GDDY
- GE
- GEHC
- GEN
- GEV
- GFI
- GFL
- GFS
- GH
- GIL
- GILD
- GIS
- GLPI
- GLW
- GM
- GMAB
- GME
- GMED
- GNRC
- GOOG
- GOOGL
- GPC
- GPN
- GRMN
- GS
- GSK
- GWRE
- HAL
- HAS
- HBAN
- HCA
- HD
- HDB
- HIG
- HL
- HLN
- HLT
- HMC
- HOLX
- HON
- HOOD
- HPE
- HPQ
- HRL
- HSBC
- HST
- HSY
- HTHT
- HUBS
- HUM
- HWM
- IAG
- IBKR
- IBM
- IBN
- ICE
- IFF
- ILMN
- INCY
- INFY
- ING
- INSM
- INTC
- INTU
- INVH
- IONQ
- IONS
- IOT
- IP
- IQV
- IR
- IREN
- IRM
- ISRG
- IT
- ITUB
- ITW
- IVZ
- JAZZ
- JBHT
- JBL
- JBS
- JCI
- JD
- JHX
- JNJ
- JPM
- KDP
- KEY
- KEYS
- KGC
- KHC
- KIM
- KKR
- KLAC
- KMB
- KMI
- KO
- KR
- KRMN
- KT
- KTOS
- KVUE
- LDOS
- LEN
- LHX
- LI
- LIN
- LITE
- LLY
- LMT
- LNG
- LNT
- LOGI
- LOW
- LRCX
- LSCC
- LTM
- LULU
- LUV
- LVS
- LYB
- LYG
- LYV
- MA
- MAA
- MAR
- MAS
- MCD
- MCHP
- MCO
- MDB
- MDLN
- MDLZ
- MDT
- MET
- META
- MFC
- MFG
- MGA
- MGM
- MKC
- MKSI
- MMM
- MNST
- MO
- MP
- MPC
- MPLX
- MRK
- MRNA
- MRSH
- MRVL
- MS
- MSFT
- MSI
- MSTR
- MT
- MTB
- MTSI
- MU
- MUFG
- NBIS
- NBIX
- NDAQ
- NEE
- NEM
- NET
- NFLX
- NI
- NIO
- NKE
- NLY
- NMR
- NOK
- NOW
- NRG
- NSC
- NTAP
- NTNX
- NTR
- NTRA
- NTRS
- NU
- NUE
- NVDA
- NVO
- NVS
- NVT
- NWG
- NWS
- NWSA
- NXPI
- NXT
- NYT
- O
- ODFL
- OHI
- OKE
- OKTA
- OMC
- ON
- ONON
- ORCL
- ORLY
- OTIS
- OVV
- OWL
- OXY
- PAA
- PAAS
- PANW
- PAYX
- PBA
- PBR
- PBR-A
- PCAR
- PCG
- PDD
- PEG
- PEN
- PEP
- PFE
- PFG
- PFGC
- PG
- PGR
- PHM
- PINS
- PLD
- PLTR
- PM
- PNC
- PNFP
- PNR
- PNW
- PPG
- PPL
- PR
- PRU
- PSA
- PSKY
- PSTG
- PSX
- PTC
- PWR
- PYPL
- Q
- QCOM
- QSR
- QXO
- RBA
- RBLX
- RBRK
- RCI
- RCL
- RDDT
- RDY
- REG
- RELX
- RF
- RIO
- RIVN
- RJF
- RKLB
- RKT
- RMBS
- RMD
- ROIV
- ROKU
- ROL
- ROP
- ROST
- RPM
- RPRX
- RRC
- RRX
- RSG
- RTO
- RTX
- RVMD
- RY
- RYAAY
- SAN
- SAP
- SATS
- SBS
- SBUX
- SCCO
- SCHW
- SCI
- SE
- SF
- SGI
- SHEL
- SHOP
- SHW
- SJM
- SKM
- SLB
- SMCI
- SMFG
- SMMT
- SN
- SNDK
- SNOW
- SNPS
- SNY
- SO
- SOFI
- SOLS
- SOLV
- SONY
- SPG
- SPGI
- SPOT
- SQM
- SRE
- SSNC
- STLA
- STLD
- STM
- STT
- STX
- STZ
- SU
- SUNB
- SUZ
- SW
- SWK
- SYF
- SYK
- SYM
- SYY
- T
- TAK
- TCOM
- TD
- TEAM
- TECK
- TEL
- TER
- TEVA
- TFC
- TGT
- TIGO
- TJX
- TKO
- TME
- TMO
- TMUS
- TOL
- TOST
- TPG
- TPR
- TRGP
- TRI
- TRMB
- TROW
- TRP
- TRU
- TRV
- TS
- TSCO
- TSEM
- TSLA
- TSM
- TSN
- TT
- TTD
- TTE
- TTWO
- TU
- TW
- TWLO
- TXN
- TXT
- UAL
- UBER
- UBS
- UDR
- UL
- ULS
- UMC
- UNH
- UNM
- UNP
- UPS
- USB
- USFD
- V
- VALE
- VEEV
- VG
- VICI
- VIK
- VLO
- VLTO
- VMC
- VNOM
- VOD
- VRSK
- VRT
- VRTX
- VST
- VTR
- VTRS
- VZ
- WAT
- WBD
- WBS
- WCN
- WDAY
- WDC
- WDS
- WEC
- WELL
- WES
- WFC
- WLK
- WM
- WMB
- WMG
- WMT
- WPC
- WPM
- WRB
- WSM
- WTRG
- WY
- WYNN
- XEL
- XOM
- XPEV
- XPO
- XYL
- XYZ
- YPF
- YUM
- YUMC
- Z
- ZBH
- ZG
- ZM
- ZS
- ZTO
- ZTS

Some files were not shown because too many files have changed in this diff Show More

Loading…
Cancel
Save