Add Oracle event_type vocabulary normalizer for fallback path
When the rule parser can't classify an 8-K and falls back to Stock Oracle's filing-events API, Oracle's vocabulary (e.g. earnings_result, shareholder_vote, regulation_fd) was being written verbatim into events.event_type. The DB has no CHECK constraint (libs/db/models.py:189), so 22 distinct Oracle values silently leaked into a column the strategy's engine filters expect to be in its 4-event vocabulary. Result: ~1,069 live rows silently dropped from strategy candidate pool. Files: - NEW libs/parser/event_type_normalizer.py: normalize_oracle_event_type() with conservative synonym map; normalize_oracle_event() additionally uses _classify_event_type from rule_parser when an item_number is present (item-code path is more reliable than Oracle's event taxonomy) - MOD apps/pipeline/event_parser/main.py: oracle-fallback branch (~line 140) now calls normalize_oracle_event before writing to DB; emits oracle_event_type_normalized log event when value changes - NEW tests/unit/test_event_type_normalizer.py: 60 tests covering identity, synonyms, case/separator insensitivity, None/empty, non-string, item_number-precedence Mapping highlights (justifications in test docstrings): earnings_result/earnings_announcement/earnings -> earnings_release guidance_revision/guidance_change/regulation_fd -> guidance_update material_definitive_agreement/definitive_agreement -> material_contract shareholder_vote/acquisition_disposition/bankruptcy/other -> other_material_event Reg FD -> guidance_update mirrors rule_parser's Item 7.01 mapping for internal consistency. Debatable but auditable. Conservative pass-through for ambiguous values (financial_obligation, articles_amendment, contract_termination, etc., 14 distinct values). Visible filter-drop > silent re-tag. Live DB counts that would reclassify on a future --reparse pass: 412 earnings_result -> earnings_release 409 shareholder_vote -> other_material_event 237 regulation_fd -> guidance_update 7 acquisition_disposition -> other_material_event 4 other -> other_material_event TOTAL 1,069 rows currently in oracle-fallback dead-zone. 60/60 normalizer tests pass; combined parser+schema validator suite 86/86. Follow-up flagged: run --reparse on historical oracle-fallback rows after extending reparse_events() to also re-normalize known oracle-fallback values (currently only re-parses event_type='unknown'). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>main
parent
cffb8d872d
commit
722e5cf6a9
@ -0,0 +1,159 @@
|
|||||||
|
"""Normalize Stock Oracle's filing-event vocabulary to the strategy's vocabulary.
|
||||||
|
|
||||||
|
Background
|
||||||
|
----------
|
||||||
|
When the rule-based parser cannot classify an 8-K (no exhibit text, low
|
||||||
|
confidence, etc.), `apps/pipeline/event_parser/main.py` falls back to the
|
||||||
|
Oracle filings-events API. Oracle's `event_type` strings are richer than
|
||||||
|
``ParserEventOutput.event_type`` — values like ``earnings_result``,
|
||||||
|
``shareholder_vote``, ``regulation_fd``, ``financial_obligation`` get written
|
||||||
|
into the ``events.event_type`` column unchanged, and the strategy's engine
|
||||||
|
filters silently drop them.
|
||||||
|
|
||||||
|
This module provides a *conservative* mapping from Oracle's vocabulary to the
|
||||||
|
strategy's: only synonyms and clear item-code parallels are mapped. Ambiguous
|
||||||
|
values are left as-is so the failure mode is honest (row drops out of the
|
||||||
|
strategy filter) rather than silently re-tagged.
|
||||||
|
|
||||||
|
Where possible, prefer ``normalize_oracle_event(value, item_number=...)``,
|
||||||
|
which routes through the existing 8-K item-code classifier in
|
||||||
|
``libs.parser.rule_parser`` for maximum consistency with the rule-based path.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Oracle vocabulary -> strategy vocabulary
|
||||||
|
#
|
||||||
|
# Conservative mapping. Each row needs an unambiguous synonym justification.
|
||||||
|
# When in doubt we leave the value untouched and let the engine filter drop
|
||||||
|
# it — that's louder and easier to debug than a quiet re-tag.
|
||||||
|
#
|
||||||
|
# Strategy vocabulary (libs/schemas/types.py ParserEventOutput.event_type):
|
||||||
|
# earnings_release, guidance_update, material_contract, management_change,
|
||||||
|
# other_material_event, regulatory_or_approval, capital_markets_or_financing,
|
||||||
|
# litigation_or_investigation, unknown
|
||||||
|
#
|
||||||
|
# Strategy *engine filter* hot path uses 4-5 of these (libs/backtest/scoring.py
|
||||||
|
# and selector.py): earnings_release, guidance_update, material_contract,
|
||||||
|
# other_material_event, management_change.
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
_ORACLE_TO_STRATEGY: dict[str, str] = {
|
||||||
|
# --- identity passes (already aligned) ---
|
||||||
|
"earnings_release": "earnings_release",
|
||||||
|
"guidance_update": "guidance_update",
|
||||||
|
"material_contract": "material_contract",
|
||||||
|
"management_change": "management_change",
|
||||||
|
"other_material_event": "other_material_event",
|
||||||
|
|
||||||
|
# --- earnings synonyms ---
|
||||||
|
# Oracle's actual 412 mislabeled rows are "earnings_result"; "earnings_announcement"
|
||||||
|
# is added defensively for any future Oracle vocabulary drift.
|
||||||
|
"earnings_result": "earnings_release",
|
||||||
|
"earnings_announcement": "earnings_release",
|
||||||
|
"earnings": "earnings_release",
|
||||||
|
|
||||||
|
# --- guidance synonyms ---
|
||||||
|
"guidance_revision": "guidance_update",
|
||||||
|
"guidance_change": "guidance_update",
|
||||||
|
# Reg FD (Item 7.01): the rule_parser already maps Item 7.01 -> guidance_update,
|
||||||
|
# so we keep parity. Reg FD often (but not always) carries guidance; this is the
|
||||||
|
# one debatable choice and is intentional for consistency with the item-code path.
|
||||||
|
"regulation_fd": "guidance_update",
|
||||||
|
"regulation_fd_disclosure": "guidance_update",
|
||||||
|
|
||||||
|
# --- material contract synonyms (Item 1.01 wording variants) ---
|
||||||
|
"material_definitive_agreement": "material_contract",
|
||||||
|
"definitive_agreement": "material_contract",
|
||||||
|
|
||||||
|
# --- other_material_event (Item 5.07, 2.01, 1.03 etc.) ---
|
||||||
|
# Item 5.07 — submission of matters to a vote of security holders.
|
||||||
|
"shareholder_vote": "other_material_event",
|
||||||
|
# Item 2.01 — completion of acquisition or disposition of assets.
|
||||||
|
"acquisition_disposition": "other_material_event",
|
||||||
|
"acquisition": "other_material_event",
|
||||||
|
"disposition": "other_material_event",
|
||||||
|
# Item 1.03 — bankruptcy or receivership.
|
||||||
|
"bankruptcy": "other_material_event",
|
||||||
|
# Generic "Other Events" (Item 8.01) — rule_parser maps 8.01 -> other_material_event.
|
||||||
|
"other_events": "other_material_event",
|
||||||
|
"other": "other_material_event",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _canonicalize(value: str) -> str:
|
||||||
|
"""Lowercase, trim, collapse separators so 'Earnings Result', 'earnings-result',
|
||||||
|
and 'EARNINGS_RESULT' all hit the same map key."""
|
||||||
|
s = value.strip().lower()
|
||||||
|
# unify separators
|
||||||
|
s = s.replace("-", "_").replace(" ", "_").replace("/", "_")
|
||||||
|
# collapse repeated underscores
|
||||||
|
while "__" in s:
|
||||||
|
s = s.replace("__", "_")
|
||||||
|
return s.strip("_")
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_oracle_event_type(oracle_value: str | None) -> str:
|
||||||
|
"""Map an Oracle ``event_type`` string to the strategy's vocabulary.
|
||||||
|
|
||||||
|
Behavior:
|
||||||
|
- ``None`` / empty / whitespace-only input -> ``"unknown"``.
|
||||||
|
- Mapped values returned in canonical strategy form.
|
||||||
|
- Unmapped values are returned untouched (preserving original casing/wording)
|
||||||
|
so the engine filter drops them visibly rather than silently re-tagging.
|
||||||
|
|
||||||
|
Examples:
|
||||||
|
>>> normalize_oracle_event_type("earnings_result")
|
||||||
|
'earnings_release'
|
||||||
|
>>> normalize_oracle_event_type("Earnings Result")
|
||||||
|
'earnings_release'
|
||||||
|
>>> normalize_oracle_event_type("financial_obligation")
|
||||||
|
'financial_obligation'
|
||||||
|
>>> normalize_oracle_event_type(None)
|
||||||
|
'unknown'
|
||||||
|
"""
|
||||||
|
if oracle_value is None:
|
||||||
|
return "unknown"
|
||||||
|
if not isinstance(oracle_value, str):
|
||||||
|
return "unknown"
|
||||||
|
if not oracle_value.strip():
|
||||||
|
return "unknown"
|
||||||
|
|
||||||
|
key = _canonicalize(oracle_value)
|
||||||
|
if key in _ORACLE_TO_STRATEGY:
|
||||||
|
return _ORACLE_TO_STRATEGY[key]
|
||||||
|
# Honest pass-through: unmapped Oracle vocab is kept verbatim.
|
||||||
|
return oracle_value
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_oracle_event(
|
||||||
|
oracle_value: str | None,
|
||||||
|
item_number: str | None = None,
|
||||||
|
) -> str:
|
||||||
|
"""Stronger normalizer that prefers 8-K ``item_number`` when available.
|
||||||
|
|
||||||
|
The existing rule-based parser already maps SEC 8-K item codes to the
|
||||||
|
strategy's vocabulary in ``libs.parser.rule_parser._ITEM_TO_EVENT_TYPE``.
|
||||||
|
Routing Oracle-fallback rows through that classifier gives us strict
|
||||||
|
parity with the rule-based path: a 2.02-tagged Oracle row becomes
|
||||||
|
``earnings_release`` *exactly* like a rule-parsed row would.
|
||||||
|
|
||||||
|
Falls back to the string synonym map when the item code is absent or
|
||||||
|
classifies as ``"unknown"``.
|
||||||
|
"""
|
||||||
|
# Prefer item-code routing — already battle-tested via rule_parser.
|
||||||
|
if item_number:
|
||||||
|
# Local import keeps this module free of heavy parser deps when only
|
||||||
|
# the string-only API is used (tests, lightweight callers).
|
||||||
|
from libs.parser.rule_parser import _classify_event_type
|
||||||
|
|
||||||
|
mapped = _classify_event_type([item_number])
|
||||||
|
if mapped != "unknown":
|
||||||
|
return mapped
|
||||||
|
|
||||||
|
return normalize_oracle_event_type(oracle_value)
|
||||||
|
|
||||||
|
|
||||||
|
__all__ = [
|
||||||
|
"normalize_oracle_event_type",
|
||||||
|
"normalize_oracle_event",
|
||||||
|
]
|
||||||
@ -0,0 +1,177 @@
|
|||||||
|
"""Unit tests for libs.parser.event_type_normalizer.
|
||||||
|
|
||||||
|
Covers:
|
||||||
|
- Each mapped Oracle value resolves to the expected strategy value.
|
||||||
|
- Unmapped values pass through unchanged (honest fail mode for engine filters).
|
||||||
|
- Empty / None / whitespace input -> "unknown".
|
||||||
|
- Case-insensitive + separator-tolerant matching.
|
||||||
|
- item_number-preferred path (uses rule_parser._ITEM_TO_EVENT_TYPE).
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from libs.parser.event_type_normalizer import (
|
||||||
|
normalize_oracle_event,
|
||||||
|
normalize_oracle_event_type,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# String-only normalizer
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"raw,expected",
|
||||||
|
[
|
||||||
|
# Identity (already aligned with strategy vocab)
|
||||||
|
("earnings_release", "earnings_release"),
|
||||||
|
("guidance_update", "guidance_update"),
|
||||||
|
("material_contract", "material_contract"),
|
||||||
|
("management_change", "management_change"),
|
||||||
|
("other_material_event", "other_material_event"),
|
||||||
|
# Earnings synonyms — the headline mapping (412 live rows)
|
||||||
|
("earnings_result", "earnings_release"),
|
||||||
|
("earnings_announcement", "earnings_release"),
|
||||||
|
("earnings", "earnings_release"),
|
||||||
|
# Guidance synonyms
|
||||||
|
("guidance_revision", "guidance_update"),
|
||||||
|
("guidance_change", "guidance_update"),
|
||||||
|
# Reg FD -> guidance_update (matches rule_parser Item 7.01 mapping)
|
||||||
|
("regulation_fd", "guidance_update"),
|
||||||
|
("regulation_fd_disclosure", "guidance_update"),
|
||||||
|
# Material contract synonyms (Item 1.01 wording)
|
||||||
|
("material_definitive_agreement", "material_contract"),
|
||||||
|
("definitive_agreement", "material_contract"),
|
||||||
|
# other_material_event aggregations
|
||||||
|
("shareholder_vote", "other_material_event"),
|
||||||
|
("acquisition_disposition", "other_material_event"),
|
||||||
|
("acquisition", "other_material_event"),
|
||||||
|
("disposition", "other_material_event"),
|
||||||
|
("bankruptcy", "other_material_event"),
|
||||||
|
("other_events", "other_material_event"),
|
||||||
|
("other", "other_material_event"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_mapped_values_normalize_correctly(raw: str, expected: str) -> None:
|
||||||
|
assert normalize_oracle_event_type(raw) == expected
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"raw",
|
||||||
|
[
|
||||||
|
# Items intentionally left unmapped — strategy intent unclear; let the
|
||||||
|
# engine filter drop them visibly rather than silently re-tagging.
|
||||||
|
"financial_obligation", # 225 live rows
|
||||||
|
"articles_amendment", # 81 live rows
|
||||||
|
"contract_termination",
|
||||||
|
"unregistered_equity_sale",
|
||||||
|
"rights_modification",
|
||||||
|
"accountant_change",
|
||||||
|
"exit_activity",
|
||||||
|
"mine_safety",
|
||||||
|
"triggering_event",
|
||||||
|
"material_impairment",
|
||||||
|
"delisting_notice",
|
||||||
|
"bylaws_amendment",
|
||||||
|
"shareholder_nomination",
|
||||||
|
"control_change",
|
||||||
|
"completely_made_up_value",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_unmapped_values_pass_through_unchanged(raw: str) -> None:
|
||||||
|
"""Honest fail mode: don't silently re-tag ambiguous values."""
|
||||||
|
assert normalize_oracle_event_type(raw) == raw
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"raw,expected",
|
||||||
|
[
|
||||||
|
("EARNINGS_RESULT", "earnings_release"),
|
||||||
|
("Earnings Result", "earnings_release"),
|
||||||
|
("earnings-result", "earnings_release"),
|
||||||
|
("Earnings Result", "earnings_release"), # double-space
|
||||||
|
(" earnings_result ", "earnings_release"), # padding
|
||||||
|
("REGULATION_FD", "guidance_update"),
|
||||||
|
("Material Definitive Agreement", "material_contract"),
|
||||||
|
("Shareholder/Vote", "other_material_event"), # alt separator
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_case_and_separator_insensitive(raw: str, expected: str) -> None:
|
||||||
|
assert normalize_oracle_event_type(raw) == expected
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("bad", [None, "", " ", "\n", "\t"])
|
||||||
|
def test_empty_and_none_become_unknown(bad: object) -> None:
|
||||||
|
assert normalize_oracle_event_type(bad) == "unknown" # type: ignore[arg-type]
|
||||||
|
|
||||||
|
|
||||||
|
def test_non_string_becomes_unknown() -> None:
|
||||||
|
# Defensive: pydantic's typed model should prevent this, but be safe.
|
||||||
|
assert normalize_oracle_event_type(123) == "unknown" # type: ignore[arg-type]
|
||||||
|
assert normalize_oracle_event_type([]) == "unknown" # type: ignore[arg-type]
|
||||||
|
|
||||||
|
|
||||||
|
def test_pass_through_preserves_original_casing() -> None:
|
||||||
|
"""Unmapped values are returned verbatim (not lowercased) so logs match
|
||||||
|
Oracle's wire format and the DB column reflects exactly what was upstream."""
|
||||||
|
assert normalize_oracle_event_type("Financial_Obligation") == "Financial_Obligation"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# item_number-preferred path
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"item_number,expected",
|
||||||
|
[
|
||||||
|
("2.02", "earnings_release"), # Results of Operations and Financial Condition
|
||||||
|
("7.01", "guidance_update"), # Regulation FD Disclosure
|
||||||
|
("5.02", "management_change"),
|
||||||
|
("1.01", "material_contract"),
|
||||||
|
("8.01", "other_material_event"),
|
||||||
|
("1.03", "other_material_event"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_item_number_takes_precedence(item_number: str, expected: str) -> None:
|
||||||
|
"""A wildly mislabeled Oracle string is overridden by a valid item_number.
|
||||||
|
|
||||||
|
Rationale: the rule-parser already trusts item codes more than text-derived
|
||||||
|
classifications; the Oracle-fallback path should follow the same convention.
|
||||||
|
"""
|
||||||
|
assert (
|
||||||
|
normalize_oracle_event(
|
||||||
|
"completely_bogus_oracle_string",
|
||||||
|
item_number=item_number,
|
||||||
|
)
|
||||||
|
== expected
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_item_number_unknown_falls_back_to_string_map() -> None:
|
||||||
|
"""Item codes outside _ITEM_TO_EVENT_TYPE shouldn't block string mapping."""
|
||||||
|
# 9.99 isn't a real 8-K item; classifier returns "unknown", so we rely on string.
|
||||||
|
assert (
|
||||||
|
normalize_oracle_event("earnings_result", item_number="9.99")
|
||||||
|
== "earnings_release"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_item_number_uses_string_map() -> None:
|
||||||
|
assert (
|
||||||
|
normalize_oracle_event("earnings_result", item_number=None)
|
||||||
|
== "earnings_release"
|
||||||
|
)
|
||||||
|
assert (
|
||||||
|
normalize_oracle_event("earnings_result", item_number="")
|
||||||
|
== "earnings_release"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_item_number_no_match_passes_through() -> None:
|
||||||
|
assert (
|
||||||
|
normalize_oracle_event("financial_obligation", item_number=None)
|
||||||
|
== "financial_obligation"
|
||||||
|
)
|
||||||
Loading…
Reference in New Issue