proactive: mission-driven employer scope + non-ortho specialty gate

The plan targets PSLF-eligible, mission-driven employers (government,
tribal, nonprofit, FQHC/community health, academic, public hospital).
For-profit DSOs (Smile Doctors, Sonrava, Specialty Dental Brands) had
slipped in through site: queries and auto-discovered boards, and their
pages produced off-scope leads — including Pediatric Dentist titles
that were never orthodontist jobs.

- employer_scope: rule-based employer classifier (.gov/.mil/.edu/
  .nsn.us hosts, configured public domains, government/tribal/
  nonprofit/academic/FQHC terms); private/unknown employers are
  FILTERED at verification (default-deny) and logged to
  workspace/candidates/*.employers.jsonl. Board catalogs declare
  employer_type; auto-discovered boards are written only when the
  SERP evidence classifies as mission-driven.
- roles: new other_specialty class (pediatric dentist, endodontist,
  oral surgeon, general dentist …) filtered before target matching;
  hidden dentist-title lanes at public employers stay.
- verifier: normalize ATS page titles ("Job Application for X at Y"
  -> "X") before storing; apply employer gate.
- discovery: scope filter + out-of-scope reporting; private auto
  catalog cleared (University of Utah Health kept as academic).
- closed markers: drop bare "filled" — federal boilerplate ("until
  the position is filled") marked open USAJOBS postings as CLOSED.
- urls: strip default :443/:80 ports so USAJOBS links dedupe.
main
I Luk Kim 3 weeks ago
parent 0fda8383ea
commit 3b6e5cdbd4

@ -96,14 +96,23 @@ Greenhouse/Lever/SmartRecruiters/USAJOBS)를 직접 크롤링해 잡는다. 대
ATS API를 probe해 치과 키워드가 실제로 있는 보드만 등록한다. 주간 deep scan에서 자동 실행되며
(`boards.discovery`), 수동 실행도 가능하다.
검증 전에 **role gate**(제목 규칙)가 보조/비임상 직무를 걸러낸다 — Dental Assistant·Hygienist·
Coordinator·Payable 등은 수집·검증 단계에서 FILTERED. ortho 신호는 페이지 전체가 아니라
검증 전에 **role gate**(제목 규칙)가 보조/비임상/타 전문과목 직무를 걸러낸다 — Dental Assistant·
Hygienist·Coordinator·Payable, 그리고 Pediatric Dentist·Endodontist 같은 비-교정 전문과목은
수집·검증 단계에서 FILTERED. 공공기관에서 교정의 공고가 일반 타이틀("Dentist II", "Staff Dentist",
"Chief Dental Officer")로 올라오는 hidden-title 레인은 유지한다. ortho 신호는 페이지 전체가 아니라
**제목 + 공고 설명 영역**에서만 확인하므로, 치과 고용주 회사소개 문구("... & Orthodontics")로 인한
오탐이 없다. 경계 타이틀은 런 후 오프라인으로 `role-audit`(gemma4)이 자동 검토한다 —
support/non_clinical 제안은 `sites/role_terms.auto.yaml`에 자동 반영되고, target 제안은 게이트
완화 위험이 있어 리포트(`workspace/manifests/role-audit-*.md`)에만 기록된다. 런타임 수집·검증은
여전히 non-AI이며, Ollama가 없으면 audit만 조용히 건너뛴다.
**Employer scope gate** — 대상은 mission-driven 고용주(정부·트라이벌·비영리·FQHC/커뮤니티헬스·
공공병원·대학)뿐이다. 영리 사기업(DSO 등)은 리드에서 제외하며, 공공/비영리 신호가 확인되지 않으면
기본 제외(default-deny)한다 — 신호가 없으면 `workspace/candidates/*.employers.jsonl`에 기록되어
나중에 검토할 수 있다. `sites/employers.yaml`의 보드는 `employer_type`을 선언하고
(예: USAJOBS → government, University of Utah Health → academic), 자동 발견 보드는 SERP 근거에서
mission-driven으로 분류된 것만 `employers.auto.yaml`에 기록된다.
```bash
uv run gimme-job proactive run # 1회 실행 (발견+검증+저장+보고)
uv run gimme-job proactive run --mode weekly # 주간 deep scan (ATS 도메인 site: 쿼리 + 보드 자동 발견)
@ -129,12 +138,13 @@ uv run gimme-job proactive list [--status NEW] # 리드 원장 조회
> DuckDuckGo 단독 모드로 전환할 수 있습니다.
> **새 고용주 ATS 보드 추가**: 보통은 자동 발견에 맡기면 된다 —
> `proactive discover-boards --write`(또는 주간 deep scan)가 검증된 보드를
> `sites/employers.auto.yaml`에 기록하고, 실행 시 두 카탈로그를 병합한다.
> 수동으로 확실히 추가하려면 `sites/employers.yaml`의 `boards:`에 항목을 추가한다 —
> `proactive discover-boards --write`(또는 주간 deep scan)가 mission-driven으로
> 분류된 보드만 `sites/employers.auto.yaml`에 기록하고, 실행 시 두 카탈로그를 병합한다.
> 수동으로 확실히 추가하려면 `sites/employers.yaml`의 `boards:`에 항목과
> `employer_type`(government/tribal/nonprofit/academic/public_health/mission)을 추가한다 —
> 각 플랫폼 토큰(Greenhouse slug / SmartRecruiters company id / ICIMS tenant / Workday
> domain·tenant·org)은 실제로 응답하는지 확인 후 반영해야 한다 (파일 상단에 예시 curl 명령).
> `smartrecruiters: {org: NATIVEHEALTH}`처럼 검증된 항목만 유지된다.
> `smartrecruiters: {org: NATIVEHEALTH, employer_type: nonprofit}`처럼 검증된 항목만 유지된다.
보고서는 매 런마다 `workspace/reports/proactive-YYYY-MM-DD.html`(스타일된 HTML 페이지)과
`.md`로 저장되고, 터미널 로그에 **클릭 가능한 링크**(OSC-8)가 함께 출력된다

@ -648,22 +648,28 @@ def proactive_discover_boards(
table = Table(title="Board Discovery", show_header=True)
table.add_column("Source", style="cyan")
table.add_column("Board", width=45)
table.add_column("Board", width=40)
table.add_column("Type", width=12)
table.add_column("Jobs", justify="right")
table.add_column("Matches", justify="right")
table.add_column("Status", width=24)
for cand in report.candidates:
if cand.reachable and cand.keyword_hits >= report.min_hits:
status = "[green]verified[/green]"
elif cand.reachable:
if not cand.reachable:
status = (
f"[red]failed: {cand.error[:18]}[/red]"
if cand.error
else "[yellow]no jobs[/yellow]"
)
elif cand.keyword_hits < report.min_hits:
status = "[yellow]no dental matches[/yellow]"
elif cand.error:
status = f"[red]failed: {cand.error[:18]}[/red]"
elif cand.employer_type in ("", "unknown", "private"):
status = "[red]out of scope (private)[/red]"
else:
status = "[yellow]no jobs[/yellow]"
status = "[green]verified[/green]"
table.add_row(
cand.source,
cand.label()[:45],
cand.label()[:40],
cand.employer_type or "-",
str(cand.total_jobs),
str(cand.keyword_hits),
status,
@ -672,7 +678,8 @@ def proactive_discover_boards(
console.print(
f"[dim]{report.queries_run} queries, {report.results_seen} results, "
f"{len(report.candidates)} new candidates, {report.known} already known, "
f"{len(report.verified())} verified[/dim]"
f"{len(report.verified())} verified, "
f"{len(report.scope_skipped())} out-of-scope[/dim]"
)
if report.blocked:
console.print("[yellow]⚠ engine blocked — rerun later or use --engine bing[/yellow]")

@ -20,13 +20,12 @@ from loguru import logger
from gimme_job.proactive.engines import SearchResult
from gimme_job.proactive.roles import (
NON_CLINICAL,
SUPPORT,
classify_role,
record_role_candidate,
UNKNOWN as ROLE_UNKNOWN,
)
from gimme_job.proactive.roles import (
UNKNOWN as ROLE_UNKNOWN,
classify_role,
is_role_excluded,
record_role_candidate,
)
DEFAULT_KEYWORDS = [
@ -60,6 +59,7 @@ class BoardConfig:
tenant: str = ""
keywords: list[str] = field(default_factory=lambda: list(DEFAULT_KEYWORDS))
source: str = ""
employer_type: str = ""
def _matches(text: str, keywords: list[str]) -> bool:
@ -74,6 +74,7 @@ def _result(
source: str,
employer: str,
query: str,
employer_type: str = "",
) -> Optional[SearchResult]:
# Strip the employer name before keyword matching — "Account Payable
# Specialist at Specialty Dental Brands" must not match on the company name.
@ -82,11 +83,12 @@ def _result(
text = re.sub(re.escape(employer), " ", text, flags=re.IGNORECASE)
if not url or not _matches(text, DEFAULT_KEYWORDS):
return None
# Role gate: support/admin postings are not orthodontist jobs.
role, role_reason = classify_role(title)
if role in (SUPPORT, NON_CLINICAL):
record_role_candidate(title, url, source, role, role_reason)
# Role gate: support/admin and non-ortho specialties are not target jobs.
excluded = is_role_excluded(title)
if excluded:
record_role_candidate(title, url, source, excluded[0], excluded[1])
return None
role, role_reason = classify_role(title)
if role == ROLE_UNKNOWN:
record_role_candidate(title, url, source, role, role_reason)
return SearchResult(
@ -96,6 +98,7 @@ def _result(
source=source,
employer_hint=employer,
query=query,
employer_type=employer_type,
)
@ -119,6 +122,7 @@ def fetch_greenhouse(cfg: BoardConfig) -> list[SearchResult]:
source=cfg.source,
employer=cfg.name,
query=f"greenhouse:{cfg.org}",
employer_type=cfg.employer_type,
)
if r:
results.append(r)
@ -149,6 +153,7 @@ def fetch_lever(cfg: BoardConfig) -> list[SearchResult]:
source=cfg.source,
employer=cfg.name,
query=f"lever:{cfg.org}",
employer_type=cfg.employer_type,
)
if r:
results.append(r)
@ -197,6 +202,7 @@ def fetch_smartrecruiters(cfg: BoardConfig) -> list[SearchResult]:
source=cfg.source,
employer=cfg.name,
query=f"smartrecruiters:{cfg.org}",
employer_type=cfg.employer_type,
)
if r:
results.append(r)
@ -249,6 +255,7 @@ def fetch_workday(cfg: BoardConfig) -> list[SearchResult]:
source=cfg.source,
employer=cfg.name,
query=f"workday:{cfg.tenant}",
employer_type=cfg.employer_type,
)
if r:
results.append(r)
@ -329,6 +336,7 @@ def fetch_icims(cfg: BoardConfig) -> list[SearchResult]:
source=cfg.source,
employer=cfg.name,
query=f"icims:{cfg.tenant}",
employer_type=cfg.employer_type,
)
if r:
results.append(r)
@ -377,6 +385,7 @@ def fetch_usajobs(cfg: BoardConfig) -> list[SearchResult]:
source=cfg.source,
employer=cfg.name,
query=f"usajobs:{keyword}",
employer_type=cfg.employer_type,
)
if r:
results.append(r)

@ -9,7 +9,9 @@ Runtime flow (pure Python + httpx + one SERP engine, no AI):
1. run ``site:{ats-domain} {keyword}`` searches (DuckDuckGo by default)
2. parse the ATS URLs into candidate boards (source + org/tenant/domain)
3. probe the public board API to verify it responds and count jobs
4. append verified boards to ``sites/employers.auto.yaml`` (machine-owned)
4. classify the employer as mission-driven (government/tribal/nonprofit/
FQHC/academic) from the SERP evidence — private/unknown boards are dropped
5. append verified boards to ``sites/employers.auto.yaml`` (machine-owned)
Hand-curated boards stay in ``sites/employers.yaml``; ``ProactivePlan.load()``
merges both catalogs and dedupes by ``board_key``.
@ -25,6 +27,7 @@ from urllib.parse import quote_plus, urlparse
from loguru import logger
from gimme_job.proactive.boards import DEFAULT_KEYWORDS, _matches, _parse_icims_jobs
from gimme_job.proactive.employer_scope import PUBLIC_TYPES, classify_employer
from gimme_job.proactive.engines import (
EngineBlockedError,
SearchEngineBase,
@ -79,11 +82,15 @@ class DiscoveredBoard:
domain: str = ""
tenant: str = ""
evidence_url: str = ""
evidence_title: str = ""
evidence_hit: bool = False
total_jobs: int = 0
keyword_hits: int = 0
reachable: bool = False
error: str = ""
# Mission-driven classification from the SERP evidence; private/unknown
# boards are never written to the machine catalog.
employer_type: str = ""
@property
def key(self) -> tuple:
@ -96,6 +103,7 @@ class DiscoveredBoard:
domain=self.domain,
tenant=self.tenant,
source=self.source,
employer_type=self.employer_type,
)
def label(self) -> str:
@ -126,7 +134,19 @@ class DiscoveryReport:
return [
c
for c in self.candidates
if c.reachable and c.keyword_hits >= self.min_hits
if c.reachable
and c.keyword_hits >= self.min_hits
and c.employer_type in PUBLIC_TYPES
]
def scope_skipped(self) -> list[DiscoveredBoard]:
"""Reachable dental boards dropped for being private/unknown employers."""
return [
c
for c in self.candidates
if c.reachable
and c.keyword_hits >= self.min_hits
and c.employer_type not in PUBLIC_TYPES
]
@ -434,6 +454,8 @@ def run_discovery(
cand = extract_candidate(r.url, r.title)
if not cand:
continue
if not cand.evidence_title:
cand.evidence_title = r.title
seen.setdefault(cand.key, cand)
for cand in seen.values():
@ -449,13 +471,22 @@ def run_discovery(
if cand.source == "workday" and cand.evidence_hit:
cand.keyword_hits = max(cand.keyword_hits, 1)
cand.error = probe.error
# Scope gate: only mission-driven employers may enter the machine
# catalog (private DSOs discovered via site: queries are dropped).
cand.employer_type, _ = classify_employer(
cand.evidence_url, f"{cand.name} {cand.evidence_title}", plan
)
report.candidates.append(cand)
skipped = report.scope_skipped()
logger.info(
f"[discovery] {report.queries_run} queries, {report.results_seen} results, "
f"{len(report.candidates)} new candidates, {report.known} known, "
f"{len(report.verified())} verified"
f"{len(report.verified())} verified, {len(skipped)} out-of-scope"
)
if skipped:
names = ", ".join(f"{c.label()} ({c.employer_type})" for c in skipped[:10])
logger.info(f"[discovery] Out-of-scope boards skipped: {names}")
return report
@ -476,6 +507,8 @@ def _board_to_dict(cfg: BoardConfig) -> dict:
data["org"] = cfg.org
elif cfg.tenant:
data["tenant"] = cfg.tenant
if cfg.employer_type:
data["employer_type"] = cfg.employer_type
return data
@ -517,7 +550,8 @@ def write_auto_boards(
header = (
"# 자동 발견 ATS 보드 — `gimme-job proactive discover-boards --write`가 생성/갱신.\n"
"# 수동 편집 금지: 직접 추가하는 항목은 sites/employers.yaml에 넣으세요.\n"
"# 검증 기준: API probe 성공 + 치과 키워드 매치 ≥ min_keyword_hits.\n\n"
"# 검증 기준: API probe 성공 + 치과 키워드 매치 ≥ min_keyword_hits\n"
"# + mission-driven 고용주 (영리 사기업/미분류는 등록하지 않음).\n\n"
)
body = yaml.dump(
{"boards": data},

@ -0,0 +1,273 @@
"""Employer scope gate — keep mission-driven employers, drop for-profit companies.
The search plan's ``setting`` terms (public health, FQHC, community health,
nonprofit, tribal, Indian Health, county health, safety net, academic, public
hospital) define the target: PSLF-eligible, mission-driven employers. Private
companies — DSOs such as Smile Doctors, Sonrava or Specialty Dental Brands —
are out of scope, and employers with no public/nonprofit signal default to
excluded (``UNKNOWN``). This is a scope decision, not a quality score.
Pure rules, no AI at runtime (CLAUDE.md). Signals:
- URL host: ``.gov`` / ``.mil`` / ``.edu`` / ``.nsn.us`` or a configured
mission-driven domain (suffix match)
- employer text: government/tribal/nonprofit/academic/FQHC terms
- board catalog: ``sites/employers.yaml`` entries carry ``employer_type``
YAML additions live in ``sites/proactive.yaml`` under
``verification.employer_terms`` / ``verification.public_domains``.
"""
from __future__ import annotations
import json
import re
from datetime import date, datetime, timedelta
from pathlib import Path
from typing import TYPE_CHECKING, Optional
from urllib.parse import urlparse
from loguru import logger
from gimme_job.utils.paths import workspace_dir
if TYPE_CHECKING:
from gimme_job.proactive.plan import ProactivePlan
GOVERNMENT = "government"
TRIBAL = "tribal"
NONPROFIT = "nonprofit"
ACADEMIC = "academic"
PUBLIC_HEALTH = "public_health"
MISSION = "mission"
PRIVATE = "private"
UNKNOWN = "unknown"
#: Employer types that pass the scope gate.
PUBLIC_TYPES = frozenset({GOVERNMENT, TRIBAL, NONPROFIT, ACADEMIC, PUBLIC_HEALTH, MISSION})
DEFAULT_EMPLOYER_TERMS: dict[str, list[str]] = {
TRIBAL: [
"tribal",
"tribe",
"indian health",
"indian health service",
"american indian",
"native american",
"alaska native",
"indian community",
"urban indian",
"navajo nation",
"cherokee nation",
"gila river",
"salt river pima",
],
GOVERNMENT: [
"veterans affairs",
"va medical",
"va hospital",
"department of health",
"health department",
"county of",
"city of",
"state of",
"military",
"army",
"navy",
"air force",
"correctional",
"prison",
"state hospital",
"public hospital",
"district hospital",
"usphs",
"commissioned corps",
],
ACADEMIC: [
"university",
"college",
"school of dentistry",
"dental school",
"academic",
"faculty",
],
PUBLIC_HEALTH: [
"federally qualified",
"fqhc",
"community health center",
"community health",
"community clinic",
"public health",
"safety net",
"county clinic",
],
NONPROFIT: [
"nonprofit",
"non-profit",
"not-for-profit",
"501(c)(3)",
],
}
#: Known mission-driven employers whose domain carries no public suffix.
DEFAULT_PUBLIC_DOMAINS: list[str] = [
"usajobs.gov",
"governmentjobs.com",
"higheredjobs.com",
"ihs.gov",
"hrsa.gov",
"va.gov",
"hhs.gov",
"tribalhealth.com",
"tribalhealth.org",
"nativehealthphoenix.org",
"gilariver.org",
"srpmic-nsn.gov",
]
#: Explicit for-profit markers (rare on pages, but unambiguous when present).
PRIVATE_TERMS: list[str] = [
"dental service organization",
"dso",
"private practice",
"for-profit",
]
_GOV_SUFFIXES = (".gov", ".mil")
_ACADEMIC_SUFFIXES = (".edu",)
_TRIBAL_SUFFIXES = (".nsn.us", "-nsn.us")
def effective_employer_terms(
plan: Optional["ProactivePlan"] = None,
) -> dict[str, list[str]]:
"""Code defaults merged with plan.verification.employer_terms (YAML additions)."""
terms = {etype: list(words) for etype, words in DEFAULT_EMPLOYER_TERMS.items()}
if plan is not None:
for etype, words in (plan.verification.employer_terms or {}).items():
current = terms.setdefault(etype, [])
for word in words or []:
if word and word not in current:
current.append(word)
return terms
def effective_public_domains(plan: Optional["ProactivePlan"] = None) -> list[str]:
domains = list(DEFAULT_PUBLIC_DOMAINS)
if plan is not None:
for domain in plan.verification.public_domains or []:
d = (domain or "").strip().lower()
if d and d not in domains:
domains.append(d)
return domains
def _host_matches(host: str, domain: str) -> bool:
d = domain.lower().lstrip(".")
return host == d or host.endswith("." + d)
def _match_term(text: str, terms: list[str]) -> Optional[str]:
lowered = text.lower()
for term in terms:
if re.search(rf"\b{re.escape(term.lower())}\b", lowered):
return term
return None
def classify_employer(
url: str,
text: str,
plan: Optional["ProactivePlan"] = None,
) -> tuple[str, str]:
"""Return (employer_type, reason) for a job page or candidate.
``text`` should be the employer-bearing text only (page title + header
region, SERP title/snippet) — job-description boilerplate is not used.
"""
host = (urlparse(url or "").netloc or "").lower().split(":")[0]
if host:
if host.endswith(_GOV_SUFFIXES):
return GOVERNMENT, f"host {host}"
if host.endswith(_ACADEMIC_SUFFIXES):
return ACADEMIC, f"host {host}"
if host.endswith(_TRIBAL_SUFFIXES):
return TRIBAL, f"host {host}"
for domain in effective_public_domains(plan):
if _host_matches(host, domain):
return MISSION, f"public domain {domain}"
lowered = (text or "").lower()
private_hit = _match_term(lowered, PRIVATE_TERMS)
if private_hit:
return PRIVATE, f"title matched '{private_hit}'"
terms = effective_employer_terms(plan)
for etype in (TRIBAL, GOVERNMENT, ACADEMIC, PUBLIC_HEALTH, NONPROFIT):
hit = _match_term(lowered, terms.get(etype, []))
if hit:
return etype, f"employer matched '{hit}'"
return UNKNOWN, "no mission-driven employer signal"
# ── Excluded candidate log (offline review input) ────────────────────────────
def employers_log_path(run_date: Optional[date] = None) -> Path:
day = run_date or date.today()
return workspace_dir() / "candidates" / f"{day.isoformat()}.employers.jsonl"
def record_employer_candidate(
url: str,
employer_type: str,
reason: str,
title: str = "",
run_date: Optional[date] = None,
) -> None:
"""Append out-of-scope employers (private/unknown) for later review."""
if employer_type in PUBLIC_TYPES:
return
path = employers_log_path(run_date)
path.parent.mkdir(parents=True, exist_ok=True)
entry = {
"title": title,
"url": url,
"employer_type": employer_type,
"reason": reason,
"logged_at": datetime.now().isoformat(timespec="seconds"),
}
with path.open("a", encoding="utf-8") as f:
f.write(json.dumps(entry, ensure_ascii=False) + "\n")
def read_employer_candidates(run_date: Optional[date] = None) -> list[dict]:
path = employers_log_path(run_date)
if not path.exists():
return []
entries: list[dict] = []
for line in path.read_text(encoding="utf-8").splitlines():
try:
entries.append(json.loads(line))
except json.JSONDecodeError:
continue
return entries
def prune_employer_logs(days: int = 30, now: Optional[date] = None) -> int:
"""Delete dated employer logs older than `days`."""
cutoff = (now or date.today()) - timedelta(days=days)
directory = workspace_dir() / "candidates"
if not directory.exists():
return 0
removed = 0
for path in directory.glob("*.employers.jsonl"):
try:
file_date = date.fromisoformat(path.name[:10])
except ValueError:
continue
if file_date < cutoff:
path.unlink(missing_ok=True)
removed += 1
if removed:
logger.info(f"[employers] pruned {removed} employer log(s) older than {days} days")
return removed

@ -21,6 +21,7 @@ from gimme_job.proactive.roles import (
)
from gimme_job.proactive.roles import (
classify_role,
normalize_title,
record_role_candidate,
)
from gimme_job.proactive.verifier import (
@ -235,7 +236,7 @@ class ProactiveEngine:
candidate = ProactiveLeadCandidate(
status=vr.status,
employer=lead["employer"],
title=vr.evidence.title or lead["title"],
title=normalize_title(vr.evidence.title) or lead["title"],
state=lead["state"],
official_url=lead["official_url"],
source_engine=lead["source_engine"],
@ -555,7 +556,12 @@ class ProactiveEngine:
total = min(len(candidates), max_pages)
logger.info(f"[proactive] Verifying {i + 1}/{total}: {cand.url}")
try:
vr = verify(page, cand.url, plan)
vr = verify(
page,
cand.url,
plan,
employer_type_hint=cand.employer_type,
)
except Exception as e:
logger.warning(f"[proactive] Verify error for {cand.url}: {e}")
result.errors.append(str(e)[:200])
@ -578,7 +584,7 @@ class ProactiveEngine:
leads.append(
{
"status": status,
"title": normalize_whitespace(vr.evidence.title) or cand.title,
"title": normalize_title(vr.evidence.title) or normalize_title(cand.title),
"url": canonical_lead_url(vr.evidence.final_url or cand.url),
"state": plan.state_for_query(cand.query),
"employer_hint": cand.employer_hint,

@ -32,6 +32,9 @@ class SearchResult:
employer_hint: Optional[str] = None
query: str = ""
secondary_urls: list[str] = field(default_factory=list)
# Mission-driven employer type when known from the board catalog
# (government/tribal/nonprofit/academic/public_health/mission).
employer_type: str = ""
class SearchEngineBase:

@ -90,6 +90,11 @@ class VerificationConfig(BaseModel):
# and sites/role_terms.auto.yaml carries machine-proposed additions.
role_terms: dict[str, list[str]] = Field(default_factory=dict)
credential_terms: list[str] = Field(default_factory=list)
# Employer scope: mission-driven employers only (government/tribal/
# nonprofit/FQHC/academic). Code defaults live in employer_scope.py;
# YAML adds project-specific terms/domains.
employer_terms: dict[str, list[str]] = Field(default_factory=dict)
public_domains: list[str] = Field(default_factory=list)
save_evidence: bool = True
@ -100,6 +105,10 @@ class BoardConfig(BaseModel):
tenant: str = ""
keywords: list[str] = Field(default_factory=list)
source: str = ""
# Mission-driven employer type declared by the catalog (employers.yaml).
# Auto-discovered boards carry the classifier result (employers.auto.yaml);
# empty means unknown → the page itself must prove scope.
employer_type: str = ""
def board_key(cfg: BoardConfig) -> tuple:
@ -218,11 +227,19 @@ class ProactivePlan(BaseModel):
if emp_path.exists():
emp_data = read_yaml(emp_path)
plan.employers = EmployersConfig.model_validate(emp_data or {})
# auto-discovered boards are machine-owned (sites/employers.auto.yaml)
# auto-discovered boards are machine-owned (sites/employers.auto.yaml);
# only mission-driven employers may live there (scope guard)
auto_path = sites_dir() / "employers.auto.yaml"
if auto_path.exists():
auto_data = read_yaml(auto_path)
plan.employers.merge(EmployersConfig.model_validate(auto_data or {}))
auto_cfg = EmployersConfig.model_validate(auto_data or {})
from gimme_job.proactive.employer_scope import PUBLIC_TYPES
for source, entries in list(auto_cfg.boards.items()):
auto_cfg.boards[source] = [
b for b in entries if b.employer_type in PUBLIC_TYPES
]
plan.employers.merge(auto_cfg)
# machine-proposed role terms (sites/role_terms.auto.yaml)
terms_path = sites_dir() / "role_terms.auto.yaml"
if terms_path.exists():

@ -180,6 +180,7 @@ def deliver_report(
reports_cfg = plan.report if plan else None
if reports_cfg is None or reports_cfg.html:
try:
from gimme_job.proactive.employer_scope import prune_employer_logs
from gimme_job.proactive.html_report import prune_reports, write_html_report
from gimme_job.proactive.roles import prune_candidate_logs
@ -187,6 +188,7 @@ def deliver_report(
html_path = write_html_report(report_text, run_date, prefix="proactive")
prune_reports(retention)
prune_candidate_logs(retention)
prune_employer_logs(retention)
except Exception as e:
logger.debug(f"[proactive] HTML report skipped: {e}")

@ -6,10 +6,11 @@ additions in ``sites/role_terms.auto.yaml`` (written by
``gimme-job proactive role-audit --write``).
Classification runs on the normalized job title:
target → orthodontist / dentist / hidden dentist titles
support → dental assistant, hygienist, technician, coordinator …
non_clinical → payable, accountant, HR, marketing, manager …
unknown → everything else (verifier then requires a dentist credential)
target → orthodontist / dentist / hidden dentist titles
support → dental assistant, hygienist, technician, coordinator …
non_clinical → payable, accountant, HR, marketing, manager …
other_specialty → pediatric dentist, endodontist, oral surgeon … (not ortho)
unknown → everything else (verifier then requires a dentist credential)
"""
from __future__ import annotations
@ -29,9 +30,37 @@ if TYPE_CHECKING:
TARGET = "target"
SUPPORT = "support"
NON_CLINICAL = "non_clinical"
OTHER_SPECIALTY = "other_specialty"
UNKNOWN = "unknown"
# Titles that are explicit non-orthodontic specialties — a pediatric dentist
# posting is not an orthodontist job even though it contains "dentist".
DEFAULT_ROLE_TERMS: dict[str, list[str]] = {
OTHER_SPECIALTY: [
"pediatric dentist",
"pediatric dentistry",
"pedodontist",
"endodontist",
"endodontics",
"periodontist",
"periodontics",
"prosthodontist",
"prosthodontics",
"oral surgeon",
"oral surgery",
"maxillofacial",
"omfs",
"oral pathologist",
"oral pathology",
"oral medicine",
"dental anesthesiologist",
"dental radiologist",
"general dentist",
"general dentistry",
"public health dentist",
"dental public health",
"orthognathic surgeon",
],
TARGET: [
"orthodontist",
"orthodontic dentist",
@ -155,12 +184,12 @@ def _match_term(text: str, terms: list[str]) -> Optional[str]:
def classify_role(title: str, plan: Optional["ProactivePlan"] = None) -> tuple[str, str]:
"""Return (role, reason). Target terms win over support/admin terms."""
"""Return (role, reason). Explicit specialties/support beat generic target terms."""
normalized = normalize_title(title)
if not normalized:
return UNKNOWN, "empty title"
terms = effective_role_terms(plan)
for role in (TARGET, SUPPORT, NON_CLINICAL):
for role in (OTHER_SPECIALTY, TARGET, SUPPORT, NON_CLINICAL):
hit = _match_term(normalized, terms.get(role, []))
if hit:
return role, f"title matched '{hit}'"
@ -170,9 +199,9 @@ def classify_role(title: str, plan: Optional["ProactivePlan"] = None) -> tuple[s
def is_role_excluded(
title: str, plan: Optional["ProactivePlan"] = None
) -> Optional[tuple[str, str]]:
"""Return (role, reason) when the posting is support/admin (pre-verify drop)."""
"""Return (role, reason) when the posting is not an orthodontist role."""
role, reason = classify_role(title, plan)
if role in (SUPPORT, NON_CLINICAL):
if role in (SUPPORT, NON_CLINICAL, OTHER_SPECIALTY):
return role, reason
return None

@ -19,12 +19,22 @@ from loguru import logger
if TYPE_CHECKING:
from playwright.sync_api import Page
from gimme_job.proactive.employer_scope import (
PUBLIC_TYPES,
classify_employer,
record_employer_candidate,
)
from gimme_job.proactive.employer_scope import (
UNKNOWN as EMPLOYER_UNKNOWN,
)
from gimme_job.proactive.plan import ProactivePlan
from gimme_job.proactive.roles import (
NON_CLINICAL,
OTHER_SPECIALTY,
SUPPORT,
UNKNOWN,
classify_role,
normalize_title,
)
from gimme_job.utils.text import normalize_whitespace
from gimme_job.utils.urls import canonical_lead_url
@ -60,6 +70,8 @@ class PageEvidence:
role: str = ""
role_reason: str = ""
credential_signal: bool = False
employer_type: str = ""
employer_type_reason: str = ""
salary_text: Optional[str] = None
fte: Optional[str] = None
license_requirement: Optional[str] = None
@ -194,6 +206,8 @@ def classify_lead(evidence: PageEvidence, plan: ProactivePlan) -> tuple[str, str
Rules (doc §15 + false-positive fixes):
- support/admin roles (dental assistant, payable, coordinator …) → FILTERED
- non-ortho specialties (pediatric dentist, endodontist …) → FILTERED
- for-profit/unknown employers (DSO …) → FILTERED (mission-driven scope)
- orthodontic signal is checked against title + job description, never
whole-page boilerplate (company blurbs mention "orthodontics" on every
posting of a dental employer)
@ -213,9 +227,14 @@ def classify_lead(evidence: PageEvidence, plan: ProactivePlan) -> tuple[str, str
role, role_reason = classify_role(evidence.title, plan)
evidence.role = role
evidence.role_reason = role_reason
if role in (SUPPORT, NON_CLINICAL):
if role in (SUPPORT, NON_CLINICAL, OTHER_SPECIALTY):
return FILTERED, f"role filter: {role} ({role_reason})"
employer_type = evidence.employer_type or EMPLOYER_UNKNOWN
if employer_type not in PUBLIC_TYPES:
reason = evidence.employer_type_reason or "no mission-driven employer signal"
return FILTERED, f"employer scope: {employer_type} ({reason})"
if not evidence.ortho_signal:
return FILTERED, "no orthodontic signal in title or job description"
@ -242,9 +261,14 @@ def classify_lead(evidence: PageEvidence, plan: ProactivePlan) -> tuple[str, str
def _extract_page_text(page: "Page") -> tuple[str, str]:
"""Return (title, body_text) from the current page."""
"""Return (normalized title, body_text) from the current page.
ATS pages embed marketing prefixes in <title> ("Job Application for
Pediatric Dentist at Specialty Dental Brands") — strip them so ledger
titles stay clean job titles.
"""
try:
title = page.title()
title = normalize_title(page.title())
except Exception:
title = ""
try:
@ -316,8 +340,17 @@ def _extract_apply_url(page: "Page") -> Optional[str]:
return None
def verify(page: "Page", url: str, plan: ProactivePlan) -> VerificationResult:
"""Open the candidate URL and classify it."""
def verify(
page: "Page",
url: str,
plan: ProactivePlan,
employer_type_hint: str = "",
) -> VerificationResult:
"""Open the candidate URL and classify it.
``employer_type_hint`` comes from the board catalog (employers.yaml) when
the employer's mission-driven status is already declared.
"""
from gimme_job.utils.paths import dom_dir, screenshots_dir
ev = PageEvidence(url=url, final_url=url)
@ -341,6 +374,15 @@ def verify(page: "Page", url: str, plan: ProactivePlan) -> VerificationResult:
page_text = normalize_whitespace(f"{ev.title} {ev.body_text}")
signal_text = normalize_whitespace(f"{ev.title} {ev.scoped_text or ev.body_text}")
if employer_type_hint in PUBLIC_TYPES:
ev.employer_type = employer_type_hint
ev.employer_type_reason = f"board catalog declared {employer_type_hint}"
else:
employer_text = normalize_whitespace(f"{ev.title} {ev.body_text[:300]}")
ev.employer_type, ev.employer_type_reason = classify_employer(
ev.final_url or url, employer_text, plan
)
ev.ats_platform = detect_ats(ev.final_url, plan.verification.ats_patterns)
ev.is_aggregator = is_aggregator(ev.final_url, plan.verification.aggregator_domains)
ev.apply_found = has_apply_path(page_text, plan.verification.apply_keywords)
@ -361,6 +403,14 @@ def verify(page: "Page", url: str, plan: ProactivePlan) -> VerificationResult:
status, reason = classify_lead(ev, plan)
if status == FILTERED and reason.startswith("employer scope"):
record_employer_candidate(
ev.final_url or url,
ev.employer_type or EMPLOYER_UNKNOWN,
ev.employer_type_reason,
title=ev.title,
)
screenshot_path: Optional[str] = None
dom_path: Optional[str] = None
if status in (VERIFY, ACTIVE) and plan.verification.save_evidence:
@ -408,6 +458,8 @@ def evidence_to_dict(ev: PageEvidence) -> dict:
"role": ev.role,
"role_reason": ev.role_reason,
"credential_signal": ev.credential_signal,
"employer_type": ev.employer_type,
"employer_type_reason": ev.employer_type_reason,
"salary_text": ev.salary_text,
"fte": ev.fte,
"license_requirement": ev.license_requirement,

@ -49,8 +49,16 @@ def canonical_lead_url(url: str) -> str:
if k.lower() not in _TRACKING_PARAMS
]
query = urlencode(keep)
scheme = parsed.scheme.lower()
netloc = parsed.netloc.lower()
# Drop default ports so USAJOBS "www.usajobs.gov:443/job/1" dedupes
# against "www.usajobs.gov/job/1".
if scheme == "https" and netloc.endswith(":443"):
netloc = netloc[: -len(":443")]
elif scheme == "http" and netloc.endswith(":80"):
netloc = netloc[: -len(":80")]
canonical = urlunparse(
(parsed.scheme.lower(), parsed.netloc.lower(), parsed.path, "", query, "")
(scheme, netloc, parsed.path, "", query, "")
).rstrip("/")
return canonical or url
except Exception:

@ -1,30 +1,10 @@
# 자동 발견 ATS 보드 — `gimme-job proactive discover-boards --write`가 생성/갱신.
# 수동 편집 금지: 직접 추가하는 항목은 sites/employers.yaml에 넣으세요.
# 검증 기준: API probe 성공 + 치과 키워드 매치 ≥ min_keyword_hits.
# 검증 기준: API probe 성공 + 치과 키워드 매치 ≥ min_keyword_hits
# + mission-driven 고용주 (영리 사기업/미분류는 등록하지 않음).
#
# 2026-09-13: 영리 DSO 보드(Smile Doctors·Sonrava·Specialty Dental Brands 등)를
# 스코프 위반으로 전부 제거. University of Utah Health(academic)는
# sites/employers.yaml로 이동.
boards:
workday:
- name: Smiledoctors
domain: smiledoctors.wd108.myworkdayjobs.com
tenant: smiledoctors
org: SD
icims:
- name: Connection
tenant: connection
- name: Sonrava
tenant: sonrava
- name: Mobile Dentists
tenant: mobiledentists
- name: Uuhc
tenant: uuhc
greenhouse:
- name: Specialty Dental Brands
org: specialtydentalbrands
- name: Willamette Dental
org: willamettedentalgroup
lever:
- name: Vitana Pediatric
org: vitana-pediatric
smartrecruiters:
- name: Epic4Specialtypartners
org: Epic4SpecialtyPartners
boards: {}

@ -2,6 +2,10 @@
# 각 엔드포인트는 실동작 검증 후 유지 — 실패/빈 결과 항목은 정리.
# 런타임에 AI 호출 없음 — 순수 공개 API/피드 + 키워드 필터.
#
# SCOPE: mission-driven 고용주만 등록한다 (정부/트라이벌/비영리/FQHC/대학).
# 영리 사기업(DSO·개인 클리닉)은 등록하지 않는다 — employer_type 참고.
# 자동 발견 보드(employers.auto.yaml)는 이 게이트를 통과한 것만 기록된다.
#
# NOTE: 이 파일은 수동 큐레이션 전용. 자동 발견된 보드는 sites/employers.auto.yaml에
# 기록되며 ProactivePlan.load()가 두 카탈로그를 병합한다.
# 자동 발견/등록: `gimme-job proactive discover-boards --write`
@ -14,6 +18,9 @@
# icims → tenant (careers-{tenant}.icims.com)
# usajobs → USAJOBS_API_KEY/EMAIL (.env) 사용, org 미사용
#
# employer_type 값: government | tribal | nonprofit | academic | public_health | mission
# (미표기 시 보드 페이지 자체가 mission-driven 신호를 증명해야 함)
#
# 새 고용주 추가 시 아래 URL로 실제 응답을 확인한 뒤 반영:
# greenhouse curl -s https://boards-api.greenhouse.io/v1/boards/{org}/jobs
# smartrecruiters curl -s https://api.smartrecruiters.com/v1/companies/{org}/postings
@ -25,7 +32,12 @@ boards:
smartrecruiters:
- name: Native Health (Phoenix, AZ)
org: NATIVEHEALTH
employer_type: nonprofit
workday: []
icims: []
icims:
- name: University of Utah Health
tenant: uuhc
employer_type: academic
usajobs:
- name: USAJOBS
- name: USAJOBS
employer_type: government

@ -238,10 +238,14 @@ verification:
- no longer accepting applications
- no longer available
- not accepting applications
- applications are no longer accepted
- job closed
- job expired
- this job is no longer open
- filled
- this announcement has closed
- this job announcement has closed
# NOTE: bare "filled" is excluded — federal boilerplate says
# "until the position is filled" on *open* announcements (USAJOBS).
ortho_signal_terms:
- orthodontist
- orthodontics
@ -293,7 +297,7 @@ verification:
- temporary coverage
- travel assignment
- temp-to-perm
# Role gate — target(교정의/치과의)만 리드 유지, support/비임상은 FILTERED.
# Role gate — target(교정의/치과의)만 리드 유지, support/비임상/타 전문과목은 FILTERED.
# 기본 용어는 gimme_job/proactive/roles.py, 자동 제안은 sites/role_terms.auto.yaml.
role_terms:
support:
@ -304,6 +308,12 @@ verification:
- insurance verification
- community outreach
- practice administrator
# Employer scope gate — 영리 사기업(DSO 등) 제외, mission-driven(정부/트라이벌/
# 비영리/FQHC/대학) 고용주만 유지. 신호가 없으면 기본 제외(default-deny).
# 기본 용어/도메인은 gimme_job/proactive/employer_scope.py에서 관리하며,
# 여기서는 프로젝트 전용 추가만 한다 (예: tribal: ["cherokee nation"]).
employer_terms: {}
public_domains: []
# unknown 타이틀(예: Orthodontic Clinician)이 ACTIVE가 되려면 필요한 자격 신호
credential_terms:
- DDS

@ -1,10 +1,11 @@
"""Shared fixtures for proactive tests."""
import pytest
from gimme_job.proactive import roles
from gimme_job.proactive import employer_scope, roles
@pytest.fixture(autouse=True)
def _isolate_candidate_logs(tmp_path, monkeypatch):
"""Keep role-candidate JSONL writes inside the test tmp dir."""
"""Keep role/employer candidate JSONL writes inside the test tmp dir."""
monkeypatch.setattr(roles, "workspace_dir", lambda: tmp_path / "workspace")
monkeypatch.setattr(employer_scope, "workspace_dir", lambda: tmp_path / "workspace")

@ -130,6 +130,50 @@ def test_boards_role_gate_drops_support_roles(monkeypatch):
assert [r.title for r in res] == ["Orthodontist"]
def test_boards_drop_non_ortho_specialties(monkeypatch):
payload = {
"jobs": [
{
"title": "Pediatric Dentist",
"absolute_url": "https://gh/j/1",
"content": "orthodontics referrals",
},
{
"title": "Orthodontist",
"absolute_url": "https://gh/j/2",
"content": "braces",
},
]
}
client = _FakeClient([_Resp(json=payload)])
_patch(monkeypatch, client)
res = fetch_greenhouse(BoardConfig(name="Acme", org="acme", source="greenhouse"))
assert [r.title for r in res] == ["Orthodontist"]
def test_boards_pass_employer_type(monkeypatch):
payload = {
"jobs": [
{
"title": "Orthodontist",
"absolute_url": "https://gh/j/1",
"content": "braces",
}
]
}
client = _FakeClient([_Resp(json=payload)])
_patch(monkeypatch, client)
res = fetch_greenhouse(
BoardConfig(
name="Navajo Nation",
org="navajo",
source="greenhouse",
employer_type="tribal",
)
)
assert res[0].employer_type == "tribal"
def test_lever_filters_dental(monkeypatch):
client = _FakeClient(
[

@ -219,7 +219,7 @@ def test_run_discovery_extracts_and_filters_known():
engine = _FakeEngine(
[
SearchResult(
title="Orthodontist - Acme",
title="Orthodontist - Gila River Health Care",
url="https://job-boards.greenhouse.io/acme/jobs/1",
),
SearchResult(
@ -247,10 +247,38 @@ def test_run_discovery_extracts_and_filters_known():
assert report.known == 1
assert len(report.candidates) == 1
assert report.candidates[0].org == "acme"
assert report.candidates[0].employer_type == "tribal"
assert report.verified()[0].total_jobs == 5
assert probed == ["acme"]
def test_run_discovery_skips_private_employers():
"""Private DSO boards are dropped even with strong dental matches."""
engine = _FakeEngine(
[
SearchResult(
title="Orthodontist - Smile Doctors",
url="https://job-boards.greenhouse.io/smiledoctors/jobs/1",
)
]
)
def fake_prober(cand):
return ProbeResult(reachable=True, total=44, matched=10)
report = run_discovery(
None,
ProactivePlan(),
engine=engine,
sources=["greenhouse"],
max_queries=1,
prober=fake_prober,
)
assert report.candidates[0].employer_type == "unknown"
assert report.verified() == []
assert [c.org for c in report.scope_skipped()] == ["smiledoctors"]
def test_run_discovery_requires_keyword_hits():
engine = _FakeEngine(
[
@ -280,10 +308,10 @@ def test_run_discovery_serp_evidence_counts_as_hit():
engine = _FakeEngine(
[
SearchResult(
title="Orthodontist - Smile Doctors",
title="Orthodontist - Navajo Nation",
url=(
"https://smiledoctors.wd108.myworkdayjobs.com/"
"en-US/SmileDoctorsCareers/job/R1"
"https://navajo.wd5.myworkdayjobs.com/"
"en-US/NavajoNationCareers/job/R1"
),
)
]
@ -300,6 +328,7 @@ def test_run_discovery_serp_evidence_counts_as_hit():
max_queries=1,
prober=fake_prober,
)
assert report.candidates[0].employer_type == "tribal"
assert report.verified()[0].keyword_hits == 1
@ -360,12 +389,18 @@ def test_write_auto_boards_roundtrip(tmp_path):
boards = [
DiscoveredBoard(
source="workday",
name="Smile Doctors",
domain="smiledoctors.wd108.myworkdayjobs.com",
tenant="smiledoctors",
org="smiledoctors",
name="Navajo Nation",
domain="navajo.wd5.myworkdayjobs.com",
tenant="navajo",
org="navajo",
employer_type="tribal",
),
DiscoveredBoard(
source="icims",
name="Gila River Health Care",
tenant="gilariver",
employer_type="tribal",
),
DiscoveredBoard(source="icims", name="Mobile Dentists", tenant="mobiledentists"),
]
_, added = write_auto_boards(boards, path=path)
assert added == 2
@ -375,8 +410,9 @@ def test_write_auto_boards_roundtrip(tmp_path):
assert added2 == 0
data = read_yaml(path)
assert data["boards"]["workday"][0]["tenant"] == "smiledoctors"
assert data["boards"]["icims"][0]["tenant"] == "mobiledentists"
assert data["boards"]["workday"][0]["tenant"] == "navajo"
assert data["boards"]["workday"][0]["employer_type"] == "tribal"
assert data["boards"]["icims"][0]["tenant"] == "gilariver"
emp = EmployersConfig.model_validate(data)
assert len(emp.all_boards()) == 2
@ -431,7 +467,12 @@ def test_engine_weekly_new_boards(monkeypatch, tmp_path):
from gimme_job.proactive.engine import ProactiveEngine, ProactiveRunResult
cand = DiscoveredBoard(
source="greenhouse", name="Acme", org="acme", reachable=True, keyword_hits=3
source="greenhouse",
name="Acme",
org="acme",
reachable=True,
keyword_hits=3,
employer_type="nonprofit",
)
report = disc.DiscoveryReport(candidates=[cand], queries_run=5)
monkeypatch.setattr(disc, "run_discovery", lambda *a, **k: report)

@ -0,0 +1,106 @@
"""Tests for the employer scope gate (mission-driven employers only)."""
from gimme_job.proactive.employer_scope import (
ACADEMIC,
GOVERNMENT,
MISSION,
NONPROFIT,
PRIVATE,
PUBLIC_HEALTH,
PUBLIC_TYPES,
TRIBAL,
UNKNOWN,
classify_employer,
read_employer_candidates,
record_employer_candidate,
)
from gimme_job.proactive.plan import ProactivePlan
def _plan() -> ProactivePlan:
plan = ProactivePlan()
plan.verification.public_domains = ["myclinic.org"]
plan.verification.employer_terms = {"tribal": ["cherokee nation"]}
return plan
def test_government_host_suffix():
assert classify_employer("https://www.ihs.gov/jobs/1", "Orthodontist")[0] == GOVERNMENT
host_type, reason = classify_employer("https://careers.example.mil/job/1", "")
assert host_type == GOVERNMENT and "host" in reason
def test_academic_and_tribal_host_suffixes():
assert classify_employer("https://dentistry.utah.edu/job/1", "")[0] == ACADEMIC
assert classify_employer("https://jobs.hopi-nsn.us/job/1", "")[0] == TRIBAL
def test_public_domain_allowlist():
etype, reason = classify_employer(
"https://careers.gilariver.org/jobs/123", "Orthodontist"
)
assert etype == MISSION and "public domain" in reason
assert classify_employer("https://www.governmentjobs.com/careers/az", "")[0] == MISSION
def test_configured_public_domain_and_terms():
plan = _plan()
assert classify_employer("https://jobs.myclinic.org/x", "", plan)[0] == MISSION
etype, reason = classify_employer(
"https://careers-cherokee.icims.com/jobs/1", "Orthodontist - Cherokee Nation", plan
)
assert etype == TRIBAL and "cherokee nation" in reason
def test_employer_keyword_types():
assert (
classify_employer("https://jobs.example.com/1", "Orthodontist - Cherokee Nation")[0]
== TRIBAL
)
assert (
classify_employer("https://jobs.example.com/2", "Navajo Area Indian Health Service")[0]
== TRIBAL
)
assert (
classify_employer("https://jobs.example.com/3", "State of Arizona Dentist")[0]
== GOVERNMENT
)
assert (
classify_employer("https://jobs.example.com/4", "University of Utah Health")[0]
== ACADEMIC
)
assert (
classify_employer("https://jobs.example.com/5", "Community Health Center Dentist")[0]
== PUBLIC_HEALTH
)
assert (
classify_employer("https://jobs.example.com/6", "Nonprofit clinic orthodontist")[0]
== NONPROFIT
)
def test_private_terms_win():
etype, reason = classify_employer(
"https://jobs.example.com/1",
"Orthodontist - Smile Doctors, a dental service organization",
)
assert etype == PRIVATE and "dental service organization" in reason
def test_unknown_is_default_deny():
etype, reason = classify_employer(
"https://job-boards.greenhouse.io/specialtydentalbrands/jobs/1",
"Job Application for Orthodontist at Specialty Dental Brands",
)
assert etype == UNKNOWN
assert etype not in PUBLIC_TYPES
assert "no mission-driven" in reason
def test_record_and_read_employer_candidates():
record_employer_candidate(
"https://gh/j/1", PRIVATE, "dental service organization", title="Orthodontist"
)
record_employer_candidate("https://gh/j/2", GOVERNMENT, "host", title="Orthodontist")
entries = read_employer_candidates()
assert len(entries) == 1
assert entries[0]["employer_type"] == PRIVATE

@ -5,6 +5,7 @@ from gimme_job.proactive import roles
from gimme_job.proactive.plan import ProactivePlan
from gimme_job.proactive.roles import (
NON_CLINICAL,
OTHER_SPECIALTY,
SUPPORT,
TARGET,
UNKNOWN,
@ -78,6 +79,35 @@ def test_non_clinical_titles():
assert role == NON_CLINICAL, title
def test_other_specialty_titles():
"""Explicit non-ortho specialties are filtered even though they say "dentist"."""
for title in (
"Job Application for Pediatric Dentist - Associate at Specialty Dental Brands",
"Pediatric Dentist",
"Pediatric Dentistry Associate",
"Endodontist",
"Periodontist",
"Prosthodontist",
"Oral Surgeon",
"General Dentist",
"Public Health Dentist",
):
role, _ = classify_role(title)
assert role == OTHER_SPECIALTY, title
def test_orthodontist_variants_stay_target():
for title in (
"Pediatric Orthodontist",
"Orthodontist",
"Orthodontic Specialist",
"Staff Dentist",
"Dentist II (Orthodontics)",
):
role, _ = classify_role(title)
assert role == TARGET, title
def test_unknown_titles():
for title in ("Clinical Lead", "Dental Program Specialist", ""):
role, _ = classify_role(title)
@ -92,6 +122,7 @@ def test_target_beats_support_on_compound_titles():
def test_is_role_excluded():
assert is_role_excluded("Dental Assistant")[0] == SUPPORT
assert is_role_excluded("Payroll Clerk")[0] == NON_CLINICAL
assert is_role_excluded("Pediatric Dentist")[0] == OTHER_SPECIALTY
assert is_role_excluded("Orthodontist") is None
assert is_role_excluded("Clinical Lead") is None

@ -54,6 +54,16 @@ def test_has_closed_markers():
assert not has_closed_markers("We are actively hiring", ["position has been filled"])
def test_open_federal_boilerplate_is_not_closed():
"""USAJOBS open announcements say "until the position is filled"."""
plan = _plan()
open_text = (
"Applications will be referred every 2 weeks until the position is filled. "
"This job is open to the public."
)
assert not has_closed_markers(open_text, plan.verification.closed_keywords)
def test_has_ortho_signal():
assert has_ortho_signal(
"Dentist II", "This role includes orthodontic treatment of children", ["orthodontic"]
@ -74,6 +84,8 @@ def _ev(**kw) -> PageEvidence:
job_signal=True,
program_signal=False,
is_aggregator=False,
employer_type="government",
employer_type_reason="host test.gov",
)
defaults.update(kw)
return PageEvidence(**defaults)
@ -183,6 +195,41 @@ def test_non_clinical_role_filtered():
assert "non_clinical" in reason
def test_non_ortho_specialty_filtered():
"""Pediatric Dentist is a different specialty — not an orthodontist job."""
plan = _plan()
ev = _ev(
title="Job Application for Pediatric Dentist - Associate at Specialty Dental Brands",
body_text="Pediatric dentistry role with orthodontic referrals. Apply now.",
ortho_signal=True,
)
status, reason = classify_lead(ev, plan)
assert status == FILTERED
assert "other_specialty" in reason
def test_private_company_filtered():
"""Real false positive: for-profit DSO posting an actual ortho job."""
plan = _plan()
ev = _ev(
title="Job Application for Orthodontist at Specialty Dental Brands",
employer_type="private",
employer_type_reason="profile is for-profit DSO",
)
status, reason = classify_lead(ev, plan)
assert status == FILTERED
assert "employer scope" in reason and "private" in reason
def test_unknown_employer_filtered():
"""Default-deny: a mission-driven signal is required to keep a lead."""
plan = _plan()
ev = _ev(title="Orthodontist", employer_type="unknown")
status, reason = classify_lead(ev, plan)
assert status == FILTERED
assert "employer scope" in reason
def test_unknown_title_requires_credential():
plan = _plan()
ev = _ev(

@ -0,0 +1,22 @@
"""Tests for URL canonicalization."""
from gimme_job.utils.urls import canonical_lead_url
def test_strips_tracking_params():
url = "https://x.com/jobs/1?utm_source=a&jobId=42&ref=homepage"
assert canonical_lead_url(url) == "https://x.com/jobs/1?jobId=42"
def test_strips_default_ports():
assert (
canonical_lead_url("https://www.usajobs.gov:443/job/879546900")
== "https://www.usajobs.gov/job/879546900"
)
assert canonical_lead_url("http://example.com:80/jobs/1") == "http://example.com/jobs/1"
def test_keeps_non_default_ports():
assert (
canonical_lead_url("http://localhost:8000/jobs/1")
== "http://localhost:8000/jobs/1"
)
Loading…
Cancel
Save