The plan targets PSLF-eligible, mission-driven employers (government,
tribal, nonprofit, FQHC/community health, academic, public hospital).
For-profit DSOs (Smile Doctors, Sonrava, Specialty Dental Brands) had
slipped in through site: queries and auto-discovered boards, and their
pages produced off-scope leads — including Pediatric Dentist titles
that were never orthodontist jobs.
- employer_scope: rule-based employer classifier (.gov/.mil/.edu/
.nsn.us hosts, configured public domains, government/tribal/
nonprofit/academic/FQHC terms); private/unknown employers are
FILTERED at verification (default-deny) and logged to
workspace/candidates/*.employers.jsonl. Board catalogs declare
employer_type; auto-discovered boards are written only when the
SERP evidence classifies as mission-driven.
- roles: new other_specialty class (pediatric dentist, endodontist,
oral surgeon, general dentist …) filtered before target matching;
hidden dentist-title lanes at public employers stay.
- verifier: normalize ATS page titles ("Job Application for X at Y"
-> "X") before storing; apply employer gate.
- discovery: scope filter + out-of-scope reporting; private auto
catalog cleared (University of Utah Health kept as academic).
- closed markers: drop bare "filled" — federal boilerplate ("until
the position is filled") marked open USAJOBS postings as CLOSED.
- urls: strip default :443/:80 ports so USAJOBS links dedupe.
After each run (daily/weekly), gemma4 reviews the day's borderline
titles from the candidate log. support/non_clinical proposals are
written to sites/role_terms.auto.yaml automatically; target proposals
stay report-only because they would loosen the role gate.
- plan: role_audit config (enabled/model/max_titles/write/apply_roles)
- engine: _run_role_audit hook after report delivery, fatal-free
(Ollama down or no candidates → silent skip), link in run output
- role_audit: stopword/short-term guard for auto-applied terms
- CLI/scheduler print the audit report link; --model defaults to config
Two false positives (Account Payable Specialist, Dental Assistant) were
classified ACTIVE because ortho signal was matched against whole-page
boilerplate — dental employers mention 'orthodontics' in every posting.
- roles.py: title-based role gate (target/support/non_clinical/unknown)
with defaults + config terms; support/admin candidates are dropped at
board fetch and SERP discovery before verification
- verifier: ortho signal now checks title + job-description container only
(greenhouse/workday/icims/lever/smartrecruiters selectors); unknown
titles need a dentist credential (DDS/DMD/license) to become ACTIVE
- Orthodontic Clinician is treated as support — verified against Smile
Doctors postings ('under close supervision of an Orthodontist')
- non-target candidates are logged to workspace/candidates/*.roles.jsonl
(30-day retention) as offline audit input
- role_audit.py + 'proactive role-audit': gemma4 reviews borderline titles
offline and proposes terms; --write appends to sites/role_terms.auto.yaml
(runtime stays non-AI per CLAUDE.md)
Board results were appended after every SERP result, so with the fixed
verify budget (60 pages) they were never verified — the summary showed
'thousands of board matches, zero new leads'. Verification now
interleaves board and SERP candidates, boards first, so hidden ATS jobs
(e.g. Smile Doctors, Sonora, Specialty Dental Brands) actually get
classified and persisted.
Also strip employer names before board keyword matching — 'Account
Payable Specialist at Specialty Dental Brands' no longer matches on the
company name ('Dental').
- render each daily/outreach report as a styled standalone HTML page
(workspace/reports/{proactive,outreach}-YYYY-MM-DD.html, dark-mode aware)
- print OSC-8 hyperlinks in run/report/outreach CLI output and in the
auto scheduler log so the page opens with one click
- prune reports older than report.retention_days (default 30)
- config: sites/proactive.yaml report.html / report.retention_days
Phase 1 — query matrix:
- every state now also gets a hidden-title query on a rotation window (doc §4)
- national pool gains recruitment-phrasing variants
- budget 80->100 queries/day, verify cap 40->60
- state_for_query picks longest state name (fixes DC/West Virginia shadowing)
Phase 2 — hidden jobs via ATS boards:
- new boards.py: Greenhouse/Lever/SmartRecruiters/Workday/ICIMS/USAJOBS adapters
over public feeds/APIs; keyword filter -> SearchResult; graceful skip on failure
- new sites/employers.yaml: employer catalog (verified endpoints only)
- engine._discover now crawls boards, feeding candidates into verify->ledger
- USAJOBS uses USAJOBS_API_KEY/EMAIL from .env (skips until key approved)
- cli summary shows ATS board matches
- bing: additional engine (secondary) for more discovery coverage
- deprioritize: locum/DSO/per-diem/temp → FILTERED (doc §25)
- verifier now extracts salary/FTE/license and the actual apply URL,
surfaced in the report alongside the posting URL
- report sorts NEW/REOPENED/STILL_OPEN by a §26 priority score
- google-unlock: headed browser opens Google Jobs; a one-time manual
CAPTCHA pass persists cookies in the JobAgent profile for future runs
- googlejobs engine now waits and reloads up to twice before declaring
a block (transient interstitial pages can clear on their own)
Google Jobs anti-bot wall fired on every query from Korea, flooding the
summary with errors. Now an engine that raises EngineBlockedError is
skipped for the rest of the run (one attempt total), and blocks are
counted separately instead of polluting the error list.
ATSU school pages were classified ACTIVE because 'APPLY' nav buttons
matched apply keywords. Now:
- job_terms (position/hiring/vacancy/...) must be present for ACTIVE
unless an official ATS platform is detected
- program_terms (residency/admissions/curriculum/tuition/...) classify
school pages as FILTERED (not a job posting)
- added tests for both cases; cleaned the two false-positive leads
- sites/proactive.yaml: search plan (50 states, terms, OCONUS, budget,
denylist for known false positives, verification rules)
- plan: query matrix generator with per-day rotation and budget caps
- engines: DuckDuckGo HTML primary + Google Jobs secondary with anti-bot
fallback
- verifier: pure-heuristic ATS detection and ACTIVE/VERIFY/CLOSED/STALE
classification (no AI at runtime)
- ledger: ProactiveLead/ProactiveRun with lifecycle transitions
(NEW → STILL_OPEN, inactive → REOPENED, VERIFY → ACTIVE = NEW)
- CLI: gimme-job proactive [--dry-run --limit --max-verify]
- BrowserConfig gains proxies list (+ legacy single proxy kept for compat)
- BrowserManager accepts proxy URLs (http/socks5 with optional auth)
- orchestrator retries scrape with next proxy on navigation error or
0-result when site has known data
- ihs: route through US proxy pool (geo-blocked from Korea)
- add README with setup and usage
Disable all sites except ihs, usajobs, tribalhealth. Drop "dentist" from usajobs (search keywords + title filter) and tribalhealth, and add an orthodontist/orthodontic title filter to ihs, so the three active sites all return orthodontist roles only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tribal Health (tribalhealth.com) is a clinician staffing agency placing providers at tribal/IHS sites — the one non-subscribed source carrying net-new tribal dental roles not on ihs.gov/usajobs. Elementor WordPress loop; collect-all + keyword-filter (mirrors nativehealth). collect_cards walks all numbered board pages internally so sparse roles on later pages aren't dropped by the orchestrator's empty-page break; paginate() returns False.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tracing was accumulating 45GB+ of data (traces) + 5.6GB (dom) + 1.3GB
(screenshots) on every run. These artifacts are only useful during
development; remove them from the production run pipeline entirely.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Apply links (Indeed, Glassdoor, Monster, etc.) are pre-rendered inside
a <template> tag within each card's share_el. <template> content is
inert in the rendered DOM — a regular querySelectorAll from outside
returns nothing, which is why the previous fix produced empty URLs.
Iterate share_el's <template> elements and search their .content for
links matching utm_campaign=google_jobs_apply. Verified on the live
Google Jobs page that each card's template contains the correct
external apply URL.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add exclude_company_keywords post-filter (case-insensitive substring
match against company field). Configure googlejobs.yaml to exclude
"U.S. Navy" / "US Navy" — these recurring military recruitment ads
aren't relevant orthodontist openings.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The htidocid-based Google Jobs deep-links require the job to be in the
current search context (location, session) to open the detail panel.
When clicked from Telegram in a different context, the link only opens
the search list page without the specific job detail.
Each card's share_el ancestor has external apply URLs (Indeed,
Glassdoor, Monster, BeBee, AAO Career Center, etc.) pre-loaded in the
DOM. These are direct, stable links to the source job posting that work
regardless of session or location.
Switch jobUrl to use the first external (non-google.com) apply URL
within the card's share_el.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previous shareUrl-based URLs included session-tied tokens (shmd, shmds,
shem) that can expire over time, causing the job detail panel to not
open when the link is clicked later from Telegram.
Switch to a minimal Google Jobs URL using only the htidocid and the new
SPA fragment format:
https://www.google.com/search?q=<q>&udm=8#vhid=vt%3D20/docid%3D<id>&vssid=jobs-detail-viewer
Verified via Playwright: this format reliably opens the specific job
detail panel without any session-tied parameters.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Use page.click(selector) instead of ElementHandle.click() so Playwright
re-queries the element at click time after AJAX page refresh.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- gilariver: parse employment type from subtitle "Location | Category | Active - Full Time"
- usajobs: add salary_text and employment_type selectors to YAML
- hrsa: filter <br> elements from td children to fix company/employment_type index offset
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- ihs: Indian Health Service Dentistry (table, keyword filter)
- gilariver: Gila River Health Care via Infor CloudSuite (Angular, click pagination)
- nativehealth: NATIVE HEALTH via SmartRecruiters (AJAX Show More, keyword filter)
- srpmic: Salt River Pima-Maricopa via GovernmentJobs company page (URL keyword search)
- bfrench: Consulting BFrench via JazzHR (simple table, keyword filter)
- govtjobs: GovernmentJobs.com main search (URL keyword search, URL pagination)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
New site adapters:
- usajobs: multi-keyword sweep (dentist/orthodontics/orthodontist), URL pagination
- docshealth, southernortho: Paylocity platform, keyword in URL
- pdshealth, saltdental: iCIMS Angular platform, URL pagination
- hospitaljobsonline: Load More button, relative URL fix
- aaoinfo: AAO Career Center, click-based AJAX pagination
- aroragroup: Load More + post-scrape keyword filter
Core improvements:
- Early stop pagination: stops when all fingerprints on a page are already in DB
- multi_keyword_mode: separate — one scrape per keyword, cross-sweep dedup
- gimme-job list command to view collected postings
- Stored column in status command
- Telegram notification (replacing KakaoTalk)
- DB repo: count_by_site(), get_existing_fingerprints()
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>