36 Commits (main)
 

Author SHA1 Message Date
I Luk Kim 3b6e5cdbd4 proactive: mission-driven employer scope + non-ortho specialty gate
The plan targets PSLF-eligible, mission-driven employers (government,
tribal, nonprofit, FQHC/community health, academic, public hospital).
For-profit DSOs (Smile Doctors, Sonrava, Specialty Dental Brands) had
slipped in through site: queries and auto-discovered boards, and their
pages produced off-scope leads — including Pediatric Dentist titles
that were never orthodontist jobs.

- employer_scope: rule-based employer classifier (.gov/.mil/.edu/
  .nsn.us hosts, configured public domains, government/tribal/
  nonprofit/academic/FQHC terms); private/unknown employers are
  FILTERED at verification (default-deny) and logged to
  workspace/candidates/*.employers.jsonl. Board catalogs declare
  employer_type; auto-discovered boards are written only when the
  SERP evidence classifies as mission-driven.
- roles: new other_specialty class (pediatric dentist, endodontist,
  oral surgeon, general dentist …) filtered before target matching;
  hidden dentist-title lanes at public employers stay.
- verifier: normalize ATS page titles ("Job Application for X at Y"
  -> "X") before storing; apply employer gate.
- discovery: scope filter + out-of-scope reporting; private auto
  catalog cleared (University of Utah Health kept as academic).
- closed markers: drop bare "filled" — federal boilerplate ("until
  the position is filled") marked open USAJOBS postings as CLOSED.
- urls: strip default :443/:80 ports so USAJOBS links dedupe.
3 weeks ago
I Luk Kim 0fda8383ea proactive: run offline role-audit daily and auto-apply safe terms
After each run (daily/weekly), gemma4 reviews the day's borderline
titles from the candidate log. support/non_clinical proposals are
written to sites/role_terms.auto.yaml automatically; target proposals
stay report-only because they would loosen the role gate.

- plan: role_audit config (enabled/model/max_titles/write/apply_roles)
- engine: _run_role_audit hook after report delivery, fatal-free
  (Ollama down or no candidates → silent skip), link in run output
- role_audit: stopword/short-term guard for auto-applied terms
- CLI/scheduler print the audit report link; --model defaults to config
3 weeks ago
I Luk Kim 555b450f0f proactive: role gate filters support/admin jobs, scoped ortho signal
Two false positives (Account Payable Specialist, Dental Assistant) were
classified ACTIVE because ortho signal was matched against whole-page
boilerplate — dental employers mention 'orthodontics' in every posting.

- roles.py: title-based role gate (target/support/non_clinical/unknown)
  with defaults + config terms; support/admin candidates are dropped at
  board fetch and SERP discovery before verification
- verifier: ortho signal now checks title + job-description container only
  (greenhouse/workday/icims/lever/smartrecruiters selectors); unknown
  titles need a dentist credential (DDS/DMD/license) to become ACTIVE
- Orthodontic Clinician is treated as support — verified against Smile
  Doctors postings ('under close supervision of an Orthodontist')
- non-target candidates are logged to workspace/candidates/*.roles.jsonl
  (30-day retention) as offline audit input
- role_audit.py + 'proactive role-audit': gemma4 reviews borderline titles
  offline and proposes terms; --write appends to sites/role_terms.auto.yaml
  (runtime stays non-AI per CLAUDE.md)
3 weeks ago
I Luk Kim 2d98afd062 proactive: give board candidates first verify slots, tighten matching
Board results were appended after every SERP result, so with the fixed
verify budget (60 pages) they were never verified — the summary showed
'thousands of board matches, zero new leads'. Verification now
interleaves board and SERP candidates, boards first, so hidden ATS jobs
(e.g. Smile Doctors, Sonora, Specialty Dental Brands) actually get
classified and persisted.

Also strip employer names before board keyword matching — 'Account
Payable Specialist at Specialty Dental Brands' no longer matches on the
company name ('Dental').
3 weeks ago
I Luk Kim 5b76a27c21 build: scope sdist to package, sites, tests, and docs
Untracked workspace/, dev/, .opencode/, and SQLite WAL files were
being swept into the source distribution (5.3 MB → 151 KB).
3 weeks ago
I Luk Kim ca62b2c3b7 proactive: daily HTML reports with clickable terminal links
- render each daily/outreach report as a styled standalone HTML page
  (workspace/reports/{proactive,outreach}-YYYY-MM-DD.html, dark-mode aware)
- print OSC-8 hyperlinks in run/report/outreach CLI output and in the
  auto scheduler log so the page opens with one click
- prune reports older than report.retention_days (default 30)
- config: sites/proactive.yaml report.html / report.retention_days
3 weeks ago
I Luk Kim a80bdaf06c proactive: auto-discover ATS boards from site: searches
Weekly deep scan now grows the employer catalog automatically:
site:{ats} SERP queries → parse Workday/ICIMS/Greenhouse/Lever/
SmartRecruiters URLs → probe public APIs → write verified boards
(min 1 dental match) to sites/employers.auto.yaml, which
ProactivePlan.load() merges with the curated employers.yaml.

Also fixes hidden-job coverage on existing boards:
- iCIMS: RSS feeds are retired — parse the iframe search page
- Workday: per-keyword searches (limit 20; tenants reject 50)
- SmartRecruiters: per-keyword API queries instead of newest 100

CLI: proactive discover-boards [--write] [--source] [--limit]
3 weeks ago
I Luk Kim 217071741f proactive: add reverify pass, cross-source dedup, and outreach list
- reverify: new 'proactive reverify' command re-opens CLOSED/STALE leads and
  transitions reopened ones via the normal ledger path (doc §15.3)
- dedup: merge cross-source duplicates by (title, employer), preferring the
  official ATS URL and storing alternates in secondary_urls (doc §19)
- outreach: new 'proactive outreach' command discovers private-practice sites
  and extracts contact email/phone into a direct-outreach report (doc §25)
- repo.get_inactive, SearchResult.secondary_urls, OutreachConfig wiring
- README + tests (repo, dedup, outreach)
3 weeks ago
I Luk Kim 5264e321fb proactive: add hidden-title state lane and direct ATS board crawling
Phase 1 — query matrix:
- every state now also gets a hidden-title query on a rotation window (doc §4)
- national pool gains recruitment-phrasing variants
- budget 80->100 queries/day, verify cap 40->60
- state_for_query picks longest state name (fixes DC/West Virginia shadowing)

Phase 2 — hidden jobs via ATS boards:
- new boards.py: Greenhouse/Lever/SmartRecruiters/Workday/ICIMS/USAJOBS adapters
  over public feeds/APIs; keyword filter -> SearchResult; graceful skip on failure
- new sites/employers.yaml: employer catalog (verified endpoints only)
- engine._discover now crawls boards, feeding candidates into verify->ledger
- USAJOBS uses USAJOBS_API_KEY/EMAIL from .env (skips until key approved)
- cli summary shows ATS board matches
3 weeks ago
I Luk Kim 942f863908 proactive: add Bing engine, deprioritize filter, and enrich lead fields
- bing: additional engine (secondary) for more discovery coverage
- deprioritize: locum/DSO/per-diem/temp → FILTERED (doc §25)
- verifier now extracts salary/FTE/license and the actual apply URL,
  surfaced in the report alongside the posting URL
- report sorts NEW/REOPENED/STILL_OPEN by a §26 priority score
4 weeks ago
I Luk Kim e8d411b01e proactive: add google-unlock and auto-retry for Google anti-bot wall
- google-unlock: headed browser opens Google Jobs; a one-time manual
  CAPTCHA pass persists cookies in the JobAgent profile for future runs
- googlejobs engine now waits and reloads up to twice before declaring
  a block (transient interstitial pages can clear on their own)
4 weeks ago
I Luk Kim e98f91036c proactive: circuit-breaker for blocked engines, track blocks separately
Google Jobs anti-bot wall fired on every query from Korea, flooding the
summary with errors. Now an engine that raises EngineBlockedError is
skipped for the rest of the run (one attempt total), and blocks are
counted separately instead of polluting the error list.
4 weeks ago
I Luk Kim 05510274a3 proactive: require job signal for ACTIVE, filter school/program pages
ATSU school pages were classified ACTIVE because 'APPLY' nav buttons
matched apply keywords. Now:
- job_terms (position/hiring/vacancy/...) must be present for ACTIVE
  unless an official ATS platform is detected
- program_terms (residency/admissions/curriculum/tuition/...) classify
  school pages as FILTERED (not a job posting)
- added tests for both cases; cleaned the two false-positive leads
4 weeks ago
I Luk Kim 4903319223 proactive: deliver report via configured notifier (telegram per global.yaml) 4 weeks ago
I Luk Kim d3fd62d51c proactive: add weekly deep scan, daily report, and scheduler
- deepscan: employer ATS domain queries (site:myworkdayjobs/icims/
  peopleadmin/governmentjobs/higheredjobs/usajobs/ihs/adea/dentalpost)
  + employer category lanes, reused through the same verify pipeline
- report: doc §23/§24 format with state coverage map, zero-result notice,
  KakaoTalk self-memo delivery with markdown fallback, optional Ollama
  qwen3.5:9b Korean summary
- scheduler: in-process daily loop (--at HH:MM), Sunday weekly deep scan
  auto-included, timezone-aware next-run computation
- CLI: proactive group with run/auto/report/list subcommands
4 weeks ago
I Luk Kim 6e9ac732db proactive: add discovery + verification core for active job search
- sites/proactive.yaml: search plan (50 states, terms, OCONUS, budget,
  denylist for known false positives, verification rules)
- plan: query matrix generator with per-day rotation and budget caps
- engines: DuckDuckGo HTML primary + Google Jobs secondary with anti-bot
  fallback
- verifier: pure-heuristic ATS detection and ACTIVE/VERIFY/CLOSED/STALE
  classification (no AI at runtime)
- ledger: ProactiveLead/ProactiveRun with lifecycle transitions
  (NEW → STILL_OPEN, inactive → REOPENED, VERIFY → ACTIVE = NEW)
- CLI: gimme-job proactive [--dry-run --limit --max-verify]
4 weeks ago
I Luk Kim cd38ece69c browser: add per-site proxy pool with failover rotation for geo-blocked sites
- BrowserConfig gains proxies list (+ legacy single proxy kept for compat)
- BrowserManager accepts proxy URLs (http/socks5 with optional auth)
- orchestrator retries scrape with next proxy on navigation error or
  0-result when site has known data
- ihs: route through US proxy pool (geo-blocked from Korea)
- add README with setup and usage
4 weeks ago
I Luk Kim ca5d229ca2 sites: focus active set on IHS/USAJOBS/tribal orthodontist roles
Disable all sites except ihs, usajobs, tribalhealth. Drop "dentist" from usajobs (search keywords + title filter) and tribalhealth, and add an orthodontist/orthodontic title filter to ihs, so the three active sites all return orthodontist roles only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
4 months ago
I Luk Kim 7086dcb0ea tribalhealth: add adapter for tribal staffing job board
Tribal Health (tribalhealth.com) is a clinician staffing agency placing providers at tribal/IHS sites — the one non-subscribed source carrying net-new tribal dental roles not on ihs.gov/usajobs. Elementor WordPress loop; collect-all + keyword-filter (mirrors nativehealth). collect_cards walks all numbered board pages internally so sparse roles on later pages aren't dropped by the orchestrator's empty-page break; paginate() returns False.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
4 months ago
I Luk Kim e13a25b3d5 runtime: disable trace/screenshot/DOM snapshot saving
Tracing was accumulating 45GB+ of data (traces) + 5.6GB (dom) + 1.3GB
(screenshots) on every run. These artifacts are only useful during
development; remove them from the production run pipeline entirely.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
5 months ago
I Luk Kim 37e5b59126 googlejobs: traverse <template> content to find apply URLs
Apply links (Indeed, Glassdoor, Monster, etc.) are pre-rendered inside
a <template> tag within each card's share_el. <template> content is
inert in the rendered DOM — a regular querySelectorAll from outside
returns nothing, which is why the previous fix produced empty URLs.

Iterate share_el's <template> elements and search their .content for
links matching utm_campaign=google_jobs_apply. Verified on the live
Google Jobs page that each card's template contains the correct
external apply URL.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
5 months ago
I Luk Kim fa748bebf9 googlejobs: filter out U.S. Navy company postings
Add exclude_company_keywords post-filter (case-insensitive substring
match against company field). Configure googlejobs.yaml to exclude
"U.S. Navy" / "US Navy" — these recurring military recruitment ads
aren't relevant orthodontist openings.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
5 months ago
I Luk Kim 2f33ea890d googlejobs: link to external apply URL instead of Google deep-link
The htidocid-based Google Jobs deep-links require the job to be in the
current search context (location, session) to open the detail panel.
When clicked from Telegram in a different context, the link only opens
the search list page without the specific job detail.

Each card's share_el ancestor has external apply URLs (Indeed,
Glassdoor, Monster, BeBee, AAO Career Center, etc.) pre-loaded in the
DOM. These are direct, stable links to the source job posting that work
regardless of session or location.

Switch jobUrl to use the first external (non-google.com) apply URL
within the card's share_el.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
5 months ago
I Luk Kim a49a6539a1 googlejobs: use clean udm=8 deep-link format for job URLs
Previous shareUrl-based URLs included session-tied tokens (shmd, shmds,
shem) that can expire over time, causing the job detail panel to not
open when the link is clicked later from Telegram.

Switch to a minimal Google Jobs URL using only the htidocid and the new
SPA fragment format:
  https://www.google.com/search?q=<q>&udm=8#vhid=vt%3D20/docid%3D<id>&vssid=jobs-detail-viewer

Verified via Playwright: this format reliably opens the specific job
detail panel without any session-tied parameters.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
5 months ago
I Luk Kim 0835522d06 aaoinfo: fix pagination click on detached DOM element
Use page.click(selector) instead of ElementHandle.click() so Playwright
re-queries the element at click time after AJAX page refresh.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 5fa28b3948 Add Google Jobs & HospitalRecruiting adapters; LinkedIn Volunteer filter; aaoinfo fixes
- Add googlejobs adapter (scroll-based extraction) and manifest
- Add hospitalrecruiting adapter (JS extraction) and manifest
- Fix googlejobs job URL: use shareUrl directly to preserve #fpstate=tldetail fragment
- Add exclude_title_keywords post-filter support (manifest + orchestrator)
- LinkedIn: exclude "volunteer" titles via exclude_title_keywords
- aaoinfo: exclude Preferred listings via :not(:has(.label-preferred)) selector
- aaoinfo: sort=start_ (descending = newest first)
- Telegram: multi-chat support, HTML escaping fixes, chunk splitting
- Various adapter, notifier, and CLI improvements

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 317f929059 LinkedIn: fix relative URLs and deduplicate accessibility text in titles
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim c4f9d5638b Fix Telegram HTML: URL & escaping, card-level chunk splitting, session scope fix
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 8ff94ad113 Configure loguru log level from LOG_LEVEL env var (default INFO)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim b28e204d80 LinkedIn: headless mode, scroll all 25 cards one-by-one
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 89ca7a4006 Fix warnings, LinkedIn selector update, Telegram notify listing, usajobs title filter
- LinkedIn: new URL with geoId, updated selectors to li[data-occludable-job-id], scroll adapter, pagination disabled
- usajobs: add include_title_keywords post-filter (dentist/orthodontist/orthodontic)
- aroragroup: fix AJAX load via doloadJBSearchList(), fix container selector
- govtjobs/srpmic: clear result_list_wait_selector to avoid 15s timeout on 0-result pages
- orchestrator: wait selector timeout WARNING → DEBUG
- extractor: no-container WARNING → DEBUG
- notifier: provider-based dispatch (telegram/kakaotalk), build_listing_messages() for Telegram
- cli notify: send job listing with links instead of Ollama summary
- global.yaml: provider set to telegram

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 22a0c6dbdc Extract employment_type for gilariver, usajobs; fix hrsa company parsing
- gilariver: parse employment type from subtitle "Location | Category | Active - Full Time"
- usajobs: add salary_text and employment_type selectors to YAML
- hrsa: filter <br> elements from td children to fix company/employment_type index offset

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 72bcffad7f Add HRSA Health Workforce Connector adapter
Form-based keyword search (PrimeNG Angular SPA) with client-side pagination.
Uses JS click to bypass headless visibility issue on the Search button.
Extracts 112 dentist opportunity cards per run.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 9d5251cead Add 6 new dental job site adapters
- ihs: Indian Health Service Dentistry (table, keyword filter)
- gilariver: Gila River Health Care via Infor CloudSuite (Angular, click pagination)
- nativehealth: NATIVE HEALTH via SmartRecruiters (AJAX Show More, keyword filter)
- srpmic: Salt River Pima-Maricopa via GovernmentJobs company page (URL keyword search)
- bfrench: Consulting BFrench via JazzHR (simple table, keyword filter)
- govtjobs: GovernmentJobs.com main search (URL keyword search, URL pagination)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 83fb8542a9 Add 8 site adapters, early stop, Telegram notifications, and list command
New site adapters:
- usajobs: multi-keyword sweep (dentist/orthodontics/orthodontist), URL pagination
- docshealth, southernortho: Paylocity platform, keyword in URL
- pdshealth, saltdental: iCIMS Angular platform, URL pagination
- hospitaljobsonline: Load More button, relative URL fix
- aaoinfo: AAO Career Center, click-based AJAX pagination
- aroragroup: Load More + post-scrape keyword filter

Core improvements:
- Early stop pagination: stops when all fingerprints on a page are already in DB
- multi_keyword_mode: separate — one scrape per keyword, cross-sweep dedup
- gimme-job list command to view collected postings
- Stored column in status command
- Telegram notification (replacing KakaoTalk)
- DB repo: count_by_site(), get_existing_fingerprints()

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago
I Luk Kim 8028d5e566 Initial implementation of gimme-job CLI
Complete Python package implementing all phases from the spec:
- Phase 0-1: Project scaffold, config, Pydantic/SQLAlchemy models, Typer CLI
- Phase 2: Runtime engine (BrowserManager, BaseAdapter/ManifestDrivenAdapter, orchestrator)
- Phase 3: Claude Code CLI integration (learn/repair modes with Jinja2 prompt templates)
- Phase 4: Ollama summarizer, KakaoTalk client, notification dispatcher
- Phase 5: Indeed adapter with manifest (sites/indeed.yaml)
- Phase 6: 36 unit tests (dates, hashing, dedupe, manifests, kakao, adapter)

gimme-job init/run/learn/repair/test/notify/status commands all wired up.
36/36 unit tests passing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
6 months ago