AI Career Observatory

Data Methodology

1. Job sources

Jobs come exclusively from public employer career pages served through public ATS APIs — currently Greenhouse and Lever. The employer's own posting is treated as the source of truth. We do not scrape gated sources or third-party aggregators.

2. Refresh frequency

Company boards are re-fetched on a rolling cycle (each company at least every ~12 hours when the sync worker runs continuously). Every job stores first_seen_at, last_seen_at and last_verified_at, shown on each job page as "Verified X hours ago".

3. Role classification

Job titles are lowercased, normalized and matched against ordered pattern rules per canonical role (30 roles, each with aliases). Matching stores the classification in the database; unmatched titles remain visible on the board but are excluded from role-level statistics. Role clusters group related titles (e.g. Forward Deployed Engineer / FDE / Deployment Engineer) so adjacent roles can be compared without pretending they're identical.

4. Skill extraction

Skills are detected by conservative pattern matching (word-boundary regexes) against job title and description text. Each match stores a confidence value and an evidence snippet. Ambiguous tokens are deliberately excluded (e.g. "rag" only matches as a standalone word, never inside "storage").

5. Salary methodology

Salary figures come only from ranges printed in the posting text, parsed with strict sanity bounds (plausible magnitude, min < max, ratio < 3×) and stored with source type job_description. We never mix in self-reported data, and we never publish medians below 5 observations.

6. Deduplication

Jobs are keyed by (source, external_id). Content hashes detect title/description changes between syncs. The same title at two locations is not merged — they are separate postings.

7. Job status lifecycle

Crawler errors never delete or close jobs — a failed sync leaves the last known state untouched and marks the company's sync as errored.

8. Limitations

The dataset covers tracked employers only — a lower bound on the market, skewed toward companies that use public ATS systems and toward the US market (salary disclosure follows US pay-transparency laws). Role and skill classification is rule-based and conservative; precision is prioritized over recall.

9. Corrections

Anyone can report an error from any page (see About). Corrections update the underlying dataset, not just the page.