Research methodology

What the catalog measures—and what it does not.

ApplyDjin observes public employer vacancy pages and turns them into comparable hiring signals. This document records the collection, normalization, quality controls, and limitations behind those signals.

Last reviewed July 13, 2026 · catalog-observation-v1

Unit
One observed role

Not one hire or one open headcount.

Demand
Distinct vacancies

A technology counts once per role.

Scope
Observed catalog

Not a census of the whole labor market.

Scope and sources

The dataset is assembled from publicly accessible employer career pages and applicant-tracking-system job feeds. A row represents a role observed in that catalog. It does not prove how many people a company plans to hire, whether an offer was made, or whether the role appears on every channel used by the employer.

ApplyDjin is independent from the companies represented. Names, logos, and source links identify where a listing was observed; they do not imply endorsement or a commercial relationship.

Collection and freshness

  • Source publication and update dates are retained when the employer feed provides them.
  • First-seen time is used as a fallback when no reliable source publication date exists.
  • A role is active only while the source or catalog state indicates that it remains open.
  • Archived roles remain useful for aggregate trend calculations but are not returned by the public vacancy API.

Normalization and duplicate control

Source titles, departments, locations, seniority labels, work formats, employment types, and technology strings vary widely. ApplyDjin maps comparable values into stable catalog categories and retains source titles and departments; new ingests also retain the raw location alongside its structured projection. Historical rows may only contain the best available normalized location. Missing values stay missing; they are not guessed merely to complete a record.

The ingestion pipeline uses stable source identifiers and normalized vacancy fingerprints to avoid re-creating the same observed role. Similar titles at one company are not automatically merged when their source identifiers or substantive fields differ.

Technology demand

Technology counts are vacancy counts, not mention counts. A technology contributes at most once to a role. Evidence is separated into core, required, supporting, and optional tiers so a background reference does not carry the same meaning as a central requirement.

Public technology totals cover only normalized tags recognized by the catalog taxonomy. Natural-language extraction can miss implicit requirements or misclassify ambiguous terms; source listings remain the final reference.

Known limitations

  • Coverage reflects selected public sources and can overrepresent companies with accessible, structured feeds.
  • One listing can represent multiple openings; multiple listings can also support one hiring plan.
  • Salary coverage is sparse because most sources do not publish a range.
  • Listings can change or expire between observations, so candidates should verify details at the source.
  • Counts describe observed demand, not applicant supply, hiring success, compensation benchmarks, or forecasted growth.
  • Small samples are volatile; directional comparisons should be read with their vacancy counts and window size.

Public access and reproducibility

The aggregate snapshot, stable JSON endpoints, CSV export, Atom feeds, and OpenAPI contract are documented on the public data page. Published API responses omit candidate data, recruiter data, contacts, applications, documents, and full scraped descriptions.

Research notes are maintained under the transparent collective byline ApplyDjin Research. Material corrections should be sent through Contact & support.

Methodology changes that alter interpretation receive a new named version. Presentation-only changes keep the current version.