What 1 Billion Real Job Records Looks Like as an AI Training Dataset

Most AI training datasets are assembled from web crawls, open corpus, or curated academic sources. Domain-specific structured data at scale is considerably rarer. In the workforce intelligence space, almost no conversation happens about what a billion real job records actually represents as a training asset, which means most models in this space are being trained on data that is structurally inferior to what is available.

Propellum’s one billion job records, sourced exclusively from employer career pages across 25 years and 15 countries, is not just a large dataset. It is a comprehensive representation of how the real labour market describes itself. Every employer naming convention for job titles. Every way a location has ever been expressed. Every salary format, every skills phrasing, every employment type variation. A model trained on this corpus has encountered every edge case the real world produces, not just the clean examples that appear in benchmark datasets.

What the Dataset Actually Contains

The value of any training dataset depends not on its size but on the diversity and quality of the signal it carries. Here is what one billion career-page-sourced job records contains across each dimension:

Normalised job titles. Employer job title conventions vary enormously and are largely arbitrary. “Senior Software Engineer,” “Sr. SWE II,” “L5 Engineer,” “Backend Engineer III,” and “Software Developer Senior” all describe the same role category. A dataset containing all variants mapped to a consistent taxonomy gives a skills classification model the full distribution of how any given role has been described in practice, rather than just the sanitised version that appears in job description templates.

Geocoded locations. Locations in raw job postings arrive as human-formatted strings: “Greater London,” “NYC area,” “Bangalore, KA,” “Remote (US only).” A dataset with location parsing applied produces geocoded coordinates and structured region hierarchies, which is the input format workforce demand forecasting models require to produce geographic granularity rather than string-matched approximations.

Salary estimates across 15 markets. Because salary data appears in fewer than 30% of job postings globally, any salary prediction model trained only on disclosed salary fields is working from a systematically biased subsample. A dataset with inference-based salary estimates applied, calibrated against the full corpus of surrounding market data, gives salary models training signals across the full distribution of roles and markets rather than only the employers who voluntarily disclosed pay.

NLP-extracted skills tags. Skills buried in free-text job descriptions are extracted and structured as discrete, queryable entities with required versus preferred distinctions preserved. The corpus covers skills phrasing as it has evolved across 25 years of market change, including the emergence of technologies and methodologies that did not exist at the dataset’s beginning. This temporal range is particularly valuable for models tracking skills demand shifts over time.

Employment type and seniority classifications. Full-time, part-time, contract, temporary, and internship designations normalised across the terminological variation different markets and employers use, combined with seniority classifications derived from title and description context rather than employer-stated levels alone.

Why Each Dimension Matters for Specific Model Types

Different model categories draw on different parts of the dataset’s value.

Skills classification models depend on the title and skills extraction dimensions. The training signal that makes a skills classifier accurate is not a clean list of skills with canonical names. It is the full distribution of how skills have been described, abbreviated, combined, and implied across millions of real postings, with the correct label attached to each variant.

Salary prediction models depend on the geocoded location data, the normalised seniority signal, and critically on the inference-based salary estimates for the 70% of postings where employers did not disclose pay. A model trained only on disclosed salaries learns from a skewed subset of the market.

Job-candidate matching algorithms depend on the relationship between normalised titles, extracted skills, and employment type classifications being consistent across the dataset. Inconsistent signal at training time produces inconsistent matching at inference time, which is the failure mode most matching systems encounter when trained on aggregator-sourced data.

Workforce demand forecasting models depend on the temporal dimension. Twenty-five years of continuous data from the same sourcing methodology, rather than snapshots assembled at different times using different collection approaches, allows a forecasting model to learn genuine trend signals rather than artefacts of methodology change.

Why the Source Matters More Than the Scale

Most job datasets on the market are assembled by pulling from other job boards and aggregators. By the time a record reaches the dataset it has passed through two or three intermediate layers, each one introducing duplicates, agency reposts, and label noise.

Propellum collects directly from employer career pages. One source. No intermediaries. Every record reflects exactly what the employer posted, at the time they posted it.

The difference is not incidental. An aggregator that pulls from other aggregators inherits every quality problem that exists upstream. Freshness degrades. Duplicates compound. Agency reposts get mixed in with real employer listings. The training signal gets noisier at every layer.

Propellum’s data has none of that. What goes into the dataset is what the employer wrote, nothing more. For supervised learning tasks where label quality directly determines model accuracy, starting from the original source is not a nice-to-have. It is the entire point.

How the Data Is Delivered

Propellum provides job training data in three formats depending on the team’s infrastructure and use case.

Bulk dataset delivery for teams that need a static snapshot for initial model training or benchmarking. Delivered in JSON or CSV with all enrichment dimensions pre-applied.

Continuous feed for teams building models that need to track market change over time, with new records delivered as they are sourced and processed rather than at scheduled intervals.

API access for teams building applications that need to query the dataset in real time, with filtering by geography at country, region, and city level, by role taxonomy node, by skills category, by seniority classification, by employment type, and by time range. Teams can combine filters to return precisely scoped subsets for geographically specific or vertically specific model training.

The Data Advantage That Compounds

Most AI teams spend the first six months of a workforce intelligence project cleaning data before they can train on it. Inconsistent titles, duplicate records, missing salary fields, aggregation noise and the pre-processing cost is often larger than the modelling cost.

Propellum’s dataset arrives pre-enriched, pre-deduplicated, and sourced from the only place in the market that produces genuinely clean ground truth: the employer’s own career page.

The models trained on it start from a better position. They encounter edge cases at training time that other datasets simply do not contain. And because the collection methodology has been consistent across 25 years, the temporal signal is genuine rather than an artefact of switching data sources halfway through.

A billion records built this way is not just more data. It is structurally different data. That difference shows up in model performance from the first training run.

Request a sample dataset →

Frequently Asked Questions

What makes a job records dataset valuable for AI training?

Scale alone is insufficient. What makes a job records dataset valuable for AI training is the combination of sourcing consistency, enrichment depth, and temporal range. A dataset sourced exclusively from employer career pages has no recruiter spam, no aggregation-introduced duplicates, and labels that reflect employer intent rather than third-party interpretation. Combined with 25 years of consistent collection methodology, the result is a training signal that aggregator-sourced datasets structurally cannot match.

How is career-page-sourced job data different from scraped aggregator data?

Career-page-sourced data is collected one hop from the employer: the original posting, as the employer wrote it, at the time they posted it. Aggregator data has passed through one or more intermediate collection and republication layers, introducing duplicates, freshness lag, agency reposts, and label noise at each layer. For supervised learning tasks where label quality determines model performance, the sourcing difference is material.

What NLP tasks is structured job data most useful for?

Skills extraction and classification, job title normalisation and taxonomy mapping, salary prediction for postings where pay is not disclosed, seniority inference from title and description signals, and employment type classification. Each of these tasks benefits from training data that reflects the full distribution of how the real labour market describes each dimension, rather than the cleaned examples that appear in academic benchmarks.

How current is the data and how frequently does it update?

Collection runs continuously from employer career pages across 15 countries. New records are available as hours-old rather than days-old, and the full corpus reflects 25 years of historical collection using consistent methodology. For teams building models that track labour market change over time, the temporal consistency of the collection approach is as important as the freshness of the most recent records.

Can the dataset be filtered by geography, industry, or role category?

Yes. Propellum’s dataset is filterable by geography at the country, region, and city level, by broad industry category, by role taxonomy node, by employment type, by seniority classification, and by time range. Teams working on geographically specific or vertically specific models can request a sample subset before committing to a full dataset delivery.