How to scrape salary ranges from Ashby job boards with Python
Pull startup job postings with pay ranges from Ashby boards, annualised to yearly amounts, with currency and remote filters and an only-new feed.
datagrit › Guides › Career Site Jobs Aggregator - ATS Salary Data
GuideMonitor open roles at a list of companies across 8 applicant tracking systems in one schema, with numeric salaries and a daily only-new feed.
Published 2026-10-01 · uses the Career Site Jobs Aggregator - ATS Salary Data Actor
If you run a niche job board, a hiring newsletter or a sales team that treats hiring as a buying signal, your input is a list of companies you care about. The output you want: their open roles in one table, salary as a number, and tomorrow only the new roles.
The difficulty is that those companies do not share a careers system. One uses Greenhouse, another Lever, a third Workday, a German startup Personio. Each has its own URL pattern, its own JSON (or XML) format, its own way of writing locations and its own idea of what a salary field looks like, if it has one at all.
This post walks through what that problem looks like in practice and how to collect the data through one API call.
Applicant tracking systems publish public job board feeds so employers can show open roles on their careers pages. The Career Site Jobs Aggregator Actor on Apify reads those feeds for eight of them: Greenhouse, Lever, Ashby, Workable, SmartRecruiters, Recruitee, Personio and Workday. It uses plain HTTP, no browser and no login.
Doing this yourself means solving a set of problems per ATS:
employmentType becomes full_time, part_time, contract and so on).184,000 USD - 287,500 USD or $28.50 - $34.00 per hour.$ is not read as USD for postings in Mexico, Latin America, the Philippines, Hong Kong or Taiwan.How much does description parsing add? In the measurements recorded on 1 October 2026, the description supplied the salary for 22 of 27 Recruitee postings (Great Minds), 8 of 10 Lever postings (Epoch AI) and 8 of 10 Workday postings (NVIDIA), where the ATS feed had no pay field.
Companies can be given in any form, mixed in one list:
{
"companies": [
"https://jobs.ashbyhq.com/ramp",
"robinhood",
"lever:palantir",
"https://greatminds.recruitee.com",
"https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite"
],
"keywords": ["engineer", "developer"],
"excludeKeywords": ["intern"],
"locations": ["united states", "remote"],
"remoteOnly": false,
"onlyWithSalary": true,
"salaryFromDescription": true,
"minAnnualSalary": 120000,
"salaryCurrencies": ["USD"],
"postedWithinDays": 30,
"onlyNewSinceLastRun": false,
"maxItemsPerCompany": 100,
"maxItems": 1000
}
Title keywords match words that start with the term, so engineer also matches "Engineering". maxItemsPerCompany stops one big employer from using up the whole maxItems budget. Workday needs the full career site URL; the other systems accept a slug.
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run_input = {
"companies": ["https://jobs.ashbyhq.com/ramp", "robinhood", "lever:palantir"],
"onlyWithSalary": True,
"salaryCurrencies": ["USD"],
"minAnnualSalary": 120000,
"maxItems": 500,
}
run = client.actor("datagrit/career-site-jobs-aggregator").call(run_input=run_input)
jobs = [j for j in client.dataset(run["defaultDatasetId"]).iterate_items() if j.get("found")]
for j in jobs[:10]:
print(j["ats"], j["company"], j["title"], j["salaryAnnualMin"], j["salaryAnnualMax"], j["salarySource"])
curl -X POST \
"https://api.apify.com/v2/acts/datagrit~career-site-jobs-aggregator/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"companies": ["https://jobs.ashbyhq.com/ramp", "robinhood"], "onlyWithSalary": true, "maxItems": 50}'
An input with no companies at all runs a small example (Ramp on Ashby, up to 10 postings) that is not charged per result, which is handy for checking the output shape.
One row per posting, the same schema for every ATS:
| Field | Example | Meaning |
|---|---|---|
company | nvidia | Company name, or board slug when the ATS has none |
ats | workday | Where the posting was read from |
title | Growth Engineer, Developer Platform | Job title |
locations | ["US, CA, Santa Clara", "US, Remote"] | Every location of the posting |
workplaceType | remote | remote, hybrid, onsite or null |
employmentType | full_time | Normalized type |
salaryMin / salaryMax | 200000 / 322000 | Range in salaryCurrency per salaryPeriod |
salaryAnnualMax | 322000 | Yearly equivalent (hour × 2080, day × 260, week × 52, month × 12) |
salarySource | description | ats or description |
salaryText | 200,000 USD - 322,000 USD | The range as published or matched |
There is also jobId, requisitionId, department, team, postedAt, daysSincePosted, url, applyUrl, salaryCurrencies, equityMentioned, bonusMentioned and, on request, the plain-text description.
Set onlyNewSinceLastRun: true and schedule a daily run on Apify. The Actor remembers delivered postings per job board and per filter combination in a storage on your account, so each run returns and charges only new postings. Changing a filter starts a separate feed. Connect the run to a webhook, Google Sheets or your CMS import.
Feed your account list (careers page URLs work) with postedWithinDays: 7 and count new roles per company. A sudden cluster of new data or security roles at an account is a timing signal for outreach.
import pandas as pd
df = pd.DataFrame(jobs)
coverage = df.groupby("ats").agg(postings=("jobId", "count"), with_salary=("hasSalary", "mean"))
print(coverage)
usd_year = df[(df["salaryCurrency"] == "USD") & df["salaryAnnualMax"].notna()]
print(usd_year.groupby("company")["salaryAnnualMax"].describe()[["count", "50%", "max"]])
salarySource is worth keeping in any analysis: you may want to treat ranges written in descriptions separately from ranges in structured ATS fields.
hasSalary is false and equityMentioned is null there.salaryCurrencies together with minAnnualSalary.ats:slug or the board URL to be explicit; the run status shows which board each slug matched.The Actor reads only public job board feeds that ATS vendors provide for employers; it does not log in or collect candidate data.
One company list in, one schema out, with salaries you can filter on: pricing is pay per result, the Actor is on Apify, and the full field reference is in the documentation.
Pull startup job postings with pay ranges from Ashby boards, annualised to yearly amounts, with currency and remote filters and an only-new feed.
Build lists of French companies from the official Sirene register, filtered by net result, revenue, NAF code and location, using Python or cURL.
Collect Greenhouse job postings with published pay ranges per location zone, annualised and filtered by currency, using Python, cURL or pandas.