How to scrape salary ranges from Ashby job boards with Python
Pull startup job postings with pay ranges from Ashby boards, annualised to yearly amounts, with currency and remote filters and an only-new feed.
datagrit › Guides › Greenhouse Salary Scraper - Job Pay Ranges
GuideCollect Greenhouse job postings with published pay ranges per location zone, annualised and filtered by currency, using Python, cURL or pandas.
Published 2026-10-01 · uses the Greenhouse Salary Scraper - Job Pay Ranges Actor
Many employers publish a pay range on their job postings, and in some regions they must. For companies hiring through Greenhouse, those ranges sit on public job boards at boards.greenhouse.io/<company>, often split into location zones. On Robinhood's board, for example, "Zone 1" covers Menlo Park and New York and "Zone 2" Denver and Chicago; many US employers add a Canadian or European zone next to them.
That is exactly the data a compensation analyst wants for benchmarking, a recruiter wants for a target-company feed, and a job seeker wants before a negotiation. It is also awkward to use as published: several ranges per role, different currencies, and no field that says whether a range is hourly or yearly.
This post shows where the data lives, what goes wrong when you collect it yourself, and how to get it as a clean table.
Greenhouse offers employers a public job board API at boards-api.greenhouse.io/v1/boards/<board>/jobs. Called with content=true&pay_transparency=true, a single request returns every open posting on the board together with its pay ranges (pay_input_ranges). No login and no browser are involved.
The friction is in the pay data itself:
The Greenhouse Salary Scraper Actor on Apify reads each board in one request, keeps every zone (count plus a one-line summary per zone), computes the band in a single currency, infers the period and leaves it empty when it cannot be determined, and annualises amounts (month × 12, week × 52, hour × 2080). If the board list stops carrying pay_input_ranges, the run fails instead of returning postings that only look unpaid.
Speed is a side effect of the one-request design: on 30 September 2026 a run over five boards (Robinhood, Airbnb, Coinbase, Mercury, Discord) returned 640 postings in 10 seconds, 543 of them with a published pay range.
Engineering roles in New York, San Francisco or remote, paying at least 150,000 USD a year, first published in the last 14 days:
{
"companies": ["robinhood", "coinbase", "mercury"],
"keywords": ["engineer"],
"departments": ["Engineering"],
"locations": ["New York", "San Francisco", "Remote"],
"onlyWithPay": true,
"minAnnualPay": 150000,
"payCurrencies": ["USD"],
"publishedWithinDays": 14,
"onlyNewSinceLastRun": false,
"includeDescription": false,
"maxItems": 300
}
companies takes board names (the last part of the careers link) or full URLs. A board that does not exist is skipped and listed in the run status; if none exists, the run fails, so a typo never looks like an empty result. Always set payCurrencies together with minAnnualPay, so the threshold is compared in a single currency.
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run_input = {
"companies": ["robinhood", "coinbase", "mercury"],
"onlyWithPay": True,
"minAnnualPay": 150000,
"payCurrencies": ["USD"],
"maxItems": 300,
}
run = client.actor("datagrit/greenhouse-salary-scraper").call(run_input=run_input)
jobs = [j for j in client.dataset(run["defaultDatasetId"]).iterate_items() if j.get("found")]
for j in jobs[:10]:
print(j["company"], j["title"], j["payAnnualMin"], j["payAnnualMax"], j["payCurrency"], j["payZones"])
curl -X POST \
"https://api.apify.com/v2/acts/datagrit~greenhouse-salary-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"companies": ["robinhood"], "onlyWithPay": true, "maxItems": 50}'
One row per posting:
| Field | Example | Meaning |
|---|---|---|
company | robinhood | Greenhouse board name |
title | AML Investigator | Job title |
departments | Financial Crimes | Departments, semicolon-separated |
location | Denver, CO; New York, NY; Westlake, TX | Location text |
payMin / payMax | 58000 / 87000 | Band across zones, in payCurrency |
payCurrency | USD | Currency of the band |
payPeriod | year | Inferred: year, month, week or hour |
payAnnualMax | 87000 | Yearly equivalent of the top of the band |
payZones | 3 | Number of published pay ranges |
payRangeSummaries | Zone 1 (Menlo Park, CA; ...): 74K-87K USD/year | Every zone on one line, pipe-separated |
Rows also include payCurrencies (all currencies used by the zones), equityMentioned, bonusMentioned and commissionMentioned (from the pay text), firstPublishedAt, daysSincePublished, requisitionId, offices and jobUrl. The plain-text description (up to 8,000 characters) is added only when includeDescription is on.
Turn on onlyNewSinceLastRun and schedule a daily run. The first run returns everything that matches; each later run returns only postings not delivered before for the same companies and filters. Postings cut off by maxItems are not remembered, so they arrive in the next run. A run with nothing new returns one status row (not charged), which makes it easy to skip empty notifications in Zapier, Make or n8n.
Set payCurrencies: ["CAD"]. A US posting with an extra Canadian zone then matches, and its band and yearly amounts are recomputed in CAD. Comparing that with a USD run of the same boards shows how each employer prices its Canadian roles.
import pandas as pd
df = pd.DataFrame(jobs)
yearly = df[(df["payPeriod"] == "year") & (df["payCurrency"] == "USD")].copy()
yearly["midpoint"] = (yearly["payAnnualMin"] + yearly["payAnnualMax"]) / 2
table = (
yearly.groupby(["company", "departments"])
.agg(roles=("id", "count"),
median_mid=("midpoint", "median"),
max_top=("payAnnualMax", "max"),
equity_share=("equityMentioned", "mean"))
.sort_values("median_mid", ascending=False)
)
print(table.head(15))
Keep in mind that payMin is the bottom of the lowest zone and payMax the top of the highest, so a midpoint across zones is a blend of locations. For location-level analysis, parse payRangeSummaries.
hasPay is false for postings without a range. Postings that lack the pay data field are skipped and counted in the run status.payPeriod and the yearly amounts are empty, and minAnnualPay skips the posting.payCurrencies and the summaries.The Actor reads only the public job board API that Greenhouse provides for employers to publish roles; it does not log in or collect candidate data. Descriptions can contain names such as a recruiter's, so they are off by default.
For pay benchmarks that respect zones and currencies, this is a straightforward way to turn Greenhouse boards into a table; pricing is pay per result, the Actor is on Apify and the field reference is in the documentation.
Pull startup job postings with pay ranges from Ashby boards, annualised to yearly amounts, with currency and remote filters and an only-new feed.
Monitor open roles at a list of companies across 8 applicant tracking systems in one schema, with numeric salaries and a daily only-new feed.
Build lists of French companies from the official Sirene register, filtered by net result, revenue, NAF code and location, using Python or cURL.