How to scrape salary ranges from Ashby job boards with Python
Pull startup job postings with pay ranges from Ashby boards, annualised to yearly amounts, with currency and remote filters and an only-new feed.
datagrit › Guides › Website Tech Stack Lookup - Technology Detector
GuideDetect CMS, ecommerce platform, analytics, CDN and hosting of many domains, plus MX mail provider, SPF and DMARC, from Python or cURL.
Published 2026-10-01 · uses the Website Tech Stack Lookup - Technology Detector Actor
"Find me every Shopify store in this list of 5,000 domains" is a typical request in B2B sales. So are its cousins: which prospects already use a competitor's product, which ones run Magento 2 and might need a migration, which domains still have no DMARC policy, which companies receive email through Microsoft 365 rather than Google Workspace.
Agencies use the answer to qualify leads, SaaS companies to size a market, deliverability consultants to audit a portfolio, and investors to understand a target's stack before a call. The information is public: a website's HTML, headers and DNS records reveal most of it. The work is in reading it reliably at scale.
This post shows what a technology detector actually looks at, where homemade versions go wrong, and how to run one over a list of domains through the Apify API.
There is no registry of "which site uses what". Detection is inference from two public sources:
Server, X-Powered-By), cookie names, meta generator tags, script URLs and HTML snippets. A /wp-content/plugins/woocommerce/ path or a wc_add_to_cart_params variable means WooCommerce._dmarc record holds the DMARC policy, NS records give the DNS provider, CNAME, A and PTR records hint at hosting and CDN, and TXT verification records show services the domain was verified with.Writing this yourself sounds like a few regexes. In practice:
https://domain/, then https://www.domain/, then http://domain/, and follows up to 8 redirects and one meta refresh.found: false with the reason blocked-by-bot-protection instead.bbc.co.uk, not www.bbc.co.uk and not co.uk. The Actor uses a public-suffix-aware domain for DNS.The Actor matches 847 technology fingerprints in 62 categories, each detection with a version when exposed, a confidence from 0 to 100 and up to three pieces of evidence.
A list of prospects, keeping only those that run one of three ecommerce platforms:
{
"domains": [
"porterandyork.com",
"coxandcox.co.uk",
"allbirds.com",
"https://example.com/pricing"
],
"categories": [],
"requireTechnologies": ["Shopify", "WooCommerce", "Magento"],
"followRedirects": true,
"includeDns": true,
"maxConcurrency": 10,
"requestTimeoutSecs": 20
}
domains accepts bare domains, hostnames or full URLs; a full URL is fetched exactly as given. Duplicates (with or without www, http or https) are merged. requireTechnologies is the lead filter: sites that do not use any of the listed technologies are skipped and not charged. categories only trims the technologies list; the summary columns are always complete.
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
with open("domains.txt") as f:
domains = [line.strip() for line in f if line.strip()]
run = client.actor("datagrit/website-tech-stack-lookup").call(run_input={
"domains": domains,
"requireTechnologies": ["Shopify", "WooCommerce"],
"maxConcurrency": 20,
})
for site in client.dataset(run["defaultDatasetId"]).iterate_items():
if site.get("found"):
print(site["domain"], site["ecommerce"], site["mailProvider"], site["dmarcPolicy"])
else:
print(site.get("input"), "skipped:", site.get("errorReason"))
curl -X POST \
"https://api.apify.com/v2/acts/datagrit~website-tech-stack-lookup/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"domains": ["shopify.com", "wordpress.org", "bbc.co.uk"]}'
Each site costs one page request plus a few DNS lookups. In a local test on 1 October 2026, 12 well-known sites took 5.9 seconds including DNS. The Actor analyses 10 sites at a time by default (up to 50) and never more than 2 URLs of the same host at once.
One row per site, with ready-made summary columns plus the full evidence list:
| Field | Example | Meaning |
|---|---|---|
domain | porterandyork.com | Registrable domain |
cms | WordPress | CMS or site builder, with version when known |
ecommerce | WooCommerce | Ecommerce platform |
cdn | Cloudflare | CDN in front of the site |
hosting | WP Engine | Hosting provider or platform |
mailProvider | Google Workspace | From MX records |
spfRecord | v=spf1 include:_spf.google.com ~all | SPF record text |
dmarcPolicy | none | none, quarantine or reject |
technologyCount | 11 | Technologies listed |
technologies | [{"name": "WooCommerce", "confidence": 99, ...}] | Name, category, version, confidence, evidence |
Other columns: analytics, tagManager, webServer, jsFramework, programmingLanguage, marketingAutomation, liveChat, payment, dnsProvider, emailServices (sending services from SPF), mxRecords, hasSpf, hasDmarc, title, language, finalUrl and statusCode. Failed sites carry errorReason, for example dns-not-found, timeout or blocked-by-bot-protection, and are not charged.
Run your prospect domains monthly with requireTechnologies: ["Shopify"]. Only Shopify stores are returned and charged; push them into your CRM with marketingAutomation and payment as segmentation fields. Stacks change slowly, so a monthly run is usually enough for a lead list.
import pandas as pd
rows = [s for s in client.dataset(run["defaultDatasetId"]).iterate_items() if s.get("found")]
df = pd.DataFrame(rows)
audit = df[["domain", "mailProvider", "hasSpf", "hasDmarc", "dmarcPolicy", "dnsChecked"]]
at_risk = audit[(audit["dnsChecked"] == True) & ((audit["hasDmarc"] == False) | (audit["dmarcPolicy"] == "none"))]
print(at_risk.sort_values("domain").to_string(index=False))
Filtering on dnsChecked matters: when a DNS lookup failed, a missing value means "unknown", not "absent".
Run a list of competitors' customers or a category of sites, then count cms, ecommerce or cdn values with df["ecommerce"].value_counts(). For per-technology analysis, explode the technologies column (or use the Actor's "Technologies with evidence" dataset view, which already has one row per technology).
found: false row. A residential proxy in proxyConfiguration often gets through. DNS lookups do not use the proxy.mailProvider: "Other" means MX records exist but point to a provider the Actor does not recognise.The Actor reads only what a browser or DNS client receives from public websites and public DNS. It does not log in, use cookies or solve captchas. If you use the results for outreach, anti-spam rules still apply.
Website stack plus mail and DNS data in one row is usually what lead qualification needs; pricing is pay per analysed website, the Actor is on Apify, and the full list of fields and categories is in the documentation.
Pull startup job postings with pay ranges from Ashby boards, annualised to yearly amounts, with currency and remote filters and an only-new feed.
Monitor open roles at a list of companies across 8 applicant tracking systems in one schema, with numeric salaries and a daily only-new feed.
Build lists of French companies from the official Sirene register, filtered by net result, revenue, NAF code and location, using Python or cURL.