datagrit

datagrit › Data › Substack Newsletter Sponsorship Prospect Finder

Data

Substack Newsletter Sponsorship Prospect Finder

Substack publications ranked for sponsorship: subscriber count, paid-tier size, plan prices, posting cadence and engagement per 1,000 subscribers.

Run it on Apify StoreUse the APIfrom $3.15 per 1,000 results + $10 per run · no code needed
from $3.15 per 1,000 results + $10 per runpay only for publication analyzeds you get
JSON · CSV · Excelexport or call via API
Scheduled runsdaily or weekly feeds with Apify schedules
v0.8updated 2026-10-02

Substack Newsletter Sponsorship Prospect Finder turns Substack category rankings, a list of newsletters you name, or the recommendation links of a newsletter into one analyzed row per publication. Each row combines the public profile (subscriber count, paid-subscriber label, monthly, annual and founding prices) with numbers measured from the publication's own post archive: posts in the last 30 and 90 days, posts per week, the share of paid-only posts, median reactions and comments per post, reactions per 1,000 subscribers, and how many newsletters started recommending it in the last 30 days.

Export the result as JSON, CSV or Excel, call it through the Apify API, or plug it into n8n, Make and AI agents through MCP.

Why use Substack Newsletter Sponsorship Prospect Finder?

Output fields

Each record contains the profile fields (publicationId, name, url, language, authorName, subscribers, paidSubscribersLabel, paidSubscribersAtLeast, paidSubscribersSource, bestsellerTier, paymentsState, plan prices), the origin of the row (foundVia, category, rank, lookalikeOf), the measured activity fields (lastPostAt, daysSinceLastPost, posts30d, posts90d, postsPerWeek, paidPostShare90d, medianReactionsPerPost, medianCommentsPerPost, reactionsPer1kSubscribers), the recommendation fields (recommendedBy30d, recommendedBy30dCapped, recommendedByLatestAt; recommendedBy30dCapped is true when the 30-day window may extend beyond the recommendations Substack returned, so read the count as a lower bound), dataGaps, sourceUrl and scrapedAt.

Substack exposes the paid-subscriber count only as a band such as "Thousands of paid subscribers", so paidSubscribersAtLeast is a lower bound. The band comes from the publication's own record: its label when Substack shows one (paidSubscribersSource = ranking-label), or the numeric order of magnitude in the same record when the label is not shown (ranking-order-of-magnitude). Only when the publication has no band of its own does the author's bestseller badge fill in (author-bestseller-tier); the badge is set per author and can sit one step above or below the publication, so it never overrides the publication's band and is always returned separately as bestsellerTier. Check paidSubscribersSource before relying on a figure. The Actor does not report who sponsors a newsletter.

Publications that hide their subscriber count have subscribers empty. When the recommendations of one publication cannot be read, its recommendation fields stay empty and dataGaps names what is missing; a publication whose post archive cannot be read is skipped without billing. If the archive or the recommendations cannot be read for most publications, the run fails instead of returning empty columns.

Rows with a status instead of data (found: false) are never billed: a requested address that does not exist or is not a Substack publication, an address that could not be reached or checked in this run (the note says it is unknown whether it is a Substack publication), a publication confirmed to exist whose profile could not be read, and a publication whose post archive could not be read, because the activity metrics are what a row is billed for. The note of each row and the run summary give the reason, and a run that returns nothing because the source could not be read says so instead of blaming your filters.

Is it legal to scrape this data?

The Actor reads only publicly available information and does not log in or bypass access controls. You are responsible for using the data in line with applicable laws

(including data protection rules) and the source's terms. If you find an issue, open it in the Issues tab; problems are answered within one business day.

Output fields

Every result is one flat record, so it drops straight into a spreadsheet, a database or a CRM.

FieldTypeDescriptionExample
foundbooleanFalse only on status rows, which explain why a requested publication produced no result. Status rows are not billed.true
inputstringPublication reference or source that a status row refers to. Empty on result rows.no-such-newsletter.substack.com
notestringReason for a status row. Empty on result rows.No Substack publication was found at no-such-newsletter.substack.com.
publicationIdintegerSubstack publication ID, unique and stable.6349492
namestringPublication name.SemiAnalysis
subdomainstringSubstack subdomain of the publication.semianalysis
customDomainstringCustom domain of the publication, when it has one.newsletter.semianalysis.com
urlstringHome page of the publication.https://newsletter.semianalysis.com
descriptionstringShort description written by the publisher.Bridging the gap between the world's most important industry, semiconductors, an
languagestringLanguage code of the publication as set by the publisher.en
authorNamestringName of the main author as shown on Substack.Dylan Patel
authorHandlestringSubstack handle of the main author.semianalysis
authorBiostringPublic bio of the main author.Bridging the gap between business and the worlds most important industry.
authorUrlstringSubstack profile page of the main author.https://substack.com/@semianalysis
subscribersintegerSubscriber count as Substack displays it, rounded by Substack (for example 318,000). Null when the publisher hides it.318000
paidSubscribersLabelstringOrder of magnitude of paid subscribers as a band label. Substack does not publish an exact number. Taken from the publication's own band (the numeric order of magnitude in its record, which matches the ranking label when one is shown). The author's bestseller tier fills in only when the publication has no band of its own, and never overrides it (see paidSubscribersSource). Null when the publication's own band is 0 or absent and the author has no bestseller badge, which does not mean it has no paid subscribers.Thousands of paid subscribers
paidSubscribersAtLeastintegerLower bound implied by the band: 1 (single digits), 10, 100, 1000, 10000, 100000 or 1000000. Null when paidSubscribersLabel is null.1000
paidSubscribersSourcestringWhere the paid-subscriber band comes from: ranking-label (the publication's label, which equals the numeric order of magnitude in the same record), ranking-order-of-magnitude (the numeric order of magnitude of the publication when its label is not shown in that context) or author-bestseller-tier (the author's bestseller badge, used only when the publication has no band of its own; set per author, so it can sit one step above or below the publication). Null when there is no band.ranking-label
bestsellerTierintegerSubstack bestseller badge of the author: 0 (none), 100, 1000 or 10000 paid subscribers. Set per author, so it can differ from the band of the publication in paidSubscribersAtLeast; it never overrides that band.1000
paymentsStatestringWhether paid subscriptions are switched on, as reported by Substack (enabled, disabled or not set up).enabled
monthlyPriceUsdnumberRegular monthly subscription price in US dollars. Null when there is no monthly plan.50
annualPriceUsdnumberRegular annual subscription price in US dollars. Null when there is no annual plan.500
foundingPriceUsdnumberPrice of the founding member plan in US dollars. Null when the publication has none.1000
annualDiscountPercentintegerDiscount of the annual plan against twelve monthly payments. Null when either plan is missing.17
firstPostAtstringISO timestamp of the first published post.2020-05-22T21:26:00.000Z
hasPodcastbooleanWhether the publication publishes a podcast.false
hasCommunitybooleanWhether the community feature is switched on.true
foundViastringHow the publication was found: leaderboard, input or lookalike.leaderboard
categorystringLeaderboard category the publication was found in. Null when found via input or lookalike.technology
rankintegerPosition in the category leaderboard, 1 = top. Null when not found via leaderboard.1
lookalikeOfstringName of the seed publication whose recommendations led to this one. Null otherwise.Lenny's Newsletter
lastPostAtstringISO timestamp of the newest post, newsletter threads excluded.2026-09-29T14:02:11.000Z
daysSinceLastPostintegerWhole days between the newest post and the run.2
posts30dintegerPosts published in the 30 days before the run.9
posts90dintegerPosts published in the 90 days before the run.31
postsPerWeeknumberPosting cadence over the last 90 days, or over the age of the publication when it is younger.2.4
paidPostShare90dnumberShare of posts of the last 90 days that are not open to everyone, between 0 and 1.0.35
medianReactionsPerPostnumberMedian likes per post over posts older than three days within the last 90 days.412
medianCommentsPerPostnumberMedian comments per post over posts older than three days within the last 90 days.38
reactionsPer1kSubscribersnumberMedian reactions per post divided by the displayed subscriber count, times 1,000. Null when either is unknown.1.3
recommendedBy30dintegerPublications that started recommending this one in the last 30 days.4
recommendedBy30dCappedbooleanTrue when all 50 newest recommendations fall within 30 days, so the real count is at least recommendedBy30d.false
recommendedByLatestAtstringISO timestamp of the newest recommendation received from another publication.2026-09-27T08:15:00.000Z
dataGapsarrayIncomplete parts of this row. recommendations: the recommendation list could not be read, so recommendedBy30d, recommendedBy30dCapped and recommendedByLatestAt are null. archive-truncated: the post archive was read up to its 400-post cap without reaching 90 days back, so the activity counts are lower bounds. A publication whose post archive cannot be read at all is skipped without billing and never appears here, so archive is not a value of this list.[]
sourceUrlstringPage the record was read from.https://substack.com/api/v1/category/public/4/all?page=0
scrapedAtstringISO timestamp of the run.2026-10-01T09:30:00.000Z

Sample record

{
  "found": true,
  "input": "no-such-newsletter.substack.com",
  "note": "No Substack publication was found at no-such-newsletter.substack.com.",
  "publicationId": 6349492,
  "name": "SemiAnalysis",
  "subdomain": "semianalysis",
  "customDomain": "newsletter.semianalysis.com",
  "url": "https://newsletter.semianalysis.com",
  "description": "Bridging the gap between the world's most important industry, semiconductors, and business.",
  "language": "en",
  "authorName": "Dylan Patel",
  "authorHandle": "semianalysis",
  "authorBio": "Bridging the gap between business and the worlds most important industry.",
  "authorUrl": "https://substack.com/@semianalysis",
  "subscribers": 318000,
  "paidSubscribersLabel": "Thousands of paid subscribers",
  "paidSubscribersAtLeast": 1000,
  "paidSubscribersSource": "ranking-label",
  "bestsellerTier": 1000,
  "paymentsState": "enabled",
  "monthlyPriceUsd": 50,
  "annualPriceUsd": 500,
  "foundingPriceUsd": 1000,
  "annualDiscountPercent": 17,
  "firstPostAt": "2020-05-22T21:26:00.000Z",
  "hasPodcast": false,
  "hasCommunity": true,
  "foundVia": "leaderboard",
  "category": "technology",
  "rank": 1,
  "lookalikeOf": "Lenny's Newsletter",
  "lastPostAt": "2026-09-29T14:02:11.000Z",
  "daysSinceLastPost": 2,
  "posts30d": 9,
  "posts90d": 31,
  "postsPerWeek": 2.4,
  "paidPostShare90d": 0.35,
  "medianReactionsPerPost": 412,
  "medianCommentsPerPost": 38,
  "reactionsPer1kSubscribers": 1.3,
  "recommendedBy30d": 4,
  "recommendedBy30dCapped": false,
  "recommendedByLatestAt": "2026-09-27T08:15:00.000Z",
  "dataGaps": [],
  "sourceUrl": "https://substack.com/api/v1/category/public/4/all?page=0",
  "scrapedAt": "2026-10-01T09:30:00.000Z"
}

Input

FieldNameTypeWhat it does
sourceSourcestringWhere the publications come from. "leaderboard" reads the Substack category rankings, "publications" analyzes the publications you list, "lookalikes" finds publications that the ones you list recommend.
categoriesCategoriesarrayLeaderboard categories to read when the source is "leaderboard". Available: culture, technology, business, us-politics, finance, food, sports, art, world-politics, health-politics, news, fashionandbeauty, music, faith, climate, science, literature, fiction, health, design, travel, parenting, philosophy, comics, international, crypto, history, humor, education, film-and-tv, home-garden, games.
leaderboardTypeLeaderboard typestringWhich ranking to read: "all" is the ranking of all publications of the category, "paid" is the ranking of publications with paid subscriptions.
maxPerCategoryPublications to scan per categoryintegerHow many leaderboard positions to scan per category before filters apply. One page of the ranking holds 25 positions.
publicationsPublicationsarrayUsed when the source is "publications": Substack URLs, custom domains or subdomains such as "lenny".
lookalikesOfSeed publicationsarrayUsed when the source is "lookalikes": publications whose recommendations are followed to find similar newsletters. Accepts URLs, custom domains or subdomains.
minSubscribersMinimum subscribersintegerKeep publications with at least this many displayed subscribers. Zero disables the filter. Publications that hide their count are dropped when a limit is set.
maxSubscribersMaximum subscribersintegerKeep publications with at most this many displayed subscribers. Zero disables the filter. Publications that hide their count are dropped when a limit is set.
requirePaidPlanOnly publications with paid plansbooleanKeep only publications that have paid subscriptions switched on and a price.
maxMonthlyPriceUsdMaximum monthly price (USD)numberKeep publications whose regular monthly plan costs at most this many US dollars. Zero disables the filter. Publications without a monthly plan are dropped when a limit is set.
languageLanguagestringKeep publications whose language code starts with this value, for example "en" or "de". Empty keeps all languages.
activeWithinDaysPosted within daysintegerKeep publications whose newest post is at most this many days old. Zero disables the filter.
minPosts30dMinimum posts in 30 daysintegerKeep publications that published at least this many posts in the last 30 days. Zero disables the filter.
maxItemsMaximum resultsintegerStop after this many publications in total.
proxyConfigurationProxy configurationobjectOptional proxy. Leave disabled unless Substack blocks datacenter traffic; residential proxy raises the platform cost of the run.

Call it from your code

Run the Actor and get the results in one request. Replace YOUR_APIFY_TOKEN with the token from your Apify account settings.

curl -X POST "https://api.apify.com/v2/acts/datagrit~substack-sponsor-prospector/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"source":"leaderboard","categories":["technology","business"],"maxPerCategory":10,"maxItems":20}'
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagrit/substack-sponsor-prospector').call({
  "source": "leaderboard",
  "categories": [
    "technology",
    "business"
  ],
  "maxPerCategory": 10,
  "maxItems": 20
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.length, items[0]);

Install with npm i apify-client.

from apify_client import ApifyClient
import os

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("datagrit/substack-sponsor-prospector").call(run_input={
  "source": "leaderboard",
  "categories": [
    "technology",
    "business"
  ],
  "maxPerCategory": 10,
  "maxItems": 20
})
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(len(items), items[0] if items else None)

Install with pip install apify-client.

Frequently asked questions

How fresh is the data?

Every run reads Substack live. Ranks change between runs, so the same leaderboard input can return slightly different publications.

How is the posting cadence calculated?

From the post archive, over the last 90 days; for publications younger than 90 days the window is their age. Reactions are medians over posts older than three days, so a fresh post does not distort them.

Why is a number empty?

Either it could not be read in this run, which `dataGaps` and the run summary name, or Substack shows no figure for that publication. An empty `paidSubscribersLabel` means the publication's record carries no paid-subscriber band (its order of magnitude is 0 or absent) and the author has no bestseller badge to fall back on; it does not mean the publication has no paid subscribers.

What happens when Substack rate-limits the run, or one newsletter's site stops answering?

The Actor retries failed requests with a growing pause and follows the `Retry-After` header, but the time a run loses to failures is limited in three layers: 30 seconds per newsletter address (the pauses between attempts plus the time spent in attempts that failed or timed out; each attempt waits at most 25 seconds for an answer), 75 seconds for Substack's own ranking and recommendation service, and 150 seconds of failed-request time for the whole run, counted as a sum over requests that run in parallel, so it is used up in less than 150 seconds of clock time. Healthy requests never count. One newsletter whose site does not answer therefore costs at most about 30 seconds of that budget and affects only that newsletter. Newsletters that have not failed yet are protected from the others: once the run budget is used up, every address that has not failed in this run still gets one attempt of up to 8 seconds per request, and an address that has failed gets no further request and no pause that no longer fits. When a post archive stays unread, the publication is skipped without billing and, in `publications` mode, gets a status row. A post archive or recommendation list that arrives as something other than JSON (for example an HTML error page) is treated like any other failed read of that one newsletter: a publication without a readable archive is skipped without billing, one without a readable recommendation list gets `dataGaps: ["recommendations"]`, and neither ends the run. A rate-limited run, or a source that accepts connections and never answers, therefore ends in minutes, not hours. The run summary and the status rows say which case it was: a request that failed is reported as a source or network problem, a publication for which no request was sent because the budget was already used up is reported as a limit of the run, not as a problem of the address. The run fails with a clear message in three cases only: most post archives or recommendation lists that were answered with an error or an unreadable body (judged from at least 5 answers, so a few silent sites at the start of a ranking do not end the run), 12 or more archive requests that got no answer at all and not one that worked, or, at the end of the run, no post archive readable at all. Addresses that do not answer are never counted as a sign that the source changed. An address whose domain does not exist is reported at once, without retries.

Can I schedule runs?

Yes, use Apify schedules or call the Actor from your own workflow.

Something looks wrong.

Open an issue with the input you used; layout changes at the source are fixed quickly.

Try Substack Newsletter Sponsorship Prospect Finder on Apify

Related Actors

Company data

French Company Finder - Sirene Financials

French company lead lists from Sirene screened by net result and revenue, with net margin, size, matching establishment and optional directors.

from $5.60 / 1,000 results
Company data

Poland KRS New Company Registrations Feed

Newly registered Polish companies, foundations and associations from the official KRS court register: NIP, address, PKD, capital, email, with filters and change detection.

from $10.50 / 1,000 results