datagrit

datagrit › Data › YouTube Transcript Scraper - Channels & Playlists

Data

YouTube Transcript Scraper - Channels & Playlists

Bulk YouTube transcripts from videos, whole channels and playlists: publish-date window, language priority with fallback, uploaded vs auto captions, video metadata and only-new-since-last-run.

Run it on Apify StoreUse the APIfrom $2.80 per 1,000 results + $10 per run · no code needed
from $2.80 per 1,000 results + $10 per runpay only for transcripts you get
JSON · CSV · Excelexport or call via API
Scheduled runsdaily or weekly feeds with Apify schedules
v0.4updated 2026-09-30

It pulls YouTube transcripts in bulk: paste any mix of single videos, whole channels and playlists, and get one row per video with the full transcript text, timestamped caption lines and the video's metadata (title, channel, exact publish date, duration, views, likes, category). It is built for people who need many transcripts at once: researchers, content and SEO teams, and anyone feeding YouTube content into an LLM, a RAG index or a search tool.

Why use it?

How it works

For every channel the Actor reads the channel's uploads list page by page (100 videos per page), newest first. For every video it reads the public video details (publish date, metadata) and the caption track list, picks the track that matches your languages and caption type, and downloads that caption file. That is three small requests per video, plus one listing request per 100 videos of a channel or playlist.

Is it legal to scrape YouTube transcripts?

The Actor reads only information that YouTube shows to any visitor without signing in: public video pages, captions and channel lists. It does not log in, does not use cookies of an account, does not solve CAPTCHAs and does not download video files. Transcripts are the work of their creators; how you use them (for example quoting, analysis, or training) is your responsibility under copyright law and YouTube's terms. This description is not legal advice.

Output fields

Every result is one flat record, so it drops straight into a spreadsheet, a database or a CRM.

FieldTypeDescriptionExample
foundbooleanTrue when the row carries a transcript (charged). False for free rows: a video without a usable transcript (see reason) or the single status row when no video matched.true
reasonstringWhy there is no transcript, only on found:false rows: noCaptions, languageNotAvailable, captionTypeNotAvailable, emptyTranscript, blocked (YouTube bot check after proxy rotation), unavailable, unplayable (private or members-only), ageRestricted, upcoming, error, or noVideos on the status row. Null when found is true.
messagestringHuman-readable detail for found:false rows (YouTube's own message when it gives one). Null when found is true.
videoIdstringYouTube video ID (11 characters). Null only on the status row.1qmF_znXxrE
titlestringVideo title.The Right to Life, Liberty — and Free Time | Danielle Roberts | TED
channelNamestringName of the channel that published the video.TED
channelIdstringYouTube channel ID (starts with UC).UCAuUUnT6oDeKwE6v1NGQxug
channelUrlstringLink to the channel page built from the channel ID.https://www.youtube.com/channel/UCAuUUnT6oDeKwE6v1NGQxug
publishedAtstringExact publish time of the video, ISO 8601 in UTC, from YouTube's video metadata.2026-09-29T15:00:05.000Z
durationSecondsintegerVideo length in seconds.233
viewCountintegerView count at the time of the run.59691
likeCountintegerLike count at the time of the run; null when the channel hides it.568
categorystringYouTube category of the video.People & Blogs
descriptionstringVideo description text.Do you feel short on time? Writer and work expert Danielle Roberts would like to
languagestringLanguage code of the caption track returned, as YouTube names it (en, de-DE, pt-BR).en
languageNamestringLanguage name of the returned caption track as shown by YouTube.English
captionTypestringmanual = captions uploaded by the channel, auto = YouTube speech recognition.manual
isLanguageFallbackbooleanTrue when none of the requested languages existed and the video's default caption track was returned instead.false
availableLanguagesarrayLanguage codes of every caption track the video has (uploaded and auto-generated). Empty when the video has no captions; null when YouTube did not return the list.["ar","en","fr","es"]
transcriptTextstringFull transcript as plain text: caption lines joined with spaces, HTML entities decoded and formatting tags removed.So it's exciting and ironic that I'm giving a TED Talk about time in three minut
wordCountintegerNumber of words in transcriptText.547
segmentCountintegerNumber of caption lines in the transcript.73
segmentsarrayCaption lines with start time and duration in seconds. Null when "Include timestamped segments" is off or when there is no transcript.[{"start":4.292,"duration":2.711,"text":"So i
inputstringThe input entry (video, channel or playlist link) this row came from; on the status row, all inputs.https://www.youtube.com/@TED
sourceTypestringKind of input entry the video came from: video, channel or playlist.channel
sourceTitlestringChannel name or playlist title of the input entry; null for single videos.TED
sourceUrlstringWatch URL of the video. Null only on the status row.https://www.youtube.com/watch?v=1qmF_znXxrE
scrapedAtstringISO 8601 time when the row was produced.2026-10-01T08:00:00.000Z

Sample record

{
  "found": true,
  "reason": null,
  "message": null,
  "videoId": "1qmF_znXxrE",
  "title": "The Right to Life, Liberty — and Free Time | Danielle Roberts | TED",
  "channelName": "TED",
  "channelId": "UCAuUUnT6oDeKwE6v1NGQxug",
  "channelUrl": "https://www.youtube.com/channel/UCAuUUnT6oDeKwE6v1NGQxug",
  "publishedAt": "2026-09-29T15:00:05.000Z",
  "durationSeconds": 233,
  "viewCount": 59691,
  "likeCount": 568,
  "category": "People & Blogs",
  "description": "Do you feel short on time? Writer and work expert Danielle Roberts would like to help you with that.",
  "language": "en",
  "languageName": "English",
  "captionType": "manual",
  "isLanguageFallback": false,
  "availableLanguages": [
    "ar",
    "en",
    "fr",
    "es"
  ],
  "transcriptText": "So it's exciting and ironic that I'm giving a TED Talk about time in three minutes.",
  "wordCount": 547,
  "segmentCount": 73,
  "segments": [
    {
      "start": 4.292,
      "duration": 2.711,
      "text": "So it's exciting and ironic"
    }
  ],
  "input": "https://www.youtube.com/@TED",
  "sourceType": "channel",
  "sourceTitle": "TED",
  "sourceUrl": "https://www.youtube.com/watch?v=1qmF_znXxrE",
  "scrapedAt": "2026-10-01T08:00:00.000Z"
}

Input

FieldNameTypeWhat it does
urlsVideos, channels or playlistsarrayAny mix of YouTube links and IDs, one per line: video links (watch, youtu.be, Shorts, live, embed) or 11-character video IDs; channel links (https://www.youtube.com/@TED, /channel/UC..., /c/..., /user/...) or bare @handles; playlist links (https://www.youtube.com/playlist?list=PL...) or playlist IDs. A channel link ending in /shorts or /streams reads that tab. A watch link that also carries &list= is read as that single video; use the /playlist?list= link for the whole playlist. If the list is empty, the Actor runs a small free example (3 videos from @TED) and says so in the status message.
maxItemsMaximum transcriptsintegerStop after this many transcripts in total across all sources. Only delivered transcripts count and are charged; rows for videos without captions are free and do not count.
maxVideosPerSourceMaximum videos per channel or playlistintegerHow many videos to take from each channel or playlist (channels newest first, playlists in playlist order). Videos without captions count toward this limit; videos outside the publish-date window, without the title keywords or already delivered by an earlier run do not. 0 = no per-source limit (the run still stops at Maximum transcripts and reads at most 50 listing pages, about 5,000 videos, per source).
channelContentChannel contentstringWhich uploads to read from a channel: long-form videos (the channel's Videos tab), Shorts, live stream replays, or all uploads mixed. A channel link ending in /shorts or /streams overrides this for that channel. YouTube lists only the newest Shorts of a channel (measured: 100 of 520 for @TED); the run status says when a list is shorter than the channel's total.
publishedAfterPublished on or afterstringOnly videos published on or after this date (UTC). Use a date like 2026-09-01 or a period back from the run start like 7 days, 2 weeks, 3 months or 1 year. Checked against each video's exact publish date. On channels the listing stops once it reaches older videos. Leave empty for no lower limit.
publishedBeforePublished beforestringOnly videos published before this date (UTC, the day itself excluded), same format as above. Leave empty for no upper limit. On a channel the newer videos still have to be listed first, so a far-back window reads more listing pages.
titleKeywordsTitle keywordsarrayOnly videos whose title contains at least one of these words or phrases (case-insensitive substring match). Leave empty to take every video.
languagesCaption languagesarrayCaption language codes in order of preference, for example en, es, pt-BR. The first language the video has wins; en also matches en-US and en-GB, and pt-BR also matches pt. Leave empty to take the video's default caption track (usually its spoken language).
captionTypeCaption typestringUploaded (manual) captions are written by the channel; auto-generated captions come from YouTube speech recognition. Choose which to prefer within each language, or accept only one kind.
languageFallbackFall back to another languagebooleanWhen a video has none of the caption languages above, return its default caption track instead (the row says isLanguageFallback: true). Turn off to get a free row with reason languageNotAvailable instead.
includeSegmentsInclude timestamped segmentsbooleanAdd the caption lines with start time and duration in seconds (segments). The full plain text (transcriptText) is always included.
onlyNewSinceLastRunOnly videos new since my last runbooleanFor scheduled runs. On a channel the Actor remembers every video an earlier run with the same settings read (with or without a transcript) and stops the listing (newest first) at the first of them, so each run returns only the uploads published since the last run; the first run returns the newest videos up to the per-channel limit. Videos in a temporary state are not remembered and are read again next time: upcoming premieres and scheduled live streams, empty caption files, YouTube bot checks, failed requests and videos without a publish date while a date window is set. On a channel they are read again only while they are newer than the stop point. On playlists and single videos, videos whose transcript was already delivered are skipped and videos without a transcript are tried again. The memory is kept per source and per combination of channel content, date window, title keywords, languages and caption settings (plus the optional memory name).
stateKeyOnly-new memory namestringOptional name that keeps a separate only-new memory, for example one per schedule or per client. Runs with the same sources, settings and name share the memory.
proxyConfigurationProxy configurationobjectYouTube is read through Apify Proxy (datacenter) by default. The Actor switches to a new proxy session when YouTube asks for a sign-in bot check, rate-limits, or returns an empty caption file. If a run fails with a block message, switch to the RESIDENTIAL group.

Call it from your code

Run the Actor and get the results in one request. Replace YOUR_APIFY_TOKEN with the token from your Apify account settings.

curl -X POST "https://api.apify.com/v2/acts/datagrit~youtube-channel-playlist-transcripts/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://www.youtube.com/@TED","https://www.youtube.com/playlist?list=PLUCPc-R61w-s"],"maxItems":20,"maxVideosPerSource":8,"languages":["en"]}'
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagrit/youtube-channel-playlist-transcripts').call({
  "urls": [
    "https://www.youtube.com/@TED",
    "https://www.youtube.com/playlist?list=PLUCPc-R61w-s"
  ],
  "maxItems": 20,
  "maxVideosPerSource": 8,
  "languages": [
    "en"
  ]
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.length, items[0]);

Install with npm i apify-client.

from apify_client import ApifyClient
import os

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("datagrit/youtube-channel-playlist-transcripts").call(run_input={
  "urls": [
    "https://www.youtube.com/@TED",
    "https://www.youtube.com/playlist?list=PLUCPc-R61w-s"
  ],
  "maxItems": 20,
  "maxVideosPerSource": 8,
  "languages": [
    "en"
  ]
})
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(len(items), items[0] if items else None)

Install with pip install apify-client.

Frequently asked questions

Why do some videos have no transcript?

Many videos have no captions at all (music, very new uploads, some Shorts). Such videos get a free row with reason `noCaptions`. The status message counts every reason.

Why does a run end without transcripts but green?

When none of the videos has captions in your languages (with fallback off) or of the caption type you asked for, every video gets a free row with the reason and the run succeeds. A run fails only when videos do have a matching caption track and YouTube still returns no text.

What does "blocked" mean?

YouTube sometimes answers cloud IP addresses with a "Sign in to confirm you're not a bot" check or an empty caption file. The Actor then switches to a new proxy session and retries a few times. If YouTube still refuses, the video gets a free row with reason `blocked`. If no transcript at all could be read although videos have captions, the run fails with a message instead of finishing green; run it again with the RESIDENTIAL proxy group. When more than half of the videos in a run are blocked, the status message warns and gives the count.

How many videos can one run read?

There is no fixed cap: the run stops at Maximum transcripts or your spending limit. The Actor reads at most 50 listing pages (about 5,000 videos) per channel or playlist, and says so in the status message when it stops there. YouTube lists only the newest Shorts of a channel (for @TED: 100 of 520), and the status message says when a list is shorter than the channel's total.

How often can I run it?

As often as you like. For a daily or weekly feed of new uploads, schedule it with **Only videos new since my last run** on: the first run returns the newest videos up to the per-channel limit, every later run only the uploads published since then (a run with nothing new returns one free row "No new videos since your last run"). A relative date such as "7 days" works too, but returns the same video again in each run that covers it.

Can "only new" miss a video?

Yes, in two cases. A video that becomes public later but keeps an older publish date (for example a private upload made public weeks later) sits below the last video read in the channel list, so the next run stops before it. Run once without "only new" and with a publish-date window to catch such videos. Also, on a channel a video that had no captions when it was read is not read again, even if captions are added later. Videos in a temporary state (an upcoming premiere, an empty caption file, a YouTube bot check) are not remembered and are read again by the next run, as long as they are newer than the last video read.

Are timestamps included?

Yes, each caption line has a start time and duration in seconds.

Can it translate transcripts?

No. It returns caption tracks that exist on YouTube (uploaded or auto-generated), in the language you ask for when the video has it.

Try YouTube Transcript Scraper - Channels & Playlists on Apify

Related Actors

Company data

French Company Finder - Sirene Financials

French company lead lists from Sirene screened by net result and revenue, with net margin, size, matching establishment and optional directors.

from $5.60 / 1,000 results
Jobs & salaries

Career Site Jobs Aggregator - ATS Salary Data

Open jobs from Greenhouse, Lever, Ashby, Workday, Workable, SmartRecruiters, Recruitee and Personio in one schema with normalized salaries.

from $2.80 / 1,000 results