TED Contract Expiry Radar - Recompete Leads
Find EU public contracts approaching expiry from TED award notices: incumbent, buyer, value, end date and renewal options.
datagrit › Data › YouTube Transcript Scraper - Channels & Playlists
DataBulk YouTube transcripts from videos, whole channels and playlists: publish-date window, language priority with fallback, uploaded vs auto captions, video metadata and only-new-since-last-run.
It pulls YouTube transcripts in bulk: paste any mix of single videos, whole channels and playlists, and get one row per video with the full transcript text, timestamped caption lines and the video's metadata (title, channel, exact publish date, duration, views, likes, category). It is built for people who need many transcripts at once: researchers, content and SEO teams, and anyone feeding YouTube content into an LLM, a RAG index or a search tool.
For every channel the Actor reads the channel's uploads list page by page (100 videos per page), newest first. For every video it reads the public video details (publish date, metadata) and the caption track list, picks the track that matches your languages and caption type, and downloads that caption file. That is three small requests per video, plus one listing request per 100 videos of a channel or playlist.
The Actor reads only information that YouTube shows to any visitor without signing in: public video pages, captions and channel lists. It does not log in, does not use cookies of an account, does not solve CAPTCHAs and does not download video files. Transcripts are the work of their creators; how you use them (for example quoting, analysis, or training) is your responsibility under copyright law and YouTube's terms. This description is not legal advice.
Every result is one flat record, so it drops straight into a spreadsheet, a database or a CRM.
| Field | Type | Description | Example |
|---|---|---|---|
found | boolean | True when the row carries a transcript (charged). False for free rows: a video without a usable transcript (see reason) or the single status row when no video matched. | true |
reason | string | Why there is no transcript, only on found:false rows: noCaptions, languageNotAvailable, captionTypeNotAvailable, emptyTranscript, blocked (YouTube bot check after proxy rotation), unavailable, unplayable (private or members-only), ageRestricted, upcoming, error, or noVideos on the status row. Null when found is true. | |
message | string | Human-readable detail for found:false rows (YouTube's own message when it gives one). Null when found is true. | |
videoId | string | YouTube video ID (11 characters). Null only on the status row. | 1qmF_znXxrE |
title | string | Video title. | The Right to Life, Liberty — and Free Time | Danielle Roberts | TED |
channelName | string | Name of the channel that published the video. | TED |
channelId | string | YouTube channel ID (starts with UC). | UCAuUUnT6oDeKwE6v1NGQxug |
channelUrl | string | Link to the channel page built from the channel ID. | https://www.youtube.com/channel/UCAuUUnT6oDeKwE6v1NGQxug |
publishedAt | string | Exact publish time of the video, ISO 8601 in UTC, from YouTube's video metadata. | 2026-09-29T15:00:05.000Z |
durationSeconds | integer | Video length in seconds. | 233 |
viewCount | integer | View count at the time of the run. | 59691 |
likeCount | integer | Like count at the time of the run; null when the channel hides it. | 568 |
category | string | YouTube category of the video. | People & Blogs |
description | string | Video description text. | Do you feel short on time? Writer and work expert Danielle Roberts would like to |
language | string | Language code of the caption track returned, as YouTube names it (en, de-DE, pt-BR). | en |
languageName | string | Language name of the returned caption track as shown by YouTube. | English |
captionType | string | manual = captions uploaded by the channel, auto = YouTube speech recognition. | manual |
isLanguageFallback | boolean | True when none of the requested languages existed and the video's default caption track was returned instead. | false |
availableLanguages | array | Language codes of every caption track the video has (uploaded and auto-generated). Empty when the video has no captions; null when YouTube did not return the list. | ["ar","en","fr","es"] |
transcriptText | string | Full transcript as plain text: caption lines joined with spaces, HTML entities decoded and formatting tags removed. | So it's exciting and ironic that I'm giving a TED Talk about time in three minut |
wordCount | integer | Number of words in transcriptText. | 547 |
segmentCount | integer | Number of caption lines in the transcript. | 73 |
segments | array | Caption lines with start time and duration in seconds. Null when "Include timestamped segments" is off or when there is no transcript. | [{"start":4.292,"duration":2.711,"text":"So i |
input | string | The input entry (video, channel or playlist link) this row came from; on the status row, all inputs. | https://www.youtube.com/@TED |
sourceType | string | Kind of input entry the video came from: video, channel or playlist. | channel |
sourceTitle | string | Channel name or playlist title of the input entry; null for single videos. | TED |
sourceUrl | string | Watch URL of the video. Null only on the status row. | https://www.youtube.com/watch?v=1qmF_znXxrE |
scrapedAt | string | ISO 8601 time when the row was produced. | 2026-10-01T08:00:00.000Z |
{
"found": true,
"reason": null,
"message": null,
"videoId": "1qmF_znXxrE",
"title": "The Right to Life, Liberty — and Free Time | Danielle Roberts | TED",
"channelName": "TED",
"channelId": "UCAuUUnT6oDeKwE6v1NGQxug",
"channelUrl": "https://www.youtube.com/channel/UCAuUUnT6oDeKwE6v1NGQxug",
"publishedAt": "2026-09-29T15:00:05.000Z",
"durationSeconds": 233,
"viewCount": 59691,
"likeCount": 568,
"category": "People & Blogs",
"description": "Do you feel short on time? Writer and work expert Danielle Roberts would like to help you with that.",
"language": "en",
"languageName": "English",
"captionType": "manual",
"isLanguageFallback": false,
"availableLanguages": [
"ar",
"en",
"fr",
"es"
],
"transcriptText": "So it's exciting and ironic that I'm giving a TED Talk about time in three minutes.",
"wordCount": 547,
"segmentCount": 73,
"segments": [
{
"start": 4.292,
"duration": 2.711,
"text": "So it's exciting and ironic"
}
],
"input": "https://www.youtube.com/@TED",
"sourceType": "channel",
"sourceTitle": "TED",
"sourceUrl": "https://www.youtube.com/watch?v=1qmF_znXxrE",
"scrapedAt": "2026-10-01T08:00:00.000Z"
}
| Field | Name | Type | What it does |
|---|---|---|---|
urls | Videos, channels or playlists | array | Any mix of YouTube links and IDs, one per line: video links (watch, youtu.be, Shorts, live, embed) or 11-character video IDs; channel links (https://www.youtube.com/@TED, /channel/UC..., /c/..., /user/...) or bare @handles; playlist links (https://www.youtube.com/playlist?list=PL...) or playlist IDs. A channel link ending in /shorts or /streams reads that tab. A watch link that also carries &list= is read as that single video; use the /playlist?list= link for the whole playlist. If the list is empty, the Actor runs a small free example (3 videos from @TED) and says so in the status message. |
maxItems | Maximum transcripts | integer | Stop after this many transcripts in total across all sources. Only delivered transcripts count and are charged; rows for videos without captions are free and do not count. |
maxVideosPerSource | Maximum videos per channel or playlist | integer | How many videos to take from each channel or playlist (channels newest first, playlists in playlist order). Videos without captions count toward this limit; videos outside the publish-date window, without the title keywords or already delivered by an earlier run do not. 0 = no per-source limit (the run still stops at Maximum transcripts and reads at most 50 listing pages, about 5,000 videos, per source). |
channelContent | Channel content | string | Which uploads to read from a channel: long-form videos (the channel's Videos tab), Shorts, live stream replays, or all uploads mixed. A channel link ending in /shorts or /streams overrides this for that channel. YouTube lists only the newest Shorts of a channel (measured: 100 of 520 for @TED); the run status says when a list is shorter than the channel's total. |
publishedAfter | Published on or after | string | Only videos published on or after this date (UTC). Use a date like 2026-09-01 or a period back from the run start like 7 days, 2 weeks, 3 months or 1 year. Checked against each video's exact publish date. On channels the listing stops once it reaches older videos. Leave empty for no lower limit. |
publishedBefore | Published before | string | Only videos published before this date (UTC, the day itself excluded), same format as above. Leave empty for no upper limit. On a channel the newer videos still have to be listed first, so a far-back window reads more listing pages. |
titleKeywords | Title keywords | array | Only videos whose title contains at least one of these words or phrases (case-insensitive substring match). Leave empty to take every video. |
languages | Caption languages | array | Caption language codes in order of preference, for example en, es, pt-BR. The first language the video has wins; en also matches en-US and en-GB, and pt-BR also matches pt. Leave empty to take the video's default caption track (usually its spoken language). |
captionType | Caption type | string | Uploaded (manual) captions are written by the channel; auto-generated captions come from YouTube speech recognition. Choose which to prefer within each language, or accept only one kind. |
languageFallback | Fall back to another language | boolean | When a video has none of the caption languages above, return its default caption track instead (the row says isLanguageFallback: true). Turn off to get a free row with reason languageNotAvailable instead. |
includeSegments | Include timestamped segments | boolean | Add the caption lines with start time and duration in seconds (segments). The full plain text (transcriptText) is always included. |
onlyNewSinceLastRun | Only videos new since my last run | boolean | For scheduled runs. On a channel the Actor remembers every video an earlier run with the same settings read (with or without a transcript) and stops the listing (newest first) at the first of them, so each run returns only the uploads published since the last run; the first run returns the newest videos up to the per-channel limit. Videos in a temporary state are not remembered and are read again next time: upcoming premieres and scheduled live streams, empty caption files, YouTube bot checks, failed requests and videos without a publish date while a date window is set. On a channel they are read again only while they are newer than the stop point. On playlists and single videos, videos whose transcript was already delivered are skipped and videos without a transcript are tried again. The memory is kept per source and per combination of channel content, date window, title keywords, languages and caption settings (plus the optional memory name). |
stateKey | Only-new memory name | string | Optional name that keeps a separate only-new memory, for example one per schedule or per client. Runs with the same sources, settings and name share the memory. |
proxyConfiguration | Proxy configuration | object | YouTube is read through Apify Proxy (datacenter) by default. The Actor switches to a new proxy session when YouTube asks for a sign-in bot check, rate-limits, or returns an empty caption file. If a run fails with a block message, switch to the RESIDENTIAL group. |
Run the Actor and get the results in one request. Replace YOUR_APIFY_TOKEN with the token from your Apify account settings.
curl -X POST "https://api.apify.com/v2/acts/datagrit~youtube-channel-playlist-transcripts/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":["https://www.youtube.com/@TED","https://www.youtube.com/playlist?list=PLUCPc-R61w-s"],"maxItems":20,"maxVideosPerSource":8,"languages":["en"]}'import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagrit/youtube-channel-playlist-transcripts').call({
"urls": [
"https://www.youtube.com/@TED",
"https://www.youtube.com/playlist?list=PLUCPc-R61w-s"
],
"maxItems": 20,
"maxVideosPerSource": 8,
"languages": [
"en"
]
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.length, items[0]);Install with npm i apify-client.
from apify_client import ApifyClient
import os
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("datagrit/youtube-channel-playlist-transcripts").call(run_input={
"urls": [
"https://www.youtube.com/@TED",
"https://www.youtube.com/playlist?list=PLUCPc-R61w-s"
],
"maxItems": 20,
"maxVideosPerSource": 8,
"languages": [
"en"
]
})
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(len(items), items[0] if items else None)Install with pip install apify-client.
Many videos have no captions at all (music, very new uploads, some Shorts). Such videos get a free row with reason `noCaptions`. The status message counts every reason.
When none of the videos has captions in your languages (with fallback off) or of the caption type you asked for, every video gets a free row with the reason and the run succeeds. A run fails only when videos do have a matching caption track and YouTube still returns no text.
YouTube sometimes answers cloud IP addresses with a "Sign in to confirm you're not a bot" check or an empty caption file. The Actor then switches to a new proxy session and retries a few times. If YouTube still refuses, the video gets a free row with reason `blocked`. If no transcript at all could be read although videos have captions, the run fails with a message instead of finishing green; run it again with the RESIDENTIAL proxy group. When more than half of the videos in a run are blocked, the status message warns and gives the count.
There is no fixed cap: the run stops at Maximum transcripts or your spending limit. The Actor reads at most 50 listing pages (about 5,000 videos) per channel or playlist, and says so in the status message when it stops there. YouTube lists only the newest Shorts of a channel (for @TED: 100 of 520), and the status message says when a list is shorter than the channel's total.
As often as you like. For a daily or weekly feed of new uploads, schedule it with **Only videos new since my last run** on: the first run returns the newest videos up to the per-channel limit, every later run only the uploads published since then (a run with nothing new returns one free row "No new videos since your last run"). A relative date such as "7 days" works too, but returns the same video again in each run that covers it.
Yes, in two cases. A video that becomes public later but keeps an older publish date (for example a private upload made public weeks later) sits below the last video read in the channel list, so the next run stops before it. Run once without "only new" and with a publish-date window to catch such videos. Also, on a channel a video that had no captions when it was read is not read again, even if captions are added later. Videos in a temporary state (an upcoming premiere, an empty caption file, a YouTube bot check) are not remembered and are read again by the next run, as long as they are newer than the last video read.
Yes, each caption line has a start time and duration in seconds.
No. It returns caption tracks that exist on YouTube (uploaded or auto-generated), in the language you ask for when the video has it.
Find EU public contracts approaching expiry from TED award notices: incumbent, buyer, value, end date and renewal options.
French company lead lists from Sirene screened by net result and revenue, with net margin, size, matching establishment and optional directors.
Ashby job postings with normalized annual salary ranges, equity flags and new-since-last-run detection.
Open jobs from Greenhouse, Lever, Ashby, Workday, Workable, SmartRecruiters, Recruitee and Personio in one schema with normalized salaries.