~/social/tiktok-transcript-scraper

TikTok Transcript & Subtitle Scraper — JSON, SRT, VTT, LLM

Extract TikTok subtitles and transcripts from video URLs, short links, or a whole @profile as JSON, SRT, VTT, plain text, or LLM-ready text. Native captions — no Whisper, no API key, about two seconds per video.

media TypeScript Global
proooxy/tiktok-transcript-scraper — spec
categorysocial / media
languageTypeScript
stackTypeScript
marketsGlobal
outputclean, RAG-ready JSON

key features

Reads TikTok's own captions instead of re-transcribing audio — about 2 seconds per video against 10–30 for a Whisper-based tool

Exact on-screen text, so nothing is misheard or invented

Five output formats — timestamped JSON, SRT, VTT, plain text, and LLM-ready text stripped of [Music] and speaker labels

One input field takes video URLs, vm./vt. short links, @username handles, and raw video IDs, mixed freely

Profile input enumerates a creator's videos, with maxVideos as the budget across the whole run

Engagement metadata per video — plays, likes, comments, shares, duration, word count, segment count

Language priority list with prefix matching (eng matches eng-US) and a toggle for auto-generated captions

Failed extractions, short-link resolution and profile enumeration are never charged

use cases

  • Building RAG or fine-tuning corpora from short-form video
  • Repurposing spoken content into blog posts, newsletters, and social copy
  • Competitive content analysis across every creator in a niche
  • Producing SRT/VTT subtitle files for reposting and localisation
  • Hook research — reading the first seconds of the top videos in a category
  • Brand-safety and disclosure review of sponsored posts at scale

input parameters

ParameterTypeRequiredDescription
urlsarrayrequiredTikTok video URLs, short links, @username handles, or raw video IDs
outputFormatstringoptionaljson, srt, vtt, text, or llm
languagesarrayoptionalPreferred caption languages in priority order; prefix matching, so eng matches eng-US
includeAutoGeneratedbooleanoptionalInclude TikTok's auto-generated captions when no manual ones exist
maxVideosintegeroptionalCap on videos processed across the whole run; 0 means unlimited
maxConcurrencyintegeroptionalParallel workers, 1–10; lower is safer against rate limits
proxyConfigurationobjectoptionalResidential proxy is strongly recommended — TikTok blocks datacenter IPs

Output Example

 1{
 2  "videoId": "7627209981670001950",
 3  "url": "https://www.tiktok.com/@tiktok/video/7627209981670001950",
 4  "title": "your TikTok grandpa @writers cramp is proud of you",
 5  "authorName": "TikTok",
 6  "authorId": "tiktok",
 7  "createTime": "2026-04-10T19:10:35.000Z",
 8  "playCount": 54600,
 9  "likeCount": 3656,
10  "commentCount": 837,
11  "shareCount": 291,
12  "availableLanguages": ["eng-US"],
13  "language": "eng-US",
14  "isAutoGenerated": true,
15  "segments": [
16    { "text": "I'm 81 years old and I'm known as the TikTok Grandpa", "start": 0.04, "end": 3.64 },
17    { "text": "my name is Ian Smith", "start": 3.641, "end": 4.721 }
18  ],
19  "text": "I'm 81 years old and I'm known as the TikTok Grandpa my name is Ian Smith...",
20  "duration": 77,
21  "wordCount": 243,
22  "segmentCount": 29,
23  "extractedAt": "2026-04-12T02:45:36.000Z",
24  "error": null
25}

Tips

Start with one video. A single URL confirms the schema and the language you get back before you point it at a profile. Set maxVideos on profile runs. It is a budget for the whole run, not per profile, so a mixed input cannot overshoot. Leave concurrency at 3. TikTok rate-limits aggressively; higher values trade reliability for a little speed. Pick llm for AI pipelines. That format drops [Music], speaker labels and annotations, which otherwise become noise in embeddings.

faq

Does this use Whisper or speech-to-text?
No. It pulls the caption files TikTok itself generated and serves from its CDN, so the text is exactly what viewers see on screen and a video takes about two seconds instead of half a minute.
What happens to a video with no captions?
It comes back with an error field set and costs nothing. You are charged per successfully extracted transcript, not per URL submitted.
Can I scrape a whole account?
Yes — pass the profile URL or @username and the Actor enumerates that creator's videos. Use maxVideos to cap the run; videos you name explicitly are taken first, then profiles fill what is left.

related in ~/social

Run TikTok Transcript & Subtitle Scraper — JSON, SRT, VTT, LLM, or get a custom build

Start extracting on Apify in minutes, or hire me to build a bespoke scraper and RAG pipeline for your exact source and schema.

run on Apify get custom data