By Erwin Wu · Updated 2026-09-21
When feeding long podcasts, interviews, and lectures into large language models, token consumption dictates context fit, processing speed, and API cost. But how many tokens does spoken audio actually produce per minute or per hour? Here is a rigorous benchmark measuring speech rate, token density, and formatting overhead across modern frontier models including GPT-5.6, Claude Opus 4.6 / 4.8, and Gemini 3.7 / 3.8.
Based on empirical testing across conversational podcasts, university lectures, and multi-speaker interviews, here is the baseline token consumption table for English audio:
| Audio Duration | Spoken Words | Tokens (Clean Markdown) | Tokens (Raw SRT) | % of Claude Opus (200k) Window |
|---|---|---|---|---|
| 1 minute | 140 – 160 words | 175 – 210 tokens | ~290 tokens | 0.1% |
| 15 minutes | 2,100 – 2,400 words | 2,600 – 3,100 tokens | ~4,400 tokens | 1.5% |
| 30 minutes | 4,200 – 4,800 words | 5,250 – 6,200 tokens | ~8,800 tokens | 3.0% |
| 60 minutes (1 hr) | 8,400 – 9,600 words | 10,500 – 12,500 tokens | ~17,800 tokens | 6.0% |
| 90 minutes | 12,600 – 14,400 words | 15,750 – 18,700 tokens | ~26,700 tokens | 9.0% |
| 120 minutes (2 hrs) | 16,800 – 19,200 words | 21,000 – 25,000 tokens | ~35,600 tokens | 11.5% |
Key takeaway: A standard 1-hour podcast produces approximately 11,500 tokens of clean text, while a 2-hour discussion averages 23,000 tokens.
The primary factor dictating token generation is speaking rate:
Many users mistakenly upload raw subtitle files (.srt or .vtt) into Claude or ChatGPT. This introduces massive, unnecessary token overhead:
1\n, 2\n), microsecond intervals (00:14:02,120 --> 00:14:05,480), and double carriage returns on every phrase. In our tests, an 11,500-token podcast ballooned to 17,800 tokens when formatted as SRT — wasting over 6,000 prompt tokens.[MM:SS]). Markdown preserves complete citation traceability while adding less than 4% token overhead compared to raw plain text.Different frontier AI architectures employ distinct Byte-Pair Encoding (BPE) tokenizers:
Context capacity insight: Because a 2-hour podcast transcript consumes roughly 23,000 tokens, it takes up only 11.5% of Claude Opus's 200k context window (and an even smaller fraction of Gemini's multi-million token window). This leaves over 85% to 90% of the model's memory free for deep reasoning, multi-turn questions, counter-argument discovery, and comprehensive synthesis.
VideoScript is purpose-built to deliver token-efficient output for AI workflows:
[MM:SS] markers.New accounts start with 120 free Credits — covering 120 minutes of media (roughly 24,000 tokens of structured text). No credit card required.
Start Transcribing Free →