How many tokens in a 1-hour podcast?

By Erwin Wu · Updated 2026-09-21

When feeding long podcasts, interviews, and lectures into large language models, token consumption dictates context fit, processing speed, and API cost. But how many tokens does spoken audio actually produce per minute or per hour? Here is a rigorous benchmark measuring speech rate, token density, and formatting overhead across modern frontier models including GPT-5.6, Claude Opus 4.6 / 4.8, and Gemini 3.7 / 3.8.

Executive benchmark: audio duration to token count

Based on empirical testing across conversational podcasts, university lectures, and multi-speaker interviews, here is the baseline token consumption table for English audio:

Audio DurationSpoken WordsTokens (Clean Markdown)Tokens (Raw SRT)% of Claude Opus (200k) Window
1 minute140 – 160 words175 – 210 tokens~290 tokens0.1%
15 minutes2,100 – 2,400 words2,600 – 3,100 tokens~4,400 tokens1.5%
30 minutes4,200 – 4,800 words5,250 – 6,200 tokens~8,800 tokens3.0%
60 minutes (1 hr)8,400 – 9,600 words10,500 – 12,500 tokens~17,800 tokens6.0%
90 minutes12,600 – 14,400 words15,750 – 18,700 tokens~26,700 tokens9.0%
120 minutes (2 hrs)16,800 – 19,200 words21,000 – 25,000 tokens~35,600 tokens11.5%

Key takeaway: A standard 1-hour podcast produces approximately 11,500 tokens of clean text, while a 2-hour discussion averages 23,000 tokens.

Speech rate variables: conversational vs. technical pacing

The primary factor dictating token generation is speaking rate:

  • Casual conversational podcasts (e.g., panel discussions, interviews, podcasts): Speakers average 150–175 words per minute. Because English spoken words average roughly 1.3 tokens in modern tokenizers, conversational audio produces ~200 tokens per minute.
  • Structured lectures & educational audio (e.g., academic lectures, scientific briefings): Pacing is more deliberate, averaging 120–140 words per minute (~160–180 tokens per minute).
  • Multilingual considerations: Spoken Mandarin Chinese averages 200–240 characters per minute. On modern high-compression multilingual tokenizers, 1 Chinese character averages 1.3 to 1.5 tokens, generating roughly 280–350 tokens per minute of audio.

The formatting overhead trap: SRT vs. clean Markdown

Many users mistakenly upload raw subtitle files (.srt or .vtt) into Claude or ChatGPT. This introduces massive, unnecessary token overhead:

  • Raw SRT files (+50% token bloat) — Subtitle files repeat numerical line indices (1\n, 2\n), microsecond intervals (00:14:02,120 --> 00:14:05,480), and double carriage returns on every phrase. In our tests, an 11,500-token podcast ballooned to 17,800 tokens when formatted as SRT — wasting over 6,000 prompt tokens.
  • Clean Markdown (optimal) — Groups speech into natural paragraphs and anchors each section with a compact timestamp ([MM:SS]). Markdown preserves complete citation traceability while adding less than 4% token overhead compared to raw plain text.

Tokenizer comparisons: GPT-5.6 vs. Claude Opus 4.6 / 4.8 vs. Gemini 3.8

Different frontier AI architectures employ distinct Byte-Pair Encoding (BPE) tokenizers:

  • OpenAI GPT-5 / GPT-5.6 — Advanced high-compression vocabulary tailored for complex multimodal inputs (~11,100 tokens for 1 hour of spoken English).
  • Anthropic Claude Opus 4.6 / 4.8 — Optimized for natural syntax and extensive document reasoning (~11,500 tokens for 1 hour of clean Markdown speech).
  • Google Gemini 3.7 / 3.8 Pro — Built for massive multimodal context ingestion, averaging ~11,300 tokens for 1 hour of spoken text.

Context capacity insight: Because a 2-hour podcast transcript consumes roughly 23,000 tokens, it takes up only 11.5% of Claude Opus's 200k context window (and an even smaller fraction of Gemini's multi-million token window). This leaves over 85% to 90% of the model's memory free for deep reasoning, multi-turn questions, counter-argument discovery, and comprehensive synthesis.

Optimizing your audio pipeline with VideoScript

VideoScript is purpose-built to deliver token-efficient output for AI workflows:

  • Filters stuttering and filler acoustic noise while keeping speech verbatim.
  • Exports directly to clean, token-efficient Markdown with compact [MM:SS] markers.
  • Avoids the 50% token penalty of raw subtitle files.

Convert your audio into token-efficient Markdown

New accounts start with 120 free Credits — covering 120 minutes of media (roughly 24,000 tokens of structured text). No credit card required.

Start Transcribing Free →

Frequently asked questions

  • Will a 2-hour podcast fit into a single Claude Opus or GPT-5.6 prompt?
    Yes. A 2-hour podcast transcript requires roughly 23,000 tokens in Markdown, easily fitting within Claude Opus (200k+ tokens), GPT-5.6 (256k+ tokens), and Gemini 3.8 (1M+ tokens).
  • Why does SRT waste so many tokens compared to Markdown?
    SRT repeats timestamp headers, millisecond codes, and line numbers on every 3 to 5 words, creating hundreds of redundant formatting tokens that waste context space.
  • How can I estimate tokens for non-English audio?
    For Spanish, French, or German audio, token counts are roughly 10–15% higher than English (~13,000 tokens/hour). For Chinese or Japanese, expect ~18,000–22,000 tokens per hour depending on character density.

Related guides

  • YouTube transcripts
  • Podcast transcripts
  • AI-ready transcripts
  • Transcript vs. summary: which do you need?
  • Turn a lecture into searchable notes
  • How VideoScript pricing works

How VideoScript works · Pricing