Audio-to-Timed-Text

Audio transcription and subtitle timing API

Extract lyrics, line-level and word-level timing, and production-ready subtitle files from public audio URLs.

View current credit pricing

Why teams evaluate this API

Designed for products that need timed lyrics, captions, karaoke data, searchable transcripts, or subtitle input for downstream video workflows.

Best-fit teams

  • Music and karaoke products that need synchronized lyrics
  • Creator tools exporting SRT, VTT, LRC, or JSON timing
  • Media pipelines that need timed text before video composition

Core capabilities

  • Lyrics extraction with line-level and word-level timestamps
  • SRT and JSON output by default, with optional VTT and LRC artifacts
  • Automatic source-duration detection for quote and task sizing
  • Asynchronous status, webhook delivery, idempotency, and signed downloads

Typical implementation flow

  1. 1Request a quote or let the service inspect the audio duration
  2. 2Create an idempotent subtitle extraction task
  3. 3Poll the task or receive a webhook when processing completes
  4. 4Read lyrics and timing, then download the requested subtitle artifacts

View current credit pricing

FAQ

How is Subtitle Sync billed?

Billing is prorated by actual audio duration and rounded up to whole credits with a minimum charge. Use the Quote API for the current amount before creating a task.

Which files can the API generate?

Every task can return lyrics and structured timing. SRT and JSON are the defaults; VTT and LRC can be requested in the same task.

Can Subtitle Sync feed a music-video workflow?

Yes. The managed MV workflow can use the same timed-text pipeline when an audio source does not already include aligned lyrics.

Is the duration hint required?

No. The service can inspect a supported public audio source before quoting or creating the task. Supplying a known duration can make a quote request faster.