← Blog

Subtitles

How to make subtitles for video automatically: SRT and VTT via API

Subtitles stopped being optional a long time ago. Up to 85% of social videos are watched with sound off – on the subway, in meetings, in bed at night. If your video has no text track, you simply lose that audience: a viewer swipes past in three seconds. In this guide I'll show how to make subtitles automatically – via the speech recognition API, with timestamps and SRT or VTT export, without hours of manual work.

Why subtitles are a must, not a nice-to-have

The first reason is silent viewing. This is the main way people consume short video: they scroll a feed in a public place, without headphones, with sound off by default. Without subtitles your video is mute and useless to them; with subtitles it becomes full content they watch to the end. The difference in completion rate can be twofold.

The second reason is accessibility. A significant part of the audience has some form of hearing difficulty, and for them subtitles are the only way to hear you. In many countries text tracks for government and educational content are required by law.

The third and fourth are search and engagement. Subtitle text gets indexed: a video on your site together with its transcript starts ranking for the phrases in the text. And a viewer who reads and listens at the same time stays engaged longer. The conclusion is simple: if you produce video systematically, you need subtitles at industrial scale – and doing them by hand is not rational.

Why the manual approach loses

Let's count the economics of manual transcription. An experienced person spends three to five hours transcribing one hour of audio – and that's without timestamp alignment. Adding timestamps takes another hour and a half. So a forty-minute podcast is nearly a full working day.

It gets worse: people get tired. On long recordings with several voices, accents, and technical vocabulary, errors pile up – missed lines, unfamiliar terms, mangled names. All of that has to be caught during proofreading, which is more time. Automatic recognition does the same volume in minutes and returns the result with word-level timestamps.

There's one more difference. Built-in auto-captions on video platforms usually don't give you a subtitle file: you see text on screen, but you can't download SRT/VTT, edit it, or embed it in your own player. An API approach returns exactly that file – portable, editable, and ready to plug into any pipeline.

How automatic subtitles work under the hood

You feed in video or audio, and three things happen. First, audio extraction: a video file contains a video track and an audio track, and recognition only needs the audio track, so it's pulled out separately – usually as WAV or MP3. This is done with a single command and takes seconds.

Second, speech recognition. The neural network cuts the audio into short windows, turns each into a spectrogram – a picture of how frequencies are distributed over time – and predicts the sequence of words from it. Then a language model adds punctuation and chooses meaningful words over random sound combinations. The output is text split into segments where every word is tied to its own second.

Third, assembling the subtitle file: the segments and timestamps are combined into a ready SRT or VTT, where each line gets a start and end mark. That's the file you attach to a player or upload to YouTube. One important note: the quality of the result depends directly on the quality of the input audio – more on that below.

SRT or VTT: which to choose

SRT (SubRip) is the oldest and most universal format. It's a plain text file where each block is a number, a line with a time range like 00:00:01,000 --> 00:00:04,000, and the subtitle text. SRT is understood by everyone: video editors, YouTube, VLC, Telegram, almost any player. If you don't know what to pick, take SRT.

VTT (WebVTT) is a newer, web-oriented format. The structure is almost the same, but it supports styling: text color, on-screen positioning, alignment. It's the HTML5 standard, so VTT is great for your own web player or app.

A practical rule: for uploading to video hosts and editors, use SRT; for your own web player or app, use VTT. Both formats are built from the same timestamps, so switching between them is a matter of one request parameter – no re-recognition needed.

Step-by-step guide: from video to a ready file

You'll need any REST client or Python, and only ffmpeg for audio extraction. Step one – pull out the audio. For Russian, the optimal choice is WAV at 16 kHz, mono:

bash
ffmpeg -i video.mp4 -vn -ac 1 -ar 16000 audio.wav

Let's decode the flags: -vn drops the video track, -ac 1 makes it mono, -ar 16000 sets the 16 kHz sample rate. If your video already has good sound, you can send MP3 too – the API accepts wav, mp3, ogg, and flac. Step two – send the audio for recognition via POST /v1/transcribe. The API key goes in the X-Api-Key header – only there, never in the URL:

Python
import requests

API = "https://ttsapi.ru/v1"
HEADERS = {"X-Api-Key": "rtt_…"}

with open("audio.wav", "rb") as f:
    r = requests.post(
        f"{API}/transcribe",
        headers=HEADERS,
        files={"audio": ("audio.wav", f, "audio/wav")},
        data={"language": "ru"},
    )

job_id = r.json()["job_id"]

The language parameter hints the recording language to the model – if you know it, pass it explicitly, it almost always improves accuracy. Recognition of long recordings runs asynchronously: you get a job_id and poll the status until it's done. Once the status is completed, fetch the ready subtitle file:

Python
import time

while True:
    status = requests.get(f"{API}/transcribe/{job_id}", headers=HEADERS).json()
    if status["status"] in ("completed", "failed"):
        break
    time.sleep(2)

srt = requests.get(
    f"{API}/transcribe/{job_id}/subtitles",
    headers=HEADERS,
    params={"format": "srt"},
)
open("subtitles.srt", "w", encoding="utf-8").write(srt.text)

For VTT, change a single parameter – params={"format": "vtt"}. The file is ready: upload it to YouTube, drop it into a Telegram bot, or attach it to the player on your site. The whole cycle for a ten-minute clip takes under a minute – manually that would be an hour and a half.

How to make subtitles high quality

Audio quality decides everything. Record at 16 kHz or higher, without heavy compression or echo. A 64 kbps MP3 is noticeably worse than WAV or 192+ kbps. One speaker per channel and minimal background noise help more than any model setting.

Pass the language explicitly and watch the timestamps. When the model doesn't spend effort detecting the language, it makes fewer mistakes on short phrases. A good API returns not just text but a confidence score for every word – highlight low-confidence words when proofreading so you check only the doubtful spots instead of reading the whole text.

Use diarization deliberately. If there are several speakers in the video – an interview, a podcast, a call – turn on speaker labeling with diarization=true and you'll get subtitles marked with who is speaking. On a monologue it's unnecessary and can add errors. And don't re-encode phone audio: send telephony recordings in their native 8 kHz mono – re-encoding adds no quality, only file size.

Common mistakes and how to avoid them

Forgetting the X-Api-Key header or confusing it with a query parameter. The key is passed only in the header; never put it in the URL – it will end up in proxy logs and browser history. If a key leaks, reissue it in the dashboard: the old one stops working instantly.

Polling without a pause and sending the whole video. Hundreds of requests per second won't speed up processing, but they will hit the 429 limit fast – add a pause between polls or use a webhook. And an extra gigabyte of video only bloats the upload: the API takes audio, so convert before sending.

Ignoring the format and not checking the result. Make sure you send wav, mp3, ogg, or flac – convert exotic containers like m4a beforehand. And remember: automation is 95% of the work, not 100%; quickly scan names, abbreviations, and terms – ten minutes of proofreading save your reputation.

What's next

Automatic subtitles are a process you can easily put on autopilot. Once the pipeline is built, you process dozens of clips a day with the same effort as one: extract audio, recognize, export SRT/VTT – and the file goes to the player or the host. For video makers it means every video ships with subtitles at no extra effort; for developers it means subtitles become another product feature, not a manual operation.

For the full list of endpoints, limits, and request examples – see the documentation.

If speech recognition is new to you, start with the basic guide – How to transcribe Russian speech to text via API: a step-by-step guide.

Start free

The free plan includes 100 audio minutes a month – no card required. Turn your first video into subtitles today.

Turn my first video into subtitles see pricing
← All articles