← Blog

Choosing an API

Speech-to-Text API: how to choose one without overpaying

Choosing a Speech-to-Text API is almost never about who is more accurate. One service produces a perfect transcript on a clean studio recording and falls apart on a phone call, another is five times cheaper but charges for every minute of an open connection, and a third transcribes beautifully yet cannot return word-level timestamps. The final cost and quality of your product depend on five or six parameters, and almost all of them are visible only in the documentation and in honest measurements. In this breakdown I will show how to compare STT APIs on substance: which metrics actually predict the result, where the hidden costs live, and how to test accuracy on your own data in a single evening.

Why the 98% accuracy claim means nothing

Almost every vendor puts up to 98% accuracy on its landing page. The number is pretty, but it has one flaw: it says almost nothing about your recordings. Accuracy is measured on a specific dataset - read speech, podcasts, or phone calls - and a model that wins on studio audio can fall apart on a conversation through a headset. The domain matters more than the brand: a call center needs narrowband 8 kHz audio and office noise handled, media needs punctuation and names, and medicine needs terminology and dosages.

The second trap is the conditions of the test. Up to 98% usually means clean audio, a single speaker, a familiar accent, and punctuation that is not counted as an error. Add a second speaker, noise, or a rare surname, and the number drops by ten to twenty percent. Treat the vendor promise as an upper bound, not as the average result you will get on your own product.

The practical conclusion is simple: choose an API by testing it on your own data, not by a marketing number. Such a test takes one evening, costs next to nothing, and on a free tier costs nothing at all, and it removes the main risk: buying a subscription for the task your engine handles worst. Below is a concrete way to do it and which parameters to compare first.

There is a subtler point too: accuracy is almost always measured on words, while a business feels quality in entities. If you need an order number, an amount, and a date out of a call, then one digit error costs more than ten errors in conjunctions and prepositions. So before measuring, define what counts as an error for you: for subtitles it is readability, for analytics it is entity correctness, for search it is hitting the key words. Then build both the metric and the test set around that criterion, not the other way round.

The metric you can trust: WER

The standard quality metric for speech recognition is WER, word error rate, the share of wrong words. It is computed by adding substitutions, insertions, and deletions and dividing by the number of words in the reference text. A WER of 0.10 means every tenth word was recognized incorrectly, and accuracy equals 1 - WER. Unlike up to 98%, WER has a definition, a formula, and reproducibility: two people measuring the same data get the same number.

The metric has nuances worth knowing before you compare. Case and punctuation are usually normalized, otherwise a comma becomes a separate error. Names, brands, and terms hurt WER the most: one rare name in a ten-word sentence produces a noticeable error share. Numbers and abbreviations are written sometimes as words and sometimes as digits, and that needs normalizing too, or the metric punishes the engine for formatting rather than for hearing. That is why comparing WER from different reports is almost meaningless: only measurements on the same reference set count.

The good news is that you can measure WER without leaving the API. The POST /v1/eval endpoint takes audio and a reference text, recognizes the recording, and returns a word-level breakdown: hits, substitutions, insertions, deletions, plus the WER and accuracy themselves. Here is a minimal Python example that is handy for running a dozen of your own clips.

Python
import requests

API = "https://ttsapi.ru/v1"
HEADERS = {"X-Api-Key": "rtt_…"}

reference_text = open("reference.txt", encoding="utf-8").read()

with open("sample.wav", "rb") as f:
    r = requests.post(
        f"{API}/eval",
        headers=HEADERS,
        files={"audio": ("sample.wav", f, "audio/wav")},
        data={
            "reference": reference_text,
            "language": "ru",
            "normalize": "true",
        },
    )

report = r.json()
print(report["wer"], report["hits"], report["substitutions"])

The full list of endpoint parameters, responses, and errors is here - see the documentation.

Russian speech: where models break most often

Russian has rich morphology, and that creates characteristic errors. Cases and endings are often heard correctly but written with the wrong grammar; homophones drift apart in meaning; a surname from a phone call turns into a random string of letters. Mixed speech is a separate story, when a single sentence contains both an English tech term and everyday Russian. The engine must not translate a technical term into an everyday word, and it must not lose it either.

The worst case for any STT engine is a narrowband phone channel. The 8 kHz band cuts off the high frequencies that distinguish sibilants, and office noise adds its own share. If your product is call quality control or support transcription, test exactly this scenario: audio from a phone, two speakers, overlapping turns, sometimes codec compression. Meetings and podcasts are easier, but echo and simultaneous speech still get in the way.

This is where keyterm prompting helps. You can pass a list of key words - names, terms, abbreviations - and the engine is prompted to look for exactly them. That is the cheapest way to raise accuracy on a domain without changing the model: instead of retraining you simply tell the service which words are definitely in the recording. For a call center these are plan names, for medicine they are drug names, for interviews they are speaker names.

Also check what the service returns: plain text only, or word-level timestamps with confidence scores as well. For Russian this is a lifesaver on rare words: you can see where the model hesitated and fix it precisely instead of re-listening to the whole file. If you produce subtitles, word-level timestamps are essential for even lines that do not flicker word by word on screen.

Streaming or deferred: what you actually pay for

Streaming recognition returns text as audio arrives. It is needed where waiting is not an option: live subtitles, a voice bot that must understand an utterance before it answers, a meeting with real-time captions. Latency is measured in seconds or less, and that is exactly what explains the higher price of streaming: the server keeps the connection open and computes results on the fly.

The deferred path works differently: the file is uploaded in full, processing happens in the background, and the result is fetched by job id or delivered by webhook. It is cheaper and scales more easily: you can queue hundreds of recordings without holding a single connection. For content, reports, and archives this is the optimal scheme - you pay for audio minutes, not for waiting time.

The third option is a compromise, and it is worth planning for early. Short files are faster to process synchronously, while long ones belong in asynchronous jobs with progress reporting. That is how our API works: POST /v1/transcribe/sync for short recordings, POST /v1/transcribe for everything else, and a long_form flag for material up to four hours with chunk-level progress. A single API covers both modes, and you can switch between them by file length.

The modes differ in price, and that is worth keeping in mind while designing. Asynchronous processing is usually cheaper than synchronous, and streaming is more expensive than both because it pays for low latency. So a sensible architecture looks like this: short interactive requests go through the synchronous mode, bulk archives go through the asynchronous one, and streaming turns on only where a user genuinely waits for text on screen. That is how you save money without hurting the product.

bash
# short files: synchronous request
curl -X POST https://ttsapi.ru/v1/transcribe/sync \
  -H "X-Api-Key: rtt_…" \
  -F "audio=clip.wav" \
  -F "language=ru"

# long files: asynchronous job with a webhook
curl -X POST https://ttsapi.ru/v1/transcribe \
  -H "X-Api-Key: rtt_…" \
  -F "audio=meeting.mp3" \
  -F "language=ru" \
  -F "long_form=true" \
  -F "webhookUrl=https://example.com/hooks/stt"

Price per minute: do the unit economics, not the price list

The price on a vendor website rarely matches the price in your budget. Start with the base rate: Yandex SpeechKit charges 0.65 rubles per minute for synchronous and streaming recognition and 0.61 for deferred, AssemblyAI is around 0.28 rubles per minute for asynchronous processing, OpenAI roughly 0.48, and ElevenLabs Scribe about 0.53. That is nearly a two-fold spread, and it is before paid add-ons.

Then the arithmetic of details begins. Diarization, PII redaction, keywords, and summaries are paid add-ons at most vendors, stacked on top of minutes, and they can easily double the bill. The minimum billing increment matters too: if billing counts 15-second blocks, ten-second utterances in a chat cost as much as full ones. Some price lists charge streaming minutes at a separate rate or open them only inside a subscription.

There is also a factor specific to teams in Russia: currency and payment method. A foreign service priced in dollars means a card that does not always go through, an exchange rate, and taxes on the buyer side; OpenAI is not even available for payment from Russia. A ruble price with card or SBP payment removes that layer of risk. So count the cost of processing an hour of real calls, not the price per minute: add-ons, minimum increments, and conversion included.

Plans, included volumes, and per-minute rates are here - see pricing.

Limits, formats, and file length: where an API hits a wall

Technical limits either fit your product or break it after the fact, when rework is expensive. The first thing to check is the file size and duration cap. For standard transcription it is usually 25 MB and 15 minutes; a one-hour meeting or a call center shift does not fit. A long-form path lifts the limit to 512 MB and four hours per recording, but it is usually available only on higher plans.

The second block is audio formats and parameters. The practical minimum a service must accept: wav, mp3, ogg, and flac, mono and stereo, and various sample rates, because files come from very different sources. It helps if stereo channels can be transcribed separately: that is a simple way to split participants of a phone call without diarization and without extra cost.

The third block is how you receive the result. Webhooks matter for integration: the service notifies you when a job is ready, and you do not have to poll the status in a loop. Also check rate limits: they are rarely mentioned on landing pages, yet they are exactly what triggers a 429 at the worst possible moment. Our limit grows with the plan, from 5 requests per second on Free to 60 on Business, with an extra burst allowance.

There is one more layer that is easy to miss: reliability. Providers behave differently under load: some queue a request and answer later, others return an error immediately. For products with peak traffic, predictable behavior at rush hour matters more than average speed: a clear error code, a header with a retry hint, and no hidden timeouts. Test that on your own traffic before release, not after the first incident.

JSON
{
  "job_id": "tr_9f31c2",
  "status": "processing",
  "progress": 42,
  "chunks_completed": 8,
  "chunks_total": 19
}

Checklist: ten questions before you pay

To avoid drowning in details, reduce the choice to ten questions. One: which quality metric was measured and on which domain. Two: is there a way to measure WER on your own data straight from the API. Three: how does the service handle a phone channel and overlapping speech. Four: are keywords and term hints supported. Five: are word-level timestamps and confidence scores available, since without them you cannot build good subtitles, search across recordings, or comfortable proofreading.

Six: how much a minute costs with add-ons and the minimum billing increment included. Seven: what the size and duration limits are and on which plans long processing unlocks. Eight: is there a streaming mode and what does it cost. Nine: how the result is delivered, webhook or polling, and what the rate limits are. Ten: can you run your own test before buying rather than after. These ten points cover nearly every reason an STT integration turns out more expensive or harder than its landing page suggested.

If a vendor cannot answer at least half of these questions in the documentation, that is a bad sign: either the parameters are inconvenient and hidden, or the team never thought about real scenarios. Good docs answer these questions with numbers and code samples, not with generic words about cutting-edge neural networks. You can check this in fifteen minutes: find the pages about timestamps, webhooks, limits, and add-on pricing. If even one is missing, price that risk into your decision.

It helps to turn the finished checklist into a table and fill it in for every candidate. Where a vendor is silent, assume the worst: no webhooks means polling, no word-level timestamps means subtitles will need manual fixes. Estimate the total cost of ownership for a quarter, including engineering time for integration: sometimes a cheap API needs three times more code, and the savings on minutes are eaten by development.

Bottom line: choose in one evening

Build a test set of 20-50 clips from your own recordings: different quality, with names and terms, with phone and studio audio. Run it through the candidates, compute WER, and compute the cost of processing an hour. Add the add-ons you actually need, then check the length limits, the streaming mode, and the request rate. Usually one or two options remain, and the choice between them comes down to price and integration convenience.

Our path is as short as it gets. The free plan includes 150 audio minutes a month, which is enough to run a test set and see WER on your own data without a card. Then Basic at 390 rubles a month gives 1,800 audio minutes and diarization, Pro at 1,490 adds streaming recognition, long recordings up to four hours, and translation, and Business at 7,990 brings the priority queue and 30,000 minutes. If you do not want a subscription, pay-as-you-go works from your balance: 0.40 rubles per minute for asynchronous transcription and 0.50 for synchronous.

And most importantly, WER is not a one-off check but a metric worth monitoring. The domain changes, new names and terms appear, and recording quality shifts. A regular measurement on a reference set shows that the engine still meets your expectations, and that is far cheaper than discovering degradation from user complaints a month after release.

If speech recognition is new to you, start with the basics - How to transcribe Russian speech to text via API: a step-by-step guide.

Start with a free measurement

The free plan includes 150 audio minutes a month - enough to run your own test set and compare WER on your recordings. No card required.

Test accuracy on my own files see pricing
← All articles