Accurate transcription for real conversations

Stream low-latency transcripts during live calls or transcribe recordings in bulk — across 23 languages, accents, and noisy telephony.

transcribe.stream
Live
Caller0:02–0:06

Hello, I'd like to check my pending fee for this month.

Telugu → text~100msword timestamps
Live demo

See it in action

Click to start transcription

English (India)

Uses your microphone — not the sample video or page audio.

Transcription output

Your transcription will appear here in real-time…
Download
Capabilities

What you get

Every capability, shown — not just described.

98%+accuracy

High accuracy

Transcription tuned for real phone audio — accents, domain terms, and spontaneous speech.

Englishहिन्दीதமிழ்తెలుగుবাংলাಕನ್ನಡ+17

23 languages

English plus 23 Indic languages, with mid-conversation language switching.

I'd like to check my pending fee

Real-time streaming

Low-latency streaming transcripts for live agents, plus fast batch for recordings.

Noise robust

Handles background noise, cross-talk, and low-bitrate telephony without falling apart.

Agent
Caller
Agent

Speaker diarization

Know who said what — agent vs caller — with word-level timing and confidence.

₹/minusage-based

Cost-effective

Usage-based pricing that scales from a handful of calls to millions of minutes.

Engines

Pick the right transcription engine

Batch accuracy at scale, or real-time transcripts for live calls.

Scribe v1 / v2

98%+ Accuracy

Industry-leading transcription powered by Scribe. Supports 90+ languages with keyterm prompting and dynamic audio tagging — built for batch transcription at scale.

  • 98%+ transcription accuracy
  • 90+ languages
  • Keyterm prompting
  • Dynamic audio tagging

Scribe v2 Realtime

Realtime

Real-time transcription with ~150ms latency and word-level timestamps. Built for live voice calls, meeting assistants, and streaming audio pipelines across 90+ languages.

  • ~150ms latency
  • 90+ languages
  • Word-level timestamps
  • Live streaming audio
Models

Choose your model

Optimized models for every use case

Pollax v1

Standard

Reliable speech recognition model optimized for general-purpose transcription

  • High accuracy
  • Multi-lingual support
  • Noise robust
  • Fast processing

Pollax v2

Advanced

Next-generation model with enhanced accuracy and advanced features

  • Best-in-class accuracy
  • Real-time streaming
  • Speaker diarization
  • Custom vocabulary
API capabilities

Every transcription control, explained

Click a capability to see how it behaves and the request shape that turns it on.

Punctuation

punctuate: true

Automatic punctuation, capitalization, and readable paragraphing — no post-processing required.

Live example

RawPunctuated

id like to check my pending fee please can you send the receipt

Request
{
  "audio_url": "https://cdn.example/call.wav",
  "language": "en-IN",
  "punctuate": true,
  "smart_format": true
}
Response
{
  "text": "I'd like to check my pending fee, please.",
  "words": [
    { "word": "I'd", "start": 0.12, "end": 0.28 },
    { "word": "like", "start": 0.30, "end": 0.48 }
  ]
}
At a glance

Production-grade transcription

0%+
Accuracy
<0ms
Streaming latency
0
Indian languages
0.0%
Uptime
Languages

22 Indian languages, 36+ in total

Native pronunciation and scripts — with the ability to switch language mid-call based on the speaker.

Indian languages

Hindiहिन्दी
English (India)
Bengaliবাংলা
Teluguతెలుగు
Marathiमराठी
Tamilதமிழ்
Gujaratiગુજરાતી
Kannadaಕನ್ನಡ
Malayalamമലയാളം
Punjabiਪੰਜਾਬੀ
Odiaଓଡ଼ିଆ
Assameseঅসমীয়া
Urduاردو
Maithiliमैथिली
Konkaniकोंकणी
Nepaliनेपाली
Kashmiriکٲشُر
Sindhiسنڌي
Dogriडोगरी
Manipuriমৈতৈলোন্
Bodoबड़ो
Santaliᱥᱟᱱᱛᱟᱲᱤ
Sanskritसंस्कृतम्

International

English (US)
English (UK)
SpanishEspañol
FrenchFrançais
GermanDeutsch
Japanese日本語
Chinese中文
PortuguesePortuguês
Korean한국어
ItalianItaliano
Arabicالعربية
RussianРусский
FAQ

Questions, answered

Real-time or batch?

Both. Stream audio for live transcripts during a call, or submit recordings for fast, accurate batch transcription.

How does it handle accents and noise?

Models are trained on real telephony audio, so they hold up across accents, background noise, and compressed call quality.

Do I get timing and speakers?

Yes — word-level timestamps, confidence scores, and speaker diarization (agent vs caller).

Which languages are supported?

English and 23 Indic languages, with the ability to detect and switch languages mid-conversation.

Transcribe your first call

Get an API key and stream your first transcript in minutes.