A speech AI platform from Soniox, Inc. It unifies speech recognition (STT), speech synthesis (TTS), and speech translation behind a single API, with support for more than 60 languages. What sets it apart is the ability to transcribe in real time at under 200 milliseconds of latency while automatically detecting language switches mid-utterance (code-switching). It also handles speaker diarization, timestamps, and voice cloning from short audio samples, and holds certifications including SOC 2 and ISO/IEC 27001. This is a developer-facing service, built to be embedded into voice agents, captioning, call center recording, and similar use cases.
Key Features
- Real-time multilingual speech recognition: Supports 60+ languages, and detects language switches automatically even when they happen mid-utterance
- Low-latency streaming: Real-time transcription with under 200 milliseconds of latency ─ suited to conversations, calls, and anything else that needs an immediate response
- Speaker diarization and timestamps: Per-speaker separation and timestamping come standard
- Speech translation: Speech translation across 3,600+ language pairs, integrable directly into the STT stream
- Voice cloning (TTS): Synthesize speech that reproduces a voice from a short audio sample
- Compliance: SOC 2 Type 2, ISO/IEC 27001:2022, HIPAA, and GDPR support, which makes enterprise adoption easier to justify
Pricing
There are no fixed monthly or annual plans ─ everything is usage-based. Core capabilities like speaker diarization, language detection, and translation are included in the usage rate at no extra charge.
| Item | Approximate cost |
|---|---|
| Speech-to-Text (async, file upload) | ~$0.10/hour |
| Speech-to-Text (real-time, streaming) | ~$0.12/hour |
| Text-to-Speech (synthesis, voice cloning) | ~$0.70/hour equivalent (input text $4.00 per million tokens, output audio $21.50 per million tokens) |
Pricing is current as of September 2026. Check the official site for the latest rates.
Pros and Cons
✅ Pros
- Handles 60+ languages, down to detecting language switches mid-utterance
- Real-time processing at under 200 milliseconds of latency
- STT, TTS, and translation live behind one API, so the integration cost of stitching multiple services together stays low
- Already certified for SOC 2, ISO/IEC 27001, HIPAA, and GDPR ─ the boxes enterprises need checked
⚠️ Cons
- Usage-based pricing only, with no fixed plans, makes monthly cost hard to forecast
- API-first rather than a GUI app, so there’s a barrier to entry for non-developers
- Little localized official material in Japanese
Comparison with Similar Services
| Criteria | Soniox | Deepgram | AssemblyAI | OpenAI Whisper |
|---|---|---|---|---|
| Provider | Soniox, Inc. | Deepgram | AssemblyAI | OpenAI |
| Primary use | Unified API for multilingual STT, TTS, and speech translation | Real-time STT API | Speech API platform focused on English-speaking markets | Speech recognition model (OSS) |
| Environment | API | API | API | Self-hosted (OSS) |
| Real-time support | Yes (under 200ms) | Yes | Yes | No |
| Distinguishing trait | Strong multilingual accuracy and code-switching detection | Low-cost, English-centric real-time API | Fewer supported languages than Soniox | Open source and highly flexible, but not real-time |
Who It’s For
- Developers automating call transcription and summarization for multilingual call centers or customer support
- Anyone looking for a foundation for real-time captioning or automatic meeting minutes
- Teams that want to wire the audio input/output layer of a voice agent (voice bot) together through an API
- Anyone streamlining speech translation or dubbing production for multilingual content
Summary
Soniox packages speech recognition, speech synthesis, and speech translation into one API, with real-time multilingual processing as its center of gravity. Code-switching detection across 60+ languages and sub-200-millisecond latency are strengths that hold up in production settings like call centers, meeting records, and voice agents. Pricing, on the other hand, is usage-based with no fixed plans, so it’s worth estimating cost up front from your expected call volume and processing hours. And since this is an API-first service with no GUI, adoption assumes you have development resources on hand.