A US-based voice dictation app that turns speech into text on the spot. It runs on macOS, Windows, and iOS, and works anywhere a text cursor can be placed — email, chat, or a code editor. Its defining feature is Avalon, an in-house speech recognition model that the company claims outperforms existing models including OpenAI’s Whisper. Because it reads context from what is on screen, it can correctly transcribe words that are hard to tell apart from sound alone, such as variable names, library names, and internal jargon. For developers, the same model is offered as the Avalon API, designed so that existing Whisper integrations can migrate with minimal changes.
Key Features
- High accuracy with the proprietary Avalon model: On AISpeak-10, a benchmark of AI and coding terminology, the company reports 97.4% accuracy, well above the 65.1% reported for Whisper Large v3. On the general-purpose OpenASR Leaderboard it records an average word error rate of 6.24%
- Screen context understanding: A Deep Context feature refers to the app in use and the text on screen to correct recognition results — leaning toward code-style notation in an editor and prose in an email. The feature is off by default and must be explicitly enabled
- Custom dictionary and writing style control: Register names, brands, and in-house terms to raise accuracy. You can also set writing style instructions such as formal or casual, so spoken input comes out as polished text
- 49 languages with automatic detection: Supports 49 languages including Japanese and detects the spoken language automatically. According to the official FAQ, more than half of daily users dictate in a language other than English
- Privacy Mode and SOC 2 Type II: Transcripts are retained by default for quality improvement, but Privacy Mode disables retention. The Business plan offers Zero Data Retention. The service is SOC 2 Type II certified
- Avalon API for developers: The same model can be called externally. For OpenAI SDK users, swapping the base URL and model name is enough to migrate, keeping the existing request structure and authentication intact
Pricing
| Plan | Price | What’s included |
|---|---|---|
| Starter | $0 | 1,000 words total to try for free. No credit card required |
| Pro | $8/month billed annually / $10/month monthly | Unlimited words, custom instructions, expanded dictionary. One account across all devices |
| Max | $24/month | Everything in Pro plus Realtime Mode, voice commands, and early access to new features |
| Team | $12/user/month (2–9 people) | Centralized billing, organization-wide settings, enforced Privacy Mode |
| Business | Contact sales (10+ people) | SSO / SAML, advanced reporting, Zero Data Retention |
| Avalon API | $0.39 per hour of audio | Per-second billing (10-second minimum). No seat fees |
A 70% student discount is available for both Pro ($3/month) and Max ($9/month).
Pricing is as of August 2026. Please check the official site for the latest pricing.
Pros & Cons
✅ Pros
- Strong on jargon, proper nouns, and code-related vocabulary, so technical dictation needs little cleanup
- Unlike built-in OS dictation, it delivers the same accuracy and settings system-wide, regardless of the app
- Style settings turn casual spoken input into readable prose
- A 1,000-word free tier lets you test it without registering a card
- The same model is available as an API, making it easy to move from personal use to your own tooling
⚠️ Cons
- Cloud processing means an internet connection is required; it cannot be used offline
- The free tier is a small 1,000 words total, so daily use effectively requires a paid plan early on
- Audio is sent to the cloud, and transcripts are used for quality improvement by default, so Privacy Mode should be configured for sensitive content
- The accuracy benchmarks are figures published by the vendor itself; results in Japanese are worth verifying separately
- Realtime Mode and voice commands are limited to the higher-tier Max plan
Comparison with Similar Services
| Criteria | Aqua Voice | Wispr Flow | superwhisper | Built-in OS dictation |
|---|---|---|---|---|
| Recognition model | Proprietary Avalon | Proprietary speech model | Choice of Whisper-family models | OS-provided model |
| Supported OS | macOS / Windows / iOS | macOS / Windows / iOS | Mainly macOS | Each OS |
| Offline use | No (cloud processing) | No | Yes (with local models) | Sometimes |
| Screen context | Yes (Deep Context) | Yes | Limited | No |
| API offering | Yes (Avalon API) | No | No | No |
| Pricing | Free tier + from $8/month | Free tier + paid plans | Free tier + paid plans | Free |
Who Is It For
- Engineers and technical writers who frequently dictate technical terms, product names, and code fragments
- Anyone who wants to draft long emails, meeting notes, or articles faster than they can type
- People looking to reduce strain on hands and shoulders, or for whom typing is the bottleneck
- Those who move between Japanese and English while dictating and want to skip manual language switching
- Developers who want accurate transcription in their own apps and are looking for an alternative to Whisper
Summary
Aqua Voice tackles head-on the classic weakness of voice input — poor handling of specialized vocabulary — through its Avalon model and screen context understanding. The 1,000-word free tier is small, so serious use assumes a paid plan, but at $8 per month it pays off quickly for anyone who writes long-form text daily. Two points are worth checking in advance: cloud processing is mandatory, and transcripts are retained by default. Start with the free tier to see how it handles Japanese and your own speaking style, then decide whether it holds up for Pro-level use.