VoiceOS is a voice operating system that lets you control your computer just by talking to it. It has two modes built for different jobs: a “dictation mode” that replaces typing, and an “agent mode” that completes tasks across apps, such as sending an email or adding a calendar event. Voice input tools and task-running agents are usually sold as separate products; VoiceOS sits between the two and puts both inside a single app.
Key Features
- Dictation mode: Instead of typing out what you literally said, it produces what you meant. Punctuation, self-corrections, and the formatting of lists and numbers are all handled automatically
- Agent mode: Sends emails, adds calendar events, logs tasks, and more across apps using nothing but voice commands. It connects to external services such as Gmail, Google Calendar, Slack, and Notion
- Screen-context awareness: Point at something on screen while you talk, and it responds based on what is currently displayed. Instructions like “summarize this” work as-is
- Custom MCP integrations: Build and connect your own MCP servers using the Python or TypeScript SDK. The official guide uses HomeKit, Spotify, and macOS system operations as examples
- No-recording policy: The company states explicitly that voice data is neither stored on its servers nor used to train models. That is a meaningful factor when work conversations are involved
Pricing
| Plan | Monthly Price | Key Features |
|---|---|---|
| Free | Free | 100 dictation sessions per week, 25 agent-mode sessions per week |
| Pro | $11.99 (billed annually) | Unlimited access to all features |
| Enterprise | Contact sales | Everything in Pro, plus security requirements such as SOC 2 Type II and SSO / SAML |
Pricing is current as of August 2026. New sign-ups get Pro-equivalent access for the first 7 days. Check the official pricing page for the latest details.
Pros and Cons
✅ Pros
- No switching tools between writing and executing. “Write this text” and “send this by email” can happen back to back inside the same app
- Because you can point at the screen while speaking, demonstratives like “this” get through as-is. There is no need to re-describe the target in text
- If you write your own MCP server, you can bring voice control to integrations no off-the-shelf product offers, such as internal company tools
⚠️ Cons
- In places where you cannot speak out loud, most of its value disappears. Shared offices, cafes, and homes with other people around limit when you can use it
- Because agent mode works across apps, cleaning up after a mid-task failure is manual. You have to check for yourself how far it got
- The free tier is counted in “N sessions per week,” but the length of a single session is not defined. That makes it hard to estimate in advance whether it will cover your usage
Comparison with Similar Services
| Criteria | VoiceOS | Wispr Flow | Superwhisper |
|---|---|---|---|
| Processing | Cloud | Cloud | On-device (works offline) |
| Supported OS | Mac / Windows | Mac / Windows / mobile | Mac / Windows / iOS |
| Scope | Dictation + cross-app task execution | Dictation and context-aware formatting | Dictation with local processing |
If you compare on voice input alone, there are plenty of options. VoiceOS stakes out its position by covering cross-app task execution on top of dictation. If you want everything processed fully offline, an on-device tool like Superwhisper is the candidate to consider.
Who Is It For?
- People who use both Mac and Windows and want the same voice controls on each
- People who want to go beyond typing less and finish emails and calendar entries entirely by voice
- Developers who want to build custom MCP integrations and bring their own work-specific tools under voice control
Summary
The decision comes down to how much you are willing to pay for agent mode. If transcription is all you need, a dedicated tool is usually cheaper. If you want to finish emails and calendar entries by voice too, the price difference earns its keep. It is also worth checking, before you commit, whether you can reliably count on an environment where speaking out loud is an option.