CoeFont is a voice AI platform that brings text-to-speech (TTS), voice conversion, and real-time speech interpretation together in one product. When it launched in April 2021 it was known as a read-aloud service built on the idea of choosing a voice the way you choose a font. Since then, its meeting interpretation feature, CoeFont Interpreter, has grown into the flagship offering, connecting with meeting tools such as Zoom, Microsoft Teams, and Google Meet to support multilingual conversations. It is developed by CoeFont Co., Ltd. in Tokyo, a startup that originated at Institute of Science Tokyo (formerly Tokyo Institute of Technology).
Key Features
- Text-to-speech with a wide range of voices: Enter text and get natural-sounding speech. You can pick a voice and adjust reading, speed, and intonation for video narration or audio content
- Voice conversion with your own voice: Create a personal AI voice from a browser recording. The system learns characteristics such as pitch, pronunciation, intonation, and speaking speed, so you can also speak another language while keeping your own voice
- Real-time interpretation (CoeFont Interpreter): Speech is converted into another language with almost no pause. The company cites a delay of around one second, comparable to a human interpreter, so conversations keep their rhythm. Supported languages include Japanese, English, Chinese, Korean, French, Spanish, German, Russian, Vietnamese, Thai, Indonesian, and Portuguese ─ more than ten in total
- Works with major meeting tools: The desktop and web app supports Zoom, Microsoft Teams, Google Meet, Webex, and Discord. The iOS and Android apps also cover face-to-face conversations
- Terminology dictionaries and automatic minutes: Register industry or in-house terms to reduce mistranslation, and generate transcripts and summaries of meetings automatically
- Security aimed at business use: SOC 2 Type 2 certified and GDPR compliant. On higher plans, excluding your data from AI training is the default setting
Pricing
| Plan | Monthly price | What’s included |
|---|---|---|
| Free | $0 | 20 min/month of interpretation, 800 characters/month of TTS, 1 project, custom voice from a browser recording |
| Standard | $20 | 5 hours/month of interpretation, 80,000 characters/month of TTS, unlimited projects |
| Plus | $350 | 8 hours/month of interpretation, 1,000,000 characters/month of TTS, TTS API, up to 5 users, data excluded from AI training by default |
| Enterprise | Contact sales | Scalable interpretation and TTS, advanced features such as custom dictionaries and speaker identification, TTS API, unlimited users, SSO |
Pricing is current as of August 2026. Check the official site for the latest information.
Pros & Cons
✅ Pros
- Text-to-speech, voice conversion, and interpretation all live in one account, so there is no need to juggle separate services
- Interpretation latency is low, letting multilingual meetings keep their conversational pace
- Broad support for major meeting tools makes it easy to adopt without changing existing workflows
- Terminology dictionaries improve accuracy for industry-specific and in-house wording
- SOC 2 Type 2 and GDPR compliance help it clear corporate security reviews
- The free plan lets you try both interpretation and text-to-speech
⚠️ Cons
- The free tier is small ─ 20 minutes of interpretation and 800 characters of TTS ─ so regular work use effectively requires a paid plan
- The gap between Standard and Plus is large, so costs jump sharply once you need API access or multiple users
- Interpretation is metered by time, so organizations with frequent long meetings need to manage the limit
- Whether synthesized speech may be used commercially depends on the conditions attached to each voice, so review the terms before using it in published content
Comparison with Similar Services
| Criteria | CoeFont | ElevenLabs | Nijivoice | DeepL Voice |
|---|---|---|---|---|
| Provider | CoeFont Co., Ltd. (Japan) | ElevenLabs (US) | Algomatic (Japan) | DeepL (Germany) |
| Main use | TTS + voice conversion + interpretation | High-quality speech synthesis and voice cloning | Japanese speech synthesis strong on emotion | Real-time meeting translation |
| Real-time interpretation | Yes (meeting tool integration) | Yes (offered as a separate feature) | No | Yes |
| Voice output | Yes | Yes | Yes | Mainly text |
| Japanese support | Native Japanese | Supported | Japanese-focused | Supported |
| Free plan | Yes | Yes | Yes | Limited |
Who Is It For
- Business users with frequent meetings involving overseas partners or teammates who want discussions to flow without a human interpreter
- Creators who need narration for video or podcasts without a recording session
- Speakers and instructors who want to present in a foreign language while keeping their own voice
- Teams working with heavy specialist terminology where general translation tools fall short on accuracy
- Companies for which SOC 2 or GDPR compliance is a condition of adoption
Summary
CoeFont started as a text-to-speech service and now centers on real-time interpretation. Covering read-aloud, voice conversion, and interpretation from a single account, along with broad support for major meeting tools, is what makes it practical day to day. Start with the free plan to check interpretation latency and voice quality, then consider Standard or higher based on how often you meet and how much text you need to convert.