Skip to content

AI model category

Speech AI models

For transcription, call notes, voice assistants and spoken replies.

  • 4 models

Speech models come in two directions. Speech-to-text turns calls, meetings and voice notes into text that other models can summarize and act on. Text-to-speech turns replies into natural spoken audio for phone assistants, audio content and accessibility.

For live voice assistants, speed matters as much as accuracy, so streaming support and response time are key checks.

Speech

All speech ai models

Speech

Deepgram Nova-3 and Flux

Deepgram

Deepgram's speech-to-text models: Nova-3 for general transcription in 60+ languages, and Flux for voice agents with built-in end-of-turn detection.

  • Deepgram
Speech

ElevenLabs Eleven v3

ElevenLabs

ElevenLabs text-to-speech: the expressive Eleven v3 in 70+ languages, a real-time conversational version, and the low-latency Flash models, with voice cloning.

  • ElevenLabs
Speech

GPT Transcribe

OpenAI

OpenAI's current speech-to-text model, replacing Whisper and the GPT-4o transcribe models, with streaming and keyword or language hints.

  • OpenAI
Speech

GPT-4o mini TTS

OpenAI

OpenAI's current text-to-speech model for turning text into natural spoken audio, at a low price level.

  • OpenAI

How to choose

How to choose in this category

Test with real recordings from your own calls, including accents, background noise and industry terms. For voice output, check the license terms for commercial use and any rules about cloned voices. Always tell callers when they are speaking with an AI.

Keep exploring

Start a project

Tell us what you want to build. We will show you a faster path.

Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.

Your privacy choices

We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.