Voice AI / Text-to-Speech
Also written as Text-to-Speech, TTS, Voice Agent, Speech Synthesis, Voice Cloning
Generating natural-sounding speech from text, and combining it with speech recognition and an LLM to build agents that hold spoken conversations, for example in call centres. ElevenLabs and OpenAI's realtime voice models are well-known examples.
Think of it like
A narrator who can read any script aloud instantly, and in a voice agent, also listen and answer back.
Junior or senior?
Voice agents are judged on delay and on handling interruptions, not only on how the voice sounds.
Senior sounds like
Talks about end-to-end response time and what happened when a caller talked over the agent.
Ask them
“How long did your voice agent take to respond, and how did it handle a caller interrupting it?”