In text, a short delay is invisible. In a phone call, even a one-second pause feels unnatural and erodes trust in the system.
Getting voice AI to feel conversational requires optimising the entire pipeline — speech recognition, reasoning, and speech synthesis — to work in near real time, often streaming partial responses before the full answer is ready.
This is one of the more active areas of our internal R&D work, because the experience bar for voice is simply higher than for text.