On Live Sales Calls, AI Gets 200 Milliseconds to Sound Human
Human turn-taking runs on a reflex clocked in milliseconds. A voice AI's real test isn't whether it can talk but whether it can take a turn on human time.

If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.

Two people having a conversation hand the floor back and forth inside a window most of them will never notice and neither of them can control. Across ten languages, from major world tongues to small indigenous communities, researchers measuring real dialogue found the gap between one speaker finishing and the next beginning clusters in a tight window between zero and 200 milliseconds, with a cross-language median right around 100. Come in much slower and the delay reads as hesitation, evasion, or a machine, and the listener feels it before they can name it.
The real test of a voice AI goes well past whether it can talk. Plenty of systems produce fluent, pleasant speech. The harder question is whether it can take a turn on human time, because the moment it misses that window the illusion of a conversation collapses into something closer to a phone tree, no matter how right the words are.
The lazy assumption treats voice as a solved problem, an API call. Pick a text-to-speech voice that sounds warm, wire it to a good model, and ship. Teams that build on that assumption end up with the voice agents that feel like leaving a voicemail and waiting for the callback. They answer correctly and a beat too late, and the beat is the part the buyer remembers.
The engineering under the voice
ElevenLabs, one of the most respected names in real-time voice, has published the anatomy of that problem, and it's a useful map precisely because it comes from a specialist rather than a pitch deck. A natural-feeling voice agent is a latency budget spread across four stages, each one a slice of the same short window: speech recognition, turn detection, the language model, and speech synthesis. The company says its own custom speech recognition adds under 100 milliseconds where the open-source standard, Whisper, adds more than 300. It reports its fastest synthesis engine runs a 75-millisecond model time and returns its first byte of audio in 135. It pegs the model step itself, the part most builders treat as the whole job, anywhere from under 350 milliseconds for a small fast model like Gemini Flash to between 700 and 1,000 for heavier ones such as GPT-4 variants and Claude. Its own conclusion is that applications should target under a second, total.
Run those four stages end to end without care and you've spent the human window before the model has finished thinking. The engineering that decides the outcome is the unglamorous kind. The stages overlap instead of running them in sequence, streaming the answer out as the first words are ready rather than holding it until the last, and detecting that the buyer has stopped talking early enough to come back the way a person would. This real work never makes it into a demo script, and it decides whether the buyer relaxes.
The test goes live
This is the bar 1mind is built to clear. Most of what gets called an AI go-to-market agent is text, a chat widget with a stronger model behind it. 1mind leads instead with a photorealistic teammate that speaks, and its Ride-Along product puts that teammate onto live Zoom, Teams, and Meet calls as a named participant that answers the buyer out loud, in real time, in a studio-quality voice. A chat window can hide a slow response inside a typing indicator. A live call has no such cover, and the buyer hears every millisecond of it.
1mind is held to the bar by the same parts everyone else builds on. It runs on a mix of frontier models, OpenAI and Google Gemini among them, which means the model step in its pipeline carries the same latency the labs' own figures describe, and the job of hitting human timing falls to the systems work wrapped around it. That layer is the one Sachin Bhat, the company's Chief Technology Officer, spent years on before this, running data and AI infrastructure at Rippling. It's the real-time plumbing where turn-taking becomes an engineering discipline in its own right.
Inside the buyer's feelings
The payoff reaches the buyer as a feeling, although it never shows up on a feature list. A voice that answers inside the human window feels unhurried and present, like someone paying attention. A voice that misses it feels like being processed, and the buyer starts performing patience instead of asking the questions they actually came with. The autonomous experience wins that person only when it's fast enough to feel human, which is why the milliseconds carry the whole weight of the product.
The buyer will never see the pipeline, and won't care that there is one. But they'll feel the 200 milliseconds, and somewhere inside that gap they'll decide whether they're talking to someone or waiting on something.
Read what the room is reading.
New pieces for growth leaders, delivered as they publish. Unsubscribe whenever.








