Humans notice silence faster on the phone
A chat interface can hide a second of thinking behind a typing indicator. A phone conversation cannot. When a caller finishes a sentence and hears empty air, the silence feels like confusion, a dropped connection, or a system that did not understand. Latency therefore changes the tone of the entire interaction even if the final answer is correct.
Delay is cumulative
A voice agent does not have one latency number. It has a chain: detecting the end of speech, transcribing the audio, sending context to the language model, receiving the first useful tokens, generating audio, and getting those samples back onto the call. Telephony transport adds another layer. Improving only one component rarely fixes the experience if the rest of the pipeline remains sequential and slow.
Cally treats latency as an architecture problem
The Cally feature sheet describes a sub-600 millisecond end-to-end turnaround across SIP streaming, VAD, STT, language-model inference, and text-to-speech. It uses on-machine FireRedVAD with an RMS fallback, real-time STT providers, and a multi-model race mode that can send inference to multiple providers and use the fastest viable first response. Automatic fallback also reduces the risk that one provider timeout freezes the call.
First response speed is not the only goal
An agent that speaks quickly but ignores interruptions still feels artificial. Low-latency systems also need to stop speaking when the caller cuts in, understand what was actually heard, and continue from the right conversational state. Cally tracks playback timing at character or word level so that when the caller interrupts, buffered audio can be cleared and only the words that were actually played remain in the history.
Measure latency by conversation, not benchmark
A synthetic benchmark can tell you how fast an isolated model returns a token. Operations teams should measure what the caller experiences: time from the end of their speech to the first audible response, interruption recovery, long-answer behavior, network variance, and fallback behavior. The best latency target is not a leaderboard number; it is a call that feels responsive under real conditions.
When evaluating a voice platform, ask for the full latency path—not only the LLM benchmark. The caller experiences the entire pipeline.