Voice assistants: caching as a conversational necessity
Voice gives you about a second before silence feels broken. Model round-trips don't fit. Cache hits do, with room to spare for the speech stack.
Voice interfaces live under a brutal latency budget: ASR, understanding, synthesis, and playback all share roughly a second before users perceive the assistant as broken. A multi-second LLM round-trip blows the budget on its own. For repeated intents, which dominate voice traffic's command-like distribution, that spend of time and tokens is doubly absurd.
In plain words. Voice users wait one second, max. Models take longer than that. The cache is how repeated requests answer instantly enough to feel like conversation.
Crowkis hands voice stacks their latency budget back: semantic hits return in under a millisecond, leaving nearly the whole second for speech processing. 'What's on my calendar', 'play the news', 'how do I get downtown' and their endless phrasings become instant, while only genuinely novel requests wait on a model.
flowchart TD Q["incoming query"] --> I["intent classifier"] I --> T["template match"] T --> V["HNSW neighbours"] V --> C["confidence gate"] C --> TR["trust + freshness"] TR -- pass --> A["answer · <1ms"] I -. veto .-> M["(nil) → your model"] T -. veto .-> M V -. veto .-> M C -. veto .-> M TR -. veto .-> M style A fill:#fbe9e8,stroke:#d62221,stroke-width:2.5px style M fill:#f3eee5
Voice phrasing variability is the semantic layer's home turf: spoken language is messier than typed, ASR adds its own noise, and exact matching is hopeless, but intent plus template plus embedding matching was built for exactly this looseness, with the confidence gate guarding against the looseness becoming wrongness.
The bottom line
Streamed cache hits complete the illusion: CGETSTREAM feeds the TTS chunk by chunk, so the assistant starts speaking immediately and naturally. Users call the product 'snappy.' The dashboard calls it a 90th-percentile latency you didn't have to engineer twice.