Can a local LLM answer as fast as the cloud?
Benchmarking full voice pipelines on consumer GPUs to see whether a private, local assistant can match the responsiveness of cloud services.
Yes. A Mixture-of-Experts model with about 3B active parameters spoke its first word in roughly 330 ms on a single RTX 5090.
The question
Cloud voice assistants reply in one to two seconds. If Kenzy is going to keep inference on your own network, it has to be at least that fast on hardware a person could actually buy. Can it be?
What was measured
The complete voice pipeline, timed end to end: speech-to-text, then the language model, then text-to-speech. Each request carried a realistic load of about 8,700 tokens of context and 31 tool definitions, with 500 requests per configuration.
- Models: a 27B dense model and a 35B Mixture-of-Experts model, plus full-duplex speech-to-speech models.
- GPUs: NVIDIA RTX 5090, RTX PRO 6000, and RTX 6000 Ada.
Result
The 35B MoE, activating roughly 3B parameters per token, reached a time-to-first-token of about 94 ms and finished the language stage in about 264 ms on an RTX 5090. End to end, the first spoken word arrived in about 330 ms and the device action in about 480 ms. Speech recognition, the model, and speech synthesis all fit on one 32 GB card.
Takeaway
Latency depends on active compute and architecture, not headline parameter count. A local assistant doesn't have to trade speed for privacy.