Can a local LLM answer as fast as the cloud?
Benchmarking full voice pipelines on consumer GPUs to see whether a private, local assistant can match the responsiveness of cloud services.
BSG / P-001 AI / Smart Home / Voice
Whole-home AI that listens, understands, and responds wherever you are.
A voice assistant is a cloud account plus a separate speaker in every room, none of which know about the others.
Why can't the whole house be one assistant, running on your own network, that knows which room you're in and who's talking?
Kenzy began in 2022 with one complaint: smart speakers are islands. Every room has its own device and its own context, and every request goes off to someone else's server. Kenzy treats the house as one system instead.
Small room nodes handle wake-word detection, capture, and playback. A central server runs the pipeline: speech-to-text, a language model, text-to-speech, and speaker identification. Each stage is a separate service that can run locally or call a cloud provider. Nothing leaves your network unless you send it.
Kenzy is split into independently deployable Python services: node, server, stt, llm, tts, and speaker, plus an experimental conversation engine. Room nodes run on single-board computers such as the Orange Pi Zero 2W, paired with ordinary USB speakerphones that have hardware echo cancellation.
Install a room node with pipx install "kenzy[node]", or use the installer on kenzy.ai.
Much of the work happens in public on research.kenzy.ai: whether consumer GPUs can match cloud response times, and how a house full of microphones decides which one should answer.