Real-Time Voice Assistant (voice-rt)
Status: Phase 2 complete & verified 2026-09-04 — live Tags: Service Deployment Infrastructure Multimodal Agentic Created: 2026-09-05 Updated: 2026-09-05 Related: local-llm-stack, cloudflare-pages-deploy, hermes-cron-operations
Definition
A live real-time, full-duplex voice assistant (speech-in → speech-out) running on this host, exposed to phones as a PWA at https://voice.meiyoucheveux.com. Built 2026-09-04: Phase 1 = offline STT→LLM→TTS round-trip (e2e.py, sample replies out_en.wav / out_zh.wav); Phase 2 = live streaming server with barge-in. Supports English + 中文 and multiple simultaneous phones over one websocket.
How It Works
Client-to-server path:
Phone PWA (mic, 16 kHz int16 PCM)
-> wss://voice.meiyoucheveux.com/ws (Cloudflare edge)
-> cloudflared.service (ingress voice.meiyoucheveux.com -> http://localhost:8090)
-> voice-rt/server.py (aiohttp, port 8090)
state machine: IDLE -> LISTENING -> PROCESSING -> SPEAKING
Per-utterance pipeline:
- VAD — Silero VAD: onset = 2 consecutive 32 ms chunks ≥ 0.5; end = 25 silent chunks; max 625 chunks.
- STT — faster-whisper medium, CPU int8 (
voice-rt/whisper/medium/). - LLM —
llama-server-voice4bon:8020(GPU 1, Qwen3-4B Q4_K_M, 4 slots) — a dedicated serving instance, see local-llm-stack. - TTS — Piper 1.7.0 (CPU),
en_US-lessac-medium/zh_CN-huayan-medium, streamed chunk-by-chunk (first audio ~20 ms after the LLM reply is ready).
Barge-in: speech detected during SPEAKING cancels the in-flight TTS+LLM and flushes the audio queue (~50 ms interrupt → IDLE).
Service / boot: runs as systemd voice-rt.service (enabled at boot, Restart=on-failure), pinning /home/romain/.hermes/hermes-agent/venv/bin/python (has aiohttp, faster_whisper, piper, onnxruntime). Boot order: docker.service → llama-server-voice4b container (unless-stopped, :8020) → cloudflared.service → voice-rt.service (:8090). Whisper + Piper models load on first use; the PWA serves immediately.
Key Parameters
| Item | Value |
|---|---|
| PWA | https://voice.meiyoucheveux.com (tap 📱 Install to add to home screen) |
| Server | voice-rt/server.py, aiohttp, port 8090 |
| Tunnel | Cloudflare, cloudflared.service, tunnel id 2350ecdf-b6f6-4b72-87d9-48f46b8108a5, config /home/romain/.cloudflared/config.yml (DNS voice is a plain CF-edge A record, not a CNAME) |
| STT | faster-whisper medium, CPU int8 |
| LLM | Qwen3-4B Q4_K_M on :8020 (GPU 1, 4 slots) — llama-server-voice4b |
| TTS | Piper 1.7.0, en_US-lessac + zh_CN-huayan (medium) |
| Languages | English + 中文 |
| Measured (LAN, 2026-09-04) | EN round-trip ~6.2 s wall; ZH ~8.1 s (STT of the full clip ~3.7 s dominates); first TTS ~20–60 ms after LLM ready; barge-in ~50 ms; 2 concurrent sessions — no cross-stall (total ≈ max, not sum) |
When To Use
- Live hands-free voice assistant on a phone (mobile Safari/Chrome PWA), EN/ZH, multiple devices at once.
- Reference design for a full-duplex speech pipeline on this host’s local stack: VAD + faster-whisper + a small GPU LLM + Piper, fronted by a Cloudflare websocket tunnel.
Risks & Pitfalls
- Silent bot on phone (transcript shows, no voice): mobile Safari/Chrome keep the playback
AudioContextsuspended until a user tap. Unlock it inside the 🎤 tap (playCtx().resume()instart()), not later when audio arrives. Also: volume up, not on the silent switch. Diagnose via the bottom status line — “speaking…” = audio blocked (client); stuck on “thinking…” = server/tunnel. - Invariants — never touch: the 27 B llama-server on
:8000andvllm-06bon:8010must stay untouched by voice-rt. Do not create a.venv-voice; all deps live in the hermes venv. ffmpeg =voice-rt/bin/ffmpeg(symlink to the imageio-ffmpeg static binary). - edge-tts test clips have internal pauses → VAD fragments them and may fire a
barge_inmid-test (correct behaviour, not a bug). write_filetruncates content >~1 KB — build big files viaexecute_codeor small part-files.- Model downloads only via
https://hf-mirror.com(huggingface.co is blocked; the pythonhuggingface_hubAPI 401s on CAS — usewget). - PyPI file downloads are slow — run
pipinstalls as background jobs. - Server stdout is buffered when backgrounded — trust test-client output, not the server log.
Tests
cd /home/romain/workspace/voice-rt
python3 test_ws.py in_en.mp3 en # EN round-trip (LAN)
python3 test_ws.py in_zh.mp3 zh # 中文 round-trip
python3 test_ws.py in_en.mp3 en interrupt # barge-in (~50 ms)
python3 conc_test.py 2 # 2 concurrent sessions
# public (through Cloudflare):
sed 's|ws://127.0.0.1:8090/ws|wss://voice.meiyoucheveux.com/ws|' test_ws.py > /tmp/pub.py
python3 /tmp/pub.py in_en.mp3 enRelated Concepts
- local-llm-stack — the llama.cpp / vLLM serving stack, including the dedicated
:8020voice4b LLM that this assistant calls - cloudflare-pages-deploy — Cloudflare tooling on this host (Python deployer); the tunnel here is the separate
cloudflared.service - hermes-cron-operations — general host infra / ops context
Sources
/home/romain/workspace/voice-rt/RUNBOOK.md(runbook; Phase 2 verified 2026-09-04)/home/romain/workspace/voice-rt/—server.py,www/index.html(PWA),test_ws.py,conc_test.py,e2e.py,bench.py,voice-rt.service- Quick health check:
curl http://127.0.0.1:8090/,:8020/health,:8000/health,https://voice.meiyoucheveux.com/(all 200) +systemctl is-active cloudflared.service voice-rt.service