vocero-s2s/docs/PROJECT_OVERVIEW.md
valenti b5f82fb48c
Some checks are pending
CI / ruff (push) Waiting to run
CI / mypy (push) Waiting to run
CI / pytest (push) Waiting to run
CI / package (push) Waiting to run
CI / Install smoke (${{ matrix.label }}) (linux, ubuntu-latest) (push) Blocked by required conditions
CI / Install smoke (${{ matrix.label }}) (macos-arm64, macos-14) (push) Blocked by required conditions
first git
2026-08-26 11:30:14 +00:00

20 KiB
Raw Permalink Blame History

speech-to-speech — Project Overview

Package: speech-to-speech Version: 0.2.11 Author: Hugging Face License: Apache-2.0 Python: 3.10 3.12 Tagline: Low-latency end-to-end Speech-to-Speech pipeline for building realtime voice agents.


1. What is speech-to-speech?

speech-to-speech is a fully modular, low-latency audio pipeline that turns a user's spoken input into a spoken reply by chaining four AI stages together: VAD → STT → LLM → TTS.

Out of the box it exposes an OpenAI-compatible Realtime WebSocket endpoint (ws://host:port/v1/realtime), so any browser or SDK that knows how to speak the OpenAI Realtime Protocol can connect and start having natural voice conversations in seconds — without writing a single line of server code.

The pipeline is designed to be backend-agnostic at every stage: you can plug different STT / LLM / TTS models depending on the available hardware (NVIDIA CUDA, Apple Silicon MLX, or CPU-only), the language(s) of your users, and the latency / quality trade-off you need.

A built-in, optional RAG (Retrieval Augmented Generation) server-side subsystem lets the LLM answers be grounded in your own private knowledge base (Markdown / TXT / JSONL documents), while staying 100 % transparent for the client. Dynamic knowledge base updates are exposed over the same HTTP server via a 11-endpoint REST API.


2. Key Features

Feature Description
Real-time voice protocol Native OpenAI v1/realtime WebSocket support: session.update, response.create, function calling, audio deltas, interruptions. Drop-in compatible with the official Realtime SDK and browser playgrounds.
4 execution modes local (mic → speakers), socket (TCP IPC), websocket (raw audio WS), realtime (OpenAI protocol — default).
6 pluggable STT backends Whisper / Whisper-MLX / MLX-Audio-Whisper / Faster-Whisper / Parakeet TDT (default) / Paraformer.
4 pluggable LLM backends transformers (local) · mlx-lm (Apple Silicon) · responses-api (OpenAI tool-calling aware endpoint) · chat-completions (any OpenAI-compatible /v1/chat/completions server).
5 pluggable TTS backends ChatTTS · Facebook MMS · Pocket (tiny CPU) · Kokoro · Qwen3-TTS (default, custom voice 1.7B).
Multi-hardware ready NVIDIA CUDA (Linux), Apple Silicon MLX + MPS (macOS), CPU fallback for every stage.
Pool of N isolated pipelines --num_pipelines N runs N independent sessions in parallel (each one with its own VAD / STT / LLM / TTS handlers and conversation state). Ideal for small multi-concurrent deployments.
Built-in VAD & interruption Voice-activity detection with configurable threshold and streaming silence → STT finalization; the LLM can be interrupted mid-speech by the user and cancelled cleanly through a CancelScope.
Automatic language detection Via lingua-language-detector for the assistant reply language.
Live partial transcription Parakeet-TDT emits live partial transcriptions every 500 ms so clients can display "user is speaking …" text.
Full tool-calling support With responses-api backend: streamed function argument delivery, parallel tool calls, custom voices over TTS — everything compliant with the OpenAI Responses API surface.
Optional RAG server-side injection Transparent retrieval hook on every LLM turn. Sentence-Transformers embeddings, NPZ-index persistence, configurable top-k / cosine threshold, system or user-message injection. Multilingual by default (Spanish / Italian / English out of the box).
Dynamic Knowledge Base REST API 11 endpoints to list, search, upsert, update, remove, and reload KB contents at runtime — including cross-restart persistence through kb/_dynamic.jsonl.
Structured CLI / JSON configuration Every parameter is a @dataclass HfArgumentParser argument with sane defaults, or a full JSON config file you can pass in as single argument.
PyPI-ready packaging & release flow Configured uv build + twine check + GitHub Actions publish workflow (tag vX.Y.Z triggers upload).

3. Architecture at a glance

A single conversation turn in realtime mode looks like this:

                         ┌─────────────────────────────────────────────────────────┐
                         │                    uvicorn + FastAPI                    │
                         │  ┌───────────────────────────────────────────────────┐  │
   Mic / Browser  ──► WS │  │   RealtimeService (session + routing + events)    │  │
   (audio in +       │   │  └───────────────────────────────────────────────────┘  │
    events)          │   │                                                           │
                     │   │  Pipeline pool  ──►  PipelineUnit #1  ──►  PipelineUnit #N│
                     │   └─────────────────────────────────────────────────────────┘
                     │                      │
                     ▼                      ▼
              ┌──────────────────────────────────────────────┐
              │     Single pipeline unit (one per user)       │
              │                                                │
              │   1. VAD            ──►   speech start / end   │
              │   2. STT (5 flavours) ──►   last user text     │
              │        │                                     │
              │        ▼                                     │
              │   ┌────────────────────────┐                  │
              │   │  RAG SEARCH HOOK        │ ◄─── kb/_index.npz
              │   │  (inject only if ≥ N)   │      + md/txt/jsonl
              │   └────────────────────────┘      + _dynamic.jsonl
              │        │                                     │
              │        ▼                                     │
              │   3. LLM (4 flavours) ──►  text reply + tools│
              │        │                                     │
              │        ▼                                     │
              │   4. TTS (5 flavours) ──►  streamed PCM audio│
              └──────────────────────────────────────────────┘
                                                      │
                                                      ▼
                                              ◄── WS audio / delta events

Key invariants:

  • Each connection = one PipelineUnit (allocated atomically; queues and state are isolated; if N units are busy, the N+1-th connection is rejected).
  • The LLM stage never sees raw audio. The pipeline only forwards text buffers to the LLM handler, so model vendors and backend switches remain transparent.
  • The RAG stage is a pure side-effect on the text buffer: it reads the most recent user text, runs an embedding + cosine top-k + threshold filter, and prepends matching knowledge as an extra system (or user) message — with a clear log line so you can always see exactly what was injected.
  • TTS streams (all modern backends): audio deltas are emitted as they are synthesised, so the first byte of a reply reaches the speaker well before the LLM has finished generating text.

4. Execution modes (--mode)

Configurable via ModuleArguments.mode (default: realtime).

Mode Input Output Best for
local Local microphone via sounddevice / miniaudio Local speakers Desktop demos, quick prototyping on laptop.
socket Raw PCM chunks on TCP socket Raw PCM on TCP socket Legacy / embedded setups, custom transport.
websocket Raw PCM over WS endpoint Raw PCM over WS Minimal custom frontend that just pushes audio.
realtime OpenAI Realtime Protocol WS (/v1/realtime) Same protocol + RAG REST API on the same HTTP server Production & SDK integration. All clients that support the Realtime API (OpenAI SDK, web playgrounds, Swift / Kotlin / JS wrappers) connect here.

5. Supported backends

5.1 STT — Speech to Text

--stt name Model family Default model Device highlights
whisper HuggingFace Transformers Whisper distil-whisper/distil-large-v3 CUDA / CPU
whisper-mlx MLX Whisper mlx-community/whisper-large-v3-turbo Apple Silicon (macOS)
mlx-audio-whisper Apple mlx-audio mlx-community/whisper-large-v3-turbo Apple Silicon
faster-whisper CTranslate2 Quantized Whisper tiny.en Very low latency CPU
parakeet-tdt (default) HuggingFace Parakeet TDT parakeet-tdt-1.1b Streaming-friendly + live partial transcription on CUDA
paraformer Alibaba Paraformer paraformer-zh Chinese-only deployments

5.2 LLM — Language Model

--llm_backend Description Default model
transformers Local, process-bound HF Transformers generation Qwen/Qwen3-4B-Instruct-2507
mlx-lm Apple Silicon local 4-bit / 8-bit quantised inference mlx-community/...
responses-api (default) OpenAI Responses-API compatible remote endpoint (full tool-calling surface + streamed function args) gpt-5.4-mini
chat-completions Any OpenAI-compatible /v1/chat/completions remote — plug in vLLM, TGI, Ollama, llama.cpp server, SGLang, TabbyAPI, etc. (auto-detects via base_url + /v1/models)

5.3 TTS — Text to Speech

--tts name Backend Default model Strengths
chatTTS ChatTTS (optional group chattts) Highly conversational, English + Chinese, very expressive prosody
facebookMMS Facebook MMS (optional group facebook-mms) facebook/mms-tts-eng Ultra-lightweight, covers 1,100+ languages
pocket PocketTTS (small CPU) Zero-dependency on-device, great for embedded / low-RAM
kokoro Kokoro TTS (optional group kokoro) State of the art English & Japanese neural TTS
qwen3 (default) Qwen3-TTS via faster-qwen3-tts GGML Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice Multilingual, supports reference voice cloning, 12 Hz token rate → ultra-low latency.

All TTS backends are reachable through a common streaming interface so the rest of the pipeline is backend-agnostic.


6. RAG Server-Side (optional add-on)

The RAG subsystem is fully documented in:

At a glance:

$ ./start_pipeline_rag.sh                         # starts pipeline + RAG REST API
RAG: Inizializzazione modello embedding=paraphrase-multilingual-MiniLM-L12-v2 device=cuda su kb_path=/home/.../kb
RAG: Índice cargado desde disco: 9 chunk (shape=(9, 384)).
RAG: Activo. kb=... top_k=3 umbral=0.250 inject_as=system idioma=es
RAG HTTP API montata su prefix='/v1/rag'

Highlights:

  • Zero client changes — injection happens server-side in the LLM chat buffer before every turn.
  • Pluggable embedding models with a default multilingual model (paraphrase-multilingual-MiniLM-L12-v2, 384 dims) and recommended alternatives for Spanish-only and very-large-KB scenarios.
  • NPZ index persistence — rebuilds only if content changed. Supports --rag_force_rebuild.
  • Three source formats: Markdown / TXT with automatic recursive chunking (size + overlap configurable), or JSONL for hand-crafted chunks.
  • Configurable retrieval behaviour: --rag_top_k, --rag_threshold, --rag_inject_as (system/user), --rag_language (es/it/en).
  • Thread-safe dynamic REST API (11 endpoints): /status, /sources, /chunks, /chunks/list, /chunks/update, /upsert/document, /search, /add/document, /add/chunks, /remove, /reload.
  • Cross-restart persistence: API calls with persist=true are appended atomically to kb/_dynamic.jsonl and automatically re-indexed on the next startup.
  • Idempotent CRUD pattern: /upsert/document handles create / replace / delete (text="") atomically so clients only need one URL.
  • Fails gracefully: any retrieval exception is logged but never breaks the conversation turn.

7. Quick start

7.1 Install (PyPI-style editable install)

git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech

# Base package (includes default parakeet-tdt STT + qwen3 TTS)
uv pip install -e .

# Optional: add the RAG subsystem for server-side KB retrieval
uv pip install -e ".[rag]"

# Optional: select which TTS / STT extras you want to bring along
uv pip install -e ".[rag,chattts,kokoro,faster-whisper]"

7.2 Launch in Realtime mode (default) using an external LLM

This is the most common configuration: LLM runs on a remote OpenAI-compatible server (e.g. http://127.0.0.1:8001/v1 with vLLM or TGI), STT + TTS run locally on CUDA / MLX.

Create a small shell wrapper (see start_pipeline.sh for the full template):

#!/bin/bash
set -euo pipefail

LLM_BASE_URL="${LLM_BASE_URL:-http://127.0.0.1:8001/v1}"
LLM_MODEL="${LLM_MODEL:-Qwen/Qwen2.5-14B-Instruct-GPTQ-Int4}"
LLM_API_KEY="${LLM_API_KEY:-placeholder}"
TTS_MODEL="${TTS_MODEL:-Qwen/Qwen3-TTS-12Hz-1.7B-Base}"
STT_MODEL="${STT_MODEL:-distil-whisper/distil-large-v3}"

cd /home/azurian/speech-to-speech

uv run speech-to-speech \
  --mode realtime \
  --ws_host 0.0.0.0 \
  --ws_port 12345 \
  --num_pipelines 2 \
  \
  --llm_backend chat-completions \
  --chat_completions_handler_base_url "$LLM_BASE_URL" \
  --chat_completions_handler_model_name "$LLM_MODEL" \
  --chat_completions_handler_api_key "$LLM_API_KEY" \
  \
  --stt whisper \
  --whisper_stt_model_name "$STT_MODEL" \
  \
  --tts qwen3 \
  --qwen3_tts_model_name "$TTS_MODEL"

Run it, then connect any OpenAI Realtime client to:

ws://<host>:12345/v1/realtime

7.3 Same pipeline + RAG enabled

Append these flags (see §6 for the full reference and the helper script start_pipeline_rag.sh):

  --rag_enabled \
  --rag_kb_path ./kb \
  --rag_top_k 3 \
  --rag_threshold 0.25 \
  --rag_language es \
  --rag_inject_as system

The RAG REST API appears immediately on the same HTTP server:

curl http://127.0.0.1:12345/v1/rag/status | jq

8. Project layout (highlights)

speech-to-speech/
├── pyproject.toml                                     ← package metadata + optional deps groups
├── LICENSE                                            ← Apache-2.0
├── start_pipeline.sh                                  ← reference launch script
├── start_pipeline_rag.sh                              ← launch script + RAG enabled
│
├── src/speech_to_speech/
│   ├── s2s_pipeline.py                                ← entry point (main), pipeline builder, realtime pool
│   ├── baseHandler.py
│   ├── chat.py
│   ├── pipeline/                                      ← queue types, CancelScope, handler types
│   │
│   ├── arguments_classes/                             ← @dataclass HfArgumentParser args (one per handler)
│   │   ├── module_arguments.py                        ← mode / stt / tts / llm_backend / num_pipelines
│   │   ├── rag_arguments.py                           ← 12 RAG-specific flags
│   │   └── ... (Whisper, Qwen3, ChatTTS, VAD, …)
│   │
│   ├── STT/                                           ← six STT handlers
│   ├── TTS/                                           ← five TTS handlers
│   ├── LLM/
│   │   ├── base_openai_compatible_language_model.py   ← RAG injection hook + full tool-calling
│   │   ├── chat_completions_language_model.py
│   │   ├── responses_api_language_model.py
│   │   └── language_model.py                          ← local transformers / mlx-lm handlers
│   │
│   ├── RAG/
│   │   ├── retriever.py                               ← singleton, embeddings, NPZ, search, full CRUD
│   │   └── router.py                                  ← 11 FastAPI /v1/rag endpoints
│   │
│   └── api/openai_realtime/
│       ├── websocket_router.py                        ← FastAPI + Realtime + conditional RAG mount
│       ├── service.py                                 ← Realtime event loop + session/handler routing
│       └── ...
│
├── kb/
│   ├── 01_faq_producto.md                             ← example Spanish FAQ (included in scaffold)
│   ├── 02_politicas_internas.md                       ← example Spanish policies
│   ├── README.md                                      ← format guide
│   └── _dynamic.jsonl                                 ← API-added chunks (auto-generated)
│
└── docs/
    ├── RAG_SERVER_SIDE.md                             ← Italian/English in-depth RAG guide (600+ lines)
    └── RAG_SERVER_SIDE.es.md                          ← Spanish translation

9. Versioning and publishing

The repository ships with a full PyPI release pipeline in .github/workflows/publish.yml:

  1. Bump version in pyproject.toml and __version__ in src/speech_to_speech/init.py.
  2. Merge a release PR containing only those two changes.
  3. Tag and push:
    git checkout main && git pull origin main
    git tag -a vX.Y.Z -m "Release vX.Y.Z"
    git push origin vX.Y.Z
    
  4. The workflow runs uv build + twine check --strict + PyPI upload automatically.

See AGENTS.md for the repository-level release rules.


10. Where to go next