Skip to main content

Local - Piper

Local text-to-speech via Piper, a fast neural TTS system optimized for local execution.

Official Documentation


Local Setup

Wyoming-Piper Server:

docker run -d \
--name piper \
-p 10200:10200 \
-v /path/to/voices:/data \
rhasspy/wyoming-piper \
--voice en_US-lessac-medium

With HTTP Server (for REST API):

# Using piper-http-server
pip install piper-http-server
piper-http-server --port 5002 --model en_US-amy-medium.onnx

Manual Installation

1. Download Piper binary:

# Linux x86_64
wget https://github.com/rhasspy/piper/releases/latest/download/piper_linux_x86_64.tar.gz
tar -xzf piper_linux_x86_64.tar.gz
cd piper

2. Download a voice model:

# Download voice model and config
wget https://huggingface.co/rhasspy/piper-voices/resolve/main/en/en_US/amy/medium/en_US-amy-medium.onnx
wget https://huggingface.co/rhasspy/piper-voices/resolve/main/en/en_US/amy/medium/en_US-amy-medium.onnx.json

3. Run HTTP server:

pip install piper-http-server
piper-http-server --port 5002 --model en_US-amy-medium.onnx

Using pip (Python)

pip install piper-tts

# Test from command line
echo "Hello world" | piper --model en_US-amy-medium.onnx --output_file test.wav

Verify

# Test TTS endpoint
curl -X POST http://localhost:5002/synthesize \
-H "Content-Type: application/json" \
-d '{"text": "Hello world"}' \
--output test.wav

# Play the audio
ffplay test.wav # or: aplay test.wav

Provider Configuration

import { PiperTTSProvider } from '@llmrtc/llmrtc-provider-local';

const tts = new PiperTTSProvider({
baseUrl: process.env.PIPER_URL || 'http://localhost:5002'
});

Environment Variables

VariableDefaultDescription
PIPER_URLhttp://localhost:5002Piper server URL

Provider Options

interface PiperConfig {
baseUrl?: string; // Server URL
voice?: string; // Voice model name
}

Available Voices

VoiceLanguageQualitySize
en_US-amy-mediumEnglish (US)Good60MB
en_US-lessac-mediumEnglish (US)Good60MB
en_US-ryan-mediumEnglish (US)Good60MB
en_GB-cori-mediumEnglish (UK)Good60MB
de_DE-thorsten-mediumGermanGood60MB
es_ES-mls_9972-mediumSpanishGood60MB
fr_FR-siwis-mediumFrenchGood60MB

Browse all 30+ languages at rhasspy.github.io/piper-samples.


Key Features

  • Fast Inference: Optimized for real-time use, even on Raspberry Pi
  • ONNX Models: Uses efficient ONNX-based VITS voice models
  • Multi-language: 30+ languages supported
  • Small Footprint: Voice models are typically 60-100MB
  • No Internet Required: Fully offline operation

Notes

  • Choose a fast Piper voice (medium quality) for real-time use
  • Works well with streaming TTS enabled in the server
  • Both .onnx model and .json config files are required
  • Python package supports 3.7-3.10
  • For lowest latency, run on a machine with good single-thread CPU performance