Skip to content

5.2.48. Speaches (unified TTS + STT engine)

Speaches is a dual-role engine — one container exposes both /v1/audio/transcriptions (STT, Faster-Whisper) and /v1/audio/speech (TTS, Kokoro + Piper voices). It is selectable via either STT_PROVIDER_SOURCE=speaches-* or TTS_PROVIDER_SOURCE=speaches-*. When both roles pick a Speaches variant, the bootstrapper dedupes to one running container.

It is documented under both aggregators:

1. Engine quick reference

  • Images:
  • CPU: ghcr.io/speaches-ai/speaches:0.9.0-rc.3-cpu
  • GPU: ghcr.io/speaches-ai/speaches:0.9.0-rc.3-cuda
  • License: MIT
  • Activation: any of
  • STT_PROVIDER_SOURCE=speaches-container-cpu
  • STT_PROVIDER_SOURCE=speaches-container-gpu
  • TTS_PROVIDER_SOURCE=speaches-container-cpu
  • TTS_PROVIDER_SOURCE=speaches-container-gpu
  • In-container port: 8000
  • Host port: ${SPEACHES_PORT} (computed from BASE_PORT)

The manifest (service.yml) and compose fragment (compose.yml) in this folder are the bootstrapper's source of truth for those values; treat this README as a pointer, not a duplicate of the aggregator docs.

2. Dependencies & Integrations

2.1. Current — Upstream (this service calls)

No upstream calls.

2.2. Current — Downstream (services that call this)

Service Category
kong infra
hermes agents
n8n agents
jupyterhub apps
open-webui apps

2.3. Architecture diagram

speaches architecture

Open the full-size diagram for a full-screen view.

2.4. Future — Missing pair integrations

No high-confidence opportunities identified.

2.5. Future — Candidate new services

No high-confidence opportunities identified.

2.6. Future — Unused features in this service

No high-confidence opportunities identified.