Self-contained HTTP service with engine, data, start script, systemd unit, and documentation. Runs independently of the Discord bot on its fixed port.
1.6 KiB
VOX TTS Server
A standalone HTTP server that produces Half-Life / Black Mesa VOX announcer speech by concatenating pre-recorded word clips. It is based on VOXGen by wphillips (https://github.com/wphillips/VOXGen).
This service was extracted from the ttstoy Discord bot cog so it can run on its own as a system service, independent of the bot.
How it works
The input text is split into words, and each word is matched to a .wav clip in
the selected voice pack. Matching clips are concatenated into a single WAV
response. Words with no matching clip are skipped.
Two packs are bundled:
vox: the classic Half-Life VOX announcer words (invox_words/)vox2: the Black Mesa announcement system words (invox2_words/)
Requirements
- Python 3.9 or newer
- No third-party packages: the server uses only the Python standard library
Running
./start.sh
The listening port defaults to 33003 and can be overridden with the PORT
environment variable.
You can also run it directly:
PORT=33003 python3 server.py
API
GET /say
Generate speech and return a WAV file.
GET /say?text=warning+reactor+core+meltdown&pack=vox
Query parameters:
text(required): words to synthesize, separated by spacespack(optional):vox(default) orvox2
Response: audio/wav (WAV bytes).
Example:
curl "http://127.0.0.1:33003/say?text=alert+intruder&pack=vox" -o vox.wav
GET /packs
Returns a comma-separated list of installed packs.
GET /health
Returns ok.
Running as a system service
A systemd unit file is provided (vox-tts-server.service). See INSTALL.md for
setup steps.