Initial standalone vox TTS server extracted from the ttstoy bot cog
Self-contained HTTP service with engine, data, start script, systemd unit, and documentation. Runs independently of the Discord bot on its fixed port.
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
# Port the VOX TTS server listens on
|
||||
PORT=33003
|
||||
@@ -0,0 +1,6 @@
|
||||
__pycache__/
|
||||
*.pyc
|
||||
venv/
|
||||
.venv/
|
||||
*.log
|
||||
*.zip
|
||||
+38
@@ -0,0 +1,38 @@
|
||||
# Installing the VOX TTS Server as a system service
|
||||
|
||||
These steps install the server as a systemd service that starts on boot and
|
||||
restarts automatically if it crashes. Commands that need root are shown with
|
||||
`sudo`; run them in your own terminal.
|
||||
|
||||
## 1. Install prerequisites
|
||||
|
||||
The server uses only the Python standard library.
|
||||
|
||||
```bash
|
||||
sudo apt install python3
|
||||
```
|
||||
|
||||
## 2. Install the systemd unit
|
||||
|
||||
```bash
|
||||
sudo cp /home/owen/vox-tts-server/vox-tts-server.service /etc/systemd/system/
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl enable --now vox-tts-server
|
||||
```
|
||||
|
||||
## 3. Verify
|
||||
|
||||
```bash
|
||||
systemctl status vox-tts-server
|
||||
curl -s http://127.0.0.1:33003/health
|
||||
```
|
||||
|
||||
The health endpoint should return `ok`.
|
||||
|
||||
## Managing the service
|
||||
|
||||
```bash
|
||||
sudo systemctl restart vox-tts-server
|
||||
sudo systemctl stop vox-tts-server
|
||||
journalctl -u vox-tts-server -f
|
||||
```
|
||||
@@ -0,0 +1,75 @@
|
||||
# VOX TTS Server
|
||||
|
||||
A standalone HTTP server that produces Half-Life / Black Mesa VOX announcer
|
||||
speech by concatenating pre-recorded word clips. It is based on VOXGen by
|
||||
wphillips (https://github.com/wphillips/VOXGen).
|
||||
|
||||
This service was extracted from the ttstoy Discord bot cog so it can run on its
|
||||
own as a system service, independent of the bot.
|
||||
|
||||
## How it works
|
||||
|
||||
The input text is split into words, and each word is matched to a `.wav` clip in
|
||||
the selected voice pack. Matching clips are concatenated into a single WAV
|
||||
response. Words with no matching clip are skipped.
|
||||
|
||||
Two packs are bundled:
|
||||
|
||||
- `vox`: the classic Half-Life VOX announcer words (in `vox_words/`)
|
||||
- `vox2`: the Black Mesa announcement system words (in `vox2_words/`)
|
||||
|
||||
## Requirements
|
||||
|
||||
- Python 3.9 or newer
|
||||
- No third-party packages: the server uses only the Python standard library
|
||||
|
||||
## Running
|
||||
|
||||
```bash
|
||||
./start.sh
|
||||
```
|
||||
|
||||
The listening port defaults to `33003` and can be overridden with the `PORT`
|
||||
environment variable.
|
||||
|
||||
You can also run it directly:
|
||||
|
||||
```bash
|
||||
PORT=33003 python3 server.py
|
||||
```
|
||||
|
||||
## API
|
||||
|
||||
### GET /say
|
||||
|
||||
Generate speech and return a WAV file.
|
||||
|
||||
```
|
||||
GET /say?text=warning+reactor+core+meltdown&pack=vox
|
||||
```
|
||||
|
||||
Query parameters:
|
||||
|
||||
- `text` (required): words to synthesize, separated by spaces
|
||||
- `pack` (optional): `vox` (default) or `vox2`
|
||||
|
||||
Response: `audio/wav` (WAV bytes).
|
||||
|
||||
Example:
|
||||
|
||||
```bash
|
||||
curl "http://127.0.0.1:33003/say?text=alert+intruder&pack=vox" -o vox.wav
|
||||
```
|
||||
|
||||
### GET /packs
|
||||
|
||||
Returns a comma-separated list of installed packs.
|
||||
|
||||
### GET /health
|
||||
|
||||
Returns `ok`.
|
||||
|
||||
## Running as a system service
|
||||
|
||||
A systemd unit file is provided (`vox-tts-server.service`). See INSTALL.md for
|
||||
setup steps.
|
||||
@@ -0,0 +1,101 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
VOX TTS Server — standalone HTTP API for Half-Life VOX engine.
|
||||
Runs on port 33003. GET /say?text=...&pack=vox returns WAV audio.
|
||||
"""
|
||||
import os
|
||||
import sys
|
||||
import io
|
||||
import wave
|
||||
from pathlib import Path
|
||||
from http.server import HTTPServer, BaseHTTPRequestHandler
|
||||
from urllib.parse import urlparse, parse_qs
|
||||
|
||||
# Engine lives alongside this server (self-contained)
|
||||
SCRIPT_DIR = Path(__file__).resolve().parent
|
||||
sys.path.insert(0, str(SCRIPT_DIR))
|
||||
|
||||
from vox_engine import VOX_PACKS, get_available_words, get_available_packs
|
||||
|
||||
PORT = int(os.environ.get("PORT", 33003))
|
||||
|
||||
|
||||
def generate_wav(text, pack="vox"):
|
||||
"""Concatenate word WAVs and return raw WAV bytes."""
|
||||
pack_dir = VOX_PACKS.get(pack)
|
||||
if not pack_dir or not pack_dir.exists():
|
||||
raise RuntimeError(f"VOX pack '{pack}' not found")
|
||||
|
||||
words = text.lower().split()
|
||||
splice = bytes()
|
||||
sample_rate = 11025
|
||||
sample_width = 1
|
||||
channels = 1
|
||||
found = False
|
||||
|
||||
for word in words:
|
||||
wav_path = pack_dir / f"{word}.wav"
|
||||
if not wav_path.exists():
|
||||
continue
|
||||
with wave.open(str(wav_path), 'rb') as snd:
|
||||
sample_rate = snd.getframerate()
|
||||
sample_width = snd.getsampwidth()
|
||||
channels = snd.getnchannels()
|
||||
splice += snd.readframes(snd.getnframes())
|
||||
found = True
|
||||
|
||||
if not found:
|
||||
raise RuntimeError("No matching VOX words found")
|
||||
|
||||
buf = io.BytesIO()
|
||||
with wave.open(buf, 'wb') as out:
|
||||
out.setnchannels(channels)
|
||||
out.setsampwidth(sample_width)
|
||||
out.setframerate(sample_rate)
|
||||
out.writeframes(splice)
|
||||
return buf.getvalue()
|
||||
|
||||
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
def do_GET(self):
|
||||
parsed = urlparse(self.path)
|
||||
if parsed.path == "/say":
|
||||
params = parse_qs(parsed.query)
|
||||
text = params.get("text", [""])[0]
|
||||
if not text:
|
||||
self.send_error(400, "Missing 'text' parameter")
|
||||
return
|
||||
pack = params.get("pack", ["vox"])[0]
|
||||
try:
|
||||
wav_bytes = generate_wav(text, pack)
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "audio/wav")
|
||||
self.send_header("Content-Length", str(len(wav_bytes)))
|
||||
self.end_headers()
|
||||
self.wfile.write(wav_bytes)
|
||||
except Exception as e:
|
||||
self.send_error(500, str(e))
|
||||
elif parsed.path == "/health":
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "text/plain")
|
||||
self.end_headers()
|
||||
self.wfile.write(b"ok")
|
||||
elif parsed.path == "/packs":
|
||||
packs = get_available_packs()
|
||||
body = ",".join(packs).encode()
|
||||
self.send_response(200)
|
||||
self.send_header("Content-Type", "text/plain")
|
||||
self.end_headers()
|
||||
self.wfile.write(body)
|
||||
else:
|
||||
self.send_error(404)
|
||||
|
||||
def log_message(self, format, *args):
|
||||
print(f"[VOX] {args[0]}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
print(f"[VOX] Starting server on port {PORT}...")
|
||||
server = HTTPServer(("0.0.0.0", PORT), Handler)
|
||||
print(f"[VOX] Ready — http://127.0.0.1:{PORT}/say?text=hello+world")
|
||||
server.serve_forever()
|
||||
@@ -0,0 +1,10 @@
|
||||
#!/bin/bash
|
||||
# Start script for the VOX TTS server.
|
||||
# The server uses only the Python standard library, so no venv is required.
|
||||
set -e
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
cd "$SCRIPT_DIR"
|
||||
|
||||
export PORT="${PORT:-33003}"
|
||||
exec python3 "$SCRIPT_DIR/server.py"
|
||||
@@ -0,0 +1,16 @@
|
||||
[Unit]
|
||||
Description=VOX TTS Server
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=owen
|
||||
WorkingDirectory=/home/owen/vox-tts-server
|
||||
Environment=PORT=33003
|
||||
ExecStart=/home/owen/vox-tts-server/start.sh
|
||||
Restart=always
|
||||
RestartSec=3
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user