Files
kingston e61a320499 Initial standalone morshu TTS server extracted from the ttstoy bot cog
Self-contained HTTP service with engine, data, start script, systemd unit,
and documentation. Runs independently of the Discord bot on its fixed port.
2026-09-14 14:31:15 -05:00

68 lines
1.7 KiB
Markdown

# Morshu TTS Server
A standalone HTTP server that generates speech in the voice of Morshu (from the
Zelda CD-i games) using phoneme matching against sampled voice lines. It is
based on MorshuTalk by jalenluorion (https://github.com/jalenluorion/MorshuTalk).
This service was extracted from the ttstoy Discord bot cog so it can run on its
own as a system service, independent of the bot.
## How it works
Input text is converted to phonemes with a grapheme-to-phoneme model, then the
engine splices together matching phoneme segments taken from a sample recording
(`morshutalk_morshu.wav`) to build the output audio.
## Requirements
- Python 3.9 or newer
- Python packages: `numpy`, `pydub`, `nltk`, `g2p_en` (see requirements.txt)
- The bundled `morshutalk_morshu.wav` sample (included in this repo)
On first import, NLTK downloads a few small data packages
(`averaged_perceptron_tagger_eng`, `cmudict`, `punkt`, `punkt_tab`).
## Running
```bash
./start.sh
```
`start.sh` creates a local virtual environment on first launch, installs the
dependencies from requirements.txt, and starts the server. The listening port
defaults to `33002` and can be overridden with the `PORT` environment variable.
You can also run it directly against an existing interpreter that has the
dependencies installed:
```bash
PORT=33002 python3 server.py
```
## API
### GET /say
Generate speech and return a WAV file.
```
GET /say?text=Hello%20it%20is%20me%20Morshu
```
Response: `audio/wav` (WAV bytes).
Example:
```bash
curl "http://127.0.0.1:33002/say?text=lamp+oil+rope+bombs" -o morshu.wav
```
### GET /health
Returns `ok`.
## Running as a system service
A systemd unit file is provided (`morshu-tts-server.service`). See INSTALL.md
for setup steps.