Morshu TTS Server
A standalone HTTP server that generates speech in the voice of Morshu (from the Zelda CD-i games) using phoneme matching against sampled voice lines. It is based on MorshuTalk by jalenluorion (https://github.com/jalenluorion/MorshuTalk).
This service was extracted from the ttstoy Discord bot cog so it can run on its own as a system service, independent of the bot.
How it works
Input text is converted to phonemes with a grapheme-to-phoneme model, then the
engine splices together matching phoneme segments taken from a sample recording
(morshutalk_morshu.wav) to build the output audio.
Requirements
- Python 3.9 or newer
- Python packages:
numpy,pydub,nltk,g2p_en(see requirements.txt) - The bundled
morshutalk_morshu.wavsample (included in this repo)
On first import, NLTK downloads a few small data packages
(averaged_perceptron_tagger_eng, cmudict, punkt, punkt_tab).
Running
./start.sh
start.sh creates a local virtual environment on first launch, installs the
dependencies from requirements.txt, and starts the server. The listening port
defaults to 33002 and can be overridden with the PORT environment variable.
You can also run it directly against an existing interpreter that has the dependencies installed:
PORT=33002 python3 server.py
API
GET /say
Generate speech and return a WAV file.
GET /say?text=Hello%20it%20is%20me%20Morshu
Response: audio/wav (WAV bytes).
Example:
curl "http://127.0.0.1:33002/say?text=lamp+oil+rope+bombs" -o morshu.wav
GET /health
Returns ok.
Running as a system service
A systemd unit file is provided (morshu-tts-server.service). See INSTALL.md
for setup steps.