Morshu TTS Server

A standalone HTTP server that generates speech in the voice of Morshu (from the Zelda CD-i games) using phoneme matching against sampled voice lines. It is based on MorshuTalk by jalenluorion (https://github.com/jalenluorion/MorshuTalk).

This service was extracted from the ttstoy Discord bot cog so it can run on its own as a system service, independent of the bot.

How it works

Input text is converted to phonemes with a grapheme-to-phoneme model, then the engine splices together matching phoneme segments taken from a sample recording (morshutalk_morshu.wav) to build the output audio.

Requirements

  • Python 3.9 or newer
  • Python packages: numpy, pydub, nltk, g2p_en (see requirements.txt)
  • The bundled morshutalk_morshu.wav sample (included in this repo)

On first import, NLTK downloads a few small data packages (averaged_perceptron_tagger_eng, cmudict, punkt, punkt_tab).

Running

./start.sh

start.sh creates a local virtual environment on first launch, installs the dependencies from requirements.txt, and starts the server. The listening port defaults to 33002 and can be overridden with the PORT environment variable.

You can also run it directly against an existing interpreter that has the dependencies installed:

PORT=33002 python3 server.py

API

GET /say

Generate speech and return a WAV file.

GET /say?text=Hello%20it%20is%20me%20Morshu

Response: audio/wav (WAV bytes).

Example:

curl "http://127.0.0.1:33002/say?text=lamp+oil+rope+bombs" -o morshu.wav

GET /health

Returns ok.

Running as a system service

A systemd unit file is provided (morshu-tts-server.service). See INSTALL.md for setup steps.

S
Description
Standalone Morshu TTS HTTP server (extracted from the ttstoy Discord bot cog)
Readme 461 KiB
Languages
Python 97%
Shell 3%