Add full ttstoy documentation to README

This commit is contained in:
2026-06-05 16:45:43 -05:00
parent 32a4689ff7
commit eed02edc55
+521
View File
@@ -529,6 +529,527 @@ The `[p]autoroom allow` command now accepts multiple users/roles in a single com
---
---
---
# TtsToy
Multi-engine text-to-speech cog with voice cloning, emoji-triggered sound effects, inline voice switching, a TTS queue system, and a companion web UI. Plays audio in voice channels and uploads MP3 files to chat.
### Engines
| Engine | Description | Requirements |
|--------|-------------|--------------|
| MiniMax | Cloud API with high-quality AI voices | MiniMax API key |
| Chatterbox | Local AI voice cloning server | Chatterbox TTS server running |
| DECTalk | Classic robotic synthesizer (1980s style) | DECTalk server (bundled) |
| Morshu | Speech from Morshu's voice lines (CD-i Zelda) | g2p_en, numpy, pydub |
| VOX | Black Mesa/Half-Life announcer system | VOX word packs (bundled) |
### Requirements
- Python 3.11+
- Red-DiscordBot 3.5.0+
- Dependencies: `requests`, `g2p_en`, `numpy`, `pydub`
- `ffmpeg` installed and on PATH
- Red's Audio cog loaded (for voice channel playback)
### Installation
```
[p]cog install scrapyard ttstoy
[p]load ttstoy
```
---
## User Commands
---
### `[p]tts <text>`
Generate TTS audio and play it in your voice channel. Also uploads the MP3 to chat.
- You must be in a voice channel (except in Chatterbox mode, which can generate audio without VC).
- The bot will auto-join your voice channel if not already connected.
- Supports emoji SFX triggers inline with text.
- Supports inline voice/engine switching with `[mode|voice]` tags.
- Multiple TTS requests in the same guild are queued and played sequentially.
- If music is playing, it is paused during TTS and resumed after.
**Inline voice switching:**
```
[p]tts [minimax|Robotnik] Pingas [dectalk] [:nh]Deep voice [chatterbox|Emily] Hello!
```
Valid mode tags: `minimax`, `chatterbox`, `dectalk`, `morshu`, `vox`
**Emoji SFX example:**
```
[p]tts Hello everyone! 🎉 Welcome to the party 😂
```
---
### `[p]ttstoy`
Top-level configuration group. Running without a subcommand shows help.
---
### `[p]ttstoy myvoice [voice_name]`
Show or set your personal voice.
- With no argument: displays your current voice and lists available voices for the active engine.
- With a voice name: sets your active voice.
- In Chatterbox mode, matches against your uploaded voice clips by display name.
- In DECTalk mode, voice selection is not used (use `[:np]`, `[:nb]`, etc. inline instead).
- In Morshu/VOX mode, no voice selection is available.
**Example:**
```
[p]ttstoy myvoice Robotnik
[p]ttstoy myvoice Emily
```
---
### `[p]ttstoy info`
DMs you detailed usage instructions for TTS Toy, including setup steps (for owners) and end-user instructions.
---
### `[p]ttstoy sfx`
List all available emoji sound effects and their triggers.
Shows each emoji/trigger mapped to its SFX folder, whether audio files are present, and usage examples.
---
### `[p]ttstoy voxwords`
List all available words in the current VOX pack. Only these words will be spoken in VOX mode; unknown words are skipped.
---
### `[p]login`
Shortcut for `[p]ttstoy login`. Gets a one-time login key for Kingston's Scrapyard web sites. The key is DM'd to you.
---
### `[p]ttstoy login`
Get a one-time login key for the TtsToy web UI and Kingston's Scrapyard homepage. Must be used in a server (the key is tied to your guild context). The key is DM'd to you and valid for 5 minutes.
---
### `[p]ttstoy webui`
Get the link to the TtsToy web UI.
---
## Chatterbox Commands
All under the `[p]chatterbox` group.
---
### `[p]chatterbox guide`
Full guide to Chatterbox TTS features, voice cloning workflow, special tokens, and per-voice tuning.
---
### `[p]chatterbox addvoice <name> [url]`
Upload a voice clip for Chatterbox voice cloning.
- Attach a `.wav` or `.mp3` file to the message, OR provide a direct URL to one.
- Recommended: 5-15 seconds of clear speech, one speaker, no background noise.
- `.wav` works best; `.mp3` accepted.
- Short clips are automatically looped to meet the minimum 5-second requirement.
- The voice is automatically set as your active voice after upload.
**Examples:**
```
[p]chatterbox addvoice CoolVoice
(with a .wav attached)
[p]chatterbox addvoice CoolVoice https://example.com/clip.wav
```
---
### `[p]chatterbox removevoice <name>`
Remove a voice from your personal voice list and delete it from the Chatterbox server (if you are the original uploader). Shared copies remain for other users.
If the removed voice was your active voice, it switches to another voice in your list or the default.
---
### `[p]chatterbox myvoices`
List all your uploaded Chatterbox voices. Shows which one is currently active. Provides quick-reference commands for switching, adding, and removing voices.
---
### `[p]chatterbox sharevoice <name> @User`
Share one of your Chatterbox voices with another user. Adds the voice to their library so they can use it with `[p]ttstoy myvoice`.
---
### `[p]chatterbox temp [value]`
Show or set the Chatterbox temperature for your current active voice.
- Range: 0.0 - 1.5
- Lower = more consistent speech, higher = more varied/random.
- Saved per voice per user (each voice remembers its own setting).
- Use `[p]chatterbox temp reset` to clear and use the server default.
---
### `[p]chatterbox exag [value]`
Show or set the Chatterbox exaggeration for your current active voice.
- Range: 0.25 - 2.0
- Higher = more expressive/dramatic delivery.
- Saved per voice per user.
- Use `[p]chatterbox exag reset` to clear.
---
### `[p]chatterbox volume [value]`
Show or set a volume offset (in dB) for your current Chatterbox voice.
- Range: -20.0 to +20.0 dB
- Allows normalizing volume across different voice clips.
- Saved per voice per user.
- Use `[p]chatterbox volume reset` to clear.
---
### `[p]chatterbox speed [value]`
Show or set the playback speed for your current Chatterbox voice.
- Range: 0.25 - 4.0
- 1.0 = normal speed. Higher = faster, lower = slower.
- Saved per voice per user.
- Use `[p]chatterbox speed reset` to clear.
---
### `[p]chatterbox reset`
Reset all per-voice parameters (temperature, exaggeration, volume, speed) for your current active voice to server defaults.
---
## Owner/Admin Commands
---
### `[p]ttstoy mode [mode]`
Show or set the global TTS engine mode.
- `minimax` - MiniMax cloud API (requires API key)
- `chatterbox` - Local Chatterbox TTS server (AI voice cloning)
- `dectalk` - DECTalk classic robotic synthesizer
- `morshu` - MorshuTalk (Morshu's CD-i voice)
- `vox` - Black Mesa VOX announcer
With no argument, shows the current mode and available modes with status info.
---
### `[p]ttstoy key`
Opens a button + modal dialog to securely enter the MiniMax API key.
---
### `[p]ttstoy model [model_name]`
Show or set the MiniMax TTS model.
Available models:
- `speech-01-turbo` - fast and cheap
- `speech-01-hd` - higher quality, slower
- `speech-02-turbo` - newer fast model
- `speech-02-hd` - newer high quality
With no argument, lists known models and the current setting.
---
### `[p]ttstoy voice [voice_name]`
Set or show the global default voice.
- In MiniMax mode: choose from preset voices (BigMan, BlueGnome, Dracafow, Gaben, Grigori, Gnome, King, Peppa, Robotnik) or pass a raw voice ID.
- In Chatterbox mode: set a filename as the global default.
- DECTalk/Morshu/VOX modes do not use voice selection.
---
### `[p]ttstoy sfxvolume [0-100]`
Show or set the volume for emoji-triggered sound effects. Default is 100.
---
### `[p]ttstoy voxpack [pack]`
Show or set the VOX voice pack.
Available packs:
- `vox` - Original Half-Life VOX
- `vox2` - Black Mesa military announcements
---
### `[p]ttstoy accessibility`
Toggle accessibility mode.
When enabled:
- Forces DECTalk engine regardless of mode setting.
- Disables emoji SFX (emojis are stripped instead of triggering sound effects).
---
### `[p]chatterbox url [url]`
Show or set the Chatterbox TTS server URL. Default: `http://127.0.0.1:8099`
Tests connectivity when showing or setting.
---
### `[p]chatterbox model [turbo|original]`
Show or switch the Chatterbox model (hot-swap).
- `turbo` - Fast (350M params), supports special tokens like `[laugh]`, `[cough]`, etc.
- `original` - Better voice cloning (0.5B params), stronger emotion control.
---
### `[p]ttstoy dectalkurl [url]`
Show or set the DECTalk API server URL. Tests connectivity.
---
### `[p]ttstoy dectalkinstall`
Install the bundled DECTalk server (runs `npm install` in the dectalk-server directory). Requires Node.js and npm.
---
### `[p]ttstoy dectalkstart`
Manually start the bundled DECTalk server process.
---
### `[p]ttstoy dectalkstop`
Stop the bundled DECTalk server process.
---
### `[p]ttstoy dectalkstatus`
Show DECTalk server status: installed, process running, API responding, auto-start setting, and URL.
---
### `[p]ttstoy dectalkautotoggle`
Toggle auto-start for the bundled DECTalk server. When enabled, the server starts automatically when switching to DECTalk mode or on cog load.
---
### `[p]ttstoy webuistatus`
Check the status of the TtsToy web UI subprocess (running, reachable, URL, PID).
---
### `[p]ttstoy webuirestart`
Restart the TtsToy web UI subprocess.
---
### `[p]ttstoy webuichannel [#channel]`
Set or show the channel where web UI TTS audio is posted in the current server. If not set, falls back to any channel named `tts`, `ttstoy`, or `tts-toy`.
---
## Emoji Sound Effects
Include supported emoji in your TTS text to trigger sound effects. SFX are spliced inline between TTS segments.
**Default emoji mappings:**
| Emoji | Sound | Emoji | Sound |
|-------|-------|-------|-------|
| 🎉 | party | 😂 | laugh |
| 🥖 | spy | 👏 | clap |
| 🔥 | fire | 💀 | skull |
| ✅ | check | ❌ | error |
| 📢 | airhorn | 🚢 | boathorn |
| 😶 | drum | 👼 | angel |
| 🥜 | cashew | 💪 | physical |
| 🧠 | intelligence | 👁 | psychic |
| ✍️ | motor | | |
Custom Discord emoji are also supported if their `:name:` is mapped in the emoji map.
The emoji-to-SFX mapping is stored in `sfx/emoji_map.json` and can be edited via the web UI.
---
## Inline Voice Switching
You can switch engines and voices mid-sentence using `[mode|voice]` tags:
```
[p]tts [minimax|Robotnik] I am the Eggman! [dectalk] [:nh]Now in DECTalk [chatterbox|Emily] And now Chatterbox
```
- `[mode]` - switch engine only, keep current voice
- `[mode|voice]` - switch engine and voice
- Valid modes: `minimax`, `chatterbox`, `dectalk`, `morshu`, `vox`
---
## DECTalk Voice Commands
When in DECTalk mode, control voices inline in your text:
| Command | Voice |
|---------|-------|
| `[:np]` | Perfect Paul (default, male) |
| `[:nb]` | Beautiful Betty (female) |
| `[:nh]` | Huge Harry (male) |
| `[:nf]` | Frail Frank (male) |
| `[:nd]` | Doctor Dennis (male) |
| `[:nk]` | Kit the Kid (child) |
| `[:nu]` | Uppity Ursula (female) |
| `[:nr]` | Rough Rita (female) |
| `[:nw]` | Whispering Wendy (female) |
You can switch voices mid-sentence: `[p]tts [:nh]Deep voice [:nb]Now female`
DECTalk also supports phoneme commands and singing.
---
## Chatterbox Special Tokens
When using the Chatterbox Turbo model, these tokens produce non-speech vocalizations:
`[laugh]` `[chuckle]` `[sigh]` `[gasp]` `[cough]` `[clear throat]` `[sniff]` `[groan]` `[shush]`
**Example:**
```
[p]tts Hey [chuckle] thanks for calling back [laugh]
```
---
## VOX Engine
The VOX engine concatenates pre-recorded word WAV files from the Black Mesa/Half-Life announcement system. Only words that exist in the VOX dictionary are spoken; unknown words are silently skipped.
Two packs are available:
- `vox` - Original Half-Life VOX (500+ words)
- `vox2` - Black Mesa military announcements (250+ words)
Use `[p]ttstoy voxwords` to see all available words for the current pack.
---
## MorshuTalk Engine
Generates speech by matching phonemes from input text to Morshu's voice samples from the CD-i Zelda games. Uses grapheme-to-phoneme conversion (g2p_en) to break text into phonemes, then concatenates matching audio snippets.
No voice selection needed. All text is always spoken in Morshu's voice.
---
## TTS Queue System
- Each guild has its own TTS queue.
- Multiple `[p]tts` requests are queued and played in order.
- If music is playing, it is paused before TTS and resumed after all queued TTS finishes.
- Queue position is shown when multiple items are queued.
---
## Web UI
TtsToy includes a companion Flask web UI for managing voices, generating TTS, and controlling settings from a browser.
- Login via `[p]ttstoy login` (DMs you a one-time key)
- Access at the configured URL (default: `https://ttstoy.kingstons-scrapyard.net`)
- Features: voice management, TTS generation, parameter tuning, SFX mapping editor
- TTS generated from the web UI is played in the user's current voice channel and posted to the configured Discord channel
---
## Health Endpoint
TtsToy exposes a health check HTTP endpoint at port 8097:
- `GET /health` returns `{"status": "ok", "service": "ttstoy"}`
---
## MiniMax Voices (Presets)
| Label | Description |
|-------|-------------|
| BigMan | - |
| BlueGnome | - |
| Dracafow | - |
| Gaben | - |
| Grigori | - |
| Gnome | Default voice |
| King | - |
| Peppa | - |
| Robotnik | - |
Custom voice IDs can also be passed directly.
---
## Credits
- MiniMax TTS API: https://www.minimaxi.com/
- Chatterbox TTS: https://github.com/resemble-ai/chatterbox
- DECTalk: Classic DEC speech synthesizer
- MorshuTalk engine by jalenluorion: https://github.com/jalenluorion/MorshuTalk
- VOX engine based on VOXGen by wphillips: https://github.com/wphillips/VOXGen
- Maintained by Scrapyard Cogworks
---
## Support
Visit us at https://homepage.kingstons-scrapyard.net/