Add full ttstoy documentation to README
This commit is contained in:
@@ -529,6 +529,527 @@ The `[p]autoroom allow` command now accepts multiple users/roles in a single com
|
||||
|
||||
---
|
||||
|
||||
---
|
||||
---
|
||||
|
||||
# TtsToy
|
||||
|
||||
Multi-engine text-to-speech cog with voice cloning, emoji-triggered sound effects, inline voice switching, a TTS queue system, and a companion web UI. Plays audio in voice channels and uploads MP3 files to chat.
|
||||
|
||||
### Engines
|
||||
|
||||
| Engine | Description | Requirements |
|
||||
|--------|-------------|--------------|
|
||||
| MiniMax | Cloud API with high-quality AI voices | MiniMax API key |
|
||||
| Chatterbox | Local AI voice cloning server | Chatterbox TTS server running |
|
||||
| DECTalk | Classic robotic synthesizer (1980s style) | DECTalk server (bundled) |
|
||||
| Morshu | Speech from Morshu's voice lines (CD-i Zelda) | g2p_en, numpy, pydub |
|
||||
| VOX | Black Mesa/Half-Life announcer system | VOX word packs (bundled) |
|
||||
|
||||
### Requirements
|
||||
|
||||
- Python 3.11+
|
||||
- Red-DiscordBot 3.5.0+
|
||||
- Dependencies: `requests`, `g2p_en`, `numpy`, `pydub`
|
||||
- `ffmpeg` installed and on PATH
|
||||
- Red's Audio cog loaded (for voice channel playback)
|
||||
|
||||
### Installation
|
||||
|
||||
```
|
||||
[p]cog install scrapyard ttstoy
|
||||
[p]load ttstoy
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## User Commands
|
||||
|
||||
---
|
||||
|
||||
### `[p]tts <text>`
|
||||
|
||||
Generate TTS audio and play it in your voice channel. Also uploads the MP3 to chat.
|
||||
|
||||
- You must be in a voice channel (except in Chatterbox mode, which can generate audio without VC).
|
||||
- The bot will auto-join your voice channel if not already connected.
|
||||
- Supports emoji SFX triggers inline with text.
|
||||
- Supports inline voice/engine switching with `[mode|voice]` tags.
|
||||
- Multiple TTS requests in the same guild are queued and played sequentially.
|
||||
- If music is playing, it is paused during TTS and resumed after.
|
||||
|
||||
**Inline voice switching:**
|
||||
```
|
||||
[p]tts [minimax|Robotnik] Pingas [dectalk] [:nh]Deep voice [chatterbox|Emily] Hello!
|
||||
```
|
||||
|
||||
Valid mode tags: `minimax`, `chatterbox`, `dectalk`, `morshu`, `vox`
|
||||
|
||||
**Emoji SFX example:**
|
||||
```
|
||||
[p]tts Hello everyone! 🎉 Welcome to the party 😂
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy`
|
||||
|
||||
Top-level configuration group. Running without a subcommand shows help.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy myvoice [voice_name]`
|
||||
|
||||
Show or set your personal voice.
|
||||
|
||||
- With no argument: displays your current voice and lists available voices for the active engine.
|
||||
- With a voice name: sets your active voice.
|
||||
- In Chatterbox mode, matches against your uploaded voice clips by display name.
|
||||
- In DECTalk mode, voice selection is not used (use `[:np]`, `[:nb]`, etc. inline instead).
|
||||
- In Morshu/VOX mode, no voice selection is available.
|
||||
|
||||
**Example:**
|
||||
```
|
||||
[p]ttstoy myvoice Robotnik
|
||||
[p]ttstoy myvoice Emily
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy info`
|
||||
|
||||
DMs you detailed usage instructions for TTS Toy, including setup steps (for owners) and end-user instructions.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy sfx`
|
||||
|
||||
List all available emoji sound effects and their triggers.
|
||||
|
||||
Shows each emoji/trigger mapped to its SFX folder, whether audio files are present, and usage examples.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy voxwords`
|
||||
|
||||
List all available words in the current VOX pack. Only these words will be spoken in VOX mode; unknown words are skipped.
|
||||
|
||||
---
|
||||
|
||||
### `[p]login`
|
||||
|
||||
Shortcut for `[p]ttstoy login`. Gets a one-time login key for Kingston's Scrapyard web sites. The key is DM'd to you.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy login`
|
||||
|
||||
Get a one-time login key for the TtsToy web UI and Kingston's Scrapyard homepage. Must be used in a server (the key is tied to your guild context). The key is DM'd to you and valid for 5 minutes.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy webui`
|
||||
|
||||
Get the link to the TtsToy web UI.
|
||||
|
||||
---
|
||||
|
||||
## Chatterbox Commands
|
||||
|
||||
All under the `[p]chatterbox` group.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox guide`
|
||||
|
||||
Full guide to Chatterbox TTS features, voice cloning workflow, special tokens, and per-voice tuning.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox addvoice <name> [url]`
|
||||
|
||||
Upload a voice clip for Chatterbox voice cloning.
|
||||
|
||||
- Attach a `.wav` or `.mp3` file to the message, OR provide a direct URL to one.
|
||||
- Recommended: 5-15 seconds of clear speech, one speaker, no background noise.
|
||||
- `.wav` works best; `.mp3` accepted.
|
||||
- Short clips are automatically looped to meet the minimum 5-second requirement.
|
||||
- The voice is automatically set as your active voice after upload.
|
||||
|
||||
**Examples:**
|
||||
```
|
||||
[p]chatterbox addvoice CoolVoice
|
||||
(with a .wav attached)
|
||||
|
||||
[p]chatterbox addvoice CoolVoice https://example.com/clip.wav
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox removevoice <name>`
|
||||
|
||||
Remove a voice from your personal voice list and delete it from the Chatterbox server (if you are the original uploader). Shared copies remain for other users.
|
||||
|
||||
If the removed voice was your active voice, it switches to another voice in your list or the default.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox myvoices`
|
||||
|
||||
List all your uploaded Chatterbox voices. Shows which one is currently active. Provides quick-reference commands for switching, adding, and removing voices.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox sharevoice <name> @User`
|
||||
|
||||
Share one of your Chatterbox voices with another user. Adds the voice to their library so they can use it with `[p]ttstoy myvoice`.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox temp [value]`
|
||||
|
||||
Show or set the Chatterbox temperature for your current active voice.
|
||||
|
||||
- Range: 0.0 - 1.5
|
||||
- Lower = more consistent speech, higher = more varied/random.
|
||||
- Saved per voice per user (each voice remembers its own setting).
|
||||
- Use `[p]chatterbox temp reset` to clear and use the server default.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox exag [value]`
|
||||
|
||||
Show or set the Chatterbox exaggeration for your current active voice.
|
||||
|
||||
- Range: 0.25 - 2.0
|
||||
- Higher = more expressive/dramatic delivery.
|
||||
- Saved per voice per user.
|
||||
- Use `[p]chatterbox exag reset` to clear.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox volume [value]`
|
||||
|
||||
Show or set a volume offset (in dB) for your current Chatterbox voice.
|
||||
|
||||
- Range: -20.0 to +20.0 dB
|
||||
- Allows normalizing volume across different voice clips.
|
||||
- Saved per voice per user.
|
||||
- Use `[p]chatterbox volume reset` to clear.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox speed [value]`
|
||||
|
||||
Show or set the playback speed for your current Chatterbox voice.
|
||||
|
||||
- Range: 0.25 - 4.0
|
||||
- 1.0 = normal speed. Higher = faster, lower = slower.
|
||||
- Saved per voice per user.
|
||||
- Use `[p]chatterbox speed reset` to clear.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox reset`
|
||||
|
||||
Reset all per-voice parameters (temperature, exaggeration, volume, speed) for your current active voice to server defaults.
|
||||
|
||||
---
|
||||
|
||||
## Owner/Admin Commands
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy mode [mode]`
|
||||
|
||||
Show or set the global TTS engine mode.
|
||||
|
||||
- `minimax` - MiniMax cloud API (requires API key)
|
||||
- `chatterbox` - Local Chatterbox TTS server (AI voice cloning)
|
||||
- `dectalk` - DECTalk classic robotic synthesizer
|
||||
- `morshu` - MorshuTalk (Morshu's CD-i voice)
|
||||
- `vox` - Black Mesa VOX announcer
|
||||
|
||||
With no argument, shows the current mode and available modes with status info.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy key`
|
||||
|
||||
Opens a button + modal dialog to securely enter the MiniMax API key.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy model [model_name]`
|
||||
|
||||
Show or set the MiniMax TTS model.
|
||||
|
||||
Available models:
|
||||
- `speech-01-turbo` - fast and cheap
|
||||
- `speech-01-hd` - higher quality, slower
|
||||
- `speech-02-turbo` - newer fast model
|
||||
- `speech-02-hd` - newer high quality
|
||||
|
||||
With no argument, lists known models and the current setting.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy voice [voice_name]`
|
||||
|
||||
Set or show the global default voice.
|
||||
|
||||
- In MiniMax mode: choose from preset voices (BigMan, BlueGnome, Dracafow, Gaben, Grigori, Gnome, King, Peppa, Robotnik) or pass a raw voice ID.
|
||||
- In Chatterbox mode: set a filename as the global default.
|
||||
- DECTalk/Morshu/VOX modes do not use voice selection.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy sfxvolume [0-100]`
|
||||
|
||||
Show or set the volume for emoji-triggered sound effects. Default is 100.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy voxpack [pack]`
|
||||
|
||||
Show or set the VOX voice pack.
|
||||
|
||||
Available packs:
|
||||
- `vox` - Original Half-Life VOX
|
||||
- `vox2` - Black Mesa military announcements
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy accessibility`
|
||||
|
||||
Toggle accessibility mode.
|
||||
|
||||
When enabled:
|
||||
- Forces DECTalk engine regardless of mode setting.
|
||||
- Disables emoji SFX (emojis are stripped instead of triggering sound effects).
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox url [url]`
|
||||
|
||||
Show or set the Chatterbox TTS server URL. Default: `http://127.0.0.1:8099`
|
||||
|
||||
Tests connectivity when showing or setting.
|
||||
|
||||
---
|
||||
|
||||
### `[p]chatterbox model [turbo|original]`
|
||||
|
||||
Show or switch the Chatterbox model (hot-swap).
|
||||
|
||||
- `turbo` - Fast (350M params), supports special tokens like `[laugh]`, `[cough]`, etc.
|
||||
- `original` - Better voice cloning (0.5B params), stronger emotion control.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy dectalkurl [url]`
|
||||
|
||||
Show or set the DECTalk API server URL. Tests connectivity.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy dectalkinstall`
|
||||
|
||||
Install the bundled DECTalk server (runs `npm install` in the dectalk-server directory). Requires Node.js and npm.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy dectalkstart`
|
||||
|
||||
Manually start the bundled DECTalk server process.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy dectalkstop`
|
||||
|
||||
Stop the bundled DECTalk server process.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy dectalkstatus`
|
||||
|
||||
Show DECTalk server status: installed, process running, API responding, auto-start setting, and URL.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy dectalkautotoggle`
|
||||
|
||||
Toggle auto-start for the bundled DECTalk server. When enabled, the server starts automatically when switching to DECTalk mode or on cog load.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy webuistatus`
|
||||
|
||||
Check the status of the TtsToy web UI subprocess (running, reachable, URL, PID).
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy webuirestart`
|
||||
|
||||
Restart the TtsToy web UI subprocess.
|
||||
|
||||
---
|
||||
|
||||
### `[p]ttstoy webuichannel [#channel]`
|
||||
|
||||
Set or show the channel where web UI TTS audio is posted in the current server. If not set, falls back to any channel named `tts`, `ttstoy`, or `tts-toy`.
|
||||
|
||||
---
|
||||
|
||||
## Emoji Sound Effects
|
||||
|
||||
Include supported emoji in your TTS text to trigger sound effects. SFX are spliced inline between TTS segments.
|
||||
|
||||
**Default emoji mappings:**
|
||||
|
||||
| Emoji | Sound | Emoji | Sound |
|
||||
|-------|-------|-------|-------|
|
||||
| 🎉 | party | 😂 | laugh |
|
||||
| 🥖 | spy | 👏 | clap |
|
||||
| 🔥 | fire | 💀 | skull |
|
||||
| ✅ | check | ❌ | error |
|
||||
| 📢 | airhorn | 🚢 | boathorn |
|
||||
| 😶 | drum | 👼 | angel |
|
||||
| 🥜 | cashew | 💪 | physical |
|
||||
| 🧠 | intelligence | 👁 | psychic |
|
||||
| ✍️ | motor | | |
|
||||
|
||||
Custom Discord emoji are also supported if their `:name:` is mapped in the emoji map.
|
||||
|
||||
The emoji-to-SFX mapping is stored in `sfx/emoji_map.json` and can be edited via the web UI.
|
||||
|
||||
---
|
||||
|
||||
## Inline Voice Switching
|
||||
|
||||
You can switch engines and voices mid-sentence using `[mode|voice]` tags:
|
||||
|
||||
```
|
||||
[p]tts [minimax|Robotnik] I am the Eggman! [dectalk] [:nh]Now in DECTalk [chatterbox|Emily] And now Chatterbox
|
||||
```
|
||||
|
||||
- `[mode]` - switch engine only, keep current voice
|
||||
- `[mode|voice]` - switch engine and voice
|
||||
- Valid modes: `minimax`, `chatterbox`, `dectalk`, `morshu`, `vox`
|
||||
|
||||
---
|
||||
|
||||
## DECTalk Voice Commands
|
||||
|
||||
When in DECTalk mode, control voices inline in your text:
|
||||
|
||||
| Command | Voice |
|
||||
|---------|-------|
|
||||
| `[:np]` | Perfect Paul (default, male) |
|
||||
| `[:nb]` | Beautiful Betty (female) |
|
||||
| `[:nh]` | Huge Harry (male) |
|
||||
| `[:nf]` | Frail Frank (male) |
|
||||
| `[:nd]` | Doctor Dennis (male) |
|
||||
| `[:nk]` | Kit the Kid (child) |
|
||||
| `[:nu]` | Uppity Ursula (female) |
|
||||
| `[:nr]` | Rough Rita (female) |
|
||||
| `[:nw]` | Whispering Wendy (female) |
|
||||
|
||||
You can switch voices mid-sentence: `[p]tts [:nh]Deep voice [:nb]Now female`
|
||||
|
||||
DECTalk also supports phoneme commands and singing.
|
||||
|
||||
---
|
||||
|
||||
## Chatterbox Special Tokens
|
||||
|
||||
When using the Chatterbox Turbo model, these tokens produce non-speech vocalizations:
|
||||
|
||||
`[laugh]` `[chuckle]` `[sigh]` `[gasp]` `[cough]` `[clear throat]` `[sniff]` `[groan]` `[shush]`
|
||||
|
||||
**Example:**
|
||||
```
|
||||
[p]tts Hey [chuckle] thanks for calling back [laugh]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## VOX Engine
|
||||
|
||||
The VOX engine concatenates pre-recorded word WAV files from the Black Mesa/Half-Life announcement system. Only words that exist in the VOX dictionary are spoken; unknown words are silently skipped.
|
||||
|
||||
Two packs are available:
|
||||
- `vox` - Original Half-Life VOX (500+ words)
|
||||
- `vox2` - Black Mesa military announcements (250+ words)
|
||||
|
||||
Use `[p]ttstoy voxwords` to see all available words for the current pack.
|
||||
|
||||
---
|
||||
|
||||
## MorshuTalk Engine
|
||||
|
||||
Generates speech by matching phonemes from input text to Morshu's voice samples from the CD-i Zelda games. Uses grapheme-to-phoneme conversion (g2p_en) to break text into phonemes, then concatenates matching audio snippets.
|
||||
|
||||
No voice selection needed. All text is always spoken in Morshu's voice.
|
||||
|
||||
---
|
||||
|
||||
## TTS Queue System
|
||||
|
||||
- Each guild has its own TTS queue.
|
||||
- Multiple `[p]tts` requests are queued and played in order.
|
||||
- If music is playing, it is paused before TTS and resumed after all queued TTS finishes.
|
||||
- Queue position is shown when multiple items are queued.
|
||||
|
||||
---
|
||||
|
||||
## Web UI
|
||||
|
||||
TtsToy includes a companion Flask web UI for managing voices, generating TTS, and controlling settings from a browser.
|
||||
|
||||
- Login via `[p]ttstoy login` (DMs you a one-time key)
|
||||
- Access at the configured URL (default: `https://ttstoy.kingstons-scrapyard.net`)
|
||||
- Features: voice management, TTS generation, parameter tuning, SFX mapping editor
|
||||
- TTS generated from the web UI is played in the user's current voice channel and posted to the configured Discord channel
|
||||
|
||||
---
|
||||
|
||||
## Health Endpoint
|
||||
|
||||
TtsToy exposes a health check HTTP endpoint at port 8097:
|
||||
- `GET /health` returns `{"status": "ok", "service": "ttstoy"}`
|
||||
|
||||
---
|
||||
|
||||
## MiniMax Voices (Presets)
|
||||
|
||||
| Label | Description |
|
||||
|-------|-------------|
|
||||
| BigMan | - |
|
||||
| BlueGnome | - |
|
||||
| Dracafow | - |
|
||||
| Gaben | - |
|
||||
| Grigori | - |
|
||||
| Gnome | Default voice |
|
||||
| King | - |
|
||||
| Peppa | - |
|
||||
| Robotnik | - |
|
||||
|
||||
Custom voice IDs can also be passed directly.
|
||||
|
||||
---
|
||||
|
||||
## Credits
|
||||
|
||||
- MiniMax TTS API: https://www.minimaxi.com/
|
||||
- Chatterbox TTS: https://github.com/resemble-ai/chatterbox
|
||||
- DECTalk: Classic DEC speech synthesizer
|
||||
- MorshuTalk engine by jalenluorion: https://github.com/jalenluorion/MorshuTalk
|
||||
- VOX engine based on VOXGen by wphillips: https://github.com/wphillips/VOXGen
|
||||
- Maintained by Scrapyard Cogworks
|
||||
|
||||
---
|
||||
|
||||
## Support
|
||||
|
||||
Visit us at https://homepage.kingstons-scrapyard.net/
|
||||
|
||||
Reference in New Issue
Block a user