diff --git a/README.md b/README.md index d942778..e046360 100644 --- a/README.md +++ b/README.md @@ -529,6 +529,527 @@ The `[p]autoroom allow` command now accepts multiple users/roles in a single com --- +--- +--- + +# TtsToy + +Multi-engine text-to-speech cog with voice cloning, emoji-triggered sound effects, inline voice switching, a TTS queue system, and a companion web UI. Plays audio in voice channels and uploads MP3 files to chat. + +### Engines + +| Engine | Description | Requirements | +|--------|-------------|--------------| +| MiniMax | Cloud API with high-quality AI voices | MiniMax API key | +| Chatterbox | Local AI voice cloning server | Chatterbox TTS server running | +| DECTalk | Classic robotic synthesizer (1980s style) | DECTalk server (bundled) | +| Morshu | Speech from Morshu's voice lines (CD-i Zelda) | g2p_en, numpy, pydub | +| VOX | Black Mesa/Half-Life announcer system | VOX word packs (bundled) | + +### Requirements + +- Python 3.11+ +- Red-DiscordBot 3.5.0+ +- Dependencies: `requests`, `g2p_en`, `numpy`, `pydub` +- `ffmpeg` installed and on PATH +- Red's Audio cog loaded (for voice channel playback) + +### Installation + +``` +[p]cog install scrapyard ttstoy +[p]load ttstoy +``` + +--- + +## User Commands + +--- + +### `[p]tts ` + +Generate TTS audio and play it in your voice channel. Also uploads the MP3 to chat. + +- You must be in a voice channel (except in Chatterbox mode, which can generate audio without VC). +- The bot will auto-join your voice channel if not already connected. +- Supports emoji SFX triggers inline with text. +- Supports inline voice/engine switching with `[mode|voice]` tags. +- Multiple TTS requests in the same guild are queued and played sequentially. +- If music is playing, it is paused during TTS and resumed after. + +**Inline voice switching:** +``` +[p]tts [minimax|Robotnik] Pingas [dectalk] [:nh]Deep voice [chatterbox|Emily] Hello! +``` + +Valid mode tags: `minimax`, `chatterbox`, `dectalk`, `morshu`, `vox` + +**Emoji SFX example:** +``` +[p]tts Hello everyone! 🎉 Welcome to the party 😂 +``` + +--- + +### `[p]ttstoy` + +Top-level configuration group. Running without a subcommand shows help. + +--- + +### `[p]ttstoy myvoice [voice_name]` + +Show or set your personal voice. + +- With no argument: displays your current voice and lists available voices for the active engine. +- With a voice name: sets your active voice. +- In Chatterbox mode, matches against your uploaded voice clips by display name. +- In DECTalk mode, voice selection is not used (use `[:np]`, `[:nb]`, etc. inline instead). +- In Morshu/VOX mode, no voice selection is available. + +**Example:** +``` +[p]ttstoy myvoice Robotnik +[p]ttstoy myvoice Emily +``` + +--- + +### `[p]ttstoy info` + +DMs you detailed usage instructions for TTS Toy, including setup steps (for owners) and end-user instructions. + +--- + +### `[p]ttstoy sfx` + +List all available emoji sound effects and their triggers. + +Shows each emoji/trigger mapped to its SFX folder, whether audio files are present, and usage examples. + +--- + +### `[p]ttstoy voxwords` + +List all available words in the current VOX pack. Only these words will be spoken in VOX mode; unknown words are skipped. + +--- + +### `[p]login` + +Shortcut for `[p]ttstoy login`. Gets a one-time login key for Kingston's Scrapyard web sites. The key is DM'd to you. + +--- + +### `[p]ttstoy login` + +Get a one-time login key for the TtsToy web UI and Kingston's Scrapyard homepage. Must be used in a server (the key is tied to your guild context). The key is DM'd to you and valid for 5 minutes. + +--- + +### `[p]ttstoy webui` + +Get the link to the TtsToy web UI. + +--- + +## Chatterbox Commands + +All under the `[p]chatterbox` group. + +--- + +### `[p]chatterbox guide` + +Full guide to Chatterbox TTS features, voice cloning workflow, special tokens, and per-voice tuning. + +--- + +### `[p]chatterbox addvoice [url]` + +Upload a voice clip for Chatterbox voice cloning. + +- Attach a `.wav` or `.mp3` file to the message, OR provide a direct URL to one. +- Recommended: 5-15 seconds of clear speech, one speaker, no background noise. +- `.wav` works best; `.mp3` accepted. +- Short clips are automatically looped to meet the minimum 5-second requirement. +- The voice is automatically set as your active voice after upload. + +**Examples:** +``` +[p]chatterbox addvoice CoolVoice +(with a .wav attached) + +[p]chatterbox addvoice CoolVoice https://example.com/clip.wav +``` + +--- + +### `[p]chatterbox removevoice ` + +Remove a voice from your personal voice list and delete it from the Chatterbox server (if you are the original uploader). Shared copies remain for other users. + +If the removed voice was your active voice, it switches to another voice in your list or the default. + +--- + +### `[p]chatterbox myvoices` + +List all your uploaded Chatterbox voices. Shows which one is currently active. Provides quick-reference commands for switching, adding, and removing voices. + +--- + +### `[p]chatterbox sharevoice @User` + +Share one of your Chatterbox voices with another user. Adds the voice to their library so they can use it with `[p]ttstoy myvoice`. + +--- + +### `[p]chatterbox temp [value]` + +Show or set the Chatterbox temperature for your current active voice. + +- Range: 0.0 - 1.5 +- Lower = more consistent speech, higher = more varied/random. +- Saved per voice per user (each voice remembers its own setting). +- Use `[p]chatterbox temp reset` to clear and use the server default. + +--- + +### `[p]chatterbox exag [value]` + +Show or set the Chatterbox exaggeration for your current active voice. + +- Range: 0.25 - 2.0 +- Higher = more expressive/dramatic delivery. +- Saved per voice per user. +- Use `[p]chatterbox exag reset` to clear. + +--- + +### `[p]chatterbox volume [value]` + +Show or set a volume offset (in dB) for your current Chatterbox voice. + +- Range: -20.0 to +20.0 dB +- Allows normalizing volume across different voice clips. +- Saved per voice per user. +- Use `[p]chatterbox volume reset` to clear. + +--- + +### `[p]chatterbox speed [value]` + +Show or set the playback speed for your current Chatterbox voice. + +- Range: 0.25 - 4.0 +- 1.0 = normal speed. Higher = faster, lower = slower. +- Saved per voice per user. +- Use `[p]chatterbox speed reset` to clear. + +--- + +### `[p]chatterbox reset` + +Reset all per-voice parameters (temperature, exaggeration, volume, speed) for your current active voice to server defaults. + +--- + +## Owner/Admin Commands + +--- + +### `[p]ttstoy mode [mode]` + +Show or set the global TTS engine mode. + +- `minimax` - MiniMax cloud API (requires API key) +- `chatterbox` - Local Chatterbox TTS server (AI voice cloning) +- `dectalk` - DECTalk classic robotic synthesizer +- `morshu` - MorshuTalk (Morshu's CD-i voice) +- `vox` - Black Mesa VOX announcer + +With no argument, shows the current mode and available modes with status info. + +--- + +### `[p]ttstoy key` + +Opens a button + modal dialog to securely enter the MiniMax API key. + +--- + +### `[p]ttstoy model [model_name]` + +Show or set the MiniMax TTS model. + +Available models: +- `speech-01-turbo` - fast and cheap +- `speech-01-hd` - higher quality, slower +- `speech-02-turbo` - newer fast model +- `speech-02-hd` - newer high quality + +With no argument, lists known models and the current setting. + +--- + +### `[p]ttstoy voice [voice_name]` + +Set or show the global default voice. + +- In MiniMax mode: choose from preset voices (BigMan, BlueGnome, Dracafow, Gaben, Grigori, Gnome, King, Peppa, Robotnik) or pass a raw voice ID. +- In Chatterbox mode: set a filename as the global default. +- DECTalk/Morshu/VOX modes do not use voice selection. + +--- + +### `[p]ttstoy sfxvolume [0-100]` + +Show or set the volume for emoji-triggered sound effects. Default is 100. + +--- + +### `[p]ttstoy voxpack [pack]` + +Show or set the VOX voice pack. + +Available packs: +- `vox` - Original Half-Life VOX +- `vox2` - Black Mesa military announcements + +--- + +### `[p]ttstoy accessibility` + +Toggle accessibility mode. + +When enabled: +- Forces DECTalk engine regardless of mode setting. +- Disables emoji SFX (emojis are stripped instead of triggering sound effects). + +--- + +### `[p]chatterbox url [url]` + +Show or set the Chatterbox TTS server URL. Default: `http://127.0.0.1:8099` + +Tests connectivity when showing or setting. + +--- + +### `[p]chatterbox model [turbo|original]` + +Show or switch the Chatterbox model (hot-swap). + +- `turbo` - Fast (350M params), supports special tokens like `[laugh]`, `[cough]`, etc. +- `original` - Better voice cloning (0.5B params), stronger emotion control. + +--- + +### `[p]ttstoy dectalkurl [url]` + +Show or set the DECTalk API server URL. Tests connectivity. + +--- + +### `[p]ttstoy dectalkinstall` + +Install the bundled DECTalk server (runs `npm install` in the dectalk-server directory). Requires Node.js and npm. + +--- + +### `[p]ttstoy dectalkstart` + +Manually start the bundled DECTalk server process. + +--- + +### `[p]ttstoy dectalkstop` + +Stop the bundled DECTalk server process. + +--- + +### `[p]ttstoy dectalkstatus` + +Show DECTalk server status: installed, process running, API responding, auto-start setting, and URL. + +--- + +### `[p]ttstoy dectalkautotoggle` + +Toggle auto-start for the bundled DECTalk server. When enabled, the server starts automatically when switching to DECTalk mode or on cog load. + +--- + +### `[p]ttstoy webuistatus` + +Check the status of the TtsToy web UI subprocess (running, reachable, URL, PID). + +--- + +### `[p]ttstoy webuirestart` + +Restart the TtsToy web UI subprocess. + +--- + +### `[p]ttstoy webuichannel [#channel]` + +Set or show the channel where web UI TTS audio is posted in the current server. If not set, falls back to any channel named `tts`, `ttstoy`, or `tts-toy`. + +--- + +## Emoji Sound Effects + +Include supported emoji in your TTS text to trigger sound effects. SFX are spliced inline between TTS segments. + +**Default emoji mappings:** + +| Emoji | Sound | Emoji | Sound | +|-------|-------|-------|-------| +| 🎉 | party | 😂 | laugh | +| 🥖 | spy | 👏 | clap | +| 🔥 | fire | 💀 | skull | +| ✅ | check | ❌ | error | +| 📢 | airhorn | 🚢 | boathorn | +| 😶 | drum | 👼 | angel | +| 🥜 | cashew | 💪 | physical | +| 🧠 | intelligence | 👁 | psychic | +| ✍️ | motor | | | + +Custom Discord emoji are also supported if their `:name:` is mapped in the emoji map. + +The emoji-to-SFX mapping is stored in `sfx/emoji_map.json` and can be edited via the web UI. + +--- + +## Inline Voice Switching + +You can switch engines and voices mid-sentence using `[mode|voice]` tags: + +``` +[p]tts [minimax|Robotnik] I am the Eggman! [dectalk] [:nh]Now in DECTalk [chatterbox|Emily] And now Chatterbox +``` + +- `[mode]` - switch engine only, keep current voice +- `[mode|voice]` - switch engine and voice +- Valid modes: `minimax`, `chatterbox`, `dectalk`, `morshu`, `vox` + +--- + +## DECTalk Voice Commands + +When in DECTalk mode, control voices inline in your text: + +| Command | Voice | +|---------|-------| +| `[:np]` | Perfect Paul (default, male) | +| `[:nb]` | Beautiful Betty (female) | +| `[:nh]` | Huge Harry (male) | +| `[:nf]` | Frail Frank (male) | +| `[:nd]` | Doctor Dennis (male) | +| `[:nk]` | Kit the Kid (child) | +| `[:nu]` | Uppity Ursula (female) | +| `[:nr]` | Rough Rita (female) | +| `[:nw]` | Whispering Wendy (female) | + +You can switch voices mid-sentence: `[p]tts [:nh]Deep voice [:nb]Now female` + +DECTalk also supports phoneme commands and singing. + +--- + +## Chatterbox Special Tokens + +When using the Chatterbox Turbo model, these tokens produce non-speech vocalizations: + +`[laugh]` `[chuckle]` `[sigh]` `[gasp]` `[cough]` `[clear throat]` `[sniff]` `[groan]` `[shush]` + +**Example:** +``` +[p]tts Hey [chuckle] thanks for calling back [laugh] +``` + +--- + +## VOX Engine + +The VOX engine concatenates pre-recorded word WAV files from the Black Mesa/Half-Life announcement system. Only words that exist in the VOX dictionary are spoken; unknown words are silently skipped. + +Two packs are available: +- `vox` - Original Half-Life VOX (500+ words) +- `vox2` - Black Mesa military announcements (250+ words) + +Use `[p]ttstoy voxwords` to see all available words for the current pack. + +--- + +## MorshuTalk Engine + +Generates speech by matching phonemes from input text to Morshu's voice samples from the CD-i Zelda games. Uses grapheme-to-phoneme conversion (g2p_en) to break text into phonemes, then concatenates matching audio snippets. + +No voice selection needed. All text is always spoken in Morshu's voice. + +--- + +## TTS Queue System + +- Each guild has its own TTS queue. +- Multiple `[p]tts` requests are queued and played in order. +- If music is playing, it is paused before TTS and resumed after all queued TTS finishes. +- Queue position is shown when multiple items are queued. + +--- + +## Web UI + +TtsToy includes a companion Flask web UI for managing voices, generating TTS, and controlling settings from a browser. + +- Login via `[p]ttstoy login` (DMs you a one-time key) +- Access at the configured URL (default: `https://ttstoy.kingstons-scrapyard.net`) +- Features: voice management, TTS generation, parameter tuning, SFX mapping editor +- TTS generated from the web UI is played in the user's current voice channel and posted to the configured Discord channel + +--- + +## Health Endpoint + +TtsToy exposes a health check HTTP endpoint at port 8097: +- `GET /health` returns `{"status": "ok", "service": "ttstoy"}` + +--- + +## MiniMax Voices (Presets) + +| Label | Description | +|-------|-------------| +| BigMan | - | +| BlueGnome | - | +| Dracafow | - | +| Gaben | - | +| Grigori | - | +| Gnome | Default voice | +| King | - | +| Peppa | - | +| Robotnik | - | + +Custom voice IDs can also be passed directly. + +--- + +## Credits + +- MiniMax TTS API: https://www.minimaxi.com/ +- Chatterbox TTS: https://github.com/resemble-ai/chatterbox +- DECTalk: Classic DEC speech synthesizer +- MorshuTalk engine by jalenluorion: https://github.com/jalenluorion/MorshuTalk +- VOX engine based on VOXGen by wphillips: https://github.com/wphillips/VOXGen +- Maintained by Scrapyard Cogworks + +--- + ## Support Visit us at https://homepage.kingstons-scrapyard.net/