Add ttstoy cog

This commit is contained in:
2026-06-05 16:41:55 -05:00
parent 0e9f4a5224
commit 6739b2d6f1
975 changed files with 8216 additions and 0 deletions
+93
View File
@@ -0,0 +1,93 @@
# TTS Toy
A RedBot cog for text-to-speech in voice chat with multiple TTS engines, emoji sound effects, and voice cloning.
## Modes
| Mode | Description | Requires |
|------|-------------|----------|
| `minimax` | MiniMax cloud API, natural voices | API key |
| `chatterbox` | Local Chatterbox server, AI voice cloning | Chatterbox server |
| `dectalk` | Classic 1980s robotic voice | DECTalk server |
| `morshu` | Morshu from CD-i Zelda | g2p_en, numpy, pydub |
| `vox` | Black Mesa VOX announcer | Built-in |
Switch modes: `[p]ttstoy mode <mode>`
## Quick Start
```
[p]ttstoy mode chatterbox
[p]chatterbox addvoice MyVoice (attach .wav/.mp3 or paste a URL)
[p]tts Hello world [laugh] this is cool
```
## Commands
Use `[p]ttstoy` to see all available commands for the current mode.
### Chatterbox Voice Management
| Command | Description |
|---------|-------------|
| `[p]chatterbox addvoice <name>` | Upload a voice clip (attach file or paste URL) |
| `[p]chatterbox removevoice <name>` | Remove a voice and delete from server |
| `[p]chatterbox sharevoice <name> @User` | Share a voice with another user |
| `[p]chatterbox myvoices` | List your voices |
| `[p]chatterbox temp [0.0-1.5]` | Per-voice randomness (per voice per user) |
| `[p]chatterbox exag [0.25-2.0]` | Per-voice expressiveness (per voice per user) |
| `[p]chatterbox volume [-20-20]` | Per-voice dB loudness offset (per voice per user) |
| `[p]chatterbox guide` | Full feature guide |
| `[p]chatterbox model [turbo\|original]` | Switch Chatterbox model (owner) |
| `[p]chatterbox url [url]` | Set Chatterbox server URL (owner) |
Append `reset` to temp, exag, or volume to clear the setting.
### General
| Command | Description |
|---------|-------------|
| `[p]tts <text>` | Speak in voice chat |
| `[p]ttstoy myvoice [name]` | Set or view your personal voice |
| `[p]ttstoy voice [name]` | Set global default voice (owner) |
| `[p]ttstoy mode [mode]` | Show or set TTS mode (owner) |
| `[p]ttstoy sfxvolume [0-100]` | Set SFX volume (owner) |
| `[p]sfx` | List available sound effects |
| `[p]ttstoy info` | DM setup instructions |
### DECTalk
| Command | Description |
|---------|-------------|
| `[p]ttstoy dectalkinstall` | Install built-in DECTalk server |
| `[p]ttstoy dectalkstart` | Start server |
| `[p]ttstoy dectalkstop` | Stop server |
| `[p]ttstoy dectalkstatus` | Check server status |
| `[p]ttstoy dectalkautotoggle` | Toggle auto-start |
| `[p]ttstoy dectalkurl [url]` | Set DECTalk API URL |
DECTalk voices are controlled in-text: `[:np]` Paul, `[:nb]` Betty, `[:nh]` Harry, `[:nf]` Frank, `[:nd]` Dennis, `[:nk]` Kit, `[:nu]` Ursula, `[:nr]` Rita, `[:nw]` Wendy.
### VOX
| Command | Description |
|---------|-------------|
| `[p]ttstoy voxpack [vox\|vox2]` | Switch VOX voice pack |
| `[p]ttstoy voxwords` | List available words |
## Emoji SFX
Include emoji in your TTS text to trigger sound effects:
```
[p]tts Hello 🎉 world 🔥
```
Use `[p]sfx` to see all available triggers. Add `.mp3` files to `sfx/<folder>/` to customize.
## Requirements
- RedBot 3.5+
- Audio cog loaded
- ffmpeg
- `requests`, `g2p_en`, `numpy`, `pydub` (pip)
+5
View File
@@ -0,0 +1,5 @@
from .ttstoy import TtsToy
async def setup(bot):
await bot.add_cog(TtsToy(bot))
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+3
View File
@@ -0,0 +1,3 @@
# DECTalk Server Configuration
PORT=3000
OUTPUT_DIR=./output
+4
View File
@@ -0,0 +1,4 @@
node_modules/
output/
*.log
.env
+83
View File
@@ -0,0 +1,83 @@
# Installing Authentic DECTalk Server
This guide shows how to install and run the authentic DECTalk server with the Moonbase Alpha voice.
## Quick Install
```bash
cd ~/devcogs/ttstoy/dectalk-server
npm install
```
This installs the `dectalk` npm package which includes the authentic DECTalk binaries.
## What You Get
The `dectalk` package provides:
- Perfect Paul voice (the Moonbase Alpha voice)
- Authentic DECTalk synthesis engine
- Support for DECTalk phoneme commands
- Same voice used by Stephen Hawking
## Start the Server
```bash
PORT=3001 npm start
```
Or use the start script:
```bash
./start.sh
```
## Test It
```bash
# Test health
curl http://127.0.0.1:3001/health
# Generate audio
curl "http://127.0.0.1:3001/say?text=aeiou" -o aeiou.mp3
ffplay aeiou.mp3
```
## Configure TTSTOY
In Discord:
```
[p]ttstoy mode dectalk
[p]ttstoy dectalkurl http://127.0.0.1:3001
[p]tts aeiou john madden
```
## Troubleshooting
### "Cannot find module 'dectalk'"
Run `npm install` in the dectalk-server directory.
### Port already in use
Use a different port:
```bash
PORT=3002 npm start
```
### Linux dependency issues
The dectalk package may need additional libraries:
```bash
sudo apt install libasound2
```
## DECTalk Commands
You can use DECTalk phoneme commands in your text:
```
[p]tts [:phoneme on] hehlow werld
[p]tts [:rate 200] speaking very fast
[p]tts [:dv ap 100] aeiou
```
These commands control pitch, rate, and other voice parameters just like in Moonbase Alpha.
+186
View File
@@ -0,0 +1,186 @@
# Integration with RedBot TTSTOY
This guide shows how to integrate the DECTalk server with your RedBot TTSTOY cog.
## Quick Start
### 1. Install and Start DECTalk Server
```bash
cd dectalk-server
npm install
npm start
```
The server will start on `http://127.0.0.1:3000`
### 2. Configure TTSTOY
In your Discord server, run these commands:
```
[p]ttstoy mode dectalk
[p]ttstoy dectalkurl http://127.0.0.1:3000
```
### 3. Test It
```
[p]tts Hello, I am DECTalk, a text to speech system.
```
## API Compatibility
The DECTalk server implements the same API format as your local-tts server:
### Endpoints
| Endpoint | Method | Purpose | Compatible With |
|----------|--------|---------|-----------------|
| `/health` | GET | Health check | local-tts |
| `/engine` | GET | Engine info | local-tts |
| `/say?text=<text>` | GET | Simple TTS | DECTalk-specific |
| `/v1/t2a_v2` | POST | MiniMax-compatible TTS | TTSTOY cog |
### Response Format
The `/v1/t2a_v2` endpoint returns the exact same format as MiniMax API:
```json
{
"base_resp": {
"status_code": 0,
"status_msg": "Success"
},
"data": {
"audio": "<hex-encoded-mp3-data>",
"audio_id": "<unique-identifier>"
}
}
```
This matches the format expected by `_save_minimax_tts()` in your TTSTOY cog.
## How It Works
1. **Text Input**: TTSTOY sends text via POST to `/v1/t2a_v2`
2. **DECTalk Generation**: Server generates WAV audio using DECTalk
3. **MP3 Conversion**: ffmpeg converts WAV to MP3 (mono, 32kHz, 128kbps)
4. **Hex Encoding**: MP3 is converted to hex string (MiniMax format)
5. **Response**: Server returns hex-encoded audio in MiniMax-compatible JSON
## Switching Between TTS Modes
You can easily switch between different TTS engines:
```bash
# Use MiniMax cloud API
[p]ttstoy mode minimax
# Use local AI TTS
[p]ttstoy mode local
[p]ttstoy localurl http://127.0.0.1:8000
# Use DECTalk
[p]ttstoy mode dectalk
[p]ttstoy dectalkurl http://127.0.0.1:3000
```
## Features Supported
✅ Text-to-speech synthesis
✅ Emoji SFX (handled by TTSTOY cog)
✅ Audio concatenation (handled by TTSTOY cog)
✅ Volume control (handled by TTSTOY cog)
❌ Voice selection (DECTalk has one voice)
❌ Speed control (not implemented yet)
## Troubleshooting
### Server won't start
Check if port 3000 is already in use:
```bash
lsof -i :3000
```
Use a different port:
```bash
PORT=8080 npm start
```
### "Cannot reach DECTalk API"
1. Check if server is running:
```bash
curl http://127.0.0.1:3000/health
```
2. Check server logs for errors
3. Verify ffmpeg is installed:
```bash
ffmpeg -version
```
### Audio quality issues
The server outputs MP3 at:
- Sample rate: 32kHz
- Bitrate: 128kbps
- Channels: Mono
This matches TTSTOY's expectations. If you need different settings, modify the `convertToMP3()` function in `server.js`.
## Production Deployment
### Using systemd
1. Edit `dectalk-server.service` with your paths
2. Copy to systemd:
```bash
sudo cp dectalk-server.service /etc/systemd/system/
```
3. Enable and start:
```bash
sudo systemctl daemon-reload
sudo systemctl enable dectalk-server
sudo systemctl start dectalk-server
sudo systemctl status dectalk-server
```
### Using PM2
```bash
npm install -g pm2
pm2 start server.js --name dectalk-server
pm2 save
pm2 startup
```
## Performance Notes
- Each TTS request generates a new MP3 file
- Files are saved to `output/` directory
- Consider adding cleanup for old files
- DECTalk is fast and lightweight compared to AI TTS
## Comparison with local-tts
| Feature | local-tts | dectalk-server |
|---------|-----------|----------------|
| Voice Quality | Natural (AI) | Robotic (classic) |
| Speed | Slower | Faster |
| Resources | High (GPU recommended) | Low |
| Voice Cloning | Yes | No |
| API Key | No | No |
| Setup Complexity | High | Low |
## Next Steps
- Replace `say` command with actual DECTalk binary for authentic voice
- Add voice parameter support (if using multi-voice DECTalk)
- Implement speed control
- Add audio file cleanup
- Add rate limiting
+237
View File
@@ -0,0 +1,237 @@
# DECTalk TTS Server
A simple Node.js server that provides DECTalk text-to-speech with MiniMax-compatible API endpoints for RedBot TTSTOY cog.
## Features
- Authentic DECTalk voice (Perfect Paul - same as Moonbase Alpha)
- MiniMax-compatible API endpoint (`/v1/t2a_v2`) for seamless integration with TTSTOY
- Simple GET endpoint (`/say?text=<text>`) for testing
- Automatic WAV to MP3 conversion using ffmpeg
- Health check endpoint (`/health`)
## Requirements
- Node.js (v14 or higher)
- npm
- ffmpeg (for audio conversion)
- Linux system (DECTalk package doesn't support macOS)
### Installing Dependencies
**Ubuntu/Debian:**
```bash
sudo apt install nodejs npm ffmpeg
```
**Fedora:**
```bash
sudo dnf install nodejs npm ffmpeg
```
**Arch Linux:**
```bash
sudo pacman -S nodejs npm ffmpeg
```
**macOS:**
```bash
brew install node ffmpeg
# macOS has built-in 'say' command, no need for espeak
```
## Installation
```bash
cd dectalk-server
npm install
```
This will install:
- `express` - Web server framework
- `dectalk` - Authentic DECTalk voice synthesis (Moonbase Alpha)
The `dectalk` package includes the actual DECTalk binaries and will work on Linux systems.
## Usage
### Start the server
```bash
npm start
```
The server will run on port 3000 by default.
### Custom port
```bash
PORT=8080 npm start
```
### Custom output directory
```bash
OUTPUT_DIR=/path/to/output npm start
```
## API Endpoints
### Health Check
```bash
GET /health
```
Returns:
```json
{"status": "ok"}
```
### Engine Info
```bash
GET /engine
```
Returns:
```json
{
"type": "dectalk",
"version": "1.0.0",
"description": "DECTalk Text-to-Speech Engine"
}
```
### Simple TTS (GET)
```bash
GET /say?text=Hello%20world
```
Returns: MP3 audio file
### MiniMax-Compatible TTS (POST)
```bash
POST /v1/t2a_v2
Content-Type: application/json
{
"text": "Hello world",
"voice_setting": {
"voice_id": "default",
"speed": 1.0
}
}
```
Returns:
```json
{
"base_resp": {
"status_code": 0,
"status_msg": "Success"
},
"data": {
"audio": "<hex-encoded-mp3>",
"audio_id": "<unique-id>"
}
}
```
## Integration with RedBot TTSTOY
1. Start the DECTalk server:
```bash
npm start
```
2. Configure TTSTOY to use DECTalk mode:
```
[p]ttstoy mode dectalk
[p]ttstoy dectalkurl http://127.0.0.1:3000
```
3. Test it:
```
[p]tts Hello, I am DECTalk
```
## Notes
- Uses authentic DECTalk (Perfect Paul voice from Moonbase Alpha)
- The `dectalk` npm package includes the actual DECTalk binaries
- Audio files are saved to the `output/` directory
- MP3 format: mono, 32kHz, 128kbps (matches TTSTOY expectations)
- Supports all 9 original DECTalk voices
- Supports DECTalk phoneme commands (e.g., `[:phoneme on]`, `[:rate 200]`)
## DECTalk Commands
DECTalk supports special commands for controlling speech:
```
[:phoneme on] - Enable phoneme mode
[:rate 200] - Set speech rate (default 180)
[:dv ap 100] - Set voice parameters
```
Example:
```bash
curl "http://127.0.0.1:3001/say?text=%5B:rate%20150%5DHello%20world&voice=paul" -o test.mp3
```
You can also use these in Discord:
```
[p]tts [:rate 200] speaking very fast
[p]tts [:phoneme on] hehlow werld
```
## DECTalk Voices
DECTalk includes 9 classic voices:
| Voice | Name | Description |
|-------|------|-------------|
| Paul | Perfect Paul | Default male voice (Moonbase Alpha, Stephen Hawking) |
| Betty | Beautiful Betty | Female voice |
| Harry | Huge Harry | Deep male voice |
| Frank | Frail Frank | Elderly male voice |
| Dennis | Doctor Dennis | Male voice |
| Kit | Kit the Kid | Child voice |
| Ursula | Uppity Ursula | Female voice |
| Rita | Rough Rita | Female voice |
| Wendy | Whispering Wendy | Soft female voice |
### Using Different Voices
**In Discord:**
```
[p]ttstoy voice Paul
[p]tts Hello, I am Perfect Paul
[p]ttstoy voice Betty
[p]tts Hello, I am Beautiful Betty
```
**Direct API:**
```bash
curl "http://127.0.0.1:3001/say?text=Hello&voice=paul" -o paul.mp3
curl "http://127.0.0.1:3001/say?text=Hello&voice=betty" -o betty.mp3
```
**List available voices:**
```bash
curl http://127.0.0.1:3001/voices
```
## Customization
The server uses the default DECTalk voice (Perfect Paul). The `dectalk` package handles all the voice synthesis internally, so no additional configuration is needed for the authentic Moonbase Alpha sound.
## Development
Run with auto-reload:
```bash
npm run dev
```
## License
MIT
+158
View File
@@ -0,0 +1,158 @@
# Testing DECTalk Server with TTSTOY
Quick testing guide to verify everything works.
## Step 1: Start the DECTalk Server
```bash
cd ~/Desktop/ttstoy/dectalk-server
PORT=3001 npm start
```
You should see:
```
DECTalk TTS Server running on http://localhost:3001
Health check: http://localhost:3001/health
Simple API: http://localhost:3001/say?text=Hello
MiniMax-compatible API: POST http://localhost:3001/v1/t2a_v2
```
## Step 2: Test the API Directly
Open a new terminal and test:
```bash
# Test health check
curl http://127.0.0.1:3001/health
# Test simple endpoint (saves to test.mp3)
curl "http://127.0.0.1:3001/say?text=Hello%20world" -o test.mp3
# Play the audio
ffplay test.mp3
# or
mpv test.mp3
# Test MiniMax-compatible endpoint
curl -X POST http://127.0.0.1:3001/v1/t2a_v2 \
-H "Content-Type: application/json" \
-d '{"text": "Testing DECTalk compatibility", "voice_setting": {"voice_id": "default"}}' \
| jq '.base_resp'
```
Expected output:
```json
{
"status_code": 0,
"status_msg": "Success"
}
```
## Step 3: Configure TTSTOY
In Discord:
```
[p]ttstoy dectalkurl http://127.0.0.1:3001
[p]ttstoy mode dectalk
[p]ttstoy mode
```
You should see:
```
Current TTS Mode: 🤖 DECTALK
DECTalk API URL: http://127.0.0.1:3001
```
## Step 4: Test TTS in Discord
```
[p]tts Hello, I am DECTalk, a text to speech system.
```
The bot should:
1. Generate audio using DECTalk
2. Upload the MP3 file to Discord
3. Play it in voice chat
## Step 5: Test with Emoji SFX
```
[p]tts Hello 🎉 this is a party 🔥
```
The bot should mix DECTalk voice with sound effects.
## Troubleshooting
### Server won't start (port in use)
```bash
# Find what's using port 3001
lsof -i :3001
# Use a different port
PORT=3002 npm start
# Update TTSTOY
[p]ttstoy dectalkurl http://127.0.0.1:3002
```
### "Cannot reach DECTalk API"
1. Check server is running:
```bash
curl http://127.0.0.1:3001/health
```
2. Check server logs in the terminal where you started it
3. Verify URL in TTSTOY:
```
[p]ttstoy dectalkurl
```
### Audio not playing
1. Make sure ffmpeg is installed:
```bash
ffmpeg -version
```
2. Check bot logs for errors
3. Test the `/say` endpoint directly and play the audio
### "say: command not found" (Linux)
The current server uses macOS `say` command. On Linux, you need to:
1. Install actual DECTalk or use espeak:
```bash
# Option 1: espeak (simple)
sudo apt install espeak
# Option 2: Install dectalk-tts
npm install -g dectalk-tts
```
2. Modify `server.js` to use the appropriate command
## Success Checklist
- [ ] DECTalk server starts without errors
- [ ] `/health` endpoint returns `{"status": "ok"}`
- [ ] `/say` endpoint generates audio file
- [ ] TTSTOY mode switches to dectalk
- [ ] `[p]tts` command works in Discord
- [ ] Audio plays in voice chat
- [ ] Emoji SFX work correctly
## Next Steps
Once everything works:
1. Set up auto-start with systemd or PM2
2. Configure firewall if needed
3. Consider using actual DECTalk binary for authentic voice
4. Add cleanup for old audio files in output/
+69
View File
@@ -0,0 +1,69 @@
# DECTalk Voices Reference
The 9 original DECTalk voices, each with unique characteristics.
## Voice List
| Voice | Description | Type |
|-------|-------------|------|
| paul | Perfect Paul (default, Moonbase Alpha) | Male |
| betty | Beautiful Betty | Female |
| harry | Huge Harry (deep voice) | Male |
| frank | Frail Frank (elderly) | Male |
| dennis | Doctor Dennis | Male |
| kit | Kit the Kid | Child |
| ursula | Uppity Ursula | Female |
| rita | Rough Rita (gravelly) | Female |
| wendy | Whispering Wendy (soft) | Female |
## Using Voices in Discord
### Set Your Voice
```
[p]ttstoy myvoice paul
[p]tts aeiou john madden
```
The voice command `[:np]` is automatically prepended to your text.
### Override Voice Mid-Text
You can use DECTalk commands to change voices within your message:
```
[p]tts [:nh]I'm gonna eat a pizza. [:dial67589340] Hi, can i order a pizza? [:nv]no! [:nh]why? [:nv] cuz you are john madden![:np]
```
Voice commands:
- `[:np]` - Paul
- `[:nb]` - Betty
- `[:nh]` - Harry
- `[:nf]` - Frank
- `[:nd]` - Dennis
- `[:nk]` - Kit
- `[:nu]` - Ursula
- `[:nr]` - Rita
- `[:nw]` - Wendy
## Testing Voices
```bash
# Test all voices
for voice in paul betty harry frank dennis kit ursula rita wendy; do
echo "Testing $voice..."
curl "http://127.0.0.1:3001/say?text=Hello%20I%20am%20$voice&voice=$voice" -o "${voice}.mp3"
done
```
## In Discord
```
# Set voice and speak
[p]ttstoy voice paul
[p]tts aeiou john madden
# Change voice
[p]ttstoy voice betty
[p]tts Hello, I am Beautiful Betty
# Override voice in text
[p]tts [:nh]Deep voice [:nb]now female voice [:np]back to paul
```
@@ -0,0 +1,16 @@
[Unit]
Description=DECTalk TTS Server
After=network.target
[Service]
Type=simple
User=YOUR_USER
WorkingDirectory=/path/to/dectalk-server
Environment=PORT=3000
Environment=OUTPUT_DIR=/path/to/dectalk-server/output
ExecStart=/usr/bin/node /path/to/dectalk-server/server.js
Restart=always
RestartSec=3
[Install]
WantedBy=multi-user.target
File diff suppressed because it is too large Load Diff
+20
View File
@@ -0,0 +1,20 @@
{
"name": "dectalk-server",
"version": "1.0.0",
"description": "DECTalk TTS server compatible with RedBot TTSTOY cog",
"main": "server.js",
"scripts": {
"start": "node server.js",
"dev": "nodemon server.js"
},
"keywords": ["dectalk", "tts", "text-to-speech", "redbot", "moonbase-alpha"],
"author": "",
"license": "MIT",
"dependencies": {
"express": "^4.18.2",
"dectalk": "^1.0.0"
},
"devDependencies": {
"nodemon": "^3.0.1"
}
}
+209
View File
@@ -0,0 +1,209 @@
#!/usr/bin/env node
/**
* DECTalk TTS Server
* Compatible with RedBot TTSTOY cog
*
* Uses authentic DECTalk (Moonbase Alpha voice)
* Provides MiniMax-compatible API endpoint at /v1/t2a_v2
* Also provides simple GET endpoint at /say?text=<text>
*/
const express = require('express');
const { say } = require('dectalk');
const { spawn } = require('child_process');
const fs = require('fs');
const path = require('path');
const crypto = require('crypto');
const app = express();
const PORT = process.env.PORT || 3000;
const OUTPUT_DIR = process.env.OUTPUT_DIR || path.join(__dirname, 'output');
// Ensure output directory exists
if (!fs.existsSync(OUTPUT_DIR)) {
fs.mkdirSync(OUTPUT_DIR, { recursive: true });
}
app.use(express.json());
// Health check endpoint
app.get('/health', (req, res) => {
res.json({ status: 'ok' });
});
// Engine info endpoint
app.get('/engine', (req, res) => {
res.json({
type: 'dectalk',
version: '1.0.0',
description: 'DECTalk Text-to-Speech Engine'
});
});
// Voices endpoint - list available DECTalk voice commands
app.get('/voices', (req, res) => {
const voices = [
{ voice_id: 'paul', command: '[:np]', description: 'Perfect Paul (default, male)' },
{ voice_id: 'betty', command: '[:nb]', description: 'Beautiful Betty (female)' },
{ voice_id: 'harry', command: '[:nh]', description: 'Huge Harry (deep male)' },
{ voice_id: 'frank', command: '[:nf]', description: 'Frail Frank (elderly male)' },
{ voice_id: 'dennis', command: '[:nd]', description: 'Doctor Dennis (male)' },
{ voice_id: 'kit', command: '[:nk]', description: 'Kit the Kid (child)' },
{ voice_id: 'ursula', command: '[:nu]', description: 'Uppity Ursula (female)' },
{ voice_id: 'rita', command: '[:nr]', description: 'Rough Rita (gravelly female)' },
{ voice_id: 'wendy', command: '[:nw]', description: 'Whispering Wendy (soft female)' }
];
res.json(voices);
});
// Simple GET endpoint for DECTalk
app.get('/say', async (req, res) => {
const text = req.query.text;
if (!text) {
return res.status(400).send('Missing text parameter');
}
try {
const wavBuffer = await generateDECTalk(text);
// Convert WAV to MP3 using ffmpeg
const mp3Buffer = await convertToMP3(wavBuffer);
res.set('Content-Type', 'audio/mpeg');
res.send(mp3Buffer);
} catch (error) {
console.error('DECTalk generation error:', error);
res.status(500).send(`TTS generation failed: ${error.message}`);
}
});
// MiniMax-compatible endpoint for RedBot TTSTOY cog
app.post('/v1/t2a_v2', async (req, res) => {
const { text } = req.body;
if (!text) {
return res.json({
base_resp: {
status_code: 1002,
status_msg: 'Text is required'
}
});
}
console.log(`Generating DECTalk TTS`);
console.log(`Text: ${text.substring(0, 100)}...`);
try {
// Generate DECTalk audio - voice is controlled by [:n*] commands in text
const wavBuffer = await generateDECTalk(text);
// Convert WAV to MP3
const mp3Buffer = await convertToMP3(wavBuffer);
// Save to file
const audioId = crypto.randomBytes(16).toString('hex');
const mp3Path = path.join(OUTPUT_DIR, `${audioId}.mp3`);
fs.writeFileSync(mp3Path, mp3Buffer);
// Convert to hex (MiniMax format)
const audioHex = mp3Buffer.toString('hex');
console.log(`✅ Generated audio: ${audioId}.mp3 (${mp3Buffer.length} bytes)`);
// Return MiniMax-compatible response
res.json({
base_resp: {
status_code: 0,
status_msg: 'Success'
},
data: {
audio: audioHex,
audio_id: audioId
}
});
} catch (error) {
console.error('DECTalk generation error:', error);
res.json({
base_resp: {
status_code: 1005,
status_msg: `TTS generation failed: ${error.message}`
}
});
}
});
/**
* Generate authentic DECTalk audio (Moonbase Alpha voice)
* @param {string} text - Text to synthesize (can include [:n*] voice commands)
* @returns {Promise<Buffer>} WAV audio buffer
*/
async function generateDECTalk(text) {
try {
// Use the authentic DECTalk package
// Voice is controlled by [:n*] commands in the text itself
const wavBuffer = await say(text);
return wavBuffer;
} catch (error) {
throw new Error(`DECTalk generation failed: ${error.message}`);
}
}
/**
* Convert WAV buffer to MP3 using ffmpeg
* @param {Buffer} wavBuffer - WAV audio buffer
* @returns {Promise<Buffer>} MP3 audio buffer
*/
function convertToMP3(wavBuffer) {
return new Promise((resolve, reject) => {
const ffmpeg = spawn('ffmpeg', [
'-f', 'wav',
'-i', 'pipe:0',
'-f', 'mp3',
'-ac', '1',
'-ar', '32000',
'-b:a', '128k',
'pipe:1'
]);
const chunks = [];
ffmpeg.stdout.on('data', (chunk) => {
chunks.push(chunk);
});
ffmpeg.stderr.on('data', (data) => {
// ffmpeg outputs progress to stderr, ignore it
});
ffmpeg.on('close', (code) => {
if (code !== 0) {
reject(new Error(`ffmpeg process exited with code ${code}`));
} else {
resolve(Buffer.concat(chunks));
}
});
ffmpeg.on('error', (err) => {
reject(new Error(`Failed to start ffmpeg: ${err.message}`));
});
// Write WAV data to ffmpeg stdin
ffmpeg.stdin.write(wavBuffer);
ffmpeg.stdin.end();
});
}
// Start server
app.listen(PORT, () => {
console.log(`🤖 DECTalk TTS Server running on http://localhost:${PORT}`);
console.log(`Health check: http://localhost:${PORT}/health`);
console.log(`Voices: http://localhost:${PORT}/voices`);
console.log(`Simple API: http://localhost:${PORT}/say?text=Hello`);
console.log(`MiniMax-compatible API: POST http://localhost:${PORT}/v1/t2a_v2`);
console.log('');
console.log('Voice commands (use in text):');
console.log(' [:np] Paul [:nb] Betty [:nh] Harry [:nf] Frank [:nd] Dennis');
console.log(' [:nk] Kit [:nu] Ursula [:nr] Rita [:nw] Wendy');
});
+33
View File
@@ -0,0 +1,33 @@
#!/bin/bash
# Quick start script for DECTalk server
echo "🤖 Starting DECTalk TTS Server..."
# Check if node_modules exists
if [ ! -d "node_modules" ]; then
echo "📦 Installing dependencies (including authentic DECTalk)..."
npm install
fi
# Check if ffmpeg is available
if ! command -v ffmpeg &> /dev/null; then
echo "❌ Error: ffmpeg is not installed"
echo "Install it with:"
echo " - Ubuntu/Debian: sudo apt install ffmpeg"
echo " - Fedora: sudo dnf install ffmpeg"
echo " - Arch: sudo pacman -S ffmpeg"
exit 1
fi
# Check if dectalk package is installed
if [ ! -d "node_modules/dectalk" ]; then
echo "❌ Error: dectalk package not installed"
echo "Run: npm install"
exit 1
fi
# Start the server
echo "✅ Starting server on port ${PORT:-3000}..."
echo "🎙️ Using authentic DECTalk (Moonbase Alpha voice)"
npm start
+48
View File
@@ -0,0 +1,48 @@
#!/bin/bash
# Test script for DECTalk server API
BASE_URL="${1:-http://127.0.0.1:3000}"
echo "🧪 Testing DECTalk Server at $BASE_URL"
echo ""
# Test 1: Health check
echo "1️⃣ Testing /health endpoint..."
curl -s "$BASE_URL/health" | jq '.' || echo "❌ Health check failed"
echo ""
# Test 2: Engine info
echo "2️⃣ Testing /engine endpoint..."
curl -s "$BASE_URL/engine" | jq '.' || echo "❌ Engine info failed"
echo ""
# Test 3: Simple GET endpoint
echo "3️⃣ Testing /say endpoint (saving to test_say.mp3)..."
curl -s "$BASE_URL/say?text=Hello%20world" -o test_say.mp3
if [ -f test_say.mp3 ] && [ -s test_say.mp3 ]; then
echo "✅ Audio saved to test_say.mp3 ($(stat -f%z test_say.mp3 2>/dev/null || stat -c%s test_say.mp3) bytes)"
else
echo "❌ Failed to generate audio"
fi
echo ""
# Test 4: MiniMax-compatible endpoint
echo "4️⃣ Testing /v1/t2a_v2 endpoint (MiniMax-compatible)..."
curl -s -X POST "$BASE_URL/v1/t2a_v2" \
-H "Content-Type: application/json" \
-d '{
"text": "Testing MiniMax compatibility",
"voice_setting": {
"voice_id": "default",
"speed": 1.0
}
}' | jq '.base_resp, .data.audio_id' || echo "❌ MiniMax endpoint failed"
echo ""
echo "✅ All tests complete!"
echo ""
echo "To test audio playback:"
echo " ffplay test_say.mp3"
echo " # or"
echo " mpv test_say.mp3"
+10
View File
@@ -0,0 +1,10 @@
{
"author": ["Kingston"],
"install_msg": "TtsToy (MiniMax) cog loaded.",
"name": "TtsToy",
"short": "MiniMax text-to-speech with emoji-triggered SFX and MorshuTalk.",
"description": "Uses MiniMax TTS, DECTalk, and MorshuTalk with optional emoji/custom-emoji SFX segments for voice channel playback.",
"requirements": ["requests", "g2p_en", "numpy", "pydub"],
"tags": ["tts", "minimax", "voice", "audio", "morshu"],
"min_bot_version": "3.5.0"
}
+306
View File
@@ -0,0 +1,306 @@
"""
Bundled MorshuTalk engine for ttstoy.
Based on MorshuTalk by jalenluorion (https://github.com/jalenluorion/MorshuTalk)
Generates speech audio from Morshu's voice lines using phoneme matching.
"""
import re
import unicodedata
from os import path
from typing import List, Tuple, Callable, Literal
import numpy as np
import random
import warnings
from pydub import AudioSegment
# ---------------------------------------------------------------------------
# G2P (Grapheme-to-Phoneme) wrapper with progress + cancel support
# ---------------------------------------------------------------------------
import nltk
for _res in ('averaged_perceptron_tagger_eng', 'cmudict', 'punkt', 'punkt_tab'):
try:
nltk.data.find(f'taggers/{_res}' if 'tagger' in _res else _res)
except LookupError:
nltk.download(_res, quiet=True)
from g2p_en.g2p import G2p, unicode, normalize_numbers, word_tokenize, pos_tag
class G2pProgress(G2p):
def __init__(self):
super().__init__()
self.cancelled = False
def cancel(self):
self.cancelled = True
def run_with_progress(self, text, callback: Callable[[int, int], None] = None):
self.cancelled = False
text = unicode(text)
text = normalize_numbers(text)
text = ''.join(char for char in unicodedata.normalize('NFD', text)
if unicodedata.category(char) != 'Mn')
text = text.lower()
text = re.sub("[^ a-z'.,?!\\-]", "", text)
text = text.replace("i.e.", "that is")
text = text.replace("e.g.", "for example")
words = word_tokenize(text)
tokens = pos_tag(words)
step = 0
total = len(tokens)
prons = []
for word, pos in tokens:
if self.cancelled:
return
if callback:
callback(step, total)
step += 1
if re.search("[a-z]", word) is None:
pron = [word]
elif word in self.homograph2features:
pron1, pron2, pos1 = self.homograph2features[word]
pron = pron1 if pos.startswith(pos1) else pron2
elif word in self.cmu:
pron = self.cmu[word][0]
else:
pron = self.predict(word)
prons.extend(pron)
prons.extend([" "])
if callback:
callback(total, total)
return prons[:-1]
# ---------------------------------------------------------------------------
# Singleton G2P instance and audio data
# ---------------------------------------------------------------------------
g2p = G2pProgress()
_WAV_PATH = path.join(path.dirname(__file__), 'morshutalk_morshu.wav')
morshu_wav = AudioSegment.from_wav(_WAV_PATH)
# Phoneme timing record from the morshu audio
morshu_rec = np.rec.array([
('', 160, 0), ('L', 250, 2), ('AE', 348, 2), ('M', 420, 2), ('P', 510, 1),
('OY', 700, 2), ('L', 835, 1), ('', 1090, 0),
('R', 1180, 2), ('OW', 1300, 2), ('', 1390, 0), ('P', 1490, 2), ('', 1850, 0),
('B', 1895, 2), ('AA', 2090, 2), ('M', 2235, 2), ('Z', 2390, 2),
('', 2780, 0), ('Y', 2840, 2), ('UW', 2960, 2),
('W', 3030, 2), ('AA', 3110, 2), ('N', 3150, 1), ('IH', 3240, 2), ('T', 3370, 2), ('', 3810, 0),
('IH', 3960, 2), ('T', 4070, 2), ('Y', 4260, 2), ('UH', 4400, 2), ('R', 4510, 2), ('Z', 4600, 2),
('M', 4675, 2), ('AY', 4810, 2), ('', 4885, 0),
('F', 4930, 2), ('R', 4980, 2), ('EH', 5100, 2), ('N', 5240, 2), ('D', 5300, 2), ('', 5520, 0),
('AE', 5630, 2), ('Z', 5740, 2), ('L', 5870, 2), ('AO', 6000, 2), ('NG', 6140, 2),
('AE', 6170, 1), ('Z', 6265, 2), ('Y', 6300, 2), ('UW', 6380, 2),
('HH', 6450, 2), ('AE', 6510, 1), ('V', 6580, 2),
('IH', 6640, 2), ('N', 6670, 2), ('AH', 6747, 2), ('F', 6855, 2),
('R', 6960, 2), ('UW', 7060, 2), ('B', 7170, 1), ('IY', 7340, 2), ('Z', 7520, 2), ('', 8236, 0),
('S', 8407, 2), ('AA', 8495, 2), ('R', 8570, 2), ('IY', 8630, 1),
('L', 8740, 2), ('IH', 8811, 2), ('NG', 8942, 2), ('K', 9014, 2), ('', 9251, 0),
('AY', 9384, 2), ('', 9467, 0), ('K', 9512, 2), ('AE', 9640, 2), ('N', 9716, 2), ('', 9844, 0),
('G', 9894, 2), ('IH', 9985, 2), ('V', 10060, 2), ('', 10149, 0),
('K', 10256, 2), ('R', 10297, 2), ('EH', 10383, 2), ('IH', 10482, 1), ('', 10564, 0), ('T', 10617, 2),
('', 10962, 0), ('K', 11019, 2), ('AH', 11100, 2), ('M', 11229, 2), ('B', 11246, 2), ('AE', 11369, 2),
('', 11511, 0), ('W', 11590, 2), ('EH', 11622, 1), ('N', 11705, 2),
('Y', 11755, 2), ('UH', 11808, 2), ('R', 11864, 2), ('AH', 11959, 2),
('L', 12095, 2), ('IH', 12202, 2), ('L', 12386, 2),
('', 12596, 0), ('M', 12748, 2), ('M', 12888, 2), ('M', 13037, 2), ('M', 13196, 2), ('', 13426, 0),
('R', 13494, 2), ('IH', 13589, 2), ('', 13632, 0), ('CH', 13773, 2), ('ER', 13991, 2), ('', 13992, 0)
], names=('phoneme', 'timing', 'priority'))
similar_phonemes = {
'AW': ['AE', 'UW'],
'DH': ['D'],
'EY': ['EH', 'IY'],
'JH': ['CH'],
'SH': ['CH'],
'TH': ['D'],
'ZH': ['CH'],
}
# ---------------------------------------------------------------------------
# Morshu TTS Engine
# ---------------------------------------------------------------------------
class Morshu:
def __init__(self):
self.input_str = ""
self.input_phonemes = []
self.stop_chars = '.,?!:;()\n'
self.space_length = 20
self.stop_length = 100
self.use_phoneme_priority = True
self.out_audio = AudioSegment.empty()
self.audio_segment_timings = np.rec.array((0, 0), names=('output', 'morshu'))
self.canceled = False
def cancel(self):
g2p.cancel()
self.canceled = True
def load_text(self, text: str = None, progress_callback: Callable[[int, int, int], None] = None) \
-> AudioSegment | Literal[False]:
"""Generate audio from text. Returns AudioSegment or False if cancelled."""
self.canceled = False
if progress_callback is None:
progress_callback = lambda major_step, minor_step, minor_total: None
if text is None:
text = self.input_str
self.input_str = text
text = text.replace('\n', ',,,')
phonemes = g2p.run_with_progress(text, lambda step, total: progress_callback(0, step, total))
if g2p.cancelled:
return False
progress_step = 0
progress_total = len(phonemes)
output = AudioSegment.empty().set_frame_rate(morshu_wav.frame_rate)
audio_out_millis = []
audio_morshu_millis = []
phoneme_segment = []
while len(phonemes) > 0:
if self.canceled:
return False
progress_callback(1, progress_step, progress_total)
progress_step += 1
p = phonemes.pop(0)
if p in g2p.phonemes:
phoneme_segment.append(p)
if p not in g2p.phonemes or len(phonemes) == 0:
output = self.append_best_morshu_phoneme_segment(
output, phoneme_segment, audio_out_millis, audio_morshu_millis)
phoneme_segment = []
if p == ' ':
output = self.append_audio_segment(
output, AudioSegment.silent(self.space_length), -1,
audio_out_millis, audio_morshu_millis)
elif p in self.stop_chars:
output = self.append_audio_segment(
output, AudioSegment.silent(self.stop_length), -1,
audio_out_millis, audio_morshu_millis)
if len(output) == 0:
warnings.warn('returned audio segment is empty', UserWarning)
self.audio_segment_timings = np.rec.array((0, 0), names=('output', 'morshu'))
else:
self.audio_segment_timings = np.rec.array(
tuple(zip(audio_out_millis, audio_morshu_millis)),
names=('output', 'morshu'))
progress_callback(1, progress_total, progress_total)
self.out_audio = output
return output
@staticmethod
def substitute_similar_phonemes(phonemes: List[str]):
i = 0
while i < len(phonemes):
if phonemes[i].endswith('0') or phonemes[i].endswith('1') or phonemes[i].endswith('2'):
phonemes[i] = phonemes[i][:len(phonemes[i]) - 1]
if phonemes[i] in similar_phonemes:
phonemes = phonemes[0:i] + similar_phonemes[phonemes[i]] + phonemes[i + 1:]
i += 1
return phonemes
@staticmethod
def append_audio_segment(audio_out: AudioSegment, audio_segment: AudioSegment,
morshu_millis_start: int,
audio_out_millis: List[int],
audio_morshu_millis: List[int]) -> AudioSegment:
audio_out_millis.append(len(audio_out))
audio_morshu_millis.append(morshu_millis_start)
audio_out += audio_segment
return audio_out
@staticmethod
def get_phoneme_sequence_occurrences(phonemes: List[str]) -> List[Tuple[int, int]]:
occurrences = []
for i in range(len(morshu_rec) - len(phonemes)):
if (morshu_rec['phoneme'][i:i + len(phonemes)] == phonemes).all():
start = morshu_rec['timing'][i - 1]
end = morshu_rec['timing'][i + len(phonemes) - 1]
occurrences.append((start, end))
return occurrences
def get_best_morshu_single_phoneme(self, phoneme: str, preceding: str = "",
succeeding: str = "") -> Tuple[AudioSegment, int]:
best_indices = []
phoneme_indices = np.where(morshu_rec['phoneme'] == phoneme)[0]
if len(phoneme_indices) == 0:
return AudioSegment.empty(), 0
highest_priority = 0
for i in phoneme_indices:
morshu_preceding = morshu_rec['phoneme'][i - 1]
priority = morshu_rec['priority'][i] if self.use_phoneme_priority else 0
if morshu_preceding == preceding:
priority += 10
elif any(c in morshu_preceding for c in "AEIOU") and any(c in preceding for c in "AEIOU"):
priority += 5
morshu_succeeding = morshu_rec['phoneme'][i + 1]
if morshu_succeeding == succeeding:
priority += 10
elif any(c in morshu_succeeding for c in "AEIOU") and any(c in succeeding for c in "AEIOU"):
priority += 1
if priority < highest_priority:
continue
if priority > highest_priority:
highest_priority = priority
best_indices = []
best_indices.append(i)
index = random.choice(best_indices)
segment = morshu_wav[morshu_rec['timing'][index - 1]: morshu_rec['timing'][index]]
return segment, morshu_rec['timing'][index - 1]
def append_best_morshu_phoneme_segment(self, output: AudioSegment, phonemes: List[str],
audio_out_millis: List[int] = None,
audio_morshu_millis: List[int] = None) -> AudioSegment:
phonemes = Morshu.substitute_similar_phonemes(phonemes)
if len(phonemes) == 1:
segment, start = self.get_best_morshu_single_phoneme(phonemes[0])
return Morshu.append_audio_segment(output, segment, start, audio_out_millis, audio_morshu_millis)
preceding = ""
while len(phonemes) > 0:
sequence_length = 1
segment = AudioSegment.empty()
start = 0
while sequence_length <= len(phonemes):
occurrences = Morshu.get_phoneme_sequence_occurrences(phonemes[:sequence_length])
if len(occurrences) == 0:
break
start, end = random.choice(occurrences)
segment = morshu_wav[start:end]
sequence_length += 1
sequence_length -= 1
if sequence_length == 1:
succeeding = phonemes[sequence_length + 1] if sequence_length + 1 < len(phonemes) else ""
segment, start = self.get_best_morshu_single_phoneme(phonemes[0], preceding, succeeding)
output = Morshu.append_audio_segment(output, segment, start, audio_out_millis, audio_morshu_millis)
preceding = phonemes[sequence_length - 1]
del phonemes[:sequence_length]
return output
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+19
View File
@@ -0,0 +1,19 @@
{
"🎉": "party",
"😂": "laugh",
"🥖": "spy",
"👏": "clap",
"🔥": "fire",
"💀": "skull",
"✅": "check",
"❌": "error",
"📢": "airhorn",
"🚢": "boathorn",
"😶": "drum",
"👼": "angel",
"🥜": "cashew",
"💪": "physical",
"🧠": "intelligence",
"👁": "psychic",
"✍️": "motor"
}
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+10
View File
@@ -0,0 +1,10 @@
{
"homepage": "/home/owen/Desktop/Homepage",
"webtable": "/home/owen/Desktop/WebTable",
"porgupicker": "/home/owen/Desktop/PorguPicker",
"galaxy": "/home/owen/Desktop/Something",
"ttstoy": "/home/owen/Desktop/ttstoy",
"ttstoy_webui": "/home/owen/Desktop/ttstoy/webui",
"cookie_domain": ".kingstons-scrapyard.net",
"session_secret": "scrapyard-shared-session-k8x7m2p4q9w1"
}
+3301
View File
File diff suppressed because it is too large Load Diff
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.

Some files were not shown because too many files have changed in this diff Show More