187 lines
3.9 KiB
Markdown
187 lines
3.9 KiB
Markdown
# Integration with RedBot TTSTOY
|
|
|
|
This guide shows how to integrate the DECTalk server with your RedBot TTSTOY cog.
|
|
|
|
## Quick Start
|
|
|
|
### 1. Install and Start DECTalk Server
|
|
|
|
```bash
|
|
cd dectalk-server
|
|
npm install
|
|
npm start
|
|
```
|
|
|
|
The server will start on `http://127.0.0.1:3000`
|
|
|
|
### 2. Configure TTSTOY
|
|
|
|
In your Discord server, run these commands:
|
|
|
|
```
|
|
[p]ttstoy mode dectalk
|
|
[p]ttstoy dectalkurl http://127.0.0.1:3000
|
|
```
|
|
|
|
### 3. Test It
|
|
|
|
```
|
|
[p]tts Hello, I am DECTalk, a text to speech system.
|
|
```
|
|
|
|
## API Compatibility
|
|
|
|
The DECTalk server implements the same API format as your local-tts server:
|
|
|
|
### Endpoints
|
|
|
|
| Endpoint | Method | Purpose | Compatible With |
|
|
|----------|--------|---------|-----------------|
|
|
| `/health` | GET | Health check | local-tts |
|
|
| `/engine` | GET | Engine info | local-tts |
|
|
| `/say?text=<text>` | GET | Simple TTS | DECTalk-specific |
|
|
| `/v1/t2a_v2` | POST | MiniMax-compatible TTS | TTSTOY cog |
|
|
|
|
### Response Format
|
|
|
|
The `/v1/t2a_v2` endpoint returns the exact same format as MiniMax API:
|
|
|
|
```json
|
|
{
|
|
"base_resp": {
|
|
"status_code": 0,
|
|
"status_msg": "Success"
|
|
},
|
|
"data": {
|
|
"audio": "<hex-encoded-mp3-data>",
|
|
"audio_id": "<unique-identifier>"
|
|
}
|
|
}
|
|
```
|
|
|
|
This matches the format expected by `_save_minimax_tts()` in your TTSTOY cog.
|
|
|
|
## How It Works
|
|
|
|
1. **Text Input**: TTSTOY sends text via POST to `/v1/t2a_v2`
|
|
2. **DECTalk Generation**: Server generates WAV audio using DECTalk
|
|
3. **MP3 Conversion**: ffmpeg converts WAV to MP3 (mono, 32kHz, 128kbps)
|
|
4. **Hex Encoding**: MP3 is converted to hex string (MiniMax format)
|
|
5. **Response**: Server returns hex-encoded audio in MiniMax-compatible JSON
|
|
|
|
## Switching Between TTS Modes
|
|
|
|
You can easily switch between different TTS engines:
|
|
|
|
```bash
|
|
# Use MiniMax cloud API
|
|
[p]ttstoy mode minimax
|
|
|
|
# Use local AI TTS
|
|
[p]ttstoy mode local
|
|
[p]ttstoy localurl http://127.0.0.1:8000
|
|
|
|
# Use DECTalk
|
|
[p]ttstoy mode dectalk
|
|
[p]ttstoy dectalkurl http://127.0.0.1:3000
|
|
```
|
|
|
|
## Features Supported
|
|
|
|
✅ Text-to-speech synthesis
|
|
✅ Emoji SFX (handled by TTSTOY cog)
|
|
✅ Audio concatenation (handled by TTSTOY cog)
|
|
✅ Volume control (handled by TTSTOY cog)
|
|
❌ Voice selection (DECTalk has one voice)
|
|
❌ Speed control (not implemented yet)
|
|
|
|
## Troubleshooting
|
|
|
|
### Server won't start
|
|
|
|
Check if port 3000 is already in use:
|
|
```bash
|
|
lsof -i :3000
|
|
```
|
|
|
|
Use a different port:
|
|
```bash
|
|
PORT=8080 npm start
|
|
```
|
|
|
|
### "Cannot reach DECTalk API"
|
|
|
|
1. Check if server is running:
|
|
```bash
|
|
curl http://127.0.0.1:3000/health
|
|
```
|
|
|
|
2. Check server logs for errors
|
|
|
|
3. Verify ffmpeg is installed:
|
|
```bash
|
|
ffmpeg -version
|
|
```
|
|
|
|
### Audio quality issues
|
|
|
|
The server outputs MP3 at:
|
|
- Sample rate: 32kHz
|
|
- Bitrate: 128kbps
|
|
- Channels: Mono
|
|
|
|
This matches TTSTOY's expectations. If you need different settings, modify the `convertToMP3()` function in `server.js`.
|
|
|
|
## Production Deployment
|
|
|
|
### Using systemd
|
|
|
|
1. Edit `dectalk-server.service` with your paths
|
|
2. Copy to systemd:
|
|
```bash
|
|
sudo cp dectalk-server.service /etc/systemd/system/
|
|
```
|
|
|
|
3. Enable and start:
|
|
```bash
|
|
sudo systemctl daemon-reload
|
|
sudo systemctl enable dectalk-server
|
|
sudo systemctl start dectalk-server
|
|
sudo systemctl status dectalk-server
|
|
```
|
|
|
|
### Using PM2
|
|
|
|
```bash
|
|
npm install -g pm2
|
|
pm2 start server.js --name dectalk-server
|
|
pm2 save
|
|
pm2 startup
|
|
```
|
|
|
|
## Performance Notes
|
|
|
|
- Each TTS request generates a new MP3 file
|
|
- Files are saved to `output/` directory
|
|
- Consider adding cleanup for old files
|
|
- DECTalk is fast and lightweight compared to AI TTS
|
|
|
|
## Comparison with local-tts
|
|
|
|
| Feature | local-tts | dectalk-server |
|
|
|---------|-----------|----------------|
|
|
| Voice Quality | Natural (AI) | Robotic (classic) |
|
|
| Speed | Slower | Faster |
|
|
| Resources | High (GPU recommended) | Low |
|
|
| Voice Cloning | Yes | No |
|
|
| API Key | No | No |
|
|
| Setup Complexity | High | Low |
|
|
|
|
## Next Steps
|
|
|
|
- Replace `say` command with actual DECTalk binary for authentic voice
|
|
- Add voice parameter support (if using multi-voice DECTalk)
|
|
- Implement speed control
|
|
- Add audio file cleanup
|
|
- Add rate limiting
|