Add ttstoy cog
This commit is contained in:
@@ -0,0 +1,186 @@
|
||||
# Integration with RedBot TTSTOY
|
||||
|
||||
This guide shows how to integrate the DECTalk server with your RedBot TTSTOY cog.
|
||||
|
||||
## Quick Start
|
||||
|
||||
### 1. Install and Start DECTalk Server
|
||||
|
||||
```bash
|
||||
cd dectalk-server
|
||||
npm install
|
||||
npm start
|
||||
```
|
||||
|
||||
The server will start on `http://127.0.0.1:3000`
|
||||
|
||||
### 2. Configure TTSTOY
|
||||
|
||||
In your Discord server, run these commands:
|
||||
|
||||
```
|
||||
[p]ttstoy mode dectalk
|
||||
[p]ttstoy dectalkurl http://127.0.0.1:3000
|
||||
```
|
||||
|
||||
### 3. Test It
|
||||
|
||||
```
|
||||
[p]tts Hello, I am DECTalk, a text to speech system.
|
||||
```
|
||||
|
||||
## API Compatibility
|
||||
|
||||
The DECTalk server implements the same API format as your local-tts server:
|
||||
|
||||
### Endpoints
|
||||
|
||||
| Endpoint | Method | Purpose | Compatible With |
|
||||
|----------|--------|---------|-----------------|
|
||||
| `/health` | GET | Health check | local-tts |
|
||||
| `/engine` | GET | Engine info | local-tts |
|
||||
| `/say?text=<text>` | GET | Simple TTS | DECTalk-specific |
|
||||
| `/v1/t2a_v2` | POST | MiniMax-compatible TTS | TTSTOY cog |
|
||||
|
||||
### Response Format
|
||||
|
||||
The `/v1/t2a_v2` endpoint returns the exact same format as MiniMax API:
|
||||
|
||||
```json
|
||||
{
|
||||
"base_resp": {
|
||||
"status_code": 0,
|
||||
"status_msg": "Success"
|
||||
},
|
||||
"data": {
|
||||
"audio": "<hex-encoded-mp3-data>",
|
||||
"audio_id": "<unique-identifier>"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This matches the format expected by `_save_minimax_tts()` in your TTSTOY cog.
|
||||
|
||||
## How It Works
|
||||
|
||||
1. **Text Input**: TTSTOY sends text via POST to `/v1/t2a_v2`
|
||||
2. **DECTalk Generation**: Server generates WAV audio using DECTalk
|
||||
3. **MP3 Conversion**: ffmpeg converts WAV to MP3 (mono, 32kHz, 128kbps)
|
||||
4. **Hex Encoding**: MP3 is converted to hex string (MiniMax format)
|
||||
5. **Response**: Server returns hex-encoded audio in MiniMax-compatible JSON
|
||||
|
||||
## Switching Between TTS Modes
|
||||
|
||||
You can easily switch between different TTS engines:
|
||||
|
||||
```bash
|
||||
# Use MiniMax cloud API
|
||||
[p]ttstoy mode minimax
|
||||
|
||||
# Use local AI TTS
|
||||
[p]ttstoy mode local
|
||||
[p]ttstoy localurl http://127.0.0.1:8000
|
||||
|
||||
# Use DECTalk
|
||||
[p]ttstoy mode dectalk
|
||||
[p]ttstoy dectalkurl http://127.0.0.1:3000
|
||||
```
|
||||
|
||||
## Features Supported
|
||||
|
||||
✅ Text-to-speech synthesis
|
||||
✅ Emoji SFX (handled by TTSTOY cog)
|
||||
✅ Audio concatenation (handled by TTSTOY cog)
|
||||
✅ Volume control (handled by TTSTOY cog)
|
||||
❌ Voice selection (DECTalk has one voice)
|
||||
❌ Speed control (not implemented yet)
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Server won't start
|
||||
|
||||
Check if port 3000 is already in use:
|
||||
```bash
|
||||
lsof -i :3000
|
||||
```
|
||||
|
||||
Use a different port:
|
||||
```bash
|
||||
PORT=8080 npm start
|
||||
```
|
||||
|
||||
### "Cannot reach DECTalk API"
|
||||
|
||||
1. Check if server is running:
|
||||
```bash
|
||||
curl http://127.0.0.1:3000/health
|
||||
```
|
||||
|
||||
2. Check server logs for errors
|
||||
|
||||
3. Verify ffmpeg is installed:
|
||||
```bash
|
||||
ffmpeg -version
|
||||
```
|
||||
|
||||
### Audio quality issues
|
||||
|
||||
The server outputs MP3 at:
|
||||
- Sample rate: 32kHz
|
||||
- Bitrate: 128kbps
|
||||
- Channels: Mono
|
||||
|
||||
This matches TTSTOY's expectations. If you need different settings, modify the `convertToMP3()` function in `server.js`.
|
||||
|
||||
## Production Deployment
|
||||
|
||||
### Using systemd
|
||||
|
||||
1. Edit `dectalk-server.service` with your paths
|
||||
2. Copy to systemd:
|
||||
```bash
|
||||
sudo cp dectalk-server.service /etc/systemd/system/
|
||||
```
|
||||
|
||||
3. Enable and start:
|
||||
```bash
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl enable dectalk-server
|
||||
sudo systemctl start dectalk-server
|
||||
sudo systemctl status dectalk-server
|
||||
```
|
||||
|
||||
### Using PM2
|
||||
|
||||
```bash
|
||||
npm install -g pm2
|
||||
pm2 start server.js --name dectalk-server
|
||||
pm2 save
|
||||
pm2 startup
|
||||
```
|
||||
|
||||
## Performance Notes
|
||||
|
||||
- Each TTS request generates a new MP3 file
|
||||
- Files are saved to `output/` directory
|
||||
- Consider adding cleanup for old files
|
||||
- DECTalk is fast and lightweight compared to AI TTS
|
||||
|
||||
## Comparison with local-tts
|
||||
|
||||
| Feature | local-tts | dectalk-server |
|
||||
|---------|-----------|----------------|
|
||||
| Voice Quality | Natural (AI) | Robotic (classic) |
|
||||
| Speed | Slower | Faster |
|
||||
| Resources | High (GPU recommended) | Low |
|
||||
| Voice Cloning | Yes | No |
|
||||
| API Key | No | No |
|
||||
| Setup Complexity | High | Low |
|
||||
|
||||
## Next Steps
|
||||
|
||||
- Replace `say` command with actual DECTalk binary for authentic voice
|
||||
- Add voice parameter support (if using multi-voice DECTalk)
|
||||
- Implement speed control
|
||||
- Add audio file cleanup
|
||||
- Add rate limiting
|
||||
Reference in New Issue
Block a user