Add ttstoy cog
This commit is contained in:
@@ -0,0 +1,237 @@
|
||||
# DECTalk TTS Server
|
||||
|
||||
A simple Node.js server that provides DECTalk text-to-speech with MiniMax-compatible API endpoints for RedBot TTSTOY cog.
|
||||
|
||||
## Features
|
||||
|
||||
- Authentic DECTalk voice (Perfect Paul - same as Moonbase Alpha)
|
||||
- MiniMax-compatible API endpoint (`/v1/t2a_v2`) for seamless integration with TTSTOY
|
||||
- Simple GET endpoint (`/say?text=<text>`) for testing
|
||||
- Automatic WAV to MP3 conversion using ffmpeg
|
||||
- Health check endpoint (`/health`)
|
||||
|
||||
## Requirements
|
||||
|
||||
- Node.js (v14 or higher)
|
||||
- npm
|
||||
- ffmpeg (for audio conversion)
|
||||
- Linux system (DECTalk package doesn't support macOS)
|
||||
|
||||
### Installing Dependencies
|
||||
|
||||
**Ubuntu/Debian:**
|
||||
```bash
|
||||
sudo apt install nodejs npm ffmpeg
|
||||
```
|
||||
|
||||
**Fedora:**
|
||||
```bash
|
||||
sudo dnf install nodejs npm ffmpeg
|
||||
```
|
||||
|
||||
**Arch Linux:**
|
||||
```bash
|
||||
sudo pacman -S nodejs npm ffmpeg
|
||||
```
|
||||
|
||||
**macOS:**
|
||||
```bash
|
||||
brew install node ffmpeg
|
||||
# macOS has built-in 'say' command, no need for espeak
|
||||
```
|
||||
|
||||
## Installation
|
||||
|
||||
```bash
|
||||
cd dectalk-server
|
||||
npm install
|
||||
```
|
||||
|
||||
This will install:
|
||||
- `express` - Web server framework
|
||||
- `dectalk` - Authentic DECTalk voice synthesis (Moonbase Alpha)
|
||||
|
||||
The `dectalk` package includes the actual DECTalk binaries and will work on Linux systems.
|
||||
|
||||
## Usage
|
||||
|
||||
### Start the server
|
||||
|
||||
```bash
|
||||
npm start
|
||||
```
|
||||
|
||||
The server will run on port 3000 by default.
|
||||
|
||||
### Custom port
|
||||
|
||||
```bash
|
||||
PORT=8080 npm start
|
||||
```
|
||||
|
||||
### Custom output directory
|
||||
|
||||
```bash
|
||||
OUTPUT_DIR=/path/to/output npm start
|
||||
```
|
||||
|
||||
## API Endpoints
|
||||
|
||||
### Health Check
|
||||
```bash
|
||||
GET /health
|
||||
```
|
||||
|
||||
Returns:
|
||||
```json
|
||||
{"status": "ok"}
|
||||
```
|
||||
|
||||
### Engine Info
|
||||
```bash
|
||||
GET /engine
|
||||
```
|
||||
|
||||
Returns:
|
||||
```json
|
||||
{
|
||||
"type": "dectalk",
|
||||
"version": "1.0.0",
|
||||
"description": "DECTalk Text-to-Speech Engine"
|
||||
}
|
||||
```
|
||||
|
||||
### Simple TTS (GET)
|
||||
```bash
|
||||
GET /say?text=Hello%20world
|
||||
```
|
||||
|
||||
Returns: MP3 audio file
|
||||
|
||||
### MiniMax-Compatible TTS (POST)
|
||||
```bash
|
||||
POST /v1/t2a_v2
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"text": "Hello world",
|
||||
"voice_setting": {
|
||||
"voice_id": "default",
|
||||
"speed": 1.0
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Returns:
|
||||
```json
|
||||
{
|
||||
"base_resp": {
|
||||
"status_code": 0,
|
||||
"status_msg": "Success"
|
||||
},
|
||||
"data": {
|
||||
"audio": "<hex-encoded-mp3>",
|
||||
"audio_id": "<unique-id>"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Integration with RedBot TTSTOY
|
||||
|
||||
1. Start the DECTalk server:
|
||||
```bash
|
||||
npm start
|
||||
```
|
||||
|
||||
2. Configure TTSTOY to use DECTalk mode:
|
||||
```
|
||||
[p]ttstoy mode dectalk
|
||||
[p]ttstoy dectalkurl http://127.0.0.1:3000
|
||||
```
|
||||
|
||||
3. Test it:
|
||||
```
|
||||
[p]tts Hello, I am DECTalk
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
- Uses authentic DECTalk (Perfect Paul voice from Moonbase Alpha)
|
||||
- The `dectalk` npm package includes the actual DECTalk binaries
|
||||
- Audio files are saved to the `output/` directory
|
||||
- MP3 format: mono, 32kHz, 128kbps (matches TTSTOY expectations)
|
||||
- Supports all 9 original DECTalk voices
|
||||
- Supports DECTalk phoneme commands (e.g., `[:phoneme on]`, `[:rate 200]`)
|
||||
|
||||
## DECTalk Commands
|
||||
|
||||
DECTalk supports special commands for controlling speech:
|
||||
|
||||
```
|
||||
[:phoneme on] - Enable phoneme mode
|
||||
[:rate 200] - Set speech rate (default 180)
|
||||
[:dv ap 100] - Set voice parameters
|
||||
```
|
||||
|
||||
Example:
|
||||
```bash
|
||||
curl "http://127.0.0.1:3001/say?text=%5B:rate%20150%5DHello%20world&voice=paul" -o test.mp3
|
||||
```
|
||||
|
||||
You can also use these in Discord:
|
||||
```
|
||||
[p]tts [:rate 200] speaking very fast
|
||||
[p]tts [:phoneme on] hehlow werld
|
||||
```
|
||||
|
||||
## DECTalk Voices
|
||||
|
||||
DECTalk includes 9 classic voices:
|
||||
|
||||
| Voice | Name | Description |
|
||||
|-------|------|-------------|
|
||||
| Paul | Perfect Paul | Default male voice (Moonbase Alpha, Stephen Hawking) |
|
||||
| Betty | Beautiful Betty | Female voice |
|
||||
| Harry | Huge Harry | Deep male voice |
|
||||
| Frank | Frail Frank | Elderly male voice |
|
||||
| Dennis | Doctor Dennis | Male voice |
|
||||
| Kit | Kit the Kid | Child voice |
|
||||
| Ursula | Uppity Ursula | Female voice |
|
||||
| Rita | Rough Rita | Female voice |
|
||||
| Wendy | Whispering Wendy | Soft female voice |
|
||||
|
||||
### Using Different Voices
|
||||
|
||||
**In Discord:**
|
||||
```
|
||||
[p]ttstoy voice Paul
|
||||
[p]tts Hello, I am Perfect Paul
|
||||
[p]ttstoy voice Betty
|
||||
[p]tts Hello, I am Beautiful Betty
|
||||
```
|
||||
|
||||
**Direct API:**
|
||||
```bash
|
||||
curl "http://127.0.0.1:3001/say?text=Hello&voice=paul" -o paul.mp3
|
||||
curl "http://127.0.0.1:3001/say?text=Hello&voice=betty" -o betty.mp3
|
||||
```
|
||||
|
||||
**List available voices:**
|
||||
```bash
|
||||
curl http://127.0.0.1:3001/voices
|
||||
```
|
||||
|
||||
## Customization
|
||||
|
||||
The server uses the default DECTalk voice (Perfect Paul). The `dectalk` package handles all the voice synthesis internally, so no additional configuration is needed for the authentic Moonbase Alpha sound.
|
||||
|
||||
## Development
|
||||
|
||||
Run with auto-reload:
|
||||
```bash
|
||||
npm run dev
|
||||
```
|
||||
|
||||
## License
|
||||
|
||||
MIT
|
||||
Reference in New Issue
Block a user