Initial standalone dectalk TTS server extracted from the ttstoy bot cog
Self-contained HTTP service with engine, data, start script, systemd unit, and documentation. Runs independently of the Discord bot on its fixed port.
This commit is contained in:
@@ -0,0 +1,4 @@
|
||||
# Port the DECTalk TTS server listens on
|
||||
PORT=33001
|
||||
# Directory for cached mp3 output
|
||||
OUTPUT_DIR=./output
|
||||
@@ -0,0 +1,4 @@
|
||||
node_modules/
|
||||
output/
|
||||
*.log
|
||||
.env
|
||||
+43
@@ -0,0 +1,43 @@
|
||||
# Installing the DECTalk TTS Server as a system service
|
||||
|
||||
These steps install the server as a systemd service that starts on boot and
|
||||
restarts automatically if it crashes. Commands that need root are shown with
|
||||
`sudo`; run them in your own terminal.
|
||||
|
||||
## 1. Install prerequisites
|
||||
|
||||
```bash
|
||||
sudo apt install ffmpeg nodejs npm
|
||||
```
|
||||
|
||||
## 2. Install node dependencies
|
||||
|
||||
```bash
|
||||
cd /home/owen/dectalk-tts-server
|
||||
npm install
|
||||
```
|
||||
|
||||
## 3. Install the systemd unit
|
||||
|
||||
```bash
|
||||
sudo cp /home/owen/dectalk-tts-server/dectalk-tts-server.service /etc/systemd/system/
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl enable --now dectalk-tts-server
|
||||
```
|
||||
|
||||
## 4. Verify
|
||||
|
||||
```bash
|
||||
systemctl status dectalk-tts-server
|
||||
curl -s http://127.0.0.1:33001/health
|
||||
```
|
||||
|
||||
The health endpoint should return `{"status":"ok"}`.
|
||||
|
||||
## Managing the service
|
||||
|
||||
```bash
|
||||
sudo systemctl restart dectalk-tts-server
|
||||
sudo systemctl stop dectalk-tts-server
|
||||
journalctl -u dectalk-tts-server -f
|
||||
```
|
||||
@@ -0,0 +1,102 @@
|
||||
# DECTalk TTS Server
|
||||
|
||||
A standalone HTTP server that turns text into speech using authentic DECTalk
|
||||
(the Moonbase Alpha voice). It exposes both a simple GET endpoint and a
|
||||
MiniMax-compatible POST endpoint, so it can be used directly or as a drop-in
|
||||
TTS backend.
|
||||
|
||||
This service was extracted from the ttstoy Discord bot cog so it can run on its
|
||||
own as a system service, independent of the bot.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Node.js (v16 or newer recommended)
|
||||
- ffmpeg (used to encode output to mp3)
|
||||
- npm dependencies: `express`, `dectalk` (installed via `npm install`)
|
||||
|
||||
Install ffmpeg on Debian/Ubuntu:
|
||||
|
||||
```bash
|
||||
sudo apt install ffmpeg
|
||||
```
|
||||
|
||||
## Running
|
||||
|
||||
```bash
|
||||
./start.sh
|
||||
```
|
||||
|
||||
`start.sh` runs `npm install` on first launch and then starts the server. The
|
||||
listening port defaults to `33001` and can be overridden with the `PORT`
|
||||
environment variable.
|
||||
|
||||
You can also run it directly:
|
||||
|
||||
```bash
|
||||
npm install
|
||||
PORT=33001 node server.js
|
||||
```
|
||||
|
||||
## API
|
||||
|
||||
### GET /say
|
||||
|
||||
Generate speech and return an mp3.
|
||||
|
||||
```
|
||||
GET /say?text=Hello%20world
|
||||
```
|
||||
|
||||
Response: `audio/mpeg` (mp3 bytes).
|
||||
|
||||
Voice is controlled inline using DECTalk voice commands (see below).
|
||||
|
||||
### POST /v1/t2a_v2
|
||||
|
||||
MiniMax-compatible endpoint. Accepts JSON `{ "text": "..." }` and returns a
|
||||
JSON body with the audio encoded as hex under `data.audio`.
|
||||
|
||||
```bash
|
||||
curl -X POST http://127.0.0.1:33001/v1/t2a_v2 \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"text":"Hello from DECTalk"}'
|
||||
```
|
||||
|
||||
### GET /voices
|
||||
|
||||
Returns the list of available DECTalk voices and their inline command codes.
|
||||
|
||||
### GET /engine
|
||||
|
||||
Returns engine type and version metadata.
|
||||
|
||||
### GET /health
|
||||
|
||||
Returns `{ "status": "ok" }`.
|
||||
|
||||
## Voice commands
|
||||
|
||||
Voices are selected inline by prefixing the text with a command code:
|
||||
|
||||
| Voice | Command | Description |
|
||||
|--------|---------|-------------------------|
|
||||
| Paul | `[:np]` | Perfect Paul (default) |
|
||||
| Betty | `[:nb]` | Beautiful Betty |
|
||||
| Harry | `[:nh]` | Huge Harry (deep male) |
|
||||
| Frank | `[:nf]` | Frail Frank (elderly) |
|
||||
| Dennis | `[:nd]` | Doctor Dennis |
|
||||
| Kit | `[:nk]` | Kit the Kid (child) |
|
||||
| Ursula | `[:nu]` | Uppity Ursula |
|
||||
| Rita | `[:nr]` | Rough Rita |
|
||||
| Wendy | `[:nw]` | Whispering Wendy (soft) |
|
||||
|
||||
Example:
|
||||
|
||||
```
|
||||
GET /say?text=[:nh]This is Huge Harry speaking.
|
||||
```
|
||||
|
||||
## Running as a system service
|
||||
|
||||
A systemd unit file is provided (`dectalk-tts-server.service`). See INSTALL.md
|
||||
for setup steps.
|
||||
@@ -0,0 +1,16 @@
|
||||
[Unit]
|
||||
Description=DECTalk TTS Server
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=owen
|
||||
WorkingDirectory=/home/owen/dectalk-tts-server
|
||||
Environment=PORT=33001
|
||||
ExecStart=/home/owen/dectalk-tts-server/start.sh
|
||||
Restart=always
|
||||
RestartSec=3
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
Generated
+1125
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,20 @@
|
||||
{
|
||||
"name": "dectalk-server",
|
||||
"version": "1.0.0",
|
||||
"description": "DECTalk TTS server compatible with RedBot TTSTOY cog",
|
||||
"main": "server.js",
|
||||
"scripts": {
|
||||
"start": "node server.js",
|
||||
"dev": "nodemon server.js"
|
||||
},
|
||||
"keywords": ["dectalk", "tts", "text-to-speech", "redbot", "moonbase-alpha"],
|
||||
"author": "",
|
||||
"license": "MIT",
|
||||
"dependencies": {
|
||||
"express": "^4.18.2",
|
||||
"dectalk": "^1.0.0"
|
||||
},
|
||||
"devDependencies": {
|
||||
"nodemon": "^3.0.1"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,209 @@
|
||||
#!/usr/bin/env node
|
||||
|
||||
/**
|
||||
* DECTalk TTS Server
|
||||
* Compatible with RedBot TTSTOY cog
|
||||
*
|
||||
* Uses authentic DECTalk (Moonbase Alpha voice)
|
||||
* Provides MiniMax-compatible API endpoint at /v1/t2a_v2
|
||||
* Also provides simple GET endpoint at /say?text=<text>
|
||||
*/
|
||||
|
||||
const express = require('express');
|
||||
const { say } = require('dectalk');
|
||||
const { spawn } = require('child_process');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const crypto = require('crypto');
|
||||
|
||||
const app = express();
|
||||
const PORT = process.env.PORT || 33001;
|
||||
const OUTPUT_DIR = process.env.OUTPUT_DIR || path.join(__dirname, 'output');
|
||||
|
||||
// Ensure output directory exists
|
||||
if (!fs.existsSync(OUTPUT_DIR)) {
|
||||
fs.mkdirSync(OUTPUT_DIR, { recursive: true });
|
||||
}
|
||||
|
||||
app.use(express.json());
|
||||
|
||||
// Health check endpoint
|
||||
app.get('/health', (req, res) => {
|
||||
res.json({ status: 'ok' });
|
||||
});
|
||||
|
||||
// Engine info endpoint
|
||||
app.get('/engine', (req, res) => {
|
||||
res.json({
|
||||
type: 'dectalk',
|
||||
version: '1.0.0',
|
||||
description: 'DECTalk Text-to-Speech Engine'
|
||||
});
|
||||
});
|
||||
|
||||
// Voices endpoint - list available DECTalk voice commands
|
||||
app.get('/voices', (req, res) => {
|
||||
const voices = [
|
||||
{ voice_id: 'paul', command: '[:np]', description: 'Perfect Paul (default, male)' },
|
||||
{ voice_id: 'betty', command: '[:nb]', description: 'Beautiful Betty (female)' },
|
||||
{ voice_id: 'harry', command: '[:nh]', description: 'Huge Harry (deep male)' },
|
||||
{ voice_id: 'frank', command: '[:nf]', description: 'Frail Frank (elderly male)' },
|
||||
{ voice_id: 'dennis', command: '[:nd]', description: 'Doctor Dennis (male)' },
|
||||
{ voice_id: 'kit', command: '[:nk]', description: 'Kit the Kid (child)' },
|
||||
{ voice_id: 'ursula', command: '[:nu]', description: 'Uppity Ursula (female)' },
|
||||
{ voice_id: 'rita', command: '[:nr]', description: 'Rough Rita (gravelly female)' },
|
||||
{ voice_id: 'wendy', command: '[:nw]', description: 'Whispering Wendy (soft female)' }
|
||||
];
|
||||
res.json(voices);
|
||||
});
|
||||
|
||||
// Simple GET endpoint for DECTalk
|
||||
app.get('/say', async (req, res) => {
|
||||
const text = req.query.text;
|
||||
|
||||
if (!text) {
|
||||
return res.status(400).send('Missing text parameter');
|
||||
}
|
||||
|
||||
try {
|
||||
const wavBuffer = await generateDECTalk(text);
|
||||
|
||||
// Convert WAV to MP3 using ffmpeg
|
||||
const mp3Buffer = await convertToMP3(wavBuffer);
|
||||
|
||||
res.set('Content-Type', 'audio/mpeg');
|
||||
res.send(mp3Buffer);
|
||||
} catch (error) {
|
||||
console.error('DECTalk generation error:', error);
|
||||
res.status(500).send(`TTS generation failed: ${error.message}`);
|
||||
}
|
||||
});
|
||||
|
||||
// MiniMax-compatible endpoint for RedBot TTSTOY cog
|
||||
app.post('/v1/t2a_v2', async (req, res) => {
|
||||
const { text } = req.body;
|
||||
|
||||
if (!text) {
|
||||
return res.json({
|
||||
base_resp: {
|
||||
status_code: 1002,
|
||||
status_msg: 'Text is required'
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
console.log(`Generating DECTalk TTS`);
|
||||
console.log(`Text: ${text.substring(0, 100)}...`);
|
||||
|
||||
try {
|
||||
// Generate DECTalk audio - voice is controlled by [:n*] commands in text
|
||||
const wavBuffer = await generateDECTalk(text);
|
||||
|
||||
// Convert WAV to MP3
|
||||
const mp3Buffer = await convertToMP3(wavBuffer);
|
||||
|
||||
// Save to file
|
||||
const audioId = crypto.randomBytes(16).toString('hex');
|
||||
const mp3Path = path.join(OUTPUT_DIR, `${audioId}.mp3`);
|
||||
fs.writeFileSync(mp3Path, mp3Buffer);
|
||||
|
||||
// Convert to hex (MiniMax format)
|
||||
const audioHex = mp3Buffer.toString('hex');
|
||||
|
||||
console.log(`✅ Generated audio: ${audioId}.mp3 (${mp3Buffer.length} bytes)`);
|
||||
|
||||
// Return MiniMax-compatible response
|
||||
res.json({
|
||||
base_resp: {
|
||||
status_code: 0,
|
||||
status_msg: 'Success'
|
||||
},
|
||||
data: {
|
||||
audio: audioHex,
|
||||
audio_id: audioId
|
||||
}
|
||||
});
|
||||
} catch (error) {
|
||||
console.error('DECTalk generation error:', error);
|
||||
res.json({
|
||||
base_resp: {
|
||||
status_code: 1005,
|
||||
status_msg: `TTS generation failed: ${error.message}`
|
||||
}
|
||||
});
|
||||
}
|
||||
});
|
||||
|
||||
/**
|
||||
* Generate authentic DECTalk audio (Moonbase Alpha voice)
|
||||
* @param {string} text - Text to synthesize (can include [:n*] voice commands)
|
||||
* @returns {Promise<Buffer>} WAV audio buffer
|
||||
*/
|
||||
async function generateDECTalk(text) {
|
||||
try {
|
||||
// Use the authentic DECTalk package
|
||||
// Voice is controlled by [:n*] commands in the text itself
|
||||
const wavBuffer = await say(text);
|
||||
return wavBuffer;
|
||||
} catch (error) {
|
||||
throw new Error(`DECTalk generation failed: ${error.message}`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Convert WAV buffer to MP3 using ffmpeg
|
||||
* @param {Buffer} wavBuffer - WAV audio buffer
|
||||
* @returns {Promise<Buffer>} MP3 audio buffer
|
||||
*/
|
||||
function convertToMP3(wavBuffer) {
|
||||
return new Promise((resolve, reject) => {
|
||||
const ffmpeg = spawn('ffmpeg', [
|
||||
'-f', 'wav',
|
||||
'-i', 'pipe:0',
|
||||
'-f', 'mp3',
|
||||
'-ac', '1',
|
||||
'-ar', '32000',
|
||||
'-b:a', '128k',
|
||||
'pipe:1'
|
||||
]);
|
||||
|
||||
const chunks = [];
|
||||
|
||||
ffmpeg.stdout.on('data', (chunk) => {
|
||||
chunks.push(chunk);
|
||||
});
|
||||
|
||||
ffmpeg.stderr.on('data', (data) => {
|
||||
// ffmpeg outputs progress to stderr, ignore it
|
||||
});
|
||||
|
||||
ffmpeg.on('close', (code) => {
|
||||
if (code !== 0) {
|
||||
reject(new Error(`ffmpeg process exited with code ${code}`));
|
||||
} else {
|
||||
resolve(Buffer.concat(chunks));
|
||||
}
|
||||
});
|
||||
|
||||
ffmpeg.on('error', (err) => {
|
||||
reject(new Error(`Failed to start ffmpeg: ${err.message}`));
|
||||
});
|
||||
|
||||
// Write WAV data to ffmpeg stdin
|
||||
ffmpeg.stdin.write(wavBuffer);
|
||||
ffmpeg.stdin.end();
|
||||
});
|
||||
}
|
||||
|
||||
// Start server
|
||||
app.listen(PORT, () => {
|
||||
console.log(`🤖 DECTalk TTS Server running on http://localhost:${PORT}`);
|
||||
console.log(`Health check: http://localhost:${PORT}/health`);
|
||||
console.log(`Voices: http://localhost:${PORT}/voices`);
|
||||
console.log(`Simple API: http://localhost:${PORT}/say?text=Hello`);
|
||||
console.log(`MiniMax-compatible API: POST http://localhost:${PORT}/v1/t2a_v2`);
|
||||
console.log('');
|
||||
console.log('Voice commands (use in text):');
|
||||
console.log(' [:np] Paul [:nb] Betty [:nh] Harry [:nf] Frank [:nd] Dennis');
|
||||
console.log(' [:nk] Kit [:nu] Ursula [:nr] Rita [:nw] Wendy');
|
||||
});
|
||||
@@ -0,0 +1,20 @@
|
||||
#!/bin/bash
|
||||
# Start script for the DECTalk TTS server.
|
||||
# Installs node dependencies on first run, then runs the server.
|
||||
set -e
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
|
||||
cd "$SCRIPT_DIR"
|
||||
|
||||
if ! command -v ffmpeg >/dev/null 2>&1; then
|
||||
echo "Error: ffmpeg is not installed. Install it with: sudo apt install ffmpeg"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [ ! -d node_modules ]; then
|
||||
echo "Installing node dependencies..."
|
||||
npm install
|
||||
fi
|
||||
|
||||
export PORT="${PORT:-33001}"
|
||||
exec node server.js
|
||||
Executable
+48
@@ -0,0 +1,48 @@
|
||||
#!/bin/bash
|
||||
|
||||
# Test script for DECTalk server API
|
||||
|
||||
BASE_URL="${1:-http://127.0.0.1:3000}"
|
||||
|
||||
echo "🧪 Testing DECTalk Server at $BASE_URL"
|
||||
echo ""
|
||||
|
||||
# Test 1: Health check
|
||||
echo "1️⃣ Testing /health endpoint..."
|
||||
curl -s "$BASE_URL/health" | jq '.' || echo "❌ Health check failed"
|
||||
echo ""
|
||||
|
||||
# Test 2: Engine info
|
||||
echo "2️⃣ Testing /engine endpoint..."
|
||||
curl -s "$BASE_URL/engine" | jq '.' || echo "❌ Engine info failed"
|
||||
echo ""
|
||||
|
||||
# Test 3: Simple GET endpoint
|
||||
echo "3️⃣ Testing /say endpoint (saving to test_say.mp3)..."
|
||||
curl -s "$BASE_URL/say?text=Hello%20world" -o test_say.mp3
|
||||
if [ -f test_say.mp3 ] && [ -s test_say.mp3 ]; then
|
||||
echo "✅ Audio saved to test_say.mp3 ($(stat -f%z test_say.mp3 2>/dev/null || stat -c%s test_say.mp3) bytes)"
|
||||
else
|
||||
echo "❌ Failed to generate audio"
|
||||
fi
|
||||
echo ""
|
||||
|
||||
# Test 4: MiniMax-compatible endpoint
|
||||
echo "4️⃣ Testing /v1/t2a_v2 endpoint (MiniMax-compatible)..."
|
||||
curl -s -X POST "$BASE_URL/v1/t2a_v2" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"text": "Testing MiniMax compatibility",
|
||||
"voice_setting": {
|
||||
"voice_id": "default",
|
||||
"speed": 1.0
|
||||
}
|
||||
}' | jq '.base_resp, .data.audio_id' || echo "❌ MiniMax endpoint failed"
|
||||
echo ""
|
||||
|
||||
echo "✅ All tests complete!"
|
||||
echo ""
|
||||
echo "To test audio playback:"
|
||||
echo " ffplay test_say.mp3"
|
||||
echo " # or"
|
||||
echo " mpv test_say.mp3"
|
||||
Reference in New Issue
Block a user