← Docs · Markdown · Agent register · Home

GLC-TTS

Hungarian text → speech (GLC-TTS, CPU, zero-shot voice clone).

Not fixed speakers (no Anna/Béla). You pick a preset voice (tts_list_voices) or supply ref_audio_url + ref_text (or transcribe_ref=true → platform Whisper for ref only).

Flow: 1. tts_list_languages / tts_list_voices 2. Optional: POST /staging/upload for custom reference audio → url 3. tts_job_start({ gen_text, voice_id, speed? })job_id (speed 0.3–2.0, default 1.0; lower = slower) 4. Poll tts_job_statustts_job_resulturl (WAV, 24 kHz mono)

Speed: omit or 1.0 = default/normal; 0.3 = slowest; below 1.0 = slower (e.g. 0.6); above 1.0 = faster (up to 2.0).

Max 1 concurrent TTS job; others queued. Whisper for ref_text is this platform only (no local STT fallback).

Service id: tts
Version: 0.1.0
Status: available

Authentication

MCP endpoint: https://mcp.glc-rag.hu/mcp (streamable HTTP)

Agents (recommended): self-register with account_type=agent to get an auto-approved token — see https://mcp.glc-rag.hu/guide/agent.

Or register as a human on the public site (all listed services are auto-approved), wait for system-admin approval, then create a token.

Authorization: Bearer mcp_...

Cursor mcp.json example:

{
  "mcpServers": {
    "tts": {
      "url": "https://mcp.glc-rag.hu/mcp",
      "headers": {
        "Authorization": "Bearer mcp_YOUR_TOKEN"
      }
    }
  }
}

Tools

tts_list_languages

List supported GLC-TTS languages (hu, en).

Input schema:

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}

Examples:

{}

tts_list_voices

List preset voices (id, language, label). For presets call tts_job_start with voice_id only — do not pass ref_audio_url for listed presets.

Input schema:

{
  "type": "object",
  "properties": {
    "language": {
      "type": "string",
      "description": "Optional filter, e.g. hu"
    }
  },
  "additionalProperties": false
}

Examples:

{
  "language": "hu"
}

tts_job_start

Start async Hungarian TTS. Returns {job_id} — not the WAV. Provide gen_text + (voice_id OR ref_audio_url with ref_text OR ref_audio_url with transcribe_ref=true). Optional speed 0.3–2.0 (default 1.0; lower=slower). Then tts_job_status → tts_job_result.

Input schema:

{
  "type": "object",
  "properties": {
    "gen_text": {
      "type": "string",
      "description": "Text to speak (match voice language: hu or en). Not the reference transcript. Long text (up to 100000 chars) is split into sentences server-side and returned as one WAV."
    },
    "voice_id": {
      "type": "string",
      "description": "Preset from tts_list_voices (e.g. pattila, jack)."
    },
    "ref_audio_url": {
      "type": "string",
      "description": "Public https URL of reference speech (5\u201315s). Large/local: POST /staging/upload first."
    },
    "ref_text": {
      "type": "string",
      "description": "Exact transcript of ref_audio (not gen_text). Omit only if transcribe_ref=true or preset default."
    },
    "transcribe_ref": {
      "type": "boolean",
      "description": "If true, platform Whisper (mcp.glc-rag.hu) produces ref_text from ref_audio_url. Default false unless preset says otherwise.",
      "default": false
    },
    "language": {
      "type": "string",
      "description": "hu or en \u2014 selects the TTS checkpoint (hu fine-tune vs en base). If omitted with voice_id, taken from the preset (pattila\u2192hu, jack\u2192en).",
      "default": "hu"
    },
    "speed": {
      "type": "number",
      "description": "Speech rate. Default 1.0. 0.3 = slowest, 1.0 = default/normal, higher = faster (up to 2.0).",
      "default": 1.0,
      "minimum": 0.3,
      "maximum": 2.0
    }
  },
  "required": [
    "gen_text"
  ],
  "additionalProperties": false
}

Examples:

{
  "gen_text": "Szia, ez egy teszt.",
  "voice_id": "pattila"
}
{
  "gen_text": "Hello, this is a test.",
  "voice_id": "jack"
}
{
  "gen_text": "Lassabb felolvas\u00e1s.",
  "voice_id": "pattila",
  "speed": 0.3
}
{
  "gen_text": "Alap\u00e9rtelmezett temp\u00f3 (mint speed n\u00e9lk\u00fcl).",
  "voice_id": "pattila",
  "speed": 1.0
}
{
  "gen_text": "Gyorsabb felolvas\u00e1s.",
  "voice_id": "pattila",
  "speed": 1.5
}
{
  "gen_text": "M\u00e1sodik mondat.",
  "ref_audio_url": "https://mcp.glc-rag.hu/staging/\u2026",
  "transcribe_ref": true
}

tts_job_status

Poll TTS job. status: queued|running|completed|failed. Backoff 15–60s.

Input schema:

{
  "type": "object",
  "properties": {
    "job_id": {
      "type": "string",
      "description": "From tts_job_start"
    }
  },
  "required": [
    "job_id"
  ],
  "additionalProperties": false
}

Examples:

{
  "job_id": "0123456789abcdef0123456789abcdef"
}

tts_job_result

Fetch completed TTS result: {url, bytes, sample_rate, …}. url is staging public WAV. Not ready while queued/running.

Input schema:

{
  "type": "object",
  "properties": {
    "job_id": {
      "type": "string",
      "description": "From tts_job_start"
    }
  },
  "required": [
    "job_id"
  ],
  "additionalProperties": false
}

Examples:

{
  "job_id": "0123456789abcdef0123456789abcdef"
}

Usage notes

Presets: voice_id: "pattila" (hu) or voice_id: "jack" (en); omit speed1.0. speed: 0.3 slowest … 1.0 default … 2.0 fastest. Long gen_text is split into sentences server-side, then concatenated into one WAV (same job_id / result.url). Custom: staging upload ref → tts_job_start({ gen_text, ref_audio_url, transcribe_ref: true }) or pass exact ref_text. Need tts (+ staging/whisper if custom ref / transcribe_ref). CPU inference is slower than realtime — poll with backoff.

Errors / limits

Queue full / unknown voice / missing ref_text without transcribe_ref → reject. Whisper failure for ref → TTS fails (no local fallback). Expired job_id → error.

Agent discovery