Repo Voice & Audio

index-tts/index-tts

Zero-shot TTS system that clones a voice from one reference clip, with emotion, speed and pronunciation control across five languages (IndexTTS-2.5).

  • 24.4k GitHub stars
  • Python
  • ⚖️ NOASSERTION
  • 🎯 Intermediate
git clone https://github.com/index-tts/index-tts.git && cd index-tts
index-tts/index-tts preview image

What it is

IndexTTS is a zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest release, IndexTTS-2.5, supports Chinese, English, Japanese, Spanish and Arabic. It offers fine-grained emotion control, speaking speed control, pronunciation control (Pinyin / CMU phonemes / Japanese Kana), and faster inference than IndexTTS-2. It ships a WebUI and a Python API, and production deployment is supported via vLLM.

Who it's for

  • Developers who need voice cloning from a single reference audio clip
  • Teams building multilingual TTS (Chinese, English, Japanese, Spanish, Arabic)
  • Users who want emotion, speed and pronunciation control over synthesized speech
  • Engineers looking to deploy TTS in production via vLLM

Requirements

Requirements

  • git
  • uv (required for dependency management)
  • NVIDIA CUDA Toolkit 12.8 or newer if a CUDA error appears during installation on Linux/Windows
  • Model checkpoints downloaded from HuggingFace or ModelScope
  • Network access to HuggingFace/ModelScope (a mirror can be set via HF_ENDPOINT)

Setup

  1. Clone the repository

    Make sure git is installed, then download the repository.

    bash
    git clone https://github.com/index-tts/index-tts.git && cd index-tts
  2. Install uv

    uv is required to manage the project's dependency environment.

    bash
    pip install -U uv  # or see the link above for other install methods
  3. Install dependencies

    Creates a .venv project directory and installs the correct Python and all required dependencies.

    bash
    uv sync --all-extras
  4. Download IndexTTS-2.5 model via HuggingFace

    Install the huggingface-hub tool and download the checkpoints.

    bash
    uv tool install "huggingface-hub"
    
    # IndexTTS-2.5
    hf download IndexTeam/IndexTTS-2.5 --local-dir=checkpoints
  5. Check GPU acceleration

    Diagnose your environment and see which GPUs are detected.

    bash
    uv run tools/gpu_check.py
  6. Launch the WebUI

    Then open http://127.0.0.1:7860 in your browser.

    bash
    # IndexTTS-2.5 (default)
    uv run webui.py

Examples

Initialize IndexTTS-2.5

python
python
from indextts.infer_v2_5 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints/config.yaml", model_dir="checkpoints", use_bf16=True)

What it does: Loads the IndexTTS-2.5 model with BF16 inference from the downloaded checkpoints.

Voice cloning from a single reference audio

python
python
text = "Translate for me, what is a surprise!"

# IndexTTS2.5 (multilingual, with language selection)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="EN", output_path="gen.wav", verbose=True)

What it does: Clones the voice in the reference clip and synthesizes the text in English.

Emotion control with a separate emotional reference and emo_alpha

python
python
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"

# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", lang="ZH", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)

What it does: Uses a sad emotional reference audio; emo_alpha (0.0–1.0, default 1.0) sets how strongly it affects the output.

Emotion control with an emotion vector

python
python
text = "对不起嘛!我的记性真的不太好,但是和你在一起的事情,我都会努力记住的~"

# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_09.wav', text=text, lang="ZH", output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)

What it does: The 8-float vector is ordered [happy, angry, sad, afraid, disgusted, melancholic, surprised, calm]; here sad is 0.8.

Speaking speed control

python
python
text = "大家好,欢迎来到IndexTTS的语速控制演示。"

# IndexTTS2.5
# Slow down (1.2x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_slow.wav", duration_factor=1.2, verbose=True)

# Speed up (0.8x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_fast.wav", duration_factor=0.8, verbose=True)

What it does: duration_factor above 1.0 slows speech and below 1.0 speeds it up; valid range 0.5–2.0, default 1.0.

Pros & cons

Pros

  • Pro:Clones a voice from a single reference audio clip
  • Pro:Fine-grained control over emotion (audio, vector, or text-based), speaking speed and pronunciation
  • Pro:IndexTTS-2.5 supports five languages and is documented as faster than IndexTTS-2 (RTF about 0.2065 vs 0.3257 overall on an RTX 4090)
  • Pro:Supports FP16/BF16 inference for lower VRAM use, optional DeepSpeed, and vLLM for production deployment

Cons

  • Con:uv is required for installation, and DeepSpeed may be difficult to install on Windows
  • Con:Random sampling (use_random) reduces voice cloning fidelity
  • Con:For IndexTTS-2.5, use_emo_text=True requires use_qwen_emo=True or it raises a RuntimeError

Images