Repo Voice & Audio

nari-labs/dia

Dia is a 1.6B-parameter English text-to-speech model from Nari Labs that generates multi-speaker dialogue with nonverbal sounds and voice cloning from a transcript.

  • 19.4k GitHub stars
  • Python
  • ⚖️ Apache-2.0
  • 🎯 Intermediate
nari-labs/dia — repo preview

What it is

Dia is a 1.6B parameter text to speech model created by Nari Labs. It directly generates highly realistic dialogue from a transcript using [S1] and [S2] speaker tags, and can be conditioned on audio for emotion and tone control. It can also produce nonverbal sounds such as laughter and coughing. Pretrained checkpoints and inference code are provided, and it currently supports English only.

Who it's for

  • Researchers and developers experimenting with open-weights dialogue text-to-speech
  • Developers who want multi-speaker conversational audio from a scripted transcript
  • Users who want voice cloning via audio prompts
  • Hugging Face Transformers users who want to run Dia through that library

Requirements

Requirements

  • GPU: tested only on GPUs (pytorch 2.0+, CUDA 12.6); CPU support is listed as to be added soon
  • About 4.4GB VRAM for bfloat16/float16 and about 7.9GB for float32 (benchmarked on RTX 4090)
  • For the Transformers path, install the main branch of transformers
  • For 5000 series GPUs, use torch 2.8 nightly (see issue #26)
  • The first run takes longer because the Descript Audio Codec must be downloaded
  • uv is needed only if using the uv-based commands

Setup

  1. Install from a clone (pip)

    Clone the repo, optionally create a virtual environment, and install in editable mode.

    bash
    # Clone this repository
    git clone https://github.com/nari-labs/dia.git
    cd dia
    
    # Optionally
    python -m venv .venv && source .venv/bin/activate
    
    # Install dia
    pip install -e .
  2. Install directly from GitHub

    Install without cloning the repository.

    bash
    # Install directly from GitHub
    pip install git+https://github.com/nari-labs/dia.git
  3. Run the simple example

    Run the bundled example after installing with pip.

    bash
    python example/simple.py
  4. Run with uv

    With uv installed, clone the repo and run the example directly.

    bash
    # Clone this repository
    git clone https://github.com/nari-labs/dia.git
    cd dia
    
    uv run example/simple.py
  5. Install Transformers main branch

    Needed to use the Hugging Face Transformers implementation of Dia.

    bash
    pip install git+https://github.com/huggingface/transformers.git
    # or install with uv
    uv pip install git+https://github.com/huggingface/transformers.git

Examples

Generate dialogue with Transformers

python
python
from transformers import AutoProcessor, DiaForConditionalGeneration


torch_device = "cuda"
model_checkpoint = "nari-labs/Dia-1.6B-0626"

text = [
    "[S1] Dia is an open weights text to dialogue model. [S2] You get full control over scripts and voices. [S1] Wow. Amazing. (laughs) [S2] Try it now on Git hub or Hugging Face."
]
processor = AutoProcessor.from_pretrained(model_checkpoint)
inputs = processor(text=text, padding=True, return_tensors="pt").to(torch_device)

model = DiaForConditionalGeneration.from_pretrained(model_checkpoint).to(torch_device)
outputs = model.generate(
    **inputs, max_new_tokens=3072, guidance_scale=3.0, temperature=1.8, top_p=0.90, top_k=45
)

outputs = processor.batch_decode(outputs)
processor.save_audio(outputs, "example.mp3")

What it does: Loads the Dia-1.6B-0626 checkpoint, generates a two-speaker script with a (laughs) tag, and saves the result as example.mp3.

Launch the Gradio UI

bash
bash
python app.py

# Or if you have uv installed
uv run app.py

What it does: Starts the repo's Gradio interface for interactive generation.

Explore the CLI

bash
bash
python cli.py --help

# Or if you have uv installed
uv run cli.py --help

What it does: Shows the command-line options available in the repo's CLI.

Script format with speaker and nonverbal tags

Prompt
prompt
[S1] Dia is an open weights text to dialogue model. [S2] You get full control over scripts and voices. [S1] Wow. Amazing. (laughs) [S2] Try it now on Git hub or Hugging Face.

Expected output: Starts with [S1], alternates between [S1] and [S2], and uses a listed nonverbal tag sparingly, as the generation guidelines recommend.

Pros & cons

Pros

  • Pro:Generates multi-speaker dialogue in one pass from a tagged transcript, with [S1]/[S2] speaker tags
  • Pro:Supports nonverbal sounds such as (laughs) and (coughs), plus voice cloning via audio prompts
  • Pro:Open pretrained checkpoints and inference code under Apache 2.0, with Gradio UI, CLI, and Transformers support
  • Pro:Modest VRAM needs (~4.4GB in bfloat16/float16) and faster-than-realtime speed on an RTX 4090

Cons

  • Con:English generation only at the moment
  • Con:Tested only on GPUs; CPU support is not yet available
  • Con:Not fine-tuned on a specific voice, so voices vary between runs unless you use an audio prompt or fix the seed
  • Con:Output is sensitive to input: very short or very long text sounds unnatural, and overusing or using unlisted nonverbal tags may cause artifacts