What it is
Dia is a 1.6B parameter text to speech model created by Nari Labs. It directly generates highly realistic dialogue from a transcript using [S1] and [S2] speaker tags, and can be conditioned on audio for emotion and tone control. It can also produce nonverbal sounds such as laughter and coughing. Pretrained checkpoints and inference code are provided, and it currently supports English only.
Who it's for
- Researchers and developers experimenting with open-weights dialogue text-to-speech
- Developers who want multi-speaker conversational audio from a scripted transcript
- Users who want voice cloning via audio prompts
- Hugging Face Transformers users who want to run Dia through that library
Requirements
Requirements
- GPU: tested only on GPUs (pytorch 2.0+, CUDA 12.6); CPU support is listed as to be added soon
- About 4.4GB VRAM for bfloat16/float16 and about 7.9GB for float32 (benchmarked on RTX 4090)
- For the Transformers path, install the main branch of transformers
- For 5000 series GPUs, use torch 2.8 nightly (see issue #26)
- The first run takes longer because the Descript Audio Codec must be downloaded
- uv is needed only if using the uv-based commands
Setup
Install from a clone (pip)
Clone the repo, optionally create a virtual environment, and install in editable mode.
bash# Clone this repository git clone https://github.com/nari-labs/dia.git cd dia # Optionally python -m venv .venv && source .venv/bin/activate # Install dia pip install -e .Install directly from GitHub
Install without cloning the repository.
bash# Install directly from GitHub pip install git+https://github.com/nari-labs/dia.gitRun the simple example
Run the bundled example after installing with pip.
bashpython example/simple.pyRun with uv
With uv installed, clone the repo and run the example directly.
bash# Clone this repository git clone https://github.com/nari-labs/dia.git cd dia uv run example/simple.pyInstall Transformers main branch
Needed to use the Hugging Face Transformers implementation of Dia.
bashpip install git+https://github.com/huggingface/transformers.git # or install with uv uv pip install git+https://github.com/huggingface/transformers.git
Examples
Generate dialogue with Transformers
pythonfrom transformers import AutoProcessor, DiaForConditionalGeneration
torch_device = "cuda"
model_checkpoint = "nari-labs/Dia-1.6B-0626"
text = [
"[S1] Dia is an open weights text to dialogue model. [S2] You get full control over scripts and voices. [S1] Wow. Amazing. (laughs) [S2] Try it now on Git hub or Hugging Face."
]
processor = AutoProcessor.from_pretrained(model_checkpoint)
inputs = processor(text=text, padding=True, return_tensors="pt").to(torch_device)
model = DiaForConditionalGeneration.from_pretrained(model_checkpoint).to(torch_device)
outputs = model.generate(
**inputs, max_new_tokens=3072, guidance_scale=3.0, temperature=1.8, top_p=0.90, top_k=45
)
outputs = processor.batch_decode(outputs)
processor.save_audio(outputs, "example.mp3")What it does: Loads the Dia-1.6B-0626 checkpoint, generates a two-speaker script with a (laughs) tag, and saves the result as example.mp3.
Launch the Gradio UI
bashpython app.py
# Or if you have uv installed
uv run app.pyWhat it does: Starts the repo's Gradio interface for interactive generation.
Explore the CLI
bashpython cli.py --help
# Or if you have uv installed
uv run cli.py --helpWhat it does: Shows the command-line options available in the repo's CLI.
Script format with speaker and nonverbal tags
Prompt[S1] Dia is an open weights text to dialogue model. [S2] You get full control over scripts and voices. [S1] Wow. Amazing. (laughs) [S2] Try it now on Git hub or Hugging Face.Expected output: Starts with [S1], alternates between [S1] and [S2], and uses a listed nonverbal tag sparingly, as the generation guidelines recommend.
Pros & cons
Pros
- Pro:Generates multi-speaker dialogue in one pass from a tagged transcript, with [S1]/[S2] speaker tags
- Pro:Supports nonverbal sounds such as (laughs) and (coughs), plus voice cloning via audio prompts
- Pro:Open pretrained checkpoints and inference code under Apache 2.0, with Gradio UI, CLI, and Transformers support
- Pro:Modest VRAM needs (~4.4GB in bfloat16/float16) and faster-than-realtime speed on an RTX 4090
Cons
- Con:English generation only at the moment
- Con:Tested only on GPUs; CPU support is not yet available
- Con:Not fine-tuned on a specific voice, so voices vary between runs unless you use an audio prompt or fix the seed
- Con:Output is sensitive to input: very short or very long text sounds unnatural, and overusing or using unlisted nonverbal tags may cause artifacts