Repo Voice & Audio

babysor/MockingBird

Mandarin-focused voice cloning toolbox forked from Real-Time-Voice-Cloning, with a web server, GUI toolbox and CLI. No longer actively updated.

  • 36.9k GitHub stars
  • Python
  • ⚖️ NOASSERTION
  • 🎯 Intermediate
babysor/MockingBird preview image

What it is

MockingBird is a PyTorch-based voice cloning project forked from Real-Time-Voice-Cloning, which only supports English. It adds Mandarin support by reusing a pretrained encoder and vocoder with a newly trained synthesizer. It can be used through a web server, a desktop toolbox or a command-line script. The author states they no longer actively update the repo.

Who it's for

  • Developers who want to experiment with Mandarin voice cloning and speech synthesis
  • Researchers training their own synthesizer or vocoder on Chinese datasets such as aidatatang_200zh, magicdata, aishell3 and data_aishell
  • Users who want to serve voice synthesis results via a web server for remote calling

Requirements

Requirements

  • Python 3.7 or higher
  • PyTorch (recommended environment: Pytorch 1.9.0 with Torchvision 0.10.0 and cudatoolkit 10.2)
  • ffmpeg
  • Packages from requirements.txt (webrtcvad-wheels if needed)
  • A synthesizer model, either trained yourself or a community-shared pretrained one
  • GPU tested: Tesla T4 and GTX 2060

Setup

  1. Install Python dependencies

    Install PyTorch and ffmpeg first, then install the remaining packages from requirements.txt. Install webrtcvad-wheels only if needed.

    bash
    pip install -r requirements.txt
    pip install webrtcvad-wheels
  2. Alternative: conda/mamba environment

    Create a virtual environment from env.yml, then activate it. env.yml covers only the necessary dependencies and does not include monotonic-align.

    bash
    conda env create -n env_name -f env.yml
    conda activate env_name
  3. Preprocess synthesizer dataset

    Download and unzip a dataset, then preprocess the audio and mel spectrograms. The --dataset parameter selects the dataset; the default is aidatatang_200zh.

    bash
    python pre.py <datasets_root>
  4. Train the synthesizer

    Train until the attention line appears and the loss meets your needs. Models are saved in synthesizer/saved_models/.

    bash
    python train.py --type=synth mandarin <datasets_root>/SV2TTS/synthesizer

Examples

Launch the web server

bash
bash
python web.py

What it does: Starts the web server, which defaults to http://localhost:8080 in the browser.

Launch the toolbox

bash
bash
python demo_toolbox.py -d <datasets_root>

What it does: Opens the demo toolbox GUI using your datasets root.

Generate voice from the command line

bash
bash
python gen_voice.py <text_file.txt> your_wav_file.wav

What it does: Generates speech from a text file. The README suggests installing cn2an (pip install cn2an) for better handling of digits.

Train a HiFi-GAN vocoder (optional)

bash
bash
python vocoder_preprocess.py <datasets_root> -m <synthesizer_model_path>
python vocoder_train.py mandarin <datasets_root> hifigan

What it does: Preprocesses data using a trained synthesizer, then trains a HiFi-GAN vocoder. The README notes vocoder training makes little difference in effect.

Pros & cons

Pros

  • Pro:Supports Mandarin and has been tested with multiple Chinese datasets
  • Pro:Reuses the pretrained encoder and vocoder, so only a synthesizer needs training
  • Pro:Offers three interfaces: web server, toolbox GUI and command line
  • Pro:Runs on Windows and Linux, with documented M1 Mac workarounds

Cons

  • Con:The author states they no longer actively update the repo
  • Con:The original synthesizer is incompatible with Chinese symbols, so demo_cli does not work and extra synthesizer models are required
  • Con:The recommended environment is pinned to older versions (Repo Tag 0.0.1, PyTorch 1.9.0), and requirements.txt may not work with newer versions

Images