What it is
MockingBird is a PyTorch-based voice cloning project forked from Real-Time-Voice-Cloning, which only supports English. It adds Mandarin support by reusing a pretrained encoder and vocoder with a newly trained synthesizer. It can be used through a web server, a desktop toolbox or a command-line script. The author states they no longer actively update the repo.
Who it's for
- Developers who want to experiment with Mandarin voice cloning and speech synthesis
- Researchers training their own synthesizer or vocoder on Chinese datasets such as aidatatang_200zh, magicdata, aishell3 and data_aishell
- Users who want to serve voice synthesis results via a web server for remote calling
Requirements
Requirements
- Python 3.7 or higher
- PyTorch (recommended environment: Pytorch 1.9.0 with Torchvision 0.10.0 and cudatoolkit 10.2)
- ffmpeg
- Packages from requirements.txt (webrtcvad-wheels if needed)
- A synthesizer model, either trained yourself or a community-shared pretrained one
- GPU tested: Tesla T4 and GTX 2060
Setup
Install Python dependencies
Install PyTorch and ffmpeg first, then install the remaining packages from requirements.txt. Install webrtcvad-wheels only if needed.
bashpip install -r requirements.txt pip install webrtcvad-wheelsAlternative: conda/mamba environment
Create a virtual environment from env.yml, then activate it. env.yml covers only the necessary dependencies and does not include monotonic-align.
bashconda env create -n env_name -f env.yml conda activate env_namePreprocess synthesizer dataset
Download and unzip a dataset, then preprocess the audio and mel spectrograms. The --dataset parameter selects the dataset; the default is aidatatang_200zh.
bashpython pre.py <datasets_root>Train the synthesizer
Train until the attention line appears and the loss meets your needs. Models are saved in synthesizer/saved_models/.
bashpython train.py --type=synth mandarin <datasets_root>/SV2TTS/synthesizer
Examples
Launch the web server
bashpython web.pyWhat it does: Starts the web server, which defaults to http://localhost:8080 in the browser.
Launch the toolbox
bashpython demo_toolbox.py -d <datasets_root>What it does: Opens the demo toolbox GUI using your datasets root.
Generate voice from the command line
bashpython gen_voice.py <text_file.txt> your_wav_file.wavWhat it does: Generates speech from a text file. The README suggests installing cn2an (pip install cn2an) for better handling of digits.
Train a HiFi-GAN vocoder (optional)
bashpython vocoder_preprocess.py <datasets_root> -m <synthesizer_model_path>
python vocoder_train.py mandarin <datasets_root> hifiganWhat it does: Preprocesses data using a trained synthesizer, then trains a HiFi-GAN vocoder. The README notes vocoder training makes little difference in effect.
Pros & cons
Pros
- Pro:Supports Mandarin and has been tested with multiple Chinese datasets
- Pro:Reuses the pretrained encoder and vocoder, so only a synthesizer needs training
- Pro:Offers three interfaces: web server, toolbox GUI and command line
- Pro:Runs on Windows and Linux, with documented M1 Mac workarounds
Cons
- Con:The author states they no longer actively update the repo
- Con:The original synthesizer is incompatible with Chinese symbols, so demo_cli does not work and extra synthesizer models are required
- Con:The recommended environment is pinned to older versions (Repo Tag 0.0.1, PyTorch 1.9.0), and requirements.txt may not work with newer versions
Images
