What it is
pyVideoTrans converts videos from one language to another through a workflow of speech recognition, subtitle translation, multi-role dubbing and audio-video synchronization. It supports local offline deployment as well as many mainstream online APIs. It ships as a Windows .exe, a source install via uv, a CLI, a WebUI and a Docker image.
Who it's for
- Users who want automated video translation with dubbing and subtitles
- Developers who need batch or headless processing via CLI or WebUI
- Windows users who want a prepackaged build without configuring Python
- Users who want to transcribe audio/video into SRT subtitles
Requirements
Requirements
- Python 3.10 recommended for source deployment
- FFmpeg installed and configured in environment variables (or ffmpeg.exe and ffprobe.exe placed in the project directory on Windows)
- Windows 10/11 for the prepackaged .exe
- For NVIDIA GPU acceleration with the packaged version: CUDA 12.8 and cuDNN 9.11
- uv recommended for package management
Setup
Windows prepackaged version
Download the latest release from the GitHub releases page, extract it to a path without Chinese characters or spaces (e.g. D:\pyVideoTrans), and double-click sp.exe. Do not run directly from within the archive.
Install uv (macOS/Linux)
Install uv if not already installed.
bashcurl -LsSf https://astral.sh/uv/install.sh | shClone and install from source
Clone the repo and sync dependencies. Whisper.net and WebUI are not installed by default.
bashgit clone https://github.com/jianchang512/pyvideotrans.git cd pyvideotrans uv syncLaunch the GUI
Run the desktop app from source.
bashuv run sp.pyInstall optional extras
Install the WebUI extra (or all optional channels with --all-extras).
bashuv sync --extra webui
Examples
Video translation via CLI
bashuv run cli.py --task vtv --name "./video.mp4" --source_language_code zh-cn --target_language_code en --voice_role "en-US-GuyNeural"What it does: Translates a Chinese video into English and dubs it with the specified voice.
Audio to subtitle
bashuv run cli.py --task stt --name "./audio.wav" --model_name large-v3What it does: Transcribes an audio file to subtitles using the large-v3 model.
Subtitle translation
bashuv run cli.py --task sts --name "./subs.srt" --target_language_code enWhat it does: Translates an existing SRT subtitle file into English.
Text to speech
bashuv run cli.py --task tts --name "./subs.srt" --voice_role "zh-CN-YunyangNeural"What it does: Generates speech from a subtitle file using the chosen voice role.
Run WebUI in Docker
bashdocker build -t pyvideotrans-webui .
docker run -d -p 7860:7860 --name pyvideotrans pyvideotrans-webuiWhat it does: Builds the image and runs the WebUI container, exposing port 7860.
Pros & cons
Pros
- Pro:Full pipeline in one tool: ASR, subtitle translation, TTS dubbing and video synthesis
- Pro:Broad model support, including local options (Faster-Whisper, Ollama, M2M100) and online APIs
- Pro:Multiple interfaces: GUI, CLI, WebUI and Docker
- Pro:Supports speaker diarization, multi-role dubbing, voice cloning and manual proofreading at each stage
Cons
- Con:Source install requires FFmpeg and manual environment setup; GPU use needs specific CUDA/cuDNN versions
- Con:Some TTS options such as GPT-SoVITS, Index-TTS and ChatTTS require local deployment
- Con:Users are responsible for legal consequences of third-party API use and copyrighted content