Skip to content

Repository files navigation

VoiceToLipSync: High-Quality Lip Sync Video Generation

PyTorch implementation of LipSync-Voice: High-Quality Lip Sync Video Generation using Wav2Lip & Openvoice models.

Tech Stack

PyTorch OpenCV Librosa GAN Wav2Lip OpenVoice


LipSync-Voice is a high-quality lip-sync video generation system that leverages deep learning to synchronize facial movements with speech. This framework processes input a sample video through LipSync.ipynb, extracting speech features and synchronizing them with facial motions to create realistic lip-sync videos—all without requiring manual installation of dependencies.

Project Duration

2025.01.01 - 2025.01.20


Yunsu Park


Myoungjin Son

Presentation

The presentation deck is available in the deck folder: LipSync_Voice_Presentation.pdf.

How to Use

  1. Click the "Open in Colab" button below.
  2. Follow the instructions in the notebook to upload your video files.
  3. Generate the lip-sync video and download the result.

VoiceToLipSync Demo

Click the badge below to run the demo

Open In Colab

Results

Reference Video

sample.mp4

Lip Synced Video

output_lip_synced.mp4

References

Lip Syncing

Speech Synthesis and Voice Conversion

  • OpenVoice: A deep learning framework for text-to-speech conversion with high-quality voice synthesis.
    Repository: https://github.com/OpenVoice

About

Voice to LipSync : High-Quality Lip Sync Video Genreation OpenVoice to Generative a Zero-Shot TTS and Wav2Lip to Generative Lip Sync Video.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages