volcengine / video-to-video-lip-sync

Make every frame speak your language — Volcengine Lip Sync API delivers pixel-perfect video-to-video lip sync and AI dubbing in minutes.

Pricing: 8 credits per second (~$0.04). High-tier top-ups (+10% bonus) bring effective pricing down to ~$0.036 per second.
Input

Click to upload or drag and drop

Supported formats: MP4, QUICKTIME, X-MATROSKA Maximum file size: 500MB

Video asset URL

Click to upload or drag and drop

Supported formats: MPEG, WAV, X-WAV, AAC, MP4, OGG Maximum file size: 10MB

Target pure vocal audio URL; used to drive video lip movements.

Service identifier

Enable vocal separation to suppress background noise.

Whether to enable scene segmentation and speaker identification. Supported only in Basic mode.

Supported in lite mode. Whether to loop the video when the audio is longer than the video.

Supported in lite mode. Whether to loop the video in reverse (backward). Requires align_audio to be set to true.

Supported in lite mode. Start time of the template video, in seconds.

Output
output typevideo

no output

README

Complete guide to using volcengine/video-to-video-lip-sync

Volcengine Video-to-Video Lip Sync API: AI Lip Sync & Video Dubbing API

Original image
Frame-Accurate Lip Synchronization
Multi-Language Video Dubbing Pipeline
Native Ecosystem Integration
High-Throughput Async Task Processing

Frequently Asked Questions About Volcengine Video-to-Video Lip Sync API