volcengine / video-to-video-lip-sync

让每一帧都用您的语言发声——火山引擎口型同步 API 仅需数分钟即可实现像素级精准的视频转视频口型同步与 AI 配音。

Pricing: 8 credits per second (~$0.04). High-tier top-ups (+10% bonus) bring effective pricing down to ~$0.036 per second.
Input

Click to upload or drag and drop

Supported formats: MP4, QUICKTIME, X-MATROSKA Maximum file size: 500MB

Video asset URL

Click to upload or drag and drop

Supported formats: MPEG, WAV, X-WAV, AAC, MP4, OGG Maximum file size: 10MB

Target pure vocal audio URL; used to drive video lip movements.

Service identifier

Enable vocal separation to suppress background noise.

Whether to enable scene segmentation and speaker identification. Supported only in Basic mode.

Supported in lite mode. Whether to loop the video when the audio is longer than the video.

Supported in lite mode. Whether to loop the video in reverse (backward). Requires align_audio to be set to true.

Supported in lite mode. Start time of the template video, in seconds.

Output
output typevideo

no output

README

Complete guide to using volcengine/video-to-video-lip-sync

火山引擎视频转视频口型同步 API:AI 口型同步与视频配音 API

Original image
帧级精准口型同步
多语种视频配音工作流
原生生态无缝集成
高吞吐量异步任务处理

火山引擎视频口型同步(Video-to-Video)API 常见问题解答