OmniHuman 1.5 ist das KI-Avatar- und Digital-Human-Modell von ByteDance, das ein einzelnes Bild und eine Audioeingabe in realistische Sprechvideos mit natürlicher Lippensynchronität, ausdrucksstarker Mimik und lebensechten Bewegungen verwandelt.

Model Type:
Pricing: 27 credits per second (~$0.135). High-tier top-ups (+10% bonus) bring effective pricing down to ~$0.122 per second.
Input

Click to upload or drag and drop

Supported formats: JPEG, PNG, WEBP Maximum file size: 10MB

Portrait image URL, supports any aspect ratio with subjects including people/pets/anime, etc.

Text
0 / ∞
No items yet. Click Add to start.

To have a specific subject in the image speak, use 'Subject Detection' to get the corresponding mask image and pass it as input.

Click to upload or drag and drop

Supported formats: MPEG, WAV, X-WAV, AAC, OGG, MP4 Maximum file size: 10MB

Audio URL. Duration must be < 60 seconds (recommended ≤15 seconds; exceeding this will cause degradation).

Prompt text, limited to Chinese/English/Japanese/Korean/Spanish/Indonesian, recommended ≤300 characters.

0/300

Output video resolution, default 1080.

Fast mode, sacrifices some quality to speed up generation.

Random seed. Default is -1 (random). When using the same positive integer and keeping all other parameters identical, the result will be highly consistent (with very high probability).

Output
output typevideo

README

Complete guide to using omnihuman-1-5

ByteDance OmniHuman 1.5 API zur Videogenerierung für KI-Avatare und digitale Menschen

Original image