wan / 2-5-image-to-video

Die Wan 2.5 API von Alibaba wurde für die filmische KI-Videoerstellung entwickelt und unterstützt sowohl Text-zu-Video (wan2.5-t2v-preview) als auch Bild-zu-Video (wan2.5-i2v-preview). Sie synchronisiert visuelle Inhalte nativ mit Dialogen, Umgebungsgeräuschen und Hintergrundmusik. Mit Unterstützung für mehrere Auflösungen (720p, 1080p) und Seitenverhältnisse (16:9, 9:16, 1:1) liefert die API vielseitige Ausgabeformate – ideal für Social Media, Werbung und kreatives Storytelling.

Model Type:
Pricing: 12 credits per second for 720p (~$0.06) and 20 credits per second for 1080p (~$0.10). High-tier top-ups (+10% bonus) bring effective pricing down to ~$0.054 per second for 720p and ~$0.09 per second for 1080p.
Input

The text prompt describing the desired video motion

0/800

Click to upload or drag and drop

Supported formats: JPEG, PNG, WEBP Maximum file size: 10MB

URL of the image to use as the first frame. Must be publicly accessible

The duration of the generated video in seconds

Video resolution. Valid values: 720p, 1080p

Negative prompt to describe content to avoid

0/500

Whether to enable prompt rewriting using LLM

Random seed for reproducibility. If None, a random seed is chosen

A configurable parameter. Defaults to true in the Playground.

There is no guarantee that everything can be filtered out; if you are not satisfied with the results, you will need to make your own arrangements.
Output
output typevideo
Examples

Explore different use cases and parameter configurations

README

Complete guide to using wan/2-5-image-to-video

Alibaba Wan 2.5 API – KI-Videoerstellung mit Audio-Synchronisation

Original image

Mit der Wan 2.5 API lassen sich Video und Audio in nur einem Schritt erzeugen. Dialoge, Umgebungsgeräusche und Hintergrundmusik werden automatisch mit den visuellen Inhalten synchronisiert – für immersive Ergebnisse ohne zusätzlichen Bearbeitungsaufwand.

Die Wan 2.5 Text-to-Video API setzt komplexe Texteingaben besonders genau um. Kameraeinstellungen, Lichtführung und Szenendynamik werden detailgetreu erfasst, sodass Entwickler darauf vertrauen können, dass jeder API-Aufruf kreative Vorgaben zuverlässig in konsistente Videoergebnisse übersetzt.

Die Wan 2.5 Preview API unterstützt eine vielfältige Auswahl visueller Stile – von filmischem Realismus bis zu Anime oder Illustration. Charakteridentität und Szenenkohärenz bleiben erhalten, sodass Entwickler mit nur einer API vielseitige Ästhetiken in ihre Anwendungen integrieren können.

Die Wan 2.5 API bietet zwei Schnittstellen: wan2.5-t2v-preview (Text-to-Video) und wan2.5-i2v-preview (Image-to-Video). Alle Modi unterstützen mehrere Auflösungen (z. B. 720p, 1080p). Für die Text-to-Video-Generierung können Sie außerdem das Seitenverhältnis wählen (16:9, 9:16, 1:1).

FunktionWan 2.5 API (Alibaba)Veo 3 (Google)
Generation ModesText-to-Video (wan2.5-t2v-preview api) & Image-to-Video (wan2.5-i2v-preview api)Text-to-Video & Image-to-Video
Audio & A/V SyncNative audio-video generation with dialogue, ambient sound, and BGMAudio available but less integrated; focus remains on visuals
Prompt AdherenceStrong fidelity to complex instructions, including camera, lighting, and motionExcellent realism, but may struggle with highly detailed or abstract prompts
Style AdaptationCinematic realism, anime, illustration; strong stylization supportFocus on cinematic realism, less flexible for stylized outputs
Multilingual SupportReliable with Chinese & minor languagesLimited; often defaults to “unknown language” in non-English prompts
Video DurationUp to 10 secondsUp to ~8 seconds
Aspect Ratio Options16:9, 9:16, 1:1 (T2V)Primarily cinematic formats; fewer ratio options

When adding speech, don’t just request “dialogue.” Instead, provide the exact words to be spoken and specify who says them. This is especially important in multi-character scenes where order and clarity matter.
For example: Character A: “We have to keep moving.” Character B: “Not until we find shelter.”
By writing dialogue this way, you ensure the API assigns the right lines to the right characters.

In some videos, the atmosphere should be driven by visuals or sound effects alone. If you don’t want dialogue, make that clear in your prompt. Adding phrases such as “no dialogue” or “no actors speaking” prevents unintended voices from appearing. This small detail keeps your output aligned with the creative vision.

Beyond dialogue, ambient sound and music set the emotional tone. Be specific about the kind of environment or soundtrack you want, whether it’s natural or dramatic.
Examples include: “soft rain tapping on windows with distant thunder” or “fast-paced action music with heavy percussion.”
The clearer you are, the better the model can synchronize visuals with sound to create an immersive result.

Wan 2.5 excels when prompts include setting, lighting, camera perspective, and mood. Instead of writing “a person walking on a road,” expand the description to capture cinematic elements.
For example: A wide shot of a mountain road at sunset, golden light flooding the sky, a cyclist racing downhill, with energetic background music in the background.
This depth of description allows the API to produce more natural, dynamic, and visually coherent videos.

4.9/ 5
37,436 ratings
Tap a star to rate