wan / 2-5-image-to-video

L'API Wan 2.5 d'Alibaba est idéale pour la création de vidéos IA cinématographiques, prenant en charge à la fois la conversion de texte en vidéo (wan2.5-t2v-preview) et d'image en vidéo (wan2.5-i2v-preview). Elle synchronise nativement les visuels avec les dialogues, les sons ambiants et la musique de fond. Avec des résolutions multiples (720p, 1080p) et des rapports d'aspect flexibles (16:9, 9:16, 1:1), l'API offre des sorties flexibles adaptées aux réseaux sociaux, à la publicité et à la narration créative.

Model Type:
Pricing: 12 credits per second for 720p (~$0.06) and 20 credits per second for 1080p (~$0.10). High-tier top-ups (+10% bonus) bring effective pricing down to ~$0.054 per second for 720p and ~$0.09 per second for 1080p.
Input

The text prompt describing the desired video motion

0/800

Click to upload or drag and drop

Supported formats: JPEG, PNG, WEBP Maximum file size: 10MB

URL of the image to use as the first frame. Must be publicly accessible

The duration of the generated video in seconds

Video resolution. Valid values: 720p, 1080p

Negative prompt to describe content to avoid

0/500

Whether to enable prompt rewriting using LLM

Random seed for reproducibility. If None, a random seed is chosen

A configurable parameter. Defaults to true in the Playground.

There is no guarantee that everything can be filtered out; if you are not satisfied with the results, you will need to make your own arrangements.
Output
output typevideo
Examples

Explore different use cases and parameter configurations

README

Complete guide to using wan/2-5-image-to-video

API Alibaba Wan 2.5 – Génération innovante de vidéos IA avec synchronisation audio

Original image

L'API Wan 2.5 permet de générer la vidéo et l'audio simultanément en une seule requête. Les dialogues, les sons ambiants et la musique de fond sont automatiquement synchronisés avec les visuels, fournissant des résultats immersifs sans nécessiter de retouches.

Avec l'API texte-en-vidéo Wan 2.5, les requêtes complexes sont mieux respectées. Les angles de caméra, les réglages d'éclairage et la dynamique des scènes sont rendus avec une plus grande précision, garantissant aux développeurs que chaque appel d'API traduira les instructions créatives en résultats vidéo cohérents.

L'API de prévisualisation Wan 2.5 supporte une large gamme de styles visuels — du réalisme cinémRe-edit traduction APIatographique à l'animation en passant par l'illustration. Elle préserve l'identité des personnages et la cohérence des scènes, permettant aux développeurs d'intégrer des styles visuels divers dans leurs applications grâce à une seule API.

L'API Wan 2.5 propose deux points de terminaison : **wan2.5-t2v-preview** (texte en vidéo) et **wan2.5-i2v-preview** (image en vidéo). Tous les modes prennent en charge plusieurs résolutions (720p, 1080p). Pour la génération texte en vidéo, vous pouvez aussi définir la proportion d'image (16:9, 9:16, 1:1).

FonctionnalitéAPI Wan 2.5 (Alibaba)Veo 3 (Google)
Generation ModesText-to-Video (wan2.5-t2v-preview api) & Image-to-Video (wan2.5-i2v-preview api)Text-to-Video & Image-to-Video
Audio & A/V SyncNative audio-video generation with dialogue, ambient sound, and BGMAudio available but less integrated; focus remains on visuals
Prompt AdherenceStrong fidelity to complex instructions, including camera, lighting, and motionExcellent realism, but may struggle with highly detailed or abstract prompts
Style AdaptationCinematic realism, anime, illustration; strong stylization supportFocus on cinematic realism, less flexible for stylized outputs
Multilingual SupportReliable with Chinese & minor languagesLimited; often defaults to “unknown language” in non-English prompts
Video DurationUp to 10 secondsUp to ~8 seconds
Aspect Ratio Options16:9, 9:16, 1:1 (T2V)Primarily cinematic formats; fewer ratio options

When adding speech, don’t just request “dialogue.” Instead, provide the exact words to be spoken and specify who says them. This is especially important in multi-character scenes where order and clarity matter.
For example: Character A: “We have to keep moving.” Character B: “Not until we find shelter.”
By writing dialogue this way, you ensure the API assigns the right lines to the right characters.

In some videos, the atmosphere should be driven by visuals or sound effects alone. If you don’t want dialogue, make that clear in your prompt. Adding phrases such as “no dialogue” or “no actors speaking” prevents unintended voices from appearing. This small detail keeps your output aligned with the creative vision.

Beyond dialogue, ambient sound and music set the emotional tone. Be specific about the kind of environment or soundtrack you want, whether it’s natural or dramatic.
Examples include: “soft rain tapping on windows with distant thunder” or “fast-paced action music with heavy percussion.”
The clearer you are, the better the model can synchronize visuals with sound to create an immersive result.

Wan 2.5 excels when prompts include setting, lighting, camera perspective, and mood. Instead of writing “a person walking on a road,” expand the description to capture cinematic elements.
For example: A wide shot of a mountain road at sunset, golden light flooding the sky, a cyclist racing downhill, with energetic background music in the background.
This depth of description allows the API to produce more natural, dynamic, and visually coherent videos.

4.9/ 5
37,436 ratings
Tap a star to rate