wan / 2-5-image-to-video

Wan 2.5 API от Alibaba разработан для создания кинематографичных AI-видео — поддерживает как text-to-video (wan2.5-t2v-preview), так и image-to-video (wan2.5-i2v-preview). Модель нативно синхронизирует визуальный ряд с речью, фоновыми звуками и музыкой. Поддерживаются разные разрешения (720p, 1080p) и соотношения сторон (16:9, 9:16, 1:1), что делает API подходящим для соцсетей, рекламы и творческих проектов.

Model Type:
Pricing: 12 credits per second for 720p (~$0.06) and 20 credits per second for 1080p (~$0.10). High-tier top-ups (+10% bonus) bring effective pricing down to ~$0.054 per second for 720p and ~$0.09 per second for 1080p.
Input

The text prompt describing the desired video motion

0/800

Click to upload or drag and drop

Supported formats: JPEG, PNG, WEBP Maximum file size: 10MB

URL of the image to use as the first frame. Must be publicly accessible

The duration of the generated video in seconds

Video resolution. Valid values: 720p, 1080p

Negative prompt to describe content to avoid

0/500

Whether to enable prompt rewriting using LLM

Random seed for reproducibility. If None, a random seed is chosen

A configurable parameter. Defaults to true in the Playground.

There is no guarantee that everything can be filtered out; if you are not satisfied with the results, you will need to make your own arrangements.
Output
output typevideo
Examples

Explore different use cases and parameter configurations

README

Complete guide to using wan/2-5-image-to-video

Alibaba Wan 2.5 API — генератор AI-видео с синхронизированной озвучкой

Original image

С помощью Wan 2.5 API можно генерировать видео и аудио в одном запросе. Диалоги, фоновые звуки и музыка автоматически синхронизируются с картинкой, создавая эффект полного погружения без дополнительного монтажа.

Wan 2.5 text-to-video API обеспечивает более точное соответствие сложным запросам. Ракурсы камеры, свет и динамика сцены передаются с высокой точностью, что дает разработчикам уверенность: каждый запрос API превращается в предсказуемый и качественный результат.

Wan 2.5 Preview API поддерживает большое разнообразие визуальных стилей — от кинематографического реализма до аниме и иллюстраций. Он сохраняет характер персонажей и целостность сцены, позволяя легко интегрировать разные визуальные форматы в приложения через единый API.

API Wan 2.5 предоставляет два типа API: wan2.5-t2v-preview (текст в видео) и wan2.5-i2v-preview (изображение в видео). Все режимы поддерживают несколько разрешений (720p, 1080p), а для генерации видео по тексту доступны варианты соотношений сторон (16:9, 9:16, 1:1).

ОсобенностьAPI Wan 2.5 (Alibaba)Veo 3 (Google)
Generation ModesText-to-Video (wan2.5-t2v-preview api) & Image-to-Video (wan2.5-i2v-preview api)Text-to-Video & Image-to-Video
Audio & A/V SyncNative audio-video generation with dialogue, ambient sound, and BGMAudio available but less integrated; focus remains on visuals
Prompt AdherenceStrong fidelity to complex instructions, including camera, lighting, and motionExcellent realism, but may struggle with highly detailed or abstract prompts
Style AdaptationCinematic realism, anime, illustration; strong stylization supportFocus on cinematic realism, less flexible for stylized outputs
Multilingual SupportReliable with Chinese & minor languagesLimited; often defaults to “unknown language” in non-English prompts
Video DurationUp to 10 secondsUp to ~8 seconds
Aspect Ratio Options16:9, 9:16, 1:1 (T2V)Primarily cinematic formats; fewer ratio options

When adding speech, don’t just request “dialogue.” Instead, provide the exact words to be spoken and specify who says them. This is especially important in multi-character scenes where order and clarity matter.
For example: Character A: “We have to keep moving.” Character B: “Not until we find shelter.”
By writing dialogue this way, you ensure the API assigns the right lines to the right characters.

In some videos, the atmosphere should be driven by visuals or sound effects alone. If you don’t want dialogue, make that clear in your prompt. Adding phrases such as “no dialogue” or “no actors speaking” prevents unintended voices from appearing. This small detail keeps your output aligned with the creative vision.

Beyond dialogue, ambient sound and music set the emotional tone. Be specific about the kind of environment or soundtrack you want, whether it’s natural or dramatic.
Examples include: “soft rain tapping on windows with distant thunder” or “fast-paced action music with heavy percussion.”
The clearer you are, the better the model can synchronize visuals with sound to create an immersive result.

Wan 2.5 excels when prompts include setting, lighting, camera perspective, and mood. Instead of writing “a person walking on a road,” expand the description to capture cinematic elements.
For example: A wide shot of a mountain road at sunset, golden light flooding the sky, a cyclist racing downhill, with energetic background music in the background.
This depth of description allows the API to produce more natural, dynamic, and visually coherent videos.

4.9/ 5
37,436 ratings
Tap a star to rate