万相3.0是阿里新一代 All-in-One 视频生成模型,专为多模态创作打造。它支持文本、图像、视频、音频及结构化信息参考,可生成一致性更强、更逼真且更可控的视频内容,全面提升叙事与编辑能力。

Pricing: 8 credits/s for 480P ($0.04/s), 16 credits/s for 720P ($0.08/s), and 32 credits/s for 1080P ($0.16/s), all 20% below official pricing. Note: The billing rule is (input video duration + output video duration) × unit price. High-tier top-ups (+10% bonus) provide an additional 10% reduction in effective pricing.
Input
Start Frame
End Frame

Click to upload or drag and drop images. Single side: [240, 8000] px; aspect ratio ≤ 8:1

Images: JPEG/JPG/PNG/WebP/BMP, Max 20MB

Text prompt used to describe the desired video content. Supports both Chinese and English, with a maximum length of 20,000 characters.

0/20000

Click to upload or drag and drop

Supported formats: JPEG, PNG, WEBP, BMP Maximum file size: 20MB; Maximum files: 10

Maximum 10 images. Mapped in array order to Image1, Image2, ... in the prompt. Specifications are the same as first_frame_url. Cannot be used together with first frame / last frame.

Click to upload or drag and drop

Supported formats: MP4, QUICKTIME Maximum file size: 100MB; Maximum files: 5

Maximum 5 clips. Each clip: 1–15s; total combined duration ≤ 15s. Mapped in array order to Video1, Video2, ... in the prompt. Formats: mp4, mov. Single side: [240, 4096] px; aspect ratio ≤ 8:1; file size per clip ≤ 100MB. Additionally, on the output side: total input video duration + output duration (duration parameter) must not exceed 30 seconds.

Click to upload or drag and drop

Supported formats: MPEG, WAV, X-WAV Maximum file size: 15MB; Maximum files: 5

Maximum 5 clips. Each clip: 1–15s; total combined duration ≤ 15s. Mapped in array order to Audio1, Audio2, ... in the prompt. Formats: wav, mp3; file size ≤ 15MB per clip. When not used as the sole media input, it is still recommended to pair with images or videos.

Text
0 / 1
No items yet. Click Add to start.

Maximum 1 publicly accessible webpage (no login required). Cannot be used together with reference_file_urls, nor with first frame / last frame.

Text
0 / 1
No items yet. Click Add to start.

Maximum 1 file. Cannot be used together with reference_link_urls, nor with first frame / last frame. Supported formats: docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md. File size ≤ 100MB. For file types such as pdf, docx, ppt, key, pages, etc., page count ≤ 50 pages.

Preset resolution levels for video generation.

Aspect ratio for generated video. "Adaptive" auto-recommends a ratio based on input media proportions and intent.

Duration of the generated video, in seconds. When no video input is provided, the value must be an integer in the range of [2, 30]. When a video input is provided, the total duration of the input video plus the output video must not exceed 30 seconds. Passing -1 enables smart duration mode, where the model automatically recommends an appropriate duration based on the input.

Whether the output video includes audio. true (default): Includes audio; false: No audio track.

Random seed for result reproducibility. 0 <= x <= 2147483647

A configurable parameter. Defaults to true in the Playground.

There is no guarantee that everything can be filtered out; if you are not satisfied with the results, you will need to make your own arrangements.
Output
output typevideo
Examples

Explore different use cases and parameter configurations

README

Complete guide to using wan/3-0-video

万相3.0 API 以更强控制力重塑多模态视频创作

Original image

A short film can develop across multiple actions, locations, camera changes, and story beats without being reduced to a single isolated moment. Wan3.0 API is well suited to sequences such as an urban chase, miniature adventure, surreal comedy, or cinematic action scene where characters, props, environments, and narrative logic need to remain connected as the story progresses.

When several characters, costumes, environments, or visual references need to remain recognizable throughout a scene, Wan 3.0 API can use those materials as persistent creative guidance. This makes it useful for fashion films, choreography sequences, character showcases, historical scenes, and other productions where appearance, styling, spatial relationships, and movement continuity all matter.

Existing video can become the starting point for controlled creative changes rather than being regenerated from scratch. Wan3.0-Video API can support tasks such as replacing a subject, changing an environment, removing an accessory, shifting the visual style, adjusting selected scene elements, or extending an existing sequence while preserving the motion and structure that should remain unchanged.

Videos built around dialogue, music, environmental sound, rhythmic motion, or expressive atmosphere can benefit from treating audio as part of the creative direction from the beginning. Wan 3.0 Video API can support content such as performance videos, cinematic commercials, immersive story scenes, and rhythm-driven motion pieces where sound, visual pacing, character performance, and scene dynamics need to feel closely synchronized.

Kie.ai 提供了一种极简的方式对接万相 3.0 API,无需繁琐的集成配置。从身份验证、请求提交、任务管理到结果获取,均采用清晰的 API 工作流,让接入与维护更加便捷。

不同的项目对生成量和复杂度的需求各不相同。Kie.ai 为万相3.0 API 提供灵活透明的计费方案,让您轻松控制使用成本,并根据实际生产需求弹性扩展生成能力。

视频生成通常涉及异步任务,需要高效地进行提交、监控和结果获取。Kie.ai 为万相3.0视频 API 提供了结构化的任务管理方案,让生成状态、处理流程和成品输出的管理变得更加轻松高效。

结构清晰的文档能显著降低 API 接入难度。Kie.ai 为万相3.0视频 API 提供了实用的开发指南,涵盖身份验证、请求结构、任务提交和结果获取等核心步骤,让集成流程一目了然、轻松落地。

Kie.ai 提供丰富的图像、视频、音频及多模态模型选择。除了万相3.0 API,您还可以根据不同的创作需求探索更多选项。一站式汇聚不断升级的生成能力,让您能够轻松为各项任务匹配最合适的模型。

在遇到 API 请求、任务状态、计费或集成等问题时,如果难以获得支持,往往会导致开发受阻。Kie.ai 为万相3.0 API 用户提供快速响应的协助,帮助其在问题发生时更高效地解决平台与工作流相关的问题。