Wan3.0 is Alibaba’s next-generation all-in-one video generation model built for multimodal creation. It supports text, image, video, audio, and structured references to produce more consistent, realistic, and controllable video content with enhanced storytelling and editing capabilities.

Pricing: 8 credits/s for 480P ($0.04/s), 16 credits/s for 720P ($0.08/s), and 32 credits/s for 1080P ($0.16/s), all 20% below official pricing. Note: The billing rule is (input video duration + output video duration) × unit price. High-tier top-ups (+10% bonus) provide an additional 10% reduction in effective pricing.
Input
Start Frame
End Frame

Click to upload or drag and drop images. Single side: [240, 8000] px; aspect ratio ≤ 8:1

Images: JPEG/JPG/PNG/WebP/BMP, Max 20MB

Text prompt used to describe the desired video content. Supports both Chinese and English, with a maximum length of 20,000 characters.

0/20000

Click to upload or drag and drop

Supported formats: JPEG, PNG, WEBP, BMP Maximum file size: 20MB; Maximum files: 10

Maximum 10 images. Mapped in array order to Image1, Image2, ... in the prompt. Specifications are the same as first_frame_url. Cannot be used together with first frame / last frame.

Click to upload or drag and drop

Supported formats: MP4, QUICKTIME Maximum file size: 100MB; Maximum files: 5

Maximum 5 clips. Each clip: 1–15s; total combined duration ≤ 15s. Mapped in array order to Video1, Video2, ... in the prompt. Formats: mp4, mov. Single side: [240, 4096] px; aspect ratio ≤ 8:1; file size per clip ≤ 100MB. Additionally, on the output side: total input video duration + output duration (duration parameter) must not exceed 30 seconds.

Click to upload or drag and drop

Supported formats: MPEG, WAV, X-WAV Maximum file size: 15MB; Maximum files: 5

Maximum 5 clips. Each clip: 1–15s; total combined duration ≤ 15s. Mapped in array order to Audio1, Audio2, ... in the prompt. Formats: wav, mp3; file size ≤ 15MB per clip. When not used as the sole media input, it is still recommended to pair with images or videos.

Text
0 / 1
No items yet. Click Add to start.

Maximum 1 publicly accessible webpage (no login required). Cannot be used together with reference_file_urls, nor with first frame / last frame.

Text
0 / 1
No items yet. Click Add to start.

Maximum 1 file. Cannot be used together with reference_link_urls, nor with first frame / last frame. Supported formats: docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md. File size ≤ 100MB. For file types such as pdf, docx, ppt, key, pages, etc., page count ≤ 50 pages.

Preset resolution levels for video generation.

Aspect ratio for generated video. "Adaptive" auto-recommends a ratio based on input media proportions and intent.

Duration of the generated video, in seconds. When no video input is provided, the value must be an integer in the range of [2, 30]. When a video input is provided, the total duration of the input video plus the output video must not exceed 30 seconds. Passing -1 enables smart duration mode, where the model automatically recommends an appropriate duration based on the input.

Whether the output video includes audio. true (default): Includes audio; false: No audio track.

Random seed for result reproducibility. 0 <= x <= 2147483647

A configurable parameter. Defaults to true in the Playground.

There is no guarantee that everything can be filtered out; if you are not satisfied with the results, you will need to make your own arrangements.
Output
output typevideo
Examples

Explore different use cases and parameter configurations

README

Complete guide to using wan/3-0-video

Wan3.0 API Redefines Multimodal Video Creation with Greater Control

Original image

A short film can develop across multiple actions, locations, camera changes, and story beats without being reduced to a single isolated moment. Wan3.0 API is well suited to sequences such as an urban chase, miniature adventure, surreal comedy, or cinematic action scene where characters, props, environments, and narrative logic need to remain connected as the story progresses.

When several characters, costumes, environments, or visual references need to remain recognizable throughout a scene, Wan 3.0 API can use those materials as persistent creative guidance. This makes it useful for fashion films, choreography sequences, character showcases, historical scenes, and other productions where appearance, styling, spatial relationships, and movement continuity all matter.

Existing video can become the starting point for controlled creative changes rather than being regenerated from scratch. Wan3.0-Video API can support tasks such as replacing a subject, changing an environment, removing an accessory, shifting the visual style, adjusting selected scene elements, or extending an existing sequence while preserving the motion and structure that should remain unchanged.

Videos built around dialogue, music, environmental sound, rhythmic motion, or expressive atmosphere can benefit from treating audio as part of the creative direction from the beginning. Wan 3.0 Video API can support content such as performance videos, cinematic commercials, immersive story scenes, and rhythm-driven motion pieces where sound, visual pacing, character performance, and scene dynamics need to feel closely synchronized.

Kie.ai provides a straightforward way to connect with Wan3.0 API without adding unnecessary complexity to the integration process. Authentication, request submission, task management, and result retrieval follow a clear API workflow that is easier to adopt and maintain.

Different projects can require very different levels of generation volume and complexity. Kie.ai offers flexible and transparent API pricing for Wan 3.0 API, making it easier to control usage costs and scale generation according to actual production needs.

Video generation often involves asynchronous tasks that need to be submitted, monitored, and retrieved efficiently. Kie.ai provides structured task handling for Wan3.0-Video API, helping keep generation status, processing flow, and completed outputs easier to manage.

Well-structured documentation makes API adoption significantly easier. Kie.ai provides practical guidance for Wan 3.0 Video API, covering authentication, request structure, task submission, result retrieval, and other essential integration steps so the workflow is easier to understand and implement.

Kie.ai provides access to a wide selection of image, video, audio, and multimodal models, giving you more options beyond Wan3.0 API for different creative requirements. This makes it easier to choose the most suitable model for each task while keeping access to a growing range of generation capabilities in one place.

Questions around API requests, task status, billing, or integration can interrupt production when support is difficult to reach. Kie.ai provides responsive assistance for Wan 3.0 API users, helping resolve platform and workflow-related issues more efficiently when they arise.