Product image → product videoGenerate a product ad from one image
Upload a product image and keep its shape, materials, logo, and user interaction consistent across the video.


Create AI videos online with Wan 3.0. Generate a native 30-second sequence up to 1080P with synchronized sound, and direct it with images, text, video, audio, documents, or web references. During the current promotion, each user can generate 100 videos for free.
Start with a product image, a written brief, or a character reference. Wan 3.0 turns the details you provide into connected shots with clear action and continuity.
Product image → product videoUpload a product image and keep its shape, materials, logo, and user interaction consistent across the video.
Written brief → story videoDescribe the goal, setting, characters, and key beats. Wan 3.0 arranges them into a video with a clear beginning, middle, and ending.
Character reference → connected shotsStart with a character reference and preserve the face, clothing, props, and space across a continuous story or series.
Wan 3.0 upgrades duration, reference breadth, audio-visual generation, and real-world fidelity in one workflow.
Build a complete narrative in one output, with longer camera movement, one-take sequences, and multi-beat pacing that can unfold naturally.
Combine up to 10 images, 5 videos, and 5 audio clips, then add text, documents, or web pages to carry the details behind the brief.
Generate sound, dialogue, rhythm, and lip-sync as part of the scene, so the picture and soundtrack move together from the first output.
Keep identity, actions, props, spaces, visual style, software UI, charts, and visible text accurate and consistent across the sequence.
Move from idea to finished direction for film, advertising, design, games, and cultural content, with Wan 3.0 Video API access now open.

Give each reference a clear job, then let Wan 3.0 carry the story, sound, and continuity through the final frame.
Wan 3.0 reads visual, verbal, sonic, structured, and web context together, so the source material can carry the details the model should preserve.
| Source | Use it to control | Best production use |
|---|---|---|
| Images | Identity, product shape, composition, style. | Characters, products, environments, and continuity. |
| Text | Action, camera language, pacing, dialogue. | Original direction and fast creative exploration. |
| Video + audio | Motion, rhythm, performance, sound intention. | Reference-led shots, music visuals, and lip-sync. |
| Documents + web | Requirements, structured facts, UI, and visible text. | Brief-led production, knowledge content, and product demos. |
Wan 3.0 can follow a large multimodal brief. Separate identity, action, sound, and output requirements so each input has a clear job.
Source — Use @Image1 for the male character, @Image2 for the female character, and @Image3 for the dojo space. STORY — A 30-second confrontation that moves from a quiet two-shot into fast hand-to-hand action, then settles on a shared final look. AUDIO — Generate natural dialogue, room tone, impact sounds, and synchronized lip movement as part of the scene. KEEP — Preserve both identities, clothing, dojo layout, props, camera axis, and visual style across every beat. Output — Use 1080P with an adaptive ratio. Review faces, actions, spatial relationships, dialogue, and visible text before export.
Answers based on the public Wan 3.0 release information and the current ArtArch promotion.
Try Wan 3.0 during the current promotion, then take the same reference-led workflow into ArtArch Studio.