Mastering Wan 3.0
Alibaba's video model, live on ZenCreator as Wan 3.0 Spicy — up to 30 seconds in one pass, optional native audio, and multiple image/video/audio references depending on the tool.
Written By Leonid
Last updated 18 days ago
What Wan 3.0 is
Wan 3.0 is Alibaba's video-generation model, now live across Text to Video, Image to Video, and Video to Video on ZenCreator, shown in the interface as Wan 3.0 Spicy — the NSFW-capable version of the model for this platform.
Its headline upgrade over earlier Wan generations and over Seedance 2.0 is scale: Wan 3.0 renders up to 30 seconds in a single pass, follows long screenplay-style prompts across multiple shots, writes its own soundtrack, and accepts images, video clips, and audio as references at once — exact limits depend on the tool (see below).
What makes it different
Much longer, multi-shot clips. Up to 30 seconds, written out by timestamp — the model handles the transitions between shots inside one generation, as long as you named them.
Optional native audio. A Generate Audio toggle (off by default) lets Wan 3.0 write its own soundtrack — dialogue, foley, and music — as part of the generation. When it's on and you haven't described the sound, the model invents it, so describe specific audio (or silence) explicitly if that matters.
No negative prompt. There's no separate "what to avoid" field — anything unwanted (extra people, camera shake, background music) has to be stated in the main prompt itself.
Multiple references at once. Images, video clips, and audio can all be attached and addressed directly in the prompt by tag — not just images and audio like Seedance 2.0. In Image to Video: up to 9 images and 3 audio clips. In Video to Video: up to 9 images and 3 video clips (plus audio). Text to Video takes no references — prompt only.
Full resolution range. 480p, 720p, or 1080p, selected per generation — price scales directly with resolution and duration.
Asynchronous generation. There's no progress bar or ETA — a request goes into a queue and a link to the result arrives when it's ready. Typical runs take 15–55 minutes; 30 seconds at 1080p can take considerably longer. Download results promptly — they don't stay available indefinitely.
The four inputs: who's responsible for what
Same mental model as with Seedance — each input type has its own job:
Text controls the spatial — appearance, mood, style, environment.
Image controls identity — exactly how a character, object, or scene looks.
Video controls the temporal — timing, rhythm, camera motion. In Video to Video, the clip you're editing is itself the video reference.
Audio controls rhythm and sound — pacing, how motion syncs to sound.
Rule of thumb: text and image for what things look like, video and audio for how they move and sound.
Reference tag syntax
When you upload reference material, tag it directly in your prompt so the model knows which one you mean — check the placeholder text in the prompt box for the exact live syntax for that tool:
Images: [img1], [img2]... (numbered by upload order)
Videos: [vid1], [vid2]...
Audio: [audio1], [audio2]...
A reference isn't just "inspiration" — it's an addressable object. Tell the prompt exactly what to take from it: a face, an outfit, a product, a style. You can mention the same reference more than once to control exactly where it applies.
Writing the prompt: match its size to the duration
Wan 3.0 was trained on detailed descriptions. A short prompt isn't brevity — it's control handed over to the model, which will fill in the plot, the lines, the music, and the shot count on its own. As a rough guide (5,000 characters is the hard ceiling):
5 seconds — roughly 700–1,200 characters, 1–2 shots: a compact paragraph covering subject, scene, and motion.
10 seconds — roughly 1,200–2,200 characters, 2–3 shots: add camera detail, lighting, and sound.
15 or 20 seconds — roughly 2,200–3,500 characters, about 4 shots: break the clip into shots by timestamp (e.g.
Shot 1 [0–3s]: ...), each with its own framing, action, camera, light, and sound.25 or 30 seconds — roughly 3,500–5,000 characters, 5–8 shots: same shot-by-timestamp structure, just more of them.
Whenever you break a clip into shots, make sure they cover the full duration with no gaps — if your last shot ends before the requested length, the model fills the remainder however it likes.
For frame-to-video generations (animating an uploaded image), describe motion and camera only — the frame already sets the subject's appearance, so restating it fights the picture and the likeness can drift.
Editing finished video
Wan 3.0 can also edit an existing clip instead of generating from scratch: add a prop, remove a person, change the lighting or style, extend the scene, or change a line of dialogue. These edits need much shorter prompts than generation — what to edit + what to do is usually enough, since the model takes everything else from the source video.
Common problems and how to fix them
Result looks generic → weak subject/action in the prompt. Put subject and one clear action in the first words.
Character changes between generations → reuse the same reference image, tag it consistently, and add "maintaining consistent [subject]."
Camera motion feels random → name the move explicitly, ideally with a number (e.g. "dolly forward at 2 ft/s") rather than an adverb like "slowly" — Wan 3.0 follows numeric speed literally.
Ending feels rushed or improvised → the shots didn't cover the full requested duration; make sure your last timestamp equals the length you ordered.
Unwanted sound or dialogue appears (with Generate Audio on) → there's no negative prompt — state explicitly what should not happen ("no dialogue," "no background music," "silence") if that's the goal.
Want to go deeper?
This article covers the essentials. For the full walkthrough — all seven prompt formulas, shot-by-shot timing, lighting and color-grading vocabulary, and a library of real Wan 3.0 examples with their prompts — see the full Wan 3.0 course inside ZenCreator.