Alibaba Wan 3.0 · 30s with sound

Wan 3.0 AI Video Generator

Type a line or drop in a photo. Get back up to 30 seconds with the sound already in it. First clip free, no card.

0 / 20,000

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

SIGN UP FREE · 1 CLIP · 480P

Artificial AnalysisIlia Trikoz — GPTunneLOur Code WorldAirMoreLoved by 5,728 creators
  • First clip free · no card
  • 30 seconds in one pass
  • Sound in the same file
  • 480P / 720P / 1080P
  • Failed runs cost nothing

ReceiptYour free clip is Wan 3.0, not 2.7model ID and task ID printed under every result

Every clip on this page shows its receipt

model
wan3.0-video
resolution
848×480
length
5.0s
frame rate
30 fps
audio
on
task
0385dc79…
See the raw response

This is the JSON Alibaba's API returned for this clip, untouched. Every field on the strip above is read out of it. If the two ever disagree, the strip is the bug.

Served through Alibaba Cloud Model Studio. TRY WAN3 is an independent interface — we are not Alibaba, and we cannot grant rights the model provider does not grant.

Prompts and what they produced

Seven Wan 3.0 prompts — each one beside the clip it describes

Every card shows a clip and the exact prompt that produced it — not a summary. One press loads it into the generator at the badge's ratio, length and resolution.

  • 6 of 7Alibaba's own, with their prompt
  • 1 of 7read back from the footage — labelled
  • 0rewritten or shortened
  • Our own runs replace these the day the key is live
A pink mech suit firing a beam at a green armoured brute across a concrete plaza16:9 · 720P · 15S

The full prompt, not a summary

A teenage girl in an oversized grey sweater and a pleated skirt walks across a campus forecourt in the middle of the day, other students in uniform passing without looking at her; her right arm is a glossy pink prosthetic. Then she stops and armour plates unfold from that arm and travel across her body until she is inside a full pink mech suit, the visor closing last. After that the camera…

Sound is written into the same prompt

Street noise and footsteps first, then servo whine and plate locks, then a low rumble under the brute, no music.

720P · 15s

One prompt, one pass, fifteen seconds — three shot sizes and no cut

  • A red supercar unfolding into a four-legged machine on wet asphalt under floodlightsText
    A car becomes an animal in four beats, and the fourth one is the ending
  • A first-person view charging across a blue glacier canyon past a wrecked ship, with ice crusting over both armsText
    Fifteen seconds of one continuous move, across the glacier
  • A silver-haired navigator holds a cyan star map opposite a black-coated operative in an orbital scrapyardReference
    Two reference subjects held across an orbital handoff
  • A ruined overgrown avenue, a child holding a light in a derelict store, and an astronaut lifting off her helmetText
    Astronaut in the ruins
  • A young man and an elderly woman play chess at a table in a busy city squareText
    Street chess turnaround
  • A small ranger reaches toward an enormous green dragon lying in a sunlit forest clearingImage
    The dragon opens its eyes
  • A masked rider on a black horse gallops alongside a steam train across a sunset desertReference
    Desert train heist

What moves a Wan 3.0 result

  • Camera

    Name one move. Not "cinematic".

  • Sound

    Say what it sounds like — silences included.

  • Change

    Describe how the frame changes across its length, not one frozen moment.

  • Dialogue

    Its own sentence, or it gets narrated instead of spoken.

Four ways in

Four ways into the Wan 3.0 AI video generator

One model covers all four. Wan 3.0 folded text, image, reference and editing into a single endpoint, so you are not switching products between jobs — and a reference set up for one shot still applies to the next.

  • Five dancers in black dresses silhouetted against a hard white backlight in ground fog

    Text to video

    One prompt, up to 20,000 characters, carries both the picture and the sound.

    20,000 characters
  • A man lowers his camera in a European square and sees himself standing on a distant clock tower

    Image to video

    A first frame, a last frame, or both — the motion in between is generated.

    First, last, or both
  • A woman in white on a lime-green sofa in a flower meadow as a butterfly lands on her arm

    Reference to video

    Ten images, five video clips and five audio tracks in one run, addressed by number.

    10 · 5 · 5
  • A woman suspended mid-air in a red-walled attic workshop as beams and chains float around her

    Video editing

    Describe a change to footage you already have; it rewrites rather than regenerates.

    Rewrite, not regenerate

Text to video

Wan 3.0 text to video: what one prompt has to carry

A prompt here does two jobs at once: the picture and the soundtrack come out of the same pass. Beside this is one prompt revealed a layer at a time — subject, camera, light, sound — with the clip changing as each lands.

One prompt, four layers

Five dancers in black chiffon stand in a bank of ground fog,
lit from directly behind by one hard white source, no fill.
They gather, they open outward, they draw back in and hold.
Breath, bare feet on a wooden floor, the rustle of fabric. No music.

  • Subject · camera · light
  • Sound — same prompt
876 / 20,000

One pass out

Five dancers in black dresses silhouetted against a hard white backlight in ground fog16:9 · 720P · 12S

wan3.0-video · 16:9 · 720P · 12s · sound in the same pass

Those four lines are the shape of the prompt behind this clip; the whole thing runs to 876 characters and is printed in full on [/prompt-ideas](/prompt-ideas). Wan 3.0 takes 20,000 — about 3,000 words, or a twelve-page script, for a clip of at most thirty seconds. The room is for being specific, not for writing more.

Image to video

Wan 3.0 image to video: first frame, last frame, or both

Give Wan 3.0 one still and it starts there. Give it two and it has to arrive — a last frame is a destination, not a suggestion, which is what makes two-ended runs the predictable way to work.

Your still

The frame the run started from: a man with a camera in a busy European square
1 image · First frame

Generated video

The man lowers the camera, turns toward a distant clock tower and sees himself standing on itWAN 3.0 · 16:9 · 720P · 5S

wan3.0-video-prime · 16:9 · 720P · 5s

First frame: the square, the camera still in his hands

Wan 3.0 generates everything between

Last frame: his face, the square thrown out of focus behind him

First frame

Last frame

  • 0 images = text to video
  • 1 image = one anchor
  • 2 images = both ends pinned
  • JPEG · PNG · WebP · BMP · ≤20MB
  • No alpha channel — flatten transparent PNGs

JPEG, PNG, WebP or BMP, up to 20MB. The input that fails without an obvious reason is a PNG carrying transparency: Wan 3.0 takes no alpha channel, so flatten it before you upload.

Reference to video

Wan 3.0 reference to video: one face across the whole take

Load up to 10 images, 5 video clips and 5 audio tracks into one generation and point at them by number — Image 1, Video 2. Beside this is one reference carried across a whole take: same subject, first second to last.

What the prompt cites

  • Character reference: the performer in the white hooded coverall

    Image 1 · Character

Generated shot

A woman in white on a lime-green sofa in a flower meadow as a butterfly lands on her arm16:9 · 5S

Same performer, same outfit · 16:9 · 5s

10 images · 5 clips · 5 audio
  • Images ≤ 10 — free
  • Clips ≤ 5 · 15s total — billed
  • Audio ≤ 5 — free
  • Reference seconds count against the 30s ceiling

Images and audio attach for free. Reference video does not: its seconds bill at your output rate and count against the same 30-second ceiling, so a 15-second reference plus a 15-second result is the longest run you can ask for.

Document to video

Turn a PPT, PDF or web page into video

The part other video models do not do. Give Wan 3.0 a deck, a report or a public URL and it reads what is inside and generates footage from the content — new video, not your slides with an avatar reading the bullets.

What went in

haulage-platform-launch.pdf3 pages · 1.1 MB
  • A long-haul chassis that reconfigures in the field
  • Eight minutes from road profile to working profile
  • First customer trials on the interstate corridor in Q4

What came back

A long-haul rig on an open highway unfolding into a walking machineSAMPLE · 16:9 · 15S

A sample render standing in for the output, and the same clip the deck-to-teaser card below uses. Our own document runs land here once the upstream key is live.

One file or one link per run · up to 100MB · up to 50 pages · doc, xls, ppt, pdf, txt, key, pages, numbers, md

It works best when the document carries a story the video can retell — a launch deck, a feature spec, a set of results worth dramatising. It works badly when the deck itself is the deliverable, or when the numbers on the slide have to appear on screen exactly as typed: Wan 3.0 renders text well for a video model, but it is generating that text, not copying it. Check anything legal or financial frame by frame before you publish.

Use cases

Six jobs people are already using Wan 3.0 for

Six jobs, each with a runnable prompt and the clip that prompt describes. The clips are sample renders rather than our own output until the upstream key is live, and where a job is a bad fit for Wan 3.0 we say so rather than leave it out.

  • A woman in a black flight suit lying back on sunset cloud in over-ear headphones
    16:9 · 30s · 960 credits

    Product cut for a launch post

    Thirty seconds of Wan 3.0 is long enough to place the product, live with it for a beat and land on the hero shot — one take, with the sound already under it.

  • A long-haul rig on an open highway unfolding into a walking machine
    16:9 · 20s · 640 credits

    Turning a deck into a teaser

    Drop the launch PDF in and Wan 3.0 builds the visuals out of what the slides claim — footage that dramatises the story, not your slides with an avatar reading them.

  • Two men in suits struggling against the mirrors of a tiled office washroom
    16:9 · 10s · 320 credits

    A talking scene that lip-syncs

    Put the line in braces and Wan 3.0 speaks it rather than narrating it. Everything outside the braces — who says it, how, and what happens next — stays description.

  • The same red-haired woman held across a wide, a lift interior and a close-up in a shopping atrium
    16:9 · 15s · 480 credits

    Keeping one character across shots

    Load the character reference once and cite it by number. The same face, hair and outfit survive a change of shot size and a change of light inside one take.

  • A hooded fighter and a young man with a glowing blue blade, in the rain
    16:9 · 8s in, 8s out · 512 credits

    Fixing a shot you already have

    Upload the footage and describe the change. Wan 3.0 rewrites it rather than regenerating, so the blocking and the timing you liked stay where they were.

  • A long-haired man turning on wet rocks at dusk as a flock of birds crosses the frame
    16:9 · 12s · 96 credits at 480P

    B-roll from a single photo

    One still becomes twelve seconds of movement. Cheapest thing Wan 3.0 does, and the fastest to iterate on.

Where Wan 3.0 is the wrong tool: hard-locked on-screen text that has to match a legal document word for word, anything longer than thirty seconds in a single pass, and 4K delivery, which it does not support. For those, generate here and finish in an editor.

Specs, checked line by line

Wan 3.0 specs, checked against Alibaba's own docs

Every row links to the page it came from and carries the date we last checked it. When Alibaba changes something, this table changes — and if you find a row that is out of date, tell us and we will fix it the same day.

Source= links to the primary page

A first-person view charging across a blue glacier canyon past a wrecked ship, with ice crusting over both armsWAN 3.0 · 16:9 · 720P · 15S
Fifteen seconds, one pass, no cut — every figure on that badge is a row in the table below
Duration

2–30 seconds, any whole number

Model Studio
Frame rate

30 fps

Model Studio
Resolution

480P · 720P · 1080P

Model Studio
4K

Not supported

1080P is the top tier. Anything sold as 4K is an upscaler.

Model Studio
Aspect ratio

adaptive, 16:9, 4:3, 1:1, 3:4, 9:16

Model Studio
Prompt limit

20,000 characters

Model Studio
Reference inputs

10 images · 5 clips · 5 audio

Model Studio
Document input

1 file or 1 link · ≤100MB · ≤50 pages

Model Studio
Audio

Generated in the same pass; on or off costs the same

Model Studio
Open weights

None published. Latest open release is Wan 2.2

Hugging Face

No, Wan 3.0 does not output 4K

Some sites advertise a 4K version. Alibaba's API reference lists three resolution tiers and 1080P is the top one. If a service hands you 4K here, that is an upscaler running after the model, not Wan 3.0 producing those pixels. We would rather lose the click than repeat a number we cannot source.

Checked by Hao Lin against the pages linked above on 25 August 2026. If a row has gone stale, tell us — this date moves when the table does.

How it compares

Wan 3.0 against the models people cross-shop it with

Published numbers only. Where a spec is not on a first-party page we leave the row out rather than fill it with a hedge — a comparison table is the easiest place on the internet to get caught guessing.

  • Max single-pass length

    Wan 3.0
    2–30s
    MiniMax H3
    4–15s
  • Frame rate

    Wan 3.0
    30 fps
    MiniMax H3
    24 fps
  • Resolution tiers

    Wan 3.0
    480P · 720P · 1080P
    MiniMax H3
    768P · 2K
  • Prompt limit

    Wan 3.0
    20,000 characters
    MiniMax H3
    7,000 characters
  • Reference inputs

    Wan 3.0
    10 images · 5 clips · 5 audio
    MiniMax H3
    9 images · 3 clips · 3 audio
  • Document or web page input

    Wan 3.0
    Yes
    MiniMax H3
    No
  • Editing a clip without regenerating it

    Wan 3.0
    Yes
    MiniMax H3
    No
  • Open weights

    Wan 3.0
    No
    MiniMax H3
    YesTheirs
  • Blind-test rank, with audio

    Wan 3.0
    #1, stated range 1–2
    MiniMax H3
    #3

And against Wan 2.7, which is what most sites are still running

Alibaba's own model table puts Wan 2.7 at 2–15 seconds, 720P and 1080P, 30 fps. Wan 3.0 doubles the ceiling to thirty seconds — but the change you will feel every day is the tier it added underneath: Wan 2.7 has nothing below 720P, and Wan 3.0 has 480P. That is where you iterate before you spend, and it is why a draft here costs a quarter of a final. The full version-to-version comparison goes through what else moved.

The honest read: pick MiniMax H3 when you need to run the weights on your own hardware, and pick Wan 3.0 when the clip has to run long, carry its own sound, or start from a document you already wrote. The three we did not table — Veo 3.1, Seedance 2.5 and Kling — each have a full comparison of their own:

Third-party blind tests

Where Wan 3.0 ranks against other video models

Artificial Analysis runs blind head-to-head preference tests and publishes the Elo, the confidence interval and the number of comparisons behind it. Read all three before you quote the rank.

  • Text-to-video, with audio

    Range 1–2

    #1

    9001500
    Elo
    1,240
    95% CI
    ±9
    Comparisons
    5,728
    Second place is 3 points behind — inside the interval
  • Second place

    #2

    9001500
    Elo
    1,237
    Gap
    −3
  • Third place

    #3

    9001500
    Elo
    1,227
    Gap
    −13

On the Artificial Analysis text-to-video arena, in the with-audio table, Wan 3.0 sits first with an Elo of 1,240 from 5,728 blind comparisons — snapshot 25 August 2026. Read that carefully before you quote it: second place is three points behind and the interval is ±9, so the leaderboard itself gives Wan 3.0 a range of 1–2 rather than a clean win.

What the table shows more clearly is price. At 720P, Wan 3.0 works out to roughly $6.00 per minute of finished video — level with the model ranked second and below the one ranked third.

Source: Artificial Analysis Video Arena, With Audio board ↗ · snapshot 2026-08-25

Pricing · one model at every tier · sound included

One model at every tier — pick the size of your month

Every plan runs the same wan3.0-video at all three resolutions, the full thirty seconds, with sound and commercial use. The tiers differ in how many seconds you get and how many takes come back at once — never in what the model is allowed to do.

Same monthly credits either way

  • Free first clipNo card

    One generation, once, on the real `wan3.0-video`. A sample rather than an allowance — every free second is cash spent before any arrives.

    $0once

    No card at any point

    Generate free clip

    An email, no card

    1 clip, once

    480P · up to 3s · with sound
    An email, no card

    Wan 3.0 has no free tier of its own — Alibaba meters it from the first second

    • The real wan3.0-video, not 2.7
    • Sound generated with the picture
    • Model ID and task ID on the receipt
    • A failed run never costs you
    • 720P and 1080PPaid
    • Clips longer than 3 secondsPaid
    • Reference, document and editing modesPaid
    • More than one clip, everPaid

    Anyone advertising unmetered free access is either serving a different model or paying a bill that will not last.

    At a glance

    Clips a month
    1, once
    Takes per run
    1
    Longest clip
    3s
    Top resolution
    480P
  • StarterSave 50%

    A launch post a week. Enough to find out whether Wan 3.0 suits the way you work, without a decision about volume.

    $9.90/mo$19.80

    $118.80 billed yearly

    Cancel anytime

    480 credits / month · ≈ 12 clips

    at 5s 480P · 6 at 720P · 3 at 1080P

    Plan credits expire at the end of the month

    • **480 credits a month** — about 12 finished clips
    • 2 takes of one prompt per run
    • Every resolution — 480P, 720P and 1080P
    • Every length — 2 to 30 seconds in one pass
    • Every input — text, image, reference, document, web page
    • Sound written in the same pass, no watermark
    • Commercial use included — publish it, sell it, bill for it
    • Wan 3.0 Prime, the fast tier, whenever you want it
    • A failed run never costs credits
    • 7-day refund while the credits are untouched
    • Cancel in two clicks — the price you joined at is locked

    Everything under the first two lines is on this plan at $9.90. The bigger plans buy seconds and takes — never a better model.

    At a glance

    Clips a month
    ≈ 12 at 5s 480P
    Takes per run
    2
    Credits per dollar
    Baseline
    Refund window
    7 days
  • Most people land here
    Pro

    One finished clip every working day, with three alternates each. The draft-at-480P, finish-at-1080P loop, sized for a working week.

    $29.90/mo$59.80

    $358.80 billed yearly

    Cancel anytime

    1,760 credits / month · ≈ 44 clips

    at 5s 480P · 22 at 720P · 11 at 1080P

    Plan credits expire at the end of the month

    • **1,760 credits a month** — 3.7× Starter
    • ≈ 44 finished clips a month, or 11 at full 1080P
    • **4 takes of one prompt per run** — pick one instead of re-rolling2× takes
    • **+32% credits per dollar** than StarterBetter rate
    • Room to draft at 480P and finish at 1080P in the same week
    • Every resolution — 480P, 720P and 1080P
    • Every length — 2 to 30 seconds in one pass
    • Every input — text, image, reference, document, web page
    • Sound written in the same pass, no watermark
    • Commercial use included — publish it, sell it, bill for it
    • Wan 3.0 Prime, the fast tier, whenever you want it
    • A failed run never costs credits
    • 7-day refund while the credits are untouched
    • Cancel in two clicks — the price you joined at is locked

    The jump from Starter is quantity, takes and rate. The model, the resolutions and the thirty seconds were already yours.

    At a glance

    Clips a month
    ≈ 44 at 5s 480P
    Takes per run
    4
    Credits per dollar
    +32% vs Starter
    Refund window
    7 days
  • StudioBest rate

    Client volume, and the plan where you stop rationing 1080P. Forty full-resolution finished clips a month, at the lowest credit price on this page.

    $79.90/mo$159.80

    $958.80 billed yearly

    Cancel anytime

    4,960 credits / month · ≈ 124 clips

    at 5s 480P · 62 at 720P · 31 at 1080P

    Plan credits expire at the end of the month

    • **4,960 credits a month** — 10× Starter, 2.8× Pro13×
    • ≈ 124 finished clips a month, or **31 at full 1080P**
    • **+65% credits per dollar** — the best rate on this pageBest rate
    • Enough 1080P that you stop drafting at 480P first
    • 4 takes of one prompt per run
    • Sized for a client roster rather than one channel
    • Top up any month without changing plan
    • Every resolution — 480P, 720P and 1080P
    • Every length — 2 to 30 seconds in one pass
    • Every input — text, image, reference, document, web page
    • Sound written in the same pass, no watermark
    • Commercial use included — publish it, sell it, bill for it
    • Wan 3.0 Prime, the fast tier, whenever you want it
    • A failed run never costs credits
    • 7-day refund while the credits are untouched
    • Cancel in two clicks — the price you joined at is locked

    Studio is volume and rate, nothing else. If you are not spending a month’s credits, Pro is the honest answer — we would rather write that here than take the difference.

    At a glance

    Clips a month
    ≈ 124 at 5s 480P
    At full 1080P
    ≈ 40 a month
    Credits per dollar
    +65% vs Starter
    Refund window
    7 days

Now check one clip

What one Wan 3.0 video actually costs

Move the slider and the number updates before you commit. Credits come off when the clip comes back, not when you press the button — a failed run costs nothing. Every plan runs the same model at all three resolutions.

5s · 720P

This clip: 80 credits

≈ $2.72 per finished 5-second 1080P clip on Pro · ≈ $0.68 at 480P

Resolution

Failed runs cost 0. Drafting at 480P costs a quarter of 1080P.

A worked example, because per-second numbers are hard to feel: a 12-second clip at 720P is the length most people land on for a product cut. A 5-second 480P draft of the same shot costs a fraction of that, which is why the sensible loop is to iterate at 480P until the framing is right and spend the 720P run once. Batch mode charges per clip, so four variants cost four times one — there is no bulk discount and we are not going to pretend there is.

Free first clip

Wan 3.0 has no free tier of its own — Alibaba meters it from the first second, and anyone offering unmetered access to Wan 3.0 is either running a different model or absorbing a bill they cannot absorb for long. Your first clip here is free because we pay for it: one generation, once, 480P, up to three seconds, on the real `wan3.0-video`. An email, no card. It is a sample rather than an allowance. After that, plans start at $9.90 a month.

  • Five seconds each; longer clips cost proportionally more
  • Yearly pricing shown — monthly is double for the same credits
  • Failed runs cost 0

Independent reviews

Four published Wan 3.0 reviews, quoted — including the unflattering ones

One benchmark and three reviews, none of them ours and none of them Alibaba’s. Every quotation below is a link to the page it was published on, with the date we read it. Two of the four are corrections rather than praise, and both stay up.

Video Arena · With Audio · snapshot 2026-08-25
Wan 3.0 tops the with-audio text-to-video board at Elo 1,240 — but on 5,728 comparisons, the fewest in the top five, and the board itself states its range as 1–2. Second place is three points behind, which is inside the error bar.
Artificial AnalysisArtificial AnalysisVideo Arena · With AudioRead 2026-08-25
Open the board

Every line above is a quotation with a link, not a summary. Alibaba's own announcement is a source too, but it is not independent, so it sits in the specs table further up instead of on this wall.

How to use

How to use the Wan 3.0 AI video generator in three steps

  1. A neon-lit street at night seen from behind a lone figure

    A neon high street at night seen from behind one walking figure, footsteps, rain on awnings and distant traffic

    PictureSound

    Describe or upload

    Type the shot, or drop in an image, a reference clip or a document. Naming the camera move and the sound gets you closer than stacking adjectives.

    Same prompt · same pass
    • 16:9 frame16:9
    • 4:3 frame4:3
    • 1:1 frame1:1
    • 3:4 frame3:4
    • 9:16 frame9:16
    23456789101112131415161718192021222324252627282930

    2–30s · whole seconds · 30 fps

    Pick length and resolution

    Anything from 2 to 30 seconds. Draft at 480P and finish at 1080P — same model, a quarter of the cost per second.

    30 fps · 480P / 720P / 1080P
  2. The finished clip playing back with its audio track
    Speech · effects · musicclip.mp4 · 1080P

    Download the MP4

    Audio is already mixed in. There is no render queue to join afterwards and no editor to open.

    Sound already in the file

About 90 seconds for a short clip, longer for 30 seconds at 1080P. You can close the tab; the result waits for you.

FAQ · Thirteen questions

Wan 3.0 questions people actually ask

Thirteen questions, answered on this page.

Is this the real Wan 3.0?

Yes. Every request goes to Alibaba's model ID wan3.0-video, and the receipt under each clip shows the model, the resolution, the length and the task ID that came back from the API. There is a button on that strip that opens the raw response if you want to read it yourself. We are an independent interface, not Alibaba.

Do I have to sign up?

Yes, for one email. There is no card at any point, and the free clip is one generation for the life of the account — a sample rather than an allowance. The three prompt tools on this site — prompt generator, video to prompt and prompt ideas — need no account at all.

Is Wan 3.0 free?

The model is not — Alibaba charges per second of output from the first second. Your first clip on this site is free because we cover the cost. Anyone advertising unmetered free access is either serving a different model behind the label or paying a bill that will not last.

Is Wan 3.0 open source? Can I download the weights?

No. There are no Wan 3.0 weights on Hugging Face, GitHub or ModelScope, and no published licence. The most recent open-weights release in the family is Wan 2.2 under Apache 2.0 — if running locally matters more to you than the 30-second ceiling, that is the one to get.

Does Wan 3.0 do 4K?

No. Three tiers: 480P, 720P and 1080P. Anything sold as 4K is an upscaler added afterwards by whoever sold it to you.

How long can one clip be?

Two to thirty seconds, any whole number. If you feed in reference video, the input length plus the output length has to stay under thirty seconds in total.

Does it generate sound?

Yes, in the same pass as the picture — speech, effects and music. You can turn audio off, and the price is the same either way.

Can I use the videos commercially?

Yes, subject to Alibaba's usage policy for Wan 3.0 and to our terms. We are not authorised to grant rights beyond what the model provider grants, and any site telling you otherwise is guessing.

What happens when a generation fails?

Credits go back automatically — no ticket, no email. If a job sits unclaimed for 24 hours it expires upstream, and we label that expired rather than failed, because they are different problems with different fixes.

How long does one generation take?

Around ninety seconds for a short 480P clip. Thirty seconds at 1080P takes noticeably longer, and queue time varies with how busy the upstream is — Wan 3.0 allows a small number of concurrent jobs per account, so at peak your request may wait before it starts. You can close the tab; the finished clip is waiting when you come back.

Can I share a clip I made?

Every finished clip gets its own page with the video, the prompt and the receipt. It stays private until you press share, and the link can be revoked afterwards.

Can I put a real person in a Wan 3.0 video?

Only someone who has agreed to it. Reference-to-video will hold a face across a whole take, which is exactly the capability that makes non-consensual use easy, so uploading a recognisable person you have no permission to use breaks our terms and Alibaba's. Reports go to a monitored inbox and takedowns are processed quickly.

How is this different from a PPT-to-video tool?

Those keep your slides and add narration over them. Wan 3.0 reads what is in the file and generates new footage from the content. Different output, different job — pick the slideshow tool when the deck itself is the deliverable.

Every answer is in the page itself — open a question to read it

One clip, no card

Make your first Wan 3.0 video

The Wan 3.0 AI video generator at the top of this page is the one paying accounts use — same model, higher ceiling. Miss the shot and it cost you nothing.

  • 480P · up to 3 seconds
  • Sound in the same file
  • An email, no card
  • A failed run is never billed

SIGN UP FREE · 1 CLIP · 480P

Not ready to hand over an email? Three things here need no account at all — the prompt generator, video to prompt and the prompt library. They are free because text is nearly free to generate, and we would rather say that than pretend the video is.

Written and maintained by the TRY WAN3 editorial teamPublished Last updated