Text, image, video and audio
Combine written direction with reference assets. Images, video and audio give the model context for the scene's look and movement.
Get Seedance 2.0 access for cinematic text-to-video and image-to-video. Direct AI video with image, video and audio references, multi-shot storytelling, and synchronized sound.
Combine written direction with reference assets. Images, video and audio give the model context for the scene's look and movement.
Generate video with dialogue, ambient sound and music in one pass, including lip-sync in 8+ languages.
Direct character movement, camera angles and environmental sound across cuts to keep the story connected.
Add up to 12 references in total, shared across all types: at most nine images, three videos and three audio clips. Video and audio references can each run up to 15 seconds.
Use @-reference tags to assign the first frame, camera reference or background music, then describe the action.
Generate a 4–15 second sequence at up to 4K and 24 fps. Choose a cinematic, square or vertical aspect ratio for delivery.
| What changes | What you direct | What the workflow avoids |
|---|---|---|
| A reference-led scene | Which image, clip, or audio matters | Repeating the full visual brief for each output |
| Audio in the generation | Dialogue, sound effects, or background music | Manual audio alignment as the first step |
| A multi-shot idea | Character movement, scene details, and camera angle | Treating every shot as an unrelated prompt |
Turn prompts and references into short-form video for TikTok, YouTube Shorts and Instagram Reels.
Previsualize scenes, explore B-roll and prototype camera sequences before a shoot.
Animate product photos with sound. Explore e-commerce workflows or compare models for your next campaign.
Get access to Seedance 2.0 and all premium AI models with one membership.
For hobbyists and explorers
For creators and pro users
For power users and teams
For teams and studios
You can use the Seedance 2.0 model directly on YouArt. Choose the model, add your prompt and optional reference assets, then start a generation.
It supports text, images, video, and audio. One generation accepts up to nine images, three video clips, and three audio clips (12 reference files in total, shared across all types).
Yes. The model generates video and audio at the same time. It supports dialogue with lip-sync in 8+ languages, ambient sound effects, and background music.
No. Start with a short direction and add only the references that are important to the scene. The model handles generation; your job is to define the creative direction clearly.
Output is available at up to 4K resolution and 24 fps. Videos can be 4–15 seconds long, with 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios.
The practical difference is simple: you direct a scene instead of generating a generic clip. Give each input a clear role. Then review the result as a coherent piece of video and audio.
Reads text, images, video, and audio together. More context for one generation
Supports multi-shot storytelling. A clearer flow from shot to shot
Generates video and audio at the same time. Less separate audio work
Uses @-reference tags to assign roles. A more direct way to describe intent
Seedance 2.0 is for creators, filmmakers, marketers, and e-commerce teams who need more than a single silent clip. It is a strong fit when you want to use reference assets and direct the story, sound, or camera role together. It is less suitable when you only need a static image. Use it when you want to turn visual inputs into a short video sequence.