coming soon

FLUX 3

Single generationUp to 20s · with audio

Stop assembling your pipeline from four different models. FLUX 3 is Black Forest Labs' answer to a workflow where the image tool doesn't know what the video tool is doing, and neither one can hear the audio: one model, trained jointly across images, video and sound from the very first step. A still becomes twenty seconds of motion with dialogue already in sync, and the same visual logic holds the whole way through.

What changes

Three shifts that change how you build

One model

Image, video and audio together

FLUX 3 covers image generation and editing, video, and sound in a single model rather than a chain of specialists — so what you establish in a still carries into the motion and the mix.

Native duration

20 seconds in one generation

Up to 20 seconds of video with audio in a single pass — long enough to carry a whole beat. Longer pieces chain clips into multi-shot sequences.

Sound that belongs

Dialogue in multiple languages

Native audio arrives with the video — multilingual dialogue included — instead of being layered on afterwards and nudged into sync.

Built for real production

Why one multimodal model matters

Every handoff between tools is a place where your look drifts. FLUX 3 collapses those handoffs into a single model, so the thing you approved in the first frame is still there in the last.

Production pressureA different model for every modality
The FLUX 3 advantageOne jointly trained model
Concrete outcomeA consistent look from still to motion to sound
Production pressureClips too short to carry a beat
The FLUX 3 advantageUp to 20 seconds in one generation
Concrete outcomeA complete moment, not a fragment
Production pressureSound bolted on in post
The FLUX 3 advantageNative audio with the generation
Concrete outcomeDialogue already sitting in the picture
Production pressureReversioning for every market
The FLUX 3 advantageMultilingual dialogue
Concrete outcomeOne shoot, several language cuts

Quick fit check

Best for

Teams tired of stitching an image model to a video model to an audio tool, who want one consistent system end to end.

Where it lands first

A scene already made with FLUX 2 on YouArt, pushed into motion with dialogue — that handoff is the one FLUX 3 is designed to remove.

Not ideal for

Work you need to ship this week — FLUX 3 isn't on YouArt yet, so FLUX 2 is the one for today.

How it works

What makes one model behave like one model

Bolting modalities together after the fact gives you a pipeline. Training them together from the start gives you a model that understands how a scene looks, moves and sounds as one thing.

Joint training

Trained across images, video and audio at once

One model, jointly trained across images, video and audio from the beginning — not separate systems wired together after the fact. That's why a character established in a still holds through the motion, and why the sound lands on the action instead of near it.

Twenty seconds, one pass

Long enough to actually be a shot

Twenty seconds of video with audio, in a single generation. That's the difference between a fragment you have to cut around and a moment that can carry a line of dialogue, a reaction and a camera move without a seam in the middle.

Longer sequences

Chain clips into multi-shot scenes

For pieces that run past a single generation, clips chain into longer multi-shot sequences — so scale comes from arranging shots, not from fighting one generation to stretch further than it should.

Workflow

From still to finished scene

The point of one multimodal model is that each step inherits the last instead of restarting from a text prompt.

1

It starts with a still

The most controllable step comes first. Character, palette and framing are settled in a single frame, before anything begins to move.

2

The frame moves

That same frame carries into up to 20 seconds of video. Because one model handles both, the look that was approved is the look that moves.

3

And it speaks

Native audio and multilingual dialogue arrive with the generation, so the scene comes with sound rather than waiting on an alignment pass.

From prompt to finished scene

Everything the creative journey needs

A finished piece needs more than a first generation — it needs a way to keep going without losing what you already had.

Video and audio continuation

FLUX 3 can continue from an input video and its audio, carrying a shot forward so pacing and sound stay coherent when a scene needs more room than the first generation gave it.

Image generation and editing

Image synthesis and editing sit in the same model as the video, so the still you refine and the frame you animate come from one system instead of two that only roughly agree.

Built for production

Built for what production actually needs

FLUX 3 is built around the parts of the job that usually cost the most time: keeping a look consistent across modalities, getting clips long enough to use, landing sound without a separate pass, and covering more than one language.

The look survives the handoff

One jointly trained model means the character, palette and lighting you set in a still are still there once the shot is moving.

Twenty seconds is a usable shot

Long enough for a line of dialogue and a reaction, instead of a fragment you have to cut around.

Sound arrives with the picture

Native audio generation removes the alignment pass that normally sits between a finished render and a finished scene.

More than one market

Multilingual dialogue means a second language version is a generation, not a reshoot.

Compared with FLUX 2

What's new versus FLUX 2

Modalities
FLUX 2Images — text-to-image and image editing in one model.
FLUX 3Image, video, audio and action prediction in one model.
Video
FLUX 2Not a video model.
FLUX 3Up to 20 seconds with audio in a single generation.
Audio
FLUX 2None.
FLUX 3Native audio generation with multilingual dialogue.
Architecture
FLUX 2Latent flow matching, coupling a 24B vision-language model with a rectified flow transformer.
FLUX 3Built on Self-Flow, jointly trained across images, video and audio.
Availability
FLUX 2Available on YouArt today, with create and edit modes.
FLUX 3Coming soon to YouArt.

Final specs land on the model page when FLUX 3 arrives on YouArt. Video opened first in early access, with the image model following.

Across industries

FLUX 3 across industries

A single model spanning image, video and audio changes the shape of the work differently depending on what you make.

Digital advertising

Agencies build a campaign frame, extend it into a twenty-second spot and deliver it in several languages, without the look drifting between the key visual and the cutdown.

Retail & e-commerce

Retailers turn a product still into a short piece of motion with sound, keeping the exact shape and colour established in the image.

Entertainment & film

Studios previsualize a beat with dialogue already in place, so a board becomes something that can actually be watched and timed.

Corporate communications

Teams produce training and internal video where a spokesperson stays identical across modules, and reversioning for another market is one more generation.

Who it's for

Who FLUX 3 is built for

One model across modalities matters most to the people currently paying the tax of moving between three.

Creative directors

Creative directors hold one visual idea across a still, a moving shot and a soundtrack, instead of re-establishing it in every tool.

Performance marketers

Performance marketers spin variants and language cuts from the same look, without a reshoot for each one.

Filmmakers

Filmmakers previsualize scenes with dialogue in place, so pacing can be judged rather than imagined.

Social creators

Social creators go from an idea to a short piece with sound already in it, without assembling a toolchain to get there.

Plans and pricing

Generate with FLUX 2, Flux Kontext and every other model on YouArt using a single credit balance.

Basic

For hobbyists and explorers

$9.99/mo
Select Plan
  • 1000 credits
  • Up to ~200 images/month
  • At least ~1000s video/month
  • Intelligent creative agent
  • Video editor
  • Latest image models, including GPT Image 2 and Nano Banana Pro
  • Latest video models, including Seedance 2
  • Realistic face uploads
  • Voice generation with ElevenLabs
  • No watermark
  • Unlimited template access

Pro

Most Popular

For creators and pro users

$29.99/mo
Select Plan
  • 3300 credits
  • Up to ~1000 images/month
  • At least ~3300s video/month
  • Intelligent creative agent
  • Video editor
  • Latest image models, including GPT Image 2 and Nano Banana Pro
  • Latest video models, including Seedance 2
  • Realistic face uploads
  • Voice generation with ElevenLabs
  • No watermark
  • Unlimited template access

Max

For power users and teams

$149.99/mo
Select Plan
  • 18000 credits
  • Up to ~6000 images/month
  • At least ~18000s video/month
  • Intelligent creative agent
  • Video editor
  • Latest image models, including GPT Image 2 and Nano Banana Pro
  • Latest video models, including Seedance 2
  • Realistic face uploads
  • Voice generation with ElevenLabs
  • No watermark
  • Unlimited template access

Team

For teams and studios

$329.99/mo
Select Plan
  • 36300 credits
  • Up to ~12000 images/month
  • At least ~36300s video/month
  • Realistic face uploads
  • Share canvas, workflows, and assets with your team
  • Up to 10 members per team
  • Up to 5 teams
  • Role-based management
  • Team credit management and spending caps
  • Per-member usage tracking

FAQ

FLUX 3 questions

What is FLUX 3?
FLUX 3 is Black Forest Labs' multimodal model, announced on 23 July 2026. It spans image, video, audio and action prediction in a single model that was trained jointly across those modalities, rather than assembled from separate systems.
Can I generate FLUX images on YouArt right now?
Yes — with FLUX 2 and Flux Kontext, both available on YouArt today. FLUX 2 handles text-to-image and reference-guided editing in create and edit modes, and Flux Kontext does instruction-based editing from a single reference image with mask-based inpainting.
How is FLUX 3 different from FLUX 2?
FLUX 2 is an image model — text-to-image plus single- and multi-reference editing in one checkpoint. FLUX 3 widens that to video, audio and action prediction in a single jointly trained model, and generates up to 20 seconds of video with audio in one pass.
Does FLUX 3 generate video with sound?
Yes — up to 20 seconds of video with audio in a single generation, including multilingual dialogue, plus continuation from an input video and its audio. Longer pieces chain clips into multi-shot sequences rather than stretching one generation.
How long can a FLUX 3 video be?
Up to 20 seconds in a single generation. Longer sequences chain individual clips into multi-shot scenes rather than extending one generation further.
How much will FLUX 3 cost on YouArt?
Pricing works on a credit system across every model on YouArt. Credit costs for FLUX 3 will be shown on its model page once it's available; in the meantime you can see plans and per-model pricing on the pricing page.
How good is FLUX 3?
In Black Forest Labs' own preference testing — preliminary results from a midtraining checkpoint, on 10-second 720p clips with audio — FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%, with narrower margins against the strongest competitors. Independent testing isn't available yet.
When will FLUX 3 come to YouArt?
FLUX 3 is in early access with Black Forest Labs right now, and it lands on YouArt as soon as that access opens up. This page is where we'll announce it. In the meantime FLUX 2 and Flux Kontext are both here today.

Coming to YouArt

Get ready for FLUX 3 on YouArt

FLUX 3 is coming to YouArt. FLUX 2 and Flux Kontext are already here, so the workflow you build today is the one FLUX 3 slots straight into.

FLUX, FLUX.2 and FLUX 3 are trademarks of Black Forest Labs Inc. YouArt is not affiliated with or endorsed by Black Forest Labs.