OpenCake Articles

FLUX 3 (Flux 3.0) Is Here: Features and Pricing

FLUX 3, often searched as Flux 3.0, is now available with 20-second AI video, native audio, keyframes, Draft mode, and HD or FHD output.

10 min read

FLUX 3 has moved from a promising early-access preview to a production AI video model creators can use. The official product name is FLUX 3, although many people search for it as Flux 3.0. Black Forest Labs now offers FLUX 3 Video through its dashboard and API, bringing up to 20-second generation, optional native audio, keyframe control, video continuation, and a lower-cost Draft workflow into one model.

This is bigger than a routine version update. Black Forest Labs built FLUX 3 as a multimodal foundation model that learns across images, video, audio, and action rather than treating each medium as an isolated task. For creators, the practical result is a model that can move between text-to-video, image-to-video, controlled transitions, and source-video continuation while keeping motion and sound connected.

What is FLUX 3?

FLUX 3 is the newest model family from Black Forest Labs, the company behind FLUX image generation. The broader family covers video, audio, image generation, and action prediction through a shared multimodal architecture. FLUX 3 Video is the production-ready part creators can access now, while FLUX 3 Image, FLUX 3 Action, and the FLUX 3 Dev open-weight backbone are separate parts of the rollout.

Black Forest Labs explains the complete model family and rollout here: Explore the official FLUX 3 model page

What is new in FLUX 3 Video?

Video clips up to 20 seconds

FLUX 3 can create clips up to 20 seconds long in a single generation. That gives a prompt enough room for a developed action, a multi-shot sequence, a short product demonstration, or a complete social hook without immediately stitching several tiny clips together.

Optional native audio

Audio can be generated with the frames and includes multilingual dialogue, synchronized speech, effects, and environmental ambience. Native audiovisual generation is useful because timing is decided in the same pass instead of being reconstructed later with unrelated sound assets.

Text, images, keyframes, and source video

The model supports text-to-video and image-to-video, plus ordered keyframes for controlled transitions. Video-to-video and continuation workflows can carry movement, framing, characters, or scene logic from an existing clip into a new result. This makes FLUX 3 useful for both blank-page creation and iterative production.

Multiple scenes and agentic chaining

A single generation can contain multiple scenes or camera angles. Individual clips can also be chained into longer sequences, with visual references helping preserve identity and continuity. This does not eliminate editing, but it gives creators more complete raw material per generation.

Draft mode before the final render

FLUX 3 Draft creates a faster HD preview at a lower price. When the prompt, motion, and composition work, the draft can be enhanced into a full-quality result using the original request. For teams producing many ad concepts, this preview-first loop can reduce the amount spent on directions that were never going to ship.

FLUX 3 specifications and pricing

FeatureFLUX 3 Video
Maximum durationUp to 20 seconds in one generation
ResolutionHD up to 1 megapixel per frame; FHD up to 2 megapixels per frame
InputsText, starting image, ordered keyframes, or source video depending on the workflow
AudioOptional native speech, dialogue, effects, and ambience
Text/image to video$0.06/sec Draft HD; $0.17/sec standard HD; $0.29/sec standard FHD
Video to video$0.12/sec Draft HD; $0.41/sec standard HD; $0.53/sec standard FHD
AvailabilityBlack Forest Labs dashboard and API; also integrated into OpenCake AI Models

The dollar amounts above are Black Forest Labs' published pay-as-you-go prices at the time of writing. OpenCake converts model usage into credits and shows the current quote before generation, so the quote in the app should be treated as the final price for an OpenCake request.

Best FLUX 3 use cases

  • Product advertisements that need a complete 10–20 second visual arc with sound.
  • Image-to-video animation built from a product hero image, campaign still, or character reference.
  • Short tutorials, explainers, and demonstrations with multiple beats or camera angles.
  • Multilingual creator videos and dialogue-led social concepts.
  • Typography, motion-design, title, and logo-reveal experiments.
  • Video continuation when an existing shot needs a coherent next moment.
  • Draft-first campaign exploration where many directions must be tested before final rendering.

How to prompt FLUX 3

Write the action as a timeline

Describe what happens first, what changes, and how the clip ends. For a product ad, that might mean: the closed package enters frame, the lid opens, the camera pushes toward the product, condensation catches the light, and the final composition holds for the brand moment.

Separate subject motion from camera motion

State what the subject does and then describe how the camera observes it. “The runner turns into the alley while the camera tracks backward at waist height” is easier to execute than a vague request for a dynamic cinematic shot.

Direct the sound explicitly

If audio is enabled, specify dialogue, voice mood, ambience, music, and important synchronized effects. If you want a clean silent visual, say so. Leaving sound undefined gives the model more freedom than many production briefs can tolerate.

Use Draft mode for uncertain directions

A draft is most valuable when the question is structural: Does the hook read? Is the camera move right? Does the story fit the duration? Once those decisions work, move to the full-quality render.

How to use FLUX 3 in OpenCake

  • Open AI Models in the OpenCake dashboard.
  • Choose FLUX 3 from the video models.
  • Select the workflow that matches your inputs: text, images or keyframes, source video, or draft enhancement.
  • Attach the relevant product, actor, image, or video references.
  • Write the visual action, camera behavior, style, and sound as one coherent brief.
  • Choose duration, resolution, aspect ratio, audio, and Draft mode where available.
  • Review the credit quote, generate, and save the result to your Library.

Frequently asked questions about FLUX 3

Is FLUX 3 available now?

Yes. FLUX 3 Video is available through the Black Forest Labs dashboard and API. The wider FLUX 3 family is rolling out in stages, so Image, Action, and open-weight Dev availability should be checked separately.

Can FLUX 3 generate audio and lip sync?

Yes. FLUX 3 supports optional native audio, including multilingual speech, synchronized dialogue, sound effects, and ambience. Quality still depends on the prompt, shot complexity, and how much dialogue must fit inside the duration.

Is FLUX 3 open source?

FLUX 3 Video is currently a hosted product. Black Forest Labs says open-weight access to the FLUX 3 Dev multimodal backbone is coming as a separate part of the rollout.

The bottom line

FLUX 3 is one of the most substantial launches yet from Black Forest Labs. Twenty-second generation, native audio, keyframe control, continuation, multiple scenes, and a practical Draft workflow make it relevant beyond eye-catching demos. The real test is how reliably it follows production constraints, but creators can now run that test with their own products, characters, campaigns, and delivery formats.

Try FLUX 3 alongside other leading image and video models: Open AI Models in OpenCake

Related articles

Related posts