FLUX 3
One Model for Image, Video and Sound
FLUX 3 is Black Forest Labs' multimodal foundation model, trained jointly on images, video, and audio in a single unified architecture. It is the company's first video model, generating up to 20 seconds with native audio in one pass. Available on Pexo.
What FLUX 3 Produces, Seen Through Pexo
Every video shown here was created through Pexo from a plain-language description — no prompt syntax and no model configuration.
What Makes FLUX 3 Different
Black Forest Labs released FLUX 3 on July 23, 2026. It is a departure from the image-only FLUX line, and the design choices behind it are unusual.

One Architecture, Not Three Models Behind One Badge
FLUX 3 is built on Self-Flow, which Black Forest Labs describes as its approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Images, video, audio, and language are trained together rather than bolted into separate models, on the premise that each modality is a projection of the same underlying reality.

Up to 20 Seconds, Audio Included by Default
This is Black Forest Labs' first video model, and it generates video with audio up to 20 seconds long in a single generation. Native audio generation is not an option you enable, it comes with every output, covering dialogue, effects, and ambience alongside the picture.

Text, Image, Video, or Keyframes as the Starting Point
FLUX 3 accepts text-to-video, image-to-video for animation or visual reference, video-to-video, video and audio continuation, and keyframe-to-video, plus multilingual dialogue. Individual generations can be chained agentically into multi-shot sequences rather than being limited to one shot at a time.
FLUX 3 vs the FLUX Models Before It
Every earlier FLUX release was an image model. FLUX 3 is the point where Black Forest Labs turned the line into a multimodal foundation model, so the comparison is less about quality scores than about what the model is for.
| Feature | FLUX 3 | FLUX.2 | FLUX.1 | FLUX.1 Dev |
|---|---|---|---|---|
| Positioning | Multimodal foundation model | Image generation | Image generation | Image generation |
| Video generation | ✓ | — | — | — |
| Native audio with the picture | ✓ | — | — | — |
| Action prediction for robotics | ✓ | — | — | — |
| Trained jointly across modalities | ✓ | — | — | — |
| Architecture | Self-Flow unified | Image flow model | Image flow model | Image flow model |
Sources: Black Forest Labs: FLUX 3 · Black Forest Labs · BFL Documentation
How to Use FLUX 3 in Pexo: Three Steps, No Setup
No account, API key, or technical knowledge required. Pexo handles model selection, generation, and delivery so you focus on what you want to make.
Type a description in plain language on web, Telegram, WhatsApp, or Discord. There's no required format and no model to select — Pexo reads your intent and handles the rest.
Pexo routes your request to FLUX 3, applies the right settings automatically, and runs the generation. You configure nothing.
Your video arrives ready to use. Want to refine it? Keep the conversation going instead of starting over.
What You Can Make with Pexo's FLUX 3
No prompt writing. No model picking. Just describe what you want.
Frequently Asked Questions About FLUX 3
What is FLUX 3?+
FLUX 3 is a multimodal foundation model that Black Forest Labs announced on July 23, 2026. It jointly learns from images, video, and audio inside one unified architecture, and it is the company's first model that generates video. Black Forest Labs frames the four modalities it covers, image, video, audio, and language, as projections of the same underlying reality.
How long can a FLUX 3 video be, and does it have sound?+
Black Forest Labs states that FLUX 3 generates video with audio up to 20 seconds in length in a single generation, and that all outputs come with native audio generation. The preliminary clips it published for analysis are 10 seconds at 720p with audio.
What is Self-Flow?+
Self-Flow is Black Forest Labs' method for efficiently aligning multimodal generation and understanding within the same underlying architecture. It is the foundation FLUX 3 is built on, and it is what allows one model to be trained across images, video, and audio at once instead of stitching separate specialist models together.
How is FLUX 3 different from FLUX.2 and FLUX.1?+
The earlier FLUX releases were image models. FLUX 3 is a multimodal foundation model covering image, video, audio, and action prediction in one architecture. Black Forest Labs notes that the FLUX 3 Image tier already shows a significant improvement over earlier versions of FLUX, specifically in handling complex prompts and generating multilingual text.
What are the FLUX 3 tiers?+
Four. FLUX 3 Video covers video and audio generation and editing. FLUX 3 Image covers image synthesis and editing. FLUX 3 Action, also called FLUX-mimic, handles action prediction and is going out through selected partners starting with mimic robotics. FLUX 3 Dev is the planned open-weight multimodal backbone.
Can I use FLUX 3 right now?+
Access is staged. Black Forest Labs opened FLUX 3 in Early Access, with FLUX 3 Video available first and FLUX 3 Image following in the weeks after. Action prediction is going through selected partners, and the open-weight Dev tier is a planned rollout. Black Forest Labs describes its published evaluation figures and specs as preliminary.
FLUX 3 Is on Pexo: Describe It and See It
The best video model available right now takes one sentence to use.





