Flux 3

FLUX 3 is a new multimodal foundation model that jointly learns from images, videos, audio, and action predictions within a unified architecture. It is now available in Early Access, showing strong capabilities in content creation and physical AI.
FLUX 3 is now available in Early Access.
FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.
No single modality provides a complete description. Each is a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process. Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.
Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.
FLUX 3 is our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments. Early results in content creation and physical AI suggest it is the right path.
FLUX 3: One model, multiple capabilities.
FLUX 3 builds on Self-Flow, our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, we significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time.
Capabilities & Early Evaluations
As a result, FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video. We are highlighting a few of the model’s key capabilities below.
Video
FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation.
Its core capabilities include the following (all outputs come with native audio generation):
- Text-to-video generation.
- Image-to-video generation, either continuing from a starting frame (“animation”) or using images as visual references.
- Video-to-video generation from a reference clip, carrying central elements of a source video - for instance the same character - into a new scene or context.
- Generative video-audio continuation from input video and audio.
- Keyframe-to-video generation for controlled transitions between defined moments.
- Multilingual dialogue.
- A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output.
- Agentic chaining of individual clips into longer, multi-shot sequences.
- High style diversity -- FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics.
- Strong typography generation and animated designs.
Across early evaluations, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons.
Image
FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. In preliminary evaluations, FLUX 3 shows significant improvement over earlier versions in complex prompts and text generation.
Action
FLUX 3's world understanding extends to action prediction. Partnering with mimic robotics, FLUX-mimic combines the FLUX 3 backbone with expertise in robot learning for dexterous manipulation.
Launch Plan
Over the coming months, capabilities will rollout including FLUX 3 Video, FLUX-mimic, FLUX 3 Image, and an open-weight FLUX 3 Dev model.
Source: Hacker News















