公司动向The Decoder5/10

Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

By Matthias Bastian

Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs
图源:The Decoder

AI 摘要

Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2.0, though independent results aren't yet available. The company ul

原文正文

Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

Key Points

- The German AI company Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, videos, and audio simultaneously.

- The model generates videos up to 20 seconds long with native audio and offers features such as text-to-video and the ability to chain individual clips together.

- In addition, the company has developed Flux-mimic, a video action model for robotics applications, which is already being tested at Audi.

German AI company Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model that learns from images, video, and audio together. In BFL's early tests, it beat several rivals in video generation.

BFL describes Flux 3 as a step toward "real-world visual intelligence," which it defines as models that can "perceive, predict, and act across physical and digital environments." The company is part of a broader push to build so-called world models.

No single modality captures reality in full, BFL argues. Images show spatial structure, video captures how it changes over time, and audio can reveal links between mechanical events and the sounds they produce. Training on all three together lets them fill gaps for one another, giving the model more information than training on each modality separately.

Flux 3 adds native audio to videos up to 20 seconds long

Flux 3 can now generate videos with native audio for the first time, with clips up to 20 seconds long. It supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips for longer multi-shot sequences. BFL says the model is especially good at human facial expressions and matching sounds to physical events.

In early evaluations using 10-second clips at 720p, BFL reports that Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent. The margins narrow against stronger competitors. Flux 3 was preferred over Kling v3 Pro 60 percent of the time, over Happy Horse v1 at 59 percent, over Happy Horse 1.1 at 57 percent, and over both Seedance 2.0 and Gemini Omni Flash at 52 percent each.

BFL says the results are preliminary, and no independent tests are available yet. Matching leading systems such as Seedance, which has already reached Hollywood, and Gemini Omni Flash would put Flux 3 among the top video models.

BFL also expects Flux 3 to improve image generation, especially for complex prompts and accurate text rendering in multiple languages. The company plans to release Flux 3 Image in early access within the next few weeks.

BFL says the model can also predict actions based on its understanding of the world. The company worked with Mimic Robotics to develop Flux-mimic, a video-action model now being tested on production tasks at Audi.

Flux 3 is based on Self-Flow, BFL's approach for teaching one model to generate and understand content at the same time. A multimodal transformer uses dedicated components to convert images, video, and audio into a shared internal representation and then turn it back into outputs.

A component for actions provides the foundation for robotics applications. BFL says this unified learning process delivers better results than the previously standard flow-matching method, both in generation quality and in the model's grasp of the physical world.

BFL plans a phased rollout with open-weight access

BFL is rolling out all capabilities in stages, with early-access phases for feedback and safety testing. Flux 3 Video is already available, and Flux 3 Image is set to follow in the coming weeks. Action prediction will initially be offered through select partners.

BFL also plans to release open-weight access to the multimodal backbone under the name "Flux 3 Dev." Longer term, the company is working on next-generation models that aim to combine perception, action, and language prediction in a single model.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

阅读原文
Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs · AI Daily