July 25, 2026 · VentureBeat
Black Forest Labs Launches FLUX 3, a Unified Multimodal Model for Images, Video, and Audio
My take: Visual content creation just took another step forward. FLUX 3 is Black Forest Labs' first model to generate images, video clips up to 20 seconds with synchronized audio, and action prediction for robotics, all from the same unified architecture. Until now, generating video with realistic sound required combining separate tools and manually syncing everything; with this model it becomes a single step.
Black Forest Labs' own benchmarks show FLUX 3 preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen 4.5 in 77%. That said, these tests were run by the company itself, so it is worth waiting for independent evaluations before taking those numbers as definitive. What is verifiable right now is the technical capability: native audio in video, multilingual dialogue, and clip continuation for longer sequences.
For now it is available only in limited early access for video, with images and the open-weight model coming later. If you produce video content for your brand, business, or creative projects, this is the category of tools that will change your production workflow sooner than you might expect. Do you have a sense of how much production time you could reclaim if video with audio were as easy as writing a text prompt?
Want to use these tools? See the unbiased reviews or back to the news.