{"id":22683,"date":"2026-07-25T04:17:10","date_gmt":"2026-07-25T04:17:10","guid":{"rendered":"https:\/\/scannn.com\/flux-3-real-world-models-towards-multimodal-flow-models-as-the-backbone-of-visual-intelligence\/"},"modified":"2026-07-25T04:17:10","modified_gmt":"2026-07-25T04:17:10","slug":"flux-3-real-world-models-towards-multimodal-flow-models-as-the-backbone-of-visual-intelligence","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/flux-3-real-world-models-towards-multimodal-flow-models-as-the-backbone-of-visual-intelligence\/","title":{"rendered":"FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence."},"content":{"rendered":"\n<div>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\"><em class=\"text-green font-medium italic\">FLUX 3 is now available in Early Access.<\/em><\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">No single modality provides a complete description. Each is a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process. Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">FLUX 3 is our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments. Early results in content creation and physical AI suggest it is the right path.<\/p>\n<h2 class=\"mx-auto max-w-[720px] text-balance text-black mt-14 mb-5 text-[1.75rem] leading-[1.2] tracking-[-0.03em] font-medium md:text-[2rem] first:mt-0\"><strong>FLUX 3: One model, multiple capabilities.<\/strong><\/h2>\n<div class=\"image-container relative mx-auto my-8 mb-4 flex w-full justify-center rounded-sm max-w-[720px]\"><\/div>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">FLUX 3 builds on <a class=\"text-green font-medium underline underline-offset-3 hover:text-black transition-colors duration-250 ease-[cubic-bezier(0.25,0.1,0.25,1)]\" target=\"_blank\" rel=\"noindex nofollow\" href=\"https:\/\/bfl.ai\/research\/self-flow\">Self-Flow,<\/a> our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, we significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time.<\/p>\n<div class=\"image-container relative mx-auto my-8 mb-4 flex w-full justify-center rounded-sm max-w-[720px]\"><img loading=\"lazy\" alt=\"\" draggable=\"false\" loading=\"lazy\" width=\"2240\" height=\"1152\" decoding=\"async\" data-nimg=\"1\" class=\"z-over-noise relative rounded-sm bg-black object-cover\" style=\"color:transparent\" src=\"https:\/\/bfl.ai\/_next\/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F2gpum2i6%2Fproduction%2F84c6a074bfe2ea965f869b948a51df493ec3637a-2240x1152.png&amp;w=3840&amp;q=75 1x\" bad-src=\"\/_next\/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F2gpum2i6%2Fproduction%2F84c6a074bfe2ea965f869b948a51df493ec3637a-2240x1152.png&amp;w=3840&amp;q=75\"\/><\/div>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\"><strong><em class=\"text-green font-medium italic\">Self-Flow vs. Flow Matching (FM). Left: generation error (Fr\u00e9chet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better).<\/em><\/strong><\/p>\n<h2 class=\"mx-auto max-w-[720px] text-balance text-black mt-14 mb-5 text-[1.75rem] leading-[1.2] tracking-[-0.03em] font-medium md:text-[2rem] first:mt-0\"><strong>Capabilities &amp; Early Evaluations<\/strong><\/h2>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">As a result, FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video. We are highlighting a few of the model\u2019s key capabilities below.<\/p>\n<h3 class=\"mx-auto max-w-[720px] text-balance text-black mt-10 mb-4 text-[1.375rem] leading-[1.3] tracking-[-0.02em] font-medium md:text-[1.5rem] first:mt-0\"><strong>Video<\/strong><\/h3>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Its core capabilities include the following (all outputs come with native audio generation):<\/p>\n<ul class=\"mx-auto max-w-[720px] marker:text-green my-6 flex list-disc flex-col gap-2.5 pl-5 md:pl-6\">\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Text-to-video generation.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Image-to-video generation, either continuing from a starting frame (\u201canimation\u201d) or using images as visual references.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Video-to-video generation from a reference clip, carrying central elements of a source video &#8211; for instance the same character &#8211; into a new scene or context.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Generative video-audio continuation from input video and audio.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Keyframe-to-video generation for controlled transitions between defined moments.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Multilingual dialogue.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Agentic chaining of individual clips into longer, multi-shot sequences.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">High style diversity &#8212; FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics.<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Strong typography generation and animated designs.<\/span><\/li>\n<\/ul>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">For the preliminary analysis below, we generated 10-second text-to-video clips in 720p with audio.<\/p>\n<div class=\"image-container relative mx-auto my-8 mb-4 flex w-full justify-center rounded-sm max-w-[720px]\"><img loading=\"lazy\" alt=\"\" draggable=\"false\" loading=\"lazy\" width=\"2880\" height=\"1800\" decoding=\"async\" data-nimg=\"1\" class=\"z-over-noise relative rounded-sm bg-black object-cover\" style=\"color:transparent\" src=\"https:\/\/bfl.ai\/_next\/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F2gpum2i6%2Fproduction%2F30bb456fef266f4b6115a9f209d4b907108ae09e-2880x1800.png&amp;w=3840&amp;q=75 1x\" bad-src=\"\/_next\/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F2gpum2i6%2Fproduction%2F30bb456fef266f4b6115a9f209d4b907108ae09e-2880x1800.png&amp;w=3840&amp;q=75\"\/><\/div>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\"><strong><em class=\"text-green font-medium italic\">Evaluations are early and we expect further improvements<\/em><\/strong><\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">As the model and the harness around it are still in development, these results are preliminary, and we expect further improvements during the early access phase. Across early evaluations, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">While still in development, FLUX 3 Video is already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities. Furthermore, these capabilities can be combined to create sequences lasting several minutes, where visual references help ensure that the characters remain consistent across all scenes.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">FLUX 3 Video is now available in <a class=\"text-green font-medium underline underline-offset-3 hover:text-black transition-colors duration-250 ease-[cubic-bezier(0.25,0.1,0.25,1)]\" target=\"_blank\" rel=\"noindex nofollow\" href=\"https:\/\/bfl.ai\/models\/flux-3\">Early Access here<\/a><\/p>\n<h3 class=\"mx-auto max-w-[720px] text-balance text-black mt-10 mb-4 text-[1.375rem] leading-[1.3] tracking-[-0.02em] font-medium md:text-[1.5rem] first:mt-0\"><strong>Image<\/strong><\/h3>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. In preliminary evaluations conducted during midtraining, FLUX 3 already shows a significant improvement over earlier versions of FLUX: its ability to handle complex prompts and text generation has improved significantly. The model produces a wide range of output styles (see the following samples), and is able to render high-accuracy text in multiple languages.<\/p>\n<div class=\"image-container relative mx-auto my-8 mb-4 flex w-full justify-center rounded-sm max-w-[720px]\"><img loading=\"lazy\" alt=\"\" draggable=\"false\" loading=\"lazy\" width=\"3400\" height=\"3659\" decoding=\"async\" data-nimg=\"1\" class=\"z-over-noise relative rounded-sm bg-black object-cover\" style=\"color:transparent\" src=\"https:\/\/bfl.ai\/_next\/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F2gpum2i6%2Fproduction%2F67e6dbce8f40db2163330184d35c8dfe4d80bb32-3400x3659.png&amp;w=3840&amp;q=75 1x\" bad-src=\"\/_next\/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F2gpum2i6%2Fproduction%2F67e6dbce8f40db2163330184d35c8dfe4d80bb32-3400x3659.png&amp;w=3840&amp;q=75\"\/><\/div>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">As with video evaluations, these are preliminary results, and we expect further improvements before release. We will open up an early access phase for FLUX 3 Image in the following weeks.<\/p>\n<h3 class=\"mx-auto max-w-[720px] text-balance text-black mt-10 mb-4 text-[1.375rem] leading-[1.3] tracking-[-0.02em] font-medium md:text-[1.5rem] first:mt-0\"><strong>Action<\/strong><\/h3>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">FLUX 3&#8217;s world understanding extends to action prediction. We have taken two routes to it: integrating native action prediction into FLUX 3 directly, scaling up our initial work in Self-Flow; and using the pretrained video backbone as a dynamics-aware foundation that specialized action models can be finetuned from with limited task-specific data.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">For the second, mimic robotics was one of the first partners to gain early access to FLUX 3. Together we developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic&#8217;s expertise in robot learning for dexterous manipulation and production deployment. <a class=\"text-green font-medium underline underline-offset-3 hover:text-black transition-colors duration-250 ease-[cubic-bezier(0.25,0.1,0.25,1)]\" target=\"_blank\" rel=\"noindex nofollow\" href=\"https:\/\/bfl.ai\/blog\/flux-3-mimic\">Read our thesis on why physical AI and content creation run on the same foundation, and how it&#8217;s being tested on real production tasks at Audi.<\/a><\/p>\n<h3 class=\"mx-auto max-w-[720px] text-balance text-black mt-10 mb-4 text-[1.375rem] leading-[1.3] tracking-[-0.02em] font-medium md:text-[1.5rem] first:mt-0\"><strong>Launch Plan<\/strong><\/h3>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Over the next few weeks and months, we will make the following capabilities available, each after an early access phase for ensuring smooth rollout, collecting feedback and rigorous safety-testing. All capabilities are built from the same underlying multimodal flow matching model. These capabilities and models include:<\/p>\n<ul class=\"mx-auto max-w-[720px] marker:text-green my-6 flex list-disc flex-col gap-2.5 pl-5 md:pl-6\">\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Video and audio generation and editing through APIs and private weight access. (\u201cFLUX 3 Video\u201d)<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Action prediction through selected research and commercial partners, beginning with mimic robotics (\u201cFLUX-mimic and FLUX 3 Action\u201d)<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Image synthesis and editing through APIs and private weight access. (\u201cFLUX 3 Image\u201d)<\/span><\/li>\n<li class=\"pl-1.5\"><span class=\"block text-[1.125rem] leading-[1.7] text-pretty text-black [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (\u201cFLUX 3 Dev\u201d)<\/span><\/li>\n<\/ul>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">We will also release more technical details on the underlying approach.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\"><a class=\"text-green font-medium underline underline-offset-3 hover:text-black transition-colors duration-250 ease-[cubic-bezier(0.25,0.1,0.25,1)]\" target=\"_blank\" rel=\"noindex nofollow\" href=\"https:\/\/tally.so\/r\/44d9NX\"><strong>Request early access here<\/strong><\/a><\/p>\n<h2 class=\"mx-auto max-w-[720px] text-balance text-black mt-14 mb-5 text-[1.75rem] leading-[1.2] tracking-[-0.03em] font-medium md:text-[2rem] first:mt-0\"><strong>What\u2019s next?<\/strong><\/h2>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">We are only beginning to scratch the surface of versatile, capable, unified multimodal models, and what they will enable. From interactive image &amp; video editing, simulation to computer use and physical AI, the frontier is wide open. While we gradually roll out these new capabilities, we are already working on the next generation models. Our goal is to unify perceptual, action and language prediction in the same unified model.<\/p>\n<p class=\"mx-auto max-w-[720px] text-[1.125rem] leading-[1.75] text-pretty text-black [&amp;+p]:mt-6 [&amp;&gt;b]:font-medium [&amp;&gt;strong]:font-medium\">If you are interested in exploring and building with FLUX 3, get in touch here. If you are interested in contributing to our mission, join us! We are <a class=\"text-green font-medium underline underline-offset-3 hover:text-black transition-colors duration-250 ease-[cubic-bezier(0.25,0.1,0.25,1)]\" target=\"_blank\" rel=\"noindex nofollow\" href=\"https:\/\/bfl.ai\/careers\">hiring<\/a> in Germany and the US.<\/p>\n<\/div>\n<p><a href=\"https:\/\/bfl.ai\/blog\/flux-3?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>FLUX 3 is now available in Early Access. FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22684,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22683","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22683","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22683"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22683\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22684"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22683"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22683"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22683"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}