{"id":22580,"date":"2010-03-29T16:20:58","date_gmt":"2010-03-29T16:20:58","guid":{"rendered":"https:\/\/scannn.com\/github-microsoft-mage-%c2%b7-github\/"},"modified":"2010-03-29T16:20:58","modified_gmt":"2010-03-29T16:20:58","slug":"github-microsoft-mage-%c2%b7-github","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/github-microsoft-mage-%c2%b7-github\/","title":{"rendered":"GitHub - microsoft\/Mage \u00b7 GitHub"},"content":{"rendered":"\n<div id=\"\">\n<div align=\"center\" dir=\"auto\">\n<a target=\"_blank\" rel=\"noopener noreferrer\" href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/assets\/mage-model-cover.png\"><\/a>\n<\/div>\n<hr\/>\n<p dir=\"auto\"><strong>Mage<\/strong> is a family of lightweight, research-friendly multimodal models built at a fixed <strong>4B-parameter<\/strong> budget. It is designed to make advanced visual <strong>understanding<\/strong> and <strong>generation<\/strong> accessible for controlled experiments, post-training research, and vertical-domain applications under realistic compute budgets.<\/p>\n<p dir=\"auto\">The family is organized around a shared <strong>codec-aligned efficiency<\/strong> philosophy \u2014 <em>spend representation capacity where the signal is<\/em> \u2014 applied to both the understanding and the generation side:<\/p>\n<p><markdown-accessiblity-table><\/p>\n<table>\n<thead>\n<tr>\n<th align=\"left\">Model<\/th>\n<th align=\"left\">Task<\/th>\n<th align=\"center\">Scale<\/th>\n<th align=\"left\">Code<\/th>\n<th align=\"left\">Report<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td align=\"left\"><strong><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/mage_vl\">Mage-VL<\/a><\/strong><\/td>\n<td align=\"left\">Image &amp; video understanding, proactive streaming<\/td>\n<td align=\"center\">4B<\/td>\n<td align=\"left\"><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/mage_vl\/README.md\"><code>mage_vl\/<\/code><\/a><\/td>\n<td align=\"left\"><em>Coming Soon<\/em><\/td>\n<\/tr>\n<tr>\n<td align=\"left\"><strong><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/mage_flow\">Mage-Flow<\/a><\/strong><\/td>\n<td align=\"left\">Text-to-image generation &amp; instruction-based editing<\/td>\n<td align=\"center\">4B<\/td>\n<td align=\"left\"><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/mage_flow\/README.md\"><code>mage_flow\/<\/code><\/a><\/td>\n<td align=\"left\"><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/assets\/mage_flow_tech_report.pdf\">PDF<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><\/markdown-accessiblity-table><\/p>\n<p dir=\"auto\">Both models are compact enough to train, fine-tune, and deploy on modest hardware, yet remain competitive with much larger open systems in their respective domains.<\/p>\n<hr\/>\n<div class=\"markdown-heading\" dir=\"auto\">\n<h2 tabindex=\"-1\" class=\"heading-element\" dir=\"auto\"> Mage-VL \u2014 codec-native streaming vision\u2013language<\/h2>\n<p><a id=\"user-content--mage-vl--codec-native-streaming-visionlanguage\" class=\"anchor\" aria-label=\"Permalink: &#x1f9e9; Mage-VL \u2014 codec-native streaming vision\u2013language\" href=\"#-mage-vl--codec-native-streaming-visionlanguage\"><svg data-component=\"Octicon\" class=\"octicon octicon-link\" viewbox=\"0 0 16 16\" version=\"1.1\" width=\"16\" height=\"16\" aria-hidden=\"true\"><path d=\"m7.775 3.275 1.25-1.25a3.5 3.5 0 1 1 4.95 4.95l-2.5 2.5a3.5 3.5 0 0 1-4.95 0 .751.751 0 0 1 .018-1.042.751.751 0 0 1 1.042-.018 1.998 1.998 0 0 0 2.83 0l2.5-2.5a2.002 2.002 0 0 0-2.83-2.83l-1.25 1.25a.751.751 0 0 1-1.042-.018.751.751 0 0 1-.018-1.042Zm-4.69 9.64a1.998 1.998 0 0 0 2.83 0l1.25-1.25a.751.751 0 0 1 1.042.018.751.751 0 0 1 .018 1.042l-1.25 1.25a3.5 3.5 0 1 1-4.95-4.95l2.5-2.5a3.5 3.5 0 0 1 4.95 0 .751.751 0 0 1-.018 1.042.751.751 0 0 1-1.042.018 1.998 1.998 0 0 0-2.83 0l-2.5 2.5a1.998 1.998 0 0 0 0 2.83Z\"\/><\/svg><\/a><\/div>\n<p dir=\"auto\"><strong>Mage-VL<\/strong> is a codec-native, proactive-streaming multimodal foundation model for image &amp; video understanding, trained <strong>entirely from scratch<\/strong> at a compact <strong>4B<\/strong> scale.<\/p>\n<p dir=\"auto\"><strong> Coming soon<\/strong> \u2014 code, checkpoints, and full details are on the way. Stay tuned.<\/p>\n<p dir=\"auto\">\u2192 <strong><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/mage_vl\/README.md\"><code>mage_vl\/README.md<\/code><\/a><\/strong><\/p>\n<div class=\"markdown-heading\" dir=\"auto\">\n<h2 tabindex=\"-1\" class=\"heading-element\" dir=\"auto\"> Mage-Flow \u2014 efficient native-resolution generation &amp; editing<\/h2>\n<p><a id=\"user-content--mage-flow--efficient-native-resolution-generation--editing\" class=\"anchor\" aria-label=\"Permalink: &#x1f3a8; Mage-Flow \u2014 efficient native-resolution generation &amp; editing\" href=\"#-mage-flow--efficient-native-resolution-generation--editing\"><svg data-component=\"Octicon\" class=\"octicon octicon-link\" viewbox=\"0 0 16 16\" version=\"1.1\" width=\"16\" height=\"16\" aria-hidden=\"true\"><path d=\"m7.775 3.275 1.25-1.25a3.5 3.5 0 1 1 4.95 4.95l-2.5 2.5a3.5 3.5 0 0 1-4.95 0 .751.751 0 0 1 .018-1.042.751.751 0 0 1 1.042-.018 1.998 1.998 0 0 0 2.83 0l2.5-2.5a2.002 2.002 0 0 0-2.83-2.83l-1.25 1.25a.751.751 0 0 1-1.042-.018.751.751 0 0 1-.018-1.042Zm-4.69 9.64a1.998 1.998 0 0 0 2.83 0l1.25-1.25a.751.751 0 0 1 1.042.018.751.751 0 0 1 .018 1.042l-1.25 1.25a3.5 3.5 0 1 1-4.95-4.95l2.5-2.5a3.5 3.5 0 0 1 4.95 0 .751.751 0 0 1-.018 1.042.751.751 0 0 1-1.042.018 1.998 1.998 0 0 0-2.83 0l-2.5 2.5a1.998 1.998 0 0 0 0 2.83Z\"\/><\/svg><\/a><\/div>\n<p dir=\"auto\"><strong>Mage-Flow<\/strong> is a compact 4B generative stack for <strong>text-to-image generation<\/strong> and <strong>instruction-based image editing<\/strong>, built from two co-designed components: <strong>Mage-VAE<\/strong> (a lightweight, high-fidelity latent tokenizer) and a <strong>Native-Resolution Multimodal Diffusion Transformer<\/strong> trained with rectified flow matching. Each task ships in <strong>Base<\/strong>, <strong>RL-aligned<\/strong>, and <strong>4-step Turbo<\/strong> variants.<\/p>\n<p dir=\"auto\"><strong>Highlights<\/strong><\/p>\n<ul dir=\"auto\">\n<li><strong>Compact &amp; competitive.<\/strong> A single 4B family for generation <em>and<\/em> editing that matches or beats much larger open systems (Qwen-Image 20B, Z-Image 6B, FLUX.2 32B, FireRed-Image-Edit 20B).<\/li>\n<li><strong>Efficient tokenizer.<\/strong> Mage-VAE matches FLUX.2-VAE reconstruction fidelity using <strong>~12\u00d7 \/ ~22\u00d7 fewer encode \/ decode MACs per pixel<\/strong>, removing the VAE high-resolution bottleneck.<\/li>\n<li><strong>Native resolution.<\/strong> One checkpoint generates from <strong>512 to 2048<\/strong> on any aspect ratio, including extreme <strong>4:1<\/strong> (e.g. <code>512\u00d72048<\/code>, <code>2048\u00d7512<\/code>).<\/li>\n<li><strong>System-level speed.<\/strong> Native-resolution packing + fused CUDA kernels cut per-step training time from <strong>~1.93 s \u2192 ~0.78 s<\/strong> (<strong>~2.5\u00d7 faster training<\/strong>); at <code>1024\u00b2<\/code> on a single A100, <strong>Mage-Flow-Turbo 0.59 s\/image<\/strong> and <strong>Mage-Flow-Edit-Turbo 1.02 s\/edit<\/strong>.<\/li>\n<li><strong>Versatile editing.<\/strong> Mage-Flow-Edit supports semantic content editing, appearance transformation, image restoration, and structure-aware outputs within a unified image-and-text-conditioned model.<\/li>\n<\/ul>\n<p dir=\"auto\">\u2192 Details, installation, Python API, CLI, and Gradio app: <strong><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/mage_flow\/README.md\"><code>mage_flow\/README.md<\/code><\/a><\/strong><\/p>\n<ul dir=\"auto\">\n<li><strong>2026-07-22<\/strong> \u2014 <strong>Mage-Flow<\/strong> checkpoints released on  <a href=\"https:\/\/huggingface.co\/collections\/microsoft\/mage\" rel=\"nofollow\">Hugging Face<\/a>: Base, RL-aligned, and 4-step Turbo variants for both text-to-image generation and image editing.<\/li>\n<li><strong>Coming soon<\/strong> \u2014 <strong>Mage-VL<\/strong>: a <strong>Base<\/strong> vision\u2013language model plus a <strong>proactive-streaming<\/strong> variant for codec-native image &amp; video understanding. Stay tuned.<\/li>\n<\/ul>\n<p dir=\"auto\"><strong>Mage-VL<\/strong> \u2014 vision\u2013language (image &amp; video understanding).<\/p>\n<p><markdown-accessiblity-table><\/p>\n<table>\n<thead>\n<tr>\n<th align=\"left\">Model<\/th>\n<th align=\"left\">Task<\/th>\n<th align=\"center\">Scale<\/th>\n<th align=\"left\">Hugging Face<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td align=\"left\"><code>Mage-VL<\/code><\/td>\n<td align=\"left\">image &amp; video understanding, proactive streaming<\/td>\n<td align=\"center\">4B<\/td>\n<td align=\"left\"> Coming soon<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><\/markdown-accessiblity-table><\/p>\n<p dir=\"auto\"><strong>Mage-Flow<\/strong> \u2014 generation &amp; editing. Each checkpoint is a self-contained diffusers-style repo (<code>transformer\/<\/code> + shared <code>vae\/<\/code>, <code>text_encoder\/<\/code>, <code>scheduler\/<\/code>).<\/p>\n<p><markdown-accessiblity-table\/><\/p>\n<p dir=\"auto\">Each model is self-contained in its own directory with a dedicated README:<\/p>\n<ul dir=\"auto\">\n<li><strong>Mage-VL<\/strong> \u2014 image\/video understanding demo, codec backends, CLI \u2192 <strong><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/mage_vl\/README.md\"><code>mage_vl\/README.md<\/code><\/a><\/strong><\/li>\n<li><strong>Mage-Flow<\/strong> \u2014 installation, Python API, CLI, prompt enhancement, Gradio app \u2192 <strong><a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/mage_flow\/README.md\"><code>mage_flow\/README.md<\/code><\/a><\/strong><\/li>\n<\/ul>\n<div class=\"highlight highlight-text-bibtex notranslate position-relative overflow-auto\" dir=\"auto\" data-snippet-clipboard-copy-content=\"@article{zhang2026mageflow,&#10;  title={Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing},&#10;  author={Zhang, Xinjie and Zhang, Peng and Zheng, Shicheng and Guo, Jinghao and Jia, Zhaoyang and Shen, Yifei and Guo, Xun and Luo, Yuxuan and Li, Jiahao and Xie, Wenxuan and Pu, Fanyi and Zhang, Xiaoyi and Zhang, Kaichen and Guo, Zongyu and Bi, Tianci and Gui, Dongnan and Liu, Zhening and Wen, Zimo and Zheng, Zihan and Yang, Senqiao and Li, Xiao and Wang, Jinglu and Li, Bin and Lu, Yan},&#10;  journal={arXiv preprint arXiv:2607.19064},&#10;  year={2026}&#10;}\">\n<pre><span class=\"pl-k\">@article<\/span>{<span class=\"pl-en\">zhang2026mageflow<\/span>,\n  <span class=\"pl-s\">title<\/span>=<span class=\"pl-s\"><span class=\"pl-pds\">{<\/span>Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing<span class=\"pl-pds\">}<\/span><\/span>,\n  <span class=\"pl-s\">author<\/span>=<span class=\"pl-s\"><span class=\"pl-pds\">{<\/span>Zhang, Xinjie and Zhang, Peng and Zheng, Shicheng and Guo, Jinghao and Jia, Zhaoyang and Shen, Yifei and Guo, Xun and Luo, Yuxuan and Li, Jiahao and Xie, Wenxuan and Pu, Fanyi and Zhang, Xiaoyi and Zhang, Kaichen and Guo, Zongyu and Bi, Tianci and Gui, Dongnan and Liu, Zhening and Wen, Zimo and Zheng, Zihan and Yang, Senqiao and Li, Xiao and Wang, Jinglu and Li, Bin and Lu, Yan<span class=\"pl-pds\">}<\/span><\/span>,\n  <span class=\"pl-s\">journal<\/span>=<span class=\"pl-s\"><span class=\"pl-pds\">{<\/span>arXiv preprint arXiv:2607.19064<span class=\"pl-pds\">}<\/span><\/span>,\n  <span class=\"pl-s\">year<\/span>=<span class=\"pl-s\"><span class=\"pl-pds\">{<\/span>2026<span class=\"pl-pds\">}<\/span><\/span>\n}<\/pre>\n<\/div>\n<p dir=\"auto\">These models are released for research purposes only and are not intended for product or service deployment. Responsible AI considerations were incorporated throughout the development process, including data selection, model training, and evaluation. The training data includes a combination of public, licensed, and internal datasets that were processed to remove clearly identifiable personal information and reduce harmful content where possible. However, as the data is largely sourced from web-scale collections, it may contain biases or uneven representation. As a result, the models may generate outputs that are inaccurate, biased, or inappropriate under certain prompts. The models should be used in controlled research settings with appropriate human oversight, and downstream users are responsible for applying additional safeguards \u2014 such as content moderation, validation, and compliance checks \u2014 before broader use.<\/p>\n<p dir=\"auto\">This project is released under the <a href=\"https:\/\/github.com\/microsoft\/Mage\/blob\/main\/LICENSE\">MIT License<\/a>.<\/p>\n<\/div>\n<p><a href=\"https:\/\/github.com\/microsoft\/Mage?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Mage is a family of lightweight, research-friendly multimodal models built at a fixed 4B-parameter budget. It is designed to make advanced visual understanding and generation accessible for controlled experiments, post-training research, and vertical-domain applications under realistic compute budgets. The family is organized around a shared codec-aligned efficiency philosophy \u2014 spend representation capacity where the signal [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22581,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22580","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22580","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22580"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22580\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22581"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22580"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22580"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22580"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}