{"id":22853,"date":"2026-07-28T08:38:20","date_gmt":"2026-07-28T08:38:20","guid":{"rendered":"https:\/\/scannn.com\/kimi-k3-architecture-notes-sebastian-raschka-phd\/"},"modified":"2026-07-28T08:38:20","modified_gmt":"2026-07-28T08:38:20","slug":"kimi-k3-architecture-notes-sebastian-raschka-phd","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/kimi-k3-architecture-notes-sebastian-raschka-phd\/","title":{"rendered":"Kimi K3 Architecture Notes | Sebastian Raschka, PhD"},"content":{"rendered":"\n<div id=\"\">\n<p>The Kimi K3 architecture figure for yesterday\u2019s big open-weight model release, along with some observations and thoughts.<\/p>\n<ol>\n<li>\n<p>Yes, it looks relatively complicated, but it\u2019s essentially a scaled-up production version of their <a href=\"http:\/\/sebastianraschka.com\/llm-architecture-gallery\/#card-kimi-linear-48b-a3b\">Kimi Linear model<\/a> they released last year (scaled up from 48B -&gt; 2.8T; K3 is by far the biggest open-weight model right now)<\/p>\n<\/li>\n<li>\n<p>The one new component compared to Kimi Linear is the <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/latent-moe\/\">LatentMoE<\/a>. I omitted it in the figure below since it\u2019s already very crowded, but that\u2019s essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/\">LLM Architecture Gallery<\/a> if you are curious). The idea here is to compress (down-project) large linear layers similar to <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/mla\/\">multi-head latent attention<\/a>.<\/p>\n<\/li>\n<li>\n<p>Kimi K3\u2019s overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/moe\/\">MoE<\/a> -&gt; LatentMoE, regular attention -&gt; multi-head latent attention and <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/hybrid-attention\/\">Kimi Delta Attention<\/a>. (I also have short tutorials and write-ups in my <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/\">gallery<\/a> if you are curious about additional details).<\/p>\n<\/li>\n<li>\n<p>The one component change that is not an efficiency tweak is <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/attention-residuals\/\">attention residuals<\/a>. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important\/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost.<\/p>\n<\/li>\n<li>\n<p>Interestingly, Kimi K3 got rid of all RoPE layers and uses <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/nope\/\">NoPE<\/a> (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like <a href=\"https:\/\/sebastianraschka.com\/llm-architecture-gallery\/swa\/\">sliding window attention<\/a>) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know.<\/p>\n<\/li>\n<li>\n<p>Kimi K3 now also has native multimodal support, which is great!<\/p>\n<\/li>\n<\/ol>\n<p>There are several other interesting training tidbits in the technical report, but that\u2019s it from the architecture front so far. A really great release overall.<\/p>\n<figure>\n<p><\/p><figcaption class=\"figure-caption\">Figure 1. Kimi K3 architecture and release-time benchmark comparisons. See <a href=\"http:\/\/sebastianraschka.com\/llm-architecture-gallery\/#card-kimi-k3\">K3<\/a> in the architecture gallery for more details.<\/figcaption><\/figure>\n<p>Source: website version of my <a href=\"https:\/\/substack.com\/@rasbt\/note\/c-303378576\">Substack note<\/a>.<\/p>\n<\/p><\/div>\n<p><a href=\"https:\/\/sebastianraschka.com\/blog\/2026\/kimi-k3-architecture-notes.html?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The Kimi K3 architecture figure for yesterday\u2019s big open-weight model release, along with some observations and thoughts. Yes, it looks relatively complicated, but it\u2019s essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -&gt; 2.8T; K3 is by far the biggest open-weight model right now) The [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22854,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22853","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22853","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22853"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22853\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22854"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22853"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22853"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22853"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}