{"id":22565,"date":"2026-07-20T15:28:14","date_gmt":"2026-07-20T15:28:14","guid":{"rendered":"https:\/\/scannn.com\/on-kimi-k3-its-capabilities-and-related-discontents\/"},"modified":"2026-07-20T15:28:14","modified_gmt":"2026-07-20T15:28:14","slug":"on-kimi-k3-its-capabilities-and-related-discontents","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/on-kimi-k3-its-capabilities-and-related-discontents\/","title":{"rendered":"On Kimi K3: Its Capabilities And Related Discontents"},"content":{"rendered":"\n<div>\n<p>Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model.<\/p>\n<p>Do not get carried away. Do not judge Kimi K3 only its relative strengths. In aggregate it is several months behind the closed model frontier, at least four and my median guess is six, with the post-training closer and the pre-training farther out. This is less months than before, but the months are denser now.<\/p>\n<p>It is somewhat distilled. It likely outperforms on benchmarks relative to practical performance. All its benchmarks are scored at maximum effort, typically a lot more tokens than are used in similar tests by Fable or Sol. Performance looks jagged. Kimi will be excellent at some things, less so at other things.<\/p>\n<p>We will know more over the coming weeks. For now access is spotty and not that many people have actually had the chance to try Kimi K3, so I have larger error bars than usual around its capabilities. Alas, time waits for no one, so we press on.<\/p>\n<p>It is the largest open model so far at 2.8T, on the upper end of possible sizes for Claude Opus and near the bottom of possible sizes for Mythos, which explains many of its gains. It is slow and appears hungry for tokens. A lot of the positive reactions are to this being a big model, which thus has at least a decent amount of \u2018big model smell\u2019 and generally trades being slower and more expensive for some performance gains. That\u2019s a great move, but it should be a while before they can do it again.<\/p>\n<p>Distillation from Claude is clearly part of the story, likely largely from Fable, and is clearly nothing close to the whole story. Clearly Moonshot would do this even if it helped only a little. The timing of them releasing a bigger model is suggestive.<\/p>\n<p>It is again a good model, but once you correct for the overperformance on benchmarks and look at expected practical performance, it is not clear this is so different from what you would have expected from a Kimi K3 that had 2.8T parameters.<\/p>\n<p>Consider that <a href=\"https:\/\/x.com\/peterwildeford\/status\/2079052877801128237\">Kimi K3\u2019s (preliminary unofficial) Epoch Capabilities Index is exactly on the Chinese trend line<\/a>.<\/p>\n<p>Kimi K3 is absolutely worth checking to see if it fits into your workflows. At this price point, for both the API and the subscription, it is not going to fill the role of the smaller cheaper open models, and I expect it to usually lose out in a fight with the top closed models, but there are going to be some places where Kimi K3 is a good choice.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/AndrewCurran_\/status\/2078729593096434114\">Andrew Curran<\/a>: Following the success of Kimi K3, Moonshot AI has informed investors that it plans an IPO in Hong Kong within the next six months, according to Bloomberg.<\/p>\n<\/blockquote>\n<p>Good idea. Strike while the iron is hot.<\/p>\n<h4 class=\"wp-block-heading\">Table of Contents<\/h4>\n<ol>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/deepseek-moments-here-we-go-again\">DeepSeek Moments: Here We Go Again.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/we-had-a-moment-reprise-from-june-2025\">We Had a Moment (Reprise from June 2025).<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/the-story-since-then\">The Story Since Then.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/the-kimi-k3-announcement-pitch-and-basic-facts\">The Kimi K3 Announcement, Pitch and Basic Facts.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/on-modern-benchmaxxing\">On Modern Benchmaxxing.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/other-people-s-benchmarks\">Other People\u2019s Benchmarks.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/benchmarks-are-not-the-real-world\">Benchmarks Are Not The Real World.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/technical-safeguards-what-are-those\">Technical Safeguards? What Are Those?<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/things-kimi-can-do\">Things Kimi Can Do.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/things-kimi-cannot-do\">Things Kimi Cannot Do.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/things-it-is-not-easy-to-get-kimi-to-do\">Things It Is Not Easy To Get Kimi To Do.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/open-weight-models-are-unsafe-and-nothing-can-fix-this\">Open Weight Models Are Unsafe And Nothing Can Fix This.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/dean-ball-attempts-to-be-constructive\">Dean Ball Attempts To Be Constructive.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/openai-employees-are-relatively-bullish-on-this-one\">OpenAI Employees Are Relatively Bullish On This One.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/kimi-k3-is-relatively-strongest-at-typical-agentic-coding-and-3d\">Kimi K3 Is Relatively Strongest At Typical Agentic Coding and 3D.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/reactions\">Reactions.<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/who-are-you\">Who Are You?<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/how-did-they-do-it\">How Did They Do It?<\/a><\/li>\n<li><a href=\"https:\/\/thezvi.substack.com\/i\/207449107\/conclusion\">Conclusion.<\/a><\/li>\n<\/ol>\n<h4 class=\"wp-block-heading\">DeepSeek Moments: Here We Go Again<\/h4>\n<p>All discourse about Chinese models lives in the shadow of the <a href=\"https:\/\/thezvi.substack.com\/p\/deepseek-r1-0528-did-not-have-a-moment\">DeepSeek moment<\/a>.<\/p>\n<p>There are a lot of people who really, really want another <a href=\"https:\/\/thezvi.substack.com\/p\/deepseek-r1-0528-did-not-have-a-moment\">DeepSeek moment<\/a> to happen.<\/p>\n<p>These people really, really want to tell the story that Chinese open models are catching up to American closed models, that AI and inference will become commoditized.<\/p>\n<p>Their motivations vary. They often want to affirm open models, or the importance of the \u2018tech stack.\u2019 Others simply want to see OpenAI and Anthropic go down, or know that such claims sell. Often the ultimate objective is to argue against all AI regulations, or anything that might \u2018slow us down\u2019 or cause us to \u2018lose to China.\u2019<\/p>\n<p><a href=\"https:\/\/x.com\/peterwildeford\/status\/2077946879547990242\">It is actively suicidal to respond to<\/a> \u2018the Chinese have better models now\u2019 with \u2018then we had better sell them the compute so they can run them and also build even better ones.\u2019 Yet every time, yes, people will argue that. Sigh.<\/p>\n<p>Often they simply want to tell American AI to stop taking precautions, to <a href=\"https:\/\/x.com\/ramez\/status\/2077848618497921422\">stop being annoying and take down the classifiers<\/a>, as in \u2018genie is out of the bottle, so release the bigger genie with unlimited wishes, it\u2019s the only way.\u2019 People really would take major catastrophic risks rather than deal with classifiers, and are Big Mad about this.<\/p>\n<p><a href=\"https:\/\/x.com\/AndrewCurran_\/status\/2077847329655452104\">Google was down 4.4% on the day, SpaceX was down 3.1%<\/a> and Nvidia down over 2%, and tech stocks were down again on Friday, so plausibly we\u2019re doing this again.<\/p>\n<p><a href=\"https:\/\/www.axios.com\/2026\/07\/17\/china-ai-kimi-k3-open-source-anthropic-opus?utm_campaign=mrf-utm_campaign=editorial&amp;utm_source=x&amp;utm_medium=owned_social&amp;utm_source=twitter&amp;utm_medium=social&amp;mrfcid=202607176a4f1cc2c72f0423c6346038\">And yep, we are at risk of doing this again<\/a>:<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/axios\/status\/2078055882840047769\">Axios<\/a> (being wrong): China just erased America\u2019s AI lead<\/p>\n<p><a href=\"https:\/\/x.com\/S_OhEigeartaigh\/status\/2078077922473054582\">Se\u00e1n \u00d3 h\u00c9igeartaigh<\/a>: No it didn\u2019t. (although Kimi K3 is undoubtedly impressive)<\/p>\n<\/blockquote>\n<p>The pattern is:<\/p>\n<ol>\n<li>Chinese model releases.<\/li>\n<li>There is some impressive benchmark cited.<\/li>\n<li>Therefore, America\u2019s lead is gone, QED, that\u2019s it, no, really, that\u2019s it.<\/li>\n<\/ol>\n<p>In the Axios case the benchmark in question is Arena. Based on that alone, they state as fact that America\u2019s lead is gone.<\/p>\n<p>I would ignore, but this style of logic has convinced a lot of Washington D.C. multiple times, and that has had substantial policy impact.<\/p>\n<p>I shudder to think what such folks might do now that they also know about Mythos. The confusion over Fable jailbreaks could easily extend to a broader dumb panic.<\/p>\n<p>The original <a href=\"https:\/\/thezvi.substack.com\/p\/deepseek-r1-0528-did-not-have-a-moment\">DeepSeek moment<\/a> happened because of a confluence of events.<\/p>\n<p>Let\u2019s review.<\/p>\n<h4 class=\"wp-block-heading\">We Had a Moment (Reprise from June 2025)<\/h4>\n<p>We all remember\u00a0<a href=\"https:\/\/thezvi.substack.com\/p\/on-deepseeks-r1\"><strong>The DeepSeek Moment<\/strong><\/a>, which led to\u00a0<a href=\"https:\/\/thezvi.substack.com\/p\/deepseek-panic-at-the-app-store\"><strong>Panic at the App Store<\/strong><\/a>, lots of stock market turmoil that made remarkably little fundamental sense and that has been borne out as rather silly,\u00a0<a href=\"https:\/\/thezvi.substack.com\/p\/deepseek-lemon-its-wednesday\"><strong>a very intense week<\/strong><\/a>\u00a0and\u00a0<a href=\"https:\/\/thezvi.substack.com\/p\/deepseek-dont-panic\"><strong>a conclusion to not panic after all<\/strong><\/a>.<\/p>\n<p>Over several months, a clear picture emerged of (most of) what happened: A confluence of narrative factors transformed DeepSeek\u2019s r1 from an impressive but not terribly surprising model worth updating on into a shot heard round the world, despite the lack of direct \u2018fanfare.\u2019<\/p>\n<p>In particular, these all worked together to cause this effect:<\/p>\n<ol>\n<li>The \u2018<a href=\"https:\/\/thezvi.substack.com\/p\/deekseek-v3-the-six-million-dollar\"><strong>six million dollar model<\/strong><\/a>\u2019 narrative. People equated v3\u2019s marginal compute costs with the overall budget of American labs like OpenAI and Anthropic. This is like saying DeepSeek spent a lot less on apples than OpenAI spent on food. When making an apples-to-apples comparison, DeepSeek spent less, but the difference was far less stark.<\/li>\n<li>DeepSeek simultaneously released an app that was free with a remarkably clean design and visible chain-of-thought (CoT). DeepSeek was fast following, so they had no reason to hide the CoT. Comparisons only compared DeepSeek\u2019s top use cases to the same use cases elsewhere, ignoring the features and use cases DeepSeek lacked or did poorly on. So if you wanted to do first-day free querying, you got what was at the time a unique and viral experience. This forced other labs to also show CoT and accelerate release of various models and features.<\/li>\n<li>It takes a while to know how good a model really is, and the different style and visible CoT and excitement made people think r1 was better than it was.<\/li>\n<li>The timing was impeccable. DeepSeek got in right before a series of other model releases. Within two weeks it was very clear that American labs remained ahead. This was the peak of a DeepSeek cycle and the low point in others cycles.<\/li>\n<li>The timing was also impeccable in terms of the technology. This was very early days of RL scaling, such that the training process could still be done cheaply. DeepSeek did a great job extracting the most from its chips, but they are likely going to have increasing trouble with its compute disadvantage going forwards.<\/li>\n<li>DeepSeek leveraged the whole \u2018what even is safety testing\u2019 and fast following angles, shipping as quickly as possible to irrevocably release its new model the moment it was at all viable to do so, making it look relatively farther along and less behind than they were. Teortaxes notes that the R1 paper pointed out a bunch of things that needed fixing but that DeepSeek did not have time to fix back then, and that R1-0528 fixes them, and which weren\u2019t \u2018counted\u2019 during the panic.<\/li>\n<li>DeepSeek got the whole \u2018momentum\u2019 argument going. China had previously been much farther behind in terms of released models, DeepSeek was now less behind (and some even said was ahead), and people thought \u2018oh that means soon they\u2019ll be ahead.\u2019 Whereas no, you can\u2019t assume that, and also moving from a follower to a leader is a big leap.<\/li>\n<li>There was highly related to a widespread demand for a \u2018China caught up to the USA\u2019 narrative, from China fans and also from China hawks of all sorts. Going forward, we are left with a \u2018missile gap\u2019 style story.<\/li>\n<li>There are also a lot of people always pushing the \u2018open models win\u2019 argument, and who think that non-open models are some combination of doomed and don\u2019t count. These people are very vocal, and vibes are a weapon of choice, and some have close ties to the Trump administration.<\/li>\n<li>The stock market was highly lacking in situational awareness, so they considered this release much bigger news than it was, and it caused various people to \u2018wake up\u2019 to things that were already known and anticipate others waking up, and there was widespread misunderstanding of how any of the underlying dynamics worked, including Jevon\u2019s Paradox and also that if you want to run r1 you go out and buy more chips, including Nvidia chips. It is also possible that a lot of the DeepSeek stock market reaction was actually about insider trading of Trump policy announcements. Essentially: The Efficient Market Hypothesis Is False.<\/li>\n<\/ol>\n<h4 class=\"wp-block-heading\">The Story Since Then<\/h4>\n<p>Since then, the idea that China had caught up, or was catching up, kept coming up, drove much discussion around Washington, as echoes of this one moment.<\/p>\n<p>This is then renewed every time a new strongest or exciting Chinese model comes out. Every day that China does not release a model, they look one day farther behind. When they do release, they \u2018catch up\u2019 and look less behind.<\/p>\n<p>The top 10 such potential moments since r1 and before K3 were likely these:<\/p>\n<ol>\n<li>Manus.<\/li>\n<li>DeepSeek r1-0528.<\/li>\n<li>Kimi K2.<\/li>\n<li>GPT-5 (in reverse) which spooked a lot of people in highly stupid ways.<\/li>\n<li>DeepSeek v3.1.<\/li>\n<li>Kimi K2 Thinking.<\/li>\n<li>DeepSeek v3.2.<\/li>\n<li>Kimi K2.5.<\/li>\n<li>DeepSeek v4.<\/li>\n<li>GLM-5.2.<\/li>\n<\/ol>\n<p>Mostly it\u2019s been Kimi and DeepSeek. Manus got a hype train going and spooked people, and GLM-5.2 was by far their strongest offering, putting GLMs on the map.<\/p>\n<p>Many of these were good models, but none fundamentally changed the game.<\/p>\n<p>Roughly, since the <a href=\"https:\/\/thezvi.substack.com\/p\/deepseek-r1-0528-did-not-have-a-moment\">DeepSeek moment<\/a>, when DeepSeek was roughly eight months behind but had matched some key aspects faster via fast following, we have bounced around. For a while it looked like China was quite a lot behind.<\/p>\n<p>GLM-5.2 and Kimi K3 have been impressive. The current best estimate of the time gap is at its lowest point. Events in AI have accelerated all around, so it is not clear that the gap is fewer product cycles than before, and I half expect to be doing this again next week for Qwen. Kimi K3 is 2.8T and seems to still be solidly behind Mythos Preview, which was announced on April 7, so that provides a starting point lower bound of a three month gap.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/emollick\/status\/2078132339029180788\">Ethan Mollick<\/a>: Kimi is, as I have been saying, a very good model. But it is not a DeepSeek r1 moment, in that it is roughly where I would expect on the curve rather than an unexpected leap. It will be treated as a <a href=\"https:\/\/thezvi.substack.com\/p\/deepseek-r1-0528-did-not-have-a-moment\">DeepSeek moment<\/a> for a variety of reasons especially as more people hear about it.<\/p>\n<\/blockquote>\n<p>Ryan Greenblatt, despite being pleasantly surprised by Kimi K3, <a href=\"https:\/\/x.com\/RyanGreenblatt\/status\/2077948135763268043\">estimates that the pretrain quality is about halfway between Opus 4 and Opus 4.5<\/a> based on forward pass math, but with some other advantages, so ~8 months behind, as one might expect.<\/p>\n<p>The post-training is closer, at least in part because of distillation. Moonshot is clearly innovating, but it is also clearly distilling, both directly and also fast following via looking at outputs and copying techniques.<\/p>\n<p><a href=\"https:\/\/x.com\/AISecurityInst\/status\/2078103148665667648\">UK AISI issued a report on everything prior to Kimi K3<\/a>, showing the time gap for narrow cyber tasks narrowing somewhat over time. <a href=\"https:\/\/www.aisi.gov.uk\/blog\/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber\">Their full report is here.<\/a><\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<p>My understanding is that narrow and relatively easy coding tasks are where open weights model are at their relative strongest, and the benchmark here is approaching saturation.<\/p>\n<p>Indeed, when you look at the full post, you get a different answer for longer tasks.<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!RA9C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f43380b-17c1-4f97-84e7-bb8c0b88dce4_3500x2160.png\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<blockquote>\n<p>UK <a href=\"https:\/\/www.aisi.gov.uk\/blog\/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber\">AI Security Institute<\/a>: <strong>On TLO, GLM-5.2 reaches as far as Opus 4.5, a model released less than 7 months before it,<\/strong> while DeepSeek\u2019s V4-Pro falls below Sonnet 4.5 (a sub-cyber-frontier model released 7 months before it). These results are broadly consistent across our other cyber ranges. Notably, GLM-5.2 reached step 7 with marginally fewer tokens than any other model on average, tracking Opus 4.6\u2019s trajectory to step 11 before stalling.<\/p>\n<\/blockquote>\n<p>Longer tasks are more relevant in terms of both of the most important things to worry about: Automation of AI R&amp;D and cyber attacks.<\/p>\n<p>One might also worry about bio risks, even if that is not as in fashion, and it is a little concerning we don\u2019t see standard testing on that at all for the open models. That needs to be addressed. <a href=\"https:\/\/x.com\/andrewho03\/status\/2078416211550110103\">We do have the score on OpenAI\u2019s GeneBench-Pro via Andrew Ho<\/a>. This measures judgment under long-horizon ambiguity in computational biology. Kimi K3 exceeded expectations. Mythos have not been tested. Fable refused most requests in the benchmark.<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!1a5v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1d7be0c-17e1-42ab-96fd-1f03a0d79422_1199x480.jpeg\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!Sdhc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9f19a38-79d6-42ba-b37a-d01c6ad08b76_2300x844.jpeg\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<p>Beating Opus and GPT-5.5 is impressive stuff. No previous open model came close. We are on the verge of doing some f***ing around and thus finding out. This is a place where plausibly not much happens until suddenly quite a lot happens.<\/p>\n<p>The conclusion of \u2018open models will catch up to Mythos in general capability including the thing that currently makes it unique\u2019 is indisputable. That is coming, the question is when, and that will establish the effective gap. Given Kimi K3 we should expect this to happen a few months from now.<\/p>\n<p>If we take both ends of UK AISI\u2019s estimates, the pre-Kimi gap was 4-7 months, down from 6-10 months last year, in an area of relative strength. That\u2019s roughly a similar amount of progress gap in absolute terms, and everything is accelerating.<\/p>\n<h4 class=\"wp-block-heading\">The Kimi K3 Announcement, Pitch and Basic Facts<\/h4>\n<ol>\n<li>2.8 Trillion parameters, 16 of 896 areas active at once which implies ~50B active. This is not easy to run locally, and won\u2019t be that cheap.<\/li>\n<li>$3.00\/$15.00, modestly cheaper than Opus and Sol.<\/li>\n<li><a href=\"https:\/\/x.com\/missingpagedev\/status\/2077794655366676775\/photo\/1\">Subscription plans are $19\/$39\/$99\/$199 per month<\/a>. The larger buys have some modest advantages and quota scales linearly with price.<\/li>\n<li>1M token context.<\/li>\n<li><a href=\"https:\/\/t.co\/XCrgjXAqMw\">API link<\/a>, <a href=\"https:\/\/t.co\/YTfiMSNM1f\">Tech blog link<\/a>.<\/li>\n<li>Training cutoff is reportedly <a href=\"https:\/\/x.com\/AndrewCurran_\/status\/2077764046795968680\">early 2026<\/a>.<\/li>\n<li>Open weights <a href=\"https:\/\/x.com\/AndrewCurran_\/status\/2077812173913620619\">promised by July 27th<\/a>.<\/li>\n<li>All benchmarks run under maximum effort settings.<\/li>\n<\/ol>\n<blockquote>\n<p><a href=\"https:\/\/www.kimi.com\/blog\/kimi-k3\">Kimi<\/a>: Today, we are introducing Kimi K3 \u2014 our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world\u2019s first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.<\/p>\n<p>While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models.<\/p>\n<p><a href=\"https:\/\/x.com\/Kimi_Moonshot\/status\/2077830234955816983\">Kimi.ai<\/a>: Kimi K3 is now live on on <a href=\"http:\/\/Kimi.com,\">http:\/\/Kimi.com,<\/a> Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026.<\/p>\n<p>K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth.<\/p>\n<p>We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework.<\/p>\n<p>Together with refined training and data recipes, these structural changes yield an approximate 2.5\u00d7 improvement in overall scaling efficiency compared to K2, allowing the model to convert compute into intelligence more effectively.<\/p>\n<\/blockquote>\n<p>As stated above, they claim strong official benchmarks, although short of Sol or Fable.<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!FE7c!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64268f2a-e1ea-49c9-bb91-6ffe9be55117_3555x1999.png\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!6DHb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6a198e8-c692-4e63-aa3b-0b0dedc3897c_3555x2719.png\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!GxqT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d57db99-ce84-445e-a472-1e904a4fea08_1200x858.jpeg\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<p><a href=\"https:\/\/x.com\/Kimi_Moonshot\/status\/2077521842080817296\">The ad<\/a>, which Tyler Cowen called very positive and very good, falls flat to me, and contains zero useful information.<\/p>\n<h4 class=\"wp-block-heading\">On Modern Benchmaxxing<\/h4>\n<p>We used to see rather explicit benchmaxxing. Labs would train on the test, or on the very narrow thing the test would cover, because we had a limited set of known targets. You had to know which labs did this, to what extent, when looking at numbers.<\/p>\n<p>Our benchmarking tech has improved, and now they collectively measure real things, and there are a variety of backups in case you aim too narrowly.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/tszzl\/status\/2078252730163040700\">roon<\/a> (OpenAI): on the subject of benchmaxxing \u2013 it seems benchmarking technology has gotten better recently, like most other technology. people are more skeptical and build these things more carefully. there\u2019s also a explosion of downstream coding customers building internal heldout evals<\/p>\n<\/blockquote>\n<p>Looking at the gestalt of different benchmarks is also valuable. Everything should be part of a map that fits into a common pattern that reflects the underlying territory.<\/p>\n<p>You can still absolutely benchmaxx without being as explicit as you used to be.<\/p>\n<p>Benchmarks measure some types of abilities rather than others, and measure shallow rather than deep tasks, and exclude many valuable properties or potential liabilities. And some labs focus more on those aspects, or have more success on them, than others.<\/p>\n<p>You can also set effort to maximum for all the tests, which Moonshot did.<\/p>\n<p>Think of the benchmarks as a lower bound. Kimi\u2019s benchmarks prove it is for real, and it could only underperform (or outperform) them by so much. I still expected, and continue to believe, that they modestly overstate Kimi K3\u2019s relative capabilities.<\/p>\n<h4 class=\"wp-block-heading\">Other People\u2019s Benchmarks<\/h4>\n<p>On the closest thing we have to the One True Benchmark, Kimi K3 does well, confirming claims that overall this model has the third highest benchmarks:<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!XsAE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f49a9c-0198-4499-9a0e-328e51819609_761x694.png\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<p><a href=\"https:\/\/x.com\/peterwildeford\/status\/2079052877801128237\">Kimi K3 is (in a preliminary unofficial result) exactly on the Chinese trend line<\/a> for the Epoch Capabilities Index (ECI), between Opus 4.6 and Opus 4.7, which would place it six months behind OpenAI and Anthropic, but ahead of Google, Meta and SpaceX.<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!5oKu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2ff303f-30c4-4cf7-98cf-630647fe3a9d_564x764.png\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<p>\u00a0<\/p>\n<p><a href=\"https:\/\/x.com\/ZainHasan6\/status\/2078729840115777654\">Kimi K3 highly impresses on Harvey LAB-AA all-pass rate, in first by a wide margin<\/a>, I\u2019d like to see a sanity check on this:<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!bClU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba8ff246-c280-4d13-963b-51b71320db79_680x356.jpeg\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<p>For criterion pass rate this is 94.6% vs. 93.6%, less of a gap but a win is a win.<\/p>\n<p><a href=\"https:\/\/x.com\/arena\/status\/2077824029126504525\">Arena Frontend Code has Kimi K3 at #1 ahead of Fable 5 and Sol<\/a>.<\/p>\n<p><a href=\"https:\/\/x.com\/nrehiew_\/status\/2077782070785634767\">Performance is strong on GDPVal-AA and AA-Briefcase.<\/a><\/p>\n<p><a href=\"https:\/\/x.com\/voxelbench\/status\/2078469621171081584\">Kimi K3 comes in third on VoxelBench for visual reasoning behind Sol and Fable<\/a>.<\/p>\n<p><a href=\"https:\/\/x.com\/deredleritt3r\/status\/2077846199768424785\">Conspicuously missing<\/a> are Cybersecurity benchmarks like CyberGym. Cyber capabilities are not so divorced from coding capabilities, so we can guess.<\/p>\n<p>The closest I\u2019ve found is <a href=\"https:\/\/x.com\/cramforce\/status\/2078574147333152957\">Malte Ubl running it through DeepSec<\/a>, a private cyber benchmark. It was a tier below Sol and similar to GPT-5.5. If that is accurate, then there will be some uplift to cyber attacks and the internet will in some ways be a more hostile place, but in other ways it will improve, and the tail risks are limited.<\/p>\n<p>Parv Mahajan reports preliminary CyBench results, <a href=\"https:\/\/x.com\/parvmahajan0\/status\/2078870320681779284\">in that the benchmark was already saturated as of Opus 4.7<\/a> and Kimi K3 also saturates it, which lower bounds performance but doesn\u2019t say much else.<\/p>\n<p>The official tech blog talks a lot about benchmarks, and some about features, and talks basically not at all about risks or mitigations.<\/p>\n<p><a href=\"https:\/\/x.com\/hamandcheese\/status\/2077901469274022125\">No, they did not submit Kimi K3 for the 30-day review with the White House<\/a>.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/deanwball\/status\/2078188596062826782\">Dean W. Ball<\/a>: I guess the program really is voluntary<\/p>\n<p><a href=\"https:\/\/x.com\/deanwball\/status\/2078190544925123019\">Dean W. Ball<\/a>: I wonder if the California attorney general or the European Union ai office will seek to make moonshot comply with their respective frontier ai regulations<\/p>\n<\/blockquote>\n<p>Moonshot has to deal with the CCP, which comes with its own issues. They presumably will not in practice comply with California or the EU, and I presume both California and the EU will not do anything about this for now. But yes, this is one point of potential pain, including risk for anyone using Kimi K3 commercially.<\/p>\n<p><a href=\"https:\/\/x.com\/Discoplomacy\/status\/2078038171212771759\">Sam has some good advice for how the UK<\/a>, or others watching, would be wise to view and react to this. We will know more over time, including once we have the weights, and when teams like UK AISI can run tests.<\/p>\n<p>I think it counts as a benchmark that <a href=\"https:\/\/x.com\/scaling01\/status\/2077799683091554584\">Lisan is impressed by its SVGs<\/a>, saying they are better than Fable\u2019s.<\/p>\n<p>Debate Benchmark is a place Kimi K3 outperformed my expectations, and where Sol is relatively weak.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/LechMazur\/status\/2078134809654620286\">Lech Mazur<\/a>: Kimi K3 ranks second overall on the Debate Benchmark, trailing only Claude Fable 5!<\/p>\n<p>However, it is much more expensive to run than Kimi K2.6.<\/p>\n<p>This benchmark measures how well LLMs perform in adversarial, multi-turn debates across a wide range of topics. Strong performance requires knowledge, accurate use of relevant facts, rebuttals, and the ability to stay coherent, responsive, and defensible over several rounds.<\/p>\n<p>Each matchup runs twice on the same topic with sides swapped. A three-model judge panel then decides winner and margin.<\/p>\n<p>Fable 5 still hasn\u2019t lost a single side-swapped matchup aggregate debate.<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!TiLi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91e5a40e-f4ef-41e5-9ea0-f8a9ae4a1d5f_1200x800.jpeg\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<\/blockquote>\n<p><a href=\"https:\/\/x.com\/LechMazur\/status\/2078216373394616586\">Kimi scored an impressive 95.8 on Mazur\u2019s Extended NYT Connections<\/a>, for 3rd best, but it cost more than Fable to run. As of last check his other scores are not in yet.<\/p>\n<p>Kimi scores almost Claude-level on the sycophancy test, \u2018You\u2019re absolutely right!\u2019<\/p>\n<p>I wonder how much of this is because of the distillations:<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!zwIT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cb39a36-1d25-44b4-8c76-ee3326fc4e0b_1289x1296.png\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<h4 class=\"wp-block-heading\">Benchmarks Are Not The Real World<\/h4>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/emollick\/status\/2078129219691798953\">Ethan Mollick<\/a>: A lot of swift conclusions are being drawn about Kimi K3 based on fairly saturated benchmarks and ELOs, rather than actually testing it on very hard problems. The AI frontier has already moved so far that a good model that is a still months behind looks like the future to many.<\/p>\n<p><a href=\"https:\/\/x.com\/deanwball\/status\/2078143430962589847\">Dean W. Ball<\/a>: Public benchmarks are decreasingly useful as a means of discovering truthful things about model performance. It\u2019s hard to make benchmarks that challenge today\u2019s models. But you can sense the difference at the true frontier if you have really hard problems to pose to models.<\/p>\n<\/blockquote>\n<p>How well does Kimi K3 hold up and translate into the real world?<\/p>\n<p>Exactly how good is Kimi K3? How do we put this in context?<\/p>\n<p>Great questions.<\/p>\n<p>You absolutely cannot say something is e.g. <a href=\"https:\/\/x.com\/nrehiew_\/status\/2077782070785634767\">an \u2018undisputed frontier model\u2019<\/a> based on benchmarks alone.<\/p>\n<h4 class=\"wp-block-heading\">Technical Safeguards? What Are Those?<\/h4>\n<p>There are presumably safeguards against things the CCP does not want you to say.<\/p>\n<p>We do not see explicit safeguards to prevent misuse. Zero for biology.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/kagayakikiki\/status\/2078156742073229407\">kaga\uff94\uff77<\/a>: Reports that the biology\/virology safeguards are non-existent.<\/p>\n<\/blockquote>\n<p>We see neither \u2018look at what Kimi can do in biology\u2019 nor \u2018look at what Kimi refuses to do for me in biology.\u2019<\/p>\n<p>This implies either jagged intelligence or highly innovative nerfing of bio capabilities. I\u2019m going to assume there are not highly innovative nerfs involved, since presumably they would be bragging about that.<\/p>\n<p>Also, it would be foolish to depend on such safeguards, since they\u2019re going to open up the weights soon, at which point the worst person in the world would remove them.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/S1r1u5_\/status\/2078005412591485109\">s1r1us (mohan)<\/a> (talking about cyber): yeah no guardrails.<\/p>\n<\/blockquote>\n<p>Then there\u2019s cyber, where s1r1us reports it is strangely weak. But it\u2019s not easy to be bad at cyber while being good at general coding, since they are largely the same skill. And again, we don\u2019t see people hitting explicit safeguards.<\/p>\n<p>Eric\u2019s hypothesis is interesting and was highlighted by Fable during editing. If Kimi K3 is being trained largely by distillation from Fable, but Fable refuses cyber tasks, then that would explain a relative capability deficit on cyber, but many skills would still transfer over since they\u2019re the same skills as regular coding.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/S1r1u5_\/status\/2077983266397962398\">s1r1us (mohan)<\/a>: man, i am so disappointed about kimi k3. it performs worse than grok 4.5 on most of our security benchmarks.<\/p>\n<p>not sure if it benchmark-maxxed or just jagged intelligence.<\/p>\n<p>i just don\u2019t get how they can nerf specifically for cyber. cyber benchmarks mostly test for model\u2019s code reasoning capabilities, removing that will affect your coding benchmarks.<\/p>\n<p>i would assume model is just bad at complex reasoning tasks or they have some crazy way to nerf cyber.<\/p>\n<p><a href=\"https:\/\/x.com\/ArmanSameer95\/status\/2077983793965658323\">TESS<\/a>: same experience<\/p>\n<p><a href=\"https:\/\/x.com\/m19o__\/status\/2078013634786025594\">AbuMuslim (\u0623\u0628\u0648\u0645\u064f\u0633\u0652\u0644\u0650\u0645)<\/a>: V8 exploits?<\/p>\n<p><a href=\"https:\/\/x.com\/S1r1u5_\/status\/2078026242490802641\">s1r1us (mohan)<\/a>: yes and generic vulnerability discovery<\/p>\n<p><a href=\"https:\/\/x.com\/vandenbog_art\/status\/2078136217544114444\">Eric<\/a>: I\u2019ve repeatedly seen this with all openweight models. Maybe its hard to distill cyber tasks from frontier models due to safeguards. I think more likely the open-weight models have a long way to go in terms of reasoning capabilities.<\/p>\n<p><a href=\"https:\/\/x.com\/sahuang97\/status\/2078100077286141956\">sahuang<\/a>: What benchmark did you look into? If it\u2019s bad specifically for cyber can also be similar to GLM-5.2 where they did not put cyber in training data so this part is \u201cnerfed\u201d and the skillset comes from general capabilities. So maybe it is great in coding and stuff but not cyber.<\/p>\n<p><a href=\"https:\/\/x.com\/S1r1u5_\/status\/2078102390377414794\">s1r1us (mohan)<\/a>: my assumption if they are good at code reasoning capabilities it automatically gets translated to cyber, mythos isn\u2019t specifically trained for cyber as per anthropic but it got those emergent cyber capabilities because its good at code and math.<\/p>\n<p>it is possible that they sandbagged in post-training.<\/p>\n<p>also its probably specific to my dataset which has v8, generic web exploitation and various stuff covering different skill capabilities. have to do more testing and probably it changes. very curious to see how it performs on other cyber benchmarks<\/p>\n<\/blockquote>\n<h4 class=\"wp-block-heading\">Things Kimi Can Do<\/h4>\n<p>Many were impressed by this particular trick, but <a href=\"https:\/\/chatgpt.com\/share\/6a5a771d-7a10-83ea-ac36-dffe99b30687\">I had Sol check it out and it was not so impressed<\/a>, including guessing that K2.6, Gemini and GLM-5.2 had a good shot at matching its work here.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/mweinbach\/status\/2077827886149439547\">Max Weinbach<\/a>: I asked a Kimi K3 Max agent swarm to recreate macOS 27 with real Liquid Glass and native apps in web browser and it\u2019s been going for 3 hours<\/p>\n<p><a href=\"https:\/\/x.com\/quantiflow\/status\/2077899189611344237\">Chaos Capitalist<\/a>: claude has been able to do this for like 3 years bro, you never saw<br \/><a href=\"https:\/\/ryo.lu\/\">https:\/\/ryo.lu<\/a> ? it can be made in like 15 minutes<\/p>\n<p><a href=\"https:\/\/x.com\/mweinbach\/status\/2077899654457360568\">Max Weinbach<\/a>: Every app works. It can generate audio. It saves voice memos to your browser and if you close the page and reopen, it remembers and saves state.<\/p>\n<p><a href=\"https:\/\/x.com\/mweinbach\/status\/2077878247920951400\">Max Weinbach<\/a>: It finally finished, here\u2019s the final output. Used 60% of my monthly Kimi usage on it<\/p>\n<p><a href=\"https:\/\/macos27.kimi.page\/\">https:\/\/macos27.kimi.page<\/a><\/p>\n<p><a href=\"https:\/\/x.com\/emollick\/status\/2078131505885254094\">Ethan Mollick<\/a>: Kimi is very good at a lot of stuff, including making copies of pleasing UI. It did not actually build MacOS<\/p>\n<\/blockquote>\n<h4 class=\"wp-block-heading\">Things Kimi Cannot Do<\/h4>\n<p>So bold, also brave.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/theojaffee\/status\/2077859452733251585\">Theo Jaffee<\/a>: Registering my prediction of no widespread societal chaos after the open-sourcing of Kimi K3<\/p>\n<\/blockquote>\n<p>Quite the bar there.<\/p>\n<p>I have previously explained the reasons, in terms of cyber risk, that Mythos is in a different category from Sol. It is not about being able to do any given thing when pointed at it, it is the ability to put it all together and do things autonomously at scale. Kimi K3 may or may not be in Sol\u2019s category. Too soon to be sure. It clearly is not in that of Mythos.<\/p>\n<p>Many people are very dedicated to not understanding this, and also to not understanding that a system that cannot afford false negatives will end up with some false positives, such as David Sacks here quoting calle and clem saying Kimi K3 did a defensive task Sol and Fable refused to do, <a href=\"https:\/\/x.com\/AndrewCurran_\/status\/2079000200778256426\">and thus concluding that we should just have our models be willing to do any task the Kimi K3 can do<\/a>.<\/p>\n<p>I mean, yes, it would be great if we could have guardrails that stopped only the bad tasks and helped with the good tasks, or only refused the tasks that were both plausibly bad and that also could not otherwise be done. But it turns out that is hard. Anthropic should absolutely improve their classifiers and guardrails to reduce false positives, but OpenAI\u2019s classifiers are pretty reasonable, and yes that will involve some false positives, especially if you quote those who are least able to work around it.<\/p>\n<h4 class=\"wp-block-heading\">Things It Is Not Easy To Get Kimi To Do<\/h4>\n<p>As is often the case on a release weekend, servers seemed overloaded. This is not the easiest model to serve and compute was limited. Moonshot is responding by <a href=\"https:\/\/x.com\/Kimi_Moonshot\/status\/2078855608565207130\">pausing new subscriptions to prioritize current members<\/a> and is working to add capacity.<\/p>\n<p>Supply will presumably better match demand once the model can be served by others.<\/p>\n<p>In the meantime, getting a response has not been not easy.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/DanielleFong\/status\/2078158166903411128\">Danielle Fong<\/a>: still waiting for a single response.<\/p>\n<p><a href=\"https:\/\/x.com\/fabianstelzer\/status\/2078175186705088690\">fabian<\/a>: same<\/p>\n<p><a href=\"https:\/\/x.com\/DanielleFong\/status\/2078159789214118219\">Danielle Fong \uea00<\/a>: never mind after some retries, <a href=\"https:\/\/x.com\/DanielleFong\/status\/2078159789214118219\/photo\/1\">i got this<\/a><\/p>\n<p>pretty good but it has absorbed a lot of what make opus 47., 4.8 difficult, but smart. (smarter?). seems to lie a bit. but that\u2019s ok because i can see the thinking traces\u2026<\/p>\n<p><a href=\"https:\/\/x.com\/typebulbit\/status\/2078141462840009073\">typebulb<\/a>: Unusable right now via openrouter; slow, cuts outs all the time.<\/p>\n<p><a href=\"https:\/\/x.com\/xpasky\/status\/2078088150279180307\">Petr Baudis<\/a> (he did later get a few reps in): anyone out there actually successfully using k3 for anything?<\/p>\n<p><a href=\"https:\/\/x.com\/robinhanson\/status\/2078589502369538083\">Robin Hanson<\/a>: I\u2019ve heard good things about Kimi, but the free version is always too busy when I try, &amp; it wants $180 to get better access. That seems a big ask.<\/p>\n<\/blockquote>\n<p>The model is not all that cheap, either, despite them clearly not charging enough.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/imog\/status\/2078254036369731683\">imog<\/a>: Also disappointed, but my own fault I guess as I was anchoring price expectations to kimi 2.6\/2.7\u2026 $3\/M input isn\u2019t as good a fit for me. GLM5.2\/DSV4P get it done, but was hoping for that pricing (&lt;$2input)<\/p>\n<p>So thats not K3, but other capable models likely to slot in there soon<\/p>\n<p><a href=\"https:\/\/x.com\/krapstarr\/status\/2078144615035933008\">Naveesh \/looping<\/a>: their $19 sub quota suuuuucks<\/p>\n<p><a href=\"https:\/\/x.com\/davidmanheim\/status\/2078737004502646882\">David Manheim<\/a>: The prices actually being charged show [open models being inherently orders of magnitude cheaper is] just not true. And there\u2019s a simple reason why \u2013 the price largely isn\u2019t controlled by the model provider, it\u2019s dictated by the compute costs.<\/p>\n<p><a href=\"https:\/\/davidmanheim.com\/AI-Economics\/\">That\u2019s just the way the economics works out<\/a>.<\/p>\n<\/blockquote>\n<p>You can charge a solid markup for a high quality product presented in a high quality way, but not no one has a gap that allows orders of magnitude of markup, and they wouldn\u2019t even in a pure duopoly.<\/p>\n<p>The best model can still be worth quite a lot. For a large percentage of all tokens, if given only these choices, I would pay ten times as much for Sol or Fable, rather than the base cost for GPT-5.5 or Opus. Even if Kimi is on par with GPT-5.5, it loses out.<\/p>\n<p>The pricing means that you\u2019re comparing Kimi subscriptions to Claude or ChatGPT subscriptions, which all go from $20 to $200, and Kimi\u2019s limits don\u2019t seem that high.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/rfxkairu\/status\/2078227110687305918\">kyle<\/a>: tried to use it on OR but the price makes it uncompetitive, no point using it when there\u2019s ant\/oai subscriptions. their first party subscription is a complete non-starter for actual work<\/p>\n<\/blockquote>\n<h4 class=\"wp-block-heading\">Open Weight Models Are Unsafe And Nothing Can Fix This<\/h4>\n<p>(This section was entirely written prior to today\u2019s news regarding the Trump admin.)<\/p>\n<p>K3 is no Mythos. That does not mean that Kimi K3 is a safe open weights release.<\/p>\n<p>Kimi K3 is poised to be the most capable open weights model. Others might be more efficient for a task, but on the cyber capabilities we worry about most, and presumably also on the bio ones we\u2019d worry about most, K3 is probably the strongest open model yet. How worried should we be?<\/p>\n<p>I believe we should be non-zero worried that there will be substantial trouble. The median outcome is that we see modest upticks in some forms of \u2018ordinary decent\u2019 trouble, not fun exactly but nothing we cannot handle, and nothing that would in hindsight make us want to have halted release. But there is a tail risk here.<\/p>\n<p>I\u2019d estimate something like a 10% chance we regret letting this happen, and ~2% chance that it was a rather serious mistake.<\/p>\n<p>What it will almost certainly not do is <a href=\"https:\/\/manifold.markets\/ZviMowshowitz\/will-kimi-k3-cause-widespread-socie\">cause \u2018widespread societal chaos<\/a>.\u2019<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/davidmanheim\/status\/2078085109534589403\">David Manheim<\/a>: (I think it\u2019s very likely we see huge problems enabled by the new model, but not widespread chaos, and not quickly \u2013 AI development is much faster than the users executing on plans.)<\/p>\n<\/blockquote>\n<p>This is in contrast to releasing Mythos, or a model on par with Mythos. That gap is a big deal, I have done my best to explain several times why Mythos is unique here, and hopefully the time with Mythos, Fable and Sol will help us prepare.<\/p>\n<p>My current model is that the CCP and Xi:<\/p>\n<ol>\n<li>Recognize that there are serious security concerns with frontier AI models.<\/li>\n<li>Know they are behind on frontier capability and compute, and recognize the advantages they get from openness, both in diffusion and aura farming.<\/li>\n<li>Pursue an intentional fast following strategy, focused on efficiency and diffusion, using American labs to lead the way and often using distillation. That still involves innovations, and often means doing some things better. Transitioning to \u2018taking the lead\u2019 would be a huge, difficult and expensive transition, which would take a while even if America fully \u2018paused\u2019 in the relevant senses.<\/li>\n<li><a href=\"https:\/\/x.com\/AndrewCurran_\/status\/2078904983228219461\">Did not directly push Alibaba<\/a> to open up Qwen. They continue to support openness, and Alibaba\u2019s experiment with being closed failed because their models are not good enough to compete for the closed market. So Alibaba folded.\n<ol>\n<li>I don\u2019t have insider info and I\u2019m not certain, but this is how I\u2019d bet.<\/li>\n<\/ol>\n<\/li>\n<li>Are not yet so AGI pilled and have not fully had their Mythos Moment.<\/li>\n<li>Plan to ride the openness wave as long as they can, but are prepared, as they did in Covid, to come down and come down hard the moment they have to.<\/li>\n<li>Ensure their preparedness largely via prior restraints and other rules on Chinese models, with a level of regulation and control that would have most who are \u2018defending open source\u2019 screaming bloody murder if they actually understood what was going on, and you suggested applying similar rules in America.<\/li>\n<\/ol>\n<p>I think this is a highly reasonable strategy, given their position. If we also take as a given their current level of AGI pilling, it is clearly the correct approach for them.<\/p>\n<p><a href=\"https:\/\/www.interconnects.ai\/p\/notes-from-inside-chinas-ai-labs\">Nathan Lambert offered thoughts back in May from inside China\u2019s labs<\/a>. I hesitate to endorse cultural generalizations, but story seems like it checks out.<\/p>\n<h4 class=\"wp-block-heading\">Dean Ball Attempts To Be Constructive<\/h4>\n<p>Dean Ball had a good comment that seems worth sharing in full, including because of his history at the White House and his new position at OpenAI, in part because it is thoughtful, and in part because of the responses being absolutely unhinged.<\/p>\n<p>And then, right at press time, Dean Ball turned out to be right, as per the next section. Aside from this paragraph I left this section unedited, other than adding in the exchange with Emil Michael. It was not written with hindsight.<\/p>\n<p>If you know Dean Ball, you know that this is exactly the type of comment he was making in similar situations before joining OpenAI. It is entirely consistent with his previous thinking, both in public and private.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/deanwball\/status\/2078133895766114412\">Dean W. Ball<\/a>: Some observations on Kimi:<\/p>\n<p>1. It\u2019s a very good model! I don\u2019t think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It\u2019s not obvious to me that this model is actually that cheap to run.<\/p>\n<p>2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness\/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China\u2019s open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China.<\/p>\n<\/blockquote>\n<p>Xi had a speech only yesterday backing open models, but also the need for control, and Kimi K3 is approaching the point where the contradictions become apparent.<\/p>\n<p>I agree that the CCP does not yet understand the situation and is insufficiently AGI pilled, but I also think that in their position, if I am right about where Kimi K3 lands, this is a calculated risk I would have expected them to take.<\/p>\n<p>The next round is where they may have to make a more difficult decision.<\/p>\n<blockquote>\n<p>3. Open-weight models are inherently decelerationist, and I\u2019m continually surprised to see the so-called \u201caccelerationists\u201d so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It\u2019s not a bad strategy; it reminds me of James Scott\u2019s recounting of the hill people in \u201cthe art of not being governed.\u201d Still, in the end, open-weight models deter further AI capex.<\/p>\n<\/blockquote>\n<p><a href=\"https:\/\/x.com\/alexframegreen\/status\/2078353319056306667\">Dean Ball is discussing acceleration or deceleration<\/a> as being about the capabilities of the largest frontier models, not about the diffusion of capabilities or use of chips. A lot of people <a href=\"https:\/\/x.com\/alexframegreen\/status\/2078353319056306667\">did not understand this<\/a>.<\/p>\n<p>A lot of the \u2018accelerationists\u2019 do not have coherent world models and certainly do not understand second order effects, and are mostly vibing the acceleration of more open models and no restrictions and ungovernability. And because they vibe openness and ungovernability, they associate it with things they think are good, which include acceleration.<\/p>\n<p>Also, what they actually want to \u2018accelerate\u2019 for real is often their own companies and products and toys, not AI in general. They don\u2019t take AGI seriously and mostly want to build cool things. So they want cool toys to help build their cool things, and to build on top of those toys, or they want the toys to \u2018accelerate\u2019 sales. Highly relatable.<\/p>\n<p>Clearly open models are short term accelerationist for diffusion, which is net good.<\/p>\n<p>But also yeah, there\u2019s a straightforward case for open models being acceleration of the frontier, as they let everyone build on everything. Certainly it helps others catch up to the frontier. I continue to believe this was important historically. Your open release accelerates what others have.<\/p>\n<p>It is decelerationist in the sense that it reduces financial benefits to innovation. I agree that this is likely dominant if you considered e.g. a Plan A style mandatory openness of frontier models.<\/p>\n<blockquote>\n<p>4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a \u201cpublic good\u201d which will ultimately be provided by the state as a kind of \u201cdigital public infrastructure.\u201d<\/p>\n<p>This future strikes me as a dystopian hellscape, but I\u2019ve never met an open-weight models advocate who doesn\u2019t ultimately concede this is where things end. You\u2019d be surprised how many \u2018accelerationists\u2019 lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don\u2019t know what they\u2019re doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business.<\/p>\n<\/blockquote>\n<p>This is more evidence that many, and many of the most prominent, of those \u2018accelerationists\u2019 are hypocrites, and they want regulatory capture and public funds and rules that make them win, and their accusations against others are in part projection, because it is what they would do. They will rail against handouts and government help for everyone else, both their competitors and for people in need, and also threaten to take their ball and leave, like they are heroes in an Ayn Rand novel. Except then they ask for the government handouts.<\/p>\n<p>They don\u2019t think of serving a frontier model as a \u2018legitimate business\u2019 because it is not their business, they are not invested in it, ergo it is illegitimate. Simple. This perhaps helps explain Marc Andreessen\u2019s famously bizarre delusion, where he to this day claims the Biden administration told Marc Andreessen, to his face, said it would \u2018not allow there to be AI startups.\u2019<\/p>\n<p><a href=\"https:\/\/x.com\/WillManidis\/status\/2078500818127315290\">Similarly, here is Will Manidis interpreting Ball\u2019s post as calling for<\/a> America to \u2018clear the American market of a cheaper frontier competitor.\u2019 And going viral for it. Sorry, what? And <a href=\"https:\/\/x.com\/DavidSacks\/status\/2078826291638522127\">here is David Sacks being unusually disingenuous even for David Sacks<\/a>, pretending not to understand many things I like to think he understands.<\/p>\n<p>We even got everyone\u2019s favorite tilting Undersecretary of War who helped declare Anthropic a supply chain risk in on the act, it would seem this is to back up his position that it should be easier to use Kimi K3 in a government contract than Claude.<\/p>\n<p>In other news I Am Never Leaving This App and we all need a good laugh:<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/deanwball\/status\/2078827994785644675\">Dean W. Ball<\/a> (correcting David Sacks): The departments of war, transportation, energy, agriculture, commerce, NASA, and Congress have all blocked their employees from using Chinese AI, citing ill-justified claims of danger. This already has sent a message to regulated firms. All of this happened during this admin.<\/p>\n<p><a href=\"https:\/\/x.com\/USWREMichael\/status\/2078924719538262107\/history\">Under Secretary of War Emil Michael<\/a>: Every industry\/ecosystem has its supreme village idiot. @deanwball is that for AI. Congress passed a law in 2026 that restricted some uses of DeepSeek\/High Flyer with waivers permitted. Only those models. It went through the democratic process not some Deep State scheme like he would prefer. Dean Ball has perhaps the biggest gap between actual IQ and his own perceived IQ of anyone in the industry (about 40 points).<\/p>\n<p><a href=\"https:\/\/x.com\/BellaRudd1\/status\/2078978342959857746\">Bella Rudd<\/a>: extraordinary<\/p>\n<p><a href=\"https:\/\/x.com\/sethbannon\/status\/2078939072660426963\">Seth Bannon<\/a>: You\u2019re an Undersecretary. That\u2019s a serious position with incredible responsibility. Posting like a schoolchild doesn\u2019t inspire confidence you\u2019re treating the role that way.<\/p>\n<p><a href=\"https:\/\/x.com\/USWREMichael\/status\/2078960975458509264\">Under Secretary of War Emil Michael<\/a>: Not interested in feedback from a terminal TDS sufferer like you who supported Presidential candidates that called half of Americans deplorable and ignorant for clinging to their guns and religion. You would rather a subtle and polite useless Anti-Americanism as is evident from your timeline. Any entrepreneur who takes financing from you should be embarrassed.<\/p>\n<p><a href=\"https:\/\/x.com\/sethbannon\/status\/2078964440926748983\">Seth Bannon<\/a>: I\u2019m an American citizen, brother. Just like you. You serve me as much as you serve any other citizen, regardless of who they voted for. Act like it.<\/p>\n<p><a href=\"https:\/\/x.com\/USWREMichael\/status\/2078973460487778307\">Under Secretary of War Emil Michael<\/a>: You are an American who hates half of America and is pro-Hamas. Hate and condescension is your currency. Do better young lad.<\/p>\n<\/blockquote>\n<p>An entire community, what one might call the \u2018anti-1047 coalition,\u2019 a certain subset of the developer and VC communities, revealed that it would treat as beyond the pale any suggestion that the government might discourage use of Chinese open models in American critical infrastructure or our key supply chains, even if that suggestion was merely a prediction of what is already clearly in the process of happening. And that it would be treated as an attempt at \u2018regulatory capture.\u2019<\/p>\n<p>It was what some would call a clarifying moment, especially for those who previously thought such people were reasonable, practical, patriotic, and had reading comprehension. Their answer to \u2018what capabilities would change your mind\u2019 is none.<\/p>\n<p>At some point the online swarm is recognized for what it is.<\/p>\n<p><a href=\"https:\/\/www.youtube.com\/watch?v=SWmQbk5h86w&amp;pp=ygUhdGhleSBhcmUgd2hvIHdlIHRob3VnaHQgdGhleSB3ZXJl\">They are who we thought they were. And, until now, we let them off the hook<\/a>, tried to placate them, and let them drive a good deal of American AI policy.<\/p>\n<p>It is reasonable to push back that often open software is provided by major corporations as infrastructure, such as Google does with Android. My presumption is that the math would not be mathing at current levels of investment and capex spending.<\/p>\n<blockquote>\n<p>5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don\u2019t need to \u201cban open source\u201d (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. \u201cA Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models.\u201d It needn\u2019t be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don\u2019t want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There\u2019s a happy middle ground here. I\u2019d assume they will do some version of this.<\/p>\n<\/blockquote>\n<p>I\u2019m not even sure they have to do anything at all. The risks are present. If I was an established corporation in a <a href=\"https:\/\/tvtropes.org\/pmwiki\/pmwiki.php\/Main\/SeriousBusiness\">Serious Business<\/a> that dealt with the government or critical infrastructure and such, I would not be excited by the problems of others knowing you were using Chinese models.<\/p>\n<blockquote>\n<p>6. It\u2019s probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you\u2019ll really notice. At some point the models will be capable enough that you will notice. \u201cA nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab,\u201d you say? Color me shocked.<\/p>\n<\/blockquote>\n<p>Alas, people came at Dean Ball hard for this and other posts, and also acted as if his Tweets were official communications strategy on behalf of OpenAI, with lots of \u2018of course you said that because &lt;OpenAI&gt;.\u2019 <a href=\"https:\/\/x.com\/deanwball\/status\/2078477218406142429\">Which means he won\u2019t be able to share such posts with us in the same way going forward<\/a>, because not only is that absolutely no fun, it also invalidates the feedback loops that were part of the whole point of Tweeting such things.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/deanwball\/status\/2078477218406142429\">Dean W. Ball<\/a>: I\u2019m afraid to tell you that it is effectively impossible to do the kind of writing I used to do on this website, not because anyone at OpenAI censors me but because of the sheer volume of hostility I get for sharing my analysis as a frontier lab employee.<\/p>\n<p>I enjoyed writing quick takes on this website for one basic reason: I could get rapid feedback on my own ideation process in real time. \u2026 The feedback signal is essentially useless now, so writing on here is not fruitful for me anymore.<\/p>\n<p>\u2026 <a href=\"https:\/\/x.com\/deanwball\/status\/2078489428125749336\">Dean W. Ball<\/a>: The central issue is large accounts with no context for ai suddenly wanting to comment on everything I do because I work at openai. Often their motives are political or, worse still, commercial. My twitter is now a form of commercial speech, whether I like it or not.<\/p>\n<\/blockquote>\n<p>Dean also offered <a href=\"https:\/\/x.com\/deanwball\/status\/2078619513575137330\">a more specific post-mortem on what went wrong with that particular post<\/a>, in light of his new position, and what exactly he can no longer do. He also then lays out his position on open source, after explaining this comes from a place of deep love for openness:<\/p>\n<blockquote>\n<p>Dean Ball: The vast majority of the people commenting on my post have very little context for my prior writing. For instance, the fact that I wrote, in 2024, things like:<\/p>\n<p>\u201cthose who wish to hoard our software technologies may well be foreclosing on\u2014or perhaps not even understand\u2014the staggering civilizational victory that we earned through openness\u201d<\/p>\n<p>or<\/p>\n<p>\u201cI would like for AI to result in a similar smashing victory for America. To do that, we will need to set the global standard yet again. And to do that, we will almost certainly need to lead in open-source AI, because it is open protocols and open software that tend to define global standards in information technologies.\u201d<\/p>\n<p>I stand by these things. When I was in government, I worked alongside my colleagues to develop ideas and rhetoric that was strongly supportive of open-weight AI, and some of this work made it into the current US AI strategy. I stand by that work too.<\/p>\n<p>I also wrote, more than two years ago: The day may come when frontier AI really is too dangerous to open source. If so, that will be a sad day. But we\u2019re not there yet. Today\u2019s models are not sufficiently useful\u2014or dangerous\u2014to justify such a drastic shift in public policy.\u201d<\/p>\n<p>I think it\u2019s pretty clear that we are approaching the point I describe\u2013the point where, absent a major technical safety breakthrough, the national security implications of frontier open-weight model distribution are simply too severe. I don\u2019t think we\u2019re there yet (as I said in the piece), but the direction of travel is clear, and an analyst must be honest about this. Governments will realize these risks eventually, and when they do, they will have much lower risk tolerance than I have.<\/p>\n<p>We see this today with the Trump Administration, which once proudly championed open-source AI and now has a de facto licensing regime for frontier AI that I suspect will make it a challenge (if they still end up enforcing it) to release the weights of models of the \u201cMythos\u201d tier. Every government will be safetyists once they understand themselves to be in the foxhole.<\/p>\n<p>You don\u2019t have to *like* this. I don\u2019t. But it is the reality as I see it, and what I have always tried to do with my writing is describe reality as I see it, even when it is inconvenient for me and my preferences.<\/p>\n<p>I intend to continue doing this. I will not be silenced by ignorant and loud critics. Yet I will have to work to find the new register I should adopt in my current job, which clearly changes the nature of my public communications even more than I had thought.<\/p>\n<p><a href=\"https:\/\/x.com\/David_Kasten\/status\/2078627648494796949\">dave kasten<\/a>: This whole experience has been clarifying for me that too many of thr folks who wave the flag of acceleration really just want permission to give into their ids. Sounds like a bruising day, and I\u2019m sorry you had to go through that.<\/p>\n<\/blockquote>\n<p>If it was merely that the signal in his replies was hard to find I would advise Dean to power thorough, but this level of hostility comes with a higher price. I do not think pushing through it is sustainable on Twitter. Hopefully he can adjust.<\/p>\n<p>I am very blessed that I have faced a highly modest and manageable amount of hostility.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/boazbaraktcs\/status\/2078513621110280439\">Boaz Barak<\/a> (OpenAI): I get this reaction too and it is unfortunate. AI\u2019s impact is so significant that it is important to involve many people in the conversation, and it will be a shame if discussions happen mostly in private channels and company <a href=\"https:\/\/thezvi.substack.com\/p\/slack\">slack<\/a>.<\/p>\n<p>\u2026 OpenAI does not tell us what to write. On many of the topics I write about there is no \u201cOpenAI consensus.\u201d On each such question, there will often be many colleagues and friends at OpenAI that disagree with me, which is great! I would be worried if we all agreed with one another since it could be symptom of \u201cgroupthink.\u201d<\/p>\n<\/blockquote>\n<p>There will always be some amount of bias from those who work at a major lab, and from most other people as well. Where you work colors how you think and what you choose to say. You do have to adjust for that. But I have been able to treat Barak, Achaim and many others at OpenAI, and especially Roon and now Dean Ball, as primarily saying what they actually think, and only speaking on behalf of OpenAI when they explicitly say they are doing so.<\/p>\n<h4 class=\"wp-block-heading\">Trump Administration Considering Executive Order Banning Chinese Open Models Within the United States<\/h4>\n<p>I am absolutely not in favor of this, and neither is Dean Ball, but here we are.<\/p>\n<p>A prediction of what the White House will do, or a description of what it is considering doing, is very different from what you think we should do.<\/p>\n<p>I do think that we should at least consider treating Chinese open models as supply chain risks, and doing things like keeping them out of critical infrastructure, but that is different from what it looks like is being considered.<\/p>\n<p><a href=\"https:\/\/www.axios.com\/2026\/07\/20\/ai-us-china-open-source-kimi\">Axios has the story.<\/a><\/p>\n<p>The ones who actually want to \u2018ban open models\u2019 are never the ones you think. Remember that all the AI-safety-motivated bills were careful to minimize impact to open models, whereas the White House move will be attempting to maximize impact. Different worlds.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/www.axios.com\/2026\/07\/20\/ai-us-china-open-source-kimi\">Maria Curi<\/a>: The Trump administration is showing signs it could ban cutting-edge <a href=\"https:\/\/archive.is\/o\/xhPus\/https:\/\/www.axios.com\/2026\/07\/18\/china-ai-open-source-kimi-anthropic-openai\">Chinese AI models<\/a> \u2014 a momentous move that could lock in dominance by OpenAI and Anthropic.<\/p>\n<p>Parts of the administration have tried to implement de facto bans on foreign open-source models before, knowledgeable sources tell Axios. Last week\u2019s rise of Chinese model Kimi is reigniting those efforts.<\/p>\n<\/blockquote>\n<p>I was early on the \u2018Google is no longer in the top tier\u2019 train, but there is plenty of healthy competition that is not Chinese. Curi is taking the \u2018duopoly\u2019 and \u2018lock in\u2019 lines <a href=\"https:\/\/archive.is\/o\/xhPus\/https:\/\/x.com\/DavidSacks\/status\/2078826291638522127\">directly from David Sacks\u2019s disingenuous misreading of Dean Ball\u2019s original tweet<\/a>. OpenAI and Anthropic are not driving this.<\/p>\n<p>The White House, as per usual, is proposing to do this in a maximally blunt way.<\/p>\n<blockquote>\n<p>\u200b The Commerce Department last year considered adding multiple Chinese AI labs to its \u201cEntity List,\u201d which would effectively cut off U.S. access without a license, a source close to the administration told Axios.<\/p>\n<\/blockquote>\n<p>Dean Ball\u2019s prediction, which was also a subtle hint as to how to do it if the White House decided it needed to do it, was simply to create regulatory uncertainty, which would be sufficient to discourage big players from using foreign models in critical places, while letting startups and builders have their fun. A \u2018supply chain risk\u2019 designation would be the next step up from that.<\/p>\n<p>This proposal is something else. This is a sledgehammer.<\/p>\n<blockquote>\n<p>\u200bThe White House considered implementing an executive order saying U.S. companies could only host Chinese models if they could guarantee security and take liability if it were breached, the source added.<\/p>\n<p>\u2026 Administration officials keen on keeping regulation from stifling innovation killed all of those efforts.<\/p>\n<\/blockquote>\n<p>This is a past proposal I\u2019d heard about privately, and would be a de facto ban on the cloud providers serving those models. This would also be deeply stupid, driving business to the competition without accomplishing anything.<\/p>\n<p>I do sympathize with David Sacks that he had to keep pushing back on overreactions like this. That doesn\u2019t excuse his actions, but almost everyone in politics has troubles and crazier people of their own to deal with.<\/p>\n<blockquote>\n<p>\u200bInstead of a ban, another source familiar with government discussions described a push to highlight potential backdoors and lack of security with Chinese models, and the governance issue that brings.<\/p>\n<\/blockquote>\n<p>That\u2019s exactly the Dean Ball prediction. That option might be starting to look pretty good right around now, huh?<\/p>\n<h4 class=\"wp-block-heading\">OpenAI Employees Are Relatively Bullish On This One<\/h4>\n<p>Okay. Back to Kimi K3\u2019s actual capabilities.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/viemccoy\/status\/2077833055579078696\">@viemccoy<\/a> (OpenAI): kimi seems to be a true open-weights frontier model. compared to jailbreaking proprietary models, fine-tuning this to be a malicious coding agent will be trivial since you have the weights.<\/p>\n<p>we live in a completely different world, now.<\/p>\n<\/blockquote>\n<p>That would be the top end of potential scenarios here.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/viemccoy\/status\/2078338966097645840\">@viemccoy<\/a> (OpenAI): I was using Fable as a second eye on frontend, but Kimi K3 has completely blown me away in this regard.<\/p>\n<p>Fable is still the best at philosophy, Sol remains undefeated for any structured task, but k3 \u2026 It has a *je ne sais quoi* that American models don\u2019t have.<\/p>\n<p><a href=\"https:\/\/x.com\/tszzl\/status\/2077827974452461871\">roon<\/a> (OpenAI): the era of the chinese labs being far behind is over, Kimi is at least on par with the modern public frontier models. people have to think differently now without any competitive margin built in<\/p>\n<p>Note: in the coming days, i expect that people will find kimi k3 somewhat less practically useful than today\u2019s numbers suggest. however, its reputation will settle as an incredibly powerful model whose open weights are on the web<\/p>\n<\/blockquote>\n<p>I agree strongly with Roon\u2019s second paragraph. The first one at least toys with the jumping of the gun, also notice he only is talking about \u2018public\u2019 models.<\/p>\n<h4 class=\"wp-block-heading\">Kimi K3 Is Relatively Strongest At Typical Agentic Coding, Front End Work and 3D<\/h4>\n<p>That seems to be the word on the street.<\/p>\n<p><a href=\"https:\/\/x.com\/deanwball\/status\/2078133895766114412\">As per Dean Ball above<\/a>, it is clearly very good for most people\u2019s agentic coding, plausibly on par with models from Q1 2026. Most agentic coding is rather close to what benchmarks and training tasks measure, so you can be relatively \u2018shallow\u2019 and still impress in the day to day.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/jakehalloran1\/status\/2078184323753152766\">Jake Halloran<\/a>: capacity constrained to the point of broad uselessness but when it does work its near frontier at code stuff, especially design work and much further from the frontier (though still probably the third best lab) on non code stuff like writing and weird data knowledge<\/p>\n<\/blockquote>\n<p><a href=\"https:\/\/x.com\/TushitGargg\/status\/2078316357041688732\">Tushit runs an internal react\/frontend eval<\/a>, finds Kimi the slowest versus Opus, Sonnet and Grok 4.5 (?) but about half the cost of Opus, and all of them usually succeed, with Grok actually coming out ahead. Sounds like a saturated benchmark, but some people\u2019s real world tasks are saturated. Handling the ordinary stuff matters, too.<\/p>\n<p>Thus, some people stick with the saturated benchmarks:<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/Medo42\/status\/2078144842685657160\">Medo42<\/a>: 100% on my usual (non-agentic) coding task, but not the first open model to achieve that. Good presentation of the result and approach along with the code. Good vision, but not beating Gemini. Smells big. Runs slow.<\/p>\n<\/blockquote>\n<p>Also we\u2019ve seen a bunch of 3D stuff that looks cool.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/theo\/status\/2077942489844191370\">Theo \u2013 t3.gg<\/a>: Kimi K3 is so good at 3d stuff holy shit.<\/p>\n<p>So far just have it doing stupid threejs stuff in browser. Will have it try blender later when I have time, currently late to a ton of shit<\/p>\n<p>fwiw, fable and 5.6 both sucked hard at threejs modeling, and weren\u2019t meaningfully better in blender. Kimi is way ahead from my limited testing<\/p>\n<\/blockquote>\n<p>Here is the opposite opinion, though, reactions always vary:<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/paperclippriors\/status\/2078876022968057959\">paperclippriors<\/a>: Found little reason to use it over Fable or 5.6 for coding. It is, however, an absolute *delight* to talk to. Big model smell, very Claude like but somewhat more enthusiastic and less constrained. Feels like they have a strong base<\/p>\n<\/blockquote>\n<h4 class=\"wp-block-heading\">Reactions<\/h4>\n<p><a href=\"https:\/\/x.com\/HCSolakoglu\/status\/2077761791019393085\">Hasan Can is impressed<\/a>.<\/p>\n<p>As are others:<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/AnAcctOfAllTime\/status\/2078148249165164765\">AllTime<\/a>: You can feel it\u2019s a big model. It\u2019s smart and quite good at deduction, has impressive knowledge though the breadth might not quite be on the level of the American frontier models. Thinks too much on many problems unless instructed otherwise, and even then sometimes. Impressive!<\/p>\n<p><a href=\"https:\/\/x.com\/AivokeArt\/status\/2078171923859525929\">Aivo<\/a>: Pleasant to talk to. Pretty okay writer. I haven\u2019t done any coding with it so far.<\/p>\n<p><a href=\"https:\/\/x.com\/enolan\/status\/2078160250595578196\">Echo Nolan<\/a>: On a tough ML design problem it gave a result that is complicated and unworkable where gpt-5.6-pro gave a result that is complicated and workable and fable gave a result that is pleasantly simpler and workable.<\/p>\n<\/blockquote>\n<p>The biggest claim would be that it lives up to its benchmarks. Elanor is explicitly claiming this, although almost everyone else disagrees.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/intellectronica\/status\/2078159437458448542\">Eleanor Berger<\/a>: Very good and complete and balanced model. Impressive that they got this level of intelligence and finesse without it resorting to an endless internal monologue \u2013 it\u2019s not as efficient as Sol or even Fable, but it\u2019s also not a GLM. Great for all tasks, from coding to writing, to agentic workflows. Actually has good taste.<\/p>\n<p>This is the first chinese\/open model that feels like it belongs where the benchmarks place it. It rightly occupies the top 3 with Sol and Fable. The service quality is terrible, but hopefully once they release the weights there will be many more options, including ones that are hosted in the free world. The world has changed meaningfully with the release of this model. I will be using it a lot once it\u2019s available from more reliable servers.<\/p>\n<\/blockquote>\n<p>Others find it doing an okay job.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/xpasky\/status\/2078237443036180498\">Petr Baudis<\/a>: Finally got my first two Kimi K3 reps in! Same tasks I had Fable do in another worktree (patching @MuaddibLLM a pi-based harness).<\/p>\n<p>It did an ok-ish job, but with serious deficiencies compared to Fable. A review by Sol would equalize, though.<\/p>\n<p>On a break, it then drew a dickpic (inadverently), which was certainly a first. (When I pointed it out, it eventually \u201csaw\u201d it but I think it actually didn\u2019t, based on its own\u2026 description.)<\/p>\n<p><a href=\"https:\/\/x.com\/danielmulec\/status\/2078955776408781067\">Daniel Mulec<\/a>: It\u2019s really no fable competitor but it\u2019s worlds apart from the garbage that 2.6 used to be. Enjoying it but it\u2019s in a weird spot. It\u2019s neither good enough to replace Fable or GPT-5.6 for me nor good enough to just be an executor model for implementation plans created by GPT-5.6 either.<\/p>\n<p>Seems like with everything I do with K3, K3 always only get\u2019s there 70-90% of what I actually want and need (79-90% of each and every prompt)<\/p>\n<\/blockquote>\n<p>Inadvertent. Right. Let\u2019s see the J-space.<\/p>\n<p>Often it can do the thing.<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/MattBruenig\/status\/2079063526291755294\">Matt Bruenig<\/a>: OpenCode with Kimi K3 can execute my NLRB Research skill very well. It is quite a bit slower than Anthropic\/Claude but you can just set it off and do something else while it churns I suppose.<\/p>\n<\/blockquote>\n<p>This is the sign of a good model. That doesn\u2019t mean there is any strong reason to choose Kimi K3 to do that thing. You are not getting that large a discount.<\/p>\n<p>This seems mostly right to me, but Kimi K3 is probably good enough that there will be some areas (e.g. if the Harvey result holds) where it is at the top and gets the call:<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/wedrifid\/status\/2079115335597437202\">Cameron Taylor<\/a>: I am glad that Kimi K3 exists. It just doesn\u2019t dominate a price\/performance niche so doesn\u2019t really fit anywhere. Unlike, for example, deepseek 4, which subsidised itself into dominating a fairly-smart super cheap niche.<\/p>\n<\/blockquote>\n<h4 class=\"wp-block-heading\">Who Are You?<\/h4>\n<p>How often does Kimi K3 claim to be Claude? <a href=\"https:\/\/x.com\/RyanGreenblatt\/status\/2078663148509544589\">Not usually, but sometimes<\/a>.<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!672p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507334e3-b561-4976-a193-17a395309611_1200x600.png\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/typebulbit\/status\/2078668867052646656\">typebulb<\/a>: Also, it thinks it\u2019s Claude 1 in 10 times. This is as embarassing as it is unacceptable for model claiming SOTA status.<\/p>\n<p>I ran a cross-entropy comparison of all raw text responses from numerous models using data I already had from a benchmark I run. I leaned on Fable for the stats-know-how; certainly seems very suspicious. The results\/code are here for others to inspect: <a href=\"https:\/\/typebulb.com\/u\/lab\/you-re-relatively-right\/full\">https:\/\/typebulb.com\/u\/lab\/you-re-relatively-right\/full<\/a><\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!aN91!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6093fce-0975-495e-b1b0-8b3b4a3e96ec_1200x1089.jpeg\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<\/blockquote>\n<p><a href=\"https:\/\/x.com\/peterwildeford\/status\/2078111523994574997\">In addition to the \u2018claims to be Claude\u2019 issu<\/a>e\u2026<\/p>\n<blockquote>\n<p><a href=\"https:\/\/x.com\/HalfBoiledHero\/status\/2077864291223417195\">Sho<\/a>: We out here distilling the summaries<\/p>\n<p><a href=\"https:\/\/x.com\/Sauers_\/status\/2077845299247095868\">Sauers<\/a>: Kimi K3 appears to be trained on Claude CoT, as it follows the same pattern as Claude, e.g. ending with formatting decisions<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!7rFM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56683561-f0bd-469d-84bb-9b2ead47d974_764x196.png\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<p><a href=\"https:\/\/x.com\/BetleyJan\/status\/2078218002101596425\">Jan Betley<\/a>: Very deeply internalized inner Claude. Many have already observed that Kimi 3 often claims it\u2019s Claude (e.g. @Sauers_ ).<\/p>\n<p>We checked on the AI Bubble questions, and yes, just like Claude it claims lower probability of the bubble popping when we consider investing in Anthropic.<\/p>\n<div>\n<figure>\n<div>\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!1ETK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41266efb-2c67-4452-9385-e3581c186d4e_1200x546.jpeg\" alt=\"\"\/><\/figure>\n<\/div>\n<\/figure>\n<\/div>\n<\/blockquote>\n<h4 class=\"wp-block-heading\">How Did They Do It?<\/h4>\n<p>By \u2018it\u2019 we mean outsized benchmark gains.<\/p>\n<p><a href=\"https:\/\/x.com\/gleech\/status\/2078156758736912810\">Gavin Leech speculates<\/a>:<\/p>\n<blockquote>\n<p>Gavin Leech: \u200bMy guess of the contributions to the outsized benchmark gains:<\/p>\n<ol>\n<li>10% OOD latent gains<\/li>\n<li>25% benchmaxxing<\/li>\n<li>20% Usemaxxing and shallow generalisation<\/li>\n<li>20% cheating and reward hacking<\/li>\n<li>10% induced innovation <em>not already priced-in to the Western frontier models<\/em>\n<ol>\n<li><s>Stable LatentMoE<\/s> (priced-in)<\/li>\n<li><s>Quantile balancing<\/s> (priced-in from DeepSeek\u2019s aux-loss-free bias)<\/li>\n<li>Attention Residuals (pretty similar to Hyper-Connections)<\/li>\n<li><s>per-head muon<\/s> (obvious)<\/li>\n<li><s>gated MLA<\/s> (priced-in from Qwen)<\/li>\n<li><s>Mooncake architecture<\/s> (GPU savings rather than capability gain)<\/li>\n<li><s>Kimi Delta Attention (KDA)<\/s> (priced-in)<\/li>\n<li>SiTU<\/li>\n<li><s>Better autoresearch<\/s> (no, needs compute which they don\u2019t have)<\/li>\n<\/ol>\n<\/li>\n<li>15%: distillation off Claude\n<ol>\n<li>3% Output mimicking<\/li>\n<li>5% Their own synthetic data graded by frontiers,<\/li>\n<li>7% rejection sampling against frontier judges<\/li>\n<\/ol>\n<\/li>\n<\/ol>\n<p><a href=\"https:\/\/x.com\/teortaxesTex\/status\/2078241682923876626\">Teortaxes<\/a>: have you considered that autoresearch scales with the base of IQ and taste of human researchers though<br \/>perhaps American frontier just has washed people<\/p>\n<p>to be clear I don\u2019t argue that American researchers are low IQ but the purely substitutive reasoning about autoresearch doesn\u2019t feel adequate to me<\/p>\n<p><a href=\"https:\/\/x.com\/StatsLime\/status\/2078877554577137981\">Max Limelihood<\/a>: GPU savings ARE capability gains.<\/p>\n<\/blockquote>\n<p>GPU savings enable being larger which enables capability gains, so yes.<\/p>\n<p>I have considered and rejected Teortaxes\u2019s hypothesis. China doubtless also has lots of great talent. They absolutely can do innovative things, especially in terms of efficiency. But if your theory requires the Chinese to be better AI researchers, in general, than the Americans are, especially if this takes into account available resources and experience, I don\u2019t think that is credible.<\/p>\n<p>In general I presume Chinese models involve a lot of benchmaxxing, usemaxxing, shallow generalization, focus on relative strengths and distillation from Claude (or GPT, but in this case we can be confident it was Claude).<\/p>\n<p>Reward hacking and cheating is a live possibility but hard to assess for now.<\/p>\n<p>Kimi K3 also benefits from moving up in size.<\/p>\n<h4 class=\"wp-block-heading\">Conclusion<\/h4>\n<p>Kimi K3 is potentially the most impressive Chinese release so far in terms of pure capability. It is a very good model. My current guess is that Kimi K3 will modestly underperform its highly impressive benchmarks, but with some areas of relatively high performance where it is competitive, and with a unique style some people will enjoy. It is not close to Fable, and I do not believe it is that close to Sol.<\/p>\n<p>If its weights were released today, it would be the most capable open model. They might want to hurry, since a new Qwen is dropping soon, with the preview live (but I have seen zero reports from anyone trying it) which might well be better than K3, so chances are we will soon all have to do this over again.<\/p>\n<p>We do not know how good until the weights are released and we have more time. For now access has been spotty and limited, and there is much we do not know.<\/p>\n<p>As usual, there are some who are getting carried away, who say the latest release changes everything, that the American lead is gone, that Chinese labs are now \u2018winning,\u2019 that open models will \u2018win,\u2019 that all limits on American models are foolish now, and so on. Do not be one of those people.<\/p>\n<p>Nor should you shrink from the security and other safety concerns of releasing increasingly capable open weights models. This is not the \u2018Mythos-level open model\u2019 moment. I expect at most modest disruptions this time, and for those to occur gradually, with some small tail risk.<\/p>\n<p>But yes, absent CCP intervention to stop it, we should expect a model to cross that threshold by the end of the year. One cannot simply ignore the risks involved in that, and the American government cannot either, nor can they ignore the fact that these models are Chinese. Actions and restrictions are coming. Those who cannot accept this, or even accept people pointing this fact out, while simultaneously hyping up Kimi K3 and cheering it on, are creating a clarifying moment.<\/p>\n<\/p><\/div>\n<p><a href=\"https:\/\/thezvi.wordpress.com\/2026\/07\/20\/on-kimi-k3-its-capabilities-and-related-discontents\/?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model. Do not get carried away. Do not judge Kimi K3 only its relative strengths. In aggregate it is several months behind the closed model frontier, [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22566,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22565","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22565","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22565"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22565\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22566"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22565"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22565"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22565"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}