{"id":22681,"date":"2026-07-23T05:09:14","date_gmt":"2026-07-23T05:09:14","guid":{"rendered":"https:\/\/scannn.com\/inside-the-model-factory-eiso-kant-poolside-ai\/"},"modified":"2026-07-23T05:09:14","modified_gmt":"2026-07-23T05:09:14","slug":"inside-the-model-factory-eiso-kant-poolside-ai","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/inside-the-model-factory-eiso-kant-poolside-ai\/","title":{"rendered":"Inside the Model Factory \u2014 Eiso Kant, Poolside AI"},"content":{"rendered":"\n<div dir=\"auto\">\n<p><span>In recent months, the open vs closed, and <\/span><a href=\"https:\/\/www.latent.space\/p\/ainews-kimi-k3-28t-a50b-the-largest\">US vs China<\/a><span> discussions on model ownership and sovereign\/local AI have heated up to a fever pitch. So it is very very good news that <\/span><strong><a href=\"https:\/\/x.com\/poolsideai\/status\/2079613777343848465\">Poolside AI<\/a><\/strong><span> are finally emerging with new models, like <\/span><strong><a href=\"https:\/\/x.com\/poolsideai\/status\/2079613777343848465\">Laguna S 2.1<\/a><\/strong><span>, that are <\/span><a href=\"https:\/\/www.latent.space\/p\/ainews-thinkys-inkling-975b-a41b\">beating Thinking Machines\u2019 recent release nearly 10 times their size<\/a><span>.<\/span><\/p>\n<div class=\"captioned-image-container\">\n<figure><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!RGsJ!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a65699-ea93-4572-addd-3c46060cb6b8_1200x822.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\"><\/p>\n<div class=\"image2-inset\"><picture><source type=\"image\/webp\" srcset=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!RGsJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a65699-ea93-4572-addd-3c46060cb6b8_1200x822.png 424w, https:\/\/substackcdn.com\/image\/fetch\/$s_!RGsJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a65699-ea93-4572-addd-3c46060cb6b8_1200x822.png 848w, https:\/\/substackcdn.com\/image\/fetch\/$s_!RGsJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a65699-ea93-4572-addd-3c46060cb6b8_1200x822.png 1272w, https:\/\/substackcdn.com\/image\/fetch\/$s_!RGsJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a65699-ea93-4572-addd-3c46060cb6b8_1200x822.png 1456w\" sizes=\"100vw\"\/><\/picture><\/div>\n<p><\/a><\/figure>\n<\/div>\n<p><a href=\"https:\/\/x.com\/robert_mchardy\/status\/2059297942301782062\">Poolside\u2019s recent tech report<\/a><span> got a lot of praise due to their level of detail, and Vibhu first covered Laguna\u2019s recent technical report on our paper club:<\/span><\/p>\n<p><a href=\"https:\/\/x.com\/eisokant\/status\/2060097309396832432?s=20\" target=\"_blank\" rel=\"noopener noreferrer\" data-component-name=\"Twitter2ToDOM\" class=\"pencraft pc-display-contents pc-reset\"><\/p>\n<div data-attrs=\"{&quot;url&quot;:&quot;https:\/\/x.com\/eisokant\/status\/2060097309396832432?s=20&quot;,&quot;full_text&quot;:&quot;Loving the &lt;span class=\\&quot;tweet-fake-link\\&quot;&gt;@latentspacepod&lt;\/span&gt; breakdown of our Laguna M.1\/XS.2 Technical Report! The Latent Space paper club just did a deep dive, and their takeaways perfectly capture what we set out to build with our Model Factory. A few quotes from the video &#x1f9f5;&#x1f447; (1\/6)\\n&lt;a class=\\&quot;tweet-url\\&quot; href=\\&quot;https:\/\/youtu.be\/QLfZamyMls0\\&quot;&gt;youtu.be\/QLfZamyMls0&lt;\/a&gt;&quot;,&quot;username&quot;:&quot;eisokant&quot;,&quot;name&quot;:&quot;Eiso Kant&quot;,&quot;profile_image_url&quot;:&quot;https:\/\/pbs.substack.com\/profile_images\/1842230143965675520\/j6mVG2Py_normal.jpg&quot;,&quot;date&quot;:&quot;2026-05-28T20:34:07.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:1,&quot;retweet_count&quot;:8,&quot;like_count&quot;:47,&quot;impression_count&quot;:11267,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:false}\" class=\"pencraft pc-display-flex pc-flexDirection-column pc-gap-12 pc-padding-16 pc-reset bg-primary-zk6FDl outline-detail-vcQLyr pc-borderRadius-md sizing-border-box-DggLA4 pressable-lg-kV7yq8 font-text-qe4AeH tweet-fWkQfo twitter-embed\">\n<div class=\"pencraft pc-display-flex pc-flexDirection-row pc-gap-12 pc-alignItems-center pc-reset\">\n<div style=\"--scale:40px;\" class=\"pencraft pc-display-flex pc-width-40 pc-height-40 pc-justifyContent-center pc-alignItems-center pc-position-relative pc-reset bg-secondary-UUD3_J flex-auto-j3S2WA outline-detail-vcQLyr pc-borderRadius-full overflow-hidden-WdpwT6 sizing-border-box-DggLA4 container-TAtrWj\">\n<div style=\"--scale:40px;\" title=\"User\" class=\"pencraft pc-display-flex pc-width-40 pc-height-40 pc-justifyContent-center pc-alignItems-center pc-position-relative pc-reset bg-secondary-UUD3_J flex-auto-j3S2WA outline-detail-vcQLyr pc-borderRadius-full overflow-hidden-WdpwT6 sizing-border-box-DggLA4 container-TAtrWj\"><picture><source type=\"image\/webp\" srcset=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!EGkE!,w_40,h_40,c_fill,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F1842230143965675520%2Fj6mVG2Py.jpg 1x, https:\/\/substackcdn.com\/image\/fetch\/$s_!EGkE!,w_80,h_80,c_fill,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F1842230143965675520%2Fj6mVG2Py.jpg 2x, https:\/\/substackcdn.com\/image\/fetch\/$s_!EGkE!,w_120,h_120,c_fill,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F1842230143965675520%2Fj6mVG2Py.jpg 3x\"\/><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!EGkE!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F1842230143965675520%2Fj6mVG2Py.jpg\" alt=\"X avatar for @eisokant\" srcset=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!EGkE!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F1842230143965675520%2Fj6mVG2Py.jpg 1x, https:\/\/substackcdn.com\/image\/fetch\/$s_!EGkE!,w_80,h_80,c_fill,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F1842230143965675520%2Fj6mVG2Py.jpg 2x, https:\/\/substackcdn.com\/image\/fetch\/$s_!EGkE!,w_120,h_120,c_fill,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F1842230143965675520%2Fj6mVG2Py.jpg 3x\" width=\"40\" height=\"40\" draggable=\"false\" class=\"img-OACg1c object-fit-cover-u4ReeV pencraft pc-reset\"\/><\/picture><\/div>\n<\/div>\n<p><span class=\"pencraft pc-reset weight-semibold-uqA4FV reset-IxiVJZ\">Eiso Kant<\/span><span class=\"pencraft pc-reset color-secondary-ls1g8s reset-IxiVJZ\">@eisokant<\/span><\/p>\n<p><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" aria-hidden=\"true\" style=\"height:20px;width:20px;\" width=\"20\" height=\"20\" viewbox=\"0 0 20 20\" fill=\"var(--color-fg-primary)\" stroke-width=\"1.8\" stroke=\"#000\"><g><path stroke=\"none\" fill-rule=\"evenodd\" clip-rule=\"evenodd\" d=\"M13.2879 19.1666L8.66337 12.575L2.87405 19.1666H0.424805L7.57674 11.0258L0.424805 0.833252H6.71309L11.0717 7.04577L16.5327 0.833252H18.982L12.1619 8.59699L19.5762 19.1666H13.2879ZM16.0154 17.3083H14.3665L3.93176 2.69159H5.58092L9.7601 8.54422L10.4828 9.55981L16.0154 17.3083Z\"\/><\/g><\/svg><\/div>\n<div class=\"pencraft pc-reset line-height-20-t4M0El font-text-qe4AeH size-15-Psle70 weight-regular-mUq6Gb reset-IxiVJZ text-aFN1BV\">Loving the <span class=\"tweet-fake-link\">@latentspacepod<\/span> breakdown of our Laguna M.1\/XS.2 Technical Report! The Latent Space paper club just did a deep dive, and their takeaways perfectly capture what we set out to build with our Model Factory. A few quotes from the video  (1\/6)<br \/>\n<a class=\"tweet-url\" href=\"https:\/\/youtu.be\/QLfZamyMls0\">youtu.be\/QLfZamyMls0<\/a><\/div>\n<div class=\"pencraft pc-display-flex pc-flexDirection-column pc-gap-8 pc-reset\">\n<p><span class=\"pencraft pc-reset reset-IxiVJZ\">8:34 PM \u00b7 May 28, 2026<\/span><span class=\"pencraft pc-reset reset-IxiVJZ\"> \u00b7 <\/span><span class=\"pencraft pc-reset reset-IxiVJZ\">11.3K Views<\/span><\/p>\n<p><span class=\"pencraft pc-reset reset-IxiVJZ\">1 Reply<\/span><span class=\"pencraft pc-reset reset-IxiVJZ\"> \u00b7 <\/span><span class=\"pencraft pc-reset reset-IxiVJZ\">8 Reposts<\/span><span class=\"pencraft pc-reset reset-IxiVJZ\"> \u00b7 <\/span><span class=\"pencraft pc-reset reset-IxiVJZ\">47 Likes<\/span><\/p>\n<\/div>\n<\/div>\n<p><\/a><\/p>\n<p><span>From spending <\/span><strong>$12 million building language models<\/strong><span> for code before the world cared to <\/span><strong>creating a Model Factory<\/strong><span> that can take a model from pre-training to release in <\/span><strong>eight weeks<\/strong><span>, Eiso Kant has spent more than a decade <\/span><strong>betting that code is the path to AGI<\/strong><span>. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he <\/span><strong>would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.<\/strong><\/p>\n<div id=\"youtube2-9_0hs2sxHHo\" data-attrs=\"{&quot;videoId&quot;:&quot;9_0hs2sxHHo&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}\" data-component-name=\"Youtube2ToDOM\" class=\"youtube-wrap\">\n<p><iframe src=\"https:\/\/www.youtube-nocookie.com\/embed\/9_0hs2sxHHo?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0\" frameborder=\"0\" loading=\"lazy\" gesture=\"media\" allow=\"autoplay; fullscreen\" allowautoplay=\"true\" allowfullscreen=\"true\" width=\"728\" height=\"409\"><\/iframe><\/p>\n<\/div>\n<p><span>We go deep on <\/span><strong>Poolside\u2019s Model Factory<\/strong><span>: the engineering systems behind 10,000\u201320,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch <\/span><strong><a href=\"https:\/\/poolside.ai\/blog\/introducing-laguna-s-2-1\">Laguna S<\/a><\/strong><span>, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.<\/span><\/p>\n<p><a href=\"https:\/\/x.com\/poolsideai\/status\/2079613777343848465\" target=\"_blank\" rel=\"noopener noreferrer\" data-component-name=\"Twitter2ToDOM\" class=\"pencraft pc-display-contents pc-reset\"><\/p>\n<div data-attrs=\"{&quot;url&quot;:&quot;https:\/\/x.com\/poolsideai\/status\/2079613777343848465&quot;,&quot;full_text&quot;:&quot;Today we're releasing Laguna S 2.1, our most capable model to date.\\n\\nIt's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.\\n\\nCapable enough to hold its own against models many &quot;,&quot;username&quot;:&quot;poolsideai&quot;,&quot;name&quot;:&quot;Poolside&quot;,&quot;profile_image_url&quot;:&quot;https:\/\/pbs.substack.com\/profile_images\/2067710536024748032\/zGfDHU4Y_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-21T17:05:36.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https:\/\/substackcdn.com\/image\/fetch\/$s_!LaJi!,w_1028,c_limit,f_auto,q_auto:best,fl_progressive:steep\/l_play_button_usfui2,w_88,e_colorize:0\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F__ss-rehost__tw-video-preview-13_2079612009239175168.jpg&quot;,&quot;link_url&quot;:&quot;https:\/\/t.co\/hJr4yQ6VzA&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:160,&quot;retweet_count&quot;:320,&quot;like_count&quot;:2184,&quot;impression_count&quot;:655658,&quot;expanded_url&quot;:null,&quot;video_url&quot;:&quot;https:\/\/video.twimg.com\/amplify_video\/2079612009239175168\/vid\/avc1\/1280x720\/CX3EZIt-noYsEA3x.mp4&quot;,&quot;video_preview_media_key&quot;:&quot;13_2079612009239175168&quot;,&quot;belowTheFold&quot;:false}\" class=\"pencraft pc-display-flex pc-flexDirection-column pc-gap-12 pc-padding-16 pc-reset bg-primary-zk6FDl outline-detail-vcQLyr pc-borderRadius-md sizing-border-box-DggLA4 pressable-lg-kV7yq8 font-text-qe4AeH tweet-fWkQfo twitter-embed\">\n<div class=\"pencraft pc-display-flex pc-flexDirection-row pc-gap-12 pc-alignItems-center pc-reset\">\n<div style=\"--scale:40px;\" class=\"pencraft pc-display-flex pc-width-40 pc-height-40 pc-justifyContent-center pc-alignItems-center pc-position-relative pc-reset bg-secondary-UUD3_J flex-auto-j3S2WA outline-detail-vcQLyr pc-borderRadius-full overflow-hidden-WdpwT6 sizing-border-box-DggLA4 container-TAtrWj\">\n<div style=\"--scale:40px;\" title=\"User\" class=\"pencraft pc-display-flex pc-width-40 pc-height-40 pc-justifyContent-center pc-alignItems-center pc-position-relative pc-reset bg-secondary-UUD3_J flex-auto-j3S2WA outline-detail-vcQLyr pc-borderRadius-full overflow-hidden-WdpwT6 sizing-border-box-DggLA4 container-TAtrWj\"><picture><source type=\"image\/webp\" srcset=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!tI_E!,w_40,h_40,c_fill,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F2067710536024748032%2FzGfDHU4Y.jpg 1x, https:\/\/substackcdn.com\/image\/fetch\/$s_!tI_E!,w_80,h_80,c_fill,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F2067710536024748032%2FzGfDHU4Y.jpg 2x, https:\/\/substackcdn.com\/image\/fetch\/$s_!tI_E!,w_120,h_120,c_fill,f_webp,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F2067710536024748032%2FzGfDHU4Y.jpg 3x\"\/><img decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!tI_E!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F2067710536024748032%2FzGfDHU4Y.jpg\" alt=\"X avatar for @poolsideai\" srcset=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!tI_E!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F2067710536024748032%2FzGfDHU4Y.jpg 1x, https:\/\/substackcdn.com\/image\/fetch\/$s_!tI_E!,w_80,h_80,c_fill,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F2067710536024748032%2FzGfDHU4Y.jpg 2x, https:\/\/substackcdn.com\/image\/fetch\/$s_!tI_E!,w_120,h_120,c_fill,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fpbs.substack.com%2Fprofile_images%2F2067710536024748032%2FzGfDHU4Y.jpg 3x\" width=\"40\" height=\"40\" draggable=\"false\" class=\"img-OACg1c object-fit-cover-u4ReeV pencraft pc-reset\"\/><\/picture><\/div>\n<\/div>\n<p><span class=\"pencraft pc-reset weight-semibold-uqA4FV reset-IxiVJZ\">Poolside<\/span><span class=\"pencraft pc-reset color-secondary-ls1g8s reset-IxiVJZ\">@poolsideai<\/span><\/p>\n<p><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" aria-hidden=\"true\" style=\"height:20px;width:20px;\" width=\"20\" height=\"20\" viewbox=\"0 0 20 20\" fill=\"var(--color-fg-primary)\" stroke-width=\"1.8\" stroke=\"#000\"><g><path stroke=\"none\" fill-rule=\"evenodd\" clip-rule=\"evenodd\" d=\"M13.2879 19.1666L8.66337 12.575L2.87405 19.1666H0.424805L7.57674 11.0258L0.424805 0.833252H6.71309L11.0717 7.04577L16.5327 0.833252H18.982L12.1619 8.59699L19.5762 19.1666H13.2879ZM16.0154 17.3083H14.3665L3.93176 2.69159H5.58092L9.7601 8.54422L10.4828 9.55981L16.0154 17.3083Z\"\/><\/g><\/svg><\/div>\n<p>Today we&#8217;re releasing Laguna S 2.1, our most capable model to date.<\/p>\n<p>It&#8217;s a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.<\/p>\n<p>Capable enough to hold its own against models many <\/p>\n<div class=\"pencraft pc-display-flex pc-flexDirection-column pc-gap-8 pc-reset\">\n<p><span class=\"pencraft pc-reset reset-IxiVJZ\">5:05 PM \u00b7 Jul 21, 2026<\/span><span class=\"pencraft pc-reset reset-IxiVJZ\"> \u00b7 <\/span><span class=\"pencraft pc-reset reset-IxiVJZ\">656K Views<\/span><\/p>\n<p><span class=\"pencraft pc-reset reset-IxiVJZ\">160 Replies<\/span><span class=\"pencraft pc-reset reset-IxiVJZ\"> \u00b7 <\/span><span class=\"pencraft pc-reset reset-IxiVJZ\">320 Reposts<\/span><span class=\"pencraft pc-reset reset-IxiVJZ\"> \u00b7 <\/span><span class=\"pencraft pc-reset reset-IxiVJZ\">2.18K Likes<\/span><\/p>\n<\/div>\n<\/div>\n<p><\/a><\/p>\n<p><span>We also discuss <\/span><strong>model-harness co-design<\/strong><span>, Poolside\u2019s path from coding agents to AGI, why Eiso thinks <\/span><strong>MCP and traditional tool calls are \u201cstupid,\u201d<\/strong><span> the real economics behind frontier-model training, <\/span><strong>Poolside\u2019s $500 million raise<\/strong><span>, open-source AI, regulation, <\/span><strong>NVIDIA and TSMC\u2019s influence<\/strong><span>, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.<\/span><\/p>\n<ul>\n<li>\n<p><span>How <\/span><strong>Andrej Karpathy\u2019s RNN work<\/strong><span> inspired Eiso to start building language models for code in 2015<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why Eiso spent four years and <\/span><strong>$12 million<\/strong><span> pursuing an idea before the market cared<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why <\/span><strong>ChatGPT felt like vindication<\/strong><span> and brought Poolside back to open source<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why Eiso would prefer <\/span><strong>100 foundation model companies<\/strong><span> over an oligopoly of five<\/span><\/p>\n<\/li>\n<li>\n<p><span>The difference between releasing open weights and publishing <\/span><strong>genuinely open research<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why Poolside deliberately built a <\/span><strong>global research organization<\/strong><span> outside the Bay Area talent war<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why model building is ultimately <\/span><strong>90% engineering<\/strong><\/p>\n<\/li>\n<li>\n<p><span>The <\/span><strong>Model Factory<\/strong><span>: Poolside\u2019s end-to-end system for rapidly training and improving models<\/span><\/p>\n<\/li>\n<li>\n<p><span>How fewer than 70 researchers run roughly <\/span><strong>10,000\u201320,000 experiments each month<\/strong><\/p>\n<\/li>\n<li>\n<p><span>How Poolside moved from six-month model cycles to <\/span><strong>five- and eight-week launches<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why <\/span><strong>streaming data directly into training<\/strong><span> unlocked faster experimentation<\/span><\/p>\n<\/li>\n<li>\n<p><span>How immutable data, versioned code, and reproducibility enable <\/span><strong>rigorous model research<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why Eiso wants capable researchers to leave their labs and <\/span><strong>become Poolside\u2019s competitors<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why 95% of model building can be reduced to <\/span><strong>better data or compute efficiency<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Laguna S and why <\/span><strong>persistence, verification, and backtracking<\/strong><span> can outperform raw intelligence<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why smaller models may handle <\/span><strong>far more knowledge work<\/strong><span> than previously expected<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why reinforcement learning will move <\/span><strong>earlier into pre-training<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why next-token prediction is still failing to extract enough knowledge from <\/span><strong>the web<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why distillation and environments have become the AI industry\u2019s favorite <\/span><strong>\u201cdrugs\u201d<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why mid-training is really an early form of <\/span><strong>curriculum design<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Low-precision training, networking bottlenecks, and the next gains in <\/span><strong>compute efficiency<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Laguna S: <\/span><strong>118 billion total parameters, 8 billion active<\/strong><span>, and eight weeks from training to launch<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why model builders can often evaluate a new checkpoint within its <\/span><strong>first 30 minutes<\/strong><\/p>\n<\/li>\n<li>\n<p><strong>Model versus harness<\/strong><span>: where agent capabilities actually come from<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why Poolside sees coding and long-horizon software tasks as a <\/span><strong>path to AGI<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why Eiso thinks <\/span><strong>MCP and traditional tool calls are \u201cstupid\u201d<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why future agents will write scripts instead of choosing from <\/span><strong>dozens of predefined tools<\/strong><\/p>\n<\/li>\n<li>\n<p><span>The case for <\/span><strong>minimal harnesses, containers, and model freedom<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why Poolside is prioritizing <\/span><strong>vision<\/strong><span> but does not expect to work on audio soon<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why language may be the most compute-efficient modality for encoding <\/span><strong>knowledge and reasoning<\/strong><\/p>\n<\/li>\n<li>\n<p><span>The real cost of model development and why the final training run is <\/span><strong>anticlimactic<\/strong><\/p>\n<\/li>\n<li>\n<p><span>The story behind the Poolside name and why it represents <\/span><strong>refusing to lower ambitions<\/strong><\/p>\n<\/li>\n<li>\n<p><span>How Poolside raised <\/span><strong>$500 million<\/strong><span> while investors still questioned whether AGI was real<\/span><\/p>\n<\/li>\n<li>\n<p><span>Why intelligence could become the world\u2019s most demanded and <\/span><strong>commoditized resource<\/strong><\/p>\n<\/li>\n<li>\n<p><span>When open models may become too capable to release <\/span><strong>without restrictions<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why <\/span><strong>unilateral AI safety<\/strong><span> does not work in a globally competitive environment<\/span><\/p>\n<\/li>\n<li>\n<p><span>How regulation could accidentally lock in an <\/span><strong>oligopoly of two or three AI companies<\/strong><\/p>\n<\/li>\n<li>\n<p><span>NVIDIA, TSMC, and the hardware systems underpinning <\/span><strong>foundation-model progress<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why reinforcement-learning wall-clock time is one of Poolside\u2019s biggest <\/span><strong>bottlenecks<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why Poolside trains models from scratch instead of simply <\/span><strong>distilling larger models<\/strong><\/p>\n<\/li>\n<li>\n<p><span>How AI changes the way companies should measure <\/span><strong>engineering productivity<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Why <\/span><strong>agency<\/strong><span> may become the most important quality for employees in the AI era<\/span><\/p>\n<\/li>\n<li>\n<p><span>How leaders align high-agency people through <\/span><strong>shared goals and clear constraints<\/strong><\/p>\n<\/li>\n<li>\n<p><span>Hiring across research, post-training, pre-training, architecture, evals, and <\/span><strong>engineering at Poolside<\/strong><\/p>\n<\/li>\n<\/ul>\n<p><strong>LinkedIn:<\/strong><span> <\/span><a href=\"https:\/\/www.linkedin.com\/in\/eisokant\">https:\/\/www.linkedin.com\/in\/eisokant<\/a><\/p>\n<p><strong>X:<\/strong><span> <\/span><a href=\"https:\/\/x.com\/eisokant\">https:\/\/x.com\/eisokant<\/a><\/p>\n<p><strong>Poolside:<\/strong><span> <\/span><a href=\"https:\/\/poolside.ai\">https:\/\/poolside.ai<\/a><\/p>\n<p><strong>00:00:00<\/strong><span> Introduction<\/span><\/p>\n<p><strong>00:00:54<\/strong><span> Karpathy, RNNs, and Building Code Models Before Transformers<\/span><\/p>\n<p><strong>00:02:26<\/strong><span> The $12M Failure and ChatGPT Vindication<\/span><\/p>\n<p><strong>00:03:39<\/strong><span> Open Source and the Case for 100 Foundation Model Companies<\/span><\/p>\n<p><strong>00:09:22<\/strong><span> Open Weights, Open Research, and Poolside\u2019s Global Team<\/span><\/p>\n<p><strong>00:16:04<\/strong><span> The Model Factory: Why Model Building Is 90% Engineering<\/span><\/p>\n<p><strong>00:20:19<\/strong><span> Agents, Automated Experiments, and Early Signs of RSI<\/span><\/p>\n<p><strong>00:24:04<\/strong><span> Streaming Data, Reproducibility, and Scientific Rigor<\/span><\/p>\n<p><strong>00:30:35<\/strong><span> Creating More Foundation Model Companies<\/span><\/p>\n<p><strong>00:36:07<\/strong><span> Laguna S: Persistence vs. Raw Intelligence<\/span><\/p>\n<p><strong>00:43:01<\/strong><span> Reinventing Pre-Training, RL, and Curriculum Design<\/span><\/p>\n<p><strong>00:52:33<\/strong><span> Low-Precision Training and Squeezing More From Smaller Models<\/span><\/p>\n<p><strong>00:58:37<\/strong><span> Model Harnesses, Coding Agents, and the Path to AGI<\/span><\/p>\n<p><strong>01:09:26<\/strong><span> Why MCP and Traditional Tool Calls Are \u201cStupid\u201d<\/span><\/p>\n<p><strong>01:13:04<\/strong><span> Vision, Multimodality, and Why Language Still Matters<\/span><\/p>\n<p><strong>01:18:15<\/strong><span> Scaling Models and the Real Economics of Training<\/span><\/p>\n<p><strong>01:20:40<\/strong><span> Why Poolside Is Called Poolside and Raising $500M<\/span><\/p>\n<p><strong>01:27:37<\/strong><span> Open Models, AI Safety, and the Risk of an Oligopoly<\/span><\/p>\n<p><strong>01:33:53<\/strong><span> NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck<\/span><\/p>\n<p><strong>01:41:52<\/strong><span> Smaller Models, Distillation, Engineering Productivity, and Hiring<\/span><\/p>\n<p><strong>Swyx [00:00:00]:<\/strong><span> All right, we\u2019re here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.<\/span><\/p>\n<p><strong>Eiso Kant [00:00:08]:<\/strong><span> Thanks. Thanks for having me, guys. Good to be here.<\/span><\/p>\n<p><strong>Swyx [00:00:10]:<\/strong><span> Yeah, fresh on the plane. You texted me, you were like, \u201cHey, I\u2019m on my way to SF.\u201d I was like, \u201cYou\u2019re on a plane right now, right?\u201d Like, hey.<\/span><\/p>\n<p><strong>Eiso Kant [00:00:16]:<\/strong><span> I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let\u2019s do it.<\/span><\/p>\n<p><strong>Swyx [00:00:23]:<\/strong><span> I mean, I think the thing I would tell guests is that they don\u2019t have to prepare that much because if you\u2019re truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don\u2019t live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we\u2019re gonna talk about. But I guess, like, what got you into democratization of AI? Like, it\u2019s not obvious from your LinkedIn or something.<\/span><\/p>\n<p><strong>Eiso Kant [00:00:57]:<\/strong><span> No, it\u2019s not at all. I don\u2019t think it\u2019s obvious how I got in this space. I owe getting into this space to Andrej Karpathy.<\/span><\/p>\n<p><strong>Eiso Kant [00:01:05]:<\/strong><span> In 2015, he wrote an article called \u201cThe Unreasonable Effectiveness of Recurrent Neural Nets.\u201d<\/span><\/p>\n<p><strong>Swyx [00:01:10]:<\/strong><span> Neural Nets, yep.<\/span><\/p>\n<p><strong>Eiso Kant [00:01:11]:<\/strong><span> And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There\u2019s a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn\u2019t. and there\u2019s a little&#8211; There\u2019s an example of code a little bit further down. Yeah, so Shakespeare.<\/span><\/p>\n<p><strong>Swyx [00:01:47]:<\/strong><span> Shakespeare.<\/span><\/p>\n<p><strong>Swyx [00:01:49]:<\/strong><span> Cool<\/span><\/p>\n<p><strong>Eiso Kant [00:01:49]:<\/strong><span> And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.<\/span><\/p>\n<p><strong>Eiso Kant [00:02:29]:<\/strong><span> Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it &#8211; it wasn\u2019t obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn\u2019t obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors\u2019 money, which was a lot back then.<\/span><\/p>\n<p><strong>Swyx [00:03:18]:<\/strong><span> Yep.<\/span><\/p>\n<p><strong>Eiso Kant [00:03:19]:<\/strong><span> You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn\u2019t really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It\u2019s like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,<\/span><\/p>\n<p><strong>Eiso Kant [00:03:56]:<\/strong><span> We really had a strong point of view at the time that, like, as you\u2019re building more capable intelligence, it should be open and open source.<\/span><\/p>\n<p><strong>Eiso Kant [00:04:04]:<\/strong><span> When we started Poolside, that wasn\u2019t the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.<\/span><\/p>\n<p><strong>Eiso Kant [00:04:23]:<\/strong><span> And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, \u201c is this really gonna work?\u201d And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn\u2019t roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.<\/span><\/p>\n<p><strong>Eiso Kant [00:04:59]:<\/strong><span> And it wasn\u2019t until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.<\/span><\/p>\n<p><strong>Eiso Kant [00:05:07]:<\/strong><span> And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.<\/span><\/p>\n<p><strong>Eiso Kant [00:05:20]:<\/strong><span> But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn\u2019t, this didn\u2019t happen overnight. It was, like, a little bit we were seeing this and we\u2019re like, \u201cOkay, The world\u2019s going down a path.\u201d And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we\u2019re working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone\u2019s trying to figure things out. You\u2019d get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.<\/span><\/p>\n<p><strong>Eiso Kant [00:06:21]:<\/strong><span> And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I\u2019m a utopian fi guy. Like, and so We took a step back and said, \u201cHey, can we play a role here?\u201d Now it was easy for us to do so because we were not at the frontier.<\/span><\/p>\n<p><strong>Eiso Kant [00:06:41]:<\/strong><span> If we were at the frontier, I don\u2019t think we could have changed our mind. and I don\u2019t mean this like it\u2019s when the moment there\u2019s too much capital involved, too much expectations, you\u2019ve built up things, right? We\u2019re a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, \u201cNo, this makes sense,\u201d Even if there\u2019s big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I\u2019ll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.<\/span><\/p>\n<p><strong>Swyx [00:08:01]:<\/strong><span> Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labs<\/span><\/p>\n<p><strong>Eiso Kant [00:08:10]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Swyx [00:08:10]:<\/strong><span> That people are now calling that. And, we\u2019re, we\u2019re doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don\u2019t have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.<\/span><\/p>\n<p><strong>Eiso Kant [00:08:36]:<\/strong><span> I really hope so, right? I think we I\u2019m, I\u2019m excited about their release, and I\u2019m excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.<\/span><\/p>\n<p><strong>Swyx [00:09:03]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:09:03]:<\/strong><span> And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don\u2019t think there\u2019s any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, \u201cHey, I am most aligned and I trust most this provider for these things.\u201d<\/span><\/p>\n<p><strong>Swyx [00:09:25]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [00:09:26]:<\/strong><span> I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there\u2019s the DeepSeek of the West. Is it today? Okay, maybe it\u2019s thinking machines reflection, but there aren\u2019t many, right? So, one of the things you guys started in France, Europe, but very much now you\u2019re taking that American standpoint and more than just that, the point is the Chinese models that we see, they\u2019re not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here\u2019s a breakdown blog, paper, technical report of here\u2019s everything for state of the art to build, frontier intelligence and you\u2019re filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.<\/span><\/p>\n<p><strong>Eiso Kant [00:10:20]:<\/strong><span> No, I appreciate it. Look, I think it\u2019s, I think it\u2019s the most meaningful contribution, right? Weights are a binary. Let\u2019s call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you\u2019re doing, right? And so now there\u2019s challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it\u2019s been haunting us for quite a few years. We from day zero were an American company.<\/span><\/p>\n<p><strong>Swyx [00:10:55]:<\/strong><span> Yeah. They moved<\/span><\/p>\n<p><strong>Swyx [00:10:56]:<\/strong><span> To France.<\/span><\/p>\n<p><strong>Eiso Kant [00:10:56]:<\/strong><span> So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, \u201cWe\u2019re not gonna hire any researchers in the Bay Area. We\u2019re gonna look for talent everywhere else in the world.\u201d and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn\u2019t fully obvious yet. I think today it very much is. And we also realized that, like, some of the world\u2019s most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we\u2019re an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company\u2019s grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it\u2019s why you\u2019re seeing like the progress, I think, on our models and the cadence at which we release, is because we didn\u2019t roll out of an existing lab. Right? we didn\u2019t, we didn\u2019t have a lot of the information that\u2019s freely flowing around here at the time. We just took this point of view as like, \u201cOkay, well, let\u2019s just work the problem. Let\u2019s just go and, like, read the few papers that are out there, and let\u2019s just figure this stuff out.\u201d And we made some hilarious mistakes in model training because of that over the years<\/span><\/p>\n<p><strong>Eiso Kant [00:12:35]:<\/strong><span> Like especially in the first 12 months. there\u2019s a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, \u201cOkay, we can do this.\u201d When we first wrote our first training code base completely from scratch, it wasn\u2019t a fork of any open source. It was just like, \u201cOkay, let\u2019s build it from scratch.\u201d I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn\u2019t get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, \u201cOh, we can do things,\u201d like if we\u2019re just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there\u2019s this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that\u2019s starting to show up in results.<\/span><\/p>\n<p><strong>Swyx [00:13:52]:<\/strong><span> Just \u2018cause we probably won\u2019t revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won\u2019t tell the solution, but we\u2019<\/span><\/p>\n<p><strong>Eiso Kant [00:14:01]:<\/strong><span> So the &#8211; This &#8211; You\u2019re gonna test my memory here,<\/span><\/p>\n<p><strong>Swyx [00:14:04]:<\/strong><span> Oh, okay<\/span><\/p>\n<p><strong>Eiso Kant [00:14:04]:<\/strong><span> So but I think<\/span><\/p>\n<p><strong>Swyx [00:14:05]:<\/strong><span> Directly<\/span><\/p>\n<p><strong>Eiso Kant [00:14:05]:<\/strong><span> I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilon<\/span><\/p>\n<p><strong>Swyx [00:14:12]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [00:14:13]:<\/strong><span> Which is, right, like in the denominator<\/span><\/p>\n<p><strong>Swyx [00:14:14]:<\/strong><span> Momentum and weights. Yeah<\/span><\/p>\n<p><strong>Eiso Kant [00:14:15]:<\/strong><span> Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don\u2019t know if it was E minus four or whatever, like a high value for epsilon.<\/span><\/p>\n<p><strong>Eiso Kant [00:14:31]:<\/strong><span> And if you think about this during training, it\u2019s like a bit weird and counterintuitive that we\u2019re adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don\u2019t recall the exact bug, but it had &#8211; What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we\u2019re like, \u201cOh, no, it has to be this way. It has to have this high value of epsilon.\u201d But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you\u2019re just trying to avoid division by zero, why can\u2019t the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.<\/span><\/p>\n<p><strong>Eiso Kant [00:15:33]:<\/strong><span> Right? Like &#8211; It\u2019s such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don\u2019t even.<\/span><\/p>\n<p><strong>Swyx [00:15:45]:<\/strong><span> Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys saw<\/span><\/p>\n<p><strong>Eiso Kant [00:15:58]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Swyx [00:15:58]:<\/strong><span> Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.<\/span><\/p>\n<p><strong>Eiso Kant [00:16:08]:<\/strong><span> So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.<\/span><\/p>\n<p><strong>Eiso Kant [00:16:18]:<\/strong><span> And I think we all know it in the industry because if you look at where\u2019s every researcher spending their time, they\u2019re spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, \u201cWell, ultimately, model building is a process.\u201d You\u2019re going from raw data, right? Like training raw material, the web, et cetera. you\u2019re doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that\u2019s, far more complex than it was three years ago. then you\u2019re training a model, which is effectively a large distributed systems problem, right? Across hardware that has still&#8211; It\u2019s become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, \u201cWell, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,\u201d right? If it\u2019s your big data pipelines, if it\u2019s your crawling ingestion of the web, if it\u2019s your, large-scale distributed training, and then you\u2019ve got your reliability. And we said, \u201cWell, why don\u2019t we take some of the world\u2019s smartest distributed systems engineers that we knew and make them part of the process of research from day zero?\u201d Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it\u2019s thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.<\/span><\/p>\n<p><strong>Eiso Kant [00:18:18]:<\/strong><span> Right? Because we don\u2019t have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.<\/span><\/p>\n<p><strong>Eiso Kant [00:18:42]:<\/strong><span> And in the. And because it\u2019s such an experimental science, ultimately, in the beginning when it wasn\u2019t that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we\u2019re a small team, right? We\u2019re less than 70 researchers, another 35 engineers. and we are running, I haven\u2019t checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it\u2019s ultimately it\u2019s, it\u2019s you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we\u2019re gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we\u2019re launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we\u2019re now training. And so the model should be an artifact of someone\u2019s process. It shouldn\u2019t be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they\u2019re rolling off, and no one is really thinking about the next launch anymore. So it\u2019s just another launch, it\u2019s another launch, another rocket comes off. And that\u2019s what we\u2019re trying to do with model building.<\/span><\/p>\n<p><strong>Eiso Kant [00:20:22]:<\/strong><span> And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It\u2019s perfect for agents.<\/span><\/p>\n<p><strong>Eiso Kant [00:20:40]:<\/strong><span> Because agents are now starting to take over more and more work in our model factory.<\/span><\/p>\n<p><strong>Vibhu [00:20:43]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:20:44]:<\/strong><span> So I look at the screens when I walk, like when we\u2019re, we come together, in our monthly, we do monthly onsites, and I walk behind people\u2019s screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They\u2019re launching the jobs. They\u2019re evaluating the results that are coming back from the model runs. They are, making the changes. And we\u2019re still in the driver\u2019s seat. We\u2019re still coming up with the ideas. We\u2019re still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it\u2019s starting to become more on the architecture side as well. You\u2019re starting to see these twinklings of what RSI is gonna look like.<\/span><\/p>\n<p><strong>Eiso Kant [00:21:27]:<\/strong><span> And that\u2019s. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn\u2019t matter if it\u2019s a training like big run or if it\u2019s now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.<\/span><\/p>\n<p><strong>Eiso Kant [00:21:57]:<\/strong><span> So there\u2019s not like a cutoff 90 days before. Like no, it\u2019s like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven\u2019t had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that\u2019s usually there\u2019s a little bit of intervention, but that\u2019s always within like call periods, right? Not on call. And I think that\u2019s starting to now compound. So the model we\u2019re releasing now, I love it. It\u2019s amazing, but we\u2019re already onto the next one. and I think that\u2019s the way it should be.<\/span><\/p>\n<p><strong>Vibhu [00:22:50]:<\/strong><span> Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, \u201cOkay, new model\u2019s dropped. Haven\u2019t heard about it.\u201d We were<\/span><\/p>\n<p><strong>Eiso Kant [00:23:02]:<\/strong><span> Yeah, we\u2019re very used to doing this every few months.<\/span><\/p>\n<p><strong>Vibhu [00:23:03]:<\/strong><span> We\u2019re, we\u2019re very much like, \u201c okay, look, it\u2019s like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it\u2019s a very cool paper on what goes into building.\u201d And then we hit this page, right? Like literally page two of tech report is, \u201cThis process allowed us to build the small model from scratch to delivery within five weeks applying the lessons\u201d. And then I\u2019m like, oh, this paper is not about here\u2019s a tech report of benchmarks and here\u2019s how many tokens it was trained on. Like for people that wanna dive more from what we\u2019re not gonna discuss on the podcast, it\u2019s all laid out here, right? From<\/span><\/p>\n<p><strong>Eiso Kant [00:23:38]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Vibhu [00:23:39]:<\/strong><span> Custom software that agents can use to interface with training code, training data.<\/span><\/p>\n<p><strong>Eiso Kant [00:23:45]:<\/strong><span> Yeah. Well, link the paper correctly, so yeah.<\/span><\/p>\n<p><strong>Vibhu [00:23:47]:<\/strong><span> Yeah. All that stuff. read the paper here, but,<\/span><\/p>\n<p><strong>Eiso Kant [00:23:50]:<\/strong><span> But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I\u2019ll just call out that Dagster just got bought by a Prefect.<\/span><\/p>\n<p><strong>Vibhu [00:24:01]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:24:01]:<\/strong><span> Isn\u2019t it fun? But yes, I\u2019m very familiar with Dagster. just anything where like they trigger some story.<\/span><\/p>\n<p><strong>Vibhu [00:24:07]:<\/strong><span> So, well, I would say, well, experiments code\u2019s obvious, but I think one of my favorite things is, I don\u2019t know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.<\/span><\/p>\n<p><strong>Vibhu [00:24:30]:<\/strong><span> And we looked at this like three years ago and we were like That makes no sense<\/span><\/p>\n<p><strong>Eiso Kant [00:24:36]:<\/strong><span> You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you\u2019ve got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren\u2019t we streaming data into training? Right? Something that\u2019s very common and like just basic<\/span><\/p>\n<p><strong>Vibhu [00:25:00]:<\/strong><span> Like just in time<\/span><\/p>\n<p><strong>Eiso Kant [00:25:01]:<\/strong><span> Just in time, like good computer science like principle. And that was one of the first things that I think unlocked &#8211; the model factory. Because the moment you start thinking about, well, a training job, it doesn\u2019t matter if it\u2019s a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it\u2019s not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we\u2019ve got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.<\/span><\/p>\n<p><strong>Vibhu [00:25:36]:<\/strong><span> But when you say AWS, it\u2019s not actual AWS, it\u2019s your internal AWS.<\/span><\/p>\n<p><strong>Eiso Kant [00:25:39]:<\/strong><span> It\u2019s our internal&#8211; No, it\u2019s our internal like just running like our infrastructure<\/span><\/p>\n<p><strong>Vibhu [00:25:42]:<\/strong><span> Site web services<\/span><\/p>\n<p><strong>Eiso Kant [00:25:43]:<\/strong><span> Exactly. Our stuff running on like an AWS account or on like any hardware, right?<\/span><\/p>\n<p><strong>Vibhu [00:25:47]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:25:48]:<\/strong><span> And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don\u2019t have to wait for the whole data set to materialize.<\/span><\/p>\n<p><strong>Eiso Kant [00:26:00]:<\/strong><span> You now all of a sudden when you\u2019re running data experiments about mixing data, it\u2019s a config. Because you\u2019ve got these data sources that are coming in, and you just &#8211; we have this service called Blender that\u2019s in the report, where we then say, \u201cOkay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,\u201d and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath&#8211; So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.<\/span><\/p>\n<p><strong>Vibhu [00:26:47]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:26:48]:<\/strong><span> And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn\u2019t understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a &#8211; We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.<\/span><\/p>\n<p><strong>Vibhu [00:27:08]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:27:09]:<\/strong><span> And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who\u2019d been working on language models since like the early 2020s, I think brought that into the company of like, \u201cHey, we wanna have even more rigor.\u201d And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start \u201cOkay, every experiment is truly an ablation. We truly need to understand it.\u201d And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there\u2019s just fun stuff like, and<\/span><\/p>\n<p><strong>Vibhu [00:28:16]:<\/strong><span> Yeah, a lot of it\u2019s fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There\u2019s like a random small paragraph in here where it\u2019s just like, \u201cOh yeah, training data, we have some, we have an auto mixer.\u201d it trains eight small models, scales them up, picks the training data set. We don\u2019t even need to look at it. I\u2019m like, \u201cWow, a lot of engineering rigor there.\u201d And there\u2019s just, there\u2019s just a lot in here.<\/span><\/p>\n<p><strong>Eiso Kant [00:28:40]:<\/strong><span> Yeah, and it\u2019- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven\u2019t earned the right for yet for a long time. So you earn the right to spend time, publishing research once you\u2019re at the frontier, because until then, you\u2019re catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn\u2019t allow us to catch up. But in this case, we said, \u201cOkay, we\u2019re gonna give ourselves.\u201d I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it\u2019s easy to like put it out. And so there\u2019s so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you\u2019ll see us like be way more proactive, and just trying to keep dropping some of those like things that we\u2019ve learned along the way that can help others like speed up.<\/span><\/p>\n<p><strong>Vibhu [00:29:40]:<\/strong><span> Which is the other cool side of this, right? It\u2019s, it\u2019s not like, back to your point, it\u2019s not just here\u2019s the benchmarks of our training. If you want to replicate, here\u2019s experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here\u2019s your system for how to do it? So it\u2019s, it\u2019s really like promoting<\/span><\/p>\n<p><strong>Eiso Kant [00:29:59]:<\/strong><span> No, thank you<\/span><\/p>\n<p><strong>Vibhu [00:29:59]:<\/strong><span> Other people can do the same.<\/span><\/p>\n<p><strong>Eiso Kant [00:30:00]:<\/strong><span> And by the way, I also wanna make clear, right, we have been incredible&#8211; Like we\u2019ve taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it\u2019s important that there\u2019s, from every country and every culture and background, including like Western companies like us, there\u2019s different models that come out that people can choose to trust. But I think we do have to give credit where credit\u2019s due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you\u2019re on the receiving end of something coming to you, I think it\u2019s, you also have an obligation to give back.<\/span><\/p>\n<p><strong>Swyx [00:30:39]:<\/strong><span> Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.<\/span><\/p>\n<p><strong>Eiso Kant [00:30:44]:<\/strong><span> That\u2019s a good question.<\/span><\/p>\n<p><strong>Swyx [00:30:45]:<\/strong><span> Moaan obviously for Therapsi. Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:30:48]:<\/strong><span> Yeah, look, I think, I think obviously everyone\u2019s been talking about Zhipu lately, with 5.2. I think what most people don\u2019t realize is when they started.<\/span><\/p>\n<p><strong>Swyx [00:30:59]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:30:59]:<\/strong><span> Right? They started years before ChatGPT.<\/span><\/p>\n<p><strong>Swyx [00:31:02]:<\/strong><span> They just rebranded. Yeah<\/span><\/p>\n<p><strong>Eiso Kant [00:31:03]:<\/strong><span> And so, I\u2019ve like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn\u2019t the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we\u2019d call, machine learning on code with some of these models. we would&#8211; people would just laugh at us, like they\u2019d be like, \u201cThis makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?\u201d And so I would say they\u2019re probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we\u2019re, that we\u2019re now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It\u2019s no longer counted in months or years. But this stuff\u2019s hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don\u2019t do so, we\u2019re, we\u2019ve got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my current<\/span><\/p>\n<p><strong>Eiso Kant [00:32:36]:<\/strong><span> Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.<\/span><\/p>\n<p><strong>Eiso Kant [00:32:45]:<\/strong><span> Like start another foundation model company because I think we need it. I think otherwise we\u2019re not gonna be in the world where, I don\u2019t want to just be the fifth or the sixth company that wins. I wanna look at a world where there\u2019s lots of choice.<\/span><\/p>\n<p><strong>Vibhu [00:32:57]:<\/strong><span> What else do people not see in starting a foundation model? it\u2019s, there\u2019s a lot of compute, there\u2019s a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there\u2019s a lot there, right? That\u2019s,<\/span><\/p>\n<p><strong>Eiso Kant [00:33:10]:<\/strong><span> Well, look, it\u2019s, I in turn&#8211; this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people\u2019s minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you\u2019re just doing two things. You\u2019re improving data or you\u2019re improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we\u2019re generating new data, we\u2019re improving data. and the only way to do that is to look at the data, right? That\u2019s a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They\u2019re bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably&#8211; if you had infinite compute, you\u2019d be at AGI probably already tomorrow.<\/span><\/p>\n<p><strong>Eiso Kant [00:34:12]:<\/strong><span> Right? Like it\u2019s not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that\u2019s, those are the main things. And to just realize that this is engineering. I think it\u2019s become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you\u2019re doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don\u2019t get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we\u2019re doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that\u2019s, it feels far for people to do so. But I\u2019ve seen in our own company, we\u2019ve seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would\u2019ve not been what I think most people assumed was possible, a couple of years ago.<\/span><\/p>\n<p><strong>Swyx [00:35:46]:<\/strong><span> Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you\u2019re not only using your own models. There\u2019s no way. But like, what\u2019s that percentage over time?<\/span><\/p>\n<p><strong>Eiso Kant [00:36:10]:<\/strong><span> This is the first model that we\u2019re releasing that is starting to meaningfully contribute to our own work. It\u2019s not a it\u2019s not state-art model yet. Fable and other, they\u2019re, they\u2019re very capable models, but Laguna S Is really interesting. I\u2019m gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it\u2019s been burned in my brain ever since because the Laguna S model, as you\u2019ll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it\u2019s just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erd\u0151s 397 independently. It\u2019s able to do complex programming tasks. It\u2019s able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it\u2019s, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I\u2019ve probably spent eight to ten hours a day with this model for the last ten days.<\/span><\/p>\n<p><strong>Eiso Kant [00:38:05]:<\/strong><span> I\u2019m not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.<\/span><\/p>\n<p><strong>Eiso Kant [00:38:12]:<\/strong><span> And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there\u2019s intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you\u2019re wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we\u2019re using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.<\/span><\/p>\n<p><strong>Eiso Kant [00:39:00]:<\/strong><span> As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.<\/span><\/p>\n<p><strong>Eiso Kant [00:39:08]:<\/strong><span> And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it\u2019s something we clearly see that as models get larger and more capable, they\u2019re able to pull more ideas and threads together that a smaller model wouldn\u2019t be able to.<\/span><\/p>\n<p><strong>Eiso Kant [00:39:36]:<\/strong><span> And we\u2019re starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I\u2019m a software developer at heart first and foremost probably, although I probably can\u2019t say it that much anymore as I don\u2019t write production code in years, is that what makes us good is our persistence. It\u2019s our ability to encounter a problem and backtrack and say, \u201cI need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.\u201d But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it\u2019s because of the behaviors. And so now the question I have, and I don\u2019t have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.<\/span><\/p>\n<p><strong>Eiso Kant [00:40:52]:<\/strong><span> So I know that at the very limit, I\u2019m not gonna use the world\u2019s largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I\u2019m starting to size down for certain tasks.<\/span><\/p>\n<p><strong>Eiso Kant [00:41:07]:<\/strong><span> So it means that there is an optimal. It means there\u2019s some curve that goes as we go up to model size for knowledge work, at some point we\u2019re at the peak, and after that, the return on investment of using a bigger model, just doesn\u2019t make sense.<\/span><\/p>\n<p><strong>Eiso Kant [00:41:22]:<\/strong><span> Now, I think the question is, before I would have thought that peak was extremely very far away.<\/span><\/p>\n<p><strong>Eiso Kant [00:41:30]:<\/strong><span> This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I\u2019m no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It\u2019s an argument that open source can win and, like, succeed in this world. And now it\u2019s of course a self-serving argument and it\u2019s a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I\u2019m against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we\u2019re gonna be king of open source small models. I think that\u2019s, It\u2019s a out. It\u2019s trying to be king of your own kingdom, but not realizing what the rest of the world\u2019s doing. All of us rather use a smarter, faster, more model. It\u2019s a sign of hope. And so I don\u2019t wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.<\/span><\/p>\n<p><strong>Vibhu [00:43:03]:<\/strong><span> Is that mostly post-training? Like<\/span><\/p>\n<p><strong>Eiso Kant [00:43:05]:<\/strong><span> Yes<\/span><\/p>\n<p><strong>Vibhu [00:43:05]:<\/strong><span> Right.<\/span><\/p>\n<p><strong>Eiso Kant [00:43:06]:<\/strong><span> It\u2019s entirely post-training.<\/span><\/p>\n<p><strong>Vibhu [00:43:08]:<\/strong><span> Are we done improving anything on training? Is, like, training done?<\/span><\/p>\n<p><strong>Eiso Kant [00:43:12]:<\/strong><span> No.<\/span><\/p>\n<p><strong>Vibhu [00:43:12]:<\/strong><span> Okay.<\/span><\/p>\n<p><strong>Eiso Kant [00:43:13]:<\/strong><span> So<\/span><\/p>\n<p><strong>Vibhu [00:43:13]:<\/strong><span> I just wanted to cover training, and then we go post-training<\/span><\/p>\n<p><strong>Eiso Kant [00:43:15]:<\/strong><span> Training is not done. I mean, look, there\u2019s a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That\u2019- but those are ultimately, engineering challenges.<\/span><\/p>\n<p><strong>Eiso Kant [00:43:31]:<\/strong><span> I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.<\/span><\/p>\n<p><strong>Vibhu [00:43:42]:<\/strong><span> Yeah, training.<\/span><\/p>\n<p><strong>Eiso Kant [00:43:44]:<\/strong><span> Not even training. Like training today, right, is, like if you look at &#8211; So we\u2019ve been working on this for years already. and I think the best&#8211; I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that\u2019s the web. and the web, I think we could arguably say probably has The totality of humanity\u2019s knowledge somewhere encoded in different places. It\u2019s a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.<\/span><\/p>\n<p><strong>Eiso Kant [00:44:39]:<\/strong><span> And so now What we are trying to figure out, and have been doing a lot of work on, and it\u2019s a place where maybe not as open as we\u2019re on other things, but we will become more over time. we\u2019ve been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there\u2019s a huge amount of gold to be found there. I think we are right now in, we\u2019ve got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they\u2019re great, and they make us feel good, and they make the models better, and like we\u2019re all addicted to them, and we\u2019ll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.<\/span><\/p>\n<p><strong>Eiso Kant [00:45:33]:<\/strong><span> I think just next token prediction during training is not enough.<\/span><\/p>\n<p><strong>Eiso Kant [00:45:36]:<\/strong><span> And<\/span><\/p>\n<p><strong>Vibhu [00:45:38]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [00:45:38]:<\/strong><span> I think we\u2019ll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think &#8211; the whole world knows how to do this now. I think we\u2019re, we\u2019re scaling it up. Everyone is. But I wonder if we need to go as far as we\u2019re going today with environments. I\u2019m not sure yet<\/span><\/p>\n<p><strong>Vibhu [00:46:01]:<\/strong><span> You mean we\u2019re going too far?<\/span><\/p>\n<p><strong>Eiso Kant [00:46:02]:<\/strong><span> I\u2019m, I\u2019m not sure if the path to AGI is just<\/span><\/p>\n<p><strong>Vibhu [00:46:06]:<\/strong><span> Is more environment<\/span><\/p>\n<p><strong>Eiso Kant [00:46:07]:<\/strong><span> More environments.<\/span><\/p>\n<p><strong>Vibhu [00:46:08]:<\/strong><span> It seems like a never-ending, \u201cOkay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?\u201d<\/span><\/p>\n<p><strong>Eiso Kant [00:46:19]:<\/strong><span> I think there is, I think there\u2019s an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we\u2019re able to get so much more out of it. It just changes a little bit the way you think about intelligence.<\/span><\/p>\n<p><strong>Vibhu [00:46:40]:<\/strong><span> Yeah. The analogy people draw often is the RL phase is where you don\u2019t learn as much new knowledge. You shift<\/span><\/p>\n<p><strong>Eiso Kant [00:46:46]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [00:46:46]:<\/strong><span> Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still just<\/span><\/p>\n<p><strong>Eiso Kant [00:47:00]:<\/strong><span> It\u2019s just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it\u2019s effectively just like,<\/span><\/p>\n<p><strong>Vibhu [00:47:06]:<\/strong><span> Second phase<\/span><\/p>\n<p><strong>Eiso Kant [00:47:07]:<\/strong><span> It\u2019s the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you\u2019d want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I\u2019m pretty sure that you\u2019ll start to see people talking soon about some other term, and there\u2019s two or &#8211; \u2018cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we\u2019re doing is we\u2019re trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It\u2019- this is, I think, a thing where&#8211; that we really try to avoid with the model factory is like Training exists because there\u2019s a training team now, right? There\u2019s people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don\u2019t, you have infinite stages. Now, we\u2019re not there. Compute\u2019s not there. Organization design is not there for it yet. but I think we\u2019ll get there. we\u2019ll look back on a couple of years and be like, \u201cOh my God, it was so cute that we did our training data like this in such a like na\u00efve way. Like we barely ordered it. We didn\u2019t really do a good job at like<\/span><\/p>\n<p><strong>Vibhu [00:48:48]:<\/strong><span> The building that curriculum will get you that in the industry.<\/span><\/p>\n<p><strong>Eiso Kant [00:48:51]:<\/strong><span> And I\u2019ll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don\u2019t have the answers, but you have some ideas.<\/span><\/p>\n<p><strong>Vibhu [00:49:02]:<\/strong><span> We have some ideas. We\u2019re not ready to talk about it yet.<\/span><\/p>\n<p><strong>Eiso Kant [00:49:05]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [00:49:05]:<\/strong><span> We\u2019ve been working on them for years, and I think that\u2019s the one thing that\u2019s also like you asked earlier about, like what\u2019s not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe works<\/span><\/p>\n<p><strong>Eiso Kant [00:49:19]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [00:49:19]:<\/strong><span> Versus like your, my crazy<\/span><\/p>\n<p><strong>Eiso Kant [00:49:22]:<\/strong><span> Pure research<\/span><\/p>\n<p><strong>Vibhu [00:49:22]:<\/strong><span> Breakthrough.<\/span><\/p>\n<p><strong>Eiso Kant [00:49:22]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [00:49:22]:<\/strong><span> Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.<\/span><\/p>\n<p><strong>Eiso Kant [00:49:31]:<\/strong><span> I mean, so like, this is a nice way. I was gonna bring up auto research at some point<\/span><\/p>\n<p><strong>Vibhu [00:49:35]:<\/strong><span> Yes<\/span><\/p>\n<p><strong>Eiso Kant [00:49:35]:<\/strong><span> As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.<\/span><\/p>\n<p><strong>Vibhu [00:49:47]:<\/strong><span> Man, it\u2019s also<\/span><\/p>\n<p><strong>Eiso Kant [00:49:48]:<\/strong><span> Like what you\u2019re looking for. You\u2019re looking for loss curves like that, like<\/span><\/p>\n<p><strong>Vibhu [00:49:51]:<\/strong><span> It\u2019s also a thing people take bets on, right? When you say more Neo labs, you\u2019re doing a version of we\u2019ll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you\u2019re right. It\u2019s all, compute efficiency, and that\u2019s the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They\u2019re not just<\/span><\/p>\n<p><strong>Eiso Kant [00:50:19]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Vibhu [00:50:19]:<\/strong><span> They\u2019re like, 99% not balancing, here\u2019s the vanilla and scale up. They\u2019re 99% on, here\u2019s novel research that\u2019ll change everything.<\/span><\/p>\n<p><strong>Eiso Kant [00:50:27]:<\/strong><span> And I think, Luke, I think you. It depends when you started as well, right?<\/span><\/p>\n<p><strong>Vibhu [00:50:30]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:50:30]:<\/strong><span> When we started, like the novel thing we did was reinforcement learning on code. No long- that\u2019s no longer novel by far, but we were like, &#8211; that\u2019s where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that\u2019s different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn\u2019t<\/span><\/p>\n<p><strong>Vibhu [00:50:53]:<\/strong><span> It\u2019s cool. this was like your original 2023 blog<\/span><\/p>\n<p><strong>Eiso Kant [00:50:57]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Vibhu [00:50:57]:<\/strong><span> Of purpose.<\/span><\/p>\n<p><strong>Eiso Kant [00:50:58]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [00:50:59]:<\/strong><span> And like you do lay it all out here.<\/span><\/p>\n<p><strong>Eiso Kant [00:51:01]:<\/strong><span> We laid<\/span><\/p>\n<p><strong>Vibhu [00:51:01]:<\/strong><span> The blog is pretty underrated, right? The whole RL on code was very early on.<\/span><\/p>\n<p><strong>Eiso Kant [00:51:06]:<\/strong><span> Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.<\/span><\/p>\n<p><strong>Vibhu [00:51:23]:<\/strong><span> Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.<\/span><\/p>\n<p><strong>Eiso Kant [00:51:35]:<\/strong><span> When we<\/span><\/p>\n<p><strong>Vibhu [00:51:35]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [00:51:36]:<\/strong><span> The date on this thing is wrong. When we published this, it was April 2023. I think this was just<\/span><\/p>\n<p><strong>Vibhu [00:51:42]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [00:51:42]:<\/strong><span> Happened on a migration, probably found it on archive.org.<\/span><\/p>\n<p><strong>Vibhu [00:51:45]:<\/strong><span> Mistral.<\/span><\/p>\n<p><strong>Eiso Kant [00:51:46]:<\/strong><span> Mistral had started, we started on the same month, right?<\/span><\/p>\n<p><strong>Vibhu [00:51:49]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:51:49]:<\/strong><span> So this wasn\u2019t even, there was only, I think, Llama out at the time<\/span><\/p>\n<p><strong>Vibhu [00:51:52]:<\/strong><span> Snell<\/span><\/p>\n<p><strong>Eiso Kant [00:51:52]:<\/strong><span> And that\u2019s it, right? And so, but I agree. I think we want, We want as many diversity of ideas, and I do think if you\u2019re starting today, you want something that gives you an edge, right? and what I do think we sometimes over.<\/span><\/p>\n<p><strong>Eiso Kant [00:52:13]:<\/strong><span> I think every archit- like at the limit, every architecture works. An RNN works, it\u2019s just not compute efficient, right? Like, say if you had infinite compute, you could probably just, take a basic RNN from back in the day, and you could get pretty far.<\/span><\/p>\n<p><strong>Eiso Kant [00:52:27]:<\/strong><span> Now there have been, meaningful breakthroughs, attention, other things that are there. but I think we\u2019re still, we\u2019re still very early in figuring these out. The things I\u2019m most excited about, I\u2019m most excited about people doing extremely low precision training, right? So like the ternary stuff that we\u2019re seeing, and it<\/span><\/p>\n<p><strong>Vibhu [00:52:47]:<\/strong><span> Oh my God<\/span><\/p>\n<p><strong>Eiso Kant [00:52:47]:<\/strong><span> Very cool. The Bonsai stuff yesterday was super cool to see. I think that if you can find tweets from me going back to 2023, which is like the notion of like, well, it\u2019s an obvious trade-off. Bigger model, lower precision equals, smaller model with higher precision, by definition, right? It\u2019s just what is, like how does that play out, right? What\u2019s the actual size limit? So you now have companies that are trying to figure that out, but those are the things that can change our industry if they\u2019re done right.<\/span><\/p>\n<p><strong>Vibhu [00:53:14]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:53:14]:<\/strong><span> Because ultimately, like our bottleneck on compute is a MatMul bottleneck, and a networking bottleneck, and the moment you start doing those things. So I\u2019m excited about that. We\u2019re not doing &#8211; I mean, we\u2019re doing the usual, like, Laguna S was trained in FP8. only thing that in this run I have to admit that wasn\u2019t FP8 was the all to all in the new run we just started yesterday. The FP8 was all to all. That was just like cut off date, like, oh, we\u2019re not perfectly comfortable wanting to do it. you\u2019ve got amazing work by Nemotron and NVFP4 training. Like, I think it\u2019s underrated what they\u2019ve done there. I\u2019m excited to get to NVFP4 training. doesn\u2019t make sense yet \u2018cause we\u2019re still training on Hoppers, right? We\u2019re like relatively small. We\u2019re 10K H200 cluster company right now. We\u2019ll be scaling to a lot more soon, but, and really a lot more if someone is thinking about applying for a job. but like the. Yes, I think it\u2019s, there\u2019s so much more juice to squeeze out of this, and hopefully Laguna S shows people that a model at this size can get a lot more and we did this thing in eight weeks. We think there\u2019s a lot more juice to squeeze out at any model size. we\u2019re now scaling up because it\u2019s the most optimal thing to do for us as a company. But if I had infinite time, I would love to push more the capabilities at other model sizes.<\/span><\/p>\n<p><strong>Vibhu [00:54:34]:<\/strong><span> I don\u2019t think we\u2019ve properly announced what your new size is. So we have XS, which was 30B-ish.<\/span><\/p>\n<p><strong>Eiso Kant [00:54:41]:<\/strong><span> Yep.<\/span><\/p>\n<p><strong>Vibhu [00:54:41]:<\/strong><span> Old medium was 200B, which is gonna be deprecated<\/span><\/p>\n<p><strong>Eiso Kant [00:54:45]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Vibhu [00:54:45]:<\/strong><span> It seems. So new Laguna Small<\/span><\/p>\n<p><strong>Eiso Kant [00:54:48]:<\/strong><span> So Laguna S, Laguna Small, 118 billion total parameters, 8B active, so very sparse. It\u2019s a scale-up of the XS architecture. It\u2019s the classic, or call it classic these days, like three-one ratio of sliding window attention to global attention. It\u2019s just, it\u2019s a nice size, for a couple of reasons. One, it\u2019s just very cost efficient. For us, it was a good way to &#8211; We wanted to get our progress out quickly. One of the things that we\u2019ve seen is that it\u2019s a balance inside a foundation model company between focus on releasing and shipping And, like, your new novel research. But with the model factory, we are able to, like, treat the release of a model as less of a time investment from the team because it\u2019s just, oh, at this moment in time, do the training run, done, apply the latest post-training. And so this is, I think, a nice weight class. It\u2019s one that also will fit on a DGX Spark, which, I have a small, like, soft spot for. I love having that little thing, like, run a good model.<\/span><\/p>\n<p><strong>Swyx [00:55:52]:<\/strong><span> Yeah, we covered it on this pod, GTC last year.<\/span><\/p>\n<p><strong>Eiso Kant [00:55:54]:<\/strong><span> Nice.<\/span><\/p>\n<p><strong>Swyx [00:55:55]:<\/strong><span> I think a OSS 120B was the first because it\u2019s a large single GPU, which was the H100, right?<\/span><\/p>\n<p><strong>Eiso Kant [00:56:02]:<\/strong><span> Exactly.<\/span><\/p>\n<p><strong>Swyx [00:56:02]:<\/strong><span> Rent one H100, now you\u2019ve got 128 gig Macs, Mac Minis, Sparks. It\u2019s, it\u2019s the home sweet spot.<\/span><\/p>\n<p><strong>Eiso Kant [00:56:10]:<\/strong><span> But I think what I\u2019m most excited about is that this model hopefully shows people what is possible in this size because, when you\u2019ll look at the benchmarks and start using it, you\u2019ll realize that we are outperforming models two or three times their size.<\/span><\/p>\n<p><strong>Swyx [00:56:24]:<\/strong><span> Yeah, and they think&#8211; So for example, today\u2019s Thinky model is like a trillion params.<\/span><\/p>\n<p><strong>Eiso Kant [00:56:28]:<\/strong><span> So yeah, exactly. And look, and by the way, I\u2019m excited about&#8211; I have&#8211; It just came out, so for those of you who are listening to this at like, I saw it on my phone<\/span><\/p>\n<p><strong>Swyx [00:56:34]:<\/strong><span> If you\u2019re, if you\u2019re listening<\/span><\/p>\n<p><strong>Eiso Kant [00:56:34]:<\/strong><span> If you\u2019re humming<\/span><\/p>\n<p><strong>Swyx [00:56:35]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [00:56:35]:<\/strong><span> Like two seconds, so I haven\u2019t even had a chance to read the post.<\/span><\/p>\n<p><strong>Swyx [00:56:39]:<\/strong><span> But somehow you are, not only you\u2019re, you\u2019re better than Thinky, which is like one of those benchmarks, but also, like, on certain benchmarks, like the \u03c4-bench one, like you\u2019re state-art.<\/span><\/p>\n<p><strong>Eiso Kant [00:56:51]:<\/strong><span> We\u2019- Look, we\u2019re doing, I\u2019m not sure if we\u2019re state-art on I mean, 3 banking, I haven\u2019t checked where we sit on the leaderboard. but I think we are, within our weight class, I feel very comfortable to say, and even in some weight classes twice larger, that we are probably state-art. I also want to caveat this, like, best model still in the world right now is definitely, give me a Fable, give me a 5.6. To your point earlier, we also use other models.<\/span><\/p>\n<p><strong>Swyx [00:57:15]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Swyx [00:57:15]:<\/strong><span> I think the, so the interesting thing you mentioned earlier is you\u2019re starting to shift a lot of your actual usage to it, right? Benchmarks are like<\/span><\/p>\n<p><strong>Eiso Kant [00:57:21]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Swyx [00:57:22]:<\/strong><span> They\u2019re good to compare, but they\u2019re not super realistic. It\u2019<\/span><\/p>\n<p><strong>Eiso Kant [00:57:24]:<\/strong><span> They have to, right? This is how they\u2019re gonna dog food benchmarking.<\/span><\/p>\n<p><strong>Eiso Kant [00:57:27]:<\/strong><span> No, you have to. Like, you have to use your own models, and you have to have your own internal evals and benchmarks. And what the funny thing is, like within first 30 minutes of a new checkpoint coming out that\u2019s, the first post-train after a train, you yourself can feel in the first 30 minutes of where this model\u2019s gonna be. Like, you don\u2019t know exactly, but like when this one came out, we were like, \u201cOh,\u201d like, \u201cthis is different.\u201d Like, and I think that\u2019s, I think that\u2019s the best example. but it\u2019s a little bit like your kids. I don\u2019t have kids, but parents, like, they see their kid and it\u2019s perfect and they love it, and then like, they don\u2019t see all the rough edges. You always get that when you build your own model. It\u2019s the most fun part is that you, like, you love a little bit every model that you do. We try to say this thing constantly, it\u2019s like, \u201cIt\u2019s the worst model we\u2019ll ever train.\u201d And so I know the team now is like already onto<\/span><\/p>\n<p><strong>Swyx [00:58:18]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [00:58:19]:<\/strong><span> The next one, as it should be, because this is a race. and this model is a moment in time that hopefully shows people that we are serious about this race, that we wanna work really hard at it, that we want feedback, right? Where is it good? Where is it not? Like, one of the nice things about having your models out in open weight and out in the world is that you get a lot of feedback.<\/span><\/p>\n<p><strong>Swyx [00:58:40]:<\/strong><span> How do you think about building it with like, working with a harness, right? So OpenCode, Codex, you have your own pool CLI tool. getting people to use it, the design of model harness, training it in.<\/span><\/p>\n<p><strong>Eiso Kant [00:58:54]:<\/strong><span> So you need to do some multi-harness training. Like if you, especially at these smaller sizes, like you wanna do a little bit of multi-harness training for these models to just get the right. And it\u2019s very little. Like, you don\u2019t need a lot, but it\u2019s just like to get the right behaviors that you see in your harness transferring to the harness that, like, you, other people might use it in. we internally have been just calling this polishing, which is like you\u2019ve got your model and you do a little bit of polishing so that, like, it\u2019s able to work well in other harnesses as it is in your own.<\/span><\/p>\n<p><strong>Eiso Kant [00:59:24]:<\/strong><span> No doubt it\u2019s going to be better in your own harness, and it\u2019s just because of like where are you putting your reinforcement learning compute, right? You\u2019re putting your RL and your synthetic data, you\u2019re putting it to your own harness because it\u2019s the one that you understand the best and you\u2019re able to push the most. because that end control is what allows you to make it better. then transferring those capabilities is more about just making sure the model, induces the right amount of reasoning and like, understands some of the maybe more complex weird tool call formats that might exist somewhere else. and so we do some multi-harness polishing, as we call it. it\u2019s not really what drives capabilities, but it does create a better experience. I think everyone probably does these days, but it is totally fair to see why your own harness is going to still be better than others. And I think we see this with all the foundation model companies. and it\u2019s just that when you are pushing capabilities, you don\u2019t really wanna trade it off by putting 10 harnesses in your RL runs because it\u2019s just complexity. It\u2019s complexity of engineering because these&#8211; When you\u2019re trying to do good science, right, when you\u2019re trying to really understand what made my model improve, you wanna make one variable change to something you understand. And a harness from someone else, you don\u2019t know or understand in the same way as you understand your own, right? They might have different agents or different prompts<\/span><\/p>\n<p><strong>Swyx [01:00:48]:<\/strong><span> Yeah,<\/span><\/p>\n<p><strong>Eiso Kant [01:00:48]:<\/strong><span> In different places<\/span><\/p>\n<p><strong>Swyx [01:00:49]:<\/strong><span> If it\u2019s open source, you can look at the source.<\/span><\/p>\n<p><strong>Eiso Kant [01:00:50]:<\/strong><span> Yeah, but it\u2019s time, right? Like I really cannot stress, like I know I\u2019m like a weird person on this because like I have friends like, \u201cCan we meet up?\u201d Or, \u201cCan we do this?\u201d Or, \u201cCan we go out?\u201d I\u2019m like, \u201cNo.\u201d Because ultimately, this is a race, and time is the only thing that matters. And if I look at our team and say, \u201cOkay What is complexity worth introducing on our general trajectory to building more capable models? Which generalized to other harnesses quickly. And by the way, our model works well on other harnesses. I really encourage people to do it. It works well. Like I\u2019we\u2019ve been testing it in OpenCode and Kilo Code and others and like, and in Claude Code.<\/span><\/p>\n<p><strong>Swyx [01:01:22]:<\/strong><span> Which just got bought today.<\/span><\/p>\n<p><strong>Eiso Kant [01:01:24]:<\/strong><span> I saw it.<\/span><\/p>\n<p><strong>Swyx [01:01:25]:<\/strong><span> I mean Honda. Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:01:25]:<\/strong><span> Exactly.<\/span><\/p>\n<p><strong>Swyx [01:01:26]:<\/strong><span> Everything\u2019s getting bought.<\/span><\/p>\n<p><strong>Eiso Kant [01:01:27]:<\/strong><span> Exactly. and I think part of that is like, and there\u2019s some amazing. I\u2019m, I\u2019m excited, like I think Hermes is a ridiculously cool harness like, and<\/span><\/p>\n<p><strong>Swyx [01:01:37]:<\/strong><span> And, part of the question was just like how much of it is model versus model plus harness, right? So new benchmarks like Agents Last Exam, it\u2019s not wanting to just measure the model. same with models getting more and more agentic. They need a harness to operate in, right?<\/span><\/p>\n<p><strong>Eiso Kant [01:01:55]:<\/strong><span> I think for when you\u2019re asking that question to a model company, I think you can separate it in two parts, which is like The harness, like we have a very slimmed down harness. When you look at it\u2019s like six tools. It\u2019s like shell and like shell kill, shell wait, write, fetch web, and like, I don\u2019t know, bash. Like I think I\u2019m missing one, but like that\u2019s effectively all the tools. And it\u2019s very simple. It\u2019s very lightweight. So it is not a harness that is designed to try to do well on a benchmark or try to do well on a certain subset of things, right? It\u2019s not a deep research harness. So I think we see incredible ability for complex harnesses that build lots of prompts around and extra data sources and other tools to really push capabilities of models forward.<\/span><\/p>\n<p><strong>Eiso Kant [01:02:41]:<\/strong><span> But our model is still better than some other harnesses who do that in coding-like tasks because it was RL\u2019d with it.<\/span><\/p>\n<p><strong>Eiso Kant [01:02:48]:<\/strong><span> Now, I do encourage people, I think our model, by the way, is perfectly fine and good on ours. The differences are probably maybe too small for anyone to notice, but we see it ultimately still on benchmarks, by a little bit. So I think it\u2019s both are true. Foundation model companies with their harnesses will really push them because it\u2019s just operationally, the best way to have scientific rigor in improving your models. But also someone who takes our model and really does a lot of work on improving a harness is going to compete us, as they should. and that\u2019s just because the harness is the stopgap between what the model is capable of And what it needs as additional instructions, and what it needs is access to data and tools, right? And that\u2019s ultimately, I think, what a harness is. It\u2019s like, is it able. As you build more capable models, you\u2019re improving the instruction following the models. And so additional harness is just saying, \u201cHey, if you encounter X, Y, or Z, behave this way.\u201d And so even if you would say that two models with two different harnesses can equally reach the same capability that you care about, a harness that is really tailored towards a capability will do it more efficiently.<\/span><\/p>\n<p><strong>Eiso Kant [01:03:58]:<\/strong><span> It\u2019s like a person who\u2019s getting a manual of how to do the task in the most efficient way with the right tools and the right data sources versus a really smart person like, \u201cGo figure it out.\u201d They\u2019ll both solve the task, but one will do it a lot more efficient. So I\u2019m a big fan of all the harness development that\u2019s happening in the world, and we want to work with more harness like creators to also make sure that like if it needs some additional training, like publishing, that we will do it.<\/span><\/p>\n<p><strong>Swyx [01:04:22]:<\/strong><span> I mean, I think when you say it\u2019s a race, there\u2019s a question of what are you racing to? are you racing to be the best coding model company or the best coding model plus harness company? I think that\u2019s a, those are different things.<\/span><\/p>\n<p><strong>Swyx [01:04:36]:<\/strong><span> Or neither.<\/span><\/p>\n<p><strong>Eiso Kant [01:04:37]:<\/strong><span> Or neither.<\/span><\/p>\n<p><strong>Swyx [01:04:37]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:04:38]:<\/strong><span> So we. I race to AGI. Coding for us since day zero of our website has been, and we\u2019ve said this over and over again, we think focusing on coding and long horizon like software tasks is a path towards AGI because it forces us to solve the hard problems. It\u2019s, it forces us to solve the ability to do extremely long horizon complex work that requires lots of reasoning, external tools, data, et cetera. And one of the things I can show you, so we\u2019ll, we\u2019ll have a web chat on with this model, and I\u2019ve loved this model for deep research, just using it in my coding harness. It was never trained for it. It was never like looked at it, but it\u2019s great at it, in my opinion. because ultimately, the skills transfer, they generalize. Now, where we are not focused on today is to make sure that the world\u2019s greatest medical knowledge is encoded in this model or the world\u2019s greatest legal knowledge. But it did. We won\u2019t be publishing this benchmark \u2018cause we didn\u2019t have time to really do proper, but it did really well on LegalBench. and at least on our first runs, and we are very rigorous. When we publish evals, we have Checked them for every little thing. We have run them many times. We\u2019ve passed, like we\u2019ve gone and we\u2019ll, like we try to be extremely honest with this, so if we haven\u2019t spent enough time on a benchmark that we use internally that is public, we just say that we won\u2019t publish it. and<\/span><\/p>\n<p><strong>Swyx [01:06:01]:<\/strong><span> I mean, the other way is just to give it to artificial analysis and let them run it.<\/span><\/p>\n<p><strong>Swyx [01:06:04]:<\/strong><span> Like third party standards.<\/span><\/p>\n<p><strong>Eiso Kant [01:06:05]:<\/strong><span> Oh, 100%, and we are gonna be doing this as well. And still it takes time and effort, right? Because you\u2019re working with people to understand like, the infra failures and like the tools they\u2019re using and like, are they set up well. But I agree. You absolutely want to. I\u2019m a big fan of companies like Vals and Artificial Analysis and like others that are doing this stuff.<\/span><\/p>\n<p><strong>Swyx [01:06:21]:<\/strong><span> I found it very nice. You\u2019re the first to bring it up.<\/span><\/p>\n<p><strong>Eiso Kant [01:06:22]:<\/strong><span> Yeah. I think they\u2019re great. They\u2019ve got like. I loved like a lot of the work they\u2019ve done and put out. and so, and there\u2019s, I think, many more, and please create more eval companies. Like create more evals. I think it\u2019s so valuable for the industry.<\/span><\/p>\n<p><strong>Swyx [01:06:34]:<\/strong><span> It\u2019s an actual monopoly I feel like. Oh, and duopoly maybe.<\/span><\/p>\n<p><strong>Eiso Kant [01:06:37]:<\/strong><span> I think it can be broken.<\/span><\/p>\n<p><strong>Swyx [01:06:39]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:06:40]:<\/strong><span> Because I think it can be broken really easily because creating an eval for many people isn\u2019t sexy work, but whoever does it, everyone is happy to get a good eval. You\u2019ve like if an eval is well constructed, everyone\u2019s celebrating it, and everyone\u2019s willing to pay for it, and everyone\u2019s willing, like the foundation model<\/span><\/p>\n<p><strong>Swyx [01:06:55]:<\/strong><span> Oh, yeah. I think creating eval, yes. But like in terms of being like we are the industry standard ones that will<\/span><\/p>\n<p><strong>Eiso Kant [01:07:01]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Swyx [01:07:01]:<\/strong><span> \u03a4-bench and make sure that you didn\u2019t, you didn\u2019t cheat<\/span><\/p>\n<p><strong>Eiso Kant [01:07:03]:<\/strong><span> Yeah, that\u2019s true<\/span><\/p>\n<p><strong>Swyx [01:07:03]:<\/strong><span> And I\u2019ll run it the same way that you run it versus your competitor run it.<\/span><\/p>\n<p><strong>Eiso Kant [01:07:05]:<\/strong><span> Yeah. That is very true, and we need that. And it\u2019s nice that\u2019s like a few standard places that we all have to like, adhere to. It keeps us all honest. I think that\u2019s super important to do so, And, but yeah, no, I think our goal is to build the world\u2019s most capable models. and right now we are focused on the coding agent capabilities, long horizon work. But what you see with that is that you get a lot for free. I\u2019ve always said it\u2019s a lot easier for us as we get to SOTA and frontier on coding to then say, \u201cOkay, now we\u2019re going to obsess in using the model factory to add more data for places that, we\u2019re not as strong on,\u201d like could be medical or legal or any other areas. and similarly, I think what we see, and we see this with reasoning models a lot, if you give models access to the right knowledge sources and they have capable ways of reasoning, they\u2019re able to go very well into domains that are less known to them or even seen less in their training data. So, but yeah. Are we a agent like model plus harness comp-? No, we\u2019re a model company. but I think models today cannot be trained without harnesses. It\u2019s not possible. So it is just like where before it was just the weights in the container, well, now there\u2019s an agent harness that\u2019s attached to it. and but I think there\u2019s a big difference in being an agent harness as a model company than someone who\u2019s truly building an agent company. I think they can do far more than we can.<\/span><\/p>\n<p><strong>Swyx [01:08:27]:<\/strong><span> Yeah. understood. Yeah. I think that is my minor pushback. If you are truly identified as a model company, then make the best model for OpenCode, right? Instead of for pool or whatever. I think that\u2019s not as, that\u2019s, that\u2019s minor compared to if the goal is AGI, make the best model for Hermes.<\/span><\/p>\n<p><strong>Swyx [01:08:45]:<\/strong><span> Right? Like just \u2018cause that is the next stage after coding.<\/span><\/p>\n<p><strong>Eiso Kant [01:08:48]:<\/strong><span> I\u2019look, and we\u2019re working like very closely with them<\/span><\/p>\n<p><strong>Swyx [01:08:52]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [01:08:52]:<\/strong><span> Because I do think like it\u2019s, and, you have to care, you have to invest in it. It\u2019s why we do the polishing and we spend time on it. and I think over time, yeah, you\u2019re, you\u2019re right that you wanna balance that out. but ultimately you just want general capabilities that everything works equally in every harness.<\/span><\/p>\n<p><strong>Swyx [01:09:10]:<\/strong><span> Just on the topic, do you guys do much with like Hermes, OpenAI Codex, NanoCodex, whatever? Pi?<\/span><\/p>\n<p><strong>Swyx [01:09:16]:<\/strong><span> Pi.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:17]:<\/strong><span> Pi.<\/span><\/p>\n<p><strong>Swyx [01:09:17]:<\/strong><span> No, Pi is different.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:18]:<\/strong><span> It\u2019s more coding.<\/span><\/p>\n<p><strong>Swyx [01:09:19]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:19]:<\/strong><span> I\u2019m a big fan of Pi, though, I have to say. I think it\u2019s a really sexy<\/span><\/p>\n<p><strong>Swyx [01:09:22]:<\/strong><span> I forgot to mention Pi.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:23]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Swyx [01:09:23]:<\/strong><span> Pi, you sound closest to Pi in terms&#8211; pool and Pi in terms of like the minimal surface<\/span><\/p>\n<p><strong>Eiso Kant [01:09:28]:<\/strong><span> In the minimal yeah.<\/span><\/p>\n<p><strong>Swyx [01:09:29]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:29]:<\/strong><span> It\u2019s because I don\u2019- I have a. Allow me for one more strong opinion.<\/span><\/p>\n<p><strong>Swyx [01:09:33]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:34]:<\/strong><span> I\u2019ve been saying this now for two years.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:37]:<\/strong><span> I think MCP and tools are stupid.<\/span><\/p>\n<p><strong>Swyx [01:09:41]:<\/strong><span> Ooh. Let\u2019s go.<\/span><\/p>\n<p><strong>Swyx [01:09:42]:<\/strong><span> You support MCP.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:43]:<\/strong><span> I support MCP and we support tools and everything. They make absolutely no sense to me.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:48]:<\/strong><span> And like, and I\u2019ll explain a little bit why and I think I can probably get people to come along on this one.<\/span><\/p>\n<p><strong>Eiso Kant [01:09:56]:<\/strong><span> If you are looking for complex tasks, increasingly longer horizon, increasingly complex tasks, doesn\u2019t matter if it\u2019s coding or something else, You are gonna be interacting with data sources, right? And you\u2019re gonna be interacting with things that are installed on some form of a virtual machine.<\/span><\/p>\n<p><strong>Eiso Kant [01:10:15]:<\/strong><span> And what we are doing is that we\u2019re putting a layer in between those things. We\u2019re putting like MCP in between, we\u2019re putting tool calls in between, and this is even more about tool calls than MCP, where the model can just write the code and interact with the system. And we\u2019re starting to see that. Like Laguna S does this a lot. You\u2019ll see this as well in like frontier models. They\u2019re increasingly no longer, \u201cHere we\u2019re gonna stuff 50 tools in the like system prompt,\u201d to \u201cNo, here\u2019s a virtual machine with these binaries installed, this code base you can operate in. Here, a folder where you can write, your memory if you want to.\u201d And the model is using code to do complex asks. And when it uses code, it is not one or two tool calls or three things that are chained together. It starts, using if statements and for loops and making things conditional. And so I think we\u2019re moving from, we already are moving from tool calls, to effectively models writing code, little scripts, and you see this a lot when you get the Python,<\/span><\/p>\n<p><strong>Swyx [01:11:15]:<\/strong><span> Code interpreter.<\/span><\/p>\n<p><strong>Eiso Kant [01:11:16]:<\/strong><span> Exactly. Like in just the arrow in, written code in the file. I don\u2019t know what you call<\/span><\/p>\n<p><strong>Swyx [01:11:21]:<\/strong><span> EOF? Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:11:22]:<\/strong><span> Yeah, exactly. Like, you already see this happening more in models because when you start training them in RL, the models wanna be free. They wanna be able to do the thing they wanna do in the most efficient possible way, and it is not calling one of the 50 tools in their like system prompt. And so I\u2019m a very big fan of Give the model a minimal harness, as minimal as possible, give it a container in which it has its own code base, right? The, got a models code base that has access to the API keys and data sources and little libraries and documentation that it needs, and just let it run free at the task. and I think that is the way we\u2019re going. I think we will, in 12 months, not see a single system prompt that is stuffed with 20 or 30 or 40 tools anymore.<\/span><\/p>\n<p><strong>Swyx [01:12:07]:<\/strong><span> No comment. no pushback there. I think there will be, it\u2019ll be supported for a long time just because that\u2019s, a lot of people are trained on that now, but maybe you guys don\u2019t have to support it in your models, going forward. So, but yeah, I mean, if you can. I do think that\u2019s, writing code is more generalist and it\u2019s a, it\u2019s a means to an end<\/span><\/p>\n<p><strong>Eiso Kant [01:12:26]:<\/strong><span> And we do support tools.<\/span><\/p>\n<p><strong>Swyx [01:12:27]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:12:27]:<\/strong><span> We support. And this is the first model we\u2019re doing parallel tool calling in which we needed to catch up on. So like that\u2019s there and like<\/span><\/p>\n<p><strong>Swyx [01:12:32]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [01:12:32]:<\/strong><span> So it\u2019s, it\u2019s there, but I,<\/span><\/p>\n<p><strong>Swyx [01:12:35]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [01:12:35]:<\/strong><span> It\u2019s a personal, nitpick. I like, I want the models to have as many degrees of freedom and just like, be free and do capable things.<\/span><\/p>\n<p><strong>Swyx [01:12:43]:<\/strong><span> Yeah. So and then, so that was on the path towards like, okay, how do you use Poolsides models and Laguna models for my Hermes or my OpenAI Codex<\/span><\/p>\n<p><strong>Eiso Kant [01:12:52]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Swyx [01:12:52]:<\/strong><span> On all those things. And so typically what I look for is, Computer use or vision. That\u2019s a, that\u2019s a very big one. You guys have a blog post on that. but then also the persistence I think is very strong value, as well as long context, which you guys have a million token context. Anything else?<\/span><\/p>\n<p><strong>Eiso Kant [01:13:08]:<\/strong><span> So for us, look, so for us, vision understanding is the next thing, right?<\/span><\/p>\n<p><strong>Swyx [01:13:11]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:13:11]:<\/strong><span> Like we don\u2019t have vision understanding.<\/span><\/p>\n<p><strong>Swyx [01:13:12]:<\/strong><span> Which I was gonna say is<\/span><\/p>\n<p><strong>Eiso Kant [01:13:14]:<\/strong><span> We don\u2019t have vision understanding in these models yet.<\/span><\/p>\n<p><strong>Swyx [01:13:16]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Swyx [01:13:17]:<\/strong><span> To<\/span><\/p>\n<p><strong>Eiso Kant [01:13:17]:<\/strong><span> And so this is something that we\u2019ve, we\u2019ve started efforts on. Like we think it\u2019s, it\u2019s super important to have visual understanding.<\/span><\/p>\n<p><strong>Swyx [01:13:23]:<\/strong><span> That\u2019s company vision.<\/span><\/p>\n<p><strong>Eiso Kant [01:13:24]:<\/strong><span> And so no, we\u2019ve got work to do there. and this is one of the things I loved about the Thinky model, like from the Two minutes I scrolled the blog post<\/span><\/p>\n<p><strong>Swyx [01:13:33]:<\/strong><span> Yep<\/span><\/p>\n<p><strong>Eiso Kant [01:13:33]:<\/strong><span> Multi, the multi<\/span><\/p>\n<p><strong>Swyx [01:13:34]:<\/strong><span> They\u2019re, they\u2019re very committed to multimodal, including audio. Yeah.<\/span><\/p>\n<p><strong>Vibhu [01:13:36]:<\/strong><span> They\u2019re state-art audio, as much as it\u2019s a trillion parameter state-art audio, but also all trained from scratch, right?<\/span><\/p>\n<p><strong>Swyx [01:13:43]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [01:13:43]:<\/strong><span> No encoder in the sense<\/span><\/p>\n<p><strong>Swyx [01:13:45]:<\/strong><span> To me, that\u2019s, that\u2019s, that\u2019s one of the strongest reasons why you need to train from scratch, is you just have a different tokenizer, you\u2019d have different<\/span><\/p>\n<p><strong>Eiso Kant [01:13:51]:<\/strong><span> I\u2019m fully aligned, like zero disagreement from me here. Like, just add the modality and don\u2019t put. keep it simple. we\u2019I don\u2019t think we\u2019ll touch audio for a very long time.<\/span><\/p>\n<p><strong>Vibhu [01:14:05]:<\/strong><span> It\u2019s in the name too, InkLink Inc.<\/span><\/p>\n<p><strong>Eiso Kant [01:14:08]:<\/strong><span> True.<\/span><\/p>\n<p><strong>Swyx [01:14:08]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Swyx [01:14:09]:<\/strong><span> I mean, what\u2019s so hard, what\u2019s so hard about audio?<\/span><\/p>\n<p><strong>Eiso Kant [01:14:11]:<\/strong><span> It\u2019s not about what\u2019s Again, it all comes down to focus.<\/span><\/p>\n<p><strong>Swyx [01:14:14]:<\/strong><span> I see.<\/span><\/p>\n<p><strong>Eiso Kant [01:14:15]:<\/strong><span> Right? Like saying no to things means that there\u2019s a research or an compute that can go to making general progress, and our view is like general progress, is going to come from the ability to push these models to far more capable reasoning, far more longer horizon tasks. I don\u2019t think audio Adds to that. I don\u2019t think it pushes us close to AGI. I think it is a necessary modality as you get closer to AGI. I think visual understanding sits in the middle of those things. I think visual understanding can absolutely, do so, but it also unlocks capabilities that are just valuable today. so but this is the point, right? You want more diversity, you want more different foundation model companies who focus on different things. I think we are just like a horse with blinders on, just like<\/span><\/p>\n<p><strong>Swyx [01:14:58]:<\/strong><span> Yeah, you have your path<\/span><\/p>\n<p><strong>Eiso Kant [01:14:59]:<\/strong><span> We have our path, we wanna catch up to the frontier, and, we don\u2019t wanna distract ourselves with anything else.<\/span><\/p>\n<p><strong>Swyx [01:15:05]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Swyx [01:15:06]:<\/strong><span> I will call out that one of the branches of research is DeepSeek OCR, which is can you just throw away the text tokenizer and just have only vision?<\/span><\/p>\n<p><strong>Eiso Kant [01:15:13]:<\/strong><span> I find this&#8211; I look, geek, the geek in me is like looks at this stuff and it\u2019s like, okay, look at this, like look at the number of bits encode<\/span><\/p>\n<p><strong>Swyx [01:15:20]:<\/strong><span> But they\u2019re right.<\/span><\/p>\n<p><strong>Eiso Kant [01:15:21]:<\/strong><span> I think it\u2019s super cool, right? But I think this is what we\u2019re gonna come back down to. Like probably works, it\u2019s just is it compute efficient enough? Is it Like I think so many of these things ultimately will work. It\u2019s just like, what\u2019s the nice thing about text? And I referenced earlier, Peng Ming and Nikolai are my two heads of applied research who are just incredible, like we wouldn\u2019t have gotten here without them and the entire team.<\/span><\/p>\n<p><strong>Eiso Kant [01:15:45]:<\/strong><span> And Nikolai have&#8211; and I have been debating, for years about like, should reasoning be in latent space? Should reasoning be in tokens? But one thing that I think him and I really agree on, and all three of us, and is that like Language is incredible because it\u2019s such an incredibly dense way to encode knowledge and information and intelligence, right? If you think about like what went into a physics paper that then is, 20 or 30 pages, like the amount of intelligence and thought and whatnot to then generate that, like in that 20-page document, like those little amount of bits, there\u2019s so much encoded. And other modalities like video and images are amazing, but they don\u2019t have the same density of like knowledge or reasoning or however, like the things that we\u2019re trying to push for that are encoded in that modality. They\u2019re there. In many cases, you can watch an incredible lecture for, 50 minutes on YouTube, but the amount&#8211; and but if you treat that as video in data versus text data, right, the bits to like signal-noise ratio, the compute efficiency of the modality is a lot less. And so we have this view as like with language you can go really far, but also when you have limited compute, limited, people, and they\u2019re very much linked to two, I think we can push language. It\u2019s the more, it\u2019s the better investment. But I want all the modalities. I find it super cool and I love what DeepSeek and others are trying. Like I can retweet them all the time, but internally we\u2019re just like, \u201cLet\u2019s stay focused.\u201d<\/span><\/p>\n<p><strong>Vibhu [01:17:17]:<\/strong><span> Which I\u2019ll say, you can see somewhat works looking at Anthropic. OpenAI has a lot of vision, multimodality. Anthropic just didn\u2019t, right? Fable\u2019s a big step up in image processing, but like they\u2019re not known as the multimodal company, right? They\u2019re the language model coding company that has multimodal capabilities that\u2019s never super flex and, goes pretty far.<\/span><\/p>\n<p><strong>Eiso Kant [01:17:42]:<\/strong><span> I look, I in this I think Anthropic, I mean, they\u2019ve done many things right, but I think this maniacal focus on just pushing capabilities, scaling up models is. I couldn\u2019t agree more. I think it\u2019s, it\u2019- that\u2019s the first hurdle, and once we get that, then we can improve a whole bunch of other things. and but at the same time, on the other end of the spectrum, it\u2019s really exciting to see people, building these spatial models, right? That are, and the world models that are being built, like for very different, use cases. but I think ultimately it all comes together at some point.<\/span><\/p>\n<p><strong>Vibhu [01:18:19]:<\/strong><span> Okay. So scaling models, this is Laguna S for small.<\/span><\/p>\n<p><strong>Eiso Kant [01:18:23]:<\/strong><span> Yes.<\/span><\/p>\n<p><strong>Vibhu [01:18:23]:<\/strong><span> You have good naming, extra small, medium, large.<\/span><\/p>\n<p><strong>Eiso Kant [01:18:26]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [01:18:26]:<\/strong><span> Still scaling?<\/span><\/p>\n<p><strong>Eiso Kant [01:18:28]:<\/strong><span> So the new medium started training, and it\u2019s much bigger than the last medium, started training yesterday. so it\u2019s a, 39-day training run. and,<\/span><\/p>\n<p><strong>Vibhu [01:18:39]:<\/strong><span> How do the days and events? Just the compute model<\/span><\/p>\n<p><strong>Eiso Kant [01:18:41]:<\/strong><span> Models factory.<\/span><\/p>\n<p><strong>Vibhu [01:18:42]:<\/strong><span> Okay.<\/span><\/p>\n<p><strong>Eiso Kant [01:18:42]:<\/strong><span> Right? And like at this point, like with the model factory, like it\u2019<\/span><\/p>\n<p><strong>Vibhu [01:18:46]:<\/strong><span> I thought it was interesting. So in the Laguna medium and extra small, you even quoted number of GPU hours for how many days and whatever for different size. And I\u2019m like, \u201cOh, you can also work backwards to how much that costs, right? What GPUs, how many hours \u201c<\/span><\/p>\n<p><strong>Eiso Kant [01:18:59]:<\/strong><span> And you realize it\u2019s not a lot.<\/span><\/p>\n<p><strong>Vibhu [01:19:00]:<\/strong><span> No, it\u2019s not.<\/span><\/p>\n<p><strong>Eiso Kant [01:19:01]:<\/strong><span> It\u2019s not a lot of money. and, you started with DeepSeek of the West and, I think that\u2019s, The DeepSeek moment, right, was a moment when people realized that you can train incredibly capable models for not a lot of money on the training run. But I think that\u2019s the falsehood, right? Like the training run is not the expensive part. The training run is a very anticlimactic event, right? Like we just had a Slack message come up yesterday saying, \u201cThe new model is training and here are the links, so you can follow along the evals,\u201d and like that\u2019s it. all the work that goes into that moment, it\u2019s like how people talk I know nothing about sports, but how, like, athletes talk about, like, it\u2019s all the preparation, it\u2019s all the going to the gym, and then the game is just a game. I think that\u2019s a little bit like with model training.<\/span><\/p>\n<p><strong>Swyx [01:19:42]:<\/strong><span> Yeah. People had over-indexed on DeepSeek was trained for $5 million or whatever it was, right? It\u2019s like there\u2019s the amount of R&amp;D before that, the infrastructure is built up. Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:19:51]:<\/strong><span> Exactly, all the things, the data. But no, so Laguna M is training, and yes, there will be an L and there will be an XL, and what you\u2019ll<\/span><\/p>\n<p><strong>Swyx [01:19:57]:<\/strong><span> Ooh.<\/span><\/p>\n<p><strong>Eiso Kant [01:19:57]:<\/strong><span> What you\u2019ll see with M, right, M is much larger than the last M, right? So these monikers are a little bit our version of the different<\/span><\/p>\n<p><strong>Swyx [01:20:04]:<\/strong><span> Yeah, he was making fun of people for saying small is 24B or something.<\/span><\/p>\n<p><strong>Swyx [01:20:08]:<\/strong><span> No, so, no. Small for Mistral now is over 100B.<\/span><\/p>\n<p><strong>Eiso Kant [01:20:12]:<\/strong><span> What?<\/span><\/p>\n<p><strong>Swyx [01:20:12]:<\/strong><span> Yeah, I can pull it up.<\/span><\/p>\n<p><strong>Eiso Kant [01:20:13]:<\/strong><span> I mean, our small, right, is 118, so I don\u2019t wanna say anything else. Like, it\u2019<\/span><\/p>\n<p><strong>Swyx [01:20:17]:<\/strong><span> I mean, I think it\u2019s also. Okay, yeah, your small is<\/span><\/p>\n<p><strong>Eiso Kant [01:20:20]:<\/strong><span> We all know that the single hardest thing for any foundation model company is naming.<\/span><\/p>\n<p><strong>Eiso Kant [01:20:25]:<\/strong><span> I don\u2019t want to say that we\u2019re good at it either. I mean, it\u2019this is Laguna S 2.1. It\u2019s, it\u2019<\/span><\/p>\n<p><strong>Swyx [01:20:32]:<\/strong><span> But at least people understand, medium is bigger than small. Until you mess that up, like<\/span><\/p>\n<p><strong>Eiso Kant [01:20:37]:<\/strong><span> Exactly<\/span><\/p>\n<p><strong>Swyx [01:20:38]:<\/strong><span> You have a pass.<\/span><\/p>\n<p><strong>Eiso Kant [01:20:38]:<\/strong><span> We try hard.<\/span><\/p>\n<p><strong>Swyx [01:20:40]:<\/strong><span> While we\u2019re on the topic of naming, this is gonna be at the end, but might as well<\/span><\/p>\n<p><strong>Eiso Kant [01:20:43]:<\/strong><span> Sure<\/span><\/p>\n<p><strong>Swyx [01:20:43]:<\/strong><span> Why Poolside? Why Laguna?<\/span><\/p>\n<p><strong>Eiso Kant [01:20:46]:<\/strong><span> So When we started the company, it was gonna be called Snowball Apps. it was after the snowball effect because we expected this company to become a snowball effect, and it definitely has been a snowball effect for us. turns out it\u2019s an Amazon trademark.<\/span><\/p>\n<p><strong>Eiso Kant [01:20:59]:<\/strong><span> I kid you not that my founder\u2019s next suggestion of a name was, \u201cLet\u2019s call it Bedrock.\u201d And so at this point it was like, \u201cOkay, no, you are amazing at naming things if you worked Amazon.\u201d and so, early on in the company, before we were incorporated, we were at an annual conference of a very big Major tech company, and we had been discussing with them. And you have to realize the company at this point is me, my founder, our CEO, Margarita. We know the first person who\u2019s gonna join us. We haven\u2019t, like, incorporated yet. and we were discussing an OpenAI Microsoft-style deal with this big tech company. Like, they were going to provide us with a lot of compute. We would give them, perpetual access, a whole bunch of things.<\/span><\/p>\n<p><strong>Eiso Kant [01:21:49]:<\/strong><span> And, we found out the name was trademarked, Snowball Labs, while we were at that conference and having this discussion that we had no right to have, right? We were a couple of guys who had nothing yet, but this big company was willing to entertain the fact that we might partner with them. And, we were discussing this, and it was in their annual conference in a public setting, and the chief scientist of that company said, \u201cPeople can hear us here. Like, we should move somewhere else. Let\u2019s go to the restaurant Poolside.\u201d And for some reason, me and Jason looked at each other in that moment and said, \u201cOh.\u201d and then later that night, &#8211; the name stuck with us. The word stuck with us, and we said, \u201cLet\u2019s call the company Poolside.\u201d And ever since, we never ended up doing that deal, and we used it as a reminder to never turn down our, round down our ambitions, because that would\u2019ve been the easy path. and the hard part was what we did, which is start and try to raise exorbitant amounts of money when you\u2019re just a couple of guys who are not even building it in Silicon Valley, who don\u2019t come from any, of the known knobs and things like this. And so everyone assumes Poolside because AGI, everyone sits Poolside, and it was a playful name, and we liked it, and it was a little bit different. But the name is, like, a reminder for us to never round down our ambitions, and whenever you\u2019re faced with those decisions to just pick the harder path.<\/span><\/p>\n<p><strong>Swyx [01:23:09]:<\/strong><span> Yeah. I mean, that\u2019s a great story. I know you\u2019ve told it before<\/span><\/p>\n<p><strong>Eiso Kant [01:23:13]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Swyx [01:23:13]:<\/strong><span> But I just wanted<\/span><\/p>\n<p><strong>Eiso Kant [01:23:14]:<\/strong><span> Right<\/span><\/p>\n<p><strong>Swyx [01:23:14]:<\/strong><span> On the record. but that\u2019s, that\u2019s what I did the first time I met you. You told me, you sat me down. You were, you, we were in the hotel somewhere.<\/span><\/p>\n<p><strong>Eiso Kant [01:23:21]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Swyx [01:23:21]:<\/strong><span> And you were like, \u201cWe\u2019re raising a $500 million.\u201d I\u2019m like. And then you gave me the whole vision, and then you did it. And I was like, well, it\u2019s, I don\u2019t have that much opportunities to ask, like, just how do you do that raise to that to those kinds of VCs? What are they looking for? like, yes, vaguely AGI, but, like, what do they want when<\/span><\/p>\n<p><strong>Eiso Kant [01:23:42]:<\/strong><span> Look, it\u2019s, the world\u2019s definitely changed, right? When we were raising that $500 million round, the majority of investor conversations were still trying to explain that these models were not just stochastic parrots and that they were gonna keep going. I\u2019ve seen the world go from OpenAI is gonna win it all and there\u2019s no one else who can build company, right? I mean, Anthropic struggled, to raise their $500 million round. That\u2019s like, well reported. They pulled it off, gladly. and so I think when we raised that, it was about a year and a half ago at this point, the world was very different than it is today. I think the world today, There\u2019s been, there\u2019s been this function where the number of people who believe AGI is real, Is probably a, an, a super linear or definitely some form of an exponential function itself.<\/span><\/p>\n<p><strong>Eiso Kant [01:24:31]:<\/strong><span> And I think this is important because if you hold the belief that we had three years ago and a year and a half ago, and we looked for people who shared that belief, which is like, this technology is gonna fundamentally underpin everything that\u2019s economically interest- or economically valuable and scientifically interesting for, like, the future, then the value function afterwards is easy to understand, which is like, hey, if you get there, you are one of the commodity, one of the players who can build this commodity. and over the years, building that commodity has become not just about building models, but also about building infrastructure and other things.<\/span><\/p>\n<p><strong>Eiso Kant [01:25:03]:<\/strong><span> And so I think today, because the number of people is bigger and the outcomes have been proven, right? I think the incredible, like, financial success that Anthropic is having right now and, like, the growth that OpenAI\u2019s had and others and Google no longer make this a question of is there product market fit, which really a couple of years ago was, like, part of the question. Like, how big can these things be? You tell people that, like, you\u2019d be at these amount of revenue numbers in our industry right now, people were still, like, would laugh you out the room.<\/span><\/p>\n<p><strong>Eiso Kant [01:25:33]:<\/strong><span> Now I think it\u2019s a function of who in the world believes that it\u2019s gonna be an oligopoly of intelligence And who believes that oligopoly can be broken by other companies. And I think that\u2019s what divides investors more than anything else. For the ones who believe in AGI, and then you\u2019ve got a whole layer that, is self-selecting out, foundation model companies because they\u2019re like, \u201cLook, I can\u2019t make &#8211; The money I put there, compared to what I can put in an application company is very different.\u201d I think there\u2019s incredible application companies, and there should be many should be built. But I do think we are still in a world right now where this is the early innings &#8211; this can still be the early innings of who is going to, be part of the set of people who win. This &#8211; Intelligence is the most, in my view, gonna be the world\u2019s most demanded commodity. It will more commoditize in margin and price. and the world wants choice and wants options. And so I think treating the world as like, \u201cOh, there\u2019s only gonna be two players,\u201d I think is very shortsighted from investors.<\/span><\/p>\n<p><strong>Eiso Kant [01:26:41]:<\/strong><span> I think that group who thought that was a lot bigger at the beginning of the year than now.<\/span><\/p>\n<p><strong>Eiso Kant [01:26:46]:<\/strong><span> I think the last couple of months have woken up a lot of people and going, \u201cHoly shit,\u201d like, the world both can use a lot more intelligence, but also, like, the world is far more complex. We should have multiple choices, more options, things that can be turned off, that can\u2019t be, that. The restrictions that people put on models now, I think, is another area of this, right?<\/span><\/p>\n<p><strong>Eiso Kant [01:27:08]:<\/strong><span> Like, the fact that We are entering into a world where model companies are saying, \u201cYou\u2019re not allowed to use me for foundation model company development.\u201d They should be allowed to do this. It\u2019s capitalism. It\u2019s their business. It\u2019s their work product.<\/span><\/p>\n<p><strong>Eiso Kant [01:27:23]:<\/strong><span> But it is insane.<\/span><\/p>\n<p><strong>Eiso Kant [01:27:25]:<\/strong><span> It is wild that we are, like, okay with that.<\/span><\/p>\n<p><strong>Swyx [01:27:30]:<\/strong><span> Do you have more problem with Anthropic saying it or the White House saying it? that&#8211; that you\u2019re picking Two different<\/span><\/p>\n<p><strong>Eiso Kant [01:27:37]:<\/strong><span> Things<\/span><\/p>\n<p><strong>Swyx [01:27:37]:<\/strong><span> Limitations and restrictions there.<\/span><\/p>\n<p><strong>Eiso Kant [01:27:39]:<\/strong><span> Look, I think I, &#8211; I\u2019ll put it this way. I think we wanna, as this technology gets more capable, for the better and worse, we do wanna yield to democracy to figure this out more and more. I think any single company making unilateral decisions, is, Is dangerous. It\u2019s a concentration of power in a small number of people with very limited checks and balances. and that has never worked out well in history, in any way, shape, or form. and this is not a criticism on the existing foundation model companies. This is just more commentary on, like, how I\u2019d like the world to be. I think in a world where the technology gets more capable, government needs to play an active role in determining, where is there real risks of misuse, right? And I do think we need to separate safety between misuse, and, doomsday scenarios that, I think No one knows if gonna, are gonna happen or not. And I think just, like, very practically, I think, I\u2019m glad to see there\u2019s a lot of conversation now starting to happen again at the government level of trying to figure this out. and now what the final decisions are, maybe I\u2019m happy about them, maybe I don\u2019t, maybe I agree, maybe not. But ultimately, like, that\u2019s democracy always, right? Like, at any given moment, I might not be perfectly happy with one or the other, but people chose to vote in someone to make those decisions. And so I think over the long run, over a 20-year time span, the world directionally goes correct and democracy does work. At least, what\u2019s the famous quote of like it\u2019s the worst of &#8211; It\u2019s the best of all the worst systems or something like that.<\/span><\/p>\n<p><strong>Swyx [01:29:26]:<\/strong><span> It\u2019s the worst form of, organization except for all the others that we\u2019ve tried.<\/span><\/p>\n<p><strong>Eiso Kant [01:29:30]:<\/strong><span> Exactly. That\u2019s the one.<\/span><\/p>\n<p><strong>Swyx [01:29:31]:<\/strong><span> You can always count on me for a Churchill quote \u2018cause I\u2019ve, studied Churchill a lot.<\/span><\/p>\n<p><strong>Eiso Kant [01:29:35]:<\/strong><span> I love that. and so that\u2019s what I hope for. Now, I do think we are in a critical moment of time, and so speaking up for anyone is important. I think, researchers who are thinking about starting their own foundation model companies start. people who wanna share their opinion and be vocal, if that\u2019s with their representatives or just out on X, like, do so.<\/span><\/p>\n<p><strong>Eiso Kant [01:29:57]:<\/strong><span> And but concretely to your point, I think we are not at a level of capability right now that we should start restricting, open models in any way, shape, or form. I think it will hurt innovation if we do so.<\/span><\/p>\n<p><strong>Swyx [01:30:14]:<\/strong><span> Is there a point at which you will change your opinion there?<\/span><\/p>\n<p><strong>Eiso Kant [01:30:17]:<\/strong><span> Yes. I mean, look, &#8211; And there has to be.<\/span><\/p>\n<p><strong>Swyx [01:30:19]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:30:20]:<\/strong><span> Right? Like, you cannot. If you sit with a straight face and say, \u201cThis can be open forever in every way, shape, or form,\u201d it is just as, I think, egregious as saying, the opposite of it all needs to be closed down right now. Like, I think at any ends of extremes of spectrums is where we go wrong.<\/span><\/p>\n<p><strong>Eiso Kant [01:30:41]:<\/strong><span> Right? In society in any way, shape, or form. And so the answer is always more nuanced, and the answer is never black and white. And so I think as we encounter, like, real world scenarios where we have to say, \u201cHey, we have to be more careful,\u201d we need to reevaluate. If that means training a model differently and opening it up, having different versions, some things that, That are restrict&#8211; I think that\u2019s totally okay because I don\u2019t think anyone should be irresponsible. What I do wanna call out is that people have been calling for the fear of misuse of these models since 2, Right? And I still remember, like, \u201cWe cannot release 2 because the whole world will get \u201c<\/span><\/p>\n<p><strong>Swyx [01:31:20]:<\/strong><span> I mean, that was Dario.<\/span><\/p>\n<p><strong>Eiso Kant [01:31:21]:<\/strong><span> And so, like, this is not a commentary on Dario, it\u2019s a commentary just in general in the space. And so We have not been very good at this so far, and we need to get better at it. And I do think that the work that\u2019s happening with, like, safety institutes and better evals and things like that is probably the right direction.<\/span><\/p>\n<p><strong>Swyx [01:31:38]:<\/strong><span> Yeah. I mean, I wanna say something in defense of this. It\u2019s better to err on the side of safety and then roll it back rather than the other way because the other way, it\u2019s a one, way decision. I think that\u2019s, I think that\u2019s true.<\/span><\/p>\n<p><strong>Vibhu [01:31:53]:<\/strong><span> The caveat there is also the competition, right? You don\u2019t have global error on the side of safety, right? You\u2019re talking<\/span><\/p>\n<p><strong>Swyx [01:32:01]:<\/strong><span> Yeah, exactly.<\/span><\/p>\n<p><strong>Vibhu [01:32:02]:<\/strong><span> So Oh, yeah<\/span><\/p>\n<p><strong>Swyx [01:32:02]:<\/strong><span> You don\u2019t get to do unilateral safety because someone else will just be more unsafe than you.<\/span><\/p>\n<p><strong>Vibhu [01:32:06]:<\/strong><span> Yeah, exactly.<\/span><\/p>\n<p><strong>Swyx [01:32:07]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [01:32:07]:<\/strong><span> You can pause innovation here. It doesn\u2019t mean it\u2019s, it\u2019s pausing elsewhere.<\/span><\/p>\n<p><strong>Swyx [01:32:11]:<\/strong><span> They\u2019ll just take over the world. It\u2019s so easy.<\/span><\/p>\n<p><strong>Eiso Kant [01:32:13]:<\/strong><span> They\u2019re, they\u2019re complex parts.<\/span><\/p>\n<p><strong>Swyx [01:32:14]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:32:15]:<\/strong><span> Right? And I think we are much better off talking about certain capabilities that we can, commonly agree on and internationally agree on that we want to, limit or not have available, than we should talk about it in black and white of models available, yes or no. Like, the moment you start getting these big blanket statements, it\u2019that\u2019s when you start getting at the risk of, like. I always think back about when we banned advertising on cigarettes. Good thing. I\u2019m not saying I\u2019m against that. But it effectively established an oligopoly of cigarette companies because no one else could ever compete. and it was the, probably the best moment to the tobacco industry that ever happened, And we don\u2019t wanna do that right now. If we pull up, walls behind innovation, and this is a self-serving comment because I\u2019m not at the frontier yet, but it\u2019s not just related to me. I think it\u2019s related to everyone in the space. you are deciding right now in 2026, based on the current capabilities of models, that this is something that only two or three companies can build, and that to me reads like chapter 14 of the most dystopian fi novel that I could read because from there I think you can play out all the scenarios that happen in the world, and none of those are the ones that make me, excited about the future. and I think that\u2019s the thing we should all think about. Like, what\u2019s the future we wanna be excited about? What do we wanna have? And I think that\u2019s a future where intelligence is a commodity. Everyone can access it. It becomes cheaper and cheaper, right? I think that\u2019s important. It can, like, impact more of the world, and it\u2019s not one where, a single company puts their thumb on their scale of both what it outputs, to or turns it on or off.<\/span><\/p>\n<p><strong>Swyx [01:33:56]:<\/strong><span> I think the one entity that has more power than the US government here is Nvidia.<\/span><\/p>\n<p><strong>Swyx [01:34:02]:<\/strong><span> Because, like, whoever gets the allocations gets the compute.<\/span><\/p>\n<p><strong>Vibhu [01:34:06]:<\/strong><span> You can take it down to TSMC or,<\/span><\/p>\n<p><strong>Swyx [01:34:09]:<\/strong><span> And TSMC below that. But I just wanna test provocative statements to see if you have any response.<\/span><\/p>\n<p><strong>Eiso Kant [01:34:18]:<\/strong><span> I need to think on that one.<\/span><\/p>\n<p><strong>Vibhu [01:34:20]:<\/strong><span> Which I think they are regulated, right? Like, you can see the government<\/span><\/p>\n<p><strong>Swyx [01:34:23]:<\/strong><span> Nvidia\u2019s not regulated.<\/span><\/p>\n<p><strong>Vibhu [01:34:24]:<\/strong><span> Can they ship to China?<\/span><\/p>\n<p><strong>Swyx [01:34:26]:<\/strong><span> Okay, but they\u2019re not China.<\/span><\/p>\n<p><strong>Eiso Kant [01:34:30]:<\/strong><span> Look, I think this industry Has existed because of what Nvidia\u2019s done.<\/span><\/p>\n<p><strong>Swyx [01:34:35]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:34:35]:<\/strong><span> Right? I know they&#8211; &#8211; People like it\u2019s easy to give them flack, but I also wanna say, like, I remember when we started Source, right? In 2015 post that capacity article. It was able for this progress to happen because we were able to put consumer GPUs in servers, and they allowed us to do so, and then, like, and you kept going further. And so this is something, like, foundation models are so closely linked to their hardware and their systems.<\/span><\/p>\n<p><strong>Swyx [01:34:58]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:34:59]:<\/strong><span> Why do we see these stepwise progress happening? We see them happening because of the next generation of networking and systems that come out, right? The difference of a model you could train on Hoppers versus GB300s is the difference between a trillion-parameter model and a five or six trillion-parameter model. And so these things really coexist, I think, very closely to each other, and I think the more interesting question, I think, for the future is going to become of, like, how do &#8211; what can we unlock in terms of model capabilities, like, as we start designing these things even more? And we\u2019re seeing that with, like, the next generation of systems. And I think the world, abhors.<\/span><\/p>\n<p><strong>Eiso Kant [01:35:42]:<\/strong><span> Like, capitalism does a really good job at trying to, like, push towards things that &#8211; that allow for more competition, right? And Nvidia allows for competition. It\u2019s not. But if a government says no one else can build foundation models effectively through the regulation, that is very different. Now, is it hard to go build an Nvidia? Absolutely. Is it hard to build a foundation model? I think it\u2019s very hard to build a foundation model. But we should, like, make the playing field one that where, if someone wakes up tomorrow and wants to do so, they are, like, allowed to do so, and they\u2019re allowed to use the tools to do so. And I think there\u2019s still a big difference between what we\u2019re seeing in the discussions around model companies versus what we\u2019re seeing with chip companies.<\/span><\/p>\n<p><strong>Vibhu [01:36:25]:<\/strong><span> The gap also seems to be the expertise in who regulates it, right? Who at the government decides what\u2019s too safe, too smart, too dangerous? but while we\u2019re throwing spicy questions out there, do you have anything that comes to top of mind that could be changed? So, should OpenAI, Anthropic, open source models? Is it open weights? Is it what we do in RL that determines, your safety barriers? Is there anything that should be done there or just spitballing?<\/span><\/p>\n<p><strong>Eiso Kant [01:36:53]:<\/strong><span> That\u2019s a good question. yes. one of the things that I\u2019m excited about that I think we\u2019re more and more talking about, I don\u2019t think anyone is doing yet, is, mix and match of hardware during RL training, right? Like, &#8211; You think about, like, the notion, and we\u2019re seeing this in inference, right? The prefill and decode<\/span><\/p>\n<p><strong>Vibhu [01:37:15]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [01:37:16]:<\/strong><span> Just work better with, a general purpose, GPU and a more specialized, like, chip, right? Like, if the Groq chip at Nvidia, the LPU and the GPU combined, and there\u2019s different versions of that in the industry. And RL is batch size constrained, Right? So, like, you are ultimately&#8211; and then you\u2019re batch size constrained because you don\u2019t have infinite tasks, right? When you\u2019ve got the entire web, you can be much more flexible in scaling up your batch size because you\u2019ve got the entire web. But for RL, you have, X millions of tasks that you are gonna be training on, and so you cannot blow up your batch size massively, which means that you can\u2019t scale compute to a certain extent with RL the same way you could scale compute with, like, training. and so I\u2019m very excited about anything that improves that. And I think one of the best ways to start improving that is the things that we\u2019re already starting to see in inference, which is the separation of the prefill and decode to different chips to come to reinforcement learning, right? and I think we\u2019ll be there soon. and I think more people should be working on this, because then all of a sudden we\u2019re able to just be way more efficient in how we train RL from a wall clock time. Again, coming back down to the fact that it\u2019s a race, right? The race is measured not in how many GPUs, but the race is measured on calendar time, and that\u2019s probably one of the biggest impacts we can have right now to speed up our industry. and so that\u2019s one, like, technically I love geeking out about and talking to people. Yeah.<\/span><\/p>\n<p><strong>Swyx [01:38:45]:<\/strong><span> Yeah, I would talk to Etched. I had a tour of their data center and, physically you can see how PD disaggregation is mapped out in the data center, and you have to own your own hardware to do that.<\/span><\/p>\n<p><strong>Eiso Kant [01:38:57]:<\/strong><span> Yeah. No, look, I think it\u2019- I think more innovation in the space is just, like, is the coolest thing.<\/span><\/p>\n<p><strong>Swyx [01:39:02]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:39:03]:<\/strong><span> And so I\u2019m, I\u2019m excited because that\u2019s like, all of us are.<\/span><\/p>\n<p><strong>Eiso Kant [01:39:09]:<\/strong><span> Like, why don\u2019t we finish, post-training this model, whatever, two weeks before release? Or no, sorry, between release, between training, then, training SFT, and then the time it takes for release. My biggest wall clock bottleneck right now is RL time.<\/span><\/p>\n<p><strong>Eiso Kant [01:39:25]:<\/strong><span> Right? And it\u2019s just because I can\u2019t scale it up further because I can\u2019t add more GPUs to it because of that batch size constraint. There\u2019s a really cool, blog post that just came out that was showing, RL done in even lower precision than any of us are doing. I thought this was really cool. So just what date is it today? We\u2019re on July 15, so this came out five days ago. and I thought this was very cool. I think, lower precision RL, while keeping it stable, we\u2019re, we\u2019re still doing this in FP8, and so, I was excited to see them sharing this work and bringing it out. it\u2019s definitely something that I\u2019m excited to be doing once we move to Blackwell GPUs.<\/span><\/p>\n<p><strong>Swyx [01:40:05]:<\/strong><span> But yeah, cool. Part of open research, you take and you give.<\/span><\/p>\n<p><strong>Eiso Kant [01:40:08]:<\/strong><span> Exactly. Yeah.<\/span><\/p>\n<p><strong>Swyx [01:40:10]:<\/strong><span> I\u2019ll just quickly mention, there was a paper that did a ablation on, levels of quantization, and they roughly concluded that four bit was the sweet spot. But I don\u2019t remember<\/span><\/p>\n<p><strong>Eiso Kant [01:40:20]:<\/strong><span> This was just a couple of years ago, right? I think I remember this.<\/span><\/p>\n<p><strong>Swyx [01:40:22]:<\/strong><span> I think one year.<\/span><\/p>\n<p><strong>Eiso Kant [01:40:23]:<\/strong><span> One year, okay.<\/span><\/p>\n<p><strong>Swyx [01:40:24]:<\/strong><span> But like, I\u2019m like, okay, maybe NVFP4 is it. You can\u2019t really&#8211; Like, the lowest you can go is ternary.<\/span><\/p>\n<p><strong>Eiso Kant [01:40:30]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Swyx [01:40:30]:<\/strong><span> That\u2019s it. Like, there\u2019s not that many.<\/span><\/p>\n<p><strong>Eiso Kant [01:40:32]:<\/strong><span> Well, I mean, there\u2019s, there\u2019s, there\u2019s still quite a difference between NVFP4 and four bit, right, in terms of what\u2019s, what\u2019s possible. But I think NVFP4 is, underrated in terms of what it is. I\u2019m, I\u2019m quite excited that &#8211; when it came out, it\u2019s, just getting that extra, like, that trade-off between range,<\/span><\/p>\n<p><strong>Swyx [01:40:51]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [01:40:51]:<\/strong><span> Is very cool.<\/span><\/p>\n<p><strong>Swyx [01:40:52]:<\/strong><span> A couple quick closing questions.<\/span><\/p>\n<p><strong>Vibhu [01:40:54]:<\/strong><span> I have a quick one.<\/span><\/p>\n<p><strong>Swyx [01:40:55]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Vibhu [01:40:55]:<\/strong><span> Okay, quick question back to technical side. So any big takeaways from XS 2.1 medium to training the new small, just general in terms of training models? You mentioned a lot in the earlier discussion about, okay, in training, there\u2019s a lot you can squeeze out, right? You can learn a lot more from the web. at the same time, you took 30B and scaled it up to 120B, right? is there any gating on how small is too small? So I\u2019m, I\u2019m just gonna ramble for a bit. I\u2019ll come to a question at the end. But, part of Carpathy\u2019s thesis was cognitive core, right? We\u2019ve seen Vipe Thinker, Nanbase, 3B, 4Bs that reason a lot, and then, the idea is you offload to a different model for the work. This, these are small reasoning models. So have you found anything interesting in model sizes, like 20, 30Bs on device, 100Bs on single GPU? can you squeeze out more there?<\/span><\/p>\n<p><strong>Eiso Kant [01:41:56]:<\/strong><span> There\u2019s a lot more to squeeze out. like, I think, not to make too many forward promises, but I think we can squeeze a lot more out of the XS size as well. and I think we learned a lot during S training that will allow us to improve XS, like, size even further. And I think already since then we have learned things that could have made S even better. I think there is a lot more still for, like, our space to squeeze out of models much smaller. I don\u2019t think that\u2019s an argument against scaling. It\u2019s just an, And one, by the way, and I think this is a nice thing that, it\u2019s really&#8211; it\u2019s not very helpful to have, a post-training recipe for a smaller model and try to apply it to a bigger model.<\/span><\/p>\n<p><strong>Vibhu [01:42:38]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:42:38]:<\/strong><span> It just, in all cases, you\u2019re gonna have to rethink most of the recipe. But, recipe for post-training for a bigger model applied to a smaller model is almost always just a really good, like, improvement and baseline. You can still tweak it more, but I don\u2019t think that\u2019s necessarily, like, obvious. and so &#8211; once you make your bigger models better, you often have a quick lever to quickly improve your smaller models again. but will we be able to squeeze a lot more out of smaller models? Laguna S gave me a lot of confidence that I think we can. and I think it\u2019s around that discussion we had earlier about that it\u2019s about the behaviors, not necessarily the raw intelligence, that you\u2019re trying to improve the models for.<\/span><\/p>\n<p><strong>Vibhu [01:43:23]:<\/strong><span> And that\u2019s on all axes of, There\u2019s like an axis of how long a model will reason, so how long can it stay agentic, then there\u2019s also efficiency, right? You wanna ideally push on both. And the thing to clarify you guys aren\u2019t doing right now, which we do see at Frontier Labs, is the distillation, right? You have a big model that you don\u2019t really ship to users, and what you put out for inference is typically distilled from that, which gets you quite a bit of gains, right?<\/span><\/p>\n<p><strong>Eiso Kant [01:43:50]:<\/strong><span> Look, I think it\u2019s, it\u2019s something we don\u2019t do right now because of, like, why we\u2019re also, like, building these models, right? These models are for us part of our research path. So we\u2019ve, Laguna Medium was much larger than the last two models that, this one and last one that we\u2019ve released and we\u2019ve trained even bigger models in the past. So there is the engineering component of, like, a bigger model and every order of magnitude size, you\u2019ll learn new things in training about stability. But at smaller model sizes, you are able to just iterate a lot quicker, like internally, right, on your research. And so, for us, distilling down to a smaller model doesn\u2019t serve the purpose. These models are. It\u2019s not the right term, but to us they\u2019re dual purpose models. They are progress for us to weigh to see did we improve in the model factory and something to put out into the world. and so that\u2019s why we don\u2019t do it. We\u2019ve done distillation experiments, and there\u2019s, like, really cool things you can do, and I think if you have lots of user data, then, you can go even further, right, in that. But I think there\u2019s something to be said in having a quick cadence of models trained end from scratch so that you as a research organization can learn the lessons and not wait. That was one of the big lessons we learned over the years when we used to have a much<\/span><\/p>\n<p><strong>Eiso Kant [01:45:09]:<\/strong><span> Longer cadence between model trainings, like six months, and we would train just, like, a big model, wait six months, train another bigger model. you would be compounding so many changes of improvements That at the by the time you\u2019re training your next model, it\u2019s a bit of a soup, and you don\u2019t really know what ingredients led to the outcomes. So when you are training far more frequently models, and this holds true for both post-training, and from training from scratch, you are much more able to get an understanding of what led to the improvements. and I think that\u2019s important. Like, ultimately, we are all still. There is no true science yet of, deep learning for large language models. but we are all, I think, trying to gain insights from our experiments because it\u2019s those insights that lead to scaling laws, that lead to the improvements that allow us to be, again, more compute efficient and get more capabilities.<\/span><\/p>\n<p><strong>Swyx [01:46:02]:<\/strong><span> Yeah. amazing. I was gonna end off with a little bit more history. you spent some time looking at, metrics for engineering team productivity. How do you think about engineering team productivity today?<\/span><\/p>\n<p><strong>Eiso Kant [01:46:14]:<\/strong><span> I mean, it\u2019s wild, right? I mean, it\u2019s the, it\u2019s like the golden age. Like, it\u2019s the fact that you can just take an idea and build something by waiting overnight for an agent to do the work.<\/span><\/p>\n<p><strong>Eiso Kant [01:46:26]:<\/strong><span> I don\u2019t know. To<\/span><\/p>\n<p><strong>Swyx [01:46:27]:<\/strong><span> Like, how do you measure when.<\/span><\/p>\n<p><strong>Swyx [01:46:28]:<\/strong><span> \u2018cause you literally in a theory<\/span><\/p>\n<p><strong>Eiso Kant [01:46:30]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Swyx [01:46:30]:<\/strong><span> You\u2019re doing this, right?<\/span><\/p>\n<p><strong>Eiso Kant [01:46:32]:<\/strong><span> Look, I think It\u2019s a good question. It\u2019s one I haven\u2019t thought about in a long time.<\/span><\/p>\n<p><strong>Swyx [01:46:36]:<\/strong><span> But, you\u2019re qual- you\u2019re pretty qualified to do it.<\/span><\/p>\n<p><strong>Eiso Kant [01:46:38]:<\/strong><span> No, I\u2019m gonna. &#8211; No, it\u2019s a fair point. Let me take a second to think about it. Look, ultimately, what is code, what is software, what is engineering is to go from something that is valuable for an end user or sets of end users, like an idea, an extra, a bug fix, a feature, to, like, delivering that value. And I think what we\u2019re doing with these models becoming more capable is that we are massively like, both cutting out middlemen and compressing the time that it takes to deliver that value. And ultimately, that iteration cycle for any startup or any company is what allows you to win, right? If you\u2019re able to solve a bug in two hours versus it staying in the back log for three weeks, if you\u2019re able to, like, be on a customer call and learn, hey, if this feature existed, it would, like, they\u2019d be willing to pay more, and it\u2019s more valuable to them, and you ship it in a week instead of in a month. And so I think ultimately, maybe the same things that we looked at years ago LLM still apply, and it\u2019s just the notion of cycle time. But in this case, it\u2019s lead time from the moment you have a valuable thing that you are looking to do for someone to the moment that it\u2019s shipped to them. Every other metric is ultimately a leading indicator for that lagging indicator, right? It doesn\u2019t matter if you\u2019re looking at amounts of code, PR, reviews, all of these things. And so I think in this case, we are starting to move so quickly in some of these things that we can just sit back and look at what was traditionally the lagging indicator. We just named it the lead time from traditionally ticket to, like, an end result. what I would look at in this new world, that maybe we didn\u2019t think about before is how much can a single person do with that,<\/span><\/p>\n<p><strong>Eiso Kant [01:48:22]:<\/strong><span> Right? One of the most, like, if you look at AI native companies, they\u2019re not designed like the engineering orgs of, LLM age. They\u2019re designed with often just the builder, right? and as close to the customer to the ability to ship. there isn\u2019t necessarily a huge team in between that sits there. And I think that is, I think, is exciting, like organizations where a single IC can just, get much closer to that. So I would look at From where the value sits that\u2019s identified to the moment it\u2019s shipped and how many people are involved in that. And you want the amount of people involved in that to be less, and you want the time end to be shorter.<\/span><\/p>\n<p><strong>Swyx [01:49:05]:<\/strong><span> Okay. is there a way to eval that when you\u2019re, interviewing somebody?<\/span><\/p>\n<p><strong>Eiso Kant [01:49:12]:<\/strong><span> Oof.<\/span><\/p>\n<p><strong>Swyx [01:49:13]:<\/strong><span> \u2018Cause that is,<\/span><\/p>\n<p><strong>Eiso Kant [01:49:14]:<\/strong><span> Look,<\/span><\/p>\n<p><strong>Swyx [01:49:14]:<\/strong><span> The most compressed version.<\/span><\/p>\n<p><strong>Eiso Kant [01:49:17]:<\/strong><span> I think the common answer to this is agency.<\/span><\/p>\n<p><strong>Swyx [01:49:20]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:49:20]:<\/strong><span> How much agency does a person have? I think in the age of AI getting more capable, agency becomes probably one of the most important qualities for anyone. and I think agency is something you can look for in, what people have done in the past because agency is something that if you have it, you are demonstrating it, right? No one has just agency and is sitting back and not, like, exercising it. The whole definition of it is that it\u2019s exercised. And so understanding, like, what were things that people did in their lives, in their professional and their personal projects that showed agency and, your personal backstory shows a ridiculous amount of agency.<\/span><\/p>\n<p><strong>Swyx [01:49:56]:<\/strong><span> Oh, dear.<\/span><\/p>\n<p><strong>Eiso Kant [01:49:58]:<\/strong><span> Like, I think that is ultimately it. It\u2019s the Silicon Valley, quota the, of the last, year and a half or so is like you can just do things, right?<\/span><\/p>\n<p><strong>Swyx [01:50:06]:<\/strong><span> Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:50:07]:<\/strong><span> That- that\u2019s I think what you\u2019re looking for.<\/span><\/p>\n<p><strong>Swyx [01:50:08]:<\/strong><span> I think then aligning high agency people is very hard because they all wanna go their own way. That\u2019s the whole point, right?<\/span><\/p>\n<p><strong>Eiso Kant [01:50:15]:<\/strong><span> They&#8211; Yeah, but I think the notion &#8211; Like, I think the notion of a good leader, right, in an organization is to be able to bring people together around, like, a common outcome. And I think what you wanna do with anyone who\u2019s high agency&#8211; I feel very lucky I\u2019ve got an organization with incredibly high agency people. Like, I mean, I\u2019m not the one who built the model, right? I cannot stress this enough. Like, it\u2019s the team that, like, achieved this, and it\u2019s a team that is incredibly high agency. And so if you look at, like, what does it take to bring that together, it\u2019s, it\u2019s ultimately a common goal and a common set of boundaries. Because if you allow to just go, \u201cYou can do everything,\u201d you become an exploration algorithm. And this is what we see in big tech, right? In research, in big tech, everything is an exploration algorithm. Everyone can do anything as long as &#8211; And then it becomes political about gathering the resources. So when you say, \u201cThis is our common goal, and these are the boundaries that we\u2019ve set,\u201d right? \u201cWe\u2019re not multimodal. We focus on RL.\u201d Like, we do these things, and you\u2019re upfront with people before they join the company, you get a lot of agency. You can run where you want, but these are the places where we<\/span><\/p>\n<p><strong>Swyx [01:51:24]:<\/strong><span> Yeah, lanes<\/span><\/p>\n<p><strong>Eiso Kant [01:51:24]:<\/strong><span> This doesn\u2019t make&#8211; This is the lanes<\/span><\/p>\n<p><strong>Swyx [01:51:25]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [01:51:25]:<\/strong><span> That makes sense. I think it gets the best out of people because, like, innovation comes from constraints.<\/span><\/p>\n<p><strong>Eiso Kant [01:51:34]:<\/strong><span> We did this with relatively little compute and relatively little money compared to some of, like, the others that are out there. and I\u2019ve thought back on that quite a bit recently and thought, it was a good thing Because those constraints forced us to become much better in certain other axes that might&#8211; others might have not, right? We purchased relatively little external data.<\/span><\/p>\n<p><strong>Swyx [01:52:01]:<\/strong><span> I was gonna ask about that. Yeah.<\/span><\/p>\n<p><strong>Eiso Kant [01:52:02]:<\/strong><span> Exactly, right. That was a constraint. but it\u2019s a constraint that pushed us to move on other areas to improve. And like, and there\u2019s lots of versions of that. So I think high agency people, you wanna empower, you wanna get them really excited about what they\u2019re doing, but you also wanna say, \u201cHey, if you join this mission, this is the outcome I need you to achieve. But these are the places that we don\u2019t go, and maybe if you care about those places, go somewhere else.\u201d<\/span><\/p>\n<p><strong>Swyx [01:52:26]:<\/strong><span> Yeah. Great. last call to action, who are you hiring?<\/span><\/p>\n<p><strong>Eiso Kant [01:52:31]:<\/strong><span> We are hiring on every possible role in applied research and engineering in the company. so from<\/span><\/p>\n<p><strong>Swyx [01:52:36]:<\/strong><span> Yeah<\/span><\/p>\n<p><strong>Eiso Kant [01:52:36]:<\/strong><span> Training all the way to evals to post-training architecture. Like, we are still in a world where, individuals can have massive impact. And I think our pitch to join us, it\u2019- We spoke a lot about the mission, how we think about things, but I think we are one of the places where it\u2019s the highest ratio to individual to impact, Right? Less than 70 people built this model. Less than 115 between engineering and researchers, like, together did this effort, and that\u2019s a very broad definition \u2018cause I put myself in the 115 list.<\/span><\/p>\n<p><strong>Eiso Kant [01:53:08]:<\/strong><span> And so being able to do this work on a mission that you\u2019re aligned with, and you can have that &#8211; every individual still has huge impact. And I think<\/span><\/p>\n<p><strong>Swyx [01:53:18]:<\/strong><span> And being able to publish, being able to open<\/span><\/p>\n<p><strong>Eiso Kant [01:53:20]:<\/strong><span> It\u2019<\/span><\/p>\n<p><strong>Swyx [01:53:20]:<\/strong><span> Open source the model.<\/span><\/p>\n<p><strong>Eiso Kant [01:53:21]:<\/strong><span> Yeah, look, all of those things are part of that. But I think ultimately, when you can today pick between joining a very large foundation model company But you are one of many.<\/span><\/p>\n<p><strong>Eiso Kant [01:53:35]:<\/strong><span> And not by any fault of them, but just by definition, the denominator has become really big. And our denominator is quite small, and so the level of impact you get to have is really high. And I think ultimately all of us, the most incredible high agency people I know, what are they optimizing for? They\u2019re optimizing for impact. they\u2019re optimizing for impact, and am I aligned with the mission? And if today you heard about the mission and aligned and you\u2019re optimizing for impact, I think we\u2019re a really good place to join.<\/span><\/p>\n<p><strong>Swyx [01:54:05]:<\/strong><span> Okay.<\/span><\/p>\n<p><strong>Eiso Kant [01:54:05]:<\/strong><span> Awesome.<\/span><\/p>\n<p><strong>Swyx [01:54:05]:<\/strong><span> I think we end it there. That\u2019s a fantastic statement. You did amazing on four hours of sleep.<\/span><\/p>\n<p><strong>Eiso Kant [01:54:11]:<\/strong><span> Thank you, guys.<\/span><\/p>\n<p><strong>Swyx [01:54:12]:<\/strong><span> So, podcast eval, definitely approved.<\/span><\/p>\n<p><strong>Eiso Kant [01:54:14]:<\/strong><span> Appreciate it. I literally wrote it down. My eyes are, like, starting to go like this. I\u2019m like, \u201cPhew.\u201d<\/span><\/p>\n<p><strong>Swyx [01:54:17]:<\/strong><span> We\u2019ll let you go. We\u2019ll let you go back.<\/span><\/p>\n<p><strong>Eiso Kant [01:54:19]:<\/strong><span> It was good to see you guys.<\/span><\/p>\n<p><strong>Swyx [01:54:19]:<\/strong><span> Thank you for setting this up. We wanted to get this in because we think it\u2019s a great model.<\/span><\/p>\n<p><strong>Eiso Kant [01:54:23]:<\/strong><span> Appreciate it.<\/span><\/p>\n<p><strong>Swyx [01:54:23]:<\/strong><span> I think a great story to tell. Thank you.<\/span><\/p>\n<\/div>\n<p><a href=\"https:\/\/www.latent.space\/p\/poolside?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign\/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines\u2019 recent release nearly 10 times [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22682,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22681","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22681","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22681"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22681\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22682"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22681"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22681"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22681"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}