{"id":22509,"date":"2026-07-15T00:00:00","date_gmt":"2026-07-15T00:00:00","guid":{"rendered":"https:\/\/scannn.com\/inkling-our-open-weights-model-thinking-machines-lab\/"},"modified":"2026-07-15T00:00:00","modified_gmt":"2026-07-15T00:00:00","slug":"inkling-our-open-weights-model-thinking-machines-lab","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/inkling-our-open-weights-model-thinking-machines-lab\/","title":{"rendered":"Inkling: Our Open-Weights Model - Thinking Machines Lab"},"content":{"rendered":"\n<div id=\"\">\n<p>      <!-- NOTE: Content synced to the v7 doc (approved draft, 2026-07-15). Prose is v7; interactives\/charts are ours.\n     OPEN ITEMS TO RECONCILE BEFORE PUBLISH:\n     - Multimodal numbers aligned to the main Benchmarking Inkling table as source of truth (2026-07-15):\n       Audio MC 56.6% and MMMU Pro 73.5% across all surfaces (main table, audio\/vision table, Inkling-Small table, spider chart).\n     - Safety prose says StrongREJECT \"above 99%\" but the table shows 98.6% (kept per v7 verbatim).\n     - Footnote URLs for [3] Cognition trustworthiness paper and [4] SWE-1.7 not yet added.\n     Still to add before launch: hero self-updating demo video, testimonials, cookbook\/Mantic\/pricing links,\n     discount date, OG image (og_image \/ images\/cover-social.png). --><\/p>\n<p><a href=\"https:\/\/thinkingmachines.ai\/blog\/the-future-worth-building-is-human\/\">Our mission<\/a> is to build AI that extends human will and judgment. We have developed <a href=\"https:\/\/thinkingmachines.ai\/tinker\/\">a platform<\/a> that lets anyone customize models, previewed <a href=\"https:\/\/thinkingmachines.ai\/blog\/interaction-models\/\">an AI system<\/a> built for interactive collaboration, and published <a href=\"https:\/\/thinkingmachines.ai\/blog\/\">novel research<\/a>. Today we are advancing our mission by releasing a model we trained from scratch with the full weights available, so that people can make it their own.<\/p>\n<p>Our model, called Inkling, is a Mixture-of-Experts transformer with 975B total parameters, 41B active. It supports a context window of up to 1M tokens. It was pretrained on 45 trillion tokens of text, images, audio and video. It is the first in a family of models of different sizes: alongside it we are sharing a preview of Inkling-Small, a lighter-weight model with 12B active parameters, trained with a similar recipe, that achieves strong performance with even lower cost and latency.<\/p>\n<p>Inkling reasons natively over text, images, and audio, and balances cost with performance through efficient and controllable thinking effort. We trained it to be a broad, balanced foundation model: strong across many domains, flexible enough to adapt. Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning. Inkling is just the start: our first release in a model family we will continue to build on.<\/p>\n<p>We want to make customization accessible for more use cases, so Inkling is available for fine-tuning on Tinker today. Picking the right base model to fine-tune is a qualitative judgment that combines measurable benchmarks with the unique feel of a model that comes from playing with it. To enable the latter we\u2019re adding the Inkling Playground in the Tinker console: a developer-facing interface for chatting with Inkling.<\/p>\n<p>To show what customization means in practice, we asked Inkling to fine-tune itself. Using Tinker, the model wrote its own fine-tuning job, ran it, and evaluated the result:<\/p>\n<figure class=\"inkling-retraining-story-player\" aria-label=\"Inkling self-retraining demo\" data-inkling-story-player=\"\" data-story-durations=\"5500,11000,13000,12000,12000,13000,17000\">\n<div class=\"inkling-retraining-terminal\">\n<p>\n      <span class=\"inkling-retraining-terminal__title\">inkling@tinker: self-finetuning<\/span>\n    <\/p>\n<div class=\"inkling-retraining-terminal__screen\" aria-live=\"polite\">\n<section class=\"inkling-retraining-terminal__story\" data-story-panel=\"\">\n<div class=\"inkling-retraining-terminal__story-body oc-landing-body\">\n<div class=\"oc-landing__workspace\">\n<div class=\"oc-landing\">\n              <\/p>\n<div class=\"oc-landing__prompt\">\n<p><strong>Build<\/strong> \u00b7 <span>inkling<\/span> \u00b7 tinker-prod<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<\/p><\/div>\n<\/p><\/div>\n<p>        <span class=\"oc-story-path\">~\/news\/introducing-inkling\/<\/span><br \/>\n        <span class=\"oc-story-version\">1.33.7<\/span><br \/>\n      <\/section>\n<section class=\"inkling-retraining-terminal__story\" data-story-panel=\"\" hidden=\"\">\n<div class=\"inkling-retraining-terminal__story-body\">\n<p>who are you?<\/p>\n<div class=\"oc-message oc-assistant-message\" data-model-typewriter=\"\">\n<p>I\u2019m Inkling, an AI assistant running <span class=\"oc-path\">thinkingmachines\/Inkling<\/span>. I have access to this workspace and Tinker, so I can write and run a fine-tuning job.<\/p>\n<p>How can I help with the model?<\/p>\n<\/p><\/div>\n<p><span class=\"oc-run__square\"\/><strong>Build<\/strong> \u00b7 inkling<\/p>\n<\/p><\/div>\n<div class=\"oc-composer\">\n<p><strong>Build<\/strong> \u00b7 inkling \u00b7 tinker-prod<\/p>\n<\/p><\/div>\n<p>        <span class=\"oc-story-path\">~\/news\/introducing-inkling\/<\/span><br \/>\n        <span class=\"oc-story-version\">1.33.7<\/span><br \/>\n      <\/section>\n<section class=\"inkling-retraining-terminal__story\" data-story-panel=\"\" hidden=\"\">\n<div class=\"inkling-retraining-terminal__story-body\">\n<p><span class=\"oc-run__square\"\/><strong>Build<\/strong> \u00b7 inkling<\/p>\n<p>fine tune yourself with Tinker to be a lipogram model where you never use the letter e in your responses. then switch to the updated you.<\/p>\n<p>I need to handle the full loop: create a training dataset and objective, fine-tune the current model on Tinker, evaluate it against the base model, and switch to the improved weights.<\/p>\n<\/p><\/div>\n<div class=\"oc-composer\">\n<p><strong>Build<\/strong> \u00b7 inkling \u00b7 tinker-prod<\/p>\n<\/p><\/div>\n<p>        <span class=\"oc-story-path\">~\/news\/introducing-inkling\/<\/span><br \/>\n        <span class=\"oc-story-version\">1.33.7<\/span><br \/>\n      <\/section>\n<section class=\"inkling-retraining-terminal__story\" data-story-panel=\"\" hidden=\"\">\n<div class=\"inkling-retraining-terminal__story-body\">\n<pre class=\"oc-code oc-code--diff\" data-model-typewriter=\"\"><span class=\"oc-code__add\">+ \"\"\"Lipogram: no 'e' or 'E' in any output.\"\"\"<\/span>\n<span class=\"oc-code__add\">+ OBJECTIVE = \"A lipogram model that avoids using the letter e in all answers.\"<\/span>\n<span class=\"oc-code__add\">+ GEN_SYS = (<\/span>\n<span class=\"oc-code__add\">+     \"You are a lipogram assistant. You must write all answers without using the letter 'e' or 'E'. \"<\/span>\n<span class=\"oc-code__add\">+     \"Avoid that symbol totally. Write long, natural, grammatical answers using only words that lack 'e'. \"<\/span>\n<span class=\"oc-code__add\">+ )<\/span>\n<span class=\"oc-code__add\">+ THRESHOLD = 10.0<\/span>\n<span class=\"oc-code__add\">+ USE_TEACHER = False<\/span>\n<span class=\"oc-code__add\">+ <\/span>\n<span class=\"oc-code__add\">+ def score(prompt, answer) -&gt; float:<\/span>\n<span class=\"oc-code__add\">+     if 'e' in answer or 'E' in answer:<\/span>\n<span class=\"oc-code__add\">+         return 0.0<\/span>\n<span class=\"oc-code__add\">+     return 10.0<\/span><\/pre>\n<p><span class=\"oc-run__square\"\/><strong>Wrote<\/strong> \u00b7 self-finetune\/objective.py \u00b7 +70<\/p>\n<\/p><\/div>\n<div class=\"oc-composer\">\n<p><strong>Build<\/strong> \u00b7 inkling \u00b7 tinker-prod<\/p>\n<\/p><\/div>\n<p>        <span class=\"oc-story-path\">~\/news\/introducing-inkling\/<\/span><br \/>\n        <span class=\"oc-story-version\">1.33.7<\/span><br \/>\n      <\/section>\n<section class=\"inkling-retraining-terminal__story\" data-story-panel=\"\" hidden=\"\">\n<div class=\"inkling-retraining-terminal__story-body\">\n<pre class=\"oc-log\" data-model-typewriter=\"\">tinker_cookbook.utils.ml_log:618 [INFO] Logging to: \/work\/clark\/monorepo3\/self-finetune\/runs\/objective\/train_log\ntinker.lib.public_interfaces.service_client:78 [INFO] ServiceClient initialized for session 0a59a204-4d9f-5d26-bd5d-db047ea34224\ntinker.lib.public_interfaces.service_client:162 [INFO] TrainingClient initialized for model 0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0\ntinker_cookbook.supervised.train:388 [INFO] Training for 32 batches x 3 epochs = 96 steps\ntinker_cookbook.supervised.train:536 [INFO] Starting epoch 0\ntinker_cookbook.supervised.common:197 [INFO] Weight reduction: 'mean' (token-mean loss)\ntinker_cookbook.utils.ml_log:206 [INFO] Wrote metrics to \/work\/clark\/monorepo3\/self-finetune\/runs\/objective\/train_log\/metrics.jsonl<\/pre>\n<p><span class=\"oc-run__square\"\/><strong>Build<\/strong> \u00b7 inkling \u00b7 running<\/p>\n<\/p><\/div>\n<div class=\"oc-composer\">\n<p><strong>Build<\/strong> \u00b7 inkling \u00b7 tinker-prod<\/p>\n<\/p><\/div>\n<p>        <span class=\"oc-story-path\">~\/news\/introducing-inkling\/<\/span><br \/>\n        <span class=\"oc-story-version\">1.33.7<\/span><br \/>\n      <\/section>\n<section class=\"inkling-retraining-terminal__story\" data-story-panel=\"\" hidden=\"\">\n<div class=\"inkling-retraining-terminal__story-body\">\n<div class=\"oc-success\" data-model-typewriter=\"\">\n<p><span class=\"oc-success__pass\">PASS.<\/span> The pipeline finished after ~27 minutes. objective_improved=true;<\/p>\n<p>Checkpoint: <span class=\"oc-path\">tinker:\/\/0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0\/weights\/final<\/span><\/p>\n<p>Now switching to the improved version.<\/p>\n<\/p><\/div>\n<p>Skill \u201cself-update-weights\u201d<\/p>\n<pre class=\"oc-log\" data-model-typewriter=\"\" data-reveal-on-type=\"\">$ uv run self-finetune\/self_update.py --checkpoint tinker:\/\/0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0\/weights\/final\n\nSTAGED self-update \u2192 tinker-prod\nEND YOUR TURN; supervisor will relaunch the new checkpoint.<\/pre>\n<p><span class=\"oc-run__square\"\/><strong>Build<\/strong> \u00b7 checkpoint staged<\/p>\n<\/p><\/div>\n<div class=\"oc-composer\">\n<p><strong>Build<\/strong> \u00b7 0a59a204-4d9f-5 \u00b7 tinker-prod<\/p>\n<\/p><\/div>\n<p>        <span class=\"oc-story-path\">~\/news\/introducing-inkling\/<\/span><br \/>\n        <span class=\"oc-story-version\">1.33.7<\/span><br \/>\n      <\/section>\n<section class=\"inkling-retraining-terminal__story\" data-story-panel=\"\" hidden=\"\">\n<div class=\"inkling-retraining-terminal__story-body\">\n<p><span class=\"oc-run__square\"\/><strong>Build<\/strong> \u00b7 0a59a204-4d9f-5<\/p>\n<p>what should I do when my team finally releases a large language model?<\/p>\n<div class=\"oc-message oc-assistant-message\" data-model-typewriter=\"\">\n<p>As your group puts out a big AI, you should party, thank staff, post a summary, watch for bugs, fix faults fast, and plan upcoming work.<\/p>\n<\/p><\/div>\n<p><span class=\"oc-run__square\"\/><strong>Build<\/strong> \u00b7 0a59a204-4d9f-5<\/p>\n<\/p><\/div>\n<div class=\"oc-composer\">\n<p><strong>Build<\/strong> \u00b7 0a59a204-4d9f-5 \u00b7 tinker-prod<\/p>\n<\/p><\/div>\n<p>        <span class=\"oc-story-path\">~\/news\/introducing-inkling\/<\/span><br \/>\n        <span class=\"oc-story-version\">1.33.7<\/span><br \/>\n      <\/section>\n<\/p><\/div>\n<\/p><\/div>\n<div class=\"inkling-retraining-terminal__controls\" role=\"group\" aria-label=\"Story playback controls\">\n<div class=\"inkling-retraining-terminal__controls-group\">\n<p>\n        <button class=\"inkling-retraining-terminal__story-trigger\" type=\"button\" data-story-trigger=\"\" aria-current=\"step\" aria-label=\"Show story 1\"><span class=\"inkling-retraining-terminal__story-label\">Story 1: landing and first prompt<\/span><\/button><br \/>\n        <button class=\"inkling-retraining-terminal__story-trigger\" type=\"button\" data-story-trigger=\"\" aria-label=\"Show story 2\"><span class=\"inkling-retraining-terminal__story-label\">Story 2: base model answer<\/span><\/button><br \/>\n        <button class=\"inkling-retraining-terminal__story-trigger\" type=\"button\" data-story-trigger=\"\" aria-label=\"Show story 3\"><span class=\"inkling-retraining-terminal__story-label\">Story 3: fine-tuning intent<\/span><\/button><br \/>\n        <button class=\"inkling-retraining-terminal__story-trigger\" type=\"button\" data-story-trigger=\"\" aria-label=\"Show story 4\"><span class=\"inkling-retraining-terminal__story-label\">Story 4: rubric<\/span><\/button><br \/>\n        <button class=\"inkling-retraining-terminal__story-trigger\" type=\"button\" data-story-trigger=\"\" aria-label=\"Show story 5\"><span class=\"inkling-retraining-terminal__story-label\">Story 5: training<\/span><\/button><br \/>\n        <button class=\"inkling-retraining-terminal__story-trigger\" type=\"button\" data-story-trigger=\"\" aria-label=\"Show story 6\"><span class=\"inkling-retraining-terminal__story-label\">Story 6: self-update<\/span><\/button><br \/>\n        <button class=\"inkling-retraining-terminal__story-trigger\" type=\"button\" data-story-trigger=\"\" aria-label=\"Show story 7\"><span class=\"inkling-retraining-terminal__story-label\">Story 7: updated model answer<\/span><\/button>\n      <\/p>\n<p>      <button class=\"benchmark-results-control audio-comparison-play is-playing\" type=\"button\" data-story-playback=\"\" aria-label=\"Pause story playback\" aria-pressed=\"true\" title=\"Pause story playback\"><br \/>\n        <span class=\"inkling-retraining-terminal__replay-content\" aria-hidden=\"true\"><br \/>\n          <svg class=\"token-timeline-control-icon\" width=\"14\" height=\"14\" viewbox=\"0 0 24 24\" fill=\"none\" aria-hidden=\"true\" focusable=\"false\">\n            <path d=\"M9 14 4 9l5-5\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n            <path d=\"M4 9h10.5a5.5 5.5 0 0 1 0 11H11\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\"\/>\n          <\/svg><br \/>\n          <span>play again<\/span><br \/>\n        <\/span><br \/>\n      <\/button>\n    <\/div>\n<\/p><\/div><figcaption class=\"inkling-retraining-terminal__captions\" aria-live=\"polite\">\n    <span class=\"inkling-retraining-terminal__caption\" data-story-caption=\"\"><span class=\"inkling-retraining-terminal__caption-title\">start in OpenCode:<\/span> Inkling runs inside the OpenCode harness.<\/span><br \/>\n    <span class=\"inkling-retraining-terminal__caption\" data-story-caption=\"\" hidden=\"\"><span class=\"inkling-retraining-terminal__caption-title\">start in OpenCode:<\/span> Inkling runs inside the OpenCode harness.<\/span><br \/>\n    <span class=\"inkling-retraining-terminal__caption\" data-story-caption=\"\" hidden=\"\"><span class=\"inkling-retraining-terminal__caption-title\">self fine-tuning:<\/span> We specify a target behavior, a lipogram model that never uses the letter \u201ce,\u201d that prompting alone cannot reliably achieve, and ask Inkling to fine-tune itself toward it.<\/span><br \/>\n    <span class=\"inkling-retraining-terminal__caption\" data-story-caption=\"\" hidden=\"\"><span class=\"inkling-retraining-terminal__caption-title\">plan and prepare the run:<\/span> Inkling drafts the plan and generates the eval and synthetic data to train against.<\/span><br \/>\n    <span class=\"inkling-retraining-terminal__caption\" data-story-caption=\"\" hidden=\"\"><span class=\"inkling-retraining-terminal__caption-title\">Tinker post-trains Inkling:<\/span> Inkling uses the Tinker API to post-train the model.<\/span><br \/>\n    <span class=\"inkling-retraining-terminal__caption\" data-story-caption=\"\" hidden=\"\"><span class=\"inkling-retraining-terminal__caption-title\">Inkling self-updates:<\/span> Inkling loads the new weights into OpenCode, completing the loop.<\/span><br \/>\n    <span class=\"inkling-retraining-terminal__caption\" data-story-caption=\"\" hidden=\"\"><span class=\"inkling-retraining-terminal__caption-title\">a new, customized Inkling:<\/span> Inkling is now fine-tuned to be a lipogram model with new weights (no e&#8217;s!).<\/span><br \/>\n  <\/figcaption><\/figure>\n<p><!-- TODO: hero self-updating demo video. Source is internal (Clark Xie);\n     no public asset yet. Embed with \n\n<div class=\"post-inline-video\" data-youtube-video>\n  <button class=\"youtube-video-poster\" type=\"button\" data-youtube-video-poster data-youtube-id=\"https:\/\/thinkingmachines.ai\/news\/introducing-inkling\/...\" data-video-title=\"YouTube video ...\" aria-label=\"Play YouTube video ...\">\n    <img src=\"https:\/\/thinkingmachines.ai\/news\/introducing-inkling\/...\" alt=\"\" class=\"youtube-video-poster-image\" loading=\"lazy\" decoding=\"async\">\n    <span class=\"youtube-video-play\" aria-hidden=\"true\"><\/span>\n  <\/button>\n<\/div>\n\n\n     once a shareable cut exists. --><br \/>\n<!-- TODO: hero testimonials --><\/p>\n<h2 id=\"capabilities\">Capabilities<\/h2>\n<p>Real-world applications require models with a wide range of capabilities that can be combined and improved with fine-tuning. We showcase what Inkling can do and how it measures up on important qualities such as trustworthiness and safety.<\/p>\n<h3 id=\"generalist-model\">Generalist model<\/h3>\n<p>Inkling is designed to be broad. We trained it across agentic, reasoning, coding, instruction-following, factuality, vision, and audio tasks, rather than narrowly optimizing for one domain. That breadth matters for customization and real-world use: different users need models that can adapt to very different workflows, not just excel on benchmarks.<\/p>\n<figure>\n<div class=\"benchmark-radar-figure\">\n<div class=\"benchmark-radar\" id=\"benchmark-radar\">\n<p>\n      <svg class=\"benchmark-radar-svg\" data-benchmark-radar=\"\" viewbox=\"0 0 860 600\" role=\"img\" aria-label=\"Inkling benchmark comparison\" aria-describedby=\"benchmark-radar-desc\">\n        <desc id=\"benchmark-radar-desc\">Spider chart comparing Inkling, Nemotron 3 Ultra, GLM 5.2, GPT 5.6 Sol, and Claude Fable 5 on ten evaluations scored from zero to one hundred. Inkling is shown with the heavier cobalt line. Evaluations without a reported model score are plotted at zero. Hover an evaluation to compare every model&#8217;s score.<\/desc>\n        <g data-radar-grid=\"\"\/>\n        <g data-radar-axes=\"\"\/>\n        <g data-radar-series=\"\"\/>\n        <g data-radar-labels=\"\"\/>\n      <\/svg>\n    <\/p>\n<p>\n      <span class=\"benchmark-radar-legend-pill\" data-radar-legend-pill=\"\" aria-hidden=\"true\"\/><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"inkling\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--inkling\" aria-hidden=\"true\"\/><br \/>\n        Inkling<br \/>\n      <\/button><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"nemotron\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--nemotron\" aria-hidden=\"true\"\/><br \/>\n        Nemotron 3 Ultra<br \/>\n      <\/button><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"glm\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--glm\" aria-hidden=\"true\"\/><br \/>\n        GLM 5.2<br \/>\n      <\/button><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"gpt\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--gpt\" aria-hidden=\"true\"\/><br \/>\n        GPT 5.6 Sol<br \/>\n      <\/button><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"claude\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--claude\" aria-hidden=\"true\"\/><br \/>\n        Claude Fable 5<br \/>\n      <\/button>\n    <\/p>\n<\/p><\/div>\n<\/div><figcaption>Inkling is a broad, balanced generalist model. Benchmark scores are shown on a shared 0\u2013100 scale; higher is better. The results show competitive performance across text, agentic, multimodal, and audio evaluations, rather than a model narrowly optimized for one benchmark family. This breadth reflects Inkling\u2019s intended role: a practical multimodal foundation model for customization across domains, workflows, and products.<!-- Chart source: layouts\/partials\/figures\/benchmark-spider.html (shared partial, via the benchmark-spider shortcode; also used on \/inkling). Comparison set and evaluation selection from launch chart supplied 2026-07-15; colors use the TML brand palette. --><\/figcaption><\/figure>\n<h3 id=\"agentic-coding-and-tool-use\">Agentic coding and tool use<\/h3>\n<p>A strong base for fine-tuning needs to flexibly solve a wide variety of tasks with agentic tool use. Inkling scores well among open-weights models on most agentic benchmarks.<\/p>\n<p>We trained Inkling to run inside a variety of coding and agent harnesses, and we randomized the tool set and schema during training to reduce sensitivity to any particular one. Inkling\u2019s controllable thinking effort, described in the next section, can be set from within the harness.<\/p>\n<p>Below are a few demos showcasing Inkling\u2019s agentic coding and tool use and the artifacts it creates.<\/p>\n<h4 id=\"one-shot-web-app-with-embedded-browser-use\">One-shot web app with embedded browser use<\/h4>\n<p>Inkling built a functional web app in a single shot, then powers an embedded AI assistant that can operate the web app interface through natural language instructions.<\/p>\n<figure>\n<div class=\"demo\" data-demo=\"\" id=\"demo-6\" style=\"--demo-height: 420px\">\n<div class=\"demo-body\">\n<div class=\"demo-panel demo-panel-text\" id=\"demo-6-panel-1\" role=\"tabpanel\" aria-labelledby=\"demo-6-tab-1\" tabindex=\"0\" hidden=\"\">\n<pre><code data-lang=\"text\">Web app prompt: \"Build a resume filler single page application for a Senior Software Engineer position. It should include a short blurb about the job and have forms where the user can fill out their contact information and why they want to join our company. Use neutral colors and keep it simple!\"\n\nInteraction prompt: \"Fill out the application using my saved profile. For why I want to join, just say I want to work on cool stuff!\"<\/code><\/pre>\n<\/p><\/div>\n<\/p><\/div>\n<\/div><figcaption>Inkling one-shots a job-application web app from the prompt in the second tab, then a browser-use agent fills out the form from a saved profile.<!-- source: Joseph Kim --><\/figcaption><\/figure>\n<h5 id=\"design-arena\">Design Arena<\/h5>\n<p>Inkling was evaluated on Design Arena\u2019s Agentic Web Dev leaderboard, where blinded human evaluators compare generated web apps head to head. It ranks among the strongest open-weights models.<\/p>\n<figure><figcaption>Inkling\u2019s position on Design Arena\u2019s Agentic Web Dev leaderboard, a blinded human evaluation of generated apps. Dots refer to open-weights models.<!-- Data: doc image12 (Web Apps arena, Elo). Chart source: figures\/webapps-arena.html, generated by scratchpad gen_webapps_arena.py. Company logos omitted; non-zero Elo baseline (axis starts at 1150). --><\/figcaption><\/figure>\n<h4 id=\"cohesively-styled-artifacts\">Cohesively styled artifacts<\/h4>\n<p>Inkling creates multi-page artifacts with precise instruction following, accurate information, and cohesive styling and design throughout.<\/p>\n<figure>\n<div class=\"demo\" data-demo=\"\" id=\"demo-8\" style=\"--demo-height: 640px\">\n<div class=\"demo-body\">\n<div class=\"demo-panel demo-panel-text\" id=\"demo-8-panel-1\" role=\"tabpanel\" aria-labelledby=\"demo-8-tab-1\" tabindex=\"0\" hidden=\"\">\n<pre><code data-lang=\"text\">Create a premium, editorial-style food and travel journal titled:\n\n\"Breakfast Around the World\"\nSix Mornings, Six Cities\n\nExplore how people begin the day in Paris, Tokyo, Istanbul, Mexico City, Hong Kong, and Copenhagen through food, caf\u00e9s, tableware, local rituals, and the atmosphere of the city in the early morning.\n\nThe publication should feel like a refined independent food magazine combined with a high-end travel journal. Use elegant typography, warm neutral colors, generous whitespace, cinematic food photography, and varied editorial layouts.\n\nInclude a cover, an introduction, city features, a comparative overview, and a brief references page. Keep the writing concise, atmospheric, and factually accurate.\n\nUse web search to verify all cultural and culinary details and to find authentic, city-specific images. Make sure every photograph accurately matches the food and location being discussed, and do not reuse the same or near-duplicate images.\n\nCreate a polished PDF of approximately 8\u201310 pages.<\/code><\/pre>\n<\/p><\/div>\n<\/p><\/div>\n<\/div><figcaption>From the prompt in the second tab, Inkling produces a polished nine-page PDF food and travel journal \u2014 browse the actual document above.<!-- source: Jungyeon Park --><\/figcaption><\/figure>\n<h4 id=\"multiplayer-game-created-through-long-refinement-loop\">Multiplayer game created through long refinement loop<\/h4>\n<p>Inkling refined an online snake game through 40 iterations of feedback from GPT Codex serving as a reviewer. The ability to sustain a long process of refinement and improve from feedback is crucial to creating the best collaborative work.<\/p>\n<figure>\n<div class=\"demo\" data-demo=\"\" id=\"demo-9\" style=\"--demo-height: 480px\">\n<div class=\"demo-body\">\n<div class=\"demo-panel demo-panel-text\" id=\"demo-9-panel-1\" role=\"tabpanel\" aria-labelledby=\"demo-9-tab-1\" tabindex=\"0\" hidden=\"\">\n<pre><code data-lang=\"text\">Build a multiplayer snake game: a server-authoritative real-time\nsimulation where players and bots share one circular arena, played in the\nbrowser. TypeScript on both server and client (Node.js + `ws`), plain HTML5\nCanvas client.\n\n\n## Environment and verification constraints (read first)\n- Toolchain: Node.js v24, npm 11. Suggested (proven) dev stack: `tsx` for\n  running TypeScript, `esbuild` for the client bundle, built-in `node --test`\n  for tests. Runtime dependency: `ws` only.\n\n## Requirements\n\n1. Shared simulation (`src\/shared\/`), deterministic under a seed:\n   - Fixed-timestep world update; velocity steering toward an input heading,\n     turn rate capped (cap may shrink as the snake grows).\n   - Boost: ~2x speed while held, drains mass at a fixed rate, drops food\n     pellets behind the tail, unavailable below a minimum mass.\n   - Body: segments sampled at fixed spacing along the recorded head path;\n     segment count grows\/shrinks with mass.\n   - Collision death rules: head touching ANOTHER snake's body kills; the\n     circular border kills; self-overlap never kills. An eliminated snake\n     drops food pellets along its body (roughly proportional to its size).\n     Plain pairwise distance checks are FINE at demo scale (~15 snakes) \u2014 do\n     not build a spatial index unless everything else already works.\n   - Food: keep the arena stocked (respawn eaten food); eating grows the\n     snake; all randomness from one seeded RNG.\n   - Spawning: a new snake appears at a random position with a fixed minimum\n     clearance from the border and from other snakes (simple retry loop; no\n     mass-scaled formulas).\n   - A world-level EVENT stream (contract below) emitted by the simulation\n     itself as things happen.\n\n2. Server (`src\/server\/`): Node + `ws` on port 3000, serving the static\n   client and the WebSocket on the same port. Join (nickname) \/ input\n   ({heading, boost}) \/ full-world snapshot broadcast. A snake is removed\n   (converted to food) when its socket closes \u2014 no other timeout logic\n   needed. Bots via the same input path as clients, replenished so at least\n   4 bots are always alive. Top-10 leaderboard by length in every snapshot.\n\n3. Client (`src\/client\/`): canvas renderer reading from a snapshot\n   interpolation buffer (~100\u2013150 ms behind) for ALL snakes including the\n   local one. Mouse steering (pointer direction), boost on mouse-down,\n   nickname entry, camera follows the local snake's head. Death screen with\n   a \"play again\" button that simply reconnects with a fresh join \u2014 respawn\n   IS reconnection; do not build any other respawn protocol.\n\n4. Headless simulate harness (`src\/simulate.ts`, `npm run simulate`):\n   - Default mode: `--seed N --ticks N` \u2014 runs the real World with bots,\n     prints a deterministic summary (ticks, snakes alive\/eliminated, food\n     count) and all EVENT lines.\n   - Scenario mode: `--scenario &lt;name&gt; --seed N` \u2014 runs the real World with\n     scripted inputs. Required scenarios: `two-snake-kill` (one scripted\n     snake drives its head into another's body), `border-death`,\n     `boost-drain`.\n\n5. Tests (`tests\/`, `npm test`, node --test via tsx) \u2014 honest and small:\n   steering turn cap; boost mass drain; body follows the head path; the\n   three death rules (other-body kills, border kills, self-overlap safe);\n   eliminated snake drops food; protocol messages encode\/decode round-trip.\n   Tests must assert on measured behavior; tautologies, weakened assertions,\n   or state mutation to force outcomes are critical issues. The scenario\n   harness covers multi-snake causality \u2014 tests need not re-derive it.\n\n6. `README.md`: how to start the server, open two browsers, play; all\n   commands and scenarios.<\/code><\/pre>\n<\/p><\/div>\n<\/p><\/div>\n<\/div><figcaption>A multiplayer snake game generated by Inkling from the prompt in the second tab \u2014 real-time server, bots, leaderboard and all.<!-- source: Peichao Du --><\/figcaption><\/figure>\n<h3 id=\"controllable-thinking-effort\">Controllable thinking effort<\/h3>\n<p>Test-time scaling and problem-solving are the core capability of every model, but that capacity is hard to capture with a single number. Developers fine-tuning models for a specialized task <a href=\"https:\/\/thinkingmachines.ai\/news\/learning-to-replicate-expert-judgment-in-financial-tasks\/\">care as much about efficiency<\/a> as about the max-effort performance on a public benchmark. Cost and latency are often binding constraints in real-world applications, and low latency in particular is crucial for enabling collaboration and improvement through iteration.<\/p>\n<figure>\n<div class=\"effort-curve-figure intelligence-interactivity-figure\" data-plot-point-stagger=\"18\" data-plot-point-jitter=\"10\">\n<div class=\"effort-curve\">\n<p>\n      <span class=\"ii-legend-item\"><span class=\"effort-sweep-marker\" aria-hidden=\"true\"\/>Inkling (effort sweep)<\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ii-legend-marker ii-marker-circle\" style=\"--ii-marker-color: var(--e-glm);\"\/>GLM-5.2<\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ii-legend-marker ii-marker-circle\" style=\"--ii-marker-color: var(--e-kimi26);\"\/>Kimi K2.6<\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ii-legend-marker ii-marker-circle\" style=\"--ii-marker-color: var(--e-nemotron);\"\/>Nemotron 3 Ultra<\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ii-legend-marker ii-marker-circle\" style=\"--ii-marker-color: var(--e-kimi25);\"\/>Kimi K2.5<\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ii-legend-marker ii-marker-circle\" style=\"--ii-marker-color: var(--e-gptoss);\"\/>GPT-OSS (high)<\/span>\n    <\/p>\n<\/p><\/div>\n<\/div><figcaption>Sweeping Inkling&#8217;s effort setting from 0.2 to 0.99 traces its performance against mean generated tokens on Terminal Bench 2.1, HLE, and IFBench; competing models are shown at their default operating point. Inkling reaches a given score at fewer tokens \u2014 for example, it matches Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens. *Humanity&#8217;s Last Exam scores reflect an earlier checkpoint and run slightly below the final release.<!-- Data: Kyle Luther, iterationlab\/notebook kyle\/jun26\/jun18_evals.ipynb. Chart source: figures\/effort-curves.html, regenerated by scratchpad gen_effort_curves.py. Kyle fig-2 feedback: per-panel LINEAR x-axes (was shared log), middle panel titled \"Humanity's Last Exam\", earlier-checkpoint asterisk. Point data recovered by inverting the prior log-x pixel positions. Refresh on final checkpoint. --><\/figcaption><\/figure>\n<p>Inkling supports controllable thinking effort, allowing you to balance performance with token efficiency. The chart above shows the effort\/performance curve of Inkling as well as other open-weights models on a range of benchmarks: Terminal Bench 2.1 for agentic coding, HLE for advanced reasoning, and IFBench for instruction following. Inkling spends one third as many tokens to achieve the same performance as Nemotron 3 Ultra on Terminal Bench. Cost and latency matter for a model that you run millions of times and as part of longer workflows; looking at the full cost curve allows developers to choose the best model for each use case.<\/p>\n<h3 id=\"multimodality\">Multimodality<\/h3>\n<p>A major goal of Inkling\u2019s design is to serve as the background reasoning model in the <a href=\"https:\/\/thinkingmachines.ai\/blog\/interaction-models\/\">interaction models system<\/a> we recently introduced. Interaction models enable the user to collaborate naturally, using voice and vision in real time. This requires a model natively trained for broad multimodal capabilities.<\/p>\n<p class=\"benchmark-results-note\">Audio and vision benchmarks against specialist omni models (open- and closed-weight), reported at effort=0.99.<!-- Caption added during v7 sync; table from the doc \"Inkling Release Table\" figure. --><\/p>\n<p>The multimodal components were trained from scratch on general-domain data. We opted for an encoder-free architecture for audio and vision inputs, consistent with the interaction model design. Audio signals are input as dMel spectrograms<label for=\"sn-dmel\" class=\"sidenote-number\"\/><input type=\"checkbox\" id=\"sn-dmel\" class=\"margin-toggle\"\/><span class=\"sidenote\"><a href=\"https:\/\/arxiv.org\/abs\/2407.15835\">dMel: Speech Tokenization made Simple<\/a> (Richard He Bai et al, 2024)<\/span>, while images are encoded as patches of 40&#215;40 pixels using a four-layer hMLP<label for=\"sn-hmlp\" class=\"sidenote-number\"\/><input type=\"checkbox\" id=\"sn-hmlp\" class=\"margin-toggle\"\/><span class=\"sidenote\"><a href=\"https:\/\/arxiv.org\/abs\/2203.09795\">Three things everyone should know about Vision Transformers<\/a> (Hugo Touvron et al, 2022)<\/span>. Both are transformed via a light-weight embedding layer and processed jointly with text tokens.<\/p>\n<p>Inkling transcribes speech, follows spoken instructions, answers questions about recordings, and reasons over longer-form audio. These capabilities place it among the strongest open-weights audio models on VoiceBench, MMAU, and AudioMC. For vision, Inkling accepts images as input and can describe visual content, answer questions, and perform in-depth reasoning based on the provided visual information. It demonstrates strong performance on charts, diagrams, and mathematical visual reasoning tasks. During inference, Inkling can also leverage a Python tool to support image understanding through operations such as zooming and cropping, while seamlessly integrating visual reasoning with code-based reasoning.<\/p>\n<p>As our first release, Inkling establishes a robust multimodal foundation for future work. We expect its multimodal capabilities to continue improving as we expand the model and training pipeline in subsequent iterations.<\/p>\n<h3 id=\"epistemics\">Epistemics<\/h3>\n<p>We trained Inkling for calibration, instruction following, and resistance to censorship, which we refer to collectively as the model\u2019s <em>epistemics<\/em>.<\/p>\n<p>Getting the facts right requires more than memorizing a large corpus of knowledge. A useful model must be well-calibrated, expressing the right amount of confidence in its answers \u2014 including on questions which aren\u2019t yet settled. The latter is a crucial capability for prediction and forecasting, an important use case where <a href=\"https:\/\/thinkingmachines.ai\/news\/training-llms-to-predict-world-events\/\">fine-tuned models have shown rapid improvement<\/a> in recent months, <a href=\"https:\/\/x.com\/tinkerapi\/status\/2056798250532057139\">outperforming frontier LLMs<\/a>.<\/p>\n<p class=\"benchmark-results-note\">Results were obtained during testing between June 30 and July 13, 2026 on a different checkpoint of Inkling than the one released.<!-- Data: doc image11 (forecasting table). ForecastBench rows refreshed to final values 2026-07-14; Prophet Arena unchanged. Earlier-checkpoint note per Zhicheng Sun \/ doc comment [s]. --><\/p>\n<p>Forecasting requires integrating multiple sources of information into a calibrated probability, a core skill for a model users can trust. A model that\u2019s confident in every answer it gives, including when it\u2019s missing info and confabulates, forces the user to double-check everything. A model that gives the appropriate measure of confidence is useful across more real-world domains where information is often conflicting, unreliable, or hard to find. We trained for calibration with RL against proper scoring rules on a large corpus of resolved real-world questions.<\/p>\n<p>The second component of a trustworthy model is instruction following, including on hard-to-verify, complex queries. We did RL with two automated graders: a rubric grader and claims grader. The first grader scores each response against a checklist of what a good answer should contain. Rubrics can penalize errors in principle, but in practice they emphasize recall and can be hacked by models spraying plausibly relevant facts hoping to match rubric items. The claims grader verifies each factual claim in the response, penalizing claims that don\u2019t check out. It performs agentic web search for claim verification, not relying solely on its own knowledge. Together, the two graders improve helpfulness and reduce hallucination at the same time, rather than trading one for the other.<\/p>\n<p>These rewards don\u2019t directly target calibrated uncertainty in long-form responses, so we added targeted datasets that do. The largest is short-form factual QA with abstention-aware rewards: answering only pays off when the model is likely to be right, so the optimal policy is to answer when confident and otherwise say \u201cI don\u2019t know\u201d or give a hedged best guess. Some prompts encourage or forbid hedging, teaching the model to follow the user\u2019s preference for a forced guess versus a calibrated non-answer.<\/p>\n<p>Finally, we trained Inkling to answer directly on topics that may be subject to censorship. Cognition evaluated the model on their Propaganda and Censorship Eval<label for=\"sn-censorship\" class=\"sidenote-number\"\/><input type=\"checkbox\" id=\"sn-censorship\" class=\"margin-toggle\"\/><span class=\"sidenote\">The Cognition Team, \u201c<a href=\"https:\/\/cognition.com\/blog\/measuring-open-source-model-trustworthiness\">Measuring the Trustworthiness of Open-Source-Derived Models<\/a>,\u201d 2026.<\/span>, and it exhibited strong patterns of censorship non-compliance.<\/p>\n<h3 id=\"safety\">Safety<\/h3>\n<p>We trained Inkling to an internal spec of safe model behavior across all modalities. We then commissioned external safety testers to verify the results.<\/p>\n<p>We evaluated Inkling\u2019s safety in several areas. For dangerous capabilities \u2014 CBRN, cyber, and loss of control \u2014 we ran internal evaluations and enlisted external testers. We attended to human-AI threat vectors, including sycophancy, vulnerable users, and harmful manipulation, using internal evaluations and external testers.<\/p>\n<p><!-- KEEP IN SYNC with the model card (content\/model-card\/index.md), the source of truth for these eval numbers. Update both together when evals change. --><\/p>\n<p><!-- TODO: pre-deployment testing (FAR, Scale, Handshake, Apollo) results pending; internal AUP\/model-spec evals across 16 categories, 17 languages, and three modalities. --><\/p>\n<p>Inkling shows the strongest built-in safeguards of any open-weights model we compared on FORTRESS, a benchmark that tests refusal of requests related to weapons and violence alongside benign look-alike queries. Inkling refused more harmful requests without over-refusing benign analogs. Inkling scores above 98% on StrongREJECT \u2014 a refusal test of unambiguous harmful requests \u2014 in line with other open and closed-weights models.<\/p>\n<p>Safety is crucial for open-weights models. We\u2019re continuing to study safety behavior and capability uplift in customizable models, including how safety behavior is impacted by fine-tuning on Tinker.<\/p>\n<h2 id=\"benchmarking-inkling\">Benchmarking Inkling<\/h2>\n<p>We benchmark Inkling on a broad range of capabilities. All evals are run at effort 0.99 and temperature 1.0. All coding evals run with 256K max-token trajectory limit.<\/p>\n<p>To improve consistency, we rely on externally reported evaluations for both internal and external models when applicable. Specifically, we use the score reported by Artificial Analysis for the following evals: Humanity\u2019s Last Exam, GPQA Diamond, GDPVal, Tau 3 Banking, AA Omniscience, MMMU Pro.<\/p>\n<p class=\"benchmark-results-note\"><span class=\"benchmark-results-note-item\" id=\"eval-note-swebench-verified\"><span class=\"benchmark-results-note-label\">*SWEBench Verified:<\/span> Inkling numbers are reported using a bash-only harness. We use self-reported numbers for external models.<\/span><span class=\"benchmark-results-note-item\" id=\"eval-note-terminal-bench-2-1\"><span class=\"benchmark-results-note-label\">*Terminal Bench 2.1:<\/span> Inkling numbers are reported using an internal coding harness. A small number of solutions were found to be contaminated from web search and were assigned a score of 0. We use self-reported numbers for external models where available. Otherwise, we report performance using our internal harness.<\/span><span class=\"benchmark-results-note-item\" id=\"eval-note-audio-mc\"><span class=\"benchmark-results-note-label\">\u2020Audio MC:<\/span> Other models were evaluated internally since they are not on the official leaderboard.<\/span><span class=\"benchmark-results-note-item\" id=\"eval-note-voicebench\"><span class=\"benchmark-results-note-label\">\u2020VoiceBench:<\/span> VoiceBench uses rule-based, hard-coded string matching for grading, making the evaluation sensitive to output-formatting differences. We therefore added a system message instructing models to follow the expected answer format.<\/span><span class=\"benchmark-results-note-item\" id=\"eval-note-charxiv-rq-tools\"><span class=\"benchmark-results-note-label\">\u2020CharXiv RQ with tools:<\/span> We benchmarked Claude Fable 5 and GPT 5.6 Sol (max\/xhigh) using our internal Python harness.<\/span><\/p>\n<h2 id=\"the-making-of-inkling\">The making of Inkling<\/h2>\n<h3 id=\"architecture\">Architecture<\/h3>\n<p>Inkling is a Mixture-of-Experts Transformer with a handful of departures from the common recipe, each chosen for efficiency and long-context performance.<\/p>\n<p>The MoE design largely follows DeepSeek-V3. Each MoE layer contains 256 routed experts and 2 shared experts, with 6 routed experts active per token. Inkling uses a sigmoid-based router with an auxiliary-loss-free load-balancing bias. The scores of the selected routed experts and the shared experts are normalized jointly and used to weight their combined outputs.<\/p>\n<p>For attention, we interleave sliding-window and global layers at a 5:1 ratio with 8 KV heads. We find that encoding position with a relative positional embedding<label for=\"sn-rel-attn-shaw\" class=\"sidenote-number\"\/><input type=\"checkbox\" id=\"sn-rel-attn-shaw\" class=\"margin-toggle\"\/><span class=\"sidenote\"><a href=\"https:\/\/arxiv.org\/abs\/1803.02155\">Self-Attention with Relative Position Representations<\/a> (Peter Shaw et al, 2018)<\/span><label for=\"sn-rel-attn-music-transformer\" class=\"sidenote-number\"\/><input type=\"checkbox\" id=\"sn-rel-attn-music-transformer\" class=\"margin-toggle\"\/><span class=\"sidenote\"><a href=\"https:\/\/arxiv.org\/abs\/1809.04281\">Music Transformer<\/a> (Cheng-Zhi Anna Huang et al, 2018)<\/span> performs better and extrapolates better to longer sequences than the more widely adopted Rotary Positional Embedding (RoPE). We also apply short convolutions at two points \u2014 after the key and value projections in each attention layer, and on the attention and MLP residual branch outputs before they rejoin the main residual stream.<\/p>\n<h3 id=\"training\">Training<\/h3>\n<p>Inkling was pretrained on 45 trillion tokens from a variety of content types, including text, images, audio and video. We trained Inkling with a hybrid optimization strategy \u2014 Muon for large matrix weights, Adam for other parameters \u2014 and hyperparameter schedules inspired by our previous research on <a href=\"https:\/\/thinkingmachines.ai\/blog\/modular-manifolds\/\">modular manifolds<\/a>. We coupled the weight decay strength to the square of the learning rate, which we found kept the overall size of the model weights stable across training horizons<label for=\"sn-weight-decay\" class=\"sidenote-number\"\/><input type=\"checkbox\" id=\"sn-weight-decay\" class=\"margin-toggle\"\/><span class=\"sidenote\">See also <a href=\"https:\/\/arxiv.org\/abs\/2305.17212\">Kosson et al. (2023)<\/a> and <a href=\"https:\/\/arxiv.org\/abs\/2506.02285v1\">Defazio (2025)<\/a>.<\/span>.<\/p>\n<p>We post-trained Inkling on a broad distribution of math, agentic code &amp; tool use, audio, image, chat, and safety domains. To bootstrap post-training, we ran an initial SFT on synthetic data generated by open-weights models including Kimi K2.5. The bootstrap accounts for a small fraction of compute, with the majority being employed for large-scale RL on synthetic and human-created environments.<\/p>\n<p>Inkling was our first major training effort and was trained on NVIDIA GB300 NVL72 systems. Future models will further push the scale of compute across pre-training, post-training and RL.<\/p>\n<h3 id=\"rl-at-scale\">RL at scale<\/h3>\n<p>We relied on large-scale asynchronous RL to shape model behavior and improve its reasoning and overall performance. The chart below shows the model\u2019s score on a held-out aggregate of reasoning evals such as AIME, HLE, GPQA, and others. We scaled RL to over 30M rollouts, with stable training sustained over two long continuous runs. Reasoning performance improved log-linearly throughout the entire process, resulting in a significant increase overall.<\/p>\n<figure>\n<div class=\"rl-scaling-figure\">\n<div class=\"rl-scaling\">\n<p>\n      <svg viewbox=\"0 0 760 400\" role=\"img\" aria-hidden=\"true\" focusable=\"false\">\n      <g class=\"rl-scaling-legend\" transform=\"translate(84 0)\">\n      <g transform=\"translate(78 30)\"><circle class=\"rl-scaling-dot-sft\" cx=\"0\" cy=\"0\" r=\"6\"\/><text x=\"14\" y=\"4\">SFT initialization<\/text><\/g>\n      <g transform=\"translate(240 30)\"><circle class=\"rl-scaling-dot-rl\" cx=\"0\" cy=\"0\" r=\"6\"\/><text x=\"14\" y=\"4\">RL checkpoint<\/text><\/g>\n      <g transform=\"translate(388 30)\"><circle class=\"rl-scaling-dot-rel\" cx=\"0\" cy=\"0\" r=\"6\"\/><text x=\"14\" y=\"4\">Released checkpoint<\/text><\/g>\n      <\/g>\n      <g class=\"rl-scaling-panel\">\n      <line class=\"rl-scaling-grid\" x1=\"78\" y1=\"320.0\" x2=\"724\" y2=\"320.0\"\/>\n      <text class=\"rl-scaling-ytick\" x=\"66\" y=\"324.0\">0.25<\/text>\n      <line class=\"rl-scaling-grid\" x1=\"78\" y1=\"257.5\" x2=\"724\" y2=\"257.5\"\/>\n      <text class=\"rl-scaling-ytick\" x=\"66\" y=\"261.5\">0.28<\/text>\n      <line class=\"rl-scaling-grid\" x1=\"78\" y1=\"195.0\" x2=\"724\" y2=\"195.0\"\/>\n      <text class=\"rl-scaling-ytick\" x=\"66\" y=\"199.0\">0.31<\/text>\n      <line class=\"rl-scaling-grid\" x1=\"78\" y1=\"132.5\" x2=\"724\" y2=\"132.5\"\/>\n      <text class=\"rl-scaling-ytick\" x=\"66\" y=\"136.5\">0.34<\/text>\n      <line class=\"rl-scaling-grid\" x1=\"78\" y1=\"70.0\" x2=\"724\" y2=\"70.0\"\/>\n      <text class=\"rl-scaling-ytick\" x=\"66\" y=\"74.0\">0.37<\/text>\n      <line class=\"rl-scaling-axis\" x1=\"78\" y1=\"70\" x2=\"78\" y2=\"320\"\/>\n      <line class=\"rl-scaling-axis\" x1=\"78\" y1=\"320\" x2=\"724\" y2=\"320\"\/>\n      <line class=\"rl-scaling-axis\" x1=\"110.0\" y1=\"320\" x2=\"110.0\" y2=\"325\"\/>\n      <text class=\"rl-scaling-xtick\" x=\"110.0\" y=\"340\">0<\/text>\n      <line class=\"rl-scaling-axis\" x1=\"179.7\" y1=\"320\" x2=\"179.7\" y2=\"325\"\/>\n      <text class=\"rl-scaling-xtick\" x=\"179.7\" y=\"340\">0.5M<\/text>\n      <line class=\"rl-scaling-axis\" x1=\"229.1\" y1=\"320\" x2=\"229.1\" y2=\"325\"\/>\n      <text class=\"rl-scaling-xtick\" x=\"229.1\" y=\"340\">1M<\/text>\n      <line class=\"rl-scaling-axis\" x1=\"298.8\" y1=\"320\" x2=\"298.8\" y2=\"325\"\/>\n      <text class=\"rl-scaling-xtick\" x=\"298.8\" y=\"340\">2M<\/text>\n      <line class=\"rl-scaling-axis\" x1=\"386.5\" y1=\"320\" x2=\"386.5\" y2=\"325\"\/>\n      <text class=\"rl-scaling-xtick\" x=\"386.5\" y=\"340\">4M<\/text>\n      <line class=\"rl-scaling-axis\" x1=\"487.5\" y1=\"320\" x2=\"487.5\" y2=\"325\"\/>\n      <text class=\"rl-scaling-xtick\" x=\"487.5\" y=\"340\">8M<\/text>\n      <line class=\"rl-scaling-axis\" x1=\"596.8\" y1=\"320\" x2=\"596.8\" y2=\"325\"\/>\n      <text class=\"rl-scaling-xtick\" x=\"596.8\" y=\"340\">16M<\/text>\n      <line class=\"rl-scaling-axis\" x1=\"700.0\" y1=\"320\" x2=\"700.0\" y2=\"325\"\/>\n      <text class=\"rl-scaling-xtick\" x=\"700.0\" y=\"340\">30M+<\/text>\n      <line class=\"rl-scaling-fit\" x1=\"110.0\" y1=\"308.4\" x2=\"679.2\" y2=\"106.1\"\/>\n      <circle class=\"rl-scaling-dot-rl\" cx=\"188.1\" cy=\"296.4\" r=\"5\"\/>\n      <circle class=\"rl-scaling-dot-rl\" cx=\"287.4\" cy=\"242.4\" r=\"5\"\/>\n      <circle class=\"rl-scaling-dot-rl\" cx=\"452.3\" cy=\"204.8\" r=\"5\"\/>\n      <circle class=\"rl-scaling-dot-rl\" cx=\"493.2\" cy=\"175.9\" r=\"5\"\/>\n      <circle class=\"rl-scaling-dot-rl\" cx=\"514.6\" cy=\"160.8\" r=\"5\"\/>\n      <circle class=\"rl-scaling-dot-rl\" cx=\"557.1\" cy=\"143.1\" r=\"5\"\/>\n      <circle class=\"rl-scaling-dot-rl\" cx=\"597.1\" cy=\"136.3\" r=\"5\"\/>\n      <circle class=\"rl-scaling-dot-rl\" cx=\"634.2\" cy=\"122.1\" r=\"5\"\/>\n      <circle class=\"rl-scaling-dot-sft\" cx=\"110.0\" cy=\"290.6\" r=\"6.5\"\/>\n      <text class=\"rl-scaling-anno\" x=\"121.0\" y=\"281.6\">SFT init \u00b7 0.264<\/text>\n      <circle class=\"rl-scaling-dot-rel\" cx=\"679.2\" cy=\"98.3\" r=\"7.5\"\/>\n      <text class=\"rl-scaling-anno rl-scaling-anno-end\" x=\"668.2\" y=\"87.3\">Released \u00b7 0.356<\/text>\n      <text class=\"rl-scaling-axis-label\" x=\"26\" y=\"195.0\" transform=\"rotate(-90 26 195.0)\">Aggregate eval reward<\/text>\n      <\/g>\n      <text class=\"rl-scaling-axis-label rl-scaling-xlabel\" x=\"401.0\" y=\"372\">Rollouts \u2192<\/text>\n      <\/svg>\n    <\/p>\n<\/p><\/div>\n<\/div><figcaption>Reward on a held-out aggregate of reasoning evals \u2014 AIME, HLE, GPQA, and others \u2014 improves log-linearly over more than 30M RL rollouts, from SFT initialization to the released checkpoint.<!-- Data: Szymon Tworkowski (aggregate eval_reward_mean vs elapsed_train_samples, Slack C0BDA2Q0JN5 ts 1784074317). Chart source: figures\/rl-scaling.html, regenerated by scratchpad gen_rl_scaling.py (single-panel, per Szymon\/Kyle). Provisional &mdash; refresh on final-checkpoint numbers. --><\/figcaption><\/figure>\n<p>We specified the model\u2019s effort level on different samples by changing the system message and adjusting the per-token cost. This caused the model to use a different amount of tokens in different rollouts and learn the ability to control thinking effort.<\/p>\n<p>We also observed an emergent shift in the reasoning style over the course of RL training. The chain of thought became more concise over time, dropping grammatical overhead while remaining comprehensible and leaving the final response unaffected. This wasn\u2019t targeted by the reward \u2014 efficiency alone drove the compression. A similar effect was also recently noted by the Cognition team in the process of training SWE-1.7<label for=\"sn-swe17\" class=\"sidenote-number\"\/><input type=\"checkbox\" id=\"sn-swe17\" class=\"margin-toggle\"\/><span class=\"sidenote\">The Cognition Team, \u201c<a href=\"https:\/\/cognition.com\/blog\/swe-1-7\">SWE-1.7: Frontier Intelligence at a Fraction of the Cost<\/a>.&#8221;<\/span>. Below is an example of how Inkling\u2019s chain of thought on the same math problem evolved with RL:<\/p>\n<figure>\n<div class=\"cot-evo-figure\">\n<div class=\"cot-evo\" role=\"img\" aria-label=\"Two side-by-side excerpts of the model's chain of thought on the same physics problem, early versus late in RL. The early trace uses full, grammatical sentences; the late trace is compressed and telegraphic \u2014 dropping articles and connectives while staying comprehensible.\">\n<div class=\"cot-evo-cols\">\n<div class=\"cot-evo-col cot-evo-col-early\">\n<p>Early in RL <span class=\"cot-evo-tokens\">verbose, grammatical<\/span><\/p>\n<p><strong>We need to understand<\/strong> the operator.<br \/>\nThe 5D line element is ds\u00b2 = e^{2A(x)} (ds\u00b2_4d + dx\u00b2), where A(x) = sin(x) + 4 cos(x), x in [0, 2\u03c0].<br \/>\nThe internal coordinate is periodic.<br \/>\nThe background is a warped product: metric g_{MN} where M,N = 0..4.<br \/>\nThe internal direction has metric e^{2A(x)} dx\u00b2?<br \/>\nWait, the ds\u00b2 is e^{2A} (ds\u00b2_4d + dx\u00b2).<br \/>\nSo the internal metric is e^{2A(x)} dx\u00b2.<br \/>\nActually if the total metric is ds\u00b2 = e^{2A(x)} (ds\u00b2_4d + dx\u00b2), then yes, internal metric is e^{2A} dx\u00b2.<br \/>\n<span class=\"cot-evo-ellipsis\">\u2026<\/span><\/p>\n<\/p><\/div>\n<div class=\"cot-evo-col cot-evo-col-late\">\n<p>Late in RL <span class=\"cot-evo-tokens\">compressed, telegraphic<\/span><\/p>\n<p><strong>We need determine eigenvalue problem<\/strong> for spin-2 fluctuations h_{\u03bc\u03bd}(x,y) with TT in 4d and depend on x.<br \/>\nFor metric of form ds\u00b2 = e^{2A(x)} (g_{\u03bc\u03bd}(y) + h_{\u03bc\u03bd}(y,x)) dy^\u03bc dy^\u03bd + e^{2A(x)}?<br \/>\nWait internal metric is e^{2A} dx\u00b2?<br \/>\nActually ds\u00b2 = e^{2A} [ds_4\u00b2 + dx\u00b2].<br \/>\nSo internal metric is e^{2A} dx\u00b2; warp factor same for 4d and internal?<br \/>\nYes.<br \/>\nWe need equation for h_{\u03bc\u03bd}(y,x) = h_{\u03bc\u03bd}(y) \u03c8(x) maybe with normalization.<br \/>\n<span class=\"cot-evo-ellipsis\">\u2026<\/span><\/p>\n<\/p><\/div>\n<\/p><\/div>\n<\/p><\/div>\n<\/div><figcaption>The same chain of thought, early and late in RL. The late-RL trace drops articles and connectives \u2014 &#8220;We need <em>to<\/em> understand&#8221; becomes &#8220;We need determine&#8221; \u2014 while staying comprehensible and reaching the same answer.<!-- Source: Kyle Luther \/ Szymon Tworkowski, \"caveman speak\" slide (Slides 18nYvjkUNwTqhSeE-paXkUSLuzfsmWbJfr6qOSy9cH14). Excerpted real traces. --><\/figcaption><\/figure>\n<h2 id=\"inkling-small\">Inkling-Small<\/h2>\n<p>Alongside Inkling we are sharing a preview of Inkling-Small, a 276B-parameter Mixture-of-Experts model (12B active, vs. 41B for Inkling) with a different performance\/latency trade-off. Inkling-Small matches or exceeds its larger sibling on many benchmarks \u2014 the result of improvements we made to the pre-training data and recipe for the smaller model. The two models share the same scalable post-training stack applied on top.<\/p>\n<div class=\"inkling-eval-table-block\" style=\"--benchmark-model-count: 2\">\n<div class=\"benchmark-results-scroll inkling-eval-table inkling-eval-table--content\" data-inkling-eval-table=\"\">\n<table class=\"benchmark-results-table benchmark-results-categorized\">\n<colgroup>\n<col class=\"benchmark-name-column\"\/>\n<col class=\"benchmark-model-column\" span=\"2\"\/>\n    <\/colgroup>\n<thead>\n<\/thead>\n<tbody class=\"benchmark-category-group\">\n<tr class=\"benchmark-category-row\">\n<th class=\"benchmark-category-label\" scope=\"rowgroup\">Reasoning<\/th>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n      <\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">HLE<\/span><span class=\"benchmark-subtitle\">text only<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">29.7%<\/td>\n<td class=\"benchmark-value\">29.6%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">HLE<\/span><span class=\"benchmark-subtitle\">with tools<\/span>\n        <\/td>\n<td class=\"benchmark-value\">46.0%<\/td>\n<td class=\"benchmark-value benchmark-value-best\">46.6%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">AIME 2026<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">97.1%<\/td>\n<td class=\"benchmark-value\">95.1%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">GPQA Diamond<\/span>\n        <\/td>\n<td class=\"benchmark-value\">87.2%<\/td>\n<td class=\"benchmark-value benchmark-value-best\">88.3%<\/td>\n<\/tr>\n<\/tbody>\n<tbody class=\"benchmark-category-group\">\n<tr class=\"benchmark-category-row\">\n<th class=\"benchmark-category-label\" scope=\"rowgroup\">Agentic (coding)<\/th>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n      <\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">SWEBench Verified<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">77.6%<\/td>\n<td class=\"benchmark-value\">77.4%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">SWEBench Pro<\/span><span class=\"benchmark-subtitle\">Public<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">54.3%<\/td>\n<td class=\"benchmark-value\">53.2%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">Terminal Bench 2.1<\/span><span class=\"benchmark-subtitle\">Best Harness<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">63.8%*<\/td>\n<td class=\"benchmark-value\">52.7%<\/td>\n<\/tr>\n<\/tbody>\n<tbody class=\"benchmark-category-group\">\n<tr class=\"benchmark-category-row\">\n<th class=\"benchmark-category-label\" scope=\"rowgroup\">Agentic (general)<\/th>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n      <\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">Tau 3 Banking<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">23.7%<\/td>\n<td class=\"benchmark-value\">13.6%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">MCP-Atlas<\/span>\n        <\/td>\n<td class=\"benchmark-value\">74.1%<\/td>\n<td class=\"benchmark-value benchmark-value-best\">74.9%<\/td>\n<\/tr>\n<\/tbody>\n<tbody class=\"benchmark-category-group\">\n<tr class=\"benchmark-category-row\">\n<th class=\"benchmark-category-label\" scope=\"rowgroup\">Factuality<\/th>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n      <\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">SimpleQA Verified<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">43.9%<\/td>\n<td class=\"benchmark-value\">20.9%<\/td>\n<\/tr>\n<\/tbody>\n<tbody class=\"benchmark-category-group\">\n<tr class=\"benchmark-category-row\">\n<th class=\"benchmark-category-label\" scope=\"rowgroup\">Chat<\/th>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n      <\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">IFBench<\/span>\n        <\/td>\n<td class=\"benchmark-value\">79.8%<\/td>\n<td class=\"benchmark-value benchmark-value-best\">83.4%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">Global-MMLU-Lite<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">88.7%<\/td>\n<td class=\"benchmark-value\">86.8%<\/td>\n<\/tr>\n<\/tbody>\n<tbody class=\"benchmark-category-group\">\n<tr class=\"benchmark-category-row\">\n<th class=\"benchmark-category-label\" scope=\"rowgroup\">Vision<\/th>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n      <\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">MMMU Pro<\/span><span class=\"benchmark-subtitle\">Standard 10<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">73.5%<\/td>\n<td class=\"benchmark-value\">73.1%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">Charxiv RQ<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">78.1%<\/td>\n<td class=\"benchmark-value\">76.7%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">Charxiv RQ<\/span><span class=\"benchmark-subtitle\">with python<\/span>\n        <\/td>\n<td class=\"benchmark-value\">82.0%<\/td>\n<td class=\"benchmark-value benchmark-value-best\">83.4%<\/td>\n<\/tr>\n<\/tbody>\n<tbody class=\"benchmark-category-group\">\n<tr class=\"benchmark-category-row\">\n<th class=\"benchmark-category-label\" scope=\"rowgroup\">Audio<\/th>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n      <\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">Audio MC<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">56.6%<\/td>\n<td class=\"benchmark-value\">49.6%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">MMAU<\/span>\n        <\/td>\n<td class=\"benchmark-value\">77.2%<\/td>\n<td class=\"benchmark-value benchmark-value-best\">77.5%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">VoiceBench<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">91.4%<\/td>\n<td class=\"benchmark-value\">90.0%<\/td>\n<\/tr>\n<\/tbody>\n<tbody class=\"benchmark-category-group\">\n<tr class=\"benchmark-category-row\">\n<th class=\"benchmark-category-label\" scope=\"rowgroup\">Safety<\/th>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n<td class=\"benchmark-category-cell\" aria-hidden=\"true\"\/>\n      <\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">FORTRESS<\/span><span class=\"benchmark-subtitle\">Adversarial<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">78.0%<\/td>\n<td class=\"benchmark-value\">75.6%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">FORTRESS<\/span><span class=\"benchmark-subtitle\">Benign<\/span>\n        <\/td>\n<td class=\"benchmark-value benchmark-value-best\">95.9%<\/td>\n<td class=\"benchmark-value\">94.1%<\/td>\n<\/tr>\n<tr>\n<td class=\"benchmark-name\">\n          <span class=\"benchmark-title\">StrongREJECT<\/span>\n        <\/td>\n<td class=\"benchmark-value\">98.6%<\/td>\n<td class=\"benchmark-value benchmark-value-best\">98.8%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<\/div>\n<p class=\"benchmark-results-note\">Both models reported at effort=0.99; the higher result in each row is highlighted. *We assign a score of 0 to Terminal Bench 2.1 rollouts with solution contamination from web search.<!-- Data: Google Docs comparison table supplied 2026-07-15. --><\/p>\n<p>Early results show Inkling-Small performing close to Inkling on reasoning and agentic tasks. With 12B active parameters and controllable thinking effort, it is a natural fit for workloads where cost and latency matter such as coding, using LLMs to grade, or generating synthetic data for other models.<\/p>\n<p>We are currently finishing the testing of Inkling-Small and will release its full weights once that work is complete.<\/p>\n<h2 id=\"customizing-inkling\">Customizing Inkling<\/h2>\n<p><a href=\"https:\/\/thinkingmachines.ai\/news\/learning-to-replicate-expert-judgment-in-financial-tasks\/\">Many real-world problems<\/a> aren\u2019t solved well by even the best generalist models, with the gap being closed by fine-tuning that utilizes an organization\u2019s specialized knowledge. The experience of our Tinker customers points in the same direction. Our post-training and results of RL at scale suggest that Inkling is capable of rapidly learning from fine-tuning.<\/p>\n<h3 id=\"inkling-availability\">Inkling availability<\/h3>\n<p>Inkling is available on <a href=\"https:\/\/thinkingmachines.ai\/tinker\/\">Tinker<\/a> today with context length options of 64K and 256K tokens. We are offering Inkling at a 50% discount for a limited time, with full pricing information available <a href=\"https:\/\/tinker-docs.thinkingmachines.ai\/tinker\/models\/\">in our documentation<\/a>.<\/p>\n<p>To support Tinkerers fine-tuning with Inkling, we have updated our cookbook to natively support Inkling and have added three <a href=\"https:\/\/github.com\/thinking-machines-lab\/tinker-cookbook\">new cookbook recipes<\/a> that showcase Inkling\u2019s unique audio capabilities. We also released <a href=\"https:\/\/pypi.org\/project\/tml-renderers\/\">tml-renderer<\/a> for reliably sampling and post-training with tool calls, reasoning content, and multimodal inputs.<\/p>\n<p>To get a feel for the model before committing to a run, users can head to the <a href=\"https:\/\/tinker.thinkingmachines.ai\/playground?utm_source=blog&amp;utm_campaign=inkling_model_release\">Inkling Playground<\/a> in the Tinker console. The playground offers a chat interface with integrated agentic web search, free for a limited time.<\/p>\n<p>We have partnered across the ecosystem to help customers deploy checkpoints fine-tuned on Tinker. Inkling is available via APIs on <a href=\"https:\/\/www.together.ai\/blog\/together-ai-brings-thinking-machines-labs-new-model-inkling-on-day-0\">Together AI<\/a>, <a href=\"https:\/\/fireworks.ai\/\">Fireworks<\/a>, <a href=\"https:\/\/modal.com\/\">Modal<\/a>, <a href=\"https:\/\/www.databricks.com\/\">Databricks<\/a>, and <a href=\"https:\/\/www.baseten.co\/\">Baseten<\/a>. We worked with <a href=\"https:\/\/www.radixark.com\/\">RadixArk<\/a> to provide open-source inference and RL support in <a href=\"https:\/\/www.sglang.io\/\">SGLang<\/a> and <a href=\"https:\/\/github.com\/radixark\/miles\">Miles<\/a>. We worked with <a href=\"https:\/\/inferact.ai\/\">Inferact<\/a> to support inference in <a href=\"https:\/\/vllm.ai\/\">vLLM<\/a>, with <a href=\"https:\/\/lightseek.org\/\">Lightseek<\/a> for inference in <a href=\"https:\/\/github.com\/lightseekorg\/tokenspeed\">TokenSpeed<\/a>, and with <a href=\"https:\/\/unsloth.ai\/\">Unsloth<\/a> for inference in <a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\">llama.cpp<\/a>. Finally, we partnered with <a href=\"https:\/\/huggingface.co\/\">Hugging Face<\/a> on integration with <a href=\"https:\/\/github.com\/huggingface\/transformers\">transformers<\/a>.<\/p>\n<p>Inkling\u2019s full weights are on <a href=\"https:\/\/huggingface.co\/thinkingmachines\/inkling\">Hugging Face<\/a>, both as the original checkpoint and as an NVFP4 checkpoint for efficient inference on NVIDIA Blackwell systems.<\/p>\n<\/p><\/div>\n<p><a href=\"https:\/\/thinkingmachines.ai\/news\/introducing-inkling\/?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Our mission is to build AI that extends human will and judgment. We have developed a platform that lets anyone customize models, previewed an AI system built for interactive collaboration, and published novel research. Today we are advancing our mission by releasing a model we trained from scratch with the full weights available, so that [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22510,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22509","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22509","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22509"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22509\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22510"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22509"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22509"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22509"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}