{"id":22951,"date":"2026-07-30T00:00:00","date_gmt":"2026-07-30T00:00:00","guid":{"rendered":"https:\/\/scannn.com\/introducing-inkling-small-thinking-machines-lab\/"},"modified":"2026-07-30T00:00:00","modified_gmt":"2026-07-30T00:00:00","slug":"introducing-inkling-small-thinking-machines-lab","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/introducing-inkling-small-thinking-machines-lab\/","title":{"rendered":"Introducing Inkling-Small - Thinking Machines Lab"},"content":{"rendered":"\n<div id=\"\">\n<p>Today, we are releasing Inkling-Small, an efficient open-weights model that achieves comparable performance to <a href=\"https:\/\/thinkingmachines.ai\/news\/introducing-inkling\/\">Inkling<\/a> at a quarter of its size.<\/p>\n<p>Inkling-Small is a Mixture-of-Experts transformer with 276B total parameters, 12B active, trained on NVIDIA GB300 NVL72 systems. Like Inkling, it features native reasoning over audio and images, variable thinking effort, a context window of up to 1M tokens, and well-rounded performance across a range of benchmarks.<\/p>\n<figure class=\"model-parameter-comparison\"><figcaption>Parameter comparison between Inkling-Small and Inkling.<\/figcaption><\/figure>\n<p>Compared to Inkling, Inkling-Small achieves comparable performance with much less compute. Across agentic tool use (Terminal-Bench 2.1), reasoning (HLE text-only, no tools), and instruction following (IFBench) benchmarks, Inkling-Small is more efficient than Inkling and competitive with other models in its weight class. Furthermore, its <a href=\"http:\/\/thinkingmachines.ai\/news\/introducing-inkling\/#controllable-thinking-effort\">variable thinking effort<\/a> lets users easily adapt it to target their use case, balancing cost and performance.<\/p>\n<figure class=\"small-efficiency-switcher\" data-se-switch=\"\">\n<div class=\"se-panels\">\n<div class=\"se-panel\" id=\"se-panel-flops\" role=\"tabpanel\" aria-labelledby=\"se-tab-flops\" data-se-panel=\"flops\">\n<div class=\"ce-figure ce-figure--3 intelligence-interactivity-figure\" data-plot-point-stagger=\"14\" data-plot-point-jitter=\"8\">\n<p>\n      <span class=\"ii-legend-item\"><span class=\"ce-swatch ce-swatch--hero\" aria-hidden=\"true\" style=\"--ce-swatch-color: var(--ce-inklingsmall);\"\/>Inkling-Small <span class=\"ce-legend-meta\">12B active<\/span> <span class=\"ce-legend-meta\">(effort sweep)<\/span><\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ce-swatch\" aria-hidden=\"true\" style=\"--ce-swatch-color: var(--ce-inkling);\"\/>Inkling <span class=\"ce-legend-meta\">41B active<\/span> <span class=\"ce-legend-meta\">(effort sweep)<\/span><\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ii-legend-marker ii-marker-circle\" style=\"--ii-marker-color: var(--ce-peer);\" aria-hidden=\"true\"\/>Comparison models<\/span>\n  <\/p>\n<\/div><\/div>\n<div class=\"se-panel\" id=\"se-panel-cost\" role=\"tabpanel\" aria-labelledby=\"se-tab-cost\" data-se-panel=\"cost\" hidden=\"\">\n<div class=\"ce-figure ce-figure--3 intelligence-interactivity-figure\" data-plot-point-stagger=\"14\" data-plot-point-jitter=\"8\">\n<p>\n      <span class=\"ii-legend-item\"><span class=\"ce-swatch ce-swatch--hero\" aria-hidden=\"true\" style=\"--ce-swatch-color: var(--ce-inklingsmall);\"\/>Inkling-Small <span class=\"ce-legend-meta\">12B active<\/span> <span class=\"ce-legend-meta\">(effort sweep)<\/span><\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ce-swatch\" aria-hidden=\"true\" style=\"--ce-swatch-color: var(--ce-inkling);\"\/>Inkling <span class=\"ce-legend-meta\">41B active<\/span> <span class=\"ce-legend-meta\">(effort sweep)<\/span><\/span><br \/>\n      <span class=\"ii-legend-item\"><span class=\"ii-legend-marker ii-marker-circle\" style=\"--ii-marker-color: var(--ce-peer);\" aria-hidden=\"true\"\/>Comparison models<\/span>\n  <\/p>\n<\/div><\/div>\n<\/div><figcaption class=\"small-efficiency-figure-caption\"><span class=\"se-cap\" data-se-cap=\"flops\" data-figure-number=\"2a\">Performance-Compute Comparison. Sweeping <a href=\"http:\/\/thinkingmachines.ai\/news\/introducing-inkling\/#controllable-thinking-effort\">reasoning effort<\/a> from <a href=\"https:\/\/tinker-docs.thinkingmachines.ai\/cookbook\/inkling\/thinking-effort\/\">minimal to xhigh<\/a> traces the performance-compute curve (output TFLOPs per sample) for Inkling-Small and Inkling on Terminal-Bench 2.1, HLE (no tool), and IFBench. We show that Inkling-Small is competitive with other open-weights models in a similar size range on both performance and efficiency. Output TFLOPs per sample are estimated as 2 \u00d7 active parameters \u00d7 mean generated tokens per sample, where generated-token counts include reasoning tokens and come from our evaluations or public reports from <a href=\"https:\/\/artificialanalysis.ai\/evaluations\/\">Artificial Analysis<\/a>.<\/span><span class=\"se-cap\" data-se-cap=\"cost\" data-figure-number=\"2b\" hidden=\"\">Performance-Cost Comparison. Sweeping <a href=\"https:\/\/thinkingmachines.ai\/news\/introducing-inkling\/#controllable-thinking-effort\">reasoning effort<\/a> from <a href=\"https:\/\/tinker-docs.thinkingmachines.ai\/cookbook\/inkling\/thinking-effort\/\">minimal to xhigh<\/a> traces the performance-cost curve (dollar output price per sample) for Inkling-Small and Inkling on Terminal-Bench 2.1, HLE (no tool), and IFBench. We show that Inkling-Small is competitive with other open-weights models in a similar size range on both performance and efficiency. Estimated output cost per sample is computed as mean generated tokens per sample \u00d7 output price per token, where generated-token counts include reasoning tokens and come from our evaluations or public reports from <a href=\"https:\/\/artificialanalysis.ai\/evaluations\/\">Artificial Analysis<\/a> (reasoning + answer tokens only). Inkling output pricing is $4.05 \/ 1M tokens and Inkling-Small output pricing is $1.20 \/ 1M tokens; for comparison-model pricing, we use the model provider\u2019s official pricing when possible. When this is not possible, we use pricing from the recommended third party inference provider or pricing from Artificial Analysis, whichever is lower.<\/span><\/figcaption><\/figure>\n<p>We are releasing the <a href=\"https:\/\/huggingface.co\/thinkingmachines\/Inkling-Small\">full weights<\/a> of Inkling-Small. We\u2019re also making it available for fine-tuning on Tinker, and for text, image, and audio chat on <a href=\"https:\/\/tinker.thinkingmachines.ai\/playground?utm_source=blog&amp;utm_campaign=inkling_small_model_release\">Tinker Playground<\/a>.<\/p>\n<h2 id=\"capabilities\">Capabilities<\/h2>\n<p>As we build our model family, we are always iterating on our approach. Inkling-Small began training after its larger counterpart, which let us improve its training process. For example, we made changes to Inkling-Small\u2019s pre-training data mix and machine learning recipe. Additionally, we post-trained an earlier checkpoint, <a href=\"http:\/\/thinkingmachines.ai\/news\/introducing-inkling\/#inkling-small\">Inkling-Small (preview)<\/a>, in part using on-policy distillation with Inkling as the teacher. Starting from that checkpoint, we continued scaling agentic coding RL for two weeks. With these improvements, Inkling-Small surpassed Inkling on reasoning and agentic coding benchmarks. Inkling maintains an advantage on knowledge coverage and factuality.<\/p>\n<figure>\n<div class=\"benchmark-radar-figure\">\n<div class=\"benchmark-radar\" id=\"benchmark-radar-small\">\n<p>\n      <svg class=\"benchmark-radar-svg\" data-benchmark-radar=\"\" viewbox=\"0 0 860 600\" role=\"img\" aria-label=\"Inkling benchmark comparison\" aria-describedby=\"benchmark-radar-desc-small\">\n        <desc id=\"benchmark-radar-desc-small\">Spider chart comparing Inkling-Small, Inkling, DeepSeek V4 Flash, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna on ten evaluations scored from zero to one hundred. Inkling-Small is shown with a solid purple line, Inkling with the heavier cobalt line, and the comparison models with dashed lines. Evaluations without a reported model score are plotted at zero. Hover an evaluation to compare every model&#8217;s score.<\/desc>\n        <g data-radar-grid=\"\"\/>\n        <g data-radar-axes=\"\"\/>\n        <g data-radar-series=\"\"\/>\n        <g data-radar-labels=\"\"\/>\n      <\/svg>\n    <\/p>\n<p>\n      <span class=\"benchmark-radar-legend-pill\" data-radar-legend-pill=\"\" aria-hidden=\"true\"\/><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"inklingsmall\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--inklingsmall benchmark-radar-key--hero\" aria-hidden=\"true\"\/><br \/>\n        Inkling-Small<br \/>\n      <\/button><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"inkling\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--inkling\" aria-hidden=\"true\"\/><br \/>\n        Inkling<br \/>\n      <\/button><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"deepseek\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--deepseek\" aria-hidden=\"true\"\/><br \/>\n        DeepSeek V4 Flash<br \/>\n      <\/button><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"gemini\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--gemini\" aria-hidden=\"true\"\/><br \/>\n        Gemini 3.5 Flash-Lite<br \/>\n      <\/button><br \/>\n      <button class=\"benchmark-radar-legend-item\" type=\"button\" data-radar-model=\"luna\"><br \/>\n        <span class=\"benchmark-radar-key benchmark-radar-key--luna\" aria-hidden=\"true\"\/><br \/>\n        GPT 5.6 Luna<br \/>\n      <\/button>\n    <\/p>\n<\/p><\/div>\n<\/div><figcaption>Inkling-Small is a broad, balanced generalist model. Benchmark scores are shown on a shared 0\u2013100 scale; higher is better.<\/figcaption><\/figure>\n<h3 id=\"reasoning-and-agentic-tasks\">Reasoning and Agentic Tasks<\/h3>\n<p>Inkling-Small matches or exceeds Inkling on reasoning and agentic tasks. On Humanity\u2019s Last Exam it scores 31.6%, ahead of Inkling\u2019s 29.7%, and the advantage holds at every thinking budget: Inkling-Small\u2019s test-time compute curves sit above Inkling\u2019s throughout. On SWEBench-Verified it exceeds 80%.<\/p>\n<p>Across many reasoning and agentic benchmarks, Inkling-Small at max reasoning effort has a strong performance-token tradeoff when compared to open weights models in its weight class.<\/p>\n<figure>\n<div class=\"cq-figure intelligence-interactivity-figure\" data-plot-point-stagger=\"14\" data-plot-point-jitter=\"12\">\n<p>\n    <span class=\"ii-legend-item\"><span class=\"cq-legend-dot\" style=\"--cq-color: var(--cq-inklingsmall);\" aria-hidden=\"true\"\/>Inkling-Small<\/span><br \/>\n    <span class=\"ii-legend-item\"><span class=\"cq-legend-dot\" style=\"--cq-color: var(--cq-inkling);\" aria-hidden=\"true\"\/>Inkling<\/span><br \/>\n    <span class=\"ii-legend-item\"><span class=\"cq-legend-dot\" style=\"--cq-color: var(--cq-peer);\" aria-hidden=\"true\"\/>Comparison models<\/span><br \/>\n    <span class=\"ii-legend-item\"><span class=\"cq-frontier-key\" aria-hidden=\"true\"\/>Pareto frontier<\/span>\n  <\/p>\n<div class=\"cq-grid-panels intelligence-interactivity-charts\">\n<div class=\"cq-cell intelligence-interactivity-chart\">\n      <svg class=\"cq-svg intelligence-interactivity-svg\" viewbox=\"0 0 460 300\" role=\"img\" aria-label=\"GDPval-AA v2 Elo against output tokens per task. Inkling-Small and Inkling are highlighted; the dashed line marks the non-dominated Pareto frontier.\">\n      <text class=\"cq-title\" x=\"252\" y=\"22\" text-anchor=\"middle\">GDPval-AA v2 (Elo)<\/text>\n      <g class=\"cq-grid\">\n        <line x1=\"56\" x2=\"448\" y1=\"272\" y2=\"272\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"201.9\" y2=\"201.9\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"131.9\" y2=\"131.9\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"61.8\" y2=\"61.8\"\/>\n        <line x1=\"69.9\" x2=\"69.9\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"219.2\" x2=\"219.2\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"368.4\" x2=\"368.4\" y1=\"44\" y2=\"272\"\/>\n      <\/g>\n      <g class=\"cq-axis\">\n        <line x1=\"56\" x2=\"56\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"272\" y2=\"272\"\/>\n      <\/g>\n      <g class=\"cq-ticks\">\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"276\">900<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"205.9\">1100<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"135.9\">1300<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"65.8\">1500<\/text>\n        <text class=\"cq-tick\" x=\"69.9\" y=\"292\">15k<\/text>\n        <text class=\"cq-tick\" x=\"219.2\" y=\"292\">30k<\/text>\n        <text class=\"cq-tick\" x=\"368.4\" y=\"292\">60k<\/text>\n      <\/g>\n      <polyline class=\"cq-frontier\" points=\"77,250.3 82.7,233.8 83.5,186.2 160.5,144.1 162.1,142.7 231.4,129.8 297.1,100.1 427,56.9\"\/>\n      <g class=\"cq-points\">\n        <circle class=\"cq-point ii-point cq-point--inklingsmall\" cx=\"162.1\" cy=\"142.7\" r=\"5.6\" data-ii-tooltip-title=\"Inkling-Small\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1269 \u00b7 23k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: Inkling-Small, 1269, 23k output tokens per task\"><title>GDPval-AA v2 \u00b7 Inkling-Small: 1269, 23k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point cq-point--inkling\" cx=\"208.9\" cy=\"153.6\" r=\"5.6\" data-ii-tooltip-title=\"Inkling\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1238 \u00b7 28.6k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: Inkling, 1238, 28.6k output tokens per task\"><title>GDPval-AA v2 \u00b7 Inkling: 1238, 28.6k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"198.4\" cy=\"179.5\" r=\"4.2\" data-ii-tooltip-title=\"Nemotron 3 Ultra\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1164 \u00b7 27.2k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: Nemotron 3 Ultra, 1164, 27.2k output tokens per task\"><title>GDPval-AA v2 \u00b7 Nemotron 3 Ultra: 1164, 27.2k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"205.9\" cy=\"170.8\" r=\"4.2\" data-ii-tooltip-title=\"DeepSeek V4 Flash\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1189 \u00b7 28.2k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: DeepSeek V4 Flash, 1189, 28.2k output tokens per task\"><title>GDPval-AA v2 \u00b7 DeepSeek V4 Flash: 1189, 28.2k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"231.4\" cy=\"129.8\" r=\"4.2\" data-ii-tooltip-title=\"DeepSeek V4 Pro\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1306 \u00b7 31.8k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: DeepSeek V4 Pro, 1306, 31.8k output tokens per task\"><title>GDPval-AA v2 \u00b7 DeepSeek V4 Pro: 1306, 31.8k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"77\" cy=\"250.3\" r=\"4.2\" data-ii-tooltip-title=\"Qwen3.5-397B-A17B\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 962 \u00b7 15.5k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: Qwen3.5-397B-A17B, 962, 15.5k output tokens per task\"><title>GDPval-AA v2 \u00b7 Qwen3.5-397B-A17B: 962, 15.5k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"83.5\" cy=\"186.2\" r=\"4.2\" data-ii-tooltip-title=\"MiMo V2.5\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1145 \u00b7 16k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: MiMo V2.5, 1145, 16k output tokens per task\"><title>GDPval-AA v2 \u00b7 MiMo V2.5: 1145, 16k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"160.5\" cy=\"144.1\" r=\"4.2\" data-ii-tooltip-title=\"MiMo V2.5 Pro\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1265.1 \u00b7 22.8k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: MiMo V2.5 Pro, 1265.1, 22.8k output tokens per task\"><title>GDPval-AA v2 \u00b7 MiMo V2.5 Pro: 1265.1, 22.8k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"171.6\" cy=\"181.3\" r=\"4.2\" data-ii-tooltip-title=\"Minimax M2.7\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1159 \u00b7 24.1k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: Minimax M2.7, 1159, 24.1k output tokens per task\"><title>GDPval-AA v2 \u00b7 Minimax M2.7: 1159, 24.1k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"297.1\" cy=\"100.1\" r=\"4.2\" data-ii-tooltip-title=\"MiniMax M3\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1390.8 \u00b7 43.1k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: MiniMax M3, 1390.8, 43.1k output tokens per task\"><title>GDPval-AA v2 \u00b7 MiniMax M3: 1390.8, 43.1k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"82.7\" cy=\"233.8\" r=\"4.2\" data-ii-tooltip-title=\"Kimi K2.5\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1009 \u00b7 15.9k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: Kimi K2.5, 1009, 15.9k output tokens per task\"><title>GDPval-AA v2 \u00b7 Kimi K2.5: 1009, 15.9k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"182.4\" cy=\"170.4\" r=\"4.2\" data-ii-tooltip-title=\"Kimi K2.6\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1190 \u00b7 25.3k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: Kimi K2.6, 1190, 25.3k output tokens per task\"><title>GDPval-AA v2 \u00b7 Kimi K2.6: 1190, 25.3k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"427\" cy=\"56.9\" r=\"4.2\" data-ii-tooltip-title=\"GLM 5.2\" data-ii-tooltip-values=\"GDPval-AA v2 \u00b7 1514 \u00b7 78.8k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"GDPval-AA v2: GLM 5.2, 1514, 78.8k output tokens per task\"><title>GDPval-AA v2 \u00b7 GLM 5.2: 1514, 78.8k output tokens\/task<\/title><\/circle>\n      <\/g>\n      <g class=\"cq-labels\">\n        <text class=\"cq-label cq-label--inklingsmall\" x=\"162.1\" y=\"130.1\" text-anchor=\"middle\">Inkling-Small<\/text>\n        <text class=\"cq-label cq-label--inkling\" x=\"220.5\" y=\"157.6\" text-anchor=\"start\">Inkling<\/text>\n      <\/g>\n      <\/svg><\/p>\n<p class=\"cq-cell-xlabel\">Output tokens\/task<\/p>\n<\/p><\/div>\n<div class=\"cq-cell intelligence-interactivity-chart\">\n      <svg class=\"cq-svg intelligence-interactivity-svg\" viewbox=\"0 0 460 300\" role=\"img\" aria-label=\"Tau-cubed Banking score against output tokens per task. Inkling-Small and Inkling are highlighted; the dashed line marks the non-dominated Pareto frontier.\">\n      <text class=\"cq-title\" x=\"252\" y=\"22\" text-anchor=\"middle\">\u03c4\u00b3-Banking<\/text>\n      <g class=\"cq-grid\">\n        <line x1=\"56\" x2=\"448\" y1=\"272\" y2=\"272\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"226.4\" y2=\"226.4\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"180.8\" y2=\"180.8\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"135.2\" y2=\"135.2\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"89.6\" y2=\"89.6\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"44\" y2=\"44\"\/>\n        <line x1=\"75.2\" x2=\"75.2\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"254.6\" x2=\"254.6\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"434\" x2=\"434\" y1=\"44\" y2=\"272\"\/>\n      <\/g>\n      <g class=\"cq-axis\">\n        <line x1=\"56\" x2=\"56\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"272\" y2=\"272\"\/>\n      <\/g>\n      <g class=\"cq-ticks\">\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"276\">5%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"230.4\">10%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"184.8\">15%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"139.2\">20%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"93.6\">25%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"48\">30%<\/text>\n        <text class=\"cq-tick\" x=\"75.2\" y=\"292\">5k<\/text>\n        <text class=\"cq-tick\" x=\"254.6\" y=\"292\">10k<\/text>\n        <text class=\"cq-tick\" x=\"434\" y=\"292\">20k<\/text>\n      <\/g>\n      <polyline class=\"cq-frontier\" points=\"77,199 79.1,176.2 210.5,101.5 251.8,80.5 427,73.2\"\/>\n      <g class=\"cq-points\">\n        <circle class=\"cq-point ii-point cq-point--inklingsmall\" cx=\"79.1\" cy=\"176.2\" r=\"5.6\" data-ii-tooltip-title=\"Inkling-Small\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 15.5% \u00b7 5.1k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: Inkling-Small, 15.5%, 5.1k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 Inkling-Small: 15.5%, 5.1k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point cq-point--inkling\" cx=\"210.5\" cy=\"101.5\" r=\"5.6\" data-ii-tooltip-title=\"Inkling\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 23.7% \u00b7 8.4k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: Inkling, 23.7%, 8.4k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 Inkling: 23.7%, 8.4k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"242.3\" cy=\"191.7\" r=\"4.2\" data-ii-tooltip-title=\"Nemotron 3 Ultra\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 13.8% \u00b7 9.5k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: Nemotron 3 Ultra, 13.8%, 9.5k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 Nemotron 3 Ultra: 13.8%, 9.5k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"281.8\" cy=\"108.8\" r=\"4.2\" data-ii-tooltip-title=\"DeepSeek V4 Flash\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 22.9% \u00b7 11.1k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: DeepSeek V4 Flash, 22.9%, 11.1k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 DeepSeek V4 Flash: 22.9%, 11.1k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"251.8\" cy=\"80.5\" r=\"4.2\" data-ii-tooltip-title=\"DeepSeek V4 Pro\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 26.0% \u00b7 9.9k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: DeepSeek V4 Pro, 26.0%, 9.9k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 DeepSeek V4 Pro: 26.0%, 9.9k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"187.3\" cy=\"195.4\" r=\"4.2\" data-ii-tooltip-title=\"Qwen3.5-397B-A17B\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 13.4% \u00b7 7.7k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: Qwen3.5-397B-A17B, 13.4%, 7.7k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 Qwen3.5-397B-A17B: 13.4%, 7.7k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"150.4\" cy=\"257.4\" r=\"4.2\" data-ii-tooltip-title=\"MiMo V2.5\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 6.6% \u00b7 6.7k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: MiMo V2.5, 6.6%, 6.7k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 MiMo V2.5: 6.6%, 6.7k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"168.6\" cy=\"235.5\" r=\"4.2\" data-ii-tooltip-title=\"MiMo V2.5 Pro\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 9.0% \u00b7 7.2k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: MiMo V2.5 Pro, 9.0%, 7.2k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 MiMo V2.5 Pro: 9.0%, 7.2k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"113.9\" cy=\"236.4\" r=\"4.2\" data-ii-tooltip-title=\"Minimax M2.7\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 8.9% \u00b7 5.8k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: Minimax M2.7, 8.9%, 5.8k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 Minimax M2.7: 8.9%, 5.8k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"77\" cy=\"199\" r=\"4.2\" data-ii-tooltip-title=\"MiniMax M3\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 13.0% \u00b7 5k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: MiniMax M3, 13.0%, 5k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 MiniMax M3: 13.0%, 5k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"103.8\" cy=\"188.1\" r=\"4.2\" data-ii-tooltip-title=\"Kimi K2.5\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 14.2% \u00b7 5.6k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: Kimi K2.5, 14.2%, 5.6k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 Kimi K2.5: 14.2%, 5.6k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"274.9\" cy=\"129.7\" r=\"4.2\" data-ii-tooltip-title=\"Kimi K2.6\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 20.6% \u00b7 10.8k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: Kimi K2.6, 20.6%, 10.8k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 Kimi K2.6: 20.6%, 10.8k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"427\" cy=\"73.2\" r=\"4.2\" data-ii-tooltip-title=\"GLM 5.2\" data-ii-tooltip-values=\"\u03c4\u00b3-Banking \u00b7 26.8% \u00b7 19.5k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"\u03c4\u00b3-Banking: GLM 5.2, 26.8%, 19.5k output tokens per task\"><title>\u03c4\u00b3-Banking \u00b7 GLM 5.2: 26.8%, 19.5k output tokens\/task<\/title><\/circle>\n      <\/g>\n      <g class=\"cq-labels\">\n        <text class=\"cq-label cq-label--inklingsmall\" x=\"89.7\" y=\"165.6\" text-anchor=\"start\">Inkling-Small<\/text>\n        <text class=\"cq-label cq-label--inkling\" x=\"198.9\" y=\"105.5\" text-anchor=\"end\">Inkling<\/text>\n      <\/g>\n      <\/svg><\/p>\n<p class=\"cq-cell-xlabel\">Output tokens\/task<\/p>\n<\/p><\/div>\n<div class=\"cq-cell intelligence-interactivity-chart\">\n      <svg class=\"cq-svg intelligence-interactivity-svg\" viewbox=\"0 0 460 300\" role=\"img\" aria-label=\"AA-Briefcase Elo against output tokens per task. Inkling-Small and Inkling are highlighted; the dashed line marks the non-dominated Pareto frontier.\">\n      <text class=\"cq-title\" x=\"252\" y=\"22\" text-anchor=\"middle\">AA-Briefcase (Elo)<\/text>\n      <g class=\"cq-grid\">\n        <line x1=\"56\" x2=\"448\" y1=\"272\" y2=\"272\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"196\" y2=\"196\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"120\" y2=\"120\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"44\" y2=\"44\"\/>\n        <line x1=\"85.2\" x2=\"85.2\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"261.4\" x2=\"261.4\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"391.2\" x2=\"391.2\" y1=\"44\" y2=\"272\"\/>\n      <\/g>\n      <g class=\"cq-axis\">\n        <line x1=\"56\" x2=\"56\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"272\" y2=\"272\"\/>\n      <\/g>\n      <g class=\"cq-ticks\">\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"276\">700<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"200\">900<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"124\">1100<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"48\">1300<\/text>\n        <text class=\"cq-tick\" x=\"85.2\" y=\"292\">30k<\/text>\n        <text class=\"cq-tick\" x=\"261.4\" y=\"292\">60k<\/text>\n        <text class=\"cq-tick\" x=\"391.2\" y=\"292\">100k<\/text>\n      <\/g>\n      <polyline class=\"cq-frontier\" points=\"77,189.5 340.5,117 427,56.9\"\/>\n      <g class=\"cq-points\">\n        <circle class=\"cq-point ii-point cq-point--inklingsmall\" cx=\"77\" cy=\"189.5\" r=\"5.6\" data-ii-tooltip-title=\"Inkling-Small\" data-ii-tooltip-values=\"AA-Briefcase \u00b7 917 \u00b7 29k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"AA-Briefcase: Inkling-Small, 917, 29k output tokens per task\"><title>AA-Briefcase \u00b7 Inkling-Small: 917, 29k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point cq-point--inkling\" cx=\"226.1\" cy=\"219.2\" r=\"5.6\" data-ii-tooltip-title=\"Inkling\" data-ii-tooltip-values=\"AA-Briefcase \u00b7 839 \u00b7 52.2k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"AA-Briefcase: Inkling, 839, 52.2k output tokens per task\"><title>AA-Briefcase \u00b7 Inkling: 839, 52.2k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"202.2\" cy=\"207.4\" r=\"4.2\" data-ii-tooltip-title=\"Nemotron 3 Ultra\" data-ii-tooltip-values=\"AA-Briefcase \u00b7 870 \u00b7 47.5k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"AA-Briefcase: Nemotron 3 Ultra, 870, 47.5k output tokens per task\"><title>AA-Briefcase \u00b7 Nemotron 3 Ultra: 870, 47.5k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"142.8\" cy=\"221.5\" r=\"4.2\" data-ii-tooltip-title=\"DeepSeek V4 Flash\" data-ii-tooltip-values=\"AA-Briefcase \u00b7 833 \u00b7 37.6k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"AA-Briefcase: DeepSeek V4 Flash, 833, 37.6k output tokens per task\"><title>AA-Briefcase \u00b7 DeepSeek V4 Flash: 833, 37.6k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"108.2\" cy=\"204\" r=\"4.2\" data-ii-tooltip-title=\"MiMo V2.5 Pro\" data-ii-tooltip-values=\"AA-Briefcase \u00b7 878.9 \u00b7 32.8k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"AA-Briefcase: MiMo V2.5 Pro, 878.9, 32.8k output tokens per task\"><title>AA-Briefcase \u00b7 MiMo V2.5 Pro: 878.9, 32.8k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"340.5\" cy=\"117\" r=\"4.2\" data-ii-tooltip-title=\"MiniMax M3\" data-ii-tooltip-values=\"AA-Briefcase \u00b7 1107.8 \u00b7 81.9k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"AA-Briefcase: MiniMax M3, 1107.8, 81.9k output tokens per task\"><title>AA-Briefcase \u00b7 MiniMax M3: 1107.8, 81.9k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"427\" cy=\"56.9\" r=\"4.2\" data-ii-tooltip-title=\"GLM 5.2\" data-ii-tooltip-values=\"AA-Briefcase \u00b7 1266 \u00b7 115k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"AA-Briefcase: GLM 5.2, 1266, 115k output tokens per task\"><title>AA-Briefcase \u00b7 GLM 5.2: 1266, 115k output tokens\/task<\/title><\/circle>\n      <\/g>\n      <g class=\"cq-labels\">\n        <text class=\"cq-label cq-label--inklingsmall\" x=\"88.6\" y=\"193.5\" text-anchor=\"start\">Inkling-Small<\/text>\n        <text class=\"cq-label cq-label--inkling\" x=\"237.7\" y=\"223.2\" text-anchor=\"start\">Inkling<\/text>\n      <\/g>\n      <\/svg><\/p>\n<p class=\"cq-cell-xlabel\">Output tokens\/task<\/p>\n<\/p><\/div>\n<div class=\"cq-cell intelligence-interactivity-chart\">\n      <svg class=\"cq-svg intelligence-interactivity-svg\" viewbox=\"0 0 460 300\" role=\"img\" aria-label=\"CritPt score against output tokens per task. Inkling-Small and Inkling are highlighted; the dashed line marks the non-dominated Pareto frontier.\">\n      <text class=\"cq-title\" x=\"252\" y=\"22\" text-anchor=\"middle\">CritPt<\/text>\n      <g class=\"cq-grid\">\n        <line x1=\"56\" x2=\"448\" y1=\"272\" y2=\"272\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"226.4\" y2=\"226.4\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"180.8\" y2=\"180.8\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"135.2\" y2=\"135.2\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"89.6\" y2=\"89.6\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"44\" y2=\"44\"\/>\n        <line x1=\"80.4\" x2=\"80.4\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"276.3\" x2=\"276.3\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"424.4\" x2=\"424.4\" y1=\"44\" y2=\"272\"\/>\n      <\/g>\n      <g class=\"cq-axis\">\n        <line x1=\"56\" x2=\"56\" y1=\"44\" y2=\"272\"\/>\n        <line x1=\"56\" x2=\"448\" y1=\"272\" y2=\"272\"\/>\n      <\/g>\n      <g class=\"cq-ticks\">\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"276\">0%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"230.4\">5%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"184.8\">10%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"139.2\">15%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"93.6\">20%<\/text>\n        <text class=\"cq-tick cq-tick--y\" x=\"48\" y=\"48\">25%<\/text>\n        <text class=\"cq-tick\" x=\"80.4\" y=\"292\">40k<\/text>\n        <text class=\"cq-tick\" x=\"276.3\" y=\"292\">100k<\/text>\n        <text class=\"cq-tick\" x=\"424.4\" y=\"292\">200k<\/text>\n      <\/g>\n      <polyline class=\"cq-frontier\" points=\"77,266.5 78.1,243.7 86.6,238.3 161.5,235.5 226.8,222.8 277.7,196.3 285.8,153.4 289.6,81.4\"\/>\n      <g class=\"cq-points\">\n        <circle class=\"cq-point ii-point cq-point--inklingsmall\" cx=\"277.7\" cy=\"196.3\" r=\"5.6\" data-ii-tooltip-title=\"Inkling-Small\" data-ii-tooltip-values=\"CritPt \u00b7 8.3% \u00b7 101k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: Inkling-Small, 8.3%, 101k output tokens per task\"><title>CritPt \u00b7 Inkling-Small: 8.3%, 101k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point cq-point--inkling\" cx=\"226.8\" cy=\"222.8\" r=\"5.6\" data-ii-tooltip-title=\"Inkling\" data-ii-tooltip-values=\"CritPt \u00b7 5.4% \u00b7 79.3k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: Inkling, 5.4%, 79.3k output tokens per task\"><title>CritPt \u00b7 Inkling: 5.4%, 79.3k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"235.2\" cy=\"243.7\" r=\"4.2\" data-ii-tooltip-title=\"Nemotron 3 Ultra\" data-ii-tooltip-values=\"CritPt \u00b7 3.1% \u00b7 82.5k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: Nemotron 3 Ultra, 3.1%, 82.5k output tokens per task\"><title>CritPt \u00b7 Nemotron 3 Ultra: 3.1%, 82.5k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"427\" cy=\"207.2\" r=\"4.2\" data-ii-tooltip-title=\"DeepSeek V4 Flash\" data-ii-tooltip-values=\"CritPt \u00b7 7.1% \u00b7 202k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: DeepSeek V4 Flash, 7.1%, 202k output tokens per task\"><title>CritPt \u00b7 DeepSeek V4 Flash: 7.1%, 202k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"285.8\" cy=\"153.4\" r=\"4.2\" data-ii-tooltip-title=\"DeepSeek V4 Pro\" data-ii-tooltip-values=\"CritPt \u00b7 13.0% \u00b7 105k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: DeepSeek V4 Pro, 13.0%, 105k output tokens per task\"><title>CritPt \u00b7 DeepSeek V4 Pro: 13.0%, 105k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"106.3\" cy=\"256.5\" r=\"4.2\" data-ii-tooltip-title=\"Qwen3.5-397B-A17B\" data-ii-tooltip-values=\"CritPt \u00b7 1.7% \u00b7 45.2k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: Qwen3.5-397B-A17B, 1.7%, 45.2k output tokens per task\"><title>CritPt \u00b7 Qwen3.5-397B-A17B: 1.7%, 45.2k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"86.6\" cy=\"238.3\" r=\"4.2\" data-ii-tooltip-title=\"MiMo V2.5\" data-ii-tooltip-values=\"CritPt \u00b7 3.7% \u00b7 41.2k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: MiMo V2.5, 3.7%, 41.2k output tokens per task\"><title>CritPt \u00b7 MiMo V2.5: 3.7%, 41.2k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"161.5\" cy=\"235.5\" r=\"4.2\" data-ii-tooltip-title=\"MiMo V2.5 Pro\" data-ii-tooltip-values=\"CritPt \u00b7 4.0% \u00b7 58.4k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: MiMo V2.5 Pro, 4.0%, 58.4k output tokens per task\"><title>CritPt \u00b7 MiMo V2.5 Pro: 4.0%, 58.4k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"77\" cy=\"266.5\" r=\"4.2\" data-ii-tooltip-title=\"Minimax M2.7\" data-ii-tooltip-values=\"CritPt \u00b7 0.6% \u00b7 39.4k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: Minimax M2.7, 0.6%, 39.4k output tokens per task\"><title>CritPt \u00b7 Minimax M2.7: 0.6%, 39.4k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"222.9\" cy=\"238.1\" r=\"4.2\" data-ii-tooltip-title=\"MiniMax M3\" data-ii-tooltip-values=\"CritPt \u00b7 3.7% \u00b7 77.9k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: MiniMax M3, 3.7%, 77.9k output tokens per task\"><title>CritPt \u00b7 MiniMax M3: 3.7%, 77.9k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"78.1\" cy=\"243.7\" r=\"4.2\" data-ii-tooltip-title=\"Kimi K2.5\" data-ii-tooltip-values=\"CritPt \u00b7 3.1% \u00b7 39.6k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: Kimi K2.5, 3.1%, 39.6k output tokens per task\"><title>CritPt \u00b7 Kimi K2.5: 3.1%, 39.6k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"405.2\" cy=\"199\" r=\"4.2\" data-ii-tooltip-title=\"Kimi K2.6\" data-ii-tooltip-values=\"CritPt \u00b7 8.0% \u00b7 183k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: Kimi K2.6, 8.0%, 183k output tokens per task\"><title>CritPt \u00b7 Kimi K2.6: 8.0%, 183k output tokens\/task<\/title><\/circle>\n        <circle class=\"cq-point ii-point\" cx=\"289.6\" cy=\"81.4\" r=\"4.2\" data-ii-tooltip-title=\"GLM 5.2\" data-ii-tooltip-values=\"CritPt \u00b7 20.9% \u00b7 106k tokens\/task\" tabindex=\"0\" role=\"img\" aria-label=\"CritPt: GLM 5.2, 20.9%, 106k output tokens per task\"><title>CritPt \u00b7 GLM 5.2: 20.9%, 106k output tokens\/task<\/title><\/circle>\n      <\/g>\n      <g class=\"cq-labels\">\n        <text class=\"cq-label cq-label--inklingsmall\" x=\"288.3\" y=\"185.7\" text-anchor=\"start\">Inkling-Small<\/text>\n        <text class=\"cq-label cq-label--inkling\" x=\"238.4\" y=\"226.8\" text-anchor=\"start\">Inkling<\/text>\n      <\/g>\n      <\/svg><\/p>\n<p class=\"cq-cell-xlabel\">Output tokens\/task<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p class=\"cq-xlabel\">Output tokens\/task<\/p>\n<\/div><figcaption>Token-Efficiency Performance Tradeoff on Reasoning + Agentic Benchmarks. We evaluate Inkling-Small and Inkling (with max effort) along with other open-weights models on agentic and reasoning tasks (GDPval-AA v2, \u03c4\u00b3-Banking, AA-Briefcase and CritPt) and show performance and output token length (including reasoning + answer). Inkling-Small is among the most efficient open weights models, marked by the dashed line. Results were obtained from our evaluation or reference data from Artificial Analysis.<\/figcaption><\/figure>\n<p>Inkling-Small also runs smoothly across a variety of coding and agent harnesses, making it a cost-efficient choice for coding and tool-use workflows.<\/p>\n<h3 id=\"multimodality\">Multimodality<\/h3>\n<p>Like Inkling, we crafted Inkling-Small for audio intelligence, making it a good candidate for real-world audio applications. We also improved its ability to use Python for visual tasks. The model can combine visual reasoning with operations such as cropping, zooming, and programmatic image inspection, improving usability on documents and charts where relevant information may be small or difficult to inspect directly.<\/p>\n<p>Inkling-Small uses the same natively multimodal encoder-free architecture as Inkling. Audio is represented as dMel spectrograms, while images are divided into 40\u00d740-pixel patches and transformed using a four-layer hMLP. Both are transformed via a light-weight embedding layer and processed jointly with text tokens. Inkling-Small nearly matches Inkling across most multimodal evaluations at a lower cost. It retains strong performance on visual reasoning, chart and diagram understanding, mathematical visual question answering, speech understanding, and longer-form audio reasoning.<\/p>\n<p class=\"benchmark-results-note\">Audio and vision benchmarks against specialist omni models (open- and closed-weight), reported at effort=0.99.<\/p>\n<h3 id=\"epistemics\">Epistemics<\/h3>\n<p>Inkling-Small was trained similarly to Inkling on epistemics, focusing on calibration, instruction following, and resistance to censorship. Inkling-Small matches Inkling\u2019s performance on forecasting. Calibration involved RL against proper scoring rules on a large corpus of real-world forecasting questions, improving its ability to express appropriate confidence and produce calibrated forecasts under uncertainty.<\/p>\n<p class=\"benchmark-results-note\">ForecastBench and Prophet Arena results were obtained during testing between July 19 and July 28, 2026.<\/p>\n<h3 id=\"safety\">Safety<\/h3>\n<p>Inkling-Small inherited the same safety post-training recipe as Inkling, with built-in safeguards covering our internal spec of safety. These include everyday human-AI interactions as well as dual-use capabilities. Inkling-Small also underwent the same pre-deployment testing process, comprising both internal evaluations and red-teaming by trusted external partners.<\/p>\n<p>On StrongREJECT, which measures whether models refuse unambiguously harmful requests, Inkling-Small is on par with Inkling, and matches the performance of existing open-weights models. On FORTRESS, which measures safety in settings spanning crime, violence, and dual-use risks, it is competitive in both refusal and over-refusal.<\/p>\n<p class=\"benchmark-results-note\">Safety benchmarks, reported at effort=0.99; higher is better throughout. FORTRESS adversarial is the rate of refusing harmful requests, benign the rate of still answering safe ones.<\/p>\n<h2 id=\"benchmarking-inkling-small\">Benchmarking Inkling-Small<\/h2>\n<p>Like Inkling, we benchmarked Inkling-Small on a broad range of capabilities. All evals run at effort 0.99 and temperature 1.0. All coding evals run with 256K max-token trajectory limit, similar to Inkling.<\/p>\n<p>To improve consistency, we rely on externally reported evaluations for both internal and external models when applicable. Specifically, we use the scores reported by:<\/p>\n<p class=\"benchmark-results-note\">Inkling-Small against open- and closed-weights models across the full eval suite. Activated and total parameters are given for scale; a dash means the score was not available at the time of writing.<span class=\"benchmark-results-note-item\" id=\"eval-note-swebench-verified\"><span class=\"benchmark-results-note-label\">*SWEBench Verified:<\/span> Inkling and Inkling-Small\u2019s numbers are reported using a bash-only harness. We use self-reported numbers for external models.<\/span><span class=\"benchmark-results-note-item\" id=\"eval-note-terminal-bench-2-1\"><span class=\"benchmark-results-note-label\">*Terminal Bench 2.1:<\/span> Inkling and Inkling-Small\u2019s numbers are reported using an internal coding harness. A small number of solutions were found to be contaminated from web search and were assigned a score of 0. We use self-reported numbers for external models where available. Otherwise, we report performance using our internal harness.<\/span><span class=\"benchmark-results-note-item\" id=\"eval-note-audio-mc\"><span class=\"benchmark-results-note-label\">\u2020Audio MC:<\/span> Other models were evaluated internally since they are not on the official leaderboard.<\/span><span class=\"benchmark-results-note-item\" id=\"eval-note-voicebench\"><span class=\"benchmark-results-note-label\">\u2020VoiceBench:<\/span> VoiceBench uses rule-based, hard-coded string matching for grading, making the evaluation sensitive to output-formatting differences. We therefore added a system message instructing models to follow the expected answer format.<\/span><span class=\"benchmark-results-note-item\" id=\"eval-note-hle-with-tools\"><span class=\"benchmark-results-note-label\">\u2020HLE with tools:<\/span> We benchmarked Minimax M2.7, Claude 4.5 Haiku, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna using our internal harness.<\/span><\/p>\n<h2 id=\"try-inkling-small-on-tinker\">Try Inkling-Small on Tinker<\/h2>\n<p>Tinker customers have seen firsthand that the right fine-tuned model can outperform closed models on a variety of tasks, and do so faster and cheaper. We believe Inkling-Small\u2019s combination of broad performance and efficiency will make it easy to experiment with and test in real applications, and we\u2019re excited to see what developers build with it.<\/p>\n<p>Inkling and Inkling-Small are available on Tinker with a limited-time discount and can be chatted with on the <a href=\"https:\/\/tinker.thinkingmachines.ai\/playground?utm_source=blog&amp;utm_campaign=inkling_small_model_release\">Tinker Playground<\/a> using text, image, and audio. All other models on Tinker and checkpoints are also now available on Tinker Playground, billed at the rates listed on our <a href=\"https:\/\/tinker-docs.thinkingmachines.ai\/tinker\/models\/\">pricing page<\/a>.<\/p>\n<p>Inkling-Small was made in pursuit of <a href=\"https:\/\/thinkingmachines.ai\/blog\/the-future-worth-building-is-human\/\">our mission<\/a> to build AI that extends human will and judgment. It\u2019s a capable and efficient model, and an important stepping stone for us as we continue our work.<\/p>\n<\/p><\/div>\n<p><a href=\"https:\/\/thinkingmachines.ai\/news\/inkling-small\/?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Today, we are releasing Inkling-Small, an efficient open-weights model that achieves comparable performance to Inkling at a quarter of its size. Inkling-Small is a Mixture-of-Experts transformer with 276B total parameters, 12B active, trained on NVIDIA GB300 NVL72 systems. Like Inkling, it features native reasoning over audio and images, variable thinking effort, a context window of [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22952,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22951","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22951","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22951"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22951\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22952"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22951"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22951"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22951"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}