{"id":23493,"date":"2026-08-25T18:27:14","date_gmt":"2026-08-25T18:27:14","guid":{"rendered":"https:\/\/scannn.com\/speculative-programmatic-tool-calling-alex-l-zhang\/"},"modified":"2026-08-25T18:27:14","modified_gmt":"2026-08-25T18:27:14","slug":"speculative-programmatic-tool-calling-alex-l-zhang","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/speculative-programmatic-tool-calling-alex-l-zhang\/","title":{"rendered":"Speculative Programmatic Tool Calling | Alex L. Zhang"},"content":{"rendered":"\n<div id=\"\">\n<link rel=\"stylesheet\" href=\"https:\/\/fonts.googleapis.com\/css2?family=IBM+Plex+Mono:wght@400;500;600&amp;family=IBM+Plex+Sans:wght@400;500;600&amp;display=swap\"\/>\n<figure id=\"sptc-teaser\">\n<div class=\"teaser\">\n<p>Fig 1. Speculative Programmatic Tool Calling (sPTC) Example<\/p>\n<div class=\"teaser-grid\">\n<div class=\"teaser-term\">\n<p>\n<span class=\"dot\" id=\"sptc-dot\"\/><span id=\"sptc-title\">the model&#8217;s turn \u2014 streaming<\/span>\n<\/p>\n<\/p><\/div>\n<p>\n      <svg viewbox=\"0 0 420 262\" preserveaspectratio=\"none\" role=\"img\" aria-label=\"Compact timeline driven by the time slider. With speculation: the decode bar grows while six thin sub-call bars and a dependent judge call run inside it, so the execute bar is a short stub of seven cache hits. Without speculation the same program runs its seven calls serially after decode, answering far later.\">\n        <rect data-grow=\"\" x=\"93.4\" data-t0=\"0\" data-t1=\"34\" y=\"28\" height=\"148\" fill=\"var(--sp-run)\" fill-opacity=\".08\"\/>\n        <rect data-grow=\"\" x=\"181.05\" data-t0=\"34\" data-t1=\"47\" y=\"28\" height=\"148\" fill=\"var(--sp-ok)\" fill-opacity=\".08\"\/>\n        <rect data-grow=\"\" x=\"93.4\" data-t0=\"0\" data-t1=\"34\" y=\"190\" height=\"36\" fill=\"var(--sp-run)\" fill-opacity=\".08\"\/>\n        <rect data-grow=\"\" x=\"181.05\" data-t0=\"34\" data-t1=\"112\" y=\"190\" height=\"36\" fill=\"var(--sp-ok)\" fill-opacity=\".06\"\/>\n        <text x=\"8\" y=\"20\" font-family=\"IBM Plex Mono, monospace\" font-size=\"9\" font-weight=\"600\" letter-spacing=\"1\" fill=\"var(--sp-spec)\">WITH SPECULATION<\/text>\n        <text data-show=\"0.5\" x=\"136.92\" y=\"20\" text-anchor=\"middle\" font-family=\"IBM Plex Mono, monospace\" font-size=\"9\" font-weight=\"600\" letter-spacing=\"1.2\" fill=\"var(--sp-run)\">DECODE<\/text>\n        <text data-show=\"34\" x=\"200.16\" y=\"20\" text-anchor=\"middle\" font-family=\"IBM Plex Mono, monospace\" font-size=\"9\" font-weight=\"600\" letter-spacing=\"1.2\" fill=\"var(--sp-ok)\">EXECUTE<\/text>\n        <line data-show=\"34\" x1=\"181.05\" y1=\"28\" x2=\"181.05\" y2=\"226\" stroke=\"var(--sp-faint)\" stroke-width=\"1\" stroke-dasharray=\"3 4\"\/>\n        <g font-family=\"IBM Plex Mono, monospace\" font-size=\"9.5\" fill=\"var(--sp-dim)\" text-anchor=\"end\">\n          <text x=\"86\" y=\"48\">prep<\/text>\n          <text x=\"86\" y=\"89\">sub-calls \u00d76<\/text>\n          <text x=\"86\" y=\"131\">judge<\/text>\n          <text x=\"86\" y=\"153\">execute<\/text>\n          <text x=\"86\" y=\"208\">baseline<\/text>\n        <\/g>\n        <rect data-grow=\"\" x=\"93.4\" data-t0=\"0\" data-t1=\"34\" y=\"38\" height=\"12\" rx=\"2\" fill=\"var(--sp-run)\" fill-opacity=\".22\" stroke=\"var(--sp-run)\" stroke-width=\".9\"\/>\n        <g fill=\"var(--sp-spec)\" fill-opacity=\".28\" stroke=\"var(--sp-spec)\" stroke-width=\".7\">\n          <rect data-grow=\"\" x=\"147.53\" data-t0=\"21\" data-t1=\"32\" y=\"58\" height=\"6\" rx=\"1.5\"\/>\n          <rect data-grow=\"\" x=\"148.28\" data-t0=\"21.3\" data-t1=\"32.3\" y=\"68\" height=\"6\" rx=\"1.5\"\/>\n          <rect data-grow=\"\" x=\"149.09\" data-t0=\"21.6\" data-t1=\"32.6\" y=\"78\" height=\"6\" rx=\"1.5\"\/>\n          <rect data-grow=\"\" x=\"149.84\" data-t0=\"21.9\" data-t1=\"32.9\" y=\"88\" height=\"6\" rx=\"1.5\"\/>\n          <rect data-grow=\"\" x=\"150.59\" data-t0=\"22.2\" data-t1=\"33.2\" y=\"98\" height=\"6\" rx=\"1.5\"\/>\n          <rect data-grow=\"\" x=\"151.4\" data-t0=\"22.5\" data-t1=\"33.5\" y=\"108\" height=\"6\" rx=\"1.5\"\/>\n        <\/g>\n        <g fill=\"var(--sp-ok)\">\n          <circle data-show=\"32\" cx=\"149.79\" cy=\"61\" r=\"2\"\/>\n          <circle data-show=\"32.3\" cx=\"150.3\" cy=\"71\" r=\"2\"\/>\n          <circle data-show=\"32.6\" cx=\"150.85\" cy=\"81\" r=\"2\"\/>\n          <circle data-show=\"32.9\" cx=\"151.36\" cy=\"91\" r=\"2\"\/>\n          <circle data-show=\"33.2\" cx=\"151.87\" cy=\"101\" r=\"2\"\/>\n          <circle data-show=\"33.5\" cx=\"152.42\" cy=\"111\" r=\"2\"\/>\n        <\/g>\n        <rect data-grow=\"\" x=\"179.76\" data-t0=\"33.5\" data-t1=\"45.5\" y=\"124\" height=\"8\" rx=\"1.5\" fill=\"var(--sp-spec)\" fill-opacity=\".28\" stroke=\"var(--sp-spec)\" stroke-width=\".7\"\/>\n        <circle data-show=\"45.5\" cx=\"173.46\" cy=\"128\" r=\"2\" fill=\"var(--sp-ok)\"\/>\n        <rect data-grow=\"\" x=\"181.05\" data-t0=\"34\" data-t1=\"46\" y=\"144\" height=\"11\" rx=\"2\" fill=\"var(--sp-ok)\" fill-opacity=\".22\" stroke=\"var(--sp-ok)\" stroke-width=\"1\"\/>\n        <circle data-show=\"46\" cx=\"174.16\" cy=\"149.5\" r=\"3\" fill=\"var(--sp-ok)\"\/>\n        <text data-show=\"46\" x=\"219.2\" y=\"153\" font-family=\"IBM Plex Mono, monospace\" font-size=\"10\" fill=\"var(--sp-ok)\">answer \u00b7 7\/7 claims<\/text>\n        <text x=\"8\" y=\"185\" font-family=\"IBM Plex Mono, monospace\" font-size=\"9\" font-weight=\"600\" letter-spacing=\"1\" fill=\"var(--sp-dim)\">WITHOUT SPECULATION \u00b7 SAME PROGRAM<\/text>\n        <text data-show=\"112\" x=\"399.4\" y=\"185\" text-anchor=\"end\" font-family=\"IBM Plex Mono, monospace\" font-size=\"10\" fill=\"var(--sp-amber)\">answer<\/text>\n        <rect data-grow=\"\" x=\"93.4\" data-t0=\"0\" data-t1=\"34\" y=\"198\" height=\"12\" rx=\"2\" fill=\"var(--sp-run)\" fill-opacity=\".22\" stroke=\"var(--sp-run)\" stroke-width=\".9\"\/>\n        <g fill=\"var(--sp-amber)\" fill-opacity=\".16\" stroke=\"var(--sp-amber)\" stroke-width=\".9\">\n          <rect data-grow=\"\" x=\"181.05\" data-t0=\"34\" data-t1=\"45\" y=\"198\" height=\"12\" rx=\"2\"\/>\n          <rect data-grow=\"\" x=\"209.41\" data-t0=\"45\" data-t1=\"56\" y=\"198\" height=\"12\" rx=\"2\"\/>\n          <rect data-grow=\"\" x=\"237.76\" data-t0=\"56\" data-t1=\"67\" y=\"198\" height=\"12\" rx=\"2\"\/>\n          <rect data-grow=\"\" x=\"266.12\" data-t0=\"67\" data-t1=\"78\" y=\"198\" height=\"12\" rx=\"2\"\/>\n          <rect data-grow=\"\" x=\"294.41\" data-t0=\"78\" data-t1=\"89\" y=\"198\" height=\"12\" rx=\"2\"\/>\n          <rect data-grow=\"\" x=\"322.76\" data-t0=\"89\" data-t1=\"100\" y=\"198\" height=\"12\" rx=\"2\"\/>\n          <rect data-grow=\"\" x=\"351.12\" data-t0=\"100\" data-t1=\"112\" y=\"198\" height=\"12\" rx=\"2\"\/>\n        <\/g>\n        <g font-family=\"IBM Plex Mono, monospace\" font-size=\"8.5\" fill=\"var(--sp-amber)\" text-anchor=\"middle\">\n          <text data-show=\"39\" x=\"195.2\" y=\"207\">c\u2081<\/text><text data-show=\"50\" x=\"223.55\" y=\"207\">c\u2082<\/text>\n          <text data-show=\"61\" x=\"251.91\" y=\"207\">c\u2083<\/text><text data-show=\"72\" x=\"280.26\" y=\"207\">c\u2084<\/text>\n          <text data-show=\"83\" x=\"308.55\" y=\"207\">c\u2085<\/text><text data-show=\"94\" x=\"336.91\" y=\"207\">c\u2086<\/text>\n          <text data-show=\"105\" x=\"366.56\" y=\"207\">judge<\/text>\n        <\/g>\n        <circle data-show=\"112\" cx=\"289.76\" cy=\"204\" r=\"3\" fill=\"var(--sp-amber)\"\/>\n        <text data-show=\"112\" x=\"296.72\" y=\"238\" text-anchor=\"middle\" font-family=\"IBM Plex Sans, sans-serif\" font-size=\"10\" fill=\"var(--sp-ok)\">execute collapses \u2014 2.4\u00d7, no async<\/text>\n        <g data-show=\"112\" stroke=\"var(--sp-ok)\" stroke-width=\"1.2\">\n          <line x1=\"211.99\" y1=\"250\" x2=\"382.06\" y2=\"250\"\/>\n          <line x1=\"211.99\" y1=\"244\" x2=\"211.99\" y2=\"256\"\/>\n          <line x1=\"382.06\" y1=\"244\" x2=\"382.06\" y2=\"256\"\/>\n        <\/g>\n      <\/svg>\n      <\/p>\n<\/p><\/div>\n<p>\n      <button id=\"sptc-play\" type=\"button\"> play<\/button><br \/>\n      <input id=\"sptc-slider\" type=\"range\" min=\"0\" max=\"115\" step=\"0.25\" value=\"0\" aria-label=\"scrub time through the turn\"\/><br \/>\n      <span class=\"readout\" id=\"sptc-readout\">Time t<\/span>\n    <\/p>\n<\/p><\/div>\n<p>  <!-- \n \n<figcaption><b>Scrub the turn.<\/b> Left: the model's raw stream &mdash; thinking tokens, then the\n  program; the cell holds only model-generated tokens. Highlights are live, not history:\n  <b style=\"color:#46c364\">green<\/b> while the fork runs a line,\n  <b style=\"color:#c39bff\">purple<\/b> while its speculated calls are in flight, and\n  <b style=\"color:#d29922\">amber<\/b> marks the line the <i>without-speculation<\/i> twin is stuck\n  on &mdash; it crawls through the same program long after the speculated run has answered. Right: the\n  same moments as time bars &mdash; six sub-calls and the judge (whose prompt needs their results) all\n  run inside decode, so execute collapses to seven cache hits; without speculation the same\n  program runs its seven calls one at a time, after decode.<\/figcaption>\n \n --><br \/>\n<\/figure>\n<p>The associated code is here: <a href=\"https:\/\/github.com\/alexzhang13\/spec-ptc\" rel=\"external nofollow noopener\" target=\"_blank\">https:\/\/github.com\/alexzhang13\/spec-ptc<\/a><\/p>\n<p>It is no surprise from my <a href=\"https:\/\/alexzhang13.github.io\/blog\/2026\/harness\/\">previous writing<\/a> and work on <a href=\"https:\/\/arxiv.org\/abs\/2512.24601\" rel=\"external nofollow noopener\" target=\"_blank\">Recursive Language Models (RLMs)<\/a> that I believe (1) code in a REPL is the only \u201ctool\u201d a system needs; (2) all other tools should be functions in that code tool. When code becomes the primary action space of your system, you need to start thinking about overlapping these specialized tool calls with the code being generated and executed.<\/p>\n<p>Inspired by <a href=\"https:\/\/en.wikipedia.org\/wiki\/Speculative_execution\" rel=\"external nofollow noopener\" target=\"_blank\">speculative execution<\/a> in CPUs and <a href=\"https:\/\/arxiv.org\/abs\/2211.17192\" rel=\"external nofollow noopener\" target=\"_blank\">speculative decoding<\/a> in LLMs, a relatively simple and useful trick I\u2019m proposing in harness design is <strong>speculative programmatic tool calling<\/strong> (<strong>sPTC<\/strong>), in which we speculate and pre-launch tool calls from partially generated REPL calls as the harness is still generating tokens, rather than waiting for it to finish the entire generation. In particular, this is mostly useful when the tool of interest is a sub-LLM or sub-agent call, such as in the RLM. If the fully generated REPL actually ends up calling these tools, they immediately return with their cached outputs from the speculated call.<\/p>\n<p>A key observation is that LLM tools, such as sub-agents or search APIs, are often high-latency and are usually the bottleneck in harnesses that rely on code execution as actions. Furthermore, the actual generation of the main context is often also a significant latency bottleneck that blocks intermediate calls from happening.<\/p>\n<p>For the rest of this blog, I\u2019ll mostly be discussing inference time savings with respect to an RLM, but note that this generally applies to harnesses that generate code (i.e. CodeAct-like or \u201ccode mode\u201d harnesses). There are two obvious areas where you save time:<\/p>\n<ol>\n<li>\n<strong>Overlapping already-generated tool calls during token streaming.<\/strong> Most harness designs will wait for the entire model generation to complete its generation before executing the tools. This design is likely a consequence of JSON-style tool-calling, where this was not really a bottleneck. Because the generation of the main context per turn is often slow, this represents a significant portion of time that can be cut, especially for models that <em>think<\/em> for a long time.<\/li>\n<li>\n<strong>Acting as a JIT compiler over REPL calls.<\/strong> Even without streaming enabled, an obvious optimization is that many REPL programs contain blocking tool calls that aren\u2019t actually blocking \u2014 e.g. two independent sub-agent calls that the code does not write as asynchronous, can still be run in parallel. <strong>sPTC<\/strong> acts as a <em>really naive<\/em> JIT compiler to prevent such cases, and likely can be improved further across languages and REPL designs.<\/li>\n<\/ol>\n<div class=\"sptc-video\">\n<p>\n    <iframe title=\"Side-by-side speculation examples\" src=\"https:\/\/alexzhang13.github.io\/assets\/video\/player.html?src=comparison.mp4\" allow=\"autoplay\"><\/iframe>\n  <\/p>\n<p class=\"caption\">Side-by-side example of sub-agents with string literal inputs, sub-agents with variable dependencies, and loops over sub-agents with variable dependencies being speculated and overlapped with the main context generation.<\/p>\n<\/div>\n<p>On locally running LLMs running one or a few chat instances at a time, your inference engine is often highly memory-bound from decoding just the main context, and speculation can help increase the arithmetic intensity. For high-volume serving systems (e.g. if you\u2019re using a frontier lab or model router API), batched requests are abstracted away to various, likely disjoint serving engines, so the gain is purely from either overlapping computation with the main context, or overlapping execution time with a slow REPL call.<\/p>\n<h2 id=\"some-comparison-numbers-on-rlms\">Some Comparison Numbers on RLMs<\/h2>\n<p>I\u2019ll start by highlighting some basic runtime numbers, although you can probably reason through the runtime benefit. The actual implementation details are quite simple and easy to follow in the codebase, which I will talk about in the latter half of the blog.<\/p>\n<p>You should keep in mind that it\u2019s very difficult to estimate the exact speed-ups because it\u2019s highly dependent on the latency of the tools, the number of tokens generated, the load of your serving engine, and the actual choices the harness makes. We can isolate some components (e.g. fix a program) to highlight obvious cases that don\u2019t really need benchmarking to understand, but the simplest example is just to replicate a setting and harness we know and benchmark the runtimes averaged over many samples.<\/p>\n<p>We run below on <a href=\"https:\/\/arxiv.org\/abs\/2511.02817\" rel=\"external nofollow noopener\" target=\"_blank\">OOLONG (trec-coarse, 132k)<\/a> and <a href=\"https:\/\/huggingface.co\/datasets\/mit-oasys\/oolong-pairs\" rel=\"external nofollow noopener\" target=\"_blank\">OOLONG-Pairs (32k)<\/a> from the original paper, with a node of 8xH100 80B running a vLLM server. We tested two cases with <code class=\"language-plaintext highlighter-rouge\">Qwen3-30B-A3B-Instruct-0527<\/code>, one with the standard temperature=0.7 and one with temperature=0.0 to control the variance of these RLM runs, whose speed is highly dependent on the trajectory it takes. We run each experiment 5 times. We report the average number of sub-calls and turns to give a relative sense for what each model is doing as well, and report across a setting where we have 4 concurrent runs, and 8 concurrent runs over the same serving engine.<\/p>\n<figure>\n<center><br \/>\n    <br \/>\n<\/center><br \/>\n<\/figure>\n<p>The speed-ups for the RLM are generally on the order of 1-1.2x. We also tested a more \u201cdeterministic\u201d suite of LM programs where we\u2019d observe hypothetical or realistic runtime speed-ups in the codebase, but we omit them here because they are too specific.<\/p>\n<h2 id=\"designing-speculative-ptc-methods\">Designing Speculative PTC Methods<\/h2>\n<p>The <a href=\"https:\/\/github.com\/alexzhang13\/spec-ptc\/\" rel=\"external nofollow noopener\" target=\"_blank\">code<\/a> is quite short and readable, but I\u2019ll talk a bit about the design philosophy below.<\/p>\n<p>The high-level design is to have imported tool calls in our REPL have a hook that can be called earlier and replaced with a cached output when needed. We want a library and frontend contract that allows us to define (1) what tools should be speculated and what shouldn\u2019t (e.g. maybe we want sub-LLM calls to be speculated, but a sub-RLM call is too costly); and (2) a mechanism for speculating on tool calls whose inputs rely on variables in memory that were computed previously, rather than just literals.<\/p>\n<p>At a high-level, the type of contract we want is a hook around functions we want speculated (e.g. the sub-call <code class=\"language-plaintext highlighter-rouge\">llm_query()<\/code> in the RLM) so that we can run it asynchronously when we parse it, save it to some global store of tool outputs, and treat it like a promise that draws from this store when it is actually invoked in the REPL. We also want to distinguish between multiple invocations of identical tool calls for when they are non-deterministic (like a sub-call).<\/p>\n<div class=\"sptc-py\">\n<pre>\n@spec.tool(speculatable=True, pure=True)\ndef tool(...) -&gt; OutputType:\n    ...\n\n# Speculated version\ndef tool_spec(...):\n    promise = launch(tool(...))\n    register_speculation(promise)\n\n# Hooked version run in real REPL\ndef tool_real(...):\n    if exists(promise, ID(...)):\n        return promise\n    else:\n        return tool(...)\n<\/pre>\n<\/div>\n<p>The contract above lets us define a \u201cshadowed\u201d namespace which invokes these modified tool calls while the LLM output is being parsed, and save them as futures in some store that can be invoked by the real tool. This gives us the relatively simple logic below:<\/p>\n<div class=\"sptc-py\">\n<pre>\nreal_ns = {**locals, real_tools}\nshadow_ns = replace_tools(real_ns)\n\n# speculate while LLM is streaming\nwhile not LLM.done:\n    code += LLM.next_tokens()\n    parse_and_peek(code, shadow_ns)  # queue without code execution\n    parse_and_speculate(code, shadow_ns)  # re-run shadow REPL\n\n# real tools now route to promised tools\nexec(code, real_ns)\n<\/pre>\n<\/div>\n<p><strong>Shadowed execution for speculation.<\/strong> The easy cases for speculation are when you can parse a tool call and infer the inputs from the tokens directly, i.e. when the inputs are literals. The nastier cases come when tool calls are embedded inside conditional or looping logic, or the inputs are dependent on some variable in memory that was computed earlier in the REPL call. In the former case, we may not know a-priori whether the conditional is satisfied, or how long the loop runs for. In the latter case, we\u2019re unsure if computing the variable input is not <em>pure<\/em> and modifies some external state, meaning we cannot compute this beforehand.<\/p>\n<p>There are many potential ways to solve this issue which could span an entire line of work, but to keep it simple, I chose to maintain a deepcopy fork of our primary code REPL, which we label as a <em>shadow REPL<\/em>, which executes the partial REPL on the fly. To prevent unwanted side-effects, most external libraries and functions like <code class=\"language-plaintext highlighter-rouge\">open<\/code> get marked as \u201cunsafe\u201d, and any speculatable tools that have input dependencies using those functions are not speculated.<\/p>\n<p><strong>Note<\/strong>: I intentionally chose not to use the partial \u201cspeculator\u201d executor as the real REPL executor, because there\u2019s a chance the entire REPL the model produces error-prone code or an incomplete tool call. In these cases, we do not want the speculator to actually modify the REPL state, and we want to treat the entire REPL cell as a unit of computation.<\/p>\n<h2 id=\"what-you-can-speculate-and-run-ahead-of-time\">What you can speculate and \u201crun\u201d ahead-of-time.<\/h2>\n<p>There\u2019s a bit of consideration into what exactly can be speculated, how aggressive you want to be, and the overall safety of running partial code. We generally care about understanding (1) do we have enough information to run the tool call ahead of time; (2) can we figure out if that tool call will actually execute, especially around conditionals.<\/p>\n<p>In PTC, we can uniquely index into a tool call by its <strong>inputs<\/strong> (and occurrence if non-deterministic). During streaming, if the inputs to each tool call are known without code execution (i.e. all inputs are literals like a string literal), we can immediately invoke this tool call asynchronously. The more common case is where the inputs have dependencies on other variables, some of which we can <em>safely<\/em> compute, and others that we cannot. For identical tool calls, such as computing a majority vote over several sub-agents, we do not want a single speculated tool call to route to each copy, so we need to track unique instances of the same tool call, unless we know the call is deterministic.<\/p>\n<p>We can look a few cases below of what currently gets <em>speculated<\/em> and what doesn\u2019t to provide some intuition:<\/p>\n<p><strong>Case 1: Literals.<\/strong> String, integer, or other literals can immediately be parsed and converted into a tool call even without shadowed execution of each line.<\/p>\n<div class=\"sptc-py\">\n<pre>\ntitle = llm_query(\"Give a title for: The Odyssey\")  # parses\nblurb = llm_query(\"One-line blurb for: The Odyssey\")  # parses\nprint(title, blurb)\n<\/pre>\n<\/div>\n<p><strong>Case 2: Input dependencies.<\/strong> When input dependencies are involved, as long as all inputs are <em>safe<\/em> (i.e. pure functions, no side effects), they can be speculated. Furthermore, tools that rely on dependencies that are speculated will wait on the dependencies to be computed first, and are then executed even if the LLM is still streaming.<\/p>\n<div class=\"sptc-py\">\n<pre>\na = llm_query(\"Triage: \" + doc)  # parses then executes\nc = llm_query(\"Summarize: \" + str(a))  # parses, waits on a, then executes\nprint(c)\n\nif len(doc) &gt; 10_000:  # will evaluate and speculate if safe\n    extra = llm_query(\"Also outline it: \" + doc)\n<\/pre>\n<\/div>\n<p><strong>Case 3: Peekable and non-peekable dependencies.<\/strong> During streaming at any step, we have a working namespace of the shadow REPL. When a complete tool is parsed, there are cases when we can speculate the inputs, even if the dependencies are variables in memory and not literals.<\/p>\n<div class=\"sptc-py\">\n<pre>\ndef gist(t):\n    return llm_query(\"One-line gist: \" + t)  # non-peekable\n\nparts = [gist(c) for c in chunks]\n\nside = llm_query(\"Give me a random title for:\", chunks[0])  # peekable\nprint(side, parts)\n<\/pre>\n<\/div>\n<p><strong>Case 4: Blocked speculation calls.<\/strong> While trying to speculate, there is an allowlist of keywords and function calls that can be used to compute dependencies for inputs. We also can specify new tools that we do want to specify are not pure functions. Any speculatable tools that have dependencies that are blocked will not be speculated.<\/p>\n<div class=\"sptc-py\">\n<pre>\na = llm_query(\"Triage: \" + doc)  # speculated\nnotes = open(\"\/tmp\/scratch.txt\").read()  # blocked\nb = llm_query(\"Annotate with notes: \" + notes)  # blocked\nc = llm_query(\"Summarize: \" + a)  # speculated after a\n<\/pre>\n<\/div>\n<p>There are many more cases, and ideally we\u2019d write a whole pseudo-compiler for this kind of thing. I suspect also various trade-offs will pop up for different harness designs, but so far this has not been heavily optimized.<\/p>\n<h3 id=\"speculation-overhead-for-ptc\">Speculation Overhead for PTC<\/h3>\n<p>Like with other parts of this trick, the extra overhead of speculative PTC is somewhat dependent on both the exact implementation and the setup of your inference engine. For this implementation, on runtime, the overhead is somewhat negligible because the speculator cheaply parses and checks whether it believes it can speculate over the partially generated REPL. On memory, we create a deepcopy of the harness code REPL, which generally is cheap relative to the allocated memory of the actual REPL variables because we tend to have few large mutable objects.<\/p>\n<p>The worst case happens when the serving engine for the tool is clogged with many concurrent and potentially extra speculated requests, which can be controlled with how aggressive speculation is and how these tools are queued.<\/p>\n<p>Speculative algorithms at the level of LLM decoding (i.e. the streamed outputs) have been explored a little bit for the older tool-calling designs, although not extensively. The likely case is that for standard tool-calling, these techniques were not that useful or the overhead was not worth the marginal latency improvements.<\/p>\n<p><a href=\"https:\/\/arxiv.org\/abs\/2406.00059\" rel=\"external nofollow noopener\" target=\"_blank\">Conveyor (Xu et al., 2024)<\/a><d-cite key=\"xu2024conveyorefficienttoolawarellm\"\/> allows users to define partial execution opportunities such as a line of code, which are parsed during decoding. In <a href=\"https:\/\/arxiv.org\/abs\/2605.13360\" rel=\"external nofollow noopener\" target=\"_blank\">Speculative Interaction Agents (Hooper et al., 2026)<\/a><d-cite key=\"hooper2026speculativeinteractionagentsbuilding\"\/>, they formally define the system proposed in Conveyor as <em>speculative tool calling<\/em>, which mainly reduces time-to-first-token (TTFT) by overlapping long thinking chains of more modern models with an invoked tool call. <a href=\"https:\/\/arxiv.org\/abs\/2605.15077\" rel=\"external nofollow noopener\" target=\"_blank\">AsyncFC (Feng et al., 2026)<\/a><d-cite key=\"feng2026concurrencymodelchangesfuturebased\"\/> instead argues that tool calls are often implemented in a blocking way, and define a contract for future-based async wrappers around function calls; this however, has the risk of not being 1:1 with the original harness trajectory.<\/p>\n<p>In the case of sPTC, the reason I\u2019d argue that adding speculation to tools in a more complex program is significantly more useful is because of the unknown runtime of the program itself. For standard tool calling, by the time the LLM has finished generating enough tokens to fully specify the tool call, there likely is not many more tokens left to generate. There are a lot more considerations in the PTC case over what you actually speculate and how, because code execution makes the actual tool call patterns significantly more complicated, leaving more room for overlap.<\/p>\n<h2 id=\"ending-note\">Ending Note<\/h2>\n<p>Speculative programmatic tool calling is a natural trick that arises from programmatic tool calling itself having more potential for overlap than traditional tool calling itself. I want to clarify that beyond overlapping with streamed token generations of the root LLM in a harness, the real value in the long run will come from more clever JIT compilation tricks to overlap tool calls in the PTC setting with the actual REPL execution itself, which may become more expensive as harnesses generate more complex programs.<\/p>\n<p>There are likely many ways to implement speculative PTC to be much faster or more aggressive, while also utilizing less overhead. We\u2019d ideally also want it to be language and harness-agnostic, which for now is mainly just {Python, bash, Bun} x {Coding harness, RLM, game agent}. The current implementation is already quite useful though, especially for locally running models and agent harnesses.<\/p>\n<p>I\u2019ve provided a simple implementation that slots directly into my RLM implementation, but it should be quite easy to build on top of and add as a plugin to your own coding agents.<\/p>\n<p><strong>Acknowledgements.<\/strong> I thank Laude for generously providing compute to run these experiments. I thank my advisor Omar Khattab for proofreading the idea.<\/p>\n<p>If you need to cite this:<\/p>\n<div class=\"sptc-py\" data-lang=\"bibtex\">\n<pre>\n@article{zhang2026sptc,\n  title   = \"Speculative Programmatic Tool Calling\",\n  author  = \"Zhang, Alex\",\n  year    = \"2026\",\n  month   = \"August\",\n  url     = \"https:\/\/alexzhang13.github.io\/blog\/2026\/spec-ptc\/\"\n}\n<\/pre>\n<\/div><\/div>\n<p><a href=\"https:\/\/alexzhang13.github.io\/blog\/2026\/spec-ptc\/?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Fig 1. Speculative Programmatic Tool Calling (sPTC) Example the model&#8217;s turn \u2014 streaming WITH SPECULATION DECODE EXECUTE prep sub-calls \u00d76 judge execute baseline answer \u00b7 7\/7 claims WITHOUT SPECULATION \u00b7 SAME PROGRAM answer c\u2081c\u2082 c\u2083c\u2084 c\u2085c\u2086 judge execute collapses \u2014 2.4\u00d7, no async play Time t The associated code is here: https:\/\/github.com\/alexzhang13\/spec-ptc It is no [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23494,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23493","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23493","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23493"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23493\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23494"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23493"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23493"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23493"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}