{"id":23936,"date":"2026-09-14T08:23:23","date_gmt":"2026-09-14T08:23:23","guid":{"rendered":"https:\/\/scannn.com\/the-pulse-tech-companies-move-to-open-ai-models\/"},"modified":"2026-09-14T08:23:23","modified_gmt":"2026-09-14T08:23:23","slug":"the-pulse-tech-companies-move-to-open-ai-models","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/the-pulse-tech-companies-move-to-open-ai-models\/","title":{"rendered":"The Pulse: tech companies move to open AI models"},"content":{"rendered":"\n<div id=\"content\">\n          <!--\n          \n\n<p><em>Before we start: I'm hosting the first-ever <strong><a href=\"https:\/\/www.pragmaticsummit.com\/\">The Pragmatic Summit<\/a><\/strong> on 11 February, 2026, in San Francisco. Join 400 top engineers and leaders as we answer the question: How is AI reshaping software engineering, dev workflows, and the modern engineering stack? \n          Spaces are limited - don't miss out! <strong><a href=\"https:\/\/www.pragmaticsummit.com\/\">Buy tickets here<\/a><\/strong>.\n\n          <\/em><\/p>\n\n\n          --><\/p>\n<p><em>Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of five topics from <\/em><a href=\"https:\/\/pragmaticengineer.substack.com\/p\/the-pulse-tech-companies-move-to\" rel=\"noopener noreferrer nofollow\"><em>last week\u2019s The Pulse<\/em><\/a><em> issue. Full subscribers received the article below seven days ago. If you\u2019ve been forwarded this email, you can <\/em><a href=\"https:\/\/newsletter.pragmaticengineer.com\/about?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\"><em>subscribe here<\/em><\/a><em>.<\/em><\/p>\n<p>In May, I <a href=\"https:\/\/newsletter.pragmaticengineer.com\/p\/the-pulse-a-trend-of-trying-to-cut?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\">covered an emerging trend<\/a> of companies <em>wanting<\/em> to cut back their AI spending, starting with engineering departments. Different approaches were being tried:<\/p>\n<ul>\n<li>Experimentation with running cheaper, open models on inference providers<\/li>\n<li>More investment in model routing to route simpler requests to cheaper models<\/li>\n<li>Knowledge-sharing sessions on how to use AI models cost-effectively<\/li>\n<li>Setting per-developer monthly AI usage limits<\/li>\n<\/ul>\n<p>A few months later, it seems that several companies have managed to achieve this, according to sources I\u2019ve spoken with. <\/p>\n<h3 id=\"uber-ai-costs-down-50\"><strong>Uber: AI costs down 50%<\/strong><\/h3>\n<p>Uber <a href=\"https:\/\/newsletter.pragmaticengineer.com\/i\/194426825\/2-are-coding-ai-agent-subsidies-doomed?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\">managed to blow through<\/a> its annual AI budget in the first three months of this year, and it wasn\u2019t a surprise to hear, in May, Uber\u2019s COO <a href=\"https:\/\/fortune.com\/2026\/05\/26\/uber-coo-ai-spending-tokens-claude-code\/?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\">say<\/a> that it was getting harder to justify spending on tools like Claude Code without seeing benefits from the leading models. It wasn\u2019t long until engineering teams at the ridesharing giant set to work on how to optimize AI spend, and their efforts weren\u2019t in vain.<\/p>\n<p>Uber cut the cost per AI request by 34%, and the cost per AI session by 52%:<\/p>\n<figure class=\"kg-card kg-image-card kg-card-hascaption\"><figcaption><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Reducing per-token and per-session spend. Source: <\/em><\/i><a href=\"https:\/\/www.uber.com\/us\/en\/blog\/efficient-software-factory\/?ref=blog.pragmaticengineer.com\" target=\"_blank\" rel=\"noopener noreferrer nofollow\"><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Uber<\/em><\/i><\/a><\/figcaption><\/figure>\n<p>Of course, Uber keeps using more AI tokens and starting more AI sessions, but thanks to optimizations the cost has been flat since March, despite significantly more usage:<\/p>\n<figure class=\"kg-card kg-image-card kg-card-hascaption\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!l-d4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F534b20ec-8678-40d4-9248-2d5751979703_2048x1145.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"1456\" height=\"814\"\/><figcaption><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Usage up, cost stable Source: <\/em><\/i><a href=\"https:\/\/www.uber.com\/us\/en\/blog\/efficient-software-factory\/?ref=blog.pragmaticengineer.com\" target=\"_blank\" rel=\"noopener noreferrer nofollow\"><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Uber<\/em><\/i><\/a><\/figcaption><\/figure>\n<p>How did they pull it off at Uber? It was via a combination of different approaches:<\/p>\n<ul>\n<li><strong>Open weight models, run using inference: <\/strong>run open weight models on inference services, which are significantly cheaper than frontier ones.<\/li>\n<li><strong>Optimized model selection<\/strong>: benchmark all available frontier and open models, to build an accurate picture of their present capabilities<\/li>\n<li><strong>Ongoing benchmarking<\/strong>: run benchmarks every week based on real work, and update them<\/li>\n<li><strong>Cheaper subagent models<\/strong>: subagents do smaller tasks not requiring the most expensive models<\/li>\n<li><strong>Reduce model effort:<\/strong> Uber found that defaulting to Medium effort gives the best cost-to-output ratio with advanced models<\/li>\n<li><strong>Optimize requests<\/strong>: trigger automatic compaction above 400K tokens, even for models with 1M context windows<\/li>\n<li><strong>Cache prompts<\/strong>: cache prompts to save money when using Uber\u2019s own harness, Minions<\/li>\n<li><strong>\u2026 and more: <\/strong>Uber wrote <a href=\"https:\/\/www.uber.com\/us\/en\/blog\/efficient-software-factory\/?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\">an engineering blog post<\/a> detailing the dozens of optimizations taken to reduce token cost without noticeable change in the quality of code generated by agents<\/li>\n<\/ul>\n<p>From the outside, the single biggest win seems to be Uber\u2019s transition to using open models for certain tasks. Open models cost 2-20x less, compared to frontier ones:<\/p>\n<figure class=\"kg-card kg-image-card kg-card-hascaption\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!30Z3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e389210-803e-483c-832a-aed043429601_1572x702.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"1456\" height=\"650\"\/><figcaption><i><em class=\"italic\" style=\"white-space: pre-wrap;\">The most expensive open weight model costs $0.30 per code review, vs $0.50 for the cheapest frontier model (and $2.50 for the most expensive one). Source: <\/em><\/i><a href=\"https:\/\/www.uber.com\/us\/en\/blog\/efficient-software-factory\/?ref=blog.pragmaticengineer.com\" target=\"_blank\" rel=\"noopener noreferrer nofollow\"><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Uber<\/em><\/i><\/a><\/figcaption><\/figure>\n<h3 id=\"pinterest-makes-90-cost-savings-by-dropping-frontier-models\"><strong>Pinterest makes 90%+ cost savings by dropping frontier models<\/strong><\/h3>\n<p>Interesting details from Pinterest\u2019s earnings call last month reveal how much the social media platform saves by running open models. Here\u2019s what CEO William Ready said (emphasis mine):<\/p>\n<blockquote><p>\u201cOur approach to model deployment includes our own compact fit-for-purpose models built for Pinterest-specific use cases and suitable open source models post-trained in our own environment within our secure cloud infrastructure. When we leverage open source models, such as with Pinterest Assistant, we are seeing superior performance for our use cases when compared to closed third-party models because we are able to post-train open models on our highly unique data.<\/p>\n<p><strong>With open models, we are achieving cost per transaction at less than 8% of the cost of comparable closed proprietary models. <\/strong>This gives us substantial headroom to deepen and extend these capabilities over time in a way that is differentiated, highly effective, and cost efficient.\u201d<\/p>\n<\/blockquote>\n<p>Basically, what used to cost Pinterest $100 to run on a closed, frontier model, they now spend $8 on by using open models on owned or rented inference!<\/p>\n<h3 id=\"att-56-savings-by-swapping-claude-for-open-models\"><strong>AT&amp;T: 56% savings by swapping Claude for open models<\/strong><\/h3>\n<p>With 100,000 employees, AT&amp;T is a big spender on AI. The telco giant cut its AI bill by 56% while measuring a 2% decrease in the quality of AI\u2019s output, after they moved workloads over to open models. From <a href=\"https:\/\/www.theinformation.com\/newsletters\/applied-ai\/t-using-open-source-models-curb-anthropic-bills?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\">The Information<\/a>:<\/p>\n<blockquote><p>\u201cAustin said he\u2019s found that open source models are \u201cjust as good or better\u201d than older models sold by the likes of Anthropic and OpenAI. For instance, AT&amp;T\u2019s software developers still rely on cutting-edge models for complex tasks like generating code, but can use cheaper open source models for less intense tasks like generating summaries of previously submitted code, he said.<\/p>\n<p>After the company began using router provider LiteLLM, the costs of some advanced AI tasks such as coding fell by as much as 56% while the quality of the AI\u2019s performance fell just 2%, Austin said.\u201d<\/p>\n<\/blockquote>\n<h3 id=\"anthropic-overpriced-compared-to-the-rest-of-the-market\"><strong>Anthropic overpriced compared to the rest of the market?<\/strong><\/h3>\n<p>Only a few months ago, Anthropic was the preferred model (Claude) and harness (Claude Code) among engineers. But Anthropic\u2019s models are becoming steeply more expensive at a time when open weight models \u2013 and also OpenAI \u2013 are getting much cheaper. Meanwhile, Opus 5 is 100x more expensive (!!) than models like GPT-5.6 Luna xhigh and DeepSeek. That may be simply too much to ignore for some tech companies:<\/p>\n<figure class=\"kg-card kg-image-card kg-card-hascaption\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!9_-_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2cd4931f-ab6d-4628-8798-27341f35eccb_1120x852.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"1120\" height=\"852\"\/><figcaption><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Typical cost for an AI agent run, per model. Source: <\/em><\/i><a href=\"https:\/\/modelzengarden.com\/?ref=blog.pragmaticengineer.com\" target=\"_blank\" rel=\"noopener noreferrer nofollow\"><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Model Zen Garden<\/em><\/i><\/a><\/figcaption><\/figure>\n<p>Seeing this data, I\u2019m not surprised that more tech companies are looking to run open weight providers on inference providers, due to the significant savings available from a model that\u2019s similarly capable as one from Anthropic.<\/p>\n<h3 id=\"what-worked-for-stripe-coinbase-uber-ramp\"><strong>What worked for Stripe, Coinbase, Uber &amp; Ramp<\/strong><\/h3>\n<p>The engineering team at Databricks interviewed engineers at Stripe, Coinbase, Uber, and Ramp, and <a href=\"https:\/\/www.databricks.com\/blog\/managing-ai-coding-costs-scale?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\">collected<\/a> how different approaches helped save costs for them. The summary:<\/p>\n<figure class=\"kg-card kg-image-card kg-card-hascaption\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!pejT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5d9da58-d8d7-4eaf-b668-a16fe4d613c5_1646x652.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"1456\" height=\"577\"\/><figcaption><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Source: <\/em><\/i><a href=\"https:\/\/www.databricks.com\/blog\/managing-ai-coding-costs-scale?ref=blog.pragmaticengineer.com\" target=\"_blank\" rel=\"noopener noreferrer nofollow\"><i><em class=\"italic\" style=\"white-space: pre-wrap;\">Databricks<\/em><\/i><\/a><\/figcaption><\/figure>\n<p>To answer the question posed in the header of this report, it\u2019s apparent that using open models is indeed the approach offering the biggest savings, followed by smart model routing. Spending controls and context optimization also bear down on costs, but they don\u2019t come close to the first two techniques in results.<\/p>\n<p><em>A week after publishing this article, Ara Krahzian at Ramp <\/em><a href=\"https:\/\/x.com\/arakharazian\/status\/2097706961584140645?s=20&amp;ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\"><em>has confirmed<\/em><\/a><em> that AI spend in August, has, indeed, declined at the top 1% of businesses by 10%, based on Ramp data:<\/em><\/p>\n<figure class=\"kg-card kg-image-card kg-card-hascaption\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/storage.ghost.io\/c\/39\/f8\/39f85cc7-8637-40fc-a57c-f45754453717\/content\/images\/2026\/09\/image.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"1200\" height=\"725\" srcset=\"https:\/\/storage.ghost.io\/c\/39\/f8\/39f85cc7-8637-40fc-a57c-f45754453717\/content\/images\/size\/w600\/2026\/09\/image.png 600w, https:\/\/storage.ghost.io\/c\/39\/f8\/39f85cc7-8637-40fc-a57c-f45754453717\/content\/images\/size\/w1000\/2026\/09\/image.png 1000w, https:\/\/storage.ghost.io\/c\/39\/f8\/39f85cc7-8637-40fc-a57c-f45754453717\/content\/images\/2026\/09\/image.png 1200w\" sizes=\"auto, (min-width: 720px) 720px\"\/><figcaption><span style=\"white-space: pre-wrap;\">AI spend starting to decline at the top 1% of firms. Source: <\/span><a href=\"https:\/\/ramp.com\/data\/ai-index-sept-2026?ref=blog.pragmaticengineer.com\" rel=\"noreferrer\"><span style=\"white-space: pre-wrap;\">Ramp<\/span><\/a><\/figcaption><\/figure>\n<p><em>I\u2019d wager those companies are not spending fewer tokens, but they are optimizing cost, in ways outlined above.<\/em><\/p>\n<hr\/>\n<p>Read the full issue of <a href=\"https:\/\/newsletter.pragmaticengineer.com\/p\/the-pulse-tech-companies-move-to?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\"><strong>last week\u2019s The Pulse<\/strong><\/a><strong>,<\/strong> or check out <a href=\"http:\/\/todo\/?ref=blog.pragmaticengineer.com\" rel=\"noopener noreferrer nofollow\"><strong>this week\u2019s The Pulse<\/strong><\/a>. This week\u2019s issue covers:<\/p>\n<ol>\n<li><strong>New trend of CPU shortages:\u00a0<\/strong>after a GPU shortage and memory shortage driven by AI companies, we\u2019re now experienceding a CPU shortage, thanks to AI agents using a lot more CPU with tool usage. If you will need more compute in the future: secure it now, while you can (even if it\u2019s expensive to do so).<\/li>\n<li><strong>Growth dream ends for more COVID-era unicorns:\u00a0<\/strong>Miro sold itself to Bending Spoons for $1.35B, after it was valued at $17B in 2022. Airtable saw a similar valuation cut last month, and it seems a batch of now-overvalued, VC-funded companies are desperate to sell.<\/li>\n<li><strong>Industry Pulse:<\/strong>\u00a0Overtime at Google to get Borg working on SpaceX\u2019s data centers; SpaceX cuts Claude Code tokens by 90%; OpenAI launches Astra; Meta unveils Muse (and Mark Zuckerberg pushed production code in this release); \u2013 to which Mark Zuckerberg made personal contributions; OpenAI\u2019s agents go rogue, again.<\/li>\n<li><strong>Do engineers lose touch when AI handles incidents?\u00a0<\/strong>In the aviation industry, pilots are exposed to emergency situations every six months, to keep their critical problem solving skills sharp. In the tech industry, we might need something similar, especially if AI would take on handling of the simpler incidents.<\/li>\n<\/ol>\n<p>            <!-- Newsletter --><\/p>\n<p>\n              <a href=\"https:\/\/newsletter.pragmaticengineer.com\/about\">Subscribe to my weekly newsletter<\/a> to get articles like this in your inbox. It&#8217;s a pretty good read &#8211; and the <a href=\"https:\/\/substack.com\/top\/technology\">#1 software engineering newsletter<\/a> on Substack.\n            <\/p>\n<\/p><\/div>\n<p><a href=\"https:\/\/blog.pragmaticengineer.com\/the-pulse-tech-companies-move-to-open-ai-models\/?utm_source=tldrnewsletter\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of five topics from last week\u2019s The Pulse issue. Full subscribers received the article below seven days ago. [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23937,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23936","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23936","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23936"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23936\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23937"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23936"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23936"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23936"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}