{"id":23557,"date":"2026-08-27T22:04:02","date_gmt":"2026-08-27T22:04:02","guid":{"rendered":"https:\/\/scannn.com\/the-plumbing-behind-metas-ai\/"},"modified":"2026-08-27T22:04:02","modified_gmt":"2026-08-27T22:04:02","slug":"the-plumbing-behind-metas-ai","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/the-plumbing-behind-metas-ai\/","title":{"rendered":"The Plumbing Behind Meta's AI"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p><span style=\"font-weight: 400;\">One of the consequences of the AI boom is that cooling servers has become a serious engineering challenge.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The more powerful the AI hardware becomes, the more heat it generates during operation. And at a certain point, simply blowing more air through a server rack stops being a particularly efficient method of cooling.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">That\u2019s why when I visited Meta\u2019s AI Infrastructure in Texas, one of the technologies I was most interested in wasn\u2019t actually the AI hardware. It was the plumbing.<\/span><\/p>\n<h2><span style=\"font-weight: 400;\">The Cooling Shift<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">Traditional data centers, that\u2019s data centers used for compute tasks such as searching for your favorite creator on Instagram or liking a post on Facebook, will more often than not use air cooling to keep the hardware at an optimal temperature.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Even a few years ago, using air cooling methods to cool AI hardware was an achievable solution. In fact, I visited a data center in Altoona, Iowa, where racks of 16 Nvidia H100s were kept cool completely through air cooling with minimal water usage. Minimal amounts of water were used at the start of the data center cooling process to cool the air during warmer months, but no water was ever being sent directly to the hardware.<\/span><\/p>\n<p><a href=\"https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg\"><img loading=\"lazy\" data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-50062\" src=\"https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?resize=960%2C960\" alt=\"Photo of Nvidia H100s in in a server rack using air cooling\" width=\"960\" height=\"960\" srcset=\"https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=1920 1920w, https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=150 150w, https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=300 300w, https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=768 768w, https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=1024 1024w, https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=1536 1536w, https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=220 220w, https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=1080 1080w, https:\/\/about.fb.com\/wp-content\/uploads\/2026\/08\/01_H100.jpg?w=600 600w\" sizes=\"auto, (max-width: 960px) 100vw, 960px\"\/><\/a><\/p>\n<p><span style=\"font-weight: 400;\">It\u2019s only more recently that newer AI hardware designs have created the demand for a newer, more optimal method of cooling.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Enter closed-loop liquid cooling.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">There\u2019s a common misconception that AI data centers are automatically big water users. The reality depends on the cooling design \u2014 Meta\u2019s data centers use a closed-loop system that recirculates water in a sealed loop, using very little on an ongoing basis.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The majority of Meta\u2019s newest AI-optimized data centers use closed-looped, liquid cooling as it is the most efficient way to cool GPU servers \u2014 both from a resources point of view, but also from an infrastructure point of view.<\/span><\/p>\n<h2><span style=\"font-weight: 400;\">What Is Closed-Loop Cooling?<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">The basic idea behind closed-loop liquid cooling is actually pretty simple.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">A liquid coolant (a mix of water and glycol) is passed through the server hardware to move heat away from the server racks. But instead of that liquid being expelled from the facility, it is pumped through a series of heat exchangers, which are used to dissipate and transfer the heat away from the liquid. Once the liquid has cooled down, it is sent back around to the server racks in a continuous looping process.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">So the same water and glycol mixture is being used over and over again to keep these chips cool. In fact, Meta expects to use these coolants for up to a decade without needing to replace them.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The heat transfer methods can vary depending on the location and environment the data center is in. When Meta needs to place liquid-cooled equipment into facilities that do not have the liquid cooling infrastructure built into the buildings, they utilize a system called Air-Assisted Liquid Cooling. This consists of racks that contain pumps and heat exchangers that essentially act similar to the large building system described above, with the same closed-loop cooling just on a smaller, more distributed scale.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Why is closed-loop liquid cooling the most ideal method? Because it\u2019s resource efficient. In fact, a typical AI-optimised data center using a closed-loop liquid cooling system with dry coolers uses less water annually than a couple of full-service restaurants. When you compare water usage to real use cases, as opposed to numbers, suddenly the low usage is actually really impressive.<\/span><\/p>\n<h2><span style=\"font-weight: 400;\">Optimizing for Efficiency<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">This innovative liquid cooling system isn\u2019t just about saving water, it\u2019s also a much more efficient use of the space inside of the racks and data centers. If you were to attempt to cool these same servers with air, you\u2019d likely need nearly double the size of the server tray in order to add in the required air cooling equipment. That means you have a much bigger tray, but still the same compute capacity, and eventually, you\u2019d reach diminishing returns with larger and larger air-cooled solutions.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">With direct-to-chip closed-loop liquid cooling, the engineers can fit many more GPUs in the same sized server rack, resulting in fewer racks required. So a facility of the same size is now able to scale its capacity without needing to take up more space.<\/span><\/p>\n<h2><span style=\"font-weight: 400;\">Meta\u2019s Open-Source Liquid Cooling Infrastructure<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">If you\u2019ve seen any of my video content on Meta\u2019s infrastructure, you\u2019ll know that they design and develop their own systems across their entire infrastructure stack, all the way from designing their own chips to the cooling systems and power infrastructure.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">So what do they do with these designs once they\u2019ve deployed them?<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Consistent with Meta\u2019s Open Compute Project legacy, these advances are being shared with the industry. The <\/span><a href=\"https:\/\/about.fb.com\/news\/2025\/10\/open-hardware-future-data-center-infrastructure\/\"><span style=\"font-weight: 400;\">Open Compute Project<\/span><\/a><span style=\"font-weight: 400;\"> (founded 2011) is an open-source hardware and software initiative that aims to make data center infrastructure more efficient, scalable, and sustainable. In 2025, Meta announced IcePack, a liquid-cooled network rack platform that\u2019s being shared openly for free via the <\/span><a href=\"https:\/\/www.opencompute.org\/\"><span style=\"font-weight: 400;\">Open Compute Project<\/span><\/a><span style=\"font-weight: 400;\">.<\/span><\/p>\n<h2><span style=\"font-weight: 400;\">Using AI to Optimize Data Center Cooling<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">The Meta Engineering teams are doing a great job of finding the optimal cooling methods for their data centers. But there\u2019s another interesting part to this story: they\u2019re also using reinforcement learning to help optimize their cooling infrastructure.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">As we found out in this short article, cooling a data center isn\u2019t as simple as choosing a temperature and leaving the system running. Conditions change. The environment changes depending on the location of the data center. The amount of work the servers are doing changes, and so does the amount of cooling. Therefore, the cooling infrastructure has to be purpose built, flexible, and deployed to address these specific considerations across every location.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">So Meta\u2019s engineering team has been experimenting with reinforcement learning to help inform the design and operations of their cooling systems. This reinforcement learning-based approach has since been scaled to the air-cooled data centers in Meta\u2019s fleet.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Rather than experimenting directly on a live data center, where getting that decision wrong could potentially cause issues with operations, Meta\u2019s engineers built a physics-based simulator of a data center environment.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The simulator can model variables such as weather conditions, server load, and the behavior of the cooling equipment. This gives the reinforcement learning model a safe environment in which to test different decisions to learn how to reduce the amount of cooling that is required while still keeping the servers within their optimal operating conditions.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">And just to be clear, while this started as an experiment, it isn\u2019t one anymore.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">In a pilot at one of Meta\u2019s data centers, this reinforcement learning-based approach reduced the amount of energy consumed by the air cooling supply fans by an average of 20% while also reducing water usage by 4% across different weather conditions.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Those aren\u2019t insignificant numbers, and when you apply those reductions across an entire data center fleet, that\u2019s a really impressive efficiency gain that doesn\u2019t go unnoticed.<\/span><\/p>\n<\/p><\/div>\n<p><script async defer crossorigin=\"anonymous\" src=\"https:\/\/connect.facebook.net\/en_US\/sdk.js#xfbml=1&#038;version=v5.0\"><\/script><br \/>\n<br \/><br \/>\n<br \/><a href=\"https:\/\/about.fb.com\/news\/2026\/08\/closed-loop-cooling-explained-the-plumbing-behind-metas-ai\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>One of the consequences of the AI boom is that cooling servers has become a serious engineering challenge. The more powerful the AI hardware becomes, the more heat it generates during operation. And at a certain point, simply blowing more air through a server rack stops being a particularly efficient method of cooling. That\u2019s why [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":23558,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[123],"tags":[],"class_list":["post-23557","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-facebook"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23557","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23557"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23557\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23558"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23557"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23557"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23557"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}