AI infrastructure moves to liquid cooling

AI infrastructure moves to liquid cooling

Key takeaways

  • Rising GPU density is pushing more heat into each rack, making liquid cooling a required part of AI server designs
  • As more components generate significant heat, the cooling architecture needs to be considered alongside the hardware itself.
  • Not all liquid cooling solves the same problem. Direct-to-chip, tank immersion and precision liquid cooling remove heat in very different ways.

Rising GPU density is pushing more heat into each rack, making liquid cooling a a required part of AI server designs. As more components generate significant heat, the cooling architecture needs to be considered alongside the hardware itself. Not all liquid cooling solves the same problem. Direct-to-chip, tank immersion and precision liquid cooling remove heat in very different ways.

What’s changed inside the AI rack

AI infrastructure is getting denser, and that means it’s getting harder to cool with air alone. More powerful GPUs are being packed into the same rack footprints, pushing more power through each server. Almost all of the energy they consume is converted into heat, making thermal management an increasingly important part of AI infrastructure.

At the highest densities, the argument over whether liquid cooling has a role is largely over. What matters now is what happens once liquid reaches the server, because “liquid cooled” can mean very different things: from cooling only the highest-power chips to capturing heat across the whole server.

Those differences determine how much heat the facility still has to remove, which cooling infrastructure needs to remain in place, what operators need to ask of their server vendors, and, increasingly, how the next generation of AI hardware needs to be designed.

What is AI infrastructure, and what are its requirements?

To understand why AI infrastructure cooling is changing, we need to start with what sits in the rack, how much power it draws and where air cooling starts to struggle.

AI infrastructure is the hardware and supporting systems needed to train and run AI workloads. That includes compute, storage, networking, power and the cooling required to keep increasingly dense hardware within its operating limits.

The components themselves are familiar to anyone who’s worked with data center infrastructure, but AI workloads have changed how much performance, and therefore how much power, is being concentrated into each rack.

The bigger issue for AI infrastructure is where that power is being concentrated. As GPU density rises, more computing power is being packed into the same physical rack footprint. A facility can accommodate a growing amount of compute by spreading it across more space, but that isn't what AI infrastructure is being designed to do. The aim in AI infrastructure design is to put more performance into fewer, denser racks.

That changes the thermal load at rack level. A rack drawing significantly more power needs to move significantly more heat away from the hardware in the same space, which turns power density into a cooling constraint. Now, cooling capacity has to be available where the load is concentrated, not simply across the facility as a whole.

Elon Musk recently argued that power and the infrastructure around it could limit how much new AI compute actually comes online. Cooling is part of that problem: the more power spent getting heat out, the less of the site’s available capacity is left to run compute.

Where air runs out of gas

Air cooling works by moving conditioned air across hot components, absorbing their heat and carrying it away. However, as rack density increases, that becomes harder to do efficiently. Fans have to move greater volumes of air through increasingly dense hardware, while the amount of heat that needs to be removed from each rack keeps rising.

At some point, it’s no longer practical to keep adding more airflow. This is where liquid cooling changes the equation. Instead of relying on air to carry heat away from the server, liquid can capture heat at the source and transfer it directly into a cooling loop. But that still leaves an important question: how much of the server does the liquid actually cool?

The market has already decided

At NVIDIA GTC 2025, AI-enabled robots were the most popular demos on the show floor. At GTC 2026, liquid cooling had taken their place and become a defining part of the infrastructure conversation.

Our internal estimates show that roughly 95% of hyperscale data centers being built today are liquid cooled. That doesn't mean every AI deployment has reached the same point, but at the highest densities, liquid cooling is being designed into the infrastructure from the outset.

The technology itself is maturing fast too. At GTC, we saw cold plates from around ten different vendors displayed side by side, with remarkably little separating one design from another. That's a useful sign of where direct-to-chip cooling sits in the market; the basic approach is widely understood and increasingly commoditized.

What matters now is where the liquid goes, how much of the server it cools and what still has to be handled elsewhere. And those choices are being made before the server ever reaches the data hall.

Liquid cooling is already doing much of the heavy lifting in high-density AI infrastructure. But how to cool AI infrastructure effectively depends on where the liquid goes and how much of the server it actually cools.

Direct-to-chip cooling targets the hottest components

Direct-to-chip cooling puts cold plates directly onto high-power components such as CPUs and GPUs. Liquid flows through those plates, absorbs heat from the chips and carries it away through the cooling loop.

It’s a well-established approach: the biggest heat sources get a dedicated path into the liquid loop, while OEMs can work with a mature ecosystem of cold plates, pumps and other components.

But the cold plate only cools what it touches. Memory, storage, power supplies and other components still generate heat, leaving around 20-30% of total system heat to be managed by air. So while the GPU may be sitting under a cold plate, the fans are still working and the air-cooling infrastructure still has a job to do.

For AI servers, that becomes harder to ignore as power density rises. Removing heat from the GPU solves a large part of the thermal problem, but it doesn’t remove the thermal load from the rest of the server.

Tank immersion puts the whole server in liquid

Tank immersion takes a different approach. Entire servers are submerged in dielectric fluid, allowing heat to be captured across the hardware rather than only at selected components.

Thermally, it’s hard to argue with. There’s no need to keep fans moving air across the server, and heat from components that direct-to-chip cooling doesn’t cover is captured in the fluid instead. The trade-off comes in how the infrastructure has to be built around it. Conventional servers move into fluid-filled tanks, which means the data hall has to accommodate a very different physical setup.

That can make perfect sense in a facility designed around immersion from the outset, but it’s an entirely different proposition - and a much bigger commitment - for operators working with standard rack infrastructure and conventional maintenance practices.

Precision liquid cooling brings whole-system cooling back into the rack

Precision liquid cooling takes the heat-capture principle of immersion and applies it inside the server chassis.

A small volume of dielectric coolant is circulated through the chassis, targeting the hottest components before flowing across the motherboard and the rest of the system. CPUs and GPUs are cooled alongside memory, storage and power components - allowing the system to capture nearly all of the heat produced inside the chassis.

The difference is that the hardware stays in a conventional rack-mounted form factor. The thermal architecture is built into the chassis. There’s no tank to accommodate or internal air loop to maintain. Heat is transferred from the internal coolant to a secondary liquid loop at the rear of the system, changing what the surrounding facility still needs to handle.

What to ask your infrastructure provider

Before committing to new infrastructure, make sure the cooling architecture matches your wider requirements.

  • Was the cooling designed into the server, or fitted around it?
  • If some heat still goes to air, what does the facility actually need to keep supporting?
  • What's the cooling plan for components beyond the CPU and GPU, particularly power delivery, memory and storage?
  • If the liquid loop is interrupted, how much time does the system give operators to respond?

Why cooling is now a server design decision

For OEMs, cooling is becoming part of the server design itself - and that changes what a cooling solution needs to do. Direct-to-chip is now standard technology. When cold plates from ten different vendors can sit side by side at GTC and look identical, it’s obvious that the tech is becoming part of the server rather than a differentiator. The real opportunity for differentiation is in how the rest of the system is cooled.

Whole-board precision liquid cooling requires the chassis, the coolant path and the thermal architecture to be designed together. It can be built around an existing server platform or designed in from the motherboard up - but either way, cooling becomes part of the hardware rather than something the facility has to figure out afterwards.

That opens the door to more than denser AI racks. A sealed, fanless chassis can make edge and harsh environments (where dust, heat and limited access make conventional air cooling difficult) a lot more viable.

And the problem won't stop with CPUs and GPUs: power supplies and power shelves are getting hotter too, while still relying heavily on fans. As more of the hardware starts producing serious heat, cooling has to be designed around the whole system.

That means the days of treating cooling as something that happens around the server are running out. Hardware and the cooling now need to be designed as one system. So which cooling architecture actually makes sense? See how precision liquid cooling compares with direct-to-chip, tank immersion and air cooling - and where each approach starts to run into its limits.

Iceotope works with OEMs to build precision liquid cooling into new and existing server platforms. If you're designing a new server platform or looking at how to cool an existing one, talk to us about your precision liquid cooling needs.

‍