Rugged edge computing challenges in thermal management

Rugged edge computing challenges in thermal management

Key takeaways:

  • Ruggedized hardware can be built to withstand extreme conditions, but fans and air paths can still introduce vulnerabilities that undermine long-term reliability.
  • Edge deployments expose cooling systems to dust, moisture, heat and vibration, making the surrounding environment part of the thermal design.
  • Precision liquid cooling removes the need to pull ambient air through the chassis, reducing exposure to contaminants and adding thermal capacity for remote deployments.

When ruggedized hardware meets the real world

Qualification testing for military deployment is more brutal than most people realize. One of the tests involves a large hammer swung into the side of the chassis, and the machine has to come out the other side still working. Servers built for the back of a Humvee go through that kind of punishment and pass. Then the same class of hardware ends up on a pole or inside a roadside cabinet, and eighteen months later a fan seizes and the site goes quiet. That’s where things get interesting.

Ruggedized equipment can be exceptionally well engineered mechanically, but the cooling system can introduce vulnerabilities that the rest of the chassis was designed to eliminate. The moment air has to pass through the chassis, the cooling system creates new points of failure: fans can seize, contaminants can enter through vents, and moving parts have to keep working reliably over years of deployment.

The hardware may be built to survive a hammer, but its cooling system still has to survive everything that comes after.

Nobody agrees what "rugged edge" means

One of the biggest challenges of edge computing is that the terminology around it is inconsistent, which makes specifying equipment harder than it should be. Two vendors can use completely different terms for the same deployment, making it difficult to know whether they’re actually talking about the same kind of environment.

For all the complexity involved in data center design, the environment around the server is relatively predictable. It’s designed around the needs of the equipment, with dedicated cooling, controlled temperature and the infrastructure needed to keep servers running.

Edge deployments are much less predictable. Equipment might sit in a retail cupboard with some climate control, inside a roadside cabinet without air conditioning, or outside in an environment exposed to dust, heat, moisture and vibration.

At the more extreme end, the hardware may need to meet IP ratings, military standards or NEBS requirements. But those labels still don’t tell you everything about the conditions the hardware will actually face.

For thermal design, the more useful question is simple: how much control do you have over the environment around the compute? Once it leaves a controlled space, temperature, dust, moisture and vibration become part of the thermal design problem.

What actually takes these deployments offline

Once compute leaves a controlled environment, the cooling system has to deal with everything the site throws at it. Dust, moisture, vibration and limited access can all turn into reliability problems. Dust is the obvious one. If a system needs to pull air through the chassis, it needs an opening to let that air in - and with that comes dust.

Salt and humidity are slower and harder to spot. Coastal sites and anywhere with large temperature swings push moisture through the same air path, bringing moisture and corrosive contaminants closer to the electronics the enclosure is supposed to protect.

That’s the contradiction in rugged edge design: the hardware might be built to withstand the environment around it, but an air-cooled system has to draw that environment through the box to keep the equipment cool.

Everything the site does, it does to the fan

Ruggedized hardware is generally good at dealing with vibration. The chassis can be reinforced, components can be shock-mounted and fasteners and the system can be qualified for conditions that would be unacceptable in a conventional data center. Then you add a fan.

Now there’s a mechanical component spinning at several thousand RPM inside that same system - potentially for years. It has its own bearings, its own failure modes and its own tolerance for the conditions around it.

The site itself creates other constraints:

  • If people are nearby, noise becomes a design constraint.
  • There may be barely enough room for the compute, let alone the additional space a conventional cooling system needs
  • At an unattended site, a failed fan can mean a site sitting offline until a technician can reach it.

If the electronics are ruggedized for the environment, the cooling system needs to be designed for it too.

The fan is the common factor

Run back through those failure modes and the same component keeps appearing: the fan.

Fans are well understood and proven technology, particularly where the equipment is accessible and maintenance is routine. At the rugged edge, the calculation is different because every moving part, maintenance visit and environmental exposure adds another variable.

Ruggedized systems can make air cooling more resilient. Higher-grade fans, better filtration and redundancy can all reduce individual failure risks, but they’re still working around the same basic architecture: moving ambient air through the enclosure to carry heat away.

For remote and exposed deployments, there comes a point where hardening the air-cooling system further adds complexity without removing the underlying dependency. And that’s where rugged edge computing calls for a cooling architecture that removes the external air path altogether.

Removing the failure mode instead of hardening it

If the air path is one of the ways the environment gets into the server, and the fan is what keeps that air moving, there’s another option: take both out of the server’s thermal architecture.

Iceotope's approach to this is precision liquid cooling - using single-phase dielectric fluid inside a sealed chassis and circulating it around the components generating the most heat. The server no longer needs ambient air moving through the enclosure, so dust, moisture and other airborne contaminants are no longer drawn through the electronics.

It's worth being precise about what fanless means here, because the term gets thrown about. The cooling system still needs fluid to circulate, but the server itself no longer relies on fans pulling outside air through the chassis. The rotating hardware associated with moving that air is removed from the server. That changes what the hardware needs from its surroundings. There’s no server airflow path to filter or maintain, and no fan inside the chassis that can become a failure point.

For rugged edge computing, that’s a different approach to thermal management. Instead of making the air path progressively harder to fail, you design the server so it doesn’t need one.

What happens in a power failure

When the power goes, the compute stops and active cooling stops with it. In a fan-cooled system, that means the airflow disappears with the power. In a liquid-cooled system, there is still a volume of coolant around the components, and that fluid keeps absorbing heat as temperatures begin to equalize.

That gives the system a period of passive thermal capacity after active circulation has stopped. How long that period lasts depends on factors including the coolant volume, starting temperature, the load immediately before the interruption and the thermal characteristics of the system.

Where backup power remains available to the rack, our internal dielectric fluid pumps can maintain system cooling for up to five minutes. The volume and heat capacity of the fluid also provide a safety buffer, extending UPS runtime before thermal limits are reached.

This principle isn’t unique to precision liquid cooling. An immersion system has a much larger volume of fluid, so there’s more thermal mass to work with. Precision liquid cooling takes a more targeted approach, putting coolant directly around the highest-heat components inside a sealed chassis.

For rugged edge deployments, that additional thermal capacity can provide useful thermal ride-through after active cooling stops, subject to the operating conditions of the system. That’s particularly useful in rugged edge AI deployments, an area we’ve looked at in more detail when comparing precision liquid cooling with air cooling.

It has to be designed in from the start

Ruggedized electronics need thermal management designed in from the start. The later thermal design enters the product development process, the fewer options there are. Once the motherboard, enclosure and other core decisions are fixed, the cooling system has fewer options and less room to work with.

That’s why Iceotope works with OEMs and ODMs at the platform level. With one telco RAN project, we took the customer’s motherboard and designed the chassis and cooling architecture around it from the ground up.

Starting there gives the thermal design a lot more freedom: the enclosure, component layout and cooling system can be designed around the conditions the product actually needs to handle. We’ve explored this approach in more detail in our article on liquid cooling for edge platforms.

But there’s also another route for customers working with an existing server platform. The For example, we recently designed a ruggedized chassis and cooling solution around a standard Dell XE7740 server. That approach can avoid a complete platform redesign while still changing how the hardware is cooled and protected.

Both have a place. The point is to involve thermal management before the hardware is boxed in by decisions made for a completely different environment - designing the cooling architecture around the deployment rather than forcing the deployment around the cooling system.

Building compute for the rugged edge? Talk to Iceotope about thermal architecture - before the environment makes the decisions for you.