Data Center Fundamentals·Cooling & HVAC Systems
Liquid Cooling: The Future of High-Density
Discover how liquid cooling enables higher rack densities for AI/ML and HPC workloads.
Introduction Last month, I walked through a new AI training facility in Northern Virginia where the racks were pushing 120 kW per cabinet.
The facility manager looked me in the eye and said, "We can't physically move enough air anymore." That's the reality hitting operators everywhere.
Moore's Law didn't stop-it just shifted from more cores to more watts.
NVIDIA's H100 GPUs pull 700W each, and we're cramming eight of them into a single 4U chassis.
Traditional air cooling taps out around 25-30 kW per rack.
Beyond that, you're fighting physics.
Liquid cooling isn't some futuristic concept anymore.
It's the bridge between what we've always done and what AI and high-performance computing demand today.
When I started in this industry 15 years ago, liquid cooling meant exotic installations in government supercomputing facilities.
Now I'm specifying it for enterprise clients running large language models.
The technology comes in two primary flavors: direct-to-chip cooling, where liquid flows through cold plates mounted directly on processors, and immersion cooling, where entire servers sit in dielectric fluid baths.
This lesson breaks down the real-world economics, the performance metrics that matter, and where each technology fits.
You'll learn how to evaluate whether your facility needs liquid cooling and which approach makes sense for different workloads.
The Physics Problem Air Cooling Can't Solve Air has a thermal conductivity of 0.024 W/m·K.
Water? About 0.6 W/m·K-roughly 25 times better.
Dielectric fluids used in immersion cooling fall somewhere in between at 0.06-0.08 W/m·K.
These aren't abstract numbers.
They represent fundamental limitations on how much heat you can remove from increasingly dense computing environments.
Here's the breakdown of what different cooling methods can realistically support:
| Cooling Method | Max Rack Density | Typical PUE | Complexity |
|---|---|---|---|
| Traditional Air (Hot/Cold Aisle) | 8-15 kW | 1.4-1.6 | Low |
| Contained Air (CRAC) | 15-25 kW | 1.3-1.5 | Medium |
| Rear-Door Heat Exchangers | 25-35 kW | 1.2-1.4 | Medium |
| Direct-to-Chip Liquid | 50-100+ kW | 1.1-1.2 | High |
| Immersion Cooling | 100-250+ kW | 1.03-1.15 | Very High |
At Microsoft's Quincy, Washington facility, I watched them retrofit a legacy air-cooled hall originally designed for 6 kW racks.
Upgrading the CRAC units, adding more cooling capacity, and reinforcing power distribution to support 20 kW racks cost them roughly $8 million for 1,200 racks.
That's $6,600 per rack just for the cooling infrastructure upgrade-and they still maxed out at 20 kW.
Compare that to a greenfield liquid cooling deployment.
Meta's AI Research SuperCluster uses direct-to-chip cooling handling 70+ kW per rack.
Their capital expenditure per rack for cooling infrastructure ran closer to $12,000, but they're supporting 3.5x the density.
The cost per kilowatt actually drops to around $170 versus $330 for their air-cooled retrofit.
Temperature also tells an important story.
Processors running cooler operate more efficiently and last longer.
With air cooling, junction temperatures on high-performance CPUs routinely hit 85-90°C under load.
Direct-to-chip cooling keeps those same processors at 55-65°C.
That 20-30 degree difference translates to roughly 10-15% better computational performance due to reduced thermal throttling.
When you're running training jobs that take weeks, that performance gain matters.
Direct-to-Chip: Targeted Heat Removal Direct-to-chip cooling places cold plates-flat metal heat exchangers-directly on top of the hottest components: CPUs, GPUs, memory modules, and sometimes voltage regulators.
Liquid flows through microchannels in these plates, absorbs heat, and carries it away to a cooling distribution unit (CDU).
I worked with Equinix on their Chicago CH3 facility when they added a liquid cooling zone specifically for AI workloads.
They installed ColdStream cold plates from CoolIT Systems on customer servers running NVIDIA A100 GPUs.
The implementation required facility-side infrastructure: a CDU on each row providing cooled water at 45°F (7°C), supply and return manifolds running along the top of each rack, and quick-disconnect fittings that customers could plug into without draining the system.
The customer experience? Plug in power, network, and two liquid cooling lines.
The cold plates handled 300W per GPU-about 75% of each GPU's thermal output.
The remaining heat from drives, power supplies, and other components still exhausted into the room and required traditional air cooling, but at dramatically reduced volumes.
Instead of moving 4,000 CFM per rack, they dropped to around 800 CFM. Key advantages of direct-to-chip:
- Works with existing server form factors (rack-mounted equipment)
- Customers can maintain familiar hardware designs
- Can be deployed incrementally-mixing liquid and air-cooled racks
- Lower risk than full immersion
- Existing IT staff can manage it with training The challenges are real too:
- Dual cooling systems still required (liquid for hot components, air for everything else)
- Leak detection and prevention becomes critical
- Regular maintenance on fittings, hoses, and pumps
- Requires facility-side liquid distribution infrastructure Google deployed direct-to-chip cooling in their custom TPU v4 pods.
They're using a water/glycol mix at 65°F (18°C) supply temperature and seeing return temperatures around 95°F (35°C).
That 30°F delta carries away approximately 90% of the pod's thermal load.
Their PUE for these specific cooling zones runs around 1.08, compared to 1.12 for their best air-cooled facilities.
Immersion Cooling: The Complete Solution Immersion cooling submerges entire servers-motherboards, processors, memory, storage, everything except fans-in thermally conductive but electrically non-conductive fluid.
Two approaches dominate: single-phase immersion, where the fluid stays liquid, and two-phase immersion, where the fluid boils and condenses in a closed loop.
I visited a two-phase immersion deployment at a cryptocurrency mining operation in Texas back in 2021.
Servers sat in open baths about four feet long, two feet wide, and 18 inches deep.
The dielectric fluid (3M Novec in this case) boiled at 122°F (50°C), turning to vapor as it absorbed heat from the servers.
The vapor rose to a condenser coil at the top of the tank where facility water cooled it back to liquid, and gravity returned it to the bath.
No pumps needed for fluid circulation.
The entire process was eerily quiet-no server fans, no screaming air handlers, just a gentle bubbling sound.
Single-phase immersion keeps the fluid liquid and pumps it through a heat exchanger.
LiquidStack, a company I've worked with, deployed single-phase systems at several HPC facilities.
Servers slide into sealed chassis that fit standard 19-inch racks.
The fluid circulates at about 10-15 gallons per minute per chassis, entering at 95°F and exiting around 115°F.
Heat exchangers in the CDU transfer thermal energy to facility water, which then goes to cooling towers or chillers. Density becomes almost absurd: A single 42U rack can support 250 kW or more.
That's like fitting the cooling capacity of ten traditional racks into one footprint.
For facilities constrained by space, this is transformative.
QTS deployed immersion cooling in their Irving, Texas facility specifically to support a customer's AI training cluster without having to build new space.
| Immersion Type | Fluid Boiling Point | Typical Cost per Liter | Pump Requirements | Complexity |
|---|---|---|---|---|
| Single-Phase |
212°F (100°C) | $15-40 | High flow pumps needed | Medium-High | | Two-Phase | 90-130°F (32-55°C) | $40-150 | Minimal (natural convection) | High | The operational differences matter.
Single-phase fluids cost less and are easier to work with-servers can be pulled and reinserted without much ceremony.
Two-phase fluids cost significantly more ($80-150 per liter vs. $15-40), but their superior heat transfer characteristics mean smaller, simpler cooling infrastructure.
When servers boil the fluid, you're working with latent heat of vaporization, which carries away far more energy than simple temperature change. Real-world challenges I've seen:
- Fluid management-tracking inventory, preventing contamination, managing top-offs
- Server compatibility-not all components tolerate immersion (some capacitors don't like it)
- Maintenance procedures-technicians need training on safe fluid handling
- End-of-life fluid disposal or reclamation
- Insurance and risk perception (convincing stakeholders that liquid and electronics mix safely) Microsoft's Project Natick proved immersion viability at scale.
They submerged an entire modular datacenter underwater off Scotland's coast, using ocean water for cooling.
After two years, they pulled it up and found server failure rates were one-eighth of their land-based facilities.
The sealed, inert environment actually improved reliability.
Where Liquid Cooling Makes Financial Sense The math changes dramatically depending on your density targets and geographic location.
I built a total cost of ownership model for a client last year comparing air versus liquid cooling for their planned AI inference facility in Phoenix-a brutally hot climate where air cooling struggles.
For a 2 MW facility at 40 kW average rack density: Air-cooled approach: 50 racks total, required massive CRAC infrastructure rated for 2.4 MW cooling (accounting for Phoenix ambient temperatures), estimated annual power cost of $1.6 million for IT load plus $800K for cooling infrastructure.
PUE projected at 1.4.
Total annual operating cost: $2.4 million. Direct-to-chip approach: Same 50 racks, reduced CRAC capacity to 0.6 MW (handling residual heat only), cooling power dropped to $240K annually, IT load unchanged at $1.6 million.
PUE of 1.12.
Total annual operating cost: $1.84 million.
Annual savings: $560K.
The capital cost difference? Air-cooled infrastructure: $4.2 million.
Direct-to-chip infrastructure: $6.8 million.
Payback period: 4.6 years.
Except we haven't factored in the clincher: the air-cooled approach would have required an additional 3,000 square feet of space for cooling equipment.
Real estate costs in that market ran $200 per square foot.
That's another $600K in construction costs, bringing payback down to 3.8 years. AI and HPC workloads hit the sweet spot for liquid cooling adoption.
Training large language models or running computational fluid dynamics simulations generates sustained, heavy thermal loads.
These aren't bursty web servers that average 30% utilization.
They run at 80-95% utilization for hours or days.
The density and consistency make liquid cooling's economics compelling.
Bitcoin mining operations were early adopters because their economics are brutal-every watt spent on cooling is a watt not earning revenue.
Immersion cooling cut their cooling overhead from 30-40% of total power draw down to 5-10%.
When bitcoin prices run high, that difference represents millions in annual revenue.
Practical Implementation Examples
Example 1: Retrofitting an Existing Facility Digital Realty tackled this at their Frankfurt facility when a financial services client needed to deploy GPU servers for quantitative trading models.
The existing infrastructure supported 12 kW per rack, but the new servers needed 45 kW.
Building out a new cooling plant would have taken 18 months and cost €4 million.
Instead, they installed rear-door heat exchangers from Motivair on six racks as a temporary measure while retrofitting a single row with direct-to-chip cooling.
They ran new facility water piping along that row, installed four CDUs (one per pair of racks), and worked with the customer to retrofit their Dell servers with cold plates from Asetek.
Timeline: four months.
Cost: €800K.
The solution handled the immediate need and proved the concept for broader deployment.
The key decision framework they used:
- Rack count under 10: Rear-door heat exchangers (lower complexity)
- Rack count 10-30: Direct-to-chip (justifies infrastructure investment)
- Rack count 30+: Evaluate immersion (economies of scale improve)
- Density over 80 kW: Immersion becomes the practical choice regardless of count
Example 2: Greenfield AI Training Facility CoreSite designed their new Silicon Valley facility (SV11) from the ground up for AI workloads.
They allocated one-third of the space (1,200 square feet) specifically for liquid-cooled infrastructure supporting 3.5 MW in 28 racks-that's 125 kW average per rack.
They chose single-phase immersion from LiquidStack based on several factors:
- Customer workloads were known to be AI training (sustained high utilization)
- Space constraints made maximum density critical
- Local utility costs favored efficiency ($0.19 per kWh)
- Customer wanted turnkey solutions, not DIY server modifications The design included three immersion cooling zones, each with its own CDU and heat rejection system.
Cooling water ran at 95°F supply, 115°F return, connecting to an adiabatic cooling tower on the roof.
Projected PUE: 1.06.
Capital cost ran $14,000 per kW-expensive, but the per-square-foot revenue at
2.9 kW per square foot made the economics work.
Example 3: Hybrid Deployment Strategy AWS takes a pragmatic approach in their custom datacenters.
Standard EC2 instance workloads run on air-cooled infrastructure optimized over 15+ years.
But their AI/ML instances (P4 and P5 families) use direct-to-chip cooling.
They manufacture custom cold plates in-house, maintain separate liquid cooling zones, and charge premium pricing for those instance types that reflects the infrastructure cost.
The calculation that drove this decision: P5 instances with eight H100 GPUs generate roughly 10 kW per 4U server.
At 10 servers per rack, that's 100 kW per rack.
Air cooling would require 2.5x the floor space to achieve the same compute density.
In Northern Virginia real estate where AWS operates multiple facilities, space costs $150+ per square foot.
The liquid cooling premium is cheaper than the real estate expansion cost.
Common Misconceptions "Liquid cooling eliminates all cooling infrastructure costs." Not even close.
You still need cooling towers or dry coolers, pumps, CDUs, monitoring systems, and often some air cooling for residual heat.
What liquid cooling does is dramatically improve the efficiency of heat transport from the source to your heat rejection equipment.
I've seen proposals claiming 80% infrastructure cost reduction-those are marketing fiction.
Realistic savings run 30-40% on cooling energy, not total infrastructure cost.
The related misconception: "Liquid cooling means no chillers needed." That depends entirely on your liquid supply temperature requirements.
Direct-to-chip systems running at 45-50°F supply temperature absolutely need chillers in most climates.
Single-phase immersion running at 95°F can use free cooling in many locations for much of the year.
Two-phase immersion with high-boiling-point fluids can potentially eliminate chillers entirely.
The specifics matter enormously. "Leaks are a constant disaster waiting to happen." Modern liquid cooling systems use quick-disconnect fittings with automatic shut-off valves.
When I disconnect a line on a direct-to-chip system, maybe 5-10 ml of fluid drips out-about two teaspoons.
Leak detection systems monitor pressure differentials and can isolate problems in seconds.
In 15 years, I've personally witnessed two significant leaks: one from an improperly torqued fitting during installation, caught immediately by pressure monitoring, and one from physical damage where a forklift hit a manifold.
Properly installed and maintained systems are remarkably reliable.
Compare this to the "air cooling never fails" mythology-I've seen far more server damage from CRAC failures causing thermal runaway than from liquid cooling incidents.
Summary & Key Takeaways
- Physics dictates the need: Air cooling maxes out around 25-30 kW per rack.
AI and HPC workloads pushing 50-150 kW per rack require liquid cooling to function at all.
- Two primary technologies serve different needs: Direct-to-chip works for moderate density increases (30-80 kW) with familiar server formats.
Immersion handles extreme density (80-250+ kW) but requires more operational changes.
- Economics hinge on density, power costs, and space constraints: Liquid cooling pays back faster in expensive real estate markets, hot climates, and when rack density exceeds 40 kW.
Calculate your specific scenario rather than assuming blanket benefits.
- PUE improvements are real but not dramatic: Expect PUE to drop from typical 1.4-1.6 (air) to 1.1-1.2 (liquid).
The bigger win is often the density enabling revenue per square foot increases.
- Implementation risk is manageable: Modern systems are engineered for datacenter reliability.
Proper training, installation, and maintenance protocols make liquid cooling operationally viable for mainstream deployments.
- AI workloads are driving mainstream adoption: What was exotic five years ago is becoming standard design practice for facilities targeting AI training, inference, and HPC customers.
Next Steps Build on this foundation by exploring thermal management calculations in the Power Distribution & Management module, where you'll learn to calculate cooling load requirements and size infrastructure appropriately.
The Efficiency & Sustainability module covers advanced topics like waste heat reuse, where liquid cooling's high-temperature heat rejection creates opportunities for district heating or absorption cooling that air-cooled systems can't match.
Understanding commissioning procedures for liquid cooling systems will prepare you for real-world deployment scenarios.