A DC Atlas Frontier
Module 5 of 7
The Operations Problem
GPUs die in one to three years and in orbit no one can swap them. Radiation-driven failure, orbital debris, and why on-orbit servicing at tens of millions per mission can't stand in for a loading-dock full of spares.
Built on published aerospace engineering figures and DC Atlas's facility data.
Suppose you have solved heat, launch and power — you have not, but suppose. You now own a 40 MW cluster of the most failure-prone hardware in the modern economy, in the one place where nothing can be repaired. This is the constraint the pitch decks never model, because it has no clean engineering fix. It is not a materials problem or a thermodynamics problem. It is that data centres are living systems, kept alive by people continuously replacing the parts that break, and in orbit there are no people and the parts still break.
Begin with how mortal AI hardware already is on the ground, where conditions are ideal. Across large GPU fleets, something like 15 to 25 percent of accelerators are lost to manufacturing attrition before they even enter steady service, a few percent more fail during burn-in, and in-service annualised failure rates sit around 9 percent. A Google architect has put the practical service life of a data-centre GPU under sustained AI load at one to three years. Terrestrial operators absorb this with continuous hot-swap maintenance: spare inventory on site, technicians on shift, orchestration software that routes work around dead nodes and a person who replaces them within hours. The entire operating model assumes hands on hardware.
How a GPU fleet dies — and why the ground survives it
Share of an AI accelerator fleet still working, with and without replacement
Illustrative, from a ~9%/yr in-service failure rate compounded (higher in orbit's radiation environment), against a terrestrial fleet held at full strength by continuous replacement. A ground operator also refreshes the whole fleet every 1–3 years for performance; an orbital operator can do neither.
Now take that fragile hardware and fly it through the low-orbit radiation environment. Commercial off-the-shelf silicon — exactly the kind you would want, because it is what makes AI economics work — is fabricated at tiny geometries and low voltages that make it acutely sensitive to radiation. Charged particles flip bits (single-event upsets) and slowly degrade transistors (total ionising dose). On the ground these effects are negligible; in orbit they become a primary driver of failure, which is why the orbit curve above bends down faster than the ground failure rate alone would suggest. You can harden or shield against them, but hardening costs performance and shielding costs mass — and mass, from Module 3, is the thing you cannot afford. So the orbital operator faces a vice: fly cheap commercial chips that fail fast in radiation, or fly hardened chips that cost more, compute less, and mass more.
| Dimension | Terrestrial data centre | Orbital data centre |
|---|---|---|
| GPU service life | 1–3 years under heavy AI load | Same — but unreplaceable |
| Annual in-service failure | ~9% per year | Higher (radiation-driven upsets + dose) |
| Replacing a dead unit | Hot-swap, hours, spares on site | On-orbit servicing: tens of $M per mission |
| Refreshing to new silicon | Every 1–3 years, routine | Effectively never — obsolete in ~18 months |
| Radiation environment | Negligible | Single-event upsets + total-dose damage to COTS silicon |
| Physical threats | None to speak of | >25,000 tracked debris >10 cm; impacts at ~10 km/s |
Finally, the neighbourhood. Low Earth orbit is not empty; it is increasingly crowded and increasingly hostile. More than 25,000 catalogued objects larger than 10 centimetres, around half a million between 1 and 10 centimetres, and over 100 million fragments larger than a millimetre share that space, all moving at closing speeds near 10 kilometres per second — fast enough that a fleck of paint hits like a bullet. A large, sprawling structure like a 40 MW data centre, with its enormous radiator and solar surfaces, is a big target that must carry shielding and spend propellant dodging conjunctions, both of which cost the mass you were already short of. And every collision risks creating more debris, in the runaway feedback known as Kessler syndrome.
Orbital Debris Frequently Asked QuestionsAdd it up and orbit is a place that breaks your hardware faster, forbids you from fixing it, and throws debris at it — which is why the last two questions are whether you could even keep it connected, and then whether any of this was ever worth doing against a building on the ground.
Questions this module answers
- Why can't you run a data centre in orbit long-term?
- GPUs fail at around 9% per year and need refreshing every one to three years. In orbit no one can swap them, and on-orbit servicing costs tens of millions of dollars per mission.
- Does radiation matter for the hardware?
- Yes. It causes single-event upsets and cumulative damage to commercial chips, making orbital failure rates worse than on the ground.
- Is orbital debris a real risk?
- There are over 25,000 tracked objects larger than 10 cm plus millions of smaller fragments, hitting at about 10 km/s. A large data centre is a big target needing shielding and avoidance propellant.