Data licensing
Licence the dataset, not a seat.
Facility-level intelligence on the world's data centres: coordinates, capacity, ownership and market context, delivered as versioned bulk files and API access, and licensed for use inside your business or inside your product.
6,800+
Facilities
Located, named and attributed to an operator
165+
Countries
Every region we can source publicly
87+ GW
Live capacity
Energised today
160+ GW
Pipeline
Under construction or planned
Every figure on this page is a floor. They are rounded below what the corpus measures so a number here does not go stale between releases.
Who licences it
Built for teams that build on data.
Research and data teams spend hours on operator websites, PDF brochures and stale spreadsheets to establish something as basic as a site's power capacity. We aggregate thousands of public sources into one standardised database and carry source references back to what each record was built from. You get the facility-level view of the global market without maintaining the collection pipeline.
Spatial and ESG data providers
Enrich your own database with facility coordinates, operator linkages and the serving utility, without manual compilation.
Investors and analysts
Benchmark operators and markets on standardised capacity, pipeline and ownership across every major region.
Operators and developers
Track competitor expansion, market saturation and announced-to-built conversion in the markets you are entering.
Enterprise site selection
Compare facilities on power, connectivity, certifications and redundancy, with source references carried on the record.
The dataset
Every facility, standardised.
Each profile carries 50+ structured fields, normalised across operators and markets so records compare cleanly, and every capacity figure carries the basis it rests on.
Location and geometry
Point coordinates on every facility — 100% of the published set — with address, campus grouping and market assignment, ready for spatial joins.
Power and capacity
Live capacity and total potential including expansion headroom, with the serving utility named on over 80% of facilities. Each megawatt is counted once.
Ownership and operator
Operator, owner and investor linkages resolving to 2,000+ company records with portfolios, parentage and market position.
Type and lifecycle status
Colocation, wholesale, hyperscale or edge, tracked from announced through construction to operational and decommissioned.
Connectivity and certifications
Carrier neutrality, interconnection, redundancy design and certification status across compliance frameworks.
Provenance
Capacity figures carry the basis they rest on — whether stated by the operator, sourced, or modelled — so the set can be filtered to your own confidence threshold.
What is actually populated
Coverage across the published set, stated as floors. The thin fields are here because a coverage table that lists only the full ones is a coverage table nobody believes.
- Coordinates
- 100%
- Capacity basis (provenance)
- 95%+
- Power capacity
- 80%+
- Serving utility
- 80%+
- Total building area
- 65%+
- Operational date
- 45%+
- Cooling type
- 40%+
- Announced date
- 35%+
What we do not hold
Geometry is a point, not a polygon.
Every published facility carries a coordinate pair in WGS84, with its address, its campus grouping and its market assignment. That is enough to join the dataset to your own spatial layers, to aggregate it by any boundary you already hold, and to place a site on a map.
It is not enough to measure a building. We hold no building footprints and no parcel boundaries. There are zero polygons in the dataset and no roadmap item that adds them. Where we have recorded how precisely a point is placed, that flag is on roughly a quarter of the set; the rest carry the coordinate without a stated precision.
If your use depends on building area, the structured field to look at is total building area, populated on 65%+ of the set. If it depends on a footprint you can compute against, we are not the source and you should not licence us for it.
How it stays current
A dataset that is watched, not just published.
Data centre records go stale in a way most reference data does not. A site is announced, permitted, financed, energised and expanded, often under three different names, and the interesting change usually appears first in a local planning notice or a regulatory filing rather than a press release. Compiling once does not survive contact with that.
- 01
Continuous monitoring
Automated collectors watch operator disclosure, regulatory and planning sources, tenders and market reporting across every country we cover, on a daily cycle rather than a research cycle.
- 02
Machine extraction
Language models read what comes back and turn it into structured candidates: which facility, what changed, what stage it moved to, what capacity was stated and by whom.
- 03
Classification and scoring
Each candidate is typed, dated at the precision its source states, and scored for confidence. Anything that conflicts with what we already hold is held back rather than overwritten.
- 04
Human adjudication
A small senior team reviews what the pipeline surfaces and decides what lands. Machines find the change; people decide whether we believe it.
What that produces is a timeline per facility — expansion, acquisition, announcement, commissioning — each entry dated at the precision its source states, and each traced to that source. 5,000+ market events are tracked across the corpus today.
Delivery
It arrives the way your stack expects it.
Bulk dataset
Full or market-scoped extracts as flat files (CSV or GeoJSON) against a stable schema, with versioned releases and a changelog between drops so you reprocess what changed rather than the whole set.
API access
Query facilities, operators and markets programmatically, and pull changes since your last sync instead of re-ingesting everything. Included with every licence.
At a glance
- Coverage
- 6,800+ facilities across 165+ countries: colocation, wholesale, hyperscale and edge, operational and pipeline
- Granularity
- Facility level, with campus grouping and operator rollups; 50+ fields per record
- Geometry
- Point coordinates on 100% of facilities (WGS84). We do not hold building footprints or parcel polygons
- Capacity basis
- Capacity figures carry the basis they rest on, so the set can be filtered by confidence. Where a source does not state whether a figure means IT load or total facility power, that is recorded rather than assumed
- Update frequency
- Database updated continuously; bulk releases on a cadence scoped to your licence, quarterly by default
- Formats
- CSV and GeoJSON as standard; Parquet on request
- Identifiers
- Stable facility and company IDs, persistent across releases for clean joins
What the dataset is growing into
It is added to continuously, so a licence covers what it becomes across the term rather than a fixed snapshot. Three areas are active now.
Interconnection queues
Grid connection applications and queue position, linked to the facilities and utilities already held.
Deeper source coverage
Regulatory, planning and permitting sources in more jurisdictions, which is where a project first becomes visible.
Market-level analysis
Extending from facility records into comparative analysis at market level.
Rights
Two licences, one question: how far does the data travel?
The dataset is the same in both. What changes is who ends up holding it. API access and platform seats come with every licence.
Internal use
It stays inside the business
Load the dataset into your own systems, join it with your own data, and use it across your team without per-seat restrictions. Nothing derived from it leaves the business.
Commercial
It reaches your clients — as conclusions, or as rows
Everything above, plus the analysis, scores and products you build on the data and sell. Where your clients see only what you concluded, that is the floor. Where the underlying records pass through to them as records, the licence is scoped on volume above it — which is the one thing that genuinely varies once the data leaves twice, and the reason the term is agreed in the conversation rather than printed here.
Licences run per organisation on a fixed term. A market-scoped licence costs less than the global set, and early licensees are on better terms than list while delivery is still hands-on. We quote in the reply rather than on the page, because a rate published today is the rate we are held to for everyone who arrives after you.
Looking for accounts on the platform instead of rights to the data? That is a different question, and the plans page answers it. A plan is people reading and downloading; a licence is the data delivered in bulk into your own systems.
Next steps
Evaluate it before you commit to it.
- 01
Sample extract
Name a market and we send a scoped dataset to evaluate in your own stack.
- 02
Schema and coverage
The full field dictionary with per-field coverage, so you know what is populated before you commit.
- 03
Licence proposal
Terms scoped to your use. No meetings required to get them.
Ask for a licence
Tell us what you want to do with the data and how much of it you need. We reply with a sample extract for a market of your choice, the field dictionary with its per-field coverage, and terms scoped to the use. No meeting required to get any of it.
Questions
Asked before you ask.
- Do you hold building footprints or parcel polygons?
- No. Geometry is a point per facility in WGS84, on 100% of the published set, with the address, campus grouping and market assignment beside it. There are no building footprints and no parcel boundaries anywhere in the dataset.
- How is the data licensed?
- Two tiers on one axis: how far the data travels. Internal use keeps everything derived inside your business. Commercial covers what reaches your clients — the analysis and products you sell them, and, scoped on volume, the underlying records passing through to them as records.
- How current is it?
- The database is updated continuously: automated collectors run daily, machine extraction turns what they find into structured candidates, and a senior team adjudicates what lands. Bulk releases go out on a cadence scoped to your licence, quarterly by default, with a changelog between drops.
- What formats do you deliver?
- CSV and GeoJSON as standard, Parquet on request, against a stable schema with facility and company IDs that persist across releases. API access is included with every licence.
Every figure DC Atlas publishes is an estimate grounded in public sources and our own corpus. See the data disclaimer.