All insights

DCFR Insight 46 / AI Factories

AI Factory Planning for GPU Superclusters and Productive Compute

An AI factory is not a conventional data hall with denser racks. It is a synchronized system of accelerators, fabric, storage, power, liquid cooling, software, and operations whose real output is reliable computational work.

AI Factory Planning for GPU Superclusters and Productive Compute

Define the product as completed work—not installed megawatts

Start with the workload: training, inference, fine-tuning, simulation, or a mixed portfolio. State model scale, job duration, concurrency, data locality, latency, checkpoint behavior, availability, security, and growth assumptions. Translate these into accelerator count, cluster topology, network bandwidth, storage throughput, power profile, heat rejection, and recovery objectives. The core acceptance measure should be productive compute—workloads completed correctly and predictably—not only energized racks, nameplate accelerator quantity, or facility power capacity.

Treat the AI factory as one vertically integrated stack

The system spans utility supply, generation and storage, medium-voltage distribution, rack power, heat capture, heat rejection, accelerators, scale-up and scale-out fabrics, storage, orchestration, observability, and workload software. Design teams need a shared model of dependencies and failure domains. A high-performance compute platform can be limited by weak storage, network oversubscription, unstable coolant temperatures, firmware drift, power-quality events, or job scheduling. Optimize total useful output and time to recovery across the stack rather than maximizing any single subsystem.

Layered AI factory stack from utility and thermal infrastructure through workload operations
Every layer must be sized and commissioned against the same workload model; a bottleneck anywhere can strand expensive accelerators.

Lock the cluster topology before the building grid

Establish the smallest fault-contained compute unit and the largest job that must remain within a low-latency domain. Map rack arrangement, copper and optical reach, cable counts, rear-door or direct-liquid-cooling connections, overhead and underfloor zones, service clearances, replacement paths, and fabric expansion. The physical topology should follow the logical network and workload architecture. Reversing this sequence can force long cable routes, fragmented clusters, blocked service access, or a building bay that cannot accept the next platform generation.

AI Factory Basis of Design

LayerDesign inputAcceptance evidenceTypical blind spot
WorkloadJob types, size, concurrency and recoveryRepresentative jobs complete at agreed service levelSynthetic tests do not reflect production behavior
Compute and fabricCluster topology, fault domain and bandwidthHealth, performance and isolation testsLogical design is forced into an unsuitable floor grid
Storage and dataCapacity, throughput, metadata and ingestionEnd-to-end data path under loadAccelerators wait on storage or permissions
PowerSteady, transient, restart and maintenance statesCoordinated system response with rack emulationNameplate values replace time-based load behavior
CoolingHeat split, temperatures, flow and water qualityBalanced clean system at all operating modesFacility and IT ownership boundary is unresolved

Requirements are platform- and workload-specific. Validate all values with current equipment suppliers and the owner’s operating model.

Engineer power for dynamic load and technology change

Model steady state, synchronized ramps, step loads, regenerative or harmonic behavior where applicable, restart sequences, maintenance states, and partial-cluster operation. Coordinate utility, on-site resources, transformers, protection, UPS architecture, busways, rack conversion, grounding, and controls against the actual platform. Preserve pathways for higher-voltage distribution and alternative rack-power architectures without assuming a future standard prematurely. Redundancy should align with job checkpointing and workload migration; identical facility tiers can produce very different business resilience.

Build a liquid-cooling operating system, not a pipe loop

Define facility-water and technology-cooling-system boundaries, supply temperatures, pressure, flow, water chemistry, materials compatibility, filtration, leak detection, make-up water, heat-exchanger approach, controls, flushing, cleanliness, sampling, and ownership. Validate behavior at minimum, normal, peak, transient, and maintenance conditions. Keep the residual air-cooling load explicit. Include connection standards and temporary cooling for technology replacement. The most dangerous ambiguity is often the interface between the facility operator and the information-technology equipment supplier.

Make data and network readiness a construction workstream

Accelerators are idle without validated fabrics, storage, software images, credentials, telemetry, and representative data. Coordinate diverse fiber entrances, carrier capacity, optical loss budgets, internal cabling, switch configuration, storage ingestion, cybersecurity, time synchronization, observability, and software release management with the physical build. Maintain an end-to-end readiness plan linking utility power, cooled racks, network domains, storage namespaces, orchestration, and user acceptance. This exposes dependencies that a traditional equipment-completion schedule misses.

Commission in layers until the workload proves the factory

Progress through component tests, factory tests, energized distribution, hydraulic cleaning and balancing, controls verification, rack emulation, network validation, platform health, and representative workloads. Test failures that cross domains: loss of a cooling distribution unit, fabric isolation, storage degradation, utility disturbance, generator transition, control-network interruption, leak response, and cluster restart. Record firmware, software, valve, protection, and control configurations at acceptance. Final demonstration should prove sustained workload performance, telemetry, recovery, and operating procedures—not a short synthetic peak alone.

Stage gates from energized infrastructure to productive AI compute
Installed megawatts are an intermediate milestone. The commercial gate is repeatable, observable workload completion at the agreed service level.

Plan the campus as a sequence of technology cohorts

AI platforms may change faster than substations, buildings, and cooling plants. Group deployment into cohorts with known rack, voltage, coolant, fabric, and service assumptions. Protect backbone capacity and routes while allowing halls or modules to diverge. Decide which systems remain common and which can be replaced independently. Maintain a controlled reference design with an exceptions process, so speed comes from reuse while new hardware is admitted through evidence rather than forcing every phase into the first generation's constraints.

From Construction Complete to Productive Compute

GateQuestionRequired proof
Infrastructure readyCan the site safely energize and reject heat?Protection, controls, hydraulic and failure-mode records
Rack readyAre power, cooling and network interfaces within specification?Measured interface conditions and configuration baseline
Cluster readyCan the platform form, isolate and recover its domains?Platform health, fabric, storage and recovery tests
Workload readyCan real jobs run repeatably and observably?Representative workload results and telemetry
Service acceptedCan operations sustain the promised outcome?Procedures, spares, escalation, monitoring and owner sign-off

Operate around failure domains and computational service levels

Create a responsibility model spanning facilities, network, platform, storage, security, and workload operations. Correlate electrical, thermal, hardware, fabric, and job telemetry on a shared timeline. Define isolation, degraded modes, checkpoint policy, spares, break-fix access, liquid-cooling response, and return-to-service gates. Measure job completion, utilization quality, failed-work loss, recovery time, energy per useful unit of computation, and change reliability. Facility availability alone can look healthy while productive compute is impaired.

Early screening checklist

What to verify before advancing this site.

  • The workload model defines productive compute and recovery objectives
  • Logical cluster and fault-domain topology drives the physical layout
  • Power studies include ramps, steps, restart, maintenance, and degraded states
  • Liquid-cooling interfaces assign chemistry, cleanliness, controls, and ownership
  • Residual air load and every heat-rejection mode are quantified
  • Network, storage, software, security, and data readiness are on the master schedule
  • Commissioning progresses from components to representative workloads
  • Cross-domain failures and return-to-service sequences are demonstrated
  • Technology cohorts can change without rebuilding the permanent backbone
  • Operations correlate facility, platform, fabric, storage, and job telemetry

What DCFR would flag

Risks surfaced at the screening stage.

DCFR would flag an AI factory defined primarily by accelerator count or facility megawatts without a workload model, topology-driven layout, liquid-cooling interface standard, dynamic power study, data-and-network readiness path, or productive-compute acceptance test.

Professional confirmation required

Items requiring licensed validation.

Confirm workloads, platform topology, power characteristics, coolant requirements, network and storage performance, software, security, commissioning, service levels, and operating responsibilities with current manufacturers, the owner, utilities, licensed professionals, integrators, contractors, and authorities.

Final takeaway

An AI factory is complete only when infrastructure, platform, data, software, and operations repeatedly convert energy into useful computational work.

Screen up to 20 candidate sites before selecting one for the full DCFR report.

Each DCFR Report Package includes a preliminary 20-site comparison PDF / export package plus one selected planning-grade feasibility report.