DCFR Insight 100 / Modular AI Infrastructure
Modular AI Infrastructure: From First Deployment to AI Factory with Speed, Scale, and Sovereignty
A deep dive into the physical systems, delivery strategy, and operating controls that turn a first AI deployment into repeatable, productive capacity—with speed, scale, and sovereignty designed in from the start.

From delivered modules to productive AI capacity
Modular artificial intelligence (AI) infrastructure can make a first deployment useful sooner and allow capacity to grow in measured increments. Its value depends on how well the project coordinates factory production, site readiness, computing architecture, and operations. A delivered module creates value only when it can run the intended workload reliably and under the required controls.
For owners and developers, the opportunity is to establish a repeatable unit of productive capacity, then expand it without repeatedly redesigning the entire facility. This requires an expansion plan from the beginning, even when the first installation is small. Sovereignty belongs in that initial brief because decisions about administrators, data flows, and software dependencies can be difficult to reverse after deployment.
The term AI factory describes infrastructure organized to train models, adapt them, and deliver AI services. NVIDIA uses the term for data centers hosting integrated accelerated computing platforms. In this article, it is an operating concept rather than a minimum building size or a certification. A large campus with poorly used processors can deliver less value than a smaller system with dependable demand and disciplined operations.
The illustrations are AI-generated architectural concepts for this article; they do not depict a named facility or establish engineered capacity or certified sovereignty.
What modularity actually includes
Modularity is the ability to assemble, operate, maintain, and expand a system through defined units and interfaces. Prefabrication is the production of components away from their final installation location. The two often overlap, but a modular AI project can use a permanent building, prefabricated infrastructure, and integrated computing racks together.
Vertiv's portfolio illustrates this breadth: it includes enclosed modular data centers, separate power modules, overhead infrastructure, and multiple-module solutions for AI. Its published description also identifies off-site production and testing in parallel with site preparation as a delivery mechanism. Those are supplier descriptions of product scope and potential benefits; they do not establish a universal project schedule or cost saving.
The most useful unit of expansion may differ across systems. A computing cluster, a cooling circuit, and an electrical block do not have to share identical boundaries. They do need a documented relationship so that an additional computing unit does not require an unexpected rebuild of its supporting infrastructure.
Four forms of modular AI infrastructure
| Form of modularity | Repeatable unit | Main planning decision |
|---|---|---|
| Integrated compute | Validated racks, networking, and management | Match the unit to the workload and interconnect |
| Facility infrastructure | Power skid, cooling plant, or distribution assembly | Define connection capacity and isolation boundaries |
| Enclosed deployment | A transportable room or group of joined modules | Resolve transport, foundations, servicing, and site integration |
| Campus development | Repeated halls or operational blocks | Reserve power, routes, and plant space for later phases |
Start with the workload and the first useful capacity
Define what the first deployment must accomplish before choosing its enclosure. An owner brief should identify the intended models, data volume, response-time targets, training or fine-tuning needs, concurrency, availability, and required location of data processing. It should also identify who will use the system and how demand has been established.
Inference runs an already trained model to produce outputs. Some inference services can expand through independent replicas. Other workloads, including very large models or distributed training, require substantial communication between processors. A site can therefore accommodate more equipment without supporting the specific job the customer wants to run. Workload testing should determine the appropriate cluster size and network architecture.
NVIDIA's DGX B200 SuperPOD reference architecture combines compute, networking, management, and storage in scalable units. Its larger configurations change switch and cable counts as the system grows. This is a concrete example of coordinated expansion: a computing unit is more than its graphics processing units (GPUs). Its reference architecture should not be treated as a universal template for every AI deployment.
For a first deployment, DCFR's planning recommendation is to select the smallest complete configuration that can serve a defined production need and demonstrate the intended expansion interfaces. A minimal proof of concept may establish that a model runs. Production acceptance must also establish that the service can be supported, recovered, secured, and measured.

Where deployment speed comes from
The credible schedule advantage comes from overlapping work that would otherwise occur sequentially. Factory assembly can proceed while foundations, utility connections, and site infrastructure are prepared. Repeated interfaces can reduce redesign. Factory tests can reveal defects before delivery. These advantages depend on a sufficiently stable design and a site program that can keep pace with manufacturing.
Track three distinct milestones: equipment ready to ship, installation ready for energization, and the first accepted production workload. A vendor's manufacturing lead time usually describes only part of that sequence. The owner should request a schedule that identifies the start condition, scope exclusions, required site inputs, testing responsibilities, and the party accountable for each interface.
At a planning level, the latest prerequisite constrains the production date. A module cannot compensate for an unavailable electrical connection, incomplete heat rejection plant, missing fiber service, or an unprepared operations team. Integration and acceptance still follow the point at which those prerequisites can work together. This is schedule logic, not a forecast of a particular project's duration.
Release manufacturing against a controlled interface package. It should specify dimensions and weights, lifting points, electrical connections, cooling connections, controls, network handoffs, and the acceptance procedure. Keep a change register that identifies whether each revision affects the factory, the site, or both. Late changes at those boundaries can consume the expected schedule benefit.
Power and cooling define the usable module
State capacity at a clear boundary. Information technology (IT) load includes servers, storage, and networking. Facility demand also includes cooling, distribution losses, and other support loads. Utility service, installed equipment ratings, and sellable IT capacity are different quantities. A feasibility study should identify all three and show the assumed operating condition behind each number.
Rack density changes the physical design. NVIDIA's documented DGX GB rack configuration is approximately 120 kilowatts (kW) and includes liquid cooling manifolds. Its processors use liquid cooling while other components retain air cooling. These values describe the published configuration; they are neither a universal AI rack allowance nor a substitute for the selected equipment's approved requirements.
Specify both the liquid and air heat loads. For the liquid system, confirm the equipment's permitted temperatures, flow, pressure, fluid chemistry, and connection requirements. A coolant distribution unit (CDU) commonly manages the interface between facility cooling and the technology circuit, but the actual topology must follow the selected system. Vertiv's technical overview identifies distribution, plumbing, material compatibility, risk mitigation, and heat rejection as separate design considerations.
The owner should require a complete thermal path from the processor to the final heat sink. Direct liquid cooling does not, by itself, establish low water consumption: that depends on the heat rejection method and operating conditions. Evaluate normal operation, the selected failure condition, maintenance isolation, and recovery after an interruption. Confirm which pumps and controls require power during the transition.
For expansion, reserve electrical ways, pipe connection points, accessible valves, and controls capacity. Specify the surviving capacity when a component is unavailable. A nominal redundancy label cannot explain the impact of a shared upstream connection or the loss of a common control system. Each phase needs a clear statement of what stays online and what must stop.
- NVIDIA: DGX GB Rack Scale Systems User Guide — Hardware
- Vertiv: Liquid cooling options for data centers

Plan the site for the next construction phase
A modular campus still needs a coherent site plan. Establish the secure operating boundary, a practical service route, delivery staging, and space for equipment replacement. Separate routine staff arrival from heavy deliveries where feasible. Check vehicle turning and backing movements against the actual delivery vehicle, gate position, and unloading method.
Reserve continuous utility corridors and identify future tie-in locations before the first phase occupies the site. Future construction should have a workable approach route and staging area that do not depend on using an operating fire route or dismantling permanent plant. Protect expansion land from becoming the default location for parking, drainage infrastructure, or temporary equipment that later proves difficult to move.
Foundations, drainage, structural loading, external equipment, noise, and fire access remain site-specific design questions. Confirm applicable requirements with the authority having jurisdiction (AHJ) and the responsible design professionals. A factory-built assembly establishes a manufacturing method; the installed project still needs its own site and approval strategy.
Scale the network and storage with the compute
Physical proximity does not establish sufficient network performance. Training jobs that communicate across many processors need an appropriate fabric, bandwidth, and congestion behavior. Inference replicas can sometimes tolerate a different topology. The expansion brief should identify which jobs span modules, which remain within a module, and how the scheduler places them.
NVIDIA's B200 reference design distinguishes compute, storage, in-band management, and out-of-band management networks. It also recommends accommodating the full scalable-unit fabric when initially deploying fewer nodes. That recommendation demonstrates an important tradeoff: some infrastructure may need to be installed ahead of immediate compute demand to preserve the intended network behavior.
Storage deserves its own expansion model. NVIDIA's storage guidance explains that training can be delayed by checkpoint writes and that required performance depends on the workload and dataset. Adding GPUs without sufficient data delivery or checkpoint performance can increase equipment count while leaving a bottleneck unchanged.
DCFR recommends reserving pathways, ports, fiber capacity, and equipment space against the intended final topology, while staging purchases where that can be done without costly rework. Test representative jobs across module boundaries. Evaluate the agreed throughput and response time under contention, during checkpoint activity, and during a planned maintenance condition.
A phased path from 1 MW to 16 MW
The following scenario illustrates planning decisions. Each building block provides 1 megawatt (MW) of IT capacity. The 1 MW, 4 MW, and 16 MW milestones are hypothetical, not vendor offerings, prescribed cluster sizes, or a forecast. At every stage, a separate engineering design must determine the facility demand and the number of processors that can actually be supported.
For a simple load illustration, assume facility support demand equals 25% of IT demand at the specified operating point. The corresponding total demands would be 1.25 MW, 5 MW, and 20 MW. These are arithmetic examples, not equipment sizing: they exclude the further analysis of peaks, reserve capacity, charging, power quality, and redundancy needed for a utility or electrical design.
The first phase should demonstrate interfaces that later phases will use. At 4 MW, attention shifts to shared services and whether one block can be maintained without unacceptable impact elsewhere. At 16 MW, operating procedures, commissioning consistency, and shared failure points deserve as much scrutiny as the repeated physical modules.
Expansion should follow evidence of demand and available infrastructure. A nominal capacity target does not establish a viable business. Compare measured use and customer commitments with the lead time for the next increment. Record which enabling works are needed early, which can wait, and the cost of deferring each decision.

A hypothetical path from first deployment to AI factory
| Stage | Operating objective | Evidence before further expansion |
|---|---|---|
| 1 MW IT | Run a defined production service and validate the first complete block | Accepted workload, stable cooling, recovery tests, measured use |
| 4 MW IT | Repeat the block while operating a shared platform | Verified inter-module performance, isolation, staffing, and demand |
| 16 MW IT | Operate several blocks with coordinated capacity and maintenance | Contracted power path, resilient shared services, repeatable acceptance |
Hypothetical IT-capacity milestones, not prescribed cluster sizes or an engineering design. The separate 25% facility-support assumption illustrates total demand only.
The operating platform turns capacity into an AI factory
A usable AI service requires provisioning, workload scheduling, identity, monitoring, model and data management, and recovery. NVIDIA's SuperPOD software documentation includes configuration, validation, administration, and GPU orchestration alongside its hardware architecture. This shows why the software operating layer must be included in the project scope from the outset.
Define ownership across the facility operator, hardware supplier, platform team, and customer. Specify who responds when a job slows down, who determines whether the cause is thermal or computational, and who can authorize a disruptive change. A unified incident process is useful even when several companies deliver the underlying systems.
Measure the outcomes customers buy. For inference, these may include response latency, throughput at an agreed model quality, and successful requests. For training, they may include time to complete an agreed workload and the fraction of time lost to interruptions. Hardware utilization alone can conceal failed jobs or unproductive work. Tie operating metrics to a stable workload definition so comparisons remain meaningful.
Sovereignty requires explicit control
For a modular AI project, sovereignty should be written as a set of required authorities, dependencies, and demonstrable controls. Data residency identifies where information is stored or processed. It does not fully describe who can administer the service, access encryption keys, change software, suspend an account, or recover the system.
A disconnected operating model is technically distinct from an ordinary connected private cloud. Google documents an air-gapped Distributed Cloud offering that can operate without connectivity to Google Cloud. Its deployment guidance also describes an offline software transfer process and a site survey. This illustrates both the possibility of disconnected operation and the continuing need to manage software supply and physical readiness.
Security also cannot be inferred from ownership or location. The National Institute of Standards and Technology (NIST) states in its zero trust architecture guidance that trust should not be granted solely because of an asset's physical or network location. Apply that principle to administrators, maintenance systems, and machine identities as well as end users.
DCFR recommends the following procurement assessment. It is a planning framework for comparing control requirements, not a legal certification or a universal definition of sovereign AI.
Separate national capability objectives from a particular organization's control requirements. A national program may seek domestic skills, local-language models, and a resilient supplier ecosystem. An enterprise may prioritize control over sensitive workloads and privileged access. Neither objective necessarily requires every component to be manufactured domestically, and neither is established by locating an enclosure inside a national border.
The appropriate degree of isolation depends on the workload and governing obligations. Greater independence can add staffing, update, integration, and recovery responsibilities. Document the remaining dependencies and the consequences if one becomes unavailable. Test the chosen operating model for an agreed duration and identify the services that degrade during the test.

Sovereignty: evidence to request
| Control area | Evidence to request |
|---|---|
| Data and metadata | Locations and permitted flows for datasets, prompts, outputs, logs, backups, and support records |
| Administrative authority | Named operator responsibilities, privileged-access rules, approvals, and auditable sessions |
| Keys and identity | Key custody, identity dependencies, recovery authority, and access revocation tests |
| Software and models | Update path, artifact verification, model rights, license dependencies, and rollback capability |
| Operational continuity | Services available during external disconnection, local monitoring, recovery procedures, and support limits |
| Supply and exit | Critical spares, replacement lead times, data export, model portability, and migration obligations |
DCFR procurement assessment. The required controls depend on the workload, operating model, and governing obligations.
Evaluate the economics over productive service
Compare alternatives over the same workload, availability target, ownership period, and scope. Include equipment, civil works, power, cooling, network, software, staffing, maintenance, commissioning, financing, and eventual refresh or removal. A low module purchase price can omit substantial site and operating costs.
For a consistent workload and hardware class, one useful measure divides total cost over a period by accepted productive GPU-hours in that period. Define productive hours carefully and avoid applying the metric across unlike models or accelerators without normalization. Inference services may be better compared by the cost of successful requests at a specified quality and latency.
Modularity can reduce exposure to unused capacity when purchases follow demand, but repeated mobilizations, duplicated plant, and small operating blocks can offset that benefit. Test slower demand, later power availability, lower service prices, and a hardware refresh before the facility reaches its planned capacity. Treat these as scenarios with stated assumptions, not predicted outcomes.
Accept the first deployment as a complete system
DCFR's recommended acceptance sequence is to verify factory results, inspect the installed systems, test their integration, and demonstrate the intended workload. Use approved commissioning procedures and qualified teams for fault simulations. The owner should receive evidence that the facility and the computing service meet the agreed conditions together.
1. Confirm the configuration
Confirm the installed configuration, connection ratings, firmware and software versions, and responsibility for open issues.
2. Verify power and thermal performance
Verify power, liquid cooling, residual air cooling, controls, and alarms under the agreed operating loads.
3. Demonstrate the workload
Demonstrate representative workloads, data movement, checkpoint recovery, and the specified performance criteria.
4. Test control and recovery
Test access controls, logging, backup restoration, and the agreed limits of disconnected operation.
5. Rehearse expansion and maintenance
Rehearse the next expansion tie-in and a maintenance event, including isolation boundaries and rollback procedures.
The owner decision should establish the first useful capacity, the infrastructure reserved for later phases, and the evidence required before additional investment. A credible modular AI strategy makes these commitments explicit. Its speed comes from coordinated delivery, its scale from compatible interfaces, and its sovereignty from controls that continue to work in operation.
Early screening checklist
What to verify before advancing this site.
- Define the first production workload and its acceptance criteria.
- Confirm the power date and complete liquid and air cooling duties.
- Reserve service routes, utility corridors, and future connection points.
- Test network, storage, recovery, and performance across modules.
- Document administrative authority, data flows, key custody, and external dependencies.
- Advance the next phase against evidenced demand and available infrastructure.
What DCFR would flag
Risks surfaced at the screening stage.
A factory delivery date presented as the production date; IT capacity confused with total demand; expansion land or utility corridors left unprotected; shared services that undermine isolation; or sovereignty claimed solely from the facility’s location.
Professional confirmation required
Items requiring licensed validation.
The responsible design, commissioning, IT, security, and operating teams should confirm the installed system, site requirements, approved equipment limits, failure conditions, and the controls required for the chosen workloads.
Final takeaway
Speed comes from coordinated delivery. Scale comes from compatible interfaces. Sovereignty comes from controls that continue to work in operation.
Screen up to 20 candidate sites before selecting one for the full DCFR report.
Each DCFR Report Package includes a preliminary 20-site comparison PDF / export package plus one selected planning-grade feasibility report.