Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS : US Pioneer Global VC DIFCHQ SFO NYC Singapore – Riyadh Swiss Our Mind

AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver. For AI inference workloads, this makes application-level performance per watt the key metric for measuring AI factory efficiency.

Not every megawatt translates to revenue-generating compute. Power distribution, cooling, networking, storage, backup, and facility overhead take a share of the power before it reaches a GPU. Static rack provisioning exacerbates this: an outdated approach to data center design allocates the maximum power draw per rack to meet worst-case peak demand, even though real workloads have different power needs and may leave some portion of that maximum power unused. Operators reserve additional capacity for failures, operational flexibility, and expansion.

In one representative power-budget view examined by NVIDIA, about 60% of delivered site power is allocated to compute for AI output.

NVIDIA DSX MaxLPS is a suite of chip, thermal, system, and software technologies that maximizes AI factory throughput within a fixed power budget. MaxLPS stands for Maximum Land Power Shell, the site-level constraints that define an AI factory: land, utility power, and the physical shell holding power, cooling, networking, and compute infrastructure.

MaxLPS is designed to optimize these three layers:

  • Dynamic power allocation: Continuously monitors and allocates unused power headroom to GPUs
  • Advanced performance per watt techniques: Software power optimization techniques that improve job-level performance at a fixed power budget
  • 45° C thermal efficiency and site design: Cuts cooling overhead through warm-water liquid cooling, improving power usage effectiveness (PUE) to convert directly into more compute within the same fixed envelope

Why static rack provisioning strands power

Traditional data center power planning reserves enough power as though every rack could draw its specified maximum power simultaneously. That protects the facility against peak demand, but it treats each rack as an isolated power island. An isolated rack provisioned with excess power cannot lend that unused power to a neighbor that could turn it into tokens.

At AI factory scale, facility overhead, rack losses, and operational inefficiencies during failures, restarts, and checkpointing reduce the power available to the AI load. Separately, static rack provisioning can strand headroom within the power allocated to racks. Power reserved for one rack’s peak demand may sit unused while another rack could use it. Dynamic power allocation targets this reclaimable rack-level headroom.

Figure 1 shows a power-budget waterfall for a 100 MW AI factory. Of the 100 MW grid input, 20 MW is allocated to facility overhead, 10 MW to rack losses, and 10 MW is unavailable for AI load because of operational inefficiency during failures, restarts, and checkpointing. This leaves 60 MW available for AI load. Each deduction is expressed as a share of the original grid input.

Training, post-training, and inference move through compute bursts, memory-bound execution, synchronization, checkpointing, prefill, decode, idle gaps, and network-bound communication. Each phase draws power differently. For stability, a rack must be provisioned to handle the workload’s peak power draw, though actual energy consumption remains below that level for meaningful periods.

How does NVIDIA DSX MaxLPS dynamic power allocation work? 

NVIDIA DSX MaxLPS replaces static power provisioning with dynamic power allocation, using Dynamic Power Software (DPS) (currently in Developer Preview) to monitor and optimize power utilization within the data center’s fixed power budget. DPS is a comprehensive power management system that models the data center topology from the utility level down to racks, nodes, and GPUs. Operators define resource groups, power budgets, and policies that govern how power can be allocated, constrained, and enforced.

Within those boundaries, DPS continuously compares allocated power against actual consumption. When GPUs or racks operate below their reserved level, DPS makes that headroom available to others in the same managed group. The site power envelope remains unchanged: DPS extracts more productivity from the available power.

DPS runs a continuous control loop, as shown in Figure 2. The system collects GPU-, rack-, and group-level power telemetry, identifies unused power capacity, reallocates power within policy, validates compliance with the approved group power budget, and responds to power events or emergency policies on a best-effort basis.

The value comes from coordinating many power decisions across the fleet rather than treating every rack as a one-off configuration. When the site budget changes, or when a grid, maintenance, or emergency event alters available power, DPS adapts operating limits across the data center without forcing operators to re-plan every rack by hand.

Figure 3 illustrates these efficiency gains by comparing the power utilization of static provisioning with that of MaxLPS dynamic provisioning. Static provisioning strands 170 kW of power, while MaxLPS dynamic provisioning reclaims this headroom, enabling an additional rack deployment within the same 540 kW site power budget.

DSX Exchange is an open source event bus for AI factory operations (currently in Developer Preview). It adds an IT/OT event bus connecting DPS and other DSX services to building management systems, electrical power monitoring systems, cooling infrastructure, grid interfaces, and compute schedulers. MaxLPS does not require DSX Exchange to function, but the integration exposes signals such as seasonal cooling headroom and facility power events that DPS can act on.

MaxLPS techniques for advanced performance per watt 

Beyond rack-level power steering, MaxLPS also includes software features that optimize how each GPU uses the power it already has.

DSX MaxLPS includes optimized workload profile power solutions (WPPS) for common data center operating modes, including inference, training, memory-bound, and compute-bound configurations. Rather than tuning power, memory, frequency, and other node behavior for every job, operators apply validated profiles that align compute behavior with the workload.

The Application Performance and Power Manager (APPM) applies the selected configuration to participating GPUs, while software such as NVIDIA Dynamo can further optimize inter-rack performance and power behavior for inference services. The principle is AI factory optimization: align GPU configuration, application behavior, and serving topology to increase fleet-wide performance and output per watt.

Inference provides a clear example because maximizing tokens per second per watt translates the gain into measured workload throughput. The same profiles apply to training and post-training.

For NVIDIA Vera Rubin NVL72 AI factories, NVIDIA projects that MaxLPS, combined with data center power planning, can enable up to 40% more Rubin GPU capacity within the same power budget. Together with measured results on NVIDIA GB200 NVL72, these results demonstrate how dynamic power management can increase AI factory productivity per megawatt.

Figure 4 shows MaxLPS results evaluated using representative inference workloads: Vera Rubin NVL72 was tested with DeepSeek-R1, while GB200 NVL72 was tested with Kimi-K2.5. MaxLPS reduces provisioned rack power from 125 kW to 90 kW on GB200 NVL72 and from 136 kW to 101 kW on Vera Rubin NVL72, enabling 39% and 35% more racks, respectively, within the same power envelope while preserving workload throughput. Performance per watt improves approximately 1.5x on GB200 NVL72 and 1.3–1.4x on Vera Rubin NVL72.

How to size an AI factory site for MaxLPS

MaxLPS infrastructure design starts with a fixed gross facility power envelope and works inward to determine the maximum number of GPU rack positions the site can support at the MaxLPS operating point. It also establishes the long-term infrastructure target for power, cooling, space, and network capacity.

Day 1 deployment can be lower than this infrastructure limit. A site that begins with a training- or post-training-heavy workload may operate at an average rack power above the MaxLPS average operating point. Because each populated rack consumes more power, fewer racks can initially be deployed within the fixed facility envelope.

MaxLPS therefore defines the lifecycle capacity target, not a requirement to populate every rack from the start. As the workload mix shifts toward inference over the hardware lifecycle, average rack power can decline toward the MaxLPS average operating point, creating headroom to populate additional rack positions without increasing the facility power envelope.

As Figure 5 shows, MaxLPS sizing anticipates this evolution from the start by selecting the GPU product family, setting the site PUE target at 45° C DLC inlet operation, sizing the east-west network for the MaxLPS GPU count, and deriving the total deployable GPU rack-position count.

The expected day-one workload mix then determines how many of those positions are initially populated. With the associated space, power distribution, cooling, and network capacity configured up front for the full MaxLPS capacity target, operators can add GPUs and racks incrementally as workloads evolve without a later facility retrofit.

Why is 45° C liquid cooling important?

Software power steering at the data center level is the largest part of the MaxLPS story, but it is not the whole system. MaxLPS also depends on chip-, thermal-, and system-level advances that allow more of the facility power budget to reach the compute infrastructure.

Vera Rubin NVL72 racks are designed for 45° C liquid-cooling inlet operation. Warmer liquid may sound counterintuitive, but the goal is to remove heat efficiently while meeting performance, reliability, and lifetime requirements. Higher coolant temperatures can enable facilities to rely more often on “free cooling,” which uses outside air or water to reject heat with minimal mechanical chilling.

Depending on the climate and site design, this method can reduce reliance on energy-intensive chillers or water-intensive adiabatic (evaporative) coolers. Chillers remain important during hot conditions and for resilience, but using them only when needed reduces overall cooling power and average annual power usage effectiveness (PUE).

Recovering that cooling power for compute is a facilities design problem as much as an operating one. Heat rejection systems are sized for worst-case scenarios, yet many sites require far less chiller compressor power under typical operating conditions, especially when dry coolers can handle a larger portion of the load. If the electrical distribution, dry coolers, and controls are sized for that flexibility, power budgeted for cooling during the worst hour becomes available to compute during much of the year.

Operating the technology cooling system loop efficiently

The technology cooling system (TCS) loop is the technology-side liquid loop that moves heat between racks, cold plates, and the coolant distribution unit (CDU). It works alongside the broader facility cooling loop, which rejects heat through chillers, dry coolers, cooling towers, or other site equipment. A well-operated TCS loop balances three constraints:

  • Adjusting the coolant flow rate to match GPU thermal demand, allowing the facility loop to absorb variation on the heat rejection side
  • Staying within the designed inlet-temperature and thermal-reliability limits
  • Responding to workload-driven thermal events without creating instability

Many facilities use proportional-integral-derivative (PID) control to manage CDU behavior. PID control is common and useful, but it is reactive—it responds only after sensors report a deviation. Large AI factories create fast, synchronized changes in thermal load, so operators often run the loop colder than necessary to preserve margin.

A conservative PID-controlled buffer, as described, wastes cooling power that could otherwise support compute. Agentic control systems, including work by Phaidra, use power telemetry and learned control policies to anticipate thermal behavior. A rack with 45 °C thermal efficiency does not require agentic control to operate: the hardware and facility design define the supported thermal envelope, and agentic control is an optimization on top of that.

Figure 6 shows the technology cooling loop versus the facility cooling loop. The TCS loop carries heat from GPU cold plates to the coolant distribution unit (CDU), which then transfers it to the facility cooling loop. Facility pumps, chillers, dry coolers, cooling towers, or other site equipment then reject the heat outside the building. Optional monitoring and predictive controls can help optimize efficiency.

Get started with NVIDIA DSX MaxLPS validation

Physical infrastructure is difficult to change once built. Whether upgrading an existing facility or planning NVIDIA Vera Rubin NVL72 sites, operators should retrofit or design for NVIDIA DSX MaxLPS at the site level, even if they do not deploy every rack on Day 1. That means validating power topology, redundancy, cooling behavior, telemetry availability, policy requirements, network capacity, rack-position optionality, and space or utility constraints before infrastructure decisions are fixed.

For teams evaluating the software path, the NVIDIA Dynamic Power Software documentation and NVL72 inference power pilot provide a starting point. A scoped validation compares an unmanaged static baseline against a MaxLPS-managed run using representative workloads, tracking throughput, latency, service error rate, power draw, utilization, and policy compliance. The results show operators when and how fast to scale as they prepare for Vera Rubin NVL72.

Facilities teams should engage NVIDIA early to validate site thermal design for 45° C inlet operation, MaxLPS rack-position optionality, and power distribution flexibility before infrastructure is fixed. The site-level design decision should be made early and does not depend on software validation. Software testing can proceed in parallel or follow at a later stage.

To turn fixed power into more AI output, power-limited AI factories need a new operating model. Static rack provisioning strands power that real workloads would otherwise turn into useful AI output.

NVIDIA DSX MaxLPS addresses this on three fronts: dynamic power allocation software that unlocks stranded rack capacity, performance-per-watt techniques that raise output per watt, and 45° C thermal and facility design that reduces cooling overhead while preserving the physical optionality to land more compute. For operators planning Vera Rubin NVL72 deployments, MaxLPS is a path to prepare sites for up to 40% more Rubin GPU capacity within fixed power budgets.

Prepare Vera Rubin NVL72 sites to provision up to 40% more Rubin GPU capacity within your fixed power budget. To get started, read the MaxLPS Overview, explore the NVIDIA Dynamic Power Software documentation, and the NVL72 inference power pilot. Engage NVIDIA early to validate the 45° C thermal design, rack-position optionality, and site-level MaxLPS readiness.

Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS