Insights

    GPU & Compute

    The bottleneck stopped being GPUs. It is now the grid.

    Devence Lab

    · 2 min read

    Share
    The bottleneck stopped being GPUs. It is now the grid.
    Photograph · Unsplash

    Gartner projects 40% of AI data centres will be power-constrained by 2027. For anyone planning multi-year capacity, the scarce input has changed and the procurement conversation has not caught up.

    In 2024 the scarce input in AI infrastructure was H100 supply. In 2026 it is the grid connection to power them. Gartner projects 40% of AI data centres will be power-constrained by 2027, and approval timelines for new grid capacity in Northern Virginia, Silicon Valley and Northern Europe now run 24 to 36 months.

    The hardware problem has eased measurably over the last eighteen months — neo-clouds and resellers now offer H100, H200 and Blackwell capacity that was unobtainable two years ago. The grid has not moved.

    The arithmetic that makes this acute

    A single 8×H100 SXM5 node draws roughly 10.1 kW under inference load: about 700W per GPU plus overhead from dual CPUs, NVLink switches, 512GB of RAM and power supplies at load. The GPUs themselves are only about 56% of that.

    Scale it and the numbers stop being a facilities concern. A hundred GPUs is around 176 kW, which standard commercial service covers. Five hundred is roughly 880 kW and needs a dedicated transformer and utility coordination. A thousand GPUs — 125 nodes — is about 1.76 MW of continuous draw including cooling at a PUE around 1.4.

    That is a modest cluster by current standards, drawing what a small industrial site draws, continuously, in one location.

    You cannot solve a grid approval backlog with more capital at the same address. The queue is the queue.

    Why this is a different category of problem

    Chip scarcity is a queue you can pay to move up, or wait out, measured in quarters. Grid interconnection is a civil engineering and regulatory process measured in years, with an outcome that money does not guarantee. A GPU shortfall delays a programme; a power shortfall can invalidate a site, and the alternative site starts its own multi-year process.

    The IEA's Energy and AI work projects global data centre electricity consumption could double by 2030, with AI workloads the majority of the incremental demand. That demand is also unusually concentrated — regional grids were planned for diffuse load, not for single sites drawing at industrial scale.

    What changes in planning

    Power availability becomes an input to architecture rather than a facilities footnote. Site selection now precedes capacity design, because where you can get power determines how much you can build. Interconnection timelines belong in the programme plan beside hardware lead times.

    And inference efficiency stops being a cost optimisation. A halving of energy per token is a doubling of what you can serve inside a fixed connection — which, when the connection is the constraint, is the only capacity you can add without a utility.

    For regulated organisations there is a sharper edge. If your data cannot leave a jurisdiction and that jurisdiction's grid is constrained, your sovereignty requirement and your capacity requirement are now competing for the same scarce thing, and neither will yield.

    Sources

    1. Power-Bound, Not GPU-Bound: AI Data Center Power Constraints Are the Real 2026 BottleneckSpheron
    2. AI Data Center Power: Grid Limits Reshape Energy in 2026Enki.AI
    3. AI data center energy in 2026dev/sustainability

    Written by the Devence Lab research team.

    Share

    Collaborate

    We share findings with partners operating in the same constraint space.

    Get in touch