by datastudy.nl

Sunday, September 13, 2026

Engineering

AI data centers destabilized the PJM grid. It is architecture

AI data center power grid stability is an architecture problem. A 2026 fault dropped 3.49 GW on PJM in seconds, twice the 2024 event. NERC has issued its top alert.

Bar chart showing two PJM grid load-drop events caused by AI data center disconnections: 1.5 gigawatts in 2024 and 3.49 gigawatts in 2026. The 2026 event was more than twice the size.
AI data center load-drop events on the PJM grid: 1.5 GW in 2024 vs 3.49 GW in 2026. Source: PJM, TechCrunch reporting.

On July 22, 2026, a single transmission line fault outside Washington, DC, knocked more than 3.1 gigawatts of data center load off the PJM grid in roughly 30 seconds. Lights flickered from Northern Virginia to Chicago. No blackout, but the voltage spike was visible across the largest electricity market in the United States. The cause was not a generation shortfall. It was a protection relay doing exactly what it was designed to do: disconnecting on the third voltage dip, as thousands of GPUs shed load in unison. This is the AI data center power grid stability problem, and it is getting worse fast. The 2026 event was more than twice the size of a near-identical incident two years earlier, when 60 facilities dropped 1.5 GW simultaneously. The North American Electric Reliability Corporation has since issued its highest urgency alert, and grid operators are now writing rules that will reshape how every AI campus connects to the grid.

What actually happened on the PJM grid?

The sequence is simple and frightening. A transmission line failed near Ashburn, Virginia, the epicenter of the world's largest data center cluster. The fault caused a voltage dip. Data centers in the region sensed the dip and switched to backup power, removing their load from the grid. As more facilities made the same split-second decision, the demand drop compounded. What started as a modest supply disruption became a 3.49 GW surplus on the grid in under a minute. PJM, which serves 67 million customers from New Jersey to Illinois, needed 11 minutes to stabilize.

The event echoed a July 2024 incident on the same grid, when a failed surge arrester triggered roughly 60 data centers to disconnect at once, pulling 1.5 GW offline. At the time, data centers accounted for about 6 percent of PJM's total load, according to Synapse Energy Economics. By 2040, they are projected to make up 24 percent. The load-drop event doubled in two years. The data center share of grid demand is on track to quadruple.

AI data center load drop events on the PJM grid: 1.5 gigawatts in 2024 versus 3.49 gigawatts in 2026. Data center share of PJM load was 6 percent in 2024 and is projected to reach 24 percent by 2040.
PJM grid load-drop events from data center disconnections, 2024 and 2026, with projected data center share of PJM load by 2040. Source: PJM, Synapse Energy Economics. Data Today analysis.

The chart above shows the two events side by side: 1.5 GW in 2024 versus 3.49 GW in 2026, alongside the projected growth in data center share of PJM load from 6 percent today to 24 percent by 2040. The trend line is the story. If the disconnection behavior does not change, the next event will be larger still.

Ali Zain Banatwala, a senior market models specialist at the Independent Electricity System Operator, told TechCrunch that when the voltage dip reached the data centers, they all decided to disconnect within a few seconds of each other. Each facility acted rationally to protect its compute. Together, at gigawatt scale, they created a grid disturbance that no one had planned for. The protection logic, written when a large load meant 50 megawatts, cannot see the grid it is now part of. So when trouble arrives upstream, it does the wrong thing: it drops out.

Why is this an architecture problem, not a supply problem?

Most of the AI power debate centers on generation: more gas turbines, more solar, more nuclear, more transmission lines. The grid needs more electrons, and that conversation is real. But the 2024 and 2026 Virginia events were not supply failures. No power plant tripped. No fuel ran short. The grid had plenty of electricity. The problem was that 3 percent of total PJM demand vanished in seconds, and the grid's frequency controls could not keep up.

The standard data center power stack has not changed in decades. Medium-voltage power arrives, transformers step it down, low-voltage UPS units condition it, and it reaches the racks. Three things break at AI scale.

First, the UPS sits deep inside the building, close to the servers. Its batteries are sized for a few minutes of backup during an outage, not for absorbing load swings that happen hundreds of times per second during training runs. An AI campus can swing 70 percent of its load in milliseconds when thousands of accelerators synchronize at job start, hit a checkpoint, or recover from a fault.

Second, the UPS spends most of its life in bypass mode. Legacy converters waste enough power that operators run them in eco-mode, feeding racks directly from the grid with no filtering. The compute's load swings go out raw, and grid transients come in unfiltered.

Third, the protection logic was written for a world of 50 MW industrial loads. When it counts voltage dips and disconnects on the third one, it removes load at the exact moment the grid needs it to stay connected. The equipment does what it was told. The load has simply outgrown the instructions.

A recent paper from researchers at multiple institutions, published on arXiv, frames this as a co-design problem. AI workloads exhibit load ramps of tens to hundreds of megawatts per second when thousands of accelerators synchronize. These swings propagate facility-wide because training jobs stay in lockstep: PDUs, substations, and the interconnection point all see the same step change. The paper notes that this behavior bypasses the grid's natural averaging and outpaces primary frequency controls, which operate on minute-scale timescales. Texas RE's director of reliability services has publicly compared AI data center load profiles to those of steel mills, with very fast and very large ramps.

What is NERC doing about it?

In 2026, the North American Electric Reliability Corporation released a Level 3 Essential Action Alert, its highest urgency notice, with a 3 August response deadline for utilities and registered entities. The alert included seven required actions and warned of customer-initiated large load reductions and significant oscillations that occur in seconds, leaving little or no room for real-time responses.

The alert is an early step toward tougher interconnection and performance rules for large data centers, similar to the standards that developed around renewable energy projects. Industry experts expect more standardized or mandatory packages for dynamic studies to begin early in the data center design process. ERCOT, the Texas grid operator, is already moving to require large loads like data centers to ride through disruptions rather than disconnecting.

These rules are not theoretical. They are direct responses to documented events. The 2024 Virginia incident is, according to the arXiv paper, the first regulator-documented instance of this failure mode at grid scale. The 2026 event is the second, and it was twice as large.

How does this change what AI builders need to plan for?

If you are building or operating AI infrastructure, the regulatory floor is shifting under you. Here is what that means in concrete terms.

  • Interconnection studies will get harder. Grid operators will require dynamic load modeling before approving new connections. Expect longer timelines and more stringent ride-through requirements. The era of connecting a data center as a dumb load is ending.
  • Protection settings need review. If your facility disconnects on the third voltage dip, you are running 1990s logic against 2020s load. That setting is now a grid liability. Auditing and updating protection schemes is a near-term must.
  • UPS architecture is the leverage point. The sponsored ON.energy piece in MIT Technology Review describes moving the UPS from 480 volts inside the building to medium voltage outside, in the path of all power flow. ON.energy says it is installing 3 GW of such systems across four campuses. Whether or not this specific architecture wins, the direction is clear: inline, medium-voltage power conditioning that absorbs swings before they reach the grid.
  • Tax credits and grid revenue change the math. Equipment that runs at medium voltage, sits outside, and stores energy can qualify for tax credits and earn revenue in grid programs like peak shaving and demand response. Backup power stops being pure insurance and starts contributing to the business case.
  • Site selection gets more complex. Regions with strong grid oversight like PJM and ERCOT will enforce ride-through rules first. Building in a loosely regulated market may buy time, but it also means higher risk of a forced retrofit later.

The Georgetown Law Technology Review makes a related argument: the very features that make AI data centers challenging, namely digital control, telemetry, and fast response, also make them candidates for improving grid reliability through coordinated flexibility and delay-tolerant workloads. The bulk power system will need to evolve from a system that manages supply to one that manages demand, and AI campuses are the largest, fastest-moving demand on it.

What should you do before the next fault?

The practical near-term tool that most experts converge on is battery energy storage. Storage can sit between a large facility and the wider grid, absorbing abrupt load swings instead of passing them through. With transmission upgrades often taking years, deployable options are attracting the most attention. Thomas Sisto, co-founder and CEO of flow battery company XL Batteries, put it bluntly: there is not one piece of equipment that fixes it. The answer will be a layered set of tools.

IONATE, a transformer intelligence company, says its technology can respond in under a millisecond, while batteries support stability over longer stretches. The combination of fast-acting power electronics and medium-duration storage is the emerging consensus architecture.

For teams planning new AI campuses, the bets to make now are straightforward. Design power conditioning at medium voltage from day one. Size battery systems for load-swing absorption, not just outage ride-through. Model your facility's dynamic load profile and share it with the utility early. And do not assume that legacy protection logic will keep you compliant as ride-through rules land.

The bets to avoid are equally clear. Do not treat grid interconnection as a paperwork exercise. Do not assume that because your facility stayed online in 2024 or 2026, it will stay online next time. The events are escalating, regulators are responding, and the facilities that cannot ride through a voltage dip will eventually be told to fix it or disconnect permanently.

The grid does not care about your training run

The core tension is this: every AI data center is engineered to protect its compute above all else. That instinct is correct for a single facility and dangerous for a grid full of them. When thousands of GPUs disconnect at the first sign of trouble, they do not protect the system they depend on. They destabilize it. The fix is not more generation. The fix is architecture that lets a data center absorb a grid fault without dropping load, and lets a grid absorb a data center without flinching. The technology exists. The rules are arriving. The only question is whether your facility is built for the grid you are connecting to, or the one that existed twenty years ago.

Sources