NVIDIA has published an account of a demand-response experiment carried out in California, where grid operator Silicon Valley Power signaled an AI data center to temporarily reduce its power draw during a summer evening spike, as air conditioning loads surged across the region. The company frames the episode as an example of an approach it wants to scale across its AI infrastructure, which it refers to as 'AI factories.'
The capability relies in part on software built by Emerald AI, a startup founded by Varun Sivaram that specializes in demand-response systems tailored to AI training and inference workloads. The underlying concept treats available compute capacity as a flexible resource: rather than running continuously at full power, data centers can throttle non-critical tasks during peak grid stress and resume them once demand eases.
For NVIDIA, coordinating power consumption with token generation addresses a growing concern among US grid operators facing surging electricity demand from AI data centers. Several states have already raised alarms about whether local grids can absorb this growth without straining reliability or driving up costs for other electricity users.
The company positions this flexibility as a way to reassure both regulators and utilities that AI infrastructure operators can act as active partners in grid management rather than rigid, always-on consumers. It remains to be seen how widely this kind of demand-responsive operation will be adopted as pressure on US energy availability for AI continues to mount.