Research Radar/Cloud computing/China · Canada
The AI setting your power company should love
One default setting in the software that serves chatbots does not lower peak power, but it cuts how fast that power swings by up to 42.6 percent. A grid would need about a fifth less emergency reserve.
What happened
Every time you send a question to a chatbot, a graphics chip in a data center wakes up, works hard, and rests. Millions of questions make the chip’s power swing up and down all day. A power grid can handle a steady load. What it hates is a load that jumps. A jump has to be matched, second by second, by power plants that stand ready for it. Grid operators call this “regulation reserve”, and they pay for it.
Four researchers measured a chatbot server’s power fifty times a second and found something surprising in a setting most engineers never think about.
The test
Pan Li of Tongji University in Shanghai, Yize Chen of the University of Alberta in Canada, and Xia Miao and Dai Wang of two Chinese energy technology companies ran a 7-billion-parameter language model on one graphics card, the kind gamers buy, using vLLM, a popular serving program. They sent it a realistic mix of requests: mostly short questions, with about 15 percent very long ones, such as a pasted document. Then they changed one setting.
The setting is called “chunked prefill”. When a long document arrives, the server can read it in one big gulp, or in small bites between other people’s requests. The small-bite option exists to keep other users from waiting. The researchers asked whether it also changes the shape of the power.
The result
The peak power did not change. The chip hit the same ceiling, about 465 to 470 watts, whichever way it read the document. What changed was how fast the power climbed and fell. With the smallest bites, the average rate of change fell from about 46 watts per second to about 30, a drop of roughly 35 percent. The busier the server, the bigger the effect: 7 percent at light load, 34.6 percent at heavy load, and up to 42.6 percent when many long documents arrived at once. The total energy used stayed the same, within two percent.
Why does a smoother climb matter if the peak is the same? This is because the grid pays for speed, not just for height. Using their real measurements, the researchers calculated how much fast-response reserve a grid operator would need for a fleet of such servers. With the small-bite setting, the answer was 20.3 to 22.7 percent less.
Think of a car. This setting does not lower the car’s top speed. It stops the driver from stamping on and off the accelerator. Therefore, the engine works just as hard, but the ride is smooth, and smooth is what the grid can plan for.
What it means
Data centers are the fastest-growing new load on grids everywhere. In Kenya, grid-connected capacity is about 3,192 megawatts, and peak demand reached 2,444 megawatts in January 2026. A planned 1,000-megawatt data center stalled in 2026 because the country could not power it. It may come back at 100 megawatts. On a small grid, a load that swings is a bigger problem than a load that is large. This paper says one free setting, already built into common software, makes AI loads easier to live with. SMOOTH BEATS SMALL.
Business ideas from this paper
- 1. Grid-friendly AI hosting for small grids
Certify data centers that run AI in a way the grid can plan for, and help them prove it with measurements like these. Small grids in Africa and Central Europe need this first.
Buyer Data-center operators and the utilities that connect them. First test Ten conversations with operators and grid planners, asking what they would pay to shorten a grid-connection approval.
- 2. A power-aware add-on for serving software
The paper shows one knob. A product would watch the grid’s stress signal and turn several knobs at once, trading a little latency for a lot of smoothness when the grid asks.
Buyer Cloud providers and companies running their own AI servers. First test An open-source plugin plus a waitlist for paid support.
- 3. Sell the smoothness back to the grid
Grid operators pay for reserve. A broker that groups AI servers and offers their controllable smoothness as a service could earn a share of what the grid saves.
Buyer Grid operators and large AI hosts. First test Letters of intent from one operator and three hosts.
How sure can you be?
This is a preprint, not yet reviewed by other scientists. The measurements are real, taken with the chip’s clock locked so nothing else changed, and repeated across ten trials. However, they were made on one consumer graphics card with one model. A follow-up in the paper’s appendix tests bigger models on two cards and sees the same pattern, but a full data-center fleet has not been measured. Read the direction as solid and the exact percentages as specific to this setup.
Do this today
If you run AI servers, check whether chunked prefill is on. If you only use AI, ask your provider one question: does your data center have a plan for the grid it sits on?
Source: Smoothing the Ramp, Not the Peak: Scheduling-Induced Power Dynamics of LLM Inference and Their Grid-Scale Consequences, arXiv preprint, August 2026. arXiv:2608.01250 · arxiv.org (preprint · not yet peer reviewed).
Just Out Tech explains new research in plain language. This article was drafted with AI assistance and checked by a human against the original source.
All measurements come from one bench setup, Qwen2.5-Coder-7B-Instruct served by vLLM on a single NVIDIA GeForce RTX 4090. The 10,000-server fleet and the 20.3-22.7% reserve figure are a bootstrap resampling of those single-GPU traces.
- When it reaches you
- A utility would have to trust this enough to buy less fast-ramping reserve. Our estimate is several years, because the authors measured a single consumer-grade card and say the magnitudes are specific to that hardware class.
- Who is building on it
- Two of the four authors hold company affiliations: Xia Miao at Fova Technology (Suzhou) Co., Ltd, Suzhou, and Dai Wang at EcoFlare Co., Ltd, Wuxi. The other two are at Tongji University and the University of Alberta. No code, data or model release is stated.