Skip to main content
Forecasting & Operations

Load forecasting for a utility

Annual average error is the number everyone quotes and almost nobody is paid on. The money sits in a handful of hours a year. That single fact changes how the forecast should be built, evaluated and used.

Four forecasts, one name

When someone says the load forecast is bad, the first job is to find out which forecast they mean. A utility runs at least four, they answer to different people, they are wrong in different ways, and improving one does nothing for the others. Buying a model before separating them is how a project ends up with a beautiful day-ahead curve delivered to a planning group that needed a ten-year peak.

HorizonDecision it feedsWho lives with itWhat accuracy actually means here
Very short term
Minutes to a few hours
Dispatch, switching, reserve postureThe control roomRamp shape and turning points, not level
Day ahead
Hourly, next one to seven days
Purchases, schedules, unit commitmentPower supply and schedulingHourly error where prices are highest, not the daily mean
Peak alerting
Which hour tops the month
Curtailment calls, demand response, load shiftingWhoever owns the wholesale billDid you call the right hour. Everything else is noise
Long term
Annual peaks, five to twenty years
Capacity, capital plan, rate case supportPlanning and financeWhether the growth assumptions are defensible in writing

A large system's day-ahead hourly forecast typically lands somewhere in the range of one and a half to three percent mean absolute percentage error. A single substation or feeder is a different animal entirely, commonly five to fifteen percent, because diversity is what makes aggregate load smooth and a feeder has very little of it. One irrigation district, one plastics plant, one data hall, and the feeder stops resembling a curve and starts resembling a switch. Anyone who quotes you a system-level number for a feeder-level product is not being careful.

You are probably here because

  • The wholesale bill jumped and it traces to two afternoons in July
  • The forecast that scored well all year missed the week that mattered
  • Afternoon load stopped behaving the way it did five years ago and nobody rebuilt the model
  • Planning and operations quote different peak numbers in the same meeting

All four point at the same thing: the forecast is being scored on the average and paid for on the extremes.

The peak hours are the entire bill

For many distribution utilities and cooperatives, a large share of annual wholesale cost is set by demand in a small number of coincident peak hours. Whether it is one hour a month or a handful of hours a year depends on your contract, but the structure is nearly always the same: an enormous amount of money riding on a tiny number of intervals. A forecast that is excellent on the other eight thousand hours and misses those has not helped.

This should change what you build. The product is not a curve. It is a call. Somebody has to decide, usually the afternoon before, whether tomorrow is a peak day and whether to spend real money reducing load: notifying interruptible customers, running distributed generation, shifting pumping, pre-cooling. Each call has a cost when it is wrong in either direction. Calling a peak that does not arrive burns goodwill with the customers who curtailed, and there are only so many times a year you can ask. Missing a peak sets a demand charge you carry for twelve months.

So the forecast should output a probability that a given day contains the monthly peak, and it should be scored against the decision, not against megawatts. Ask your current vendor a simple question: over the last three years, how many of the actual peak hours did the forecast rank in its top three candidates for that month? That number is worth more than any error statistic, and most programs have never computed it.

The forecast that is excellent on eight thousand hours and misses the twelve that set your bill has not helped anyone.

Weather is most of the signal and you do not control it

On a hot-summer system, temperature and humidity explain the overwhelming majority of day-ahead variance. Which means your forecast error is mostly inherited weather forecast error, and no amount of modelling skill recovers it. If tomorrow's high is off by three degrees, your load will be off, and the honest ceiling on your accuracy is set by your weather provider rather than by your algorithm.

Practical consequences. Buy more than one weather feed and score them against each other on your own service territory rather than on national skill scores, because vendors differ locally and the local difference is the only one you pay for. Use the spread across providers as an uncertainty signal: when three forecasts disagree by four degrees on Thursday, the load forecast for Thursday deserves a wider band and a human should see it. And use the right variables. Dry-bulb temperature alone is a weak proxy for cooling load; humidity matters, and building thermal mass means the previous two days matter too. A model with lagged and smoothed temperature terms usually beats one with today's high, and that gap is larger than most algorithm choices.

Weather stations are their own small trap. A station that moved, a sensor that drifted, or a gap that your provider silently interpolated will show up as an unexplained model error weeks later. Keep the raw observations and record which station each reading came from.

Behind-the-meter solar broke the old temperature curve

Load used to rise monotonically with afternoon temperature, which made the modelling comfortable. On systems with meaningful rooftop solar, what the utility sees is net load, and net load falls in the middle of the day and then climbs steeply as output drops through the late afternoon. The peak migrates later. The steep evening ramp becomes the operationally interesting part of the day. And cloud cover becomes a first-class input, because a cloudy hot day and a clear hot day now have different shapes rather than different levels.

Two things follow. First, if your training data spans the period in which local solar adoption doubled, the relationship you are fitting is not stable and the model will systematically misread recent years. Either model gross load and subtract an estimated solar production series, or include an installed-capacity term that lets the fit move. Second, you probably do not know your installed capacity accurately. Interconnection records lag, and small systems get missed. Estimating aggregate behind-the-meter output from clear-sky irradiance and a capacity estimate, then reconciling against observed midday net load, is unglamorous work that produces more improvement than any change of algorithm.

Where day-ahead error comes from — a summer-peaking system, our decomposition

Weather forecast error passed straight through
44
Distributed solar output and cloud cover
19
Large industrial and irrigation customers
15
Calendar effects: holidays, school terms, local events
11
Model form and hyperparameters
11

Illustrative decomposition for a summer-peaking system with visible rooftop solar. Yours will differ; the last row is usually the smallest and gets the most attention.

Meter data arrives late, and then it changes

Interval meter data is the best input a utility has and the most misunderstood. Reads arrive over hours to days, some meters miss a communication window and backfill later, and the validation and estimation process substitutes values for gaps and then revises them. A read you pulled on Tuesday is not necessarily the read that will be in the system on Friday.

This ruins a model quietly. If training rows are built from settled, fully revised data and the production model runs on whatever arrived by six in the evening, you have trained on information the model will never have. The result is a system that validates beautifully and underperforms in service, and the gap is often larger than any modelling improvement you were chasing.

The fix is to store data with two timestamps: when the interval occurred, and when you learned the value. Then build training rows using only what was known at the equivalent hour on that historical day. It is more work than reading a table, and it is the difference between a number you can trust and a number you cannot. The same discipline applies to estimated versus actual reads, which should never be silently interchangeable, and to meter changes, which produce a discontinuity that looks exactly like a behaviour change.

Evaluation Note

A backcast on observed weather is not a forecast

The single most common way a load model is oversold is by evaluating it with the weather that actually happened rather than the weather that was forecast at the time. This removes the largest error source in the problem, and it is not unusual for it to make a model look roughly twice as accurate as it will be in service. Insist on evaluation against archived weather forecasts, at the same lead time the production system will use. If your vendor cannot supply archived forecasts, that is itself an answer.

We will score your current forecast honestly, for free.

Send two years of hourly actuals, the matching forecasts your system produced at the time, and your peak-hour definition to contact@precisionfederal.com. You get back the error broken out by hour of day, by temperature band, and by whether the day contained a peak — plus how often the peak hour was ranked correctly. One business day.

contact@precisionfederal.com

The scheduler who overrides the model is telling you something

Every utility has someone who adjusts the forecast before it is used. Often they have done it for twenty years, often the adjustment is right, and often the modelling team treats it as an insult. It is not. It is a free labelled dataset about what the model does not know.

Log every override: the hour, the direction, the size, and one sentence on why. Within a season you will have a list of the model's blind spots in the operator's own words, and they are usually specific and fixable. The county fair. The plant that runs a Saturday shift in harvest season. The first genuinely cold morning of the autumn, when heating load behaves differently than it does at the same temperature in February. Most of those become features. A few become a note in the interface saying the model has no information about this, which is also a useful thing for the interface to say.

Do not remove the override. A forecast a control room cannot adjust is a forecast a control room will stop opening, and you will hear about that when it is far too late.

What we would not build

  • A single model for every horizon. Dispatch and capital planning need different inputs and different owners
  • A deep network on three years of hourly data when gradient-boosted trees with proper weather features are within noise and can be explained
  • A point forecast with no band for a decision that is fundamentally about risk
  • Customer-level forecasts when nobody has named a decision that changes because of them
  • A retraining job that pulls whatever is in the meter table today, mixing revised and unrevised reads without noticing
  • An accuracy dashboard that reports the annual average and nothing about peak hours

And a case for doing nothing at all. If you buy power under a fixed all-in rate with no demand component and no market exposure, better accuracy has almost no dollar value to you, and the correct answer is to leave the forecast alone and spend the money on something that does. We have said this to utilities that were expecting a proposal. Ask what a one percent improvement is worth in your contract before anyone builds anything; if nobody can answer, that is the first piece of work.

What a first build looks like

Ten weeks, in the order we would do it

1
Write down what a one percent accuracy gain is worth under your actual contract
Week 1
2
Rebuild history with two timestamps so training only uses what was knowable
Weeks 2–4
3
Score the existing forecast against archived weather forecasts, split out by peak days
Week 4
4
Build the baseline: calendar, lagged and smoothed weather, solar estimate, holidays
Weeks 5–7
5
Add the peak-day probability output and a band, and put both in front of the operator
Weeks 7–9
6
Run in parallel with the incumbent, publish both, change nothing operationally
Week 10 onward

Step three frequently ends the project in the best possible way. Sometimes the incumbent forecast is already good, the peak-hour ranking is already strong, and the honest recommendation is to keep it and spend the budget on the meter data foundation instead. That is a cheap answer to have bought.

Step six matters more than it looks. A forecast earns its way into operational use by being visibly right for a season while nothing depends on it. Cutting the parallel run to save six weeks is how a good model gets switched off after one bad Tuesday, because nobody had built any reason to trust it yet.

A forecast earns its way into a control room by being visibly right for a season while nothing depends on it.

Before you commission anything

  • You know which of the four forecasts the project is for
  • Someone has stated in dollars what an accuracy gain is worth in your contract
  • Evaluation uses archived weather forecasts at production lead time
  • Error is reported separately for peak days and peak hours
  • Training data respects when each value became known
  • Estimated and actual meter reads are distinguishable everywhere
  • Behind-the-meter solar is modelled explicitly, not absorbed into a trend
  • The output carries a band, and the peak-day call carries a probability
  • Operator overrides are logged with a reason and reviewed each season

Bottom line

Load forecasting rewards care about inputs far more than cleverness about models. The gains that are actually available sit in weather feed selection and local scoring, in modelling distributed solar rather than letting it corrupt a temperature relationship, in reconstructing history so the model trains on what was knowable, and in reporting error where the money is instead of where it is convenient. Aim the product at the decision: for most utilities that decision is which afternoon to act on, not what the curve looks like on an ordinary Tuesday in April.

Frequently asked questions

What accuracy should we expect from a day-ahead load forecast?

At system level on a large service territory, roughly one and a half to three percent mean absolute percentage error is the usual range, with worse performance on extreme days and holidays. Individual feeders and substations are much harder, commonly five to fifteen percent, because they lack the diversity that smooths an aggregate. Treat any single quoted figure as meaningless until you know the aggregation level, the horizon, and whether it was measured against forecast weather.

Is machine learning better than a regression for this?

Usually a little, and much less than the marketing suggests. Gradient-boosted trees on well-constructed weather and calendar features tend to beat a classical regression by a modest margin and are hard to beat further without adding data rather than model capacity. The larger gains almost always come from better weather inputs, explicit treatment of distributed solar, and honest handling of late-arriving meter data. Model choice is real but it is not where the headroom is.

How much history do we need?

Three to five years of hourly load with matching weather is a reasonable working minimum, mainly so the model has seen several summers and several cold snaps. More history helps only if the system it describes still exists. If your territory added substantial rooftop solar, electric vehicle charging or a large new customer during that window, older years describe a different system and should be weighted or corrected rather than trusted.

Can a forecast tell us which day will set the monthly peak?

It can rank candidates well and it cannot be certain, which is why the output should be a probability rather than a call. The practical measure is how often the actual peak hour appeared in the model's top few candidates for that month, evaluated over several years. That number is directly comparable to what your team achieves today and is the only accuracy statistic that maps onto the decision you actually make.

Do we need our own weather forecasts?

You need more than one source and a record of how each performs on your own territory, which is not the same as buying a premium feed. Score the providers you can get against your own stations for a season, use the disagreement between them as an uncertainty signal, and keep the raw observations so a moved station or a drifting sensor can be found later rather than showing up as unexplained model error.

1 business day response

Wondering whether your forecast is as good as it scores?

We rebuild forecast histories so they train on what was knowable, model distributed solar explicitly, and report error where the money is. If the incumbent forecast is already good enough, that is what the assessment will say.

CapabilitiesMore insights →Email an engineer or email bo@precisionfederal.com
ForecastingUtility AnalyticsData EngineeringDecision Support