The four questions that come before the model
Most forecasting work we are asked to review is technically sound and operationally dead. The backtest is honest, the error is competitive with anything published on a similar dataset, the pipeline runs on schedule. And the planner still opens a spreadsheet on Monday morning and types over the number. That is not a modeling failure and it will not be fixed by a better model. It is what happens when a forecast is built as an estimate of the future instead of as an input to a decision that already exists, with a deadline, an owner, and a cost of being wrong in each direction.
You are probably here because
- The model beat the old one on the backtest and the planning team still overrides it every cycle
- Somebody asks how accurate the forecast is and there are three answers depending on who is in the room
- The forecast is a single number and every consumer of it privately adds their own buffer on top
- Accuracy has improved twice in eighteen months and no operational metric moved with it
All four come from the same place: the forecast was specified as a prediction rather than as an input to a decision. The sections on horizon, grain and quantiles are where that gets repaired.

Before anyone touches the data, four questions. Which decision does this change? Not which report it appears in. Which commitment gets made differently. Who makes that decision, and on what day? A name and a recurring calendar entry. How much would the number have to move to change the answer? If demand could double and the order stays the same because the supplier ships in fixed pallets, precision below a pallet is decoration. What does being wrong in each direction cost? In currency, roughly, per unit, and separately for high and low.
If those four have answers, the modeling work is straightforward and the project usually lands. If they do not, no amount of accuracy will save it, and we say so before quoting the work rather than after.
Lead time sets the horizon. Nothing else does.
The horizon is not a preference and it is not inherited from the data. It is the elapsed time between making a commitment and living with its consequence. If a purchase order takes six weeks to arrive, a forecast of next week is a hobby: whatever it says, the stock on hand was decided a month and a half ago. If a staffing roster locks ten days out and cannot be changed without overtime penalties, forecasting day three changes nothing that day-eleven does not already decide.
The usual defect is a horizon inherited from data granularity. Daily data exists, so the team forecasts thirty days daily, and nobody notices that the only commitment in the calendar happens quarterly at a category level. The horizon should come from the commitment calendar, and it should include the review period. If orders are placed weekly and the supplier takes six weeks, the exposure window is seven weeks, not six, because whatever you order today has to cover demand until the next order arrives.
Timing matters as much as span. A weekly forecast delivered Tuesday for a Monday ordering cycle is a forecast for the following week, whatever the header says. Publish before the decision with enough margin for a person to read it and ask a question, or the forecast is documentation rather than input.
Grain is an arithmetic choice, and the arithmetic is unkind
Aggregation buys accuracy. For series that are roughly independent, summing n of them shrinks relative error by something close to the square root of n: twenty-five items rolled into a category can show relative error near a fifth of the item-level figure. Real demand is correlated, so the gain is smaller in practice, but the direction is reliable and it explains why the category forecast always looks so much better than the item forecast built from the same model.
That is also why an aggregate accuracy number is close to meaningless as a claim about a system that makes item-level decisions. If replenishment happens per item per location, that is the grain the forecast has to work at, and a company-level accuracy figure is a comfort metric. Forecast at the decision grain, then reconcile upward so the numbers in the operating plan add up to the numbers in the orders.
| Decision | Horizon comes from | Grain | Useful output |
|---|---|---|---|
| Replenishment order | Supplier lead time plus the review period | Item × location | A quantile at the service level you fund |
| Shift staffing | Roster lock, typically 7–21 days | Interval × site | Quantile plus the intraday shape |
| Working capital | The financing cycle, often 13 weeks | Weekly, per entity | A distribution with the downside named |
| Hiring or capacity | Time to productive, 2–4 quarters | Monthly, per team | Two or three scenarios, not one line |
| Committed volume contract | The contract term | Annual total | Distribution against the break-even |
A point forecast answers a question nobody asked
The mean is the right answer when the cost of being high equals the cost of being low and both grow linearly. Almost nothing in operations looks like that. A stockout costs a lost sale, an expedite fee, and sometimes a customer. A unit of excess costs storage and capital for a few weeks. When those two numbers differ by a factor of ten, an unbiased forecast is a systematically expensive one.
The arithmetic is old and it is one line. The optimal service level is the cost of being under divided by the sum of the costs of being under and over. Put a stockout at $40 of margin and holding at $4 per unit over the cycle, and the ratio is 40/44, which is 0.91. The right order quantity sits at the ninety-first percentile of the demand distribution, not at the mean. Publish the mean and the planner has to reconstruct that number themselves, in their head, under time pressure, without the residual distribution in front of them.
There are three honest ways to produce quantiles and the cheapest one is usually good enough to start. Fit quantile models directly against pinball loss. Fit a distributional model and take its percentiles. Or take the empirical distribution of your own backtest errors at each horizon and add it to the point forecast. The third requires no new model and it has a property the first two often lack: it includes the errors your process actually makes, not the errors your model assumes it makes.
Then check coverage before shipping. Model-derived intervals are usually too narrow because they account for noise and not for parameter uncertainty, regime change, or the promotion nobody logged. If the nominal ninety percent interval contains the truth sixty-two percent of the time in backtest, the honest move is to widen it and say why, not to relabel it.
Where the usable improvement comes from — our ranking
How we rank the levers across forecasting work we are brought into. Judgment, not a benchmark. The bottom row is where most projects start.
The baseline you have to beat, and the one nobody scores
Before any model is evaluated, run the seasonal naive: this week equals the same week last year, or this Tuesday equals last Tuesday, whichever cycle dominates. On stable series it is hard to beat by much, and a model that improves on it by three percent has told you something useful about how much signal is there to find.
Then score every step of the process against that baseline, including the human ones. This is forecast value added, and it is uncomfortable in a productive way. The statistical model is one step. The planner override is another. The commercial adjustment made in the monthly meeting is a third. Each gets scored against the naive baseline and against the step before it. It is common to find that one of those steps subtracts accuracy while consuming several days of skilled labour every month, and it is almost never the step people expect.
On numbers, plainly, so you have something to calibrate against. Aggregated weekly demand on a stable business, one to four weeks out, commonly lands in the range of ten to twenty percent weighted absolute percentage error. The same business at item and location grain routinely runs thirty to sixty percent. Slow-moving items with many zero weeks can exceed one hundred percent and percentage error stops being meaningful at all. Any accuracy claim quoted without its grain and its horizon is unfalsifiable, and unfalsifiable numbers are how forecasting programs lose the trust of the operators who have to live with them.
Match the metric to the money
If you optimise squared error you have asserted that being high and being low cost the same. If that is false, the metric is quietly arguing against the business. Optimise pinball loss at the quantile the decision needs, and report it.
Weighting matters just as much. Mean absolute percentage error treats a two-unit item and a twenty-thousand-unit item as equals, so a program can improve its headline metric by getting better at products that do not matter. Weighted absolute percentage error, computed as total absolute error over total actuals, weights by volume and is much closer to the currency. For intermittent series, percentage errors divide by zero and the symmetric variants behave badly near zero; a scaled error such as MASE, which compares against the naive one-step error, stays defined and stays comparable across series.
Ship the decision rule in the same artifact as the number
A forecast row that reads p50: 420, p91: 610 is data. The same row followed by order to p91 unless the item is on the discontinue list is a decision. Put the rule beside the number, in the same table, version it, and log which rule produced each order. When someone later asks why the warehouse holds six weeks of a slow item, the answer is retrievable in a minute instead of being reconstructed from memory in a meeting.
Send the backtest and we will tell you where the error is coming from.
Email your backtest output at decision grain, the horizon you forecast, and the decision it feeds to contact@precisionfederal.com. You get back a short written note: how you compare to the seasonal naive at that horizon, whether your intervals are calibrated, and the three changes we would make first. One business day. No charge and no meeting.
contact@precisionfederal.comThe forecast and the plan are different objects
The forecast is what you believe will happen. The plan is what you have committed to deliver. They are allowed to disagree, and the disagreement is the most useful number in the document, because it is the gap the business has to close with action rather than arithmetic.
What destroys this is the quarterly ritual where the forecast is adjusted upward until it matches the plan. After that neither object means anything: the forecast is no longer a belief and the plan is no longer a commitment against a belief. Keep two columns, label them honestly, and let the delta be the agenda item. The teams that do this argue less, because the argument has somewhere to land.
Store forecasts as immutable facts with an as-of date rather than overwriting a current-state table. Without that you cannot answer what the forecast said when the order was placed, which is the only question that matters in a post-mortem. That storage discipline and the leakage traps that go with it are their own subject, covered in time-series forecasting in production.
Keep the overrides. Measure them.
Planner overrides are not a failure of the model. A person often holds information the model cannot have: a customer that churned last Thursday, a competitor's promotion starting Monday, a site closing for renovation. Removing their ability to intervene produces a worse system and a resentful team.
What you do owe is measurement. Log the pre-override value, the post-override value, the reason code, and the actual. Score both. The pattern that shows up most often is a small set of series where overrides add real accuracy because the person genuinely knows something, and a long tail where overrides are pattern-matching against noise and subtract value. Once that is visible per person and per series, the conversation stops being about trust and becomes about which series each planner should be spending their limited attention on. Require a reason code for any override past a threshold, and review the largest ones each cycle rather than all of them.
Cost to change once consumers depend on the output
Difficulty as we rank it, driven by how much downstream logic each change breaks. Settle the top two before the first publication.
The failures we are called in to fix
- A horizon copied from the data cadence rather than derived from the commitment calendar
- Aggregate accuracy quoted for a system that makes item-level decisions
- A point forecast with no distribution, and six consumers each adding a private buffer
- Intervals from the model likelihood, never checked for coverage against backtest
- Percentage error on intermittent demand, silently dropping every zero-demand period
- The forecast quietly reconciled to the plan, after which neither number means anything
- Overrides forbidden or unmeasured — the two ways to learn nothing from them
- A feature computed from a table that gets restated, so the backtest saw numbers the live model never will
A four-week pass that usually settles it
Forecast Design Pass
Step five is the one that changes minds. Replaying last year's decisions under the new rule converts an accuracy argument into a currency argument: this many stockouts avoided, this much stock not held. It is also where you discover the constraints nobody mentioned, like a supplier minimum that makes half the recommendations impossible to execute.
Before you publish
- The decision, its owner and its calendar day are written down
- Horizon equals lead time plus review period, not the data cadence
- Grain matches the decision, with reconciliation upward to the plan
- A seasonal naive baseline exists at the same grain and horizon
- Every process step, human ones included, is scored against that baseline
- Output carries quantiles, and coverage has been checked in backtest
- The metric is weighted by volume, and is defined when demand is zero
- Forecasts are stored as facts with an as-of date, never overwritten
- Overrides are permitted, logged with a reason, and scored
- The forecast and the plan are two columns that are allowed to disagree
When a forecast is the wrong answer
Some of the requests that arrive as forecasting projects should not be forecasting projects, and saying so early is cheaper for everyone than saying it in month three.
The series is a policy, not a phenomenon. If a number is set by a decision you control — a price, a quota, a headcount cap — forecasting it is forecasting your own behaviour. Model the driver and the response instead.
The history is too short for the seasonality that matters. Annual seasonality needs at least two, preferably three complete cycles before a model can separate it from trend. With fourteen months of data and a strong yearly pattern, a model will confidently report a trend that is one holiday season. Borrow structure from similar series or use a simple profile until the history exists.
The variance is dominated by a few known events. When three customers are forty percent of volume and their orders are negotiated, the forecast worth having is a call sheet, not a model. Sales knows more than the time series does, and the useful work is capturing what they know in a structured way and scoring it.
Nothing downstream can move. If capacity is fixed, the supplier ships fixed pallets and the roster is set by a contract, an improved forecast changes nothing this year. That is a legitimate finding. It is also an argument for spending the budget on the constraint rather than on the prediction.
Our accuracy is bad. Should we start with better data or a better model?
Data, almost always, and specifically the point-in-time correctness of it. The most common cause of a backtest that flatters and a production system that disappoints is a feature computed from a table that gets restated after the fact. The model saw a version of history that was not available on the day it would have had to predict. That defect is invisible in offline scores and fatal in production.
How much history do we need?
For weekly data with annual seasonality, two full years is a working minimum and three is comfortable. For daily data with weekly seasonality, a few months can be enough. Below that, use a simple seasonal profile borrowed from a comparable group of series and be explicit that it is borrowed, so nobody quotes its confidence intervals as if they were earned.
Can one model serve planning, finance and operations?
One model, often. One output, rarely. Finance wants monthly totals in currency, operations wants weekly units at item and location, and planning wants a quarterly view aligned to the commitment cycle. Produce them from one reconciled hierarchy so they add up, then publish three views with the grain and horizon stated on each. The alternative is three teams maintaining three forecasts that disagree in a meeting once a month.
How often should the model be retrained?
Refit on a schedule that matches how fast the process changes, and monitor rather than guess. A monthly refit is a reasonable default for demand that shifts gradually. What matters more than cadence is a guard: compare the new fit against the incumbent on a held-out recent window and refuse the promotion if it is worse, so an automated retrain cannot quietly ship a regression.
Bottom line
Forecast accuracy is a means, and treating it as the goal is what produces technically good work that nobody uses. Lead time tells you the horizon. The decision tells you the grain. The asymmetry between the cost of too much and the cost of too little tells you which quantile to publish. Get those three right with an ordinary model and you will beat a better model published at the wrong shape, because one of them can be acted on and the other one has to be reinterpreted by hand every cycle by someone who is already busy.
Frequently asked questions
It depends entirely on grain and horizon, which is why a single number is usually a warning sign. Aggregated weekly demand one to four weeks out often lands between ten and twenty percent weighted absolute percentage error. The same business at item and location grain commonly runs thirty to sixty percent. Slow movers with frequent zeros can exceed one hundred percent, and percentage error stops being meaningful there. The useful target is a stated improvement over the seasonal naive at your own grain and horizon.
A range, in almost every operational case, because the cost of being high and the cost of being low are rarely equal. Compute the service level as the cost of a shortage divided by the sum of shortage and holding costs, and publish that quantile alongside the median. The cheapest honest way to get there is the empirical distribution of your own backtest errors at each horizon.
Log the value before the override, the value after, and the actual, then score both against the seasonal naive. This is forecast value added. Expect a mixed answer: overrides usually add accuracy on a small set of series where the person has private information and subtract it on a long tail where they are reacting to noise. Reported per person and per series, it becomes a workload conversation rather than a trust conversation.
Weighted absolute percentage error for anything with meaningful volume, because it weights by units and stays close to the currency. A scaled error such as MASE for intermittent series, because percentage errors are undefined at zero. Pinball loss at the quantile you actually order to, if you have moved to quantile output. Report all three at decision grain rather than one at company level.
For a few hundred series with clear seasonality, well-tuned statistical methods are hard to beat and much cheaper to operate. Learned models start to pay when you have thousands of related series that can share structure, or strong known-future drivers such as price and promotions. In the work we see, the model family is rarely the binding constraint; horizon, grain and input correctness usually are.
