What you are actually deciding
Conversations about restaurant forecasting usually start with accuracy and end nowhere. The decision underneath is narrower and more awkward than accuracy. On a Tuesday morning a general manager commits named people to named hours for a week that has not started yet, and once that schedule is posted, changing it costs money, goodwill, or both. Everything about the model follows from that commitment. A forecast that is two points better and arrives after the schedule is posted is worth nothing at all.

That framing changes what you build. You are not predicting sales. You are producing a number that a manager will use to decide whether Tuesday dinner gets three servers or four, ten days from now, knowing that one of them will call out and that the weather forecast that far ahead is close to worthless. The model has to be right at that distance, at that grain, in that unit.
You are probably here because
- Labor as a percent of sales is drifting up and nobody can say which stores are causing it
- You bought a scheduling tool with a forecast in it and the managers override the forecast every week
- Overtime keeps appearing on Saturday night and nobody decided to spend it
- Dinner runs short-staffed for the first ninety minutes and then everyone stands around
The first two are usually process problems. The last one is almost always a timestamp problem, and it is the cheapest fix in this article.
The posting deadline sets the horizon, not the model
Predictive scheduling ordinances now cover a meaningful share of restaurant labor. Seattle, San Francisco, New York City, Philadelphia, Chicago, Emeryville and the state of Oregon all require the schedule to be posted somewhere between ten and fourteen days ahead, and all of them require premium pay when it changes after posting, typically one to four hours of pay per affected change. Even in markets with no ordinance at all, a group that posts late loses people, because a server with two jobs schedules the other one first.
So the horizon is not a modeling choice. If the rule is fourteen days and you post a seven-day week, the far end of your forecast is twenty-one days out. Most vendor demos evaluate at seven days, where the numbers look good, and most internal backtests inherit that habit. Accuracy at twenty-one days is materially worse than at seven, and the gap is largest exactly where it hurts: local events, weather and promotions. Backtest at the horizon you will actually run at, and report both, so nobody is surprised in month three.
There is a second, softer horizon that matters just as much. The day before, a manager can still call someone in or send someone home. That is a different decision with a different model: it uses same-day sales through lunch, and it is often the highest-return piece of the whole system because it is the only place where new information can still change an outcome.
Forecast the arrival, not the check
This is the mistake that produces the most confident wrong answers. In most point-of-sale exports, the timestamp on a transaction is when the check closed. In full service that is forty-five minutes to nearly two hours after the party sat down. Build an hourly demand curve from check-close times and the model will ask for servers an hour after the servers were needed, which is exactly what a manager means when they say the forecast is wrong at dinner but right on the day total.
The fix is to use a timestamp that reflects arrival. Check-open time is usually available and is close enough. The first item fired to the kitchen display is better for back-of-house labor. Reservation and waitlist seating times from the host stand are best of all where they exist. In quick service and drive-thru the gap between order and close is a couple of minutes and none of this matters, which is why a vendor whose product grew up in quick service can be genuinely surprised when it lands badly in a casual dining group.
Channels need the same care. Third-party delivery consumes kitchen labor, packaging and expo time, and almost no front-of-house labor. A group where delivery went from four percent of sales to twenty-two percent and never split the forecast is staffing the dish pit against dine-in covers. Forecast dine-in, takeout, delivery and catering as separate series, then convert each to labor with its own coefficients.
What actually moves the number
Teams tend to arrive convinced that weather is the thing that will fix the forecast. It is worth adding, it is cheap, and it is not that thing. Here is the honest ranking from store-level work, with the size of the swing a practitioner would recognise.
| Driver | Typical effect on a store-day | Worth modeling? |
|---|---|---|
| Day of week and daypart | The dominant structure; 40–60% of the variation in a stable store | Yes. This is the base and a seasonal-naive model gets most of it |
| Season and the local school calendar | ±10–25%, larger near campuses and in resort markets | Yes, and a calendar feed beats a learned seasonal term |
| Local events — stadium, convention, concert | ±15–80% on the handful of stores in range | Yes, as a feed. Do not ask a model to learn a concert schedule |
| Weather | Usually ±3–12%; much larger for patios, walk-up windows and delivery mix | Yes, but expect a point or two of accuracy, not a transformation |
| Promotions and limited-time offers | ±5–30% | Only if promo history is recorded cleanly, and it usually is not |
| Pay cycles and month boundaries | ±3–8% in some markets, near zero in others | Cheap to test, easy to drop |
| A closed lane, a remodel next door, a competitor opening | Large, sustained, and invisible in the data until after the fact | No. This is what the general manager knows and you do not |
That last row is the reason the override conversation matters more than the feature engineering conversation.
The schedule is an assignment problem wearing a forecast's clothes
Suppose the demand curve is right. Turning it into a posted schedule is a constrained assignment problem, and the constraints are not optional. Stated availability. Minimum shift length. Meal and rest breaks, which in California means a meal period before the end of the fifth hour and a rest period per four hours, with an hour of premium pay owed for each one missed. Hour limits for minors during the school year. Alcohol service and food handler certifications. Maximum consecutive days. Rest between a close and the next morning's open. Overtime at forty hours in the week, and after eight hours in the day in several states.
And floors. You cannot run a line below a minimum viable crew no matter what a Tuesday-in-February forecast says, and a system that suggests it will be ignored once and then ignored forever. Encode the floors and ceilings per store, per daypart, and let the optimizer work between them.
A tool that hands a manager a labor-hour target and no schedule has done the easy half of the work and left the hard half. The hour target is arithmetic. The schedule is the thing that takes a manager four hours on a Tuesday.
Where the value sits in a labor project — our working split
Weights sum to 100. Our starting allocation for a multi-unit group, not a measurement. Argue with it before you plan the work.
The manager override is data, not a defect
Every system of this kind gets overridden. The model says four servers; the general manager schedules five. The instinct is to lock the schedule down. That instinct is wrong twice over: managers hold information the model does not have, and a manager who cannot override will simply stop using the tool and rebuild the schedule in a notebook.
Instead, make the override cheap and make it recorded. Two fields: what they changed, and why, from a short controlled list. A forty-top on the book. A new server on their third shift. The fryer is down. Road closed. After a season of that you can measure whether overrides help. The honest finding in most groups is that overrides help on event days and hurt on ordinary Tuesdays, and that four or five managers account for most of the damage and most of the value. That is a coaching conversation with evidence behind it, which is worth more than the two points of accuracy that started the project.
Measure it properly: compare the naive baseline, the model, and the posted schedule as three separate forecasts of the same week. If the posted schedule beats the model, the model is not ready. If the model beats the posted schedule and the managers still override, you have a trust problem, not an accuracy problem, and shipping a better model will not fix it.
What a point of accuracy is worth
Do this arithmetic before the project, not after. Twenty stores, one unnecessary scheduled labor hour per store per day, at a fully loaded rate of eighteen to twenty-two dollars, is roughly one hundred and thirty to one hundred and sixty thousand dollars a year. That is the ceiling, not the prize. Some of that hour is deliberate service insurance and you do not want it back.
A realistic capture from a good demand model plus schedule generation plus a disciplined posting process is half a point to a point and a half of labor as a percent of sales, and in most groups the first half of that comes from fixing the process rather than from the model: posting on time, using earned-hours targets consistently, and killing the standing overtime that appears in the last thirty-six hours of the week when someone calls out. Overtime typically runs two to six percent of hours in a group with no controls and it is almost never a scheduling decision anyone made on purpose.
Write the dollar number down before you start, and write down who owns it
If the projected saving is a hundred and forty thousand a year across a group and the project costs three hundred thousand plus a subscription, the honest answer is to do the process work and skip the model. We would rather tell an operator that in week one than in month nine. The projects that pay back are the ones where somebody in operations, not analytics, signed up for the number.
When you do not need any of this
One restaurant with a settled pattern and a general manager who has been there four years does not need a forecasting model. The average of the last four matching weekdays, adjusted for what the manager knows about the week ahead, lands within a couple of points of anything you can buy. Two to six restaurants with an owner who reads every schedule is the same story. Spending forecasting money there instead of on the ordering process is a straightforward waste.
The value appears when nobody can look at every store every week. Call it eight or ten locations, or any group where a regional manager carries more stores than there are days in the week, or a group opening enough new stores that the pattern keeps resetting. Below that line, the right deliverable is a good spreadsheet and a written posting discipline. We have said exactly that to operators who came in expecting a build.
Measure the thing you are actually paying for
Do not report error metrics to operators. Report hours and dollars. The dashboard that gets used has six lines on it.
- Scheduled hours against earned hours against actual worked, by store-week, with the gaps named
- Sales per labor hour by daypart, because the daily number hides the shape
- Peak intervals below the service floor — the one metric that protects guests from the cost programme
- Overtime hours and where they originated, since almost all of it is created in the final two days
- Premium pay triggered by post-publication changes, per store, per month
- Baseline, model and posted schedule scored against the same week, so the override conversation has numbers in it
How this actually gets built
A first pass for a multi-unit group
The shadow period is the one that gets cut for schedule and the one that decides adoption. Six weeks of the tool being visibly right, next to what the manager actually did, buys more trust than any amount of explanation. It also surfaces the stores where the data is broken, and there are always two or three: a store whose business day rolls at a different hour, a store that changed its close time and never told the data team, a store where the delivery tablet never made it into the sales feed.
The mistakes we get called about
- Demand built on check-close timestamps, staffing dinner an hour late in every full-service store
- One blended forecast across dine-in, takeout and delivery, so kitchen labor tracks the wrong series
- Backtested at seven days and deployed against a fourteen-day posting rule
- An hour target with no schedule, leaving the four-hour task exactly where it was
- Overrides blocked, followed by managers rebuilding the schedule outside the tool
- New stores scored against mature-store error targets for their first two quarters
- No service floor, so the first quiet month produces a shift nobody could run
- Accuracy reported to operators who needed hours, dollars and a reason
Bottom line
The forecast is the least interesting part of a restaurant labor system, and it is where nearly all the attention goes. The horizon comes from the posting rule. The demand curve comes from when guests arrived, not when checks closed. The schedule comes from constraints that are written in law and in the physical limits of a kitchen. The override is information you should be collecting rather than a defect you should be suppressing. And under roughly eight locations, the correct recommendation is usually a spreadsheet and a firmer posting discipline. Everything in this article is cheaper than the model, and all of it has to be right before the model matters.
Frequently asked questions
Two to three years per store is comfortable, mainly so the model sees each holiday twice and can separate seasonality from trend. One year works with a pooled model that borrows structure across stores. A store with under six months of history should be forecast from its cohort, not from itself, and should be flagged in the interface so nobody treats its number as settled.
For a stable store at a one-week horizon, an error in the eight to fifteen percent range is normal, and a seasonal-naive baseline is often within two or three points of a good model. At fifteen-minute grain the error is much larger, frequently thirty percent and up, which is why labor decisions should be made on daypart blocks rather than on quarter-hour buckets.
Usually yes, because a basic forecast feed is inexpensive, and usually much less than expected. It matters most for patios, walk-up windows, and the dine-in versus delivery split. It matters least for a suburban drive-thru at dinner. Add it after the calendar work is done, and be prepared to find it worth a point.
Yes, with the change and a short reason recorded. Managers know things the data does not contain, and a locked schedule guarantees the tool gets abandoned. Once you have a season of override records you can measure which overrides help, which do not, and which managers to coach, which is more valuable than the accuracy the lockdown was protecting.
Below roughly eight locations there is normally someone who can still look at every schedule every week, and their judgment plus a four-week weekday average is competitive with a purchased system. The threshold is organisational rather than statistical: build when nobody can hold the whole group in their head any more.
