“What is a good MAPE?” is one of the most common questions in demand planning, and it is the wrong question in a specific and expensive way. A single accuracy percentage compresses thousands of individual misses into one number, and the compression is not neutral: it hides some kinds of failure and exaggerates others. An operation can report a respectable accuracy figure every month while the items carrying the margin are wrong, in the same direction, all year. Here is what the number conceals, and how to grade a forecast so it means something.
What is a good MAPE, and why is that the wrong question?
MAPE is the average percentage by which your forecasts miss. For each item and period, take the gap between forecast and actual, express it as a percentage of the actual, and average those percentages.
There is no universal threshold for a good one, and any figure quoted as an industry standard deserves suspicion, because the achievable error depends almost entirely on how volatile the thing you are forecasting is. A steady high-volume line and a lumpy specialty item are not in the same event. Comparing your number against another company’s mostly tells you their demand is calmer or wilder than yours, not that their planning is better.
The question that does have an answer is comparative. Is this forecast better than the cheapest thing you could have done instead? “Same as last month”, “same as this month last year”, or a rolling average of the last few periods are all free. If your forecasting process cannot beat them, it is not earning its keep, whatever its accuracy score says. Scoring a forecast against what a naive guess would have achieved on the same data was proposed for exactly this reason in Hyndman and Koehler, “Another look at measures of forecast accuracy”, International Journal of Forecasting 22(4), 2006. Ask your team for that comparison. It is one extra column and it settles the argument.
Why does MAPE break on slow-moving items?
Because it divides by the actual, and the actual is sometimes zero. Hyndman and Athanasopoulos put it plainly in Forecasting: Principles and Practice: percentage errors have “the disadvantage of being infinite or undefined” when the actual is zero, “and having extreme values” when it is close to zero.
Watch what that does in a real planning meeting. A week with no sales cannot be scored at all, so the row gets dropped. A week where one unit sold and three were forecast produces a 200% error that dominates everything around it. Neither is a judgement about your forecast; both are artefacts of the formula. And the response is almost always the same: the slow movers come out of the accuracy report because they “distort” it.
Once that happens, the reported number describes only the part of the catalogue that was easy to forecast in the first place, and the exclusion is rarely written anywhere the executive reading the number can see it. Slow items are usually most of your item count. Removing them from the report does not stop them consuming cash.
Why does MAPE punish over-forecasting more than under-forecasting?
This is the property with the largest consequences and the least discussion. The same source notes that percentage errors “put a heavier penalty on negative errors than on positive errors”, and you can see why with arithmetic you can do in your head.
Take an item where 100 units actually sold. Forecast it at zero, the worst possible under-call, and you score a 100% error. That is the ceiling; no under-forecast can do worse. Now forecast the same item at 400 and you score 300%. Forecast it at 1,000 and you score 900%. On the over-forecast side there is no ceiling at all.
So the penalty is lopsided, and a process tuned to make the number look good learns to forecast low. Nobody decides this. It emerges over a few review cycles as planners work out which kind of miss gets them criticised.
The operational result is an operation that runs systematically short: stockouts, expedited freight, and the steady friction of sales chasing availability, all while the accuracy report improves. If your accuracy number has been getting better and your service level has not, this is the first place to look.
Why can the headline number be healthy while the wrong items are wrong?
Two mechanisms produce this, and most planning reports contain both.
One item, one vote. MAPE averages across items without weighting them. An item worth a fraction of a percent of revenue counts exactly as much as one worth a tenth of it. A catalogue with a long tail therefore reports a number that is mostly about items nobody makes decisions on. The fix is to weight by volume or by value: total the misses in units or shekels and divide by total actuals, so each unit counts once instead of each item counting once. It is not more sophisticated, it is just honest about what you care about.
Errors cancel when you add them up. Report accuracy at the total level and opposite misses offset: one item over by a hundred units and another under by a hundred net to zero at the top. That is arithmetic, not luck, and it means the group number always looks better than the item numbers. It is not false, it is answering a question about the total, which is right for cash planning and wrong for purchasing, because you buy per item and the surplus of one will not cover the shortage of another.
Both point the same way: grade the forecast at the level where the decision gets made. If purchase orders go out per item, grade per item. If capacity is planned per family, grade per family. Grade above the decision and you are measuring something nobody acts on.
What is forecast bias, and why does it matter more than accuracy?
Accuracy measures how big your misses are. Bias measures which way they lean. They are independent, and the second is usually the one costing you money.
To find it, average the misses with their signs instead of ignoring the signs. Near zero means you are wrong in both directions roughly equally. Consistently positive or consistently negative means you are wrong the same way every single cycle, which is a correctable habit rather than irreducible noise.
The distinction matters because the two failures look identical on an accuracy report and behave nothing alike in the warehouse:
| Leans neither way | Leans the same way every time | |
|---|---|---|
| Small misses | The forecast is doing its job. | A consistent, correctable gap. Adjust for it and the forecast improves immediately. |
| Big misses | A genuinely volatile item, honestly forecast. Manage it with buffers, not with a better model. | The worst case. Wrong every cycle, and wrong the same way. Usually a process problem, not a model problem. |
Persistent lean is almost always human and structural rather than statistical: a sales team whose forecast is also their target, a planner who was blamed for a stockout two years ago, a commercial ambition being entered as a demand forecast. No model fixes that, and no accuracy metric reveals it, because taking the absolute value throws the direction away before the average is taken. If you add one number to your pack this quarter, add the one that keeps its sign.
How do you grade a forecast honestly?
The mechanics are ordinary. The discipline is the hard part.
- Freeze the forecast when it is made. Store it, dated, at the level and horizon where it drives a decision. A forecast quietly overwritten by next month’s version cannot be graded, and this alone is why most organisations cannot name their accuracy at all.
- Grade both horizons you actually use. The short one that drives replenishment and the long one that drives capacity and cash are different promises and fail differently.
- Report four numbers, not one: how big the misses are, which way they lean, the same pair weighted by value, and the comparison against the naive baseline. Four columns, four questions.
- Break it out by segment, fast against slow, top revenue against the tail, per family or per region. The total is for the board; the breakout is where the decisions live.
- Say which rows were excluded and why. Dropping zero-demand periods because the formula cannot handle them is a legitimate decision and an illegitimate silence.
- Put it in front of the same people every cycle. The value is not in the calculation. It is in the same people seeing the same number often enough to recognise a pattern.
None of this requires a model, a vendor, or a data project. It requires last year’s forecasts, last year’s actuals, and the willingness to let the first few months be embarrassing.
We have written separately about what happens when companies actually grade their forecasts, including the published industry evidence and the monthly grading ledger we ran for three years on a real engagement. Two companion guides cover the neighbouring mistakes: why your average is lying to you on the flat number hiding a moving business, and demand is not a bell curve on why stock buffers get set wrong on exactly the items that need them most.