Demand Forecasting

Why Do Companies Grade Their Salespeople but Never Their Forecasts?

Dan Malka · Updated:

Most companies run on forecasts, and almost none of them grade those forecasts against what actually happened. Sales quotas get reviewed monthly; the forecast that set the inventory, the hiring, and the cash plan usually gets quietly replaced by next month’s forecast, errors unexamined. The published forecasting successes all share one unglamorous trait: somebody measured the error, on a schedule, and let the number be embarrassing until it improved. The examples below are real industry evidence, labeled by source type. They are not our client work; the last section describes, without any invented numbers, what the grading practice looks like when we run it.

What does a graded forecast improvement actually look like?

Per an AWS vendor case study with the customer named, More Retail, one of India’s large grocery chains, wired machine-learning forecasts directly into automated store ordering, and tracked the results: forecast accuracy rose from 24% to 76%, fresh-produce wastage fell about 30%, in-stock availability climbed from 80% to 90%, and gross profit improved 25%.

Two things deserve attention beyond the headline. First, they knew their starting accuracy was 24%. Most organizations cannot name their number at all, which means they cannot know if anything improved. Second, the forecast fed a decision, automated ordering, so every error surfaced as a concrete cost, not a slide. Measurement plus consequence is the whole mechanism.

Can grading fix chronic over-forecasting?

Over-forecasting is the polite failure mode: shelves look full, nobody misses a sale, and the cost hides in working capital and waste. Per Zebra’s vendor case study with Walgreens as the named customer, across 20,000 plus SKUs per location, AI-driven demand forecasting cut forecast error by 15 percentage points and reduced over-forecasting from 15% to 1%, while sustaining 98% in-stock rates during demand surges.

The instructive part is the metric choice. Walgreens did not just track “accuracy”; it tracked the direction of its errors. Over-forecast and under-forecast have different costs, and only an operation that grades its forecasts ever finds out which side of the ledger it habitually bleeds on.

Why do probabilistic forecasts change the conversation?

The academic anchor here is DeepAR, the probabilistic forecasting method published by Amazon researchers (peer-reviewed, later productized in AWS), which improved accuracy roughly 15% over the prior state of the art. But its deeper contribution is philosophical: DeepAR outputs a distribution, not a single number. The forecast says “most likely 1,000 units, with a 90% chance of falling between 700 and 1,400.”

That framing makes honest grading natural. A point forecast invites the question “was it right?”, which it never exactly is, and the conversation dies. A probabilistic forecast invites “did reality land inside the range as often as promised?”, which is a checkable, gradeable claim. It also maps directly onto real decisions: safety stock is precisely a bet on the upper tail, and you cannot size that bet from a single number.

What does a grading practice look like month after month?

Here is the practice itself, from our own work, stated without any invented outcome figures. For three years, DataWise ran an embedded forecasting engagement for a US-market industrial manufacturer. Every month, new forecasts went into the leadership planning meeting, and, just as importantly, last cycle’s forecasts were graded against actuals. The grades lived in a running yearly ledger tracking error at the 3-month and the 12-month horizon. Predictions were graded, not just made.

We will not decorate that with a percentage, because the honest summary is not a number; it is that the ledger existed, leadership saw it every month for three years, and the forecasts stayed an input to strategic planning for that entire period. Forecasting services that survive thirty-six consecutive gradings are rare. That survival is the credential.

What should you take away?

  • You cannot improve an ungraded forecast: More Retail’s journey started from knowing the number was 24% (vendor case study, named customer).
  • Grade the direction, not just the size, of errors: Walgreens’ over-forecast metric fell from 15% to 1% (vendor case study, named customer).
  • Prefer ranges to points: probabilistic methods like DeepAR (peer-reviewed) make forecasts checkable promises instead of guesses.
  • Put the grade in front of decision-makers on a schedule; a ledger nobody sees changes nothing.
  • Judge forecasting vendors by whether they volunteer their error history. If they only show wins, they are not grading.

What does this mean for a mid-size Israeli operation?

You almost certainly already forecast, in a spreadsheet, in Priority, or in someone’s head, and you almost certainly do not grade it. The good news is that the first step costs nearly nothing: take the forecasts you made over the last year, put them next to actuals, and compute the error by month and by product family. No model required. That single table tells you where the money leaks: chronic over-purchasing on some lines, stockouts on others, and usually two or three products causing most of the damage.

Only then does AI enter the picture, because now it has a baseline to beat and a ledger to be judged by. That ordering, measurement first, models second, is the difference between the cases above and the far more numerous forecasting projects that quietly disappear.

If you would like help building that first error ledger from your own data, before any modeling commitment, that is a conversation we genuinely enjoy having.