Decision Analytics

Why Your Average Is Lying to You

Dan Malka · Updated:

Your operating reports are mostly averages: average margin, average deal size, average cost per lead, average lead time, average days to collect. The average is not a lie. It is arithmetically correct, and it answers one question well. The trouble is that it gets used to answer a different question, and almost nobody notices the swap. Here are the three ways that costs real money, and one check you can run this week with an export and a spreadsheet.

Why can a flat average hide a business that is changing underneath it?

Gross margin holds flat for four quarters. Nobody investigates a flat line, so it never reaches the agenda. Underneath it, your high-margin segment is shrinking and a low-margin one is growing just fast enough to fill the revenue hole. Nothing in the average moved. The business changed completely.

This is the most expensive version of the mistake, precisely because the number looks healthy. Marketing meets it as a steady blended cost per lead while a cheap channel quietly saturates and an expensive one takes over the volume. Sales meets it as a flat win rate covering a shift from small fast deals to large slow ones. Manufacturing meets it as a stable scrap rate that is falling on one line and climbing on another.

The cleanest documented proof of this is not commercial, but it is exact. Reviewing 1973 graduate admissions at the University of California, Berkeley, the aggregate figures showed men admitted at about 44% and women at about 35%. Broken out by department, the pooled and corrected data showed a small but statistically significant bias in favour of women. The total and the parts disagreed in direction, because the two groups were applying to departments with very different acceptance rates. The study is Bickel, Hammel and O’Connell, “Sex Bias in Graduate Admissions: Data From Berkeley”, Science 187 (1975). If someone on your team calls this Simpson’s paradox, that is what they mean: the parts can point one way while the total points the other.

The check is free. Whenever an aggregate is stable or improving, recompute it inside each segment. If the direction survives at segment level, the total was telling you the truth. If it reverses, you were looking at a change in mix, not a change in performance.

Why is the average wrong about a typical week?

Most operational numbers are lopsided: many small values, a few very large ones, and nothing at all below zero. Weekly demand on a slower-moving item is the standard case. Most weeks are quiet, and then one customer orders a pallet. That single week drags the average up above almost every week you actually live through.

So the average is wrong twice at the same time. It is too high for the normal week, which makes the item look busier than it is, and far too low for the spike week, which is the one your buffer exists for.

Sales knows this shape as average deal size, pulled upward by one enterprise contract until it describes nobody on the team and quietly distorts every quota built from it. Finance knows it as average collection days, where a handful of chronic late payers move the number while most customers pay on time.

The diagnostic takes a minute and costs nothing. Put the average and the median side by side. The median is simply the middle value: half your weeks above it, half below. If the two are close, the average is doing honest work. If the average sits well above the median, you have a long tail on the high side, and every plan built on that average is under-sizing your exposure to it. The size of the gap tells you how badly.

A rule that survives contact with a real business:

  • Budgeting a total (annual purchase volume, total spend, total hours): use the average. Totals add up, and the average is built for adding up.
  • Describing a typical case (a normal week, a normal invoice, a normal collection time): use the median. It is not dragged around by one outlier.
  • Sizing a buffer or a promise (safety stock, staffing for peak, a service-level commitment): use neither. Use a percentile, because what you are really deciding is how often you are willing to be wrong.

Why do plans built on averages go wrong in a predictable direction?

Because the things that cost you money do not move in a straight line with the thing you averaged. Overtime, expedited freight, a lost sale, a customer who stops calling: none of those scale politely with volume. A busy week costs disproportionately more than a quiet week saves. Run the plan at average demand and you do not get the average outcome, you get a systematically optimistic one. Sam Savage put this to a management audience as the flaw of averages in Harvard Business Review, November 2002: plans built on average assumptions usually go wrong, and they go wrong in a predictable direction.

An importer feels this most sharply in lead time. There is a hard floor, because production plus sailing plus customs clearance cannot go below some physical minimum. There is no ceiling at all. A missed sailing, a port backlog, a quality hold, a supplier shutting down for a fortnight: every one of them adds time and not one of them gives any back. Set your reorder point from the average lead time and you are covered in the cycles that behave and exposed in the ones that do not, and because the long delays are the ones that run longest, those are also the cycles where you are short by the most.

What can you check on your own numbers this week?

An export and a spreadsheet. No model, no vendor, no project.

  1. Segment the flat number. Take one aggregate that has been stable and unquestioned: gross margin, cost per lead, win rate, scrap rate, on-time delivery. Recompute it inside your three or four natural segments and put those lines next to the total. This is the highest-value hour on the list.
  2. Add a median column. For your top twenty items by revenue, or your reps, or your channels, compute the average and the median of the same series and sort by the gap between them. The top of that list is where averages are misleading you most.
  3. Compare the average against the 90th percentile for the same item. That gap is what your buffer is silently being asked to absorb. Put it next to what you actually hold, and you will usually find some items holding far more and some far less, with neither decided on purpose.
  4. Plot one histogram. Take the worst offender from step 2 and plot it. Two separate humps mean you are managing two different behaviours under one name, for example a steady replenishment customer and an occasional project order, and they should be planned as two things. Do this once even where the summary numbers look reassuring: researchers have built whole collections of data sets that share an average and a spread while looking nothing like each other.

None of this needs statistical training and none of it needs us. It is a morning’s work, and it will tell you whether the averages running your business are describing it or hiding it.

If you then want a second pair of eyes on which parts of the picture are worth modelling and which should stay on simple rules, our forecasting practice starts with exactly that assessment on your real data, before anyone commits to building anything. Two companion guides go deeper where this bites hardest: demand is not a bell curve on why standard stock buffers are set wrong, and what forecast accuracy metrics hide on why a healthy accuracy score can sit on top of the wrong items being wrong.