Blog » Metric Design
Forecast Accuracy in a KPI Tree: Use WAPE, Not MAPE
· 12 min read
Forecast accuracy KPI tree: every SKU cut its error rate, total demand held at 126,000 units, and 4,560 more units were forecast wrong.
Every SKU Got More Accurate and the Plan Got Worse
A forecast accuracy KPI tree usually breaks at the top node, and the top node is usually MAPE. The figure is correct. It is also an unweighted average of ratios, and an unweighted average of ratios moves for two reasons that look identical on a dashboard.
Take one demand plan, four SKUs, two quarters. In the first quarter SKU A sold 100,000 units against a forecast of 92,000, SKU B sold 20,000 against 18,400, SKU C sold 5,000 against 4,400, and SKU D sold 1,000 against 700. Total demand was 126,000 units and 10,500 of them were forecast wrong.
In the second quarter every SKU improved: A ran 7.00 percent error against 8.00, B ran 7.00 against 8.00, C ran 11.00 against 12.00, and D ran 27.00 against 30.00. Total demand was 126,000 units again.
The units forecast wrong rose to 15,060. MAPE fell from 14.50 percent to 13.00. The cost of the error rose 98,100 dollars.
Nothing there is a rounding artifact. Every number reconciles, and the metric the demand review watched moved the wrong way for a reason it cannot express.
What Is a Forecast Accuracy KPI Tree?
A forecast accuracy KPI tree decomposes a planning miss into the summable quantities that produce it: actual demand, signed forecast error, absolute forecast error and the cost of that error in dollars. Accuracy rates are derived at each node by dividing two of those columns. No row stores a percentage, because percentages do not aggregate.
The quantities are standard. Forecast error is the difference between forecast and actual for one item in one period, and every accuracy measure in common use is built from that difference.¹ They differ in one respect that matters here: whether the result survives aggregation.
Absolute error survives. The absolute errors of four SKUs sum to the group's, which is what lets a parent node be the arithmetic consequence of its children rather than a label above them.
Percentage error does not survive. Dividing before summing discards the quantity that made each item matter, and the division cannot be undone one level up. The same failure appears wherever a tree stores a rate instead of its parts.
The planning side is documented practice: the Association for Financial Professionals publishes an FP&A guide on driver-based models² and tracks integrated planning in its 2026 benchmarking survey,³ and the annual FP&A Trends survey covers the same ground.⁴ None of them settles which accuracy column the tree should hold.
Absolute Error Is the Spine of the Tree
Three columns carry a forecast accuracy tree, and all three are additive. Every rate in the tree is derived from them.
Actual demand is additive. The four SKUs sum to 126,000 units in both quarters, which makes the comparison fair before any rate is displayed.
Absolute error is additive. 8,000 plus 1,600 plus 600 plus 300 gives the first quarter's 10,500 units. 4,900 plus 1,400 plus 660 plus 8,100 gives the second quarter's 15,060.
Signed error is additive and carries different information. The first quarter ran short on every SKU, so its signed errors sum to minus 10,500. The second quarter over-forecast SKU A by 4,900 and ran short on the other three, so its signed errors sum to minus 5,260.
Weighted absolute percentage error is the quotient of the first two. 10,500 over 126,000 is 8.33 percent. 15,060 over 126,000 is 11.95. Both are computed at the level being read, from columns summed to that level, which is why neither can disagree with the rows underneath it.
Why Did MAPE Improve While the Forecast Got Worse?
Because MAPE weights every SKU equally and the demand moved to the inaccurate one. SKU D grew from 1,000 units to 30,000 while carrying the worst error rate in the file. MAPE counts D once, the same as SKU A. Weighted absolute percentage error counts D by its 30,000 units, so it rose 3.62 points.
Run the bridge in weight terms, because the weights are what MAPE throws away.
Weighted absolute percentage error is the demand-weighted average of each SKU's absolute percentage error. In the first quarter SKU A carried 79.37 percent of the weight and SKU D carried 0.79. In the second quarter A carried 55.56 and D carried 23.81, on the same 126,000 units.
Split the 3.62 point rise into two terms. The mix term applies the new weights to the old error rates and contributes plus 5.10 points, of which SKU D supplies 6.90 by itself and SKU A returns 1.90. The rate term applies the improvement every SKU delivered to the new weights and contributes minus 1.48 points. The two sum to 3.62 with nothing left over.
MAPE cannot run that bridge, because it has no weights to move. It reported the rate term, which was real, and was blind to a mix term three and a half times larger.
Total demand was 126,000 units in both quarters, every SKU cut its error rate, and 4,560 more units were forecast wrong. A tree that stores only MAPE reports the 1.50 point improvement and can account for none of it.
| Measure | What it divides | Q1 | Q2 | What it misses |
|---|---|---|---|---|
| MAPE | Each SKU's error rate, averaged over the SKU count | 14.50% | 13.00% | The weights, so demand moving into a weak SKU is invisible |
| WAPE | Summed absolute error over summed actual demand | 8.33% | 11.95% | The direction of the error |
| Net bias, units | Summed signed error, no denominator | -10,500 | -5,260 | Offsetting misses, which cancel inside the sum |
| Absolute error, units | Summed absolute error, no denominator | 10,500 | 15,060 | What a unit of error is worth, which differs by SKU |
| Error cost, dollars | Summed absolute error priced per SKU | 71,100 | 169,200 | Nothing in the plan, which is why it belongs at the root |
Net Bias Hid Two Thirds of the Error
Bias and error answer different questions, and a tree needs both columns.
Net forecast bias is the sum of signed errors. In the second quarter it was minus 5,260 units against the first quarter's minus 10,500, an apparent improvement of 5,240. Read alone it says the plan got closer.
It did not. SKU A was over-forecast by 4,900 units and the other three were short by 10,160. The two directions cancelled inside the sum, 15,060 units were still wrong, and net bias reported about a third of it.
The two directions also cost different things. An over-forecast unit becomes inventory and an eventual markdown. An under-forecast unit becomes a stockout. Netting them prices the two as though they offset, and they do not.
Store both columns. Signed error answers whether the plan leans and absolute error answers what it costs. The two agree only when every item misses in the same direction, which is what the first quarter happened to do. Offsetting of this kind is what the MECE checks on a tree exist to expose.
Where Did the 98,100 Dollars Go?
Into demand moving to the expensive-to-miss SKUs, not into forecasting technique. The volume term adds 117,900 dollars and the rate term returns 19,800, which is what the improvement across all four SKUs was worth. SKU D supplies 130,500 of the volume term by itself, growing from 1,000 units to 30,000 at 15 dollars per unit missed.
Price the error before bridging it, because units of error are not interchangeable.
In the worked plan SKU A and SKU B cost 6.00 dollars per unit missed and SKU C and SKU D cost 15.00. Slower-moving, higher-value items carry more holding cost when over-forecast and more lost margin when short. First quarter error cost 71,100 dollars. Second quarter cost 169,200.
- Volume term. Apply each SKU's first quarter error rate to its change in units. SKU D contributes 8,700 units at 15 dollars, or 130,500. SKU C adds 120 units, or 1,800. SKU A returns 2,400 units at 6 dollars, or minus 14,400. The subtotal is 117,900 dollars on 6,420 units.
- Rate term. Apply each SKU's error rate improvement to its second quarter units. The four return 1,860 units between them, worth 19,800 dollars. The subtotal is minus 19,800.
The two terms sum to 98,100 dollars, and the unit bridge underneath them sums to the 4,560 extra units that were wrong.
Should a KPI Tree Store MAPE at All?
No. MAPE is an average of ratios, so it is not the quotient of two summable columns at any level of the tree, and a node that cannot be recomputed from its children is a label rather than a result. Report MAPE beside the tree if a contract requires it, and never let a parent node inherit it.
Three structural problems follow from one cause.
It cannot be rolled up. The MAPE of a category is not the average of its SKUs' MAPEs unless every SKU sold the same number of units, which no real file does. Regrouping from SKU to category changes the figure, and both figures are arguable, which is worse than one being wrong.
It breaks on small denominators. A SKU with 12 units of demand and a 20 unit forecast scores 66.67 percent error, and a SKU with zero demand in the period returns no figure at all. The M5 competition, run on retail data with heavily intermittent demand, scored entrants on a scaled error rather than a percentage one for exactly this reason.⁵
It is asymmetric. An under-forecast cannot score worse than 100 percent, while an over-forecast has no ceiling, so the measure quietly rewards planning low.⁶
WAPE carries none of the three. It is one division of two sums, so it recomputes at every level and cannot disagree with itself.
Four Tests Before You Trust a Forecast Accuracy Tree
- Additivity test. The sum of SKU absolute errors must equal the group absolute error in every period, before any rate is displayed. A tree that stores percentages cannot run this test at all.
- Recomputation test. Regroup from SKU to category and confirm each accuracy node recomputes as summed absolute error over summed actual demand. If the category figure equals the simple average of its SKUs, the tree is averaging ratios and every parent above it is wrong. The same trap closes on a tree built in a spreadsheet.
- Sign test. Compare the sum of signed errors with the sum of absolute errors in the same period. A gap as wide as the 5,260 against 15,060 above means offsetting misses, and the netted figure should never be shown without the absolute one beside it.
- Constant-mix test. Recompute the current period's accuracy on the prior period's demand weights. The gap between that figure and the reported one is the mix term. If it is larger than the rate term, the move is a volume story and no change of method will fix it.
Where a Forecast Accuracy Tree Is the Wrong Instrument
Four situations defeat this decomposition.
New and discontinued items. A SKU with one period of history has no prior weight, so the mix term absorbs its whole error and reports a change that is really an addition. Restate both periods on the items present in both, then handle the newcomers separately.
Intermittent demand. A SKU that sells in 9 weeks out of 13 produces percentage errors on a denominator that is often zero, which is the condition the forecasting literature flags as a known defect of percentage measures.⁵ Aggregate to a level or a period where demand is continuous, or hold absolute error in units alone.
Horizon questions. The tree explains which items carried the error. It cannot say whether a 13 week lag was the right horizon.
Cause attribution. The tree names the SKU and prices the miss. Whether a promotion, a competitor or a late shipment caused it sits outside the columns, the same boundary that separates a plan model from a driver tree.
Building the Tree in kpitree.io
kpitree.io builds this decomposition from one uploaded CSV. The file needs one row per SKU per period and five summable columns: actual units, forecast units, signed error, absolute error and error cost in dollars.
Inside the tree, parent to child relationships are addition and subtraction, which is what lets four SKU absolute errors produce the group absolute error as a real arithmetic identity rather than a layout convention. Weighted absolute percentage error is a derived node dividing two summed columns at the level being read: absolute error over actual units. Regroup SKU to category to region and it recomputes from the parts each time, because no row holds a ratio.
Error cost is the node worth putting at the root. It is additive, it is denominated in the unit a demand review argues in, and it ranks SKUs by what fixing them is worth rather than by how wrong they look. The smallest useful version is two periods and the dozen SKUs the review already argues about.
Frequently Asked Questions
What is the difference between MAPE and WAPE? MAPE averages each item's error rate, weighting every item equally. WAPE divides summed absolute error by summed actual demand, so each item is weighted by its volume. On the worked second quarter MAPE reads 13.00 percent and WAPE reads 11.95, and only WAPE agrees with the 15,060 units that were wrong.
Is WAPE the same as weighted MAPE? Yes, when the weights are actual demand. The two names describe one calculation.
Should accuracy be reported as accuracy or as error? Store error and derive accuracy. One minus WAPE gives 88.05 percent for the second quarter, but the stored column has to be the error, because error is what sums.
Why does category accuracy change when the tree is regrouped? Because something in the tree is averaging ratios. A node computed as summed absolute error over summed actual demand returns the same figure however the rows are grouped.
Does a lower MAPE mean a better plan? Not by itself. MAPE fell 1.50 points in the worked quarter while the error bill rose 98,100 dollars.
Which accuracy measure belongs at the top of the tree? The cost of the error in dollars. It is additive, and it ranks items by what they are worth fixing.
Closing: Store Error Units, Not Error Rates
MAPE and net bias are the two figures a demand review reads, and neither can explain itself. One averages ratios across items of wildly different size, the other cancels opposite misses inside a sum, and both improved in a quarter where 4,560 more units were forecast wrong.
Store five summable columns per SKU per period. Derive every accuracy rate by dividing two of them at the level being read. Bridge any move into a mix term and a rate term, in units first and dollars second, and check that the two terms sum to the reported change with nothing left over.
Done that way, a quarter where every SKU improved and the error bill rose 98,100 dollars stops being a contradiction. It becomes 29,000 units of demand that moved to the SKU nobody could forecast, priced at 15 dollars a unit missed, next to a rate improvement worth 19,800 that was real and was never going to be enough.
Upload one CSV with your SKUs, two periods and five summable columns, and decompose your own forecast error in kpitree.io.
Sources
- Rob Hyndman on measuring forecast accuracy
- Driver-Based Modelling, FP&A Guide
- 2026 AFP FP&A Benchmarking Survey Report: Integrated Planning
- FP&A Trends Survey 2025: From Ambition to Execution
- M5 accuracy competition: Results, findings and conclusions
- Forecast Evaluation for Data Scientists: Common Pitfalls and Best Practices