Results: where we won, and where we did not
Across 13 public case studies, our forecasts had less error than simply repeating last season in 13 of 13, with a median of 25% less error. They did not win everywhere: an existing grid forecast was more accurate in one case, simpler methods won some reality checks, and one plan cost more than a rule of thumb. All of it is below.
- 13 of 13 beat repeating last season in testing on past data
- 11 of 13 won by the AI forecasting model; classic statistics won the other 2
- 3 of 4 plans saved money against a rule of thumb when replayed on real data (example costs)
Every case study
| Case study | Selected method | Typical error | Repeat last season | Improvement | Reality check | Plan vs rule of thumb |
|---|---|---|---|---|---|---|
| Yogurt production planGreece · daily, 28 days ahead | AI model with extra inputs | 19.3% | 26.8% | 28% | Selected method held up (23%) | Saved $21,369 (9.5%) |
| Fresh food reorder planChina · daily, 7 days ahead | AI model with extra inputs | 13% | 14.9% | 13% | Theta method did better (10.4% vs 11.2%) | Saved $102 (0.2%) |
| Pizza shop ingredient orders SyntheticUnited States (synthetic) · daily, 14 days ahead | Exponential smoothing | 11.9% | 15.5% | 23% | Selected method held up (10.2%) | Cost $494 (9.7%) more |
| Contact center staffing planSan Francisco, United States · daily, 42 days ahead | AI forecasting model | 9.8% | 13.4% | 27% | Exponential smoothing did better (9.9% vs 10.3%) | Saved $9,091 (2.1%) |
| Grid demand and battery scheduleUnited States · hourly, 7 days (168 hours) ahead | AI forecasting model | 5.7% | 7.2% | 21% | Selected method held up (7%) | 29.2% more value |
| Online store revenueUnited Kingdom · daily, 30 days ahead | AI forecasting model | 31.4% | 39.5% | 20% | Selected method held up (17.2%) | No plan |
| Wholesale orders for 500 store-product pairs Flagged for reviewIowa, United States · weekly, 8 weeks ahead | AI forecasting model | 64.6% | 86.5% | 25% | Not run | No plan |
| Hotel room-nightsPortugal · daily, 60 days ahead | AI forecasting model | 29.6% | 46.8% | 37% | Not run | No plan |
| Airport passengers by airlineSan Francisco, United States · monthly, 12 months ahead | AI forecasting model | 11.8% | 19.1% | 38% | Not run | No plan |
| Bike rentals with weatherWashington, DC, United States · hourly, 7 days (168 hours) ahead | AI model with extra inputs | 28.6% | 62.7% | 54% | Selected method held up (13.3%) | No plan |
| Messy POS export Synthetic Messy fileUnited States (synthetic) · daily, 14 days ahead | Exponential smoothing | 15% | 19.5% | 23% | Not run | No plan |
| Messy monthly airline file Messy fileSan Francisco, United States · monthly, 12 months ahead | AI forecasting model | 12.1% | 19.1% | 36% | Not run | No plan |
| Messy daily rentals file Messy fileWashington, DC, United States · daily, 14 days ahead | AI forecasting model | 27.4% | 31.7% | 14% | Not run | No plan |
Typical error is the share of actual volume the forecast missed by (WAPE), averaged over several tests on past data. Improvement is relative to repeating the last season (for example, the same weekday last week). Reality checks use a final stretch of data that no method saw. Plan amounts use example costs; for the battery, the percentage is relative to the value the rule of thumb creates.
Where the simple approach won
These are the results a sales page would leave out. We think they are the reason to trust the rest.
- Grid demand and battery schedule: the existing forecast was more accurate
The grid operators' own day-ahead forecast had 3% typical error against 5.7% for ours. It is issued about a day ahead, while our test forecasts look up to 7 days ahead, so the comparison favors it.
- Pizza shop ingredient orders: the rule of thumb cost less than our plan
Replayed on the 14 held-back days, ordering the forecast plus 20% cost $494 (9.7%) less than the recommended plan, although the plan expected to save money. The data is synthetic and the costs are examples.
- Fresh food reorder plan: a simpler method won the reality check
On data no method had seen, theta method had 10.4% error against 11.2% for the method we selected from repeated tests.
- Contact center staffing plan: a simpler method won the reality check
On data no method had seen, exponential smoothing had 9.9% error against 10.3% for the method we selected from repeated tests.
- Fresh food reorder plan: the plan barely beat the rule
The plan saved $102 (0.2%) on held-back data. When a simple rule is already close to optimal, we will tell you to keep it.
- Pizza shop ingredient orders: classic statistics beat the AI model
Exponential smoothing had the lowest error in testing (11.9%). We pick whatever tests best on your data, not the newest model.
- Messy POS export: classic statistics beat the AI model
Exponential smoothing had the lowest error in testing (15%). We pick whatever tests best on your data, not the newest model.
- Wholesale orders for 500 store-product pairs: flagged for review instead of sent
137 rows have a negative value (for example returns or refunds). They were netted against the other rows in the same period.
How to read these numbers
A typical error of 10% means that, added up over the period, the forecast missed by 10% of what really happened. What counts as good depends on the business: daily sales of a single fresh item are much harder to forecast than monthly passengers at an airport. That is why we always compare with a simple rule on the same data, and why we test on your history before you rely on anything.
The full method is on the methodology page. The raw outputs for every case are in the 4castPlannr repository.