Results: where we won, and where we did not

Across 13 public case studies, our forecasts had less error than simply repeating last season in 13 of 13, with a median of 25% less error. They did not win everywhere: an existing grid forecast was more accurate in one case, simpler methods won some reality checks, and one plan cost more than a rule of thumb. All of it is below.

Every case study

Forecast accuracy in testing (typical error, lower is better) and plan results on held-back data
Case studySelected methodTypical errorRepeat last seasonImprovementReality checkPlan vs rule of thumb
Yogurt production planGreece · daily, 28 days ahead AI model with extra inputs 19.3% 26.8% 28% Selected method held up (23%) Saved $21,369 (9.5%)
Fresh food reorder planChina · daily, 7 days ahead AI model with extra inputs 13% 14.9% 13% Theta method did better (10.4% vs 11.2%) Saved $102 (0.2%)
Pizza shop ingredient orders SyntheticUnited States (synthetic) · daily, 14 days ahead Exponential smoothing 11.9% 15.5% 23% Selected method held up (10.2%) Cost $494 (9.7%) more
Contact center staffing planSan Francisco, United States · daily, 42 days ahead AI forecasting model 9.8% 13.4% 27% Exponential smoothing did better (9.9% vs 10.3%) Saved $9,091 (2.1%)
Grid demand and battery scheduleUnited States · hourly, 7 days (168 hours) ahead AI forecasting model 5.7% 7.2% 21% Selected method held up (7%) 29.2% more value
Online store revenueUnited Kingdom · daily, 30 days ahead AI forecasting model 31.4% 39.5% 20% Selected method held up (17.2%) No plan
Wholesale orders for 500 store-product pairs Flagged for reviewIowa, United States · weekly, 8 weeks ahead AI forecasting model 64.6% 86.5% 25% Not run No plan
Hotel room-nightsPortugal · daily, 60 days ahead AI forecasting model 29.6% 46.8% 37% Not run No plan
Airport passengers by airlineSan Francisco, United States · monthly, 12 months ahead AI forecasting model 11.8% 19.1% 38% Not run No plan
Bike rentals with weatherWashington, DC, United States · hourly, 7 days (168 hours) ahead AI model with extra inputs 28.6% 62.7% 54% Selected method held up (13.3%) No plan
Messy POS export Synthetic Messy fileUnited States (synthetic) · daily, 14 days ahead Exponential smoothing 15% 19.5% 23% Not run No plan
Messy monthly airline file Messy fileSan Francisco, United States · monthly, 12 months ahead AI forecasting model 12.1% 19.1% 36% Not run No plan
Messy daily rentals file Messy fileWashington, DC, United States · daily, 14 days ahead AI forecasting model 27.4% 31.7% 14% Not run No plan

Typical error is the share of actual volume the forecast missed by (WAPE), averaged over several tests on past data. Improvement is relative to repeating the last season (for example, the same weekday last week). Reality checks use a final stretch of data that no method saw. Plan amounts use example costs; for the battery, the percentage is relative to the value the rule of thumb creates.

Where the simple approach won

These are the results a sales page would leave out. We think they are the reason to trust the rest.

How to read these numbers

A typical error of 10% means that, added up over the period, the forecast missed by 10% of what really happened. What counts as good depends on the business: daily sales of a single fresh item are much harder to forecast than monthly passengers at an airport. That is why we always compare with a simple rule on the same data, and why we test on your history before you rely on anything.

The full method is on the methodology page. The raw outputs for every case are in the 4castPlannr repository.

See how it does on your data