How we test forecasts before you rely on them
We measure accuracy on your own history before you see a forecast: we hide recent data, forecast it using only what came before, and compare with what happened, several times over. Every method, from repeating last season to an AI forecasting model, is tested the same way, and the one with the least error is used. Plans are then checked against your current rule of thumb.
Testing on your own history
Forecasts are tested the way they will be used. If you plan 4 weeks ahead, each test hides 4 weeks, forecasts them from the data before, and scores the result. We repeat this over several starting points (rolling tests), so one lucky or unlucky month does not decide the choice. Data from the hidden stretch is never used to fit or choose a model.
Comparing with simple rules
A forecast is only useful if it beats what you could do without it. Every test includes simple rules: repeating the last value, and repeating the last season, such as the same weekday last week or the same month last year. Our reports show the improvement over repeating the last season. When nothing beats the simple rule, we say so and recommend the simple rule.
The methods we test
- Simple rules: repeat the last value; repeat the last season.
- Classic statistics: exponential smoothing, the Theta method, and Croston's method for items that sell only now and then.
- An AI forecasting model: a pretrained time series model (Chronos-2) that has learned patterns from a large collection of public and synthetic series and is applied to your data without training on it. When useful inputs exist, such as promotions, prices, holidays or weather, it can use them too.
Reality checks on data nobody saw
In our case studies we also set aside a final stretch of data before doing anything. After the method is chosen, we score every method on it. This catches choices that looked good in testing but do not hold up. We publish these results even when a simpler method wins, as it did in some of our case studies.
Ranges, not just one number
Every forecast comes with a range that should contain the actual value about 8 times out of 10. We check how often it really does in testing and report it, so you know whether the range is honest.
From forecast to plan
When your question is a decision, we build an optimization model of it. It knows what each choice costs: running short, holding stock, wasting product, overtime, changeovers, battery wear. It knows your limits: pack sizes, delivery days, line capacity, shift patterns, shelf life. Instead of planning for a single forecast, it plans across many possible futures drawn from the forecast range, and picks the plan with the lowest expected total cost. Every plan is compared with your current rule of thumb, both on the expected numbers and, where the data allows, replayed on what actually happened.
Before a plan goes out, automatic checks confirm it respects every limit and that the math solved properly. Inputs we had to assume are labeled as assumptions in the report.
When we flag a result for review
Some data is not ready for an automatic answer: very short history, many missing periods, items that almost never sell, or accuracy that is too low to act on. Then the result is marked for review and a person looks at it before you do. Our Iowa wholesale case study shows what that looks like.
Terms we use, for the curious
- Typical error (WAPE)
- Weighted absolute percentage error: the total of all forecast misses divided by the total actual volume. 10% means the forecast missed by 10% of what really happened. It weights big days and big items more, which matches how they affect your business.
- MASE
- Mean absolute scaled error: the forecast error divided by the error of a simple rule. Below 1 means better than the simple rule. We report it alongside WAPE.
- Seasonal naive
- The "repeat last season" rule: the forecast for next Tuesday is last Tuesday.
- Backtest
- Testing a forecasting method on past data as if it were the future.
- Holdout
- A final stretch of data set aside before any testing, used once at the end as a reality check.
- 80% prediction interval
- The range the forecast expects to contain the actual value 8 times out of 10.
- Linear and mixed-integer programming (LP and MILP)
- The math behind the plans. A solver finds the best decision (for example, how many cases to order on each delivery day) subject to your limits. Mixed-integer means some decisions must be whole numbers, like full cases or whole shifts.
- Scenario-based planning
- Instead of one forecast, the plan is evaluated across many possible demand paths sampled from the forecast range, so it is robust when demand is higher or lower than expected.