Forecasting from a messy POS export: mixed dates, duplicates and refunds

Case study on a synthetic public dataset from United States · Updated

We forecast daily dollars in sales 14 days ahead from a synthetic public dataset. In testing on past data, the selected method (exponential smoothing) was off by 15% of actual volume on average, 23% less error than repeating the same weekday last week.

Read this first

  • This dataset is synthetic: it was generated to look like a year of sales at a US pizza place. It is not a real business.
  • This file was deliberately broken the way real exports break, to test that the engine cleans it without help.

The request

"Here's our POS export. Forecast daily sales in dollars for the next two weeks."

The request, written the way a business owner would ask it.

The forecast

We expect about 32,014 dollars in sales over the next 14 days, about 2,287 per day on average.

Line chart of daily dollars in sales: the actual history, the forecast for the next 14 days with its 80% range.
The forecast for the next 14 days, with the range it expects 8 times out of 10.

How accurate was it?

We hid the most recent stretch of history, forecast it using only what came before, and compared with what really happened. We did that 5 times, each time 14 days ahead, for every method below. The typical error is the share of actual volume the forecast missed by.

The selected method, exponential smoothing, a classic statistical method, was off by 15% on average, against 19.5% for simply repeating the same weekday last week.

Typical error in testing, lower is better
MethodKindTypical errorActual inside 80% range
Exponential smoothing SelectedClassic15%74%
AI forecasting modelAI15.3%60%
Theta methodClassic15.9%79%
Repeat the last valueSimple rule18.6%94%
Repeat last seasonSimple rule19.5%84%
Line chart comparing past forecasts with actual dollars in sales over 5 test runs of 14 days each, made using only the data available at the time.
Back-testing: 5 times we hid the next 14 days, forecast them, and compared with what happened.
Bar chart of the typical error of each forecasting method in testing. The selected method had 15% error; repeating the same weekday last week had 19.5%.
Typical error of each method in testing, as a share of actual volume. Lower is better.

What the engine noticed in the data

  • Dates were written in 4 different formats; all were converted.
  • 478 rows have a negative value (for example returns or refunds). They were netted against the other rows in the same period.
  • Removed 256 exact duplicate rows (column 'rownames' should be unique).
  • 24 days with no rows were filled by interpolation, because dollars in sales is otherwise never close to zero (more likely missing data than a closed business). If you were closed on those days, tell us and we will treat them as zero.
  • 24 days (6.6%) had no data and were filled as described above.

Technical details

Data
Pizza Place Sales (gt::pizzaplace), United States (synthetic). 1 series, daily, from 2015-01-01 to 2015-12-31.
How we prepared the data
Intentionally dirtied copy of pizza-place-sales. What was broken on purpose: Semicolon separators and decimal commas; four different date formats mixed row by row (2015-01-01, 01/01/2015, 01 Jan 2015, Jan 01 2015); 5% of trading days missing; about 2% of lines duplicated exactly (double scans); about 1% refunds as negative-price lines. The spec is the same kind of spec a clean file would get: no per-case cleaning code.
Testing
5 rolling tests, each 14 days ahead. Selection metric: WAPE (weighted absolute percentage error, the "typical error" above). Also reported: MASE 1.06 for the selected method.
Methods
Exponential smoothing (ETS), Chronos-2 (AI foundation model), Theta, Naive (last value), Seasonal naive.
Reproduce
The case folder, spec and outputs are in the 4castPlannr repository under cases/pizza-place-sales-dirty/.

Data source and license

pizzaplace dataset from the gt R package, Copyright (c) 2018-2026 Posit Software, PBC, MIT License.

License: MIT (gt R package by Posit; data is part of the package). Source: original data. The forecasts and charts on this page are derived from that data and carry the same attribution.

Have a decision like this?

This case is an example of restaurant demand for restaurants and food service. Send us your own export and question, and we will run the same tests on your data.

Related case studies

  • Pizza shop ingredient orders

    A synthetic pizza shop: a 14-day sales forecast and a two-week order plan for dough, cheese, sauce and boxes. The rule of thumb won the reality check.

    United States (synthetic) · 11.9% error · 23% better than repeating last season · synthetic data