Forecasting from a messy spreadsheet: mixed date styles and missing values

Case study on a public dataset from Washington, DC, United States · Updated

We forecast daily bike rentals 14 days ahead from a public dataset. In testing on past data, the selected method (the AI forecasting model) was off by 27.4% of actual volume on average, 14% less error than repeating the same weekday last week.

Read this first

  • This file was deliberately broken the way real exports break, to test that the engine cleans it without help.

The request

"Forecast daily rentals for the next 2 weeks. Sorry, the file is a bit of a mess."

The request, written the way a business owner would ask it.

The forecast

We expect about 53,850 bike rentals over the next 14 days, about 3,846 per day on average.

Line chart of daily bike rentals: the actual history, the forecast for the next 14 days with its 80% range.
The forecast for the next 14 days, with the range it expects 8 times out of 10.

How accurate was it?

We hid the most recent stretch of history, forecast it using only what came before, and compared with what really happened. We did that 5 times, each time 14 days ahead, for every method below. The typical error is the share of actual volume the forecast missed by.

The selected method, an AI forecasting model that was not trained on this data, was off by 27.4% on average, against 31.7% for simply repeating the same weekday last week.

Typical error in testing, lower is better
MethodKindTypical errorActual inside 80% range
AI forecasting model SelectedAI27.4%74%
Repeat the last valueSimple rule28.4%91%
AI model with extra inputsAI28.9%70%
Theta methodClassic29.2%61%
Exponential smoothingClassic30.4%73%
Repeat last seasonSimple rule31.7%69%
Line chart comparing past forecasts with actual bike rentals over 5 test runs of 14 days each, made using only the data available at the time.
Back-testing: 5 times we hid the next 14 days, forecast them, and compared with what happened.
Bar chart of the typical error of each forecasting method in testing. The selected method had 27.4% error; repeating the same weekday last week had 31.7%.
Typical error of each method in testing, as a share of actual volume. Lower is better.

What the engine noticed in the data

  • Dates were written in 2 different formats; all were converted.
  • Removed 21 exact duplicate rows.

Technical details

Data
Bike Sharing Dataset (UCI ML Repository #275; Capital Bikeshare, Washington DC). 1 series, daily, from 2011-01-01 to 2012-12-31.
How we prepared the data
Intentionally dirtied copy of uci-bike-sharing. What was broken on purpose: Tab separated; dates switch from 2011-01-01 style to day/month/year (31/12/2012) halfway through; 4% of days missing; 3% of rows duplicated; 5% of temperatures are 'N/A'; counts written with thousands separators ('1,234'). The spec is the same kind of spec a clean file would get: no per-case cleaning code.
Testing
5 rolling tests, each 14 days ahead. Selection metric: WAPE (weighted absolute percentage error, the "typical error" above). Also reported: MASE 1.42 for the selected method.
Methods
Chronos-2 (AI foundation model), Naive (last value), Chronos-2 with covariates, Theta, Exponential smoothing (ETS), Seasonal naive.
Reproduce
The case folder, spec and outputs are in the 4castPlannr repository under cases/uci-bike-sharing-dirty/.

Data source and license

Fanaee-T, H. & Gama, J. (2013). Event labeling combining ensemble detectors and background knowledge. Progress in AI, doi:10.1007/s13748-013-0040-3. Data: UCI ML Repository, https://doi.org/10.24432/C5W894, CC BY 4.0. Raw trip data: Capital Bikeshare.

License: CC BY 4.0. Source: original data. The forecasts and charts on this page are derived from that data and carry the same attribution.

Have a decision like this?

This case is an example of staffing plans for call centers and service businesses. Send us your own export and question, and we will run the same tests on your data.

Related case studies

  • Contact center staffing plan

    San Francisco 311 public data: daily phone, web and app requests forecast six weeks ahead, and an agent shift plan replayed on the weeks that followed.

    San Francisco, United States · 9.8% error · 27% better than repeating last season

  • Hotel room-nights

    A public dataset from two hotels: daily room-nights forecast 60 days ahead per hotel, with the accuracy we measured in testing.

    Portugal · 29.6% error · 37% better than repeating last season

  • Airport passengers by airline

    San Francisco International Airport's public statistics: monthly passengers for the 10 largest airlines forecast a year ahead.

    San Francisco, United States · 11.8% error · 38% better than repeating last season