Google has released TimesFM-3, a model it says forecasts your sales, stock, and revenue without ever training on your data. The claim is bold: next month's table with zero training on your own records. The video puts that claim to the test on daily sales data from an ice-cream company.
TimesFM is not a language model but a time-series foundation model. The logic is familiar: where Gemini guesses the next word from half a sentence, TimesFM looks at a chart it has never seen and draws what comes next. The first version arrived in 2024, fed on 100 billion time points: trend searches, page views, traffic, weather, sales, and synthetic data.
The new version scales two things up: 330 million parameters and a training corpus past one trillion points. The real break is scope: earlier versions read a single signal, while TimesFM-3 jointly uses related product sales, past foot traffic, and known future information such as weather outlooks, holidays, and planned promos.
To push the model until it breaks, the presenter wrote his own app: 156 days of ice-cream sales history, the first 128 days as context and the last 28 as the forecast horizon. The columns carry scoops sold, cone and syrup sales, a promo flag, temperature, and store traffic. Scores are tracked with MAPE, MAE, and WAPE, and each extra feature's revenue delta is computed.
With promo, weather, and traffic signals switched off, the model settles around 4.6 to 4.7 MAPE. The chart looks decent but misses the sharp spikes from mango promos. Past sales alone cannot explain campaign weeks.
Flipping the promo signal on changes the picture: the model catches the campaign spikes, MAPE drops visibly, and the promos' contribution to revenue comes into view. Adding weather and traffic moves the forecast again, and the presenter notes how September-October cold suppresses ice-cream sales. Combined signals clearly beat any single one.
In the head-to-head, zero-shot TimesFM scores 4.73 MAPE while Holt-Winters lands at 8.89 and the seasonal naive method at 6.72. The gaps are clean: 4.16 points better than Holt-Winters, 1.99 better than naive. And a single forecast arrives in 203 milliseconds, with zero training.
The inside of the box gets its own tour: instead of reading day by day, the model compresses 128 days into four 32-day chunks with a technique called continuous patching. The analogy lands: reading sentence by sentence instead of letter by letter. Noisy daily wiggles get filtered out, the broad trend steps forward, and that is where the speed comes from.
The architecture runs on two-way attention: first every row studies only its own past, mathematically barred from peeking ahead, which is how the organic baseline is learned. Then attention turns down the column: with September sales still empty, the planned promo discount and the weather outlook flow upward from the same column and lift the forecast. The cycle repeats twenty times through the transformer stack — the presenter likens it to a planner reading past charts, then glancing at the campaign schedule on the desk.
The verdict stays balanced: for a company with no forecasting team, this is more than an upgrade — it is the difference between having forecasts and not having them. But where half a point of error costs millions, large retailers will keep training custom models, because supply constraints, product cannibalization, and pricing strategy stay invisible to this model. Practical notes: the code is on GitHub, the weights ship under a non-commercial license, and the commercial route goes through the BigQuery integration.
| Model | MAPE | Gap (pts) |
|---|---|---|
| TimesFM-3 (zero-shot) | 4.73 | — |
| Seasonal naive | 6.72 | +1.99 |
| Holt-Winters | 8.89 | +4.16 |
AI commentary
"I stress-tested the zero-shot claim so you do not have to: the core holds, the showcase extras depend on your data. My rule is simple — set the baseline with TimesFM first, spend your team only where it beats that line."
AI assessment
The strongest objection first, in its most generous form: this duel ran on one ice-cream dataset, in one forecast window, inside the presenter's own app. An independent test with rolling windows and more than twenty thousand forecasts confirms the zero-shot claim — yet on that tester's data the new multivariate extras made no difference. So the core is solid, the showcase features are data-dependent.
What the video leaves untested is a long list: supply constraints, products cannibalizing each other, and pricing decisions never enter the model. Single-window luck matters too — in the independent test, shifting the window moved classical scores by ten percent. Cost at scale, latency budgets, and the calibration of probabilistic outputs are all left open.
On verification there are two layers: the code is Apache-2.0, but the weights are limited to non-commercial use, and the commercial route runs through Google's data platform. The demo MAPE and the promos' revenue contribution both deserve an independent rerun on your own data before they enter a decision meeting.
My practical take: for a business that cannot staff a forecasting team, this model is an instant baseline; for a team already on BigQuery, it is a forecast in a few lines of SQL. But where half a point of error writes millions, custom training stays mandatory. I would set the baseline with this model first, then spend the team's time only on work that beats it.
Sources
7 links; 2 of them also cited by 1 other story. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
- @youtube AI with Surya — TimesFM-3 vs Classical Forecasting
- @research.google https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/
Also cited by: Top 10 Trending GitHub Repos: From Archify to Omarchy and TimesFM — Visual Code to Agentic Linux
- @github https://github.com/google-research/timesfm
Also cited by: Top 10 Trending GitHub Repos: From Archify to Omarchy and TimesFM — Visual Code to Agentic Linux
- @cloud.google https://cloud.google.com/blog/products/data-analytics/timesfm-models-in-bigquery-and-alloydb
- @cloud.google https://docs.cloud.google.com/bigquery/docs/timesfm-model
- @dev.to https://dev.to/han-co/i-tested-googles-timesfm-3-the-zero-shot-claim-holds-the-new-features-dont-4ol4
- @andrew.ooo https://andrew.ooo/posts/timesfm-google-time-series-foundation-model-review/
timesfm · forecasting · google · ai · time series · bigquery