level: technical
ibm released granite time series patchtst-fm-r2, a 385 million parameter foundation model for zero-shot forecasting. it handles demand, prices, energy loads, traffic, and telemetry without per-dataset training. the model supports context lengths up to 8,192 steps, flexible forecast horizons, and probabilistic outputs via a 99-quantile prediction head. it is dual-licensed under apache 2.0 and openmdw 1.0, both permissive for commercial use.
on the gift-eval benchmark, patchtst-fm-r2 ranks second among replicable zero-shot models for crps and mase as of september 8, 2026. it achieves a geometric-mean crps of 0.467 and mase of 0.6846. among models with permissive licenses, it is the top performer. even when compared to pretrained models allowed to use benchmark training data, it ranks third for crps and fourth for mase, outperforming larger models like chronos-2 and timer-s1.
the architecture replaces standard transformer layers with conformer blocks that combine self-attention and temporal convolution. this captures both long-range and local patterns, with attention focusing on distant relationships while convolution handles short-term structure. the model uses 50% overlapping patches with hamming-window weighting and overlap-and-add forecasting to smooth predictions. training data includes gift-eval pretrain subsets, synthetic data, and tsmixup corpora, with documentation to support governance reviews.
why it matters: data science teams can deploy a single pretrained model for many forecasting tasks without training separate models, reducing cost and time to production.