Adapting a compact time series foundation model for realized volatility
VolaTTM fine-tunes one IBM Granite Tiny Time Mixer on cross-asset realized measures from VOLARE. The model improves macro RMSE and QLIKE over rolling HAR and Log-HAR at 1, 5, and 22 sessions, while Log-HAR retains lower MAE.
Abstract
This study evaluates whether a compact pretrained time series model can improve cross-asset realized-volatility forecasts relative to rolling econometric benchmarks. A single IBM Granite Tiny Time Mixer R2.1 model is fine-tuned on 5-minute realized variance and related end-of-day measures from the VOLARE database. Training targets end before 2024, model selection uses 2024, and evaluation covers January 2025 through June 2026. VolaTTM obtains lower macro RMSE and QLIKE than HAR-RV and Log-HAR at all three horizons. The QLIKE result is broad across assets, but the experiment is not a trading backtest and its evaluation period is not a fully project-blind holdout.
Motivation
Realized volatility is persistent, heterogeneous across horizons, asymmetric, and sensitive to market regime. HAR models encode daily, weekly, and monthly components explicitly and remain difficult baselines to beat. Time series foundation models offer a different prior: reusable temporal representations learned before the target task, followed by inexpensive adaptation.
The relevant question is not whether a neural model can fit volatility. It is whether one compact pretrained model can add robust out-of-sample information beyond direct rolling HAR forecasts under a defensible temporal protocol.
Data and target
The experiment uses the realized-variance archives distributed by VOLARE. The download contains 158,887 daily observations across 40 equities, 5 foreign-exchange rates, and 5 futures, ending on 30 June 2026. VOLARE derives the measures from high-frequency tick data using an asset-specific cleaning and sampling pipeline.
The target is rv5, realized variance computed from 5-minute returns. Forecasts are point targets at 1, 5, and 22 observed sessions, not averages over future windows. The model observes eleven additional origin-day channels including realized kernel, bipower variation, semivariances, quarticity, range, returns, volume, and trade counts.
Model and training
VolaTTM starts from IBM's daily-frequency TTM R2.1 branch with a 512-step context. The forecast-channel-mixing decoder is enabled so that the target can use the other realized-measure channels. This remains one sub-million-parameter model, not an ensemble or mixture of experts.
The objective assigns weight 0.50 to Smooth L1, 0.30 to path QLIKE, and 0.20 to QLIKE at the reported horizons. Training uses AdamW, OneCycle scheduling, mixed precision, gradient clipping, asset-balanced sampling, and three random seeds. Seed 17 at epoch 3 was selected from 2024 only.
Econometric comparisons
HAR-RV and Log-HAR are the primary baselines. Separate direct models are estimated for every asset and horizon, with parameters re-estimated at each forecast origin on the latest 1,000 eligible sessions. Persistence is a naive reference. EWMA and GARCH are secondary controls because they use close-to-close returns rather than the same realized-measure information set. Corporate-action discontinuities in archive close prices further limit their interpretation.
Results
Scores below are macro averages across assets on 2025 through June 2026. MAE and RMSE use annualized volatility. QLIKE uses variance. Lower is better.
| Horizon | Model | MAE | RMSE | QLIKE |
|---|---|---|---|---|
| 1 | HAR-RV | 0.05164 | 0.08592 | 0.22188 |
| 1 | Log-HAR | 0.04958 | 0.08610 | 0.23887 |
| 1 | VolaTTM | 0.05192 | 0.08505 | 0.21483 |
| 5 | HAR-RV | 0.06134 | 0.10099 | 0.33504 |
| 5 | Log-HAR | 0.05832 | 0.10030 | 0.36285 |
| 5 | VolaTTM | 0.06173 | 0.09996 | 0.31054 |
| 22 | HAR-RV | 0.06686 | 0.10678 | 0.39199 |
| 22 | Log-HAR | 0.06433 | 0.10644 | 0.43810 |
| 22 | VolaTTM | 0.06466 | 0.10262 | 0.34682 |
Against Log-HAR, VolaTTM has lower QLIKE for 50 of 50 assets at one session, 49 of 50 at five sessions, and 47 of 50 at 22 sessions. Diebold-Mariano tests with Newey-West correction identify 30, 23, and 26 significant improvements at the 5 percent level and no significant losses. These p-values are not adjusted for multiple comparisons.
Negative finding
A multiplicative level calibration was fitted per asset and horizon using 2024 only. It reduced some absolute errors but degraded QLIKE after the regime change. It was rejected and is not packaged with the released model. Validation-period recalibration should not be assumed to transfer robustly.
Leakage audit and limitations
Executable tests enforce disjoint target dates across training, validation, and evaluation. Inputs end at the forecast origin and targets begin on the next session. HAR training outcomes are observed no later than the current origin. The released checkpoint is selected without reading its evaluation loss.
Further limitations include universe-selection and survivorship effects, heterogeneous market calendars, unadjusted close prices in secondary return baselines, uncorrected multiple tests, and the absence of transaction costs or portfolio utility. Forecast accuracy does not imply trading profitability.
Conclusion
This experiment provides evidence that a focused, sub-million-parameter time series foundation model can complement classical volatility models. VolaTTM improves RMSE and QLIKE consistently across the reported horizons, while Log-HAR remains preferable under MAE. The result supports prospective evaluation rather than a claim of universal dominance.
Resources and references
- VolaTTM model and model card
- Source code and leakage audit
- Ekambaram et al. (2024), Tiny Time Mixers.
- Cipollini et al. (2026), VOLatility Archive for Realized Estimates.
- Ambach et al. (2026), Forecasting Realized Volatility with Time Series Foundation Models.
