# Position: The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting

Qinwei Ma Affiliation: Tsinghua University Correspondence to: <qinweimartin@gmail.com>    Jingzhe Shi Affiliation: Tsinghua University Correspondence to: <shi-jz21@mails.tsinghua.edu.cn>    Jiahao Qiu Affiliation: Princeton University    Zaiwen Yang Affiliation: Tsinghua University

###### Abstract

Recent work has questioned the effectiveness and robustness of neural network architectures for time series forecasting tasks. We summarize these concerns and analyze groundly their inherent limitations: i.e. the irreconcilable conflict between single (or few similar) domains SOTA and generalizability over general domains for time series forecasting neural network architecture designs. Moreover, neural networks architectures for general domain time series forecasting are becoming more and more complicated and their performance has almost saturated in recent years. As a result, network architectures developed aiming at fitting general time series domains are almost not inspiring for real world practices for certain single (or few similar) domains such as Finance, Weather, Traffic, etc: each specific domain develops their own methods that rarely utilize advances in neural network architectures of time series community in recent 2-3 years. As a result, we call for the time series community to shift focus away from research on time series neural network architectures for general domains: these researches have become saturated and away from domain-specific SOTAs over time. We should either (1) focus on deep learning methods for certain specific domain(s), or (2) turn to the development of meta-learning methods for general domains.

###### Keywords:

Machine Learning, ICML <sup>††</sup>affiliationnotice: \* Equal contribution and equal correspondence.

## 1 Introduction

Time series forecasting is a cornerstone of decision-making in high-stakes environments, including financial markets, energy grid management, climate science, and healthcare. Driven by the transformative success of large-scale pre-training and universal architectures in Natural Language Processing (NLP) and Computer Vision (CV), the time series research community has recently pivoted toward a similar paradigm ([Liang et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib24)). The prevailing trend assumes that increasing model capacity and designing complex, “general-purpose” neural architectures will eventually yield a foundation model for forecasting ([Gao et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib25); [Das et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib30)).

However, we argue that this pursuit overlooks a fundamental structural difference between time series and other common domains. Despite the proliferation of sophisticated architectures, a significant gap remains between academic benchmarks and real-world deployment: when state-of-the-art (SOTA) performance is required for specific domains, practitioners consistently favor domain-specific models over the “universal” architectures promoted in the literature. This suggests that the current research trajectory may be reaching a point of diminishing returns ([Brigato et al., 2026](https://arxiv.org/html/2602.01736v1#bib.bib73); [Wang et al., 2025b](https://arxiv.org/html/2602.01736v1#bib.bib74)).

The central thesis of this paper is that the community faces a fundamental conflict of domain-specific excellence and cross-domain generalizability. Moreover, such conflict is irreconcilable for general-domain time series forecasting. This tension is rooted in two primary constraints:

1.  1\.

    Inherent Domain Heterogeneity: Unlike in NLP where each domain shares similar linguistic structures, time series data across different sectors, such as wind power fluctuations versus high-frequency limit order books, are governed by entirely different physical processes and causal structures. A search for a single architecture that is capable of capturing these disparate “logics” optimally is likely a misplaced endeavor ([Pfister, 2024](https://arxiv.org/html/2602.01736v1#bib.bib31)).

2.  2\.

    Statistical Limits of Temporal Data: Time series data is intrinsically constrained by the temporal window of observation. While CV and NLP benefit from nearly infinite spatial or linguistic scaling, time series scaling is bounded by history. Theoretically, for a data duration $T$, the generalization error is lower-bounded by a factor proportional to $1/\sqrt{T}$. No architectural innovation can circumvent this statistical bottleneck ([Kuznetsov and Mohri, 2014](https://arxiv.org/html/2602.01736v1#bib.bib62); [Shi et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib16)).

![Figure 1 : Left: A mind map of the core arguments in this position paper. We propose that there exists a inherent, irreconcilable conflict between general-domain and single-domain performance for time series forecasting neural network architecture designing. We analyze its causes, consequences and advocated better research directions. We discuss more about Alternative Views in Section 6 . Right: Error reached/reachable for a specific domain (e.g. weather forecasting) when utilizing different types of methods with different available domain-specific context and properties.](figures/intro_figure_2.png)

Figure 1: Left: A mind map of the core arguments in this position paper. We propose that there exists a inherent, irreconcilable conflict between general-domain and single-domain performance for time series forecasting neural network architecture designing. We analyze its causes, consequences and advocated better research directions. We discuss more about Alternative Views in Section [6](https://arxiv.org/html/2602.01736v1#S6). Right: Error reached/reachable for a specific domain (e.g. weather forecasting) when utilizing different types of methods with different available domain-specific context and properties.

We argue that it is such conflict that has caused many peculiar phenomena observed in time series communities ([Bergmeir, 2024b](https://arxiv.org/html/2602.01736v1#bib.bib79); [Brigato et al., 2026](https://arxiv.org/html/2602.01736v1#bib.bib73)), and has been the reason why advances in general-domain TSF neural networks barely benefits researchers in those specific domains (e.g. Traffic, Weather, Electricity, etc.). In this paper, we advocate for a strategic redirection of the field. As illustrated in Figure [1](https://arxiv.org/html/2602.01736v1#S1.F1), we summarize our core arguments and the resulting research redirection. Specifically, we propose that:

- •

  The community should acknowledge the “seesaw” conflict between domain SOTA and cross-domain generalization. General-purpose architectures that are “SOTA” on benchmark suites still lag domain-specific models; given that practitioners can often deploy domain SOTA directly, we should stop optimizing general-purpose architectures as an end in itself.

- •

  Research effort should shift to either (1) online meta-learning and continual adaptation that enable models to “learn how to learn” from limited domain context (e.g., LLM-as-scientist), or (2) domain-specific forecasting models that explicitly incorporate domain context and mechanisms (finance, weather, traffic, etc.).

By pivoting from architecture search to meta-adaptation, our community can bridge the gap between theoretical research and the practical necessity of domain-specific precision.

## 2 Background

Since 2021, the research community has made great effort in developing general domain time series forecasting neural network architectures. Starting from paradigm and datasets proposed in Informer ([Zhou et al., 2021](https://arxiv.org/html/2602.01736v1#bib.bib1)), many neural networks architectures have been proposed for general domain time series forecasting performance. A common practice for these papers is to propose a neural network architecture to achieve as high performance as possible across a wide range of datasets of time series like Electricity, Traffic, Weather, etc.

### 2.1 Neural Network Methods for General-domain Time Series Forecasting

In 2021, Informer ([Zhou et al., 2021](https://arxiv.org/html/2602.01736v1#bib.bib1)) set up a multi-dataset paradigm (e.g. ETT, ECL, etc.) for neural network based general domain long-horizon forecasting, after which a burst of “general-domain time series forecasting” architectures were proposed under a mostly unified task protocol. Beyond Transformer variants (decomposition, frequency, patching, or de-stationary attention) ([Wu et al., 2022b](https://arxiv.org/html/2602.01736v1#bib.bib12); [Zhou et al., 2022](https://arxiv.org/html/2602.01736v1#bib.bib13); [Woo et al., 2022](https://arxiv.org/html/2602.01736v1#bib.bib17); [Liu et al., 2022](https://arxiv.org/html/2602.01736v1#bib.bib21); [Liu et al., 2024b](https://arxiv.org/html/2602.01736v1#bib.bib14); [Wu et al., 2022a](https://arxiv.org/html/2602.01736v1#bib.bib18); [Nie et al., 2022](https://arxiv.org/html/2602.01736v1#bib.bib19); [Cirstea et al., 2022](https://arxiv.org/html/2602.01736v1#bib.bib20); [Grigsby et al., 2021](https://arxiv.org/html/2602.01736v1#bib.bib22)), the community also explored competitive non-attention backbones, including lightweight linear variants (e.g. FITS, OLinear, CrossLinear, GLinear) ([Xu et al., 2024b](https://arxiv.org/html/2602.01736v1#bib.bib7); [Yue et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib8); [Zhou et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib23); [Rizvi et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib9)), modern CNN-style sequence models (e.g. ModernTCN, PatchMixer) ([donghao and wang xue, 2024](https://arxiv.org/html/2602.01736v1#bib.bib15); [Gong et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib77)), and MLP-mixer style models (e.g. TSMixer, TimeMixer) ([Ekambaram et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib76); [Wang et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib75)).

Crucially, some early works already exposed warning signs about the “general-domain” time series. DLinear ([Zeng et al., 2022](https://arxiv.org/html/2602.01736v1#bib.bib2)) shows that, under the same benchmark protocol, a single linear layer can match or nearly match many increasingly elaborate architectures; previous context-length scaling analyses ([Shi et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib16)) further suggest that simply extending the history window can even *hurt* performance on small datasets. These pieces of work in previous year also questioned the effectiveness of proposed methods and metrics.

### 2.2 Large Multi-Domain Time Series Datasets and Large Foundational Time Series Models

Large Datasets. A common practice is to propose large datasets consisted of several small datasets, represented by recent benchmarking efforts such as TFB ([Qiu et al., 2024b](https://arxiv.org/html/2602.01736v1#bib.bib10); [Qiu et al., 2024a](https://arxiv.org/html/2602.01736v1#bib.bib35)), in which $8068$ uni-variate time series and $25$ multi-variate time series datasets are included and used for benchmarking methods. A unified end-to-end evaluating pipeline is offered with standardized metrics and protocols, which has also motivated follow-up analyses and surveys on evaluation sensitivity and saturation. ([Kim et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib57); [Brigato et al., 2026](https://arxiv.org/html/2602.01736v1#bib.bib73); [Wang et al., 2025b](https://arxiv.org/html/2602.01736v1#bib.bib74))

Large Models. Inspired from pretrained foundational models for NLP, foundational models for time series have been proposed. Early representatives include unified or decoder-only forecasting backbones trained on multi-source collections ([Gao et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib25); [Das et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib30)). More recently, stronger 2025-era foundation models and training objectives have been explored, e.g., Sundial ([Liu et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib3)), LightGTS ([Wang et al., 2025a](https://arxiv.org/html/2602.01736v1#bib.bib4)), and UDE ([Wang et al., 2025d](https://arxiv.org/html/2602.01736v1#bib.bib5)), as well as in-context fine-tuning for time-series foundation models. ([Faw et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib6)) In parallel, a growing line of work explicitly investigates *LLM+time series* hybrids: digit-tokenization / zero-shot extrapolation with off-the-shelf LLMs ([Gruver et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib85)), prompt-based GPT-style forecasters ([Cao et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib26); [Liu et al., 2024a](https://arxiv.org/html/2602.01736v1#bib.bib86); [Pan et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib87); [Bumb et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib89)), “reprogramming” pretrained LLMs for forecasting ([Jin et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib27)), and alignment-style approaches that adapt LLM representations to numerical series ([Chang et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib28)). More recently, integrated heterogeneous prompting and cross-modal alignment frameworks have also been explored ([Wang et al., 2025c](https://arxiv.org/html/2602.01736v1#bib.bib88)). Relatedly, language-model-like sequence backbones have also been explored for time series tasks beyond Transformers (e.g., RWKV-style models) ([Hou and Yu, 2024](https://arxiv.org/html/2602.01736v1#bib.bib29); [Liang et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib24)).

Previous Calls for Better Large Datasets and Better Practice. Previously there have been calls for more scientific and better Large Datasets. For example,  ([Bergmeir, 2024a](https://arxiv.org/html/2602.01736v1#bib.bib32)) states the fact that such a foundational model might perform well on average on multiple datasets at the cost of poorer performance on certain subsets of datasets (similar to the James-Stein Paradox ([Stein, 1956](https://arxiv.org/html/2602.01736v1#bib.bib33); [James and Stein, 1961](https://arxiv.org/html/2602.01736v1#bib.bib34))).  ([Bergmeir, 2024a](https://arxiv.org/html/2602.01736v1#bib.bib32); [Pfister, 2024](https://arxiv.org/html/2602.01736v1#bib.bib31)) also call for future time series datasets with multi-modality data, for the reason that time series alone (especially uni-variate time series) do not have enough information (e.g. electricity and market time series might have similar trend in input, but might evolve in very different ways, hence more information is needed). More recently, some work states other anomalies of the time series forecasting community:  ([Brigato et al., 2026](https://arxiv.org/html/2602.01736v1#bib.bib73)) shows the sensitivity of evaluation results of network architectures to hyperparameters,  ([Wang et al., 2025b](https://arxiv.org/html/2602.01736v1#bib.bib74)) shows some forecasting benchmarks have been almost saturated.

In this paper, we analyze the root cause of these phenomena ground-based. We propose that these phenomena are caused by an incompatible conflict between Domain SOTA and Cross-Domain Generalizability, which is especially irreconcilable for the general domain time series forecasting tasks the community has been working on. Because of this irreconcilable conflict, current general-domain time series forecasting neural networks lag behind domain-specificially designed time series neural networks by enough margin that these domains develop their own SOTA neural networks (that are irrelevant to general domain time series neural networks).

## 3 The Irreconcilable Conflict: Single-Domain vs. General-Domain SOTA for Time Series Neural Networks

The central thesis of this paper is that any single architecture in Time Series Forecasting (TSF) faces an irreconcilable conflict between achieving domain-specific State-of-the-Art (SOTA) performance and maintaining cross-domain generalizability. While fields like NLP have converged toward a “one-size-fits-all” foundation model, TSF remains stubbornly fragmented. We argue that this is not a temporary limitation of current algorithms, but a fundamental conflict rooted in the nature of time series forecasting. Specifically, a model that generalizes must discard the unique contexts and structural priors that are essential for optimal performance in specialized fields. In the following sections, we demonstrate the irreconcilability of this conflict by analyzing two primary catalysts: the inherent cross-domain diversity of time series and the natural constraints of data scarcity that prevent traditional scaling.

### 3.1 Cross-Domain Diversity

One leading cause of the conflict comes from the diversity among different time series domains.

First, Time series data is rarely “self-contained”; its behavior is usually driven by external “context” that varies in structure and modality. For instance, weather forecasting relies on spatial-temporal context like geographic topography or atmospheric pressure grids, whereas financial forecasting depends on unstructured modalities like news sentiment or limit order book dynamics. Because it is architecturally difficult to build a single model capable of accepting such diverse auxiliary inputs, most existing works ignore these “extra” contexts to maintain a unified, sequence-only structure. Even for some “Multimodal TSF Models” which accept more modalities as input ([Bergmeir, 2024b](https://arxiv.org/html/2602.01736v1#bib.bib79)), they can not include all domain-specific modalities. For example, geographic topography for weather, stock market in another country for financial market prediction, etc., cannot be well-aligned as input to a single “Multimodal TSF Model”. By treating all domains as simple numerical sequences, these models suffer from an inevitable Bayes error. They discard the domain-specific information necessary to reach the theoretical limit of predictability, trading precision for architectural uniformity.

Furthermore, time series from different domains are governed by distinct mathematical and physical laws, requiring different feature and structural priors. For example, stock prices and tradings in financial markets can be affected by Macroeconomic Indicators (e.g. policies, government reports) ([Chen et al., 1986](https://arxiv.org/html/2602.01736v1#bib.bib90)), Cross-Assets and Inter-Market Correlations (e.g. gold price rises when there is a fear in stock market) ([Longin and Solnik, 2001](https://arxiv.org/html/2602.01736v1#bib.bib91)), Corporate and Fundamental Events, etc; while a city’s temperature can be affected by Geographic Baseline, Urban Structure, Anthropogenic Heat, Blue and Green Infrastructure, etc. ([Oke, 1982](https://arxiv.org/html/2602.01736v1#bib.bib92); [Stewart and Oke, 2012](https://arxiv.org/html/2602.01736v1#bib.bib93)) Domain experts design specific rules, features, neural network architectures and structures to take into account these factors. An architecture’s specific design may become a liability in another domain. Consequently, a single architecture cannot be simultaneously optimal for the rigid cycles of physical systems and the stochastic volatility of human-driven markets.

### 3.2 Data Scarcity

While fields like Natural Language Processing (NLP) and Computer Vision (CV) have thrived by scaling model parameters alongside massive datasets, TSF faces a fundamental bottleneck in data availability. This scarcity is not merely a logistical hurdle or budget limitation, but a natural constraint that prevents a “one-size-fits-all” scaling approach.

#### Natural Scarcity and Augmentation Challenge

Compared to the trillions of tokens available in language corpora, high-quality time series data is naturally scarce. Many critical domains, such as macroeconomics or long-term climate cycles, produce only a single data point per month or year. Furthermore, traditional data augmentation techniques used in CV (like flipping or cropping) often destroy the temporal dependencies and causal structures inherent in time series, making it difficult to synthetically expand training sets without introducing significant bias.

Furthermore, there is a significant imbalance in dataset sizes across different application domains. For example, the traffic and electricity domains each have nearly 200,000 total timestamps. In contrast, four of the ten domains, including prominent fields like stock and healthcare, have an aggregate dataset size of no more than 2,500 timestamps. Figure [2](https://arxiv.org/html/2602.01736v1#S3.F2) visualizes this disparity in total timestamps across different fields for TFB ([Qiu et al., 2024b](https://arxiv.org/html/2602.01736v1#bib.bib10)).

![Figure 2 : Left: Number of total timestamps across different domains for multivariate forecasting datasets in TFB dataset ( Qiu et al., 2024b ) ; Right: Number of total timestamps across different domains forecasting tasks used by most TSF papers represented by ModernTCN ( donghao and wang xue, 2024 ) . All are small compared to typical NLP datasets ( Yenduri et al., 2023 ) .](figures/tfb_total_timestamps.png)

Figure 2: Left: Number of total timestamps across different domains for multivariate forecasting datasets in TFB dataset ([Qiu et al., 2024b](https://arxiv.org/html/2602.01736v1#bib.bib10)); Right: Number of total timestamps across different domains forecasting tasks used by most TSF papers represented by ModernTCN ([donghao and wang xue, 2024](https://arxiv.org/html/2602.01736v1#bib.bib15)). All are small compared to typical NLP datasets ([Yenduri et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib61)).

#### The Temporal Boundary of Approximation Error

More importantly, TSF is bounded by a unique temporal constraint: even if one increases sampling frequency or collects data from more sensors within a fixed time span $T$, the information gain is limited. As we demonstrate in Appendix [A](https://arxiv.org/html/2602.01736v1#A1), for a given observation window, the approximation error of the underlying process is fundamentally bounded by $O(1/\sqrt{T})$. Unlike other domains where data can be gathered in parallel across the internet, time series data is tethered to the linear progression of time. This creates an inevitable estimation error that cannot be overcome simply by increasing sample density.

#### Failure of Scaling and the Need of Human Priors

The consequence of this data limitation is that TSF cannot rely on the “scaling laws” that define modern AI. In the absence of infinite data to “wash away” architectural inefficiencies, there is an almost-inevitable estimation error for any single domain. The only way to address it is by injecting domain-specific human priors into the model. For example, in power grid forecasting, incorporating physical constraints like Kirchhoff’s laws or seasonal load patterns allows a model to achieve high precision with limited data. However, these priors are strictly domain-dependent This necessity for specific inductive biases reinforces the irreconcilable gap between a unified, general-purpose architecture and a domain-specific SOTA model.

## 4 Consequences of the conflict and focusing on a ‘General Domain Time Series Forecasting’

### 4.1 Peculiar observations in Time Series Forecasting Community

In our TSF community, there have been peculiar observations and warning signals. These observations can actually be explained starting from the irreconcilable conflict we proposed. We provide two representative examples.

TSF Neural Networks do not always beat traditional methods. One specific observation by previous work ([Makridakis et al., 2018](https://arxiv.org/html/2602.01736v1#bib.bib78); [Bergmeir, 2024b](https://arxiv.org/html/2602.01736v1#bib.bib79)) is that traditional statistical forecasters (e.g., ARIMA ([Box et al., 2015](https://arxiv.org/html/2602.01736v1#bib.bib80)), exponential smoothing/ETS ([Hyndman et al., 2002](https://arxiv.org/html/2602.01736v1#bib.bib81)), and the Theta method ([Assimakopoulos and Nikolopoulos, 2000](https://arxiv.org/html/2602.01736v1#bib.bib82))) excel at general-domain time series forecasting tasks even compared with SOTA neural network methods. This is because traditional TSF methods are designed to fit general properties of general time series, and have been utilizing most of the features and properties shared across general domain time series. Some deep learning neural network methods are actually inspired by these traditional methods, for example, FITS ([Xu et al., 2024b](https://arxiv.org/html/2602.01736v1#bib.bib7)) proposes an FFT-LowPassFilter-Linear-iFFT learning pipeline, which connects to classic adaptive/Wiener filtering ideas ([Wiener, 1949](https://arxiv.org/html/2602.01736v1#bib.bib83); [Widrow et al., 1975](https://arxiv.org/html/2602.01736v1#bib.bib84)) that date back more than half a century.

Saturation of metrics of current general domain TSF neural network methods. Another observation is that the newly proposed neural networks have been making less and less progress on benchmarks. For example, ([Wang et al., 2025b](https://arxiv.org/html/2602.01736v1#bib.bib74)) proposed Accuracy Law, stating that accuracies for ETT, Electricity and Weather datasets are unlikely to be improved. As a result shown in ([Brigato et al., 2026](https://arxiv.org/html/2602.01736v1#bib.bib73)), hyperparameters have great impact on general time series forecasting models’ rankings: result variance is too large that almost all 8 general domain TSF neural networks tested can be made “SOTA on an averaged metric” with a “properly set” experiment setting, even on large, general-domain datasets that include many time series data samples such as TFB ([Qiu et al., 2024b](https://arxiv.org/html/2602.01736v1#bib.bib10)) and Gift-eval ([Aksu et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib94)). This indicates that simply increasing dataset size would not resolve the saturation issue. The devil lies in the inherent irreconcilable conflict between single domains performance and performance across general domains: a neural network designed targeting general time series performance can only leverage general domain features and properties, thus such direction can be easily exploited by hunderds of researcher years.

### 4.2 Diverse SOTA proposed by different task Communities irrelevant to the time series community

The gap between general-domain time series forecasting neural networks and domain-specific TS neural networks is so huge that, in most cases, single time series domain usually utilizes or develops their own SOTA methods. These methods are usually developed from general deep learning or machine learning methods, and unfortunately, are usually irrelevant to general domain time series forecasting neural networks, especially those developed in recent 3-4 years.

For example, in Table [1](https://arxiv.org/html/2602.01736v1#S4.T1) we list some single-domain time series forecasting challenges and competitions online. The top winning methods are based on, either statistical methods, or feature engineering and ML-based methods, or ‘plain’ deep-learning methods (rather than DLTSF methods, which are architectures designed for general domain raw time series inputs).

| Forecasting Competitions   | \#Teams | Top-1/Top-N Methods                                                                                                    |
|----------------------------|---------|------------------------------------------------------------------------------------------------------------------------|
| Kaggle & Optiver:          | 3225    | Top1: Feature Engineering +                                                                                            |
| Trading at Close           |         | Ensembled Model of CatBoost,GRU,Transformers ([hyd,](https://arxiv.org/html/2602.01736v1#bib.bib55) )                  |
| Kaggle & Jane Street:      | 4085    | Top1: AutoEncoder+MLP+XGBoost ([Wang,](https://arxiv.org/html/2602.01736v1#bib.bib56) )                                |
| Market Prediction          |         |                                                                                                                        |
| M6 Forecasting Competition | 226     | Top9: 1Judgement, 4TS-based, 3ML-based, 1NS ([Makridakis et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib36)) |

Table 1: Examples of Challenges and Competitions related to time series forecasting. For Kaggle-Optiver Competition: we further check top4 open-discussion solutions (all in top 20), which are all based on Feature Engineering and ensemble of traditional ML methods/basic DL methods. For Kaggle-Jane Street Competition: we further check top4 open-discussion solutions (all in top 20), which are all based on simple nets like MLPs, GBMs, etc. For M6 competition: Judgment: either pure judgmental or judgment-informed; TS-based: (traditional) time series approaches, but also their combinations; ML-based: ML approaches integrated with TS and combinations; NS: not specified. Though not specifically discussed by M6 ([Makridakis et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib36)) whether the teams use NNs proposed by time series forecasting community, but it seems most still use traditional deep-learning approaches, represented by ([Mashlakov,](https://arxiv.org/html/2602.01736v1#bib.bib58) ).

Apart from Kaggle time series-related competitions, we also check the relevant academic work in domains that are related to time series, and appear in the datasets from  ([Zhou et al., 2021](https://arxiv.org/html/2602.01736v1#bib.bib1)). We selected 9 representative academic papers on domains including Finance, Weather, Health, Traffic from domain-top journals and conferences, and record the detailed ML/DL methods they use. Results are shown in Table [2](https://arxiv.org/html/2602.01736v1#S4.T2). From the table we observe that: researchers in these domains develop their own SOTA methods that are built from basic building blocks (e.g. MLP, Transformers, etc.) and are not utilizing nor adapted from general domain time series methods.

| Work                                                                                                                                                                                                             | Published at | Domain  | Method Used/Discussed                |
|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------|---------|--------------------------------------|
|  ([Kaniel et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib65)) ([Ferson et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib64))                                                                   | JFE,JFQA     | Finance | Features+MLP                         |
|  ([Bi et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib37)) ([Price et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib63))                                                                        | Nature       | Weather | 3D Transformer,DDPM                  |
|  ([Zhang et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib69)) ([Li et al., 2025a](https://arxiv.org/html/2602.01736v1#bib.bib70))                                                                       | ML Conf.     | Health  | Foundation model,Transformer         |
|  ([Jungel et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib66)) ([Fang et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib67)) ([Li et al., 2025b](https://arxiv.org/html/2602.01736v1#bib.bib71)) | ML Conf.     | Traffic | Game Theory,Diffusion RL,Transformer |

Table 2: Examples of academic work published at top journals and conferences, studying typical problems with deep learning in different domains, within the past 3 years. ML Conf. refers to ICML, NeurIPS and ICLR. Our time series community has designed general time series forecasting neural networks, and they are tested on related datasets like Weather, Traffic ([Zhou et al., 2021](https://arxiv.org/html/2602.01736v1#bib.bib1)); however, academic work in corresponding domains are not using these general time series forecasting neural nets designed by time series community, implying the capability of currently designed general-domain TSF neural networks.

## 5 Possible Solutions

Given the fundamental barriers of diversity and scarcity, achieving SOTA performance requires a strategic departure from universalism. We propose several solutions which may overcome the underlying limitation.

### 5.1 Domain-specific Neural Network

The most direct solution is the development of Domain-Specific Neural Networks (DSNNs)—architectures intentionally narrow in scope but deep in domain integration. Unlike general foundation models that prioritize zero-shot transfer, DSNNs are engineered to internalize the unique contexts, constraints, and data-generating mechanisms of a specific field. In practice, this usually means that the model is not only trained *on* a domain, but also designed *with* domain structure in mind. A few representative examples illustrate why such specialization often dominates real deployments:

Physics-Informed Architectures (Weather/Climate). In weather and climate forecasting, many important models ([Kochkov et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib52); [Verma et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib53); [Bi et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib37); [Chen et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib39); [Nguyen et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib38)) embed physical structure (e.g., conservation laws, dynamical constraints, spatial fields) into the modeling pipeline, which can outperform generic sequence backbones that must infer these priors from limited observations.

Microstructure- and Regime-Aware Models (Finance). Financial time series are heavily non-stationary and driven by market microstructure and regime shifts ([Xu et al., 2024a](https://arxiv.org/html/2602.01736v1#bib.bib40); [Gajamannage et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib41); [Rahimikia et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib42)); prior work shows that explicitly designed modules for such properties can yield clear gains, whereas scaling generic architectures provides limited benefit ([Huang et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib54)).

Task-Specific Efficiency Under Scarcity (Energy/Operations). In many operational forecasting problems (e.g., load forecasting), supervised models tailored to the target horizon, covariates, and constraints can match or exceed large pretrained models under limited data budgets ([Tang et al., 2025b](https://arxiv.org/html/2602.01736v1#bib.bib45); [Wu et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib43)).

These examples suggest a practical lesson: when domain context is available and the goal is domain-level SOTA, inductive bias and domain integration matter more than universal architectural novelty. DSNNs therefore serve as a realistic anchor for what practitioners optimize, and motivate system-level solutions (e.g., meta-learning) that can reuse such experts across domains.

### 5.2 Meta-Learning: A Systemic Alternative to Unified Architectures

Given that the conflict between single-domain precision and universal generalizability is irreconcilable within a single architecture, we argue that true cross-domain “generalization” should not be sought at the neural network architecture level, but at the meta-system level. Instead of a fixed model trying to absorb all domain nuances, we propose meta-learning frameworks that can dynamically discover and deploy the optimal architecture for a given task.

#### Agentic Meta-Learning: the ‘LLM Scientist’

The most flexible approach to this problem is an agentic framework, often referred to as an ‘LLM Scientist’. Rather than acting as a forecaster itself, where it often struggles with numerical precision, the LLM acts as an orchestrator that analyzes the “context” and “properties” we identified in Section [3.1](https://arxiv.org/html/2602.01736v1#S3.SS1). Works like ([Zhao et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib49); [Jin et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib44); [Tang et al., 2025a](https://arxiv.org/html/2602.01736v1#bib.bib46)) typically employ specialized agents to perform LLM-guided diagnostics on a dataset’s statistics, subsequently proposing targeted preprocessing and selecting the best-fit model architecture from a library of experts. By utilizing the reasoning capabilities of LLMs to “plan” the forecasting strategy, these systems can inject human-like domain knowledge, such as identifying whether a series is stationary or cyclical, without requiring the underlying numerical model to be universal. This solves the “Bayes error” by allowing the system to use the most context-appropriate tools for each unique domain.

#### Learning-Based Meta-Selection

A more automated alternative is learning-based meta-learning, where a “meta-model” is trained to predict the performance of various base-learners across a wide array of datasets. Recent frameworks like ([Abdallah et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib50); [Norton et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib51); [Talagala et al., 2022](https://arxiv.org/html/2602.01736v1#bib.bib47)) demonstrate this by mapping the statistical characteristics of a new, unseen time series to a “performance matrix” of hundreds of specialized forecasting models. By learning which architectures excel on specific data properties (e.g., high-frequency vs. seasonal), these systems can recommend a SOTA-level model for a new domain in a zero-shot or few-shot manner. Unlike traditional scaling, which attempts to build one massive model, this approach builds a meta-recommender that leverages the efficiency of domain-specific models, thereby bypassing the data scarcity bottleneck by focusing on “learning to select” rather than “learning to forecast” from scratch.

#### Meta-Learning for Automated Feature Selection

In many high-stakes domains, engineered covariates and feature selection remain essential because raw temporal sequences often hide the true causal drivers. Yet manual feature design is costly and domain-specific. meta-learning offers a scalable alternative by learning, from prior tasks, which feature constructions tend to help under which data regimes.

For example, Meta-MSGL ([Khan et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib48)) uses a meta-learning controller to score high-dimensional meta-features and estimate the marginal utility of feature groups across datasets. Related meta-learned feature-importance mechanisms have also been explored in clinical time-series settings such as EHR modeling ([Aguiar et al., 2022](https://arxiv.org/html/2602.01736v1#bib.bib72)). Together, these approaches can prioritize transformations (e.g., seasonal Fourier terms) or indicators (e.g., trend/volatility measures) consistent with a series’ global statistics, mitigating the Bayes error by tailoring the input space to the task without manual intervention.

## 6 Alternative Views

### 6.1 The Universal Scaling Hypothesis

A prominent alternative view is that the conflict between domain-specific SOTA and generalizability is not fundamental, but rather a symptom of insufficient model capacity and data volume ([Kaplan et al., 2020](https://arxiv.org/html/2602.01736v1#bib.bib59)). Proponents of this view argue that if we scale architectures (e.g., to billions of parameters) and training corpora (to trillions of time points) sufficiently, a “super-model” would eventually internalize the diverse properties and latent contexts of all domains ([Vaswani et al., 2023](https://arxiv.org/html/2602.01736v1#bib.bib68); [Brown et al., 2020](https://arxiv.org/html/2602.01736v1#bib.bib11)). Under this hypothesis, the model would effectively learn to “switch” its internal logic based on the input signal, rendering domain-specific architectures obsolete.

Our Response: While the universal scaling hypothesis has proven transformative for NLP and CV ([Kaplan et al., 2020](https://arxiv.org/html/2602.01736v1#bib.bib59); [Krizhevsky et al., 2012](https://arxiv.org/html/2602.01736v1#bib.bib60)), we argue it is fundamentally hindered in TSF by the approximation error bound discussed in Section [3.2](https://arxiv.org/html/2602.01736v1#S3.SS2) ([Kuznetsov and Mohri, 2014](https://arxiv.org/html/2602.01736v1#bib.bib62)). In text or images, ‘more data’ often translates to higher semantic coverage; in time series, the dataset size is often limited naturally, and even if we manage to collect more data manually, the available data is strictly limited by the temporal span $T$. Scaling the model size without a corresponding linear increase in the temporal horizon does not resolve the $O(1/\sqrt{T})$ bottleneck; instead, it frequently exacerbates overfitting to domain-specific noise ([Shi et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib16)). Moreover, as we have mentioned, for time series domain diversity is more irreconcilable compared to CV/NLP: different time series domains follow different or even contradictory evolving rules, hence adding more domains/more data might even hurt the performance ([Bergmeir, 2024a](https://arxiv.org/html/2602.01736v1#bib.bib32); [Pfister, 2024](https://arxiv.org/html/2602.01736v1#bib.bib31)).

### 6.2 DL methods are more human-efficient

A reasonable counterargument is that deep learning methods are *human-efficient*: once a generic architecture and training pipeline are established, practitioners can often achieve strong performance on a new dataset with minimal manual modeling effort (e.g., no hand-crafted feature engineering or explicit physical constraints) ([Liang et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib24)). From this perspective, the goal of building universal architectures is not necessarily to beat the best domain-expert systems in every field, but to lower the barrier to entry and provide a reliable “default” forecaster across many domains.

Our Response. We agree that general-purpose neural forecasting systems can be attractive as a low-effort baseline. However, in high-stakes domains, the objective is typically *near-ceiling* performance under domain constraints and covariates; the limiting factor is usually domain context and valid inductive bias, not the backbone architecture. Thus, a universal sequence-only model can be human-efficient but information-inefficient ([Bergmeir, 2024a](https://arxiv.org/html/2602.01736v1#bib.bib32)).

Moreover, meta-learning can be *at least as human-efficient* while offering a more plausible path to higher performance: it automates the selection/adaptation of specialized experts (and their priors) from data, instead of forcing one backbone to fit all domains ([Abdallah et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib50); [Norton et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib51)). Therefore, the human-efficiency argument strengthens our position: prioritize domain-specific models and meta-learning systems over further “universal” architecture engineering.

### 6.3 General SOTA is also important

Another alternative view is that it is misguided to compare against domain-specific SOTA at all: what the TSF community should optimize for is *general* SOTA—a model that performs best *on average* across a broad suite of datasets under a unified protocol ([Qiu et al., 2024b](https://arxiv.org/html/2602.01736v1#bib.bib10); [Aksu et al., 2024](https://arxiv.org/html/2602.01736v1#bib.bib94)). Under this lens, incremental gains on cross-domain benchmarks are still meaningful, even if a domain expert can engineer a stronger system for a particular field. Moreover, for researchers attempting to work on a specific domain, they might try some simple, neat, effective general-domain methods, then try to build or get some more complicated, complex single-domain SOTA methods. It is very questionable whether a complicated, complex, hyperparameter-sensitive “general-domain SOTA” is helpful for these specific-domain researchers ([Brigato et al., 2026](https://arxiv.org/html/2602.01736v1#bib.bib73)).

Our Response. We agree that strong cross-domain baselines and fair unified benchmarks are valuable—and that “generalizability” is a legitimate scientific goal. Our concern is that the current “general SOTA” game is increasingly dominated by evaluation sensitivity (hyperparameters, preprocessing, dataset curation) and yields diminishing practical returns: leaderboard gains on averaged benchmarks rarely translate to the settings practitioners care about, where domain context is available and success is measured against a domain-specific ceiling ([Brigato et al., 2026](https://arxiv.org/html/2602.01736v1#bib.bib73)).

Meta-learning preserves generalizability (a unified interface and reuse across datasets) while better matching domain practice: rather than forcing one backbone to fit all domains, it learns to *select/adapt* specialized pipelines and inductive biases ([Abdallah et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib50); [Norton et al., 2025](https://arxiv.org/html/2602.01736v1#bib.bib51)). This shifts progress from chasing average-score gains to learning better rules for choosing the right expert.

## 7 Conclusion

![Figure 3 : Suggested practices: from pursuing a neural network architecture that has huge gaps with domain SOTAs to (1) developing single-domain SOTA neural network methods, or (2) developing meta-learning methods that are generalizable across domains.](figures/draft_practice.png)

Figure 3: Suggested practices: from pursuing a neural network architecture that has huge gaps with domain SOTAs to (1) developing single-domain SOTA neural network methods, or (2) developing meta-learning methods that are generalizable across domains.

Time Series Forecasting is a critical task across industries, from weather prediction and electricity demand to quantitative trading. Motivated by these applications, the community has pursued neural architectures that achieve SOTA across broad, general-domain benchmarks. Yet, as we argue, an irreconcilable tension between specific-domain and general-domain SOTA means these models often fall short of domain-specific ceilings and are rarely used by practitioners who can leverage domain context and tailored inductive biases. Because researchers are not aware of this fundamental, irreconcilable conflict is previously overlooked (e.g. some opinion believes scaling up/better designs can solve the problem), substantial effort continues to go into general-domain architecture design.

We suggest that these efforts should be switched to the development of more useful and effective topics: as proposed in our work, we call for researchers to focus on (1) developing domain-specific time series forecasting neural networks or (2) developing meta-learning methods that are more generalizable than a certain neural network architecture.

## References

- Abdallah et al. (2025) M. Abdallah, R. A. Rossi, K. Mahadik, S. Kim, H. Zhao, and S. Bagchi Evaluation-free time-series forecasting model selection via meta-learning. ACM Transactions on Knowledge Discovery from Data 19 (3), pp. 1–41. Cited by: [§5.2](https://arxiv.org/html/2602.01736v1#S5.SS2.SSS0.Px2.p1.1), [§6.2](https://arxiv.org/html/2602.01736v1#S6.SS2.p3.1), [§6.3](https://arxiv.org/html/2602.01736v1#S6.SS3.p3.1).
- Aguiar et al. (2022) H. Aguiar, M. Santos, P. Watkinson, and T. Zhu Learning of cluster-based feature importance for electronic health record time-series. In Proceedings of the 39th International Conference on Machine Learning, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (Eds.), Proceedings of Machine Learning Research, Vol. 162, pp. 161–179. External Links: [Link](https://proceedings.mlr.press/v162/aguiar22a.html) Cited by: [§5.2](https://arxiv.org/html/2602.01736v1#S5.SS2.SSS0.Px3.p2.1).
- Aksu et al. (2024) T. Aksu, G. Woo, J. Liu, X. Liu, C. Liu, S. Savarese, C. Xiong, and D. Sahoo GIFT-eval: a benchmark for general time series forecasting model evaluation. External Links: 2410.10393, [Link](https://arxiv.org/abs/2410.10393) Cited by: [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p3.1), [§6.3](https://arxiv.org/html/2602.01736v1#S6.SS3.p1.1).
- Assimakopoulos and Nikolopoulos (2000) V. Assimakopoulos and K. Nikolopoulos The theta model: a decomposition approach to forecasting. International Journal of Forecasting 16 (4), pp. 521–530. Cited by: [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p2.1).
- Bergmeir (2024a) C. Bergmeir Invited talk by christoph bergmeir - fundamental limitations of foundational forecasting models: the need for multimodality and rigorous evaluation. In NeurIPS 2024 Workshop on Time Series in the Age of Large Models, NeurIPS ’24 Workshop. External Links: [Link](https://neurips.cc/virtual/2024/108471) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p3.1), [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p2.1), [§6.2](https://arxiv.org/html/2602.01736v1#S6.SS2.p2.1).
- Bergmeir (2024b) C. Bergmeir Invited talk by christoph bergmeir - fundamental limitations of foundational forecasting models: the need for multimodality and rigorous evaluation. In NeurIPS 2024 Workshop on Time Series in the Age of Large Models, External Links: [Link](https://neurips.cc/virtual/2024/108471) Cited by: [§1](https://arxiv.org/html/2602.01736v1#S1.p5.1), [§3.1](https://arxiv.org/html/2602.01736v1#S3.SS1.p2.1), [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p2.1).
- Bi et al. (2023) K. Bi, L. Xie, H. Zhang, X. Chen, X. Gu, and Q. Tian Accurate medium-range global weather forecasting with 3d neural networks. Nature 619 (7970), pp. 533–538. Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.3.1), [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p2.1).
- Box et al. (2015) G. E. P. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung Time series analysis: forecasting and control. 5 edition, Wiley. Cited by: [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p2.1).
- Brigato et al. (2026) L. Brigato, R. Morand, K. Strømmen, M. Panagiotou, M. Schmidt, and S. Mougiakakou There are no champions in supervised long-term time series forecasting. External Links: 2502.14045, [Link](https://arxiv.org/abs/2502.14045) Cited by: [§1](https://arxiv.org/html/2602.01736v1#S1.p2.1), [§1](https://arxiv.org/html/2602.01736v1#S1.p5.1), [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p1.1), [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p3.1), [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p3.1), [§6.3](https://arxiv.org/html/2602.01736v1#S6.SS3.p1.1), [§6.3](https://arxiv.org/html/2602.01736v1#S6.SS3.p2.1).
- Brown et al. (2020) T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei Language models are few-shot learners. External Links: 2005.14165, [Link](https://arxiv.org/abs/2005.14165) Cited by: [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p1.1).
- Bumb et al. (2025) M. Bumb, A. Vemulapalli, S. H. V. P. Jella, A. Gupta, A. La, R. A. Rossi, H. Chen, F. Dernoncourt, N. K. Ahmed, and Y. Wang Forecasting time series with llms via patch-based prompting and decomposition. External Links: 2506.12953, [Link](https://arxiv.org/abs/2506.12953) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Cao et al. (2024) D. Cao, F. Jia, S. O. Arik, T. Pfister, Y. Zheng, W. Ye, and Y. Liu TEMPO: prompt-based generative pre-trained transformer for time series forecasting. External Links: 2310.04948, [Link](https://arxiv.org/abs/2310.04948) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Chang et al. (2025) C. Chang, W. Wang, W. Peng, and T. Chen LLM4TS: aligning pre-trained llms as data-efficient time-series forecasters. External Links: 2308.08469, [Link](https://arxiv.org/abs/2308.08469) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Chen et al. (2023) L. Chen, X. Zhong, F. Zhang, Y. Cheng, Y. Xu, Y. Qi, and H. Li FuXi: a cascade machine learning forecasting system for 15-day global weather forecast. npj climate and atmospheric science 6 (1), pp. 190. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p2.1).
- Chen et al. (1986) N. Chen, R. Roll, and S. A. Ross Economic forces and the stock market. Journal of Business 59 (3), pp. 383–403. Cited by: [§3.1](https://arxiv.org/html/2602.01736v1#S3.SS1.p3.1).
- Cirstea et al. (2022) R. Cirstea, C. Guo, B. Yang, T. Kieu, X. Dong, and S. Pan Triformer: triangular, variable-specific attentions for long sequence multivariate time series forecasting–full version. External Links: 2204.13767, [Link](https://arxiv.org/abs/2204.13767) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Das et al. (2024) A. Das, W. Kong, R. Sen, and Y. Zhou A decoder-only foundation model for time-series forecasting. External Links: 2310.10688, [Link](https://arxiv.org/abs/2310.10688) Cited by: [§1](https://arxiv.org/html/2602.01736v1#S1.p1.1), [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- donghao and wang xue (2024) L. donghao and wang xue ModernTCN: a modern pure convolution structure for general time series analysis. In The Twelfth International Conference on Learning Representations, External Links: [Link](https://openreview.net/forum?id=vpJMJerXHU) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1), [Figure 2](https://arxiv.org/html/2602.01736v1#S3.F2), [Figure 2](https://arxiv.org/html/2602.01736v1#S3.F2.4).
- Ekambaram et al. (2023) V. Ekambaram, A. Jati, N. Nguyen, P. Sinthong, and J. Kalagnanam TSMixer: lightweight mlp-mixer model for multivariate time series forecasting. External Links: 2306.09364, [Link](https://arxiv.org/abs/2306.09364) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Fang et al. (2024) Y. Fang, Y. Liang, B. Hui, Z. Shao, L. Deng, X. Liu, X. Jiang, and K. Zheng Efficient large-scale traffic forecasting with transformers: a spatial data management perspective. External Links: 2412.09972, [Link](https://arxiv.org/abs/2412.09972) Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.5.1).
- Faw et al. (2025) M. Faw, R. Sen, Y. Zhou, and A. Das In-context fine-tuning for time-series foundation models. In International Conference on Machine Learning (ICML), External Links: [Link](https://icml.cc/virtual/2025/poster/43707) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Ferson et al. (2025) W. E. Ferson, A. F. Siegel, and J. L. Wang Factor model comparisons with conditioning information. Journal of Financial and Quantitative Analysis 60 (3), pp. 1401–1426. External Links: [Document](https://dx.doi.org/10.1017/S002210902400005X) Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.2.1).
- Gajamannage et al. (2023) K. Gajamannage, Y. Park, and D. I. Jayathilake Real-time forecasting of time series in financial markets using sequentially trained dual-lstms. Expert Systems with Applications 223, pp. 119879. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p3.1).
- Gao et al. (2024) S. Gao, T. Koker, O. Queen, T. Hartvigsen, T. Tsiligkaridis, and M. Zitnik UniTS: a unified multi-task time series model. External Links: 2403.00131, [Link](https://arxiv.org/abs/2403.00131) Cited by: [§1](https://arxiv.org/html/2602.01736v1#S1.p1.1), [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Gong et al. (2023) Z. Gong, Y. Tang, and J. Liang PatchMixer: a patch-mixing architecture for long-term time series forecasting. External Links: 2310.00655, [Link](https://arxiv.org/abs/2310.00655) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Grigsby et al. (2021) J. Grigsby, Z. Wang, N. Nguyen, and Y. Qi Long-range transformers for dynamic spatiotemporal forecasting. External Links: 2109.12218, [Link](https://arxiv.org/abs/2109.12218) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Gruver et al. (2023) N. Gruver, M. Finzi, S. Qiu, and A. G. Wilson Large language models are zero-shot time series forecasters. In Advances in Neural Information Processing Systems, Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Hou and Yu (2024) H. Hou and F. R. Yu RWKV-ts: beyond traditional recurrent neural network for time series tasks. External Links: 2401.09093, [Link](https://arxiv.org/abs/2401.09093) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Huang et al. (2024) H. Huang, M. Chen, and X. Qiao Generative learning for financial time series with irregular and scale-invariant patterns. In The Twelfth International Conference on Learning Representations, Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p3.1).
- \[30\] hyd Optiver - trading at the close 1st solution. In Kaggle Optiver - Trading at the Close Competition, Kaggle Competition. External Links: [Link](https://www.kaggle.com/competitions/optiver-trading-at-the-close/discussion/487446) Cited by: [Table 1](https://arxiv.org/html/2602.01736v1#S4.T1.3.3.3).
- Hyndman et al. (2002) R. J. Hyndman, A. B. Koehler, R. D. Snyder, and S. Grose A state space framework for automatic forecasting using exponential smoothing methods. International Institute of Forecasters. Cited by: [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p2.1).
- James and Stein (1961) W. James and C. Stein Estimation with quadratic loss. In Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Vol. 1, pp. 361–379. Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p3.1).
- Jin et al. (2023) M. Jin, S. Wang, L. Ma, Z. Chu, J. Y. Zhang, X. Shi, P. Chen, Y. Liang, Y. Li, S. Pan, et al. Time-llm: time series forecasting by reprogramming large language models. arXiv preprint arXiv:2310.01728. Cited by: [§5.2](https://arxiv.org/html/2602.01736v1#S5.SS2.SSS0.Px1.p1.1).
- Jin et al. (2024) M. Jin, S. Wang, L. Ma, Z. Chu, J. Y. Zhang, X. Shi, P. Chen, Y. Liang, Y. Li, S. Pan, and Q. Wen Time-llm: time series forecasting by reprogramming large language models. External Links: 2310.01728, [Link](https://arxiv.org/abs/2310.01728) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Jungel et al. (2024) K. Jungel, D. Paccagnan, A. Parmentier, and M. Schiffer WardropNet: traffic flow predictions via equilibrium-augmented learning. External Links: 2410.06656, [Link](https://arxiv.org/abs/2410.06656) Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.5.1).
- Kaniel et al. (2023) R. Kaniel, Z. Lin, M. Pelger, and S. Van Nieuwerburgh Machine-learning the skill of mutual fund managers. Journal of Financial Economics 150 (1), pp. 94–138. External Links: ISSN 0304-405X, [Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jfineco.2023.07.004), [Link](https://www.sciencedirect.com/science/article/pii/S0304405X23001253) Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.2.1).
- Kaplan et al. (2020) J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei Scaling laws for neural language models. External Links: 2001.08361, [Link](https://arxiv.org/abs/2001.08361) Cited by: [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p1.1), [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p2.1).
- Khan et al. (2025) I. Khan, X. Zhang, R. K. Ayyasamy, S. M. Alhashmi, and A. Rahim Enhancing classification algorithm recommendation in automated machine learning: a meta-learning approach using multivariate sparse group lasso. Computer Modeling in Engineering & Sciences (CMES) 142 (2). Cited by: [§5.2](https://arxiv.org/html/2602.01736v1#S5.SS2.SSS0.Px3.p2.1).
- Kim et al. (2025) J. Kim, H. Kim, H. Kim, D. Lee, and S. Yoon A comprehensive survey of deep learning for time series forecasting: architectural diversity and open challenges. External Links: 2411.05793, [Link](https://arxiv.org/abs/2411.05793) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p1.1).
- Kochkov et al. (2024) D. Kochkov, J. Yuval, I. Langmore, P. Norgaard, J. Smith, G. Mooers, M. Klöwer, J. Lottes, S. Rasp, P. Düben, et al. Neural general circulation models for weather and climate. Nature 632 (8027), pp. 1060–1066. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p2.1).
- Krizhevsky et al. (2012) A. Krizhevsky, I. Sutskever, and G. E. Hinton ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, F. Pereira, C.J. Burges, L. Bottou, and K.Q. Weinberger (Eds.), Vol. 25, pp. . External Links: [Link](https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf) Cited by: [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p2.1).
- Kuznetsov and Mohri (2014) V. Kuznetsov and M. Mohri Generalization bounds for time series prediction with non-stationary processes. In International conference on algorithmic learning theory, pp. 260–274. Cited by: [Appendix A](https://arxiv.org/html/2602.01736v1#A1.p1.1), [item 2](https://arxiv.org/html/2602.01736v1#S1.I1.i2.p1.1), [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p2.1).
- Li et al. (2025a) H. Li, B. Deng, C. Xu, Z. Feng, V. Schlegel, Y. Huang, Y. Sun, J. Sun, K. Yang, Y. Yu, and J. Bian MIRA: medical time series foundation model for real-world health data. External Links: 2506.07584, [Link](https://arxiv.org/abs/2506.07584) Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.4.1).
- Li et al. (2025b) M. Li, J. Wang, G. Yu, X. Wang, Q. Chen, W. Ni, L. Li, and H. Peng RobustLight: improving robustness via diffusion reinforcement learning for traffic signal control. In Proceedings of the 42nd International Conference on Machine Learning, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research, Vol. 267, pp. 36192–36214. External Links: [Link](https://proceedings.mlr.press/v267/li25cs.html) Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.5.1).
- Liang et al. (2024) Y. Liang, H. Wen, Y. Nie, Y. Jiang, M. Jin, D. Song, S. Pan, and Q. Wen Foundation models for time series analysis: a tutorial and survey. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, pp. 6555–6565. External Links: [Link](http://dx.doi.org/10.1145/3637528.3671451), [Document](https://dx.doi.org/10.1145/3637528.3671451) Cited by: [§1](https://arxiv.org/html/2602.01736v1#S1.p1.1), [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1), [§6.2](https://arxiv.org/html/2602.01736v1#S6.SS2.p1.1).
- Liu et al. (2024a) H. Liu, Z. Zhao, J. Wang, H. Kamarthi, and B. A. Prakash LSTPrompt: large language models as zero-shot time series forecasters by long-short-term prompting. External Links: 2402.16132, [Link](https://arxiv.org/abs/2402.16132) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Liu et al. (2024b) Y. Liu, T. Hu, H. Zhang, H. Wu, S. Wang, L. Ma, and M. Long ITransformer: inverted transformers are effective for time series forecasting. External Links: 2310.06625, [Link](https://arxiv.org/abs/2310.06625) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Liu et al. (2025) Y. Liu, G. Qin, Z. Shi, Z. Chen, C. Yang, X. Huang, J. Wang, and M. Long Sundial: a family of highly capable time series foundation models. External Links: 2502.00816, [Link](https://arxiv.org/abs/2502.00816) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Liu et al. (2022) Y. Liu, H. Wu, J. Wang, and M. Long Non-stationary transformers: exploring the stationarity in time series forecasting. External Links: 2205.14415, [Link](https://arxiv.org/abs/2205.14415) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Longin and Solnik (2001) F. Longin and B. Solnik Extreme correlation of international equity markets. The Journal of Finance 56 (2), pp. 649–676. Cited by: [§3.1](https://arxiv.org/html/2602.01736v1#S3.SS1.p3.1).
- Makridakis et al. (2018) S. Makridakis, E. Spiliotis, and V. Assimakopoulos The M4 competition: results, findings, conclusion and way forward. International Journal of Forecasting 34 (4), pp. 802–808. Cited by: [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p2.1).
- Makridakis et al. (2023) S. Makridakis, E. Spiliotis, R. Hollyman, F. Petropoulos, N. Swanson, and A. Gaba The m6 forecasting competition: bridging the gap between forecasting and investment decisions. External Links: 2310.13357, [Link](https://arxiv.org/abs/2310.13357) Cited by: [Table 1](https://arxiv.org/html/2602.01736v1#S4.T1), [Table 1](https://arxiv.org/html/2602.01736v1#S4.T1.3.6.3).
- \[53\] A. Mashlakov Never have i ever … won m6 competition?. In Medium, Medium. External Links: [Link](https://medium.com/@alekseimashlakov/never-have-i-ever-won-m6-competition-53267e894c63) Cited by: [Table 1](https://arxiv.org/html/2602.01736v1#S4.T1).
- Nguyen et al. (2023) T. Nguyen, J. Brandstetter, A. Kapoor, J. K. Gupta, and A. Grover Climax: a foundation model for weather and climate. arXiv preprint arXiv:2301.10343. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p2.1).
- Nie et al. (2022) Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam A time series is worth 64 words: long-term forecasting with transformers. External Links: 2211.14730, [Link](https://arxiv.org/abs/2211.14730) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Norton et al. (2025) D. A. Norton, E. Ott, A. Pomerance, B. Hunt, and M. Girvan Tailored forecasting from short time series via meta-learning. arXiv preprint arXiv:2501.16325. Cited by: [§5.2](https://arxiv.org/html/2602.01736v1#S5.SS2.SSS0.Px2.p1.1), [§6.2](https://arxiv.org/html/2602.01736v1#S6.SS2.p3.1), [§6.3](https://arxiv.org/html/2602.01736v1#S6.SS3.p3.1).
- Oke (1982) T. R. Oke The energetic basis of the urban heat island. Quarterly Journal of the Royal Meteorological Society 108 (455), pp. 1–24. Cited by: [§3.1](https://arxiv.org/html/2602.01736v1#S3.SS1.p3.1).
- Pan et al. (2024) Z. Pan, Y. Jiang, S. Garg, A. Schneider, Y. Nevmyvaka, and D. Song $\textbf{S}^{2}$ip-LLM: semantic space informed prompt learning with llm for time series forecasting. External Links: 2403.05798, [Link](https://arxiv.org/abs/2403.05798) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Pfister (2024) T. Pfister Invited talk by tomas pfister - multimodal time series modeling. In NeurIPS 2024 Workshop on Time Series in the Age of Large Models, NeurIPS ’24 Workshop. External Links: [Link](https://neurips.cc/virtual/2024/108469) Cited by: [item 1](https://arxiv.org/html/2602.01736v1#S1.I1.i1.p1.1), [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p3.1), [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p2.1).
- Price et al. (2024) I. Price, A. Sanchez-Gonzalez, F. Alet, T. Andersson, A. El-Kadi, D. Masters, T. Ewalds, J. Stott, S. Mohamed, P. Battaglia, R. Lam, and M. Willson Probabilistic weather forecasting with machine learning. Nature 637, pp. 84–90. External Links: [Document](https://dx.doi.org/10.1038/s41586-024-08252-9) Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.3.1).
- Qiu et al. (2024a) X. Qiu, J. Hu, L. Zhou, X. Wu, J. Du, B. Zhang, C. Guo, A. Zhou, C. S. Jensen, Z. Sheng, and B. Yang TFB: towards comprehensive and fair benchmarking of time series forecasting methods. External Links: 2403.20150, [Link](https://arxiv.org/abs/2403.20150) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p1.1).
- Qiu et al. (2024b) X. Qiu, J. Hu, L. Zhou, X. Wu, J. Du, B. Zhang, C. Guo, A. Zhou, C. S. Jensen, Z. Sheng, and B. Yang TFB: towards comprehensive and fair benchmarking of time series forecasting methods. Proc. VLDB Endow. 17 (9), pp. 2363–2377. External Links: [Link](https://www.vldb.org/pvldb/vol17/p2363-hu.pdf) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p1.1), [Figure 2](https://arxiv.org/html/2602.01736v1#S3.F2), [Figure 2](https://arxiv.org/html/2602.01736v1#S3.F2.4), [§3.2](https://arxiv.org/html/2602.01736v1#S3.SS2.SSS0.Px1.p2.1), [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p3.1), [§6.3](https://arxiv.org/html/2602.01736v1#S6.SS3.p1.1).
- Rahimikia et al. (2025) E. Rahimikia, H. Ni, and W. Wang Re (visiting) time series foundation models in finance. arXiv preprint arXiv:2511.18578. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p3.1).
- Rizvi et al. (2025) S. T. H. Rizvi, N. Kanwal, M. Naeem, A. Cuzzocrea, and A. Coronato Bridging simplicity and sophistication using glinear: a novel architecture for enhanced time series prediction. External Links: 2501.01087, [Link](https://arxiv.org/abs/2501.01087) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Shi et al. (2024) J. Shi, Q. Ma, H. Ma, and L. Li Scaling law for time series forecasting. External Links: 2405.15124, [Link](https://arxiv.org/abs/2405.15124) Cited by: [item 2](https://arxiv.org/html/2602.01736v1#S1.I1.i2.p1.1), [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p2.1), [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p2.1).
- Stein (1956) C. Stein Inadmissibility of the usual estimator for the mean of a multivariate normal distribution. In Proceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, Vol. 1, pp. 197–206. Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p3.1).
- Stewart and Oke (2012) I. D. Stewart and T. R. Oke Local climate zones for urban temperature studies. Bulletin of the American Meteorological Society 93 (12), pp. 1879–1900. Cited by: [§3.1](https://arxiv.org/html/2602.01736v1#S3.SS1.p3.1).
- Talagala et al. (2022) T. S. Talagala, F. Li, and Y. Kang FFORMPP: feature-based forecast model performance prediction. International Journal of Forecasting 38 (3), pp. 920–943. Cited by: [§5.2](https://arxiv.org/html/2602.01736v1#S5.SS2.SSS0.Px2.p1.1).
- Tang et al. (2025a) H. Tang, C. Zhang, M. Jin, Q. Yu, Z. Wang, X. Jin, Y. Zhang, and M. Du Time series forecasting with llms: understanding and enhancing model capabilities. ACM SIGKDD Explorations Newsletter 26 (2), pp. 109–118. Cited by: [§5.2](https://arxiv.org/html/2602.01736v1#S5.SS2.SSS0.Px1.p1.1).
- Tang et al. (2025b) P. Tang, Z. Li, X. Wang, X. Liu, and P. Mou Time series data augmentation for energy consumption data based on improved timegan. Sensors 25 (2), pp. 493. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p4.1).
- Vaswani et al. (2023) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin Attention is all you need. External Links: 1706.03762, [Link](https://arxiv.org/abs/1706.03762) Cited by: [§6.1](https://arxiv.org/html/2602.01736v1#S6.SS1.p1.1).
- Verma et al. (2024) Y. Verma, M. Heinonen, and V. Garg Climode: climate and weather forecasting with physics-informed neural odes. arXiv preprint arXiv:2404.10024. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p2.1).
- \[73\] M. Wang JaneStreet - market prediction. In Kaggle JaneStreet Market Prediction, Kaggle Competition. External Links: [Link](https://www.kaggle.com/c/jane-street-market-prediction/discussion/229623) Cited by: [Table 1](https://arxiv.org/html/2602.01736v1#S4.T1.3.4.3).
- Wang et al. (2024) S. Wang, H. Wu, X. Shi, T. Hu, H. Luo, L. Ma, J. Y. Zhang, and J. Zhou TimeMixer: decomposable multiscale mixing for time series forecasting. External Links: 2405.14616, [Link](https://arxiv.org/abs/2405.14616) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Wang et al. (2025a) Y. Wang, Y. Qiu, P. Chen, Y. Shu, Z. Rao, L. Pan, B. Yang, and C. Guo LightGTS: a lightweight general time series forecasting model. External Links: 2506.06005, [Link](https://arxiv.org/abs/2506.06005) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Wang et al. (2025b) Y. Wang, H. Wu, Y. Ma, Y. Fang, Z. Zhang, Y. Liu, S. Wang, Z. Ye, Y. Xiang, J. Wang, and M. Long Accuracy law for the future of deep time series forecasting. External Links: 2510.02729, [Link](https://arxiv.org/abs/2510.02729) Cited by: [§1](https://arxiv.org/html/2602.01736v1#S1.p2.1), [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p1.1), [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p3.1), [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p3.1).
- Wang et al. (2025c) Z. Wang, L. Lan, and Y. Li Time-prompt: integrated heterogeneous prompts for unlocking llms in time series forecasting. External Links: 2506.17631, [Link](https://arxiv.org/abs/2506.17631) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Wang et al. (2025d) Z. Wang, P. Tao, J. Shi, R. Bao, R. Liu, and L. Chen A time-series foundation model by universal delay embedding. External Links: 2509.12080, [Link](https://arxiv.org/abs/2509.12080) Cited by: [§2.2](https://arxiv.org/html/2602.01736v1#S2.SS2.p2.1).
- Widrow et al. (1975) B. Widrow, J. R. Glover, J. M. McCool, J. Kaunitz, C. S. Williams, R. H. Hearn, J. R. Zeidler, E. Dong, and R. C. Goodlin Adaptive noise cancelling: principles and applications. Proceedings of the IEEE 63 (12), pp. 1692–1716. Cited by: [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p2.1).
- Wiener (1949) N. Wiener Extrapolation, interpolation, and smoothing of stationary time series. MIT Press. Cited by: [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p2.1).
- Woo et al. (2022) G. Woo, C. Liu, D. Sahoo, A. Kumar, and S. Hoi ETSformer: exponential smoothing transformers for time-series forecasting. External Links: 2202.01381, [Link](https://arxiv.org/abs/2202.01381) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Wu et al. (2022a) H. Wu, T. Hu, Y. Liu, H. Zhou, J. Wang, and M. Long TimesNet: temporal 2d-variation modeling for general time series analysis. External Links: 2210.02186, [Link](https://arxiv.org/abs/2210.02186) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Wu et al. (2022b) H. Wu, J. Xu, J. Wang, and M. Long Autoformer: decomposition transformers with auto-correlation for long-term series forecasting. External Links: 2106.13008, [Link](https://arxiv.org/abs/2106.13008) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Wu et al. (2025) S. Wu, Z. Wang, X. Liu, Y. Zhao, Y. Hu, and Y. Huang Temporal structure-preserving transformer for industrial load forecasting. Neural Networks, pp. 107887. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p4.1).
- Xu et al. (2024a) K. Xu, L. Chen, J. Patenaude, and S. Wang Rhine: a regime-switching model with nonlinear representation for discovering and forecasting regimes in financial markets. In Proceedings of the 2024 SIAM International Conference on Data Mining (SDM), pp. 526–534. Cited by: [§5.1](https://arxiv.org/html/2602.01736v1#S5.SS1.p3.1).
- Xu et al. (2024b) Z. Xu, A. Zeng, and Q. Xu FITS: modeling time series with $10k$ parameters. In International Conference on Learning Representations (ICLR), External Links: [Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/701251e1db4a2e4dd2ef23f5265d5936-Abstract-Conference.html) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1), [§4.1](https://arxiv.org/html/2602.01736v1#S4.SS1.p2.1).
- Yenduri et al. (2023) G. Yenduri, R. M, C. S. G, S. Y, G. Srivastava, P. K. R. Maddikunta, D. R. G, R. H. Jhaveri, P. B, W. Wang, A. V. Vasilakos, and T. R. Gadekallu Generative pre-trained transformer: a comprehensive review on enabling technologies, potential applications, emerging challenges, and future directions. External Links: 2305.10435, [Link](https://arxiv.org/abs/2305.10435) Cited by: [Figure 2](https://arxiv.org/html/2602.01736v1#S3.F2), [Figure 2](https://arxiv.org/html/2602.01736v1#S3.F2.4).
- Yue et al. (2025) W. Yue, Y. Liu, H. Li, H. Wang, X. Ying, R. Guo, B. Xing, and J. Shi OLinear: a linear model for time series forecasting in orthogonally transformed domain. External Links: 2505.08550, [Link](https://arxiv.org/abs/2505.08550) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Zeng et al. (2022) A. Zeng, M. Chen, L. Zhang, and Q. Xu Are transformers effective for time series forecasting?. External Links: 2205.13504, [Link](https://arxiv.org/abs/2205.13504) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p2.1).
- Zhang et al. (2023) X. Zhang, S. Li, Z. Chen, X. Yan, and L. Petzold Improving medical predictions by irregular multimodal electronic health records modeling. External Links: 2210.12156, [Link](https://arxiv.org/abs/2210.12156) Cited by: [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2.3.4.1).
- Zhao et al. (2025) H. Zhao, X. Zhang, J. Wei, Y. Xu, Y. He, S. Sun, and C. You Timeseriesscientist: a general-purpose ai agent for time series analysis. arXiv preprint arXiv:2510.01538. Cited by: [§5.2](https://arxiv.org/html/2602.01736v1#S5.SS2.SSS0.Px1.p1.1).
- Zhou et al. (2021) H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang Informer: beyond efficient transformer for long sequence time-series forecasting. External Links: 2012.07436, [Link](https://arxiv.org/abs/2012.07436) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1), [§2](https://arxiv.org/html/2602.01736v1#S2.p1.1), [§4.2](https://arxiv.org/html/2602.01736v1#S4.SS2.p3.1), [Table 2](https://arxiv.org/html/2602.01736v1#S4.T2).
- Zhou et al. (2025) P. Zhou, Y. Liu, J. Liang, Q. Song, and X. Li CrossLinear: plug-and-play cross-correlation embedding for time series forecasting with exogenous variables. External Links: 2505.23116, [Link](https://arxiv.org/abs/2505.23116) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).
- Zhou et al. (2022) T. Zhou, Z. Ma, Q. Wen, X. Wang, L. Sun, and R. Jin FEDformer: frequency enhanced decomposed transformer for long-term series forecasting. External Links: 2201.12740, [Link](https://arxiv.org/abs/2201.12740) Cited by: [§2.1](https://arxiv.org/html/2602.01736v1#S2.SS1.p1.1).

## Appendix A Proof for temporal boundary of estimation error

This result is mainly based on  ([Kuznetsov and Mohri, 2014](https://arxiv.org/html/2602.01736v1#bib.bib62)), which bounded the generalization loss for non-stationary process. More specifically, let $((X_{1},Y_{1}),\dots,(X_{m},Y_{m}))$ be a sample of pairs from the distribution $\mathcal{Z}=\mathcal{X}\times\mathcal{Y}$, and $H=\{h:\mathcal{X}\rightarrow\mathcal{Y}\}$ be a class of hypothesis functions that admits some loss function $L:\mathcal{X}\times\mathcal{X}\rightarrow\mathbb{R}^{+}$. Assume data distribution will converge to some stationary distribution:

$$\beta(a)=\sup_{t}\mathbb{E}{\left[\|\mathbf{P}_{t+a}(\cdot|\mathbf{Z}_{-\infty}^{t})-\Pi\|_{TV}\right]}\rightarrow 0\ \text{when}\ a\rightarrow+\infty$$ (1)

where $\mathbf{Z}_{a}^{b}=(Z_{a},\dots Z_{b})$ denotes a vector of data samples, $\mathbf{P}_{t}$ denotes the distribution data sample at time $t$. Then we have the following statement:

###### Theorem A.1.

Let $L$ be a loss function bounded by $M$ and $H$ be any set of hypotheses functions. Suppose $T=ma$ for some $m,a>0$. Then for any $\delta>a(m-1)\beta(a)$ with probability $1-\delta$:

$$\mathcal{L}_{\Pi}(h)=\mathbb{E}_{\Pi}[\ell(h,Z)]\leq\frac{1}{T}\sum_{t=1}^{T}\ell(h,Z_{t})+2\mathfrak{R}_{m}(H,\Pi)+M\sqrt{\frac{\log\frac{a}{\delta^{\prime}}}{2m}},$$ (2)

holds for arbitrary hypothesis function $h\in H$, where $\delta^{\prime}=\delta-a(m-1)\mathsf{\beta}(a)$ and $\mathfrak{R}_{m}(H,\Pi)=\frac{1}{m}\mathbb{E}[\sup_{h\in H}\sum_{i=1}^{m}\sigma_{i}\ell(h,\widetilde{Z}_{\Pi,i})]$ with $\sigma_{i}$ a sequence of Rademacher random variables.

We define the series to be *exponentially $\beta$-mixing* if $\beta(a)\leq Ce^{-da}$ for some positive constants $C,d$. Then, the argument in Section [3.2](https://arxiv.org/html/2602.01736v1#S3.SS2) is actually a direct consequence of Theorem [A.1](https://arxiv.org/html/2602.01736v1#A1.Thmtheorem1). A proof sketch is provided as follows;

###### Proof.

To maintain high probability, $\delta\rightarrow 0$, which gives $a(m-1)\beta(a)\rightarrow 0$. This implies $T\beta(a)\rightarrow 0$. Thus, by the exponentially $\beta$-mixing condition, we have $a=O(\log T)$, then $\sqrt{\frac{\log\frac{a}{\delta^{\prime}}}{2m}}=\tilde{O}(1/\sqrt{T})$. ∎

## Appendix B LLM Usage

LLMs are used in this paper for polishing writings.
