Position: The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting

Source

  • Preprint: arXiv:2602.01736v1, submitted 2 February 2026. The PDF itself is dated 1 February 2026.
  • Authors: Qinwei Ma, Jingzhe Shi, Jiahao Qiu, Zaiwen Yang. PDF affiliations: Tsinghua University and Princeton University; Ma and Shi are equal contributors.
  • Raw Markdown · Rendered PDF · Canonical HTML.
  • Status checked 28 September 2026: the unversioned arXiv record lists only v1 and no accepted venue. Treat as an arXiv position preprint, not an ICML acceptance; the ICML template/keywords do not establish peer review. Institutional affiliations provide provenance, not validation of the argument. Included at Alex’s explicit request.
  • Searches by exact title/arXiv ID did not verify official code, a project/blog announcement, or an author X thread. No new model, checkpoint, or dataset release is claimed here.
  • Conversion note: canonical HTML was used to recover the full bibliography and readable tables after TeX conversion left template/raw-LaTeX material. All three figures are retained from the source PDF assets; the original PDF and source archive are preserved. HTML calls the appendix theorem A.1; the PDF calls it Theorem 1.

Core Claim

Authors’ position: optimizing one sequence-only forecasting architecture across heterogeneous domains trades away domain context and specialized inductive biases. They advocate domain-specific systems or a meta-system that selects/adapts experts rather than further general-purpose backbone competition.

Wiki assessment: the useful argument is about the information interface and evaluation target, not a demonstrated impossibility of shared architectures or foundation models. A shared model conditioned on context, routed experts, or task adaptation is not equivalent to a context-blind fixed forecaster. The paper does not isolate these alternatives experimentally.

Argument And Evidence

ComponentWhat v1 suppliesEvidential boundary
Missing context (§3.1)Weather geography/physical structure versus financial news, microstructure, and regimes; numeric histories alone may omit predictive information.This motivates context-aware interfaces; it does not prove that a common backbone cannot consume them.
Data scarcity (§3.2, Fig. 2)Unequal timestamp counts in selected TFB/ModernTCN benchmark domains and a mixing-based bound in Appendix A.Selected benchmark coverage is not the total available time-series corpus. Timestamp count, duration, channel count, independent systems, and effective samples are different units.
Domain practice (§4, Tables 1–2)Three competition examples and nine selected papers across finance, weather, health, and traffic. Optiver’s listed winner combines feature engineering, CatBoost, GRU, and Transformers; Jane Street lists autoencoder, MLP, and XGBoost.Descriptive selection, not a new matched-input, matched-budget cross-domain forecasting experiment. Some cited traffic work is control/RL rather than the same forecasting task. Non-adoption does not establish a performance lower bound.
Alternatives (§5)Domain-specific priors; LLM-orchestrated diagnostics/model selection; learned expert selection; meta-learned feature selection.Proposed directions illustrated with prior work, not an implemented or benchmarked new meta-learning system.
Counterarguments (§6)Discusses scaling, human-efficiency, and average cross-domain SOTA; concedes the value of convenient general-purpose baselines.The recommendation weights domain-ceiling accuracy more heavily than engineering cost, reuse, or low-effort deployment. Those are separate objectives.

There is no new controlled benchmark demonstrating that specialized models universally dominate TSFMs, no new state-of-the-art forecaster, and no direct evaluation of action-conditioned world models.

The Temporal-Bound Caveat

The introduction describes a lower bound proportional to and uses it to argue that architecture/scaling cannot overcome limited temporal duration. Appendix A actually displays an upper bound on stationary-target risk:

The setup assumes a bounded loss and convergence toward a stationary distribution; the proof sketch further invokes exponential -mixing. It obtains a scaling for the displayed concentration term through logarithmic block separation. It does not derive a universal minimax lower bound, eliminate the hypothesis-class/Rademacher term, or prove architecture-independent approximation error. Moreover, in the sum is an observation index count: translating it into fixed physical duration while varying sampling rate requires additional dependence assumptions.

Consequently, the paper supports caution about dependence and effective data, but not the stronger conclusion that extra systems, context, priors, synthetic trajectories, or pretraining cannot help. The claimed irreconcilability remains a position, not an established theorem. See the tension ledger.

Relation To Existing Sources

  • Context is Key is the cleaner empirical anchor for missing-context forecasts. Its implication is to supply and evaluate context, not necessarily abandon shared representations.
  • Alex’s latent-state time-series position also rejects a narrow raw-observation forecasting frame, but targets useful latent state and broader system capabilities rather than concluding that cross-domain foundation models are impossible.
  • Scaling Law for Time Series Forecasting, coauthored by Jingzhe Shi, is cited by this paper. It makes look-back horizon/data regime part of scaling; it is not evidence that all scaling is futile. Scaling-laws for Large Time-series Models supplies empirical positive scaling evidence under its own forecasting setup.
  • MIRA is directly cited as a specialized medical foundation-model example: domain-specific and foundation-model are not opposites. Moirai-MoE is a useful existing shared-backbone specialization comparator, not a rebuttal experiment performed here.
  • Awesome Agentic Time Series maps the orchestration alternative. Selecting a forecaster does not itself provide learned next-state dynamics or action-conditioned planning.

Foundation TSFM Relevance

Agenda slotVerdictEvidenceMissing pieces
Context interfacewarningHighlights information discarded by sequence-only interfaces.Controlled identical-input comparison of shared, routed, adapted, and specialized systems.
Data diversity, curriculum, and long tailwarningHighlights dependence and imbalance; Appendix A sharpens effective-sample questions.No universal lower bound or full-corpus scaling test; distinguish duration, systems, channels, regimes, and repeated samples.
BenchmarkswarningSeparates average benchmark gains from domain practice and domain-ceiling goals.Matched task/context/budget baselines; per-domain and worst-domain metrics, adaptation and human effort.
Dynamic compute allocationadjacentMeta-selection could choose specialized pipelines.No implemented router, budget-aware selection study, or measured inference/training trade-off.
Causal structure, counterfactuals, and controlinsufficient evidenceDomain mechanisms are discussed; some cited applications involve control.No explicit action/control-input model, causal identification, or closed-loop experiment.

The source is a design/evaluation warning for the Foundation TSFM agenda, not a slot-closing method. It does not warrant a new model entity or an added row in the agenda’s selective coverage matrix.

Open Questions

  1. Under identical information and total adaptation budgets, where do context-conditioned shared backbones lose to domain-specific experts, and where do they win on engineering cost or low-data transfer?
  2. Can a routed/shared model preserve worst-domain and rare-regime performance without losing cross-domain sample efficiency?
  3. How should effective data be measured separately for longer observation duration, faster sampling, more channels, and more independently observed systems?
  4. Does learned model selection beat a simple validation-selected expert library once diagnostics, tuning, selection errors, and orchestration costs are included?