Use What You Know: Causal Foundation Models with Partial Graphs

Source

Metadata and credibility. Arik Reuter, Anish Dhir, Cristiana Diaconu, Jake Robertson, Ole Ossen, Frank Hutter, Adrian Weller, Mark van der Wilk, and Bernhard Schölkopf. First submitted 2026-02-16, revised 2026-06-25; accepted ICML 2026 poster. Affiliations include Cambridge, MPI for Intelligent Systems, UCL, Prior Labs, Freiburg, ELLIS Institute Tübingen, The Alan Turing Institute, and Oxford. Technical claims below refer to the actual arXiv v2 PDF/source, not an assumed publisher-PDF match.

The repository was inspected at commit 37c3d325408dc857b294866bdc2f5e3acf907649 (2026-04-20), which predates v2. A README and focused implementation snapshot are preserved with provenance under the raw paper directory. No separate official blog or author X announcement was verified; authenticated recent-X search returned no matches and public search did not establish an older thread. This is a discovery limit, not proof that no announcement exists.

Core Claim

A causal foundation model should accept what the user knows about causal structure without forcing them to supply a complete graph or retrain a separate model. This paper turns that idea into an explicit three-state interface: a variable is a known ancestor, a known non-ancestor, or its relationship is unknown.

A PFN trained on synthetic structural causal models takes an observational dataset, an intervention query, and this partial structural context. Soft feature-attention biases plus a GCN/AdaLN conditioning path improve causal-effect estimation in the studied synthetic and semi-synthetic settings. The central practical result is unknown is safer than wrong: hiding uncertain relationships is substantially less damaging than flipping them.

This is tabular causal inference over i.i.d. samples, not a time-series foundation model, latent-dynamics rollout model, or demonstrated controller.

Mechanism

Partial Knowledge Is A Distinct Input

For variables , the partially known ancestor matrix (PAM) is

An ancestor may be a direct or indirect cause. The PAM is therefore not a binary adjacency matrix with missing edges represented by zero, nor is it a standard partial ancestral graph for latent-confounder discovery. In this paper, all modeled variables are assumed observed.

The target is a conditional interventional distribution, schematically

Here denotes the intervention variable/value, not time. The paper calls it a treatment; in the wiki’s terminology it is an intervention on an SCM variable. Training samples fresh noise under the intervened SCM, rather than performing individual-level abduction and shared-noise counterfactual trajectory replay.

How The Structure Enters The Network

  • Embed numeric values, variable roles, and observational/query roles.
  • Alternate attention across variables within each sample and attention across observational samples. Queries attend to the observational context.
  • Add a per-head, per-layer bias to feature-attention logits: encourage attention from an effect to its known ancestors, discourage known non-ancestors, and apply zero structural bias for unknown relationships.
  • Encode broader graph structure with a GCN and modulate hidden features through adaptive layer normalization.
  • Decode the query outcome into a discretized bar distribution. Train by negative log-likelihood on synthetic observational/interventional tasks while varying how much ancestry is hidden.

An implementation-consistent attention expression is

Notation caveat: Section 4.3 first defines as normalized attention, then describes as a logit modification; Appendix D also switches the ancestor indices. The pinned implementation resolves the intended soft-attention path: it transposes the relationship matrix and supplies an additive float mask to PyTorch multi-head attention, i.e. before softmax. The two bias arrays are unconstrained parameters in the inspected code; their signs are not guaranteed to stay positive despite the appendix’s positive-scalar wording. Do not copy the displayed post-softmax addition as an algorithm.

Evidence And Its Boundaries

Controlled Synthetic Tests

  • Linear-Gaussian SCMs create cases where observational data do not identify direction. Conditioning supplies information that cannot simply be recovered by fitting the same observations harder.
  • Soft attention and GCN + soft attention outperform the hard-mask and GCN-only alternatives in the reported comparisons. Full ancestry performs similarly to full adjacency; this does not mean they encode identical finite-sample information.
  • A model trained across varying knowledge levels largely matches models specialized to complete or absent knowledge on MSE and ; Appendix H.2 reports a clearer NLL advantage for the fully informed specialist when all information is available. This is not uniform parity across metrics.
  • Complex-prior tests use 30,000 datasets. Benefits in absolute metrics are modest; the prior often makes observational and interventional outcome distributions quite similar. Gains depend on how informative the causal structure actually is.
  • Appendix K compares hiding entries with sign-flipping on 100,000 samples: incorrect structural assertions degrade results much faster than explicit unknowns.

Semi-Synthetic RealCause Results

Table 1 compares models built on the same causal prior. Lower is better for both root PEHE and relative ATE error. The following are reported central values, with uncertainty retained in the raw table:

DatasetRoot PEHE: no ancestry → ancestryRelative ATE error: no ancestry → ancestry
IHDP6.28 → 5.490.67 → 0.49
ACIC3.47 → 2.790.46 → 0.17
CPS12800 → 112130.99 → 0.70
PSID, original imbalanced sample13096 → 129750.98 → 1.09

The PSID ATE result is a counterexample to a universal improvement claim. The authors rebalance by retaining 141 treated individuals and subsampling 500 untreated individuals from 2266; Table 2 then reports better results with ancestry. These are different evaluation populations, so absolute root-PEHE values should not be read as a before/after improvement on a fixed target population.

Not overall causal-inference SOTA. Appendix J/Table 3 reports CausalPFN best overall on RealCause; for IHDP its root PEHE is , versus for the graph-conditioned model. The authors attribute this to prior alignment, but that explanation is not a matched-prior isolation experiment. Evaluation on the authors’ own synthetic prior likewise does not remove train–test prior advantages relative to external models.

Limitations And Gotchas

  1. Causal sufficiency is assumed throughout. No hidden confounders, independent exogenous noises, and acyclic SCMs delimit the claim. Partial knowledge does not remove causal-identification assumptions.
  2. Bayesian uncertainty is prior-dependent. Correct partial knowledge may identify a query, but otherwise posterior uncertainty remains dependent on the SCM prior. The consistency discussion concerns an ideal posterior under regularity conditions; it does not establish consistency or calibration of a finite trained network.
  3. No guarantee of robustness to false knowledge. The model can confidently answer the wrong causal question when the supplied graph is wrong. Missing-context and corrupted-context tests must be separate.
  4. Temporal ordering needs semantic care. The paper motivates non-ancestry from temporal precedence. For real time series, ordering should refer to the underlying state/event time, not merely delayed measurement, reporting, or database ingestion timestamps; this is a transfer caveat, not an evaluated result.
  5. Effect estimation is not rollout. There is no maintained latent state, multistep action sequence, irregular-stream evaluation, or closed-loop control result.
  6. Reproduction is not turnkey. The pinned repository contains training/model/prior code and checkpoint pointers, but its README’s root run.py does not exist in that tree and no repository license file was found. Paper CC BY 4.0 does not imply code or weight reuse permission. No training or inference reproduction was attempted during ingest.
  7. Metric equation typo. Appendix F.2.1 prints CATE as the integral of a difference of normalized densities without multiplying by , which would integrate to zero, and then uses the estimated-effect symbol for the ground truth. The intended quantity is a difference of expectations, . This is a documentation caveat, not evidence that the reported metric implementation used the printed typo.

Foundation TSFM Relevance

Agenda slotVerdictEvidenceMissing pieces
Context interfaceadjacentAn explicit known / non-ancestor / unknown graph interface conditions a single pretrained model at inference.No numeric time-series context experiment; needs variable identity, lag, provenance, and uncertainty semantics.
Causal structureadjacentAncestral context improves intervention-effect estimation under fully observed SCM assumptions.Temporal SCMs, hidden confounding, evolving structure, real intervention validation.
Control and counterfactualsinsufficient evidencePredicts interventional outcomes in synthetic and semi-synthetic tabular tasks.Action-sequence rollout, individual-level counterfactual semantics, feedback control, off-policy and overlap audits.
Benchmarks and evaluation hygienewarningWrong knowledge is worse than unknown; prior alignment and PSID imbalance change comparative conclusions.Matched-prior temporal tests, effect-strength stratification, uncertainty calibration under shift.

Transfer hypothesis: supply partial causal/structural context to a multivariate time-series model rather than either ignoring domain knowledge or forcing a complete topology. A clean first test compares identical numeric-history models with no structure, correct partial structure, unknown-marked omissions, and deliberately corrupted structure, at matched architecture and training budget. Score probabilistic forecasts and intervention-sensitive outcomes separately; do not count an observational forecasting gain as causal identification.

Open Questions

  • Can partial, lag-indexed ancestry context help a time-series model without collapsing unknown relationships into absent links?
  • How should a model represent disputed, stale, or confidence-weighted causal assertions instead of treating every supplied sign as correct?
  • Can a similar interface retain calibrated uncertainty with hidden confounding and changing system structure?
  • Does the benefit survive matched priors, intervention-effect-strength stratification, and evaluation on real intervention logs?