Your autoregressive model already reveals the causal graph
- Autoregressive models trained via next-token prediction implicitly learn the conditional independence structure of their data-generating process. We exploit this observation to perform scalable causal discovery from a single observed sequence of discrete events—without any task-specific retraining. Such single-stream settings arise naturally in vehicle diagnostics, manufacturing systems, and patient trajectories, yet they remain largely unsolved: the absence of repeated samples, massive event vocabularies, and long-range temporal dependencies render existing methods either inaccurate or computationally intractable. We introduce TRACE, a framework that repurposes any pretrained autoregressive model as a density estimator for conditional mutual information, the fundamental primitive for conditional independence testing. By constructing parallelized CI tests on GPUs, TRACE recovers both the sample-level time causal graph and its summary projection, scaling linearly with the vocabularyAutoregressive models trained via next-token prediction implicitly learn the conditional independence structure of their data-generating process. We exploit this observation to perform scalable causal discovery from a single observed sequence of discrete events—without any task-specific retraining. Such single-stream settings arise naturally in vehicle diagnostics, manufacturing systems, and patient trajectories, yet they remain largely unsolved: the absence of repeated samples, massive event vocabularies, and long-range temporal dependencies render existing methods either inaccurate or computationally intractable. We introduce TRACE, a framework that repurposes any pretrained autoregressive model as a density estimator for conditional mutual information, the fundamental primitive for conditional independence testing. By constructing parallelized CI tests on GPUs, TRACE recovers both the sample-level time causal graph and its summary projection, scaling linearly with the vocabulary size while naturally handling delayed causal effects. Crucially, we prove that minimizing the standard cross-entropy pretraining loss directly minimizes an upper bound on the causal identification error, establishing a duality between sequence prediction and causal discovery. On nonlinear SCMs () and real-world vehicle diagnostic logs (), TRACE is the first applicable method at this scale, outperforming the strongest baseline by over 20 F1 points.…


| Author: | Hugo MathORCiDGND, Rainer LienhartORCiDGND |
|---|---|
| Frontdoor URL | https://opus.bibliothek.uni-augsburg.de/opus4/131260 |
| URL: | https://openreview.net/forum?id=Q66hINx9fA |
| Parent Title (English): | ICML 2026 Workshop on Structured Probabilistic Inference & Generative Modeling (SPIGM), 10-11 July 2026, Seoul, South Korea |
| Type: | Conference Proceeding |
| Language: | English |
| Date of Publication (online): | 2026/06/19 |
| Year of first Publication: | 2026 |
| Publishing Institution: | Universität Augsburg |
| Release Date: | 2026/06/22 |
| Institutes: | Fakultät für Angewandte Informatik |
| Fakultät für Angewandte Informatik / Institut für Informatik | |
| Fakultät für Angewandte Informatik / Institut für Informatik / Lehrstuhl für Maschinelles Lernen und Maschinelles Sehen | |
| Dewey Decimal Classification: | 0 Informatik, Informationswissenschaft, allgemeine Werke / 00 Informatik, Wissen, Systeme / 004 Datenverarbeitung; Informatik |
| Latest Publications (not yet published in print): | Aktuelle Publikationen (noch nicht gedruckt erschienen) |


