CausalVAE as a Plug-in for World Models: Towards Reliable Counterfactual Dynamics
Abstract
Latent world models roll observations forward accurately onfactual transitions, but their predictive latents are typically entangledand weakly aligned with the underlying causal factors; as a result, theydegrade under interventions and on counterfactual queries. We addressthis with CausalVAE, a plug-in structural module that attaches to di-verse encoder–transition backbones and organizes their latents througha learned directed acyclic graph (DAG) over causal factors, withoutaltering the base prediction architecture. Across four benchmarks andeight backbones, the plug-in largely preserves factual retrieval while im-proving intervention-aware counterfactual retrieval on most backbone–benchmark pairs; the gains are benchmark- and backbone-dependentrather than universal. Improvements are largest on the Physics bench-mark: for a graph-network backbone trained with a negative-log-likelihoodobjective, counterfactual H@1 (CF-H@1) rises from 11.0 to 41.0, withseveral other backbones gaining +10 to +30 CF-H@1 points. A causalanalysis further shows that the learned structure recovers meaningfulfirst-order physical interaction trends, supporting the interpretability ofthe latent causal structure.