Thermo-JEPA: Learning a Geometry-Grounded Thermal World Model via Cross-Modal Privileged Masking
Abstract
World models have driven remarkable progress in physicalperception and dynamic forecasting, yet current mainstream implemen-tations predominantly rely on visible-light (RGB) observations, whichlimits their all-weather reliability. Developing a thermal world model isessential, but it is hindered by the unique distribution of thermal imageryand the extreme scarcity of continuous video datasets. In this paper, weidentify a fundamental challenge in extending standard masked modelingto the thermal domain: intra-frame spatial homogenization. Due to ther-mal equilibrium, distinct physical entities often exhibit near-zero temper-ature gradients. Consequently, unconstrained self-attention blurs seman-tic boundaries and collapses object topologies, severely degrading down-stream dynamics forecasting. To overcome this, we propose Thermo-JEPA, a geometry-grounded pre-training framework based on Learn-ing Using Privileged Information (LUPI). We extract intra-frame spatialtopologies from aligned RGB data and inject them as a confidence-gatedstructural prior into the thermal teacher’s attention mechanism. Thisexplicitly enforces semantic segregation without corrupting fine-grainedintra-object thermal gradients. Concurrently, we release RGBT-World,the largest unified, spatiotemporally aligned RGB-Thermal video corpustailored for generative pre-training. Extensive experiments demonstratethat Thermo-JEPA effectively prevents representation collapse, achiev-ing state-of-the-art zero-shot dynamics alignment on pure thermal inputsand significantly outperforming massive generic video foundation mod-els.