Infinite Gaze Generation for Videos with Autoregressive Diffusion
Abstract
Predicting human gaze in video is fundamental to advanc-ing scene understanding and multimodal interaction. While traditionalsaliency maps provide spatial probability distributions and scanpaths of-fer ordered fixations, both abstractions often collapse the fine-grainedtemporal dynamics of raw gaze. Furthermore, existing models are typi-cally constrained to short-term windows (≈ 3–5s), failing to capture thelong-range behavioral dependencies inherent in real-world content. Wepresent a generative framework for infinite-horizon raw gaze predictionin videos of arbitrary length. By leveraging an autoregressive diffusionmodel, we synthesize gaze trajectories characterized by continuous spa-tial coordinates and high-resolution timestamps. Our model is condi-tioned on a saliency-aware visual latent space. Quantitative and qualita-tive evaluations demonstrate that our approach significantly outperformsexisting approaches in long-range spatio-temporal accuracy and trajec-tory realism. Project website: https://www.immersivecomputinglab.org/publication/infinite-gaze/.