Egocentric World Model for Photorealistic Hand Object Interaction Synthesis
Abstract
To serve as a scalable data source for embodied AI, worldmodels should act as true simulators that infer interaction dynamicsstrictly from user actions, rather than mere conditional video genera-tors relying on privileged future object states. In this context, egocentricHand Object Interaction (HOI) world models are critical for predictingphysically grounded first-person rollouts. However, building such modelsis profoundly challenging due to rapid head motions, severe occlusions,and high-DoF hand articulations that abruptly alter contact topolo-gies. Consequently, existing approaches often circumvent these physicschallenges by resorting to conditional video generation with access toknown future object trajectories. We introduce EgoHOI, an egocentricHOI world model that breaks away from this shortcut to simulate pho-torealistic, contact-consistent interactions from action signals alone. Toensure physical accuracy without future-state inputs, EgoHOI distillsgeometric and kinematic priors from 3D estimates into physics-informedembeddings. These embeddings regularize the egocentric rollouts towardphysically valid dynamics. Experiments on the HOT3D dataset demon-strate consistent gains over strong baselines, and ablations validate theeffectiveness of our physics-informed design.