ECHO: Ego-centric Modeling of Human-Object Interactions
Abstract
Modeling human-object interactions (HOI) from an egocen-tric perspective is a critical yet challenging task, particularly when re-lying on sparse signals from wearable devices like smart glasses andwatches. We present ECHO, the first unified framework to jointly recoverhuman pose, object motion, and contact dynamics solely from head andwrist tracking. To tackle the underconstrained nature of this problem,we introduce a novel tri-variate diffusion process with independent noiseschedules that models the mutual dependencies between the human, ob-ject, and interaction modalities. This formulation allows ECHO to oper-ate with flexible input configurations, making it robust to intermittenttracking and capable of leveraging partial observations. Crucially, it en-ables training on a combination of large-scale human motion datasetsand smaller HOI collections, learning strong priors while capturing in-teraction nuances. Furthermore, we employ a smooth inpainting inferencemechanism that enables the generation of temporally consistent interac-tions for arbitrarily long sequences. Extensive evaluations demonstratethat ECHO achieves state-of-the-art performance, significantly outper-forming existing methods lacking such flexibility.