Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
Abstract
How can we infer a 3D hand pose when large portions of the handare heavily occluded by itself or by objects? Humans often resolve such ambigui-ties by leveraging contextual knowledge—such as affordances, where an object’sshape and function suggest how the object is typically grasped. Inspired by thisobservation, we propose a generative prior for 3D hand pose modeling guidedby affordance-aware textual descriptions of hand-object interactions (HOI). Ourmethod employs a diffusion-based generative model that learns the distributionof plausible hand poses conditioned on contextual signals, such as affordancedescriptions and image features. The affordance descriptions are designed to rep-resent the semantic intent and geometric structure of HOI, using the reasoningfrom a vision-language model (VLM) and grasp classification. We leverage thediffusion prior to refine the 3D pose predictions in hand reconstruction into moreaccurate and functionally coherent estimation. Our experiments demonstrate thatour affordance-guided refinement significantly improves 3D hand pose estima-tion performance on 3D hand affordance datasets, HOGraspNet and HO3D, overstate-of-the-art methods such as foundation models for hand reconstruction andthe latest diffusion priors for 3D hands.