Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
Abstract
Generating talking avatars is a fundamental task in videogeneration. Although existing methods can generate full-body talkingavatars with simple human motion, extending this task to groundedhuman-object interaction (GHOI) remains an open challenge, requiringthe avatar to perform text-aligned interactions with surrounding objects.This challenge stems from the need for environmental perception and thecontrol-quality dilemma in GHOI generation. To address this, we proposea novel dual-stream framework, InteractAvatar, which decouples per-ception and planning from video synthesis for grounded human-object in-teraction. Leveraging detection to enhance environmental perception, weintroduce a Perception and Interaction Module (PIM) to generate text-aligned interaction motions. Additionally, an Audio-Interaction Aware† ⋆Equal contribution. Corresponding author.Generation Module (AIM) is proposed to synthesize vivid talking avatarsperforming object interactions. With a specially designed motion-to-video aligner, PIM and AIM share a similar network structure and enableparallel co-generation of motions and plausible videos, effectively miti-gating the control-quality dilemma. Finally, we establish a benchmark,GroundInter, for evaluating GHOI video generation. Extensive exper-iments and comparisons demonstrate the effectiveness of our method ingenerating grounded human-object interactions for talking avatars.