STEP: Spatial Thinking and Egocentric Pointing for Embodied Instruction Following
Abstract
Embodied Instruction Following requires agents to executenatural-language tasks via holistic perception and sound planning. Cur-rent LLM-driven agents suffer from two primary bottlenecks: SpatialMyopia, where limited egocentric views hinder global topological under-standing, and Granularity Imbalance, where planning is polarized be-tween inefficient atomic actions that lack interpretability and ambiguoussubgoals that fail to align with low-level control. To address these is-sues, we propose STEP, a framework that empowers agents with SpatialThinking and Egocentric Pointing. For perception, STEP integrates Hy-brid Map-Egocentric Perception, fusing multi-view streams into Bird’s-Eye-View maps to maintain both global context and local visual cues.For planning, STEP introduces Point-Level Planning, which derives spa-tial waypoints through explicit reasoning chains, achieving highly inter-pretable, intermediate-granularity execution. We also design STEP-CoT,a scalable data engine, to generate a 400k Chain-of-Thought datasetthat aligns diverse modalities—views, maps, actions, and instructions—through reasoning chains. Results on ALFRED and AI2-THOR-Nav,along with real-world deployments, demonstrate STEP’s superior perfor-mance and generalization. Ultimately, STEP achieves a unified paradigm:Look at the View, Think with the Map, and Point to the Goal.