Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
Abstract
While recent video world models can generate highly real-istic videos, their ability to perform semantic reasoning and planningremains unclear and unquantified. We introduce Target-Bench, thefirst benchmark that enables comprehensive evaluation of video worldmodels’ semantic reasoning, spatial estimation, and planning capabil-ities. Target-Bench provides 450 robot-collected scenarios spanning 47semantic categories, with SLAM-based trajectories serving as motiontendency references. Our benchmark reconstructs motion from generatedvideos with a metric scale recovery mechanism, enabling the evaluationof planning performance with five complementary metrics that focus ontarget-approaching capability and directional consistency. Our evalua-tion result shows that the best off-the-shelf model achieves only a 0.368overall score, revealing a significant gap between realistic visual genera-tion and semantic reasoning in current video world models. Furthermore,we demonstrate that fine-tuning process on a relatively small real-worldrobot dataset can significantly improve task-level planning performance.