Thinking Ahead: Foresight Intelligence in MLLMs and World Model
Abstract
In this work, we introduce FSU-QA, a VQA dataset for au-tonomous driving scenarios designed to advance research on ForesightIntelligence—the ability to anticipate and reason about complex, long-horizon futures. Unlike existing benchmarks that mainly focus on imme-diate perception or reactive planning, FSU-QA evaluates future-orienteddriving understanding through multi-agent-aware and rule-grounded coun-terfactual QA. Rather than holding surrounding agents fixed or target-ing geometric path generation, our benchmark requires models to infersemantic future outcomes from front-view historical observations andpast ego trajectories. A comprehensive evaluation on the accompanyingFSU-Bench reveals that state-of-the-art VLMs still face significant chal-lenges in anticipating future events. Furthermore, beyond model perfor-mance, we examine whether WM-generated predictions remain seman-tically consistent by using VLM-based proxy judges, and validate thisevaluation protocol through shuffled control experiments. Fine-tuningmodels on FSU-QA leads to substantial improvements in foresight under-standing, demonstrating the dataset’s effectiveness and offering a prin-cipled foundation for future research.