Cambrian-P: Pose-Grounded Video Understanding
Abstract
Camera pose matters. The position and orientation of eachviewpoint define a shared spatial coordinate frame that relates observa-tions across video frames. Yet this signal is absent from multimodal LLMs(MLLMs) for video understanding, which process frames as isolated 2Dsnapshots, instead of the persistent scene humans perceive. We revisitpose as a lightweight supervisory signal and introduce Cambrian-P , avideo MLLM augmented with per-frame learnable camera tokens and apose regression head. With a carefully designed sampling scheme, themodel achieves substantial gains of 4.5–6.5% on spatial reasoning bench-marks such as VSI-Bench, generalizes across eight additional spatial andgeneral video QA benchmarks, and, as a byproduct, achieves state ofthe art streaming pose estimation on ScanNet. Surprisingly, training onpseudo-annotated poses from in-the-wild video further improves generalvideo QA benchmarks, showing pose helps beyond spatial reasoning. To-gether, these results position camera pose as a fundamental signal forvideo models that reason about the physical world.