Inference-time Motion Calibration for Video Generation
Abstract
Modern video generation models, despite high-quality frame synthesis, still struggle with motion fidelity and temporal consistency, causing artifacts and semantic drift; current solutions typically require extensive backbone retraining. We introduce a training-free framework for inference-time motion calibration that intervenes in the sampling trajectory of pretrained models to correct nascent inconsistencies efficiently. Our framework has two complementary components. First, Dynamic Latent Control (DLC) at selected steps estimates a clean latent, evaluates a VFM-based motion–semantic loss on proxy frames, and applies a corrective update to steer the trajectory. To control overhead, Amortization Trajectory Rectification (ATR) interleaves sparse calibration with fast amortized steps by caching a rectification field. Second, Dynamic RoPE Control (DRC) modulates temporal RoPE in attention blocks from the same motion signals, strengthening long-range coherence or sharpening short-range dynamics when flicker is detected. Extensive experiments on VBench and VBench 2.0 show that our method improves advanced motion fidelity metrics while incurring a small trade-off on superficial temporal smoothness for some backbones.