FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
Abstract
Visual Autoregressive modeling (VAR) has emerged as ahighly efficient alternative to diffusion-based frameworks, achieving com-parable synthesis quality. However, as this paradigm extends to Space-time Autoregressive modeling (STAR) for video generation, scaling res-olution and frame counts leads to a "token explosion" that creates amassive computational bottleneck in the final refinement stages. To ad-dress this, we propose FastSTAR, a training-free acceleration frame-work designed for high-quality video generation. Our core method, Spa-tiotemporal Token Pruning, identifies essential tokens by integratingtwo specialized terms: (1) Spatial similarity, which evaluates structuralconvergence across hierarchical scales to skip computations in regionswhere further refinement becomes redundant, and (2) Temporal similar-ity, which identifies active motion trajectories by assessing feature-levelvariations relative to the preceding clip. Combined with a Partial Up-date mechanism, FastSTAR ensures that only non-converged regions arerefined, maintaining fluid motion while bypassing redundant computa-tions. Experimental results on InfinityStar demonstrate that FastSTARachieves up to a 2.01× speedup with a PSNR of 28.29 and less than 1%performance degradation, proving a superior efficiency-quality trade-offfor STAR-based video synthesis.