Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
Abstract
Latent diffusion models pair VAEs with diffusion backbones,and the structure of VAE latents strongly influences the difficulty of dif-fusion training. However, existing video VAEs typically focus on recon-struction fidelity, overlooking latent structure. We present a statisticalanalysis of video VAE latent spaces and identify two spectral propertiesessential for diffusion training: a channel-wise eigenspectrum dominatedby a few modes, and a spatio-temporal frequency spectrum biased towardlow frequencies. To induce these properties, we propose two lightweight,backbone-agnostic regularizers: Latent Masked Reconstruction and Lo-cal Correlation Regularization. Experiments show that our Spectral-Structured VAE (SSVAE) achieves a 3× speedup in text-to-video gen-eration convergence and a 10% gain in video reward, outperformingstrong open-source VAEs. Code is available at: https://github.com/zai-org/SSVAE.