TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy
Abstract
While Video Virtual Try-on (VVT) has achieved remark-able progress in synthesizing realistic garment overlays on dynamic sub-jects, existing paradigms remains fundamentally constrained by a pas-sive dependency on source camera trajectories, failing to accommodatethe requisite interactive freedom for omnidirectional viewpoint explo-ration. To address this limitation, we define a pioneering research fron-tier: Camera-controllable Video Virtual Try-on (CaM-VVT). Unlike con-ventional VVT, CaM-VVT not only necessitates viewpoint-agnostic tex-ture hallucination but also strict structural synchronization between non-rigid human dynamics and background contexts under arbitrary, un-constrained camera movements. To tackle these challenges, we presentTryOnCrafter, the first unified DiT-based framework specifically archi-tected for the CaM-VVT task. Departing from implicit pixel-space ma-nipulation, we introduce a Renderable 4D Try-on Proxy that explicitlydecouples the human subject from the environment. This is achieved bydistilling high-fidelity 2D try-on priors into a clothed 3DGS-based avatar,which is subsequently animated via SMPL-X sequences and metric-alignedinto a reconstructed background point cloud. This proxy establishes a ro-bust structural foundation with superior texture density and motion in-tegrity. Our Proxy-Anchored Video DiT leverages this robust structuralfoundation as a primary geometric anchor, ensuring that the synthesizedphotorealistic videos are strictly constrained by prescribed trajectoriesand physically plausible deformations. Benefiting from the inherent ed-itability of the 4D proxy, TryOnCrafter facilitates diverse downstreamapplications, including human relocalization, “bullet time” effects, and360-degree orbital viewing. Extensive experiments on our establishedCaM-VVTBench demonstrate that TryOnCrafter significantly outper-forms existing baselines in preserving structural consistency and garmentidentity across complex camera maneuvers.