FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On
Abstract
Image-based Virtual Try-On (IVTON) has greatly advancedthrough diffusion models, yet existing methods require many samplingsteps and depend on masks with costly auxiliary networks. In addition,the absence of large-scale mask-free paired datasets further limits thedevelopment of mask-free IVTON. We propose FDM-MFVT, a few-stepdiffusion model for mask-free IVTON, integrating an Outfit-aware NoiseOptimization Module (OANO) and an Instruction-driven Try-on Module(IDT) to enhance efficiency and flexibility.The OANO module initializesthe alignment space with noise using the input image and only needs 6steps to generate a higher-fidelity try-on image compared to 30 steps.TheIDT module uses virtual try-on prompts and efficient adaptation to gen-erate high-quality results from garment and person images alone. Wefurther introduce MFVT, a 30,000-pair mask-free IVTON dataset. Ex-periments show that FDM-MFVT achieves superior quantitative andqualitative results with fewer inference steps than mask-based and mask-free baseline methods.