Steering 3D Generations: Preference Alignment via Direct Reward and Preference Optimization
Abstract
Generating high-fidelity and controllable 3D assets from a single image remains a significant challenge. We propose a novel framework that integrates preference optimization into a rectified flow-based diffusion model operating on a latent set representation. This approach enhances geometric fidelity and user alignment, enabling the efficient synthesis of diverse 3D formats. Our core contribution is a versatile reward-based optimization strategy for aligning generation with human-centric criteria like surface quality. We introduce Direct Reward Optimization (DRO), which fine-tunes the model using only absolute quality feedback, removing the need for pairwise data. The framework also seamlessly incorporates Direct Preference Optimization (DPO) when paired data is available. Parameter-efficient control is achieved via Low-Rank Adaptation (LoRA). To ensure data quality, we also develop a robust preprocessing pipeline using a neural SDF reconstruction method with Gaussian curvature constraints to produce high-quality watertight meshes. Extensive experiments show our framework achieves state-of-the-art performance. Quantitative evaluations demonstrate significant improvements in Chamfer Distance and F-Score, while qualitative results confirm superior geometric fidelity and user preference alignment compared to existing methods.