UniCSG: Unified High-Fidelity content-constrained style-driven generation via Staged Semantic and Frequency Disentanglement
Abstract
Style transfer must match a target style while preservingcontent semantics. DiT-based diffusion models often suffer from con-tent–style entanglement, leading to reference-content leakage and unsta-ble generation. We present UniCSG, a unified framework for content-constrained, style-driven generation in both text-guided and reference-guided settings. UniCSG employs staged training: (i) a latent-space se-mantic disentanglement stage that combines low-frequency preprocess-ing with conditioning corruption to encourage content–style separation,and (ii) a latent-space frequency-aware detail reconstruction stage thatrefines details via multi-scale frequency supervision. We further incorpo-rate pixel-space reward learning to align latent objectives with percep-tual quality after decoding. Experiments demonstrate improved contentfaithfulness, style alignment, and robustness in both settings.