TextFace: Compositional Text-Guided Identity Preserving Face Synthesis for Face Recognition
Abstract
Identity preserving face synthesis (IPFS) aims to generate large-scale virtual face datasets to train robust face recognition (FR) models while mitigating privacy and bias issues of web-crawled data. While diffusion-based IPFS methods have made huge progress, FR models trained on synthetic data still underperform those trained on real images. This gap stems from limited style diversity and insufficient intraidentity variation. Through analysis, we find that identity embeddings from real images primarily encode core facial features and inherently contain intra-identity variation. To generate more realistic virtual images, we propose TextFace, a diffusion-based IPFS framework that combines compositional text control for overall style diversity and a principled score-space identity blending strategy to enhance intra-identity variation. The identity-irrelevant condition factors are modeled by compositional language attributes, while identity-inherent factors are diversified by identity blending, which injects small, stable structural variations during denoising without sacrificing identity consistency. Extensive experiments show that FR models trained on TextFace synthetic data consistently outperform prior IPFS pipelines and substantially narrow the gap to real-data training.