Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
Abstract
Deep generative models, while revolutionizing image gener-ation, largely operate as opaque “black boxes”, hindering human un-derstanding, control, and alignment. While methods like sparse autoen-coders (SAEs) show remarkable empirical success, they often lack theo-retical guarantees, risking subjective insights. Our primary objective is toestablish a principled foundation for interpretable generative models. Wedemonstrate that the principle of causal minimality – favoring the sim-plest causal explanation – can endow the latent representations of gener-ative models with clear causal interpretation and robust, component-wiseidentifiable control. We introduce a novel theoretical framework for hi-erarchical selection models, where higher-level concepts emerge from theconstrained composition of lower-level variables, better capturing thecomplex dependencies in data generation. Under theoretically derivedminimality conditions that manifest as sparsity constraints, we show thatlearned representations can be equivalent to the true latent variables ofthe data-generating process. Empirically, applying these constraints totext-to-image diffusion models allows us to extract their innate hierarchi-cal concept graphs, offering fresh insights into their internal knowledgeorganization. Furthermore, these causally grounded concepts serve aslevers for fine-grained model steering, paving the way for transparent,reliable systems.