QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
Abstract
Slot Attention, an approach that binds different objects in ascene to a set of "slots", has become a leading method in unsupervisedobject-centric learning. Most methods assume a fixed slot count K, andto better accommodate the dynamic nature of object cardinality, a fewworks have explored K-adaptive variants. However, existing K-adaptivemethods still suffer from two limitations. First, they do not explicitlyconstrain slot-binding quality, so low-quality slots lead to ambiguousfeature attribution. Second, adding a slot-count penalty to the recon-struction objective creates conflicting optimization goals between reduc-ing the number of active slots and maintaining reconstruction fidelity.As a result, they still lag significantly behind strong K-fixed baselines.To address these challenges, we propose Quality-Guided K-AdaptiveSlot Attention (QASA). First, we decouple slot selection from recon-struction, eliminating the mutual constraints between the two objectives.Then, we propose an unsupervised Slot-Quality metric to assess per-slotquality, providing a principled signal for fine-grained slot–object binding.Based on this metric, we design a Quality-Guided Slot Selection schemethat dynamically selects a subset of high-quality slots and feeds them intoour newly designed gated decoder for reconstruction during training. Atinference, token-wise competition on slot attention yields a K-adaptiveoutcome. We conduct experiments on both object discovery and objectproperty prediction. Results show that QASA substantially outperformsexisting K-adaptive methods on both real and synthetic datasets. More-over, on real-world datasets, QASA surpasses K-fixed methods. The codeis available at https://github.com/ouyangtianran/QASA-tianran.Fig. 1: mBOi vs. slot count K on COCO [25]. K-fixed methods, such as SPOT [20]and DINOSAUR [33], show large performance fluctuations as K varies. By contrast,K-adaptive methods do not rely on a carefully tuned K. Our method substantiallyoutperforms existing K-adaptive baselines, including MetaSlot [26] and AdaSlot [16],and even surpasses strong K-fixed baselines.