Bottom-up modeling of repeated elements via single image analysis-by-synthesis
Abstract
We address the problem of discovering repeated elementsfrom a single image. In contrast to existing approaches that depend onlarge annotated datasets, curated multi-image collections, or object seg-mentation masks, we show that a single image can suffice to learn ameaningful object model in a completely bottom-up fashion, without anyprior knowledge beyond a coarse scale prior. Our method learns a tunableimage-space prototype of the repeated elements through a reconstructionobjective, enabling the model to identify and synthesize consistent objectinstances within the same image. Experiments on 116 real images fromthe FSC-147 dataset demonstrate that our method successfully learnscoherent element models and captures intra-category variation on chal-lenging images. Qualitative results reveal superior reconstructions andinterpretable decompositions compared to classical decomposition, jointalignment, and 3D object modeling methods, while maintaining a simple2D formulation. These results suggest that meaningful object discoverycan emerge from single-image learning alone.