SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
Abstract
We present SynCity 3000, a framework for generating 3Dscenes that are globally coherent while enabling fine-grained layout con-trol. Building on the ability of current image-to-3D generators to pro-duce complex 3D assets from a single image, we extend this capabilityto the scale of entire scenes by adapting the generator to be applicableas a convolutional operator. We achieve this by fine-tuning the modelon scene-like data generated by a new synthetic data engine, which wepropose to address the scarcity of 3D scene data for training. The convo-lutional generator is then applied to a dimetric image of the entire scene,generated from the user prompt, resulting in 3D scenes of arbitrary sizeand complexity. Across diverse prompts and layouts, SynCity 3000 pro-duces large, coherent, and detailed scenes, addressing the shortcomingsof prior approaches to 3D scene generation.