WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
Abstract
Recent progress in image and video synthesis has inspiredtheir use in advancing 3D scene generation. However, we observe thattext-to-image and -video approaches struggle to maintain scene- andobject-level consistency beyond a limited environment scale without apersistent, explicit geometric representation. We thus present a geometry-first approach that decouples this complex problem of large-scale 3Dscene synthesis into its structural composition, represented as a meshscaffold, and realistic appearance synthesis, which leverages powerful im-age synthesis models conditioned on the mesh scaffold. From an inputtext description, we first construct a mesh capturing the environment’sgeometry (walls, floors, etc.), and then use image synthesis, segmentationand object reconstruction to populate the mesh structure with objects inrealistic layouts. This mesh scaffold is then rendered to condition imagesynthesis, providing a structural backbone for consistent appearance gen-eration. This enables scalable, arbitrarily-sized 3D scenes of high objectrichness and diversity, combining robust 3D consistency with photoreal-istic detail. We believe this marks a significant step toward generatingtruly environment-scale, immersive 3D worlds.