Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures
Abstract
We propose HeadsUp, a scalable feed-forward method forreconstructing high-quality 3D Gaussian heads from large-scale multi-camera setups. Our method employs an efficient encoder-decoder archi-tecture that compresses input views into a compact latent representation.This latent representation is then decoded into a set of UV-parameterized3D Gaussians anchored to a neutral head template. This UV represen-tation decouples the number of 3D Gaussians from the number andresolution of input images, enabling training with many high-resolutioninput views. We train and evaluate our model on an internal dataset withmore than 10 000 subjects, which is an order of magnitude larger than ex-isting multi-view human head datasets. HeadsUp achieves state-of-the-artreconstruction quality and generalizes to novel identities without test-timeoptimization. We extensively analyze the scaling behavior of our modelacross identities, views, and model capacity, revealing practical insightsfor quality-compute trade-offs. Finally, we highlight the strength of ourlatent space by showcasing two downstream applications: generating novel3D identities and animating the 3D heads with expression blendshapes.