SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction
Abstract
Deep learning models have become the standard tool for doc-ument rectification and illumination correction, yet their performance isfundamentally bound by their training data. For nearly a decade, the com-munity has heavily relied on Doc3D, a pioneering but increasingly limiteddocument unwarping dataset in terms of scale and quality. To address thisbottleneck, we introduce SyntheticDoc, a massive, high-quality datasetdesigned to push the boundaries of document unwarping. SyntheticDocis composed of 1,000,000 high-resolution procedurally generated trainingsamples, alongside extensive validation and test sets. Each sample ispaired with rich, pixel-perfect annotations, including UV maps, normalmaps, albedo and shading. To ensure physical accuracy and photorealism,the paper geometries are generated via a physics-based simulator andrendered using a path tracer. To demonstrate the benefit of our dataset,we train a simple baseline model on SyntheticDoc and report on itsperformance in comparison to state-of-the-art methods on both documentunwarping and illumination correction tasks. Our dataset is available athttps://igl.ethz.ch/projects/SyntheticDoc/ and the code used to generateit at https://github.com/tanguymagne/SyntheticDoc.