Video Can Teach PAN-Sharpening: PSF-Aware Cross-Domain Supervision
Abstract
Most PAN-sharpening networks are trained on synthetically degraded panchromatic (PAN) and multispectral (MS) pairs because real high-resolution MS (HRMS) labels are rarely available. However, real satellite acquisitions often suffer from cross-modal misregistration and sensor-dependent optical responses that are absent from synthetic training pipelines, leading to spectral distortions and double-edge artifacts at test time. This suggests that the key bottleneck lies not only in network design but also in the supervision paradigm. In this work, we show that natural videos can serve as a spatial teacher for PAN-sharpening: nearby HR frame pairs provide abundant high-frequency structures and exhibit frame-to-frame shifts that mimic realistic misregistration, enabling strong spatial supervision without HRMS labels. Based on this insight, we propose ViPS, a cross-domain training framework under the paradigm of “Video can teach PAN-Sharpening”. ViPS disentangles supervision sources by learning spatial detail restoration from video-derived pseudo PAN–MS pairs, while enforcing spectral fidelity from real satellite PAN–MS pairs. To align optical degradations across domains, we utilize point spread function (PSF) banks and randomly sample a kernel during training, enabling physically grounded cross-domain learning. Extensive experiments across multiple sensors show that ViPS outperforms very recent state-of-the-art methods under both reducedand full-resolution settings, while maintaining fast inference.