NanoVSR: Towards Real-Time Video Super-Resolution on Edge Devices
Abstract
Recent Video Super-Resolution (VSR) methods rely heavilyon transformers and explicit optical flow, creating computational over-head and custom operations that hinder deployment on hardware acceler-ators like TensorRT. To address this, we introduce NanoVSR, a scalable,fully convolutional architecture designed for resource-constrained edgedevices. Using structural reparameterization, NanoVSR collapses intostandard convolutions during inference, ensuring seamless hardware com-patibility and negligible runtime overhead. Furthermore, despite lackingexplicit motion compensation, it maintains competitive restoration qual-ity by implicitly learning spatio-temporal alignments through progres-sive training. Evaluated on the REDS4 benchmark, NanoVSR demon-strates an exceptional balance between accuracy and computational ef-ficiency, significantly improving the trade-off for compact architectures.Our NanoVSR-644k baseline yields 28.64 dB PSNR on REDS4 whiledelivering 27.20 FPS on the NVIDIA Jetson Orin NX 16GB (25W),offering massive speed gains over heavier models. The scaled NanoVSR-1.7M variant reaches 29.15 dB with a throughput of 19.58 FPS, providingsuperior, edge-optimized upscaling.