RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
Abstract
Recent advances in generative super-resolution (SR) havegreatly improved visual realism, yet existing evaluation and optimizationframeworks remain misaligned with human perception. Full-Referenceand No-Reference metrics often fail to reflect perceptual preference, ei-ther penalizing semantically plausible details due to pixel misalignmentor favoring visually sharp but inconsistent artifacts. Moreover, most SRmethods rely on ground-truth (GT)–dependent distribution matching,which does not necessarily correspond to human judgments. In this work,we propose RefReward-SR, a low-resolution (LR) reference-aware rewardmodel for preference-aligned SR. Instead of relying on GT supervisionor NR evaluation, RefReward-SR assesses high-resolution (HR) recon-structions conditioned on their LR inputs, treating the LR image asa semantic anchor. Leveraging the visual–linguistic priors of a Multi-modal Large Language Model (MLLM), it evaluates semantic consistencyand plausibility in a reasoning-aware manner. To support this paradigm,we construct RefSR-18K, the first large-scale LR-conditioned preferencedataset for SR, providing pairwise rankings based on LR–HR consistencyand HR naturalness. We fine-tune the MLLM with Group Relative Pol-icy Optimization (GRPO) using LR-conditioned ranking rewards, andfurther integrate GRPO into SR model training with RefReward-SR asthe core reward signal for preference-aligned generation. Extensive ex-periments show that our framework achieves substantially better align-ment with human judgments, producing reconstructions that preservesemantic consistency while enhancing perceptual plausibility and visualnaturalness.