NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment
Abstract
We introduce the Noise-Tilted Reverse Kernel (NTRK), areward-guided diffusion sampler that injects reward gradients through thenoise term, leaving the pretrained reverse kernel unchanged and requiringonly a single sample per step. Reward-guided sampling at inference timehas greatly expanded the versatility of pretrained diffusion models. Yet ex-isting methods face a trade-off. Gradient-based guidance shifts the reversemean, steering generation but pushing intermediate states outside theregion that the model was trained on and degrading quality. Search-basedmethods preserve quality but gain no gradient signal. No prior methodachieves both. NTRK resolves this by keeping the reverse mean fixed andbiasing the noise term toward high reward. This is enabled by a whiteningoperator, the central mechanism behind NTRK, which converts rewardgradients into noise-compatible perturbations without losing their guidingsignal. Across various reward alignment tasks, NTRK outperforms recentstate-of-the-art baselines without losing sample quality. Remarkably, onaesthetic generation, NTRK surpasses the reward of the best baseline at500 NFEs using only 25 NFEs, a 20× reduction in compute.