Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
Abstract
Visual autoregressive (VAR) models have recently emerged as a promising alternative to diffusion models, achieving competitive performance in text-to-image generation while offering substantially faster inference. Although conditional image generation has been widely studied, training-free prompt-guided image editing remains largely unexplored despite its importance for practical applications. In this paper, we present Visual AutoRegressive Inverse Noise (VARIN), the first noise inversion framework for training-free text-based image editing in visual autoregressive models. VARIN introduces Location-aware Argmax Inversion (LAI), a pseudo-inverse function for the argmax operator that reconstructs the inverse Gumbel noise used during discrete autoregressive sampling. The recovered inverse noise enables accurate reconstruction of the source image while providing controllable guidance for promptdriven image editing. Extensive experiments demonstrate that VARIN produces edits that closely follow the target prompts while effectively preserving the background and structural details of the source image. These results establish VARIN as a simple, efficient, and practical image editing framework for visual autoregressive models.