MotionEditGS: Editing Motion and Appearance of 4D Scenes from Monocular Video via Semantically Anchored Gaussians
Abstract
Editing dynamic 3D scenes reconstructed from monocularvideos is important for applications such as content creation and dataaugmentation, but remains challenging. Prior works adapt image diffu-sion models across frames while enforcing spatial–temporal consistency,yet they often produce uniform, appearance-only edits. We instead lever-age a video diffusion model to enable photorealistic 4D edits and control-lable motion changes. In monocular settings, however, fitting dynamic3D Gaussians with pure photometric loss is under-constrained, lead-ing to shallow geometric reconstructions and temporal artifacts in theedited scene. We address this with a novel 4D editing pipeline that usescompact latents distilled from DINO features as additional supervisionduring scene optimization and preserves them during editing as a se-mantic identity signal. Such semantic consistency regularization providesstrong geometric and motion cues, enabling robust scene editing evenunder large or complex edits from the video diffusion model. Further-more, to promote semantically novel content generation during editing,we apply a gradient-based feature lifting method based on accumulatedfeature gradients during Gaussian densification. Experiments show im-proved spatial-temporal realism and text alignment over image editingbaselines, and demonstrate compositional motion-appearance edits withsignificant geometry changes that prior methods cannot achieve.