Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets
Abstract
Dataset distillation (DD) condenses large corpora into com-pact, information-rich subsets for efficient training and reuse. However,under noisy supervision, DD risks condensing corrupted associations to-gether with useful signals, degrading robustness. Conventional noisy-label remedies (sample selection, loss weighting, label correction) tightlycouple noise estimation with model optimization, often require clean an-chors, and can amplify confirmation bias—assumptions that are mis-aligned with DD’s goal of compact, plug-and-play supervision. We there-fore propose a trajectory-based DD framework that jointly suppressesnoise and preserves transferable knowledge without relabeling or cleansubsets. It comprises two complementary components: Selective Guid-ance Reweighting (SGR), which fuses global forgetting patterns (second-split forgetting) with local neighborhood consistency into a progressivereweighting scheme that prioritizes clean supervision along the teachertrajectory; and Teacher-Inspired Auxiliary Targets (TIAT), which injectauxiliary residual guidance distilled from intermediate teacher dynamicsto reinforce informative signals while remaining internally consistent. To-gether, SGR and TIAT produce distilled datasets with cleaner and richerrepresentations under noisy supervision. The framework is robust, label-preserving, computationally lightweight, and broadly applicable, yield-ing consistent gains over state-of-the-art DD baselines across symmetric,asymmetric, and real-world noise.