MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning
Abstract
Memorization in machine learning models enables high per-formance on rare in-distribution samples by capturing their atypicalpatterns. However, it also causes harmful retention of noise and out-liers, degrading generalization. While memorization has been extensivelystudied in both supervised and self-supervised learning in the visiondomain, it remains unexplored in multi-modal contrastive learning. Weaddress this gap by introducing MultiMem, the first metric designed toquantify memorization in multi-modal contrastive learning. Through oursystematic analysis, we demonstrate that cross-modal semantic misalign-ment has the strongest influence on memorization, with text being thedominant modality driving memorization, followed by video, image, andaudio. We show that targeted augmentations applied across all modalitieseffectively reduce memorization as measured by our MultiMem metricand improve model performance. Overall, this work establishes the firstframework for measuring and mitigating memorization in multi-modalcontrastive learning, preventing harmful data retention and contributingto higher-performing models. Full version with appendix available athttp://arxiv.org/abs/2606.22220