Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning

Published in ICLR 2025 — The Thirteenth International Conference on Learning Representations, Singapore, 2025

Masked image modeling has become a standard recipe for vision self-supervised learning, but random masking treats every region of an image as equally informative. This paper introduces a frequency-guided masking strategy: the frequency characteristics of an image determine which regions are hidden, so the pretext task concentrates on the components that carry the most structural information. The result is a self-supervised objective with markedly better data and compute efficiency — competitive downstream performance from a fraction of the pre-training budget usually required — and stronger transfer to fine-grained downstream vision tasks.