Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Foveation as self-supervision signal: Blur masks beat blank masks for learning robust visual representations

Created on 02 Oct 2026

Authors

Maruya, A., Adeli Jelodar, H., Zheng, T., Kriegeskorte, N., Qian, N.

Abstract

The primate retina is strikingly non-uniform, with receptor density falling off sharply from fovea to periphery. Every eye-movement brings a previously low-acuity peripheral area to high-acuity foveal processing. We hypothesize that this process provides a natural self-supervision signal, enabling representation learning by predicting fine details in the periphery. We implement this in a SimMIM-style vision transformer, replacing its blank masks (ViT-Blank) with a uniform blur mask (ViT-Blur) and further with retinal masks in which a foveal patch is randomly selected, and all other patches are progressively blurred based on eccentricity from the foveal patch (ViT-Retina). ViTs were pre-trained to reconstruct the full-resolution image from the corrupted input and then fine-tuned on clean images for classification. ViT-Blur and ViT-Retina outperform ViT-Blank on the pre-training reconstruction task while controlling for retained information from the input across corruption types. Critically, they also outperform ViT-Blank on downstream classification, showing that they learn representations that generalize better. The ViT-Blur and ViT-Retina are also more robust in impoverishment conditions such as higher mask ratios, fewer pre-training epochs, smaller training sets, and test time image corruptions. Using additive band-limited noise, we discovered that ViT-Blank relies on higher spatial frequencies compared to ViT-Blur and ViT-Retina, and we show this to be the likely cause of poorer generalization: the centroid of a model's spatial frequency tuning is negatively correlated with its classification accuracy ($r = -0.73$). We further show that a closed-form linear-ridge encoder derived from the same reconstruction objective reproduces this tuning ($r = 0.88$). Overall, our experiments show the promise of a biologically-inspired self-supervision objective for learning robust visual representations.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 02 Oct 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 16
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement