Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Systematic image perturbations reveal persistent gaps between human and machine vision

Created on 07 Aug 2026

Authors

Kato, M., He, B. J.

Abstract

Deep neural networks (DNNs) are promising computational models for understanding visual object recognition. Yet, whether DNNs use similar visual cues for object recognition as humans do remains unknown. We created an image set that systematically untangles global shape, internal parts, and texture information, and compared human recognition behavior against >200 DNNs spanning diverse architectures, training diets, and training objectives. No DNNs replicated humans' cue-reliance profile, including those with recurrence or specialized training. Fine-tuned text-image contrastive-trained models, regardless of architecture, were most human-like overall, but lost their human-alignment when the global shape was disrupted. Strikingly, all DNNs substantially underperformed humans when the global shape cue alone was critical to object recognition. Furthermore, alignment with ventral stream neural recordings in an existing database did not predict alignment to human behavior, and model performance does not always predict its human-alignment. Together, these findings reveal systematic and persistent differences between human and machine vision.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 07 Aug 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 37
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement