Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Behaviorally prioritized entity-relation structure captures human visual cortical representations of natural scenes

Created on 19 Sep 2026

Authors

Wu, Y., Jiang, W., Li, S.

Abstract

Understanding natural scenes requires identifying visible entities and representing how those entities are related. Recent studies have shown that artificial neural networks (ANNs), large language models (LLMs), and vision language models (VLMs) can predict visual cortical responses to natural images. However, the neural organization of relational scene meaning remains poorly understood, in part because these models typically encode scene content in global feature spaces that are difficult to decompose into separable entity and relation components. Here, we combined scene-graph annotations, behavioral measurements, and large-scale neural datasets to characterize structured relational representations during natural vision. We used RotatE, a knowledge-graph embedding model, to represent head-relation-tail triplets annotated for images from the 7T Natural Scenes Dataset. Triplet embeddings reliably captured cortical representational structure across the visual hierarchy. Behavioral judgments further revealed systematic differences in triplet accessibility associated with visual, relational, and graph properties. Prioritizing more behaviorally accessible triplets improved neural correspondence and explained unique variance beyond object co-occurrence, ANN image features, and LLM caption embeddings. Decomposing triplet representations into entity and relation components revealed partially dissociable cortical contributions, with lateral parietal cortex showing sensitivity to both. Triplet-based semantic information also remained spatially grounded: visual-field-specific triplet models preferentially predicted voxels with matching retinotopic preferences. Finally, cross-species comparison indicated that triplet-based semantic features were relatively more aligned with human high-level visual cortex than with macaque inferotemporal cortex. Together, these findings provide new insights into the representation of semantic relational information in the human visual cortex during natural scene perception.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 19 Sep 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 17
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement