Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

RNA-guided contrastive learning enhances patient-level prediction from histology

Created on 01 Oct 2026

Authors

Zou, A., Huang, E., Wu, Q., Barnard, M. E., Zhang, C.

Abstract

Predicting molecular receptor status, including estrogen receptor (ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2), directly from routine hematoxylin and eosin (H&E)-stained whole-slide images (WSIs) could reduce reliance on costly immunohistochemistry, but image-only models have no access to the transcriptomic programs that define these subtypes. Here we show that coupling a whole-slide vision transformer (GigaPath) with bulk RNA-sequencing profiles through a training-time-only contrastive objective improves receptor-status classification from WSIs alone, without requiring any RNA-seq data at inference. Rather than training a dedicated bulk-transcriptomics encoder, we purpose the gene-sentence encoding and pretrained text encoder from BioBERT1 and apply them to bulk RNA-seq profiles from The Cancer Genome Atlas breast cancer cohort (TCGA-BRCA). Under five-fold cross-validation, this RNA-guided pre-training consistently improved AUROC for ER, PR, and HER2 relative to a contrastive-pre-training-free baseline and reproducibly introduced receptor-relevant structure into the learned slide representation. However, evaluation on independent external cohorts revealed a striking, receptor-dependent divergence: gains for ER and PR, both governed by broad, multi-gene luminal transcriptional programs, generalized robustly, whereas the apparent AUROC improvement for HER2, driven by focal ERBB2 amplification rather than a coordinated transcriptional signature, masked a severe loss of classification sensitivity at the operating threshold. These findings establish training-time transcriptomic supervision as a practical route to enhancing histology-only molecular subtyping, while showing that gene-selection strategy, and not model architecture alone, determines which molecular alterations such supervision can teach a vision model to recognize.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 01 Oct 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 13
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement