Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Prediction of microsatellite instability in colorectal cancer based on tissue phenotypes inferred from pathological whole slide images using self-distillation.

Created on 04 Aug 2026

Authors

Zhiwu Wang, Yankun Liu, Wei Xiong, Lei Wang, ShuXue Xi, Changcheng Lu, Yan Wu, Qingke Li, Chunling Liu, Jingwu Li, Yufeng Li

Published in

Pathology, research and practice. Volume 286. Pages 156637. Jul 30, 2026. Epub Jul 30, 2026.

Abstract

The application of Multiple Instance Learning (MIL) for classifying Whole Slide Images (WSIs) has gained extensive use in recent years, primarily due to the high cost and time consumption associated with pixel-level annotation of WSIs, which is challenging to accomplish. The advancements in MIL for WSIs have predominantly concentrated on two fronts: the development of superior feature extractors (for instance, utilizing self-supervised learning for training feature extractors) and the formulation of enhanced instance aggregation strategies. Regrettably, the majority of the most advanced approaches have neglected phenotypic variances among instances when employing attention mechanisms. To capitalize on the disparities between instance tissues, we have introduced a phenotypic self-distillation approach to MIL. Our framework is composed of three components: i) a self-supervised feature extractor based on contrastive learning and a phenotype extractor pre-trained on the Kather100K dataset, which automatically provides 9-class tissue phenotype labels (e.g., tumor epithelium, stroma, lymphocytes) without requiring manual annotation, ii) the incorporation of a self-distillation loss between the features of instances and their phenotypes to augment the informational content of both perspectives, and iii) the aggregation of MIL instances for the final MSI prediction. The efficacy of this framework was evaluated on two datasets: the TCGA-CRC dataset was used for training and internal testing with a fixed 70%/30% split, while the Tangshan People's Hospital cohort served as an independent external validation set. On the TCGA-CRC dataset (n = 360; 65 MSI-H, 295 MSS), our model achieved an AUC of 0.8846 and an accuracy of 0.84, using a fixed 70%/30% train-test split. On the Tangshan People's Hospital dataset (n = 472; 56 MSI-H, 426 MSS), the model attained an AUC of 0.7258 and an accuracy of 0.70.

PMID:
42546403
Bibliographic data and abstract were imported from PubMed on 04 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 3
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement