Authors
Yankovitz, G., Gat-Viks, I.
Abstract
Transcriptomic classification is often hindered by the small number of samples relative to the high dimensionality of gene expression data. We introduce General Gene Expression (GGE), a deep meta-learning framework designed to support robust classification in this limited-sample setting. By training across 5,220 distinct biomedical prediction objectives drawn from 1,779 different human datasets, GGE learns a model initialization that captures biological patterns shared across heterogeneous classification tasks. This learned initialization has two key advantages. First, it enables improved performance to new datasets using only a small number of labeled samples. Second, because it is learned across diverse prediction objectives, it can be applied to a broad range of biomedical problems. We show that GGE outperforms established classifiers in data-limited settings across a wide range of applications, including datasets generated using different RNA-seq platforms and preprocessing pipelines. In addition, attention-based analysis identifies recurrent genes that contribute to performance across multiple biological objectives, providing insight into shared determinants of human biological states. Together, these results establish GGE is a general-purpose framework for human transcriptome-based classification in biomedical settings where labeled data are scarce.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 03 Oct 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 14
- Comments 0