Authors
Miller, D. M., Bordin, N., Jeyananthan, J., Waman, V., Heinzinger, M., Orengo, C.
Abstract
Protein structure prediction has expanded structural databases to hundreds of millions of domains. Classifying these domains into homologous superfamilies reveals evolutionary and functional relationships that can persist despite low sequence similarity. As the size of structural databases continues to grow, homology classification requires methods that combine scalability with accuracy. Here we present ContrasTED, which uses CATH-supervised center-contrastive learning to project structure-aware embeddings into a domain-level metric space for nearest-centroid superfamily assignment. On a sequence-filtered S20 benchmark (n = 1,028), superfamily assignment accuracy reached 92.9% (1-NN) and 91.4% (nearest centroid), exceeding sequence search, profile HMMs, Foldseek, and a classifier trained on embeddings. The learned latent space separates superfamilies while retaining structural information below 20% sequence identity, with the largest gains among sparsely represented superfamilies. ContrasTED produces 4.67 million new candidate assignments across 3,796 superfamilies in The Encyclopedia of Domains (TED), extending annotation coverage beyond previous structure-based methods.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 04 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 8
- Comments 0