Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

ContrasTED: contrastive domain embeddings for scalable remote homology classification

Created on 04 Sep 2026

Authors

Miller, D. M., Bordin, N., Jeyananthan, J., Waman, V., Heinzinger, M., Orengo, C.

Abstract

Protein structure prediction has expanded structural databases to hundreds of millions of domains. Classifying these domains into homologous superfamilies reveals evolutionary and functional relationships that can persist despite low sequence similarity. As the size of structural databases continues to grow, homology classification requires methods that combine scalability with accuracy. Here we present ContrasTED, which uses CATH-supervised center-contrastive learning to project structure-aware embeddings into a domain-level metric space for nearest-centroid superfamily assignment. On a sequence-filtered S20 benchmark (n = 1,028), superfamily assignment accuracy reached 92.9% (1-NN) and 91.4% (nearest centroid), exceeding sequence search, profile HMMs, Foldseek, and a classifier trained on embeddings. The learned latent space separates superfamilies while retaining structural information below 20% sequence identity, with the largest gains among sparsely represented superfamilies. ContrasTED produces 4.67 million new candidate assignments across 3,796 superfamilies in The Encyclopedia of Domains (TED), extending annotation coverage beyond previous structure-based methods.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 04 Sep 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 8
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement