Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Leveraging latent space models for enzyme discovery and sampling.

Created on 02 Sep 2026

Authors

Chang-Hwa Chiang, Daniel Ong, Alison R H Narayan, Charles L Brooks

Published in

Proceedings of the National Academy of Sciences of the United States of America. Volume 123. Issue 36. Pages e2608891123. Sep 08, 2026. Epub Sep 01, 2026.

Abstract

The exponential growth in available protein sequence data has broadened enzyme discovery opportunities but simultaneously highlighted a significant gap between sequence and function information. Traditional tools like phylogenetic trees and sequence similarity networks (SSNs) are widely adopted for sampling enzymes for novel transformations. However, their utility suffers from inherent limitations, which are exacerbated for large enzyme families. Phylogenetic trees, while useful for studying evolutionary relationships, become computationally intensive and difficult to visualize for larger protein datasets. SSNs, on the other hand, are sensitive to user-defined thresholds for sequence clustering and easily fail to capture more distant relationships between clusters. Additionally, both tools are alignment based and cannot capture higher-order interactions between residues. In this study, we address these limitations by optimizing a variational autoencoder (VAE)-based latent space model to visualize and explore enzyme sequence-function landscapes. By training our models on simulated datasets and real enzyme families, such as cyclases and flavin-dependent monooxygenases (FDMOs), we demonstrated that the optimized latent space effectively preserves phylogenetic relationships and enables high-resolution clustering for functionally distinct enzymes. The models further outperform traditional SSNs in capturing local and global relationships in a continuous two-dimensional space, enabling the discovery of multiple uncharacterized FDMOs for oxidative dearomatization and decarboxylative hydroxylation that illustrates their application. Our findings show that low-dimensional latent spaces can serve as valuable tools for enzyme discovery, allowing for interpolation and extrapolation to guide novel enzyme sampling for biocatalytic reactions.

PMID:
42679024
Bibliographic data and abstract were imported from PubMed on 02 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 11
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement