Authors
Nelson, D. R., Plouviez, M., Daakour, S., Jaiswal, A., Fu, W., Amin, S. A., Salehi-Ashtiani, K.
Abstract
Phytoplankton drive ~46% of global primary production, yet whether their protein repertoires encode ocean conditions - and whether reference databases capture the environmentally responsive fraction - is unknown. We used a protein language model to classify 447.7 million proteins from 2,357 samples and analyzed domain profiles for 231.7 million algal proteins across 2,044 samples. Under spatial block cross-validation, environment predicted individual domain abundances at R2 up to 0.59, while domain profiles predicted sea surface temperature at R2 = 0.38. The strongest coupling lay beyond annotated sequence space: 33,950 families clustered from 201 million Pfam-dark proteins coupled to environment 2.29-fold more strongly than Pfam domains. Cross-taxonomic selection, protein language modeling, and AlphaFold 3 predictions of 138 well-folded, InterPro-unannotated representatives support the Pfam-dark proteome as a biologically structured evolutionary compartment. Thus, the proteins most tightly coupled to ocean state are those least represented in reference databases.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 29 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 14
- Comments 0