Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Active Learning Enables Efficient Directed Evolution of a Far-Red Fluorescent Protein with Minimal Experimental Data

Created on 14 Aug 2026

Authors

Brown, D. V., Cross, R. S., Zhu, S., Hill, T., Sok, C. L., Jenkins, M. R., Dramicanin, M., Bowden, R.

Abstract

Fluorescent proteins are fundamental tools for cellular imaging. Most fluorescent proteins in routine use, including GFP, are derived from the jellyfish Aequorea victoria and emit blue-green light, which is strongly absorbed and scattered by tissue, limiting imaging depth. Far-red and near-infrared fluorescent proteins, engineered from bacteriophytochromes, address this limitation because far-red light penetrates tissue considerably further. However, these proteins are typically much dimmer than their A. victoria-derived counterparts. Improving brightness by conventional directed evolution requires screening large random mutant libraries, a process that is slow, labor-intensive, and often impractical outside specialized laboratories. We utilized an active-learning-guided directed evolution workflow that identified improved variants from substantially less data than conventional screening. Each round coupled automated, miniaturized cell-free protein expression directly from a DNA template without cloning or cell culture, with a machine-learning model retrained on cumulative sequence--function data to nominate the most informative variants for the next round. Applied to miRFP670nano3, this workflow screened 120 variants across successive rounds and identified twelve with improved brightness, the best four-fold brighter in bacterial systems. However, these gains did not translate when the variants were evaluated in mammalian cells, indicating that performance can be strongly dependent on cellular context. Retrospective simulation across benchmark datasets from ProteinGym showed that performing more experimental batches with fewer samples per batch consistently accelerated convergence to high-fitness sequences. Incorporating protein-language-model derived zero-shot fitness priors also accelerated convergence, but only in proportion to how well each prior score correlated with the true fitness landscape. Together, these findings established generalizable design rules, favoring smaller acquisition batches and confidence-weighted priors, for engineering proteins from minimal experimental data.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 14 Aug 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 5
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement