Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

SemVac: A Semantic Vaccinology Paradigm Powered by LLMs for Antigen Discovery

Created on 18 Jul 2026

Authors

Zhao, Y., Shu, Y., Shu, L., Lv, P., Chi, X., Li, D., Zhang, J., Huang, Z., Ren, H., Xu, J., Zai, X., Chen, W.

Abstract

Reverse vaccinology has enabled sequence-based antigen discovery, but it overlooks the rich semantic knowledge embedded in the biomedical literature. Here we establish Semantic Vaccinology (SemVac), a paradigm that leverages large language models (LLMs) to predict protective antigens directly from scientific text. Benchmarking 14 state-of-the-art LLMs on a curated antigen dataset shows that text-reasoning-based approaches match or exceed specialized deep learning models in precision, while offering superior robustness on functionally ambiguous proteins. Intriguingly, explicit reasoning modes (e.g., chain-of-thought) increase recall but consistently reduce precision, revealing an over-reasoning pitfall in biological discovery tasks. Applied to the complete proteome of Mpox virus, SemVac recapitulates known protective antigens and identifies previously unrecognized candidates such as B20R, which our semantic analysis links to immune evasion and structural exposure. This work establishes literature-driven semantic reasoning as a powerful complement to conventional vaccinology, with broad implications for AI-aided scientific discovery.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 18 Jul 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 7
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement