Authors
Liu, A., Ho, A., Droste, A. M., Martin, D., Wong, E., Zhou, E., Zhou, I., Park, J., Jiao, J., Skelly, K.-R., Kim, K., Li, J., Rao, K., Uehara, M., Marion, M., Fitzgerald, N., Dias, R., Shringarpure, S., Yuan, Y., Wang, Y.
Abstract
We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. The majority of existing life sciences benchmarks have a narrow scope or are purely knowledge-based, and therefore fail to capture the complexity of real-world research, which often involves ambiguities and requires the accurate execution of multiple dependent judgment calls. Additionally, almost all existing benchmarks span at best a small collection of subdomains within the life sciences; there is at present no existing life sciences benchmark with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional research settings. LifeSciBench addresses this gap by spanning seven representative scientific workflows and seven life science domains, with each constituent task paired with a human expert-written rubric. Across five frontier and domain-specialized models, GPT-Rosalind performs best, with a task-weighted mean normalized rubric score of 0.576 and a task-weighted response pass rate of 36.1% (response-level values are first averaged within each task, and the resulting task-level values are then averaged with equal weight). LifeSciBench remains unsaturated, with 171 tasks (22.8%) having no observed passing response from any evaluated model and 261 tasks (34.8%) having a best-model pass rate below 20%. LifeSciBench therefore serves as a high-resolution evaluation of practical scientific reasoning and operational decision-making in the life sciences.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 19 Aug 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 25
- Comments 0