Authors
Carli, F., Rusina, P., Ai, L., Kuechenhoff, L., To, P. K. P., McDonagh, E. M., Lobentanzer, S., Petroni, F., Dugourd, A., Ochoa, D., Saez-Rodriguez, J.
Abstract
Language models and agents are increasingly used in biomedicine, but current benchmarks reward correct answers even when the underlying reasoning is flawed. Here we introduce Karenina, an open-source framework that turns expert knowledge into multi-dimensional evaluations of questions, conversations and autonomous agents. Illustrated in Question-Answer pairs, multi-turn conversations and autonomous data-analysis, these dimensions together moves evaluation beyond scoring, enabling trustworthy decision-making with AI in biomedicine.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 06 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 20
- Comments 0