Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

GroundAnnot: a closed-vocabulary contract for grounding LLM gene-set annotation in live enrichment backends

Created on 03 Oct 2026

Authors

Malima, M. B.

Abstract

Motivation: LLM agents increasingly draft functional interpretations of gene lists, but can cite Gene Ontology (GO) terms that no current enrichment backend returned for that list, and can pair real GO accessions with fabricated labels. Results: We present GroundAnnot, a client for PANTHER, Enrichr, and g:Profiler that returns a closed vocabulary of GO term IDs and their backend labels, and enforces two contracts: Contract A (no ID or label outside the backend payload) and Contract B (no enrichment claim outside the backend's significant set). Across three open-weight models (Qwen2.5-7B local; Qwen3.8-27B and GPT-OSS-120B via Groq) on six curated disease gene lists, mean valid-enriched rates were 15.3%, 20.1%, and 35.1%. All unsupported IDs were real GO terms classified by QuickGO as wrong-biology, obsolete, or wrong-branch (zero fabricated accessions). A distinct failure mode emerged: for the Alzheimer's gene list, the local 7B model paired every one of its ten stable picks with a label that does not match the current GO term for that accession (10/10 mismatches), producing a coherent synaptic-signalling narrative that the IDs do not support. Mean overlap with the backends' FDR top-10 union was 0.33-0.67/10. On 50 MSigDB Hallmark gene sets, the strongest model named the defining pathway in 21/44 (47.7%) directly named cases. When an enriched shortlist was placed in the prompt, all three models complied fully (60/60 for each model; 180/180 overall), showing that a bounded vocabulary removes these output classes when the downstream agent is required to use it.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 03 Oct 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 8
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement