Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Sino-US-DrugQA: A benchmark for evaluating large language models in cross-jurisdictional pharmaceutical regulation.

Created on 16 Sep 2026

Authors

Xuejing Fu, Zhen Chen, Wentao Lu

Published in

PloS one. Volume 21. Issue 9. Pages e0343858. Epub Sep 15, 2026.

Abstract

Cross-jurisdictional pharmaceutical compliance requires comparison of regulatory requirements across administrative systems such as the US Food and Drug Administration and China's National Medical Products Administration. Although large language models (LLMs) are increasingly explored for healthcare and regulatory applications, their performance in cross-jurisdictional pharmaceutical regulation has not been systematically evaluated using a dedicated benchmark. We introduce Sino-US-DrugQA, a bilingual multiple-choice benchmark covering Monolingual, Comparative, and Parallel regulatory question-answer tasks. The 11,871 candidate items underwent deterministic structural validation and full-dataset semantic quality screening, followed by risk-stratified independent review of 1,432 items by two regulatory experts. The final release comprised 11,444 items, including 10,122 classified as pass and 1,322 as borderline. Among 500 items sampled from the semantic screen-negative population, 18 were subsequently classified as material errors, corresponding to an observed residual material-error proportion of 3.60% (Wilson 95% confidence interval, 2.29%-5.62%). Four representative LLMs-gpt-5.6-terra, gemini-3.6-flash, deepseek-v4-flash, and qwen-3.5-max-were evaluated under a standardized zero-shot protocol. Overall accuracy ranged from 83.21% to 86.43%. Comparative accuracy was consistently lower than Monolingual accuracy, with absolute differences of 4.42-8.98 percentage points across models. The two highest-scoring models, GPT and Gemini, did not differ significantly after adjustment for multiple comparisons. These results indicate that explicit comparison across non-equivalent regulatory systems remains more challenging than single-jurisdiction question answering. Sino-US-DrugQA provides a validated and reproducible resource for evaluating bilingual regulatory reasoning. The findings support further investigation of expert-supervised decision-support workflows rather than autonomous regulatory interpretation. The stable dataset release and evaluation resources are available at https://github.com/DodgeLU/Sino-US-DrugQA.

PMID:
42743305
Bibliographic data and abstract were imported from PubMed on 16 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 12
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement