Authors
Xuejing Fu, Zhen Chen, Wentao Lu
Published in
PloS one. Volume 21. Issue 9. Pages e0343858. Epub Sep 15, 2026.
Abstract
Cross-jurisdictional pharmaceutical compliance requires comparison of regulatory requirements across administrative systems such as the US Food and Drug Administration and China's National Medical Products Administration. Although large language models (LLMs) are increasingly explored for healthcare and regulatory applications, their performance in cross-jurisdictional pharmaceutical regulation has not been systematically evaluated using a dedicated benchmark. We introduce Sino-US-DrugQA, a bilingual multiple-choice benchmark covering Monolingual, Comparative, and Parallel regulatory question-answer tasks. The 11,871 candidate items underwent deterministic structural validation and full-dataset semantic quality screening, followed by risk-stratified independent review of 1,432 items by two regulatory experts. The final release comprised 11,444 items, including 10,122 classified as pass and 1,322 as borderline. Among 500 items sampled from the semantic screen-negative population, 18 were subsequently classified as material errors, corresponding to an observed residual material-error proportion of 3.60% (Wilson 95% confidence interval, 2.29%-5.62%). Four representative LLMs-gpt-5.6-terra, gemini-3.6-flash, deepseek-v4-flash, and qwen-3.5-max-were evaluated under a standardized zero-shot protocol. Overall accuracy ranged from 83.21% to 86.43%. Comparative accuracy was consistently lower than Monolingual accuracy, with absolute differences of 4.42-8.98 percentage points across models. The two highest-scoring models, GPT and Gemini, did not differ significantly after adjustment for multiple comparisons. These results indicate that explicit comparison across non-equivalent regulatory systems remains more challenging than single-jurisdiction question answering. Sino-US-DrugQA provides a validated and reproducible resource for evaluating bilingual regulatory reasoning. The findings support further investigation of expert-supervised decision-support workflows rather than autonomous regulatory interpretation. The stable dataset release and evaluation resources are available at https://github.com/DodgeLU/Sino-US-DrugQA.
PMID:
42743305
Bibliographic data and abstract were imported from PubMed on 16 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 12
- Comments 0