Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Fine-tuning a compact large language model on real-world cases yields diagnostic performance comparable to flagship models in rheumatology.

Created on 21 Sep 2026

Authors

Guanhong Yao, Ut-Kei Wong, Haihong Yao, Wuji Zhang, Yingxi Zhu, Guanghao Shen, Zhanguo Li, Hui Gao

Published in

International journal of medical informatics. Volume 222. Pages 106729. Sep 16, 2026. Epub Sep 16, 2026.

Abstract

Large language models (LLMs) show strong potential in medical applications. However, most leading models are closed-source and accessible only via cloud-based APIs (application programming interfaces), raising privacy and compliance concerns when handling sensitive clinical data, while high-performing open-source models are often prohibitively large and costly to deploy and smaller models show limited diagnostic capability. Although fine-tuning with domain-specific data may improve diagnostic performance, systematic evaluations on real-world clinical cases remain lacking, leaving real-world diagnostic performance unclear.
We fine-tuned a general-purpose LLM (Qwen3) using 19,682 real-world rheumatology cases, examining the effects of training sample size and model scale on performance. The best-performing setting (Qwen3-8B with full training dataset) was further fine-tuned and compared with other models (GPT-5.6 Sol, GPT-5.6 Luna DeepSeek-V4-Pro, DeepSeek-V4-Flash, Baichuan-M2-32B) on diagnostic performance (hit1), with a separate assessment of local deployment requirements and serving capacity for Qwen3-8B and the larger open-source models. Diagnostic performance was evaluated on a held-out internal test set of real-world cases, with hit1 defined as the percentage of cases where the model's top prediction matched the physician-recorded primary diagnosis, adjudicated by GPT-Judge. Blinded physician adjudication was performed in randomly sampled test cases to assess model performance, while manual review characterized the patterns underlying the improvement from fine-tuning.
Fine-tuning increased Qwen3-8B's hit1 from 68.41% to 79.84%, the highest among the evaluated models on the internal test set. In the 1000-case physician-adjudicated comparison, its Adjudicated Hit1 was 91.6% versus 93.5% for GPT-5.6 Sol. Validation loss consistently decreased with larger training sets and model scales. The fine-tuned model substantially lowered the upfront hardware barrier to local deployment compared with the assessed configurations for larger open-source models.
Fine-tuning Qwen3 on real-world rheumatology cases substantially improved diagnostic performance, yielding a compact model that lowers the upfront hardware barrier to local deployment. This approach offers a practical route to locally deployed diagnostic support, allowing clinical data and model weights to remain within the institution.

PMID:
42763976
Bibliographic data and abstract were imported from PubMed on 21 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 6
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement