Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Assessment of the efficacy of ChatGPT responses to bacterial species-specific questions in microbiology.

Created on 08 Aug 2026

Authors

Withanage Dona Manushi Dinasha, Nissanka Mudiyanselage Tanuri Ayanga Nissanka, Chamudhi Prabashi Wickramasinghe, Warnakulasuriya Palakuttige Pasindu Damsara Fernando, Vindya Perera, Hiripitiyage Gayan Danushka Gunatilake

Published in

Access microbiology. Volume 8. Issue 8. Epub Aug 04, 2026.

Abstract

Background. ChatGPT, an OpenAI chatbot, serves as a valuable tool in the present for learning and education. It also offers information on microbiology, as its popularity grows among students. However, assessing the accuracy of ChatGPT's responses is essential due to the potential for 'hallucinations' in large language models (LLMs). Objectives. This study focused on evaluating the accuracy of ChatGPT's responses to general questions on bacterial species and assessing whether the responses included key microbiological terms typically expected in academic or examination settings. Methodology. Questions were designed to reflect interactions at three language proficiency levels, including low, moderate and high. A clinical microbiologist finalized a list of 15 bacterial species, each with 18 specific questions of both local and international relevance. These questions were then prompted to ChatGPT 3.5 and 4.1 mini models, simulating real user interactions. Responses were evaluated using a microbiology reference guide and categorized as accurate, mixed/incomplete or inaccurate. Results. Results revealed average scores of 1.5%, 58.1% and 40.4% and 0.5%, 43.2% and 56.3% for inaccurate, mixed/incomplete and accurate answers for 3.5 and 4.1 mini models, respectively. While high proficiency demonstrated a higher percentage of accurate responses, all other results were either mixed/incomplete or inaccurate. Conclusion. Findings suggest that precise questions yielded more accurate responses, while imprecise questions often led to partially correct responses. Notably, ChatGPT 4.1 mini gave clearer and more reliable answers than ChatGPT 3.5. The study emphasizes the influence of question formulation on response accuracy, recommending further research to explore more advanced LLMs like ChatGPT-4o and ChatGPT-o3 models.

PMID:
42569113
Bibliographic data and abstract were imported from PubMed on 08 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 8
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement