Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Accuracy of ChatGPT, Gemini, Claude and DeepSeek in Carbohydrate Counting.

Created on 13 Apr 2026

Authors

Luca Zagaroli, Nicholas Caione, Sabina Zara, Federica Guerra, Antonella Zugaro, Marco Giorgio Baroni, Maria Laura Iezzi, Maurizio Delvecchio

Published in

Diabetes, obesity & metabolism. Apr 13, 2026. Epub Apr 13, 2026.

Abstract

To evaluate the accuracy of four general-purpose artificial intelligence (AI) models-ChatGPT (OpenAI), Gemini (Google), Claude (Anthropic) and DeepSeek (DeepSeek AI)-in calculating the carbohydrate content of meals compared with clinicians-calculated reference values.
The primary endpoint was equivalence between clinicians and AI-generated calculations within an error margin of ±5%. One-hundred twenty-four meals were analysed, equally distributed among breakfast, lunch, dinner and snacks. Carbohydrate contents were jointly determined by two paediatric diabetologists and one clinical nutritionist using the USDA FoodData Central and CREA Italian Food Composition Tables. Each AI model received identical, standardized prompts in English describing the meals. Statistical analyses included the Two One-Sided Tests procedure, the Bland-Altman plots, the Wilcoxon signed-rank and the Spearman correlations.
The clinicians' median carbohydrate content was 30.32 g. Model medians were 30.75 g (ChatGPT), 30.40 g (Gemini), 29.75 g (DeepSeek) and 29.25 g (Claude). ChatGPT showed the smallest bias, the narrowest limits of agreement, and the highest correlation with clinicians' calculation. Only ChatGPT met the predefined ±5% equivalence criterion, whereas Gemini and DeepSeek achieved equivalence within a ±10% margin. Claude displayed the largest negative bias and the widest dispersion.
ChatGPT most accurately approximated clinicians' carbohydrate calculation among the tested AI models and fulfilled strict clinical equivalence criteria. Although the other models tended to underestimate carbohydrate content, their mean deviations remained within clinically acceptable limits. These findings suggest that AI tools, particularly ChatGPT, may serve as useful adjuncts for carbohydrate counting for people with type 1 diabetes, supporting self-management.

PMID:
41969183
Bibliographic data and abstract were imported from PubMed on 13 Apr 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 218
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement