Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Characterization and Validation of EHR Computable Phenotypes for Long COVID Using Patient-Reported Symptoms: Insights from the Nationwide RECOVER Program.

Created on 24 Jul 2026

Authors

Victor M Castro, Vivian Gainer, Nich Wattanasin, Andrew Cagan, Ana Holzbach, James Chan, Leora Horwitz, Rachel Kenney, Ivan Diaz, Hannah Mandel, Shannon Wuller, Mady Hornig, Lisa O'Brien, Andrew Wylam, James Doster, Richard A Moffitt, Emily Pfaff, Mark G Weiner, Sajjad Abedian, Michael Koropsak, Sairam Parthasarathy, Hanieh Razzaghi, Justin Manjourides, Elizabeth W Karlson, Shawn N Murphy

Published in

Journal of the American Medical Informatics Association : JAMIA. Jul 24, 2026. Epub Jul 24, 2026.

Abstract

Long COVID (LC) remains poorly understood, and there is a critical need for advanced computational tools to better identify and characterize patients. In this study, we use summarized symptom reports by RECOVER-Adult cohort participants linked to EHR data to characterize patients and train a computable phenotype algorithm of LC.
The study included adult participants with linked FHIR-sourced EHR data. We characterized EHR diagnoses, procedures, medications, lab tests, and vital sign features associated with LC. A computable phenotyping algorithm was trained and validated against patient-reported symptoms.
We assessed model discrimination and calibration in a held-out test set. We describe important model features and evaluate model discrimination and calibration.
The study included 1,501 RECOVER-Adult cohort participants with linked EHR data. 376 (25%) met criteria for highly symptomatic LC based on the RECOVER Long COVID Research Index (LCRI). EHR features associated with LC included clinician diagnosis of shortness of breath, malaise and fatigue, and cardiac dysrhythmias; documented treatment with albuterol, gabapentin, or duloxetine; or elevated heart rate. The algorithm identifying patients with highly symptomatic LC had an AUROC of 0.80 (95% confidence interval (CI) 0.74-0.85), and AUPRC of 0.58 (95% CI, 0.47-0.69).
These findings demonstrate that, using EHR data, a machine-learning model can accurately select patients with sets of self-reported LC symptoms. The model could help identify patients within a health system with the highest probability of the condition and facilitate screening, recruitment for clinical trials, and etiologic studies.

PMID:
42496643
Bibliographic data and abstract were imported from PubMed on 24 Jul 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 12
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement