Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Use of machine learning to detect Escherichia coli in drinking water in Bangladesh.

Created on 21 Aug 2026

Authors

Iqramul Haq, Md Yusuf Hossain Ador, Diego Nobrega

Published in

PloS one. Volume 21. Issue 8. Pages e0343606. Epub Aug 20, 2026.

Abstract

Escherichia coli (E. coli) is a key indicator of fecal contamination in freshwater and can signal the presence of other harmful bacteria and viruses. The aim of the study is to evaluate the performance of machine learning (ML) tools to detect E. coli in drinking water in Bangladesh using surveillance data under two scenarios: an imbalanced dataset and a balanced dataset. We utilized data from the 2019 Bangladesh Multiple Indicator Cluster Survey, which included a total of 6,069 household drinking water samples. We used agglomerative hierarchical clustering with Ward's linkage to identify district-level hotspots. Extreme Gradient Boosting with SHapley Additive exPlanations values were used for feature selection, and the Synthetic Minority Over-sampling Technique (SMOTE) was used to address class imbalance in the classification task. We applied nine classical ML models in this study: Adaptive Boosting (AdaBoost), Decision Trees (DT), Gradient Boosting Algorithm (GBA), k-Nearest Neighbors (KNN), Light Gradient-Boosting Machine (LightGBM), Logistic Regression (LR), Naïve Bayes (NB), Random Forest (RF), and Support Vector Machine (SVM), along with a Deep Learning Multi-Layer Perceptron (DL-MLP) model to predict the risk of E. coli contamination (REcC) in water. Model performance was evaluated using accuracy, precision, recall, F1 score, Cohen Kappa, area under the curve (AUC), and a violin plot. E. coli contamination in drinking water was detected in 39.2% (95% CI: 37.4-41.2) of households. Bandarban district had the highest REcC. After applying SMOTE and 10-fold cross-validation with hyperparameter tuning, model performance was more consistent across algorithms. In terms of model evaluation, AdaBoost slightly outperformed the others with an accuracy of 81.6%, Cohen kappa statistic of 19.4%, precision of 82.2%, recall of 99%, F1-score of 89.8%, and an AUC of 68.6%. Ensemble model for example AdaBoost and GBA models had ability to accurately classify drinking water samples with respect to the presence of E. coli using surveillance data than others selected model in this study.

PMID:
42623399
Bibliographic data and abstract were imported from PubMed on 21 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 2
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement