Authors
Seunghyun Lim, Ezekiel Ahn, Dapeng Zhang, Lyndel W Meinhardt, Sunchung Park
Published in
Bioinformatics advances. Volume 6. Issue 1. Pages vbag224. Epub Aug 06, 2026.
Abstract
Early detection of plant pathogens is essential for timely disease management, but diagnosis at low infection levels remains difficult because pathogen-derived sequences are often masked by abundant host DNA. K-mer-based machine learning offers a potentially sensitive, alignment-free feature-based approach for detecting infection and estimating pathogen abundance from sequencing data, but its performance across infection levels, biological backgrounds, and preprocessing strategies remains insufficiently characterized.
We evaluated a k-mer-based machine-learning framework using simulated Illumina short-read data from two Coffea arabica cultivars (ET39 and Typica) infected with Hemileia vastatrix or Fusarium xylarioides across a broad range of infection rates. Among ten models tested on raw and host-depleted k-mer profiles, logistic regression (LR) and linear support vector machine (LSVM) were the most sensitive. Host depletion with KrakenUniq substantially improved low-infection-rate detection, lowering the practical detection threshold to 0.05%. Models trained at lower infection rates performed better when tested across different infection rates than models trained at higher rates, although performance declined when the test infection rate fell below 0.05% because infected samples were increasingly misclassified as healthy. External validation showed limited biological transferability: low-infection-rate performance depended strongly on host background, whereas high-infection-rate performance depended more on pathogen identity. To improve weak-signal detection, we developed a two-stage framework combining elastic-net logistic regression for detection with ridge regression for infection rate prediction. This framework maintained near-perfect detection at 0.05% and above, improved detection at 0.01% under mixed-infection-rate training, and showed strong agreement between true and predicted infection rates (Spearman's ρ = 0.95). Predictive k-mers were extensively shared between LR and LSVM and showed distinct compositional differences between healthy- and infected-associated features.
Code is available at GitHub (https://github.com/TropicalBreeding/coffee-pathogen-kmer-ml).
PMID:
42633195
Bibliographic data and abstract were imported from PubMed on 23 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 6
- Comments 0