Authors
Haworth, R. M., Commichaux, S., Pop, M.
Abstract
Motivation: High-throughput sequencing technologies have driven rapid growth of biological sequence databases. Public repositories must therefore rely on automated computational heuristics to screen submitted sequences for errors and low quality. For example, SILVA SSU Ref, which exploits the conserved nature of 16S and 18S rRNA sequences, applies strict algorithmic quality controls yet still accepts sequences with up to 30% of their nucleotides deviating from any previously accepted sequence. This permissiveness creates opportunities for the admission of modified sequences, such as biologically plausible sequences generated by DNA foundation models. The vulnerability of public sequence databases to becoming polluted or poisoned with modified sequences necessitates the development of methods to detect such sequences. Results: We present the first investigation, to our knowledge, of the detectability of modified sequences. We consider simple computationally-modified 16S rRNA sequences that pass the quality control inclusion criteria of the SILVA database, which we generate via random substitutions. We present classifiers that can distinguish such modified sequences from natural 16S rRNA using conserved motifs. Our best classifier achieves over 90% sensitivity and specificity on our testing set when using a 5% artificial mutation rate. One feature used in our classifiers, gapped k-mers constructed from universally conserved nucleotides, was conserved across all three domains of life despite relying on exact matches to patterns found in E. coli, advancing our understanding of conserved grammatical structure in small subunit rRNA sequences. Availability and Implementation: Our source code is available at https://github.com/rainhaworth/16S-Mutation-Classifiers.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 30 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 8
- Comments 0