Authors
O'Brien, A., Gardette, A.
Abstract
1. Background. Fungal internal transcribed spacer (ITS) analyses usually infer taxonomy and novelty from similarity to homologous reference sequences. Percent identity is a strong signal in that setting, but it cannot directly compare non-homologous barcode views and does not explicitly optimize the higher-rank placement of genera absent from the reference collection. We asked whether a compact sequence encoder could learn that open-world structure and whether maximum embedding cosine similarity could improve novelty detection. 2. Design. From the UNITE dynamic release of 19 February 2025 we built ITS-core, ITS1 and ITS2 views and a genus-separated development, calibration and test design. A convolutional encoder mapped sequences to 256-dimensional L2-normalized embeddings. A controlled ladder compared genus-proxy supervision (M0), cross-view invariance (M1), hierarchical taxonomic geometry (M2), leave-one-genus family episodes (M3), and the combined objective (M4). M4 was frozen before CAL_KNOWN and TEST were opened. A secondary historical ITS2 benchmark was audited for leakage and scored only after a fixed-recipe, benchmark-safe retraining that excluded all held-out benchmark taxa. 3. Results. Cross-view supervision raised controlled ITS2 development AUROC from 0.688 to 0.717 and novel-family placement from 22.1% to 33.0%. Episodic training increased novel-family placement to 41.0% but reduced novelty discrimination at some views. The combined M4 model recovered both behaviours, reaching development AUROCs of 0.778, 0.755 and 0.751 for ITS-core, ITS1 and ITS2, with novel-family placement of 57.5%, 48.7% and 41.3%. On untouched TEST data, ITS2 reached AUROC 0.763 (95% CI 0.748-0.779) and placed unseen genera correctly at family, order and class for 52.0%, 81.1% and 92.9% of queries; development AUROCs lay above the TEST intervals at ITS-core and ITS1, indicating selection optimism of roughly 0.02-0.04 AUROC at those two views. At alpha = 0.05, conformal calibration gave a 4.7% observed false-novelty rate and detected 11.2% of novel-genus ITS2 queries. On the paired leakage-safe historical benchmark, percent identity remained stronger than M4 cosine: AUROC 0.770 versus 0.702 and 27.8% versus 8.6% novel-query detection at approximately a 5% false-novelty rate. 4. Interpretation. Structured fungal ITS embeddings add higher-rank placement and cross-view retrieval that are not available from a same-locus identity score alone. Maximum embedding cosine does not replace direct sequence identity as the novelty statistic when homologous ITS2 references are available; whether a different score read from the same embedding would narrow that gap remains untested.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 24 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 0
- Comments 0