Authors
Bajiya, N., Mehta, N. K., Raghava, G. P. S.
Abstract
Over the past decade, aptamers have emerged as promising therapeutics, with cancer therapeutics as a major research area. Existing computational methods, however, either predict aptamer-target pairs or design aptamers against a specific target. In this study, we present AntiCapt, a computational method for identifying single-stranded DNA (ssDNA) anticancer aptamers (ACAs) using sequence descriptors and fine-tuned nucleotide language models. The main dataset comprises 1,021 experimentally validated ACAs and an equal number of non-aptamer sequences. Comparative analysis revealed distinct patterns in the nucleotide, dinucleotide, and trinucleotide compositions associated with ACAs. We developed machine learning (ML) models using composition, autocorrelation, binary profiles and structural features. Among these, the best-performing composition-based model achieved an AUC of 0.93 on an independent dataset, while chemical descriptor-and structural feature-based models achieved a maximum AUC of 0.89. ML models using pretrained and fine-tuned NLM embeddings were also developed, with fine-tuned HyenaDNA embeddings achieving the highest performance, with an independent AUC of 0.94. Furthermore, we developed a model for discriminating anticancer aptamers from general aptamers. The best-performing models were integrated into AntiCapt, a web server and a standalone tool for predicting, designing, and genome-scale scanning anticancer aptamers (https://webs.iiitd.edu.in/raghava/anticapt/).
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 08 Oct 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 6
- Comments 0