Authors
Caitlin Whitter, Aurora Evelyn Clark, Alex Pothen, Rajiv Khanna
Published in
Journal of chemical information and modeling. Volume 66. Issue 17. Pages 10593-10608. Sep 14, 2026.
Abstract
We describe MolSelector, a machine learning framework for selecting representative subsets of molecular science data sets for efficient and accurate neural network training. As part of MolSelector, we introduce the interpretable Atypicality Score algorithm, which identifies molecules that are typical versus atypical of their data set and selects a representative sample of the data set on that basis. These subsets can be used for subsequent neural network training, leading to fast model training with reduced memory requirements while still maintaining low test set error across several molecular property prediction tasks. Our experiments demonstrated that the Atypicality Score subsets resulted in errors close to the errors obtained when the entire training set is used, while achieving a 3× or greater training time speedup compared to this baseline. Additionally, we analyzed the Atypicality Score algorithm's typical and atypical molecule assignments to gain insight into the molecular characteristics the algorithm determined most beneficial for subset selection.
PMID:
42734560
Bibliographic data and abstract were imported from PubMed on 14 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 6
- Comments 0