Authors
Nael, M. A., Elokely, K.
Abstract
Background: Subtype-selectivity predictions are scored against measured selectivity and judged against an assumed noise ceiling. We asked what an 2-adrenergic benchmark rewards and which controls change its interpretation. Research design and methods: On a frozen benchmark of 586 paired 2A/2C compounds we evaluated Glide SP docking, CNN rescoring, ligand-only fingerprint models, receptor descriptors and pose contacts, with dopamine D3/D2 as comparator, applying five controls: a measured ceiling, a cluster-identity null, a nonselective reference, a same-receptor floor and a trivial-descriptor baseline. Results: Five descriptors from SMILES reached Spearman 0.645, 72% of the measured ceiling, against 0.071 for Glide SP and 0.188 for CNN rescoring; receptor properties and pose contacts reduced to size under control, while a non-size signal of 0.258 survived. Measured rather than propagated noise raised that ceiling from 0.704 to 0.897; cluster identity alone reached R2 0.499 on D3/D2 and none on 2; a nonselective reference received +1.43 to +4.79 kcal/mol where zero is expected; and a same-receptor floor reached 1.77-fold against 1.88-fold across subtypes. Conclusions: Such benchmarks reward molecular size first; a method must exceed 0.645 before its score indicates structural reasoning. The controls are inexpensive; conclusions rest on two receptor pairs, a three-pair floor and static structures.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 24 Aug 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 8
- Comments 0