Authors
Harms, C. M., Klein-Seetharaman, J.
Abstract
Identifying protein targets from phenotype-first or mechanism-uncertain compounds remains difficult because proteome-scale docking can generate thousands of structurally plausible interactions per compound. We developed a workflow that converts ranked proteome-scale docking profiles into stable Gene Ontology (GO) Biological Process enrichment signatures and evaluates whether those signatures can reduce the candidate target space while preferentially retaining known drug-target relationships. Docking targets were retained at the PDB-chain level, mapped to unique human gene identities, and analyzed with PANTHER overrepresentation against the screened structural gene universe. GO enrichment was evaluated from the top 25 through 505 ranked proteins in increments of 10, and a stable compound-level GO profile was selected using a Jaccard stability threshold of 0.80 across three consecutive transitions. Benchmarking used Yamanishi drug-target interactions, with 682 mapped compounds assigned to a prespecified development/held-out split (552/130) and evaluated at canonical, mechanistic, and fine-mechanistic biological resolutions. In the full corrected benchmark, 642 compounds with usable canonical GO profiles showed greater within-class than between-class similarity (0.1019 versus 0.0931; delta = 0.0088; 100,000-permutation p = 0.00222), and canonical class explained 1.34% of multivariate GO-profile variation by PERMANOVA (p < 1e-4). Fine-mechanistic labels showed stronger organization in the full dataset (delta = 0.0245; PERMANOVA R-squared = 0.0976; both p < 1e-4). Held-out validation was more modest and metric-dependent: mechanistic labels were significant by PERMANOVA (R-squared = 0.0626, p = 0.0437), whereas the frozen fine-mechanistic analysis showed greater within-class similarity (0.1350 versus 0.1110; p = 0.038) and significant nearest-neighbor recovery (p = 0.0495), but not significant PERMANOVA (p = 0.119). The principal held-out search-space experiment evaluated 107 compounds, 753,492 candidate protein rows, and 424 represented gold-standard targets. A direct GO gate retained 1.26% of candidates while retaining 11.32% of known targets (8.96-fold enrichment); ontology-propagated GO associations retained 5.10% of candidates and 21.46% of known targets (4.21-fold enrichment). At matched candidate-space sizes, GO-associated prioritization retained 27.59% versus 22.41% of known targets at approximately 5% of candidates, 39.39% versus 36.08% at 10%, and 58.73% versus 56.13% at 20%. By contrast, additive protein-level GO reranking was heterogeneous: among 605 evaluable compounds, 21.7% improved their best known-target rank, but mean reciprocal rank decreased from 0.0276 to 0.0189. These results support GO enrichment as an intermediate biological search-space reduction and prioritization layer rather than a universal direct target-scoring function. Keywords: proteome-scale docking; Gene Ontology; target discovery; targetome; PANTHER; target fishing; biological filtering; search-space reduction; Yamanishi benchmark
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 23 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 1
- Comments 0