Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Chemically Interpretable Explanations for Molecular Property Prediction via Fragment-Level Shapley Values.

Created on 14 Sep 2026

Authors

Jannik P Roth

Published in

Journal of chemical information and modeling. Volume 66. Issue 17. Pages 10661-10673. Sep 14, 2026.

Abstract

Machine learning has emerged as a powerful approach for molecular property prediction and drug discovery. However, the black-box nature of many machine learning models limits their interpretability, trustworthiness, and adoption in interdisciplinary research settings. This is especially the case in molecular machine learning, where explanations of model predictions are used for informed decision-making and downstream tasks. Shapley values, originating from cooperative game theory, provide a principled framework for attributing model predictions to individual input features. However, existing Shapley value-based explanations for molecular machine learning often rely on sampling-based approximations or operate at the level of abstract features, which can reduce attribution stability and limit chemical interpretability and actionability. Here, we introduce a fragment-level Shapley value framework that enables the exact computation of feature contributions at the level of chemically meaningful fragments for molecular property predictions without relying on sampling or feature imputation. By decomposing molecules into fragments, the proposed approach yields actionable explanations that can be directly related to established chemical concepts. We apply the method post hoc to random forest and graph convolutional network models using common molecular representations, including extended-connectivity fingerprints and molecular graphs. The approach is evaluated across three representative property prediction tasks: aqueous solubility, mutagenicity, and antiviral potency. Fragment-level Shapley values reproduce well-established chemical trends, identify known toxicophores, and enable guided molecular optimization. In addition, the method provides insights into model learning characteristics and helps delineate the applicability domain, particularly in settings with limited and structurally biased data. Overall, this work demonstrates that adapting Shapley values to chemically meaningful fragments enables interpretable explanations for molecular machine learning models, supporting molecular optimization and model validation.

PMID:
42734503
Bibliographic data and abstract were imported from PubMed on 14 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 10
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement