Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Mitigating Goodhart's law in epitope-conditioned TCR generation using plug-and-play reward designs.

Created on 24 Aug 2026

Authors

Pengfei Zhang, Xiaoyi He, Fredo Guan, Hao Mei, Gloria Grama, Seojin Bang, Heewook Lee

Published in

Bioinformatics (Oxford, England). Volume 42. Issue Supplement_2. Aug 01, 2026.

Abstract

Epitope-conditioned T cell receptor (TCR) generation extends protein language modeling to the design of therapeutically relevant receptors. Reinforcement learning (RL) post-training with surrogate binding predictors can improve generation controllability, but it is vulnerable to Goodhart's Law: optimizing an imperfect surrogate reward can lead to reward inflation, distributional drift, and biologically implausible or nonspecific sequences. We investigate whether reward hacking can be mitigated through improved reward formulations without modifying the generator architecture or training pipeline and without requiring additional training data.
We introduce a plug-and-play reward-design framework for RL-based TCR generation that combines heuristic biological priors, model ensembling, and binding-specificity objectives based on max-margin and contrastive formulations. These components suppress degenerate sequences, reduce model-specific biases, and discourage cross-epitope binding. During RL fine-tuning, the proposed rewards stabilize optimization, limit surrogate-reward inflation, preserve canonical CDR3β sequence patterns and repertoire diversity, and maintain closer alignment with experimentally validated TCR-binder distributions. In evaluations on unseen epitopes, the specificity-aware reward formulations provide the strongest overall performance, improving diversity and ground-truth distributional alignment while retaining biological authenticity and predicted target binding. These findings demonstrate that Goodhart-resistant reward design improves the reliability, controllability, and generalization of epitope-conditioned TCR generation.
Code and models are available in a public repository (https://github.com/Lee-CBG/TCRRobustRewardDesign).

PMID:
42635193
Bibliographic data and abstract were imported from PubMed on 24 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 18
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement