Authors
Pengfei Zhang, Xiaoyi He, Fredo Guan, Hao Mei, Gloria Grama, Seojin Bang, Heewook Lee
Published in
Bioinformatics (Oxford, England). Volume 42. Issue Supplement_2. Aug 01, 2026.
Abstract
Epitope-conditioned T cell receptor (TCR) generation extends protein language modeling to the design of therapeutically relevant receptors. Reinforcement learning (RL) post-training with surrogate binding predictors can improve generation controllability, but it is vulnerable to Goodhart's Law: optimizing an imperfect surrogate reward can lead to reward inflation, distributional drift, and biologically implausible or nonspecific sequences. We investigate whether reward hacking can be mitigated through improved reward formulations without modifying the generator architecture or training pipeline and without requiring additional training data.
We introduce a plug-and-play reward-design framework for RL-based TCR generation that combines heuristic biological priors, model ensembling, and binding-specificity objectives based on max-margin and contrastive formulations. These components suppress degenerate sequences, reduce model-specific biases, and discourage cross-epitope binding. During RL fine-tuning, the proposed rewards stabilize optimization, limit surrogate-reward inflation, preserve canonical CDR3β sequence patterns and repertoire diversity, and maintain closer alignment with experimentally validated TCR-binder distributions. In evaluations on unseen epitopes, the specificity-aware reward formulations provide the strongest overall performance, improving diversity and ground-truth distributional alignment while retaining biological authenticity and predicted target binding. These findings demonstrate that Goodhart-resistant reward design improves the reliability, controllability, and generalization of epitope-conditioned TCR generation.
Code and models are available in a public repository (https://github.com/Lee-CBG/TCRRobustRewardDesign).
PMID:
42635193
Bibliographic data and abstract were imported from PubMed on 24 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 18
- Comments 0