Authors
Rostami, Z., Choubdar Parvin, H., Iyer, E. S., Vitaro, P., Bohme, R., Heiling, M., Ebitz, R. B., Mayo, L. M., Bagot, R. C.
Abstract
Adaptive behavior requires that organisms learn which actions are rewarded and to update action selection when the environment changes. While human and non-human animals exhibit adaptive behavior, whether apparently similar behavior reflects common decision strategies remains unclear. Probabilistic reversal learning provides a cross-species assay of reward-guided choice, yet standard metrics such as accuracy or reward rate can obscure underlying strategies that generate choices. Here, we applied parallel probabilistic reversal learning tasks in mice and humans and used a generalized linear model-hidden Markov model to infer latent decision strategies from trial-by-trial behavior. Across species, choices were organized into stable behavioral states with differing reliance on choice history, reward history, and response bias. Among these latent states, we identify a conserved reward-learning strategy in mice and humans characterized by the greatest feedback sensitivity, reward efficiency, and adaptation after reversal. Simulating choice behavior using state-specific decision policies reproduced the empirical hierarchy of performance, confirming that the latent states capture meaningful behavioral strategies. Although mice and humans differ in the temporal dynamics of reward learning, both species ultimately converge on the same optimized strategy. These findings identify a conserved latent reward-learning strategy in mice and humans, defining a translational framework for studying how adaptive decision-making is shaped by task experience, stress, affective processes, and neural circuit function.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 09 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 9
- Comments 0