Authors
Goto, Y., Yoshino, S., Kita, C., Won, M., Lee, Y.-A.
Abstract
Larceny imposes profound societal and economic burdens; however, punitive judicial measures frequently fail to deter recidivism. The neurobehavioral mechanisms driving habitual offending, whether instrumental or kleptomanic, in theft recidivists remain poorly understood. In this study, we investigated explore-exploit decision-making and underlying reinforcement learning architectures in theft recidivists with a 4-arm bandit task while prefrontal cortex (PFC) hemodynamics were continuously monitored using functional near-infrared spectroscopy (fNIRS). Model-free behavioral analyses revealed that non-kleptomanic (TR-K) but not kleptomanic (TR+K) theft recidivists accumulated significantly higher cumulative regret and made fewer optimal choices compared to control individuals without criminal records (CT). Model-based analyses through Variational Bayesian Analysis identified Q-learning with decay model as the optimal computational fit. Parameter extraction using the model demonstrated that the TR-K group exhibited a lower learning rate than CT and TR+K groups, indicating a learning deficit in updating action values following environmental feedback. fNIRS tracking of trial-by-trial latent reinforcement variables revealed that while PFC activity was modulated by these variables, group differences were characterized by static baseline hemodynamic shifts rather than rewirings of value-tracking neural circuits. These findings suggest that non-kleptomanic recurrent thefts may be associated with impaired reinforcement learning mechanisms, challenging current punitive deterrence models.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 17 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 10
- Comments 0