Authors
Aloor, J., Sit, T. P., Gauld, O. M., Warren, J., Mower, M., Lee, D., Duan, C. A.
Abstract
Real-life decisions are rarely made in isolation; outcomes depend not only on an individual's choices but also on the actions of others. In competitive settings, optimal decision-making often requires unpredictability: for example, a penalty taker in football randomises their kicks to avoid being predicted by the goalkeeper. Such stochastic strategies prevent opponents from exploiting predictable patterns. To probe the neural mechanisms underlying these game-theoretic behaviours, we trained head-fixed mice in a zero-sum game against a competitive computer opponent, and compared their choice patterns and dorsal cortical dynamics across learning. We fit an unsupervised hidden Markov model to characterise their strategy on each trial, revealing a shift from structured to stochastic strategies as animals learned to avoid predictable choice patterns. We validated this model using behavioural data from monkeys playing the same game against different levels of computer opponents, enabling direct cross-species comparison of decision strategies. This revealed fundamental computations shared across species as well as species-specific differences. We tracked dorsal cortical dynamics while mice transitioned from naive to expert players, and found that reward history information was reduced during stochastic choices compared to structured, reward-guided choices. Furthermore, the influence of immediate reward information on guiding upcoming choices was significantly diminished during stochastic choices. Together, these findings reveal a putative neural mechanism for behavioural stochasticity, where reward information is decoupled from future choice to maximise behavioural unpredictability and long-term gains in a competitive multi-agent game.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 05 Aug 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 19
- Comments 0