---
title: An introduction to reinforcement learning for neuroscience
url: https://www.ml-quant.com/papers/arxiv/2311.07315/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2311.07315
source_url: https://arxiv.org/abs/2311.07315
featured: 2024-08-07
citations: 10
topic: ML & AI Methods
---


# An introduction to reinforcement learning for neuroscience

The review explores the development of reinforcement learning in neuroscience, drawing comparisons between machine learning techniques and neuroscience, and introduces modern deep reinforcement learning methods.

- Source: https://arxiv.org/abs/2311.07315
- Identifier: arXiv:2311.07315
- Released: 2023-11-13
- First featured: Quant Letter No. 60 (2024-08-07): https://www.ml-quant.com/issues/2024-08-07/
- Citations (Semantic Scholar): 10
- Published in: Neurons, Behavior, Data analysis, and Theory
- Topic: ML & AI Methods

## Related

- [Mastering Diverse Domains through World Models](https://www.ml-quant.com/papers/arxiv/2301.04104/): Algorithm Mastery: DreamerV3, a universal algorithm, excels in over 150 varied tasks, including diamond collection in Minecraft without human input, expanding the scope of reinforcement learning.
- [SimPO: Simple Preference Optimization with a Reference-Free Reward](https://www.ml-quant.com/papers/arxiv/2405.14734/): Simple Preference Optimization: SimPO improves reinforcement learning from human feedback by using the average log probability of a sequence as the implicit reward, enhancing training stability and computational efficiency.
- [DPO Meets PPO: Reinforced Token Optimization for RLHF](https://www.ml-quant.com/papers/arxiv/2404.18922/): A new framework is introduced that models Reinforcement Learning from Human Feedback as a Markov decision process, using an algorithm that learns from preference data.
- [Settling the Sample Complexity of Model-Based Offline Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2204.05275/): A paper reveals that a model-based approach can achieve optimal sample complexity without burn-in cost in offline reinforcement learning for tabular Markov decision processes, providing an efficient solution for sample-starved applications.
- [Critique-out-Loud Reward Models](https://www.ml-quant.com/papers/arxiv/2408.11791/): The article presents CLoud reward models that use human feedback to improve reinforcement learning, enhancing accuracy and win rate in ArenaHard.
- [Empirical Design in Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2304.01315/): The article highlights the importance of proper statistical evidence and avoiding common errors in empirical design for effective reinforcement learning experiments.
