---
title: The Exploratory Multi-Asset Mean-Variance Portfolio Selection using Reinforcement Learning
url: https://www.ml-quant.com/papers/arxiv/2505.07537/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2505.07537
source_url: http://arxiv.org/abs/2505.07537v1
featured: 2025-05-14
citations: 0
topic: Portfolio & Allocation
---


# The Exploratory Multi-Asset Mean-Variance Portfolio Selection using Reinforcement Learning

The use of the soft actor-critic (SAC) algorithm in multi-asset portfolio selection is explored, showing superior performance in both simulated and real markets.

- Source: http://arxiv.org/abs/2505.07537v1
- Identifier: arXiv:2505.07537
- Released: 2025-05-12
- First featured: Quant Letter No. 97 (2025-05-14): https://www.ml-quant.com/issues/2025-05-14/
- Citations (Semantic Scholar): 0
- Published in: not yet
- Topic: Portfolio & Allocation

## Related

- [Adaptive and Regime-Aware RL for Portfolio Optimization](https://www.ml-quant.com/papers/arxiv/2509.14385/): The study presents a new reinforcement learning framework for portfolio optimization, which performs well under financial stress and supports dynamic asset allocation.
- [Diffusion-Augmented Reinforcement Learning for Robust Portfolio Optimization under Stress Scenarios](https://www.ml-quant.com/papers/arxiv/2510.07099/): The research introduces a framework called DARL that combines DDPMs with DRL for portfolio management, improving its ability to withstand crises.
- [Discrete-Time Mean-Variance Strategy Based on Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2312.15385/): The article discusses a new reinforcement learning-based model for analyzing real-world data, which is more applicable than the continuous-time model.
- [Evaluation of Deep Reinforcement Learning Algorithms for Portfolio Optimisation](https://www.ml-quant.com/papers/arxiv/2307.07694/): The study finds that PPO and A2C deep reinforcement learning algorithms are more effective for portfolio optimization due to their noise handling and policy derivation capabilities, despite their high sample complexity.
- [Combining Reinforcement Learning and Barrier Functions for Adaptive Risk Management in Portfolio Optimization](https://www.ml-quant.com/papers/arxiv/2306.07013/): Reinforcement learning and barrier functions used in portfolio management framework
- [A Comparative Analysis of Portfolio Optimization Using Mean-Variance, Hierarchical Risk Parity, and Reinforcement Learning Approaches on the Indian Stock Market](https://www.ml-quant.com/papers/arxiv/2305.17523/): Comparing portfolio optimization approaches on stock data
