---
title: An adaptive dual-level reinforcement learning approach for optimal trade execution
url: https://www.ml-quant.com/papers/arxiv/2307.10649/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2307.10649
source_url: https://arxiv.org/abs/2307.10649
featured: 2023-07-26
citations: 6
topic: Trading, Microstructure & Execution
---


# An adaptive dual-level reinforcement learning approach for optimal trade execution

The research introduces a reinforcement learning strategy that accurately tracks the daily volume-weighted average price of stocks, using a dual-level architecture for better results.

- Source: https://arxiv.org/abs/2307.10649
- Identifier: arXiv:2307.10649
- Released: 2023-07-20
- First featured: Quant Letter No. 9 (2023-07-26): https://www.ml-quant.com/issues/2023-07-26/
- Citations (Semantic Scholar): 6
- Published in: Expert Syst. Appl.
- Topic: Trading, Microstructure & Execution

## Related

- [Deep Reinforcement Learning for Active High Frequency Trading](https://www.ml-quant.com/papers/arxiv/2101.07107/): A new Deep Reinforcement Learning framework has been developed for high frequency stock trading, showing potential for profitable long-term strategies.
- [JAX-LOB: A GPU-Accelerated limit order book simulator to unlock large scale reinforcement learning for trading](https://www.ml-quant.com/papers/arxiv/2308.13289/): JAX-LOB: The paper introduces JAX-LOB, the first GPU-powered limit order book simulator capable of processing multiple books simultaneously, designed for efficient large-scale simulations of LOB dynamics for research, calibration, and reinforcement learning training.
- [Domain-adapted Learning and Imitation: DRL for Power Arbitrage](https://www.ml-quant.com/papers/arxiv/2301.08360/): Leveraging Expertise: A dual-agent reinforcement learning approach can optimize European power arbitrage trading, improving training convergence and performance, and tripling profit and loss.
- [Deep reinforcement trading with predictable returns](https://www.ml-quant.com/papers/arxiv/2104.14683/): The performance of model-free deep reinforcement learning traders in a market environment with different mean-reverting factors is investigated.
- [Gray-box Adversarial Attack of Deep Reinforcement Learning-based Trading Agents*](https://www.ml-quant.com/papers/arxiv/2309.14615/): A study has shown that a gray-box method can significantly reduce the profits of a Deep Reinforcement Learning-based trading agent, highlighting the need for stronger automated trading systems.
- [The QLBS Model Within the Presence of Feedback Loops Through the Impacts of a Large Trader](https://www.ml-quant.com/papers/arxiv/2311.06790/): The QLBS model is expanded to include a large trader's impact on exchange rates and contingent claim prices, using reinforcement learning to find an optimal hedging strategy, reducing transaction costs and aligning with the trader's fair price.
