---
title: Deviations from the Nash equilibrium in a two-player optimal execution game with reinforcement learning
url: https://www.ml-quant.com/papers/arxiv/2408.11773/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2408.11773
source_url: https://arxiv.org/abs/2408.11773
featured: 2024-08-28
citations: 1
topic: Trading, Microstructure & Execution
---


# Deviations from the Nash equilibrium in a two-player optimal execution game with reinforcement learning

Autonomous trading bots using advanced algorithms can disrupt markets by deviating from traditional predictions, often favoring optimal solutions over equilibrium.

- Source: https://arxiv.org/abs/2408.11773
- Identifier: arXiv:2408.11773
- Released: 2024-08-21
- First featured: Quant Letter No. 63 (2024-08-28): https://www.ml-quant.com/issues/2024-08-28/
- Citations (Semantic Scholar): 1
- Published in: Annals of Operations Research
- Topic: Trading, Microstructure & Execution

## Related

- [Reinforcement Learning for Optimal Execution When Liquidity Is Time-Varying](https://www.ml-quant.com/papers/arxiv/2402.12049/): Research shows Double Deep Q-learning, a Reinforcement Learning technique, can effectively learn optimal trading strategies in fluctuating liquidity conditions.
- [The QLBS Model Within the Presence of Feedback Loops Through the Impacts of a Large Trader](https://www.ml-quant.com/papers/arxiv/2311.06790/): The QLBS model is expanded to include a large trader's impact on exchange rates and contingent claim prices, using reinforcement learning to find an optimal hedging strategy, reducing transaction costs and aligning with the trader's fair price.
- [Closed‐loop Nash competition for liquidity](https://www.ml-quant.com/papers/arxiv/2112.02961/): In a multi-player stochastic differential game, excessive trading can be reduced through coordination and the existence of a closed-loop Nash equilibrium, especially when the price impact parameter is small.
- [Reinforcement Learning for Optimal Execution](https://www.ml-quant.com/papers/ssrn/4720833/): A new actor-critic reinforcement learning algorithm is introduced for optimal execution problem, featuring a recalibration step for convergence and showing linear convergence under appropriate conditions.
- [Reinforcement Learning for Arbitrage in Decentralized Exchanges](https://www.ml-quant.com/papers/ssrn/4708173/): The article presents a game-theory model to explain the tactics of various players in decentralized exchanges with automated market makers.
- [Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2609.22785/): The research extends adversarial reinforcement learning for market making to handle self-exciting order arrivals and price impact, using an LSTM module to improve robustness in complex microstructure environments.
