---
title: Gray-box Adversarial Attack of Deep Reinforcement Learning-based Trading Agents*
url: https://www.ml-quant.com/papers/arxiv/2309.14615/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2309.14615
source_url: https://arxiv.org/abs/2309.14615
featured: 2023-09-28
citations: 5
topic: Trading, Microstructure & Execution
---


# Gray-box Adversarial Attack of Deep Reinforcement Learning-based Trading Agents*

A study has shown that a gray-box method can significantly reduce the profits of a Deep Reinforcement Learning-based trading agent, highlighting the need for stronger automated trading systems.

- Source: https://arxiv.org/abs/2309.14615
- Identifier: arXiv:2309.14615
- Released: 2023-09-26
- First featured: Quant Letter No. 17 (2023-09-28): https://www.ml-quant.com/issues/2023-09-28/
- Citations (Semantic Scholar): 5
- Published in: 2023 International Conference on Machine Learning and Applications (ICMLA)
- Topic: Trading, Microstructure & Execution

## Related

- [Deep Reinforcement Learning for Active High Frequency Trading](https://www.ml-quant.com/papers/arxiv/2101.07107/): A new Deep Reinforcement Learning framework has been developed for high frequency stock trading, showing potential for profitable long-term strategies.
- [Domain-adapted Learning and Imitation: DRL for Power Arbitrage](https://www.ml-quant.com/papers/arxiv/2301.08360/): Leveraging Expertise: A dual-agent reinforcement learning approach can optimize European power arbitrage trading, improving training convergence and performance, and tripling profit and loss.
- [JAX-LOB: A GPU-Accelerated limit order book simulator to unlock large scale reinforcement learning for trading](https://www.ml-quant.com/papers/arxiv/2308.13289/): JAX-LOB: The paper introduces JAX-LOB, the first GPU-powered limit order book simulator capable of processing multiple books simultaneously, designed for efficient large-scale simulations of LOB dynamics for research, calibration, and reinforcement learning training.
- [The QLBS Model Within the Presence of Feedback Loops Through the Impacts of a Large Trader](https://www.ml-quant.com/papers/arxiv/2311.06790/): The QLBS model is expanded to include a large trader's impact on exchange rates and contingent claim prices, using reinforcement learning to find an optimal hedging strategy, reducing transaction costs and aligning with the trader's fair price.
- [An adaptive dual-level reinforcement learning approach for optimal trade execution](https://www.ml-quant.com/papers/arxiv/2307.10649/): The research introduces a reinforcement learning strategy that accurately tracks the daily volume-weighted average price of stocks, using a dual-level architecture for better results.
- [Deep Reinforcement Learning: Policy Gradients for US Equities Trading](https://www.ml-quant.com/papers/ssrn/4645453/): The study shows that Deep Reinforcement Learning can effectively interpret synthetic alpha signals in financial trading, outperforming the market benchmark.
