---
title: Tail-Safe Hedging: Explainable Risk-Sensitive Reinforcement Learning with a White-Box CBF-QP Safety Layer in Arbitrage-Free Markets
url: https://www.ml-quant.com/papers/arxiv/2510.04555/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2510.04555
source_url: http://arxiv.org/abs/2510.04555v1
featured: 2025-10-09
citations: 0
topic: Derivatives & Volatility
---


# Tail-Safe Hedging: Explainable Risk-Sensitive Reinforcement Learning with a White-Box CBF-QP Safety Layer in Arbitrage-Free Markets

The research presents Tail-Safe, a derivative hedging framework that blends reinforcement learning with a safety layer designed for financial constraints, ensuring robust forward invariance of the safe set under limited model mismatch.

- Source: http://arxiv.org/abs/2510.04555v1
- Identifier: arXiv:2510.04555
- Released: 2025-10-06
- First featured: Quant Letter No. 115 (2025-10-09): https://www.ml-quant.com/issues/2025-10-09/
- Citations (Semantic Scholar): 0
- Published in: not yet
- Topic: Derivatives & Volatility

## Related

- [ARL-Based Multi-Action Market Making with Hawkes Processes and Variable Volatility](https://www.ml-quant.com/papers/doi/10-1145-3677052-3698695/): The study combines Adversarial Reinforcement Learning, Hawkes Processes, and variable volatility to enhance market-making strategies, showing improved adaptability in high-volatility conditions and better market simulations.
- [Deep Reinforcement Learning Algorithms for Option Hedging](https://www.ml-quant.com/papers/arxiv/2504.05521/): A comparison of eight Deep Reinforcement Learning algorithms for dynamic hedging found that Monte Carlo Policy Gradient and Proximal Policy Optimization performed best, with the former outperforming the Black-Scholes delta hedge baseline.
- [To Hedge or Not to Hedge: Optimal Strategies for Stochastic Trade Flow Management](https://www.ml-quant.com/papers/arxiv/2503.02496/): The paper proposes using reinforcement learning methods to manage stochastic trade flows, offering an alternative to traditional grid-based numerical PDE techniques.
- [A Risk Sensitive Contract-unified Reinforcement Learning Approach for Option Hedging](https://www.ml-quant.com/papers/arxiv/2411.09659/): The paper proposes a risk-sensitive reinforcement learning approach for dynamic hedging of options, reducing tail risk using historical market data.
- [Solving The Dynamic Volatility Fitting Problem: A Deep Reinforcement Learning Approach](https://www.ml-quant.com/papers/arxiv/2410.11789/): The article discusses the use of Deep Reinforcement Learning in solving volatility issues in equity derivatives, showing its effectiveness and adaptability in handling complex functions and online learning.
- [EX-DRL: Hedging Against Heavy Losses with EXtreme Distributional Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2408.12446/): The article introduces EXtreme DRL (EX-DRL), a new method to improve the accuracy of extreme quantile predictions in Distributional Reinforcement Learning, improving financial risk management.
