ML-QuantSubscribe

arXivML & AI Methods

$\epsilon$-Policy Gradient for Online Pricing

The article introduces an ε-policy gradient algorithm for online pricing learning tasks. This algorithm merges model-based and model-free reinforcement learning methods, and optimizes regret by balancing exploration and exploitation costs. It is expected to achieve a regret of order √T over T trials.

Featured in No. 48 on 8 May 2024 · 2 days after release · 1 citation today

Released
6 May 2024
First featured
No. 48 · 8 May 2024
Citations (Semantic Scholar)
1
Influential citations
0
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
3
Identifier
arXiv:2405.03624

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page