ML-QuantSubscribe

arXivLLMs & Text

FinRankGRPO: Optimizing LLMs for Listwise Financial Asset Ranking via Group Relative Policy Optimization

Develops a two-stage framework that fine-tunes language models for listwise asset ranking using Spearman rank correlation rewards, achieving a Sharpe ratio of 0.636 on asset allocation.

Featured in No. 132 on 25 Sep 2026 · 3 days after release · 0 citations today

The two-stage construction framework of our FinRankGRPO, Stage 1 is SFT in high quality distill CoT datasets, Stage 2
Figure 2: The two-stage construction framework of our FinRankGRPO, Stage 1 is SFT in high quality distill CoT datasets, Stage 2 is trained by our FinRankGRPO, which is a Spearman-related ranking reward.
Released
22 Sep 2026
First featured
No. 132 · 25 Sep 2026
Citations (Semantic Scholar)
0
Influential citations
0
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
2 of 5
Identifier
arXiv:2609.24175
Authors
Ningyuan Deng et al.

Abstract

From arXiv (CC0).

While Large Language Models (LLMs) excel at understanding unstructured financial contexts, their direct use in portfolio optimization is limited by a mismatch between next-token prediction and the listwise ranking objectives required for asset allocation. They also struggle with precise numerical forecasting, leading to instability and arithmetic hallucinations. To bridge this gap, we propose FinRankGRPO, a framework that shifts LLM based portfolio construction from direct numerical prediction to listwise ranking of financial assets. We introduce a two-stage training process, supervised finetuning on Chain-of-Thought reasoning data, followed by our Financial Asset Ranking via Group Relative Policy Optimization with a Spearman rank correlation reward that aligns generated asset rankings with ground truth market orderings. The second stage uses a novel Spearman rank correlation reward to explicitly align the model's generative preferences with ground truth market orderings. Experimental results show that FinRankGRPO outperforms traditional quantitative and state-of-the-art commercial models, achieving a Sharpe ratio of 0.636 and a Spearman correlation of 0.023.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page