---
title: Large Language Models in Token Space
url: https://www.ml-quant.com/papers/ssrn/5140817/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: SSRN 5140817
source_url: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5140817
featured: 2025-03-05
citations: unknown
topic: LLMs & Text
---


# Large Language Models in Token Space

The paper presents a framework that treats large language models as reinforcement learning agents in token space, offering theoretical insights for more effective language models.

- Source: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5140817
- Identifier: SSRN 5140817
- Released: 2025-02-17
- First featured: Quant Letter No. 87 (2025-03-05): https://www.ml-quant.com/issues/2025-03-05/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: LLMs & Text

## Related

- [Training Language Models to Self-Correct via Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2409.12917/): SCoRe, a new online reinforcement learning approach, enhances the self-correction ability of large language models, showing top performance with Gemini 1.0 Pro and 1.5 Flash models.
- [Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models](https://www.ml-quant.com/papers/arxiv/2501.09686/): The article discusses advancements in Large Language Models (LLMs) reasoning, emphasizing the use of reinforcement learning and thought simulation for complex reasoning, and the potential of scaling during training and testing.
- [Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF](https://www.ml-quant.com/papers/arxiv/2405.19320/): The study presents a unified approach to reinforcement learning from human feedback for large language models, offering theoretical guarantees and practical effectiveness.
- [Can large language models explore in-context?](https://www.ml-quant.com/papers/arxiv/2403.15371/): Large Language Models such as GPT-3.5, GPT-4, and Llama2 struggle to explore in reinforcement learning environments without significant interventions, indicating the need for algorithmic interventions in complex decision-making scenarios.
- [Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation](https://www.ml-quant.com/papers/arxiv/2502.17011/): The paper introduces a new method for predicting bond yields using Causal Generative Adversarial Networks and reinforcement learning, which improves forecasting performance by generating synthetic bond yield data.
- [FLAG-Trader: Fusion LLM-Agent with Gradient-based Reinforcement Learning for Financial Trading](https://www.ml-quant.com/papers/arxiv/2502.11433/): Fusion LLM-Agent: FLAG-Trader, a new architecture combining linguistic processing and reinforcement learning, has been proposed to enhance decision-making in interactive financial markets.
