---
title: LLMs & Text
url: https://www.ml-quant.com/topics/llms-text/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
---


# LLMs & Text

Large language models, agents, sentiment and text as data in finance.

577 papers featured; 19 in the last 12 months.

Papers featured per quarter: 2023 Q2 10, 2023 Q3 39, 2023 Q4 46, 2024 Q1 47, 2024 Q2 75, 2024 Q3 107, 2024 Q4 90, 2025 Q1 68, 2025 Q2 55, 2025 Q3 21, 2025 Q4 13, 2026 Q1 0, 2026 Q2 1, 2026 Q3 5

## Most cited

- [DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models](https://www.ml-quant.com/papers/arxiv/2402.03300/): 9003 citations. Advancing Math Reasoning in Language Models: DeepSeekMath7B is a new language model that uses web data and Group Relative Policy Optimization for advanced mathematical reasoning, scoring high on the MATH benchmark.
- [Mistral 7B](https://www.ml-quant.com/papers/arxiv/2310.06825/): 3920 citations. Superior Language Model: Mistral 7B v0.1 is a language model with 7 billion parameters that excels in reasoning, mathematics, and code generation, and has a version specifically designed to follow instructions.
- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): 2189 citations. The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration](https://www.ml-quant.com/papers/arxiv/2306.00978/): 1800 citations. The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): 1775 citations. The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [s1: Simple test-time scaling](https://www.ml-quant.com/papers/arxiv/2501.19393/): 1462 citations. The research presents a method called budget forcing, which uses a small dataset to achieve test-time scaling and improved reasoning performance in language modeling, particularly in competition math questions.
- [DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model](https://www.ml-quant.com/papers/arxiv/2405.04434/): 1459 citations. MoE Language Model: DeepSeek-V2, a language model with 236B parameters, offers enhanced performance and cost efficiency compared to its predecessor, ranking high among open-source models.
- [MemGPT: Towards LLMs as Operating Systems](https://www.ml-quant.com/papers/arxiv/2310.08560/): 1373 citations. Extended Context in LLMs: MemGPT is a system that manages different memory levels, providing extended context within large language models' limited context windows, enhancing document analysis and multi-session chat performance.
- [SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models](https://www.ml-quant.com/papers/arxiv/2303.08896/): 1236 citations. Hallucination Detection for LLMs: The paper presents SelfCheckGPT, a new approach for fact-checking black-box model responses without an external database, proving its superior ability to detect and rank factual and non-factual sentences.
- [A Simple and Effective Pruning Approach for Large Language Models](https://www.ml-quant.com/papers/arxiv/2306.11695/): 981 citations. Wanda, a new method, efficiently prunes weights in Large Language Models without retraining, offering a more efficient approach to inducing sparsity in pretrained models.
- [TinyLlama: An Open-Source Small Language Model](https://www.ml-quant.com/papers/arxiv/2401.02385/): 902 citations. Small Open-Source Language Model: The article presents TinyLlama, a compact 1.1B language model that performs remarkably well in various tasks despite its small size, having been pretrained on around 1 trillion tokens.
- [TÜLU 3: Pushing Frontiers in Open Language Model Post-Training](https://www.ml-quant.com/papers/arxiv/2411.15124/): 888 citations. Open Language Model Post-Training: The Tulu 3 model, a top-tier post-trained language model, is introduced, outperforming other models and providing a detailed guide for its use and adaptation.

## Latest

- [Financial Language Models as Applied Artificial Intelligence Systems for News-Based Trading under Market Frictions](https://www.ml-quant.com/papers/arxiv/2609.23703/) (2026-09-25): Introduces a market-friction-aware framework that converts timestamped financial news into auditable trading decisions while accounting for execution timing, transaction costs, and liquidity constraints.
- [From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication](https://www.ml-quant.com/papers/arxiv/2609.25034/) (2026-09-25): The study shows that how monetary policy sentiment unfolds across a press conference, not just its average tone, predicts rate changes and shapes forecaster expectations at the ECB and Fed.
- [FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering](https://www.ml-quant.com/papers/arxiv/2609.24002/) (2026-09-25): A benchmark reveals that language models answer financial questions above 90 percent with clarification but only 28.9 percent when they must elicit it themselves, exposing model ambiguity resolution.
- [FinRankGRPO: Optimizing LLMs for Listwise Financial Asset Ranking via Group Relative Policy Optimization](https://www.ml-quant.com/papers/arxiv/2609.24175/) (2026-09-25): Develops a two-stage framework that fine-tunes language models for listwise asset ranking using Spearman rank correlation rewards, achieving a Sharpe ratio of 0.636 on asset allocation.
- [LLM-Based Semantic Surprises in FOMC Communication: Asset Prices and Financial-Market Stress](https://www.ml-quant.com/papers/ssrn/7519200/) (2026-09-25): Semantic surprises extracted from Federal Reserve statements predict subsequent financial-stress dynamics and reduce forecast error by up to 23%, particularly when initial stress is high or during recessions.
- [Transforming the Voice of the Customer: Large Language Models for Identifying Customer Needs](https://www.ml-quant.com/papers/arxiv/2503.01870/) (2026-04-16): Large Language Models are streamlining the process of identifying customer needs, letting analysts concentrate on more valuable work while still delivering precise insights.
- [Twitter Sentiment and Financial Trends](https://www.ml-quant.com/papers/ssrn/4467949/) (2025-12-28): A new financial sentiment index derived from Twitter data shows strong links to market conditions and can forecast stock market returns, particularly in response to changes in U.S. monetary policy.
- [Measuring Corruption from Text Data](https://www.ml-quant.com/papers/arxiv/2512.09652/) (2025-12-14): An automated corruption index using Brazilian municipal audit reports is efficient and more reliable than manual methods in detecting corruption.
- [Reasoning Models Ace the CFA Exams](https://www.ml-quant.com/papers/arxiv/2512.08270/) (2025-12-14): An evaluation of reasoning models on CFA mock exams shows that models like Gemini 3.0 Pro and GPT-5 perform well, achieving high pass rates in professional testing.
- [Standard Occupation Classifier - A Natural Language Processing Approach](https://www.ml-quant.com/papers/arxiv/2511.23057/) (2025-12-01): A project successfully developed a natural language processing model that classifies job ads with 72% accuracy by using an ensemble approach.
- [Measuring economic outlook in the news](https://www.ml-quant.com/papers/arxiv/2511.04299/) (2025-11-12): We build an interpretable, privacy‑friendly sentiment indicator from Swiss news using ML and LLMs, and it improves short‑term GDP forecasts.
- [Aligning Multilingual News for Stock Return Prediction](https://www.ml-quant.com/papers/arxiv/2510.19203/) (2025-10-27): Uses optimal-transport to align English–Japanese stock news, producing clearer signals that better predict returns.
- [Black Box Absorption: LLMs Undermining Innovative Ideas](https://www.ml-quant.com/papers/arxiv/2510.20612/) (2025-10-27): LLM platforms can quietly absorb users’ ideas, creating power imbalances; the paper proposes governance and technical fixes to protect creators.
- [Integrating Transparent Models, LLMs, and Practitioner-in-the-Loop: A Case of Nonprofit Program Evaluation](https://www.ml-quant.com/papers/arxiv/2510.19799/) (2025-10-27): Combining transparent decision trees, LLMs, and practitioner input yields accurate, explainable case-level predictions for public and nonprofit programs.
- [News-Aware Direct Reinforcement Trading for Financial Markets](https://www.ml-quant.com/papers/arxiv/2510.19173/) (2025-10-27): - News-RL Trading - News-Driven Trading - News-Based RL - RL for News Trading - News-Powered RL If you want the shortest single choice: News-RL Trading.: Short summary: Feeding news sentiment extracted by large language models together with raw price and volume into a sequence-model reinforcement learning system boosts cryptocurrency trading performance, removing the need for handcrafted trading rules.
- [Abstract Classification: SVM vs BERT vs GPT-3.5](https://www.ml-quant.com/papers/repec/spr-scient-v-130-y-2025-i-1-d-10-1007-s11192-024-05217-7/) (2025-10-27): SVM vs BERT vs GPT-3.5: Compares SVM, SPECTER, BERT, and GPT-3.5 for classifying abstracts: BERT performs best, while GPT-3.5 is inconsistent with limited training data.
- [From Classical Rationality to Contextual Reasoning: Quantum Logic as a New Frontier for Human-Centric AI in Finance](https://www.ml-quant.com/papers/arxiv/2510.05475/) (2025-10-09): The potential of quantum logic in advancing artificial intelligence applications in financial modeling is discussed.
- [From News to Returns: A Granger-Causal Hypergraph Transformer on the Sphere](https://www.ml-quant.com/papers/arxiv/2510.04357/) (2025-10-09): The research suggests the CSHT, a new architecture for financial time-series forecasting that considers the impact of financial news and sentiment on asset returns, providing robust generalisation across market regimes and clear attribution pathways.
- [Extracting the Structure of Press Releases for Predicting Earnings Announcement Returns](https://www.ml-quant.com/papers/doi/10-1145-3768292-3770344/) (2025-10-03): The research explores the predictive power of textual features in earnings press releases on stock returns, concluding that press release content is as informative as earnings surprise, with FinBERT being the most predictive.
- [The (Short-Term) Effects of Large Language Models on Unemployment and Earnings](https://www.ml-quant.com/papers/arxiv/2509.15510/) (2025-09-22): Large Language Models like ChatGPT have boosted earnings for workers in certain occupations without affecting unemployment rates, indicating they enhance income rather than replace jobs.
- [Context-Aware Language Models for Forecasting Market Impact from Sequences of Financial News](https://www.ml-quant.com/papers/arxiv/2509.12519/) (2025-09-22): The study suggests using large language models to process financial news and small models to encode historical context, resulting in improved simulated investment performance.
- [Trading-R1: Financial Trading with LLM Reasoning via Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2509.11420/) (2025-09-22): The article presents Trading-R1, a finance-focused AI model that aligns with trading principles, showing it offers better risk-adjusted returns and fewer drawdowns than other models.
- [FinReflectKG: Agentic Construction and Evaluation of Financial Knowledge Graphs](https://www.ml-quant.com/papers/arxiv/2508.17906/) (2025-08-29): Financial Knowledge Graph Construction: The paper introduces a large-scale financial knowledge graph dataset from SEC 10-K filings of S and P 100 companies, with a reflection-agent-based mode providing the best balance of efficiency, accuracy, and reliability.
- [Bias-Adjusted LLM Agents for Human-Like Decision-Making via Behavioral Economics](https://www.ml-quant.com/papers/arxiv/2508.18600/) (2025-08-29): A persona-based approach using individual-level data from behavioral economics shows potential in adjusting biases in large language models, enabling them to simulate human-like decision patterns.
- [AlphaAgents: Large Language Model based Multi-Agents for Equity Portfolio Constructions](https://www.ml-quant.com/papers/arxiv/2508.11152/) (2025-08-20): The article investigates the application and effectiveness of role-based multi-agent AI systems in equity research and portfolio management, including their advantages and challenges.
- [Note on Selection Bias in Observational Estimates of Algorithmic Progress](https://www.ml-quant.com/papers/arxiv/2508.11033/) (2025-08-20): Criticisms have been raised about Ho et. al's 2024 study on the growing efficiency of language models, including potential selection bias in assessing algorithmic quality.
- [Interpreting the Interpreter: Can We Model post-ECB Conferences Volatility with LLM Agents?](https://www.ml-quant.com/papers/arxiv/2508.13635/) (2025-08-20): A new method using a Large Language Model can predict financial market responses to European Central Bank press conferences, aiding in maintaining financial stability.
- [A Multi-Task Evaluation of LLMs' Processing of Academic Text Input](https://www.ml-quant.com/papers/arxiv/2508.11779/) (2025-08-20): Large language models like Google's Gemini struggle with processing academic text, showing reliable summarizing and paraphrasing skills but poor text grading and reflection abilities, hence their unchecked use in peer reviews is discouraged.
- [Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models](https://www.ml-quant.com/papers/arxiv/2508.10192/) (2025-08-20): The paper presents Semantic Divergence Metrics (SDM), a new system that improves the detection of significant deviations in Large Language Models' responses from the input context by measuring consistency across various semantically equivalent paraphrases.
- [Event-Aware Sentiment Factors from LLM-Augmented Financial Tweets: A Transparent Framework for Interpretable Quant Trading](https://www.ml-quant.com/papers/arxiv/2508.07408/) (2025-08-12): The study shows how large language models can be used in financial semantic annotation and alpha signal discovery, indicating that social media sentiment can be a useful predictor in financial forecasting.
