---
title: The Memorization Problem: Can We Trust LLMs'Economic Forecasts?
url: https://www.ml-quant.com/papers/arxiv/2504.14765/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2504.14765
source_url: http://arxiv.org/abs/2504.14765v1
featured: 2025-04-23
citations: 36
topic: LLMs & Text
---


# The Memorization Problem: Can We Trust LLMs'Economic Forecasts?

The study shows that large language models can remember exact economic figures from before their knowledge cutoff dates, which may skew their predictive abilities in forecasting and backtesting trading strategies.

- Source: http://arxiv.org/abs/2504.14765v1
- Identifier: arXiv:2504.14765
- Released: 2025-04-21
- First featured: Quant Letter No. 94 (2025-04-23): https://www.ml-quant.com/issues/2025-04-23/
- Citations (Semantic Scholar): 36
- Published in: not yet
- Topic: LLMs & Text

## Related

- [Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?](https://www.ml-quant.com/papers/arxiv/2505.07078/): FINSABER, a backtesting framework, shows that Large Language Models' effectiveness in stock trading decreases over longer periods and larger symbol universes, emphasizing the need for trend detection and risk controls.
- [Look-Ahead Bias in Stock Return Predictions](https://www.ml-quant.com/papers/ssrn/4586726/): Large language models like ChatGPT can generate profitable trading signals from news sentiment, but backtesting can yield biased results due to overlapping periods.
- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration](https://www.ml-quant.com/papers/arxiv/2306.00978/): The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
- [MemGPT: Towards LLMs as Operating Systems](https://www.ml-quant.com/papers/arxiv/2310.08560/): Extended Context in LLMs: MemGPT is a system that manages different memory levels, providing extended context within large language models' limited context windows, enhancing document analysis and multi-session chat performance.
