---
title: Memorisation or Alpha? Detecting Look-Ahead Contamination in Cross-Sectional Equity Signals
url: https://www.ml-quant.com/papers/ssrn/7490302/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: SSRN 7490302
source_url: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7490302
featured: 2026-09-25
citations: unknown
topic: ML & AI Methods
---


# Memorisation or Alpha? Detecting Look-Ahead Contamination in Cross-Sectional Equity Signals

Testing whether a large language model ranks stocks by forecasting or memory, the study finds a significant information-coefficient gap of 0.185 inside versus outside its training window, suggesting substantial look-ahead contamination.

- Source: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7490302
- Identifier: SSRN 7490302
- Released: 2026-09-22
- First featured: Quant Letter No. 132 (2026-09-25): https://www.ml-quant.com/issues/2026-09-25/
- Citations (Semantic Scholar): not tracked
- Published in: not yet
- Topic: ML & AI Methods
- Authors: Bach Nguyen

## Related

- [AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments](https://www.ml-quant.com/papers/arxiv/2405.07960/): AI Evaluation in Clinical Environments: The paper introduces AgentClinic, a benchmark for assessing large language models in simulated clinical environments, highlighting the significant impact of biases on diagnostic accuracy and patient interactions.
- [Learning Performance-Improving Code Edits](https://www.ml-quant.com/papers/arxiv/2302.07867/): The research presents a framework for optimizing programs using large language models, achieving a mean speedup of 6.86, outperforming average individual programmers.
- [Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia](https://www.ml-quant.com/papers/arxiv/2312.03664/): Concordia is a library designed to help build and operate Generative Agent-Based Models (GABMs), using Large Language Models (LLMs) to simulate physical or digital environments.
- [Jamba-1.5: Hybrid Transformer-Mamba Models at Scale](https://www.ml-quant.com/papers/arxiv/2408.12570/): Transformer-Mamba Models: Jamba-1.5 is a new large language model with enhanced conversational and instruction-following capabilities, featuring a unique quantization technique for cost-effective inference.
- [MindSearch: Mimicking Human Minds Elicits Deep AI Searcher](https://www.ml-quant.com/papers/arxiv/2407.20183/): Mimicking Human Minds for Search: MindSearch is a Large Language Model-based framework that simulates human cognitive processes for web information seeking, greatly enhancing response quality.
- [From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets](https://www.ml-quant.com/papers/arxiv/2509.06069/): Large language models (LLMs) in expert services have pros and cons, with human markets being more efficient, but LLMs potentially reducing trust and overshadowing experts' preferences, while also improving experts' communication of their goals.
