Machine learningLLMs & Text
Do Large Language Model Benchmarks Test Reliability?
The article highlights the need for reliable large language models, criticizes current benchmarks for their inadequacy, and suggests the use of platinum benchmarks to reduce label errors and ambiguity.
Featured in No. 92 on 9 Apr 2025 · · 54 citations today
- Released
- 5 Feb 2025
- First featured
- No. 92 · 9 Apr 2025
- Citations (Semantic Scholar)
- 54
- Influential citations
- 6
- Published in
- Not yet, as far as Semantic Scholar knows
- Shares when featured
- 73
- Identifier
- arXiv:2502.03461
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).