Machine learningLLMs & Text
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Large language models may not be truly reasoning but overfitting to specific datasets, as shown by decreased accuracy on new benchmarks.
Featured in No. 48 on 8 May 2024 · 7 days after release · 238 citations today · published in Neural Information Processing Systems
- Released
- 1 May 2024
- First featured
- No. 48 · 8 May 2024
- Citations (Semantic Scholar)
- 238
- Influential citations
- 12
- Published in
- Neural Information Processing Systems
- Shares when featured
- 1,161
- Identifier
- arXiv:2405.00332
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).