Machine learningLLMs & Text
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
A study shows that FP8 weight and activation quantization in large language models (LLMs) is lossless across all scales, while INT8 and INT4 quantizations also perform well with proper tuning, offering guidelines for deploying quantized LLMs.
Featured in No. 73 on 6 Nov 2024 · 2 days after release · 44 citations today · published in Annual Meeting of the Association for Computational Linguistics
- Released
- 4 Nov 2024
- First featured
- No. 73 · 6 Nov 2024
- Citations (Semantic Scholar)
- 44
- Influential citations
- 4
- Published in
- Annual Meeting of the Association for Computational Linguistics
- Shares when featured
- 50
- Identifier
- arXiv:2411.02355
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).