ML-QuantSubscribe

Machine learningLLMs & Text

"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization

A study shows that FP8 weight and activation quantization in large language models (LLMs) is lossless across all scales, while INT8 and INT4 quantizations also perform well with proper tuning, offering guidelines for deploying quantized LLMs.

Featured in No. 73 on 6 Nov 2024 · 2 days after release · 44 citations today · published in Annual Meeting of the Association for Computational Linguistics

Released
4 Nov 2024
First featured
No. 73 · 6 Nov 2024
Citations (Semantic Scholar)
44
Influential citations
4
Published in
Annual Meeting of the Association for Computational Linguistics
Shares when featured
50
Identifier
arXiv:2411.02355

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page