ML-QuantSubscribe

Machine learningLLMs & Text

LLM Pruning and Distillation in Practice: The Minitron Approach

The article discusses the successful compression of Llama 3.1 8B and MistralNeMo 12B models to smaller parameters using pruning and distillation strategies, with the results tested on common benchmarks and the base model weights made available on Hugging Face.

Featured in No. 63 on 28 Aug 2024 · 7 days after release · 111 citations today

Released
21 Aug 2024
First featured
No. 63 · 28 Aug 2024
Citations (Semantic Scholar)
111
Influential citations
11
Published in
Not yet, as far as Semantic Scholar knows
Shares when featured
372
Identifier
arXiv:2408.11796

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page