Machine learningLLMs & Text
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
SimLayerKV is a technique that minimizes memory usage in large language models by identifying and reducing cache in lazy layers, achieving significant cache compression with minimal performance loss.
Featured in No. 71 on 23 Oct 2024 · 6 days after release · 13 citations today · published in Trans. Mach. Learn. Res.
- Released
- 17 Oct 2024
- First featured
- No. 71 · 23 Oct 2024
- Citations (Semantic Scholar)
- 13
- Influential citations
- 0
- Published in
- Trans. Mach. Learn. Res.
- Shares when featured
- 8
- Identifier
- arXiv:2410.13846
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).