Machine learningLLMs & Text
RefreshKV: Updating Small KV Cache During Long-form Generation
Recycled Attention is a method for large language models that alternates attention between full context and a subset of input tokens, improving performance and reducing computational load in long-context tasks.
Featured in No. 74 on 13 Nov 2024 · 5 days after release · 12 citations today · published in Annual Meeting of the Association for Computational Linguistics
- Released
- 8 Nov 2024
- First featured
- No. 74 · 13 Nov 2024
- Citations (Semantic Scholar)
- 12
- Influential citations
- 1
- Published in
- Annual Meeting of the Association for Computational Linguistics
- Shares when featured
- 14
- Identifier
- arXiv:2411.05787
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).