ML-QuantSubscribe

Machine learningML & AI Methods

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

The article introduces Cross-Layer Attention (CLA), a new attention design that minimizes the key-value cache size, allowing for longer sequence lengths and larger batch sizes during inference.

Featured in No. 51 on 28 May 2024 · 7 days after release · 140 citations today · published in Neural Information Processing Systems

Released
21 May 2024
First featured
No. 51 · 28 May 2024
Citations (Semantic Scholar)
140
Influential citations
10
Published in
Neural Information Processing Systems
Shares when featured
206
Identifier
arXiv:2405.12981

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page