ML-QuantSubscribe

Machine learningLLMs & Text

You Only Cache Once: Decoder-Decoder Architectures for Language Models

YOCO architecture improves large language models by reducing GPU memory usage and speeding up the prefill stage, outperforming the Transformer model.

Featured in No. 49 on 15 May 2024 · 7 days after release · 155 citations today · published in Neural Information Processing Systems

Released
8 May 2024
First featured
No. 49 · 15 May 2024
Citations (Semantic Scholar)
155
Influential citations
14
Published in
Neural Information Processing Systems
Shares when featured
272
Identifier
arXiv:2405.05254

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page