Machine learningML & AI Methods
Gated Linear Attention Transformers with Hardware-Efficient Training
Efficient Training of Gated Linear Attention Transformers: The research introduces a more hardware-efficient version of gated linear attention Transformers that performs well against other models, especially in training on longer sequences.
Featured in No. 29 on 13 Dec 2023 · 2 days after release · 539 citations today · published in International Conference on Machine Learning
- Released
- 11 Dec 2023
- First featured
- No. 29 · 13 Dec 2023
- Citations (Semantic Scholar)
- 539
- Influential citations
- 85
- Published in
- International Conference on Machine Learning
- Shares when featured
- 150
- Identifier
- arXiv:2312.06635
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).