Machine learningLLMs & Text
Not All Language Model Features Are One-Dimensionally Linear
The study uses sparse autoencoders to explore the multi-dimensional nature of language model representations in GPT-2 and Mistral 7B, and identifies tasks where these features solve computational problems.
Featured in No. 51 on 28 May 2024 · 5 days after release · 207 citations today · published in International Conference on Learning Representations
- Released
- 23 May 2024
- First featured
- No. 51 · 28 May 2024
- Citations (Semantic Scholar)
- 207
- Influential citations
- 15
- Published in
- International Conference on Learning Representations
- Shares when featured
- 6
- Identifier
- arXiv:2405.14860
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).