Machine learningLLMs & Text
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
Research suggests that Large Language Models (LLMs) can be compromised by benign data, proposing a bi-directional anchoring method to identify such data and maintain model safety.
Featured in No. 62 on 21 Aug 2024 · · 124 citations today
- Released
- 1 Apr 2024
- First featured
- No. 62 · 21 Aug 2024
- Citations (Semantic Scholar)
- 124
- Influential citations
- 13
- Published in
- Not yet, as far as Semantic Scholar knows
- Shares when featured
- 149
- Identifier
- arXiv:2404.01099
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).