Machine learningLLMs & Text
Data Selection for Language Models via Importance Resampling
The Data Selection with Importance Resampling (DSIR) framework is a new method for selecting subsets of large raw unlabeled datasets, surpassing manual curation and heuristic filtering methods in both specific and general language models.
Featured in No. 23 on 25 Oct 2023 · · 371 citations today · published in Neural Information Processing Systems
- Released
- 6 Feb 2023
- First featured
- No. 23 · 25 Oct 2023
- Citations (Semantic Scholar)
- 371
- Influential citations
- 42
- Published in
- Neural Information Processing Systems
- Shares when featured
- 279
- Identifier
- arXiv:2302.03169
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).