ML-QuantSubscribe

Machine learningLLMs & Text

Data Selection for Language Models via Importance Resampling

The Data Selection with Importance Resampling (DSIR) framework is a new method for selecting subsets of large raw unlabeled datasets, surpassing manual curation and heuristic filtering methods in both specific and general language models.

Featured in No. 23 on 25 Oct 2023 · · 371 citations today · published in Neural Information Processing Systems

Released
6 Feb 2023
First featured
No. 23 · 25 Oct 2023
Citations (Semantic Scholar)
371
Influential citations
42
Published in
Neural Information Processing Systems
Shares when featured
279
Identifier
arXiv:2302.03169

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page