Machine learningLLMs & Text
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
The SUPER benchmark tests Large Language Models' ability to set up and execute tasks from research repositories, revealing that current models struggle with these tasks, suggesting a need for further advancements in this field.
Featured in No. 66 on 18 Sep 2024 · 7 days after release · 54 citations today · published in Conference on Empirical Methods in Natural Language Processing
- Released
- 11 Sep 2024
- First featured
- No. 66 · 18 Sep 2024
- Citations (Semantic Scholar)
- 54
- Influential citations
- 6
- Published in
- Conference on Empirical Methods in Natural Language Processing
- Shares when featured
- 11
- Identifier
- arXiv:2409.07440
Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).