ML-QuantSubscribe

Machine learningLLMs & Text

SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories

The SUPER benchmark tests Large Language Models' ability to set up and execute tasks from research repositories, revealing that current models struggle with these tasks, suggesting a need for further advancements in this field.

Featured in No. 66 on 18 Sep 2024 · 7 days after release · 54 citations today · published in Conference on Empirical Methods in Natural Language Processing

Released
11 Sep 2024
First featured
No. 66 · 18 Sep 2024
Citations (Semantic Scholar)
54
Influential citations
6
Published in
Conference on Empirical Methods in Natural Language Processing
Shares when featured
11
Identifier
arXiv:2409.07440

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page