---
title: SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
url: https://www.ml-quant.com/papers/arxiv/2409.07440/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2409.07440
source_url: https://arxiv.org/abs/2409.07440
featured: 2024-09-18
citations: 54
topic: LLMs & Text
---


# SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories

The SUPER benchmark tests Large Language Models' ability to set up and execute tasks from research repositories, revealing that current models struggle with these tasks, suggesting a need for further advancements in this field.

- Source: https://arxiv.org/abs/2409.07440
- Identifier: arXiv:2409.07440
- Released: 2024-09-11
- First featured: Quant Letter No. 66 (2024-09-18): https://www.ml-quant.com/issues/2024-09-18/
- Citations (Semantic Scholar): 54
- Published in: Conference on Empirical Methods in Natural Language Processing
- Topic: LLMs & Text

## Related

- [Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters](https://www.ml-quant.com/papers/arxiv/2408.03314/): The research investigates enhancing Large Language Models' (LLMs) performance using more test-time computation, suggesting a compute-optimal scaling strategy based on prompt difficulty.
- [AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration](https://www.ml-quant.com/papers/arxiv/2306.00978/): The study suggests Activation-aware Weight Quantization (AWQ), a hardware-friendly method for quantizing large language models that reduces error and improves performance on various benchmarks.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/): The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [MemGPT: Towards LLMs as Operating Systems](https://www.ml-quant.com/papers/arxiv/2310.08560/): Extended Context in LLMs: MemGPT is a system that manages different memory levels, providing extended context within large language models' limited context windows, enhancing document analysis and multi-session chat performance.
- [SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models](https://www.ml-quant.com/papers/arxiv/2303.08896/): Hallucination Detection for LLMs: The paper presents SelfCheckGPT, a new approach for fact-checking black-box model responses without an external database, proving its superior ability to detect and rank factual and non-factual sentences.
- [A Simple and Effective Pruning Approach for Large Language Models](https://www.ml-quant.com/papers/arxiv/2306.11695/): Wanda, a new method, efficiently prunes weights in Large Language Models without retraining, offering a more efficient approach to inducing sparsity in pretrained models.
