---
title: Stronger Models are NOT Stronger Teachers for Instruction Tuning
url: https://www.ml-quant.com/papers/arxiv/2411.07133/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2411.07133
source_url: https://arxiv.org/abs/2411.07133
featured: 2024-11-13
citations: 17
topic: LLMs & Text
---


# Stronger Models are NOT Stronger Teachers for Instruction Tuning

A study introduces Compatibility-Adjusted Reward (CAR), a new metric to evaluate the effectiveness of language models, challenging the belief that larger models are better for instruction tuning.

- Source: https://arxiv.org/abs/2411.07133
- Identifier: arXiv:2411.07133
- Released: 2024-11-11
- First featured: Quant Letter No. 74 (2024-11-13): https://www.ml-quant.com/issues/2024-11-13/
- Citations (Semantic Scholar): 17
- Published in: not yet
- Topic: LLMs & Text

## Related

- [Llemma: An Open Language Model For Mathematics](https://www.ml-quant.com/papers/arxiv/2310.10631/): Open Math Language Model: The article discusses Llemma, a superior language model for mathematics that can prove theorems without additional fine-tuning.
- [Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?](https://www.ml-quant.com/papers/arxiv/2405.05904/): Research shows that large language models have difficulty acquiring new factual knowledge through fine-tuning, learning new information slower than consistent knowledge, and are more likely to hallucinate, indicating the risks of introducing new facts through fine-tuning.
- [TableLlama: Towards Open Large Generalist Models for Tables](https://www.ml-quant.com/papers/arxiv/2311.09206/): The paper presents TableLlama, an open-source large language model fine-tuned for table-based tasks, and introduces a new dataset, TableInstruct, which improves the performance and generalizability of these models.
- [Designing Heterogeneous LLM Agents for Financial Sentiment Analysis](https://www.ml-quant.com/papers/arxiv/2401.05799/): A study suggests using large language models without fine-tuning for financial sentiment analysis, offering a design framework that enhances accuracy.
- [Instruct-FinGPT: Financial Sentiment Analysis by Instruction Tuning of General-Purpose Large Language Models](https://www.ml-quant.com/papers/arxiv/2306.12659/): A new approach improves financial sentiment analysis by addressing limitations of language models.
- [AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation](https://www.ml-quant.com/papers/arxiv/2408.00764/): Enhancing LLM Planning: The study improves the planning abilities of Large Language Models (LLMs) using instruction tuning and a framework called AgentGen, which generates diverse environments and planning tasks.
