---
title: No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
url: https://www.ml-quant.com/papers/arxiv/2407.10964/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2407.10964
source_url: https://arxiv.org/abs/2407.10964
featured: 2024-11-13
citations: 6
topic: LLMs & Text
---


# No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations

The FUNGI method improves transformer encoders' features using self-supervised gradients, enhancing performance in vision, natural language processing, and audio tasks and datasets.

- Source: https://arxiv.org/abs/2407.10964
- Identifier: arXiv:2407.10964
- Released: 2024-07-15
- First featured: Quant Letter No. 74 (2024-11-13): https://www.ml-quant.com/issues/2024-11-13/
- Citations (Semantic Scholar): 6
- Published in: not yet
- Topic: LLMs & Text

## Related

- [player2vec: A Language Modeling Approach to Understand Player Behavior in Games](https://www.ml-quant.com/papers/arxiv/2404.04234/): Player Behavior in Games: A new technique for learning hidden user profiles from player behavior data in video and mobile games is presented, utilizing a long-range Transformer model from natural language processing, showing promising results in matching behavior event distribution.
- [SliceGPT: Compress Large Language Models by Deleting Rows and Columns](https://www.ml-quant.com/papers/arxiv/2401.15024/): Compressing Language Models: The paper introduces SliceGPT, a post-training sparsification scheme for large language models that reduces the network's embedding dimension, maintains high performance, reduces inference computation, and reveals computational invariance in transformer networks.
- [You Only Cache Once: Decoder-Decoder Architectures for Language Models](https://www.ml-quant.com/papers/arxiv/2405.05254/): YOCO architecture improves large language models by reducing GPU memory usage and speeding up the prefill stage, outperforming the Transformer model.
- [Planning In Natural Language Improves LLM Search For Code Generation](https://www.ml-quant.com/papers/arxiv/2409.03733/): PLANSEARCH is a new search algorithm that creates diverse solutions for natural language problems, outperforming traditional methods in various benchmarks.
- [Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together](https://www.ml-quant.com/papers/arxiv/2407.10930/): The article explores a method to enhance Natural Language Processing systems by simultaneously optimizing language model weights and prompting strategies, leading to significant improvements in tasks like multi-hop QA and mathematical reasoning.
- [AmbigNLG: Addressing Task Ambiguity in Instruction for NLG](https://www.ml-quant.com/papers/arxiv/2402.17717/): Task Ambiguity in NLG: AmbigNLG, a new task and dataset, tackles task ambiguity in instructions for Natural Language Generation, improving the alignment of generated text with user expectations and boosting the performance of Large Language Models.
