---
title: Emergent mechanisms for long timescales depend on training curriculum and affect performance in memory tasks
url: https://www.ml-quant.com/papers/arxiv/2309.12927/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
identifier: arXiv:2309.12927
source_url: https://arxiv.org/abs/2309.12927
featured: 2024-10-31
citations: 13
topic: ML & AI Methods
---


# Emergent mechanisms for long timescales depend on training curriculum and affect performance in memory tasks

The research shows that recurrent neural networks improve performance and generalization by adapting timescales for memory-dependent tasks.

- Source: https://arxiv.org/abs/2309.12927
- Identifier: arXiv:2309.12927
- Released: 2023-09-22
- First featured: Quant Letter No. 72 (2024-10-31): https://www.ml-quant.com/issues/2024-10-31/
- Citations (Semantic Scholar): 13
- Published in: International Conference on Learning Representations
- Topic: ML & AI Methods

## Related

- [Mamba: Linear-Time Sequence Modeling with Selective State Spaces](https://www.ml-quant.com/papers/arxiv/2312.00752/): Sequence Modeling: Mamba, a neural network architecture that doesn't use attention or MLP blocks, provides faster inference and better performance in language, audio, and genomics than Transformers.
- [Graph Mamba: Towards Learning on Graphs with State Space Models](https://www.ml-quant.com/papers/arxiv/2402.08678/): Graph Mamba Networks, a new type of Graph Neural Networks, have been introduced, which achieve excellent performance in various benchmark datasets despite lower computational cost.
- [Edge Directionality Improves Learning on Heterophilic Graphs](https://www.ml-quant.com/papers/arxiv/2305.10498/): The study presents Directed Graph Neural Network (Dir-GNN), a new deep learning framework for directed graphs that surpasses traditional models in heterophilic benchmarks.
- [Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks](https://www.ml-quant.com/papers/arxiv/2310.02244/): Deep Residual Network Feature Learning: The research explores depthwise parametrizations in deep residual networks, pinpointing Depth-$\mu$P as the best parametrization for maximizing feature learning and diversity, but notes its limitations in deeper networks.
- [Modular Duality in Deep Learning](https://www.ml-quant.com/papers/arxiv/2410.21265/): The article presents a new theory of modular dualization for general neural networks, providing a theoretical basis for fast and scalable training algorithms, potentially leading to a new generation of optimizers for neural architectures.
- [Arabic Music Classification and Generation using Deep Learning](https://www.ml-quant.com/papers/arxiv/2410.19719/): The study suggests a machine learning method using a convolutional neural network for classifying and creating new and traditional Egyptian music by composer, with 81.4% accuracy.
