Library
Papers with code
Papers that shipped their code, from the Papers with Code feed (2023-25). 837 items, newest first.
- 20 May 2026219shares
SkillsVote: Governance for Agent Skills
Governance for Agent Skills: SkillsVote is a system that helps long-term AI agents manage and improve their skills over time.
Papers with codeTrendingIn No. 131
- 20 May 2026101shares
MMSkills: Multimodal Skills for Agents
Multimodal Skills for Agents: Multimodal procedural knowledge frameworks enhance visual agents by merging text and visuals for improved decision-making.
Papers with codeTrendingIn No. 131
- 20 May 202659shares
SelfDistilled RL: Enhancing Training with Self-Distillation
Enhancing Training with Self-Distillation: SDAR optimizes training for multiturn agents in reinforcement learning by boosting positive feedback and minimizing negative feedback.
Papers with codeTrendingIn No. 131
- 20 May 202648shares
AI Research Trust
AI excels in structured tasks but needs human oversight in scientific research to handle new ideas and judgment effectively.
Papers with codeRisingIn No. 131
- 20 May 202642shares
Olympiad Reasoning Scaling
A new systematic approach boosts reasoning models, making them top contenders in math and physics competitions through advanced learning methods.
Papers with codeRisingIn No. 131
- 20 May 202633shares
Dynamic NPC Steering
ReactiveGWM improves game play by separating player controls from NPC actions, enhancing strategic options in various games.
Papers with codeRisingIn No. 131
- 20 May 202631shares
DexJoCo for Manipulation
DexJoCo provides a benchmark and toolkit for evaluating dexterous manipulation skills, along with a budget-friendly data collection system.
Papers with codeRisingIn No. 131
- 16 Apr 2026376shares
SkillClaw: Skill Evolution
Skill Evolution: SkillClaw enhances multiuser AI systems by leveraging group interactions to strengthen shared abilities.
Papers with codeTrendingIn No. 130
- 16 Apr 2026367shares
ClawGUI: Unified GUI
Unified GUI: ClawGUI is an open-source tool that streamlines the development of GUI agents using unified reinforcement learning across different platforms.
Papers with codeTrendingIn No. 130
- 16 Apr 2026200shares
DDTree: Speculative Decoding
Speculative Decoding: DDTree improves speculative decoding by generating draft trees from data distributions and validating multiple paths at once.
Papers with codeTrendingIn No. 130
- 16 Apr 202655shares
Strips as Tokens
SATO introduces a new way to order tokens in transformers that improves mesh generation by preserving edge flow using triangle strip sequences.
Papers with codeTrendingIn No. 130
- 16 Apr 202653shares
Introspective Consistency
Introspective Diffusion Language Models enhance autoregressive models by refining their output consistency through advanced decoding and optimized methods.
Papers with codeTrendingIn No. 130
- 16 Apr 202646shares
HabitatGS Navigation
HabitatGS enhances HabitatSim with 3D Gaussian Splatting for realistic visuals and dynamic avatars, improving AI agent navigation and generalization.
Papers with codeTrendingIn No. 130
- 16 Apr 202646shares
KnowUBench: Mobile Agent Evaluation
Mobile Agent Evaluation: KnowUBench assesses how well personalized mobile agents can understand user preferences and provide helpful assistance in real-world graphical user interfaces.
Papers with codeRisingIn No. 130
- 16 Apr 202641shares
KnowRL: LLM Reasoning Enhancement
LLM Reasoning Enhancement: KnowRL improves the reasoning abilities of language models through a framework that uses reinforcement learning to provide better, guided interactions.
Papers with codeRisingIn No. 130
- 16 Apr 202635shares
OnPolicy Distillation in Language Models
Effective distillation in large language models depends on matching thought processes between teacher and student models, with teachers needing to impart new skills.
Papers with codeRisingIn No. 130
- 16 Apr 202634shares
Parallel Decoding for Diffusion Models
DMax introduces a new technique for diffusion language models that reduces mistakes in parallel decoding.
Papers with codeRisingIn No. 130
- 16 Apr 202633shares
Autonomous ML with AiScientist
AiScientist develops a system that boosts long-term machine learning research by improving coordination and project management.
Papers with codeRisingIn No. 130
- 16 Apr 202622shares
Benchmarking LLMs for Human Behavior Simulation
The OmniBehavior benchmark reveals that large language models have difficulty mimicking complex human behaviors due to biases and limited diversity.
Papers with codeRisingIn No. 130
- 3 Apr 2026475shares
Latent Space as Foundation
Latent space improves language models by creating a continuous representation that minimizes redundancy and boosts efficiency.
Papers with codeTrendingIn No. 129
- 3 Apr 2026280shares
Unified Multimodal Processing
The Discrete Native Autoregressive framework enables integrated handling of various data types through a common discrete space and innovative visual transformer design.
Papers with codeTrendingIn No. 129
- 3 Apr 2026117shares
Generative World Renderer: Enhanced AAA Game Rendering
Enhanced AAA Game Rendering: A new dataset from AAA games enhances rendering quality and better evaluation methods that match human perception.
Papers with codeRisingIn No. 129
- 3 Apr 202665shares
SKILL0: RL for Skill Internalization
RL for Skill Internalization: SKILL0 empowers LLM agents to autonomously learn and execute tasks, boosting their effectiveness with a flexible training process.
Papers with codeRisingIn No. 129
- 3 Apr 202630shares
GEMS Multimodal Generation Framework with Memory
GEMS introduces a multimodal framework that helps agents refine their skills and memory, leading to improved performance across different tasks.
Papers with codeRisingIn No. 129
- 4 Mar 20261,883shares
dLLM Diffusion Framework
A new open-source framework has been created to make it easier and more consistent to develop and customize diffusion language models.
Papers with codeTrendingIn No. 128
- 4 Mar 202694shares
MobilityBench: Route Planning Evaluation
Route Planning Evaluation: MobileBench offers a scalable way to evaluate route-planning agents using large language models with real-world anonymized query testing.
Papers with codeTrendingIn No. 128
- 4 Mar 202679shares
Trinity of World Model Consistency
To improve World Models, three key principles—modal, spatial, and temporal—are identified, along with a benchmark for assessing multimodal learning systems.
Papers with codeTrendingIn No. 128
- 4 Mar 202622shares
BeyondSWE: Evaluating Code Agents
Evaluating Code Agents: BeyondSWE and SearchSWE are new benchmarks that improve the evaluation of coding agents by including real-life reasoning and knowledge.
Papers with codeRisingIn No. 128
- 4 Mar 202616shares
ARLArena: Stable Reinforcement Learning
Stable Reinforcement Learning: The ARLArena framework boosts reinforcement learning training stability through the SAMPO method for more consistent policy optimization.
Papers with codeRisingIn No. 128
- 4 Mar 202613shares
MediXR1: Medical RL for Clinical Reasoning
Medical RL for Clinical Reasoning: MediXR1 is a flexible reinforcement learning framework for medical language models, enhancing clinical reasoning with varied rewards and assessments.
Papers with codeRisingIn No. 128
- 12 Feb 20261,238shares
Flash Frontier: Sparse Agentic Intelligence
Sparse Agentic Intelligence: Step 3.5 Flash is an AI model designed for advanced intelligence using efficient parameters and attention mechanisms.
Papers with codeTrendingIn No. 127
- 12 Feb 2026866shares
InternAgent 1.5: Autonomous Discovery System
Autonomous Discovery System: InternAgent1.5 is a system that integrates computational modeling with experimental research for autonomous scientific discovery.
Papers with codeTrendingIn No. 127
- 12 Feb 2026133shares
CodeWorld: GUI Agents for Visual States
GUI Agents for Visual States: Code2World enables autonomous GUI agents to predict visual outcomes by generating renderable code, improving navigation.
Papers with codeTrendingIn No. 127
- 12 Feb 202692shares
SkillRL: Evolving RL Agents
Evolving RL Agents: SkillRL enhances large language models by enabling skill discovery and policy evolution while conserving computational resources.
Papers with codeTrendingIn No. 127
- 12 Feb 202660shares
UniAudio 2.0: Unified Audio Model
Unified Audio Model: Researchers developed ReasoningCodec for improved audio processing and UniAudio 2.0, a strong model for various audio tasks using a large dataset.
Papers with codeTrendingIn No. 127
- 12 Feb 202639shares
WeakDriven Learning: Improving Agents with Weak Checkpoints
Improving Agents with Weak Checkpoints: WMSS boosts large language models by using weak checkpoints to close learning gaps for better results.
Papers with codeRisingIn No. 127
- 12 Feb 202626shares
Agent World Model: Advancing RL in Synthetic Environments
Advancing RL in Synthetic Environments: Large language models in synthetic settings exceed traditional models in adapting to new situations.
Papers with codeRisingIn No. 127
- 12 Feb 202620shares
ROCKET: Model Compression via Knapsack Optimization
Model Compression via Knapsack Optimization: ROCKET introduces an efficient model compression technique that treats the process like solving a knapsack problem.
Papers with codeRisingIn No. 127
- 12 Feb 202619shares
Prism: Enhanced Block-Sparse Attention for LLMs
Enhanced Block-Sparse Attention for LLMs: Prism enhances long-context LLM performance by optimizing which blocks to select in blocksparse attention.
Papers with codeRisingIn No. 127
- 12 Feb 202617shares
LOCAbench: Benchmarking Language Agents in Extreme Contexts
Benchmarking Language Agents in Extreme Contexts: LOCAbench is a benchmark designed to evaluate language agents handling long-context tasks.
Papers with codeRisingIn No. 127
- 2 Feb 2026176shares
AI Safety Framework
AI agents require better safety protocols when using tools and interacting with their surroundings.
Papers with codeTrendingIn No. 126
- 2 Feb 2026145shares
SelfEvolving Systems
Adaptive agents that learn from feedback can adapt more effectively to changing environments and enhance knowledge sharing.
Papers with codeTrendingIn No. 126
- 2 Feb 202682shares
Enhancing Math Reasoning
MathForge enhances AI's mathematical reasoning by adjusting question difficulty and reformulating problems.
Papers with codeTrendingIn No. 126
- 2 Feb 202680shares
ASTRA: Automated Decision Framework
Automated Decision Framework: ASTRA trains language models with synthetic data to improve their ability to make complex decisions.
Papers with codeRisingIn No. 126
- 2 Feb 202642shares
AdaReasoner: Visual Reasoning Tool
Visual Reasoning Tool: AdaReasoner enables multimodal models to learn how to use tools effectively for enhanced visual reasoning through scalable data and adaptive methods.
Papers with codeRisingIn No. 126
- 2 Feb 202631shares
Token Filtering for Models
Token filtering during pretraining minimizes undesirable traits in language models and enhances performance while managing noisy labels.
Papers with codeRisingIn No. 126
- 2 Feb 202614shares
SelfDistillation in RL
SelfDistillation Policy Optimization (SDPO) enhances reinforcement learning by incorporating textual feedback to make training language models more efficient and accurate.
Papers with codeRisingIn No. 126
- 16 Jan 20262,083shares
Scalable Conditional Memory
Conditional memory in Transformer models improves knowledge retrieval and reasoning by efficiently managing data sparsity.
Papers with codeTrendingIn No. 125
- 16 Jan 2026600shares
Unified Multimodal Retrieval
The Qwen3VLEmbedding and Qwen3VLReranker models form a precise multimodal search system using cross-attention techniques.
Papers with codeTrendingIn No. 125
- 16 Jan 202668shares
Advancements in 3D Orientation
Orient Anything V2 improves understanding of 3D orientation through new asset synthesis and rotation prediction methods.
Papers with codeTrendingIn No. 125
- 16 Jan 202663shares
Optimized Code Generation
The Controlled SelfEvolution method boosts code generation by employing genetic evolution and better exploration strategies.
Papers with codeTrendingIn No. 125
- 16 Jan 202661shares
Automated Research Evaluation
DeepResearchEval uses adaptable agents to automate complex research tasks, verifying facts without relying on citations.
Papers with codeRisingIn No. 125
- 16 Jan 202647shares
Linear Attention Enhancement
MultiHead Linear Attention boosts performance by balancing representational diversity and computational efficiency while increasing expressive capabilities.
Papers with codeRisingIn No. 125
- 16 Jan 202638shares
Synthetic Data for Programming
Code LLMs trained on synthetic data excel in competitive programming compared to traditional models, reducing dependence on real datasets.
Papers with codeRisingIn No. 125
- 16 Jan 202634shares
SelfEvolving Reasoning Agents
A data-free self-evolution framework enables large language models to improve reasoning by generating their own questions, achieving performance similar to supervised learning.
Papers with codeRisingIn No. 125
- 16 Jan 202626shares
Stable Sinkhorn-Knopp Iterations
Hyperconnections with dynamic residual matrices improve convergence stability through a new reparameterization technique that ensures precise doubly stochasticity.
Papers with codeRisingIn No. 125
- 28 Dec 2025328shares
Agentic AI Adaptation Framework
The paper presents a framework for improving AI systems by adapting agents and tools, focusing on design strategies and challenges to boost AI performance.
Papers with codeTrendingIn No. 124
- 28 Dec 202550shares
Evaluating LLMs for Science
A new framework for Scientific General Intelligence (SGI) is introduced, enhancing existing models with TestTime Reinforcement Learning to improve their scientific capabilities.
Papers with codeRisingIn No. 124
- 28 Dec 202519shares
Uncovering Language Model Policies
The research analyzes large language model policies by dissecting them into modular parts, revealing reasoning patterns and advocating for Bottom-up Policy Optimization to enhance complex reasoning tasks.
Papers with codeRisingIn No. 124
- 19 Dec 20251,450shares
StepGUI: Optimizing GUI Automation
Optimizing GUI Automation: A new training system boosts GUI automation by making it more efficient, accurate, and private for real-world use.
Papers with codeTrendingIn No. 123
- 19 Dec 2025119shares
AI Agents: Memory Research Overview
Memory Research Overview: The survey examines agent memory research, detailing its types, functions, and future research directions.
Papers with codeTrendingIn No. 123
- 19 Dec 202531shares
Jacobi Forcing: Parallel Decoding Efficiency
Parallel Decoding Efficiency: Jacobi Forcing enables quicker and more efficient parallel decoding of transformer models while maintaining performance.
Papers with codeTrendingIn No. 123
- 19 Dec 202527shares
LitePT: Efficient 3D Model
Efficient 3D Model: LitePT is a 3D point cloud model that uses a combination of convolutions and attention to enhance efficiency.
Papers with codeRisingIn No. 123
- 19 Dec 202521shares
ErrorFree Attention: Optimized Language Model
Optimized Language Model: ErrorFree Linear Attention (EFLA) is a quick and dependable attention method that outperforms DeltaNet in language processing tasks.
Papers with codeRisingIn No. 123
- 19 Dec 202520shares
Universal Reasoning: Enhanced Transformer
Enhanced Transformer: The Universal Reasoning Model enhances Universal Transformers to improve reasoning abilities on ARCAGI challenges.
Papers with codeRisingIn No. 123
- 14 Dec 202511,739shares
DeepCode: Doc-to-Code Synthesis
Doc-to-Code Synthesis: DeepCode is an autonomous framework that improves the process of turning documents into code, surpassing human experts with its advanced optimization techniques.
Papers with codeTrendingIn No. 122
- 14 Dec 202539shares
GRAPE: Positional Encoding Framework
Positional Encoding Framework: GRAPE is a new framework for positional encoding that combines rotations and logit biases to enhance the performance of existing methods like RoPE and ALiBi.
Papers with codeTrendingIn No. 122
- 14 Dec 202538shares
Textto-3D Reinforcement Learning
ARDR1 is a groundbreaking model that uses reinforcement learning to generate 3D content from text, featuring new reward systems and optimization methods.
Papers with codeRisingIn No. 122
- 14 Dec 202522shares
Procedural Terrain Diffusion
Terrain Diffusion leverages diffusion models and InfiniteDiffusion to produce realistic, infinitely expandable environments that can be accessed quickly.
Papers with codeRisingIn No. 122
- 1 Dec 2025312shares
Evolving Cost-Effective Deep Learning Rubrics
Reinforcement Learning with Evolving Rubrics (RLER) enhances the affordable training of deep learning models for lengthy tasks and outperforms current methods.
Papers with codeTrendingIn No. 121
- 1 Dec 202587shares
Efficient Memory
GAM is a framework that improves memory efficiency and task performance through JIT compilation, a lightweight memorizer, and reinforcement learning.
Papers with codeRisingIn No. 121
- 1 Dec 202555shares
Collaborative Representation
LatentMAS boosts collaboration among LLM agents by utilizing latent space representations, enhancing reasoning quality while reducing computational expenses.
Papers with codeRisingIn No. 121
- 19 Nov 2025743shares
MiroThinker — Scalable Research Agent
MiroThinker v1.0 — An open-source agent that boosts reasoning by training models to interact more with tools and environments, enabling long multi-step workflows and strong benchmark results.
Papers with codeTrendingIn No. 120
- 19 Nov 202544shares
P RL Physics Olympiad Champion
P A family of open-source, reinforcement-learned models that excel at physics reasoning (winning top medals at IPhO 2025) and also perform well on math and coding.
Papers with codeTrendingIn No. 120
- 19 Nov 202533shares
PartXMLLM — 3D Part-aware LLM
PartXMLLM: Turns a colored 3D point cloud and a text prompt into a compact program that lists part-level boxes, labels, and edit commands to drive geometry tools.
Papers with codeRisingIn No. 120
- 19 Nov 202512shares
LoopTool — Robust LLM Call Loop
LoopTool: An automated loop that repeatedly generates tool-using examples, refines data and models, and retrains LLMs to improve their tool-use and overall performance.
Papers with codeRisingIn No. 120
- 12 Nov 2025937shares
DeepEyesV2: Agentic Multimodal
Agentic Multimodal: DeepEyesV2 — a two-stage-trained multimodal agent that learns to use external tools to solve real-world reasoning tasks.
Papers with codeTrendingIn No. 119
- 12 Nov 202535shares
Diffusion LMs: Low-Data Learners
Low-Data Learners: Diffusion language models — they predict tokens via iterative denoising and flexible token order, outperforming autoregressive models with limited data and keeping advantages as they scale.
Papers with codeTrendingIn No. 119
- 12 Nov 202518shares
Adaptive Verifiable Environments for LM Reasoning
RLVE Gradually increases problem difficulty during training so language models learn stronger reasoning, outperforming fixed environments and standard RL.
Papers with codeRisingIn No. 119
- 12 Nov 202514shares
Retrofitted Recurrence for Deeper LM Reasoning
Recurrence curriculum: Convert pretrained nonrecurrent models to depth-recurrent ones and train with an increasing-recurrence curriculum to get better results for the same compute.
Papers with codeRisingIn No. 119
- 4 Nov 2025500shares
Kimi Linear: Hybrid Linear Attention
Hybrid Linear Attention: Kimi Linear: a new hybrid linear‑attention model that beats full attention while using far less memory, decoding much faster, and with code and checkpoints open‑sourced.
Papers with codeTrendingIn No. 118
- 4 Nov 202588shares
FP Fixes Training–Inference Mismatch in RL Fine-tuning
Switching RL fine‑tuning from BF16 to FP16: fixes numerical mismatches between training and inference, making training more stable and faster.
Papers with codeTrendingIn No. 118
- 4 Nov 202533shares
Toolathlon: Long‑Horizon Agent Benchmark
Long‑Horizon Agent Benchmark: Toolathlon — a benchmark across 32 apps and 604 tools that tests agents on long, realistic workflows and shows current models often fail to complete them.
Papers with codeRisingIn No. 118
- 4 Nov 202519shares
CALM Continuous Next‑Vector Models
CALM — a method that compresses multiple tokens into one continuous vector and predicts those vectors instead of tokens, cutting generation steps and making generation much faster.
Papers with codeRisingIn No. 118
- 27 Oct 202577shares
DeepAgent — Autonomous Tool Agent
DeepAgent — an end-to-end agent that finds and uses tools on its own and compresses long interaction histories so it solves complex, long-horizon tasks better.
Papers with codeTrendingIn No. 117
- 27 Oct 202566shares
LightMem — Efficient Memory for LLMs
LightMem — a lightweight three-stage memory that filters, groups, and consolidates past interactions to improve LLM accuracy while greatly cutting token, API, and runtime costs.
Papers with codeTrendingIn No. 117
- 27 Oct 202558shares
DeepAnalyze — Auto Data-Science LLM
DeepAnalyze8B: an 8B-parameter model trained progressively to run full data-science pipelines and produce analyst-quality research reports.
Papers with codeRisingIn No. 117
- 27 Oct 202516shares
PBSAttn — Block-Sparse Attention
PBSAttn: rearranges attention into block-sparse patterns to speed long-context LLM processing while keeping accuracy close to full attention.
Papers with codeRisingIn No. 117
- 27 Oct 202511shares
EDR Steerable Multi-Agent Analytics
EnterpriseDeepResearch (EDR): a multiagent system that plans, searches, uses tools, visualizes results, and reflects to automatically produce high-quality enterprise research reports.
Papers with codeRisingIn No. 117
- 25 Jul 20254,954shares
Transformer Explainer
Despite the significant impact of Transformers on machine learning, their workings remain unclear to many people.
Papers with codeTrendingIn No. 107
- 25 Jul 20254,612shares
Superhuman Reasoning
The article discusses the challenge of surpassing human cognitive limitations in the training of Large Language Models (LLMs).
Papers with codeTrendingIn No. 107
- 25 Jul 20252,204shares
Workflow Generation Models
The article presents a curated dataset of 4K workflows used to create comprehensive reasoning data, including aspects like node selection, workflow planning, and code-level workflow representation.
Papers with codeTrendingIn No. 107
- 25 Jul 2025833shares
GLM4 vs Qwen2.5VL7B
The article discusses a model that performs better than Qwen2. 5VL7B in 28 public benchmarks and equals or surpasses Qwen2. 5VL72B in 18 benchmarks.
Papers with codeTrendingIn No. 107
- 25 Jul 2025605shares
NVIDIA Audio Flamingo
The article emphasizes the importance of enhancing large language models to understand audio, including non-speech sounds and non-verbal speech, for various real-world applications.
Papers with codeTrendingIn No. 107
- 25 Jul 2025399shares
Streaming 4D VGT
The article introduces a streaming 4D visual geometry transformer, similar to autoregressive large language models, designed to support interactive and real-time applications.
Papers with codeTrendingIn No. 107
- 25 Jul 2025365shares
IQLearn: Inverse softQ
Inverse softQ: The article highlights the availability of human or expert data in sequential decision-making tasks like robotics control and game playing, which offers valuable task-related information.
Papers with codeTrendingIn No. 107
- 25 Jul 2025233shares
Benchmarking Autonomous Agents
The article presents REAL, a tool for assessing multiturn agents' performance on simulations of actual websites.
Papers with codeRisingIn No. 107
- 25 Jul 2025210shares
Mathematical Deep Learning Introduction
The book offers a beginner-friendly guide to understanding deep learning algorithms.
Papers with codeRisingIn No. 107
- 25 Jul 2025169shares
HackSynth LLM Agent
The article introduces HackSynth, an agent based on Large Language Model for independent penetration testing.
Papers with codeRisingIn No. 107
- 25 Jul 2025119shares
PhysX Asset Generation
The article explores the shift of 3D modeling from a virtual environment to a physical one.
Papers with codeRisingIn No. 107