REMORY: Learning Residual Memory for Context Compaction Paper • 2610.11287 • Published 3 days ago • 10
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors Paper • 2609.34344 • Published 13 days ago • 5
view article Article YODAS v3: A 1 Million Hour Dataset for the Next Generation of Open Voice AI Research espnet • 14 days ago • 36
EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory Paper • 2610.10533 • Published 4 days ago • 7
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization Paper • 2609.38169 • Published 12 days ago • 118
Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective Paper • 2610.03185 • Published 9 days ago • 29
ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation Paper • 2609.39306 • Published 11 days ago • 33
Questioning the Questions: Sustaining Self-Evolution in Reasoning Models Paper • 2610.04299 • Published 8 days ago • 70
LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment Paper • 2603.00042 • Published May 3 • 2
Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution Paper • 2610.06804 • Published 6 days ago • 6
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches Paper • 2610.06647 • Published 6 days ago • 97
HuatuoGPT-3: RL-Only Domain Adaptation from Base Models Paper • 2610.05966 • Published 6 days ago • 43
Labels Override Definitions in Jev-Style Typed Decision Models Paper • 2610.02586 • Published 10 days ago • 7
Stepped MoE: Segment-Level Routing with Configurable Inference Complexity Paper • 2610.07348 • Published 6 days ago • 4
OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation Paper • 2610.02781 • Published 9 days ago • 16