Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction Paper • 2610.12299 • Published 3 days ago • 47
GRACE: Generation-aware latent compression for efficient video generation Paper • 2610.10524 • Published 4 days ago • 76
EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos Paper • 2609.39378 • Published 11 days ago • 74
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 11 days ago • 120
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering Paper • 2609.38177 • Published 12 days ago • 74
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control Paper • 2608.24115 • Published Aug 25 • 14
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Paper • 2606.18216 • Published Jun 16 • 65
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models Paper • 2512.03045 • Published Dec 2, 2025 • 3
Better Source, Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching Paper • 2602.05951 • Published Feb 5
Repurposing Geometric Foundation Models for Multi-view Diffusion Paper • 2603.22275 • Published Mar 23 • 50