It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them Paper • 2609.37863 • Published 6 days ago • 38 • 3
Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models Paper • 2609.33355 • Published 8 days ago • 46 • 3
TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent Paper • 2609.27277 • Published 12 days ago • 32 • 3
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion Paper • 2609.32540 • Published 9 days ago • 34 • 2
Learning from Teacher Continuations at Student States Paper • 2609.36246 • Published 7 days ago • 40 • 2
RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement Paper • 2609.39045 • Published 5 days ago • 85 • 2
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models Paper • 2609.38827 • Published 5 days ago • 57 • 3
Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents Paper • 2609.33772 • Published 8 days ago • 35 • 7
EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making Paper • 2609.38334 • Published 4 days ago • 74 • 3
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks Paper • 2609.38288 • Published 6 days ago • 134 • 3
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 5 days ago • 280 • 3
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents Paper • 2609.40325 • Published 5 days ago • 97 • 3
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 6 days ago • 539 • 3
APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants Paper • 2609.37559 • Published 6 days ago • 46 • 3
Selecting Diverse SFT Traces Improves Post-RL Generalization Paper • 2609.33780 • Published 8 days ago • 38 • 2
RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation Paper • 2609.18703 • Published 19 days ago • 55 • 3
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 10 days ago • 159 • 3
Harness-Zero: Harness Distillation via Agent-as-Harness Paper • 2609.24974 • Published 14 days ago • 37 • 3
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue Paper • 2609.26780 • Published 13 days ago • 103 • 5