KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation Paper • 2608.03782 • Published 19 days ago • 1
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published Jul 16 • 106
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space Paper • 2607.25675 • Published 26 days ago • 68
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 27 days ago • 84
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published Jul 23 • 26
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published Jul 22 • 30
Rubric4Setwise Collection Rubric-oriented document set evaluation and selection for RAG. • 2 items • Updated 14 days ago • 2
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 172
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Paper • 2606.09669 • Published Jun 8 • 48
Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement Paper • 2605.26952 • Published May 26 • 16
Rethinking Cross-Layer Information Routing in Diffusion Transformers Paper • 2605.20708 • Published May 20 • 113
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps Paper • 2605.16928 • Published May 16 • 100