DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 14 days ago • 81
Steering Geometry: Validating Human Value Geometry in LLM Steering Space Paper • 2609.06289 • Published 14 days ago • 31
Running Featured 789 Agent Memory Leaderboard 🧠 789 Unified memory evaluation · Results expected August 12.
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models Paper • 2608.04701 • Published Aug 5 • 9
trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 Text Generation • 2.44M • Updated 11 days ago • 13.7M • 48
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 157
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Paper • 2607.21072 • Published Jul 23 • 35
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published Jul 22 • 111
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published Jul 21 • 114