StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 7 days ago • 432
OpenFineTrain Collection Very fine 55 trillion+ tokens worth of ungated huggingface datasets for LLM training across 80+ datasets. • 87 items • Updated 3 days ago • 2
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants Paper • 2607.26611 • Published 24 days ago • 32
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models Paper • 2606.18142 • Published Jun 17 • 2
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning Paper • 2605.16301 • Published Jun 3 • 1
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models Paper • 2607.01774 • Published Jul 20 • 39
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning Paper • 2607.17599 • Published Jul 20 • 2
Trajectory-aware Cross-view Geo-localization with Sequential Observations Paper • 2607.15491 • Published Jul 16 • 8
Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training Paper • 2607.19058 • Published Jul 21 • 7
Delineate Anything v2: A Global Foundation Model for Field Delineation Paper • 2607.19069 • Published Jul 21 • 6
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Paper • 2607.19344 • Published Jul 21 • 4
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration Paper • 2607.18529 • Published Jul 20 • 5