RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation Paper • 2609.16900 • Published 24 days ago • 48
RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation Paper • 2609.16900 • Published 24 days ago • 48
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding Paper • 2608.06501 • Published Aug 6 • 4
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events Paper • 2608.06485 • Published Aug 6 • 5
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding Paper • 2608.06501 • Published Aug 6 • 4
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events Paper • 2608.06485 • Published Aug 6 • 5
Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm Paper • 2607.27851 • Published Jul 30 • 3
Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm Paper • 2607.27851 • Published Jul 30 • 3
BubbleRAG: Evidence-Driven Retrieval-Augmented Generation for Black-Box Knowledge Graphs Paper • 2603.20309 • Published Mar 19 • 21
SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning Paper • 2410.17238 • Published Oct 22, 2024 • 1
BubbleRAG: Evidence-Driven Retrieval-Augmented Generation for Black-Box Knowledge Graphs Paper • 2603.20309 • Published Mar 19 • 21
SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model Paper • 2602.21818 • Published Feb 25 • 56
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines Paper • 2602.14296 • Published Feb 15 • 51
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines Paper • 2602.14296 • Published Feb 15 • 51
Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch Paper • 2512.02395 • Published Dec 2, 2025 • 52
Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model Paper • 2509.04548 • Published Sep 4, 2025 • 6
ALPHA: AnomaLous Physiological Health Assessment Using Large Language Models Paper • 2311.12524 • Published Nov 21, 2023 • 1
REPAIR: Robust Editing via Progressive Adaptive Intervention and Reintegration Paper • 2510.01879 • Published Oct 2, 2025 • 8
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Paper • 2505.24120 • Published May 30, 2025 • 50