I Have a Stream: Making Self-Supervised Learning Work on Continuous Video Paper • 2609.40333 • Published 11 days ago • 15
I Have a Stream: Making Self-Supervised Learning Work on Continuous Video Paper • 2609.40333 • Published 11 days ago • 15
Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning Paper • 2609.31199 • Published 16 days ago • 14
view post Post 5991 who's working on an NVFP4 version of Kimi-K3? See translation 4 replies · 🤗 11 11 👍 6 6 🚀 4 4 🔥 4 4 😔 1 1 + Reply
Self-Supervised Learning of Structured Dynamics from Videos Paper • 2607.21576 • Published Jul 23 • 21
Self-Supervised Learning of Structured Dynamics from Videos Paper • 2607.21576 • Published Jul 23 • 21
DataComp-VLM: Improved Open Datasets for Vision-Language Models Paper • 2606.28551 • Published Jun 26 • 52
PEEK: Picking Essential frames via Efficient Knowledge distillation Paper • 2605.31029 • Published May 29 • 18
Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models Paper • 2602.10382 • Published Feb 12 • 2
Language-Switching Triggers Take a Latent Detour Through Language Models Paper • 2605.18646 • Published May 18 • 4
TAPS: Task Aware Proposal Distributions for Speculative Sampling Paper • 2603.27027 • Published Mar 27 • 146
TAPS: Task Aware Proposal Distributions for Speculative Sampling Paper • 2603.27027 • Published Mar 27 • 146
Conditioned Prompt-Optimization for Continual Deepfake Detection Paper • 2407.21554 • Published Jul 31, 2024 • 1
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models Paper • 2603.19466 • Published Mar 19 • 41
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models Paper • 2603.19466 • Published Mar 19 • 41
Visual Memory Injection Attacks for Multi-Turn Conversations Paper • 2602.15927 • Published Feb 17 • 3
Visual Memory Injection Attacks for Multi-Turn Conversations Paper • 2602.15927 • Published Feb 17 • 3