Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 7 days ago • 128
A Latent Variable Framework for Scaling Laws in Large Language Models Paper • 2512.06553 • Published Jun 2
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Paper • 2606.18089 • Published Jul 5
SPRI: Aligning Large Language Models with Context-Situated Principles Paper • 2502.03397 • Published Feb 5, 2025 • 1
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families Paper • 2412.06540 • Published Dec 9, 2024 • 2
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence Paper • 2502.09927 • Published Feb 14, 2025 • 1
Out-of-Distribution Detection using Synthetic Data Generation Paper • 2502.03323 • Published Feb 5, 2025
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content Paper • 2410.10783 • Published Oct 14, 2024 • 26
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead Paper • 2407.00066 • Published Jun 17, 2024
Distributional Preference Alignment of LLMs via Optimal Transport Paper • 2406.05882 • Published Jun 9, 2024 • 2
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content Paper • 2410.10783 • Published Oct 14, 2024 • 26
Cluster & Tune: Boost Cold Start Performance in Text Classification Paper • 2203.10581 • Published Mar 20, 2022