-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 110 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
Cinny
cinnybun02
AI & ML interests
None yet
Recent Activity
new activity about 18 hours ago
zai-org/GLM-5.3-Flash:Could've kept the active params 9b or below. new activity about 18 hours ago
zai-org/GLM-5.3-Flash:No 4bit QAT like deepseek v4 flash?????Organizations
None yet