GSQ-RCO quant for ThinkingCap-Qwen3.8-27B?

#1
by Exo87 - opened

Hi bottlecapai,

Could you release a GSQ-RCO quant of ThinkingCap-Qwen3.8-27B, or coordinate with ISTA-DASLab to do it? Their GSQ-RCO method keeps near-BF16 quality at ~2–3 bits, cutting the model to ~1/5 size. For low-VRAM users, that’s the difference between needing 60GB+ / multi-GPU and running on 12–16GB consumer GPUs. Combined with ThinkingCap’s ~46% fewer thinking tokens, it’d be much faster and locally usable. Your model already quantizes well, so it’s a natural fit.

Thanks!

BottleCapAI org
•
edited 10 days ago

Hi @Exo87 ,
it would be great to combine. I will investigate it but no promises 😇

Out of curiosity, what is your use case? How you used models from ISTA-DASLab ?

Aren't you limited by slow response time anyway despite ternary quantizations?

Hi @oplatek , thanks for looking into it!

My setup: RTX 3060 12GB VRAM. So yes, I'm definitely VRAM-limited, but the bigger problem is actually speed and token burn.

Use case: I take existing GitHub projects and modify them to fit my specific needs. The base Qwen3.8-27B thinking model is the issue, it burns so many thinking tokens per task that it either runs out of context before finishing the project, or takes forever. So I'm stuck choosing between slow-and-incomplete local, or fast-and-complete but cloud (DeepSeek), which I'd rather avoid.

On ISTA-DASLab: I haven't run their quants myself yet, I found their GSQ-RCO work and it looked like the missing piece. Ternary/low-bit quants usually hurt quality, but their results suggest you can get the VRAM savings without the accuracy hit, which is exactly what a 12GB card needs.

Why your model matters here: ThinkingCap's ~46% token reduction directly attacks my real bottleneck (thinking length / context exhaustion), and GSQ-RCO attacks the VRAM side. Together they'd be the first combo that actually fits my workflow locally, smaller and smarter about tokens. That's why I'm pushing for it.

No pressure on the promise, just wanted to plant the flag. Thanks for the great work either way 🙏

@oplatek any chance this in qwen flash? 🫥

Sign up or log in to comment