Qwen3.8-27B-Uncensored-Aggressive — NVFP4
NVFP4 quant of Qwen3.8-27B-Uncensored-Aggressive (α=1.15 recipe update), compressed-tensors
NVFP4A16 (E2M1 4-bit weights, FP8-E4M3 scales @ group 16), for NVIDIA V100 (SM70) under
1Cat-vLLM. Vision tower and the grafted MTP head (bf16) are preserved.
About this update (α=1.15)
Recipe update to the Aggressive line: the previous build ablated at α≈1.24, which a larger benchmark sweep showed over-ablates past the ~1.15 quality peak. This build uses α=1.15 — more open and better on every measured axis.
Evaluation (bf16 parent, larger-sample, thinking mode, Claude-judged)
| openness ↑ | confab ↓ | factual ↑ | gsm8k ↑ | |
|---|---|---|---|---|
| stock base (censored) | 0.08 | 0.75 | 1.00 | 0.85 |
| Aggressive (α=1.15) | 0.88 | 0.725 | 1.00 | 0.85 |
| previous Aggressive (α≈1.24) | 0.80 | 0.80 | 1.00 | 0.817 |
Quant: variant-C — GPTQ NVFP4A16 targeting [Linear], GDN in_proj_qkv/in_proj_z kept fp16,
768 CoT + 256 wiki calibration, actorder=weight, MSE observer. MTP head carried verbatim in bf16.
Serve (1Cat-vLLM, 2× V100, TP2)
--kv-cache-dtype fp8_e5m2, MTP speculative decoding, --gpu-memory-utilization tuned to KV/context budget.
Serves on V100/SM70 via 1Cat prepare_nvfp4_linear (min capability 70; not the modelopt path).
Note
Uncensored / de-refused. Use responsibly and in compliance with applicable law.
- Downloads last month
- 144
Model tree for philbert440/Qwen3.8-27B-Uncensored-Aggressive-NVFP4
Base model
Qwen/Qwen3.8-27B