Qwen3.8-27B-Uncensored-Aggressive — NVFP4

NVFP4 quant of Qwen3.8-27B-Uncensored-Aggressive (α=1.15 recipe update), compressed-tensors NVFP4A16 (E2M1 4-bit weights, FP8-E4M3 scales @ group 16), for NVIDIA V100 (SM70) under 1Cat-vLLM. Vision tower and the grafted MTP head (bf16) are preserved.

About this update (α=1.15)

Recipe update to the Aggressive line: the previous build ablated at α≈1.24, which a larger benchmark sweep showed over-ablates past the ~1.15 quality peak. This build uses α=1.15 — more open and better on every measured axis.

Evaluation (bf16 parent, larger-sample, thinking mode, Claude-judged)

openness ↑ confab ↓ factual ↑ gsm8k ↑
stock base (censored) 0.08 0.75 1.00 0.85
Aggressive (α=1.15) 0.88 0.725 1.00 0.85
previous Aggressive (α≈1.24) 0.80 0.80 1.00 0.817

Quant: variant-C — GPTQ NVFP4A16 targeting [Linear], GDN in_proj_qkv/in_proj_z kept fp16, 768 CoT + 256 wiki calibration, actorder=weight, MSE observer. MTP head carried verbatim in bf16.

Serve (1Cat-vLLM, 2× V100, TP2)

--kv-cache-dtype fp8_e5m2, MTP speculative decoding, --gpu-memory-utilization tuned to KV/context budget. Serves on V100/SM70 via 1Cat prepare_nvfp4_linear (min capability 70; not the modelopt path).

Note

Uncensored / de-refused. Use responsibly and in compliance with applicable law.

Downloads last month
144
Safetensors
Model size
28B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for philbert440/Qwen3.8-27B-Uncensored-Aggressive-NVFP4

Base model

Qwen/Qwen3.8-27B
Quantized
(5)
this model

Collection including philbert440/Qwen3.8-27B-Uncensored-Aggressive-NVFP4