How to use from the
Use from the
MLX library
# Make sure mlx-lm is installed
# pip install --upgrade mlx-lm

# Generate text with mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("Johneeee/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-oQ4-fa6-qkv6-tail6-e6-b8")

prompt = "Write a story about Einstein"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True
)

text = generate(model, tokenizer, prompt=prompt, verbose=True)

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU — oQ4 MLX Quant

oMLX oQ4 recipe-driven quantization of the DavidAU Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU merge (Qwen3.6-27B base). Per-tensor hard floors preserved on all 298 recipe-pinned tensors; only the base tier was downgraded to oQ4 to land a smaller file.

Quantization summary

Quant method oMLX oQ, enhanced
Base level oQ4 — base bits 4, group size 64, affine
Recipe pins 298 tensors (187 × 8-bit, 111 × 6-bit) — preserved verbatim
oQ4 self-boosts 45 extra tensors lifted to 5-bit (linear-attention, calibration-driven)
Final override map 343 tensors (187 × 8-bit, 111 × 6-bit, 45 × 5-bit)
Model size 20.16 GB on disk (5 shards)
Language model ~21.6 GB dry-run estimate at oQ6e reference; actual oQ4 build ~19.7 GB LM
Text-only vision encoder stripped
MTP stripped
Dtype float16

Recipe / pinned floors

  • FA6 — 6-bit floor on all 16 full-attention layers (self_attn + MLP)
  • QKV6 — 6-bit floor on all linear_attn.in_proj_qkv
  • Tail6 — 6-bit floor on layers 56–58
  • E6 — embed_tokens at 6-bit
  • b8 — 8-bit bump set (187 tensors, incl. lm_head)

Verified: recipe → config override map matches exactly (0 missing, 0 bit mismatches); on-disk safetensors dtypes confirm pinned tensors are packed-quantized (U32), not fp16.

Use with mlx

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-oQ4-fa6-qkv6-tail6-e6-b8")

prompt = "hello"

if tokenizer.chat_template is not None:
    messages = [{"role": "user", "content": prompt}]
    prompt = tokenizer.apply_chat_template(
        messages, add_generation_prompt=True, return_dict=False,
    )

response = generate(model, tokenizer, prompt=prompt, verbose=True)

About the base

This MLX quant is derived from the fp source Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU (Qwen3.6-27B base, Fable-Fusion-711 merge family, uncensored / abliterated, DavidAU). 64-layer hybrid architecture: 48 linear-attention + 16 full-attention layers. Original model license: apache-2.0.

Downloads last month
220
Safetensors
Model size
27B params
Tensor type
U32
·
F16
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Johneeee/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-oQ4-fa6-qkv6-tail6-e6-b8

Base model

Qwen/Qwen3.6-27B
Quantized
(742)
this model