Instructions to use ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-Flash-Next for mlx-serve, iQ-MLX 3.3 bpw (64 GB Macs)
mlx-serve pack of Qwen/Qwen3.8-Flash-Next,
the Qwen4 preview architecture (model_type: qwen4_exp), quantized to 3.3 bits per
weight with imatrix calibration so it fits a 64 GB Mac with room for 64k of context.
52 GB resident. Includes the MTP head and the vision tower (image and video input).
Download MLXServe.com
sudo sysctl iogpu.wired_limit_mb=58000 # 64 GB Macs, once per boot; see below
mlx-serve --model ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw --serve --ctx-size 65536 --skip-mem-preflight
(launching mlx-serve via UI is also fine, make sure to enable --skip-mem-preflight via settings)
If you have 128 GB, use ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit instead. It is 75 GB resident and a few points better.
Quality
Measured against the bf16 checkpoint on 1200 held-out next-token positions (agent, code, prose, math; documents the calibration never saw). Top-1 is how often the pack's greedy token equals bf16's; mass is the bf16 probability of the pack's pick, with bf16's own pick as the ceiling.
| pack | resident | top-1 vs bf16 | mass (bf16 ceiling 0.762) |
|---|---|---|---|
| mixed-4-8bit | 73 GB | 89.1% | 0.738 |
| iQ-MLX 3.3 bpw (this pack) | 52 GB | 85.6% | 0.732 |
Same positions, same server, cold prefills. The 3.3 bpw pack keeps 99% of the 4-8 pack's probability mass at 21 GB less.
Variants tried and rejected: more expert bytes at the cost of a 6-bit spine (worse), forcing the down projections to 3-bit (a wash), 3.6x more calibration data at 44 GB (no change in measured error; bytes moved the needle, calibration volume did not).
Memory on a 64 GB Mac
macOS caps what the GPU may wire at about 75% of RAM (48 GB on a 64 GB Mac), which is not enough. Raise it once per boot:
sudo sysctl iogpu.wired_limit_mb=58000
With that, 52 GB of weights plus a 64k KV cache (1.6 GB) plus the prefill working set
fit. Measured on this pack: 52.2 GB active after load, 55.6 GB peak during a 36k-token
prefill. Keep --ctx-size at 65536; --kv-quant 8 halves the cache if you need more
headroom. The 32 GB n-gram table is not in that number (next section).
What is different about this model
This is not a Qwen3.5-style pack. Three things around the usual GDN + MoE trunk:
- Gated residual streams. The residual is 4 streams wide (4 x 2560). Every block reads a sigmoid-mixed average of the normalized streams and writes back through per-stream scalar gates. The final mixer replaces the usual final norm.
- N-gram embedding (51B parameters). A second embedding table indexed by hashed bigrams and trigrams of the token ids: 16 heads, each a prime-sized bucket space of ~20M rows, 160 dims per row, injected once before layer 1. It is a lookup, no compute, which is why Qwen quotes the model as 125B: the full checkpoint is 125B trunk + 51B n-gram + 4B MTP = 180B (360 GB bf16).
- Qwen Sparse Attention. Past 2048 tokens each attention layer only reads the 512 most relevant 4-token blocks per query (picked by a small indexer), plus the query's own partial block. Attention cost stays flat with context. Native 262k context.
How this pack stores the n-gram table
The 51B table is NOT in the safetensors shards. It is one merged 4-bit table
in ngram_table.bin (32.0 GB, safetensors format, .bin so nothing
mlx-loads it). mlx-serve mmaps the file and, per token, dequantizes the 16 rows
it needs on the CPU (16 x 80 bytes) and uploads only the resulting 2560-vector.
The table never becomes resident: its cost is page cache, which the OS evicts
as needed. On a 64 GB Mac that page cache competes with everything else, so the
first long prompt after boot pays random SSD reads (a few seconds per 8k tokens on
a cold cache); warm is free.
Widths
| tensors | width |
|---|---|
| routed experts (512 x 48 layers, the 121B) | per-layer, imatrix-measured allocation on layers 0-45: 3-bit group 128 on 67 gate-up/down groups, 2-bit group 128 on 23, 2-bit group 64 on 2; 4-bit group 64 on layers 46-47 and the MTP head. 3.14 bpw over the experts |
| attention, GDN, hyper-connections, indexer, shared experts | 8-bit, group 64 |
| lm_head | 8-bit, group 64 |
| embed_tokens | 4-bit, group 64 |
| n-gram table | 4-bit, group 32 (row width 160) |
| routers, inject gates, norms, convs, SSM state | bf16 |
| MTP head | experts 4-bit group 64, projections 8-bit group 64 |
Every quantized weight with an imatrix entry uses an activation-weighted (scale, bias)
search instead of min-max, byte-compatible with mx.quantize's affine layout, so the
engine reads it like any other affine pack. Every (1 + w) RMSNorm has the +1
folded into the stored weight; depthwise convs are transposed to MLX's [C, K, 1];
experts.gate_up_proj is split into switch_mlp.gate_proj / up_proj. The vision
tower ships dense bf16 in model-vision.safetensors (~0.9 GB).
Serving notes
- Speed. M4 Max, llmprobe
--bench-only --rungs 2k: 52.6 tok/s decode, 762 tok/s prefill, 318 ms TTFT; the mixed-4-8bit pack measured 55.5 / 754 / 278 ms in the same run. 3-bit experts ride MLX's stock gather kernel, 2-bit ones mlx-serve's fused ones. - MTP. The checkpoint's own 1-layer speculative head is in the pack (
--mtpor per-request"enable_mtp": true). Default-off; not re-measured on this pack. - Thinking is on by default (
"enable_thinking": falseturns it off). Tools use Qwen3.8's XML call format; mlx-serve parses and schema-coerces it. - Images and video go through the Qwen3-VL-style tower (
model.visual.*, dense bf16). MTP is declined on image turns (serial decode).
Calibration and conversion
All in the mlx-serve repo, tests/:
qwen38_flash_next_imatrix_collect.pycollects per-input-channel activation statistics on the exact bf16 checkpoint (482k tokens of agent traffic, code, prose and math, rendered with the model's own chat template) by streaming one decoder layer at a time through HF transformers, so the 360 GB never has to be resident. The same script writes the held-out bf16 logits the quality table is scored on.qwen38_flash_next_iq_allocate.pymeasures the imatrix-weighted error of every layer's gate-up and down banks at 2/3/4 bits and spends a 47 GB budget where it buys the most.convert_qwen38_flash_next.py --imatrix --allocwrites the pack.qwen38_flash_next_score.pyscores any served pack against the reference.
- Downloads last month
- 3,731
4-bit
Model tree for ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-3.3bpw
Base model
Qwen/Qwen3.8-Flash-Next