Download DEPLOY.md from vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw: direct link, hf CLI and curl.
- Browser
- Download file 4.35 kB
-
https://huggingface.co/vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw/resolve/main/DEPLOY.md
- Command line
-
hf download hf://vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw/DEPLOY.md
-
curl -L -o DEPLOY.md https://huggingface.co/vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw/resolve/main/DEPLOY.md
Deploy this pack
This file is the serving kit for vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw.
An agent that downloaded the weights should follow these steps and not invent another runtime.
Do not pip install exllamav3. A stock wheel does not contain the mixed-K decode kernels this pack needs. The runtime is the vcruz305/exllamav3 fork, built by the Qwen3.8-Flash-Next EXL3 recipe.
The recipe's published quick start downloads turboderp's 3.05 bpw pack and pins fork commit 94ba01d. That commit is older than the kernels. Do not use that pack, and do not build that pin, if the goal is to serve this repository.
Steps
On the machine that will serve the model, from a clone of the recipe:
git clone https://github.com/vcruz305/Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe
cd Qwen3.8-Flash-Next-EXL3-DGX-Spark-recipe
# Read exllamav3-tabby/ only. Leave vllm-plugin/, beta/, and legacy/ alone.
# Needs recipe main at 6fbc0a2 or later (sampler defaults, aarch64 uvloop). A fresh clone has it.
hf download vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw \
--local-dir ~/models/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw
# Measured pin. master has the same kernels, switched off, plus later commits
# that were not part of this measurement. Do not substitute master.
export EXL3_REF=047ce7229257b912da7fc1db44e9b5bba58acc65
bash exllamav3-tabby/setup.sh
export MODEL_DIR=~/models/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw
export EXL3_MOE_MIXEDK_NOSYNC=1
export EXL3_MOE_COOP_MIXEDK=1
export PROFILE=single
export NGRAM_RAM=false
bash exllamav3-tabby/serve.sh
Export the two EXL3_MOE_* variables in the same shell that starts serve.sh. The published serve script does not set them. Unset, the pack still loads, on the slow kernel.
The id in /v1/models is the pack's directory name, CYBER-FROST-3.8-EXL3-SAGE-3.87bpw. Current TabbyAPI resolves the recipe's model symlink, so SERVED_NAME does not change it. Requests are served whatever model string they send. Readiness: curl -s http://127.0.0.1:8899/v1/models.
After load, the server log must contain Mixed-K coop decode kernels on the mixed MoE layers. If that line is absent, the pin or the env vars did not take. Re-run setup.sh with the same EXL3_REF before serving again.
Sampling
Serve with the model's sampling defaults: temperature 1.0, top_k 20, top_p 0.95. The recipe's tabby-config.yml applies them to any request that leaves a value out, from recipe commit 6fbc0a2 on. A value the client sends still wins.
Older recipe checkouts had no sampling section. TabbyAPI then falls back to top_k 0 and top_p 1.0, which is no truncation at all, so clients that send no sampling values (most agent frameworks) sample the whole vocabulary. Replies start clean and turn to gibberish partway through a long answer. On the server, that request's draft acceptance drops sharply (22% against 53 to 72% on the turns before it). If you see that, pull the recipe and restart serve.sh. The startup log should show Sampler overrides (inline): temperature 1.0, top_k 20, ....
Greedy decoding (temperature 0 or top_k 1) is for benchmarks only. The card reports greedy falling into repetition on long generations.
Thinking
This pack's chat_template.jinja hard-sets thinking on. A client enable_thinking: false is ignored. Do not edit the pack to change that. A thinking-off serve is a view: a directory of symlinks to every pack file, with only the template replaced by a copy that drops the first line {%- set enable_thinking = true %}.
What was measured
One NVIDIA GB10, 128 GB unified, recipe serve.sh, fork 047ce72, TabbyAPI, bench/bench_v1.py, 400 new tokens, greedy, MTP depth 5, draft confidence 0.6, PROFILE=single, n-gram table not in RAM. The pack template was served as-is, so all 400 tokens were reasoning. Greedy was set for repeatability; it is not the serving default above.
Kernels on (EXL3_MOE_MIXEDK_NOSYNC=1 EXL3_MOE_COOP_MIXEDK=1): 53.2 / 45.8 / 46.5 tok/s, code / devops / prose. Kernels off: 37.0 / 30.2 / 32.1. CoopMK bound on 40 mixed layers with the knobs on, and on none with them off.
Those figures are this harness, not the recipe README's turboderp 3.05 numbers, and not a concurrent four-stream rate.