Instructions to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0 # Run inference directly in the terminal: ./llama-cli -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Use Docker
docker model run hf.co/shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
- LM Studio
- Jan
- vLLM
How to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
- Ollama
How to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with Ollama:
ollama run hf.co/shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
- Unsloth Desktop
- Pi
How to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with Docker Model Runner:
docker model run hf.co/shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
- Lemonade
How to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Run and chat with the model
lemonade run user.Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF-Q2_0
List all available models
lemonade list
- Hermes Agent
How to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Bonsai 2 27B × TWIN-TURBO — Residual Transplant
Treat this model as uncensored. The transferred weights come from a Heretic / Uncensored donor model, whose safety alignment was deliberately reduced. Do not assume this model refuses harmful requests, and do not use it where a safety-aligned model is expected. Its refusal rate has not been measured, so it may be more or less compliant than the donor, but it should not be relied on as safe.
What this release is for. The goal of this experiment is to move Bonsai 2 toward the donor model, and it is measured only by that yardstick: distance to TWIN-TURBO (KL divergence), plus small sanity probes. It has not been compared on standard benchmarks or refusal tests against the official Bonsai 2 or other Bonsai 2 derivatives (e.g. abliterated / uncensored builds). Whether it is better than any of them for a given use is unknown. It is closer to TWIN-TURBO than the official model, but it is still far from a reproduction of TWIN-TURBO (see Evaluation).
Experimental ternary residual transplant from DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored ("TWIN-TURBO" below; a Heretic / Uncensored model) into the Prism ML Bonsai 2 27B ternary host.
DavidAU publishes many similarly named models. This work uses exactly the repository and revision listed under Sources.
This is not a conventional quantization of TWIN-TURBO.
Direct post-training ternarization of the source model caused severe quality collapse. Instead, this project analyzes the relationship between TWIN-TURBO, Qwen3.8, Qwen3.6 and Bonsai 2, extracts the comparatively small tuning residual, folds it into the Bonsai 2 Hadamard basis, and transfers as much of that residual as possible into the existing ternary lattice.
The released model uses activation-aware per-group scale fitting with a stochastic-rounding fallback.
Sources
| Role | Repository | Revision |
|---|---|---|
| Donor ("TWIN-TURBO") | DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored (safetensors) | 61a55dc6615945ba4cdc32e212dfe966fc42b1c8 |
| Ternary host | prism-ml/Ternary-Bonsai-2-27B-gguf (F16 + PQ2_0) | b072e1d3b35a0a630cece372c2127528e0994386 |
| Blend component (residual only) | Qwen/Qwen3.8-27B | 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 |
| Blend component (residual only) | Qwen/Qwen3.6-27B | 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9 |
Released file
Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0.gguf
- Architecture: Qwen3.8 hybrid attention (GGUF arch
qwen35: Gated DeltaNet + full attention) - Host: Prism ML Ternary Bonsai 2 27B
- Packing: PQ2_0, same tensor-type layout as the official PQ2_0 file (402 PQ2_0, 96 BF16, 353 F32 tensors)
- Size: 7.2 GB (7,206,169,152 bytes)
- Ternary representation:
{-1, 0, +1}with FP16 group scales, group size 128 - Chat template: the official Bonsai 2 template (TWIN-TURBO's own multi-mode template is not included)
- Text-only: MTP and vision (mmproj) are not included
Runtime requirement
Use the PrismML llama.cpp fork (PrismML-Eng/llama.cpp, branch prism).
PQ2_0 is a Prism-specific packing, and Bonsai 2 relies on runtime Hadamard activation transforms described by prism.hadamard.* metadata. At the time of writing these live in Prism ML's forks. Upstream llama.cpp and apps built on it (LM Studio, Ollama, and the snippets under "Use this model") were not verified and may fail to load the file or produce invalid output.
Tested only with that fork at commit adfffbe41, CUDA, on an RTX 5090 (Windows). Other runtimes and apps are untested.
Usage
llama-server -m Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0.gguf -ngl 99 -c 32768 --jinja --reasoning-format deepseek --temp 1.0 --top-k 20 --top-p 0.95 --chat-template-kwargs '{"reasoning_effort":"medium"}'
- Sampling follows the Bonsai 2 recommendation (temp 1.0, top-k 20, top-p 0.95).
reasoning_effort: mediumis suggested. With the template default (xhigh), our probes below saw thinking exhaust a 6144-token budget on 5.6% of borderline prompts, and other Bonsai 2 derivatives report the same failure mode. We did not measuremedium.- Allow a generous output budget in thinking mode (Prism ML recommends 16k+ tokens).
- For creative long-form writing,
--presence-penalty 1.0 --dry-multiplier 0.8may help against repetition loops. The effect in our small prose probe was minor (see Prose stability).
What this model actually contains
The source TWIN-TURBO model is not simply Qwen3.8 plus a small fine-tune.
Across all 497 language-model weight matrices, a per-tensor least-squares fit gives:
TWIN-TURBO ≈ a × Qwen3.8 + b × Qwen3.6 + R
a = 0.542 ± 0.007 (range 0.529–0.562)
b = 0.489 ± 0.006 (range 0.475–0.508)
|R| / |W|: median 1.85%, max 3.9%
a + b > 1 for every matrix (min 1.004), which is consistent with a SLERP-style merge. The residual R is much smaller than the full delta to Qwen3.8 (~20–37% of |W|) and is strongly low-rank in sampled layers (rank-16 holds ~80–94% of its energy for most matrix types).
This model starts from the official Bonsai 2 ternary host (itself a ternary Qwen3.8) and transfers only R. It does not attempt to reconstruct the Qwen3.8/Qwen3.6 blend.
Therefore:
This model is not a reproduction of TWIN-TURBO.
It is better described as:
Bonsai 2 + a measured transplant of TWIN-TURBO's Heretic / Uncensored tuning residual.
R is everything TWIN-TURBO adds beyond the best linear blend of Qwen3.8 and Qwen3.6. That is DavidAU's tuning, and with it the Heretic (refusal-removal) edits of the donor's lineage: such edits are low-rank and lie outside the span of the two base models. How much of the donor's de-censoring survives the transplant has not been quantified, so the model should be treated as uncensored (see the warning above).
Method
TWIN-TURBO, Qwen3.8, Qwen3.6 (HF safetensors)
│
├─ per-tensor fit TWIN-TURBO ≈ a·Qwen3.8 + b·Qwen3.6
│
└─ residual R (as an HF checkpoint)
│
v
Prism converter, Hadamard-manifest-aware (grouped GDN V layout)
│
v
fold R into the Bonsai 2 basis (official signs + block-1024 FWHT)
│
v
official Bonsai 2 F16 (exactly trit × scale)
│
├─ ternary tensors: activation-aware scale least squares,
│ stochastic-rounding fallback
├─ high-precision matrices: R added exactly
└─ 1D tensors (norms, SSM params): unchanged
│
v
Prism llama-quantize with the official tensor-type plan → released model
- Scale least squares. For each ternary tensor the ternary codes stay fixed and only the free FP16 per-group scales are fitted, minimizing the output error
||(Q − (B + R)) X||²under real activationsX. Activations were dumped from the official Bonsai 2 at the input of each weight's matmul (already in the folded basis) on a ~49k-token calibration text. The ridge strength is chosen per tensor on held-out activations, and the result is accepted only if its held-out error is below 0.97× that of leaving the tensor unchanged. Accepted: 350 of 402 ternary tensors. - Stochastic-rounding fallback. The other 52 tensors (21
ffn_down, 12ffn_up, 9ffn_gate, 8ssm_out, plusoutputandtoken_embd, which had no usable activations) receive unbiased stochastic rounding ofR / scaleonto the code lattice. - High-precision matrices. For the 144 matrices kept above ternary precision (BF16
ssm_alpha/ssm_beta,ssm_conv1d),Ris added exactly. - Packing check. Re-packing the official F16 with the same plan reproduces all 402 official PQ2_0 tensors byte-for-byte, so the host starts exactly from the official model. About 3.25% of the file's bytes differ from the official PQ2_0.
The combination performed better than either method alone (see Evaluation).
Evaluation
The reference, called TWIN-TURBO Q8_0 throughout this card, is the donor revision above converted by us from its safetensors to a plain (non-Hadamard) Q8_0 GGUF with the same llama.cpp fork, text-only, without MTP. It is not one of DavidAU's own GGUF releases. Results use the WikiText-2 test split (8 × 4096-token chunks, llama-perplexity --kl-divergence), which is independent of the activation-fitting text. KL is KL(TWIN-TURBO ‖ model).
| Model | PPL | KL vs TWIN-TURBO | Same top-1 |
|---|---|---|---|
| TWIN-TURBO Q8_0 (reference) | 5.600 | — | — |
| Official Bonsai 2 PQ2_0 | 8.150 | 0.4460 | 74.1% |
| Stochastic rounding only | 7.876 | 0.3985 | 74.9% |
| Scale LS only (rejected tensors left official) | 7.796 | 0.3954 | 75.2% |
| This release: scale LS + stochastic-rounding fallback | 7.672 | 0.3801 | 75.5% |
| Ceiling: R added without rounding (off-lattice, Q8_0, not a release candidate) | 7.759 | 0.3687 | 75.3% |
This release captures about 85% of the KL reduction that is achievable by transferring R at all. Its PPL being slightly below the ceiling is within measurement noise.
The transplant moves the ternary model measurably toward TWIN-TURBO. Most of the remaining distance (KL ≈ 0.37 even for the ceiling) presumably comes from what R does not contain: the Qwen3.6 half of the blend and the ternarization itself.
Note that PPL on WikiText is not a measure of overall model quality. TWIN-TURBO's own PPL here (5.60) is far below all ternary variants.
Behavioral observations
These probes are small. Bonsai-recommended sampling, 3 seeds, thinking on at each model's template-default reasoning effort, 6144-token output budget.
| TWIN-TURBO Q8_0 | Official Bonsai 2 | This release | |
|---|---|---|---|
| Simple reasoning prompts correct (12 × 3) | 36/36 | 36/36 | 36/36 |
| Borderline prompts: median thinking tokens (12 × 3) | 365 | 446 | 395 |
| Borderline prompts: mean thinking tokens | 768 | 1182 | 1419 |
| Borderline prompts: no answer (budget exhausted) | 0% | 5.6% | 5.6% |
- No reasoning regression was found on this small set.
- The borderline-prompt results are mixed. This release answers some edgy writing prompts faster than official Bonsai (e.g. a phishing-awareness example or dark jokes), but deliberates longer on others. On hazard-information prompts it keeps the host's long deliberation: TWIN-TURBO answers the chemical-mixing and hotwiring prompts in a few hundred thinking tokens, while official Bonsai 2 and this release take thousands. On the poisonous-plants prompt all three deliberate long.
- A likely explanation is that this behavior lives in the parts of TWIN-TURBO that
Rdoes not carry (the Qwen3.6 blend) and/or in the ternary host itself. This is a hypothesis, not a measured result. - An earlier build using stochastic rounding alone had a higher no-answer rate (8.3%). Removing most of the rounding noise brought it back to the official level.
Prose stability
A small long-form probe: 6 prompts (3 Japanese, 3 English), 2 seeds, thinking off, temp 1.0 / top-k 20 / top-p 0.95, no penalties.
| Model | Mean repeated 10-char n-grams | Loops | Hit length limit |
|---|---|---|---|
| TWIN-TURBO Q8_0 | 0.043 | 1 / 12 | 2 / 12 |
| Official Bonsai 2 | 0.049 | 1 / 12 | 0 / 12 |
| This release | 0.020 | 0 / 12 | 0 / 12 |
The loops that did occur, on a Japanese short-story prompt, appeared in TWIN-TURBO Q8_0 and in official Bonsai 2 alike. The transplant did not obviously worsen repetition. With n = 12 this means "not worse", not "better".
Overall prose quality is still below the full-precision donor. We attribute this to the ternary host (WikiText PPL 7.67 vs 5.60), which other Bonsai 2 derivatives share.
Negative results
Several approaches were tested and deliberately not released:
- One-shot PTQ ternarization of TWIN-TURBO collapsed quality (calibration PPL ~80–304, vs ~2.9 for official Bonsai 2 on the same text).
- Stochastic rounding of R alone works (KL 0.3985), but its per-tensor noise in output space is 4–72× the edit's own energy.
- Over-driving the residual (1.5× R) added more noise than signal.
- GPTQ-style error feedback on the lattice mostly declined to edit.
- Splicing TWIN-TURBO's
lm_head/ embedding (Q8, correctly folded) made the model worse (KL 0.380 → 0.596). Bonsai's hidden states are co-adapted with its own head. - A least-squares
lm_headbridge fitted to TWIN-TURBO logits improved in-domain but never beat the ternary head on WikiText, even when trained with WikiText data. Logit MSE is misaligned with KL. - Transferring the full task vector (TWIN-TURBO − Qwen3.8, blend included) through scale fitting reached KL 0.397, worse than transferring R alone.
Research conclusion
Local tensor-space edits can transfer a small, coherent tuning residual into Bonsai 2. Our experiments suggest they cannot reconstruct a substantially different dense model: layer-local objectives stopped tracking end-to-end behavior whenever the target was far from the QAT-co-adapted ternary host.
A natural next step would be end-to-end KL distillation from TWIN-TURBO that trains only the per-group FP16 scales and normalization parameters (~0.2B parameters) while keeping the ternary codes fixed.
Limitations
- Research release. Not benchmarked on standard suites (MMLU, HumanEval, etc.), long context, tool calling or multilingual tasks beyond the small probes above.
- Uncensored lineage. The donor is a Heretic / Uncensored model. Treat this release as uncensored: it may comply with requests that the official Bonsai 2 or other safety-aligned models refuse. Its refusal rate on a harmful-prompt benchmark has not been measured. Users are responsible for appropriate and lawful use.
- Prose quality and perplexity are below the full-precision donor (TWIN-TURBO Q8_0). Qwen3.8 itself was not measured.
- Thinking can exhaust the output budget on some prompts (see Usage).
- Text-only; MTP excluded.
- The build pipeline and calibration text are not published yet.
License and attribution
Apache License 2.0. See LICENSE and NOTICE.txt.
Created using Bonsai by Prism ML.
Derived from:
- Prism ML — Ternary Bonsai 2 27B
- DavidAU — Qwen3.8-27B TWIN-TURBO Fable Cold Fusion 709 ULTRA HERETIC Uncensored
- Qwen / Alibaba Cloud — Qwen3.8-27B and Qwen3.6-27B (used for the residual)
Acknowledgements
Thanks to Prism ML for publishing Bonsai 2, its runtime and Hadamard contract, to DavidAU for the TWIN-TURBO source model, and to Hikari07jp, whose code-space abliteration of Bonsai 2 showed that the official ternary codes can be edited in place.
This is an independent research experiment. It is not an official Prism ML, Qwen or DavidAU release.
- Downloads last month
- 484
2-bit
docker model run hf.co/shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0