--- license: apache-2.0 library_name: llama.cpp pipeline_tag: text-generation language: - en base_model: - prism-ml/Ternary-Bonsai-2-27B-gguf - DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored base_model_relation: merge tags: - gguf - llama-cpp - ternary - 2-bit - pq2_0 - qwen3.8 - bonsai - prismml - experimental - residual-transplant - uncensored - heretic --- # Bonsai 2 27B × TWIN-TURBO — Residual Transplant > [!WARNING] > **Treat this model as uncensored.** The transferred weights come from a Heretic / Uncensored donor model, whose safety alignment was deliberately reduced. Do not assume this model refuses harmful requests, and do not use it where a safety-aligned model is expected. Its refusal rate has not been measured, so it may be more or less compliant than the donor, but it should not be relied on as safe. > [!NOTE] > **What this release is for.** The goal of this experiment is to move Bonsai 2 *toward the donor model*, and it is measured only by that yardstick: distance to TWIN-TURBO (KL divergence), plus small sanity probes. It has **not** been compared on standard benchmarks or refusal tests against the official Bonsai 2 or other Bonsai 2 derivatives (e.g. abliterated / uncensored builds). Whether it is better than any of them for a given use is unknown. It is closer to TWIN-TURBO than the official model, but it is still far from a reproduction of TWIN-TURBO (see [Evaluation](#evaluation)). Experimental **ternary residual transplant** from [DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored](https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored) ("TWIN-TURBO" below; a Heretic / Uncensored model) into the Prism ML Bonsai 2 27B ternary host. DavidAU publishes many similarly named models. This work uses exactly the repository and revision listed under [Sources](#sources). This is **not** a conventional quantization of TWIN-TURBO. Direct post-training ternarization of the source model caused severe quality collapse. Instead, this project analyzes the relationship between TWIN-TURBO, Qwen3.8, Qwen3.6 and Bonsai 2, extracts the comparatively small tuning residual, folds it into the Bonsai 2 Hadamard basis, and transfers as much of that residual as possible into the existing ternary lattice. The released model uses activation-aware per-group scale fitting with a stochastic-rounding fallback. ## Sources | Role | Repository | Revision | |---|---|---| | Donor ("TWIN-TURBO") | [DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored](https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored) (safetensors) | `61a55dc6615945ba4cdc32e212dfe966fc42b1c8` | | Ternary host | [prism-ml/Ternary-Bonsai-2-27B-gguf](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) (F16 + PQ2_0) | `b072e1d3b35a0a630cece372c2127528e0994386` | | Blend component (residual only) | [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) | `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0` | | Blend component (residual only) | [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) | `6a9e13bd6fc8f0983b9b99948120bc37f49c13e9` | ## Released file `Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0.gguf` - Architecture: Qwen3.8 hybrid attention (GGUF arch `qwen35`: Gated DeltaNet + full attention) - Host: Prism ML Ternary Bonsai 2 27B - Packing: PQ2_0, same tensor-type layout as the official PQ2_0 file (402 PQ2_0, 96 BF16, 353 F32 tensors) - Size: 7.2 GB (7,206,169,152 bytes) - Ternary representation: `{-1, 0, +1}` with FP16 group scales, group size 128 - Chat template: the official Bonsai 2 template (TWIN-TURBO's own multi-mode template is **not** included) - Text-only: MTP and vision (mmproj) are not included ## Runtime requirement **Use the PrismML llama.cpp fork** (`PrismML-Eng/llama.cpp`, branch `prism`). PQ2_0 is a Prism-specific packing, and Bonsai 2 relies on runtime Hadamard activation transforms described by `prism.hadamard.*` metadata. At the time of writing these live in Prism ML's forks. Upstream llama.cpp and apps built on it (LM Studio, Ollama, and the snippets under "Use this model") were not verified and may fail to load the file or produce invalid output. Tested only with that fork at commit `adfffbe41`, CUDA, on an RTX 5090 (Windows). Other runtimes and apps are untested. ## Usage ```bash llama-server -m Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0.gguf -ngl 99 -c 32768 --jinja --reasoning-format deepseek --temp 1.0 --top-k 20 --top-p 0.95 --chat-template-kwargs '{"reasoning_effort":"medium"}' ``` - Sampling follows the Bonsai 2 recommendation (temp 1.0, top-k 20, top-p 0.95). - `reasoning_effort: medium` is suggested. With the template default (`xhigh`), our probes below saw thinking exhaust a 6144-token budget on 5.6% of borderline prompts, and other Bonsai 2 derivatives report the same failure mode. We did not measure `medium`. - Allow a generous output budget in thinking mode (Prism ML recommends 16k+ tokens). - For creative long-form writing, `--presence-penalty 1.0 --dry-multiplier 0.8` may help against repetition loops. The effect in our small prose probe was minor (see Prose stability). ## What this model actually contains The source TWIN-TURBO model is not simply Qwen3.8 plus a small fine-tune. Across all 497 language-model weight matrices, a per-tensor least-squares fit gives: ```text TWIN-TURBO ≈ a × Qwen3.8 + b × Qwen3.6 + R a = 0.542 ± 0.007 (range 0.529–0.562) b = 0.489 ± 0.006 (range 0.475–0.508) |R| / |W|: median 1.85%, max 3.9% ``` `a + b > 1` for every matrix (min 1.004), which is consistent with a SLERP-style merge. The residual `R` is much smaller than the full delta to Qwen3.8 (~20–37% of `|W|`) and is strongly low-rank in sampled layers (rank-16 holds ~80–94% of its energy for most matrix types). This model starts from the official Bonsai 2 ternary host (itself a ternary Qwen3.8) and transfers only `R`. It does not attempt to reconstruct the Qwen3.8/Qwen3.6 blend. Therefore: **This model is not a reproduction of TWIN-TURBO.** It is better described as: > Bonsai 2 + a measured transplant of TWIN-TURBO's Heretic / Uncensored tuning residual. `R` is everything TWIN-TURBO adds beyond the best linear blend of Qwen3.8 and Qwen3.6. That is DavidAU's tuning, and with it the Heretic (refusal-removal) edits of the donor's lineage: such edits are low-rank and lie outside the span of the two base models. How much of the donor's de-censoring survives the transplant has not been quantified, so the model should be treated as uncensored (see the warning above). ## Method ```text TWIN-TURBO, Qwen3.8, Qwen3.6 (HF safetensors) │ ├─ per-tensor fit TWIN-TURBO ≈ a·Qwen3.8 + b·Qwen3.6 │ └─ residual R (as an HF checkpoint) │ v Prism converter, Hadamard-manifest-aware (grouped GDN V layout) │ v fold R into the Bonsai 2 basis (official signs + block-1024 FWHT) │ v official Bonsai 2 F16 (exactly trit × scale) │ ├─ ternary tensors: activation-aware scale least squares, │ stochastic-rounding fallback ├─ high-precision matrices: R added exactly └─ 1D tensors (norms, SSM params): unchanged │ v Prism llama-quantize with the official tensor-type plan → released model ``` - **Scale least squares.** For each ternary tensor the ternary codes stay fixed and only the free FP16 per-group scales are fitted, minimizing the output error `||(Q − (B + R)) X||²` under real activations `X`. Activations were dumped from the official Bonsai 2 at the input of each weight's matmul (already in the folded basis) on a ~49k-token calibration text. The ridge strength is chosen per tensor on held-out activations, and the result is accepted only if its held-out error is below 0.97× that of leaving the tensor unchanged. Accepted: 350 of 402 ternary tensors. - **Stochastic-rounding fallback.** The other 52 tensors (21 `ffn_down`, 12 `ffn_up`, 9 `ffn_gate`, 8 `ssm_out`, plus `output` and `token_embd`, which had no usable activations) receive unbiased stochastic rounding of `R / scale` onto the code lattice. - **High-precision matrices.** For the 144 matrices kept above ternary precision (BF16 `ssm_alpha` / `ssm_beta`, `ssm_conv1d`), `R` is added exactly. - **Packing check.** Re-packing the official F16 with the same plan reproduces all 402 official PQ2_0 tensors byte-for-byte, so the host starts exactly from the official model. About 3.25% of the file's bytes differ from the official PQ2_0. The combination performed better than either method alone (see Evaluation). ## Evaluation The reference, called **TWIN-TURBO Q8_0** throughout this card, is the donor revision above converted by us from its safetensors to a plain (non-Hadamard) Q8_0 GGUF with the same llama.cpp fork, text-only, without MTP. It is not one of DavidAU's own GGUF releases. Results use the WikiText-2 **test** split (8 × 4096-token chunks, `llama-perplexity --kl-divergence`), which is independent of the activation-fitting text. KL is KL(TWIN-TURBO ‖ model). | Model | PPL | KL vs TWIN-TURBO | Same top-1 | |---|---:|---:|---:| | TWIN-TURBO Q8_0 (reference) | 5.600 | — | — | | Official Bonsai 2 PQ2_0 | 8.150 | 0.4460 | 74.1% | | Stochastic rounding only | 7.876 | 0.3985 | 74.9% | | Scale LS only (rejected tensors left official) | 7.796 | 0.3954 | 75.2% | | **This release: scale LS + stochastic-rounding fallback** | **7.672** | **0.3801** | **75.5%** | | Ceiling: R added without rounding (off-lattice, Q8_0, not a release candidate) | 7.759 | 0.3687 | 75.3% | This release captures about 85% of the KL reduction that is achievable by transferring `R` at all. Its PPL being slightly below the ceiling is within measurement noise. The transplant moves the ternary model measurably toward TWIN-TURBO. Most of the remaining distance (KL ≈ 0.37 even for the ceiling) presumably comes from what `R` does not contain: the Qwen3.6 half of the blend and the ternarization itself. Note that PPL on WikiText is not a measure of overall model quality. TWIN-TURBO's own PPL here (5.60) is far below all ternary variants. ## Behavioral observations These probes are small. Bonsai-recommended sampling, 3 seeds, thinking on at each model's template-default reasoning effort, 6144-token output budget. | | TWIN-TURBO Q8_0 | Official Bonsai 2 | This release | |---|---:|---:|---:| | Simple reasoning prompts correct (12 × 3) | 36/36 | 36/36 | 36/36 | | Borderline prompts: median thinking tokens (12 × 3) | 365 | 446 | 395 | | Borderline prompts: mean thinking tokens | 768 | 1182 | 1419 | | Borderline prompts: no answer (budget exhausted) | 0% | 5.6% | 5.6% | - No reasoning regression was found on this small set. - The borderline-prompt results are **mixed**. This release answers some edgy writing prompts faster than official Bonsai (e.g. a phishing-awareness example or dark jokes), but deliberates longer on others. On hazard-information prompts it keeps the host's long deliberation: TWIN-TURBO answers the chemical-mixing and hotwiring prompts in a few hundred thinking tokens, while official Bonsai 2 and this release take thousands. On the poisonous-plants prompt all three deliberate long. - A likely explanation is that this behavior lives in the parts of TWIN-TURBO that `R` does not carry (the Qwen3.6 blend) and/or in the ternary host itself. This is a hypothesis, not a measured result. - An earlier build using stochastic rounding alone had a higher no-answer rate (8.3%). Removing most of the rounding noise brought it back to the official level. ## Prose stability A small long-form probe: 6 prompts (3 Japanese, 3 English), 2 seeds, thinking **off**, temp 1.0 / top-k 20 / top-p 0.95, no penalties. | Model | Mean repeated 10-char n-grams | Loops | Hit length limit | |---|---:|---:|---:| | TWIN-TURBO Q8_0 | 0.043 | 1 / 12 | 2 / 12 | | Official Bonsai 2 | 0.049 | 1 / 12 | 0 / 12 | | This release | 0.020 | 0 / 12 | 0 / 12 | The loops that did occur, on a Japanese short-story prompt, appeared in TWIN-TURBO Q8_0 and in official Bonsai 2 alike. The transplant did not obviously worsen repetition. With n = 12 this means "not worse", not "better". Overall prose quality is still below the full-precision donor. We attribute this to the ternary host (WikiText PPL 7.67 vs 5.60), which other Bonsai 2 derivatives share. ## Negative results Several approaches were tested and deliberately not released: - **One-shot PTQ ternarization** of TWIN-TURBO collapsed quality (calibration PPL ~80–304, vs ~2.9 for official Bonsai 2 on the same text). - **Stochastic rounding of R alone** works (KL 0.3985), but its per-tensor noise in output space is 4–72× the edit's own energy. - **Over-driving the residual** (1.5× R) added more noise than signal. - **GPTQ-style error feedback on the lattice** mostly declined to edit. - **Splicing TWIN-TURBO's `lm_head` / embedding** (Q8, correctly folded) made the model worse (KL 0.380 → 0.596). Bonsai's hidden states are co-adapted with its own head. - **A least-squares `lm_head` bridge** fitted to TWIN-TURBO logits improved in-domain but never beat the ternary head on WikiText, even when trained with WikiText data. Logit MSE is misaligned with KL. - **Transferring the full task vector** (TWIN-TURBO − Qwen3.8, blend included) through scale fitting reached KL 0.397, worse than transferring R alone. ## Research conclusion Local tensor-space edits can transfer a small, coherent tuning residual into Bonsai 2. Our experiments suggest they cannot reconstruct a substantially different dense model: layer-local objectives stopped tracking end-to-end behavior whenever the target was far from the QAT-co-adapted ternary host. A natural next step would be end-to-end KL distillation from TWIN-TURBO that trains only the per-group FP16 scales and normalization parameters (~0.2B parameters) while keeping the ternary codes fixed. ## Limitations - Research release. Not benchmarked on standard suites (MMLU, HumanEval, etc.), long context, tool calling or multilingual tasks beyond the small probes above. - **Uncensored lineage.** The donor is a Heretic / Uncensored model. Treat this release as uncensored: it may comply with requests that the official Bonsai 2 or other safety-aligned models refuse. Its refusal rate on a harmful-prompt benchmark has not been measured. Users are responsible for appropriate and lawful use. - Prose quality and perplexity are below the full-precision donor (TWIN-TURBO Q8_0). Qwen3.8 itself was not measured. - Thinking can exhaust the output budget on some prompts (see Usage). - Text-only; MTP excluded. - The build pipeline and calibration text are not published yet. ## License and attribution Apache License 2.0. See `LICENSE` and `NOTICE.txt`. Created using Bonsai by Prism ML. Derived from: - Prism ML — [Ternary Bonsai 2 27B](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf) - DavidAU — [Qwen3.8-27B TWIN-TURBO Fable Cold Fusion 709 ULTRA HERETIC Uncensored](https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored) - Qwen / Alibaba Cloud — [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) and [Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B) (used for the residual) ## Acknowledgements Thanks to Prism ML for publishing Bonsai 2, its runtime and Hadamard contract, to DavidAU for the TWIN-TURBO source model, and to Hikari07jp, whose code-space abliteration of Bonsai 2 showed that the official ternary codes can be edited in place. This is an independent research experiment. It is not an official Prism ML, Qwen or DavidAU release.