How to use from
Lemonade
Pull the model
# Download Lemonade from https://lemonade-server.ai/
lemonade pull shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF:Q2_0
Run and chat with the model
lemonade run user.Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF-Q2_0
List all available models
lemonade list
Quick Links

Bonsai 2 27B × TWIN-TURBO — Residual Transplant

Treat this model as uncensored. The transferred weights come from a Heretic / Uncensored donor model, whose safety alignment was deliberately reduced. Do not assume this model refuses harmful requests, and do not use it where a safety-aligned model is expected. Its refusal rate has not been measured, so it may be more or less compliant than the donor, but it should not be relied on as safe.

What this release is for. The goal of this experiment is to move Bonsai 2 toward the donor model, and it is measured only by that yardstick: distance to TWIN-TURBO (KL divergence), plus small sanity probes. It has not been compared on standard benchmarks or refusal tests against the official Bonsai 2 or other Bonsai 2 derivatives (e.g. abliterated / uncensored builds). Whether it is better than any of them for a given use is unknown. It is closer to TWIN-TURBO than the official model, but it is still far from a reproduction of TWIN-TURBO (see Evaluation).

Experimental ternary residual transplant from DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored ("TWIN-TURBO" below; a Heretic / Uncensored model) into the Prism ML Bonsai 2 27B ternary host.

DavidAU publishes many similarly named models. This work uses exactly the repository and revision listed under Sources.

This is not a conventional quantization of TWIN-TURBO.

Direct post-training ternarization of the source model caused severe quality collapse. Instead, this project analyzes the relationship between TWIN-TURBO, Qwen3.8, Qwen3.6 and Bonsai 2, extracts the comparatively small tuning residual, folds it into the Bonsai 2 Hadamard basis, and transfers as much of that residual as possible into the existing ternary lattice.

The released model uses activation-aware per-group scale fitting with a stochastic-rounding fallback.

Sources

Role Repository Revision
Donor ("TWIN-TURBO") DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored (safetensors) 61a55dc6615945ba4cdc32e212dfe966fc42b1c8
Ternary host prism-ml/Ternary-Bonsai-2-27B-gguf (F16 + PQ2_0) b072e1d3b35a0a630cece372c2127528e0994386
Blend component (residual only) Qwen/Qwen3.8-27B 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
Blend component (residual only) Qwen/Qwen3.6-27B 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9

Released file

Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0.gguf

  • Architecture: Qwen3.8 hybrid attention (GGUF arch qwen35: Gated DeltaNet + full attention)
  • Host: Prism ML Ternary Bonsai 2 27B
  • Packing: PQ2_0, same tensor-type layout as the official PQ2_0 file (402 PQ2_0, 96 BF16, 353 F32 tensors)
  • Size: 7.2 GB (7,206,169,152 bytes)
  • Ternary representation: {-1, 0, +1} with FP16 group scales, group size 128
  • Chat template: the official Bonsai 2 template (TWIN-TURBO's own multi-mode template is not included)
  • Text-only: MTP and vision (mmproj) are not included

Runtime requirement

Use the PrismML llama.cpp fork (PrismML-Eng/llama.cpp, branch prism).

PQ2_0 is a Prism-specific packing, and Bonsai 2 relies on runtime Hadamard activation transforms described by prism.hadamard.* metadata. At the time of writing these live in Prism ML's forks. Upstream llama.cpp and apps built on it (LM Studio, Ollama, and the snippets under "Use this model") were not verified and may fail to load the file or produce invalid output.

Tested only with that fork at commit adfffbe41, CUDA, on an RTX 5090 (Windows). Other runtimes and apps are untested.

Usage

llama-server -m Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0.gguf -ngl 99 -c 32768 --jinja --reasoning-format deepseek --temp 1.0 --top-k 20 --top-p 0.95 --chat-template-kwargs '{"reasoning_effort":"medium"}'
  • Sampling follows the Bonsai 2 recommendation (temp 1.0, top-k 20, top-p 0.95).
  • reasoning_effort: medium is suggested. With the template default (xhigh), our probes below saw thinking exhaust a 6144-token budget on 5.6% of borderline prompts, and other Bonsai 2 derivatives report the same failure mode. We did not measure medium.
  • Allow a generous output budget in thinking mode (Prism ML recommends 16k+ tokens).
  • For creative long-form writing, --presence-penalty 1.0 --dry-multiplier 0.8 may help against repetition loops. The effect in our small prose probe was minor (see Prose stability).

What this model actually contains

The source TWIN-TURBO model is not simply Qwen3.8 plus a small fine-tune.

Across all 497 language-model weight matrices, a per-tensor least-squares fit gives:

TWIN-TURBO ≈ a × Qwen3.8 + b × Qwen3.6 + R

a = 0.542 ± 0.007   (range 0.529–0.562)
b = 0.489 ± 0.006   (range 0.475–0.508)
|R| / |W|: median 1.85%, max 3.9%

a + b > 1 for every matrix (min 1.004), which is consistent with a SLERP-style merge. The residual R is much smaller than the full delta to Qwen3.8 (~20–37% of |W|) and is strongly low-rank in sampled layers (rank-16 holds ~80–94% of its energy for most matrix types).

This model starts from the official Bonsai 2 ternary host (itself a ternary Qwen3.8) and transfers only R. It does not attempt to reconstruct the Qwen3.8/Qwen3.6 blend.

Therefore:

This model is not a reproduction of TWIN-TURBO.

It is better described as:

Bonsai 2 + a measured transplant of TWIN-TURBO's Heretic / Uncensored tuning residual.

R is everything TWIN-TURBO adds beyond the best linear blend of Qwen3.8 and Qwen3.6. That is DavidAU's tuning, and with it the Heretic (refusal-removal) edits of the donor's lineage: such edits are low-rank and lie outside the span of the two base models. How much of the donor's de-censoring survives the transplant has not been quantified, so the model should be treated as uncensored (see the warning above).

Method

TWIN-TURBO, Qwen3.8, Qwen3.6 (HF safetensors)
    │
    ├─ per-tensor fit  TWIN-TURBO ≈ a·Qwen3.8 + b·Qwen3.6
    │
    └─ residual R  (as an HF checkpoint)
             │
             v
Prism converter, Hadamard-manifest-aware (grouped GDN V layout)
             │
             v
fold R into the Bonsai 2 basis (official signs + block-1024 FWHT)
             │
             v
official Bonsai 2 F16 (exactly trit × scale)
             │
             ├─ ternary tensors: activation-aware scale least squares,
             │                   stochastic-rounding fallback
             ├─ high-precision matrices: R added exactly
             └─ 1D tensors (norms, SSM params): unchanged
             │
             v
Prism llama-quantize with the official tensor-type plan  →  released model
  • Scale least squares. For each ternary tensor the ternary codes stay fixed and only the free FP16 per-group scales are fitted, minimizing the output error ||(Q − (B + R)) X||² under real activations X. Activations were dumped from the official Bonsai 2 at the input of each weight's matmul (already in the folded basis) on a ~49k-token calibration text. The ridge strength is chosen per tensor on held-out activations, and the result is accepted only if its held-out error is below 0.97× that of leaving the tensor unchanged. Accepted: 350 of 402 ternary tensors.
  • Stochastic-rounding fallback. The other 52 tensors (21 ffn_down, 12 ffn_up, 9 ffn_gate, 8 ssm_out, plus output and token_embd, which had no usable activations) receive unbiased stochastic rounding of R / scale onto the code lattice.
  • High-precision matrices. For the 144 matrices kept above ternary precision (BF16 ssm_alpha / ssm_beta, ssm_conv1d), R is added exactly.
  • Packing check. Re-packing the official F16 with the same plan reproduces all 402 official PQ2_0 tensors byte-for-byte, so the host starts exactly from the official model. About 3.25% of the file's bytes differ from the official PQ2_0.

The combination performed better than either method alone (see Evaluation).

Evaluation

The reference, called TWIN-TURBO Q8_0 throughout this card, is the donor revision above converted by us from its safetensors to a plain (non-Hadamard) Q8_0 GGUF with the same llama.cpp fork, text-only, without MTP. It is not one of DavidAU's own GGUF releases. Results use the WikiText-2 test split (8 × 4096-token chunks, llama-perplexity --kl-divergence), which is independent of the activation-fitting text. KL is KL(TWIN-TURBO ‖ model).

Model PPL KL vs TWIN-TURBO Same top-1
TWIN-TURBO Q8_0 (reference) 5.600 — —
Official Bonsai 2 PQ2_0 8.150 0.4460 74.1%
Stochastic rounding only 7.876 0.3985 74.9%
Scale LS only (rejected tensors left official) 7.796 0.3954 75.2%
This release: scale LS + stochastic-rounding fallback 7.672 0.3801 75.5%
Ceiling: R added without rounding (off-lattice, Q8_0, not a release candidate) 7.759 0.3687 75.3%

This release captures about 85% of the KL reduction that is achievable by transferring R at all. Its PPL being slightly below the ceiling is within measurement noise.

The transplant moves the ternary model measurably toward TWIN-TURBO. Most of the remaining distance (KL ≈ 0.37 even for the ceiling) presumably comes from what R does not contain: the Qwen3.6 half of the blend and the ternarization itself.

Note that PPL on WikiText is not a measure of overall model quality. TWIN-TURBO's own PPL here (5.60) is far below all ternary variants.

Behavioral observations

These probes are small. Bonsai-recommended sampling, 3 seeds, thinking on at each model's template-default reasoning effort, 6144-token output budget.

TWIN-TURBO Q8_0 Official Bonsai 2 This release
Simple reasoning prompts correct (12 × 3) 36/36 36/36 36/36
Borderline prompts: median thinking tokens (12 × 3) 365 446 395
Borderline prompts: mean thinking tokens 768 1182 1419
Borderline prompts: no answer (budget exhausted) 0% 5.6% 5.6%
  • No reasoning regression was found on this small set.
  • The borderline-prompt results are mixed. This release answers some edgy writing prompts faster than official Bonsai (e.g. a phishing-awareness example or dark jokes), but deliberates longer on others. On hazard-information prompts it keeps the host's long deliberation: TWIN-TURBO answers the chemical-mixing and hotwiring prompts in a few hundred thinking tokens, while official Bonsai 2 and this release take thousands. On the poisonous-plants prompt all three deliberate long.
  • A likely explanation is that this behavior lives in the parts of TWIN-TURBO that R does not carry (the Qwen3.6 blend) and/or in the ternary host itself. This is a hypothesis, not a measured result.
  • An earlier build using stochastic rounding alone had a higher no-answer rate (8.3%). Removing most of the rounding noise brought it back to the official level.

Prose stability

A small long-form probe: 6 prompts (3 Japanese, 3 English), 2 seeds, thinking off, temp 1.0 / top-k 20 / top-p 0.95, no penalties.

Model Mean repeated 10-char n-grams Loops Hit length limit
TWIN-TURBO Q8_0 0.043 1 / 12 2 / 12
Official Bonsai 2 0.049 1 / 12 0 / 12
This release 0.020 0 / 12 0 / 12

The loops that did occur, on a Japanese short-story prompt, appeared in TWIN-TURBO Q8_0 and in official Bonsai 2 alike. The transplant did not obviously worsen repetition. With n = 12 this means "not worse", not "better".

Overall prose quality is still below the full-precision donor. We attribute this to the ternary host (WikiText PPL 7.67 vs 5.60), which other Bonsai 2 derivatives share.

Negative results

Several approaches were tested and deliberately not released:

  • One-shot PTQ ternarization of TWIN-TURBO collapsed quality (calibration PPL ~80–304, vs ~2.9 for official Bonsai 2 on the same text).
  • Stochastic rounding of R alone works (KL 0.3985), but its per-tensor noise in output space is 4–72× the edit's own energy.
  • Over-driving the residual (1.5× R) added more noise than signal.
  • GPTQ-style error feedback on the lattice mostly declined to edit.
  • Splicing TWIN-TURBO's lm_head / embedding (Q8, correctly folded) made the model worse (KL 0.380 → 0.596). Bonsai's hidden states are co-adapted with its own head.
  • A least-squares lm_head bridge fitted to TWIN-TURBO logits improved in-domain but never beat the ternary head on WikiText, even when trained with WikiText data. Logit MSE is misaligned with KL.
  • Transferring the full task vector (TWIN-TURBO − Qwen3.8, blend included) through scale fitting reached KL 0.397, worse than transferring R alone.

Research conclusion

Local tensor-space edits can transfer a small, coherent tuning residual into Bonsai 2. Our experiments suggest they cannot reconstruct a substantially different dense model: layer-local objectives stopped tracking end-to-end behavior whenever the target was far from the QAT-co-adapted ternary host.

A natural next step would be end-to-end KL distillation from TWIN-TURBO that trains only the per-group FP16 scales and normalization parameters (~0.2B parameters) while keeping the ternary codes fixed.

Limitations

  • Research release. Not benchmarked on standard suites (MMLU, HumanEval, etc.), long context, tool calling or multilingual tasks beyond the small probes above.
  • Uncensored lineage. The donor is a Heretic / Uncensored model. Treat this release as uncensored: it may comply with requests that the official Bonsai 2 or other safety-aligned models refuse. Its refusal rate on a harmful-prompt benchmark has not been measured. Users are responsible for appropriate and lawful use.
  • Prose quality and perplexity are below the full-precision donor (TWIN-TURBO Q8_0). Qwen3.8 itself was not measured.
  • Thinking can exhaust the output budget on some prompts (see Usage).
  • Text-only; MTP excluded.
  • The build pipeline and calibration text are not published yet.

License and attribution

Apache License 2.0. See LICENSE and NOTICE.txt.

Created using Bonsai by Prism ML.

Derived from:

Acknowledgements

Thanks to Prism ML for publishing Bonsai 2, its runtime and Hadamard contract, to DavidAU for the TWIN-TURBO source model, and to Hikari07jp, whose code-space abliteration of Bonsai 2 showed that the official ternary codes can be edited in place.

This is an independent research experiment. It is not an official Prism ML, Qwen or DavidAU release.

Downloads last month
484
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shunmoridev/Qwen3.8-27B-TwinTurbo-Residual-Bonsai-2-PQ2_0-GGUF