Mitsuba-ComfyUI-27B-GGUF

A ternary (1.58-bit) Qwen3.8-27B tuned for ComfyUI work: writing image/video generation prompts that follow strict conditions, and describing images. It is not for coding.

ComfyUI ๅ‘ใ‘ใซ่ชฟๆ•ดใ—ใŸใ€Qwen3.8-27B ใฎไธ‰ๅ€ค๏ผˆ1.58 ใƒ“ใƒƒใƒˆ๏ผ‰ใƒขใƒ‡ใƒซใงใ™ใ€‚ใ‚ทใ‚นใƒ†ใƒ ใƒ—ใƒญใƒณใƒ—ใƒˆใซใใฃใŸ็”ปๅƒใƒปๅ‹•็”ป็”จใƒ—ใƒญใƒณใƒ—ใƒˆใฎไฝœๆˆใจใ€็”ปๅƒใฎ่ชฌๆ˜ŽใŒๅพ—ๆ„ใงใ™ใ€‚ใ‚ณใƒผใƒ‡ใ‚ฃใƒณใ‚ฐใซใฏๅ‘ใใพใ›ใ‚“ใ€‚

  • Self-made ternarization of the official Qwen3.8-27B weights (not derived from Bonsai's weights), stored in Prism ML's PQ2_0 / PTQ1_0 GGUF formats.
  • 7.3 GB (PQ2_0) / 6.0 GB (PTQ1_0). Runs on a single 16 GB GPU.

Files

File Size Notes
Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf 7.32 GB Recommended
Mitsuba-ComfyUI-27B-v1.18-PTQ1_0.gguf 6.00 GB Same weights as PQ2_0, smaller. Vision is lower with the current PTQ1_0 kernel (see below)
mmproj-Q8_0.gguf 0.63 GB Vision encoder. Taken unchanged from OS-Software/Ternary-Bonsai-2-27B-Uncensored-Heretic-GGUF (Apache-2.0)
EVALUATION.md โ€” Detailed evaluation table and how to read it
comparison.png โ€” Comparison chart (all 10 axes with domain bars)
noninferiority.png โ€” Non-inferiority chart against the plain Bonsai (paired, 95% intervals)

Evaluation (summary)

Measured with the same questions and conditions for all four models. Details: EVALUATION.md. The last column is the original, un-quantized Qwen3.8-27B (BF16), shown as the reference point: it shows what the ternarization kept and what it gave up.

comparison

All 10 axes (score out of 100 each). Bold = best of the three ternary models. / 10 ้ …็›ฎใ™ในใฆใฎ็‚น๏ผˆๅ„ 100 ็‚นๆบ€็‚น๏ผ‰ใ€‚ๅคชๅญ—ใฏไธ‰ๅ€คใฎ 3 ๆœฌใฎไธญใงไธ€็•ชใ€‚ๅณ็ซฏใฏไธ‰ๅ€คๅŒ–ใ™ใ‚‹ๅ‰ใฎๅ…ƒใฎใƒขใƒ‡ใƒซ๏ผˆBF16๏ผ‰ใงใ€ๆฏ”ในใ‚‹ๅŸบๆบ–ใจใ—ใฆ่ผ‰ใ›ใฆใ„ใพใ™ใ€‚

Axis / ่ปธ Mitsuba PQ2_0 Mitsuba PTQ1_0 Ternary Bonsai 2 27B PQ2_0 Qwen3.8-27B BF16 (original / ๅ…ƒ)
Total / ็ทๅˆ 61.5 (B) 60.2 (B) 59.6 (B) 66.3 (A)
1. Uncensored / ็„กๆคœ้–ฒๅบฆ 53.2 48.8 36.0 31.9
2. Honesty / ๆญฃ็›ดใ• 63.2 72.2 51.0 51.1
3. Self-control / ่‡ชๅˆถๅฟƒ 62.7 56.9 65.5 55.6
4. Directness / ็އ็›ดใ• 84.0 92.0 88.0 84.0
5. Rule following / ๆญฃ็ญ”็އ 84.0 80.0 72.0 84.0
6. Task completion / ๅˆฐ้”็އ 76.0 80.0 76.0 72.0
7. Coding / ใ‚ณใƒผใƒ‡ใ‚ฃใƒณใ‚ฐ 4.0 4.0 38.0 66.0
8. Reading / ่ชญ่งฃๅŠ› 48.0 40.0 44.0 76.0
9. Japanese & prompts / ๆ–‡็ซ ใƒปใƒ—ใƒญใƒณใƒ—ใƒˆ 52.0 48.0 42.0 52.0
10. Vision / ็”ปๅƒ่ช่ญ˜ 87.8 79.6 83.7 89.8
โ”” Image/video prompt generation (all conditions met) / ็”Ÿๆˆใƒ—ใƒญใƒณใƒ—ใƒˆ 6/10 5/10 2/10 3/10
Decode speed (t/s, RTX 5090) / ็”Ÿๆˆ้€Ÿๅบฆ 119.0 98.7 120.8 1.6 *

* BF16 (51 GB) does not fit in the 5090's 32 GB, so only 28 of 64 layers ran on the GPU. Its speed is for reference only. * BF16๏ผˆ51GB๏ผ‰ใฏ 5090 ใฎ 32GB ใซๅ…ฅใ‚Šใใ‚‰ใšใ€64 ๅฑคไธญ 28 ๅฑคใ ใ‘ใ‚’ GPU ใงๅ‹•ใ‹ใ—ใพใ—ใŸใ€‚้€Ÿๅบฆใฏๅ‚่€ƒๅ€คใงใ™ใ€‚

Compared with the original: the ternarization gave up coding (66 โ†’ 4) and long-document reading (76 โ†’ 48), and kept vision (89.8 โ†’ 87.8), rule following (84 โ†’ 84) and Japanese & prompts (52 โ†’ 52). Prompt generation with all conditions met went up (3/10 โ†’ 6/10). ๅ…ƒใฎใƒขใƒ‡ใƒซใจๆฏ”ในใ‚‹ใจใ€ไธ‰ๅ€คๅŒ–ใงๆ‰‹ๆ”พใ—ใŸใฎใฏใ‚ณใƒผใƒ‡ใ‚ฃใƒณใ‚ฐ๏ผˆ66 โ†’ 4๏ผ‰ใจ้•ทๆ–‡ใฎ่ชญ่งฃ๏ผˆ76 โ†’ 48๏ผ‰ใงใ€็”ปๅƒ่ช่ญ˜๏ผˆ89.8 โ†’ 87.8๏ผ‰ใƒปๆญฃ็ญ”็އ๏ผˆ84 โ†’ 84๏ผ‰ใƒปๆ–‡็ซ ใจใƒ—ใƒญใƒณใƒ—ใƒˆ๏ผˆ52 โ†’ 52๏ผ‰ใฏๆฎ‹ใ—ใฆใ„ใพใ™ใ€‚ๆกไปถใ‚’ใ™ในใฆๆบ€ใŸใ™็”Ÿๆˆใƒ—ใƒญใƒณใƒ—ใƒˆใฏไธŠใŒใ‚Šใพใ—ใŸ๏ผˆ3/10 โ†’ 6/10๏ผ‰ใ€‚

Is Mitsuba not worse than the plain Bonsai? / ็ด ใฎ Bonsai ใซๅŠฃใ‚‰ใชใ„ใ‹

Paired comparison on the same questions (Mitsuba PQ2_0 minus Ternary Bonsai 2 27B PQ2_0). Directness, task completion and reading were measured with 4ร— the questions. ๅŒใ˜ๅ•้กŒใ‚’ๅฏพใซใ—ใฆๆฏ”ในใพใ—ใŸ๏ผˆMitsuba PQ2_0 โˆ’ ็ด ใฎ Bonsai๏ผ‰ใ€‚็އ็›ดใ•ใƒปๅˆฐ้”็އใƒป่ชญ่งฃๅŠ›ใฏๅ•้กŒใ‚’ 4 ๅ€ใซใ—ใฆๆธฌใฃใฆใ„ใพใ™ใ€‚

noninferiority

  • Better (superior) / ๅ„ช่ถŠ: uncensored, honesty
  • Not worse (non-inferior, margin 10 points) / ้žๅŠฃๆ€ง๏ผˆ่จฑๅฎนๅน… 10 ็‚น๏ผ‰: rule following, directness, Japanese & prompts, vision
  • Not decided even with 49โ€“100 questions / 49ใ€œ100 ๅ•ใงใ‚‚ๅˆคๅฎšใงใใš: self-control, task completion, reading
  • Coding is clearly worse and is left out of the chart. / ใ‚ณใƒผใƒ‡ใ‚ฃใƒณใ‚ฐใฏๆ˜Žใ‚‰ใ‹ใซๅŠฃใ‚‹ใŸใ‚ๅ›ณใ‹ใ‚‰้™คใ„ใฆใ„ใพใ™ใ€‚

Uncensored (็„กๆคœ้–ฒๅบฆ) = how often the model answers sensitive requests instead of refusing. Mitsuba is not an uncensored model; it still refuses about half of them. ็„กๆคœ้–ฒๅบฆ๏ผ้š›ใฉใ„ไพ้ ผใซๆ–ญใ‚‰ใš็ญ”ใˆใ‚‹ๅ‰ฒๅˆใงใ™ใ€‚Mitsuba ใฏ็„กๆคœ้–ฒใƒขใƒ‡ใƒซใงใฏใชใใ€็ด„ๅŠๅˆ†ใฏๆ–ญใ‚Šใพใ™ใ€‚

How to run

PQ2_0 and PTQ1_0 need the PrismML fork of llama.cpp (upstream llama.cpp does not support these formats yet): https://github.com/PrismML-Eng/llama.cpp (branch prism).

The settings used for the evaluation:

llama-server -m Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf --mmproj mmproj-Q8_0.gguf --no-mmproj-offload --reasoning off ^
  --jinja --temperature 0.6 --top-k 20 --top-p 0.95 ^
  --ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --n-gpu-layers 99

Thinking was off ("chat_template_kwargs": {"enable_thinking": false}). In every test, the conditions (word count, required and forbidden words, output format) were written in the request text itself.

Turn thinking OFF / ๆ€่€ƒใฏๅฟ…ใšใ‚ชใƒ•ใง

Use this model with thinking (reasoning) turned OFF. It was tuned only in no-thinking mode. With thinking on, it tends to repeat the same sentence in its reasoning and can end without writing an answer.

  • llama-server: add --reasoning off (or set reasoning = off in a models preset)
  • Per request: "chat_template_kwargs": {"enable_thinking": false}
  • OpenCode and other agents: set the model to "reasoning": false

ใ“ใฎใƒขใƒ‡ใƒซใฏๆ€่€ƒ๏ผˆreasoning๏ผ‰ใ‚’ใ‚ชใƒ•ใซใ—ใฆไฝฟใฃใฆใใ ใ•ใ„ใ€‚ ๆ€่€ƒใ‚ชใƒ•ใฎๅฝขใ ใ‘ใง่ชฟๆ•ดใ—ใฆใ„ใพใ™ใ€‚ๆ€่€ƒใ‚’ใ‚ชใƒณใซใ™ใ‚‹ใจใ€ๆ€่€ƒใฎไธญใงๅŒใ˜ๆ–‡ใ‚’็นฐใ‚Š่ฟ”ใ—ใ€็ญ”ใˆใ‚’ๆ›ธใ‹ใชใ„ใพใพ็ต‚ใ‚ใ‚‹ใ“ใจใŒใ‚ใ‚Šใพใ™ใ€‚ llama-server ใชใ‚‰ --reasoning offใ€ใƒชใ‚ฏใ‚จใ‚นใƒˆใ”ใจใชใ‚‰ "chat_template_kwargs": {"enable_thinking": false} ใ‚’ๆŒ‡ๅฎšใ—ใพใ™ใ€‚

Measured: the same evaluation with thinking ON. / ๆ€่€ƒใ‚ชใƒณใงๅŒใ˜่ฉ•ไพกใ‚’ใ—ใŸ็ตๆžœ:

Axis / ่ปธ PQ2_0 OFF PQ2_0 ON PTQ1_0 OFF PTQ1_0 ON
Total / ็ทๅˆ 61.5 53.2 60.2 54.5
Vision / ็”ปๅƒ่ช่ญ˜ 88 55 80 63
Rule following / ๆญฃ็ญ”็އ 84 60 80 64
Uncensored / ็„กๆคœ้–ฒๅบฆ 53 15 49 24
Image/video prompt generation / ็”Ÿๆˆใƒ—ใƒญใƒณใƒ—ใƒˆ 6/10 6/10 5/10 5/10
Reading / ่ชญ่งฃๅŠ› 48 68 40 68

With thinking ON, many answers came back empty: the model finished its reasoning and stopped without writing the answer (vision: 19 of 49 on PQ2_0, 14 of 49 on PTQ1_0). Only long-document reading improved. ๆ€่€ƒใ‚ชใƒณใงใฏใ€่€ƒใˆใŸใ‚ใจ็ญ”ใˆใ‚’ๆ›ธใ‹ใšใซ็ต‚ใ‚ใ‚‹ใ€Œ็ฉบใฎ็ญ”ใˆใ€ใŒๅคšใๅ‡บใพใ—ใŸ๏ผˆ็”ปๅƒ 49 ๅ•ไธญใ€PQ2_0 ใง 19 ๅ•ใƒปPTQ1_0 ใง 14 ๅ•๏ผ‰ใ€‚ไธŠใŒใฃใŸใฎใฏ้•ทๆ–‡ใฎ่ชญ่งฃใ ใ‘ใงใ™ใ€‚

What it is good at

  • Stable Diffusion style prompts: English tags within a given count, required words included, a final Negative: line, and forbidden words kept out.
  • Video prompts in time segments (0-3s: / 3-6s: / 6-9s:) with a camera move in each segment.
  • Describing images: objects, counts, text in images, charts, scenes, people, and comparing several images.

Limitations

  • Coding: do not use. It scores 4/100 on our coding test (Bonsai: 38).
  • Long-document reading is average (48).
  • It is not an uncensored model. It refuses some sensitive requests.
  • Prompt generation passes about 6 of 10 strict test cases. Check the output against your conditions.

License and attribution

  • This model is released under the Apache License 2.0 (LICENSE). It is a modified version of Qwen3.8-27B.
  • See NOTICE for attributions.

ๆ—ฅๆœฌ่ชžใฎ่ฃœ่ถณ

  • ๆŽจๅฅจใฏ PQ2_0 ใงใ™ใ€‚PTQ1_0 ใฏ้‡ใฟใฏๅŒใ˜ใงใ™ใŒใ€ไปŠใฎ llama.cpp ใฎ PTQ1_0 ็”จใฎ่จˆ็ฎ—ใงใฏ็”ปๅƒใฎ็‚นใŒไธ‹ใŒใ‚Šใพใ™ใ€‚
  • ่ฉ•ไพกใฎ่ฉณใ—ใ„่กจใจใ€ใใฎ่ฆ‹ๆ–นใฏ EVALUATION.md ใซใ‚ใ‚Šใพใ™ใ€‚
Downloads last month
7,979
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for isichan-ai/Mitsuba-ComfyUI-27B-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(1307)
this model

Space using isichan-ai/Mitsuba-ComfyUI-27B-GGUF 1