Haverbex-Muse-Glimmer-30B ยท Q4_K-mixed

A per-tensor mixed-precision GGUF of meta-models/Muse-Glimmer-30B. It occupies 17.51 GB instead of 55.73 GB, and on a paired 1,396-item MMLU screen its score is not distinguishable from the BF16 original.

The language stack holds 27.85 B parameters, so the file works out to 5.03 bits per weight. It fits a 24 GB consumer card with room for a long context.

Two builds are published, at two operating points:

file size bpw MMLU vs BF16
haverbex-muse-glimmer-30b-Q4_K-mixed.gguf 17.51 GB 5.03 โˆ’0.36 pp, not distinguishable
haverbex-muse-glimmer-30b-IQ2_XS-mixed.gguf 10.23 GB 3.04 โˆ’4.51 pp, real and measured

The names state the build: the feed-forward tensors โ€” 69.6% of the bytes โ€” are Q4_K in the first and IQ2_XS in the second, with everything around them held higher. mixed marks a per-tensor allocation rather than one of llama.cpp's uniform presets.

Take the 17.51 GB build unless you specifically need to fit under ~11 GB. The smaller one gives up real accuracy, and the section below says exactly how much and where the curve breaks.

Quality

Every rung was scored against the same BF16 GGUF, item by item, through one llama.cpp runtime. The comparison is paired, so the McNemar test and the interval below describe the difference directly rather than two independent samples.

build size bpw MMLU (1,396 paired) difference McNemar p 95% CI
BF16 GGUF 55.73 GB 16.0 82.45% โ€” โ€” โ€”
Q4_K-mixed (published) 17.51 GB 5.03 82.09% โˆ’0.36 pp 0.42 [โˆ’1.24, +0.52]
Q6_K-attention variant 17.89 GB 5.14 81.95% โˆ’0.50 pp 0.24 [โˆ’1.33, +0.33]
IQ3_XXS variant 14.17 GB 4.07 80.59% โˆ’1.86 pp 0.0051 [โˆ’3.16, โˆ’0.56]
IQ2_S variant 11.28 GB 3.35 77.22% โˆ’5.23 pp 4e-10 [โˆ’6.95, โˆ’3.51]
IQ2_XS-mixed (published) 10.23 GB 3.04 77.94% โˆ’4.51 pp 1.4e-07 [โˆ’6.17, โˆ’2.85]
IQ2_XXS variant 9.58 GB 2.85 73.93% โˆ’8.52 pp 2.6e-16 [โˆ’10.51, โˆ’6.53]

Where the curve breaks

Cost per bit surrendered, from the measured points:

interval pp of MMLU per bit
5.03 โ†’ 4.07 bpw 1.56
4.07 โ†’ 3.04 bpw 2.57
3.04 โ†’ 2.85 bpw 21.1

The wall is just under 3 bits on this model. IQ2_XS-mixed sits immediately above it; the 9.58 GB build below it loses 8.52 pp for 0.65 GB, which is why nothing under 10 GB is published. That shape is why a 10 GB target and a 2 pp budget cannot both be met by post-training quantization here: reaching 10 GB means 2.87 bits on a 27.85 B dense stack, and the accuracy at that density is already spent.

Note that the 11.28 GB IQ2_S variant scored below the smaller 10.23 GB build (โˆ’5.23 vs โˆ’4.51 pp). The two are within each other's confidence intervals, so the honest reading is that they are indistinguishable and the larger one bought nothing โ€” its extra bits went to k/v, the attention gate and the output head rather than to the feed-forward tensors.

The unpublished variants are reported for the same reason: the Q6_K-attention variant costs 0.38 GB more than the 17.51 GB build and did not score better.

Checking the conversion before trusting the quantization

No BF16-versus-GGUF reference is published for this model. The vendor reports one aggregate figure averaged over fifteen benchmarks, which cannot set a pass line for a different harness. Without a reference, a low score has two possible causes that look identical: quantization loss, or a converter that mishandled one of this architecture's unusual pieces (logit softcapping at 20.0, an output multiplier of 0.196, a QK scale of 3.87, NoPE on full-attention layers, a sigmoid gate on attention output).

So both stacks were scored on the same 299 items, with the same prompts and the same rule for reading the answer:

stack accuracy items with no answer letter
transformers BF16 85.95% 0
llama.cpp BF16 GGUF 85.95% 0

They agreed on 296 of 299 items individually. The converter and the runtime reproduce the reference implementation, which means the numbers in the first table are quantization loss and nothing else.

Speed

Measured with llama-bench on one RTX 4090 (24 GB), llama.cpp 0b1bad14f, all layers on GPU, f16 KV cache, 3 repetitions.

context depth decode (tok/s) prefill (tok/s)
0 50.34 ยฑ 0.05 3,525 ยฑ 171
1,024 49.77 ยฑ 0.08 3,399 ยฑ 173
4,096 49.01 ยฑ 0.06 3,303 ยฑ 111
16,384 48.67 ยฑ 0.05 3,103 ยฑ 88
32,768 47.87 ยฑ 0.15 2,844 ยฑ 110

Decode falls 4.9% between an empty context and 32k tokens. The architecture explains that: three of every four layers use a 2,048-token sliding window, and grouped-query attention is 32:2, so the KV cache stays small and reading it back never dominates.

Two combined figures, which is what a request actually costs:

workload tok/s
2,048-token prompt then 128 generated 682.9 ยฑ 3.6
8,192-token prompt then 128 generated 1,676.6 ยฑ 2.9

Quantizing the KV cache to q8_0 at 16k depth gives 47.18 tok/s against 48.67 for f16. It buys memory headroom at a small cost in speed, not a speedup.

Why decode sits near 50 tok/s

Decoding one token requires reading the weights once. At 17.51 GB and 50.34 tok/s the model is moving 881 GB/s, and the RTX 4090 tops out near 1,008 GB/s. The measurement is at 87.4% of the hardware limit, so there is very little left to win by tuning the runtime.

That bounds what any speedup has to do. It has to move fewer bytes per token, read fewer weights per token, or produce more than one token per pass.

Moving fewer bytes means a smaller quantization, and this card already measures what that costs: the IQ3_XXS variant is 19% smaller and gives up 1.86 pp, and the published IQ2_XS build is 42% smaller and gives up 4.51 pp. Reading fewer weights per token means a sparse or mixture-of-experts architecture, which is a property of the base model rather than of the quantization.

The remaining option is the practical one. Speculative decoding produces several tokens per verification pass and leaves the output distribution unchanged, so it does not trade quality for speed. llama.cpp supports it for this architecture through DFlash, and upstream reports for Muse Glimmer put it between 38.96 and 84.64 tok/s in one case (PR #26842) and around 137 tok/s in another (issue #26894). Unsloth measure it at 3.1ร— on an RTX 5090 (74.9 โ†’ 233.4 tok/s) and ship the draft model that does it โ€” dflash-kquant.gguf, 1.63 GB, in unsloth/Muse-Glimmer-30B-GGUF. It is built for this same base model, so it is the first thing to try with either build here. That path is not measured here. Issue #26894 also records a blocker worth knowing about before trying it: DFlash fails to bind against GGUFs that store sliding_window_pattern as a per-layer array, and rewriting that key to the scalar form is what made it work.

For serving rather than single-stream chat, concurrency is the other lever. Prefill already runs at roughly 60 times decode, so a server with several slots raises aggregate throughput well past the single-stream number above.

How it compares to other builds

Nobody publishes decode speed on the same hardware, so this table carries a hardware column and must not be read down as a ranking: an RTX 5090 has roughly 1.8ร— the memory bandwidth of a 4090, and an M4 Max roughly a quarter of it.

build size hardware decode with speculation source
Haverbex Q4_K-mixed 17.51 GB RTX 4090 50.3 tok/s not measured measured here
Haverbex IQ2_XS-mixed 10.23 GB โ€” not measured not measured โ€”
Unsloth Muse-Glimmer-30B-GGUF quant not stated RTX 5090 74.9 tok/s 233.4 tok/s (3.1ร—) their card
Unsloth, same quant not stated Apple M5 Max 26.6 tok/s 50.2 tok/s (1.8ร—) their card
Unsloth, same quant not stated Apple M4 Max 23.7 tok/s 37.8 tok/s (1.5ร—) their card
Ternary Bonsai 27B 7.17 GB H100 98.0 tok/s โ€” their card
Ternary Bonsai 27B 7.17 GB Apple M5 Pro 26.2 tok/s โ€” their card
Mach-1-Additive-35B 7 GB "consumer laptops" up to 120 tok/s โ€” announcement post

Unsloth do not say which of their fourteen quantizations produced the 74.9 and 233.4 figures, and those files span 10.75 GB to 32.30 GB, so those rows cannot be normalised to bandwidth the way the 4090 row above can.

Prefill, where it is published at all:

build hardware prefill
Haverbex Q4_K-mixed RTX 4090 3,525 tok/s empty context, 2,844 at 32k
Ternary Bonsai 27B H100 2,596 tok/s
Ternary Bonsai 27B Apple M5 Pro 393 tok/s

Unsloth publish no prefill figures.

The 10.23 GB build is not benchmarked. At the same 87.4% bandwidth utilisation it should decode roughly 1.7ร— faster than the 17.51 GB one, purely because there are fewer bytes to read per token. That follows from the bandwidth bound but it is an inference, so it is left out of the table rather than estimated into it.

Unsloth also ship a vision projector (mmproj, 1.4โ€“3.85 GB) and fourteen quantizations from 10.75 GB to 32.30 GB. This repository ships two builds, text-only, each with paired MMLU measured against BF16 โ€” which is the difference in what the two sets of files are for.

Known issue when generating

Under concurrent serving, a share of completions come back empty. On GSM8K through lm-evaluation-harness with 8 to 32 server slots, 28โ€“41% of responses carried no answer. The same items served one at a time are fine: 32 of 32 single-stream requests returned complete answers, and every item that had failed inside the concurrent run answered correctly when replayed alone.

An earlier version of this card attributed this to the converted tokenizer metadata. That was wrong, and the correction matters because it changes what you should do about it. The GGUF stores <|start|> and <|message|> as CONTROL, as it should. llama.cpp demotes them to USER_DEFINED at load time for every model it runs, under a code path commented as a workaround for a different model family, which is why the load-time warning appears. The warning is not the bug.

Ruled out by measurement, each on the same 100 items: sampling temperature (0/8 truncated at the server default, at 0.8 and at 0.0), the harness stop strings (four stop-list variants all complete), slot reuse by prefix similarity (-sps 0: 39% vs 41% baseline), and the sliding-window cache (--swa-full: 39%). Raising max_gen_toks from the task default of 256 to 2048 is a separate and real fix โ€” this model emits reasoning_content before its answer โ€” but it moved the score only 61.4% to 62.7%.

What remains is an intermittent fault in concurrent serving. It reproduces without any harness: eight identical-prefix requests sent through a thread pool produced one empty completion where the same eight sent sequentially produced none.

If you are benchmarking or doing anything correctness-critical, serve with --parallel 1. For interactive use the effect is invisible, because single-stream generation is unaffected.

No generative benchmark score is published for this model, because every number measured so far was produced under the concurrent configuration. The MMLU figures above are unaffected: they score one token at the answer position and never enter this path.

Precision map

tensors type share of file
ffn_{gate,up,down}, 156 Q4_K 69.6%
attn_{q,output}, 104 Q5_K 9.6%
attn_{k,v}, 104 Q8_0 0.6%
attn_gate, 52 Q6_K 4.8%
token_embd Q6_K 4.5%
output Q8_0 4.5%
313 norm tensors quantizer default negligible

k and v stay at Q8_0 because 32:2 grouped-query attention leaves them at 0.6% of the file, so keeping them precise is close to free. attn_gate is a fifth projection inside attention that most architectures do not have, and it was held at Q6_K on the assumption that a gate error propagates further than a weight error.

The 313 protected tensors include 104 that the converter creates rather than reads: attn_q_norm filled with the QK scale of 3.87 and attn_k_norm filled with 1.0. This architecture performs no runtime QK multiply, so that scale lives entirely in those weights.

Quantization used an importance matrix built over a calibration mix that excludes both the MMLU items and the perplexity corpus used for evaluation.

Generation budget

This model emits reasoning_content before its answer. With a small max_tokens the reasoning consumes the budget and content comes back empty โ€” verified on the 10.23 GB build: a "write a Fibonacci function" prompt returns nothing at 512 tokens and correct code at 2048. Set max_tokens to at least 2048, and higher for anything that needs long reasoning.

Usage

Needs llama.cpp at 0b1bad14f or later. Earlier builds report unknown model architecture: 'muse-glimmer'.

llama-server -m haverbex-muse-glimmer-30b-Q4_K-mixed.gguf -ngl 999 -c 32768 \
  --host 127.0.0.1 --port 8080 --temp 0
import openai
client = openai.OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="-")
reply = client.chat.completions.create(
    model="haverbex-muse-glimmer-30b",
    messages=[{"role": "user", "content": "Explain grouped-query attention briefly."}],
)
print(reply.choices[0].message.content)

Set the temperature explicitly for tool-calling work. The llama-server default of 0.8 was enough, in earlier testing on a different model, to turn a deterministic task into a loop.

Note that -c is the total context divided across --parallel slots, not the context per slot.

Provenance

  • base meta-models/Muse-Glimmer-30B at revision 46ac57a4502be771dc743ac85866351c5629ca71, with both weight shards verified by SHA-256 before conversion
  • llama.cpp 0b1bad14ff204627636aeb1de22ddcd5acb859d4
  • quality measured on an A100-SXM4-80GB, speed on an RTX 4090
  • per-item scores, the equivalence verdict, the recipe and the run receipts are in measurements/

Limits

The MMLU screen covers 1,396 of 14,042 items, so it is a screen and not a full benchmark. MMLU measures retained knowledge and says nothing about whether the compressed model still behaves well when it has to act. That matters more than usual here, because the base model is built as an agent model, and its agentic behaviour was not evaluated.

The build is text-only. No mmproj was produced, and vision, DFlash, multi-GPU tensor split and KV-cache quantization were all excluded from the measured configuration. Every open upstream issue against this architecture at the time of the run concerned one of those paths.

Speed figures come from one GPU model. Decode is bandwidth-bound, so a card with different memory bandwidth will land somewhere else, roughly in proportion.

Citation

@misc{haverbexmuseglimmer2026,
  title        = {Haverbex-Muse-Glimmer-30B: quality-first mixed-precision
                  quantization of Muse-Glimmer-30B at 5.03 bits per weight},
  author       = {topabaem},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/topabaem/Haverbex-Muse-Glimmer-30B}},
  note         = {GGUF build \texttt{haverbex-muse-glimmer-30b-Q4\_K-mixed.gguf};
                  base model meta-models/Muse-Glimmer-30B at revision
                  46ac57a4502be771dc743ac85866351c5629ca71;
                  quantized and evaluated with llama.cpp 0b1bad14f}
}

Please also cite the base model and the tools this build depends on:

@misc{museglimmer30b,
  title        = {Muse-Glimmer-30B},
  author       = {{Meta}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/meta-models/Muse-Glimmer-30B}}
}

@software{llamacpp,
  title  = {llama.cpp},
  author = {Gerganov, Georgi and {llama.cpp contributors}},
  url    = {https://github.com/ggml-org/llama.cpp}
}

Models referenced in the comparison: Mach-1-Additive-35B (Syzygy Research) and Ternary-Bonsai-27B-gguf (Prism ML). Figures attributed to them come from their own model cards, except Mach-1's bits-per-weight and file size, which are from its announcement post.

Speed figures for other builds are taken from their own model cards: unsloth/Muse-Glimmer-30B-GGUF (Unsloth), Ternary-Bonsai-27B-gguf (Prism ML), and Mach-1-Additive-35B (Syzygy Research), except Mach-1's, which is from its announcement post.

Downloads last month
162
GGUF
Model size
28B params
Architecture
muse-glimmer
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for topabaem/Haverbex-Muse-Glimmer-30B

Quantized
(176)
this model