Instructions to use topabaem/Haverbex-Muse-Glimmer-30B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use topabaem/Haverbex-Muse-Glimmer-30B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS # Run inference directly in the terminal: llama cli -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS # Run inference directly in the terminal: llama cli -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS # Run inference directly in the terminal: ./llama-cli -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Use Docker
docker model run hf.co/topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
- LM Studio
- Jan
- vLLM
How to use topabaem/Haverbex-Muse-Glimmer-30B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "topabaem/Haverbex-Muse-Glimmer-30B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "topabaem/Haverbex-Muse-Glimmer-30B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
- Ollama
How to use topabaem/Haverbex-Muse-Glimmer-30B with Ollama:
ollama run hf.co/topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
- Unsloth Desktop
- Pi
How to use topabaem/Haverbex-Muse-Glimmer-30B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use topabaem/Haverbex-Muse-Glimmer-30B with Docker Model Runner:
docker model run hf.co/topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
- Lemonade
How to use topabaem/Haverbex-Muse-Glimmer-30B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Run and chat with the model
lemonade run user.Haverbex-Muse-Glimmer-30B-IQ2_XS
List all available models
lemonade list
- Hermes Agent
How to use topabaem/Haverbex-Muse-Glimmer-30B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use topabaem/Haverbex-Muse-Glimmer-30B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "topabaem/Haverbex-Muse-Glimmer-30B:IQ2_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Haverbex-Muse-Glimmer-30B ยท Q4_K-mixed
A per-tensor mixed-precision GGUF of meta-models/Muse-Glimmer-30B. It occupies
17.51 GB instead of 55.73 GB, and on a paired 1,396-item MMLU screen its score is
not distinguishable from the BF16 original.
The language stack holds 27.85 B parameters, so the file works out to 5.03 bits per weight. It fits a 24 GB consumer card with room for a long context.
Two builds are published, at two operating points:
| file | size | bpw | MMLU vs BF16 |
|---|---|---|---|
haverbex-muse-glimmer-30b-Q4_K-mixed.gguf |
17.51 GB | 5.03 | โ0.36 pp, not distinguishable |
haverbex-muse-glimmer-30b-IQ2_XS-mixed.gguf |
10.23 GB | 3.04 | โ4.51 pp, real and measured |
The names state the build: the feed-forward tensors โ 69.6% of the bytes โ are
Q4_K in the first and IQ2_XS in the second, with everything around them held
higher. mixed marks a per-tensor allocation rather than one of llama.cpp's
uniform presets.
Take the 17.51 GB build unless you specifically need to fit under ~11 GB. The smaller one gives up real accuracy, and the section below says exactly how much and where the curve breaks.
Quality
Every rung was scored against the same BF16 GGUF, item by item, through one llama.cpp runtime. The comparison is paired, so the McNemar test and the interval below describe the difference directly rather than two independent samples.
| build | size | bpw | MMLU (1,396 paired) | difference | McNemar p | 95% CI |
|---|---|---|---|---|---|---|
| BF16 GGUF | 55.73 GB | 16.0 | 82.45% | โ | โ | โ |
| Q4_K-mixed (published) | 17.51 GB | 5.03 | 82.09% | โ0.36 pp | 0.42 | [โ1.24, +0.52] |
| Q6_K-attention variant | 17.89 GB | 5.14 | 81.95% | โ0.50 pp | 0.24 | [โ1.33, +0.33] |
| IQ3_XXS variant | 14.17 GB | 4.07 | 80.59% | โ1.86 pp | 0.0051 | [โ3.16, โ0.56] |
| IQ2_S variant | 11.28 GB | 3.35 | 77.22% | โ5.23 pp | 4e-10 | [โ6.95, โ3.51] |
| IQ2_XS-mixed (published) | 10.23 GB | 3.04 | 77.94% | โ4.51 pp | 1.4e-07 | [โ6.17, โ2.85] |
| IQ2_XXS variant | 9.58 GB | 2.85 | 73.93% | โ8.52 pp | 2.6e-16 | [โ10.51, โ6.53] |
Where the curve breaks
Cost per bit surrendered, from the measured points:
| interval | pp of MMLU per bit |
|---|---|
| 5.03 โ 4.07 bpw | 1.56 |
| 4.07 โ 3.04 bpw | 2.57 |
| 3.04 โ 2.85 bpw | 21.1 |
The wall is just under 3 bits on this model. IQ2_XS-mixed sits immediately
above it; the 9.58 GB build below it loses 8.52 pp for 0.65 GB, which is why
nothing under 10 GB is published. That shape is why a 10 GB target and a 2 pp
budget cannot both be met by post-training quantization here: reaching 10 GB
means 2.87 bits on a 27.85 B dense stack, and the accuracy at that density is
already spent.
Note that the 11.28 GB IQ2_S variant scored below the smaller 10.23 GB build (โ5.23 vs โ4.51 pp). The two are within each other's confidence intervals, so the honest reading is that they are indistinguishable and the larger one bought nothing โ its extra bits went to k/v, the attention gate and the output head rather than to the feed-forward tensors.
The unpublished variants are reported for the same reason: the Q6_K-attention variant costs 0.38 GB more than the 17.51 GB build and did not score better.
Checking the conversion before trusting the quantization
No BF16-versus-GGUF reference is published for this model. The vendor reports one aggregate figure averaged over fifteen benchmarks, which cannot set a pass line for a different harness. Without a reference, a low score has two possible causes that look identical: quantization loss, or a converter that mishandled one of this architecture's unusual pieces (logit softcapping at 20.0, an output multiplier of 0.196, a QK scale of 3.87, NoPE on full-attention layers, a sigmoid gate on attention output).
So both stacks were scored on the same 299 items, with the same prompts and the same rule for reading the answer:
| stack | accuracy | items with no answer letter |
|---|---|---|
transformers BF16 |
85.95% | 0 |
llama.cpp BF16 GGUF |
85.95% | 0 |
They agreed on 296 of 299 items individually. The converter and the runtime reproduce the reference implementation, which means the numbers in the first table are quantization loss and nothing else.
Speed
Measured with llama-bench on one RTX 4090 (24 GB), llama.cpp 0b1bad14f, all
layers on GPU, f16 KV cache, 3 repetitions.
| context depth | decode (tok/s) | prefill (tok/s) |
|---|---|---|
| 0 | 50.34 ยฑ 0.05 | 3,525 ยฑ 171 |
| 1,024 | 49.77 ยฑ 0.08 | 3,399 ยฑ 173 |
| 4,096 | 49.01 ยฑ 0.06 | 3,303 ยฑ 111 |
| 16,384 | 48.67 ยฑ 0.05 | 3,103 ยฑ 88 |
| 32,768 | 47.87 ยฑ 0.15 | 2,844 ยฑ 110 |
Decode falls 4.9% between an empty context and 32k tokens. The architecture explains that: three of every four layers use a 2,048-token sliding window, and grouped-query attention is 32:2, so the KV cache stays small and reading it back never dominates.
Two combined figures, which is what a request actually costs:
| workload | tok/s |
|---|---|
| 2,048-token prompt then 128 generated | 682.9 ยฑ 3.6 |
| 8,192-token prompt then 128 generated | 1,676.6 ยฑ 2.9 |
Quantizing the KV cache to q8_0 at 16k depth gives 47.18 tok/s against 48.67 for f16. It buys memory headroom at a small cost in speed, not a speedup.
Why decode sits near 50 tok/s
Decoding one token requires reading the weights once. At 17.51 GB and 50.34 tok/s the model is moving 881 GB/s, and the RTX 4090 tops out near 1,008 GB/s. The measurement is at 87.4% of the hardware limit, so there is very little left to win by tuning the runtime.
That bounds what any speedup has to do. It has to move fewer bytes per token, read fewer weights per token, or produce more than one token per pass.
Moving fewer bytes means a smaller quantization, and this card already measures what that costs: the IQ3_XXS variant is 19% smaller and gives up 1.86 pp, and the published IQ2_XS build is 42% smaller and gives up 4.51 pp. Reading fewer weights per token means a sparse or mixture-of-experts architecture, which is a property of the base model rather than of the quantization.
The remaining option is the practical one. Speculative decoding produces several
tokens per verification pass and leaves the output distribution unchanged, so it
does not trade quality for speed. llama.cpp supports it for this architecture
through DFlash, and upstream reports for Muse Glimmer put it between 38.96 and
84.64 tok/s in one case (PR #26842)
and around 137 tok/s in another (issue #26894).
Unsloth measure it at 3.1ร on an RTX 5090 (74.9 โ 233.4 tok/s) and ship the
draft model that does it โ dflash-kquant.gguf, 1.63 GB, in
unsloth/Muse-Glimmer-30B-GGUF.
It is built for this same base model, so it is the first thing to try with either
build here. That path is not measured here. Issue #26894 also records a blocker
worth knowing about before trying it: DFlash fails to bind against GGUFs that
store sliding_window_pattern as a per-layer array, and rewriting that key to
the scalar form is what made it work.
For serving rather than single-stream chat, concurrency is the other lever. Prefill already runs at roughly 60 times decode, so a server with several slots raises aggregate throughput well past the single-stream number above.
How it compares to other builds
Nobody publishes decode speed on the same hardware, so this table carries a hardware column and must not be read down as a ranking: an RTX 5090 has roughly 1.8ร the memory bandwidth of a 4090, and an M4 Max roughly a quarter of it.
| build | size | hardware | decode | with speculation | source |
|---|---|---|---|---|---|
| Haverbex Q4_K-mixed | 17.51 GB | RTX 4090 | 50.3 tok/s | not measured | measured here |
| Haverbex IQ2_XS-mixed | 10.23 GB | โ | not measured | not measured | โ |
| Unsloth Muse-Glimmer-30B-GGUF | quant not stated | RTX 5090 | 74.9 tok/s | 233.4 tok/s (3.1ร) | their card |
| Unsloth, same | quant not stated | Apple M5 Max | 26.6 tok/s | 50.2 tok/s (1.8ร) | their card |
| Unsloth, same | quant not stated | Apple M4 Max | 23.7 tok/s | 37.8 tok/s (1.5ร) | their card |
| Ternary Bonsai 27B | 7.17 GB | H100 | 98.0 tok/s | โ | their card |
| Ternary Bonsai 27B | 7.17 GB | Apple M5 Pro | 26.2 tok/s | โ | their card |
| Mach-1-Additive-35B | 7 GB | "consumer laptops" | up to 120 tok/s | โ | announcement post |
Unsloth do not say which of their fourteen quantizations produced the 74.9 and 233.4 figures, and those files span 10.75 GB to 32.30 GB, so those rows cannot be normalised to bandwidth the way the 4090 row above can.
Prefill, where it is published at all:
| build | hardware | prefill |
|---|---|---|
| Haverbex Q4_K-mixed | RTX 4090 | 3,525 tok/s empty context, 2,844 at 32k |
| Ternary Bonsai 27B | H100 | 2,596 tok/s |
| Ternary Bonsai 27B | Apple M5 Pro | 393 tok/s |
Unsloth publish no prefill figures.
The 10.23 GB build is not benchmarked. At the same 87.4% bandwidth utilisation it should decode roughly 1.7ร faster than the 17.51 GB one, purely because there are fewer bytes to read per token. That follows from the bandwidth bound but it is an inference, so it is left out of the table rather than estimated into it.
Unsloth also ship a vision projector (mmproj, 1.4โ3.85 GB) and fourteen
quantizations from 10.75 GB to 32.30 GB. This repository ships two builds,
text-only, each with paired MMLU measured against BF16 โ which is the difference
in what the two sets of files are for.
Known issue when generating
Under concurrent serving, a share of completions come back empty. On GSM8K through lm-evaluation-harness with 8 to 32 server slots, 28โ41% of responses carried no answer. The same items served one at a time are fine: 32 of 32 single-stream requests returned complete answers, and every item that had failed inside the concurrent run answered correctly when replayed alone.
An earlier version of this card attributed this to the converted tokenizer
metadata. That was wrong, and the correction matters because it changes what you
should do about it. The GGUF stores <|start|> and <|message|> as CONTROL, as
it should. llama.cpp demotes them to USER_DEFINED at load time for every model it
runs, under a code path commented as a workaround for a different model family,
which is why the load-time warning appears. The warning is not the bug.
Ruled out by measurement, each on the same 100 items: sampling temperature
(0/8 truncated at the server default, at 0.8 and at 0.0), the harness stop
strings (four stop-list variants all complete), slot reuse by prefix similarity
(-sps 0: 39% vs 41% baseline), and the sliding-window cache (--swa-full:
39%). Raising max_gen_toks from the task default of 256 to 2048 is a separate
and real fix โ this model emits reasoning_content before its answer โ but it
moved the score only 61.4% to 62.7%.
What remains is an intermittent fault in concurrent serving. It reproduces without any harness: eight identical-prefix requests sent through a thread pool produced one empty completion where the same eight sent sequentially produced none.
If you are benchmarking or doing anything correctness-critical, serve with
--parallel 1. For interactive use the effect is invisible, because
single-stream generation is unaffected.
No generative benchmark score is published for this model, because every number measured so far was produced under the concurrent configuration. The MMLU figures above are unaffected: they score one token at the answer position and never enter this path.
Precision map
| tensors | type | share of file |
|---|---|---|
ffn_{gate,up,down}, 156 |
Q4_K | 69.6% |
attn_{q,output}, 104 |
Q5_K | 9.6% |
attn_{k,v}, 104 |
Q8_0 | 0.6% |
attn_gate, 52 |
Q6_K | 4.8% |
token_embd |
Q6_K | 4.5% |
output |
Q8_0 | 4.5% |
| 313 norm tensors | quantizer default | negligible |
k and v stay at Q8_0 because 32:2 grouped-query attention leaves them at
0.6% of the file, so keeping them precise is close to free. attn_gate is a
fifth projection inside attention that most architectures do not have, and it was
held at Q6_K on the assumption that a gate error propagates further than a weight
error.
The 313 protected tensors include 104 that the converter creates rather than
reads: attn_q_norm filled with the QK scale of 3.87 and attn_k_norm filled
with 1.0. This architecture performs no runtime QK multiply, so that scale lives
entirely in those weights.
Quantization used an importance matrix built over a calibration mix that excludes both the MMLU items and the perplexity corpus used for evaluation.
Generation budget
This model emits reasoning_content before its answer. With a small
max_tokens the reasoning consumes the budget and content comes back empty โ
verified on the 10.23 GB build: a "write a Fibonacci function" prompt returns
nothing at 512 tokens and correct code at 2048. Set max_tokens to at least
2048, and higher for anything that needs long reasoning.
Usage
Needs llama.cpp at 0b1bad14f or later. Earlier builds report unknown model architecture: 'muse-glimmer'.
llama-server -m haverbex-muse-glimmer-30b-Q4_K-mixed.gguf -ngl 999 -c 32768 \
--host 127.0.0.1 --port 8080 --temp 0
import openai
client = openai.OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="-")
reply = client.chat.completions.create(
model="haverbex-muse-glimmer-30b",
messages=[{"role": "user", "content": "Explain grouped-query attention briefly."}],
)
print(reply.choices[0].message.content)
Set the temperature explicitly for tool-calling work. The llama-server default of 0.8 was enough, in earlier testing on a different model, to turn a deterministic task into a loop.
Note that -c is the total context divided across --parallel slots, not the
context per slot.
Provenance
- base
meta-models/Muse-Glimmer-30Bat revision46ac57a4502be771dc743ac85866351c5629ca71, with both weight shards verified by SHA-256 before conversion - llama.cpp
0b1bad14ff204627636aeb1de22ddcd5acb859d4 - quality measured on an A100-SXM4-80GB, speed on an RTX 4090
- per-item scores, the equivalence verdict, the recipe and the run receipts are in
measurements/
Limits
The MMLU screen covers 1,396 of 14,042 items, so it is a screen and not a full benchmark. MMLU measures retained knowledge and says nothing about whether the compressed model still behaves well when it has to act. That matters more than usual here, because the base model is built as an agent model, and its agentic behaviour was not evaluated.
The build is text-only. No mmproj was produced, and vision, DFlash, multi-GPU tensor split and KV-cache quantization were all excluded from the measured configuration. Every open upstream issue against this architecture at the time of the run concerned one of those paths.
Speed figures come from one GPU model. Decode is bandwidth-bound, so a card with different memory bandwidth will land somewhere else, roughly in proportion.
Citation
@misc{haverbexmuseglimmer2026,
title = {Haverbex-Muse-Glimmer-30B: quality-first mixed-precision
quantization of Muse-Glimmer-30B at 5.03 bits per weight},
author = {topabaem},
year = {2026},
howpublished = {\url{https://huggingface.co/topabaem/Haverbex-Muse-Glimmer-30B}},
note = {GGUF build \texttt{haverbex-muse-glimmer-30b-Q4\_K-mixed.gguf};
base model meta-models/Muse-Glimmer-30B at revision
46ac57a4502be771dc743ac85866351c5629ca71;
quantized and evaluated with llama.cpp 0b1bad14f}
}
Please also cite the base model and the tools this build depends on:
@misc{museglimmer30b,
title = {Muse-Glimmer-30B},
author = {{Meta}},
year = {2026},
howpublished = {\url{https://huggingface.co/meta-models/Muse-Glimmer-30B}}
}
@software{llamacpp,
title = {llama.cpp},
author = {Gerganov, Georgi and {llama.cpp contributors}},
url = {https://github.com/ggml-org/llama.cpp}
}
Models referenced in the comparison: Mach-1-Additive-35B (Syzygy Research) and Ternary-Bonsai-27B-gguf (Prism ML). Figures attributed to them come from their own model cards, except Mach-1's bits-per-weight and file size, which are from its announcement post.
Speed figures for other builds are taken from their own model cards: unsloth/Muse-Glimmer-30B-GGUF (Unsloth), Ternary-Bonsai-27B-gguf (Prism ML), and Mach-1-Additive-35B (Syzygy Research), except Mach-1's, which is from its announcement post.
- Downloads last month
- 162
2-bit
Model tree for topabaem/Haverbex-Muse-Glimmer-30B
Base model
meta-models/Muse-Glimmer-30B