Instructions to use ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Muse-Glimmer-30B, MLX 8-bit for mlx-serve
8-bit MLX conversion of meta-models/Muse-Glimmer-30B, built for mlx-serve — the native Zig MLX server for Apple Silicon (mlxserve.com).
Conversion
- Affine 8-bit, group size 64, on every matmul-read weight: text attention (incl. the sigmoid output gate projection), MLPs, lm_head, and the vision tower.
- Kept bf16: embed_tokens and the vision position embedding table (gather-read), the patch embedder (input dim not divisible by 64), all norms and biases.
- Weights, config and tokenizer are otherwise verbatim from the base repo. Trunk 32.9 GB, plus the 2.7 GB drafter below — 35.7 GB total.
DFlash drafter included
drafter/ holds a DFlash block-drafter for this trunk, so speculative decoding
works with no flags and no second download. mlx-serve probes that subdirectory
when the model loads, which also means a hot model switch brings the drafter
with it. --no-drafter opts out; an explicit --drafter <dir> still wins.
It is a 5-layer assistant (block_size 16, mask_token_id 201818, reading the
trunk at layers 1/13/25/37/49), shipped pre-packed at 8-bit so the server serves
it as-is. One assistant pass drafts a whole block and the trunk verifies it in a
single forward.
Speed on mlx-serve
All numbers below are mlx-serve (mlxserve.com) on an M4 Max (128 GB), temperature 0, median of 3, taken from the server's own reported decode rate:
| coding / agent workload | mlx-serve + drafter | mlx-serve, drafter off | |
|---|---|---|---|
| write a class from scratch | 41.4 tok/s | 15.9 tok/s | 2.61x |
| edit a file and re-emit it | 51.7 tok/s | 26.9 tok/s | 1.92x |
| emit a tool call | 42.1 tok/s | 18.6 tok/s | 2.26x |
62-86% of drafted tokens accepted. Code is where this pays: indentation, closing brackets, repeated identifiers and schema keys are all predictable enough to draft several ahead. Free-form prose accepts closer to 35% and gains proportionally less.
The effective block is a hardware property, not a checkpoint constant: the
verify width a machine can run depends on its qmm lanes, so on Apple silicon
without the wide lane the server caps the block (it logs
capped (no wide verify lane) at load) and uses the full 16 where the lane
exists. Greedy output is byte-identical to decoding without the drafter — the
drafter changes speed, never the tokens.
Serving with mlx-serve
Install and run — see mlxserve.com:
mlx-serve --model <this-repo> --serve
Text, tool calling and thinking are served. The model always reasons in a to=self channel; mlx-serve returns it as reasoning_content and parses the ATEM (<atem:invoke>) tool-call format natively, on both the OpenAI and Anthropic APIs. Vision weights are included in this conversion but image input is not served yet.
Needs about 38 GB of memory with the drafter resident.
- Downloads last month
- 548
Quantized
Model tree for ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit
Base model
meta-models/Muse-Glimmer-30B