Instructions to use here-be-dragons-ai/Kolibri-1-MLX-3bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("here-be-dragons-ai/Kolibri-1-MLX-3bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "here-be-dragons-ai/Kolibri-1-MLX-3bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "here-be-dragons-ai/Kolibri-1-MLX-3bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default here-be-dragons-ai/Kolibri-1-MLX-3bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use here-be-dragons-ai/Kolibri-1-MLX-3bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "here-be-dragons-ai/Kolibri-1-MLX-3bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "here-be-dragons-ai/Kolibri-1-MLX-3bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Kolibri-1 MLX 3-bit
A mixed 3/6-bit MLX quantization of Aleph-Alpha/Kolibri-1, Aleph Alpha's 78B-A3.5B mixture-of-experts reasoning model for German and English, sized to run on a 48 GB Apple Silicon Mac.
This is a community conversion by here-be-dragons.ai, not an official Aleph Alpha release. For the model itself (training, evaluations, intended use, limitations) see the original model card and the tech report.
Quantization
| Part | Precision |
|---|---|
| Routed experts (75.5B of 78.1B parameters) | 3 bit affine, group size 64 |
| Attention, shared expert, embedding, LM head | 6 bit affine, group size 64 |
MoE router (mlp.gate) |
bf16, as in the release; expert_bias stays fp32 |
3.61 bits per weight, 33 GiB on disk. The block-FP8 release weights were dequantized to bf16 and quantized once, with no intermediate format. A uniform 4-bit version would be about 44 GB and does not fit on a 48 GB machine.
Requirements
The
kolibri1architecture is not yet part of a released mlx-vlm or mlx-lm. Until the port is merged upstream, install mlx-vlm from thekolibri1branch of our fork:pip install git+https://github.com/here-be-dragons-ai/mlx-vlm@kolibri1mlx-lm support is pending upstream review.
The same weights load in both mlx-vlm and mlx-lm.
- Apple Silicon with 48 GB unified memory or more
- On 48 GB, raise the GPU wired-memory limit, since the weights alone are 32.8 GiB:
sudo sysctl -w iogpu.wired_limit_mb=40960 - Do not run another large model at the same time.
Usage (mlx-vlm)
from mlx_vlm import load, stream_generate
model, processor = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
tokenizer = processor.tokenizer
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
for chunk in stream_generate(model, processor, prompt, max_tokens=2048,
temperature=1.0, top_p=0.97, top_k=128):
print(chunk.text, end="", flush=True)
OpenAI-compatible server:
python -m mlx_vlm.server --model here-be-dragons-ai/Kolibri-1-MLX-3bit --port 8080
Usage (mlx-lm)
from mlx_lm import load, stream_generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
sampler = make_sampler(temp=1.0, top_p=0.97, top_k=128)
for chunk in stream_generate(model, tokenizer, prompt, max_tokens=2048, sampler=sampler):
print(chunk.text, end="", flush=True)
Server notes (mlx-vlm)
Reasoning is controlled with the top-level request fields reasoning_effort or
enable_thinking; the server ignores them inside chat_template_kwargs. The thinking
comes back in reasoning, separate from content.
Recommended sampling, from the original release: temperature=1.0, top_p=0.97, top_k=128
(also set in generation_config.json).
The chat template supports Kolibri's reasoning mode: pass reasoning_effort
(none, low, medium, high) to apply_chat_template. Without it, the model does not think.
Measurements
On an M5 Pro with 48 GB:
- Decode: ~70 tokens/s at short context, 57 t/s at 23k, 40 t/s at 96k tokens
- Peak memory: 35.3 GB
- Needle retrieval succeeded at 23k and 96k tokens of context
- Tool calls work through
<tool_call>+ JSON
The port's forward pass matches a reference implementation of the vLLM semantics to 1e-5 (on CPU).
License
Apache 2.0, same as the original model. See LICENSE. Kolibri 1 was developed by Aleph Alpha Research GmbH.
- Downloads last month
- 215
3-bit