Kolibri-1 MLX 3-bit

A mixed 3/6-bit MLX quantization of Aleph-Alpha/Kolibri-1, Aleph Alpha's 78B-A3.5B mixture-of-experts reasoning model for German and English, sized to run on a 48 GB Apple Silicon Mac.

This is a community conversion by here-be-dragons.ai, not an official Aleph Alpha release. For the model itself (training, evaluations, intended use, limitations) see the original model card and the tech report.

Quantization

Part Precision
Routed experts (75.5B of 78.1B parameters) 3 bit affine, group size 64
Attention, shared expert, embedding, LM head 6 bit affine, group size 64
MoE router (mlp.gate) bf16, as in the release; expert_bias stays fp32

3.61 bits per weight, 33 GiB on disk. The block-FP8 release weights were dequantized to bf16 and quantized once, with no intermediate format. A uniform 4-bit version would be about 44 GB and does not fit on a 48 GB machine.

Requirements

The kolibri1 architecture is not yet part of a released mlx-vlm or mlx-lm. Until the port is merged upstream, install mlx-vlm from the kolibri1 branch of our fork:

pip install git+https://github.com/here-be-dragons-ai/mlx-vlm@kolibri1

mlx-lm support is pending upstream review.

The same weights load in both mlx-vlm and mlx-lm.

  • Apple Silicon with 48 GB unified memory or more
  • On 48 GB, raise the GPU wired-memory limit, since the weights alone are 32.8 GiB: sudo sysctl -w iogpu.wired_limit_mb=40960
  • Do not run another large model at the same time.

Usage (mlx-vlm)

from mlx_vlm import load, stream_generate

model, processor = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
tokenizer = processor.tokenizer
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
for chunk in stream_generate(model, processor, prompt, max_tokens=2048,
                             temperature=1.0, top_p=0.97, top_k=128):
    print(chunk.text, end="", flush=True)

OpenAI-compatible server:

python -m mlx_vlm.server --model here-be-dragons-ai/Kolibri-1-MLX-3bit --port 8080

Usage (mlx-lm)

from mlx_lm import load, stream_generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load("here-be-dragons-ai/Kolibri-1-MLX-3bit")
messages = [{"role": "user", "content": "Erkläre kurz, was ein Mixture-of-Experts-Modell ist."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, reasoning_effort="low"
)
sampler = make_sampler(temp=1.0, top_p=0.97, top_k=128)
for chunk in stream_generate(model, tokenizer, prompt, max_tokens=2048, sampler=sampler):
    print(chunk.text, end="", flush=True)

Server notes (mlx-vlm)

Reasoning is controlled with the top-level request fields reasoning_effort or enable_thinking; the server ignores them inside chat_template_kwargs. The thinking comes back in reasoning, separate from content.

Recommended sampling, from the original release: temperature=1.0, top_p=0.97, top_k=128 (also set in generation_config.json).

The chat template supports Kolibri's reasoning mode: pass reasoning_effort (none, low, medium, high) to apply_chat_template. Without it, the model does not think.

Measurements

On an M5 Pro with 48 GB:

  • Decode: ~70 tokens/s at short context, 57 t/s at 23k, 40 t/s at 96k tokens
  • Peak memory: 35.3 GB
  • Needle retrieval succeeded at 23k and 96k tokens of context
  • Tool calls work through <tool_call> + JSON

The port's forward pass matches a reference implementation of the vLLM semantics to 1e-5 (on CPU).

License

Apache 2.0, same as the original model. See LICENSE. Kolibri 1 was developed by Aleph Alpha Research GmbH.

Downloads last month
215
Safetensors
Model size
78B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for here-be-dragons-ai/Kolibri-1-MLX-3bit

Quantized
(9)
this model