Text Generation
MLX
Safetensors
English
llama
optiq
quantized
mixed-precision
4bit
8bit
apple-silicon
edit-prediction
next-edit-suggestion
code
autocomplete
fim
4-bit precision
Instructions to use bouroo/zeta-2.1-OptiQ-5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use bouroo/zeta-2.1-OptiQ-5 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("bouroo/zeta-2.1-OptiQ-5") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use bouroo/zeta-2.1-OptiQ-5 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "bouroo/zeta-2.1-OptiQ-5" --prompt "Once upon a time"
- Atomic Chat
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -11,6 +11,8 @@ tags:
|
|
| 11 |
- edit-prediction
|
| 12 |
- next-edit-suggestion
|
| 13 |
- code
|
|
|
|
|
|
|
| 14 |
license: apache-2.0
|
| 15 |
language:
|
| 16 |
- en
|
|
@@ -20,13 +22,13 @@ pipeline_tag: text-generation
|
|
| 20 |
|
| 21 |
# bouroo/zeta-2.1-OptiQ-5
|
| 22 |
|
| 23 |
-
A 5.0-bits-per-weight target, mixed-precision MLX conversion of [zed-industries/zeta-2.1](https://huggingface.co/zed-industries/zeta-2.1),
|
| 24 |
|
| 25 |
> "4-bit" / "8-bit" here are per-tensor choices from a mixed-precision allocation, not a global bit-width. Achieved BPW is the size-weighted average.
|
| 26 |
|
| 27 |
## Variants
|
| 28 |
|
| 29 |
-
|
| 30 |
|
| 31 |
| Variant | Target BPW | Achieved BPW | 8-bit tensors | 4-bit tensors | Size |
|
| 32 |
|---|---|---|---|---|---|
|
|
@@ -44,34 +46,50 @@ This model is part of a target-BPW family. Pick the size/quality trade-off you w
|
|
| 44 |
| Candidate bits | 4, 8 |
|
| 45 |
| Tensors 8-bit (sensitive) | 104 |
|
| 46 |
| Tensors 4-bit (robust) | 121 |
|
| 47 |
-
| Total quantized tensors | 225 |
|
| 48 |
| Group size | 64 |
|
| 49 |
| Reference signal | bf16 |
|
| 50 |
| Calibration samples | 8 (OptiQ mix) |
|
| 51 |
| Size on disk | 5.75 GB |
|
| 52 |
|
| 53 |
-
Per-tensor
|
| 54 |
|
| 55 |
## About the base model
|
| 56 |
|
| 57 |
-
[Zeta 2.1](https://huggingface.co/zed-industries/zeta-2.1) is a **code edit
|
| 58 |
|
| 59 |
-
##
|
| 60 |
|
| 61 |
-
|
| 62 |
|
| 63 |
```
|
| 64 |
-
<[fim-suffix]>
|
| 65 |
-
<
|
| 66 |
-
|
| 67 |
-
-
|
| 68 |
-
+++ b/some_file.py
|
| 69 |
-
-old
|
| 70 |
-
+new
|
| 71 |
-
<filename>path/to/target_file.py code before editable region <|marker_1|> code needs to<|user_cursor|> rewritten <|marker_2|> <[fim-middle]>
|
| 72 |
```
|
| 73 |
|
| 74 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 75 |
|
| 76 |
## Usage
|
| 77 |
|
|
@@ -80,7 +98,7 @@ Expected output (no backticks): `<|marker_1|> revised content of the editable re
|
|
| 80 |
```python
|
| 81 |
from mlx_lm import load, generate
|
| 82 |
model, tokenizer = load("bouroo/zeta-2.1-OptiQ-5")
|
| 83 |
-
|
| 84 |
```
|
| 85 |
|
| 86 |
### OptiQ serve
|
|
@@ -92,14 +110,24 @@ optiq serve --model bouroo/zeta-2.1-OptiQ-5
|
|
| 92 |
|
| 93 |
### LM Studio
|
| 94 |
|
| 95 |
-
LM Studio supports MLX models on Apple Silicon:
|
| 96 |
-
|
| 97 |
```bash
|
| 98 |
lms get https://huggingface.co/bouroo/zeta-2.1-OptiQ-5
|
| 99 |
-
lms load
|
| 100 |
```
|
| 101 |
|
| 102 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 103 |
|
| 104 |
---
|
| 105 |
|
|
|
|
| 11 |
- edit-prediction
|
| 12 |
- next-edit-suggestion
|
| 13 |
- code
|
| 14 |
+
- autocomplete
|
| 15 |
+
- fim
|
| 16 |
license: apache-2.0
|
| 17 |
language:
|
| 18 |
- en
|
|
|
|
| 22 |
|
| 23 |
# bouroo/zeta-2.1-OptiQ-5
|
| 24 |
|
| 25 |
+
A 5.0-bits-per-weight target, mixed-precision MLX conversion of [zed-industries/zeta-2.1](https://huggingface.co/zed-industries/zeta-2.1), tuned for **code edit-prediction / autocomplete**. Sensitive tensors stay 8-bit; robust tensors drop to 4-bit.
|
| 26 |
|
| 27 |
> "4-bit" / "8-bit" here are per-tensor choices from a mixed-precision allocation, not a global bit-width. Achieved BPW is the size-weighted average.
|
| 28 |
|
| 29 |
## Variants
|
| 30 |
|
| 31 |
+
Pick the size/quality trade-off:
|
| 32 |
|
| 33 |
| Variant | Target BPW | Achieved BPW | 8-bit tensors | 4-bit tensors | Size |
|
| 34 |
|---|---|---|---|---|---|
|
|
|
|
| 46 |
| Candidate bits | 4, 8 |
|
| 47 |
| Tensors 8-bit (sensitive) | 104 |
|
| 48 |
| Tensors 4-bit (robust) | 121 |
|
|
|
|
| 49 |
| Group size | 64 |
|
| 50 |
| Reference signal | bf16 |
|
| 51 |
| Calibration samples | 8 (OptiQ mix) |
|
| 52 |
| Size on disk | 5.75 GB |
|
| 53 |
|
| 54 |
+
Per-tensor allocation is in `optiq_metadata.json` and `config.json` under `quantization`. A `generation_config.json` ships autocomplete-optimized defaults (`temperature` 0.2, `top_p` 0.95, `max_new_tokens` 128, `eos_token_id` 2).
|
| 55 |
|
| 56 |
## About the base model
|
| 57 |
|
| 58 |
+
[Zeta 2.1](https://huggingface.co/zed-industries/zeta-2.1) is a **code edit-prediction model** (next-edit suggestion) finetuned from `ByteDance-Seed/Seed-Coder-8B-Base` — an 8B dense Llama (32 layers, GQA x8 KV heads, 32k context, BF16). Given context + an editable region, it predicts the rewritten region.
|
| 59 |
|
| 60 |
+
## Prompt format (edit-prediction / FIM)
|
| 61 |
|
| 62 |
+
This is a **completion model** (no chat template). It uses Zeta's SPM-style format with markers. Minimal insertion prompt (empty region):
|
| 63 |
|
| 64 |
```
|
| 65 |
+
<[fim-suffix]>{code after cursor}
|
| 66 |
+
<[fim-prefix]><filename>{file_path}
|
| 67 |
+
{code before cursor}<|marker_1|><|marker_2|>
|
| 68 |
+
<[fim-middle]>
|
|
|
|
|
|
|
|
|
|
|
|
|
| 69 |
```
|
| 70 |
|
| 71 |
+
For an **edit** (rewrite an existing region), wrap the current region content with the markers and put `<|user_cursor|>` where the cursor lands:
|
| 72 |
+
|
| 73 |
+
```
|
| 74 |
+
<[fim-suffix]>{code after region}
|
| 75 |
+
<[fim-prefix]><filename>{path}
|
| 76 |
+
{related files / edit_history, each prefixed with <filename>}
|
| 77 |
+
<filename>{path}
|
| 78 |
+
{before}<|marker_1|>{current part A}<|user_cursor|>{current part B}<|marker_2|>{after}
|
| 79 |
+
<[fim-middle]>
|
| 80 |
+
```
|
| 81 |
+
|
| 82 |
+
The model generates `<|marker_1|>{predicted part A}<|user_cursor|>{part B}<|marker_2|>` and **stops at EOS** (`<[end_of_sentence]>`, id 2). Stop on `<|marker_2|>` / EOS.
|
| 83 |
+
|
| 84 |
+
Minimal `mlx_lm` example:
|
| 85 |
+
|
| 86 |
+
```python
|
| 87 |
+
from mlx_lm import load, generate
|
| 88 |
+
model, tokenizer = load("bouroo/zeta-2.1-OptiQ-5")
|
| 89 |
+
prompt = "<[fim-suffix]>{suffix}\n<[fim-prefix]><filename>app.py\n{prefix}<|marker_1|><|marker_2|>\n<[fim-middle]>"
|
| 90 |
+
out = generate(model, tokenizer, prompt=prompt, max_tokens=128, temp=0.2, top_p=0.95)
|
| 91 |
+
# out starts with <|marker_1|>, stop at <|marker_2|>
|
| 92 |
+
```
|
| 93 |
|
| 94 |
## Usage
|
| 95 |
|
|
|
|
| 98 |
```python
|
| 99 |
from mlx_lm import load, generate
|
| 100 |
model, tokenizer = load("bouroo/zeta-2.1-OptiQ-5")
|
| 101 |
+
out = generate(model, tokenizer, prompt=fim_prompt, max_tokens=128, temp=0.2, top_p=0.95)
|
| 102 |
```
|
| 103 |
|
| 104 |
### OptiQ serve
|
|
|
|
| 110 |
|
| 111 |
### LM Studio
|
| 112 |
|
|
|
|
|
|
|
| 113 |
```bash
|
| 114 |
lms get https://huggingface.co/bouroo/zeta-2.1-OptiQ-5
|
| 115 |
+
lms load zeta-2.1-optiq-5 # LM Studio normalizes the key (no namespace); run `lms ls` to confirm
|
| 116 |
```
|
| 117 |
|
| 118 |
+
### Edit-prediction / autocomplete client
|
| 119 |
+
|
| 120 |
+
This repo ships a working client under [`examples/`](./tree/main/examples):
|
| 121 |
+
|
| 122 |
+
- `examples/zeta_fim.py` — builds the exact FIM prompt, calls a local LM Studio server (`/v1/completions`), parses the markers, prints the predicted edit.
|
| 123 |
+
- `examples/continue-config.json` — Continue.dev config pointing at the LM Studio server with a custom FIM template.
|
| 124 |
+
|
| 125 |
+
> **LM Studio's built-in autocomplete is simple FIM and cannot format Zeta's multi-marker edit-prediction prompt.** For real autocomplete/edit-prediction, use the `examples/` client or the Zed editor against the server; the built-in autocomplete UI is not suitable.
|
| 126 |
+
|
| 127 |
+
|
| 128 |
+
## Verification
|
| 129 |
+
|
| 130 |
+
Confirmed with `mlx_lm` (load + edit-prediction generation) and LM Studio (`lms get` + `lms load`).
|
| 131 |
|
| 132 |
---
|
| 133 |
|