bouroo commited on
Commit
0b7717e
·
verified ·
1 Parent(s): 4ef0072

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +49 -21
README.md CHANGED
@@ -11,6 +11,8 @@ tags:
11
  - edit-prediction
12
  - next-edit-suggestion
13
  - code
 
 
14
  license: apache-2.0
15
  language:
16
  - en
@@ -20,13 +22,13 @@ pipeline_tag: text-generation
20
 
21
  # bouroo/zeta-2.1-OptiQ-5
22
 
23
- A 5.0-bits-per-weight target, mixed-precision MLX conversion of [zed-industries/zeta-2.1](https://huggingface.co/zed-industries/zeta-2.1), produced with [OptiQ](https://mlx-optiq.com/) using calibration-driven KL-divergence sensitivity analysis. Sensitive tensors are kept at 8-bit; robust tensors are quantized to 4-bit. A higher target BPW keeps more tensors at 8-bit (larger, higher fidelity); a lower target pushes more to 4-bit (smaller).
24
 
25
  > "4-bit" / "8-bit" here are per-tensor choices from a mixed-precision allocation, not a global bit-width. Achieved BPW is the size-weighted average.
26
 
27
  ## Variants
28
 
29
- This model is part of a target-BPW family. Pick the size/quality trade-off you want:
30
 
31
  | Variant | Target BPW | Achieved BPW | 8-bit tensors | 4-bit tensors | Size |
32
  |---|---|---|---|---|---|
@@ -44,34 +46,50 @@ This model is part of a target-BPW family. Pick the size/quality trade-off you w
44
  | Candidate bits | 4, 8 |
45
  | Tensors 8-bit (sensitive) | 104 |
46
  | Tensors 4-bit (robust) | 121 |
47
- | Total quantized tensors | 225 |
48
  | Group size | 64 |
49
  | Reference signal | bf16 |
50
  | Calibration samples | 8 (OptiQ mix) |
51
  | Size on disk | 5.75 GB |
52
 
53
- Per-tensor bit allocation is recorded in `optiq_metadata.json` and embedded in `config.json` under `quantization`. All three variants share the **same** sensitivity analysis (one calibration sweep); only the allocation target differs.
54
 
55
  ## About the base model
56
 
57
- [Zeta 2.1](https://huggingface.co/zed-industries/zeta-2.1) is a **code edit prediction model** (next-edit suggestion) finetuned from `ByteDance-Seed/Seed-Coder-8B-Base`. Given code context, edit history and an editable region around the cursor, it predicts the rewritten content of the region. It is an 8B-parameter dense Llama-architecture model (32 layers, 4096 hidden, GQA with 8 KV heads), trained in BF16 with a 32k context window.
58
 
59
- ### Prompt format
60
 
61
- The model uses **SPM (suffix-prefix-middle)** style prompting with numbered multi-region markers for editable regions.
62
 
63
  ```
64
- <[fim-suffix]> code after editable region <[fim-prefix]>
65
- <filename>related/file.py related file content <filename>
66
- edit_history
67
- --- a/some_file.py
68
- +++ b/some_file.py
69
- -old
70
- +new
71
- <filename>path/to/target_file.py code before editable region <|marker_1|> code needs to<|user_cursor|> rewritten <|marker_2|> <[fim-middle]>
72
  ```
73
 
74
- Expected output (no backticks): `<|marker_1|> revised content of the editable region <|marker_2|>`.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
75
 
76
  ## Usage
77
 
@@ -80,7 +98,7 @@ Expected output (no backticks): `<|marker_1|> revised content of the editable re
80
  ```python
81
  from mlx_lm import load, generate
82
  model, tokenizer = load("bouroo/zeta-2.1-OptiQ-5")
83
- print(generate(model, tokenizer, prompt="def fibonacci(n):\n ", max_tokens=128))
84
  ```
85
 
86
  ### OptiQ serve
@@ -92,14 +110,24 @@ optiq serve --model bouroo/zeta-2.1-OptiQ-5
92
 
93
  ### LM Studio
94
 
95
- LM Studio supports MLX models on Apple Silicon:
96
-
97
  ```bash
98
  lms get https://huggingface.co/bouroo/zeta-2.1-OptiQ-5
99
- lms load bouroo/zeta-2.1-OptiQ-5
100
  ```
101
 
102
- Or search `bouroo/zeta-2.1-OptiQ-5` in the LM Studio in-app Hugging Face browser.
 
 
 
 
 
 
 
 
 
 
 
 
103
 
104
  ---
105
 
 
11
  - edit-prediction
12
  - next-edit-suggestion
13
  - code
14
+ - autocomplete
15
+ - fim
16
  license: apache-2.0
17
  language:
18
  - en
 
22
 
23
  # bouroo/zeta-2.1-OptiQ-5
24
 
25
+ A 5.0-bits-per-weight target, mixed-precision MLX conversion of [zed-industries/zeta-2.1](https://huggingface.co/zed-industries/zeta-2.1), tuned for **code edit-prediction / autocomplete**. Sensitive tensors stay 8-bit; robust tensors drop to 4-bit.
26
 
27
  > "4-bit" / "8-bit" here are per-tensor choices from a mixed-precision allocation, not a global bit-width. Achieved BPW is the size-weighted average.
28
 
29
  ## Variants
30
 
31
+ Pick the size/quality trade-off:
32
 
33
  | Variant | Target BPW | Achieved BPW | 8-bit tensors | 4-bit tensors | Size |
34
  |---|---|---|---|---|---|
 
46
  | Candidate bits | 4, 8 |
47
  | Tensors 8-bit (sensitive) | 104 |
48
  | Tensors 4-bit (robust) | 121 |
 
49
  | Group size | 64 |
50
  | Reference signal | bf16 |
51
  | Calibration samples | 8 (OptiQ mix) |
52
  | Size on disk | 5.75 GB |
53
 
54
+ Per-tensor allocation is in `optiq_metadata.json` and `config.json` under `quantization`. A `generation_config.json` ships autocomplete-optimized defaults (`temperature` 0.2, `top_p` 0.95, `max_new_tokens` 128, `eos_token_id` 2).
55
 
56
  ## About the base model
57
 
58
+ [Zeta 2.1](https://huggingface.co/zed-industries/zeta-2.1) is a **code edit-prediction model** (next-edit suggestion) finetuned from `ByteDance-Seed/Seed-Coder-8B-Base` — an 8B dense Llama (32 layers, GQA x8 KV heads, 32k context, BF16). Given context + an editable region, it predicts the rewritten region.
59
 
60
+ ## Prompt format (edit-prediction / FIM)
61
 
62
+ This is a **completion model** (no chat template). It uses Zeta's SPM-style format with markers. Minimal insertion prompt (empty region):
63
 
64
  ```
65
+ <[fim-suffix]>{code after cursor}
66
+ <[fim-prefix]><filename>{file_path}
67
+ {code before cursor}<|marker_1|><|marker_2|>
68
+ <[fim-middle]>
 
 
 
 
69
  ```
70
 
71
+ For an **edit** (rewrite an existing region), wrap the current region content with the markers and put `<|user_cursor|>` where the cursor lands:
72
+
73
+ ```
74
+ <[fim-suffix]>{code after region}
75
+ <[fim-prefix]><filename>{path}
76
+ {related files / edit_history, each prefixed with <filename>}
77
+ <filename>{path}
78
+ {before}<|marker_1|>{current part A}<|user_cursor|>{current part B}<|marker_2|>{after}
79
+ <[fim-middle]>
80
+ ```
81
+
82
+ The model generates `<|marker_1|>{predicted part A}<|user_cursor|>{part B}<|marker_2|>` and **stops at EOS** (`<[end_of_sentence]>`, id 2). Stop on `<|marker_2|>` / EOS.
83
+
84
+ Minimal `mlx_lm` example:
85
+
86
+ ```python
87
+ from mlx_lm import load, generate
88
+ model, tokenizer = load("bouroo/zeta-2.1-OptiQ-5")
89
+ prompt = "<[fim-suffix]>{suffix}\n<[fim-prefix]><filename>app.py\n{prefix}<|marker_1|><|marker_2|>\n<[fim-middle]>"
90
+ out = generate(model, tokenizer, prompt=prompt, max_tokens=128, temp=0.2, top_p=0.95)
91
+ # out starts with <|marker_1|>, stop at <|marker_2|>
92
+ ```
93
 
94
  ## Usage
95
 
 
98
  ```python
99
  from mlx_lm import load, generate
100
  model, tokenizer = load("bouroo/zeta-2.1-OptiQ-5")
101
+ out = generate(model, tokenizer, prompt=fim_prompt, max_tokens=128, temp=0.2, top_p=0.95)
102
  ```
103
 
104
  ### OptiQ serve
 
110
 
111
  ### LM Studio
112
 
 
 
113
  ```bash
114
  lms get https://huggingface.co/bouroo/zeta-2.1-OptiQ-5
115
+ lms load zeta-2.1-optiq-5 # LM Studio normalizes the key (no namespace); run `lms ls` to confirm
116
  ```
117
 
118
+ ### Edit-prediction / autocomplete client
119
+
120
+ This repo ships a working client under [`examples/`](./tree/main/examples):
121
+
122
+ - `examples/zeta_fim.py` — builds the exact FIM prompt, calls a local LM Studio server (`/v1/completions`), parses the markers, prints the predicted edit.
123
+ - `examples/continue-config.json` — Continue.dev config pointing at the LM Studio server with a custom FIM template.
124
+
125
+ > **LM Studio's built-in autocomplete is simple FIM and cannot format Zeta's multi-marker edit-prediction prompt.** For real autocomplete/edit-prediction, use the `examples/` client or the Zed editor against the server; the built-in autocomplete UI is not suitable.
126
+
127
+
128
+ ## Verification
129
+
130
+ Confirmed with `mlx_lm` (load + edit-prediction generation) and LM Studio (`lms get` + `lms load`).
131
 
132
  ---
133