developerjeremylive abenzerps commited on
Commit
9d5ab63
·
0 Parent(s):

Duplicate from abenzerps/Ornith-1.5-9B-DFlash-GGUF

Browse files

Co-authored-by: Ahmet Benzer <abenzerps@users.noreply.huggingface.co>

.gitattributes ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ *.gguf filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Ornith AI
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
22
+
23
+ ---
24
+
25
+ ### Additional Attribution & Notices
26
+ - **Upstream Model:** Ornith-1.5-9B by Ornith AI (https://huggingface.co/ornith-ai/Ornith-1.5-9B)
27
+ - **Official GGUF Release:** Ornith-1.5-9B-GGUF (https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF)
28
+ - **Quantization Framework:** llama.cpp / GGML (MIT License, Copyright (c) 2023-2026 Georgi Gerganov and llama.cpp contributors)
29
+ - **Quantization Recipe:** GSQ-RCO optimization applied by the GSQ-RCO research team (2026).
Ornith-1.5-9B-DFlash-BF16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3b2fc5a11a05c738fbe713dd98cf958eea0cd85e60ec750a529d9ad962268c06
3
+ size 2594873632
Ornith-1.5-9B-DFlash-Q4_K_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a6730c2bc06cd855946decd88ac392c413f74cdfe25903ec5dd6b922c2c173e1
3
+ size 765960480
Ornith-1.5-9B-DFlash-Q6_K.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d64f74d08cadad5c08ecff45c895006d1f8f772ea93865508ee1a2b9ce5a2af2
3
+ size 1070899488
Ornith-1.5-9B-DFlash-Q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e81e57c83de0ef6e49548c740f0e274b96447b53ffa3de18b06b8d625138e338
3
+ size 1383768352
README.md ADDED
@@ -0,0 +1,59 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: mit
5
+ library_name: gguf
6
+ tags:
7
+ - text-generation
8
+ - dflash
9
+ - speculative-decoding
10
+ - draft-model
11
+ - block-diffusion
12
+ - llama.cpp
13
+ - quantized
14
+ - reasoning
15
+ - ornith
16
+ base_model: ornith-ai/Ornith-1.5-9B-GGUF
17
+ base_model_relation: quantized
18
+ pipeline_tag: text-generation
19
+ ---
20
+
21
+ # Ornith-1.5-9B DFlash GGUF
22
+
23
+ Official GGUF quantizations of [ornith-ai/Ornith-1.5-9B-DFlash](https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash), the speculative decoding draft model designed to accelerate [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B) inference in `llama.cpp`.
24
+
25
+ ## Overview
26
+
27
+ This is a **speculative decoding draft model** (Block-Diffusion / DFlash architecture), not a standalone language model.
28
+
29
+ It runs alongside [Ornith-1.5-9B GGUF](https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF) to increase generation throughput by drafting up to 16 tokens per step (`block_size = 16`).
30
+
31
+ ![Ornith-1.5 w/ Different Decoding Acceleration Techniques](assets/dflash-acceleration.jpg)
32
+
33
+ ## GGUF Files
34
+
35
+ | File | Quantization | Size |
36
+ | :--- | :---: | :---: |
37
+ | [Ornith-1.5-9B-DFlash-Q4_K_M.gguf](Ornith-1.5-9B-DFlash-Q4_K_M.gguf) | Q4_K_M *(Recommended)* | 730.48 MiB |
38
+ | [Ornith-1.5-9B-DFlash-Q6_K.gguf](Ornith-1.5-9B-DFlash-Q6_K.gguf) | Q6_K | 1021.29 MiB |
39
+ | [Ornith-1.5-9B-DFlash-Q8_0.gguf](Ornith-1.5-9B-DFlash-Q8_0.gguf) | Q8_0 | 1319.66 MiB |
40
+ | [Ornith-1.5-9B-DFlash-BF16.gguf](Ornith-1.5-9B-DFlash-BF16.gguf) | BF16 | 2474.66 MiB |
41
+
42
+ ## Usage
43
+
44
+ Pair this draft model with your target Ornith-1.5-9B model in `llama.cpp` using the `-md` (model draft) flag:
45
+
46
+ ```bash
47
+ llama-cli \
48
+ -m Ornith-1.5-9B-Q4_K_M.gguf \
49
+ -md Ornith-1.5-9B-DFlash-Q4_K_M.gguf \
50
+ --spec-type draft-dflash \
51
+ -ngl 99 -c 4096 --temp 0.6 \
52
+ -p "<|im_start|>user\nWrite a quicksort in Python.<|im_end|>\n<|im_start|>assistant\n"
53
+ ```
54
+
55
+ ## Attribution & License
56
+
57
+ - Distributed under the [MIT License](LICENSE), matching upstream [ornith-ai/Ornith-1.5-9B-DFlash](https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash).
58
+ - Target Model: [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B).
59
+ - Reference GGUF release: [ornith-ai/Ornith-1.5-9B-GGUF](https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUF).
assets/dflash-acceleration.jpg ADDED