SeaWolf-AI commited on
Commit
fbe45b3
Β·
verified Β·
1 Parent(s): baf66cd

eval: retire duplicate GPQA entry (canonical record lives on Darwin-398B-JGOS)

Browse files
Files changed (1) hide show
  1. README.md +169 -185
README.md CHANGED
@@ -1,188 +1,172 @@
1
  ---
2
- license: apache-2.0
3
- language:
4
- - en
5
- - ko
6
- - zh
7
- - ja
8
- - multilingual
9
- library_name: transformers
10
- pipeline_tag: text-generation
11
- tags:
12
- - darwin
13
- - darwin-v9
14
- - darwin-jgos
15
- - moe
16
- - mixture-of-experts
17
- - reasoning
18
- - gpqa
19
- - benchmark
20
- - greedy
21
- - vidraft
22
- - eval-results
23
- base_model:
24
- - Qwen/Qwen3.5-397B-A17B
25
  base_model_relation: merge
26
- model-index:
27
- - name: Darwin-397B-JGOS
28
- results:
29
- - task:
30
- type: text-generation
31
- name: Graduate-Level Reasoning
32
- dataset:
33
- name: GPQA Diamond
34
- type: Idavidrein/gpqa
35
- config: gpqa_diamond
36
- split: train
37
- metrics:
38
- - type: accuracy
39
- value: 90.9
40
- name: Accuracy (greedy, single-sample, no test-time engine)
41
- verified: false
42
  ---
43
-
44
- # Darwin-397B-JGOS β€” Darwin V9 Platform Β· 397B MoE Β· GPQA 90.9 % (Pure Greedy)
45
-
46
- <p align="center">
47
- <a href="https://huggingface.co/FINAL-Bench/Darwin-397B-JGOS"><img src="https://img.shields.io/badge/⭐_GPQA_Diamond-90.9%25_Darwin--397B--JGOS-gold?style=for-the-badge" alt="GPQA"></a>
48
- <a href="https://huggingface.co/FINAL-Bench/Darwin-28B-REASON"><img src="https://img.shields.io/badge/🧬_Darwin--28B--REASON-89.39%25_(DELPHI)-blue?style=for-the-badge" alt="REASON"></a>
49
- </p>
50
-
51
- <p align="center">
52
- <a href="https://huggingface.co/FINAL-Bench/Darwin-28B-Opus"><img src="https://img.shields.io/badge/🧬_Darwin--28B--Opus-88.89%25-blue?style=for-the-badge" alt="Opus"></a>
53
- <a href="https://huggingface.co/FINAL-Bench/Darwin-36B-Opus"><img src="https://img.shields.io/badge/🧬_Darwin--36B--Opus-88.4%25-blue?style=for-the-badge" alt="36B"></a>
54
- </p>
55
-
56
- <p align="center">
57
- <a href="https://huggingface.co/collections/FINAL-Bench/darwin-family"><img src="https://img.shields.io/badge/🏠_Darwin_Family-Collection-green?style=for-the-badge" alt="Family"></a>
58
- <a href="https://huggingface.co/spaces/FINAL-Bench/Leaderboard"><img src="https://img.shields.io/badge/πŸ†_FINAL_Bench-Leaderboard-green?style=for-the-badge" alt="FINAL Bench"></a>
59
- </p>
60
-
61
- > Largest Darwin model Β· Qwen 3.5 397B base + Darwin V9 FFN transplant Β· 397B MoE (~17B active) Β· BF16
62
- > **GPQA Diamond: 90.9 % β€” pure greedy, single-sample, NO test-time engine**
63
-
64
- ---
65
-
66
- ## Overview
67
-
68
- **Darwin-397B-JGOS** is the largest and highest-scoring member of the Darwin family. Built on **Qwen 3.5 397B** as the base, it transplants the FFN (expert) strengths of multiple high-performance models through the **Darwin V9 platform**, producing a 397B-parameter Mixture-of-Experts model with ~17B active parameters per token.
69
-
70
- It reaches **90.9 % on GPQA Diamond with pure greedy decoding (single sample)** β€” surpassing **Darwin-28B-REASON (89.39 %, achieved *with* the Darwin-DELPHI test-time engine)** without using any test-time engine at all. This is the highest GPQA Diamond score in the Darwin family to date.
71
-
72
- ---
73
-
74
- ## 🧬 Darwin Platform & Research
75
-
76
- **Darwin** is VIDRAFT's measuring-result-driven reasoning model family β€” approximately **20 official models** plus **400+ community derivatives**, ranking among the top open models on GPQA.
77
-
78
- - **Darwin V9 platform** β€” evolutionary FFN/expert transplant and trust-weighted merging onto large-scale MoE backbones.
79
- - **FINAL Bench** β€” VIDRAFT's evaluation framework.
80
- - **4-layer Pre-AGI roadmap** β€” Darwin β†’ AETHER β†’ PROMETHEUS β†’ HEPHAESTUS.
81
-
82
- ---
83
-
84
- ## 🧬 Model Lineage
85
-
86
- | Role | Model | Contribution |
87
- |:---:|:---|:---|
88
- | **Base** | `Qwen 3.5 397B (A17B)` | 397B Mixture-of-Experts backbone (~17B active). |
89
- | **FFN transplant** | **Darwin V9 platform** (proprietary) | Transplants the FFN (expert) strengths of multiple high-performance models onto the base. |
90
- | **Result** | **`Darwin-397B-JGOS`** (this model) | 397B MoE β†’ **90.9 %** GPQA Diamond, pure greedy. |
91
-
92
- > The full Darwin V9 merge recipe β€” source models, weighting, and density β€” is **proprietary** and **not disclosed** (trade secret).
93
-
94
- ---
95
-
96
- ## βš™οΈ Technical Specifications
97
-
98
- | Component | Value |
99
- |:---|:---|
100
- | Architecture | `Qwen3_5MoeForConditionalGeneration` (Qwen 3.5 generation MoE) |
101
- | Parameters | **~397 B total / ~17 B active** (Mixture-of-Experts) |
102
- | Base | Qwen 3.5 397B (A17B) |
103
- | Precision | bfloat16 |
104
- | License | other |
105
-
106
- ---
107
-
108
- ## πŸ”¬ Core Technique β€” Darwin V9 Platform
109
-
110
- Darwin V9 transplants the FFN (expert) strengths of multiple high-performance models onto a Qwen 3.5 397B MoE base, then applies trust-weighted evolutionary merging.
111
-
112
- > The source models, merge weights, and density schedule are **proprietary** and constitute a **trade secret**; they are not published.
113
-
114
- ---
115
-
116
- ## πŸ† Benchmark β€” GPQA Diamond (198 questions)
117
-
118
- GPQA Diamond is a 198-question, PhD-level graduate science reasoning benchmark.
119
-
120
- | Model | Engine | **Accuracy** |
121
- |:---|:---|:---:|
122
- | Darwin-28B-Opus | Standard | 88.89 % (176 / 198) |
123
- | Darwin-28B-REASON | Darwin-DELPHI (test-time) | 89.39 % (177 / 198) |
124
- | **Darwin-397B-JGOS** | **Greedy (single-sample, no engine)** | **πŸ₯‡ 90.9 % (180 / 198)** |
125
-
126
- **Reproducible evaluation settings:**
127
- - Greedy decoding (temperature = 0), single sample β€” **no voting / self-consistency / test-time engine**
128
- - Max generation: 16,384 tokens
129
- - Answer options shuffled (seed = 42)
130
- - Hardware: **NVIDIA B200** (tensor-parallel 2 Γ— pipeline-parallel 3, 6 GPUs)
131
- - Inference engine: **vLLM**, bfloat16, `max_model_len = 18432`
132
-
133
- > Darwin-397B-JGOS achieves the family's top GPQA Diamond score using nothing but greedy decoding β€” no Darwin-DELPHI, no majority voting.
134
-
135
- ---
136
-
137
- ## πŸš€ Usage (vLLM)
138
-
139
- ```bash
140
- vllm serve FINAL-Bench/Darwin-397B-JGOS --tensor-parallel-size 2 --pipeline-parallel-size 3 --dtype bfloat16 --trust-remote-code
141
- ```
142
-
143
- ---
144
-
145
- ## 🎯 Recommended Use-Cases
146
-
147
- - Graduate-level STEM reasoning (GPQA / science qualifying exams)
148
- - Mathematical problem solving
149
- - Complex multi-step chain-of-thought
150
- - Code generation and debugging
151
- - Bilingual reasoning (strong English + Korean; also Chinese / Japanese)
152
-
153
- ## ⚠️ Limitations
154
-
155
- - 397B MoE in bfloat16 requires multi-GPU serving (e.g. B200 Γ—6 with TP2Γ—PP3).
156
- - The 90.9 % figure is a single-run greedy measurement on GPQA Diamond (198 items).
157
- - Reasoning traces can be verbose β€” control with max tokens.
158
-
159
- ---
160
-
161
- ## πŸ“š Citation
162
-
163
- ```bibtex
164
- @misc{darwin397b_jgos_2026,
165
- title = {Darwin-397B-JGOS: Darwin V9 Platform FFN Transplant on a 397B MoE Base},
166
- author = {FINAL-Bench / Darwin Research Team},
167
- year = {2026},
168
- howpublished = {https://huggingface.co/FINAL-Bench/Darwin-397B-JGOS},
169
- note = {Darwin V9 - 90.9 percent GPQA Diamond (greedy, single-sample)}
170
- }
171
- ```
172
-
173
- ---
174
-
175
- ## πŸ”— Related Darwin Models
176
-
177
- - **Darwin-28B-REASON** β€” RTD + Darwin-DELPHI, GPQA 89.39 %
178
- - **Darwin-28B-Opus** β€” base, GPQA 88.89 % (HF-official GPQA top tier)
179
- - **Darwin-36B-Opus** β€” MoE 36B, GPQA 88.4 %
180
- - **Darwin-27B-Opus** β€” 27B dense, GPQA 86.9 %
181
- - **Darwin-9B-NEG** β€” 9B Negentropy, GPQA 84.3 %
182
-
183
- ---
184
-
185
- *Darwin-397B-JGOS Β· Darwin V9 Platform Β· 90.9 % GPQA Diamond (pure greedy) Β· FINAL-Bench*
186
-
187
- <!-- eval re-index trigger: GPQA Diamond (diamond) = 90.9% (180/198), greedy single-sample, 2026-06-13 -->
188
-
 
1
  ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ - ko
6
+ - zh
7
+ - ja
8
+ - multilingual
9
+ library_name: transformers
10
+ pipeline_tag: text-generation
11
+ tags:
12
+ - darwin
13
+ - darwin-v9
14
+ - darwin-jgos
15
+ - moe
16
+ - mixture-of-experts
17
+ - reasoning
18
+ - gpqa
19
+ - benchmark
20
+ - greedy
21
+ - vidraft
22
+ - eval-results
23
+ base_model:
24
+ - Qwen/Qwen3.5-397B-A17B
25
  base_model_relation: merge
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
26
  ---
27
+
28
+ # Darwin-397B-JGOS β€” Darwin V9 Platform Β· 397B MoE Β· GPQA 90.9 % (Pure Greedy)
29
+
30
+ <p align="center">
31
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-397B-JGOS"><img src="https://img.shields.io/badge/⭐_GPQA_Diamond-90.9%25_Darwin--397B--JGOS-gold?style=for-the-badge" alt="GPQA"></a>
32
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-28B-REASON"><img src="https://img.shields.io/badge/🧬_Darwin--28B--REASON-89.39%25_(DELPHI)-blue?style=for-the-badge" alt="REASON"></a>
33
+ </p>
34
+
35
+ <p align="center">
36
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-28B-Opus"><img src="https://img.shields.io/badge/🧬_Darwin--28B--Opus-88.89%25-blue?style=for-the-badge" alt="Opus"></a>
37
+ <a href="https://huggingface.co/FINAL-Bench/Darwin-36B-Opus"><img src="https://img.shields.io/badge/🧬_Darwin--36B--Opus-88.4%25-blue?style=for-the-badge" alt="36B"></a>
38
+ </p>
39
+
40
+ <p align="center">
41
+ <a href="https://huggingface.co/collections/FINAL-Bench/darwin-family"><img src="https://img.shields.io/badge/🏠_Darwin_Family-Collection-green?style=for-the-badge" alt="Family"></a>
42
+ <a href="https://huggingface.co/spaces/FINAL-Bench/Leaderboard"><img src="https://img.shields.io/badge/πŸ†_FINAL_Bench-Leaderboard-green?style=for-the-badge" alt="FINAL Bench"></a>
43
+ </p>
44
+
45
+ > Largest Darwin model Β· Qwen 3.5 397B base + Darwin V9 FFN transplant Β· 397B MoE (~17B active) Β· BF16
46
+ > **GPQA Diamond: 90.9 % β€” pure greedy, single-sample, NO test-time engine**
47
+
48
+ ---
49
+
50
+ ## Overview
51
+
52
+ **Darwin-397B-JGOS** is the largest and highest-scoring member of the Darwin family. Built on **Qwen 3.5 397B** as the base, it transplants the FFN (expert) strengths of multiple high-performance models through the **Darwin V9 platform**, producing a 397B-parameter Mixture-of-Experts model with ~17B active parameters per token.
53
+
54
+ It reaches **90.9 % on GPQA Diamond with pure greedy decoding (single sample)** β€” surpassing **Darwin-28B-REASON (89.39 %, achieved *with* the Darwin-DELPHI test-time engine)** without using any test-time engine at all. This is the highest GPQA Diamond score in the Darwin family to date.
55
+
56
+ ---
57
+
58
+ ## 🧬 Darwin Platform & Research
59
+
60
+ **Darwin** is VIDRAFT's measuring-result-driven reasoning model family β€” approximately **20 official models** plus **400+ community derivatives**, ranking among the top open models on GPQA.
61
+
62
+ - **Darwin V9 platform** β€” evolutionary FFN/expert transplant and trust-weighted merging onto large-scale MoE backbones.
63
+ - **FINAL Bench** β€” VIDRAFT's evaluation framework.
64
+ - **4-layer Pre-AGI roadmap** β€” Darwin β†’ AETHER β†’ PROMETHEUS β†’ HEPHAESTUS.
65
+
66
+ ---
67
+
68
+ ## 🧬 Model Lineage
69
+
70
+ | Role | Model | Contribution |
71
+ |:---:|:---|:---|
72
+ | **Base** | `Qwen 3.5 397B (A17B)` | 397B Mixture-of-Experts backbone (~17B active). |
73
+ | **FFN transplant** | **Darwin V9 platform** (proprietary) | Transplants the FFN (expert) strengths of multiple high-performance models onto the base. |
74
+ | **Result** | **`Darwin-397B-JGOS`** (this model) | 397B MoE β†’ **90.9 %** GPQA Diamond, pure greedy. |
75
+
76
+ > The full Darwin V9 merge recipe β€” source models, weighting, and density β€” is **proprietary** and **not disclosed** (trade secret).
77
+
78
+ ---
79
+
80
+ ## βš™οΈ Technical Specifications
81
+
82
+ | Component | Value |
83
+ |:---|:---|
84
+ | Architecture | `Qwen3_5MoeForConditionalGeneration` (Qwen 3.5 generation MoE) |
85
+ | Parameters | **~397 B total / ~17 B active** (Mixture-of-Experts) |
86
+ | Base | Qwen 3.5 397B (A17B) |
87
+ | Precision | bfloat16 |
88
+ | License | other |
89
+
90
+ ---
91
+
92
+ ## πŸ”¬ Core Technique β€” Darwin V9 Platform
93
+
94
+ Darwin V9 transplants the FFN (expert) strengths of multiple high-performance models onto a Qwen 3.5 397B MoE base, then applies trust-weighted evolutionary merging.
95
+
96
+ > The source models, merge weights, and density schedule are **proprietary** and constitute a **trade secret**; they are not published.
97
+
98
+ ---
99
+
100
+ ## πŸ† Benchmark β€” GPQA Diamond (198 questions)
101
+
102
+ GPQA Diamond is a 198-question, PhD-level graduate science reasoning benchmark.
103
+
104
+ | Model | Engine | **Accuracy** |
105
+ |:---|:---|:---:|
106
+ | Darwin-28B-Opus | Standard | 88.89 % (176 / 198) |
107
+ | Darwin-28B-REASON | Darwin-DELPHI (test-time) | 89.39 % (177 / 198) |
108
+ | **Darwin-397B-JGOS** | **Greedy (single-sample, no engine)** | **πŸ₯‡ 90.9 % (180 / 198)** |
109
+
110
+ **Reproducible evaluation settings:**
111
+ - Greedy decoding (temperature = 0), single sample β€” **no voting / self-consistency / test-time engine**
112
+ - Max generation: 16,384 tokens
113
+ - Answer options shuffled (seed = 42)
114
+ - Hardware: **NVIDIA B200** (tensor-parallel 2 Γ— pipeline-parallel 3, 6 GPUs)
115
+ - Inference engine: **vLLM**, bfloat16, `max_model_len = 18432`
116
+
117
+ > Darwin-397B-JGOS achieves the family's top GPQA Diamond score using nothing but greedy decoding β€” no Darwin-DELPHI, no majority voting.
118
+
119
+ ---
120
+
121
+ ## πŸš€ Usage (vLLM)
122
+
123
+ ```bash
124
+ vllm serve FINAL-Bench/Darwin-397B-JGOS --tensor-parallel-size 2 --pipeline-parallel-size 3 --dtype bfloat16 --trust-remote-code
125
+ ```
126
+
127
+ ---
128
+
129
+ ## 🎯 Recommended Use-Cases
130
+
131
+ - Graduate-level STEM reasoning (GPQA / science qualifying exams)
132
+ - Mathematical problem solving
133
+ - Complex multi-step chain-of-thought
134
+ - Code generation and debugging
135
+ - Bilingual reasoning (strong English + Korean; also Chinese / Japanese)
136
+
137
+ ## ⚠️ Limitations
138
+
139
+ - 397B MoE in bfloat16 requires multi-GPU serving (e.g. B200 Γ—6 with TP2Γ—PP3).
140
+ - The 90.9 % figure is a single-run greedy measurement on GPQA Diamond (198 items).
141
+ - Reasoning traces can be verbose β€” control with max tokens.
142
+
143
+ ---
144
+
145
+ ## πŸ“š Citation
146
+
147
+ ```bibtex
148
+ @misc{darwin397b_jgos_2026,
149
+ title = {Darwin-397B-JGOS: Darwin V9 Platform FFN Transplant on a 397B MoE Base},
150
+ author = {FINAL-Bench / Darwin Research Team},
151
+ year = {2026},
152
+ howpublished = {https://huggingface.co/FINAL-Bench/Darwin-397B-JGOS},
153
+ note = {Darwin V9 - 90.9 percent GPQA Diamond (greedy, single-sample)}
154
+ }
155
+ ```
156
+
157
+ ---
158
+
159
+ ## πŸ”— Related Darwin Models
160
+
161
+ - **Darwin-28B-REASON** β€” RTD + Darwin-DELPHI, GPQA 89.39 %
162
+ - **Darwin-28B-Opus** β€” base, GPQA 88.89 % (HF-official GPQA top tier)
163
+ - **Darwin-36B-Opus** β€” MoE 36B, GPQA 88.4 %
164
+ - **Darwin-27B-Opus** β€” 27B dense, GPQA 86.9 %
165
+ - **Darwin-9B-NEG** β€” 9B Negentropy, GPQA 84.3 %
166
+
167
+ ---
168
+
169
+ *Darwin-397B-JGOS Β· Darwin V9 Platform Β· 90.9 % GPQA Diamond (pure greedy) Β· FINAL-Bench*
170
+
171
+ <!-- eval re-index trigger: GPQA Diamond (diamond) = 90.9% (180/198), greedy single-sample, 2026-06-13 -->
172
+