ConorWang commited on
Commit
42e99c8
·
verified ·
1 Parent(s): a5c1ca2

Upload README.md

Browse files
Files changed (1) hide show
  1. README.md +420 -0
README.md CHANGED
@@ -1,3 +1,423 @@
1
  ---
 
 
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ library_name: transformers
3
+ pipeline_tag: text-generation
4
  license: apache-2.0
5
+ base_model:
6
+ - Qwen/Qwen3.8-27B
7
+ base_model_relation: finetune
8
+ language:
9
+ - en
10
+ - zh
11
+ tags:
12
+ - veriloop
13
+ - post-training
14
+ - coding-agent
15
+ - software-engineering
16
+ - mathematical-reasoning
17
+ - scientific-reasoning
18
+ - tool-use
19
+ - long-context
20
+ - safetensors
21
+ - vllm
22
+ - apache-2.0
23
  ---
24
+
25
+ <div align="center">
26
+
27
+ <img src="https://huggingface.co/tsinghua-sigs-robot-lab/veriloop-coder-e1/resolve/main/veriloop_logo.png" width="154" alt="VeriLoop logo">
28
+
29
+ # VeriLoop E2
30
+
31
+ ### Open 27B Post-Trained Model for Code Agents, Mathematical Reasoning, and Scientific Problem Solving
32
+
33
+ **Built on Qwen3.8-27B · 262K native context · Apache License 2.0**
34
+
35
+ **Developed by Tsinghua SIGS Robot Lab · Libo Wang**
36
+
37
+ <p>
38
+ <a href="https://www.apache.org/licenses/LICENSE-2.0"><img src="https://img.shields.io/badge/License-Apache--2.0-2F80ED?style=flat-square" alt="License: Apache 2.0"></a>
39
+ <img src="https://img.shields.io/badge/Base-Qwen3.8--27B-5B5BD6?style=flat-square" alt="Base: Qwen3.8-27B">
40
+ <img src="https://img.shields.io/badge/Context-262K-1F6FEB?style=flat-square" alt="Native context: 262K">
41
+ <img src="https://img.shields.io/badge/Serving-vLLM%200.17.0-0A7F6F?style=flat-square" alt="Serving: vLLM 0.17.0">
42
+ <img src="https://img.shields.io/badge/Stage-Post--Training-7C3AED?style=flat-square" alt="Stage: Post-Training">
43
+ </p>
44
+
45
+ <p>
46
+ <a href="https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2"><strong>Model</strong></a> ·
47
+ <a href="xxxx"><strong>Technical Report</strong></a> ·
48
+ <a href="https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence"><strong>Evaluation Evidence</strong></a> ·
49
+ <a href="xxxx"><strong>Scientific Artifacts</strong></a> ·
50
+ <a href="xxxx"><strong>GitHub</strong></a> ·
51
+ <a href="xxxx"><strong>Zenodo</strong></a>
52
+ </p>
53
+
54
+ </div>
55
+
56
+ ---
57
+
58
+ ## Overview
59
+
60
+ **VeriLoop E2** is an open 27B post-trained model built on **Qwen3.8-27B**. It is designed as the model component of the broader VeriLoop reasoning system, with emphasis on three workloads: **software-engineering agents, mathematical reasoning, and scientific problem solving**.
61
+
62
+ The system separates generative intelligence from verification authority. **VeriLoop E2** is responsible for proposal generation, abstraction, diagnosis, repair hypotheses, and structured reasoning. The **VeriLoop Harness** governs evidence admission, deterministic checks, external verification, commit/rollback, stopping, and evidence-state persistence. The model therefore does not self-certify its own progress.
63
+
64
+ This release focuses on the model weights, public inference path, evaluation record, and public functional description of the Harness. The production Harness implementation itself is **not included** in this repository.
65
+
66
+ ### Release highlights
67
+
68
+ - **27B open-weight post-trained model** derived from Qwen3.8-27B.
69
+ - **262,144-token native context window** in the released tokenizer configuration.
70
+ - Strong release results across nine code-agent, mathematics, and science benchmarks, including **76.2% SWE-bench Pro**, **88.8% Terminal-Bench 2.1**, **98.3% AIME 2026**, **93.9% GPQA Diamond**, and **89.6% Apex 2025**.
71
+ - A reproducible scientific-reasoning program built around verifier-governed recurrence rather than unconstrained retry.
72
+ - Two public-facing scientific demonstrations: a strict finite-dimensional **Riemann ζ zero-proportion certificate at 67.350003708785593%**, and **Asymptotic Graviton Tomography** for the black-hole information problem.
73
+ - OpenAI-compatible serving through **vLLM 0.17.0** with a validated 131,072-token serving configuration.
74
+
75
+ ---
76
+
77
+ ## Model Summary
78
+
79
+ | Property | VeriLoop E2 |
80
+ |---|---|
81
+ | Model family | VeriLoop E2 |
82
+ | Base model | Qwen3.8-27B |
83
+ | Parameter class | 27B |
84
+ | HF architecture class | `Qwen3_5ForConditionalGeneration` |
85
+ | Training stage | Post-Training |
86
+ | Primary domains | Code agents, software engineering, mathematics, scientific reasoning |
87
+ | Public post-training corpus accounting | 1,841,831 records |
88
+ | Native context length | 262,144 tokens |
89
+ | Validated vLLM serving length | 131,072 tokens |
90
+ | Tokenizer class | `Qwen2Tokenizer` |
91
+ | Weight format | `safetensors` |
92
+ | Languages | English, Chinese |
93
+ | Recommended serving engine | vLLM 0.17.0 |
94
+ | Model-weight license | Apache License 2.0 |
95
+ | Release year | 2026 |
96
+
97
+ The post-training mix spans repository-level software engineering, terminal and tool use, mathematical reasoning, scientific reasoning, verifier-sensitive repair, and recurrence-oriented training. Exact data construction, filtering, and training methodology are documented in the technical report rather than duplicated here.
98
+
99
+ ---
100
+
101
+ ## Benchmark Results
102
+
103
+ The README reports the **frozen release scores** for VeriLoop E2. Agentic benchmarks use the E2 checkpoint inside the frozen evaluation workflow, including the internal VeriLoop Harness where required by the task, benchmark-native tools, and the benchmark's official or designated evaluator. Exact per-benchmark protocols, task-level outputs, evaluator receipts, and integrity metadata are published separately in the **Evaluation Evidence** package.
104
+
105
+ > **Attribution boundary.** The reported results characterize the evaluated E2 system configuration. They should not be interpreted as evidence that an untouched Qwen3.8-27B base checkpoint, or the E2 checkpoint outside the evaluated runtime, reproduces the same numbers.
106
+
107
+ <p align="center">
108
+ <img src="assets/veriloop_e2_benchmark_grid.png" width="100%" alt="VeriLoop E2 benchmark comparison across nine public benchmarks">
109
+ </p>
110
+
111
+ <p align="center">
112
+ <sub><strong>Figure 1.</strong> VeriLoop E2 release snapshot across nine public benchmarks. Higher is better. Provider colors are fixed across panels; exact public model variants are shown in the comparison tables below. Full protocol and source provenance: <a href="https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence">Evaluation Evidence</a>.</sub>
113
+ </p>
114
+
115
+ ### Code and agentic benchmarks
116
+
117
+ | Benchmark | **VeriLoop E2** | OpenAI | Anthropic | Kimi | GLM | Qwen | DeepSeek |
118
+ |---|---:|---:|---:|---:|---:|---:|---:|
119
+ | **SWE-bench Pro** | **[76.2](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/swe-bench-pro)** | GPT-5.6 Sol 64.6 | Claude Fable 5.1 81.2 | — | GLM-5.2 Max 62.1 | Qwen3.8-Max 67.7 | DeepSeek V4 Pro Max 55.4 |
120
+ | **Terminal-Bench 2.1** | **[88.8](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-2.1)** | GPT-5.6 Sol 88.8 | — | Kimi K3 88.3 | GLM-5.3 88.2 | Qwen3.8-Max 86.6 | DeepSeek V4 Pro 87.9 |
121
+ | **DeepSWE v1.1** | **[64.6](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/deepswe-1.1)** | GPT-5.6 Sol 72.7 | Claude Fable 5 69.7 | Kimi K3 67.5 | GLM-5.3 66.9 | — | DeepSeek V4 Pro 62.7 |
122
+ | **Terminal-Bench 3.0** | **[29.7](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-3.0)** | GPT-5.6 Sol 34.6 | Claude Fable 5 33.7 | Kimi K3 17.4 | GLM-5.3 28.3 | — | — |
123
+ | **Terminal-Bench 4.0** | **[37.9](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-4.0)** | GPT-6 Astra 59.6 | Claude Fable 5.1 55.1 | — | GLM-5.3 41.8 | — | — |
124
+ | **SWE-Marathon v1.1** | **45.0** | GPT-5.6 Sol 42.5 | Claude Opus 4.8 48.8 | Kimi K3 48.1 | GLM-5.3 42.5 | — | — |
125
+
126
+ ### Mathematics and science benchmarks
127
+
128
+ | Benchmark | **VeriLoop E2** | OpenAI | Anthropic | Kimi | GLM | DeepSeek | Gemini |
129
+ |---|---:|---:|---:|---:|---:|---:|---:|
130
+ | **AIME 2026** | **[98.3](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/aime-2026)** | GPT-5.5 100.0 | Claude Opus 4.8 100.0 | Kimi K3 97.0 | — | DeepSeek V4 Pro 97.0 | Gemini 3.1 Pro 98.0 |
131
+ | **GPQA Diamond** | **[93.9](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/gpqa-diamond)** | GPT-5.6 Sol 94.1 | Claude Fable 5 92.6 | Kimi K3 93.5 | GLM-5.2 Max 91.2 | — | — |
132
+ | **Apex 2025** | **[89.6](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/apex-2025)** | GPT-5.5 80.0 | Claude Opus 4.8 81.0 | Kimi K3 66.0 | — | DeepSeek V4 Pro 28.0 | Gemini 3.1 Pro 61.0 |
133
+
134
+ A dash means that the release figure does not include a public comparison point for that provider on that benchmark. Each linked VeriLoop E2 score above resolves directly to its benchmark-specific public evidence directory. External reference values mirror the frozen comparison set used in Figure 1; harness notes and protocol caveats are retained in the evaluation ledger rather than duplicated here. **SWE-Marathon v1.1 (45.0)** remains reported on the model page, but is intentionally not represented by a Hugging Face Native Benchmark `.eval_results` entry because a stable HF benchmark registration/task identifier is not currently available.
135
+
136
+ ### Evaluation evidence
137
+
138
+ The public evidence package is intended to make the benchmark record inspectable rather than merely declarative. Where available, each task record binds:
139
+
140
+ ```text
141
+ task identity
142
+ ↓
143
+ model / system output
144
+ ↓
145
+ benchmark-native execution or evaluator record
146
+ ↓
147
+ score / pass-fail decision
148
+ ↓
149
+ integrity metadata and provenance
150
+ ```
151
+
152
+ **Evidence repository:** [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence)
153
+
154
+ | Benchmark | Public evaluation source |
155
+ |---|---|
156
+ | SWE-bench Pro | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/swe-bench-pro](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/swe-bench-pro) |
157
+ | Terminal-Bench 2.1 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-2.1](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-2.1) |
158
+ | DeepSWE v1.1 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/deepswe-1.1](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/deepswe-1.1) |
159
+ | Terminal-Bench 3.0 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-3.0](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-3.0) |
160
+ | Terminal-Bench 4.0 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-4.0](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-4.0) |
161
+ | AIME 2026 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/aime-2026](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/aime-2026) |
162
+ | GPQA Diamond | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/gpqa-diamond](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/gpqa-diamond) |
163
+ | Apex 2025 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/apex-2025](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/apex-2025) |
164
+
165
+ The structured Hugging Face evaluation descriptors for these eight registered benchmarks are published in the model repository under [`/.eval_results/`](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2/tree/main/.eval_results).
166
+
167
+ ---
168
+
169
+ ## VeriLoop Harness
170
+
171
+ VeriLoop is not designed around the idea that a model should announce its own improvement. The Harness treats model output as a **candidate state** that must earn admission through external evidence.
172
+
173
+ The current public abstraction is **VeriLoop-Governed Recurrence (VGR)**:
174
+
175
+ ```text
176
+ Current request
177
+ ↓
178
+ Contract compilation
179
+ ↓
180
+ VeriLoop E2 proposes a candidate
181
+ ↓
182
+ External / deterministic verification
183
+ ↓
184
+ Protected evidence state comparison
185
+ ├── no protected regression + at least one strict improvement → COMMIT
186
+ ├── otherwise → ROLLBACK
187
+ └── zero-rank certificate → STOP
188
+ ↓
189
+ Verified evidence becomes the next recurrence state
190
+ ```
191
+
192
+ The key boundary is deliberate:
193
+
194
+ - **Model authority:** propose, reason, abstract, diagnose, synthesize, repair.
195
+ - **Harness authority:** admit evidence, execute deterministic checks, verify, commit, roll back, stop, and persist verified state.
196
+ - **Benchmark / domain authority:** define task truth through native evaluators, tests, formal checks, numerical certificates, or other domain-specific validators.
197
+
198
+ This architecture is intended to preserve capability while preventing self-reported success from becoming system state. In software engineering, that means tests and execution receipts dominate plausible-looking patches. In mathematics and physics, it means a retained derivation must survive the relevant symbolic, numerical, or formal checks before it is promoted.
199
+
200
+ The production implementation contains private orchestration, routing, thresholds, prompt compilation, evidence-state machinery, repair arbitration, and deployment controls. Those implementation details are not part of this open model release. The README exposes the **functional contract**, not the proprietary runtime.
201
+
202
+ ---
203
+
204
+ ## Public 14-Rule Engineering Contract
205
+
206
+ The public Golden Rules are the model-visible execution discipline used to keep long-horizon work bounded, testable, and auditable.
207
+
208
+ | # | Rule | Public meaning |
209
+ |---:|---|---|
210
+ | 1 | **Current Request Supremacy** | The current request and exact output contract override stale memory, templates, and unrelated context. |
211
+ | 2 | **Evidence Before Escalation** | Search, tools, reverse analysis, or repair are triggered by concrete missing evidence or observed failure, not instinct. |
212
+ | 3 | **Read Before Rewrite** | Inspect the relevant entry points, interfaces, tests, conventions, and failure signals before editing. |
213
+ | 4 | **Minimal Sufficient Implementation** | Produce the smallest complete artifact that satisfies the task and preserves required interfaces. |
214
+ | 5 | **Surgical Repair, Not Blind Regeneration** | Repair the broken invariant; broaden the rewrite only when evidence shows local repair is insufficient. |
215
+ | 6 | **Intent Tests Beat Cosmetic Tests** | Syntax and formatting matter, but functional intent is the decisive acceptance criterion. |
216
+ | 7 | **Fail Loud, Never Fake Success** | Unknowns, skipped checks, degraded states, and failures remain explicit; unexecuted validation is never reported as success. |
217
+ | 8 | **Deterministic Logic Belongs in Code** | Parsing, scoring, structural checks, transformations, and reproducible validation should be deterministic whenever possible. |
218
+ | 9 | **Budget Is a First-Class Contract** | Token, tool, time, and compute budgets are part of the task contract rather than afterthoughts. |
219
+ | 10 | **Tool Use Must Be Typed and Accountable** | Every tool action has a trigger, expected output, and defined downstream consumer. |
220
+ | 11 | **Checkpoint Long Tasks** | Persist useful candidate state, validation receipts, repair records, and progress across long-running work. |
221
+ | 12 | **Follow Local Conventions** | Respect repository, benchmark, filename, API, language, and artifact conventions unless the request explicitly changes them. |
222
+ | 13 | **One Final Deliverable, Full Evidence Bundle** | Select one deliverable while retaining the evidence needed to explain and reproduce why it was selected. |
223
+ | 14 | **Separate Core Discipline from Domain Overlays** | Domain- or benchmark-specific rules apply only when relevant and cannot override the current request. |
224
+
225
+ ---
226
+
227
+ ## Scientific Demonstrations
228
+
229
+ The scientific demonstrations are not presented as isolated chat transcripts. They are examples of how the E2 model and the internal Harness can divide a difficult research problem into candidate derivations, falsifiable subclaims, executable checks, and retained evidence.
230
+
231
+ ### Riemann ζ: 67.350003708785593% strict finite-dimensional certificate
232
+
233
+ The released Riemann artifact reports a frozen assembly value of
234
+
235
+ \[
236
+ \kappa = 67.350003708785593\%.
237
+ \]
238
+
239
+ The current strict package closes **3/3 local inequalities**, resolves **327/327 difficult wells**, executes **190,375,830 strict branch-and-bound nodes**, and passes the final **exact rational assembly** check.
240
+
241
+ What this result **does** establish within the released artifact is a strict finite-dimensional computer-assisted certificate under its stated analytic setup and imported assumptions. What it **does not** establish is equally important:
242
+
243
+ - it is **not a proof of the Riemann Hypothesis**;
244
+ - it is **not yet an end-to-end Lean/nanoda kernel proof** of the complete upstream analytic chain;
245
+ - imported analytic normalization steps must remain clearly separated from the finite-dimensional certificate until the formal bridge and complete replay are closed.
246
+
247
+ The public package therefore emphasizes claim discipline, reproducibility, and certificate structure rather than treating a numerical percentage as a substitute for mathematical provenance.
248
+
249
+ **Artifact:** `xxxx` · **Technical note:** `xxxx` · **Zenodo:** `xxxx`
250
+
251
+ ### Black-hole information problem: Asymptotic Graviton Tomography
252
+
253
+ **Asymptotic Graviton Tomography** is the second scientific reasoning demonstration. It studies an information-reconstruction route through asymptotic gravitational observables, using the Harness to separate retained derivations from rejected or insufficiently supported branches.
254
+
255
+ The public claim is intentionally bounded: this is a **research demonstration of a verifier-governed theoretical-physics derivation**, not a declaration that the black-hole information paradox has been solved. The artifact is intended to expose the derivation structure, assumptions, checks, and remaining theoretical boundaries clearly enough for external scientific criticism.
256
+
257
+ **Artifact:** `xxxx` · **Technical note:** `xxxx`
258
+
259
+ > The Riemann and black-hole demo artifacts are released separately under **research-only, non-commercial terms**. They are not covered by the Apache-2.0 grant for the model weights and public inference utilities unless a specific file explicitly says otherwise.
260
+
261
+ ---
262
+
263
+ ## Inference
264
+
265
+ ### Recommended environment
266
+
267
+ The following stack is the validated reference environment for the public serving path:
268
+
269
+ | Component | Version / setting |
270
+ |---|---|
271
+ | Python | 3.12.x |
272
+ | vLLM | 0.17.0 |
273
+ | PyTorch | 2.10.0 + CUDA 12.9 build |
274
+ | CUDA runtime | 12.9 |
275
+ | Transformers | 4.57.6 |
276
+ | Triton | 3.6.0 |
277
+ | Dtype | `bfloat16` |
278
+ | Validated serving context | 131,072 tokens |
279
+
280
+ The released tokenizer advertises a native maximum length of **262,144 tokens**. The 131,072-token value above is the **validated public serving configuration**, not a redefinition of the model's native context length. Longer serving windows require appropriate accelerator memory and KV-cache planning.
281
+
282
+ ### vLLM server
283
+
284
+ Use the tokenizer and chat template shipped with the model repository.
285
+
286
+ ```bash
287
+ python -m pip install "vllm==0.17.0"
288
+
289
+ MODEL="<MODEL_PATH_OR_HF_ID>"
290
+
291
+ vllm serve "${MODEL}" \
292
+ --served-model-name veriloop-e2 \
293
+ --dtype bfloat16 \
294
+ --model-impl vllm \
295
+ --language-model-only \
296
+ --max-model-len 131072 \
297
+ --max-num-seqs 16 \
298
+ --gpu-memory-utilization 0.92 \
299
+ --generation-config vllm \
300
+ --disable-uvicorn-access-log \
301
+ --host 127.0.0.1 \
302
+ --port 8001
303
+ ```
304
+
305
+ ### OpenAI-compatible request
306
+
307
+ ```bash
308
+ curl http://127.0.0.1:8001/v1/chat/completions \
309
+ -H "Content-Type: application/json" \
310
+ -d '{
311
+ "model": "veriloop-e2",
312
+ "messages": [
313
+ {"role": "user", "content": "Reply exactly: 8001_OK"}
314
+ ],
315
+ "temperature": 0,
316
+ "max_tokens": 16,
317
+ "stream": false
318
+ }'
319
+ ```
320
+
321
+ A validated smoke test returns:
322
+
323
+ ```text
324
+ 8001_OK
325
+ ```
326
+
327
+ ### Python client
328
+
329
+ ```python
330
+ from openai import OpenAI
331
+
332
+ client = OpenAI(
333
+ base_url="http://127.0.0.1:8001/v1",
334
+ api_key="EMPTY",
335
+ )
336
+
337
+ response = client.chat.completions.create(
338
+ model="veriloop-e2",
339
+ messages=[
340
+ {"role": "user", "content": "Explain why rollback matters in verifier-governed reasoning."}
341
+ ],
342
+ temperature=0.2,
343
+ max_tokens=1024,
344
+ )
345
+
346
+ print(response.choices[0].message.content)
347
+ ```
348
+
349
+ ### Protocol note
350
+
351
+ For tool use or long multi-turn reasoning, treat the repository-shipped tokenizer and chat template as part of the model protocol. Replacing role delimiters, tool-call syntax, stop semantics, or reasoning-history behavior can change observed system behavior even when the weights are unchanged.
352
+
353
+ ---
354
+
355
+ ## Release Boundary and Artifact Terms
356
+
357
+ Different artifacts intentionally carry different permissions. Do not infer that the model-weight license automatically applies to separately published scientific artifacts or private system components.
358
+
359
+ | Artifact | Public status | Terms |
360
+ |---|---|---|
361
+ | **VeriLoop E2 model weights, tokenizer, configuration** | Public | **Apache License 2.0** |
362
+ | **Public vLLM launch / inference utilities** | Public | **Apache License 2.0** |
363
+ | **Benchmark results and evaluation evidence** | Public / separately published | Reuse permitted with attribution to **VeriLoop E2 / Libo Wang**; upstream benchmark assets retain their original terms |
364
+ | **Riemann ζ scientific artifact** | Public / separately published | **Research-only, non-commercial**; see artifact-specific terms |
365
+ | **Asymptotic Graviton Tomography artifact** | Public / separately published | **Research-only, non-commercial**; see artifact-specific terms |
366
+ | **Production VeriLoop Harness implementation** | Not included | Not licensed by this release |
367
+
368
+ The open model license does not disclose or license unpublished Harness orchestration, prompt compilation, verifier routing, private evidence-state schemas, repair arbitration, deployment infrastructure, private training data, or other non-distributed internal systems.
369
+
370
+ ---
371
+
372
+ ## Limitations
373
+
374
+ - VeriLoop E2 is a post-trained model component; the complete production VeriLoop Harness is not part of this release.
375
+ - Reported system benchmarks may depend on benchmark-native tools, sandbox behavior, evaluator versions, and the frozen Harness configuration described in the evidence package.
376
+ - The model can still produce incorrect code, invalid proofs, physically unsupported arguments, insecure commands, or incomplete analyses.
377
+ - A plausible-looking derivation is not equivalent to a verified result; domain-native verification remains necessary.
378
+ - Long-context performance depends on serving configuration, accelerator memory, KV-cache budget, and workload shape.
379
+ - Community-modified templates, stop rules, parsers, or client logic can materially change observed tool-use and reasoning behavior.
380
+ - Scientific demonstration artifacts have explicit scope boundaries and should not be generalized beyond the claims actually certified by their released evidence.
381
+
382
+ ---
383
+
384
+ ## Links
385
+
386
+ | Resource | Link |
387
+ |---|---|
388
+ | Hugging Face model | [https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2) |
389
+ | Technical report | `xxxx` |
390
+ | Evaluation evidence | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence) |
391
+ | GitHub | `xxxx` |
392
+ | Riemann ζ artifact | `xxxx` |
393
+ | Black-hole / Asymptotic Graviton Tomography artifact | `xxxx` |
394
+ | Zenodo DOI | `xxxx` |
395
+
396
+ Unresolved links in this table remain placeholders until their corresponding public artifacts are released. The Hugging Face model and evaluation-evidence links above are final public identifiers.
397
+
398
+ ---
399
+
400
+ ## Citation
401
+
402
+ If you use VeriLoop E2 in research, please cite the model release and the relevant evaluation or scientific artifact separately.
403
+
404
+ ```bibtex
405
+ @misc{wang2026veriloope2,
406
+ title = {VeriLoop E2: A 27B Post-Trained Model for Code, Mathematics, and Scientific Reasoning},
407
+ author = {Wang, Libo},
408
+ year = {2026},
409
+ note = {Tsinghua Shenzhen International Graduate School (SIGS)},
410
+ howpublished = {Open model release},
411
+ url = {https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2}
412
+ }
413
+ ```
414
+
415
+ For benchmark figures or evaluation records, attribution should identify **VeriLoop E2 / Libo Wang** and link to the public evaluation evidence package: [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence).
416
+
417
+ ---
418
+
419
+ ## Acknowledgements
420
+
421
+ VeriLoop E2 builds on **Qwen3.8-27B** and the broader open-source model-serving, evaluation, and scientific-computing ecosystem. We thank the communities behind Qwen, Transformers, vLLM, Safetensors, software-engineering benchmarks, mathematical evaluation suites, and reproducible scientific computation.
422
+
423
+ The model, benchmark evidence, and scientific artifacts are published with explicit boundaries so that capability claims can be inspected at the level at which they were actually produced and verified.