Text Generation
Transformers
Safetensors
English
Chinese
qwen3_5
image-text-to-text
conversational
veriloop
post-training
coding-agent
software-engineering
mathematical-reasoning
scientific-reasoning
tool-use
long-context
vllm
apache-2.0
Eval Results
Instructions to use tsinghua-sigs-robot-lab/VeriLoop-E2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tsinghua-sigs-robot-lab/VeriLoop-E2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tsinghua-sigs-robot-lab/VeriLoop-E2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("tsinghua-sigs-robot-lab/VeriLoop-E2") model = AutoModelForMultimodalLM.from_pretrained("tsinghua-sigs-robot-lab/VeriLoop-E2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tsinghua-sigs-robot-lab/VeriLoop-E2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tsinghua-sigs-robot-lab/VeriLoop-E2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tsinghua-sigs-robot-lab/VeriLoop-E2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2
- SGLang
How to use tsinghua-sigs-robot-lab/VeriLoop-E2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tsinghua-sigs-robot-lab/VeriLoop-E2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tsinghua-sigs-robot-lab/VeriLoop-E2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tsinghua-sigs-robot-lab/VeriLoop-E2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tsinghua-sigs-robot-lab/VeriLoop-E2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tsinghua-sigs-robot-lab/VeriLoop-E2 with Docker Model Runner:
docker model run hf.co/tsinghua-sigs-robot-lab/VeriLoop-E2
Upload README.md
Browse files
README.md
CHANGED
|
@@ -1,3 +1,423 @@
|
|
| 1 |
---
|
|
|
|
|
|
|
| 2 |
license: apache-2.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
library_name: transformers
|
| 3 |
+
pipeline_tag: text-generation
|
| 4 |
license: apache-2.0
|
| 5 |
+
base_model:
|
| 6 |
+
- Qwen/Qwen3.8-27B
|
| 7 |
+
base_model_relation: finetune
|
| 8 |
+
language:
|
| 9 |
+
- en
|
| 10 |
+
- zh
|
| 11 |
+
tags:
|
| 12 |
+
- veriloop
|
| 13 |
+
- post-training
|
| 14 |
+
- coding-agent
|
| 15 |
+
- software-engineering
|
| 16 |
+
- mathematical-reasoning
|
| 17 |
+
- scientific-reasoning
|
| 18 |
+
- tool-use
|
| 19 |
+
- long-context
|
| 20 |
+
- safetensors
|
| 21 |
+
- vllm
|
| 22 |
+
- apache-2.0
|
| 23 |
---
|
| 24 |
+
|
| 25 |
+
<div align="center">
|
| 26 |
+
|
| 27 |
+
<img src="https://huggingface.co/tsinghua-sigs-robot-lab/veriloop-coder-e1/resolve/main/veriloop_logo.png" width="154" alt="VeriLoop logo">
|
| 28 |
+
|
| 29 |
+
# VeriLoop E2
|
| 30 |
+
|
| 31 |
+
### Open 27B Post-Trained Model for Code Agents, Mathematical Reasoning, and Scientific Problem Solving
|
| 32 |
+
|
| 33 |
+
**Built on Qwen3.8-27B · 262K native context · Apache License 2.0**
|
| 34 |
+
|
| 35 |
+
**Developed by Tsinghua SIGS Robot Lab · Libo Wang**
|
| 36 |
+
|
| 37 |
+
<p>
|
| 38 |
+
<a href="https://www.apache.org/licenses/LICENSE-2.0"><img src="https://img.shields.io/badge/License-Apache--2.0-2F80ED?style=flat-square" alt="License: Apache 2.0"></a>
|
| 39 |
+
<img src="https://img.shields.io/badge/Base-Qwen3.8--27B-5B5BD6?style=flat-square" alt="Base: Qwen3.8-27B">
|
| 40 |
+
<img src="https://img.shields.io/badge/Context-262K-1F6FEB?style=flat-square" alt="Native context: 262K">
|
| 41 |
+
<img src="https://img.shields.io/badge/Serving-vLLM%200.17.0-0A7F6F?style=flat-square" alt="Serving: vLLM 0.17.0">
|
| 42 |
+
<img src="https://img.shields.io/badge/Stage-Post--Training-7C3AED?style=flat-square" alt="Stage: Post-Training">
|
| 43 |
+
</p>
|
| 44 |
+
|
| 45 |
+
<p>
|
| 46 |
+
<a href="https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2"><strong>Model</strong></a> ·
|
| 47 |
+
<a href="xxxx"><strong>Technical Report</strong></a> ·
|
| 48 |
+
<a href="https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence"><strong>Evaluation Evidence</strong></a> ·
|
| 49 |
+
<a href="xxxx"><strong>Scientific Artifacts</strong></a> ·
|
| 50 |
+
<a href="xxxx"><strong>GitHub</strong></a> ·
|
| 51 |
+
<a href="xxxx"><strong>Zenodo</strong></a>
|
| 52 |
+
</p>
|
| 53 |
+
|
| 54 |
+
</div>
|
| 55 |
+
|
| 56 |
+
---
|
| 57 |
+
|
| 58 |
+
## Overview
|
| 59 |
+
|
| 60 |
+
**VeriLoop E2** is an open 27B post-trained model built on **Qwen3.8-27B**. It is designed as the model component of the broader VeriLoop reasoning system, with emphasis on three workloads: **software-engineering agents, mathematical reasoning, and scientific problem solving**.
|
| 61 |
+
|
| 62 |
+
The system separates generative intelligence from verification authority. **VeriLoop E2** is responsible for proposal generation, abstraction, diagnosis, repair hypotheses, and structured reasoning. The **VeriLoop Harness** governs evidence admission, deterministic checks, external verification, commit/rollback, stopping, and evidence-state persistence. The model therefore does not self-certify its own progress.
|
| 63 |
+
|
| 64 |
+
This release focuses on the model weights, public inference path, evaluation record, and public functional description of the Harness. The production Harness implementation itself is **not included** in this repository.
|
| 65 |
+
|
| 66 |
+
### Release highlights
|
| 67 |
+
|
| 68 |
+
- **27B open-weight post-trained model** derived from Qwen3.8-27B.
|
| 69 |
+
- **262,144-token native context window** in the released tokenizer configuration.
|
| 70 |
+
- Strong release results across nine code-agent, mathematics, and science benchmarks, including **76.2% SWE-bench Pro**, **88.8% Terminal-Bench 2.1**, **98.3% AIME 2026**, **93.9% GPQA Diamond**, and **89.6% Apex 2025**.
|
| 71 |
+
- A reproducible scientific-reasoning program built around verifier-governed recurrence rather than unconstrained retry.
|
| 72 |
+
- Two public-facing scientific demonstrations: a strict finite-dimensional **Riemann ζ zero-proportion certificate at 67.350003708785593%**, and **Asymptotic Graviton Tomography** for the black-hole information problem.
|
| 73 |
+
- OpenAI-compatible serving through **vLLM 0.17.0** with a validated 131,072-token serving configuration.
|
| 74 |
+
|
| 75 |
+
---
|
| 76 |
+
|
| 77 |
+
## Model Summary
|
| 78 |
+
|
| 79 |
+
| Property | VeriLoop E2 |
|
| 80 |
+
|---|---|
|
| 81 |
+
| Model family | VeriLoop E2 |
|
| 82 |
+
| Base model | Qwen3.8-27B |
|
| 83 |
+
| Parameter class | 27B |
|
| 84 |
+
| HF architecture class | `Qwen3_5ForConditionalGeneration` |
|
| 85 |
+
| Training stage | Post-Training |
|
| 86 |
+
| Primary domains | Code agents, software engineering, mathematics, scientific reasoning |
|
| 87 |
+
| Public post-training corpus accounting | 1,841,831 records |
|
| 88 |
+
| Native context length | 262,144 tokens |
|
| 89 |
+
| Validated vLLM serving length | 131,072 tokens |
|
| 90 |
+
| Tokenizer class | `Qwen2Tokenizer` |
|
| 91 |
+
| Weight format | `safetensors` |
|
| 92 |
+
| Languages | English, Chinese |
|
| 93 |
+
| Recommended serving engine | vLLM 0.17.0 |
|
| 94 |
+
| Model-weight license | Apache License 2.0 |
|
| 95 |
+
| Release year | 2026 |
|
| 96 |
+
|
| 97 |
+
The post-training mix spans repository-level software engineering, terminal and tool use, mathematical reasoning, scientific reasoning, verifier-sensitive repair, and recurrence-oriented training. Exact data construction, filtering, and training methodology are documented in the technical report rather than duplicated here.
|
| 98 |
+
|
| 99 |
+
---
|
| 100 |
+
|
| 101 |
+
## Benchmark Results
|
| 102 |
+
|
| 103 |
+
The README reports the **frozen release scores** for VeriLoop E2. Agentic benchmarks use the E2 checkpoint inside the frozen evaluation workflow, including the internal VeriLoop Harness where required by the task, benchmark-native tools, and the benchmark's official or designated evaluator. Exact per-benchmark protocols, task-level outputs, evaluator receipts, and integrity metadata are published separately in the **Evaluation Evidence** package.
|
| 104 |
+
|
| 105 |
+
> **Attribution boundary.** The reported results characterize the evaluated E2 system configuration. They should not be interpreted as evidence that an untouched Qwen3.8-27B base checkpoint, or the E2 checkpoint outside the evaluated runtime, reproduces the same numbers.
|
| 106 |
+
|
| 107 |
+
<p align="center">
|
| 108 |
+
<img src="assets/veriloop_e2_benchmark_grid.png" width="100%" alt="VeriLoop E2 benchmark comparison across nine public benchmarks">
|
| 109 |
+
</p>
|
| 110 |
+
|
| 111 |
+
<p align="center">
|
| 112 |
+
<sub><strong>Figure 1.</strong> VeriLoop E2 release snapshot across nine public benchmarks. Higher is better. Provider colors are fixed across panels; exact public model variants are shown in the comparison tables below. Full protocol and source provenance: <a href="https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence">Evaluation Evidence</a>.</sub>
|
| 113 |
+
</p>
|
| 114 |
+
|
| 115 |
+
### Code and agentic benchmarks
|
| 116 |
+
|
| 117 |
+
| Benchmark | **VeriLoop E2** | OpenAI | Anthropic | Kimi | GLM | Qwen | DeepSeek |
|
| 118 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 119 |
+
| **SWE-bench Pro** | **[76.2](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/swe-bench-pro)** | GPT-5.6 Sol 64.6 | Claude Fable 5.1 81.2 | — | GLM-5.2 Max 62.1 | Qwen3.8-Max 67.7 | DeepSeek V4 Pro Max 55.4 |
|
| 120 |
+
| **Terminal-Bench 2.1** | **[88.8](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-2.1)** | GPT-5.6 Sol 88.8 | — | Kimi K3 88.3 | GLM-5.3 88.2 | Qwen3.8-Max 86.6 | DeepSeek V4 Pro 87.9 |
|
| 121 |
+
| **DeepSWE v1.1** | **[64.6](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/deepswe-1.1)** | GPT-5.6 Sol 72.7 | Claude Fable 5 69.7 | Kimi K3 67.5 | GLM-5.3 66.9 | — | DeepSeek V4 Pro 62.7 |
|
| 122 |
+
| **Terminal-Bench 3.0** | **[29.7](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-3.0)** | GPT-5.6 Sol 34.6 | Claude Fable 5 33.7 | Kimi K3 17.4 | GLM-5.3 28.3 | — | — |
|
| 123 |
+
| **Terminal-Bench 4.0** | **[37.9](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-4.0)** | GPT-6 Astra 59.6 | Claude Fable 5.1 55.1 | — | GLM-5.3 41.8 | — | — |
|
| 124 |
+
| **SWE-Marathon v1.1** | **45.0** | GPT-5.6 Sol 42.5 | Claude Opus 4.8 48.8 | Kimi K3 48.1 | GLM-5.3 42.5 | — | — |
|
| 125 |
+
|
| 126 |
+
### Mathematics and science benchmarks
|
| 127 |
+
|
| 128 |
+
| Benchmark | **VeriLoop E2** | OpenAI | Anthropic | Kimi | GLM | DeepSeek | Gemini |
|
| 129 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 130 |
+
| **AIME 2026** | **[98.3](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/aime-2026)** | GPT-5.5 100.0 | Claude Opus 4.8 100.0 | Kimi K3 97.0 | — | DeepSeek V4 Pro 97.0 | Gemini 3.1 Pro 98.0 |
|
| 131 |
+
| **GPQA Diamond** | **[93.9](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/gpqa-diamond)** | GPT-5.6 Sol 94.1 | Claude Fable 5 92.6 | Kimi K3 93.5 | GLM-5.2 Max 91.2 | — | — |
|
| 132 |
+
| **Apex 2025** | **[89.6](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/apex-2025)** | GPT-5.5 80.0 | Claude Opus 4.8 81.0 | Kimi K3 66.0 | — | DeepSeek V4 Pro 28.0 | Gemini 3.1 Pro 61.0 |
|
| 133 |
+
|
| 134 |
+
A dash means that the release figure does not include a public comparison point for that provider on that benchmark. Each linked VeriLoop E2 score above resolves directly to its benchmark-specific public evidence directory. External reference values mirror the frozen comparison set used in Figure 1; harness notes and protocol caveats are retained in the evaluation ledger rather than duplicated here. **SWE-Marathon v1.1 (45.0)** remains reported on the model page, but is intentionally not represented by a Hugging Face Native Benchmark `.eval_results` entry because a stable HF benchmark registration/task identifier is not currently available.
|
| 135 |
+
|
| 136 |
+
### Evaluation evidence
|
| 137 |
+
|
| 138 |
+
The public evidence package is intended to make the benchmark record inspectable rather than merely declarative. Where available, each task record binds:
|
| 139 |
+
|
| 140 |
+
```text
|
| 141 |
+
task identity
|
| 142 |
+
↓
|
| 143 |
+
model / system output
|
| 144 |
+
↓
|
| 145 |
+
benchmark-native execution or evaluator record
|
| 146 |
+
↓
|
| 147 |
+
score / pass-fail decision
|
| 148 |
+
↓
|
| 149 |
+
integrity metadata and provenance
|
| 150 |
+
```
|
| 151 |
+
|
| 152 |
+
**Evidence repository:** [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence)
|
| 153 |
+
|
| 154 |
+
| Benchmark | Public evaluation source |
|
| 155 |
+
|---|---|
|
| 156 |
+
| SWE-bench Pro | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/swe-bench-pro](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/swe-bench-pro) |
|
| 157 |
+
| Terminal-Bench 2.1 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-2.1](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-2.1) |
|
| 158 |
+
| DeepSWE v1.1 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/deepswe-1.1](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/deepswe-1.1) |
|
| 159 |
+
| Terminal-Bench 3.0 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-3.0](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-3.0) |
|
| 160 |
+
| Terminal-Bench 4.0 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-4.0](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/terminal-bench-4.0) |
|
| 161 |
+
| AIME 2026 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/aime-2026](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/aime-2026) |
|
| 162 |
+
| GPQA Diamond | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/gpqa-diamond](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/gpqa-diamond) |
|
| 163 |
+
| Apex 2025 | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/apex-2025](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence/tree/main/apex-2025) |
|
| 164 |
+
|
| 165 |
+
The structured Hugging Face evaluation descriptors for these eight registered benchmarks are published in the model repository under [`/.eval_results/`](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2/tree/main/.eval_results).
|
| 166 |
+
|
| 167 |
+
---
|
| 168 |
+
|
| 169 |
+
## VeriLoop Harness
|
| 170 |
+
|
| 171 |
+
VeriLoop is not designed around the idea that a model should announce its own improvement. The Harness treats model output as a **candidate state** that must earn admission through external evidence.
|
| 172 |
+
|
| 173 |
+
The current public abstraction is **VeriLoop-Governed Recurrence (VGR)**:
|
| 174 |
+
|
| 175 |
+
```text
|
| 176 |
+
Current request
|
| 177 |
+
↓
|
| 178 |
+
Contract compilation
|
| 179 |
+
↓
|
| 180 |
+
VeriLoop E2 proposes a candidate
|
| 181 |
+
↓
|
| 182 |
+
External / deterministic verification
|
| 183 |
+
↓
|
| 184 |
+
Protected evidence state comparison
|
| 185 |
+
├── no protected regression + at least one strict improvement → COMMIT
|
| 186 |
+
├── otherwise → ROLLBACK
|
| 187 |
+
└── zero-rank certificate → STOP
|
| 188 |
+
↓
|
| 189 |
+
Verified evidence becomes the next recurrence state
|
| 190 |
+
```
|
| 191 |
+
|
| 192 |
+
The key boundary is deliberate:
|
| 193 |
+
|
| 194 |
+
- **Model authority:** propose, reason, abstract, diagnose, synthesize, repair.
|
| 195 |
+
- **Harness authority:** admit evidence, execute deterministic checks, verify, commit, roll back, stop, and persist verified state.
|
| 196 |
+
- **Benchmark / domain authority:** define task truth through native evaluators, tests, formal checks, numerical certificates, or other domain-specific validators.
|
| 197 |
+
|
| 198 |
+
This architecture is intended to preserve capability while preventing self-reported success from becoming system state. In software engineering, that means tests and execution receipts dominate plausible-looking patches. In mathematics and physics, it means a retained derivation must survive the relevant symbolic, numerical, or formal checks before it is promoted.
|
| 199 |
+
|
| 200 |
+
The production implementation contains private orchestration, routing, thresholds, prompt compilation, evidence-state machinery, repair arbitration, and deployment controls. Those implementation details are not part of this open model release. The README exposes the **functional contract**, not the proprietary runtime.
|
| 201 |
+
|
| 202 |
+
---
|
| 203 |
+
|
| 204 |
+
## Public 14-Rule Engineering Contract
|
| 205 |
+
|
| 206 |
+
The public Golden Rules are the model-visible execution discipline used to keep long-horizon work bounded, testable, and auditable.
|
| 207 |
+
|
| 208 |
+
| # | Rule | Public meaning |
|
| 209 |
+
|---:|---|---|
|
| 210 |
+
| 1 | **Current Request Supremacy** | The current request and exact output contract override stale memory, templates, and unrelated context. |
|
| 211 |
+
| 2 | **Evidence Before Escalation** | Search, tools, reverse analysis, or repair are triggered by concrete missing evidence or observed failure, not instinct. |
|
| 212 |
+
| 3 | **Read Before Rewrite** | Inspect the relevant entry points, interfaces, tests, conventions, and failure signals before editing. |
|
| 213 |
+
| 4 | **Minimal Sufficient Implementation** | Produce the smallest complete artifact that satisfies the task and preserves required interfaces. |
|
| 214 |
+
| 5 | **Surgical Repair, Not Blind Regeneration** | Repair the broken invariant; broaden the rewrite only when evidence shows local repair is insufficient. |
|
| 215 |
+
| 6 | **Intent Tests Beat Cosmetic Tests** | Syntax and formatting matter, but functional intent is the decisive acceptance criterion. |
|
| 216 |
+
| 7 | **Fail Loud, Never Fake Success** | Unknowns, skipped checks, degraded states, and failures remain explicit; unexecuted validation is never reported as success. |
|
| 217 |
+
| 8 | **Deterministic Logic Belongs in Code** | Parsing, scoring, structural checks, transformations, and reproducible validation should be deterministic whenever possible. |
|
| 218 |
+
| 9 | **Budget Is a First-Class Contract** | Token, tool, time, and compute budgets are part of the task contract rather than afterthoughts. |
|
| 219 |
+
| 10 | **Tool Use Must Be Typed and Accountable** | Every tool action has a trigger, expected output, and defined downstream consumer. |
|
| 220 |
+
| 11 | **Checkpoint Long Tasks** | Persist useful candidate state, validation receipts, repair records, and progress across long-running work. |
|
| 221 |
+
| 12 | **Follow Local Conventions** | Respect repository, benchmark, filename, API, language, and artifact conventions unless the request explicitly changes them. |
|
| 222 |
+
| 13 | **One Final Deliverable, Full Evidence Bundle** | Select one deliverable while retaining the evidence needed to explain and reproduce why it was selected. |
|
| 223 |
+
| 14 | **Separate Core Discipline from Domain Overlays** | Domain- or benchmark-specific rules apply only when relevant and cannot override the current request. |
|
| 224 |
+
|
| 225 |
+
---
|
| 226 |
+
|
| 227 |
+
## Scientific Demonstrations
|
| 228 |
+
|
| 229 |
+
The scientific demonstrations are not presented as isolated chat transcripts. They are examples of how the E2 model and the internal Harness can divide a difficult research problem into candidate derivations, falsifiable subclaims, executable checks, and retained evidence.
|
| 230 |
+
|
| 231 |
+
### Riemann ζ: 67.350003708785593% strict finite-dimensional certificate
|
| 232 |
+
|
| 233 |
+
The released Riemann artifact reports a frozen assembly value of
|
| 234 |
+
|
| 235 |
+
\[
|
| 236 |
+
\kappa = 67.350003708785593\%.
|
| 237 |
+
\]
|
| 238 |
+
|
| 239 |
+
The current strict package closes **3/3 local inequalities**, resolves **327/327 difficult wells**, executes **190,375,830 strict branch-and-bound nodes**, and passes the final **exact rational assembly** check.
|
| 240 |
+
|
| 241 |
+
What this result **does** establish within the released artifact is a strict finite-dimensional computer-assisted certificate under its stated analytic setup and imported assumptions. What it **does not** establish is equally important:
|
| 242 |
+
|
| 243 |
+
- it is **not a proof of the Riemann Hypothesis**;
|
| 244 |
+
- it is **not yet an end-to-end Lean/nanoda kernel proof** of the complete upstream analytic chain;
|
| 245 |
+
- imported analytic normalization steps must remain clearly separated from the finite-dimensional certificate until the formal bridge and complete replay are closed.
|
| 246 |
+
|
| 247 |
+
The public package therefore emphasizes claim discipline, reproducibility, and certificate structure rather than treating a numerical percentage as a substitute for mathematical provenance.
|
| 248 |
+
|
| 249 |
+
**Artifact:** `xxxx` · **Technical note:** `xxxx` · **Zenodo:** `xxxx`
|
| 250 |
+
|
| 251 |
+
### Black-hole information problem: Asymptotic Graviton Tomography
|
| 252 |
+
|
| 253 |
+
**Asymptotic Graviton Tomography** is the second scientific reasoning demonstration. It studies an information-reconstruction route through asymptotic gravitational observables, using the Harness to separate retained derivations from rejected or insufficiently supported branches.
|
| 254 |
+
|
| 255 |
+
The public claim is intentionally bounded: this is a **research demonstration of a verifier-governed theoretical-physics derivation**, not a declaration that the black-hole information paradox has been solved. The artifact is intended to expose the derivation structure, assumptions, checks, and remaining theoretical boundaries clearly enough for external scientific criticism.
|
| 256 |
+
|
| 257 |
+
**Artifact:** `xxxx` · **Technical note:** `xxxx`
|
| 258 |
+
|
| 259 |
+
> The Riemann and black-hole demo artifacts are released separately under **research-only, non-commercial terms**. They are not covered by the Apache-2.0 grant for the model weights and public inference utilities unless a specific file explicitly says otherwise.
|
| 260 |
+
|
| 261 |
+
---
|
| 262 |
+
|
| 263 |
+
## Inference
|
| 264 |
+
|
| 265 |
+
### Recommended environment
|
| 266 |
+
|
| 267 |
+
The following stack is the validated reference environment for the public serving path:
|
| 268 |
+
|
| 269 |
+
| Component | Version / setting |
|
| 270 |
+
|---|---|
|
| 271 |
+
| Python | 3.12.x |
|
| 272 |
+
| vLLM | 0.17.0 |
|
| 273 |
+
| PyTorch | 2.10.0 + CUDA 12.9 build |
|
| 274 |
+
| CUDA runtime | 12.9 |
|
| 275 |
+
| Transformers | 4.57.6 |
|
| 276 |
+
| Triton | 3.6.0 |
|
| 277 |
+
| Dtype | `bfloat16` |
|
| 278 |
+
| Validated serving context | 131,072 tokens |
|
| 279 |
+
|
| 280 |
+
The released tokenizer advertises a native maximum length of **262,144 tokens**. The 131,072-token value above is the **validated public serving configuration**, not a redefinition of the model's native context length. Longer serving windows require appropriate accelerator memory and KV-cache planning.
|
| 281 |
+
|
| 282 |
+
### vLLM server
|
| 283 |
+
|
| 284 |
+
Use the tokenizer and chat template shipped with the model repository.
|
| 285 |
+
|
| 286 |
+
```bash
|
| 287 |
+
python -m pip install "vllm==0.17.0"
|
| 288 |
+
|
| 289 |
+
MODEL="<MODEL_PATH_OR_HF_ID>"
|
| 290 |
+
|
| 291 |
+
vllm serve "${MODEL}" \
|
| 292 |
+
--served-model-name veriloop-e2 \
|
| 293 |
+
--dtype bfloat16 \
|
| 294 |
+
--model-impl vllm \
|
| 295 |
+
--language-model-only \
|
| 296 |
+
--max-model-len 131072 \
|
| 297 |
+
--max-num-seqs 16 \
|
| 298 |
+
--gpu-memory-utilization 0.92 \
|
| 299 |
+
--generation-config vllm \
|
| 300 |
+
--disable-uvicorn-access-log \
|
| 301 |
+
--host 127.0.0.1 \
|
| 302 |
+
--port 8001
|
| 303 |
+
```
|
| 304 |
+
|
| 305 |
+
### OpenAI-compatible request
|
| 306 |
+
|
| 307 |
+
```bash
|
| 308 |
+
curl http://127.0.0.1:8001/v1/chat/completions \
|
| 309 |
+
-H "Content-Type: application/json" \
|
| 310 |
+
-d '{
|
| 311 |
+
"model": "veriloop-e2",
|
| 312 |
+
"messages": [
|
| 313 |
+
{"role": "user", "content": "Reply exactly: 8001_OK"}
|
| 314 |
+
],
|
| 315 |
+
"temperature": 0,
|
| 316 |
+
"max_tokens": 16,
|
| 317 |
+
"stream": false
|
| 318 |
+
}'
|
| 319 |
+
```
|
| 320 |
+
|
| 321 |
+
A validated smoke test returns:
|
| 322 |
+
|
| 323 |
+
```text
|
| 324 |
+
8001_OK
|
| 325 |
+
```
|
| 326 |
+
|
| 327 |
+
### Python client
|
| 328 |
+
|
| 329 |
+
```python
|
| 330 |
+
from openai import OpenAI
|
| 331 |
+
|
| 332 |
+
client = OpenAI(
|
| 333 |
+
base_url="http://127.0.0.1:8001/v1",
|
| 334 |
+
api_key="EMPTY",
|
| 335 |
+
)
|
| 336 |
+
|
| 337 |
+
response = client.chat.completions.create(
|
| 338 |
+
model="veriloop-e2",
|
| 339 |
+
messages=[
|
| 340 |
+
{"role": "user", "content": "Explain why rollback matters in verifier-governed reasoning."}
|
| 341 |
+
],
|
| 342 |
+
temperature=0.2,
|
| 343 |
+
max_tokens=1024,
|
| 344 |
+
)
|
| 345 |
+
|
| 346 |
+
print(response.choices[0].message.content)
|
| 347 |
+
```
|
| 348 |
+
|
| 349 |
+
### Protocol note
|
| 350 |
+
|
| 351 |
+
For tool use or long multi-turn reasoning, treat the repository-shipped tokenizer and chat template as part of the model protocol. Replacing role delimiters, tool-call syntax, stop semantics, or reasoning-history behavior can change observed system behavior even when the weights are unchanged.
|
| 352 |
+
|
| 353 |
+
---
|
| 354 |
+
|
| 355 |
+
## Release Boundary and Artifact Terms
|
| 356 |
+
|
| 357 |
+
Different artifacts intentionally carry different permissions. Do not infer that the model-weight license automatically applies to separately published scientific artifacts or private system components.
|
| 358 |
+
|
| 359 |
+
| Artifact | Public status | Terms |
|
| 360 |
+
|---|---|---|
|
| 361 |
+
| **VeriLoop E2 model weights, tokenizer, configuration** | Public | **Apache License 2.0** |
|
| 362 |
+
| **Public vLLM launch / inference utilities** | Public | **Apache License 2.0** |
|
| 363 |
+
| **Benchmark results and evaluation evidence** | Public / separately published | Reuse permitted with attribution to **VeriLoop E2 / Libo Wang**; upstream benchmark assets retain their original terms |
|
| 364 |
+
| **Riemann ζ scientific artifact** | Public / separately published | **Research-only, non-commercial**; see artifact-specific terms |
|
| 365 |
+
| **Asymptotic Graviton Tomography artifact** | Public / separately published | **Research-only, non-commercial**; see artifact-specific terms |
|
| 366 |
+
| **Production VeriLoop Harness implementation** | Not included | Not licensed by this release |
|
| 367 |
+
|
| 368 |
+
The open model license does not disclose or license unpublished Harness orchestration, prompt compilation, verifier routing, private evidence-state schemas, repair arbitration, deployment infrastructure, private training data, or other non-distributed internal systems.
|
| 369 |
+
|
| 370 |
+
---
|
| 371 |
+
|
| 372 |
+
## Limitations
|
| 373 |
+
|
| 374 |
+
- VeriLoop E2 is a post-trained model component; the complete production VeriLoop Harness is not part of this release.
|
| 375 |
+
- Reported system benchmarks may depend on benchmark-native tools, sandbox behavior, evaluator versions, and the frozen Harness configuration described in the evidence package.
|
| 376 |
+
- The model can still produce incorrect code, invalid proofs, physically unsupported arguments, insecure commands, or incomplete analyses.
|
| 377 |
+
- A plausible-looking derivation is not equivalent to a verified result; domain-native verification remains necessary.
|
| 378 |
+
- Long-context performance depends on serving configuration, accelerator memory, KV-cache budget, and workload shape.
|
| 379 |
+
- Community-modified templates, stop rules, parsers, or client logic can materially change observed tool-use and reasoning behavior.
|
| 380 |
+
- Scientific demonstration artifacts have explicit scope boundaries and should not be generalized beyond the claims actually certified by their released evidence.
|
| 381 |
+
|
| 382 |
+
---
|
| 383 |
+
|
| 384 |
+
## Links
|
| 385 |
+
|
| 386 |
+
| Resource | Link |
|
| 387 |
+
|---|---|
|
| 388 |
+
| Hugging Face model | [https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2](https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2) |
|
| 389 |
+
| Technical report | `xxxx` |
|
| 390 |
+
| Evaluation evidence | [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence) |
|
| 391 |
+
| GitHub | `xxxx` |
|
| 392 |
+
| Riemann ζ artifact | `xxxx` |
|
| 393 |
+
| Black-hole / Asymptotic Graviton Tomography artifact | `xxxx` |
|
| 394 |
+
| Zenodo DOI | `xxxx` |
|
| 395 |
+
|
| 396 |
+
Unresolved links in this table remain placeholders until their corresponding public artifacts are released. The Hugging Face model and evaluation-evidence links above are final public identifiers.
|
| 397 |
+
|
| 398 |
+
---
|
| 399 |
+
|
| 400 |
+
## Citation
|
| 401 |
+
|
| 402 |
+
If you use VeriLoop E2 in research, please cite the model release and the relevant evaluation or scientific artifact separately.
|
| 403 |
+
|
| 404 |
+
```bibtex
|
| 405 |
+
@misc{wang2026veriloope2,
|
| 406 |
+
title = {VeriLoop E2: A 27B Post-Trained Model for Code, Mathematics, and Scientific Reasoning},
|
| 407 |
+
author = {Wang, Libo},
|
| 408 |
+
year = {2026},
|
| 409 |
+
note = {Tsinghua Shenzhen International Graduate School (SIGS)},
|
| 410 |
+
howpublished = {Open model release},
|
| 411 |
+
url = {https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2}
|
| 412 |
+
}
|
| 413 |
+
```
|
| 414 |
+
|
| 415 |
+
For benchmark figures or evaluation records, attribution should identify **VeriLoop E2 / Libo Wang** and link to the public evaluation evidence package: [https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence](https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence).
|
| 416 |
+
|
| 417 |
+
---
|
| 418 |
+
|
| 419 |
+
## Acknowledgements
|
| 420 |
+
|
| 421 |
+
VeriLoop E2 builds on **Qwen3.8-27B** and the broader open-source model-serving, evaluation, and scientific-computing ecosystem. We thank the communities behind Qwen, Transformers, vLLM, Safetensors, software-engineering benchmarks, mathematical evaluation suites, and reproducible scientific computation.
|
| 422 |
+
|
| 423 |
+
The model, benchmark evidence, and scientific artifacts are published with explicit boundaries so that capability claims can be inspected at the level at which they were actually produced and verified.
|