Instructions to use sora42y/Sev-2B-preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sora42y/Sev-2B-preview with Transformers:
# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("sora42y/Sev-2B-preview") model = AutoModelForCausalLM.from_pretrained("sora42y/Sev-2B-preview", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Sev-2B-preview
A 2B model that makes multiple-choice decisions in one forward pass, with no reasoning tokens. This is a preview; a full release and a tech report are coming.
- Base: Qwen3.5-2B-Base, all parameters fine-tuned · by Sora · weights under CC BY-NC 4.0
- Output: one probability per lettered option, read from the next-token logits after
Answer: - Fast mode: stop after layer 16 of 24, 1.5x faster, accuracy within about a point
Tree Forcing
A reasoning model (the teacher) writes a chain of thought for each training problem. We turn it into a tree of intermediate judgments and train the student on every node, with the true parents' conclusions written into the prompt: teacher forcing on the reasoning graph instead of on tokens. At inference only the final question is asked.
The released weights are the fine-tuned model blended with the base (alpha = 0.7, chosen on held-out selection sets).
Results
Every model scored on the same items by the same harness; accuracy.
| set (items) | Sev-2B-preview | decider-2b v11 | Strands Decider 2B v19 | Jeff-Qwen3.5-2B | Qwen3.5-2B-Base |
|---|---|---|---|---|---|
| Knights & Knaves diagnostic (1,600) | .486 | .236 | .189 | .221 | – |
| xkk, Knights & Knaves (612) | .495 | .260 | .261 | .283 | .198 |
| ProverQA (1,451) | .704 | .581 | .667 | .565 | .456 |
| LSAT-AR (61) | .426 | .426 | .295 | .344 | .230 |
| ZebraLogic mc (2,757) | .461 | .396 | .519 | .475 | .440 |
| LogiQA 2.0 (1,522) | .529 | .565 | .444 | .459 | .511 |
| MuSR (756) | .541 | .530 | .522 | .573 | .530 |
| BBH (5,508) | .502 | .512 | .450 | .617 | .453 |
| JevBench public (231) | .693 | .766 | .736 | .766 | .610 |
| TD, Typed Decisions (2,000) | .586 | .615 | .620 | .547 | .475 |
- Ahead of all three comparators (paired bootstrap, 95% interval above 0) on the Knights & Knaves diagnostic, xkk and ProverQA. Behind on JevBench and BBH.
- One training seed; the Knights & Knaves diagnostic comes from our own generator (held-out programs).
Use
# hf download sora42y/Sev-2B-preview sev_infer.py --local-dir .
from sev_infer import decide, load, load_exit_map
model, tok = load("sora42y/Sev-2B-preview")
problem = (
'Ticket: "I was charged twice for order 5521. Please refund one of the charges."\n\n'
"Question: Which team should handle this ticket?\n"
"Options:\n"
"A) billing: Charges, refunds and invoices.\n"
"B) shipping: Delivery dates and lost parcels.\n"
"C) technical: Log-in problems and app errors."
)
decide(model, tok, problem) # {"A": ..., "B": ..., "C": ...}
decide(model, tok, problem, exit_map=load_exit_map("sora42y/Sev-2B-preview", model)) # fast mode
- Prompt: the problem ending in its
Options:block, then\n\nAnswer:(added bydecide). No chat template. - Batching: right-pad and read each row at its last real token. Never left-pad: the linear-attention layers ignore the attention mask.
- Command line:
python sev_infer.py --model sora42y/Sev-2B-preview --problem-file problem.txt [--fast]
Fast mode
A100 80GB, batch size 1, eager PyTorch, bf16:
| decoder only | end to end | |
|---|---|---|
| full, 24 layers | 45.2 ms | 55.8 ms |
| exit after layer 16 | 29.9 ms | 38.1 ms |
The layer-16 state goes through a tuned lens (tuned_lens_L16.safetensors, fitted without labels) and the LM head.
Learn-then-Test certifies layer 16 on held-out items: at most 5% disagreement with the full model (measured 4.1%);
over eight evaluation sets the mean accuracy change is +0.1 points.
Calibration (optional)
Divide the letter logits by T before the softmax: T = 1.05 for most questions, 2.00 for ordinal ratings, 0.95 for
yes/no decisions about a stated case (plus +0.25 on the Yes logit). decide(..., temperature=T, yes_letter=..., yes_bias=...).
Data and limits
- 84,939 training items from 1,075 sources: public datasets, our own program generators and exam-style multiple
choice. Every source and its licence:
data_card.csv. Some sources are non-commercial, hence CC BY-NC. - Teacher: DeepSeek's API model; no teacher output is distributed.
- Evaluation items sharing a 13-gram with the training data were removed before scoring. Related training families
(LogiQA, LSAT, intent datasets) are listed in
details.md. - English only. Not for text generation, questions without options, or high-stakes decisions without human review.
- The full card, with every table, the training settings and all disclosures:
details.md.
Citation
@misc{yang2026sev2bpreview,
title = {Sev-2B-preview: a one-pass decision model trained with Tree Forcing},
author = {Haotian Yang},
year = {2026},
url = {https://huggingface.co/sora42y/Sev-2B-preview}
}
- Downloads last month
- 20
Model tree for sora42y/Sev-2B-preview
Base model
Qwen/Qwen3.5-2B-Base