Sev-2B-preview

A 2B model that makes multiple-choice decisions in one forward pass, with no reasoning tokens. This is a preview; a full release and a tech report are coming.

  • Base: Qwen3.5-2B-Base, all parameters fine-tuned · by Sora · weights under CC BY-NC 4.0
  • Output: one probability per lettered option, read from the next-token logits after Answer:
  • Fast mode: stop after layer 16 of 24, 1.5x faster, accuracy within about a point

Tree Forcing

A reasoning model (the teacher) writes a chain of thought for each training problem. We turn it into a tree of intermediate judgments and train the student on every node, with the true parents' conclusions written into the prompt: teacher forcing on the reasoning graph instead of on tokens. At inference only the final question is asked.

The released weights are the fine-tuned model blended with the base (alpha = 0.7, chosen on held-out selection sets).

Results

Every model scored on the same items by the same harness; accuracy.

set (items) Sev-2B-preview decider-2b v11 Strands Decider 2B v19 Jeff-Qwen3.5-2B Qwen3.5-2B-Base
Knights & Knaves diagnostic (1,600) .486 .236 .189 .221 –
xkk, Knights & Knaves (612) .495 .260 .261 .283 .198
ProverQA (1,451) .704 .581 .667 .565 .456
LSAT-AR (61) .426 .426 .295 .344 .230
ZebraLogic mc (2,757) .461 .396 .519 .475 .440
LogiQA 2.0 (1,522) .529 .565 .444 .459 .511
MuSR (756) .541 .530 .522 .573 .530
BBH (5,508) .502 .512 .450 .617 .453
JevBench public (231) .693 .766 .736 .766 .610
TD, Typed Decisions (2,000) .586 .615 .620 .547 .475
  • Ahead of all three comparators (paired bootstrap, 95% interval above 0) on the Knights & Knaves diagnostic, xkk and ProverQA. Behind on JevBench and BBH.
  • One training seed; the Knights & Knaves diagnostic comes from our own generator (held-out programs).

Use

# hf download sora42y/Sev-2B-preview sev_infer.py --local-dir .
from sev_infer import decide, load, load_exit_map

model, tok = load("sora42y/Sev-2B-preview")
problem = (
    'Ticket: "I was charged twice for order 5521. Please refund one of the charges."\n\n'
    "Question: Which team should handle this ticket?\n"
    "Options:\n"
    "A) billing: Charges, refunds and invoices.\n"
    "B) shipping: Delivery dates and lost parcels.\n"
    "C) technical: Log-in problems and app errors."
)
decide(model, tok, problem)                                                  # {"A": ..., "B": ..., "C": ...}
decide(model, tok, problem, exit_map=load_exit_map("sora42y/Sev-2B-preview", model))   # fast mode
  • Prompt: the problem ending in its Options: block, then \n\nAnswer: (added by decide). No chat template.
  • Batching: right-pad and read each row at its last real token. Never left-pad: the linear-attention layers ignore the attention mask.
  • Command line: python sev_infer.py --model sora42y/Sev-2B-preview --problem-file problem.txt [--fast]

Fast mode

A100 80GB, batch size 1, eager PyTorch, bf16:

decoder only end to end
full, 24 layers 45.2 ms 55.8 ms
exit after layer 16 29.9 ms 38.1 ms

The layer-16 state goes through a tuned lens (tuned_lens_L16.safetensors, fitted without labels) and the LM head. Learn-then-Test certifies layer 16 on held-out items: at most 5% disagreement with the full model (measured 4.1%); over eight evaluation sets the mean accuracy change is +0.1 points.

Calibration (optional)

Divide the letter logits by T before the softmax: T = 1.05 for most questions, 2.00 for ordinal ratings, 0.95 for yes/no decisions about a stated case (plus +0.25 on the Yes logit). decide(..., temperature=T, yes_letter=..., yes_bias=...).

Data and limits

  • 84,939 training items from 1,075 sources: public datasets, our own program generators and exam-style multiple choice. Every source and its licence: data_card.csv. Some sources are non-commercial, hence CC BY-NC.
  • Teacher: DeepSeek's API model; no teacher output is distributed.
  • Evaluation items sharing a 13-gram with the training data were removed before scoring. Related training families (LogiQA, LSAT, intent datasets) are listed in details.md.
  • English only. Not for text generation, questions without options, or high-stakes decisions without human review.
  • The full card, with every table, the training settings and all disclosures: details.md.

Citation

@misc{yang2026sev2bpreview,
  title  = {Sev-2B-preview: a one-pass decision model trained with Tree Forcing},
  author = {Haotian Yang},
  year   = {2026},
  url    = {https://huggingface.co/sora42y/Sev-2B-preview}
}
Downloads last month
20
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sora42y/Sev-2B-preview

Finetuned
(97)
this model