Instructions to use vllm-sr/Decision-2.0-Lux-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vllm-sr/Decision-2.0-Lux-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="vllm-sr/Decision-2.0-Lux-9B", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vllm-sr/Decision-2.0-Lux-9B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("vllm-sr/Decision-2.0-Lux-9B", trust_remote_code=True, device_map="auto")Decision-2.0-Lux-9B
Decision-2.0-Lux-9B is the 9B model of Decision 2.0, the decision models of vLLM Semantic Router. Give it an input (text or JSON) and the questions you need answered: pick one of several options, say yes or no, or rate on a scale. It answers them all at once and returns a probability for every answer, without generating text.
| Parameters | 7.94B |
| Context length | 16,384 tokens |
| Decision types | Choice Β· Yes / No Β· Score |
| License | Apache-2.0 |
Highlights
- Top JevArena score of its size: 68.1, ahead of the 2 other same-size models compared.
- Ahead of Decision 1.0 Lux: +2.3 on JevArena and +2.8 on the Jev Decision Index.
- Speed: a median of 18.4 ms per single-question request on a single GPU.
- Many questions, one pass: Choice, Yes / No and Score questions about the same input are answered together in one forward pass, with a probability for every option.
Quickstart
pip install "transformers>=5.17" torch safetensors
import json
from transformers import AutoModel
model = AutoModel.from_pretrained("vllm-sr/Decision-2.0-Lux-9B", trust_remote_code=True)
result = model.system_one(
state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"returns": "Refunds, replacements and damaged deliveries",
"billing": "Payments, invoices and charges",
"technical": "Product setup and faults"
}
},
"receipt": {
"type": "noul",
"instructions": "Does the customer have a receipt?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this request?",
"criteria": [
"Routine",
"Soon",
"Today"
]
}
},
)
print(json.dumps(result["answers"], indent=2))
# Or as a pipeline:
# transformers.pipeline("decision", model="vllm-sr/Decision-2.0-Lux-9B", trust_remote_code=True)(state=..., questions=...)
Evaluation
| Model | JevArena β | Human-labelled transfer β | Jev Decision Index β |
|---|---|---|---|
| Decision-2.0-Lux-9B | 68.1 | 56.2 | 46.3 |
| Decision 1.0 Lux | 65.8 | 55.8 | 43.5 |
| Nimble v2 | 62.1 | 53.6 | β |
JevArena
Every model answers the same frozen prompts, scored the same way; missing or invalid answers count as errors. Human-labelled transfer is the median macro-F1 over 15 human-labelled tasks (Γ100).
Jev Decision Index
Decision 2.0: independent reproduction with the official 0.2.1 kit on the released weights; others: public board snapshot, 2026-09-28. Training data audited at row level against all Index test items.
License
Apache-2.0 (LICENSE).
Citation
@misc{decision_2_0_lux_9b_2026,
title = {{Decision-2.0-Lux-9B}: A Decision 2.0 Model for Structured Decisions},
author = {{vLLM Semantic Router Team}},
year = {2026},
howpublished = {\url{https://huggingface.co/vllm-sr/Decision-2.0-Lux-9B}}
}
- Downloads last month
- 257





# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="vllm-sr/Decision-2.0-Lux-9B", trust_remote_code=True)