cloudyu commited on
Commit
5997495
·
verified ·
1 Parent(s): ea13f62

System 1 only by default (thinking: "off"); adaptive thinking is opt-in

Browse files
Files changed (2) hide show
  1. README.md +4 -4
  2. serve_decide.py +3 -3
README.md CHANGED
@@ -70,15 +70,15 @@ Details: `reports/decision_index_adaptive.json`, `reports/adaptive_latency_summa
70
 
71
  ## Overview
72
 
73
- **GEV-26B-Decide answers typed questions with a calibrated probability for every option, and thinks only when it needs
74
- to.** System 1 decides in one forward pass (about 45 ms). When its leading option is uncertain, System 2 (the same
75
  backbone in Gemma-4 thinking mode) reasons over the question, and the reasoning is folded into the final probabilities.
76
  One set of weights, one vLLM engine, for text and images.
77
 
78
  | | what it does | output |
79
  |---|---|---|
80
  | **System 1** | typed decisions: yes/no · pick one of 2–256 options · rate 0–5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
81
- | **Adaptive thinking** | System 1 first; below 0.8 confidence, System 2 thinks and its answer is folded in | calibrated probabilities |
82
  | **System 2** | the unmodified `google/gemma-4-26B-A4B-it`, optionally thinking step by step, text and images | text / reasoning |
83
 
84
  GEV-26B-Decide was previously published as `autotrust/JEV-Gemma4-26B-A4B`; the weights are the same.
@@ -252,7 +252,7 @@ curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
252
  | `state` | what the decision is about: a string, a JSON object, or a list mixing text and images `["Photo: ", {"image": "https://… or data:…"}]` |
253
  | `question` | one question about the state |
254
  | `options` | `choice` only: 2–256 strings |
255
- | `thinking` | `"auto"` (default: adaptive), `"on"` (always think), `"off"` (System 1 only); `noul` and `choice`. Use `"auto"` for reasoning questions and `"off"` for classification, retrieval and tool routing |
256
  | `threshold` | System 1 confidence below which `"auto"` thinks (default 0.8) |
257
  | `think_budget` | maximum thinking tokens; default: no cap beyond the context window |
258
  | `chat_template_kwargs` | passed to the base model's chat template, as in its chat API (Gemma-4 has thinking on/off only, so there is no `reasoning_effort` setting) |
 
70
 
71
  ## Overview
72
 
73
+ **GEV-26B-Decide answers typed questions with a calibrated probability for every option; with thinking switched on, it
74
+ thinks only when it needs to.** System 1 decides in one forward pass (about 45 ms). When its leading option is uncertain, System 2 (the same
75
  backbone in Gemma-4 thinking mode) reasons over the question, and the reasoning is folded into the final probabilities.
76
  One set of weights, one vLLM engine, for text and images.
77
 
78
  | | what it does | output |
79
  |---|---|---|
80
  | **System 1** | typed decisions: yes/no · pick one of 2–256 options · rate 0–5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
81
+ | **Adaptive thinking** (opt-in: `thinking: "auto"`) | System 1 first; below 0.8 confidence, System 2 thinks and its answer is folded in | calibrated probabilities |
82
  | **System 2** | the unmodified `google/gemma-4-26B-A4B-it`, optionally thinking step by step, text and images | text / reasoning |
83
 
84
  GEV-26B-Decide was previously published as `autotrust/JEV-Gemma4-26B-A4B`; the weights are the same.
 
252
  | `state` | what the decision is about: a string, a JSON object, or a list mixing text and images `["Photo: ", {"image": "https://… or data:…"}]` |
253
  | `question` | one question about the state |
254
  | `options` | `choice` only: 2–256 strings |
255
+ | `thinking` | `"off"` (default: System 1 only), `"auto"` (adaptive), `"on"` (always think); `noul` and `choice`. Switch on `"auto"` for reasoning questions; keep `"off"` for classification, retrieval and tool routing |
256
  | `threshold` | System 1 confidence below which `"auto"` thinks (default 0.8) |
257
  | `think_budget` | maximum thinking tokens; default: no cap beyond the context window |
258
  | `chat_template_kwargs` | passed to the base model's chat template, as in its chat API (Gemma-4 has thinking on/off only, so there is no `reasoning_effort` setting) |
serve_decide.py CHANGED
@@ -12,8 +12,8 @@ POST /v1/decide
12
  options list of strings (choice only)
13
  strategy choices with more than 16 options: "single" (one pass, labels A-P then Q-Z, AA, ...), "tournament"
14
  (groups of <=16 + a final of 16), "permute" (single pass over 4 option orders, averaged); default per model
15
- thinking "off", "auto" (think only when the leading option is below `threshold`), "on" (always think);
16
- default per model (GEV: "auto"; JEV-27B: "off")
17
  threshold System 1 confidence below which "auto" switches thinking on (default per model)
18
  reasoning controls, as for the base model's own chat API:
19
  chat_template_kwargs passed to the base model's chat template (e.g. Qwen3.8: {"reasoning_effort": "low"})
@@ -65,7 +65,7 @@ PROFILES = {
65
  "qwen": {"prefix": "", "image": "<|vision_start|><|image_pad|><|vision_end|>", "end_think": "</think>",
66
  "after_think": "\n\n", "strategy": "single", "threshold": 0.8, "mix": 0.5, "thinking": "off"},
67
  "gemma": {"prefix": "<bos>", "image": "<|image|>", "end_think": "<channel|>", "after_think": "",
68
- "strategy": "tournament", "threshold": 0.8, "mix": 0.5, "thinking": "auto"},
69
  }
70
  # Read the full distribution: override generation_config defaults (top_k/top_p) that would truncate processed logprobs.
71
  READ = dict(max_tokens=1, temperature=1.0, top_p=1.0, top_k=0, min_p=0.0, repetition_penalty=1.0,
 
12
  options list of strings (choice only)
13
  strategy choices with more than 16 options: "single" (one pass, labels A-P then Q-Z, AA, ...), "tournament"
14
  (groups of <=16 + a final of 16), "permute" (single pass over 4 option orders, averaged); default per model
15
+ thinking "off" (default: System 1 only), "auto" (think only when the leading option is below `threshold`),
16
+ "on" (always think)
17
  threshold System 1 confidence below which "auto" switches thinking on (default per model)
18
  reasoning controls, as for the base model's own chat API:
19
  chat_template_kwargs passed to the base model's chat template (e.g. Qwen3.8: {"reasoning_effort": "low"})
 
65
  "qwen": {"prefix": "", "image": "<|vision_start|><|image_pad|><|vision_end|>", "end_think": "</think>",
66
  "after_think": "\n\n", "strategy": "single", "threshold": 0.8, "mix": 0.5, "thinking": "off"},
67
  "gemma": {"prefix": "<bos>", "image": "<|image|>", "end_think": "<channel|>", "after_think": "",
68
+ "strategy": "tournament", "threshold": 0.8, "mix": 0.5, "thinking": "off"},
69
  }
70
  # Read the full distribution: override generation_config defaults (top_k/top_p) that would truncate processed logprobs.
71
  READ = dict(max_tokens=1, temperature=1.0, top_p=1.0, top_k=0, min_p=0.0, repetition_penalty=1.0,