8ball-laya-mtl2
A System One decision model fine-tuned to be a Magic 8-Ball: it answers one
20-way choice question β "the asker posed the question in the state; answer it
with the classic Magic 8-Ball response that fits best" β with a calibrated
probability distribution over the canonical 20 answers (10 affirmative /
5 non-committal / 5 negative), each mapped to a signed alignment value in [-1, 1].
421M parameters (ModernBERT-large encoder + decision head), Apache-2.0, runs on
CPU (330β390 ms/question) or a single consumer GPU (5β10 ms). Loads like any
Laya checkpoint:
import laya
agent = laya.load("RockmSockmJesus/8ball-laya-mtl2")
result = agent.predict("Should I deploy on Friday?", {
"answer": {
"type": "choice",
"instructions": "The asker posed the question found in the state. Answer it with the classic Magic 8-Ball response that fits best.",
"criteria": { # the canonical 20, verbatim
"It is certain": "Strong affirmative: no doubt whatsoever",
# ... see eightball.answers.CRITERIA or the repo
},
}
})
Or with the full CLI (sampling, deflection, alignment score):
uv run 8ball --provider laya-local "Should I deploy on Friday?"
Results
The 8-ball persona (held-out eval, n=620)
| metric | value |
|---|---|
| Sampled bucket mix | 49.7 / 25.1 / 25.2 (canonical: 50 / 25 / 25) |
| Brier vs persona gold | 0.0031 (teacher: 0.0095, uniform: 0.0175) |
| Mean p_max | 0.106 (p5 0.084 Β· p50 0.100 Β· p95 0.145) β deliberately humble |
| Deflection rate @0.09 | 18.7% of shakes fall back to the non-committal answers |
General decision skill (S1MB english-v1, 137 benchmarks, 26,269 judgments)
Evaluated with the official S1MB evaluator
(baseline-adjusted scores 0β100; extended-input condition --max-len 9216).
| model | noul | choice | score | Task Avg |
|---|---|---|---|---|
| TypeSafe Jev 1.13 | 64.77 | 67.31 | 46.69 | 59.59 |
| laya-typed-decisions (base, published) | 20.06 | 18.93 | 5.99 | 15.00 |
| 8ball v1 (persona-only fine-tune) | 18.77 | 17.14 | 2.14 | 12.69 |
| 8ball v2 (MTL r1, 47k items) | 28.87 | 21.47 | 19.18 | 23.17 |
| 8ball v3 (this model, 143k items, 6 epochs) | 33.85 | 24.82 | 21.43 | 26.70 |
The multi-task training traded ~1.7 Task Avg vs the base teacher for the persona (mix, alignment scale, humility) β and recovered general skill 12.69 β 26.70 along the way. Full methodology: github.com/RockmSockmJesus/8ball.
Training
- Recipe: RLCD (upstream Laya loop) β proper-scoring-rule policy gradient (spherical 0.75 + ranked probability 1.0) over noisy logit projections, plus soft cross-entropy against gold distributions. 6 epochs total, fp32, per-type temperature calibration after the final epoch.
- Corpus (143,514 items, 0 skipped):
- Open-Jev
release-v2-redistributabletrain (79k, CC0) β all 12 sources, uncapped - 14 public NLP train splits converted to S1MB decision formats (30.1k): SNLI, HANS, PAWS, CREAK, SciTail, ETHOS, Civil Comments, DBpedia, AG News, poem_sentiment, ARC, OpenBookQA, AQuA-RAT, Banking77
- 8-ball persona gold (11.6k effective): laya-typed-decisions teacher distributions over 6,200 whimsical questions, bucket-rescaled to the canonical 50/25/25 mix, blended 70/30 with uniform (humility baked into targets)
- Qwen3.8-27B-generated coding/logic/finance decision cases (4.4k)
- Open-Jev
- Never trained on S1MB test rows β NLP conversion transfers decision formats (instructions/criteria/state shape read from local test schemas) applied to train-split rows only.
Limitations
- No world knowledge: it judges the state you hand it. Without evidence in the state it is deliberately near-uniform β an honest 8-ball, not an oracle.
- The persona is trained-in: this checkpoint is near-uniform on almost everything.
For general typed decisions use
convaiinnovations/laya-typed-decisions. - Score-primitive skill is weak (21.4) relative to frontier decision models.
- Trained for the fixed 20-option 8-ball choice; other option sets work (request-time declaration) but are not calibrated for.
Intended use
A calibrated, whimsical decision layer: "should Iβ¦" questions in β an 8-ball answer with a signed alignment score between β1 and +1 (direction from the sampled answer, magnitude from the probability mass). See the repo for the sampling + deflection layer (deflection threshold calibrated to this checkpoint: 0.09).
- Downloads last month
- -
Model tree for PythiaFinance/8ball
Base model
convaiinnovations/laya-typed-decisions