8ball-laya-mtl2

A System One decision model fine-tuned to be a Magic 8-Ball: it answers one 20-way choice question β€” "the asker posed the question in the state; answer it with the classic Magic 8-Ball response that fits best" β€” with a calibrated probability distribution over the canonical 20 answers (10 affirmative / 5 non-committal / 5 negative), each mapped to a signed alignment value in [-1, 1].

421M parameters (ModernBERT-large encoder + decision head), Apache-2.0, runs on CPU (330–390 ms/question) or a single consumer GPU (5–10 ms). Loads like any Laya checkpoint:

import laya

agent = laya.load("RockmSockmJesus/8ball-laya-mtl2")
result = agent.predict("Should I deploy on Friday?", {
    "answer": {
        "type": "choice",
        "instructions": "The asker posed the question found in the state. Answer it with the classic Magic 8-Ball response that fits best.",
        "criteria": {  # the canonical 20, verbatim
            "It is certain": "Strong affirmative: no doubt whatsoever",
            # ... see eightball.answers.CRITERIA or the repo
        },
    }
})

Or with the full CLI (sampling, deflection, alignment score):

uv run 8ball --provider laya-local "Should I deploy on Friday?"

Results

The 8-ball persona (held-out eval, n=620)

metric value
Sampled bucket mix 49.7 / 25.1 / 25.2 (canonical: 50 / 25 / 25)
Brier vs persona gold 0.0031 (teacher: 0.0095, uniform: 0.0175)
Mean p_max 0.106 (p5 0.084 Β· p50 0.100 Β· p95 0.145) β€” deliberately humble
Deflection rate @0.09 18.7% of shakes fall back to the non-committal answers

General decision skill (S1MB english-v1, 137 benchmarks, 26,269 judgments)

Evaluated with the official S1MB evaluator (baseline-adjusted scores 0–100; extended-input condition --max-len 9216).

model noul choice score Task Avg
TypeSafe Jev 1.13 64.77 67.31 46.69 59.59
laya-typed-decisions (base, published) 20.06 18.93 5.99 15.00
8ball v1 (persona-only fine-tune) 18.77 17.14 2.14 12.69
8ball v2 (MTL r1, 47k items) 28.87 21.47 19.18 23.17
8ball v3 (this model, 143k items, 6 epochs) 33.85 24.82 21.43 26.70

The multi-task training traded ~1.7 Task Avg vs the base teacher for the persona (mix, alignment scale, humility) β€” and recovered general skill 12.69 β†’ 26.70 along the way. Full methodology: github.com/RockmSockmJesus/8ball.

Training

  • Recipe: RLCD (upstream Laya loop) β€” proper-scoring-rule policy gradient (spherical 0.75 + ranked probability 1.0) over noisy logit projections, plus soft cross-entropy against gold distributions. 6 epochs total, fp32, per-type temperature calibration after the final epoch.
  • Corpus (143,514 items, 0 skipped):
    • Open-Jev release-v2-redistributable train (79k, CC0) β€” all 12 sources, uncapped
    • 14 public NLP train splits converted to S1MB decision formats (30.1k): SNLI, HANS, PAWS, CREAK, SciTail, ETHOS, Civil Comments, DBpedia, AG News, poem_sentiment, ARC, OpenBookQA, AQuA-RAT, Banking77
    • 8-ball persona gold (11.6k effective): laya-typed-decisions teacher distributions over 6,200 whimsical questions, bucket-rescaled to the canonical 50/25/25 mix, blended 70/30 with uniform (humility baked into targets)
    • Qwen3.8-27B-generated coding/logic/finance decision cases (4.4k)
  • Never trained on S1MB test rows β€” NLP conversion transfers decision formats (instructions/criteria/state shape read from local test schemas) applied to train-split rows only.

Limitations

  • No world knowledge: it judges the state you hand it. Without evidence in the state it is deliberately near-uniform β€” an honest 8-ball, not an oracle.
  • The persona is trained-in: this checkpoint is near-uniform on almost everything. For general typed decisions use convaiinnovations/laya-typed-decisions.
  • Score-primitive skill is weak (21.4) relative to frontier decision models.
  • Trained for the fixed 20-option 8-ball choice; other option sets work (request-time declaration) but are not calibrated for.

Intended use

A calibrated, whimsical decision layer: "should I…" questions in β†’ an 8-ball answer with a signed alignment score between βˆ’1 and +1 (direction from the sampled answer, magnitude from the probability mass). See the repo for the sampling + deflection layer (deflection threshold calibrated to this checkpoint: 0.09).

Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for PythiaFinance/8ball

Finetuned
(8)
this model