SatellaJev-100M v0.1

SatellaJev-100M v0.1 is a 101.5M-parameter, non-generative transformer trained from scratch for typed probabilistic decision scoring.

Instead of generating free-form text, SatellaJev receives a state/context, a question, a decision type, and a set of candidate answers. It assigns a scalar score to each candidate and converts the candidate scores into a probability distribution with softmax.

The model was trained on the release-v2-redistributable configuration of ZefanCai/Open-Jev.

What SatellaJev does

For a decision with candidates c1 ... cn, SatellaJev learns a shared scoring function:

score_i = f(state, question, kind, candidate_i)

The candidate scores are converted into a probability distribution:

P(candidate_i) = exp(score_i / T) / sum_j exp(score_j / T)

where T is the fitted post-hoc calibration temperature.

This makes SatellaJev closer to a probabilistic decision/ranking model than a conventional language model. It has no autoregressive language-model head and does not generate arbitrary text.

Supported Open-Jev decision kinds:

  • noul
  • choice
  • score

Candidate sets are variable-length; the output dimensionality is determined by the options supplied at inference time rather than by a fixed classification head.


Architecture

Component Value
Parameters 101,495,040 (101.50M)
Transformer layers 16
Hidden size 640
Attention heads 10
Head dimension 64
FFN SwiGLU
FFN hidden size 2304
Vocabulary 7,000
Maximum sequence length 1,024
Positional encoding RoPE
Normalization RMSNorm
Attention Bidirectional self-attention
Dropout 0.10
Output head Shared scalar candidate scorer
Generative LM head None

The tokenizer is a custom 7,000-token SentencePiece BPE tokenizer trained only on the Open-Jev training split. Byte fallback is enabled.

Each candidate is encoded independently with the shared transformer. The hidden state at the final <DECIDE> token is passed through a scalar scoring head. Scores for candidates belonging to the same decision are then normalized together.


Input format

A candidate is presented to the model in the following form:

<KIND_CHOICE>
<STATE>
...
</STATE>
<QUESTION>
...
</QUESTION>
<CANDIDATE>
...
</CANDIDATE>
<DECIDE>

For a decision with several options, the model evaluates one candidate-formatted sequence per option and applies softmax across their scalar scores.


Training data

Dataset: ZefanCai/Open-Jev
Configuration: release-v2-redistributable

Split Records Use
Train 79,116 Weight optimization + tokenizer training
Calibration 4,672 Post-hoc temperature fitting only
Validation 3,723 Checkpoint selection/evaluation
Test 10,356 Frozen final evaluation
OOD 15,701 Frozen out-of-distribution evaluation

The validation, test, and OOD splits were not used to optimize model weights. The calibration split was used only after training to fit a single scalar temperature.

Open-Jev targets are reference probability distributions supplied by the dataset. SatellaJev probabilities therefore represent predictions relative to those targets and should not automatically be interpreted as empirical real-world probabilities.


Training

SatellaJev-100M v0.1 was trained from scratch for six epochs.

Hyperparameter Value
Epochs 6
Batch size 64 decision records
Gradient accumulation 2
Effective decision batch 128 records
Optimizer AdamW
Peak learning rate 3e-4
Weight decay 0.10
Warmup ratio 0.03
LR schedule Cosine decay
Gradient clipping 1.0
Loss Soft-target CE + 0.1 Γ— Brier
Checkpoint criterion Validation cross-entropy

The optimization objective combines soft-target cross-entropy with a Brier term:

Loss = CrossEntropy + 0.1 Γ— Brier

CrossEntropy = -sum_i target_i Γ— log(prediction_i)

Brier = (1 / N) Γ— sum_i (prediction_i - target_i)^2

where target_i is the Open-Jev target probability and prediction_i is SatellaJev's predicted probability for candidate i.

Validation learning curve

Validation cross-entropy decreased consistently across training:

Epoch Validation CE
1 0.6043
2 0.5534
3 0.5006
4 0.4638
5 0.4352
6 0.4161

The epoch-6 checkpoint was selected as the final model.


Calibration

After training, a single temperature parameter was fitted on the dedicated calibration split while all model weights remained frozen.

Fitted temperature: 0.980080

A temperature close to 1 indicates that only a small post-hoc rescaling of the raw candidate logits was selected in-distribution.


Results

Overall

Split Accuracy ↑ Cross-Entropy ↓ Brier ↓ ECE ↓
Calibration 86.71% 0.3140 0.0545 0.0211
Validation 81.65% 0.4170 0.0749 0.0316
Test 82.03% 0.3930 0.0759 0.0287
OOD 72.09% 0.7604 0.1308 0.0997

The frozen test set closely matches validation performance, while the OOD split shows a clear but non-catastrophic degradation.


Test-set performance by decision type

Decision type Records Accuracy ↑ Cross-Entropy ↓ Brier ↓ ECE ↓
noul 6,364 89.13% 0.2455 0.0787 0.0185
choice 2,408 62.79% 0.7717 0.0826 0.0524
score 1,584 82.77% 0.4100 0.0546 0.0526

choice is the hardest in-distribution task, while noul and score are substantially stronger.


OOD performance by decision type

Decision type Records Accuracy ↑ Cross-Entropy ↓ Brier ↓ ECE ↓
noul 9,929 81.55% 0.5204 0.1405 0.0697
choice 3,219 49.89% 1.3901 0.1060 0.1947
score 2,553 63.30% 0.8996 0.1240 0.1403

The largest generalization weakness is OOD multi-option choice reasoning. This is also where calibration degrades most strongly.

These results suggest that the model transfers binary and score-style decision behavior more reliably than arbitrary multi-candidate comparisons under distribution shift.


Example inference

The repository includes inference.py, which provides SatellaJevPredictor.

from inference import SatellaJevPredictor

model = SatellaJevPredictor(
    "EphAsad/SatellaJev-100M-v0.1"
)

result = model.predict(
    state={
        "temperature": 8,
        "door": "closed",
        "alarm": False,
    },
    question="Which action should be taken?",
    options=[
        "Continue monitoring",
        "Escalate",
    ],
    kind="choice",
)

print(result["distribution"])

Example output shape:

{
    "Continue monitoring": 0.82,
    "Escalate": 0.18
}

The values shown above are illustrative; actual outputs depend on the supplied state, question, candidates, and model checkpoint.


Intended use

SatellaJev is intended for research into:

  • small probabilistic decision models;
  • candidate ranking and reranking;
  • constrained decision support;
  • soft-label prediction;
  • uncertainty-aware option scoring;
  • domain adaptation of compact decision models;
  • comparison of in-distribution and out-of-distribution decision behavior.

Potential downstream systems may use SatellaJev as a constrained scoring component after another system has produced a set of candidate actions or hypotheses.

It is not intended to replace a general-purpose language model for free-form generation.


Limitations

Candidate-dependent operation

SatellaJev does not generate candidate answers. Candidate options must be provided by the calling application or another model.

Independent candidate encoding

In v0.1, each candidate is encoded independently with the shared context. Candidates do not directly attend to one another inside the transformer. Their interaction occurs only when scalar scores are normalized across the candidate set.

This design supports arbitrary candidate counts, but it may limit comparative reasoning between closely related options.

OOD degradation

OOD accuracy falls from 82.03% on the frozen test set to 72.09% on OOD, and calibration error increases from 0.0287 to 0.0997.

The largest weakness is OOD choice:

  • accuracy: 49.89%
  • ECE: 0.1947

Applications should not assume that in-distribution calibration transfers unchanged to unfamiliar distributions.

Probability interpretation

Outputs are learned relative to Open-Jev target distributions. They are not guaranteed to correspond to real-world event frequencies, risk estimates, or human consensus.

High-impact decisions

SatellaJev-100M v0.1 is an experimental research model. It should not be used as an autonomous decision maker in medical, legal, financial, safety-critical, or other high-impact settings.


Research observations from v0.1

The first SatellaJev experiment supports several useful observations:

  1. A ~100M-parameter transformer can learn the Open-Jev typed decision objective from scratch without a generative language-model head.
  2. Frozen test performance closely tracks validation performance.
  3. The model retains meaningful performance under the provided OOD split, but with substantial degradation.
  4. OOD multi-option choice reasoning is the clearest failure mode.
  5. In-distribution probability calibration is comparatively strong; the fitted temperature (0.980080) remains close to the unscaled value of 1.

These observations motivate future work on explicit candidate interaction, comparative/pairwise objectives, stronger OOD calibration, and domain-specific adaptation.


Files

A complete release may include:

model.safetensors
config.json
tokenizer.model
tokenizer.vocab
tokenizer_config.json
training_config.json
metrics.json
inference.py
README.md

Version

SatellaJev-100M v0.1

Architecture and evaluation values in this model card correspond to the final six-epoch checkpoint described above.

Downloads last month
36
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train EphAsad/SatellaJev-100M-v0.1