SatellaJev-100M v0.1
SatellaJev-100M v0.1 is a 101.5M-parameter, non-generative transformer trained from scratch for typed probabilistic decision scoring.
Instead of generating free-form text, SatellaJev receives a state/context, a question, a decision type, and a set of candidate answers. It assigns a scalar score to each candidate and converts the candidate scores into a probability distribution with softmax.
The model was trained on the release-v2-redistributable configuration of ZefanCai/Open-Jev.
What SatellaJev does
For a decision with candidates c1 ... cn, SatellaJev learns a shared scoring function:
score_i = f(state, question, kind, candidate_i)
The candidate scores are converted into a probability distribution:
P(candidate_i) = exp(score_i / T) / sum_j exp(score_j / T)
where T is the fitted post-hoc calibration temperature.
This makes SatellaJev closer to a probabilistic decision/ranking model than a conventional language model. It has no autoregressive language-model head and does not generate arbitrary text.
Supported Open-Jev decision kinds:
noulchoicescore
Candidate sets are variable-length; the output dimensionality is determined by the options supplied at inference time rather than by a fixed classification head.
Architecture
| Component | Value |
|---|---|
| Parameters | 101,495,040 (101.50M) |
| Transformer layers | 16 |
| Hidden size | 640 |
| Attention heads | 10 |
| Head dimension | 64 |
| FFN | SwiGLU |
| FFN hidden size | 2304 |
| Vocabulary | 7,000 |
| Maximum sequence length | 1,024 |
| Positional encoding | RoPE |
| Normalization | RMSNorm |
| Attention | Bidirectional self-attention |
| Dropout | 0.10 |
| Output head | Shared scalar candidate scorer |
| Generative LM head | None |
The tokenizer is a custom 7,000-token SentencePiece BPE tokenizer trained only on the Open-Jev training split. Byte fallback is enabled.
Each candidate is encoded independently with the shared transformer. The hidden state at the final <DECIDE> token is passed through a scalar scoring head. Scores for candidates belonging to the same decision are then normalized together.
Input format
A candidate is presented to the model in the following form:
<KIND_CHOICE>
<STATE>
...
</STATE>
<QUESTION>
...
</QUESTION>
<CANDIDATE>
...
</CANDIDATE>
<DECIDE>
For a decision with several options, the model evaluates one candidate-formatted sequence per option and applies softmax across their scalar scores.
Training data
Dataset: ZefanCai/Open-Jev
Configuration: release-v2-redistributable
| Split | Records | Use |
|---|---|---|
| Train | 79,116 | Weight optimization + tokenizer training |
| Calibration | 4,672 | Post-hoc temperature fitting only |
| Validation | 3,723 | Checkpoint selection/evaluation |
| Test | 10,356 | Frozen final evaluation |
| OOD | 15,701 | Frozen out-of-distribution evaluation |
The validation, test, and OOD splits were not used to optimize model weights. The calibration split was used only after training to fit a single scalar temperature.
Open-Jev targets are reference probability distributions supplied by the dataset. SatellaJev probabilities therefore represent predictions relative to those targets and should not automatically be interpreted as empirical real-world probabilities.
Training
SatellaJev-100M v0.1 was trained from scratch for six epochs.
| Hyperparameter | Value |
|---|---|
| Epochs | 6 |
| Batch size | 64 decision records |
| Gradient accumulation | 2 |
| Effective decision batch | 128 records |
| Optimizer | AdamW |
| Peak learning rate | 3e-4 |
| Weight decay | 0.10 |
| Warmup ratio | 0.03 |
| LR schedule | Cosine decay |
| Gradient clipping | 1.0 |
| Loss | Soft-target CE + 0.1 Γ Brier |
| Checkpoint criterion | Validation cross-entropy |
The optimization objective combines soft-target cross-entropy with a Brier term:
Loss = CrossEntropy + 0.1 Γ Brier
CrossEntropy = -sum_i target_i Γ log(prediction_i)
Brier = (1 / N) Γ sum_i (prediction_i - target_i)^2
where target_i is the Open-Jev target probability and prediction_i is SatellaJev's predicted probability for candidate i.
Validation learning curve
Validation cross-entropy decreased consistently across training:
| Epoch | Validation CE |
|---|---|
| 1 | 0.6043 |
| 2 | 0.5534 |
| 3 | 0.5006 |
| 4 | 0.4638 |
| 5 | 0.4352 |
| 6 | 0.4161 |
The epoch-6 checkpoint was selected as the final model.
Calibration
After training, a single temperature parameter was fitted on the dedicated calibration split while all model weights remained frozen.
Fitted temperature: 0.980080
A temperature close to 1 indicates that only a small post-hoc rescaling of the raw candidate logits was selected in-distribution.
Results
Overall
| Split | Accuracy β | Cross-Entropy β | Brier β | ECE β |
|---|---|---|---|---|
| Calibration | 86.71% | 0.3140 | 0.0545 | 0.0211 |
| Validation | 81.65% | 0.4170 | 0.0749 | 0.0316 |
| Test | 82.03% | 0.3930 | 0.0759 | 0.0287 |
| OOD | 72.09% | 0.7604 | 0.1308 | 0.0997 |
The frozen test set closely matches validation performance, while the OOD split shows a clear but non-catastrophic degradation.
Test-set performance by decision type
| Decision type | Records | Accuracy β | Cross-Entropy β | Brier β | ECE β |
|---|---|---|---|---|---|
noul |
6,364 | 89.13% | 0.2455 | 0.0787 | 0.0185 |
choice |
2,408 | 62.79% | 0.7717 | 0.0826 | 0.0524 |
score |
1,584 | 82.77% | 0.4100 | 0.0546 | 0.0526 |
choice is the hardest in-distribution task, while noul and score are substantially stronger.
OOD performance by decision type
| Decision type | Records | Accuracy β | Cross-Entropy β | Brier β | ECE β |
|---|---|---|---|---|---|
noul |
9,929 | 81.55% | 0.5204 | 0.1405 | 0.0697 |
choice |
3,219 | 49.89% | 1.3901 | 0.1060 | 0.1947 |
score |
2,553 | 63.30% | 0.8996 | 0.1240 | 0.1403 |
The largest generalization weakness is OOD multi-option choice reasoning. This is also where calibration degrades most strongly.
These results suggest that the model transfers binary and score-style decision behavior more reliably than arbitrary multi-candidate comparisons under distribution shift.
Example inference
The repository includes inference.py, which provides SatellaJevPredictor.
from inference import SatellaJevPredictor
model = SatellaJevPredictor(
"EphAsad/SatellaJev-100M-v0.1"
)
result = model.predict(
state={
"temperature": 8,
"door": "closed",
"alarm": False,
},
question="Which action should be taken?",
options=[
"Continue monitoring",
"Escalate",
],
kind="choice",
)
print(result["distribution"])
Example output shape:
{
"Continue monitoring": 0.82,
"Escalate": 0.18
}
The values shown above are illustrative; actual outputs depend on the supplied state, question, candidates, and model checkpoint.
Intended use
SatellaJev is intended for research into:
- small probabilistic decision models;
- candidate ranking and reranking;
- constrained decision support;
- soft-label prediction;
- uncertainty-aware option scoring;
- domain adaptation of compact decision models;
- comparison of in-distribution and out-of-distribution decision behavior.
Potential downstream systems may use SatellaJev as a constrained scoring component after another system has produced a set of candidate actions or hypotheses.
It is not intended to replace a general-purpose language model for free-form generation.
Limitations
Candidate-dependent operation
SatellaJev does not generate candidate answers. Candidate options must be provided by the calling application or another model.
Independent candidate encoding
In v0.1, each candidate is encoded independently with the shared context. Candidates do not directly attend to one another inside the transformer. Their interaction occurs only when scalar scores are normalized across the candidate set.
This design supports arbitrary candidate counts, but it may limit comparative reasoning between closely related options.
OOD degradation
OOD accuracy falls from 82.03% on the frozen test set to 72.09% on OOD, and calibration error increases from 0.0287 to 0.0997.
The largest weakness is OOD choice:
- accuracy: 49.89%
- ECE: 0.1947
Applications should not assume that in-distribution calibration transfers unchanged to unfamiliar distributions.
Probability interpretation
Outputs are learned relative to Open-Jev target distributions. They are not guaranteed to correspond to real-world event frequencies, risk estimates, or human consensus.
High-impact decisions
SatellaJev-100M v0.1 is an experimental research model. It should not be used as an autonomous decision maker in medical, legal, financial, safety-critical, or other high-impact settings.
Research observations from v0.1
The first SatellaJev experiment supports several useful observations:
- A ~100M-parameter transformer can learn the Open-Jev typed decision objective from scratch without a generative language-model head.
- Frozen test performance closely tracks validation performance.
- The model retains meaningful performance under the provided OOD split, but with substantial degradation.
- OOD multi-option choice reasoning is the clearest failure mode.
- In-distribution probability calibration is comparatively strong; the fitted temperature (
0.980080) remains close to the unscaled value of 1.
These observations motivate future work on explicit candidate interaction, comparative/pairwise objectives, stronger OOD calibration, and domain-specific adaptation.
Files
A complete release may include:
model.safetensors
config.json
tokenizer.model
tokenizer.vocab
tokenizer_config.json
training_config.json
metrics.json
inference.py
README.md
Version
SatellaJev-100M v0.1
Architecture and evaluation values in this model card correspond to the final six-epoch checkpoint described above.
- Downloads last month
- 36