Sol Pro
Sol Pro is a 138,037,594-parameter language model from Sol Labs. It combines recurrent transformer blocks with tensorized n-gram memory and a 2,048-token context.
The released checkpoint scores 28.63 on the normalized Intelligence Index and 24.13 on the raw accuracy Index.
Model
| Setting | Value |
|---|---|
| Parameters | 138,037,594 |
| Context | 2,048 tokens |
| Vocabulary | 4,096-token byte-level BPE |
| Hidden width | 768 |
| Stored / effective blocks | 12 / 16 |
| Attention | 24 query heads, 8 KV heads, head dimension 32 |
| Feed-forward layers | SwiGLU, width 3,968 |
| Position encoding | RoPE, theta 20,000 |
| Memory | Rank-299 tensorized 2-, 3-, and 4-gram memory |
| Recurrence | Learned pass embeddings and loop gates |
| Embeddings | Tied input and output weights |
| Inference weights | BF16 safetensors |
Four middle blocks run twice. Causal grouped-query attention uses learned Q/K RMSNorm and XSA value-direction subtraction. The n-gram module shares token-position factors across its three orders.
Load and generate
Install the dependencies:
pip install torch safetensors tokenizers huggingface_hub
import sys
from huggingface_hub import snapshot_download
model_dir = snapshot_download(
"solintellegence/sol-pro",
allow_patterns=[
"model.safetensors", "modeling_sol_pro.py",
"config.json", "tokenizer.json",
],
)
sys.path.insert(0, model_dir)
from modeling_sol_pro import load_model, generate
model, tokenizer = load_model(model_dir)
print(generate(
model,
tokenizer,
"The best way to learn something new is",
max_new_tokens=64,
))
The repository includes standalone PyTorch model code. Generation recomputes the context at each step; KV caching is not implemented. Keep prompts and continuations within 2,048 tokens.
Evaluation
| Benchmark | Examples | Normalized accuracy | Raw accuracy |
|---|---|---|---|
| HellaSwag | 10,042 | 46.76% | 36.88% |
| ARC Easy | 2,376 | 54.08% | 55.35% |
| ARC Challenge | 1,172 | 32.00% | 29.35% |
| PIQA | 1,838 | 70.95% | 69.26% |
| Arithmark3 | 1,000 | 36.00% | 37.20% |
| Intelligence Index | 28.63 | 24.13 |
These are zero-shot float32 measurements on complete evaluation splits. HellaSwag, ARC Easy, ARC Challenge, and PIQA use lm-evaluation-harness 0.4.12 with batch size 64. Arithmark3 uses the official AxiomicLabs script with its default batch size 32 and 1,024-token context, explicitly set to float32. The Index uses the leaderboard's chance-adjusted formula and each task's normalized accuracy.
This is a task-adapted checkpoint using public benchmark training splits. Evaluation used separate held-out splits, with matching evaluation contexts excluded from the adaptation data; Arithmark evaluation examples were not used for adaptation. The scores have not been independently verified and do not establish a leaderboard position.
WikiText-2 validation cross-entropy is 2.7677 over 366,592 tokens. Exact results, package versions, hashes, and commands are in the float32 evaluation summary. Raw outputs are available for lm-eval and ArithMark-3.
Use and limits
Sol Pro supports text-completion and small-model research. Multiple-choice scores do not establish reliable free-form reasoning or conversational behavior. Outputs can be incorrect, repetitive, or inconsistent.
Files
| File | Contents |
|---|---|
model.safetensors |
Released inference weights |
modeling_sol_pro.py |
Architecture, loader, and text generator |
config.json |
Architecture configuration |
tokenizer.json |
Original tokenizer |
evaluation/standard_float32/ |
Standard-harness evaluation results and reproducibility details |
banner.png |
Sol Pro artwork |
License
Sol Pro is released under Apache 2.0. Dataset licenses and terms remain with their respective owners.
- Downloads last month
- -
