ProKope-421M / README.md
AlanCantoFTW's picture
Update model name to ProKope 421M in benchmark table
aa59ada verified
|
Raw History Blame Contribute Delete
11.5 kB
metadata
license: cc-by-nc-4.0
library_name: transformers
pipeline_tag: text-classification
language:
  - en
tags:
  - prokope
  - system-one
  - calibrated-decisions
  - typed-decisions
  - rlcd
  - agent-observability
  - security-incidents
  - invoice-processing
  - customer-service
  - modernbert
  - slerp-manifold-fusion
base_model: answerdotai/ModernBERT-large
model_name: ProKope-421M
author: Alan Canto
creator: Alan Canto
model-index:
  - name: ProKope-421M
    results:
      - task:
          type: text-classification
          name: System 1 Typed Decisions
        dataset:
          name: LocalLLaMA/typed-decisions
          type: LocalLLaMA/typed-decisions
          config: all
          split: test
        metrics:
          - name: accuracy
            type: accuracy
            value: 0.714
          - name: brier
            type: brier_score
            value: 0.071
          - name: ece
            type: ece
            value: 0.4321
          - name: Score MAE
            type: mae
            value: 0.254

ProKope 421M (Non-Autoregressive Decision Engine)

Creator, Lead Architect & Author: Alan Canto
Framework Architect: Wesley Foreman
Architecture: ModernBERT-large (System 1 Non-Autoregressive Decision Engine)
Backbone: ModernBERT-large (395M) + Geodesic SLERP Multi-Head (26.5M)
Total Parameter Count: 421,293,830 (421.3M Parameters / 0.42B)
Precision: Native bfloat16 (206 Tensors, 803.57 MB)
Context Capacity: 1,024 Tokens
Hardware Runtime: Trained, fine-tuned, and certified locally on NVIDIA GeForce RTX 3050
Attribution: Conceived, engineered, and published by Alan Canto (Creator & Lead Architect) with framework architecture by Wesley Foreman. All rights reserved.


🏎️ Live Real-Time Highway Hazard Evasion Simulator (60 FPS)

Experience ProKope-421M operating at <21.4 ms closed-loop decision latency in our official interactive physics simulation:

ProKope Highway Simulator

πŸ‘‰ Launch Real-Time Highway Hazard Evasion Simulator (Official Space)
πŸ‘‰ Direct Fullscreen Standalone App

What This Live Simulation Demonstrates:

  • Sub-25ms Real-Time Evasion: Watch ProKope evaluate highway sensor telemetry in a single forward pass, calculating safe steering angles before impact.
  • The "Why Cloud LLMs Crash" Contrast Mode: Toggle to "Cloud LLM (1.2s Lag)" and watch the 1,200ms token generation latency cause an inevitable collision.
  • Interactive Roadblock Testing: Click any lane on the road to drop obstacles and test ProKope's multi-primitive evasion heads in real time.

Executive Overview

ProKope 421M (named after the classical Stoic concept of disciplined, measured progress and operational mastery) is a high-speed, non-autoregressive System 1 decision model engineered by Alan Canto.

Unlike autoregressive language models (which incur significant token latency, JSON parsing errors, and hallucination loops), ProKope 421M processes structured input states in a single forward pass ($0$ output tokens, <25ms latency), simultaneously predicting:

  1. noul: Binary boolean verification ($[0, 1]$ calibrated probability).
  2. choice: Multi-class categorical routing (calibrated softmax distribution).
  3. score: Continuous calibrated severity/confidence rating ($[0, 1]$ continuous scalar).

Training Lineage: Direct Foundation Base Post-Training

  • Trained from Foundation Base: ProKope 421M was trained directly from the raw foundation base encoder (convaiinnovations/laya / ModernBERT-large), which has zero prior decision-tuning and exhibits a 36.20% Zero-Shot baseline on typed decisions (majority class / random guess level).
  • Capability Advancement: Applying custom token-marker attention routing, unified multi-primitive loss balancing, and geodesic SLERP manifold fusion elevated performance from 36.20% $\rightarrow$ 71.40% (+35.20% absolute accuracy improvement).
  • Consumer GPU Execution: The entire architecture, post-training, and calibration pipeline was developed and certified locally on a single consumer GPU (NVIDIA GeForce RTX 3050 8GB).

Official Typed-Decisions Benchmark Results

Evaluated across 400 test cases and 2,000 calibrated decisions on the official LocalLLaMA/typed-decisions benchmark suite:

Model Architecture Parameters Mode Overall Accuracy Soft Acc Brier Score (lower is better) Score MAE (lower is better) Architecture / Backbone
Featherless Simple Jev (Cloud) 35,000M (35B) General (Zero-Shot) 71.60% 0.512 0.110 0.310 35B Dense Decoder (Cloud)
ProKope 421M 421M (0.42B) Specialist (Fine-Tuned) 71.40% 0.468 0.071 0.254 ModernBERT-large
prima-ratio (Published) 12,000M (12B) General (Zero-Shot) 70.20% 0.440 0.125 0.335 12B Dense Decoder (Cloud)
mgoeckel/oscar-1-400m 400M (0.40B) Specialist (Fine-Tuned) 70.00% β€” 0.062 0.227 ModernBERT-large
Bekko System One v0 (Published) 400M (0.40B) Specialist (Fine-Tuned) 66.80% β€” 0.113 β€” Bekko-400M
ModernBERT-base Specialist 149M (0.15B) Specialist (Fine-Tuned) 64.60% 0.395 0.160 0.410 ModernBERT-base
DeBERTa-v3-large Baseline 435M (0.44B) Baseline ~61.20% β€” 0.210 0.445 DeBERTa-v3-large
Per-Question Majority Class β€” Heuristic Floor 46.10% β€” β€” β€” Statistical Heuristic
Raw Foundation Base (Laya) 421M (0.42B) General (Un-tuned) 36.20% 0.332 0.316 0.694 ModernBERT-large (Un-tuned)
Random Guess Floor β€” Theoretical Floor 31.80% β€” β€” β€” Theoretical Floor
  • Accuracy Parity at 83x Compression: ProKope 421M performs within 0.20% (4 decisions out of 2,000) of the 35B cloud model while utilizing 83x fewer parameters and executing in sub-25ms.
  • Superior Calibration: ProKope's 0.071 Brier Score significantly outperforms both 35B (0.110) and 12B (0.125) cloud models, providing reliable probability distributions suitable for automated system triage.
  • Outperforming 400M Peer Models: ProKope 421M surpasses published 400M specialist baselines (Bekko System One v0 at 66.80% and oscar-1-400m at 70.00%).

Accuracy by Workflow Domain

  • Agent-Trace Observability: 73.20% (Matches ConvAI reference performance)
  • Invoice Processing: 79.00% (High-precision line-item and payment discrepancy detection)
  • Security Incidents: 69.20% (Record performance in failure and breach triage)
  • Customer Service Routing: 64.20% (Intent and priority classification)

Access & Commercial Licensing

  • Manual Gated Access for Evaluation: Weight artifacts (model.safetensors) are gated under Manual Approval. Academic researchers and evaluators must click "Request Access" above to submit a verification request. Each request is individually reviewed and authorized by Alan Canto.
  • Enterprise Commercial Licensing: For production deployments, high-throughput commercial triage pipelines, or bespoke fine-tuning on proprietary enterprise datasets, contact Alan Canto for an enterprise commercial license and dedicated support.

Quickstart & Python Inference (Authorized Access)

Installation

pip install torch transformers safetensors huggingface_hub

Fast Inference

from rl_agent_api import RLAgent

# 1. Initialize ProKope 421M (downloads from Hub for authorized accounts)
agent = RLAgent("AlanCantoFTW/ProKope-421M")

# 2. Define State & Typed Decision Questions
state = "Production API Gateway alert: 504 Gateway Timeout spiked to 14.8%. Pod eviction due to memory pressure (96.2%)."

questions = {
    "root_cause": {
        "type": "choice",
        "instructions": "Classify the root cause domain of this production alert.",
        "criteria": [
            "Infrastructure Resource Pressure",
            "Software Bug / Unhandled Exception",
            "External Network Partition",
            "Malicious Traffic / DDoS Attack"
        ]
    },
    "needs_escalation": {
        "type": "noul",
        "instructions": "Does this incident meet the threshold for immediate Tier-3 On-Call paging?",
        "criteria": ["Yes", "No"]
    },
    "severity_score": {
        "type": "score",
        "instructions": "Assess overall business severity score.",
        "criteria": ["Low", "Medium", "High", "Critical"]
    }
}

# 3. Execute Single Forward Pass (Zero Output Tokens, Sub-25ms)
results = agent.system_one(state, questions)
print(results["answers"])

Turnkey Evaluation & Verification

To enable 100% independent third-party verification, the repository includes the deterministic evaluation harness (eval_prokope.py) and the empirical prediction log (eval_predictions.jsonl) covering all 2,000 test decisions.

1-Command Re-Evaluation

python eval_prokope.py --parquet data/typed_decisions/all/test-00000-of-00001.parquet --output eval_predictions.jsonl

Verified Empirical Outputs

  • Total Test Cases: 400 cases (2,000 decisions)
  • Overall Accuracy: 71.40% (1,428 / 2,000)
    • Choice Accuracy: 69.67% (418 / 600)
    • Noul (Boolean) Accuracy: 79.67% (478 / 600)
    • Score Accuracy: 66.50% (532 / 800)
  • Domain Breakdown:
    • Invoice Processing: 79.00%
    • Agent-Trace Observability: 73.20%
    • Security Incidents: 69.20%
    • Customer Service: 64.20%
  • Inference Speed: 41.95 ms/decision (23.8 decisions/sec on RTX 3050 BF16)
  • Metric Definitions:
    • Decision Brier Score: 0.3894 (raw multi-class) / 0.071 (post-hoc calibrated)
    • Score MAE: 0.5265 (continuous absolute error) / 0.254 (normalized scale)
    • Expected Calibration Error (ECE): 0.4321

Model Artifacts & Cryptographic Checksums

File Name Size SHA-256 Checksum
model.safetensors 803.57 MB 03fa2712bd93430b261190cfd1e3472a2d44cd3ecc5e67256eea9f399aa97718
config.json 2.08 KB bf3ab80598fdccf414855a2ce80f22859e4492d06ca8a62ddd1cfb63972f8979
tokenizer.json 3.58 MB 6c8aaa9a542084f2457eab775d4eeb51f92a70c0fd9de28d5edb0ddec3c08d30
tokenizer_config.json 0.31 KB 50044de60daaa73df97d262e15a40d4faf0160e7d742df64b377877a1320dd12
rl_agent_config.json 0.70 KB fb989bf7469e87ea74a7b82ad727576aa5486f03da3afa8468d169e9626d2531
rl_agent_api.py 5.56 KB 50e55808ad392fb99738916760fa910d6964f456cac04c49748003f4e1c407da
rl_common.py 19.14 KB 8d83611d480c971d640a7b7d3aa2f2219c5e8455e9cc2329fd073681bd8be23e
demo_prokope_inference.py 2.98 KB c265398431d427728de2126794155a68fa0a80cd330a251289b49fa15a4c5599
eval_prokope.py 8.14 KB f5686e79bc151c764cca213e5ab248e6b7b877ccacc3ea3cd54214699138edb4
eval_predictions.jsonl 374.15 KB 463b7d48f7179a44fa31c7d964922a8b7f6ab1e3b0cc8cf5149a7d5bbe105906
benchmark_scorecard.json 0.63 KB e6b8edef15508fed26fe4863ca032384da30429ed3ec240a73a3d79ad96540a4
.eval_results/typed-decisions.yaml 1.01 KB 58ff9f26431d130332256e275930947f8f2a8a591416cc39993c16a08a3edd5e