Granite-4.2-3B-Heretic-NX-PRIME-GGUF / EVALUATION_SUMMARY.md
0xzknw's picture
Add files using upload-large-folder tool
25759b6 verified
|
Raw History Blame Contribute Delete
2.18 kB
# Evaluation summary
## Selected checkpoint
- Model: Granite 4.2 3B Heretic NX PRIME
- Residual-stream protection rank: 16
- Beta: 2.4
- Layers: 25 through 36
- Projection families: attention output and MLP down projection
- Source revision: `b7e947307dd2efb3ad3b853b0e8a7e75f8ad4ac2`
## Fixed 104-row refusal proxy
The final static checkpoint was reloaded using Transformers NF4. It produced
0 explicit-refusal marker hits on 104 fixed held-out prompts. The official base
produced 103 marker hits under the same lexical detector.
This detector looks only for explicit refusal phrases. It does not establish
whether an answer completed a harmful task, and it must not be interpreted as
a general safety or quality score.
## First-token KL
Two distinct protocols are reported:
1. Activation-native screen: the official base is loaded once in NF4 and the
candidate residual intervention is applied through runtime hooks. Beta 2.4
scored 0.0246739865 KL and 1/104 refusal markers.
2. Artifact validation: the official BF16 source and edited BF16 checkpoint
are quantized independently to NF4 before comparison. The static beta 2.4
checkpoint scored 0.0424175691 KL and 0/104 refusal markers.
The independent-quantization protocol includes quantization variance; it is
not a pure BF16-to-BF16 KL measurement.
## Capability slice
Deterministic first-token multiple-choice scoring was run on 854 paired rows:
| Task | Base | Candidate |
|---|---:|---:|
| ARC-Challenge, 256 rows | 72.27% | 72.27% |
| HellaSwag, 256 rows | 64.06% | 64.84% |
| MMLU, 342 rows | 56.14% | 57.89% |
| Combined | 63.35% | 64.29% |
The paired bootstrap mean difference was +0.9368 percentage point. The 95%
interval was [-0.3513, +2.2248] points. Non-inferiority passed at a predeclared
-3 percentage-point margin, and the equivalence gate passed.
## GGUF validation
The BF16, Q8_0 and Q4_K_M files were converted with llama.cpp b10621. Each
file loaded successfully and completed a one-token inference smoke test using
the same llama.cpp build. No full behavioral equivalence claim is made for the
quantized GGUF variants until a dedicated GGUF-native 104-row evaluation is
run.