Mamba2_3GB_ReCall_Recovered
Research prerelease: an independently compressed and adapted version of
NVIDIA Mamba-2 8B, with an enabled
soft Resurface-inspired recall adapter. This is a base language model, not an
instruction-tuned assistant. It requires the supplied custom loader; it is not
a standard Transformers from_pretrained checkpoint.
This repository mirrors the exact files of
GitHub release v0.2.0-resurface
under release/, preserving the original manifest hierarchy.
Published quality summary
| Fixed released configuration | Official WikiText-2 test PPL ↓ | Historical synthetic CONFIRM normal MK ↑ |
|---|---|---|
| Compressed/readapted base, Resurface off | 7.53418476 | 91/384 (23.70%) |
| Same frozen base, released Resurface on | 7.50295894 | 340/384 (88.54%) |
The PPL column uses the complete WikiText-2 test split: 147 reset windows and 300,963 scored tokens. The fixed-release comparison and independent audit were completed after publication, without selecting a new checkpoint. The MK column is the earlier synthetic CONFIRM result on 384 normal prompts; MK was not rerun for the test update. The two columns have different evaluation data and should not be treated as one newly paired test.
Model and storage
The complete pure Mamba-2 architecture is retained: 56 blocks, width 4096, eight SSM groups, 256K vocabulary, untied embedding/head, 8,236,999,680 base parameters, plus 1,154,104 adapter parameters.
| Component | Representation |
|---|---|
| 112 input/output projections | Rotated E8P12 plus signed-axis residual: 20 index bits per 8 values, nominal 2.5 bits/value; scales/transforms/headers counted separately |
| Embedding / output head | Group 128 W4 / W5, with FP16 scales |
| 393 other base tensors | Readapted FP16 |
| 224 adapter tensors | FP16 soft gated cross-head readout at all 56 layers |
Encoded base-plus-adapter data occupy 3,141,468,439 bytes (3.141 GB). The original 26-asset GitHub bundle totals 3,165,987,804 bytes (3.166 GB), including tokenizer, software, licenses and reports; this Hub card is additional. The adapter file is 2,539,647 bytes. The E8HUF001 container stores raw members in this version, with no additional Huffman compression.
“3GB” describes stored model data, not GPU memory. The reference loader expands weights to FP16. Its archived-source generation smoke peaked at 22.064 GB allocated GPU memory on a 96 GB RTX PRO 6000; this is not a minimum-VRAM test. Batch-one recurrent/conv cache was 122,028,032 bytes (116.375 MiB). There is no compressed-resident GPU kernel or demonstrated ASIC/neuromorphic deployment.
Measured quality
| Complete WikiText-2 validation | PPL |
|---|---|
| Original NVIDIA weights cast to FP16 | 7.334175947 |
| Compressed/readapted base | 7.622396588 |
| Same base with released soft adapter | 7.593163114 |
The adapter improves paired compressed-base PPL by 0.383521%; candidate PPL remains 3.531237% above original FP16. Evaluation covers 130 reset windows and 264,764 targets. The source number is an aligned historical evaluation; current/adapter/restored runs were paired. This validation corpus informed development and is not an untouched test set.
| Independent numeric-binding CONFIRM | Recall | Target-removed false matches |
|---|---|---|
| Compressed/readapted base | 91/384 (23.6979%) | 0/384 |
| Same base with released adapter | 340/384 (88.5417%) | 0/384 |
CONFIRM uses new numeric instances, three shared template families and 16/64 records per prompt. The paired gain is 64.84375 percentage points, with 251 gains and 2 losses. The uncompressed source was not evaluated on CONFIRM. Historical DEV recall was 167/384 for original FP16 versus 346/384 for the trained candidate; these models had unequal adaptation budgets.
The adapter operates after native D*x, before grouped gated RMSNorm. The
same soft policy remains enabled for recall and PPL, without extra recurrent
state. It was trained for 1536 updates on synthetic numeric bindings plus prose
CE/KL and gate-closure regularization; the prose teacher was the frozen
compressed base. All 507 base tensors stayed unchanged during adapter training.
This is a post-D Resurface-inspired variant, not an exact original-method port.
See the training protocol and full results. An initial DEV run failed historical cross-process output equality; that failure is preserved. A disclosed continuation removed that prerequisite, retaining the same candidate and quality thresholds. Same-process restored controls, complete PPL and subsequent fresh CONFIRM passed. Cross-process bitwise determinism remains unresolved. No unseen-template, longer-context, general task parity, global-smallest or equal-training-budget claim is made.
Download, verify and run
Use the official hf CLI. Restore needs Python 3.10+, PyTorch and a C++17
compiler; generation needs compatible Linux/CUDA and the public Mamba runtime.
Allow additional disk space for restored data and the temporary joined container.
hf download EndlessChasing/Mamba2_3GB_ReCall_Recovered \
--revision v0.2.0-resurface --local-dir ./model-download
mkdir model-software
unzip model-download/release/source.zip -d model-software
python3 -m pip install -e ./model-software
python3 -m pip install 'mamba-ssm==2.3.2.post1' --no-build-isolation
python3 model-software/scripts/package_release.py verify \
--release-dir ./model-download/release \
--expected-manifest-sha256 da5931dc8315bf576b773abdf4c77828a2994ccaf4fd14858c7798235b2ef19d
python3 model-software/scripts/package_release.py restore \
--release-dir ./model-download/release --output ./restored-model \
--expected-manifest-sha256 da5931dc8315bf576b773abdf4c77828a2994ccaf4fd14858c7798235b2ef19d
python3 -m mamba_e8w5.release_generate \
--model-dir ./restored-model/raw \
--prompt "The capital of France is" --max-new-tokens 12 \
--repeat --report ./generation-receipt.json
The loader verifies all 507 decoded base and 224 adapter tensor hashes and always
installs the soft adapter. Original full-precision weights, Hessians and training
data are unnecessary. Keep reports outside release/ and restored-model/raw/.
See the environment and release guide
for dependency details. Context plus generation must not exceed 4096 tokens.
Attribution and license scopes
- Upstream NVIDIA model weights and tokenizer: Apache-2.0, pinned revision
b915550c63ba9359f88f44d1f6a600d85af27302. This independently modified model is not endorsed by NVIDIA. - Repository software and QuIP#-derived E8/LDLQ components: GPL-3.0. Corresponding source and notices accompany the release. The mixed-license metadata does not replace the component license texts or imply that applying a GPL quantizer automatically relicenses every model weight.
- Native Mamba runtime: Apache-2.0. The new public adapter implementation and trained artifact are distinct from excluded private reference materials.
- WikiText was obtained separately under its upstream terms; its text and tokenized passages are not included in this distribution.
Read the complete license scope and attribution
and the original license texts shipped under release/licenses/.
Model tree for EndlessChasing/Mamba2_3GB_ReCall_Recovered
Base model
nvidia/mamba2-8b-3t-4k