Mamba2_3GB_ReCall_Recovered

Research prerelease: an independently compressed and adapted version of NVIDIA Mamba-2 8B, with an enabled soft Resurface-inspired recall adapter. This is a base language model, not an instruction-tuned assistant. It requires the supplied custom loader; it is not a standard Transformers from_pretrained checkpoint.

This repository mirrors the exact files of GitHub release v0.2.0-resurface under release/, preserving the original manifest hierarchy.

Published quality summary

Fixed released configuration Official WikiText-2 test PPL ↓ Historical synthetic CONFIRM normal MK ↑
Compressed/readapted base, Resurface off 7.53418476 91/384 (23.70%)
Same frozen base, released Resurface on 7.50295894 340/384 (88.54%)

The PPL column uses the complete WikiText-2 test split: 147 reset windows and 300,963 scored tokens. The fixed-release comparison and independent audit were completed after publication, without selecting a new checkpoint. The MK column is the earlier synthetic CONFIRM result on 384 normal prompts; MK was not rerun for the test update. The two columns have different evaluation data and should not be treated as one newly paired test.

Model and storage

The complete pure Mamba-2 architecture is retained: 56 blocks, width 4096, eight SSM groups, 256K vocabulary, untied embedding/head, 8,236,999,680 base parameters, plus 1,154,104 adapter parameters.

Component Representation
112 input/output projections Rotated E8P12 plus signed-axis residual: 20 index bits per 8 values, nominal 2.5 bits/value; scales/transforms/headers counted separately
Embedding / output head Group 128 W4 / W5, with FP16 scales
393 other base tensors Readapted FP16
224 adapter tensors FP16 soft gated cross-head readout at all 56 layers

Encoded base-plus-adapter data occupy 3,141,468,439 bytes (3.141 GB). The original 26-asset GitHub bundle totals 3,165,987,804 bytes (3.166 GB), including tokenizer, software, licenses and reports; this Hub card is additional. The adapter file is 2,539,647 bytes. The E8HUF001 container stores raw members in this version, with no additional Huffman compression.

“3GB” describes stored model data, not GPU memory. The reference loader expands weights to FP16. Its archived-source generation smoke peaked at 22.064 GB allocated GPU memory on a 96 GB RTX PRO 6000; this is not a minimum-VRAM test. Batch-one recurrent/conv cache was 122,028,032 bytes (116.375 MiB). There is no compressed-resident GPU kernel or demonstrated ASIC/neuromorphic deployment.

Measured quality

Complete WikiText-2 validation PPL
Original NVIDIA weights cast to FP16 7.334175947
Compressed/readapted base 7.622396588
Same base with released soft adapter 7.593163114

The adapter improves paired compressed-base PPL by 0.383521%; candidate PPL remains 3.531237% above original FP16. Evaluation covers 130 reset windows and 264,764 targets. The source number is an aligned historical evaluation; current/adapter/restored runs were paired. This validation corpus informed development and is not an untouched test set.

Independent numeric-binding CONFIRM Recall Target-removed false matches
Compressed/readapted base 91/384 (23.6979%) 0/384
Same base with released adapter 340/384 (88.5417%) 0/384

CONFIRM uses new numeric instances, three shared template families and 16/64 records per prompt. The paired gain is 64.84375 percentage points, with 251 gains and 2 losses. The uncompressed source was not evaluated on CONFIRM. Historical DEV recall was 167/384 for original FP16 versus 346/384 for the trained candidate; these models had unequal adaptation budgets.

The adapter operates after native D*x, before grouped gated RMSNorm. The same soft policy remains enabled for recall and PPL, without extra recurrent state. It was trained for 1536 updates on synthetic numeric bindings plus prose CE/KL and gate-closure regularization; the prose teacher was the frozen compressed base. All 507 base tensors stayed unchanged during adapter training. This is a post-D Resurface-inspired variant, not an exact original-method port.

See the training protocol and full results. An initial DEV run failed historical cross-process output equality; that failure is preserved. A disclosed continuation removed that prerequisite, retaining the same candidate and quality thresholds. Same-process restored controls, complete PPL and subsequent fresh CONFIRM passed. Cross-process bitwise determinism remains unresolved. No unseen-template, longer-context, general task parity, global-smallest or equal-training-budget claim is made.

Download, verify and run

Use the official hf CLI. Restore needs Python 3.10+, PyTorch and a C++17 compiler; generation needs compatible Linux/CUDA and the public Mamba runtime. Allow additional disk space for restored data and the temporary joined container.

hf download EndlessChasing/Mamba2_3GB_ReCall_Recovered \
  --revision v0.2.0-resurface --local-dir ./model-download
mkdir model-software
unzip model-download/release/source.zip -d model-software

python3 -m pip install -e ./model-software
python3 -m pip install 'mamba-ssm==2.3.2.post1' --no-build-isolation

python3 model-software/scripts/package_release.py verify \
  --release-dir ./model-download/release \
  --expected-manifest-sha256 da5931dc8315bf576b773abdf4c77828a2994ccaf4fd14858c7798235b2ef19d

python3 model-software/scripts/package_release.py restore \
  --release-dir ./model-download/release --output ./restored-model \
  --expected-manifest-sha256 da5931dc8315bf576b773abdf4c77828a2994ccaf4fd14858c7798235b2ef19d

python3 -m mamba_e8w5.release_generate \
  --model-dir ./restored-model/raw \
  --prompt "The capital of France is" --max-new-tokens 12 \
  --repeat --report ./generation-receipt.json

The loader verifies all 507 decoded base and 224 adapter tensor hashes and always installs the soft adapter. Original full-precision weights, Hessians and training data are unnecessary. Keep reports outside release/ and restored-model/raw/. See the environment and release guide for dependency details. Context plus generation must not exceed 4096 tokens.

Attribution and license scopes

  • Upstream NVIDIA model weights and tokenizer: Apache-2.0, pinned revision b915550c63ba9359f88f44d1f6a600d85af27302. This independently modified model is not endorsed by NVIDIA.
  • Repository software and QuIP#-derived E8/LDLQ components: GPL-3.0. Corresponding source and notices accompany the release. The mixed-license metadata does not replace the component license texts or imply that applying a GPL quantizer automatically relicenses every model weight.
  • Native Mamba runtime: Apache-2.0. The new public adapter implementation and trained artifact are distinct from excluded private reference materials.
  • WikiText was obtained separately under its upstream terms; its text and tokenized passages are not included in this distribution.

Read the complete license scope and attribution and the original license texts shipped under release/licenses/.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EndlessChasing/Mamba2_3GB_ReCall_Recovered

Quantized
(2)
this model

Dataset used to train EndlessChasing/Mamba2_3GB_ReCall_Recovered