EndlessChasing's picture
Publish verified Mamba2-8B E8/W5 Resurface model
d719dc2 verified
|
Raw History Blame Contribute Delete
3.66 kB

Mamba2-8B axis E8 + soft Resurface — v0.2.0-resurface

Public research prerelease of NVIDIA Mamba2-8B with 112 E8-plus-axis projections, W4 input embeddings, W5 output head, 393 readapted FP16 small tensors and the same enabled soft Resurface-inspired adapter for every task. No external weight base or original checkpoint is needed for inference. The adapter is mandatory.

Measured quality

Full WikiText-2 validation: PPL 7.593163113563, compared with 7.622396587826 for the readapted compressed base without the adapter; 130 windows, 264,764 predicted tokens, 2048-token windows with the declared final tail. Validation informed development; these are not untouched test results.

The aligned historical original FP16 source scored 7.334175947319; this release is +3.531237% higher (worse) in PPL. The original source was not rerun in the soft-adapter evaluation process. This is not an equal-quality claim.

Fresh synthetic multi-key confirmation: 340/384 normal exact answers versus 91/384 without the adapter. Target-removed controls: 0/384 versus 0/384. These task-specific results do not establish general reasoning performance. Original-source MK was not measured on this fresh confirmation split; earlier DEV comparisons are separately documented in the repository.

The original strict DEV run stopped on cross-process historical MK mismatch. A separately declared continuation retained the same final adapter and required same-process baseline restoration, full PPL nonregression and fresh confirmation. The release does not claim that the original strict prerequisite passed. Exact reports and hashes are shipped with this distribution.

Storage and execution

The base model data is 3,138,928,792 bytes; the actual serialized adapter is 2,539,647 bytes, totaling 3,141,468,439 bytes before tokenizer and manifests. The adapter has 1,154,104 parameters (2,308,208 FP16 payload bytes). All shipped assets and metadata are counted in release_manifest.json.

This E8HUF001 envelope uses raw members: it is exact to the quantized files and does not add entropy compression. Quantization itself is lossy relative to the original model. The native quality runtime expands base weights to FP16; the download size is not its GPU memory usage. Native Mamba dependencies and a compatible CUDA environment are required; no optimized compressed GPU kernel, ASIC performance or globally smallest-model claim is made.

Download, restore, run

Use scripts/download_release.py with tag v0.2.0-resurface, then extract source.zip. Run the archived scripts/package_release.py restore with the independently published release-manifest SHA to restore exact files. Load the restored raw directory only through mamba_e8w5.release_runtime.load_model, or use archived python -m mamba_e8w5.release_generate --help. That entry point verifies the soft adapter and all model identities; directly calling the old base loader would omit it.

See archived docs/DOWNLOAD.md and docs/RESURFACE_RELEASE.md for exact commands. Corresponding source is in source.zip; license notices are in licenses/. The upstream NVIDIA model/tokenizer and native Mamba runtime are Apache-2.0; this repository's software, including the QuIP-derived codec, is GPL-3.0. Read the shipped notices for the component-specific terms. Only selected pinned evidence is archived. For other report links in the docs, use the public repository at tag v0.2.0-resurface; these reports are not inference inputs. The original checkpoint is needed only to reproduce quantization/training, not to download, restore or use this package.