DeepSeek-V4-Flash-0731-lossless-CSF
A lossless MXFP4-CSF container of deepseek-ai/DeepSeek-V4-Flash-0731 at revision 9e165c30e2704aec5d9d593cce3eebd58bbef1cb.
CSF stores the routed-expert block-scale planes in a compressed form. Nothing else changes: FP4 weight nibbles, every other tensor, the tensor names, dtypes, shapes and the source metadata are all byte-identical. Decoding is integer arithmetic on bytes, with no requantization or fitting. The original safetensors shards can be restored bit-exactly.
Routed experts (43 layers x 256) - MXFP4 (E2M1, UE8M0 scale per 32)
Attention, shared experts - FP8 E4M3 (128 x 128 blocks, UE8M0 scales)
MTP layers (including their routed experts), embeddings, norms - source format
Routed-expert block scales - lossless MXFP4-CSF (row-base-offset1-u24-exceptions/1)
Sizes
- Weight files: 166.89 GB in the source, 159.44 GB here (7.45 GB saved).
- Compressed scales: 33,024 matrices (UE8M0 block scales of the 43 x 256 x 3 main-layer routed-expert projections): 8.66 GB -> 1.21 GB (13.9%).
Provenance
- Source:
deepseek-ai/DeepSeek-V4-Flash-0731revision9e165c30e2704aec5d9d593cce3eebd58bbef1cb, 48 index-referenced safetensors shards. During export, each source shard's SHA-256 was checked against its Hugging Face LFS hash. - Built with trellis-quant
trellis_quant.lossless_scale_checkpoint(commit60feca330087, familydeepseek_v4_flash). verification.json: every original shard was rebuilt from this container, and both the rebuilt and the stored files matched their SHA-256 (33,024 scale matrices, 48 shards, passed).- Hub
main(7872f01b) differs from 9e165c30 only in README.md.
Layout
lil-mxfp4-csf-checkpoint/1, codec row-base-offset1-u24-exceptions/1:
tensors/- the source shard names; each routed-expert scale<name>is stored as<name>.mxfp4_csf_fixed(uint8) plus<name>.mxfp4_csf_exceptions(uint32)metadata/- byte copies of the source's config, tokenizer, index, README and LICENSEmanifest.json,build-contract.json,receipts/(per-shard source headers and hashes),verification.json,LICENSE
There is no top-level model.safetensors.index.json, so a plain safetensors loader will not open this directory by mistake.
Serving
Use vLLM with the MXFP4-CSF reader: --quantization mxfp4_csf --load-format mxfp4_csf. Point vLLM at a serving directory that holds the files of metadata/, with config.json's quantization_config replaced by:
{
"...": "every key of metadata/config.json quantization_config",
"quant_method": "mxfp4_csf",
"format_version": 1,
"checkpoint_root": "/path/to/this/checkpoint"
}
The weights are read from checkpoint_root; the serving directory holds only metadata.
Serving needs a runtime whose MXFP4-CSF reader knows the deepseek_v4_flash family.
Verify or restore
PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint verify \
--checkpoint /path/to/this/checkpoint --workers 16
PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint restore \
--checkpoint /path/to/this/checkpoint --destination /path/to/fresh/dir --workers 16
verify rebuilds every original shard and compares SHA-256 values. restore writes the original shards and metadata files into a fresh directory, checking every hash before a file is published. The result is an ordinary copy of the source checkpoint.
License
Same license as the source; LICENSE is copied unchanged from deepseek-ai/DeepSeek-V4-Flash-0731.
Model tree for local-inference-lab/DeepSeek-V4-Flash-0731-lossless-CSF
Base model
deepseek-ai/DeepSeek-V4-Flash-0731