DeepSeek-V4.1-Flash-lossless-CSF
A lossless MXFP4-CSF container of deepseek-ai/DeepSeek-V4.1-Flash at revision dba1be0a40aa45a94ad051997016db3960a90277.
CSF stores the routed-expert block-scale planes in a compressed form. Nothing else changes: FP4 weight nibbles, every other tensor, the tensor names, dtypes, shapes and the source metadata are all byte-identical. Decoding is integer arithmetic on bytes, with no requantization or fitting. The original safetensors shards can be restored bit-exactly.
Routed experts (40 layers x 384) - MXFP4 (E2M1, UE8M0 scale per 32)
Engram tables, attention, shared experts - FP8 E4M3 with UE8M0 scales
Vision encoder, aligner, MTP (including its routed experts) - source format
Routed-expert block scales - lossless MXFP4-CSF (row-base-offset1-u24-exceptions/1)
Sizes
- Weight files: 510.30 GB in the source, 495.63 GB here (14.67 GB saved).
- Compressed scales: 46,080 matrices (UE8M0 block scales of the 40 x 384 x 3 main-layer routed-expert projections): 16.99 GB -> 2.31 GB (13.6%).
Provenance
- Source:
deepseek-ai/DeepSeek-V4.1-Flashrevisiondba1be0a40aa45a94ad051997016db3960a90277, 48 index-referenced safetensors shards. During export, each source shard's SHA-256 was checked against its Hugging Face LFS hash. - Built with trellis-quant
trellis_quant.lossless_scale_checkpoint(commit60feca330087, familydeepseek_v41). verification.json: every original shard was rebuilt from this container, and both the rebuilt and the stored files matched their SHA-256 (46,080 scale matrices, 48 shards, passed).- Hub
main(2cba9e42) addschat_template.jinjaand is otherwise byte-identical to dba1be0a. Every tensor and config file matches. tensors/is byte-identical (all 48 compressed-shard SHA-256 values) tolocal-inference-lab/DeepSeek-V4.1-Flash-MXFP4-CSFrevision 872da235. This build adds the source LICENSE tometadata/and records the source revision.
Layout
lil-mxfp4-csf-checkpoint/1, codec row-base-offset1-u24-exceptions/1:
tensors/- the source shard names; each routed-expert scale<name>is stored as<name>.mxfp4_csf_fixed(uint8) plus<name>.mxfp4_csf_exceptions(uint32)metadata/- byte copies of the source's config, tokenizer, index, README and LICENSEmanifest.json,build-contract.json,receipts/(per-shard source headers and hashes),verification.json,LICENSE
There is no top-level model.safetensors.index.json, so a plain safetensors loader will not open this directory by mistake.
Serving
Use vLLM with the MXFP4-CSF reader: --quantization mxfp4_csf --load-format mxfp4_csf. Point vLLM at a serving directory that holds the files of metadata/, with config.json's quantization_config replaced by:
{
"...": "every key of metadata/config.json quantization_config",
"quant_method": "mxfp4_csf",
"format_version": 1,
"checkpoint_root": "/path/to/this/checkpoint"
}
The weights are read from checkpoint_root; the serving directory holds only metadata.
Verify or restore
PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint verify \
--checkpoint /path/to/this/checkpoint --workers 16
PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint restore \
--checkpoint /path/to/this/checkpoint --destination /path/to/fresh/dir --workers 16
verify rebuilds every original shard and compares SHA-256 values. restore writes the original shards and metadata files into a fresh directory, checking every hash before a file is published. The result is an ordinary copy of the source checkpoint.
License
Same license as the source; LICENSE is copied unchanged from deepseek-ai/DeepSeek-V4.1-Flash.
Model tree for local-inference-lab/DeepSeek-V4.1-Flash-lossless-CSF
Base model
deepseek-ai/DeepSeek-V4.1-Flash