DeepSeek-V4-Flash-0731-lossless-CSF

A lossless MXFP4-CSF container of deepseek-ai/DeepSeek-V4-Flash-0731 at revision 9e165c30e2704aec5d9d593cce3eebd58bbef1cb.

CSF stores the routed-expert block-scale planes in a compressed form. Nothing else changes: FP4 weight nibbles, every other tensor, the tensor names, dtypes, shapes and the source metadata are all byte-identical. Decoding is integer arithmetic on bytes, with no requantization or fitting. The original safetensors shards can be restored bit-exactly.

Routed experts (43 layers x 256) - MXFP4 (E2M1, UE8M0 scale per 32)
Attention, shared experts - FP8 E4M3 (128 x 128 blocks, UE8M0 scales)
MTP layers (including their routed experts), embeddings, norms - source format

Routed-expert block scales - lossless MXFP4-CSF (row-base-offset1-u24-exceptions/1)

Sizes

  • Weight files: 166.89 GB in the source, 159.44 GB here (7.45 GB saved).
  • Compressed scales: 33,024 matrices (UE8M0 block scales of the 43 x 256 x 3 main-layer routed-expert projections): 8.66 GB -> 1.21 GB (13.9%).

Provenance

  • Source: deepseek-ai/DeepSeek-V4-Flash-0731 revision 9e165c30e2704aec5d9d593cce3eebd58bbef1cb, 48 index-referenced safetensors shards. During export, each source shard's SHA-256 was checked against its Hugging Face LFS hash.
  • Built with trellis-quant trellis_quant.lossless_scale_checkpoint (commit 60feca330087, family deepseek_v4_flash).
  • verification.json: every original shard was rebuilt from this container, and both the rebuilt and the stored files matched their SHA-256 (33,024 scale matrices, 48 shards, passed).
  • Hub main (7872f01b) differs from 9e165c30 only in README.md.

Layout

lil-mxfp4-csf-checkpoint/1, codec row-base-offset1-u24-exceptions/1:

  • tensors/ - the source shard names; each routed-expert scale <name> is stored as <name>.mxfp4_csf_fixed (uint8) plus <name>.mxfp4_csf_exceptions (uint32)
  • metadata/ - byte copies of the source's config, tokenizer, index, README and LICENSE
  • manifest.json, build-contract.json, receipts/ (per-shard source headers and hashes), verification.json, LICENSE

There is no top-level model.safetensors.index.json, so a plain safetensors loader will not open this directory by mistake.

Serving

Use vLLM with the MXFP4-CSF reader: --quantization mxfp4_csf --load-format mxfp4_csf. Point vLLM at a serving directory that holds the files of metadata/, with config.json's quantization_config replaced by:

{
  "...": "every key of metadata/config.json quantization_config",
  "quant_method": "mxfp4_csf",
  "format_version": 1,
  "checkpoint_root": "/path/to/this/checkpoint"
}

The weights are read from checkpoint_root; the serving directory holds only metadata.

Serving needs a runtime whose MXFP4-CSF reader knows the deepseek_v4_flash family.

Verify or restore

PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint verify \
  --checkpoint /path/to/this/checkpoint --workers 16
PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint restore \
  --checkpoint /path/to/this/checkpoint --destination /path/to/fresh/dir --workers 16

verify rebuilds every original shard and compares SHA-256 values. restore writes the original shards and metadata files into a fresh directory, checking every hash before a file is published. The result is an ordinary copy of the source checkpoint.

License

Same license as the source; LICENSE is copied unchanged from deepseek-ai/DeepSeek-V4-Flash-0731.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for local-inference-lab/DeepSeek-V4-Flash-0731-lossless-CSF

Quantized
(196)
this model