Instructions to use finbase0530/MiniMax-H3-FL2VA-MLX-Serve-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use finbase0530/MiniMax-H3-FL2VA-MLX-Serve-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir MiniMax-H3-FL2VA-MLX-Serve-8bit finbase0530/MiniMax-H3-FL2VA-MLX-Serve-8bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
|
Download MODIFICATIONS.md from finbase0530/MiniMax-H3-FL2VA-MLX-Serve-8bit: direct link, hf CLI and curl.
- Browser
- Download file 1.08 kB
-
https://huggingface.co/finbase0530/MiniMax-H3-FL2VA-MLX-Serve-8bit/resolve/main/MODIFICATIONS.md
- Command line
-
hf download hf://finbase0530/MiniMax-H3-FL2VA-MLX-Serve-8bit/MODIFICATIONS.md
-
curl -L -o MODIFICATIONS.md https://huggingface.co/finbase0530/MiniMax-H3-FL2VA-MLX-Serve-8bit/resolve/main/MODIFICATIONS.md
1.08 kB
Modifications to MiniMax H3
These files are MODIFIED versions of the MiniMax H3 Works, redistributed under the MiniMax H3 Community License Agreement (see LICENSE and NOTICE).
Modified by: mlx-serve (https://github.com/ddalcu/mlx-serve)
What changed:
transformer.safetensorsandtext_encoder.safetensorsare QUANTIZED from the original bfloat16 releases to MLX affine 8-bit, group size 64. Gathered embedding tables and the checkpoint's fp32 islands (patch projections, output heads, time embedder, rope inverse frequencies) are left dense.video_vae.safetensorsandaudio_vae.safetensorsare byte-for-byte copies of the originals, unmodified.- The tokenizer files are byte-for-byte copies from
MiniMaxAI/MiniMax-H3(FL2VA/processor/), relocated into this directory so the model is self-contained. config.jsonis new, written by mlx-serve's converter to describe the layout above. It is not from the original release.
No weights were retrained, distilled, pruned or otherwise altered beyond the numeric quantization described above.