Instructions to use aselea/Kev-4B-MLX-Serve-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use aselea/Kev-4B-MLX-Serve-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download aselea/Kev-4B-MLX-Serve-8bit --local-dir Kev-4B-MLX-Serve-8bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Kev-4B for mlx-serve
Kev-4B (a LoRA on
Qwen3.5-4B-Base with a pointer head) packed for
mlx-serve's POST /v1/decisions.
Kev answers typed questions about a piece of text (choice, noul, score) with calibrated
probabilities. It never generates text.
The pack folds the LoRA into the base the way kev does on MLX, quantizes the trunk to 8-bit
(affine, group 64; a bf16 build comes from --q-bits 0), and stores the pointer head as
kev_head.safetensors with the calibration temperature in kev_config.json. No PyTorch or pickle
file is needed to serve it.
curl -s localhost:11234/v1/decisions -H 'content-type: application/json' -d '{
"model": "<this model id>",
"state": {"subject": "Charged twice", "body": "Please refund the duplicate today."},
"questions": {"refund": {"type": "noul", "instructions": "Does the customer ask for money back?"}}
}'
Built with tests/convert_kev_weights.py from the mlx-serve repo. Kev and Qwen3.5 are Apache-2.0.
- Downloads last month
- 117
Model size
4B params
Tensor type
U32
路
BF16 路
F32 路
Hardware compatibility
Log In to add your hardware
8-bit