Kev-4B for mlx-serve

Kev-4B (a LoRA on Qwen3.5-4B-Base with a pointer head) packed for mlx-serve's POST /v1/decisions.

Kev answers typed questions about a piece of text (choice, noul, score) with calibrated probabilities. It never generates text.

The pack folds the LoRA into the base the way kev does on MLX, quantizes the trunk to 8-bit (affine, group 64; a bf16 build comes from --q-bits 0), and stores the pointer head as kev_head.safetensors with the calibration temperature in kev_config.json. No PyTorch or pickle file is needed to serve it.

curl -s localhost:11234/v1/decisions -H 'content-type: application/json' -d '{
  "model": "<this model id>",
  "state": {"subject": "Charged twice", "body": "Please refund the duplicate today."},
  "questions": {"refund": {"type": "noul", "instructions": "Does the customer ask for money back?"}}
}'

Built with tests/convert_kev_weights.py from the mlx-serve repo. Kev and Qwen3.5 are Apache-2.0.

Downloads last month
117
Safetensors
Model size
4B params
Tensor type
U32
路
BF16
路
F32
路
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for aselea/Kev-4B-MLX-Serve-8bit

Quantized
(9)
this model