Instructions to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4 # Run inference directly in the terminal: llama cli -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4 # Run inference directly in the terminal: llama cli -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4 # Run inference directly in the terminal: ./llama-cli -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4 # Run inference directly in the terminal: ./build/bin/llama-cli -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Use Docker
docker model run hf.co/BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
- LM Studio
- Jan
- vLLM
How to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
- Ollama
How to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with Ollama:
ollama run hf.co/BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
- Unsloth Desktop
- Pi
How to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with Docker Model Runner:
docker model run hf.co/BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
- Lemonade
How to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Run and chat with the model
lemonade run user.Swift-1.5-Qwen3.8-27B-NVFP4-GGUF-NVFP4
List all available models
lemonade list
- Hermes Agent
How to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4Run Hermes
hermesSwift-1.5-Qwen3.8-27B-NVFP4 GGUF (Q8mix)
A llama.cpp GGUF conversion of ukisai/Swift-1.5-Qwen3.8-27b-NVFP4,
UkisAI's calibrated ModelOpt NVFP4 + FP8 checkpoint of Swift-1.5-Qwen3.8-27B.
The 193 NVFP4 tensors are not requantized — they are repacked into GGML's
native NVFP4 super-block layout with their original per-16 UE4M3 scales and
per-tensor weight_scale_2 factors preserved bit-for-bit.
| File | Size | BPW |
|---|---|---|
Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix.gguf |
19.74 GB | 5.78 (27.32 B params) |
mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf |
885 MB | official UkisAI projector |
Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix.tensor-types.txt |
41 KB | exact per-tensor layout used by llama-quantize |
Tensor layout
| Tensors | Type | Source |
|---|---|---|
MLP ffn_gate/ffn_up/ffn_down for all 64 layers + output (lm_head) — 193 tensors |
NVFP4 | ModelOpt NVFP4, byte-identical to the checkpoint (incl. 386 scale tensors) |
Linear-attention attn_qkv, attn_gate, ssm_out (48 layers); full-attention attn_q/k/v/output (16 layers) — 208 tensors |
Q8_0 | FP8 (E4M3) in the checkpoint, dequantized then Q8_0 |
token_embd |
Q8_0 | BF16 in the checkpoint |
MTP head (blk.64, 15 tensors) |
Q6_K attn / Q4_K FFN, F32 norms | unsloth/Qwen3.8-27B-GGUF MTP/mtp-Qwen3.8-27B-Q4_0.gguf (UkisAI did not fine-tune the MTP head) |
ssm_alpha, ssm_beta, ssm_conv1d, norms, per-tensor scales |
F32 | BF16 in the checkpoint (lossless) |
The MTP head is included (qwen35.nextn_predict_layers = 1), so --spec-type draft-mtp
works as a self-speculative draft.
How it was made
llama.cpp commit 2145525a4081d66ff1a87cf43ef809f95a85ac0c (2026-09-26):
convert_hf_to_gguf.py --outtype bf16 --fp8-as-q8— the converter natively repacks the ModelOptNVFP4group into GGML NVFP4 and turns the FP8 group into Q8_0.llama-quantize --tensor-type-file <the included .tensor-types.txt>with base typeq8_0. NVFP4 and Q8_0 tensors map to their own type and are copied byte-for-byte; the BF16 leftovers become Q8_0; recurrent gates stay F32.blk.64.*(MTP head) spliced byte-for-byte from unsloth's Q4_0 MTP file with a byte-level GGUF rewrite (swap_mtp.pyin the build workspace): metadata and all other 1237 tensors untouched, 185 MB smaller.
Verification
- 1147/1252 tensors byte-identical between the raw conversion and the final file
(everything
llama-quantizecopied). - All 193 NVFP4 tensors independently repacked from the source safetensors and compared byte-for-byte: exact match.
- GGML
nvfp4decode vs ModelOptweight * weight_scale * weight_scale_2: exact (max abs diff 0.0 on sampled tensors). - Metadata/tensor names vs UkisAI's own GGUF: all 41 metadata keys match
(except cosmetic
general.version), names = official set + 386 scale companions, all shapes identical. - Inference smoke test on RTX PRO 4000 Blackwell (24 GB): 182.6 t/s prompt, 20.0 t/s generation; coherent output.
- MTP splice: metadata identical, 15/15
blk.64.*tensors byte-identical to unsloth's file, all other 1237 tensors byte-identical to the first build.
MTP / speculative decoding benchmark
RTX PRO 4000 Blackwell (24 GB), 3 prompts x 256 tokens, greedy (temperature 0),
--spec-type draft-mtp:
| MTP head | gen t/s | speedup | acceptance | mean accepted |
|---|---|---|---|---|
| no spec | 19.8 | 1.00x | - | - |
| NVFP4 (RTN, all 8 weights) | 40.2 | 2.03x | 72.0% | 3.16/4 |
| Q8_0 (original build) | 41.0 | 2.07x | 73.7% | 3.21/4 |
| Q6_K attn / Q4_K FFN (this file) | 41.3 | 2.09x | 73.7% | 3.21/4 |
--spec-draft-n-max 3 is the optimum; asking for more drafts costs throughput
(4/5/6 drafts: 40.5 / 38.1 / 37.2 t/s) because acceptance falls with position
faster than the extra drafts pay off.
An RTN-NVFP4 MTP head (all 8 weights, no calibration) was also built and measured interleaved at n_max 3/4/5, two runs each: it is equal or behind the Q4_K/Q6_K head at every n_max (40.3 / 39.8 / 37.1 t/s average; acceptance 72.0 / 64.4 / 55.6%), so the K-quant head is kept.
Usage
llama-server -m Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix.gguf \
--mmproj mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf \
--spec-type draft-mtp --spec-draft-n-max 3 \
-c 262144 -fa on --jinja
NVFP4 runs natively on NVIDIA Blackwell. Chat template: reasoning_effort
accepts xhigh (default), medium, low. Sampling: temp 1.0, top_p 0.95,
top_k 20 (stored in the GGUF).
License
Swift weights are under the Swift Open License v1.0 (see the source model card). Qwen3.8-27B is Apache 2.0. All credit for the model and the NVFP4 calibration goes to UkisAI.
- Downloads last month
- 873
4-bit
Model tree for BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF
Base model
Qwen/Qwen3.8-27B
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf BroLaurens/Swift-1.5-Qwen3.8-27B-NVFP4-GGUF:NVFP4