Instructions to use Ninnix96/Qwengram-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Ninnix96/Qwengram-4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Ninnix96/Qwengram-4B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Ninnix96/Qwengram-4B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Ninnix96/Qwengram-4B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Ninnix96/Qwengram-4B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Ninnix96/Qwengram-4B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Ninnix96/Qwengram-4B:Q4_K_M
Use Docker
docker model run hf.co/Ninnix96/Qwengram-4B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Ninnix96/Qwengram-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ninnix96/Qwengram-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ninnix96/Qwengram-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Ninnix96/Qwengram-4B:Q4_K_M
- Ollama
How to use Ninnix96/Qwengram-4B with Ollama:
ollama run hf.co/Ninnix96/Qwengram-4B:Q4_K_M
- Unsloth Desktop
- Pi
How to use Ninnix96/Qwengram-4B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Ninnix96/Qwengram-4B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Ninnix96/Qwengram-4B with Docker Model Runner:
docker model run hf.co/Ninnix96/Qwengram-4B:Q4_K_M
- Lemonade
How to use Ninnix96/Qwengram-4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Ninnix96/Qwengram-4B:Q4_K_M
Run and chat with the model
lemonade run user.Qwengram-4B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Ninnix96/Qwengram-4B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Ninnix96/Qwengram-4B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Ninnix96/Qwengram-4B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Ninnix96/Qwengram-4B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Ninnix96/Qwengram-4B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download runtime/README.md from Ninnix96/Qwengram-4B: direct link, hf CLI and curl.
- Browser
- Download file 3.4 kB
-
https://huggingface.co/Ninnix96/Qwengram-4B/resolve/main/runtime/README.md
- Command line
-
hf download hf://Ninnix96/Qwengram-4B/runtime/README.md
-
curl -L -o README.md https://huggingface.co/Ninnix96/Qwengram-4B/resolve/main/runtime/README.md
3.4 kB
| # Qwengram-4B GGUF runtime validation | |
| Canonical checkpoint: **REAL-15M + linear750**. See the [frozen decision](../evaluation/decision.md) for the independent 10M/15M comparison. | |
| | Precision | Stock NLL | Qwengram NLL | Reader gain [95% CI] | Gain retention [95% CI] | Perplexity reduction vs stock | | |
| | --- | ---: | ---: | ---: | ---: | ---: | | |
| | BF16 | 2.225750 | 2.198449 | 0.027301 [0.022650, 0.031917] | 100% | 2.69% | | |
| | Q8_0 | 2.226990 | 2.199469 | 0.027521 [0.022892, 0.032133] | 100.8% [97.1%, 104.7%] | 2.71% | | |
| | Q6_K | 2.230865 | 2.200943 | 0.029922 [0.023478, 0.036439] | 109.6% [99.4%, 118.6%] | 2.95% | | |
| | Q4_K_M | 2.262792 | 2.231828 | 0.030964 [0.025884, 0.036537] | 113.4% [100.6%, 127.6%] | 3.05% | | |
| ## Matched test | |
| The matched CPU test scores 8,128 tokens from the first 64 consecutive 256-token WikiText-2 raw test chunks, scoring the last 127 tokens per chunk. All runs use eight threads and context/batch/microbatch 256, with no warmup. Reader gain is `NLL(stock) - NLL(Qwengram)`; retention divides each quantized gain by the BF16 gain. Paired 95% intervals use 10,000 resamples of 16 consecutive four-chunk blocks, seed 1234. The external PLE is Ivan Fioravanti's Q4_1 sidecar. | |
| BF16 has the lowest absolute Qwengram NLL in this test. Gain retention measures the added PLE benefit within each precision. | |
| All eight stock/Qwengram runs were measured with the same patched runtime. Per-chunk scores and model hashes are in [results.json](results.json); commands, timings and logs are in [logs](logs/). | |
| All 441 backbone tensors and nine tokenizer fields match exactly within each stock/Qwengram pair. All 11 reader/arbiter tensors remain bit-exact FP32 across BF16, Q8_0, Q6_K and Q4_K_M and match the verified reader and matching arbiter. See [packaging](packaging.json) and [verification](verification.json). | |
| The [PLE sidecar check](sidecar-validation.json) compares 4,096 addressed rows with independent reference addressing and dequantization. Full prefill, split prefill, token decode, EOS boundaries and repeated reset have max absolute difference zero. The PLE remains unchanged. | |
| This CPU test uses a quantized Q4_1 PLE and a WikiText-2 slice. The frozen Kaggle study uses the original FP8 PLE and different evaluation streams. These scores are separate benchmarks. | |
| ## Generation checks | |
| On the tested AMD BC-250, BF16, Q8_0, Q4_K_M with full Vulkan offload (`-ngl 99`) matched the corresponding CPU eight-token greedy continuation for `The capital of France is`. These short checks do not establish broad GPU parity. Q6_K differed at full offload but matched with `-ngl 33`, which keeps the first decoder layer on CPU. Use `-ngl 33` for Q6_K on this tested device, or CPU (`-ngl 0`). The fallback is also a short generation check. | |
| The loader rejected missing and invalid PLE sidecars and conflicting injection indices. The prior 0.8B and 2B Q8_0 CPU continuations remain exact with the updated runtime. See [smoke.json](smoke.json), [Q6 output](q6-smoke.json), [Q6 fallback](q6-fallback.json) and [legacy checks](legacy-smoke.json). | |
| ## Runtime source | |
| Published fork: [3616a858f2326e87ad8b48e1341a4e341b3dad73](https://github.com/Ninnix/llama.cpp-qwengram/commit/3616a858f2326e87ad8b48e1341a4e341b3dad73). The tested base commit, source patch and binary hashes are recorded in verification.json; [the source patch](llama-qwengram-4b.patch) reproduces the tested source. | |