Instructions to use Accio-Lab/occamy-1.0-APEX-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Accio-Lab/occamy-1.0-APEX-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Accio-Lab/occamy-1.0-APEX-GGUF # Run inference directly in the terminal: llama cli -hf Accio-Lab/occamy-1.0-APEX-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Accio-Lab/occamy-1.0-APEX-GGUF # Run inference directly in the terminal: llama cli -hf Accio-Lab/occamy-1.0-APEX-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Accio-Lab/occamy-1.0-APEX-GGUF # Run inference directly in the terminal: ./llama-cli -hf Accio-Lab/occamy-1.0-APEX-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Accio-Lab/occamy-1.0-APEX-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf Accio-Lab/occamy-1.0-APEX-GGUF
Use Docker
docker model run hf.co/Accio-Lab/occamy-1.0-APEX-GGUF
- LM Studio
- Jan
- vLLM
How to use Accio-Lab/occamy-1.0-APEX-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Accio-Lab/occamy-1.0-APEX-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Accio-Lab/occamy-1.0-APEX-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Accio-Lab/occamy-1.0-APEX-GGUF
- Ollama
How to use Accio-Lab/occamy-1.0-APEX-GGUF with Ollama:
ollama run hf.co/Accio-Lab/occamy-1.0-APEX-GGUF
- Unsloth Desktop
- Pi
How to use Accio-Lab/occamy-1.0-APEX-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Accio-Lab/occamy-1.0-APEX-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Accio-Lab/occamy-1.0-APEX-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Accio-Lab/occamy-1.0-APEX-GGUF with Docker Model Runner:
docker model run hf.co/Accio-Lab/occamy-1.0-APEX-GGUF
- Lemonade
How to use Accio-Lab/occamy-1.0-APEX-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Accio-Lab/occamy-1.0-APEX-GGUF
Run and chat with the model
lemonade run user.occamy-1.0-APEX-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use Accio-Lab/occamy-1.0-APEX-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Accio-Lab/occamy-1.0-APEX-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Accio-Lab/occamy-1.0-APEX-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Accio-Lab/occamy-1.0-APEX-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Accio-Lab/occamy-1.0-APEX-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Accio-Lab/occamy-1.0-APEX-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Occamy-1.0 · APEX GGUF
Three plain profiles · four calibrated profiles, including Mini
Occamy-1.0 quantized with the LocalAI team's APEX profiles. APEX allocates GGUF precision by tensor role and layer position: routed experts, shared experts and attention receive different allocations. The three plain profiles and four calibrated I- profiles use the pinned upstream 40-layer configurations. Every file is generated directly from the same BF16 source. The calibrated group adds I-Compact, I-Quality, I-Balanced and I-Mini.
Downloads
| Profile | Exact bytes | GB | GiB | Download |
|---|---|---|---|---|
| Compact | 16,538,851,072 | 16.54 | 15.403 | GGUF |
| Quality | 22,819,401,472 | 22.82 | 21.252 | GGUF |
| Balanced | 25,335,983,872 | 25.34 | 23.596 | GGUF |
Calibrated profiles
| Profile | Exact bytes | GB | GiB | Download |
|---|---|---|---|---|
| I-Compact | 16,538,851,360 | 16.54 | 15.403 | GGUF |
| I-Quality | 22,819,401,760 | 22.82 | 21.252 | GGUF |
| I-Balanced | 25,335,984,160 | 25.34 | 23.596 | GGUF |
| I-Mini | 13,467,210,784 | 13.47 | 12.542 | GGUF |
Profile names describe allocation strategies; file size does not follow the name order. These are weight-file sizes, not peak memory requirements. Allow additional memory for the runtime, context and image input.
Precision allocation
| Profile | Routed experts: layers 0–4 / 5–9 / 10–29 / 30–34 / 35–39 | Shared experts | Attention / SSM projections |
|---|---|---|---|
| Compact | Q4_K / Q3_K / Q3_K / Q3_K / Q4_K | Q6_K | Q4_K |
| Quality | Q6_K / Q5_K / IQ4_XS / Q5_K / Q6_K | Q8_0 | Q6_K |
| Balanced | Q6_K / Q5_K / Q5_K / Q5_K / Q6_K | Q8_0 | Q6_K |
| I-Mini | Q3_K / Q3_K / IQ2_S / Q3_K / Q3_K | Q5_K in first/last 5 layers; Q4_K elsewhere | Q4_K in first/last 3 layers; Q3_K elsewhere |
I-Compact, I-Quality and I-Balanced use the same allocations as their plain counterparts, with the validated importance matrix. I-Mini uses the upstream mini profile, which requires calibration.
Embeddings, output and tensors outside the profile overrides follow the pinned stock quantizer's Q4_K_M base recipe for Compact and Q6_K for Quality/Balanced and their I variants; I-Mini uses Q3_K_M. I-Compact uses Q4_K_M. Small tensors that the quantizer excludes retain their source type. All actual tensor types are recorded in artifact checks; this is mixed precision, not a uniform bit width.
Run with llama.cpp
Use a recent llama.cpp build with Qwen3.5 MoE support. The validated revision is 972d2313bc0bf0a45f634f77d95c9fb03aeab12c.
hf download Accio-Lab/occamy-1.0-APEX-GGUF occamy-1.0-APEX-Compact.gguf \
--local-dir ./occamy-gguf
llama-server -m ./occamy-gguf/occamy-1.0-APEX-Compact.gguf \
-ngl 999 -c 8192 -np 1 -fa on --jinja \
--host 127.0.0.1 --port 8000 --alias occamy
Replace Compact with Quality, Balanced, I-Compact, I-Quality, I-Balanced or I-Mini to use another profile. Input content should be NFC-normalized; all files store the source-compatible qwen2 pre-tokenizer metadata. See the documented tokenizer protocol.
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"occamy","messages":[{"role":"user","content":"Compute 2+2. Answer briefly."}],"temperature":0,"max_tokens":128,"chat_template_kwargs":{"enable_thinking":false}}'
For vision, download the matching existing F16 projector and add --mmproj ./occamy-gguf/mmproj-occamy-1.0-F16.gguf. It is a separate 899,282,944-byte file in the standard GGUF repository. Image inference was not retested for this APEX release.
hf download Accio-Lab/occamy-1.0-GGUF mmproj-occamy-1.0-F16.gguf \
--local-dir ./occamy-gguf
Measured validation
| Model | Authored functional cases | WikiText subset PPL |
|---|---|---|
| BF16 | 19/20 | 8.2963 +/- 0.36372 |
| APEX-Compact | 20/20 | 8.4567 +/- 0.36569 |
| APEX-Quality | 19/20 | 8.2568 +/- 0.36068 |
| APEX-Balanced | 19/20 | 8.2994 +/- 0.36385 |
Calibrated profiles, measured with the same protocol:
| Model | Authored functional cases | WikiText subset PPL |
|---|---|---|
| APEX-I-Compact | 19/20 | 8.5074 +/- 0.37314 |
| APEX-I-Quality | 19/20 | 8.3743 +/- 0.36829 |
| APEX-I-Balanced | 19/20 | 8.3482 +/- 0.36702 |
| APEX-I-Mini | 20/20 | 9.5643 +/- 0.43416 |
Calibration did not improve the paired WikiText subset PPL for the three matched profiles in this run. I-Compact, I-Quality and I-Balanced fail the same Chinese JSON case (json_4) as BF16; no forced JSON grammar was used. I-Mini is an experimental size/quality tradeoff: its subset PPL is higher, while all 20 authored fixtures pass. All stored code outputs for the eight models were rescored after adding the ordinary isinstance builtin to the restricted Python harness; generation and assertions were unchanged. Original execution results and the rescoring receipt are preserved.
All seven files pass complete finite-value checks for all 733 tensors, logical shape and vocabulary/template preservation, and 16/16 NFC tokenizer probes with special-token parsing. The functional checks include English/Chinese instructions, arithmetic, executed Python, strict JSON, conversation memory and native tool round trips. Plain-profile evidence remains in validation. Calibrated per-case outputs, failures, exact file hashes and reproduction are in calibrated validation.
The BF16 reference is reused from the preceding run with the same pinned source, runtime and inputs. These are 20 authored fixtures and a small WikiText subset, not full quality benchmarks. Upstream APEX results on Qwen are not Occamy measurements. CPU-only and Apple Silicon execution, image inference and long-context quality were not tested in this run.
Calibration and MTP
The four I- profiles use a 192,232,448-byte importance matrix collected from BF16. The fixed calibration mix has 608 training conversations: 192 UltraChat, 96 CodeFeedback, 128 Glaive function-calling, 96 GSM8K and 96 authored multilingual/state-tracking examples. Sources, revisions, licenses, selected rows and hashes are recorded in the manifest. Source inputs were NFC-normalized and rendered with the pinned chat template, with thinking disabled and a maximum of 1,536 tokens per conversation. No held-out WikiText or functional-fixture prompt was used for calibration.
Four disjoint shards were processed with stock llama-imatrix, context 512, special-token parsing and output-projection collection. The combined matrix covers 712 chunks / 364,544 processed tokens. All 120 routed projection matrices cover all 256 experts: 30,720/30,720 expert-projection pairs observed, with a minimum activation count of 46. Stored sums/counts are finite and nonnegative. This is observed coverage, not equal per-expert sampling.
The matrix is GGUF importance-matrix data saved with a .bin suffix, not a model checkpoint. Its SHA256 is 81a0e2944b9c2290f924926a96d98b9ddc636ad1fb8accdfd596e99e6f29f160. Raw training conversations are not redistributed here; the pinned preparation recipe, manifest and coverage audit are included. See reproduction.
Calibration does not guarantee a quality improvement. Compare the measured subset results below; upstream APEX benchmark results on other checkpoints are not Occamy measurements.
The three plain profiles do not use an importance matrix. The four I- profiles are calibrated. There is no bundled MTP head in these files. The Occamy MTP checkpoint remains separate and does not establish llama.cpp speculative-decoding support for this release.
Source and license
Source revision: 8f8e0e58a3c9df042be1a3fa2c191fd8047acfb8. APEX profile revision: 636cec7e3f8d308e4162ab2bd5ec4b56807dd806. The source BF16 conversion is reused without tensor changes; tokenizer metadata is set during quantization. Full generation commands and upstream matched-tensor allocations are included in validation.
The model weights retain Apache 2.0. Credit for the APEX quantization strategy goes to the LocalAI team, whose toolchain is MIT licensed. Accio-Lab produced and validated these Occamy files.
Citation
The checkpoint accompanies Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work. Please cite the original report when using the model:
@misc{chen2026occamy10openparetofrontier35b,
title={Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work},
author={Wenhui Chen and Shiwen Cheng and Hao Dong and Chenda Duan and Ruixiang Feng and Zhong Guan and Boqiang Guo and Xueyuan Han and Haojie Hao and Liangmeng Huang and Zhelong Huang and Xinke Kong and Hongyu Li and Jiazheng Li and Junbo Li and Qingchuan Li and Yukun Lian and Chang Liu and Tianyu Liu and Zicheng Liu and Shuyi Ouyang and Yijun Pan and Kunyu Shi and Xiaojun Tang and Bingquan Wang and Kesu Wang and Yuchen Wang and Sibo Wei and Sicong Xie and Xiaoying Xing and Yi Xu and Zhijun Xu and Hongwei Xue and Qingcheng Zeng and Di Zhang and Guannan Zhang and Haochen Zhang and Tianlong Zhang and Tianyu Zhao and Tianyu Zhao and Yanjun Zheng and Jialong Zhu and Zijian Zou},
year={2026},
eprint={2609.11977},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.11977},
}
Standard GGUF · BF16 · Collection · Explorer
- Downloads last month
- 1,532
We're not able to determine the quantization variants.