Accio

Occamy-1.0 · APEX GGUF

Three plain profiles · four calibrated profiles, including Mini

Collection · Checkpoint explorer · Project · Paper

Occamy-1.0 quantized with the LocalAI team's APEX profiles. APEX allocates GGUF precision by tensor role and layer position: routed experts, shared experts and attention receive different allocations. The three plain profiles and four calibrated I- profiles use the pinned upstream 40-layer configurations. Every file is generated directly from the same BF16 source. The calibrated group adds I-Compact, I-Quality, I-Balanced and I-Mini.

Downloads

Profile Exact bytes GB GiB Download
Compact 16,538,851,072 16.54 15.403 GGUF
Quality 22,819,401,472 22.82 21.252 GGUF
Balanced 25,335,983,872 25.34 23.596 GGUF

Calibrated profiles

Profile Exact bytes GB GiB Download
I-Compact 16,538,851,360 16.54 15.403 GGUF
I-Quality 22,819,401,760 22.82 21.252 GGUF
I-Balanced 25,335,984,160 25.34 23.596 GGUF
I-Mini 13,467,210,784 13.47 12.542 GGUF

Profile names describe allocation strategies; file size does not follow the name order. These are weight-file sizes, not peak memory requirements. Allow additional memory for the runtime, context and image input.

Precision allocation

Profile Routed experts: layers 0–4 / 5–9 / 10–29 / 30–34 / 35–39 Shared experts Attention / SSM projections
Compact Q4_K / Q3_K / Q3_K / Q3_K / Q4_K Q6_K Q4_K
Quality Q6_K / Q5_K / IQ4_XS / Q5_K / Q6_K Q8_0 Q6_K
Balanced Q6_K / Q5_K / Q5_K / Q5_K / Q6_K Q8_0 Q6_K
I-Mini Q3_K / Q3_K / IQ2_S / Q3_K / Q3_K Q5_K in first/last 5 layers; Q4_K elsewhere Q4_K in first/last 3 layers; Q3_K elsewhere

I-Compact, I-Quality and I-Balanced use the same allocations as their plain counterparts, with the validated importance matrix. I-Mini uses the upstream mini profile, which requires calibration.

Embeddings, output and tensors outside the profile overrides follow the pinned stock quantizer's Q4_K_M base recipe for Compact and Q6_K for Quality/Balanced and their I variants; I-Mini uses Q3_K_M. I-Compact uses Q4_K_M. Small tensors that the quantizer excludes retain their source type. All actual tensor types are recorded in artifact checks; this is mixed precision, not a uniform bit width.

Run with llama.cpp

Use a recent llama.cpp build with Qwen3.5 MoE support. The validated revision is 972d2313bc0bf0a45f634f77d95c9fb03aeab12c.

hf download Accio-Lab/occamy-1.0-APEX-GGUF occamy-1.0-APEX-Compact.gguf \
  --local-dir ./occamy-gguf

llama-server -m ./occamy-gguf/occamy-1.0-APEX-Compact.gguf \
  -ngl 999 -c 8192 -np 1 -fa on --jinja \
  --host 127.0.0.1 --port 8000 --alias occamy

Replace Compact with Quality, Balanced, I-Compact, I-Quality, I-Balanced or I-Mini to use another profile. Input content should be NFC-normalized; all files store the source-compatible qwen2 pre-tokenizer metadata. See the documented tokenizer protocol.

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"occamy","messages":[{"role":"user","content":"Compute 2+2. Answer briefly."}],"temperature":0,"max_tokens":128,"chat_template_kwargs":{"enable_thinking":false}}'

For vision, download the matching existing F16 projector and add --mmproj ./occamy-gguf/mmproj-occamy-1.0-F16.gguf. It is a separate 899,282,944-byte file in the standard GGUF repository. Image inference was not retested for this APEX release.

hf download Accio-Lab/occamy-1.0-GGUF mmproj-occamy-1.0-F16.gguf \
  --local-dir ./occamy-gguf

Measured validation

Model Authored functional cases WikiText subset PPL
BF16 19/20 8.2963 +/- 0.36372
APEX-Compact 20/20 8.4567 +/- 0.36569
APEX-Quality 19/20 8.2568 +/- 0.36068
APEX-Balanced 19/20 8.2994 +/- 0.36385

Calibrated profiles, measured with the same protocol:

Model Authored functional cases WikiText subset PPL
APEX-I-Compact 19/20 8.5074 +/- 0.37314
APEX-I-Quality 19/20 8.3743 +/- 0.36829
APEX-I-Balanced 19/20 8.3482 +/- 0.36702
APEX-I-Mini 20/20 9.5643 +/- 0.43416

Calibration did not improve the paired WikiText subset PPL for the three matched profiles in this run. I-Compact, I-Quality and I-Balanced fail the same Chinese JSON case (json_4) as BF16; no forced JSON grammar was used. I-Mini is an experimental size/quality tradeoff: its subset PPL is higher, while all 20 authored fixtures pass. All stored code outputs for the eight models were rescored after adding the ordinary isinstance builtin to the restricted Python harness; generation and assertions were unchanged. Original execution results and the rescoring receipt are preserved.

All seven files pass complete finite-value checks for all 733 tensors, logical shape and vocabulary/template preservation, and 16/16 NFC tokenizer probes with special-token parsing. The functional checks include English/Chinese instructions, arithmetic, executed Python, strict JSON, conversation memory and native tool round trips. Plain-profile evidence remains in validation. Calibrated per-case outputs, failures, exact file hashes and reproduction are in calibrated validation.

The BF16 reference is reused from the preceding run with the same pinned source, runtime and inputs. These are 20 authored fixtures and a small WikiText subset, not full quality benchmarks. Upstream APEX results on Qwen are not Occamy measurements. CPU-only and Apple Silicon execution, image inference and long-context quality were not tested in this run.

Calibration and MTP

The four I- profiles use a 192,232,448-byte importance matrix collected from BF16. The fixed calibration mix has 608 training conversations: 192 UltraChat, 96 CodeFeedback, 128 Glaive function-calling, 96 GSM8K and 96 authored multilingual/state-tracking examples. Sources, revisions, licenses, selected rows and hashes are recorded in the manifest. Source inputs were NFC-normalized and rendered with the pinned chat template, with thinking disabled and a maximum of 1,536 tokens per conversation. No held-out WikiText or functional-fixture prompt was used for calibration.

Four disjoint shards were processed with stock llama-imatrix, context 512, special-token parsing and output-projection collection. The combined matrix covers 712 chunks / 364,544 processed tokens. All 120 routed projection matrices cover all 256 experts: 30,720/30,720 expert-projection pairs observed, with a minimum activation count of 46. Stored sums/counts are finite and nonnegative. This is observed coverage, not equal per-expert sampling.

The matrix is GGUF importance-matrix data saved with a .bin suffix, not a model checkpoint. Its SHA256 is 81a0e2944b9c2290f924926a96d98b9ddc636ad1fb8accdfd596e99e6f29f160. Raw training conversations are not redistributed here; the pinned preparation recipe, manifest and coverage audit are included. See reproduction.

Calibration does not guarantee a quality improvement. Compare the measured subset results below; upstream APEX benchmark results on other checkpoints are not Occamy measurements.

The three plain profiles do not use an importance matrix. The four I- profiles are calibrated. There is no bundled MTP head in these files. The Occamy MTP checkpoint remains separate and does not establish llama.cpp speculative-decoding support for this release.

Source and license

Source revision: 8f8e0e58a3c9df042be1a3fa2c191fd8047acfb8. APEX profile revision: 636cec7e3f8d308e4162ab2bd5ec4b56807dd806. The source BF16 conversion is reused without tensor changes; tokenizer metadata is set during quantization. Full generation commands and upstream matched-tensor allocations are included in validation.

The model weights retain Apache 2.0. Credit for the APEX quantization strategy goes to the LocalAI team, whose toolchain is MIT licensed. Accio-Lab produced and validated these Occamy files.

Citation

The checkpoint accompanies Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work. Please cite the original report when using the model:

@misc{chen2026occamy10openparetofrontier35b,
      title={Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work}, 
      author={Wenhui Chen and Shiwen Cheng and Hao Dong and Chenda Duan and Ruixiang Feng and Zhong Guan and Boqiang Guo and Xueyuan Han and Haojie Hao and Liangmeng Huang and Zhelong Huang and Xinke Kong and Hongyu Li and Jiazheng Li and Junbo Li and Qingchuan Li and Yukun Lian and Chang Liu and Tianyu Liu and Zicheng Liu and Shuyi Ouyang and Yijun Pan and Kunyu Shi and Xiaojun Tang and Bingquan Wang and Kesu Wang and Yuchen Wang and Sibo Wei and Sicong Xie and Xiaoying Xing and Yi Xu and Zhijun Xu and Hongwei Xue and Qingcheng Zeng and Di Zhang and Guannan Zhang and Haochen Zhang and Tianlong Zhang and Tianyu Zhao and Tianyu Zhao and Yanjun Zheng and Jialong Zhu and Zijian Zou},
      year={2026},
      eprint={2609.11977},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.11977}, 
}

Standard GGUF · BF16 · Collection · Explorer

Downloads last month
1,532
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Accio-Lab/occamy-1.0-APEX-GGUF

Quantized
(31)
this model

Space using Accio-Lab/occamy-1.0-APEX-GGUF 1

Collection including Accio-Lab/occamy-1.0-APEX-GGUF

Paper for Accio-Lab/occamy-1.0-APEX-GGUF