Holotron4-30B-A3B GGUF

GGUF quantizations of Hcompany/Holotron4-30B-A3B, a 30B hybrid NemotronH MoE vision-language model (VLM) for Computer Use, tool-driven work, and agentic workflows.

For vision and audio input, use the included F16 multimodal projector (mmproj-Holotron4-30B-f16.gguf).

Upstream benchmarks

Holotron4 improves over its base model, Nemotron 3 Nano Omni, on GUI workflows and in environments with MCP tools, APIs, or code sandboxes. Gains are absolute percentage points.

Benchmark Interface Nemotron 3 Nano Omni Holotron4-30B-A3B Gain
OSWorld GUI 21.0 76.3 +55.3
OSWorld 2.0 GUI and code 0.2 7.9 +7.7
AutomationBench MCP 19.4 35.6 +16.2
PinchBench Terminal 84.7 88.6 +3.9
ALE (Linux, code) Terminal 0.6 8.5 +7.9

GGUF files

Due to the model's MoE intermediate dimensions (1856), standard K-quants (Q4_K, Q6_K) trigger internal fallbacks in llama.cpp to Q5/Q8 types. Therefore, clean 32-block quantizations (Q4_0 and Q8_0) are provided for optimal efficiency and accuracy.

Quantization File Size Notes
Q4_0 Holotron4-30B-Q4_0.gguf 18.0 GB Recommended balanced default (clean 4-bit)
Q8_0 Holotron4-30B-Q8_0.gguf 33.6 GB Near-lossless high-precision reference
Multimodal projector mmproj-Holotron4-30B-f16.gguf 2.99 GB Required for image and audio input

Usage

Use a current llama.cpp build with the included chat template.

llama-cli \
  -m Holotron4-30B-Q4_0.gguf \
  -c 4096 -n 512 --temp 0.6 --top-p 0.95 \
  --jinja --chat-template-file chat_template.jinja \
  -p "<|im_start|>user\nWhat is computer use in AI?<|im_end|>\n<|im_start|>assistant\n"

For multimodal / vision and audio input:

llama-mtmd-cli \
  -m Holotron4-30B-Q4_0.gguf \
  -mm mmproj-Holotron4-30B-f16.gguf \
  --image screenshot.png \
  -p "Describe the interface in this screenshot."

Source

Downloads last month
250
GGUF
Model size
32B params
Architecture
nemotron_h_moe
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for abenzerps/Holotron4-30B-A3B-GGUF

Quantized
(5)
this model

Collection including abenzerps/Holotron4-30B-A3B-GGUF