Qwen-AgentWorld-35B-A3B-GPTQ-Pro-Int4

A compact 4-bit GPTQ-Pro quantization of Qwen-AgentWorld-35B-A3B — built for local agent simulation, tool-use reasoning, terminal/world-state prediction, and long-context agent workflows.

This repository contains an INT4 GPTQ-Pro quantized version of Qwen/Qwen-AgentWorld-35B-A3B, preserving the core architecture and behavior of the original language world model while dramatically reducing the weight footprint for practical local serving.

The upstream model is not just another chat model with a tool-use sticker slapped on it. Qwen-AgentWorld is trained as a language world model: given an agent’s action and interaction history, it predicts the next environment observation/state. That makes it especially interesting for agentic workflows where the model needs to reason about what should happen next inside tools, terminals, browsers, operating systems, software projects, and simulated environments.

This quantized build is intended for people who want to actually run the thing, not frame the BF16 checkpoint and admire it from across the room.


What this model is good at

This model is best suited for:

  • Agentic environment simulation
  • Tool-call and MCP-style interaction modeling
  • Terminal state prediction
  • Software-engineering workflow simulation
  • Browser / web / OS interaction reasoning
  • Synthetic trajectory generation for agents
  • Long-context planning and replay
  • Local agent research
  • OpenAI-compatible backend serving

The upstream AgentWorld model covers seven major interaction domains:

Domain Use case
MCP / Tool calling Predicting tool results and interaction flow
Search Simulating search-based information environments
Terminal Predicting shell command outputs and filesystem-like state
SWE Software-engineering task environments
Android Mobile GUI interaction simulation
Web Browser and website interaction simulation
OS Desktop / operating-system style workflows

Model details

Field Value
Base model Qwen/Qwen-AgentWorld-35B-A3B
Quantized model groxaxo/Qwen-AgentWorld-35B-A3B-GPTQ-Pro-Int4
Architecture Qwen3.5 MoE / Language World Model
Total parameters 35B
Active parameters ~3B active per token
Quantization GPTQ-Pro INT4
Format Safetensors
Context length Up to 262,144 tokens, depending on serving stack and VRAM
License Apache-2.0, following upstream

Why this quant exists

The original BF16 checkpoint is powerful, but heavy. This GPTQ-Pro INT4 build is aimed at practical deployment where you want the AgentWorld behavior without needing a datacenter-class inference box.

Use this when you want:

  • Smaller disk and VRAM footprint
  • Faster local experimentation
  • OpenAI-compatible serving
  • Multi-GPU consumer inference
  • Agent simulation without calling external APIs
  • A world-model backend for agent training, evaluation, or synthetic rollout generation

Quantization is not magic. INT4 can affect fidelity, especially for very long-context workloads. For serious benchmark or production use, compare against the BF16 upstream model on your own tasks.


Recommended serving: vLLM

Use a recent vLLM build. Qwen3.5 MoE support is new enough that old wheels may fail or silently behave like garbage. Delightful hobby, this one.

pip install -U vllm
Downloads last month
208
Safetensors
Model size
35B params
Tensor type
BF16
·
I32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for groxaxo/Qwen-AgentWorld-35B-A3B-GPTQ-Pro-Int4

Quantized
(73)
this model