AREX

AREX-2 27B · INT8 W8A16

Symmetric group-128 weight-only quantization for vLLM.

Base model · Paper · Project · vLLM

Format W8A16 Weights INT8 Activations BF16 License Apache 2.0

This is a numerical quantization of BAAI/AREX-2, not a fine-tune. It was exported from official revision d4e3502f92d9e889031c2d04e387ab2eb3520268. Model, architecture, training-data, and upstream evaluation credit belongs to BAAI and the upstream contributors. Results from the upstream BF16 model do not measure this INT8 checkpoint.

Model summary

AREX-2 is a 27B-parameter, Qwen3.8-compatible multimodal model for long-horizon agent tasks. The upstream checkpoint declares a native context length of 262,144 tokens. Refer to the upstream model card for architecture, capabilities, evaluation protocols, and usage guidance.

Dual RTX 3090 deployment

Validation scope: the checkpoint loaded and passed a short text-chat smoke test in the pinned vLLM 0.29.0 deployment image on 2×RTX 3090, and completed the synthetic serving matrix below. No task-quality or multimodal evaluation was run; the native 262,144-token limit was not tested.

Runtime property Tested profile
GPUs 2×RTX 3090 (Ampere sm_86)
Tensor parallelism 2
Activation dtype BF16
KV cache FP8 E4M3
CPU KV offload 28 GiB, using the deployment's vLLM backport
Maximum model length 262,144 configured; tested up to 128,003 prompt tokens + 1,024 generated tokens
Maximum active sequences 1
Optional speculation External DFlash2 W4A16 drafter; not included in this repository

Memory and capacity

Metric Result
Indexed tensor payload 30.77 GB / 28.65 GiB
vLLM GPU memory after short probes 22,688 MiB used and 1,490 MiB free per GPU
Long-context capacity 128K prompt requests benchmarked; the configured 262,144-token limit was not tested

GPU memory is the observed full serving-process allocation, not model weights alone. The longest benchmark request had 128,003 prompt tokens and 1,024 generated tokens; this does not establish capacity at the configured 262,144-token limit.

Measured serving performance

The matrix used synthetic, salted prompts at approximately 1K, 8K, 32K, 64K, and 128K tokens, with exact 512- or 1,024-token greedy completions. There were two cold-prefix runs per cell; TTFT is time to first streamed token, and prompt usage comes from the server. The table reports medians; the decode column also shows the two-run range. DFlash2 remained enabled (7 draft tokens, 15-token verify ceiling, lookup and chains on), with the production KV-cache and CPU-offload profile.

Prompt tokens New tokens Median TTFT (s) Prefill (tok/s) Decode median (tok/s; min–max)
1,027 512 0.74 1389 64.6 (58.4–70.7)
1,028 1,024 0.73 1405 243.1 (236.6–249.7)
8,194 512 6.22 1317 65.1 (62.6–67.6)
8,196 1,024 6.34 1293 202.0 (198.4–205.6)
32,002 512 27.52 1163 141.6 (140.4–142.9)
32,002 1,024 29.98 1071 61.9 (61.7–62.1)
64,002 512 77.72 824 67.4 (63.1–71.7)
64,002 1,024 84.52 758 80.7 (72.4–89.0)
128,002 512 196.28 652 35.6 (35.4–35.8)
128,003 1,024 197.15 649 63.9 (63.2–64.6)

Quantization fidelity

This export uses data-free, symmetric round-to-nearest (RTN) quantization. No calibration dataset, perplexity comparison, logit/KLD evaluation, or task-quality evaluation has been run for this derivative. The source model's scores should not be interpreted as results for this checkpoint.

Checkpoint profile

Property Value
Base checkpoint BAAI/AREX-2, revision d4e3502f92d9e889031c2d04e387ab2eb3520268
Quantization Symmetric INT8 W8A16, group size 128, RTN; no calibration
Quantized modules 400 two-dimensional Linear weights
Runtime format compressed-tensors / pack-quantized
Kernel dispatch CompressedTensorsWNA16 → MarlinLinearKernel in the tested runtime
Scale storage FP16
Preserved source precision Vision tower, token embeddings, output head, recurrent GDN gates, and other excluded/non-Linear tensors
Runtime vLLM with compressed-tensors support; this is not a GGUF checkpoint

The quantization ignore rules are inherited from the validated W8A16 profile. The AREX-2 source checkpoint contains no MTP tensors; any speculative-decoding drafter must be supplied separately.

Quantization design

Component Treatment
400 eligible two-dimensional Linear weights Symmetric INT8, group size 128, round-to-nearest
Vision tower, embeddings, output head, recurrent GDN gates Retained at source precision
Other non-Linear and excluded tensors Retained at source precision

The export uses the same exact eligible-module and ignore profile as the local W8A16 reference. It does not fine-tune or calibrate the model.

Why W8A16

W8A16 stores the selected Linear weights in 8-bit integers with group-wise scales while keeping activations at 16-bit precision. This repository is intended as a smaller serving checkpoint for compatible vLLM deployments. No claim is made that INT8 is faster or higher-quality than BF16, FP8, or other quantization formats for AREX-2; benchmark on the intended hardware and workload.

Operational notes

Question Guidance
Which runtime is validated? The pinned vLLM 0.29.0 deployment image used for the profile above. Other versions and kernels need separate validation.
Can llama.cpp load this checkpoint? No. It uses the vLLM compressed-tensors format, not GGUF.
Is vision validated? The source-precision vision weights are retained, but the smoke test was text-only; multimodal quality and memory use are unmeasured.
Is 262K serving validated? No. The source model's context setting is retained, but long-context capacity and quality were not tested.
Is DFlash2 included? No. The deployment smoke used an external drafter; this model repository contains only AREX-2 checkpoint assets.

Files and provenance

The repository contains the sharded SafeTensors checkpoint, model.safetensors.index.json, model/tokenizer configuration, and the upstream Apache-2.0 LICENSE. Quantized tensors are represented by packed weights, per-group scales, and original weight shapes in the index.

Acknowledgements and license

This repository repackages numerical weights derived from BAAI/AREX-2. It does not claim authorship of the base model, architecture, training data, or upstream evaluations. The upstream Apache-2.0 license is included; retain the source attribution and follow the upstream terms.

Downloads last month
15
Safetensors
Model size
27B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for numsu/AREX-2-27B-INT8-W8A16

Base model

Qwen/Qwen3.8-27B
Finetuned
BAAI/AREX-2
Quantized
(7)
this model

Paper for numsu/AREX-2-27B-INT8-W8A16