Download README.md from numsu/AREX-2-27B-INT8-W8A16: direct link, hf CLI and curl.
- Browser
- Download file 7.85 kB
-
https://huggingface.co/numsu/AREX-2-27B-INT8-W8A16/resolve/main/README.md
- Command line
-
hf download hf://numsu/AREX-2-27B-INT8-W8A16/README.md
-
curl -L -o README.md https://huggingface.co/numsu/AREX-2-27B-INT8-W8A16/resolve/main/README.md
license: apache-2.0
base_model: BAAI/AREX-2
base_model_relation: quantized
library_name: vllm
pipeline_tag: image-text-to-text
tags:
- qwen3_5
- arex-2
- compressed-tensors
- w8a16
- int8
- quantized
- vllm
- agent
- reasoning
- tool-use
- vision
- long-context
AREX-2 27B · INT8 W8A16
Symmetric group-128 weight-only quantization for vLLM.
Base model · Paper · Project · vLLM
This is a numerical quantization of BAAI/AREX-2, not a fine-tune. It was exported from official revision
d4e3502f92d9e889031c2d04e387ab2eb3520268. Model, architecture, training-data, and upstream evaluation credit belongs to BAAI and the upstream contributors. Results from the upstream BF16 model do not measure this INT8 checkpoint.
Model summary
AREX-2 is a 27B-parameter, Qwen3.8-compatible multimodal model for long-horizon agent tasks. The upstream checkpoint declares a native context length of 262,144 tokens. Refer to the upstream model card for architecture, capabilities, evaluation protocols, and usage guidance.
Dual RTX 3090 deployment
Validation scope: the checkpoint loaded and passed a short text-chat smoke test in the pinned vLLM 0.29.0 deployment image on 2×RTX 3090, and completed the synthetic serving matrix below. No task-quality or multimodal evaluation was run; the native 262,144-token limit was not tested.
| Runtime property | Tested profile |
|---|---|
| GPUs | 2×RTX 3090 (Ampere sm_86) |
| Tensor parallelism | 2 |
| Activation dtype | BF16 |
| KV cache | FP8 E4M3 |
| CPU KV offload | 28 GiB, using the deployment's vLLM backport |
| Maximum model length | 262,144 configured; tested up to 128,003 prompt tokens + 1,024 generated tokens |
| Maximum active sequences | 1 |
| Optional speculation | External DFlash2 W4A16 drafter; not included in this repository |
Memory and capacity
| Metric | Result |
|---|---|
| Indexed tensor payload | 30.77 GB / 28.65 GiB |
| vLLM GPU memory after short probes | 22,688 MiB used and 1,490 MiB free per GPU |
| Long-context capacity | 128K prompt requests benchmarked; the configured 262,144-token limit was not tested |
GPU memory is the observed full serving-process allocation, not model weights alone. The longest benchmark request had 128,003 prompt tokens and 1,024 generated tokens; this does not establish capacity at the configured 262,144-token limit.
Measured serving performance
The matrix used synthetic, salted prompts at approximately 1K, 8K, 32K, 64K, and 128K tokens, with exact 512- or 1,024-token greedy completions. There were two cold-prefix runs per cell; TTFT is time to first streamed token, and prompt usage comes from the server. The table reports medians; the decode column also shows the two-run range. DFlash2 remained enabled (7 draft tokens, 15-token verify ceiling, lookup and chains on), with the production KV-cache and CPU-offload profile.
| Prompt tokens | New tokens | Median TTFT (s) | Prefill (tok/s) | Decode median (tok/s; min–max) |
|---|---|---|---|---|
| 1,027 | 512 | 0.74 | 1389 | 64.6 (58.4–70.7) |
| 1,028 | 1,024 | 0.73 | 1405 | 243.1 (236.6–249.7) |
| 8,194 | 512 | 6.22 | 1317 | 65.1 (62.6–67.6) |
| 8,196 | 1,024 | 6.34 | 1293 | 202.0 (198.4–205.6) |
| 32,002 | 512 | 27.52 | 1163 | 141.6 (140.4–142.9) |
| 32,002 | 1,024 | 29.98 | 1071 | 61.9 (61.7–62.1) |
| 64,002 | 512 | 77.72 | 824 | 67.4 (63.1–71.7) |
| 64,002 | 1,024 | 84.52 | 758 | 80.7 (72.4–89.0) |
| 128,002 | 512 | 196.28 | 652 | 35.6 (35.4–35.8) |
| 128,003 | 1,024 | 197.15 | 649 | 63.9 (63.2–64.6) |
Quantization fidelity
This export uses data-free, symmetric round-to-nearest (RTN) quantization. No calibration dataset, perplexity comparison, logit/KLD evaluation, or task-quality evaluation has been run for this derivative. The source model's scores should not be interpreted as results for this checkpoint.
Checkpoint profile
| Property | Value |
|---|---|
| Base checkpoint | BAAI/AREX-2, revision d4e3502f92d9e889031c2d04e387ab2eb3520268 |
| Quantization | Symmetric INT8 W8A16, group size 128, RTN; no calibration |
| Quantized modules | 400 two-dimensional Linear weights |
| Runtime format | compressed-tensors / pack-quantized |
| Kernel dispatch | CompressedTensorsWNA16 → MarlinLinearKernel in the tested runtime |
| Scale storage | FP16 |
| Preserved source precision | Vision tower, token embeddings, output head, recurrent GDN gates, and other excluded/non-Linear tensors |
| Runtime | vLLM with compressed-tensors support; this is not a GGUF checkpoint |
The quantization ignore rules are inherited from the validated W8A16 profile. The AREX-2 source checkpoint contains no MTP tensors; any speculative-decoding drafter must be supplied separately.
Quantization design
| Component | Treatment |
|---|---|
| 400 eligible two-dimensional Linear weights | Symmetric INT8, group size 128, round-to-nearest |
| Vision tower, embeddings, output head, recurrent GDN gates | Retained at source precision |
| Other non-Linear and excluded tensors | Retained at source precision |
The export uses the same exact eligible-module and ignore profile as the local W8A16 reference. It does not fine-tune or calibrate the model.
Why W8A16
W8A16 stores the selected Linear weights in 8-bit integers with group-wise scales while keeping activations at 16-bit precision. This repository is intended as a smaller serving checkpoint for compatible vLLM deployments. No claim is made that INT8 is faster or higher-quality than BF16, FP8, or other quantization formats for AREX-2; benchmark on the intended hardware and workload.
Operational notes
| Question | Guidance |
|---|---|
| Which runtime is validated? | The pinned vLLM 0.29.0 deployment image used for the profile above. Other versions and kernels need separate validation. |
| Can llama.cpp load this checkpoint? | No. It uses the vLLM compressed-tensors format, not GGUF. |
| Is vision validated? | The source-precision vision weights are retained, but the smoke test was text-only; multimodal quality and memory use are unmeasured. |
| Is 262K serving validated? | No. The source model's context setting is retained, but long-context capacity and quality were not tested. |
| Is DFlash2 included? | No. The deployment smoke used an external drafter; this model repository contains only AREX-2 checkpoint assets. |
Files and provenance
The repository contains the sharded SafeTensors checkpoint, model.safetensors.index.json, model/tokenizer configuration, and the upstream Apache-2.0 LICENSE. Quantized tensors are represented by packed weights, per-group scales, and original weight shapes in the index.
Acknowledgements and license
This repository repackages numerical weights derived from BAAI/AREX-2. It does not claim authorship of the base model, architecture, training data, or upstream evaluations. The upstream Apache-2.0 license is included; retain the source attribution and follow the upstream terms.