GeoRefine Qwen3.8-27B TBE
Developed by Tsotchke Corporation. The GeoRefine codec and independent verifier are Apache-2.0 source releases: corporate GitHub repository, glc-loader 1.1.1, and georefine-verify 1.1.2. The model weights retain Qwen’s upstream license and attribution.
This is a bit-exact storage encoding of Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0. It contains the text model, vision model, and MTP weights. Decoding reproduces every parent BF16 tensor byte for byte. No weights were trained or altered.
Format: georefine.tbe.serve.v1. The files named shards/*.safetensors are codec containers. They are not ordinary Hugging Face model shards; transformers.from_pretrained cannot load this repository directly. A portable CPU exporter can reconstruct a standard BF16 Transformers checkpoint from this bundle without downloading the parent. That compatibility path restores BF16 memory and storage use. Compressed inference requires a compatible GeoRefine TBE backend. The independent CPU verifier checks this bundle against Qwen's published shards; it is not an inference server.
Verified result
| Measurement | Result | Scope |
|---|---|---|
| BF16 tensor bytes | 55,562,855,904 | Parent tensor payload |
| GeoRefine encoded tensor bytes | 39,832,462,895 | 1.395× smaller tensor payload |
| Entire bundle on disk | 39,856,482,250 bytes | 20 codec shards plus manifest and sidecars |
| Independently matched tensors | 1,199/1,199 | 492 codec tensors and 707 raw tensors; exact bytes |
| Matched sidecars | 10/10 | Includes tokenizer, config, vision preprocessors, and parent LICENSE |
| Compressed Transformers load | 1,184 base tensors; 484 coded Linear weights | CPU reference on dadbox; encoded weight arrays remain resident |
| Full-model CPU forwards | Text and minimal synthetic image inputs produced finite logits | CPU reference on dadbox; not a generation or speed benchmark |
| Same-engine autoregressive decode | 37.84 vs 28.29 tokens/s | Codec vs BF16, batch 1, one RTX PRO 6000 Blackwell 96 GB |
| Same-engine speculative decode, k=5 | 108.6 vs 92.65 tokens/s | Codec vs BF16, GSM8K test rows 0–49, 128 generated tokens |
| Peak GPU memory in measured run | 49.9–50.0 vs 65.4–72.8 GB | Codec vs BF16, NVML high-water |
The independent verifier checked all 18 parent shards against Hugging Face's published LFS SHA-256 values, all 20 bundle shards against the manifest, and every decoded tensor against its parent. On the measured engine, codec and BF16 logits matched bit for bit for 711/711 tested rows. This does not mean different inference engines produce identical logits: operation order and BF16 rounding can differ. Speculative and plain greedy decoding matched on 442/442 codec runs. These are measured scopes, not guarantees for every future runtime.
The speed figures are from our engine on one RTX PRO 6000 Blackwell, batch 1. They are workload and hardware specific. Batch sizes above 1, 262k context, a 24 GB GPU fit, and cost per token have not been established for this bundle. This is lossless weight encoding, not a 4-bit model.
Runtime status (2026-09-28). The standard glc-serve HTTP generation route is not the measured fast engine. In a same-RTX eight-prompt check it reached 14.95 tok/s end to end in fused mode versus 17.50 tok/s for BF16 and matched 5/8 response texts. Its exact mode matched 8/8 texts and one synthetic image response but reached 7.99 tok/s. These are results for that route, not a limit of the codec.
The source-free FastSession/FastDecoder wheel with the measured RTX PRO 6000 SM120 tune bundled inside it did reproduce the faster engine result: 37.880 decode tok/s versus 28.288 for BF16 (1.339×), and 32.314 versus 25.762 tok/s end to end (1.254×). Generated token IDs and text matched 8/8 fixed prompts; one synthetic image probe matched too. A separate 128-token text probe found plain and spec_k=5 token sequences identical at 37.743 versus 68.820 tok/s, with image-token parity as well. The Apache source and wheel, complete runtime receipts, and public source repository accompany this model. The 1.1.0 wheel's kernels, tune, and FastSession code are byte-identical to the GPU-tested candidate; only the package version constant and metadata changed. Eight prompts and one image are a runtime proof, not a general capability certificate, and fast results on other hardware remain unverified.
Transformers compatibility
The Apache-2.0 reference packages include the compressed PyTorch loader, independent verifier, CPU exporter, source archives, wheels, and checksums. The fused RTX PRO benchmark used a separate engine that is not in these reference wheels.
The georefine-export CPU tool reconstructs 20 ordinary safetensors shards and model.safetensors.index.json, then copies Qwen's original model and processor files. The output loads through AutoModelForMultimodalLM.from_pretrained on a system with enough memory to run Qwen3.8-27B. Plan for about 56 GB of output storage and one codec shard of temporary scratch. A small synthetic fixture passes the exporter and a real safetensors reader; a full 27B export and Transformers load are not yet verified. The public model will need that test before broad compatibility is claimed.
The glc_loader.load_compressed_transformers CPU reference keeps TBE arrays encoded as resident safetensors data and replaces the base Qwen model's 484 coded Linear weights with compressed modules. It decodes one weight only while that layer runs. On dadbox, the full 1,184-tensor Qwen graph loaded from this bundle and produced finite text logits [1,6,248320] and minimal-image logits [1,2,248320]. The text run took 2,415 seconds in total, including a 261-second load, and peaked at about 37.5 GiB process RSS; the separate image run took 1,585 seconds in total and peaked at about 38.8 GiB. This CPU path establishes full-size compatibility, not practical serving speed. Same-system BF16 logits parity, generation, native-kernel dispatch outside our measured RTX PRO engine, and platform-specific performance remain unverified. The 15 MTP weights remain in the bundle for native speculative serving.
Independent verification
The CPU-only georefine-verify package takes this bundle and the exact Qwen revision above. A full run needs downloads of roughly 40 GB of this bundle and 56 GB of parent weights; --evict-reference limits parent scratch use. The verified command was:
georefine-verify \
--bundle . \
--reference Qwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 \
--expect-manifest-sha256 10028416e07c802b47e4dc93a86e9f0dd35ce6021f0e07b96c70bd9cb7a72bc3 \
--work-dir ./georefine-verify-work --evict-reference
The public manifest removes three local build paths while retaining every tensor, shard, sidecar and accounting record. Its SHA-256 is 10028416e07c802b47e4dc93a86e9f0dd35ce6021f0e07b96c70bd9cb7a72bc3. This sanitized bundle passed a fresh independent check on a Linux CPU host: 1,199/1,199 tensors bit-identical, 10/10 sidecars identical, every parent LFS and bundle manifest shard hash checked, with no caveats. The complete verification receipt has SHA-256 1f231be78c7a3231e8a7960e55c655db8ebba944086b4be3129d9eb517d8625f. Its original receipt SHA-256 is 571c5e99311ad740e73b49cf15b7d1e42614050df2b931250d5c66fcc3630017; only machine-local paths and the host label were changed for publication. The Apache-2.0 verifier source and wheel are in software/reference.
Provenance and license
Parent: Qwen/Qwen3.8-27B, revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0, Apache-2.0. GeoRefine transcodes its BF16 tensors into reversible TBE containers, leaving the decoded tensors and 10 accompanying files unchanged. The parent LICENSE is included in the bundle. The GeoRefine codec loader and independent verifier are licensed Apache-2.0; their source and wheels are separate from these model weights.
- Downloads last month
- 49
Model tree for Tsotchke-Corporation/Qwen3.8-27B-GeoRefine-TBE
Base model
Qwen/Qwen3.8-27B