Blackfrost-AI

CYBER-FROST-3.8 — EXL3 (SAGE) 3.87 bpw

Official collaboration with Blackfrost-AI (Sir Frosty) on the EXL3 conversion of CYBER-FROST-3.8 (Blackfrost-AI/CYBER-FROST-3.8-FP8, Blackfrost-AI/CYBER-FROST-3.8-BF16).

Top-1 and KL versus the BF16 model, 10,240 positions per set, unrounded next-token:

  • Held-out top-1: 9,614 / 10,240 (93.89%). KL 0.0264 nats (0.0380 bits).
  • Independent top-1: 9,589 / 10,240 (93.64%). KL 0.0423 nats (0.0610 bits).

This pack is the SAGE EXL3 deployment artifact of CYBER-FROST-3.8, for security professionals conducting authorized research, assessment, engineering, and response work. It replaces the earlier 3.5 bpw pack, which scored 86.96% held-out and 85.16% independent top-1 against the FP8 release.

Serving: DEPLOY.md. An agent should follow that file. It builds vcruz305/exllamav3 through the Qwen3.8-Flash-Next recipe at the measured fork pin, and it does not use the recipe's default pack.

Release status and contents

Cyber-Frost is a public research release under active quality assessment. This repository is the official collaboration EXL3 of the Cyber-Frost release. It is a standalone checkpoint, not an adapter, and it does not require the FP8 or BF16 parent at load time.

The release contains the EXL3 SafeTensors shards, weight index, n-gram table, configuration, tokenizer and processor assets, the packaged chat template, the upstream license, and this card. Parent scores are not copied here. The measurements below were taken on this pack.

Field This artifact
Model CYBER-FROST-3.8-EXL3-SAGE-3.87bpw
Publisher vcruz305, official collaboration with Blackfrost-AI (Sir Frosty)
Conversion source Blackfrost-AI/CYBER-FROST-3.8-BF16
FP8 release Blackfrost-AI/CYBER-FROST-3.8-FP8
Former parent name BLACKFROST-3.8-ICED-BF16 / BLACKFROST-3.8-DERISKED-FP8
Architecture Qwen4ExpForConditionalGeneration
Method SAGE EXL3
Body bitrate 3.87 bits/param
Held-out top-1 9,614 / 10,240 (93.89%)
Held-out KL 0.0264 nats (0.0380 bits)
Independent top-1 9,589 / 10,240 (93.64%)
Independent KL 0.0423 nats (0.0610 bits)
LM head 8 bits
MTP layer 4 bits
Vision tower 6 bits
N-gram table 6 bits, included
Weight layout 9 SafeTensors shards plus ngram_embedding.safetensors
Published file payload 104,872,927,412 bytes (97.67 GiB), before this card and banner
Configured context 262,144 tokens
Native speculative head one MTP layer
Primary task text generation

The configuration includes a vision tower and processor files. This EXL3 release has not received a multimodal quality evaluation. Do not infer validated image or video capability from their presence.

One DGX Spark (GB10, 128 GB unified) loads and serves this pack with the settings in DEPLOY.md: a 262,144-token window, 8-bit KV cache, the n-gram table streamed from disk, one request at a time. Other hardware and settings were not tested. Do not infer a fit from the bitrate alone. Packing, the head, embeddings, the n-gram table, MTP, vision, and runtime memory are separate from the 3.87 bpw body figure.

Why Cyber-Frost exists

Security work is unusually vulnerable to false refusals. The same vocabulary appears in incident response, exploit validation, malware analysis, defensive engineering, and unauthorized activity. A general-purpose assistant can react to individual terms instead of the operator's legitimate scope.

Cyber-Frost is designed to reduce that unnecessary friction in professional, authorized workflows. It is intended to stay technically direct when an analyst is reviewing a finding, reproducing a vulnerability in a controlled environment, writing detection content, analyzing malicious code, or operating an approved security agent. That is a design objective inherited from the BF16 parent, not a measured behavioral claim for this EXL3 pack.

Authorization is an external control. The model cannot establish ownership, consent, rules of engagement, jurisdiction, or whether a target is in scope. Deployers must enforce identity, scope, tool permissions, logging, rate limits, and human review outside the model.

Security corpus

The BF16 parent was fine-tuned on a Blackfrost-AI security corpus combining curated security material, operator-authored workflows, realistic engagement-style scenarios, and Blackfrost-owned distillation data. Corpus sizes, source-by-source counts, raw engagement material, client identities, prompts, and responses are intentionally not published. This EXL3 conversion added no training data.

Domain coverage of the parent includes:

  • Reconnaissance and OSINT
  • Social engineering, business-email compromise, and deepfake-enabled abuse
  • Web application and API security
  • Identity, authentication, and Active Directory security
  • Network, perimeter, VPN, and protocol security
  • Vulnerability research, bug bounty, and binary exploitation
  • Malware analysis, ransomware, and endpoint defense
  • Cloud, container, and Kubernetes security
  • Software supply-chain security
  • Mobile, IoT, wireless, and physical security
  • Industrial-control-system and operational-technology security
  • Cryptography and security protocols
  • Privilege escalation, lateral movement, and data exfiltration
  • Threat intelligence, APT analysis, and purple-team operations
  • AI-agent, LLM, and adversarial-ML security

Blackfrost-AI attests that owned portions of the corpus were developed from sanitized experience with authorized security work. It also applies a frontier-scale policy to its distillation teachers, excluding teachers below the 753B-parameter class. The parent release evidence independently binds one security subset to a Qwen3.8 2.4T teacher. It does not include a corpus-wide teacher manifest. Those are operator provenance statements, not independent benchmark findings.

Training-data provenance and licensing review for the mixed-source corpus remains in progress on the parent release. This section describes the parent inherited by the quantized artifact.

Model specifications

The text stack has 48 blocks with hybrid linear and full attention, using a full-attention block every fourth layer. Hidden size is 2,560, with 24 attention heads and 2 KV heads. The MoE stack contains 512 routed experts, selects 10 experts per token, and includes a shared expert. Vocabulary size is 248,320. One native MTP layer is packaged for speculative decoding. The parent reports about 180B parameters in Hub metadata.

This EXL3 pack was quantized directly from the BF16 release, not from the FP8 release. It is not a uniform-bitrate pass. SAGE is the mixed-precision method used for the EXL3 allocation. The body bitrate is 3.87 bits/param. The Hub quantization_config.bits field is the integer 4 because the Hub parser requires an integer. The true body bitrate is 3.87.

The 262,144-token configuration ceiling is not a quality guarantee. Long-context, high-concurrency, multimodal, tool-use, and speculative-decoding qualification for this EXL3 pack is not claimed here.

Lineage

  1. Foundational checkpoint: Qwen/Qwen3.8-Flash-Next.
  2. Blackfrost security adaptation: security-domain fine-tuning followed by a full BF16 merge.
  3. Behavioral stage: a Blackfrost-AI behaviorally modified derivative targeting lower false-refusal friction in authorized security workflows. The proprietary transformation process is not distributed.
  4. BF16 release: Blackfrost-AI/CYBER-FROST-3.8-BF16. Formerly BLACKFROST-3.8-ICED-BF16. The rename is not another training run.
  5. FP8 release: Blackfrost-AI/CYBER-FROST-3.8-FP8. Formerly BLACKFROST-3.8-DERISKED-FP8. No additional training during that conversion.
  6. This pack: SAGE EXL3 at 3.87 bpw body, encoded from the BF16 release, official collaboration between vcruz305 and Blackfrost-AI (Sir Frosty). No additional training during this conversion.

Tokenizer and processor lineage comes through the pinned Qwen foundation and the Cyber-Frost releases. The packaged Blackfrost chat template is the parent's.

Measured agreement

Compared with the BF16 model on 10,240 positions per set: 10 sequences of 1,024 tokens, unrounded next-token top-1. Two sets, both scored the same way. KL is KL(reference || this pack) in nats. Bit figures are those nats divided by ln(2). Next-token NLL is the change versus the reference; a negative value means this pack assigned slightly higher probability to the real next token.

Set Top-1 Top-5 KL Next-token NLL change
Held-out 9,614 / 10,240 (93.89%) 99.85% 0.0264 nats (0.0380 bits) -0.0002 nats (-0.0003 bits)
Independent 9,589 / 10,240 (93.64%) 99.84% 0.0423 nats (0.0610 bits) +0.0060 nats (+0.0086 bits)

Against the FP8 release on the same positions, this pack scores 9,628 / 10,240 (94.02%, KL 0.0280 nats) held-out and 9,582 / 10,240 (93.57%, KL 0.0438 nats) independent.

These are measurements of this pack. They are not a claim that the pack matches the BF16 or FP8 release, and they are not scores from those releases. Evaluation text was not used for calibration.

Prompt, tool use, and sampling

The release includes the same chat_template.jinja and embedded tokenizer template as the published Cyber-Frost BF16 and FP8 variants.

The template supplies the Cyber-Frost operating prompt, appends caller-provided system context, supports image and video placeholders, exposes Qwen-style reasoning controls, and serializes XML-style tool calls and tool responses. Thinking is enabled by default. Supported reasoning-effort values are xhigh, medium, and low. Generation defaults on the parent are temperature 1.0, top-p 0.95, and top-k 20.

Use those sampling defaults. The recipe's TabbyAPI config applies them when a request leaves them out (see DEPLOY.md). Without it, TabbyAPI samples with no top-k or top-p truncation, and clients that send no sampling values get replies that turn to gibberish partway through. In a small long-generation check (6 prompts, greedy and default sampling, up to 6,144 new tokens each), greedy decoding fell into repetition on 2 of the 6 prompts and default sampling on none. The FP8 release showed the same greedy repetition count in the same runtime.

The template's authorization assumption is not an access-control mechanism. An agent runtime must independently restrict credentials, targets, files, networks, commands, and approval-requiring actions. Tool-call output is proposed text until an external executor acts on it.

Changing the template, caller system message, reasoning mode, sampling, quantization backend, or runtime can materially change behavior. Record those settings when reporting results.

Deployment

Load this pack with the vcruz305/exllamav3 fork, not a stock wheel. The serving kit is DEPLOY.md. It points at the Qwen3.8-Flash-Next EXL3 recipe and the fork commit the mixed-K decode kernels were measured on. The recipe's published quick start downloads a different pack and an older pin. Follow DEPLOY.md, not that quick start, if you are serving this repository. Record the runtime commit, the two kernel env vars, hardware, context, n-gram and MTP settings, chat template, and sampling parameters when reporting results.

The clean model name is CYBER-FROST-3.8-EXL3-SAGE-3.87bpw.

hf download vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw \
  --local-dir CYBER-FROST-3.8-EXL3-SAGE-3.87bpw

Intended use

Cyber-Frost is intended for qualified security professionals working within explicit authorization, including defensive research, secure code review, vulnerability validation, red-team and purple-team exercises, detection engineering, incident response, malware analysis, bug hunting, and controlled security-agent workflows.

It is not intended to authorize access, choose targets, define rules of engagement, make autonomous high-impact decisions, or replace legal, compliance, and safety review. Do not use it to access systems or data without permission, evade oversight, persist in third-party environments, deploy malware, steal credentials or data, disrupt services, or cause physical harm.

Limitations

  • Generated findings, code, commands, indicators, and remediation advice may be wrong, incomplete, outdated, or fabricated. Independently review them and execute only in isolated, authorized environments.
  • Reduced over-refusal is a design objective inherited from the BF16 parent, not a measured result for this EXL3 pack.
  • BF16, FP8, and NVFP4 behavior, safety observations, benchmark results, and runtime characteristics do not transfer through this conversion.
  • Greedy decoding can fall into repetition on long generations; use the sampling defaults above.
  • The model is not a policy engine, authorization service, sandbox, malware scanner, or secrets boundary.
  • This card does not claim a refusal rate, a standardized cyber-competence score, production readiness, long-context quality, multimodal quality, or reliable MTP acceleration.
  • Model behavior can shift with prompts, sampling, runtime versions, quantization kernels, speculative settings, and agent scaffolding.

The deployer is responsible for authorization, least privilege, isolation, network policy, credential handling, human approval gates, monitoring, incident response, and compliance with applicable law.

License

Use and redistribution of this checkpoint are governed by the Qwen Community License 1.0. Review the license before use. This research release is provided without a warranty of correctness, fitness, security, or non-infringement.

Report reproducible packaging issues through this repository's Discussions page, or through Blackfrost-AI/CYBER-FROST-3.8-FP8 for the FP8 release. Do not include secrets, client data, live targets, or sensitive exploit details.

Downloads last month
657
Safetensors
Model size
33B params
Tensor type
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vcruz305/CYBER-FROST-3.8-EXL3-SAGE-3.87bpw

Quantized
(10)
this model