OpenJev-GGUF / convert.log
ggerganov's picture
ggerganov HF Staff
Upload folder using huggingface_hub
10840f3 verified
Raw History Blame Contribute Delete
458 kB
+ OUTPUT_DIR=./upload-OpenJev
+ LLAMA_CPP=llama.cpp
+ DISPLAY_NAME=OpenJev
+ QUANTIZE=llama.cpp/build/bin/llama-quantize
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-OpenJev-PRIMARY --outtype bf16 --outfile ./upload-OpenJev/OpenJev-BF16.gguf --no-mtp --model-name OpenJev
INFO:hf-to-gguf:Loading model: model-temp-OpenJev-PRIMARY
INFO:hf-to-gguf:gguf: detected OpenJev checkpoint
[transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial'
INFO:hf-to-gguf:Model architecture: OpenJevModel
INFO:hf-to-gguf:gguf: detected OpenJev checkpoint
[transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial'
INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00012.safetensors'
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
INFO:hf-to-gguf:Exporting model...
INFO:hf-to-gguf:output.weight, torch.bfloat16 --> BF16, shape = {5120, 248320}
INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> BF16, shape = {5120, 248320}
INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.0.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.0.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.0.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.0.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.0.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.0.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.0.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.0.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.0.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.1.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.1.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.1.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.1.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.1.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.1.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.1.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.1.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.1.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.2.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.2.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.2.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.2.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.2.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.2.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.2.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.2.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.2.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.3.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.3.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.3.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.3.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.4.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.4.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.4.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.4.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.4.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.4.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.4.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.4.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.4.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.4.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.5.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.5.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.5.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.5.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.5.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.5.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.5.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.5.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.5.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.5.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.6.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.6.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.6.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.6.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.6.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.6.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.6.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.6.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.6.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.6.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.7.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.7.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.7.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.7.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.8.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.8.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.8.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.8.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.8.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.8.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.8.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.8.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.8.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.8.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.9.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.9.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.9.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.9.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.9.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.9.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.9.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.9.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.10.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.10.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.10.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.10.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.10.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.10.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.10.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.10.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.10.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.11.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.11.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.11.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.12.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.12.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.12.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.12.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.12.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.12.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.12.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.12.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.12.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.13.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.13.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.13.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.13.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.13.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.13.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.13.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.13.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.13.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.14.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.14.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.14.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.14.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.14.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.14.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.14.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.14.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.14.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.15.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.15.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.15.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.16.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.16.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.16.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.16.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.16.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.9.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.9.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.16.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.16.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.16.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.16.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.17.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.17.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.17.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.17.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.17.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.17.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.17.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.17.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.17.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.18.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.18.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.18.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.18.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.18.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.18.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.18.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.18.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.18.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.19.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.19.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.19.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.20.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.20.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.20.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.20.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.20.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.20.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.20.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.20.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.20.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.21.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.21.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.21.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.21.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.21.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.21.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.21.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.21.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.21.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.22.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.22.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.22.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.22.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.22.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.22.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.22.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.22.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.22.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.23.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.23.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.23.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.23.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.24.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.24.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.24.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.24.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.24.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.24.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.24.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.24.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.24.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.24.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.25.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.25.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.25.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.25.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.25.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.25.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.25.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.25.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.25.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.25.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.26.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.26.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.26.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.26.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.26.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.26.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.26.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.26.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.26.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.26.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.26.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.27.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.27.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.27.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.27.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.27.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.27.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.27.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.27.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.27.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.27.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.27.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.28.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.28.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.28.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.28.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.28.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.28.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.28.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.28.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.28.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.28.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.28.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.28.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.28.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.28.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.29.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.29.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.29.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.29.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.29.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.29.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.29.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.29.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.29.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.29.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.29.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.29.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.29.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.29.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.30.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.30.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.30.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.30.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.30.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.30.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.30.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.30.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.30.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.30.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.30.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.30.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.30.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.30.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.31.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.31.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.31.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.31.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.31.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.31.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.31.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.31.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.31.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.31.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.31.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.32.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.32.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.32.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.32.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.32.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.32.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.32.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.32.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.32.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.32.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.32.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.32.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.32.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.32.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.33.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.33.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.33.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.33.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.33.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.33.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.33.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.33.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.33.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.33.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.33.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.33.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.33.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.33.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.34.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.34.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.34.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.34.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.34.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.34.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.34.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.34.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.34.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.34.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.34.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.34.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.34.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.34.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.35.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.35.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.35.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.35.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.35.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.35.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.35.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.35.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.35.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.35.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.35.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.36.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.36.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.36.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.36.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.36.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.36.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.36.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.36.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.36.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.36.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.36.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.36.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.36.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.36.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.37.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.37.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.37.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.37.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.37.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.37.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.37.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.37.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.37.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.37.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.37.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.37.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.37.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.37.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.38.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.38.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.38.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.38.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.38.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.38.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.38.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.38.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.38.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.38.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.38.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.38.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.38.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.38.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.39.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.39.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.39.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.39.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.39.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.39.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.39.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.39.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.39.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.39.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.39.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.40.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.40.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.40.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.40.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.40.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.40.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.40.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.40.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.40.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.40.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.40.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.40.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.40.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.40.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.41.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.41.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.41.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.41.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.41.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.41.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.41.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.41.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.41.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.41.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.41.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.41.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.41.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.41.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.42.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.42.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.42.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.42.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.42.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.42.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.42.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.42.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.42.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.42.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.42.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.42.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.42.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.42.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.43.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.43.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.43.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.43.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.43.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.43.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.43.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.43.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.43.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.43.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.43.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.44.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.44.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.44.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.44.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.44.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.44.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.44.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.44.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.44.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.44.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.44.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.44.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.44.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.44.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.45.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.45.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.45.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.45.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.45.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.45.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.45.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.45.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.45.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.45.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.45.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.45.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.45.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.45.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.46.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.46.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.46.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.46.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.46.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.46.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.46.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.46.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.46.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.46.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.46.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.46.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.46.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.46.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.47.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.47.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.47.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.47.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.47.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.47.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.47.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.47.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.47.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.47.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.47.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.48.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.48.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.48.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.48.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.48.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.48.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.48.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.48.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.48.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.48.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.48.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.48.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.48.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.48.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.49.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.49.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.49.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.49.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.49.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.49.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.49.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.49.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.49.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.49.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.49.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.49.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.49.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.49.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.50.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.50.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.50.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.50.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.50.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.50.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.50.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.50.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.50.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.50.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.50.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.50.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.50.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.50.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.51.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.51.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.51.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.51.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.51.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.51.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.51.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.51.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.51.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.51.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.51.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.52.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.52.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.52.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.52.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.52.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.52.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.52.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.52.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.52.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.52.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.52.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.52.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.52.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.52.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.53.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.53.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.53.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.53.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.53.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.53.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.53.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.53.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.53.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.53.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.53.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.53.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.53.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.53.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.54.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.54.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.54.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.54.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.54.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.54.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.54.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.54.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.54.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.54.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.54.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.54.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.54.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.54.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.55.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.55.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.55.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.55.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.55.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.55.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.55.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.55.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.55.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.55.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.55.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.56.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.56.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.56.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.56.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.56.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.56.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.56.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.56.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.56.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.56.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.56.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.56.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.56.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.56.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.57.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.57.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.57.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.57.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.57.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.57.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.57.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.57.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.57.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.57.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.57.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.57.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.57.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.57.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.58.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.58.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.58.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.58.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.58.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.58.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.58.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.58.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.58.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.58.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.58.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.58.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.58.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.58.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.59.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.59.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.59.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.59.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.59.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.59.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.59.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.59.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.59.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.59.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.59.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.60.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.60.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.60.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.60.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.60.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.60.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.60.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.60.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.60.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.60.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.60.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.60.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.60.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.60.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.61.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.61.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.61.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.61.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.61.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.61.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.61.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.61.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.61.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.61.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.61.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.61.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.61.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.61.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.62.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.62.ssm_a, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.62.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240}
INFO:hf-to-gguf:blk.62.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48}
INFO:hf-to-gguf:blk.62.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.62.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48}
INFO:hf-to-gguf:blk.62.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240}
INFO:hf-to-gguf:blk.62.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144}
INFO:hf-to-gguf:blk.62.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128}
INFO:hf-to-gguf:blk.62.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.62.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.62.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.62.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.62.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.63.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.63.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120}
INFO:hf-to-gguf:blk.63.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.63.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408}
INFO:hf-to-gguf:blk.63.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:blk.63.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.63.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:blk.63.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120}
INFO:hf-to-gguf:blk.63.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256}
INFO:hf-to-gguf:blk.63.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288}
INFO:hf-to-gguf:blk.63.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024}
INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:Set meta model
INFO:hf-to-gguf:Set model parameters
INFO:hf-to-gguf:gguf: context length = 262144
INFO:hf-to-gguf:gguf: embedding length = 5120
INFO:hf-to-gguf:gguf: feed forward length = 17408
INFO:hf-to-gguf:gguf: head count = 24
INFO:hf-to-gguf:gguf: key-value head count = 4
WARNING:hf-to-gguf:Unknown RoPE type: default
INFO:hf-to-gguf:gguf: rope scaling type = NONE
INFO:hf-to-gguf:gguf: mrope sections: [11, 11, 10, 0]
INFO:hf-to-gguf:gguf: rope theta = 10000000
INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-06
INFO:hf-to-gguf:gguf: file type = 32
INFO:hf-to-gguf:Set model quantization version
INFO:hf-to-gguf:Set model tokenizer
[transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial'
INFO:gguf.vocab:Adding 247587 merge(s).
INFO:gguf.vocab:Setting special token type eos to 248046
INFO:gguf.vocab:Setting special token type pad to 248044
INFO:gguf.vocab:Setting special token type bos to 248044
INFO:gguf.vocab:Setting add_bos_token to False
INFO:gguf.vocab:Setting add_eos_token to False
INFO:gguf.vocab:Setting chat_template to {%- set image_count = namespace(value=0) %}
{%- set video_count = namespace(value=0) %}
{%- macro render_content(content, do_vision_count, is_system_content=false) %}
{%- if content is string %}
{{- content }}
{%- elif content is iterable and content is not mapping %}
{%- for item in content %}
{%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
{%- if is_system_content %}
{{- raise_exception('System message cannot contain images.') }}
{%- endif %}
{%- if do_vision_count %}
{%- set image_count.value = image_count.value + 1 %}
{%- endif %}
{%- if add_vision_id %}
{{- 'Picture ' ~ image_count.value ~ ': ' }}
{%- endif %}
{{- '<|vision_start|><|image_pad|><|vision_end|>' }}
{%- elif 'video' in item or item.type == 'video' %}
{%- if is_system_content %}
{{- raise_exception('System message cannot contain videos.') }}
{%- endif %}
{%- if do_vision_count %}
{%- set video_count.value = video_count.value + 1 %}
{%- endif %}
{%- if add_vision_id %}
{{- 'Video ' ~ video_count.value ~ ': ' }}
{%- endif %}
{{- '<|vision_start|><|video_pad|><|vision_end|>' }}
{%- elif 'text' in item %}
{{- item.text }}
{%- else %}
{{- raise_exception('Unexpected item type in content.') }}
{%- endif %}
{%- endfor %}
{%- elif content is none or content is undefined %}
{{- '' }}
{%- else %}
{{- raise_exception('Unexpected content type.') }}
{%- endif %}
{%- endmacro %}
{%- if not messages %}
{{- raise_exception('No messages provided.') }}
{%- endif %}
{%- set reasoning_instructions = '' %}
{%- if enable_thinking is undefined or enable_thinking is true %}
{%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}
{%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}
{{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}
{%- endif %}
{%- if resolved_reasoning_effort == 'xhigh' %}
{%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}
{%- elif resolved_reasoning_effort == 'low' %}
{%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}
{%- endif %}
{%- endif %}
{%- if tools and tools is iterable and tools is not mapping %}
{{- '<|im_start|>system\n' }}
{%- if reasoning_instructions %}
{{- reasoning_instructions + '\n\n' }}
{%- endif %}
{{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>" }}
{{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
{%- if messages[0].role == 'system' %}
{%- set content = render_content(messages[0].content, false, true)|trim %}
{%- if content %}
{{- '\n\n' + content }}
{%- endif %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- else %}
{%- if messages[0].role == 'system' %}
{%- set content = render_content(messages[0].content, false, true)|trim %}
{%- if content %}
{{- '<|im_start|>system\n' + (reasoning_instructions + '\n\n' if reasoning_instructions else '') + content + '<|im_end|>\n' }}
{%- elif reasoning_instructions %}
{{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }}
{%- endif %}
{%- elif reasoning_instructions %}
{{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
{%- for message in messages[::-1] %}
{%- set index = (messages|length - 1) - loop.index0 %}
{%- if ns.multi_step_tool and message.role == "user" %}
{%- set content = render_content(message.content, false)|trim %}
{%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
{%- set ns.multi_step_tool = false %}
{%- set ns.last_query_index = index %}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if ns.multi_step_tool %}
{{- raise_exception('No user query found in messages.') }}
{%- endif %}
{%- for message in messages %}
{%- set content = render_content(message.content, true)|trim %}
{%- if message.role == "system" %}
{%- if not loop.first %}
{{- raise_exception('System message must be at the beginning.') }}
{%- endif %}
{%- elif message.role == "user" %}
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{%- set reasoning_content = '' %}
{%- if message.reasoning_content is string %}
{%- set reasoning_content = message.reasoning_content %}
{%- endif %}
{%- set reasoning_content = reasoning_content|trim %}
{%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
{%- else %}
{{- '<|im_start|>' + message.role + '\n' + content }}
{%- endif %}
{%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{%- if loop.first %}
{%- if content|trim %}
{{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- else %}
{{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- endif %}
{%- else %}
{{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
{%- endif %}
{%- if tool_call.arguments is defined and tool_call.arguments != '' %}
{%- for args_name, args_value in tool_call.arguments|items %}
{{- '<parameter=' + args_name + '>\n' }}
{%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}
{{- args_value }}
{{- '\n</parameter>\n' }}
{%- endfor %}
{%- endif %}
{{- '</function>\n</tool_call>' }}
{%- endfor %}
{%- endif %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if loop.previtem and loop.previtem.role != "tool" %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- content }}
{{- '\n</tool_response>' }}
{%- if not loop.last and loop.nextitem.role != "tool" %}
{{- '<|im_end|>\n' }}
{%- elif loop.last %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- else %}
{{- raise_exception('Unexpected message role.') }}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- if enable_thinking is defined and enable_thinking is false %}
{{- '<think>\n\n</think>\n\n' }}
{%- else %}
{{- '<think>\n' }}
{%- endif %}
{%- endif %}
INFO:gguf.gguf_writer:Writing the following files:
INFO:gguf.gguf_writer:upload-OpenJev/OpenJev-BF16.gguf: n_tensors = 851, total_size = 53.8G
Writing: 0%| | 0.00/53.8G [00:00<?, ?byte/s]
Writing: 5%|▍ | 2.54G/53.8G [00:05<01:52, 454Mbyte/s]
Writing: 9%|β–‰ | 5.09G/53.8G [00:11<01:50, 442Mbyte/s]
Writing: 11%|β–ˆ | 5.67G/53.8G [00:13<01:52, 427Mbyte/s]
Writing: 12%|β–ˆβ– | 6.26G/53.8G [00:14<01:49, 435Mbyte/s]
Writing: 13%|β–ˆβ–Ž | 6.79G/53.8G [00:15<01:46, 440Mbyte/s]
Writing: 14%|β–ˆβ–Ž | 7.39G/53.8G [00:16<01:42, 453Mbyte/s]
Writing: 15%|β–ˆβ– | 7.92G/53.8G [00:17<01:39, 461Mbyte/s]
Writing: 16%|β–ˆβ–Œ | 8.54G/53.8G [00:18<01:36, 470Mbyte/s]
Writing: 17%|β–ˆβ–‹ | 9.07G/53.8G [00:20<01:35, 468Mbyte/s]
Writing: 18%|β–ˆβ–Š | 9.66G/53.8G [00:21<01:32, 476Mbyte/s]
Writing: 19%|β–ˆβ–‰ | 10.3G/53.8G [00:22<01:31, 475Mbyte/s]
Writing: 20%|β–ˆβ–ˆ | 10.8G/53.8G [00:23<01:29, 479Mbyte/s]
Writing: 21%|β–ˆβ–ˆ | 11.3G/53.8G [00:24<01:28, 479Mbyte/s]
Writing: 22%|β–ˆβ–ˆβ– | 11.9G/53.8G [00:26<01:26, 484Mbyte/s]
Writing: 23%|β–ˆβ–ˆβ–Ž | 12.5G/53.8G [00:27<01:27, 475Mbyte/s]
Writing: 24%|β–ˆβ–ˆβ– | 13.1G/53.8G [00:28<01:24, 480Mbyte/s]
Writing: 25%|β–ˆβ–ˆβ–Œ | 13.7G/53.8G [00:29<01:23, 482Mbyte/s]
Writing: 27%|β–ˆβ–ˆβ–‹ | 14.3G/53.8G [00:30<01:21, 485Mbyte/s]
Writing: 28%|β–ˆβ–ˆβ–Š | 14.8G/53.8G [00:32<01:21, 477Mbyte/s]
Writing: 29%|β–ˆβ–ˆβ–Š | 15.4G/53.8G [00:33<01:19, 483Mbyte/s]
Writing: 30%|β–ˆβ–ˆβ–‰ | 16.0G/53.8G [00:34<01:18, 481Mbyte/s]
Writing: 31%|β–ˆβ–ˆβ–ˆ | 16.5G/53.8G [00:35<01:17, 484Mbyte/s]
Writing: 32%|β–ˆβ–ˆβ–ˆβ– | 17.1G/53.8G [00:36<01:15, 489Mbyte/s]
Writing: 33%|β–ˆβ–ˆβ–ˆβ–Ž | 17.7G/53.8G [00:37<01:14, 483Mbyte/s]
Writing: 34%|β–ˆβ–ˆβ–ˆβ– | 18.2G/53.8G [00:39<01:14, 476Mbyte/s]
Writing: 35%|β–ˆβ–ˆβ–ˆβ– | 18.8G/53.8G [00:40<01:12, 483Mbyte/s]
Writing: 36%|β–ˆβ–ˆβ–ˆβ–Œ | 19.4G/53.8G [00:41<01:11, 480Mbyte/s]
Writing: 37%|β–ˆβ–ˆβ–ˆβ–‹ | 19.9G/53.8G [00:42<01:10, 482Mbyte/s]
Writing: 38%|β–ˆβ–ˆβ–ˆβ–Š | 20.4G/53.8G [00:43<01:09, 482Mbyte/s]
Writing: 39%|β–ˆβ–ˆβ–ˆβ–‰ | 21.1G/53.8G [00:44<01:07, 486Mbyte/s]
Writing: 40%|β–ˆβ–ˆβ–ˆβ–ˆ | 21.7G/53.8G [00:46<01:06, 481Mbyte/s]
Writing: 41%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 22.3G/53.8G [00:47<01:06, 478Mbyte/s]
Writing: 42%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 22.8G/53.8G [00:48<01:04, 480Mbyte/s]
Writing: 43%|β–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 23.3G/53.8G [00:49<01:03, 482Mbyte/s]
Writing: 45%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 23.9G/53.8G [00:50<01:01, 482Mbyte/s]
Writing: 46%|β–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 24.5G/53.8G [00:52<01:01, 477Mbyte/s]
Writing: 47%|β–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 25.1G/53.8G [00:53<01:01, 468Mbyte/s]
Writing: 48%|β–ˆβ–ˆβ–ˆβ–ˆβ–Š | 25.7G/53.8G [00:54<00:59, 472Mbyte/s]
Writing: 49%|β–ˆβ–ˆβ–ˆβ–ˆβ–Š | 26.2G/53.8G [00:55<00:58, 474Mbyte/s]
Writing: 50%|β–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 26.8G/53.8G [00:57<00:56, 477Mbyte/s]
Writing: 51%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 27.3G/53.8G [00:58<00:56, 469Mbyte/s]
Writing: 52%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 27.9G/53.8G [00:59<00:54, 474Mbyte/s]
Writing: 53%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 28.5G/53.8G [01:00<00:53, 473Mbyte/s]
Writing: 54%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 29.1G/53.8G [01:01<00:51, 476Mbyte/s]
Writing: 55%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 29.5G/53.8G [01:02<00:51, 475Mbyte/s]
Writing: 56%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 30.2G/53.8G [01:04<00:49, 479Mbyte/s]
Writing: 57%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 30.8G/53.8G [01:05<00:48, 472Mbyte/s]
Writing: 58%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 31.4G/53.8G [01:06<00:47, 469Mbyte/s]
Writing: 59%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 31.9G/53.8G [01:07<00:46, 471Mbyte/s]
Writing: 60%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 32.5G/53.8G [01:09<00:45, 473Mbyte/s]
Writing: 61%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 33.1G/53.8G [01:10<00:43, 471Mbyte/s]
Writing: 63%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 33.7G/53.8G [01:11<00:43, 466Mbyte/s]
Writing: 64%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 34.2G/53.8G [01:12<00:42, 457Mbyte/s]
Writing: 65%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 34.8G/53.8G [01:14<00:41, 461Mbyte/s]
Writing: 66%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 35.3G/53.8G [01:15<00:40, 462Mbyte/s]
Writing: 67%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 35.9G/53.8G [01:16<00:38, 463Mbyte/s]
Writing: 68%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 36.5G/53.8G [01:17<00:38, 455Mbyte/s]
Writing: 69%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 37.1G/53.8G [01:19<00:36, 459Mbyte/s]
Writing: 70%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 37.7G/53.8G [01:20<00:35, 457Mbyte/s]
Writing: 71%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 38.2G/53.8G [01:21<00:34, 459Mbyte/s]
Writing: 72%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 38.7G/53.8G [01:22<00:33, 458Mbyte/s]
Writing: 73%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 39.2G/53.8G [01:23<00:31, 461Mbyte/s]
Writing: 74%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 39.8G/53.8G [01:24<00:30, 458Mbyte/s]
Writing: 75%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 40.2G/53.8G [01:26<00:29, 453Mbyte/s]
Writing: 76%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 40.7G/53.8G [01:27<00:28, 457Mbyte/s]
Writing: 77%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 41.2G/53.8G [01:28<00:27, 459Mbyte/s]
Writing: 78%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 41.7G/53.8G [01:29<00:26, 458Mbyte/s]
Writing: 78%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 42.2G/53.8G [01:30<00:25, 461Mbyte/s]
Writing: 80%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 42.8G/53.8G [01:31<00:24, 456Mbyte/s]
Writing: 80%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 43.3G/53.8G [01:32<00:23, 452Mbyte/s]
Writing: 81%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 43.7G/53.8G [01:33<00:22, 457Mbyte/s]
Writing: 82%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 44.3G/53.8G [01:34<00:20, 459Mbyte/s]
Writing: 83%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž | 44.8G/53.8G [01:35<00:19, 458Mbyte/s]
Writing: 84%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 45.3G/53.8G [01:36<00:18, 462Mbyte/s]
Writing: 85%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 45.8G/53.8G [01:38<00:17, 457Mbyte/s]
Writing: 86%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ | 46.3G/53.8G [01:39<00:16, 452Mbyte/s]
Writing: 87%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹ | 46.8G/53.8G [01:40<00:15, 456Mbyte/s]
Writing: 88%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š | 47.3G/53.8G [01:41<00:14, 458Mbyte/s]
Writing: 89%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 47.8G/53.8G [01:42<00:13, 457Mbyte/s]
Writing: 90%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰ | 48.3G/53.8G [01:43<00:11, 461Mbyte/s]
Writing: 91%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 48.9G/53.8G [01:44<00:10, 457Mbyte/s]
Writing: 92%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 49.3G/53.8G [01:45<00:09, 452Mbyte/s]
Writing: 93%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 49.8G/53.8G [01:47<00:08, 457Mbyte/s]
Writing: 94%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Ž| 50.4G/53.8G [01:48<00:07, 459Mbyte/s]
Writing: 95%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–| 50.9G/53.8G [01:49<00:06, 458Mbyte/s]
Writing: 95%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ| 51.3G/53.8G [01:50<00:05, 462Mbyte/s]
Writing: 97%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹| 51.9G/53.8G [01:51<00:04, 460Mbyte/s]
Writing: 97%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‹| 52.4G/53.8G [01:52<00:03, 458Mbyte/s]
Writing: 98%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š| 52.9G/53.8G [01:53<00:01, 462Mbyte/s]
Writing: 99%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰| 53.4G/53.8G [01:54<00:00, 468Mbyte/s] Writing: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 53.8G/53.8G [01:55<00:00, 466Mbyte/s]
INFO:hf-to-gguf:Model successfully exported to upload-OpenJev/OpenJev-BF16.gguf
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-OpenJev-PRIMARY --outtype bf16 --outfile ./upload-OpenJev/mmproj-OpenJev-BF16.gguf --mmproj --model-name OpenJev
INFO:hf-to-gguf:Loading model: model-temp-OpenJev-PRIMARY
INFO:hf-to-gguf:gguf: detected OpenJev checkpoint
[transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial'
INFO:hf-to-gguf:Model architecture: OpenJevModel
INFO:hf-to-gguf:gguf: detected OpenJev checkpoint
[transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial'
INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00012.safetensors'
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
INFO:hf-to-gguf:Exporting model...
INFO:hf-to-gguf:v.blk.0.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.0.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.0.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.0.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.0.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.1.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.1.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.1.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.1.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.10.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.10.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.10.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.10.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.11.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.11.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.11.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.11.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.11.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.12.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.12.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.12.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.12.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.13.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.13.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.13.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.13.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.14.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.14.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.14.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.14.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.15.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.15.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.15.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.15.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.15.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.16.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.16.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.16.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.16.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.17.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.17.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.17.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.17.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.18.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.18.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.18.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.18.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.19.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.19.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.19.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.19.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.19.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.2.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.2.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.2.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.2.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.20.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.20.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.20.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.20.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.21.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.21.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.21.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.21.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.22.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.22.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.22.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.22.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.23.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.23.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.23.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.23.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.23.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.24.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.24.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.24.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.24.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.25.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.25.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.25.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.25.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.26.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.26.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.26.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.26.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.3.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.3.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.3.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.3.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.3.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.4.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.4.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.4.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.4.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.5.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.5.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.5.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.5.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.6.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.6.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.6.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.6.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.7.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.7.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.7.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.7.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.7.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.8.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.8.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.8.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.8.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.9.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.9.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.9.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.9.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:mm.0.bias, torch.bfloat16 --> F32, shape = {4608}
INFO:hf-to-gguf:mm.0.weight, torch.bfloat16 --> BF16, shape = {4608, 4608}
INFO:hf-to-gguf:mm.2.bias, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:mm.2.weight, torch.bfloat16 --> BF16, shape = {4608, 5120}
INFO:hf-to-gguf:v.post_ln.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.post_ln.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.patch_embd.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.patch_embd.weight, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152}
INFO:hf-to-gguf:v.patch_embd.weight.1, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152}
INFO:hf-to-gguf:v.position_embd.weight, torch.bfloat16 --> F32, shape = {1152, 2304}
INFO:hf-to-gguf:Set meta model
INFO:hf-to-gguf:Set model parameters
INFO:hf-to-gguf:Set model quantization version
INFO:gguf.gguf_writer:Writing the following files:
INFO:gguf.gguf_writer:upload-OpenJev/mmproj-OpenJev-BF16.gguf: n_tensors = 334, total_size = 931.1M
Writing: 0%| | 0.00/931M [00:00<?, ?byte/s]
Writing: 48%|β–ˆβ–ˆβ–ˆβ–ˆβ–Š | 448M/931M [00:01<00:01, 445Mbyte/s]
Writing: 98%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š| 913M/931M [00:02<00:00, 449Mbyte/s] Writing: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 931M/931M [00:02<00:00, 448Mbyte/s]
INFO:hf-to-gguf:Model successfully exported to upload-OpenJev/mmproj-OpenJev-BF16.gguf
+ FLAGS_Q4_K_M='--pure --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ssm_=q8_0'
+ llama.cpp/build/bin/llama-quantize ./upload-OpenJev/OpenJev-BF16.gguf ./upload-OpenJev/OpenJev-Q8_0.gguf Q8_0
version: 0.5.0-dev (build 11354, commit 54a0c5da9)
built with GNU 14.2.0 for Linux x86_64
llama_quantize: quantizing './upload-OpenJev/OpenJev-BF16.gguf' to './upload-OpenJev/OpenJev-Q8_0.gguf' as Q8_0
llama_model_loader: loaded meta data with 48 key-value pairs and 851 tensors from ./upload-OpenJev/OpenJev-BF16.gguf (version GGUF V3 (latest))
llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
llama_model_loader: - kv 0: general.architecture str = qwen35
llama_model_loader: - kv 1: general.type str = model
llama_model_loader: - kv 2: general.sampling.top_k i32 = 20
llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.950000
llama_model_loader: - kv 4: general.sampling.temp f32 = 1.000000
llama_model_loader: - kv 5: general.name str = OpenJev
llama_model_loader: - kv 6: general.size_label str = 27B
llama_model_loader: - kv 7: general.license str = cc-by-nc-4.0
llama_model_loader: - kv 8: general.tags arr[str,8] = ["decision-model", "zero-shot-classif...
llama_model_loader: - kv 9: general.languages arr[str,6] = ["en", "de", "fr", "hi", "zh", "ja"]
llama_model_loader: - kv 10: qwen35.block_count u32 = 64
llama_model_loader: - kv 11: qwen35.context_length u32 = 262144
llama_model_loader: - kv 12: qwen35.embedding_length u32 = 5120
llama_model_loader: - kv 13: qwen35.feed_forward_length u32 = 17408
llama_model_loader: - kv 14: qwen35.attention.head_count u32 = 24
llama_model_loader: - kv 15: qwen35.attention.head_count_kv u32 = 4
llama_model_loader: - kv 16: qwen35.rope.dimension_sections arr[i32,4] = [11, 11, 10, 0]
llama_model_loader: - kv 17: qwen35.rope.freq_base f32 = 10000000.000000
llama_model_loader: - kv 18: qwen35.attention.layer_norm_rms_epsilon f32 = 0.000001
llama_model_loader: - kv 19: qwen35.attention.key_length u32 = 256
llama_model_loader: - kv 20: qwen35.attention.value_length u32 = 256
llama_model_loader: - kv 21: general.file_type u32 = 32
llama_model_loader: - kv 22: qwen35.ssm.conv_kernel u32 = 4
llama_model_loader: - kv 23: qwen35.ssm.state_size u32 = 128
llama_model_loader: - kv 24: qwen35.ssm.group_count u32 = 16
llama_model_loader: - kv 25: qwen35.ssm.time_step_rank u32 = 48
llama_model_loader: - kv 26: qwen35.ssm.inner_size u32 = 6144
llama_model_loader: - kv 27: qwen35.attention.recurrent_layers arr[bool,64] = [true, true, true, false, true, true,...
llama_model_loader: - kv 28: qwen35.full_attention_interval u32 = 4
llama_model_loader: - kv 29: qwen35.rope.dimension_count u32 = 64
llama_model_loader: - kv 30: qwen35.decision.type str = openjev
llama_model_loader: - kv 31: qwen35.decision.temperature.choice f32 = 0.850000
llama_model_loader: - kv 32: qwen35.decision.temperature.score f32 = 0.850000
llama_model_loader: - kv 33: qwen35.decision.temperature.noul f32 = 1.554713
llama_model_loader: - kv 34: general.quantization_version u32 = 2
llama_model_loader: - kv 35: tokenizer.ggml.model str = gpt2
llama_model_loader: - kv 36: tokenizer.ggml.pre str = qwen35
llama_model_loader: - kv 37: tokenizer.ggml.tokens arr[str,248320] = ["!", "\"", "#", "$", "%", "&", "'", ...
llama_model_loader: - kv 38: tokenizer.ggml.token_type arr[i32,248320] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
llama_model_loader: - kv 39: tokenizer.ggml.merges arr[str,247587] = ["Δ  Δ ", "Δ Δ  Δ Δ ", "i n", "Δ  t",...
llama_model_loader: - kv 40: tokenizer.ggml.eos_token_id u32 = 248046
llama_model_loader: - kv 41: tokenizer.ggml.padding_token_id u32 = 248044
llama_model_loader: - kv 42: tokenizer.ggml.bos_token_id u32 = 248044
llama_model_loader: - kv 43: tokenizer.ggml.add_bos_token bool = false
llama_model_loader: - kv 44: tokenizer.ggml.add_eos_token bool = false
llama_model_loader: - kv 45: tokenizer.chat_template str = {%- set image_count = namespace(value...
llama_model_loader: - kv 46: tokenizer.chat_template.systemone str = {% set letters = 'ABCDEFGHIJKLMNOPQRS...
llama_model_loader: - kv 47: tokenizer.chat_templates arr[str,1] = ["systemone"]
llama_model_loader: - type f32: 353 tensors
llama_model_loader: - type bf16: 498 tensors
[ 1/ 851] output.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q8_0 .. size = 2425.00 MiB -> 1288.28 MiB
[ 2/ 851] output_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 3/ 851] token_embd.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q8_0 .. size = 2425.00 MiB -> 1288.28 MiB
[ 4/ 851] blk.0.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 5/ 851] blk.0.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 6/ 851] blk.0.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 7/ 851] blk.0.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 8/ 851] blk.0.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 9/ 851] blk.0.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 10/ 851] blk.0.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 11/ 851] blk.0.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 12/ 851] blk.0.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 13/ 851] blk.0.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 14/ 851] blk.0.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 15/ 851] blk.0.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 16/ 851] blk.0.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 17/ 851] blk.0.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 18/ 851] blk.1.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 19/ 851] blk.1.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 20/ 851] blk.1.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 21/ 851] blk.1.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 22/ 851] blk.1.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 23/ 851] blk.1.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 24/ 851] blk.1.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 25/ 851] blk.1.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 26/ 851] blk.1.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 27/ 851] blk.1.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 28/ 851] blk.1.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 29/ 851] blk.1.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 30/ 851] blk.1.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 31/ 851] blk.1.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 32/ 851] blk.2.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 33/ 851] blk.2.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 34/ 851] blk.2.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 35/ 851] blk.2.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 36/ 851] blk.2.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 37/ 851] blk.2.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 38/ 851] blk.2.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 39/ 851] blk.2.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 40/ 851] blk.2.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 41/ 851] blk.2.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 42/ 851] blk.2.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 43/ 851] blk.2.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 44/ 851] blk.2.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 45/ 851] blk.2.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 46/ 851] blk.3.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 47/ 851] blk.3.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 48/ 851] blk.3.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 49/ 851] blk.3.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 50/ 851] blk.3.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 51/ 851] blk.3.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 52/ 851] blk.3.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 53/ 851] blk.3.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 54/ 851] blk.3.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 55/ 851] blk.3.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 56/ 851] blk.3.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 57/ 851] blk.4.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 58/ 851] blk.4.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 59/ 851] blk.4.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 60/ 851] blk.4.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 61/ 851] blk.4.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 62/ 851] blk.4.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 63/ 851] blk.4.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 64/ 851] blk.4.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 65/ 851] blk.4.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 66/ 851] blk.4.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 67/ 851] blk.4.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 68/ 851] blk.4.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 69/ 851] blk.4.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 70/ 851] blk.4.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 71/ 851] blk.5.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 72/ 851] blk.5.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 73/ 851] blk.5.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 74/ 851] blk.5.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 75/ 851] blk.5.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 76/ 851] blk.5.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 77/ 851] blk.5.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 78/ 851] blk.5.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 79/ 851] blk.5.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 80/ 851] blk.5.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 81/ 851] blk.5.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 82/ 851] blk.5.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 83/ 851] blk.5.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 84/ 851] blk.5.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 85/ 851] blk.6.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 86/ 851] blk.6.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 87/ 851] blk.6.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 88/ 851] blk.6.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 89/ 851] blk.6.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 90/ 851] blk.6.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 91/ 851] blk.6.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 92/ 851] blk.6.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 93/ 851] blk.6.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 94/ 851] blk.6.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 95/ 851] blk.6.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 96/ 851] blk.6.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 97/ 851] blk.6.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 98/ 851] blk.6.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 99/ 851] blk.7.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 100/ 851] blk.7.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 101/ 851] blk.7.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 102/ 851] blk.7.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 103/ 851] blk.7.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 104/ 851] blk.7.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 105/ 851] blk.7.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 106/ 851] blk.7.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 107/ 851] blk.7.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 108/ 851] blk.7.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 109/ 851] blk.7.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 110/ 851] blk.8.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 111/ 851] blk.8.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 112/ 851] blk.8.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 113/ 851] blk.8.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 114/ 851] blk.8.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 115/ 851] blk.8.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 116/ 851] blk.8.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 117/ 851] blk.8.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 118/ 851] blk.8.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 119/ 851] blk.8.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 120/ 851] blk.8.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 121/ 851] blk.8.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 122/ 851] blk.8.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 123/ 851] blk.8.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 124/ 851] blk.9.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 125/ 851] blk.9.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 126/ 851] blk.9.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 127/ 851] blk.9.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 128/ 851] blk.9.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 129/ 851] blk.9.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 130/ 851] blk.9.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 131/ 851] blk.9.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 132/ 851] blk.9.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 133/ 851] blk.9.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 134/ 851] blk.9.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 135/ 851] blk.9.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 136/ 851] blk.9.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 137/ 851] blk.9.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 138/ 851] blk.10.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 139/ 851] blk.10.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 140/ 851] blk.10.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 141/ 851] blk.10.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 142/ 851] blk.10.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 143/ 851] blk.10.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 144/ 851] blk.10.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 145/ 851] blk.10.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 146/ 851] blk.10.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 147/ 851] blk.10.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 148/ 851] blk.10.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 149/ 851] blk.10.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 150/ 851] blk.10.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 151/ 851] blk.10.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 152/ 851] blk.11.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 153/ 851] blk.11.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 154/ 851] blk.11.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 155/ 851] blk.11.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 156/ 851] blk.11.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 157/ 851] blk.11.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 158/ 851] blk.11.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 159/ 851] blk.11.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 160/ 851] blk.11.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 161/ 851] blk.11.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 162/ 851] blk.11.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 163/ 851] blk.12.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 164/ 851] blk.12.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 165/ 851] blk.12.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 166/ 851] blk.12.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 167/ 851] blk.12.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 168/ 851] blk.12.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 169/ 851] blk.12.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 170/ 851] blk.12.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 171/ 851] blk.12.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 172/ 851] blk.12.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 173/ 851] blk.12.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 174/ 851] blk.12.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 175/ 851] blk.12.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 176/ 851] blk.12.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 177/ 851] blk.13.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 178/ 851] blk.13.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 179/ 851] blk.13.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 180/ 851] blk.13.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 181/ 851] blk.13.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 182/ 851] blk.13.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 183/ 851] blk.13.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 184/ 851] blk.13.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 185/ 851] blk.13.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 186/ 851] blk.13.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 187/ 851] blk.13.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 188/ 851] blk.13.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 189/ 851] blk.13.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 190/ 851] blk.13.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 191/ 851] blk.14.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 192/ 851] blk.14.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 193/ 851] blk.14.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 194/ 851] blk.14.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 195/ 851] blk.14.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 196/ 851] blk.14.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 197/ 851] blk.14.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 198/ 851] blk.14.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 199/ 851] blk.14.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 200/ 851] blk.14.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 201/ 851] blk.14.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 202/ 851] blk.14.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 203/ 851] blk.14.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 204/ 851] blk.14.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 205/ 851] blk.15.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 206/ 851] blk.15.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 207/ 851] blk.15.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 208/ 851] blk.15.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 209/ 851] blk.15.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 210/ 851] blk.15.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 211/ 851] blk.15.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 212/ 851] blk.15.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 213/ 851] blk.15.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 214/ 851] blk.15.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 215/ 851] blk.15.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 216/ 851] blk.16.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 217/ 851] blk.16.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 218/ 851] blk.16.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 219/ 851] blk.16.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 220/ 851] blk.16.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 221/ 851] blk.16.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 222/ 851] blk.16.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 223/ 851] blk.16.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 224/ 851] blk.16.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 225/ 851] blk.16.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 226/ 851] blk.16.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 227/ 851] blk.16.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 228/ 851] blk.16.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 229/ 851] blk.16.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 230/ 851] blk.17.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 231/ 851] blk.17.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 232/ 851] blk.17.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 233/ 851] blk.17.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 234/ 851] blk.17.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 235/ 851] blk.17.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 236/ 851] blk.17.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 237/ 851] blk.17.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 238/ 851] blk.17.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 239/ 851] blk.17.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 240/ 851] blk.17.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 241/ 851] blk.17.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 242/ 851] blk.17.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 243/ 851] blk.17.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 244/ 851] blk.18.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 245/ 851] blk.18.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 246/ 851] blk.18.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 247/ 851] blk.18.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 248/ 851] blk.18.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 249/ 851] blk.18.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 250/ 851] blk.18.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 251/ 851] blk.18.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 252/ 851] blk.18.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 253/ 851] blk.18.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 254/ 851] blk.18.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 255/ 851] blk.18.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 256/ 851] blk.18.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 257/ 851] blk.18.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 258/ 851] blk.19.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 259/ 851] blk.19.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 260/ 851] blk.19.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 261/ 851] blk.19.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 262/ 851] blk.19.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 263/ 851] blk.19.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 264/ 851] blk.19.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 265/ 851] blk.19.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 266/ 851] blk.19.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 267/ 851] blk.19.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 268/ 851] blk.19.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 269/ 851] blk.20.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 270/ 851] blk.20.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 271/ 851] blk.20.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 272/ 851] blk.20.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 273/ 851] blk.20.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 274/ 851] blk.20.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 275/ 851] blk.20.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 276/ 851] blk.20.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 277/ 851] blk.20.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 278/ 851] blk.20.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 279/ 851] blk.20.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 280/ 851] blk.20.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 281/ 851] blk.20.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 282/ 851] blk.20.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 283/ 851] blk.21.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 284/ 851] blk.21.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 285/ 851] blk.21.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 286/ 851] blk.21.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 287/ 851] blk.21.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 288/ 851] blk.21.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 289/ 851] blk.21.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 290/ 851] blk.21.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 291/ 851] blk.21.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 292/ 851] blk.21.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 293/ 851] blk.21.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 294/ 851] blk.21.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 295/ 851] blk.21.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 296/ 851] blk.21.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 297/ 851] blk.22.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 298/ 851] blk.22.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 299/ 851] blk.22.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 300/ 851] blk.22.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 301/ 851] blk.22.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 302/ 851] blk.22.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 303/ 851] blk.22.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 304/ 851] blk.22.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 305/ 851] blk.22.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 306/ 851] blk.22.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 307/ 851] blk.22.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 308/ 851] blk.22.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 309/ 851] blk.22.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 310/ 851] blk.22.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 311/ 851] blk.23.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 312/ 851] blk.23.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 313/ 851] blk.23.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 314/ 851] blk.23.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 315/ 851] blk.23.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 316/ 851] blk.23.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 317/ 851] blk.23.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 318/ 851] blk.23.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 319/ 851] blk.23.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 320/ 851] blk.23.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 321/ 851] blk.23.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 322/ 851] blk.24.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 323/ 851] blk.24.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 324/ 851] blk.24.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 325/ 851] blk.24.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 326/ 851] blk.24.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 327/ 851] blk.24.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 328/ 851] blk.24.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 329/ 851] blk.24.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 330/ 851] blk.24.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 331/ 851] blk.24.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 332/ 851] blk.24.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 333/ 851] blk.24.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 334/ 851] blk.24.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 335/ 851] blk.24.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 336/ 851] blk.25.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 337/ 851] blk.25.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 338/ 851] blk.25.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 339/ 851] blk.25.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 340/ 851] blk.25.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 341/ 851] blk.25.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 342/ 851] blk.25.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 343/ 851] blk.25.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 344/ 851] blk.25.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 345/ 851] blk.25.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 346/ 851] blk.25.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 347/ 851] blk.25.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 348/ 851] blk.25.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 349/ 851] blk.25.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 350/ 851] blk.26.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 351/ 851] blk.26.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 352/ 851] blk.26.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 353/ 851] blk.26.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 354/ 851] blk.26.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 355/ 851] blk.26.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 356/ 851] blk.26.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 357/ 851] blk.26.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 358/ 851] blk.26.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 359/ 851] blk.26.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 360/ 851] blk.26.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 361/ 851] blk.26.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 362/ 851] blk.26.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 363/ 851] blk.26.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 364/ 851] blk.27.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 365/ 851] blk.27.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 366/ 851] blk.27.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 367/ 851] blk.27.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 368/ 851] blk.27.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 369/ 851] blk.27.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 370/ 851] blk.27.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 371/ 851] blk.27.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 372/ 851] blk.27.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 373/ 851] blk.27.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 374/ 851] blk.27.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 375/ 851] blk.28.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 376/ 851] blk.28.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 377/ 851] blk.28.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 378/ 851] blk.28.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 379/ 851] blk.28.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 380/ 851] blk.28.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 381/ 851] blk.28.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 382/ 851] blk.28.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 383/ 851] blk.28.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 384/ 851] blk.28.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 385/ 851] blk.28.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 386/ 851] blk.28.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 387/ 851] blk.28.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 388/ 851] blk.28.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 389/ 851] blk.29.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 390/ 851] blk.29.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 391/ 851] blk.29.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 392/ 851] blk.29.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 393/ 851] blk.29.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 394/ 851] blk.29.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 395/ 851] blk.29.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 396/ 851] blk.29.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 397/ 851] blk.29.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 398/ 851] blk.29.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 399/ 851] blk.29.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 400/ 851] blk.29.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 401/ 851] blk.29.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 402/ 851] blk.29.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 403/ 851] blk.30.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 404/ 851] blk.30.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 405/ 851] blk.30.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 406/ 851] blk.30.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 407/ 851] blk.30.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 408/ 851] blk.30.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 409/ 851] blk.30.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 410/ 851] blk.30.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 411/ 851] blk.30.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 412/ 851] blk.30.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 413/ 851] blk.30.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 414/ 851] blk.30.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 415/ 851] blk.30.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 416/ 851] blk.30.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 417/ 851] blk.31.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 418/ 851] blk.31.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 419/ 851] blk.31.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 420/ 851] blk.31.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 421/ 851] blk.31.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 422/ 851] blk.31.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 423/ 851] blk.31.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 424/ 851] blk.31.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 425/ 851] blk.31.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 426/ 851] blk.31.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 427/ 851] blk.31.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 428/ 851] blk.32.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 429/ 851] blk.32.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 430/ 851] blk.32.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 431/ 851] blk.32.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 432/ 851] blk.32.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 433/ 851] blk.32.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 434/ 851] blk.32.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 435/ 851] blk.32.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 436/ 851] blk.32.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 437/ 851] blk.32.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 438/ 851] blk.32.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 439/ 851] blk.32.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 440/ 851] blk.32.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 441/ 851] blk.32.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 442/ 851] blk.33.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 443/ 851] blk.33.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 444/ 851] blk.33.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 445/ 851] blk.33.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 446/ 851] blk.33.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 447/ 851] blk.33.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 448/ 851] blk.33.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 449/ 851] blk.33.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 450/ 851] blk.33.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 451/ 851] blk.33.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 452/ 851] blk.33.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 453/ 851] blk.33.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 454/ 851] blk.33.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 455/ 851] blk.33.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 456/ 851] blk.34.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 457/ 851] blk.34.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 458/ 851] blk.34.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 459/ 851] blk.34.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 460/ 851] blk.34.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 461/ 851] blk.34.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 462/ 851] blk.34.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 463/ 851] blk.34.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 464/ 851] blk.34.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 465/ 851] blk.34.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 466/ 851] blk.34.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 467/ 851] blk.34.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 468/ 851] blk.34.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 469/ 851] blk.34.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 470/ 851] blk.35.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 471/ 851] blk.35.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 472/ 851] blk.35.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 473/ 851] blk.35.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 474/ 851] blk.35.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 475/ 851] blk.35.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 476/ 851] blk.35.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 477/ 851] blk.35.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 478/ 851] blk.35.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 479/ 851] blk.35.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 480/ 851] blk.35.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 481/ 851] blk.36.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 482/ 851] blk.36.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 483/ 851] blk.36.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 484/ 851] blk.36.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 485/ 851] blk.36.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 486/ 851] blk.36.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 487/ 851] blk.36.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 488/ 851] blk.36.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 489/ 851] blk.36.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 490/ 851] blk.36.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 491/ 851] blk.36.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 492/ 851] blk.36.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 493/ 851] blk.36.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 494/ 851] blk.36.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 495/ 851] blk.37.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 496/ 851] blk.37.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 497/ 851] blk.37.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 498/ 851] blk.37.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 499/ 851] blk.37.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 500/ 851] blk.37.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 501/ 851] blk.37.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 502/ 851] blk.37.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 503/ 851] blk.37.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 504/ 851] blk.37.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 505/ 851] blk.37.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 506/ 851] blk.37.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 507/ 851] blk.37.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 508/ 851] blk.37.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 509/ 851] blk.38.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 510/ 851] blk.38.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 511/ 851] blk.38.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 512/ 851] blk.38.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 513/ 851] blk.38.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 514/ 851] blk.38.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 515/ 851] blk.38.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 516/ 851] blk.38.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 517/ 851] blk.38.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 518/ 851] blk.38.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 519/ 851] blk.38.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 520/ 851] blk.38.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 521/ 851] blk.38.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 522/ 851] blk.38.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 523/ 851] blk.39.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 524/ 851] blk.39.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 525/ 851] blk.39.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 526/ 851] blk.39.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 527/ 851] blk.39.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 528/ 851] blk.39.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 529/ 851] blk.39.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 530/ 851] blk.39.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 531/ 851] blk.39.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 532/ 851] blk.39.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 533/ 851] blk.39.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 534/ 851] blk.40.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 535/ 851] blk.40.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 536/ 851] blk.40.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 537/ 851] blk.40.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 538/ 851] blk.40.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 539/ 851] blk.40.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 540/ 851] blk.40.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 541/ 851] blk.40.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 542/ 851] blk.40.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 543/ 851] blk.40.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 544/ 851] blk.40.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 545/ 851] blk.40.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 546/ 851] blk.40.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 547/ 851] blk.40.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 548/ 851] blk.41.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 549/ 851] blk.41.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 550/ 851] blk.41.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 551/ 851] blk.41.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 552/ 851] blk.41.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 553/ 851] blk.41.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 554/ 851] blk.41.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 555/ 851] blk.41.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 556/ 851] blk.41.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 557/ 851] blk.41.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 558/ 851] blk.41.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 559/ 851] blk.41.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 560/ 851] blk.41.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 561/ 851] blk.41.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 562/ 851] blk.42.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 563/ 851] blk.42.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 564/ 851] blk.42.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 565/ 851] blk.42.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 566/ 851] blk.42.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 567/ 851] blk.42.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 568/ 851] blk.42.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 569/ 851] blk.42.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 570/ 851] blk.42.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 571/ 851] blk.42.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 572/ 851] blk.42.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 573/ 851] blk.42.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 574/ 851] blk.42.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 575/ 851] blk.42.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 576/ 851] blk.43.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 577/ 851] blk.43.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 578/ 851] blk.43.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 579/ 851] blk.43.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 580/ 851] blk.43.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 581/ 851] blk.43.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 582/ 851] blk.43.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 583/ 851] blk.43.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 584/ 851] blk.43.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 585/ 851] blk.43.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 586/ 851] blk.43.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 587/ 851] blk.44.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 588/ 851] blk.44.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 589/ 851] blk.44.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 590/ 851] blk.44.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 591/ 851] blk.44.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 592/ 851] blk.44.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 593/ 851] blk.44.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 594/ 851] blk.44.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 595/ 851] blk.44.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 596/ 851] blk.44.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 597/ 851] blk.44.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 598/ 851] blk.44.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 599/ 851] blk.44.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 600/ 851] blk.44.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 601/ 851] blk.45.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 602/ 851] blk.45.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 603/ 851] blk.45.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 604/ 851] blk.45.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 605/ 851] blk.45.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 606/ 851] blk.45.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 607/ 851] blk.45.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 608/ 851] blk.45.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 609/ 851] blk.45.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 610/ 851] blk.45.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 611/ 851] blk.45.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 612/ 851] blk.45.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 613/ 851] blk.45.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 614/ 851] blk.45.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 615/ 851] blk.46.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 616/ 851] blk.46.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 617/ 851] blk.46.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 618/ 851] blk.46.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 619/ 851] blk.46.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 620/ 851] blk.46.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 621/ 851] blk.46.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 622/ 851] blk.46.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 623/ 851] blk.46.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 624/ 851] blk.46.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 625/ 851] blk.46.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 626/ 851] blk.46.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 627/ 851] blk.46.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 628/ 851] blk.46.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 629/ 851] blk.47.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 630/ 851] blk.47.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 631/ 851] blk.47.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 632/ 851] blk.47.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 633/ 851] blk.47.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 634/ 851] blk.47.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 635/ 851] blk.47.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 636/ 851] blk.47.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 637/ 851] blk.47.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 638/ 851] blk.47.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 639/ 851] blk.47.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 640/ 851] blk.48.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 641/ 851] blk.48.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 642/ 851] blk.48.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 643/ 851] blk.48.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 644/ 851] blk.48.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 645/ 851] blk.48.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 646/ 851] blk.48.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 647/ 851] blk.48.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 648/ 851] blk.48.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 649/ 851] blk.48.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 650/ 851] blk.48.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 651/ 851] blk.48.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 652/ 851] blk.48.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 653/ 851] blk.48.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 654/ 851] blk.49.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 655/ 851] blk.49.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 656/ 851] blk.49.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 657/ 851] blk.49.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 658/ 851] blk.49.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 659/ 851] blk.49.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 660/ 851] blk.49.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 661/ 851] blk.49.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 662/ 851] blk.49.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 663/ 851] blk.49.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 664/ 851] blk.49.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 665/ 851] blk.49.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 666/ 851] blk.49.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 667/ 851] blk.49.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 668/ 851] blk.50.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 669/ 851] blk.50.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 670/ 851] blk.50.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 671/ 851] blk.50.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 672/ 851] blk.50.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 673/ 851] blk.50.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 674/ 851] blk.50.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 675/ 851] blk.50.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 676/ 851] blk.50.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 677/ 851] blk.50.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 678/ 851] blk.50.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 679/ 851] blk.50.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 680/ 851] blk.50.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 681/ 851] blk.50.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 682/ 851] blk.51.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 683/ 851] blk.51.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 684/ 851] blk.51.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 685/ 851] blk.51.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 686/ 851] blk.51.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 687/ 851] blk.51.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 688/ 851] blk.51.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 689/ 851] blk.51.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 690/ 851] blk.51.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 691/ 851] blk.51.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 692/ 851] blk.51.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 693/ 851] blk.52.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 694/ 851] blk.52.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 695/ 851] blk.52.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 696/ 851] blk.52.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 697/ 851] blk.52.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 698/ 851] blk.52.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 699/ 851] blk.52.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 700/ 851] blk.52.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 701/ 851] blk.52.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 702/ 851] blk.52.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 703/ 851] blk.52.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 704/ 851] blk.52.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 705/ 851] blk.52.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 706/ 851] blk.52.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 707/ 851] blk.53.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 708/ 851] blk.53.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 709/ 851] blk.53.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 710/ 851] blk.53.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 711/ 851] blk.53.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 712/ 851] blk.53.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 713/ 851] blk.53.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 714/ 851] blk.53.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 715/ 851] blk.53.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 716/ 851] blk.53.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 717/ 851] blk.53.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 718/ 851] blk.53.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 719/ 851] blk.53.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 720/ 851] blk.53.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 721/ 851] blk.54.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 722/ 851] blk.54.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 723/ 851] blk.54.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 724/ 851] blk.54.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 725/ 851] blk.54.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 726/ 851] blk.54.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 727/ 851] blk.54.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 728/ 851] blk.54.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 729/ 851] blk.54.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 730/ 851] blk.54.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 731/ 851] blk.54.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 732/ 851] blk.54.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 733/ 851] blk.54.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 734/ 851] blk.54.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 735/ 851] blk.55.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 736/ 851] blk.55.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 737/ 851] blk.55.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 738/ 851] blk.55.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 739/ 851] blk.55.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 740/ 851] blk.55.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 741/ 851] blk.55.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 742/ 851] blk.55.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 743/ 851] blk.55.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 744/ 851] blk.55.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 745/ 851] blk.55.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 746/ 851] blk.56.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 747/ 851] blk.56.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 748/ 851] blk.56.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 749/ 851] blk.56.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 750/ 851] blk.56.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 751/ 851] blk.56.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 752/ 851] blk.56.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 753/ 851] blk.56.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 754/ 851] blk.56.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 755/ 851] blk.56.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 756/ 851] blk.56.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 757/ 851] blk.56.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 758/ 851] blk.56.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 759/ 851] blk.56.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 760/ 851] blk.57.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 761/ 851] blk.57.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 762/ 851] blk.57.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 763/ 851] blk.57.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 764/ 851] blk.57.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 765/ 851] blk.57.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 766/ 851] blk.57.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 767/ 851] blk.57.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 768/ 851] blk.57.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 769/ 851] blk.57.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 770/ 851] blk.57.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 771/ 851] blk.57.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 772/ 851] blk.57.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 773/ 851] blk.57.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 774/ 851] blk.58.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 775/ 851] blk.58.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 776/ 851] blk.58.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 777/ 851] blk.58.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 778/ 851] blk.58.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 779/ 851] blk.58.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 780/ 851] blk.58.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 781/ 851] blk.58.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 782/ 851] blk.58.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 783/ 851] blk.58.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 784/ 851] blk.58.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 785/ 851] blk.58.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 786/ 851] blk.58.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 787/ 851] blk.58.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 788/ 851] blk.59.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 789/ 851] blk.59.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 790/ 851] blk.59.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 791/ 851] blk.59.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 792/ 851] blk.59.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 793/ 851] blk.59.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 794/ 851] blk.59.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 795/ 851] blk.59.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 796/ 851] blk.59.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 797/ 851] blk.59.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 798/ 851] blk.59.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 799/ 851] blk.60.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 800/ 851] blk.60.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 801/ 851] blk.60.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 802/ 851] blk.60.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 803/ 851] blk.60.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 804/ 851] blk.60.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 805/ 851] blk.60.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 806/ 851] blk.60.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 807/ 851] blk.60.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 808/ 851] blk.60.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 809/ 851] blk.60.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 810/ 851] blk.60.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 811/ 851] blk.60.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 812/ 851] blk.60.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 813/ 851] blk.61.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 814/ 851] blk.61.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 815/ 851] blk.61.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 816/ 851] blk.61.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 817/ 851] blk.61.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 818/ 851] blk.61.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 819/ 851] blk.61.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 820/ 851] blk.61.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 821/ 851] blk.61.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 822/ 851] blk.61.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 823/ 851] blk.61.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 824/ 851] blk.61.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 825/ 851] blk.61.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 826/ 851] blk.61.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 827/ 851] blk.62.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 828/ 851] blk.62.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 829/ 851] blk.62.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 830/ 851] blk.62.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 831/ 851] blk.62.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 832/ 851] blk.62.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 833/ 851] blk.62.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 834/ 851] blk.62.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 835/ 851] blk.62.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 836/ 851] blk.62.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 837/ 851] blk.62.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 838/ 851] blk.62.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 839/ 851] blk.62.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 840/ 851] blk.62.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 841/ 851] blk.63.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 842/ 851] blk.63.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 843/ 851] blk.63.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 844/ 851] blk.63.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 845/ 851] blk.63.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 846/ 851] blk.63.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 847/ 851] blk.63.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 848/ 851] blk.63.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 849/ 851] blk.63.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 850/ 851] blk.63.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB
[ 851/ 851] blk.63.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
llama_model_quantize_impl: model size = 51305.09 MiB (16.00 BPW)
llama_model_quantize_impl: quant size = 27260.56 MiB (8.50 BPW)
llama_quantize: quantize time = 22409.99 ms
llama_quantize: total time = 22409.99 ms
+ llama.cpp/build/bin/llama-quantize --pure --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ssm_=q8_0 ./upload-OpenJev/OpenJev-BF16.gguf ./upload-OpenJev/OpenJev-Q4_K_M.gguf Q4_K_M
version: 0.5.0-dev (build 11354, commit 54a0c5da9)
built with GNU 14.2.0 for Linux x86_64
llama_quantize: quantizing './upload-OpenJev/OpenJev-BF16.gguf' to './upload-OpenJev/OpenJev-Q4_K_M.gguf' as Q4_K_M
llama_model_loader: loaded meta data with 48 key-value pairs and 851 tensors from ./upload-OpenJev/OpenJev-BF16.gguf (version GGUF V3 (latest))
llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
llama_model_loader: - kv 0: general.architecture str = qwen35
llama_model_loader: - kv 1: general.type str = model
llama_model_loader: - kv 2: general.sampling.top_k i32 = 20
llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.950000
llama_model_loader: - kv 4: general.sampling.temp f32 = 1.000000
llama_model_loader: - kv 5: general.name str = OpenJev
llama_model_loader: - kv 6: general.size_label str = 27B
llama_model_loader: - kv 7: general.license str = cc-by-nc-4.0
llama_model_loader: - kv 8: general.tags arr[str,8] = ["decision-model", "zero-shot-classif...
llama_model_loader: - kv 9: general.languages arr[str,6] = ["en", "de", "fr", "hi", "zh", "ja"]
llama_model_loader: - kv 10: qwen35.block_count u32 = 64
llama_model_loader: - kv 11: qwen35.context_length u32 = 262144
llama_model_loader: - kv 12: qwen35.embedding_length u32 = 5120
llama_model_loader: - kv 13: qwen35.feed_forward_length u32 = 17408
llama_model_loader: - kv 14: qwen35.attention.head_count u32 = 24
llama_model_loader: - kv 15: qwen35.attention.head_count_kv u32 = 4
llama_model_loader: - kv 16: qwen35.rope.dimension_sections arr[i32,4] = [11, 11, 10, 0]
llama_model_loader: - kv 17: qwen35.rope.freq_base f32 = 10000000.000000
llama_model_loader: - kv 18: qwen35.attention.layer_norm_rms_epsilon f32 = 0.000001
llama_model_loader: - kv 19: qwen35.attention.key_length u32 = 256
llama_model_loader: - kv 20: qwen35.attention.value_length u32 = 256
llama_model_loader: - kv 21: general.file_type u32 = 32
llama_model_loader: - kv 22: qwen35.ssm.conv_kernel u32 = 4
llama_model_loader: - kv 23: qwen35.ssm.state_size u32 = 128
llama_model_loader: - kv 24: qwen35.ssm.group_count u32 = 16
llama_model_loader: - kv 25: qwen35.ssm.time_step_rank u32 = 48
llama_model_loader: - kv 26: qwen35.ssm.inner_size u32 = 6144
llama_model_loader: - kv 27: qwen35.attention.recurrent_layers arr[bool,64] = [true, true, true, false, true, true,...
llama_model_loader: - kv 28: qwen35.full_attention_interval u32 = 4
llama_model_loader: - kv 29: qwen35.rope.dimension_count u32 = 64
llama_model_loader: - kv 30: qwen35.decision.type str = openjev
llama_model_loader: - kv 31: qwen35.decision.temperature.choice f32 = 0.850000
llama_model_loader: - kv 32: qwen35.decision.temperature.score f32 = 0.850000
llama_model_loader: - kv 33: qwen35.decision.temperature.noul f32 = 1.554713
llama_model_loader: - kv 34: general.quantization_version u32 = 2
llama_model_loader: - kv 35: tokenizer.ggml.model str = gpt2
llama_model_loader: - kv 36: tokenizer.ggml.pre str = qwen35
llama_model_loader: - kv 37: tokenizer.ggml.tokens arr[str,248320] = ["!", "\"", "#", "$", "%", "&", "'", ...
llama_model_loader: - kv 38: tokenizer.ggml.token_type arr[i32,248320] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
llama_model_loader: - kv 39: tokenizer.ggml.merges arr[str,247587] = ["Δ  Δ ", "Δ Δ  Δ Δ ", "i n", "Δ  t",...
llama_model_loader: - kv 40: tokenizer.ggml.eos_token_id u32 = 248046
llama_model_loader: - kv 41: tokenizer.ggml.padding_token_id u32 = 248044
llama_model_loader: - kv 42: tokenizer.ggml.bos_token_id u32 = 248044
llama_model_loader: - kv 43: tokenizer.ggml.add_bos_token bool = false
llama_model_loader: - kv 44: tokenizer.ggml.add_eos_token bool = false
llama_model_loader: - kv 45: tokenizer.chat_template str = {%- set image_count = namespace(value...
llama_model_loader: - kv 46: tokenizer.chat_template.systemone str = {% set letters = 'ABCDEFGHIJKLMNOPQRS...
llama_model_loader: - kv 47: tokenizer.chat_templates arr[str,1] = ["systemone"]
llama_model_loader: - type f32: 353 tensors
llama_model_loader: - type bf16: 498 tensors
llama_tensor_get_type: output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.0.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.0.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.0.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.0.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.0.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.1.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.1.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.1.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.1.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.1.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.2.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.2.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.2.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.2.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.2.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.3.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.3.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.3.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.3.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.4.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.4.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.4.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.4.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.4.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.5.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.5.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.5.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.5.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.5.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.6.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.6.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.6.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.6.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.6.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.7.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.7.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.7.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.7.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.8.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.8.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.8.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.8.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.8.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.9.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.9.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.9.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.9.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.9.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.10.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.10.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.10.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.10.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.10.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.11.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.11.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.11.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.11.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.12.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.12.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.12.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.12.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.12.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.13.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.13.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.13.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.13.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.13.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.14.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.14.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.14.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.14.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.14.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.15.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.15.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.15.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.15.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.16.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.16.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.16.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.16.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.16.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.17.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.17.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.17.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.17.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.17.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.18.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.18.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.18.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.18.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.18.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.19.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.19.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.19.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.19.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.20.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.20.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.20.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.20.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.20.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.21.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.21.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.21.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.21.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.21.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.22.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.22.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.22.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.22.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.22.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.23.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.23.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.23.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.23.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.24.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.24.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.24.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.24.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.24.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.25.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.25.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.25.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.25.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.25.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.26.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.26.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.26.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.26.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.26.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.27.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.27.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.27.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.27.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.28.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.28.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.28.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.28.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.28.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.29.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.29.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.29.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.29.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.29.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.30.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.30.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.30.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.30.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.30.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.31.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.31.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.31.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.31.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.32.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.32.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.32.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.32.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.32.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.33.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.33.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.33.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.33.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.33.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.34.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.34.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.34.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.34.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.34.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.35.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.35.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.35.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.35.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.36.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.36.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.36.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.36.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.36.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.37.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.37.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.37.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.37.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.37.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.38.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.38.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.38.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.38.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.38.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.39.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.39.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.39.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.39.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.40.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.40.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.40.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.40.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.40.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.41.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.41.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.41.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.41.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.41.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.42.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.42.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.42.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.42.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.42.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.43.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.43.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.43.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.43.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.44.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.44.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.44.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.44.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.44.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.45.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.45.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.45.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.45.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.45.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.46.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.46.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.46.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.46.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.46.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.47.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.47.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.47.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.47.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.48.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.48.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.48.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.48.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.48.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.49.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.49.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.49.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.49.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.49.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.50.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.50.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.50.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.50.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.50.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.51.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.51.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.51.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.51.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.52.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.52.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.52.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.52.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.52.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.53.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.53.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.53.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.53.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.53.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.54.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.54.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.54.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.54.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.54.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.55.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.55.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.55.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.55.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.56.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.56.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.56.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.56.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.56.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.57.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.57.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.57.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.57.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.57.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.58.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.58.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.58.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.58.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.58.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.59.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.59.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.59.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.59.attn_v.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.60.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.60.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.60.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.60.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.60.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.61.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.61.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.61.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.61.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.61.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.62.attn_gate.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.62.attn_qkv.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.62.ssm_alpha.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.62.ssm_beta.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.62.ssm_out.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.63.attn_k.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.63.attn_output.weight - applying manual override: q4_K -> q6_K
llama_tensor_get_type: blk.63.attn_q.weight - applying manual override: q4_K -> q8_0
llama_tensor_get_type: blk.63.attn_v.weight - applying manual override: q4_K -> q8_0
[ 1/ 851] output.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q6_K .. size = 2425.00 MiB -> 994.63 MiB
[ 2/ 851] output_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 3/ 851] token_embd.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q4_K .. size = 2425.00 MiB -> 682.03 MiB
[ 4/ 851] blk.0.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 5/ 851] blk.0.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 6/ 851] blk.0.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 7/ 851] blk.0.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 8/ 851] blk.0.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 9/ 851] blk.0.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 10/ 851] blk.0.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 11/ 851] blk.0.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 12/ 851] blk.0.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 13/ 851] blk.0.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 14/ 851] blk.0.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 15/ 851] blk.0.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 16/ 851] blk.0.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 17/ 851] blk.0.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 18/ 851] blk.1.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 19/ 851] blk.1.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 20/ 851] blk.1.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 21/ 851] blk.1.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 22/ 851] blk.1.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 23/ 851] blk.1.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 24/ 851] blk.1.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 25/ 851] blk.1.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 26/ 851] blk.1.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 27/ 851] blk.1.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 28/ 851] blk.1.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 29/ 851] blk.1.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 30/ 851] blk.1.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 31/ 851] blk.1.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 32/ 851] blk.2.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 33/ 851] blk.2.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 34/ 851] blk.2.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 35/ 851] blk.2.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 36/ 851] blk.2.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 37/ 851] blk.2.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 38/ 851] blk.2.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 39/ 851] blk.2.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 40/ 851] blk.2.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 41/ 851] blk.2.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 42/ 851] blk.2.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 43/ 851] blk.2.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 44/ 851] blk.2.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 45/ 851] blk.2.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 46/ 851] blk.3.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 47/ 851] blk.3.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 48/ 851] blk.3.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 49/ 851] blk.3.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 50/ 851] blk.3.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 51/ 851] blk.3.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 52/ 851] blk.3.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 53/ 851] blk.3.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 54/ 851] blk.3.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 55/ 851] blk.3.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 56/ 851] blk.3.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 57/ 851] blk.4.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 58/ 851] blk.4.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 59/ 851] blk.4.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 60/ 851] blk.4.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 61/ 851] blk.4.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 62/ 851] blk.4.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 63/ 851] blk.4.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 64/ 851] blk.4.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 65/ 851] blk.4.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 66/ 851] blk.4.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 67/ 851] blk.4.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 68/ 851] blk.4.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 69/ 851] blk.4.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 70/ 851] blk.4.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 71/ 851] blk.5.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 72/ 851] blk.5.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 73/ 851] blk.5.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 74/ 851] blk.5.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 75/ 851] blk.5.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 76/ 851] blk.5.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 77/ 851] blk.5.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 78/ 851] blk.5.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 79/ 851] blk.5.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 80/ 851] blk.5.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 81/ 851] blk.5.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 82/ 851] blk.5.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 83/ 851] blk.5.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 84/ 851] blk.5.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 85/ 851] blk.6.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 86/ 851] blk.6.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 87/ 851] blk.6.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 88/ 851] blk.6.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 89/ 851] blk.6.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 90/ 851] blk.6.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 91/ 851] blk.6.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 92/ 851] blk.6.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 93/ 851] blk.6.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 94/ 851] blk.6.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 95/ 851] blk.6.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 96/ 851] blk.6.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 97/ 851] blk.6.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 98/ 851] blk.6.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 99/ 851] blk.7.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 100/ 851] blk.7.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 101/ 851] blk.7.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 102/ 851] blk.7.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 103/ 851] blk.7.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 104/ 851] blk.7.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 105/ 851] blk.7.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 106/ 851] blk.7.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 107/ 851] blk.7.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 108/ 851] blk.7.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 109/ 851] blk.7.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 110/ 851] blk.8.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 111/ 851] blk.8.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 112/ 851] blk.8.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 113/ 851] blk.8.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 114/ 851] blk.8.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 115/ 851] blk.8.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 116/ 851] blk.8.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 117/ 851] blk.8.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 118/ 851] blk.8.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 119/ 851] blk.8.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 120/ 851] blk.8.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 121/ 851] blk.8.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 122/ 851] blk.8.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 123/ 851] blk.8.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 124/ 851] blk.9.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 125/ 851] blk.9.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 126/ 851] blk.9.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 127/ 851] blk.9.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 128/ 851] blk.9.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 129/ 851] blk.9.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 130/ 851] blk.9.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 131/ 851] blk.9.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 132/ 851] blk.9.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 133/ 851] blk.9.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 134/ 851] blk.9.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 135/ 851] blk.9.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 136/ 851] blk.9.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 137/ 851] blk.9.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 138/ 851] blk.10.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 139/ 851] blk.10.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 140/ 851] blk.10.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 141/ 851] blk.10.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 142/ 851] blk.10.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 143/ 851] blk.10.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 144/ 851] blk.10.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 145/ 851] blk.10.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 146/ 851] blk.10.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 147/ 851] blk.10.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 148/ 851] blk.10.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 149/ 851] blk.10.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 150/ 851] blk.10.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 151/ 851] blk.10.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 152/ 851] blk.11.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 153/ 851] blk.11.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 154/ 851] blk.11.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 155/ 851] blk.11.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 156/ 851] blk.11.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 157/ 851] blk.11.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 158/ 851] blk.11.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 159/ 851] blk.11.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 160/ 851] blk.11.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 161/ 851] blk.11.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 162/ 851] blk.11.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 163/ 851] blk.12.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 164/ 851] blk.12.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 165/ 851] blk.12.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 166/ 851] blk.12.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 167/ 851] blk.12.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 168/ 851] blk.12.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 169/ 851] blk.12.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 170/ 851] blk.12.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 171/ 851] blk.12.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 172/ 851] blk.12.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 173/ 851] blk.12.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 174/ 851] blk.12.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 175/ 851] blk.12.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 176/ 851] blk.12.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 177/ 851] blk.13.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 178/ 851] blk.13.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 179/ 851] blk.13.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 180/ 851] blk.13.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 181/ 851] blk.13.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 182/ 851] blk.13.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 183/ 851] blk.13.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 184/ 851] blk.13.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 185/ 851] blk.13.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 186/ 851] blk.13.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 187/ 851] blk.13.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 188/ 851] blk.13.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 189/ 851] blk.13.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 190/ 851] blk.13.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 191/ 851] blk.14.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 192/ 851] blk.14.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 193/ 851] blk.14.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 194/ 851] blk.14.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 195/ 851] blk.14.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 196/ 851] blk.14.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 197/ 851] blk.14.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 198/ 851] blk.14.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 199/ 851] blk.14.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 200/ 851] blk.14.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 201/ 851] blk.14.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 202/ 851] blk.14.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 203/ 851] blk.14.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 204/ 851] blk.14.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 205/ 851] blk.15.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 206/ 851] blk.15.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 207/ 851] blk.15.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 208/ 851] blk.15.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 209/ 851] blk.15.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 210/ 851] blk.15.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 211/ 851] blk.15.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 212/ 851] blk.15.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 213/ 851] blk.15.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 214/ 851] blk.15.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 215/ 851] blk.15.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 216/ 851] blk.16.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 217/ 851] blk.16.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 218/ 851] blk.16.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 219/ 851] blk.16.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 220/ 851] blk.16.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 221/ 851] blk.16.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 222/ 851] blk.16.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 223/ 851] blk.16.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 224/ 851] blk.16.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 225/ 851] blk.16.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 226/ 851] blk.16.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 227/ 851] blk.16.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 228/ 851] blk.16.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 229/ 851] blk.16.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 230/ 851] blk.17.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 231/ 851] blk.17.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 232/ 851] blk.17.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 233/ 851] blk.17.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 234/ 851] blk.17.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 235/ 851] blk.17.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 236/ 851] blk.17.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 237/ 851] blk.17.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 238/ 851] blk.17.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 239/ 851] blk.17.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 240/ 851] blk.17.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 241/ 851] blk.17.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 242/ 851] blk.17.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 243/ 851] blk.17.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 244/ 851] blk.18.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 245/ 851] blk.18.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 246/ 851] blk.18.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 247/ 851] blk.18.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 248/ 851] blk.18.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 249/ 851] blk.18.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 250/ 851] blk.18.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 251/ 851] blk.18.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 252/ 851] blk.18.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 253/ 851] blk.18.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 254/ 851] blk.18.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 255/ 851] blk.18.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 256/ 851] blk.18.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 257/ 851] blk.18.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 258/ 851] blk.19.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 259/ 851] blk.19.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 260/ 851] blk.19.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 261/ 851] blk.19.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 262/ 851] blk.19.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 263/ 851] blk.19.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 264/ 851] blk.19.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 265/ 851] blk.19.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 266/ 851] blk.19.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 267/ 851] blk.19.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 268/ 851] blk.19.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 269/ 851] blk.20.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 270/ 851] blk.20.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 271/ 851] blk.20.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 272/ 851] blk.20.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 273/ 851] blk.20.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 274/ 851] blk.20.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 275/ 851] blk.20.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 276/ 851] blk.20.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 277/ 851] blk.20.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 278/ 851] blk.20.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 279/ 851] blk.20.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 280/ 851] blk.20.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 281/ 851] blk.20.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 282/ 851] blk.20.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 283/ 851] blk.21.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 284/ 851] blk.21.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 285/ 851] blk.21.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 286/ 851] blk.21.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 287/ 851] blk.21.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 288/ 851] blk.21.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 289/ 851] blk.21.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 290/ 851] blk.21.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 291/ 851] blk.21.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 292/ 851] blk.21.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 293/ 851] blk.21.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 294/ 851] blk.21.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 295/ 851] blk.21.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 296/ 851] blk.21.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 297/ 851] blk.22.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 298/ 851] blk.22.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 299/ 851] blk.22.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 300/ 851] blk.22.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 301/ 851] blk.22.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 302/ 851] blk.22.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 303/ 851] blk.22.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 304/ 851] blk.22.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 305/ 851] blk.22.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 306/ 851] blk.22.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 307/ 851] blk.22.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 308/ 851] blk.22.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 309/ 851] blk.22.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 310/ 851] blk.22.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 311/ 851] blk.23.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 312/ 851] blk.23.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 313/ 851] blk.23.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 314/ 851] blk.23.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 315/ 851] blk.23.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 316/ 851] blk.23.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 317/ 851] blk.23.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 318/ 851] blk.23.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 319/ 851] blk.23.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 320/ 851] blk.23.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 321/ 851] blk.23.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 322/ 851] blk.24.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 323/ 851] blk.24.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 324/ 851] blk.24.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 325/ 851] blk.24.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 326/ 851] blk.24.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 327/ 851] blk.24.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 328/ 851] blk.24.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 329/ 851] blk.24.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 330/ 851] blk.24.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 331/ 851] blk.24.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 332/ 851] blk.24.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 333/ 851] blk.24.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 334/ 851] blk.24.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 335/ 851] blk.24.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 336/ 851] blk.25.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 337/ 851] blk.25.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 338/ 851] blk.25.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 339/ 851] blk.25.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 340/ 851] blk.25.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 341/ 851] blk.25.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 342/ 851] blk.25.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 343/ 851] blk.25.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 344/ 851] blk.25.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 345/ 851] blk.25.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 346/ 851] blk.25.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 347/ 851] blk.25.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 348/ 851] blk.25.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 349/ 851] blk.25.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 350/ 851] blk.26.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 351/ 851] blk.26.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 352/ 851] blk.26.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 353/ 851] blk.26.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 354/ 851] blk.26.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 355/ 851] blk.26.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 356/ 851] blk.26.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 357/ 851] blk.26.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 358/ 851] blk.26.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 359/ 851] blk.26.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 360/ 851] blk.26.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 361/ 851] blk.26.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 362/ 851] blk.26.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 363/ 851] blk.26.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 364/ 851] blk.27.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 365/ 851] blk.27.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 366/ 851] blk.27.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 367/ 851] blk.27.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 368/ 851] blk.27.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 369/ 851] blk.27.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 370/ 851] blk.27.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 371/ 851] blk.27.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 372/ 851] blk.27.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 373/ 851] blk.27.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 374/ 851] blk.27.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 375/ 851] blk.28.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 376/ 851] blk.28.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 377/ 851] blk.28.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 378/ 851] blk.28.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 379/ 851] blk.28.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 380/ 851] blk.28.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 381/ 851] blk.28.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 382/ 851] blk.28.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 383/ 851] blk.28.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 384/ 851] blk.28.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 385/ 851] blk.28.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 386/ 851] blk.28.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 387/ 851] blk.28.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 388/ 851] blk.28.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 389/ 851] blk.29.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 390/ 851] blk.29.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 391/ 851] blk.29.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 392/ 851] blk.29.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 393/ 851] blk.29.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 394/ 851] blk.29.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 395/ 851] blk.29.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 396/ 851] blk.29.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 397/ 851] blk.29.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 398/ 851] blk.29.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 399/ 851] blk.29.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 400/ 851] blk.29.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 401/ 851] blk.29.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 402/ 851] blk.29.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 403/ 851] blk.30.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 404/ 851] blk.30.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 405/ 851] blk.30.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 406/ 851] blk.30.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 407/ 851] blk.30.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 408/ 851] blk.30.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 409/ 851] blk.30.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 410/ 851] blk.30.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 411/ 851] blk.30.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 412/ 851] blk.30.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 413/ 851] blk.30.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 414/ 851] blk.30.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 415/ 851] blk.30.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 416/ 851] blk.30.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 417/ 851] blk.31.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 418/ 851] blk.31.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 419/ 851] blk.31.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 420/ 851] blk.31.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 421/ 851] blk.31.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 422/ 851] blk.31.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 423/ 851] blk.31.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 424/ 851] blk.31.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 425/ 851] blk.31.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 426/ 851] blk.31.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 427/ 851] blk.31.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 428/ 851] blk.32.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 429/ 851] blk.32.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 430/ 851] blk.32.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 431/ 851] blk.32.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 432/ 851] blk.32.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 433/ 851] blk.32.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 434/ 851] blk.32.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 435/ 851] blk.32.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 436/ 851] blk.32.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 437/ 851] blk.32.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 438/ 851] blk.32.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 439/ 851] blk.32.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 440/ 851] blk.32.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 441/ 851] blk.32.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 442/ 851] blk.33.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 443/ 851] blk.33.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 444/ 851] blk.33.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 445/ 851] blk.33.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 446/ 851] blk.33.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 447/ 851] blk.33.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 448/ 851] blk.33.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 449/ 851] blk.33.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 450/ 851] blk.33.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 451/ 851] blk.33.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 452/ 851] blk.33.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 453/ 851] blk.33.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 454/ 851] blk.33.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 455/ 851] blk.33.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 456/ 851] blk.34.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 457/ 851] blk.34.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 458/ 851] blk.34.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 459/ 851] blk.34.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 460/ 851] blk.34.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 461/ 851] blk.34.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 462/ 851] blk.34.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 463/ 851] blk.34.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 464/ 851] blk.34.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 465/ 851] blk.34.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 466/ 851] blk.34.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 467/ 851] blk.34.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 468/ 851] blk.34.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 469/ 851] blk.34.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 470/ 851] blk.35.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 471/ 851] blk.35.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 472/ 851] blk.35.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 473/ 851] blk.35.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 474/ 851] blk.35.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 475/ 851] blk.35.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 476/ 851] blk.35.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 477/ 851] blk.35.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 478/ 851] blk.35.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 479/ 851] blk.35.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 480/ 851] blk.35.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 481/ 851] blk.36.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 482/ 851] blk.36.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 483/ 851] blk.36.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 484/ 851] blk.36.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 485/ 851] blk.36.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 486/ 851] blk.36.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 487/ 851] blk.36.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 488/ 851] blk.36.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 489/ 851] blk.36.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 490/ 851] blk.36.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 491/ 851] blk.36.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 492/ 851] blk.36.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 493/ 851] blk.36.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 494/ 851] blk.36.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 495/ 851] blk.37.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 496/ 851] blk.37.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 497/ 851] blk.37.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 498/ 851] blk.37.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 499/ 851] blk.37.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 500/ 851] blk.37.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 501/ 851] blk.37.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 502/ 851] blk.37.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 503/ 851] blk.37.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 504/ 851] blk.37.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 505/ 851] blk.37.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 506/ 851] blk.37.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 507/ 851] blk.37.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 508/ 851] blk.37.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 509/ 851] blk.38.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 510/ 851] blk.38.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 511/ 851] blk.38.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 512/ 851] blk.38.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 513/ 851] blk.38.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 514/ 851] blk.38.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 515/ 851] blk.38.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 516/ 851] blk.38.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 517/ 851] blk.38.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 518/ 851] blk.38.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 519/ 851] blk.38.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 520/ 851] blk.38.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 521/ 851] blk.38.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 522/ 851] blk.38.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 523/ 851] blk.39.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 524/ 851] blk.39.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 525/ 851] blk.39.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 526/ 851] blk.39.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 527/ 851] blk.39.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 528/ 851] blk.39.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 529/ 851] blk.39.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 530/ 851] blk.39.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 531/ 851] blk.39.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 532/ 851] blk.39.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 533/ 851] blk.39.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 534/ 851] blk.40.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 535/ 851] blk.40.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 536/ 851] blk.40.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 537/ 851] blk.40.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 538/ 851] blk.40.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 539/ 851] blk.40.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 540/ 851] blk.40.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 541/ 851] blk.40.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 542/ 851] blk.40.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 543/ 851] blk.40.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 544/ 851] blk.40.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 545/ 851] blk.40.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 546/ 851] blk.40.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 547/ 851] blk.40.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 548/ 851] blk.41.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 549/ 851] blk.41.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 550/ 851] blk.41.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 551/ 851] blk.41.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 552/ 851] blk.41.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 553/ 851] blk.41.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 554/ 851] blk.41.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 555/ 851] blk.41.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 556/ 851] blk.41.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 557/ 851] blk.41.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 558/ 851] blk.41.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 559/ 851] blk.41.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 560/ 851] blk.41.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 561/ 851] blk.41.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 562/ 851] blk.42.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 563/ 851] blk.42.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 564/ 851] blk.42.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 565/ 851] blk.42.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 566/ 851] blk.42.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 567/ 851] blk.42.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 568/ 851] blk.42.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 569/ 851] blk.42.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 570/ 851] blk.42.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 571/ 851] blk.42.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 572/ 851] blk.42.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 573/ 851] blk.42.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 574/ 851] blk.42.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 575/ 851] blk.42.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 576/ 851] blk.43.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 577/ 851] blk.43.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 578/ 851] blk.43.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 579/ 851] blk.43.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 580/ 851] blk.43.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 581/ 851] blk.43.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 582/ 851] blk.43.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 583/ 851] blk.43.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 584/ 851] blk.43.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 585/ 851] blk.43.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 586/ 851] blk.43.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 587/ 851] blk.44.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 588/ 851] blk.44.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 589/ 851] blk.44.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 590/ 851] blk.44.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 591/ 851] blk.44.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 592/ 851] blk.44.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 593/ 851] blk.44.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 594/ 851] blk.44.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 595/ 851] blk.44.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 596/ 851] blk.44.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 597/ 851] blk.44.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 598/ 851] blk.44.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 599/ 851] blk.44.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 600/ 851] blk.44.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 601/ 851] blk.45.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 602/ 851] blk.45.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 603/ 851] blk.45.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 604/ 851] blk.45.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 605/ 851] blk.45.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 606/ 851] blk.45.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 607/ 851] blk.45.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 608/ 851] blk.45.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 609/ 851] blk.45.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 610/ 851] blk.45.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 611/ 851] blk.45.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 612/ 851] blk.45.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 613/ 851] blk.45.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 614/ 851] blk.45.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 615/ 851] blk.46.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 616/ 851] blk.46.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 617/ 851] blk.46.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 618/ 851] blk.46.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 619/ 851] blk.46.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 620/ 851] blk.46.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 621/ 851] blk.46.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 622/ 851] blk.46.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 623/ 851] blk.46.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 624/ 851] blk.46.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 625/ 851] blk.46.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 626/ 851] blk.46.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 627/ 851] blk.46.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 628/ 851] blk.46.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 629/ 851] blk.47.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 630/ 851] blk.47.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 631/ 851] blk.47.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 632/ 851] blk.47.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 633/ 851] blk.47.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 634/ 851] blk.47.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 635/ 851] blk.47.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 636/ 851] blk.47.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 637/ 851] blk.47.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 638/ 851] blk.47.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 639/ 851] blk.47.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 640/ 851] blk.48.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 641/ 851] blk.48.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 642/ 851] blk.48.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 643/ 851] blk.48.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 644/ 851] blk.48.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 645/ 851] blk.48.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 646/ 851] blk.48.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 647/ 851] blk.48.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 648/ 851] blk.48.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 649/ 851] blk.48.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 650/ 851] blk.48.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 651/ 851] blk.48.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 652/ 851] blk.48.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 653/ 851] blk.48.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 654/ 851] blk.49.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 655/ 851] blk.49.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 656/ 851] blk.49.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 657/ 851] blk.49.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 658/ 851] blk.49.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 659/ 851] blk.49.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 660/ 851] blk.49.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 661/ 851] blk.49.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 662/ 851] blk.49.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 663/ 851] blk.49.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 664/ 851] blk.49.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 665/ 851] blk.49.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 666/ 851] blk.49.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 667/ 851] blk.49.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 668/ 851] blk.50.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 669/ 851] blk.50.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 670/ 851] blk.50.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 671/ 851] blk.50.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 672/ 851] blk.50.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 673/ 851] blk.50.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 674/ 851] blk.50.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 675/ 851] blk.50.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 676/ 851] blk.50.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 677/ 851] blk.50.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 678/ 851] blk.50.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 679/ 851] blk.50.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 680/ 851] blk.50.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 681/ 851] blk.50.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 682/ 851] blk.51.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 683/ 851] blk.51.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 684/ 851] blk.51.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 685/ 851] blk.51.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 686/ 851] blk.51.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 687/ 851] blk.51.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 688/ 851] blk.51.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 689/ 851] blk.51.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 690/ 851] blk.51.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 691/ 851] blk.51.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 692/ 851] blk.51.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 693/ 851] blk.52.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 694/ 851] blk.52.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 695/ 851] blk.52.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 696/ 851] blk.52.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 697/ 851] blk.52.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 698/ 851] blk.52.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 699/ 851] blk.52.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 700/ 851] blk.52.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 701/ 851] blk.52.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 702/ 851] blk.52.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 703/ 851] blk.52.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 704/ 851] blk.52.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 705/ 851] blk.52.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 706/ 851] blk.52.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 707/ 851] blk.53.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 708/ 851] blk.53.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 709/ 851] blk.53.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 710/ 851] blk.53.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 711/ 851] blk.53.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 712/ 851] blk.53.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 713/ 851] blk.53.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 714/ 851] blk.53.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 715/ 851] blk.53.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 716/ 851] blk.53.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 717/ 851] blk.53.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 718/ 851] blk.53.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 719/ 851] blk.53.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 720/ 851] blk.53.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 721/ 851] blk.54.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 722/ 851] blk.54.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 723/ 851] blk.54.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 724/ 851] blk.54.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 725/ 851] blk.54.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 726/ 851] blk.54.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 727/ 851] blk.54.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 728/ 851] blk.54.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 729/ 851] blk.54.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 730/ 851] blk.54.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 731/ 851] blk.54.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 732/ 851] blk.54.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 733/ 851] blk.54.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 734/ 851] blk.54.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 735/ 851] blk.55.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 736/ 851] blk.55.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 737/ 851] blk.55.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 738/ 851] blk.55.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 739/ 851] blk.55.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 740/ 851] blk.55.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 741/ 851] blk.55.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 742/ 851] blk.55.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 743/ 851] blk.55.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 744/ 851] blk.55.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 745/ 851] blk.55.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 746/ 851] blk.56.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 747/ 851] blk.56.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 748/ 851] blk.56.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 749/ 851] blk.56.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 750/ 851] blk.56.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 751/ 851] blk.56.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 752/ 851] blk.56.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 753/ 851] blk.56.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 754/ 851] blk.56.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 755/ 851] blk.56.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 756/ 851] blk.56.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 757/ 851] blk.56.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 758/ 851] blk.56.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 759/ 851] blk.56.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 760/ 851] blk.57.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 761/ 851] blk.57.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 762/ 851] blk.57.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 763/ 851] blk.57.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 764/ 851] blk.57.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 765/ 851] blk.57.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 766/ 851] blk.57.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 767/ 851] blk.57.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 768/ 851] blk.57.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 769/ 851] blk.57.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 770/ 851] blk.57.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 771/ 851] blk.57.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 772/ 851] blk.57.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 773/ 851] blk.57.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 774/ 851] blk.58.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 775/ 851] blk.58.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 776/ 851] blk.58.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 777/ 851] blk.58.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 778/ 851] blk.58.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 779/ 851] blk.58.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 780/ 851] blk.58.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 781/ 851] blk.58.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 782/ 851] blk.58.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 783/ 851] blk.58.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 784/ 851] blk.58.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 785/ 851] blk.58.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 786/ 851] blk.58.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 787/ 851] blk.58.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 788/ 851] blk.59.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 789/ 851] blk.59.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 790/ 851] blk.59.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 791/ 851] blk.59.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 792/ 851] blk.59.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 793/ 851] blk.59.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 794/ 851] blk.59.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 795/ 851] blk.59.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 796/ 851] blk.59.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 797/ 851] blk.59.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 798/ 851] blk.59.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 799/ 851] blk.60.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 800/ 851] blk.60.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 801/ 851] blk.60.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 802/ 851] blk.60.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 803/ 851] blk.60.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 804/ 851] blk.60.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 805/ 851] blk.60.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 806/ 851] blk.60.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 807/ 851] blk.60.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 808/ 851] blk.60.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 809/ 851] blk.60.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 810/ 851] blk.60.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 811/ 851] blk.60.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 812/ 851] blk.60.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 813/ 851] blk.61.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 814/ 851] blk.61.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 815/ 851] blk.61.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 816/ 851] blk.61.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 817/ 851] blk.61.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 818/ 851] blk.61.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 819/ 851] blk.61.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 820/ 851] blk.61.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 821/ 851] blk.61.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 822/ 851] blk.61.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 823/ 851] blk.61.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 824/ 851] blk.61.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 825/ 851] blk.61.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 826/ 851] blk.61.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 827/ 851] blk.62.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 828/ 851] blk.62.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 829/ 851] blk.62.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB
[ 830/ 851] blk.62.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 831/ 851] blk.62.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 832/ 851] blk.62.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 833/ 851] blk.62.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 834/ 851] blk.62.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 835/ 851] blk.62.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 836/ 851] blk.62.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB
[ 837/ 851] blk.62.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB
[ 838/ 851] blk.62.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB
[ 839/ 851] blk.62.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB
[ 840/ 851] blk.62.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB
[ 841/ 851] blk.63.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 842/ 851] blk.63.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 843/ 851] blk.63.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
[ 844/ 851] blk.63.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB
[ 845/ 851] blk.63.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB
[ 846/ 851] blk.63.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB
[ 847/ 851] blk.63.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB
[ 848/ 851] blk.63.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 849/ 851] blk.63.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 850/ 851] blk.63.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB
[ 851/ 851] blk.63.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB
llama_model_quantize_impl: model size = 51305.09 MiB (16.00 BPW)
llama_model_quantize_impl: quant size = 18084.41 MiB (5.64 BPW)
llama_quantize: quantize time = 93084.12 ms
llama_quantize: total time = 93084.12 ms
+ python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-OpenJev-PRIMARY --outtype q8_0 --outfile ./upload-OpenJev/mmproj-OpenJev-Q8_0.gguf --mmproj --model-name OpenJev
INFO:hf-to-gguf:Loading model: model-temp-OpenJev-PRIMARY
INFO:hf-to-gguf:gguf: detected OpenJev checkpoint
[transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial'
INFO:hf-to-gguf:Model architecture: OpenJevModel
INFO:hf-to-gguf:gguf: detected OpenJev checkpoint
[transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial'
INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json'
INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00012.safetensors'
INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00012.safetensors'
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
INFO:hf-to-gguf:Exporting model...
INFO:hf-to-gguf:v.blk.0.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.0.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.0.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.0.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.0.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.0.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.0.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.0.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.0.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.1.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.1.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.1.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.1.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.1.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.1.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.1.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.1.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.10.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.10.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.10.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.10.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.10.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.10.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.10.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.10.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.11.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.11.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.11.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.11.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.11.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.11.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.11.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.11.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.12.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.12.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.12.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.12.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.12.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.12.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.12.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.12.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.13.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.13.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.13.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.13.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.13.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.13.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.13.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.13.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.14.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.14.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.14.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.14.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.14.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.14.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.14.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.14.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.15.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.15.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.15.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.15.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.15.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.15.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.15.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.15.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.16.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.16.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.16.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.16.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.16.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.16.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.16.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.16.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.17.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.17.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.17.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.17.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.17.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.17.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.17.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.17.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.18.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.18.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.18.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.18.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.18.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.18.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.18.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.18.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.19.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.19.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.19.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.19.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.19.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.19.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.19.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.19.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.2.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.2.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.2.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.2.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.2.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.2.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.2.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.2.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.20.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.20.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.20.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.20.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.20.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.20.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.20.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.20.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.21.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.21.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.21.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.21.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.21.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.21.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.21.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.21.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.22.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.22.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.22.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.22.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.22.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.22.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.22.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.22.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.23.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.23.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.23.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.23.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.23.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.23.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.23.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.23.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.24.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.24.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.24.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.24.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.24.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.24.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.24.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.24.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.25.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.25.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.25.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.25.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.25.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.25.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.25.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.25.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.26.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.26.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.26.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.26.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.26.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.26.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.26.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.26.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.3.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.3.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.3.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.3.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.3.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.3.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.3.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.3.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.4.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.4.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.4.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.4.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.4.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.4.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.4.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.4.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.5.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.5.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.5.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.5.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.5.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.5.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.5.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.5.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.6.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.6.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.6.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.6.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.6.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.6.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.6.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.6.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.7.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.7.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.7.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.7.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.7.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.7.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.7.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.7.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.8.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.8.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.8.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.8.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.8.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.8.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.8.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.8.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.attn_out.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152}
INFO:hf-to-gguf:v.blk.9.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456}
INFO:hf-to-gguf:v.blk.9.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456}
INFO:hf-to-gguf:v.blk.9.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304}
INFO:hf-to-gguf:v.blk.9.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304}
INFO:hf-to-gguf:v.blk.9.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152}
WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16
INFO:hf-to-gguf:v.blk.9.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152}
INFO:hf-to-gguf:v.blk.9.ln1.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.ln1.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.ln2.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.blk.9.ln2.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:mm.0.bias, torch.bfloat16 --> F32, shape = {4608}
INFO:hf-to-gguf:mm.0.weight, torch.bfloat16 --> Q8_0, shape = {4608, 4608}
INFO:hf-to-gguf:mm.2.bias, torch.bfloat16 --> F32, shape = {5120}
INFO:hf-to-gguf:mm.2.weight, torch.bfloat16 --> Q8_0, shape = {4608, 5120}
INFO:hf-to-gguf:v.post_ln.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.post_ln.weight, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.patch_embd.bias, torch.bfloat16 --> F32, shape = {1152}
INFO:hf-to-gguf:v.patch_embd.weight, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152}
INFO:hf-to-gguf:v.patch_embd.weight.1, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152}
INFO:hf-to-gguf:v.position_embd.weight, torch.bfloat16 --> F32, shape = {1152, 2304}
INFO:hf-to-gguf:Set meta model
INFO:hf-to-gguf:Set model parameters
INFO:hf-to-gguf:Set model quantization version
INFO:gguf.gguf_writer:Writing the following files:
INFO:gguf.gguf_writer:upload-OpenJev/mmproj-OpenJev-Q8_0.gguf: n_tensors = 334, total_size = 629.2M
Writing: 0%| | 0.00/629M [00:00<?, ?byte/s]
Writing: 42%|β–ˆβ–ˆβ–ˆβ–ˆβ– | 262M/629M [00:01<00:01, 262Mbyte/s]
Writing: 84%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ– | 528M/629M [00:02<00:00, 261Mbyte/s] Writing: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 629M/629M [00:02<00:00, 257Mbyte/s]
INFO:hf-to-gguf:Model successfully exported to upload-OpenJev/mmproj-OpenJev-Q8_0.gguf
+ echo OpenJev-BF16.gguf
+ echo OpenJev-Q8_0.gguf
+ echo OpenJev-Q4_K_M.gguf
+ echo mmproj-OpenJev-BF16.gguf
+ echo mmproj-OpenJev-Q8_0.gguf