Instructions to use ggml-org/OpenJev-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ggml-org/OpenJev-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ggml-org/OpenJev-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ggml-org/OpenJev-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ggml-org/OpenJev-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ggml-org/OpenJev-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ggml-org/OpenJev-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ggml-org/OpenJev-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ggml-org/OpenJev-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ggml-org/OpenJev-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ggml-org/OpenJev-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ggml-org/OpenJev-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ggml-org/OpenJev-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ggml-org/OpenJev-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ggml-org/OpenJev-GGUF:Q4_K_M
- Ollama
How to use ggml-org/OpenJev-GGUF with Ollama:
ollama run hf.co/ggml-org/OpenJev-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use ggml-org/OpenJev-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/OpenJev-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ggml-org/OpenJev-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ggml-org/OpenJev-GGUF with Docker Model Runner:
docker model run hf.co/ggml-org/OpenJev-GGUF:Q4_K_M
- Lemonade
How to use ggml-org/OpenJev-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ggml-org/OpenJev-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.OpenJev-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ggml-org/OpenJev-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/OpenJev-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ggml-org/OpenJev-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ggml-org/OpenJev-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/OpenJev-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ggml-org/OpenJev-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download convert.log from ggml-org/OpenJev-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 458 kB
-
https://huggingface.co/ggml-org/OpenJev-GGUF/resolve/main/convert.log
- Command line
-
hf download hf://ggml-org/OpenJev-GGUF/convert.log
-
curl -L -o convert.log https://huggingface.co/ggml-org/OpenJev-GGUF/resolve/main/convert.log
458 kB
| + OUTPUT_DIR=./upload-OpenJev | |
| + LLAMA_CPP=llama.cpp | |
| + DISPLAY_NAME=OpenJev | |
| + QUANTIZE=llama.cpp/build/bin/llama-quantize | |
| + python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-OpenJev-PRIMARY --outtype bf16 --outfile ./upload-OpenJev/OpenJev-BF16.gguf --no-mtp --model-name OpenJev | |
| INFO:hf-to-gguf:Loading model: model-temp-OpenJev-PRIMARY | |
| INFO:hf-to-gguf:gguf: detected OpenJev checkpoint | |
| [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' | |
| INFO:hf-to-gguf:Model architecture: OpenJevModel | |
| INFO:hf-to-gguf:gguf: detected OpenJev checkpoint | |
| [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' | |
| INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00012.safetensors' | |
| INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only | |
| INFO:hf-to-gguf:Exporting model... | |
| INFO:hf-to-gguf:output.weight, torch.bfloat16 --> BF16, shape = {5120, 248320} | |
| INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> BF16, shape = {5120, 248320} | |
| INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.0.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.0.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.0.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.0.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.0.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.0.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.0.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.0.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.0.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.1.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.1.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.1.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.1.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.1.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.1.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.1.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.1.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.1.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.2.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.2.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.2.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.2.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.2.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.2.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.2.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.2.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.2.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.3.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.3.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.3.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.3.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.3.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.3.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.3.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.3.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.3.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.4.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.4.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.4.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.4.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.4.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.4.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.4.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.4.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.4.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.4.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.4.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.5.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.5.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.5.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.5.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.5.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.5.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.5.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.5.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.5.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.5.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.5.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.6.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.6.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.6.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.6.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.6.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.6.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.6.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.6.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.6.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.6.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.6.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.7.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.7.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.7.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.7.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.7.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.7.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.7.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.7.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.7.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.8.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.8.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.8.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.8.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.8.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.8.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.8.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.8.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.8.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.8.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.8.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.9.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.9.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.9.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.9.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.9.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.9.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.9.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.9.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.9.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.10.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.10.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.10.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.10.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.10.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.10.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.10.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.10.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.10.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.11.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.11.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.11.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.12.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.12.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.12.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.12.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.12.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.12.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.12.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.12.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.12.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.13.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.13.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.13.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.13.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.13.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.13.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.13.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.13.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.13.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.14.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.14.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.14.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.14.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.14.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.14.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.14.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.14.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.14.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.15.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.15.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.15.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.16.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.16.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.16.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.16.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.16.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.9.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.9.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.16.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.16.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.16.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.16.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.17.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.17.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.17.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.17.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.17.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.17.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.17.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.17.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.17.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.18.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.18.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.18.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.18.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.18.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.18.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.18.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.18.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.18.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.19.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.19.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.19.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.20.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.20.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.20.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.20.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.20.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.20.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.20.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.20.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.20.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.21.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.21.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.21.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.21.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.21.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.21.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.21.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.21.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.21.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.22.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.22.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.22.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.22.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.22.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.22.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.22.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.22.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.22.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.23.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.23.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.23.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.23.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.23.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.23.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.23.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.23.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.24.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.24.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.24.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.24.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.24.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.24.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.24.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.24.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.24.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.24.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.24.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.25.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.25.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.25.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.25.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.25.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.25.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.25.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.25.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.25.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.25.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.25.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.26.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.26.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.26.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.26.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.26.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.26.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.26.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.26.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.26.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.26.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.26.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.27.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.27.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.27.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.27.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.27.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.27.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.27.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.27.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.27.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.27.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.27.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.28.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.28.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.28.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.28.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.28.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.28.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.28.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.28.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.28.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.28.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.28.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.28.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.28.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.28.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.29.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.29.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.29.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.29.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.29.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.29.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.29.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.29.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.29.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.29.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.29.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.29.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.29.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.29.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.30.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.30.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.30.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.30.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.30.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.30.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.30.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.30.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.30.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.30.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.30.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.30.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.30.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.30.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.31.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.31.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.31.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.31.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.31.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.31.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.31.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.31.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.31.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.31.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.31.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.32.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.32.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.32.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.32.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.32.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.32.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.32.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.32.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.32.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.32.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.32.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.32.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.32.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.32.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.33.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.33.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.33.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.33.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.33.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.33.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.33.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.33.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.33.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.33.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.33.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.33.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.33.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.33.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.34.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.34.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.34.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.34.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.34.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.34.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.34.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.34.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.34.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.34.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.34.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.34.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.34.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.34.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.35.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.35.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.35.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.35.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.35.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.35.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.35.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.35.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.35.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.35.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.35.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.36.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.36.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.36.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.36.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.36.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.36.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.36.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.36.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.36.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.36.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.36.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.36.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.36.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.36.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.37.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.37.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.37.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.37.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.37.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.37.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.37.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.37.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.37.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.37.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.37.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.37.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.37.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.37.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.38.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.38.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.38.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.38.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.38.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.38.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.38.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.38.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.38.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.38.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.38.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.38.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.38.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.38.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.39.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.39.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.39.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.39.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.39.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.39.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.39.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.39.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.39.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.39.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.39.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.40.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.40.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.40.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.40.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.40.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.40.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.40.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.40.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.40.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.40.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.40.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.40.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.40.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.40.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.41.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.41.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.41.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.41.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.41.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.41.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.41.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.41.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.41.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.41.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.41.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.41.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.41.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.41.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.42.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.42.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.42.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.42.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.42.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.42.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.42.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.42.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.42.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.42.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.42.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.42.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.42.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.42.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.43.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.43.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.43.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.43.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.43.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.43.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.43.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.43.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.43.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.43.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.43.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.44.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.44.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.44.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.44.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.44.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.44.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.44.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.44.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.44.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.44.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.44.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.44.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.44.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.44.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.45.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.45.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.45.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.45.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.45.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.45.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.45.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.45.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.45.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.45.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.45.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.45.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.45.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.45.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.46.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.46.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.46.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.46.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.46.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.46.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.46.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.46.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.46.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.46.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.46.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.46.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.46.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.46.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.47.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.47.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.47.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.47.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.47.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.47.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.47.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.47.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.47.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.47.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.47.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.48.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.48.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.48.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.48.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.48.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.48.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.48.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.48.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.48.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.48.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.48.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.48.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.48.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.48.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.49.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.49.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.49.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.49.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.49.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.49.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.49.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.49.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.49.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.49.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.49.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.49.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.49.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.49.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.50.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.50.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.50.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.50.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.50.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.50.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.50.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.50.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.50.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.50.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.50.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.50.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.50.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.50.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.51.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.51.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.51.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.51.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.51.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.51.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.51.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.51.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.51.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.51.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.51.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.52.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.52.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.52.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.52.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.52.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.52.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.52.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.52.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.52.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.52.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.52.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.52.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.52.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.52.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.53.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.53.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.53.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.53.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.53.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.53.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.53.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.53.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.53.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.53.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.53.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.53.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.53.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.53.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.54.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.54.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.54.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.54.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.54.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.54.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.54.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.54.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.54.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.54.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.54.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.54.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.54.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.54.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.55.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.55.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.55.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.55.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.55.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.55.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.55.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.55.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.55.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.55.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.55.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.56.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.56.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.56.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.56.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.56.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.56.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.56.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.56.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.56.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.56.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.56.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.56.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.56.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.56.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.57.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.57.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.57.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.57.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.57.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.57.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.57.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.57.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.57.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.57.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.57.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.57.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.57.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.57.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.58.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.58.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.58.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.58.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.58.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.58.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.58.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.58.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.58.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.58.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.58.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.58.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.58.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.58.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.59.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.59.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.59.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.59.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.59.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.59.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.59.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.59.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.59.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.59.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.59.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.60.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.60.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.60.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.60.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.60.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.60.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.60.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.60.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.60.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.60.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.60.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.60.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.60.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.60.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.61.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.61.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.61.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.61.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.61.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.61.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.61.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.61.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.61.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.61.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.61.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.61.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.61.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.61.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.62.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.62.ssm_a, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.62.ssm_conv1d.weight, torch.bfloat16 --> F32, shape = {4, 10240} | |
| INFO:hf-to-gguf:blk.62.ssm_dt.bias, torch.bfloat16 --> F32, shape = {48} | |
| INFO:hf-to-gguf:blk.62.ssm_alpha.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.62.ssm_beta.weight, torch.bfloat16 --> BF16, shape = {5120, 48} | |
| INFO:hf-to-gguf:blk.62.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {5120, 10240} | |
| INFO:hf-to-gguf:blk.62.attn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 6144} | |
| INFO:hf-to-gguf:blk.62.ssm_norm.weight, torch.bfloat16 --> F32, shape = {128} | |
| INFO:hf-to-gguf:blk.62.ssm_out.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.62.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.62.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.62.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.62.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.63.attn_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.63.ffn_down.weight, torch.bfloat16 --> BF16, shape = {17408, 5120} | |
| INFO:hf-to-gguf:blk.63.ffn_gate.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.63.ffn_up.weight, torch.bfloat16 --> BF16, shape = {5120, 17408} | |
| INFO:hf-to-gguf:blk.63.post_attention_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:blk.63.attn_k_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.63.attn_k.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:blk.63.attn_output.weight, torch.bfloat16 --> BF16, shape = {6144, 5120} | |
| INFO:hf-to-gguf:blk.63.attn_q_norm.weight, torch.bfloat16 --> F32, shape = {256} | |
| INFO:hf-to-gguf:blk.63.attn_q.weight, torch.bfloat16 --> BF16, shape = {5120, 12288} | |
| INFO:hf-to-gguf:blk.63.attn_v.weight, torch.bfloat16 --> BF16, shape = {5120, 1024} | |
| INFO:hf-to-gguf:output_norm.weight, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:Set meta model | |
| INFO:hf-to-gguf:Set model parameters | |
| INFO:hf-to-gguf:gguf: context length = 262144 | |
| INFO:hf-to-gguf:gguf: embedding length = 5120 | |
| INFO:hf-to-gguf:gguf: feed forward length = 17408 | |
| INFO:hf-to-gguf:gguf: head count = 24 | |
| INFO:hf-to-gguf:gguf: key-value head count = 4 | |
| WARNING:hf-to-gguf:Unknown RoPE type: default | |
| INFO:hf-to-gguf:gguf: rope scaling type = NONE | |
| INFO:hf-to-gguf:gguf: mrope sections: [11, 11, 10, 0] | |
| INFO:hf-to-gguf:gguf: rope theta = 10000000 | |
| INFO:hf-to-gguf:gguf: rms norm epsilon = 1e-06 | |
| INFO:hf-to-gguf:gguf: file type = 32 | |
| INFO:hf-to-gguf:Set model quantization version | |
| INFO:hf-to-gguf:Set model tokenizer | |
| [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' | |
| INFO:gguf.vocab:Adding 247587 merge(s). | |
| INFO:gguf.vocab:Setting special token type eos to 248046 | |
| INFO:gguf.vocab:Setting special token type pad to 248044 | |
| INFO:gguf.vocab:Setting special token type bos to 248044 | |
| INFO:gguf.vocab:Setting add_bos_token to False | |
| INFO:gguf.vocab:Setting add_eos_token to False | |
| INFO:gguf.vocab:Setting chat_template to {%- set image_count = namespace(value=0) %} | |
| {%- set video_count = namespace(value=0) %} | |
| {%- macro render_content(content, do_vision_count, is_system_content=false) %} | |
| {%- if content is string %} | |
| {{- content }} | |
| {%- elif content is iterable and content is not mapping %} | |
| {%- for item in content %} | |
| {%- if 'image' in item or 'image_url' in item or item.type == 'image' %} | |
| {%- if is_system_content %} | |
| {{- raise_exception('System message cannot contain images.') }} | |
| {%- endif %} | |
| {%- if do_vision_count %} | |
| {%- set image_count.value = image_count.value + 1 %} | |
| {%- endif %} | |
| {%- if add_vision_id %} | |
| {{- 'Picture ' ~ image_count.value ~ ': ' }} | |
| {%- endif %} | |
| {{- '<|vision_start|><|image_pad|><|vision_end|>' }} | |
| {%- elif 'video' in item or item.type == 'video' %} | |
| {%- if is_system_content %} | |
| {{- raise_exception('System message cannot contain videos.') }} | |
| {%- endif %} | |
| {%- if do_vision_count %} | |
| {%- set video_count.value = video_count.value + 1 %} | |
| {%- endif %} | |
| {%- if add_vision_id %} | |
| {{- 'Video ' ~ video_count.value ~ ': ' }} | |
| {%- endif %} | |
| {{- '<|vision_start|><|video_pad|><|vision_end|>' }} | |
| {%- elif 'text' in item %} | |
| {{- item.text }} | |
| {%- else %} | |
| {{- raise_exception('Unexpected item type in content.') }} | |
| {%- endif %} | |
| {%- endfor %} | |
| {%- elif content is none or content is undefined %} | |
| {{- '' }} | |
| {%- else %} | |
| {{- raise_exception('Unexpected content type.') }} | |
| {%- endif %} | |
| {%- endmacro %} | |
| {%- if not messages %} | |
| {{- raise_exception('No messages provided.') }} | |
| {%- endif %} | |
| {%- set reasoning_instructions = '' %} | |
| {%- if enable_thinking is undefined or enable_thinking is true %} | |
| {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %} | |
| {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %} | |
| {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }} | |
| {%- endif %} | |
| {%- if resolved_reasoning_effort == 'xhigh' %} | |
| {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %} | |
| {%- elif resolved_reasoning_effort == 'low' %} | |
| {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %} | |
| {%- endif %} | |
| {%- endif %} | |
| {%- if tools and tools is iterable and tools is not mapping %} | |
| {{- '<|im_start|>system\n' }} | |
| {%- if reasoning_instructions %} | |
| {{- reasoning_instructions + '\n\n' }} | |
| {%- endif %} | |
| {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }} | |
| {%- for tool in tools %} | |
| {{- "\n" }} | |
| {{- tool | tojson }} | |
| {%- endfor %} | |
| {{- "\n</tools>" }} | |
| {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }} | |
| {%- if messages[0].role == 'system' %} | |
| {%- set content = render_content(messages[0].content, false, true)|trim %} | |
| {%- if content %} | |
| {{- '\n\n' + content }} | |
| {%- endif %} | |
| {%- endif %} | |
| {{- '<|im_end|>\n' }} | |
| {%- else %} | |
| {%- if messages[0].role == 'system' %} | |
| {%- set content = render_content(messages[0].content, false, true)|trim %} | |
| {%- if content %} | |
| {{- '<|im_start|>system\n' + (reasoning_instructions + '\n\n' if reasoning_instructions else '') + content + '<|im_end|>\n' }} | |
| {%- elif reasoning_instructions %} | |
| {{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }} | |
| {%- endif %} | |
| {%- elif reasoning_instructions %} | |
| {{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }} | |
| {%- endif %} | |
| {%- endif %} | |
| {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %} | |
| {%- for message in messages[::-1] %} | |
| {%- set index = (messages|length - 1) - loop.index0 %} | |
| {%- if ns.multi_step_tool and message.role == "user" %} | |
| {%- set content = render_content(message.content, false)|trim %} | |
| {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %} | |
| {%- set ns.multi_step_tool = false %} | |
| {%- set ns.last_query_index = index %} | |
| {%- endif %} | |
| {%- endif %} | |
| {%- endfor %} | |
| {%- if ns.multi_step_tool %} | |
| {{- raise_exception('No user query found in messages.') }} | |
| {%- endif %} | |
| {%- for message in messages %} | |
| {%- set content = render_content(message.content, true)|trim %} | |
| {%- if message.role == "system" %} | |
| {%- if not loop.first %} | |
| {{- raise_exception('System message must be at the beginning.') }} | |
| {%- endif %} | |
| {%- elif message.role == "user" %} | |
| {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }} | |
| {%- elif message.role == "assistant" %} | |
| {%- set reasoning_content = '' %} | |
| {%- if message.reasoning_content is string %} | |
| {%- set reasoning_content = message.reasoning_content %} | |
| {%- endif %} | |
| {%- set reasoning_content = reasoning_content|trim %} | |
| {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %} | |
| {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }} | |
| {%- else %} | |
| {{- '<|im_start|>' + message.role + '\n' + content }} | |
| {%- endif %} | |
| {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %} | |
| {%- for tool_call in message.tool_calls %} | |
| {%- if tool_call.function is defined %} | |
| {%- set tool_call = tool_call.function %} | |
| {%- endif %} | |
| {%- if loop.first %} | |
| {%- if content|trim %} | |
| {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }} | |
| {%- else %} | |
| {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }} | |
| {%- endif %} | |
| {%- else %} | |
| {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }} | |
| {%- endif %} | |
| {%- if tool_call.arguments is defined and tool_call.arguments != '' %} | |
| {%- for args_name, args_value in tool_call.arguments|items %} | |
| {{- '<parameter=' + args_name + '>\n' }} | |
| {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %} | |
| {{- args_value }} | |
| {{- '\n</parameter>\n' }} | |
| {%- endfor %} | |
| {%- endif %} | |
| {{- '</function>\n</tool_call>' }} | |
| {%- endfor %} | |
| {%- endif %} | |
| {{- '<|im_end|>\n' }} | |
| {%- elif message.role == "tool" %} | |
| {%- if loop.previtem and loop.previtem.role != "tool" %} | |
| {{- '<|im_start|>user' }} | |
| {%- endif %} | |
| {{- '\n<tool_response>\n' }} | |
| {{- content }} | |
| {{- '\n</tool_response>' }} | |
| {%- if not loop.last and loop.nextitem.role != "tool" %} | |
| {{- '<|im_end|>\n' }} | |
| {%- elif loop.last %} | |
| {{- '<|im_end|>\n' }} | |
| {%- endif %} | |
| {%- else %} | |
| {{- raise_exception('Unexpected message role.') }} | |
| {%- endif %} | |
| {%- endfor %} | |
| {%- if add_generation_prompt %} | |
| {{- '<|im_start|>assistant\n' }} | |
| {%- if enable_thinking is defined and enable_thinking is false %} | |
| {{- '<think>\n\n</think>\n\n' }} | |
| {%- else %} | |
| {{- '<think>\n' }} | |
| {%- endif %} | |
| {%- endif %} | |
| INFO:gguf.gguf_writer:Writing the following files: | |
| INFO:gguf.gguf_writer:upload-OpenJev/OpenJev-BF16.gguf: n_tensors = 851, total_size = 53.8G | |
| Writing: 0%| | 0.00/53.8G [00:00<?, ?byte/s][A | |
| Writing: 5%|β | 2.54G/53.8G [00:05<01:52, 454Mbyte/s][A | |
| Writing: 9%|β | 5.09G/53.8G [00:11<01:50, 442Mbyte/s][A | |
| Writing: 11%|β | 5.67G/53.8G [00:13<01:52, 427Mbyte/s][A | |
| Writing: 12%|ββ | 6.26G/53.8G [00:14<01:49, 435Mbyte/s][A | |
| Writing: 13%|ββ | 6.79G/53.8G [00:15<01:46, 440Mbyte/s][A | |
| Writing: 14%|ββ | 7.39G/53.8G [00:16<01:42, 453Mbyte/s][A | |
| Writing: 15%|ββ | 7.92G/53.8G [00:17<01:39, 461Mbyte/s][A | |
| Writing: 16%|ββ | 8.54G/53.8G [00:18<01:36, 470Mbyte/s][A | |
| Writing: 17%|ββ | 9.07G/53.8G [00:20<01:35, 468Mbyte/s][A | |
| Writing: 18%|ββ | 9.66G/53.8G [00:21<01:32, 476Mbyte/s][A | |
| Writing: 19%|ββ | 10.3G/53.8G [00:22<01:31, 475Mbyte/s][A | |
| Writing: 20%|ββ | 10.8G/53.8G [00:23<01:29, 479Mbyte/s][A | |
| Writing: 21%|ββ | 11.3G/53.8G [00:24<01:28, 479Mbyte/s][A | |
| Writing: 22%|βββ | 11.9G/53.8G [00:26<01:26, 484Mbyte/s][A | |
| Writing: 23%|βββ | 12.5G/53.8G [00:27<01:27, 475Mbyte/s][A | |
| Writing: 24%|βββ | 13.1G/53.8G [00:28<01:24, 480Mbyte/s][A | |
| Writing: 25%|βββ | 13.7G/53.8G [00:29<01:23, 482Mbyte/s][A | |
| Writing: 27%|βββ | 14.3G/53.8G [00:30<01:21, 485Mbyte/s][A | |
| Writing: 28%|βββ | 14.8G/53.8G [00:32<01:21, 477Mbyte/s][A | |
| Writing: 29%|βββ | 15.4G/53.8G [00:33<01:19, 483Mbyte/s][A | |
| Writing: 30%|βββ | 16.0G/53.8G [00:34<01:18, 481Mbyte/s][A | |
| Writing: 31%|βββ | 16.5G/53.8G [00:35<01:17, 484Mbyte/s][A | |
| Writing: 32%|ββββ | 17.1G/53.8G [00:36<01:15, 489Mbyte/s][A | |
| Writing: 33%|ββββ | 17.7G/53.8G [00:37<01:14, 483Mbyte/s][A | |
| Writing: 34%|ββββ | 18.2G/53.8G [00:39<01:14, 476Mbyte/s][A | |
| Writing: 35%|ββββ | 18.8G/53.8G [00:40<01:12, 483Mbyte/s][A | |
| Writing: 36%|ββββ | 19.4G/53.8G [00:41<01:11, 480Mbyte/s][A | |
| Writing: 37%|ββββ | 19.9G/53.8G [00:42<01:10, 482Mbyte/s][A | |
| Writing: 38%|ββββ | 20.4G/53.8G [00:43<01:09, 482Mbyte/s][A | |
| Writing: 39%|ββββ | 21.1G/53.8G [00:44<01:07, 486Mbyte/s][A | |
| Writing: 40%|ββββ | 21.7G/53.8G [00:46<01:06, 481Mbyte/s][A | |
| Writing: 41%|βββββ | 22.3G/53.8G [00:47<01:06, 478Mbyte/s][A | |
| Writing: 42%|βββββ | 22.8G/53.8G [00:48<01:04, 480Mbyte/s][A | |
| Writing: 43%|βββββ | 23.3G/53.8G [00:49<01:03, 482Mbyte/s][A | |
| Writing: 45%|βββββ | 23.9G/53.8G [00:50<01:01, 482Mbyte/s][A | |
| Writing: 46%|βββββ | 24.5G/53.8G [00:52<01:01, 477Mbyte/s][A | |
| Writing: 47%|βββββ | 25.1G/53.8G [00:53<01:01, 468Mbyte/s][A | |
| Writing: 48%|βββββ | 25.7G/53.8G [00:54<00:59, 472Mbyte/s][A | |
| Writing: 49%|βββββ | 26.2G/53.8G [00:55<00:58, 474Mbyte/s][A | |
| Writing: 50%|βββββ | 26.8G/53.8G [00:57<00:56, 477Mbyte/s][A | |
| Writing: 51%|βββββ | 27.3G/53.8G [00:58<00:56, 469Mbyte/s][A | |
| Writing: 52%|ββββββ | 27.9G/53.8G [00:59<00:54, 474Mbyte/s][A | |
| Writing: 53%|ββββββ | 28.5G/53.8G [01:00<00:53, 473Mbyte/s][A | |
| Writing: 54%|ββββββ | 29.1G/53.8G [01:01<00:51, 476Mbyte/s][A | |
| Writing: 55%|ββββββ | 29.5G/53.8G [01:02<00:51, 475Mbyte/s][A | |
| Writing: 56%|ββββββ | 30.2G/53.8G [01:04<00:49, 479Mbyte/s][A | |
| Writing: 57%|ββββββ | 30.8G/53.8G [01:05<00:48, 472Mbyte/s][A | |
| Writing: 58%|ββββββ | 31.4G/53.8G [01:06<00:47, 469Mbyte/s][A | |
| Writing: 59%|ββββββ | 31.9G/53.8G [01:07<00:46, 471Mbyte/s][A | |
| Writing: 60%|ββββββ | 32.5G/53.8G [01:09<00:45, 473Mbyte/s][A | |
| Writing: 61%|βββββββ | 33.1G/53.8G [01:10<00:43, 471Mbyte/s][A | |
| Writing: 63%|βββββββ | 33.7G/53.8G [01:11<00:43, 466Mbyte/s][A | |
| Writing: 64%|βββββββ | 34.2G/53.8G [01:12<00:42, 457Mbyte/s][A | |
| Writing: 65%|βββββββ | 34.8G/53.8G [01:14<00:41, 461Mbyte/s][A | |
| Writing: 66%|βββββββ | 35.3G/53.8G [01:15<00:40, 462Mbyte/s][A | |
| Writing: 67%|βββββββ | 35.9G/53.8G [01:16<00:38, 463Mbyte/s][A | |
| Writing: 68%|βββββββ | 36.5G/53.8G [01:17<00:38, 455Mbyte/s][A | |
| Writing: 69%|βββββββ | 37.1G/53.8G [01:19<00:36, 459Mbyte/s][A | |
| Writing: 70%|βββββββ | 37.7G/53.8G [01:20<00:35, 457Mbyte/s][A | |
| Writing: 71%|βββββββ | 38.2G/53.8G [01:21<00:34, 459Mbyte/s][A | |
| Writing: 72%|ββββββββ | 38.7G/53.8G [01:22<00:33, 458Mbyte/s][A | |
| Writing: 73%|ββββββββ | 39.2G/53.8G [01:23<00:31, 461Mbyte/s][A | |
| Writing: 74%|ββββββββ | 39.8G/53.8G [01:24<00:30, 458Mbyte/s][A | |
| Writing: 75%|ββββββββ | 40.2G/53.8G [01:26<00:29, 453Mbyte/s][A | |
| Writing: 76%|ββββββββ | 40.7G/53.8G [01:27<00:28, 457Mbyte/s][A | |
| Writing: 77%|ββββββββ | 41.2G/53.8G [01:28<00:27, 459Mbyte/s][A | |
| Writing: 78%|ββββββββ | 41.7G/53.8G [01:29<00:26, 458Mbyte/s][A | |
| Writing: 78%|ββββββββ | 42.2G/53.8G [01:30<00:25, 461Mbyte/s][A | |
| Writing: 80%|ββββββββ | 42.8G/53.8G [01:31<00:24, 456Mbyte/s][A | |
| Writing: 80%|ββββββββ | 43.3G/53.8G [01:32<00:23, 452Mbyte/s][A | |
| Writing: 81%|βββββββββ | 43.7G/53.8G [01:33<00:22, 457Mbyte/s][A | |
| Writing: 82%|βββββββββ | 44.3G/53.8G [01:34<00:20, 459Mbyte/s][A | |
| Writing: 83%|βββββββββ | 44.8G/53.8G [01:35<00:19, 458Mbyte/s][A | |
| Writing: 84%|βββββββββ | 45.3G/53.8G [01:36<00:18, 462Mbyte/s][A | |
| Writing: 85%|βββββββββ | 45.8G/53.8G [01:38<00:17, 457Mbyte/s][A | |
| Writing: 86%|βββββββββ | 46.3G/53.8G [01:39<00:16, 452Mbyte/s][A | |
| Writing: 87%|βββββββββ | 46.8G/53.8G [01:40<00:15, 456Mbyte/s][A | |
| Writing: 88%|βββββββββ | 47.3G/53.8G [01:41<00:14, 458Mbyte/s][A | |
| Writing: 89%|βββββββββ | 47.8G/53.8G [01:42<00:13, 457Mbyte/s][A | |
| Writing: 90%|βββββββββ | 48.3G/53.8G [01:43<00:11, 461Mbyte/s][A | |
| Writing: 91%|βββββββββ | 48.9G/53.8G [01:44<00:10, 457Mbyte/s][A | |
| Writing: 92%|ββββββββββ| 49.3G/53.8G [01:45<00:09, 452Mbyte/s][A | |
| Writing: 93%|ββββββββββ| 49.8G/53.8G [01:47<00:08, 457Mbyte/s][A | |
| Writing: 94%|ββββββββββ| 50.4G/53.8G [01:48<00:07, 459Mbyte/s][A | |
| Writing: 95%|ββββββββββ| 50.9G/53.8G [01:49<00:06, 458Mbyte/s][A | |
| Writing: 95%|ββββββββββ| 51.3G/53.8G [01:50<00:05, 462Mbyte/s][A | |
| Writing: 97%|ββββββββββ| 51.9G/53.8G [01:51<00:04, 460Mbyte/s][A | |
| Writing: 97%|ββββββββββ| 52.4G/53.8G [01:52<00:03, 458Mbyte/s][A | |
| Writing: 98%|ββββββββββ| 52.9G/53.8G [01:53<00:01, 462Mbyte/s][A | |
| Writing: 99%|ββββββββββ| 53.4G/53.8G [01:54<00:00, 468Mbyte/s][A Writing: 100%|ββββββββββ| 53.8G/53.8G [01:55<00:00, 466Mbyte/s] | |
| INFO:hf-to-gguf:Model successfully exported to upload-OpenJev/OpenJev-BF16.gguf | |
| + python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-OpenJev-PRIMARY --outtype bf16 --outfile ./upload-OpenJev/mmproj-OpenJev-BF16.gguf --mmproj --model-name OpenJev | |
| INFO:hf-to-gguf:Loading model: model-temp-OpenJev-PRIMARY | |
| INFO:hf-to-gguf:gguf: detected OpenJev checkpoint | |
| [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' | |
| INFO:hf-to-gguf:Model architecture: OpenJevModel | |
| INFO:hf-to-gguf:gguf: detected OpenJev checkpoint | |
| [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' | |
| INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00012.safetensors' | |
| INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only | |
| INFO:hf-to-gguf:Exporting model... | |
| INFO:hf-to-gguf:v.blk.0.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.0.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.0.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.0.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.0.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.0.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.0.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.1.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.1.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.1.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.1.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.1.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.1.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.10.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.10.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.10.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.10.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.10.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.10.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.11.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.11.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.11.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.11.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.11.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.11.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.12.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.12.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.12.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.12.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.12.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.12.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.13.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.13.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.13.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.13.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.13.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.13.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.14.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.14.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.14.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.14.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.14.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.14.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.15.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.15.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.15.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.15.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.15.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.15.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.16.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.16.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.16.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.16.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.16.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.16.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.17.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.17.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.17.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.17.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.17.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.17.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.18.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.18.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.18.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.18.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.18.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.18.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.19.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.19.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.19.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.19.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.19.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.19.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.2.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.2.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.2.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.2.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.2.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.2.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.20.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.20.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.20.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.20.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.20.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.20.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.21.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.21.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.21.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.21.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.21.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.21.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.22.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.22.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.22.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.22.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.22.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.22.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.23.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.23.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.23.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.23.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.23.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.23.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.24.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.24.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.24.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.24.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.24.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.24.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.25.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.25.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.25.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.25.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.25.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.25.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.26.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.26.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.26.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.26.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.26.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.26.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.3.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.3.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.3.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.3.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.3.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.3.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.4.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.4.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.4.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.4.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.4.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.4.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.5.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.5.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.5.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.5.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.5.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.5.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.6.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.6.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.6.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.6.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.6.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.6.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.7.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.7.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.7.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.7.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.7.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.7.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.8.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.8.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.8.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.8.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.8.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.8.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.attn_out.weight, torch.bfloat16 --> BF16, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.9.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.9.attn_qkv.weight, torch.bfloat16 --> BF16, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.9.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.9.ffn_up.weight, torch.bfloat16 --> BF16, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.9.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.ffn_down.weight, torch.bfloat16 --> BF16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.9.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:mm.0.bias, torch.bfloat16 --> F32, shape = {4608} | |
| INFO:hf-to-gguf:mm.0.weight, torch.bfloat16 --> BF16, shape = {4608, 4608} | |
| INFO:hf-to-gguf:mm.2.bias, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:mm.2.weight, torch.bfloat16 --> BF16, shape = {4608, 5120} | |
| INFO:hf-to-gguf:v.post_ln.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.post_ln.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.patch_embd.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.patch_embd.weight, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152} | |
| INFO:hf-to-gguf:v.patch_embd.weight.1, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152} | |
| INFO:hf-to-gguf:v.position_embd.weight, torch.bfloat16 --> F32, shape = {1152, 2304} | |
| INFO:hf-to-gguf:Set meta model | |
| INFO:hf-to-gguf:Set model parameters | |
| INFO:hf-to-gguf:Set model quantization version | |
| INFO:gguf.gguf_writer:Writing the following files: | |
| INFO:gguf.gguf_writer:upload-OpenJev/mmproj-OpenJev-BF16.gguf: n_tensors = 334, total_size = 931.1M | |
| Writing: 0%| | 0.00/931M [00:00<?, ?byte/s][A | |
| Writing: 48%|βββββ | 448M/931M [00:01<00:01, 445Mbyte/s][A | |
| Writing: 98%|ββββββββββ| 913M/931M [00:02<00:00, 449Mbyte/s][A Writing: 100%|ββββββββββ| 931M/931M [00:02<00:00, 448Mbyte/s] | |
| INFO:hf-to-gguf:Model successfully exported to upload-OpenJev/mmproj-OpenJev-BF16.gguf | |
| + FLAGS_Q4_K_M='--pure --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ssm_=q8_0' | |
| + llama.cpp/build/bin/llama-quantize ./upload-OpenJev/OpenJev-BF16.gguf ./upload-OpenJev/OpenJev-Q8_0.gguf Q8_0 | |
| version: 0.5.0-dev (build 11354, commit 54a0c5da9) | |
| built with GNU 14.2.0 for Linux x86_64 | |
| llama_quantize: quantizing './upload-OpenJev/OpenJev-BF16.gguf' to './upload-OpenJev/OpenJev-Q8_0.gguf' as Q8_0 | |
| llama_model_loader: loaded meta data with 48 key-value pairs and 851 tensors from ./upload-OpenJev/OpenJev-BF16.gguf (version GGUF V3 (latest)) | |
| llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output. | |
| llama_model_loader: - kv 0: general.architecture str = qwen35 | |
| llama_model_loader: - kv 1: general.type str = model | |
| llama_model_loader: - kv 2: general.sampling.top_k i32 = 20 | |
| llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.950000 | |
| llama_model_loader: - kv 4: general.sampling.temp f32 = 1.000000 | |
| llama_model_loader: - kv 5: general.name str = OpenJev | |
| llama_model_loader: - kv 6: general.size_label str = 27B | |
| llama_model_loader: - kv 7: general.license str = cc-by-nc-4.0 | |
| llama_model_loader: - kv 8: general.tags arr[str,8] = ["decision-model", "zero-shot-classif... | |
| llama_model_loader: - kv 9: general.languages arr[str,6] = ["en", "de", "fr", "hi", "zh", "ja"] | |
| llama_model_loader: - kv 10: qwen35.block_count u32 = 64 | |
| llama_model_loader: - kv 11: qwen35.context_length u32 = 262144 | |
| llama_model_loader: - kv 12: qwen35.embedding_length u32 = 5120 | |
| llama_model_loader: - kv 13: qwen35.feed_forward_length u32 = 17408 | |
| llama_model_loader: - kv 14: qwen35.attention.head_count u32 = 24 | |
| llama_model_loader: - kv 15: qwen35.attention.head_count_kv u32 = 4 | |
| llama_model_loader: - kv 16: qwen35.rope.dimension_sections arr[i32,4] = [11, 11, 10, 0] | |
| llama_model_loader: - kv 17: qwen35.rope.freq_base f32 = 10000000.000000 | |
| llama_model_loader: - kv 18: qwen35.attention.layer_norm_rms_epsilon f32 = 0.000001 | |
| llama_model_loader: - kv 19: qwen35.attention.key_length u32 = 256 | |
| llama_model_loader: - kv 20: qwen35.attention.value_length u32 = 256 | |
| llama_model_loader: - kv 21: general.file_type u32 = 32 | |
| llama_model_loader: - kv 22: qwen35.ssm.conv_kernel u32 = 4 | |
| llama_model_loader: - kv 23: qwen35.ssm.state_size u32 = 128 | |
| llama_model_loader: - kv 24: qwen35.ssm.group_count u32 = 16 | |
| llama_model_loader: - kv 25: qwen35.ssm.time_step_rank u32 = 48 | |
| llama_model_loader: - kv 26: qwen35.ssm.inner_size u32 = 6144 | |
| llama_model_loader: - kv 27: qwen35.attention.recurrent_layers arr[bool,64] = [true, true, true, false, true, true,... | |
| llama_model_loader: - kv 28: qwen35.full_attention_interval u32 = 4 | |
| llama_model_loader: - kv 29: qwen35.rope.dimension_count u32 = 64 | |
| llama_model_loader: - kv 30: qwen35.decision.type str = openjev | |
| llama_model_loader: - kv 31: qwen35.decision.temperature.choice f32 = 0.850000 | |
| llama_model_loader: - kv 32: qwen35.decision.temperature.score f32 = 0.850000 | |
| llama_model_loader: - kv 33: qwen35.decision.temperature.noul f32 = 1.554713 | |
| llama_model_loader: - kv 34: general.quantization_version u32 = 2 | |
| llama_model_loader: - kv 35: tokenizer.ggml.model str = gpt2 | |
| llama_model_loader: - kv 36: tokenizer.ggml.pre str = qwen35 | |
| llama_model_loader: - kv 37: tokenizer.ggml.tokens arr[str,248320] = ["!", "\"", "#", "$", "%", "&", "'", ... | |
| llama_model_loader: - kv 38: tokenizer.ggml.token_type arr[i32,248320] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ... | |
| llama_model_loader: - kv 39: tokenizer.ggml.merges arr[str,247587] = ["Δ Δ ", "Δ Δ Δ Δ ", "i n", "Δ t",... | |
| llama_model_loader: - kv 40: tokenizer.ggml.eos_token_id u32 = 248046 | |
| llama_model_loader: - kv 41: tokenizer.ggml.padding_token_id u32 = 248044 | |
| llama_model_loader: - kv 42: tokenizer.ggml.bos_token_id u32 = 248044 | |
| llama_model_loader: - kv 43: tokenizer.ggml.add_bos_token bool = false | |
| llama_model_loader: - kv 44: tokenizer.ggml.add_eos_token bool = false | |
| llama_model_loader: - kv 45: tokenizer.chat_template str = {%- set image_count = namespace(value... | |
| llama_model_loader: - kv 46: tokenizer.chat_template.systemone str = {% set letters = 'ABCDEFGHIJKLMNOPQRS... | |
| llama_model_loader: - kv 47: tokenizer.chat_templates arr[str,1] = ["systemone"] | |
| llama_model_loader: - type f32: 353 tensors | |
| llama_model_loader: - type bf16: 498 tensors | |
| [ 1/ 851] output.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q8_0 .. size = 2425.00 MiB -> 1288.28 MiB | |
| [ 2/ 851] output_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 3/ 851] token_embd.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q8_0 .. size = 2425.00 MiB -> 1288.28 MiB | |
| [ 4/ 851] blk.0.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 5/ 851] blk.0.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 6/ 851] blk.0.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 7/ 851] blk.0.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 8/ 851] blk.0.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 9/ 851] blk.0.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 10/ 851] blk.0.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 11/ 851] blk.0.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 12/ 851] blk.0.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 13/ 851] blk.0.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 14/ 851] blk.0.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 15/ 851] blk.0.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 16/ 851] blk.0.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 17/ 851] blk.0.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 18/ 851] blk.1.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 19/ 851] blk.1.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 20/ 851] blk.1.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 21/ 851] blk.1.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 22/ 851] blk.1.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 23/ 851] blk.1.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 24/ 851] blk.1.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 25/ 851] blk.1.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 26/ 851] blk.1.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 27/ 851] blk.1.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 28/ 851] blk.1.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 29/ 851] blk.1.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 30/ 851] blk.1.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 31/ 851] blk.1.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 32/ 851] blk.2.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 33/ 851] blk.2.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 34/ 851] blk.2.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 35/ 851] blk.2.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 36/ 851] blk.2.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 37/ 851] blk.2.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 38/ 851] blk.2.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 39/ 851] blk.2.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 40/ 851] blk.2.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 41/ 851] blk.2.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 42/ 851] blk.2.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 43/ 851] blk.2.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 44/ 851] blk.2.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 45/ 851] blk.2.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 46/ 851] blk.3.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 47/ 851] blk.3.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 48/ 851] blk.3.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 49/ 851] blk.3.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 50/ 851] blk.3.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 51/ 851] blk.3.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 52/ 851] blk.3.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 53/ 851] blk.3.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 54/ 851] blk.3.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 55/ 851] blk.3.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 56/ 851] blk.3.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 57/ 851] blk.4.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 58/ 851] blk.4.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 59/ 851] blk.4.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 60/ 851] blk.4.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 61/ 851] blk.4.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 62/ 851] blk.4.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 63/ 851] blk.4.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 64/ 851] blk.4.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 65/ 851] blk.4.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 66/ 851] blk.4.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 67/ 851] blk.4.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 68/ 851] blk.4.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 69/ 851] blk.4.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 70/ 851] blk.4.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 71/ 851] blk.5.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 72/ 851] blk.5.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 73/ 851] blk.5.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 74/ 851] blk.5.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 75/ 851] blk.5.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 76/ 851] blk.5.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 77/ 851] blk.5.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 78/ 851] blk.5.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 79/ 851] blk.5.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 80/ 851] blk.5.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 81/ 851] blk.5.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 82/ 851] blk.5.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 83/ 851] blk.5.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 84/ 851] blk.5.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 85/ 851] blk.6.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 86/ 851] blk.6.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 87/ 851] blk.6.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 88/ 851] blk.6.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 89/ 851] blk.6.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 90/ 851] blk.6.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 91/ 851] blk.6.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 92/ 851] blk.6.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 93/ 851] blk.6.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 94/ 851] blk.6.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 95/ 851] blk.6.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 96/ 851] blk.6.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 97/ 851] blk.6.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 98/ 851] blk.6.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 99/ 851] blk.7.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 100/ 851] blk.7.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 101/ 851] blk.7.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 102/ 851] blk.7.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 103/ 851] blk.7.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 104/ 851] blk.7.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 105/ 851] blk.7.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 106/ 851] blk.7.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 107/ 851] blk.7.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 108/ 851] blk.7.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 109/ 851] blk.7.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 110/ 851] blk.8.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 111/ 851] blk.8.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 112/ 851] blk.8.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 113/ 851] blk.8.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 114/ 851] blk.8.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 115/ 851] blk.8.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 116/ 851] blk.8.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 117/ 851] blk.8.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 118/ 851] blk.8.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 119/ 851] blk.8.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 120/ 851] blk.8.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 121/ 851] blk.8.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 122/ 851] blk.8.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 123/ 851] blk.8.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 124/ 851] blk.9.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 125/ 851] blk.9.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 126/ 851] blk.9.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 127/ 851] blk.9.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 128/ 851] blk.9.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 129/ 851] blk.9.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 130/ 851] blk.9.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 131/ 851] blk.9.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 132/ 851] blk.9.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 133/ 851] blk.9.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 134/ 851] blk.9.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 135/ 851] blk.9.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 136/ 851] blk.9.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 137/ 851] blk.9.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 138/ 851] blk.10.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 139/ 851] blk.10.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 140/ 851] blk.10.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 141/ 851] blk.10.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 142/ 851] blk.10.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 143/ 851] blk.10.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 144/ 851] blk.10.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 145/ 851] blk.10.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 146/ 851] blk.10.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 147/ 851] blk.10.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 148/ 851] blk.10.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 149/ 851] blk.10.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 150/ 851] blk.10.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 151/ 851] blk.10.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 152/ 851] blk.11.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 153/ 851] blk.11.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 154/ 851] blk.11.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 155/ 851] blk.11.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 156/ 851] blk.11.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 157/ 851] blk.11.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 158/ 851] blk.11.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 159/ 851] blk.11.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 160/ 851] blk.11.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 161/ 851] blk.11.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 162/ 851] blk.11.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 163/ 851] blk.12.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 164/ 851] blk.12.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 165/ 851] blk.12.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 166/ 851] blk.12.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 167/ 851] blk.12.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 168/ 851] blk.12.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 169/ 851] blk.12.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 170/ 851] blk.12.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 171/ 851] blk.12.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 172/ 851] blk.12.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 173/ 851] blk.12.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 174/ 851] blk.12.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 175/ 851] blk.12.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 176/ 851] blk.12.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 177/ 851] blk.13.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 178/ 851] blk.13.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 179/ 851] blk.13.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 180/ 851] blk.13.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 181/ 851] blk.13.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 182/ 851] blk.13.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 183/ 851] blk.13.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 184/ 851] blk.13.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 185/ 851] blk.13.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 186/ 851] blk.13.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 187/ 851] blk.13.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 188/ 851] blk.13.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 189/ 851] blk.13.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 190/ 851] blk.13.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 191/ 851] blk.14.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 192/ 851] blk.14.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 193/ 851] blk.14.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 194/ 851] blk.14.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 195/ 851] blk.14.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 196/ 851] blk.14.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 197/ 851] blk.14.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 198/ 851] blk.14.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 199/ 851] blk.14.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 200/ 851] blk.14.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 201/ 851] blk.14.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 202/ 851] blk.14.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 203/ 851] blk.14.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 204/ 851] blk.14.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 205/ 851] blk.15.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 206/ 851] blk.15.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 207/ 851] blk.15.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 208/ 851] blk.15.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 209/ 851] blk.15.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 210/ 851] blk.15.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 211/ 851] blk.15.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 212/ 851] blk.15.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 213/ 851] blk.15.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 214/ 851] blk.15.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 215/ 851] blk.15.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 216/ 851] blk.16.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 217/ 851] blk.16.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 218/ 851] blk.16.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 219/ 851] blk.16.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 220/ 851] blk.16.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 221/ 851] blk.16.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 222/ 851] blk.16.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 223/ 851] blk.16.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 224/ 851] blk.16.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 225/ 851] blk.16.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 226/ 851] blk.16.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 227/ 851] blk.16.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 228/ 851] blk.16.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 229/ 851] blk.16.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 230/ 851] blk.17.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 231/ 851] blk.17.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 232/ 851] blk.17.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 233/ 851] blk.17.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 234/ 851] blk.17.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 235/ 851] blk.17.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 236/ 851] blk.17.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 237/ 851] blk.17.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 238/ 851] blk.17.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 239/ 851] blk.17.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 240/ 851] blk.17.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 241/ 851] blk.17.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 242/ 851] blk.17.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 243/ 851] blk.17.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 244/ 851] blk.18.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 245/ 851] blk.18.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 246/ 851] blk.18.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 247/ 851] blk.18.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 248/ 851] blk.18.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 249/ 851] blk.18.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 250/ 851] blk.18.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 251/ 851] blk.18.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 252/ 851] blk.18.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 253/ 851] blk.18.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 254/ 851] blk.18.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 255/ 851] blk.18.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 256/ 851] blk.18.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 257/ 851] blk.18.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 258/ 851] blk.19.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 259/ 851] blk.19.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 260/ 851] blk.19.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 261/ 851] blk.19.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 262/ 851] blk.19.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 263/ 851] blk.19.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 264/ 851] blk.19.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 265/ 851] blk.19.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 266/ 851] blk.19.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 267/ 851] blk.19.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 268/ 851] blk.19.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 269/ 851] blk.20.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 270/ 851] blk.20.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 271/ 851] blk.20.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 272/ 851] blk.20.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 273/ 851] blk.20.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 274/ 851] blk.20.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 275/ 851] blk.20.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 276/ 851] blk.20.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 277/ 851] blk.20.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 278/ 851] blk.20.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 279/ 851] blk.20.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 280/ 851] blk.20.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 281/ 851] blk.20.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 282/ 851] blk.20.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 283/ 851] blk.21.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 284/ 851] blk.21.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 285/ 851] blk.21.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 286/ 851] blk.21.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 287/ 851] blk.21.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 288/ 851] blk.21.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 289/ 851] blk.21.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 290/ 851] blk.21.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 291/ 851] blk.21.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 292/ 851] blk.21.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 293/ 851] blk.21.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 294/ 851] blk.21.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 295/ 851] blk.21.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 296/ 851] blk.21.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 297/ 851] blk.22.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 298/ 851] blk.22.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 299/ 851] blk.22.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 300/ 851] blk.22.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 301/ 851] blk.22.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 302/ 851] blk.22.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 303/ 851] blk.22.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 304/ 851] blk.22.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 305/ 851] blk.22.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 306/ 851] blk.22.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 307/ 851] blk.22.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 308/ 851] blk.22.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 309/ 851] blk.22.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 310/ 851] blk.22.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 311/ 851] blk.23.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 312/ 851] blk.23.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 313/ 851] blk.23.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 314/ 851] blk.23.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 315/ 851] blk.23.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 316/ 851] blk.23.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 317/ 851] blk.23.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 318/ 851] blk.23.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 319/ 851] blk.23.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 320/ 851] blk.23.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 321/ 851] blk.23.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 322/ 851] blk.24.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 323/ 851] blk.24.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 324/ 851] blk.24.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 325/ 851] blk.24.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 326/ 851] blk.24.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 327/ 851] blk.24.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 328/ 851] blk.24.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 329/ 851] blk.24.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 330/ 851] blk.24.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 331/ 851] blk.24.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 332/ 851] blk.24.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 333/ 851] blk.24.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 334/ 851] blk.24.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 335/ 851] blk.24.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 336/ 851] blk.25.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 337/ 851] blk.25.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 338/ 851] blk.25.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 339/ 851] blk.25.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 340/ 851] blk.25.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 341/ 851] blk.25.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 342/ 851] blk.25.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 343/ 851] blk.25.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 344/ 851] blk.25.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 345/ 851] blk.25.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 346/ 851] blk.25.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 347/ 851] blk.25.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 348/ 851] blk.25.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 349/ 851] blk.25.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 350/ 851] blk.26.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 351/ 851] blk.26.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 352/ 851] blk.26.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 353/ 851] blk.26.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 354/ 851] blk.26.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 355/ 851] blk.26.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 356/ 851] blk.26.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 357/ 851] blk.26.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 358/ 851] blk.26.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 359/ 851] blk.26.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 360/ 851] blk.26.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 361/ 851] blk.26.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 362/ 851] blk.26.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 363/ 851] blk.26.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 364/ 851] blk.27.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 365/ 851] blk.27.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 366/ 851] blk.27.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 367/ 851] blk.27.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 368/ 851] blk.27.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 369/ 851] blk.27.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 370/ 851] blk.27.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 371/ 851] blk.27.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 372/ 851] blk.27.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 373/ 851] blk.27.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 374/ 851] blk.27.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 375/ 851] blk.28.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 376/ 851] blk.28.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 377/ 851] blk.28.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 378/ 851] blk.28.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 379/ 851] blk.28.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 380/ 851] blk.28.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 381/ 851] blk.28.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 382/ 851] blk.28.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 383/ 851] blk.28.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 384/ 851] blk.28.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 385/ 851] blk.28.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 386/ 851] blk.28.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 387/ 851] blk.28.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 388/ 851] blk.28.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 389/ 851] blk.29.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 390/ 851] blk.29.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 391/ 851] blk.29.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 392/ 851] blk.29.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 393/ 851] blk.29.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 394/ 851] blk.29.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 395/ 851] blk.29.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 396/ 851] blk.29.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 397/ 851] blk.29.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 398/ 851] blk.29.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 399/ 851] blk.29.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 400/ 851] blk.29.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 401/ 851] blk.29.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 402/ 851] blk.29.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 403/ 851] blk.30.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 404/ 851] blk.30.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 405/ 851] blk.30.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 406/ 851] blk.30.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 407/ 851] blk.30.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 408/ 851] blk.30.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 409/ 851] blk.30.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 410/ 851] blk.30.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 411/ 851] blk.30.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 412/ 851] blk.30.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 413/ 851] blk.30.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 414/ 851] blk.30.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 415/ 851] blk.30.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 416/ 851] blk.30.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 417/ 851] blk.31.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 418/ 851] blk.31.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 419/ 851] blk.31.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 420/ 851] blk.31.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 421/ 851] blk.31.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 422/ 851] blk.31.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 423/ 851] blk.31.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 424/ 851] blk.31.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 425/ 851] blk.31.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 426/ 851] blk.31.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 427/ 851] blk.31.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 428/ 851] blk.32.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 429/ 851] blk.32.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 430/ 851] blk.32.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 431/ 851] blk.32.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 432/ 851] blk.32.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 433/ 851] blk.32.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 434/ 851] blk.32.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 435/ 851] blk.32.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 436/ 851] blk.32.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 437/ 851] blk.32.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 438/ 851] blk.32.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 439/ 851] blk.32.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 440/ 851] blk.32.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 441/ 851] blk.32.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 442/ 851] blk.33.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 443/ 851] blk.33.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 444/ 851] blk.33.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 445/ 851] blk.33.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 446/ 851] blk.33.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 447/ 851] blk.33.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 448/ 851] blk.33.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 449/ 851] blk.33.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 450/ 851] blk.33.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 451/ 851] blk.33.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 452/ 851] blk.33.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 453/ 851] blk.33.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 454/ 851] blk.33.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 455/ 851] blk.33.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 456/ 851] blk.34.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 457/ 851] blk.34.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 458/ 851] blk.34.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 459/ 851] blk.34.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 460/ 851] blk.34.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 461/ 851] blk.34.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 462/ 851] blk.34.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 463/ 851] blk.34.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 464/ 851] blk.34.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 465/ 851] blk.34.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 466/ 851] blk.34.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 467/ 851] blk.34.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 468/ 851] blk.34.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 469/ 851] blk.34.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 470/ 851] blk.35.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 471/ 851] blk.35.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 472/ 851] blk.35.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 473/ 851] blk.35.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 474/ 851] blk.35.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 475/ 851] blk.35.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 476/ 851] blk.35.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 477/ 851] blk.35.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 478/ 851] blk.35.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 479/ 851] blk.35.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 480/ 851] blk.35.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 481/ 851] blk.36.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 482/ 851] blk.36.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 483/ 851] blk.36.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 484/ 851] blk.36.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 485/ 851] blk.36.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 486/ 851] blk.36.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 487/ 851] blk.36.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 488/ 851] blk.36.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 489/ 851] blk.36.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 490/ 851] blk.36.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 491/ 851] blk.36.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 492/ 851] blk.36.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 493/ 851] blk.36.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 494/ 851] blk.36.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 495/ 851] blk.37.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 496/ 851] blk.37.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 497/ 851] blk.37.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 498/ 851] blk.37.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 499/ 851] blk.37.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 500/ 851] blk.37.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 501/ 851] blk.37.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 502/ 851] blk.37.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 503/ 851] blk.37.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 504/ 851] blk.37.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 505/ 851] blk.37.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 506/ 851] blk.37.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 507/ 851] blk.37.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 508/ 851] blk.37.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 509/ 851] blk.38.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 510/ 851] blk.38.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 511/ 851] blk.38.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 512/ 851] blk.38.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 513/ 851] blk.38.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 514/ 851] blk.38.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 515/ 851] blk.38.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 516/ 851] blk.38.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 517/ 851] blk.38.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 518/ 851] blk.38.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 519/ 851] blk.38.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 520/ 851] blk.38.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 521/ 851] blk.38.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 522/ 851] blk.38.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 523/ 851] blk.39.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 524/ 851] blk.39.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 525/ 851] blk.39.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 526/ 851] blk.39.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 527/ 851] blk.39.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 528/ 851] blk.39.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 529/ 851] blk.39.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 530/ 851] blk.39.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 531/ 851] blk.39.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 532/ 851] blk.39.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 533/ 851] blk.39.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 534/ 851] blk.40.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 535/ 851] blk.40.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 536/ 851] blk.40.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 537/ 851] blk.40.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 538/ 851] blk.40.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 539/ 851] blk.40.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 540/ 851] blk.40.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 541/ 851] blk.40.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 542/ 851] blk.40.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 543/ 851] blk.40.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 544/ 851] blk.40.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 545/ 851] blk.40.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 546/ 851] blk.40.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 547/ 851] blk.40.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 548/ 851] blk.41.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 549/ 851] blk.41.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 550/ 851] blk.41.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 551/ 851] blk.41.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 552/ 851] blk.41.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 553/ 851] blk.41.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 554/ 851] blk.41.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 555/ 851] blk.41.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 556/ 851] blk.41.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 557/ 851] blk.41.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 558/ 851] blk.41.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 559/ 851] blk.41.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 560/ 851] blk.41.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 561/ 851] blk.41.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 562/ 851] blk.42.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 563/ 851] blk.42.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 564/ 851] blk.42.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 565/ 851] blk.42.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 566/ 851] blk.42.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 567/ 851] blk.42.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 568/ 851] blk.42.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 569/ 851] blk.42.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 570/ 851] blk.42.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 571/ 851] blk.42.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 572/ 851] blk.42.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 573/ 851] blk.42.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 574/ 851] blk.42.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 575/ 851] blk.42.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 576/ 851] blk.43.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 577/ 851] blk.43.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 578/ 851] blk.43.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 579/ 851] blk.43.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 580/ 851] blk.43.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 581/ 851] blk.43.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 582/ 851] blk.43.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 583/ 851] blk.43.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 584/ 851] blk.43.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 585/ 851] blk.43.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 586/ 851] blk.43.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 587/ 851] blk.44.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 588/ 851] blk.44.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 589/ 851] blk.44.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 590/ 851] blk.44.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 591/ 851] blk.44.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 592/ 851] blk.44.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 593/ 851] blk.44.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 594/ 851] blk.44.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 595/ 851] blk.44.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 596/ 851] blk.44.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 597/ 851] blk.44.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 598/ 851] blk.44.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 599/ 851] blk.44.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 600/ 851] blk.44.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 601/ 851] blk.45.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 602/ 851] blk.45.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 603/ 851] blk.45.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 604/ 851] blk.45.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 605/ 851] blk.45.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 606/ 851] blk.45.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 607/ 851] blk.45.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 608/ 851] blk.45.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 609/ 851] blk.45.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 610/ 851] blk.45.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 611/ 851] blk.45.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 612/ 851] blk.45.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 613/ 851] blk.45.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 614/ 851] blk.45.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 615/ 851] blk.46.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 616/ 851] blk.46.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 617/ 851] blk.46.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 618/ 851] blk.46.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 619/ 851] blk.46.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 620/ 851] blk.46.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 621/ 851] blk.46.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 622/ 851] blk.46.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 623/ 851] blk.46.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 624/ 851] blk.46.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 625/ 851] blk.46.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 626/ 851] blk.46.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 627/ 851] blk.46.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 628/ 851] blk.46.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 629/ 851] blk.47.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 630/ 851] blk.47.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 631/ 851] blk.47.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 632/ 851] blk.47.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 633/ 851] blk.47.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 634/ 851] blk.47.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 635/ 851] blk.47.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 636/ 851] blk.47.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 637/ 851] blk.47.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 638/ 851] blk.47.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 639/ 851] blk.47.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 640/ 851] blk.48.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 641/ 851] blk.48.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 642/ 851] blk.48.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 643/ 851] blk.48.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 644/ 851] blk.48.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 645/ 851] blk.48.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 646/ 851] blk.48.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 647/ 851] blk.48.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 648/ 851] blk.48.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 649/ 851] blk.48.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 650/ 851] blk.48.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 651/ 851] blk.48.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 652/ 851] blk.48.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 653/ 851] blk.48.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 654/ 851] blk.49.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 655/ 851] blk.49.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 656/ 851] blk.49.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 657/ 851] blk.49.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 658/ 851] blk.49.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 659/ 851] blk.49.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 660/ 851] blk.49.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 661/ 851] blk.49.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 662/ 851] blk.49.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 663/ 851] blk.49.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 664/ 851] blk.49.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 665/ 851] blk.49.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 666/ 851] blk.49.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 667/ 851] blk.49.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 668/ 851] blk.50.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 669/ 851] blk.50.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 670/ 851] blk.50.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 671/ 851] blk.50.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 672/ 851] blk.50.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 673/ 851] blk.50.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 674/ 851] blk.50.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 675/ 851] blk.50.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 676/ 851] blk.50.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 677/ 851] blk.50.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 678/ 851] blk.50.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 679/ 851] blk.50.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 680/ 851] blk.50.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 681/ 851] blk.50.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 682/ 851] blk.51.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 683/ 851] blk.51.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 684/ 851] blk.51.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 685/ 851] blk.51.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 686/ 851] blk.51.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 687/ 851] blk.51.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 688/ 851] blk.51.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 689/ 851] blk.51.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 690/ 851] blk.51.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 691/ 851] blk.51.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 692/ 851] blk.51.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 693/ 851] blk.52.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 694/ 851] blk.52.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 695/ 851] blk.52.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 696/ 851] blk.52.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 697/ 851] blk.52.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 698/ 851] blk.52.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 699/ 851] blk.52.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 700/ 851] blk.52.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 701/ 851] blk.52.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 702/ 851] blk.52.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 703/ 851] blk.52.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 704/ 851] blk.52.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 705/ 851] blk.52.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 706/ 851] blk.52.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 707/ 851] blk.53.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 708/ 851] blk.53.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 709/ 851] blk.53.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 710/ 851] blk.53.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 711/ 851] blk.53.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 712/ 851] blk.53.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 713/ 851] blk.53.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 714/ 851] blk.53.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 715/ 851] blk.53.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 716/ 851] blk.53.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 717/ 851] blk.53.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 718/ 851] blk.53.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 719/ 851] blk.53.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 720/ 851] blk.53.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 721/ 851] blk.54.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 722/ 851] blk.54.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 723/ 851] blk.54.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 724/ 851] blk.54.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 725/ 851] blk.54.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 726/ 851] blk.54.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 727/ 851] blk.54.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 728/ 851] blk.54.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 729/ 851] blk.54.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 730/ 851] blk.54.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 731/ 851] blk.54.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 732/ 851] blk.54.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 733/ 851] blk.54.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 734/ 851] blk.54.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 735/ 851] blk.55.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 736/ 851] blk.55.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 737/ 851] blk.55.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 738/ 851] blk.55.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 739/ 851] blk.55.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 740/ 851] blk.55.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 741/ 851] blk.55.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 742/ 851] blk.55.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 743/ 851] blk.55.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 744/ 851] blk.55.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 745/ 851] blk.55.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 746/ 851] blk.56.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 747/ 851] blk.56.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 748/ 851] blk.56.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 749/ 851] blk.56.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 750/ 851] blk.56.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 751/ 851] blk.56.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 752/ 851] blk.56.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 753/ 851] blk.56.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 754/ 851] blk.56.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 755/ 851] blk.56.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 756/ 851] blk.56.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 757/ 851] blk.56.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 758/ 851] blk.56.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 759/ 851] blk.56.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 760/ 851] blk.57.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 761/ 851] blk.57.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 762/ 851] blk.57.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 763/ 851] blk.57.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 764/ 851] blk.57.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 765/ 851] blk.57.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 766/ 851] blk.57.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 767/ 851] blk.57.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 768/ 851] blk.57.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 769/ 851] blk.57.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 770/ 851] blk.57.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 771/ 851] blk.57.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 772/ 851] blk.57.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 773/ 851] blk.57.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 774/ 851] blk.58.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 775/ 851] blk.58.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 776/ 851] blk.58.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 777/ 851] blk.58.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 778/ 851] blk.58.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 779/ 851] blk.58.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 780/ 851] blk.58.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 781/ 851] blk.58.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 782/ 851] blk.58.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 783/ 851] blk.58.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 784/ 851] blk.58.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 785/ 851] blk.58.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 786/ 851] blk.58.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 787/ 851] blk.58.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 788/ 851] blk.59.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 789/ 851] blk.59.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 790/ 851] blk.59.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 791/ 851] blk.59.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 792/ 851] blk.59.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 793/ 851] blk.59.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 794/ 851] blk.59.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 795/ 851] blk.59.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 796/ 851] blk.59.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 797/ 851] blk.59.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 798/ 851] blk.59.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 799/ 851] blk.60.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 800/ 851] blk.60.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 801/ 851] blk.60.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 802/ 851] blk.60.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 803/ 851] blk.60.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 804/ 851] blk.60.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 805/ 851] blk.60.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 806/ 851] blk.60.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 807/ 851] blk.60.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 808/ 851] blk.60.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 809/ 851] blk.60.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 810/ 851] blk.60.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 811/ 851] blk.60.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 812/ 851] blk.60.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 813/ 851] blk.61.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 814/ 851] blk.61.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 815/ 851] blk.61.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 816/ 851] blk.61.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 817/ 851] blk.61.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 818/ 851] blk.61.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 819/ 851] blk.61.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 820/ 851] blk.61.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 821/ 851] blk.61.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 822/ 851] blk.61.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 823/ 851] blk.61.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 824/ 851] blk.61.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 825/ 851] blk.61.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 826/ 851] blk.61.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 827/ 851] blk.62.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 828/ 851] blk.62.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 829/ 851] blk.62.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 830/ 851] blk.62.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 831/ 851] blk.62.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 832/ 851] blk.62.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 833/ 851] blk.62.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 834/ 851] blk.62.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 835/ 851] blk.62.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 836/ 851] blk.62.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 837/ 851] blk.62.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 838/ 851] blk.62.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 839/ 851] blk.62.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 840/ 851] blk.62.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 841/ 851] blk.63.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 842/ 851] blk.63.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 843/ 851] blk.63.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 844/ 851] blk.63.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 845/ 851] blk.63.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 846/ 851] blk.63.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 847/ 851] blk.63.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 848/ 851] blk.63.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 849/ 851] blk.63.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 850/ 851] blk.63.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q8_0 .. size = 170.00 MiB -> 90.31 MiB | |
| [ 851/ 851] blk.63.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| llama_model_quantize_impl: model size = 51305.09 MiB (16.00 BPW) | |
| llama_model_quantize_impl: quant size = 27260.56 MiB (8.50 BPW) | |
| llama_quantize: quantize time = 22409.99 ms | |
| llama_quantize: total time = 22409.99 ms | |
| + llama.cpp/build/bin/llama-quantize --pure --tensor-type output.weight=q6_k --tensor-type attn_=q8_0 --tensor-type ssm_=q8_0 ./upload-OpenJev/OpenJev-BF16.gguf ./upload-OpenJev/OpenJev-Q4_K_M.gguf Q4_K_M | |
| version: 0.5.0-dev (build 11354, commit 54a0c5da9) | |
| built with GNU 14.2.0 for Linux x86_64 | |
| llama_quantize: quantizing './upload-OpenJev/OpenJev-BF16.gguf' to './upload-OpenJev/OpenJev-Q4_K_M.gguf' as Q4_K_M | |
| llama_model_loader: loaded meta data with 48 key-value pairs and 851 tensors from ./upload-OpenJev/OpenJev-BF16.gguf (version GGUF V3 (latest)) | |
| llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output. | |
| llama_model_loader: - kv 0: general.architecture str = qwen35 | |
| llama_model_loader: - kv 1: general.type str = model | |
| llama_model_loader: - kv 2: general.sampling.top_k i32 = 20 | |
| llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.950000 | |
| llama_model_loader: - kv 4: general.sampling.temp f32 = 1.000000 | |
| llama_model_loader: - kv 5: general.name str = OpenJev | |
| llama_model_loader: - kv 6: general.size_label str = 27B | |
| llama_model_loader: - kv 7: general.license str = cc-by-nc-4.0 | |
| llama_model_loader: - kv 8: general.tags arr[str,8] = ["decision-model", "zero-shot-classif... | |
| llama_model_loader: - kv 9: general.languages arr[str,6] = ["en", "de", "fr", "hi", "zh", "ja"] | |
| llama_model_loader: - kv 10: qwen35.block_count u32 = 64 | |
| llama_model_loader: - kv 11: qwen35.context_length u32 = 262144 | |
| llama_model_loader: - kv 12: qwen35.embedding_length u32 = 5120 | |
| llama_model_loader: - kv 13: qwen35.feed_forward_length u32 = 17408 | |
| llama_model_loader: - kv 14: qwen35.attention.head_count u32 = 24 | |
| llama_model_loader: - kv 15: qwen35.attention.head_count_kv u32 = 4 | |
| llama_model_loader: - kv 16: qwen35.rope.dimension_sections arr[i32,4] = [11, 11, 10, 0] | |
| llama_model_loader: - kv 17: qwen35.rope.freq_base f32 = 10000000.000000 | |
| llama_model_loader: - kv 18: qwen35.attention.layer_norm_rms_epsilon f32 = 0.000001 | |
| llama_model_loader: - kv 19: qwen35.attention.key_length u32 = 256 | |
| llama_model_loader: - kv 20: qwen35.attention.value_length u32 = 256 | |
| llama_model_loader: - kv 21: general.file_type u32 = 32 | |
| llama_model_loader: - kv 22: qwen35.ssm.conv_kernel u32 = 4 | |
| llama_model_loader: - kv 23: qwen35.ssm.state_size u32 = 128 | |
| llama_model_loader: - kv 24: qwen35.ssm.group_count u32 = 16 | |
| llama_model_loader: - kv 25: qwen35.ssm.time_step_rank u32 = 48 | |
| llama_model_loader: - kv 26: qwen35.ssm.inner_size u32 = 6144 | |
| llama_model_loader: - kv 27: qwen35.attention.recurrent_layers arr[bool,64] = [true, true, true, false, true, true,... | |
| llama_model_loader: - kv 28: qwen35.full_attention_interval u32 = 4 | |
| llama_model_loader: - kv 29: qwen35.rope.dimension_count u32 = 64 | |
| llama_model_loader: - kv 30: qwen35.decision.type str = openjev | |
| llama_model_loader: - kv 31: qwen35.decision.temperature.choice f32 = 0.850000 | |
| llama_model_loader: - kv 32: qwen35.decision.temperature.score f32 = 0.850000 | |
| llama_model_loader: - kv 33: qwen35.decision.temperature.noul f32 = 1.554713 | |
| llama_model_loader: - kv 34: general.quantization_version u32 = 2 | |
| llama_model_loader: - kv 35: tokenizer.ggml.model str = gpt2 | |
| llama_model_loader: - kv 36: tokenizer.ggml.pre str = qwen35 | |
| llama_model_loader: - kv 37: tokenizer.ggml.tokens arr[str,248320] = ["!", "\"", "#", "$", "%", "&", "'", ... | |
| llama_model_loader: - kv 38: tokenizer.ggml.token_type arr[i32,248320] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ... | |
| llama_model_loader: - kv 39: tokenizer.ggml.merges arr[str,247587] = ["Δ Δ ", "Δ Δ Δ Δ ", "i n", "Δ t",... | |
| llama_model_loader: - kv 40: tokenizer.ggml.eos_token_id u32 = 248046 | |
| llama_model_loader: - kv 41: tokenizer.ggml.padding_token_id u32 = 248044 | |
| llama_model_loader: - kv 42: tokenizer.ggml.bos_token_id u32 = 248044 | |
| llama_model_loader: - kv 43: tokenizer.ggml.add_bos_token bool = false | |
| llama_model_loader: - kv 44: tokenizer.ggml.add_eos_token bool = false | |
| llama_model_loader: - kv 45: tokenizer.chat_template str = {%- set image_count = namespace(value... | |
| llama_model_loader: - kv 46: tokenizer.chat_template.systemone str = {% set letters = 'ABCDEFGHIJKLMNOPQRS... | |
| llama_model_loader: - kv 47: tokenizer.chat_templates arr[str,1] = ["systemone"] | |
| llama_model_loader: - type f32: 353 tensors | |
| llama_model_loader: - type bf16: 498 tensors | |
| llama_tensor_get_type: output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.0.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.0.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.0.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.0.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.0.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.1.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.1.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.1.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.1.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.1.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.2.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.2.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.2.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.2.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.2.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.3.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.3.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.3.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.3.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.4.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.4.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.4.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.4.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.4.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.5.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.5.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.5.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.5.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.5.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.6.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.6.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.6.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.6.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.6.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.7.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.7.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.7.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.7.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.8.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.8.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.8.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.8.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.8.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.9.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.9.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.9.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.9.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.9.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.10.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.10.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.10.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.10.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.10.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.11.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.11.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.11.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.11.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.12.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.12.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.12.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.12.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.12.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.13.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.13.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.13.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.13.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.13.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.14.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.14.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.14.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.14.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.14.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.15.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.15.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.15.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.15.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.16.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.16.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.16.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.16.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.16.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.17.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.17.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.17.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.17.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.17.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.18.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.18.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.18.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.18.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.18.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.19.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.19.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.19.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.19.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.20.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.20.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.20.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.20.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.20.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.21.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.21.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.21.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.21.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.21.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.22.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.22.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.22.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.22.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.22.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.23.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.23.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.23.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.23.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.24.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.24.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.24.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.24.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.24.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.25.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.25.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.25.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.25.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.25.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.26.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.26.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.26.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.26.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.26.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.27.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.27.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.27.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.27.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.28.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.28.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.28.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.28.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.28.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.29.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.29.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.29.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.29.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.29.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.30.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.30.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.30.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.30.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.30.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.31.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.31.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.31.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.31.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.32.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.32.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.32.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.32.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.32.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.33.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.33.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.33.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.33.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.33.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.34.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.34.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.34.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.34.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.34.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.35.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.35.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.35.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.35.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.36.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.36.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.36.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.36.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.36.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.37.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.37.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.37.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.37.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.37.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.38.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.38.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.38.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.38.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.38.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.39.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.39.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.39.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.39.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.40.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.40.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.40.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.40.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.40.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.41.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.41.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.41.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.41.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.41.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.42.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.42.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.42.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.42.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.42.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.43.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.43.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.43.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.43.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.44.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.44.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.44.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.44.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.44.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.45.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.45.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.45.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.45.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.45.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.46.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.46.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.46.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.46.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.46.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.47.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.47.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.47.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.47.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.48.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.48.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.48.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.48.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.48.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.49.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.49.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.49.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.49.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.49.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.50.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.50.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.50.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.50.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.50.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.51.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.51.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.51.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.51.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.52.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.52.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.52.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.52.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.52.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.53.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.53.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.53.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.53.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.53.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.54.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.54.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.54.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.54.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.54.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.55.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.55.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.55.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.55.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.56.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.56.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.56.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.56.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.56.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.57.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.57.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.57.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.57.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.57.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.58.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.58.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.58.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.58.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.58.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.59.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.59.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.59.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.59.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.60.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.60.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.60.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.60.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.60.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.61.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.61.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.61.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.61.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.61.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.62.attn_gate.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.62.attn_qkv.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.62.ssm_alpha.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.62.ssm_beta.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.62.ssm_out.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.63.attn_k.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.63.attn_output.weight - applying manual override: q4_K -> q6_K | |
| llama_tensor_get_type: blk.63.attn_q.weight - applying manual override: q4_K -> q8_0 | |
| llama_tensor_get_type: blk.63.attn_v.weight - applying manual override: q4_K -> q8_0 | |
| [ 1/ 851] output.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q6_K .. size = 2425.00 MiB -> 994.63 MiB | |
| [ 2/ 851] output_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 3/ 851] token_embd.weight - [ 5120, 248320, 1, 1], type = bf16, converting to q4_K .. size = 2425.00 MiB -> 682.03 MiB | |
| [ 4/ 851] blk.0.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 5/ 851] blk.0.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 6/ 851] blk.0.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 7/ 851] blk.0.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 8/ 851] blk.0.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 9/ 851] blk.0.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 10/ 851] blk.0.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 11/ 851] blk.0.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 12/ 851] blk.0.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 13/ 851] blk.0.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 14/ 851] blk.0.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 15/ 851] blk.0.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 16/ 851] blk.0.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 17/ 851] blk.0.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 18/ 851] blk.1.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 19/ 851] blk.1.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 20/ 851] blk.1.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 21/ 851] blk.1.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 22/ 851] blk.1.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 23/ 851] blk.1.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 24/ 851] blk.1.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 25/ 851] blk.1.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 26/ 851] blk.1.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 27/ 851] blk.1.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 28/ 851] blk.1.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 29/ 851] blk.1.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 30/ 851] blk.1.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 31/ 851] blk.1.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 32/ 851] blk.2.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 33/ 851] blk.2.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 34/ 851] blk.2.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 35/ 851] blk.2.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 36/ 851] blk.2.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 37/ 851] blk.2.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 38/ 851] blk.2.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 39/ 851] blk.2.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 40/ 851] blk.2.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 41/ 851] blk.2.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 42/ 851] blk.2.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 43/ 851] blk.2.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 44/ 851] blk.2.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 45/ 851] blk.2.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 46/ 851] blk.3.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 47/ 851] blk.3.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 48/ 851] blk.3.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 49/ 851] blk.3.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 50/ 851] blk.3.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 51/ 851] blk.3.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 52/ 851] blk.3.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 53/ 851] blk.3.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 54/ 851] blk.3.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 55/ 851] blk.3.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 56/ 851] blk.3.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 57/ 851] blk.4.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 58/ 851] blk.4.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 59/ 851] blk.4.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 60/ 851] blk.4.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 61/ 851] blk.4.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 62/ 851] blk.4.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 63/ 851] blk.4.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 64/ 851] blk.4.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 65/ 851] blk.4.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 66/ 851] blk.4.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 67/ 851] blk.4.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 68/ 851] blk.4.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 69/ 851] blk.4.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 70/ 851] blk.4.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 71/ 851] blk.5.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 72/ 851] blk.5.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 73/ 851] blk.5.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 74/ 851] blk.5.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 75/ 851] blk.5.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 76/ 851] blk.5.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 77/ 851] blk.5.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 78/ 851] blk.5.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 79/ 851] blk.5.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 80/ 851] blk.5.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 81/ 851] blk.5.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 82/ 851] blk.5.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 83/ 851] blk.5.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 84/ 851] blk.5.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 85/ 851] blk.6.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 86/ 851] blk.6.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 87/ 851] blk.6.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 88/ 851] blk.6.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 89/ 851] blk.6.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 90/ 851] blk.6.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 91/ 851] blk.6.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 92/ 851] blk.6.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 93/ 851] blk.6.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 94/ 851] blk.6.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 95/ 851] blk.6.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 96/ 851] blk.6.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 97/ 851] blk.6.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 98/ 851] blk.6.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 99/ 851] blk.7.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 100/ 851] blk.7.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 101/ 851] blk.7.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 102/ 851] blk.7.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 103/ 851] blk.7.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 104/ 851] blk.7.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 105/ 851] blk.7.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 106/ 851] blk.7.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 107/ 851] blk.7.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 108/ 851] blk.7.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 109/ 851] blk.7.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 110/ 851] blk.8.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 111/ 851] blk.8.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 112/ 851] blk.8.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 113/ 851] blk.8.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 114/ 851] blk.8.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 115/ 851] blk.8.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 116/ 851] blk.8.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 117/ 851] blk.8.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 118/ 851] blk.8.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 119/ 851] blk.8.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 120/ 851] blk.8.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 121/ 851] blk.8.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 122/ 851] blk.8.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 123/ 851] blk.8.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 124/ 851] blk.9.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 125/ 851] blk.9.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 126/ 851] blk.9.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 127/ 851] blk.9.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 128/ 851] blk.9.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 129/ 851] blk.9.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 130/ 851] blk.9.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 131/ 851] blk.9.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 132/ 851] blk.9.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 133/ 851] blk.9.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 134/ 851] blk.9.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 135/ 851] blk.9.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 136/ 851] blk.9.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 137/ 851] blk.9.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 138/ 851] blk.10.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 139/ 851] blk.10.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 140/ 851] blk.10.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 141/ 851] blk.10.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 142/ 851] blk.10.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 143/ 851] blk.10.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 144/ 851] blk.10.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 145/ 851] blk.10.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 146/ 851] blk.10.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 147/ 851] blk.10.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 148/ 851] blk.10.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 149/ 851] blk.10.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 150/ 851] blk.10.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 151/ 851] blk.10.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 152/ 851] blk.11.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 153/ 851] blk.11.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 154/ 851] blk.11.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 155/ 851] blk.11.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 156/ 851] blk.11.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 157/ 851] blk.11.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 158/ 851] blk.11.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 159/ 851] blk.11.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 160/ 851] blk.11.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 161/ 851] blk.11.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 162/ 851] blk.11.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 163/ 851] blk.12.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 164/ 851] blk.12.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 165/ 851] blk.12.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 166/ 851] blk.12.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 167/ 851] blk.12.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 168/ 851] blk.12.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 169/ 851] blk.12.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 170/ 851] blk.12.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 171/ 851] blk.12.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 172/ 851] blk.12.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 173/ 851] blk.12.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 174/ 851] blk.12.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 175/ 851] blk.12.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 176/ 851] blk.12.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 177/ 851] blk.13.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 178/ 851] blk.13.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 179/ 851] blk.13.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 180/ 851] blk.13.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 181/ 851] blk.13.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 182/ 851] blk.13.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 183/ 851] blk.13.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 184/ 851] blk.13.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 185/ 851] blk.13.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 186/ 851] blk.13.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 187/ 851] blk.13.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 188/ 851] blk.13.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 189/ 851] blk.13.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 190/ 851] blk.13.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 191/ 851] blk.14.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 192/ 851] blk.14.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 193/ 851] blk.14.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 194/ 851] blk.14.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 195/ 851] blk.14.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 196/ 851] blk.14.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 197/ 851] blk.14.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 198/ 851] blk.14.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 199/ 851] blk.14.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 200/ 851] blk.14.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 201/ 851] blk.14.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 202/ 851] blk.14.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 203/ 851] blk.14.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 204/ 851] blk.14.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 205/ 851] blk.15.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 206/ 851] blk.15.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 207/ 851] blk.15.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 208/ 851] blk.15.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 209/ 851] blk.15.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 210/ 851] blk.15.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 211/ 851] blk.15.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 212/ 851] blk.15.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 213/ 851] blk.15.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 214/ 851] blk.15.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 215/ 851] blk.15.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 216/ 851] blk.16.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 217/ 851] blk.16.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 218/ 851] blk.16.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 219/ 851] blk.16.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 220/ 851] blk.16.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 221/ 851] blk.16.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 222/ 851] blk.16.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 223/ 851] blk.16.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 224/ 851] blk.16.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 225/ 851] blk.16.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 226/ 851] blk.16.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 227/ 851] blk.16.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 228/ 851] blk.16.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 229/ 851] blk.16.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 230/ 851] blk.17.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 231/ 851] blk.17.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 232/ 851] blk.17.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 233/ 851] blk.17.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 234/ 851] blk.17.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 235/ 851] blk.17.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 236/ 851] blk.17.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 237/ 851] blk.17.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 238/ 851] blk.17.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 239/ 851] blk.17.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 240/ 851] blk.17.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 241/ 851] blk.17.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 242/ 851] blk.17.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 243/ 851] blk.17.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 244/ 851] blk.18.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 245/ 851] blk.18.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 246/ 851] blk.18.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 247/ 851] blk.18.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 248/ 851] blk.18.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 249/ 851] blk.18.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 250/ 851] blk.18.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 251/ 851] blk.18.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 252/ 851] blk.18.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 253/ 851] blk.18.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 254/ 851] blk.18.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 255/ 851] blk.18.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 256/ 851] blk.18.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 257/ 851] blk.18.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 258/ 851] blk.19.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 259/ 851] blk.19.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 260/ 851] blk.19.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 261/ 851] blk.19.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 262/ 851] blk.19.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 263/ 851] blk.19.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 264/ 851] blk.19.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 265/ 851] blk.19.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 266/ 851] blk.19.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 267/ 851] blk.19.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 268/ 851] blk.19.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 269/ 851] blk.20.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 270/ 851] blk.20.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 271/ 851] blk.20.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 272/ 851] blk.20.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 273/ 851] blk.20.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 274/ 851] blk.20.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 275/ 851] blk.20.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 276/ 851] blk.20.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 277/ 851] blk.20.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 278/ 851] blk.20.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 279/ 851] blk.20.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 280/ 851] blk.20.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 281/ 851] blk.20.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 282/ 851] blk.20.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 283/ 851] blk.21.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 284/ 851] blk.21.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 285/ 851] blk.21.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 286/ 851] blk.21.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 287/ 851] blk.21.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 288/ 851] blk.21.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 289/ 851] blk.21.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 290/ 851] blk.21.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 291/ 851] blk.21.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 292/ 851] blk.21.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 293/ 851] blk.21.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 294/ 851] blk.21.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 295/ 851] blk.21.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 296/ 851] blk.21.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 297/ 851] blk.22.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 298/ 851] blk.22.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 299/ 851] blk.22.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 300/ 851] blk.22.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 301/ 851] blk.22.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 302/ 851] blk.22.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 303/ 851] blk.22.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 304/ 851] blk.22.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 305/ 851] blk.22.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 306/ 851] blk.22.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 307/ 851] blk.22.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 308/ 851] blk.22.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 309/ 851] blk.22.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 310/ 851] blk.22.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 311/ 851] blk.23.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 312/ 851] blk.23.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 313/ 851] blk.23.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 314/ 851] blk.23.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 315/ 851] blk.23.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 316/ 851] blk.23.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 317/ 851] blk.23.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 318/ 851] blk.23.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 319/ 851] blk.23.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 320/ 851] blk.23.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 321/ 851] blk.23.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 322/ 851] blk.24.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 323/ 851] blk.24.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 324/ 851] blk.24.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 325/ 851] blk.24.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 326/ 851] blk.24.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 327/ 851] blk.24.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 328/ 851] blk.24.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 329/ 851] blk.24.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 330/ 851] blk.24.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 331/ 851] blk.24.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 332/ 851] blk.24.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 333/ 851] blk.24.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 334/ 851] blk.24.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 335/ 851] blk.24.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 336/ 851] blk.25.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 337/ 851] blk.25.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 338/ 851] blk.25.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 339/ 851] blk.25.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 340/ 851] blk.25.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 341/ 851] blk.25.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 342/ 851] blk.25.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 343/ 851] blk.25.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 344/ 851] blk.25.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 345/ 851] blk.25.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 346/ 851] blk.25.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 347/ 851] blk.25.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 348/ 851] blk.25.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 349/ 851] blk.25.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 350/ 851] blk.26.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 351/ 851] blk.26.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 352/ 851] blk.26.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 353/ 851] blk.26.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 354/ 851] blk.26.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 355/ 851] blk.26.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 356/ 851] blk.26.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 357/ 851] blk.26.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 358/ 851] blk.26.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 359/ 851] blk.26.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 360/ 851] blk.26.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 361/ 851] blk.26.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 362/ 851] blk.26.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 363/ 851] blk.26.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 364/ 851] blk.27.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 365/ 851] blk.27.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 366/ 851] blk.27.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 367/ 851] blk.27.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 368/ 851] blk.27.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 369/ 851] blk.27.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 370/ 851] blk.27.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 371/ 851] blk.27.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 372/ 851] blk.27.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 373/ 851] blk.27.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 374/ 851] blk.27.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 375/ 851] blk.28.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 376/ 851] blk.28.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 377/ 851] blk.28.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 378/ 851] blk.28.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 379/ 851] blk.28.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 380/ 851] blk.28.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 381/ 851] blk.28.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 382/ 851] blk.28.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 383/ 851] blk.28.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 384/ 851] blk.28.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 385/ 851] blk.28.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 386/ 851] blk.28.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 387/ 851] blk.28.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 388/ 851] blk.28.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 389/ 851] blk.29.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 390/ 851] blk.29.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 391/ 851] blk.29.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 392/ 851] blk.29.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 393/ 851] blk.29.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 394/ 851] blk.29.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 395/ 851] blk.29.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 396/ 851] blk.29.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 397/ 851] blk.29.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 398/ 851] blk.29.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 399/ 851] blk.29.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 400/ 851] blk.29.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 401/ 851] blk.29.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 402/ 851] blk.29.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 403/ 851] blk.30.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 404/ 851] blk.30.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 405/ 851] blk.30.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 406/ 851] blk.30.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 407/ 851] blk.30.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 408/ 851] blk.30.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 409/ 851] blk.30.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 410/ 851] blk.30.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 411/ 851] blk.30.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 412/ 851] blk.30.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 413/ 851] blk.30.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 414/ 851] blk.30.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 415/ 851] blk.30.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 416/ 851] blk.30.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 417/ 851] blk.31.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 418/ 851] blk.31.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 419/ 851] blk.31.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 420/ 851] blk.31.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 421/ 851] blk.31.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 422/ 851] blk.31.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 423/ 851] blk.31.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 424/ 851] blk.31.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 425/ 851] blk.31.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 426/ 851] blk.31.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 427/ 851] blk.31.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 428/ 851] blk.32.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 429/ 851] blk.32.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 430/ 851] blk.32.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 431/ 851] blk.32.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 432/ 851] blk.32.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 433/ 851] blk.32.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 434/ 851] blk.32.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 435/ 851] blk.32.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 436/ 851] blk.32.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 437/ 851] blk.32.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 438/ 851] blk.32.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 439/ 851] blk.32.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 440/ 851] blk.32.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 441/ 851] blk.32.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 442/ 851] blk.33.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 443/ 851] blk.33.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 444/ 851] blk.33.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 445/ 851] blk.33.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 446/ 851] blk.33.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 447/ 851] blk.33.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 448/ 851] blk.33.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 449/ 851] blk.33.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 450/ 851] blk.33.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 451/ 851] blk.33.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 452/ 851] blk.33.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 453/ 851] blk.33.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 454/ 851] blk.33.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 455/ 851] blk.33.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 456/ 851] blk.34.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 457/ 851] blk.34.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 458/ 851] blk.34.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 459/ 851] blk.34.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 460/ 851] blk.34.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 461/ 851] blk.34.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 462/ 851] blk.34.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 463/ 851] blk.34.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 464/ 851] blk.34.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 465/ 851] blk.34.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 466/ 851] blk.34.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 467/ 851] blk.34.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 468/ 851] blk.34.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 469/ 851] blk.34.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 470/ 851] blk.35.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 471/ 851] blk.35.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 472/ 851] blk.35.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 473/ 851] blk.35.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 474/ 851] blk.35.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 475/ 851] blk.35.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 476/ 851] blk.35.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 477/ 851] blk.35.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 478/ 851] blk.35.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 479/ 851] blk.35.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 480/ 851] blk.35.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 481/ 851] blk.36.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 482/ 851] blk.36.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 483/ 851] blk.36.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 484/ 851] blk.36.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 485/ 851] blk.36.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 486/ 851] blk.36.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 487/ 851] blk.36.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 488/ 851] blk.36.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 489/ 851] blk.36.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 490/ 851] blk.36.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 491/ 851] blk.36.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 492/ 851] blk.36.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 493/ 851] blk.36.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 494/ 851] blk.36.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 495/ 851] blk.37.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 496/ 851] blk.37.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 497/ 851] blk.37.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 498/ 851] blk.37.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 499/ 851] blk.37.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 500/ 851] blk.37.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 501/ 851] blk.37.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 502/ 851] blk.37.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 503/ 851] blk.37.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 504/ 851] blk.37.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 505/ 851] blk.37.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 506/ 851] blk.37.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 507/ 851] blk.37.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 508/ 851] blk.37.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 509/ 851] blk.38.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 510/ 851] blk.38.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 511/ 851] blk.38.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 512/ 851] blk.38.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 513/ 851] blk.38.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 514/ 851] blk.38.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 515/ 851] blk.38.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 516/ 851] blk.38.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 517/ 851] blk.38.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 518/ 851] blk.38.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 519/ 851] blk.38.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 520/ 851] blk.38.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 521/ 851] blk.38.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 522/ 851] blk.38.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 523/ 851] blk.39.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 524/ 851] blk.39.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 525/ 851] blk.39.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 526/ 851] blk.39.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 527/ 851] blk.39.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 528/ 851] blk.39.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 529/ 851] blk.39.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 530/ 851] blk.39.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 531/ 851] blk.39.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 532/ 851] blk.39.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 533/ 851] blk.39.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 534/ 851] blk.40.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 535/ 851] blk.40.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 536/ 851] blk.40.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 537/ 851] blk.40.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 538/ 851] blk.40.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 539/ 851] blk.40.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 540/ 851] blk.40.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 541/ 851] blk.40.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 542/ 851] blk.40.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 543/ 851] blk.40.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 544/ 851] blk.40.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 545/ 851] blk.40.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 546/ 851] blk.40.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 547/ 851] blk.40.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 548/ 851] blk.41.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 549/ 851] blk.41.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 550/ 851] blk.41.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 551/ 851] blk.41.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 552/ 851] blk.41.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 553/ 851] blk.41.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 554/ 851] blk.41.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 555/ 851] blk.41.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 556/ 851] blk.41.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 557/ 851] blk.41.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 558/ 851] blk.41.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 559/ 851] blk.41.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 560/ 851] blk.41.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 561/ 851] blk.41.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 562/ 851] blk.42.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 563/ 851] blk.42.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 564/ 851] blk.42.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 565/ 851] blk.42.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 566/ 851] blk.42.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 567/ 851] blk.42.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 568/ 851] blk.42.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 569/ 851] blk.42.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 570/ 851] blk.42.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 571/ 851] blk.42.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 572/ 851] blk.42.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 573/ 851] blk.42.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 574/ 851] blk.42.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 575/ 851] blk.42.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 576/ 851] blk.43.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 577/ 851] blk.43.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 578/ 851] blk.43.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 579/ 851] blk.43.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 580/ 851] blk.43.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 581/ 851] blk.43.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 582/ 851] blk.43.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 583/ 851] blk.43.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 584/ 851] blk.43.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 585/ 851] blk.43.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 586/ 851] blk.43.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 587/ 851] blk.44.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 588/ 851] blk.44.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 589/ 851] blk.44.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 590/ 851] blk.44.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 591/ 851] blk.44.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 592/ 851] blk.44.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 593/ 851] blk.44.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 594/ 851] blk.44.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 595/ 851] blk.44.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 596/ 851] blk.44.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 597/ 851] blk.44.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 598/ 851] blk.44.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 599/ 851] blk.44.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 600/ 851] blk.44.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 601/ 851] blk.45.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 602/ 851] blk.45.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 603/ 851] blk.45.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 604/ 851] blk.45.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 605/ 851] blk.45.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 606/ 851] blk.45.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 607/ 851] blk.45.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 608/ 851] blk.45.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 609/ 851] blk.45.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 610/ 851] blk.45.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 611/ 851] blk.45.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 612/ 851] blk.45.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 613/ 851] blk.45.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 614/ 851] blk.45.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 615/ 851] blk.46.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 616/ 851] blk.46.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 617/ 851] blk.46.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 618/ 851] blk.46.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 619/ 851] blk.46.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 620/ 851] blk.46.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 621/ 851] blk.46.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 622/ 851] blk.46.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 623/ 851] blk.46.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 624/ 851] blk.46.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 625/ 851] blk.46.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 626/ 851] blk.46.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 627/ 851] blk.46.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 628/ 851] blk.46.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 629/ 851] blk.47.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 630/ 851] blk.47.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 631/ 851] blk.47.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 632/ 851] blk.47.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 633/ 851] blk.47.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 634/ 851] blk.47.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 635/ 851] blk.47.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 636/ 851] blk.47.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 637/ 851] blk.47.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 638/ 851] blk.47.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 639/ 851] blk.47.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 640/ 851] blk.48.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 641/ 851] blk.48.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 642/ 851] blk.48.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 643/ 851] blk.48.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 644/ 851] blk.48.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 645/ 851] blk.48.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 646/ 851] blk.48.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 647/ 851] blk.48.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 648/ 851] blk.48.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 649/ 851] blk.48.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 650/ 851] blk.48.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 651/ 851] blk.48.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 652/ 851] blk.48.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 653/ 851] blk.48.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 654/ 851] blk.49.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 655/ 851] blk.49.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 656/ 851] blk.49.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 657/ 851] blk.49.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 658/ 851] blk.49.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 659/ 851] blk.49.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 660/ 851] blk.49.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 661/ 851] blk.49.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 662/ 851] blk.49.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 663/ 851] blk.49.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 664/ 851] blk.49.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 665/ 851] blk.49.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 666/ 851] blk.49.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 667/ 851] blk.49.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 668/ 851] blk.50.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 669/ 851] blk.50.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 670/ 851] blk.50.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 671/ 851] blk.50.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 672/ 851] blk.50.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 673/ 851] blk.50.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 674/ 851] blk.50.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 675/ 851] blk.50.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 676/ 851] blk.50.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 677/ 851] blk.50.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 678/ 851] blk.50.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 679/ 851] blk.50.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 680/ 851] blk.50.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 681/ 851] blk.50.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 682/ 851] blk.51.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 683/ 851] blk.51.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 684/ 851] blk.51.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 685/ 851] blk.51.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 686/ 851] blk.51.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 687/ 851] blk.51.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 688/ 851] blk.51.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 689/ 851] blk.51.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 690/ 851] blk.51.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 691/ 851] blk.51.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 692/ 851] blk.51.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 693/ 851] blk.52.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 694/ 851] blk.52.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 695/ 851] blk.52.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 696/ 851] blk.52.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 697/ 851] blk.52.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 698/ 851] blk.52.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 699/ 851] blk.52.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 700/ 851] blk.52.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 701/ 851] blk.52.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 702/ 851] blk.52.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 703/ 851] blk.52.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 704/ 851] blk.52.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 705/ 851] blk.52.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 706/ 851] blk.52.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 707/ 851] blk.53.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 708/ 851] blk.53.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 709/ 851] blk.53.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 710/ 851] blk.53.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 711/ 851] blk.53.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 712/ 851] blk.53.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 713/ 851] blk.53.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 714/ 851] blk.53.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 715/ 851] blk.53.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 716/ 851] blk.53.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 717/ 851] blk.53.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 718/ 851] blk.53.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 719/ 851] blk.53.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 720/ 851] blk.53.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 721/ 851] blk.54.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 722/ 851] blk.54.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 723/ 851] blk.54.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 724/ 851] blk.54.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 725/ 851] blk.54.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 726/ 851] blk.54.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 727/ 851] blk.54.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 728/ 851] blk.54.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 729/ 851] blk.54.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 730/ 851] blk.54.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 731/ 851] blk.54.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 732/ 851] blk.54.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 733/ 851] blk.54.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 734/ 851] blk.54.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 735/ 851] blk.55.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 736/ 851] blk.55.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 737/ 851] blk.55.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 738/ 851] blk.55.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 739/ 851] blk.55.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 740/ 851] blk.55.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 741/ 851] blk.55.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 742/ 851] blk.55.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 743/ 851] blk.55.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 744/ 851] blk.55.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 745/ 851] blk.55.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 746/ 851] blk.56.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 747/ 851] blk.56.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 748/ 851] blk.56.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 749/ 851] blk.56.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 750/ 851] blk.56.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 751/ 851] blk.56.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 752/ 851] blk.56.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 753/ 851] blk.56.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 754/ 851] blk.56.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 755/ 851] blk.56.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 756/ 851] blk.56.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 757/ 851] blk.56.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 758/ 851] blk.56.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 759/ 851] blk.56.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 760/ 851] blk.57.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 761/ 851] blk.57.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 762/ 851] blk.57.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 763/ 851] blk.57.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 764/ 851] blk.57.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 765/ 851] blk.57.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 766/ 851] blk.57.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 767/ 851] blk.57.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 768/ 851] blk.57.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 769/ 851] blk.57.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 770/ 851] blk.57.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 771/ 851] blk.57.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 772/ 851] blk.57.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 773/ 851] blk.57.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 774/ 851] blk.58.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 775/ 851] blk.58.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 776/ 851] blk.58.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 777/ 851] blk.58.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 778/ 851] blk.58.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 779/ 851] blk.58.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 780/ 851] blk.58.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 781/ 851] blk.58.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 782/ 851] blk.58.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 783/ 851] blk.58.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 784/ 851] blk.58.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 785/ 851] blk.58.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 786/ 851] blk.58.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 787/ 851] blk.58.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 788/ 851] blk.59.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 789/ 851] blk.59.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 790/ 851] blk.59.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 791/ 851] blk.59.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 792/ 851] blk.59.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 793/ 851] blk.59.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 794/ 851] blk.59.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 795/ 851] blk.59.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 796/ 851] blk.59.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 797/ 851] blk.59.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 798/ 851] blk.59.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 799/ 851] blk.60.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 800/ 851] blk.60.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 801/ 851] blk.60.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 802/ 851] blk.60.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 803/ 851] blk.60.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 804/ 851] blk.60.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 805/ 851] blk.60.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 806/ 851] blk.60.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 807/ 851] blk.60.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 808/ 851] blk.60.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 809/ 851] blk.60.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 810/ 851] blk.60.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 811/ 851] blk.60.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 812/ 851] blk.60.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 813/ 851] blk.61.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 814/ 851] blk.61.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 815/ 851] blk.61.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 816/ 851] blk.61.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 817/ 851] blk.61.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 818/ 851] blk.61.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 819/ 851] blk.61.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 820/ 851] blk.61.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 821/ 851] blk.61.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 822/ 851] blk.61.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 823/ 851] blk.61.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 824/ 851] blk.61.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 825/ 851] blk.61.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 826/ 851] blk.61.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 827/ 851] blk.62.attn_gate.weight - [ 5120, 6144, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 828/ 851] blk.62.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 829/ 851] blk.62.attn_qkv.weight - [ 5120, 10240, 1, 1], type = bf16, converting to q8_0 .. size = 100.00 MiB -> 53.12 MiB | |
| [ 830/ 851] blk.62.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 831/ 851] blk.62.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 832/ 851] blk.62.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 833/ 851] blk.62.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 834/ 851] blk.62.ssm_a - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 835/ 851] blk.62.ssm_alpha.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 836/ 851] blk.62.ssm_beta.weight - [ 5120, 48, 1, 1], type = bf16, converting to q8_0 .. size = 0.47 MiB -> 0.25 MiB | |
| [ 837/ 851] blk.62.ssm_conv1d.weight - [ 4, 10240, 1, 1], type = f32, size = 0.156 MiB | |
| [ 838/ 851] blk.62.ssm_dt.bias - [ 48, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 839/ 851] blk.62.ssm_norm.weight - [ 128, 1, 1, 1], type = f32, size = 0.000 MiB | |
| [ 840/ 851] blk.62.ssm_out.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q8_0 .. size = 60.00 MiB -> 31.88 MiB | |
| [ 841/ 851] blk.63.attn_k.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 842/ 851] blk.63.attn_k_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 843/ 851] blk.63.attn_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| [ 844/ 851] blk.63.attn_output.weight - [ 6144, 5120, 1, 1], type = bf16, converting to q6_K .. size = 60.00 MiB -> 24.61 MiB | |
| [ 845/ 851] blk.63.attn_q.weight - [ 5120, 12288, 1, 1], type = bf16, converting to q8_0 .. size = 120.00 MiB -> 63.75 MiB | |
| [ 846/ 851] blk.63.attn_q_norm.weight - [ 256, 1, 1, 1], type = f32, size = 0.001 MiB | |
| [ 847/ 851] blk.63.attn_v.weight - [ 5120, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 10.00 MiB -> 5.31 MiB | |
| [ 848/ 851] blk.63.ffn_down.weight - [ 17408, 5120, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 849/ 851] blk.63.ffn_gate.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 850/ 851] blk.63.ffn_up.weight - [ 5120, 17408, 1, 1], type = bf16, converting to q4_K .. size = 170.00 MiB -> 47.81 MiB | |
| [ 851/ 851] blk.63.post_attention_norm.weight - [ 5120, 1, 1, 1], type = f32, size = 0.020 MiB | |
| llama_model_quantize_impl: model size = 51305.09 MiB (16.00 BPW) | |
| llama_model_quantize_impl: quant size = 18084.41 MiB (5.64 BPW) | |
| llama_quantize: quantize time = 93084.12 ms | |
| llama_quantize: total time = 93084.12 ms | |
| + python3 llama.cpp/convert_hf_to_gguf.py ./model-temp-OpenJev-PRIMARY --outtype q8_0 --outfile ./upload-OpenJev/mmproj-OpenJev-Q8_0.gguf --mmproj --model-name OpenJev | |
| INFO:hf-to-gguf:Loading model: model-temp-OpenJev-PRIMARY | |
| INFO:hf-to-gguf:gguf: detected OpenJev checkpoint | |
| [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' | |
| INFO:hf-to-gguf:Model architecture: OpenJevModel | |
| INFO:hf-to-gguf:gguf: detected OpenJev checkpoint | |
| [transformers] Missing validation function in 'RotaryEmbeddingConfigMixin' for 'rope_type'='axial' | |
| INFO:hf-to-gguf:gguf: loading model weight map from 'model.safetensors.index.json' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00001-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00002-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00003-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00004-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00005-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00006-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00007-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00008-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00009-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00010-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00011-of-00012.safetensors' | |
| INFO:hf-to-gguf:gguf: indexing model part 'model-00012-of-00012.safetensors' | |
| INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only | |
| INFO:hf-to-gguf:Exporting model... | |
| INFO:hf-to-gguf:v.blk.0.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.0.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.0.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.0.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.0.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.0.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.0.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.0.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.0.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.1.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.1.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.1.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.1.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.1.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.1.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.1.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.1.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.10.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.10.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.10.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.10.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.10.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.10.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.10.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.10.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.11.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.11.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.11.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.11.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.11.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.11.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.11.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.11.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.12.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.12.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.12.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.12.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.12.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.12.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.12.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.12.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.13.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.13.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.13.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.13.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.13.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.13.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.13.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.13.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.14.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.14.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.14.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.14.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.14.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.14.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.14.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.14.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.15.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.15.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.15.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.15.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.15.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.15.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.15.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.15.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.16.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.16.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.16.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.16.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.16.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.16.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.16.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.16.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.17.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.17.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.17.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.17.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.17.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.17.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.17.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.17.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.18.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.18.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.18.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.18.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.18.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.18.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.18.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.18.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.19.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.19.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.19.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.19.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.19.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.19.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.19.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.19.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.2.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.2.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.2.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.2.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.2.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.2.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.2.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.2.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.20.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.20.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.20.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.20.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.20.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.20.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.20.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.20.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.21.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.21.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.21.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.21.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.21.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.21.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.21.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.21.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.22.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.22.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.22.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.22.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.22.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.22.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.22.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.22.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.23.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.23.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.23.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.23.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.23.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.23.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.23.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.23.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.24.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.24.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.24.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.24.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.24.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.24.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.24.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.24.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.25.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.25.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.25.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.25.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.25.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.25.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.25.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.25.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.26.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.26.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.26.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.26.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.26.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.26.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.26.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.26.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.3.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.3.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.3.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.3.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.3.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.3.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.3.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.3.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.4.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.4.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.4.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.4.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.4.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.4.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.4.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.4.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.5.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.5.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.5.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.5.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.5.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.5.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.5.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.5.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.6.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.6.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.6.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.6.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.6.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.6.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.6.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.6.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.7.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.7.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.7.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.7.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.7.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.7.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.7.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.7.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.8.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.8.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.8.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.8.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.8.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.8.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.8.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.8.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.attn_out.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.attn_out.weight, torch.bfloat16 --> Q8_0, shape = {1152, 1152} | |
| INFO:hf-to-gguf:v.blk.9.attn_qkv.bias, torch.bfloat16 --> F32, shape = {3456} | |
| INFO:hf-to-gguf:v.blk.9.attn_qkv.weight, torch.bfloat16 --> Q8_0, shape = {1152, 3456} | |
| INFO:hf-to-gguf:v.blk.9.ffn_up.bias, torch.bfloat16 --> F32, shape = {4304} | |
| INFO:hf-to-gguf:v.blk.9.ffn_up.weight, torch.bfloat16 --> Q8_0, shape = {1152, 4304} | |
| INFO:hf-to-gguf:v.blk.9.ffn_down.bias, torch.bfloat16 --> F32, shape = {1152} | |
| WARNING:hf-to-gguf:Can't quantize tensor with shape (1152, 4304) to Q8_0, falling back to F16 | |
| INFO:hf-to-gguf:v.blk.9.ffn_down.weight, torch.bfloat16 --> F16, shape = {4304, 1152} | |
| INFO:hf-to-gguf:v.blk.9.ln1.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.ln1.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.ln2.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.blk.9.ln2.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:mm.0.bias, torch.bfloat16 --> F32, shape = {4608} | |
| INFO:hf-to-gguf:mm.0.weight, torch.bfloat16 --> Q8_0, shape = {4608, 4608} | |
| INFO:hf-to-gguf:mm.2.bias, torch.bfloat16 --> F32, shape = {5120} | |
| INFO:hf-to-gguf:mm.2.weight, torch.bfloat16 --> Q8_0, shape = {4608, 5120} | |
| INFO:hf-to-gguf:v.post_ln.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.post_ln.weight, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.patch_embd.bias, torch.bfloat16 --> F32, shape = {1152} | |
| INFO:hf-to-gguf:v.patch_embd.weight, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152} | |
| INFO:hf-to-gguf:v.patch_embd.weight.1, torch.bfloat16 --> F32, shape = {16, 16, 3, 1152} | |
| INFO:hf-to-gguf:v.position_embd.weight, torch.bfloat16 --> F32, shape = {1152, 2304} | |
| INFO:hf-to-gguf:Set meta model | |
| INFO:hf-to-gguf:Set model parameters | |
| INFO:hf-to-gguf:Set model quantization version | |
| INFO:gguf.gguf_writer:Writing the following files: | |
| INFO:gguf.gguf_writer:upload-OpenJev/mmproj-OpenJev-Q8_0.gguf: n_tensors = 334, total_size = 629.2M | |
| Writing: 0%| | 0.00/629M [00:00<?, ?byte/s][A | |
| Writing: 42%|βββββ | 262M/629M [00:01<00:01, 262Mbyte/s][A | |
| Writing: 84%|βββββββββ | 528M/629M [00:02<00:00, 261Mbyte/s][A Writing: 100%|ββββββββββ| 629M/629M [00:02<00:00, 257Mbyte/s] | |
| INFO:hf-to-gguf:Model successfully exported to upload-OpenJev/mmproj-OpenJev-Q8_0.gguf | |
| + echo OpenJev-BF16.gguf | |
| + echo OpenJev-Q8_0.gguf | |
| + echo OpenJev-Q4_K_M.gguf | |
| + echo mmproj-OpenJev-BF16.gguf | |
| + echo mmproj-OpenJev-Q8_0.gguf | |