This model is a straight conversion of the original

Here is the code I used to convert it. Tried it in the latest comfyui and it works. See the included workflow file for an example. On my 3090, and using the lightvaew2_1 autoencoder it takes 25 seconds to run the workflow.

import torch
from safetensors.torch import save_file

SRC = '/home/user/convert model/model_iter6000.pt'
DST = '/home/user/convert model/wan2.1-t2v-1.3b-4step-distill.safetensors'

sd = torch.load(SRC, map_location='cpu')

def conv(k):
    if k.startswith('blocks.'):
        idx, rest = k[len('blocks.'):].split('.', 1)
        if rest.startswith('attn1.'):
            s = rest[len('attn1.'):].replace('to_out.0', 'o').replace('to_q', 'q').replace('to_k', 'k').replace('to_v', 'v')
            return f'blocks.{idx}.self_attn.{s}'
        if rest.startswith('attn2.'):
            s = rest[len('attn2.'):].replace('to_out.0', 'o').replace('to_q', 'q').replace('to_k', 'k').replace('to_v', 'v')
            return f'blocks.{idx}.cross_attn.{s}'
        if rest.startswith('ffn.net.0.proj.'):
            return f'blocks.{idx}.ffn.0.' + rest[len('ffn.net.0.proj.'):]
        if rest.startswith('ffn.net.2.'):
            return f'blocks.{idx}.ffn.2.' + rest[len('ffn.net.2.'):]
        if rest.startswith('norm2.'):
            return f'blocks.{idx}.norm3.' + rest[len('norm2.'):]
        if rest == 'scale_shift_table':
            return f'blocks.{idx}.modulation'
        return None
    if k in ('patch_embedding.weight', 'patch_embedding.bias'):
        return k
    if k == 'proj_out.weight':
        return 'head.head.weight'
    if k == 'proj_out.bias':
        return 'head.head.bias'
    if k == 'scale_shift_table':
        return 'head.modulation'
    if k.startswith('condition_embedder.time_embedder.linear_1.'):
        return 'time_embedding.0.' + k.split('.')[-1]
    if k.startswith('condition_embedder.time_embedder.linear_2.'):
        return 'time_embedding.2.' + k.split('.')[-1]
    if k.startswith('condition_embedder.time_proj.'):
        return 'time_projection.1.' + k.split('.')[-1]
    if k.startswith('condition_embedder.text_embedder.linear_1.'):
        return 'text_embedding.0.' + k.split('.')[-1]
    if k.startswith('condition_embedder.text_embedder.linear_2.'):
        return 'text_embedding.2.' + k.split('.')[-1]
    return None

out = {}
for k, v in sd.items():
    nk = conv(k)
    if nk is None:
        raise SystemExit(f'UNMAPPED: {k}')
    out[nk] = v.contiguous()

save_file(out, DST, metadata={'format': 'pt', 'framework': 'pt'})
print(f'Saved {len(out)} tensors -> {DST}')

GGUFs made with city96's stuff

Q3_K_S is for the desperate

Models from Q3_K_M upward start looking 'fine'.

The ggufs work great with comfyui. They are also compatible with leejet's stable-diffusion.cpp, but theyre buggy, video outputs seem messed up somewhat, and ram usage spikes on VAE decode to like 60 gb+ on a 5 second video using the TAE, if you use VAE and tile it its fine and vram usage is low. Might look into it later.

Downloads last month
248
GGUF
Model size
1B params
Architecture
wan
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for anonymous1436778242134213535654/Wan2.1-T2V-1.3B-Distill-Comfy

Quantized
(1)
this model