recoilme's picture
comfyui/README: add a from-scratch quickstart
26a3064
|
Raw History Blame Contribute Delete
3.94 kB

ComfyUI workflow for the Qwen3-0.6B adapter

This folder contains a ComfyUI workflow that drives FLUX.2-klein-4B with the Qwen3-0.6B + adapter text encoder instead of its native Qwen3-4B encoder.

Quickstart (from scratch)

From an empty machine to a first image β€” verified end to end on ComfyUI 0.36 (~8.4 GiB of downloads, ~1.4 GiB more on the first run):

# 1. ComfyUI itself β€” skip if you already have one
git clone https://github.com/comfyanonymous/ComfyUI
cd ComfyUI
python -m venv .venv && . .venv/bin/activate       # Python 3.10+
pip install -r requirements.txt

# 2. the custom node that provides the two KleinAdapter nodes
cd custom_nodes
git clone https://github.com/recoilme/klein-qwen3-adapter-comfyui
cd ..

# 3. models (paths are relative to ComfyUI/)
mkdir -p models/klein_adapter
curl -L -o models/diffusion_models/flux-2-klein-4b.safetensors \
  https://huggingface.co/Comfy-Org/flux2-klein/resolve/main/split_files/diffusion_models/flux-2-klein-4b.safetensors
curl -L -o models/vae/flux2-vae.safetensors \
  https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/vae/flux2-vae.safetensors
curl -L -o models/klein_adapter/adapter_v14_bal.safetensors \
  https://huggingface.co/AiArtLab/qwen3-0.6b-4b-adapter/resolve/main/adapter_v14_bal.safetensors
# Qwen3-0.6B (1.4 GiB) needs no manual step: it is pulled from the Hub on the first
# run. Offline machine: fetch it beforehand with `hf download Qwen/Qwen3-0.6B`.

# 4. start ComfyUI
python main.py                                      # then open the URL it prints (:8188)

# 5. in the UI: Workflow -> Open ->
#    custom_nodes/klein-qwen3-adapter-comfyui/workflows/flux2_klein_qwen3_06b_adapter.json
#    (or drag this folder's JSON onto the canvas), type a prompt, press Run.

768Γ—1280 with the distilled model takes ~2 s per image on an RTX 5090 and peaks at ~12.6 GiB of VRAM (mostly the klein DiT itself).

Models

ComfyUI/models/
β”œβ”€β”€ diffusion_models/
β”‚   └── flux-2-klein-4b.safetensors        # distilled klein DiT (stock), 7.2 GiB
β”œβ”€β”€ vae/
β”‚   └── flux2-vae.safetensors              # stock Flux2 VAE (untouched), 321 MiB
└── klein_adapter/                          # created automatically by the node
    └── adapter_v14_bal.safetensors        # this adapter, 840 MiB
  • flux-2-klein-4b.safetensors: the klein DiT in ComfyUI format, from Comfy-Org/flux2-klein (split_files/diffusion_models/). The base model flux-2-klein-base-4b.safetensors works too (see the settings below).
  • flux2-vae.safetensors: from Comfy-Org/flux2-dev (split_files/vae/).
  • adapter_v14_bal.safetensors: from this repo (see above).
  • Qwen3-0.6B is downloaded automatically from HF on first use.

Run

  1. Load flux2_klein_qwen3_06b_adapter.json in ComfyUI.
  2. Set your prompt in the Positive prompt node.
  3. Queue.

Settings in the workflow: distilled klein, 4 steps, guidance 1.0, euler, 768Γ—1280. For the base model use flux-2-klein-base-4b.safetensors, 50 steps, guidance 4.0.

Notes

The node computes the text conditioning identically to example.py in this repo (same chat template, same layer taps 2,9,14,18,23,27, same fp32 adapter, same drop_first=5) β€” checked numerically, the two tensors are bit-equal ((1, 251, 7680), max abs diff 0.0). The image itself is then produced by ComfyUI's own sampler (its noise, scheduler, text-position ids and 512-token padding of the conditioning), so it is not pixel-identical to diffusers β€” the same framework difference you'd see with the native klein encoder.