Qwen-Image-2.1 FG Extract LoRA

A rank-32 LoRA that extracts one element from an image as an RGBA cutout at the same size and position. You give it a rough mask hint and it copies the element pixel for pixel, with no redrawing or moving.

It was trained for and runs on the distilled turbo model Viggle/Qwen-Image-2.1-viggle-turbo at 4 steps with no CFG. That takes about 2 s per extraction at 512². The text encoder and VAE come from Qwen/Qwen-Image-2.1. It is meant to be used with an alpha-capable VAE (see trmz/flux2-klein-alpha).

Each pair shows the input (the composite with the mask applied, so everything outside the hint is dimmed) and the RGBA foreground the model returns:

preview 1 preview 2

Files

file notes
qwen-image-2.1-fg-extract-v4.safetensors Recommended. 1500 steps, trained on per-element instructions (prompt_style=instruct).
qwen-image-2.1-fg-extract-v3.safetensors 1000 steps on a fixed prompt. Works well as a fallback.

Keys are in diffusers format (transformer.<module>.lora_A/B.weight), with alpha folded into B.

How to use

⚠️ Only the weights are provided here, with no pipeline or mask code. You have to build the mask conditioning yourself (Picture 1 below) and supply the mask yourself, e.g. from SAM, a box or a detector. The model does not find the object for you.

  1. Load Viggle/Qwen-Image-2.1-viggle-turbo with the text encoder and VAE from Qwen/Qwen-Image-2.1 (use an alpha VAE if you want RGBA output), then load the LoRA:
    pipe.load_lora_weights("trmz/qwen-image-2.1-fg-extract-lora", weight_name="qwen-image-2.1-fg-extract-v4.safetensors")
    
  2. Build the two condition images from your image and a mask of the element:
    import numpy as np
    from PIL import Image, ImageFilter
    
    def dimmed(img, mask, dim=0.2, grow=2):
        a = np.asarray(img.convert("RGB")).astype(np.float32)
        m = mask.convert("L").filter(ImageFilter.MaxFilter(2 * grow + 1))
        m = (np.asarray(m) > 127)[..., None]
        return Image.fromarray((a * m + a * (1 - m) * dim).astype(np.uint8))
    
    images = [dimmed(img, mask), img.convert("RGB")]   # Picture 1, Picture 2
    
  3. Run it with the prompt below at 4 steps, CFG off, and an output size equal to the input size. The result is the element on a transparent canvas, in place.

Inputs

  • Picture 1: the image with everything outside the element's mask dimmed to 20%.
  • Picture 2: the original image.
  • The mask can be rough (too small, too big or shifted). Don't pass a black-and-white mask as an image, because the model copies it into the output.

Prompt for v4 ({el} is the element name, e.g. "pink flowers"):

Extract the {el} from Picture 2 and output it alone as an RGBA image with a transparent background. Picture 1 is Picture 2 with everything except the {el} darkened. The bright area is only a rough hint: it can be too small, too big or shifted, so trust this description over the highlight and keep the whole {el}, no more and no less. Copy the {el} pixel for pixel: same position, same size, same colours, including its border, shading and any symbol or text that belongs to it. Do not redraw, restyle, move or resize it. Everything else, including the background and any neighbouring elements, must become fully transparent. The image has alpha channel and the background is transparent.

Prompt for v3:

Extract the highlighted element from Picture 2 and output it alone as an RGBA image with a transparent background, pixel-identical and at the same position and size. The image has alpha channel and the background is transparent.

Recommended settings are 4 steps on the Viggle turbo. At 512² an extraction takes about 2 s. 1024² gives slightly cleaner edges but takes about 9 s.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for trmz/qwen-image-2.1-fg-extract-lora

Adapter
(1)
this model