Instructions to use MarwenBellili/canolaxray-qwen-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use MarwenBellili/canolaxray-qwen-lora with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("MarwenBellili/canolaxray-qwen-lora") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
CanolaXray-Qwen-LoRA β caption-conditioned adaptation of Qwen-Image to seed radiographs
A LoRA adapter for Qwen-Image (20 B multimodal diffusion transformer, flow matching) that synthesises soft X-ray radiographs of individual canola (Brassica napus) seeds at 1024 Γ 1024, conditioned on a natural-language caption derived automatically from measured damage morphology.
This is Approach 2 of a two-approach study. The counterpart is
canolaxray-conditional-DDPM, a 19.1 M-parameter conditional DDPM trained from
scratch on pixels and conditioned on a spatially registered damage map built from the same
measurements.
What it does
The methodological content sits in the captions, not the adapter. If every image in a class shares one prompt string, the conditional expectation the model learns is the class mean, and sampling returns a smoothed average rather than a plausible specimen. So each training image gets its own caption, rendered deterministically from measurements of that specimen:
- crack count, type, length, contrast, location and orientation
- cavity extent, seed-coat delamination, fragmentation
- texture and framing descriptors computed from the exported pixels
Each measured quantity maps through a fitted ordinal vocabulary into a phrase. Only the
severity phrase (no damage, low damage, medium damage, high damage) is held
lexically invariant within a grade, because that is the intended control variable. No
per-image annotation is required, so the procedure applies unchanged to a corpus of any size,
and the detector's own parameters are fitted without annotation by exploiting the fact that
the grades are ordered.
Model details
| Base model | Qwen-Image β frozen Qwen2.5-VL text encoder (7 B), frozen VAE (8Γ spatial, 16 latent channels), 20 B MMDiT |
| Adapted component | MMDiT only, via LoRA |
| Rank / alpha | r = 16, Ξ± = 8 (effective scale 0.5) |
| Trainable parameters | β 10β40 M |
| Resolution | 1024 Γ 1024 (latent 128 Γ 128 Γ 16, 4096 image tokens at patch factor 2) |
| Formulation | flow matching, straight-line path, velocity regression |
| Timestep sampling | shift, discrete_flow_shift = 2.2 |
| Learning rate | 5 Γ 10β»β΅ (1 Γ 10β»β΄ destabilises the adapter) |
| Precision | bf16, optional fp8 base quantisation |
| Memory | gradient checkpointing + block swap |
| Trainer | kohya-ss/musubi-tuner |
The adapter is saved in musubi-tuner's native LoRA format. Two practical notes carried over from the pipeline:
- Merge precision. Adapter weights are stored fp32, base weights bf16, so a naive merge promotes to fp32 and then meets bf16 activations at inference. Cast the adapter to bf16 before merging.
- Training and sampling must use the same
discrete_flow_shift.
Training data
The same 4180-image corpus and the same stratified 3553 / 627 split as the DDPM model, fixed by seed so the two approaches saw identical data.
| grade | total | train | held out |
|---|---|---|---|
| ND | 1565 | 1330 | 235 |
| LD | 990 | 842 | 148 |
| MD | 1092 | 928 | 164 |
| HD | 533 | 453 | 80 |
Images are exported at the training resolution rather than at the 256 px source resolution. The reason is the VAE: mean measured crack width is β 2.1 px at 256 px, which lands at β 0.26 latent pixels after 8Γ downsampling β below the representable scale, so the structure would be attenuated before the denoiser ever sees it. Unsharp masking preserves edge acutance through resampling.
Results (held out, 22β25 images per grade, 1024 Γ 1024)
Every nearest-neighbour measurement is reported against a real-to-real control, obtained by holding out real images and matching them against the rest of the corpus. The control, not 0 or 1, is the reference point.
Distributional. Per-grade FID: HD 0.012, ND 0.024, LD 0.068, MD 0.204; overall 0.057. No real-versus-real floor was computed at this resolution, so the per-grade ordering is the informative part rather than the absolute scale. Background luminance β€ 6 Γ 10β»βΈ.
Nearest-neighbour SSIM.
| grade | best | median | worst | real-to-real control |
|---|---|---|---|---|
| ND | 0.97 | 0.95 | 0.89 | 0.953 |
| LD | 0.96 | 0.94 | 0.88 | 0.951 |
| MD | 0.94 | 0.92 | 0.81 | 0.943 |
| HD | 0.93 | 0.91 | 0.89 | 0.907 |
Similarity falls with severity, but so does the control by a comparable amount, so the apparent ranking between grades largely dissolves once each grade is judged against its own reference. Residual concentrates in a bright ring at the seed-coat boundary (highest-gradient structure; sub-pixel misalignment against a physically distinct seed) with weaker filamentary structure along fissures β position and contrast approximately but not exactly right, which is the limitation a caption predicts: it can request a fissure of a given length in a given region, but not at a coordinate.
No evidence of memorisation. The memorisation gap (synthetic best-match similarity minus real-to-real control) is negative in all four grades: β0.001 HD, β0.049 MD, β0.009 LD, β0.002 ND.
Spectral fidelity. Power-law exponent 3.49β3.70, essentially matching real images, so overall spectral decay is correct. The deficit is confined to the top of the band: 75β85 % below real above β 0.7 of Nyquist, partially compensated by excess just below it. Crucially the deficit is grade-dependent in the reverse of what complexity would predict β HD only β 6 % below real above half-Nyquist, ND β 76 % below. In HD specimens high-frequency energy is genuine structure (crack edges, fragment boundaries) and is reproduced; in ND it is almost entirely detector noise. The model reproduces structure but omits noise.
Distribution match β the central negative finding.
- Feature-space neighbour overlap 58 % (HD), 59 % (MD), 63 % (LD), 63 % (ND), each p < 0.001, against a control of 91β94 % and chance of 92 %. Synthetic images cluster with each other rather than mixing into the real distribution.
- A two-sample classifier separates real from synthetic at ROC-AUC 0.991β1.000.
- PRDC coverage 0.08β0.19 at precision 0.52β0.88 β the mode-collapse corner: plausible samples from a narrow region. KID 0.044β0.070. Intensity histogram shift (Wasserstein on 0β255) 2.0 grey levels for HD, 8.7 MD, 5.6 LD, 4.2 ND.
The VAE is not the bottleneck. The obvious hypothesis was that the autoencoder's information floor explains the gap. It does not: real images that are simply encoded and decoded score 0.85β0.86 on the neighbour statistic, at or above the real control of 0.75β0.82, while synthetic images score 0.54β0.59. The gap is attributable to the adapter or the sampler. One qualification: Inception features are insensitive to the highest spatial frequencies, precisely where the spectral deficit lives, so the VAE may still limit high-frequency detail while contributing little to the feature-space gap. That test has not been run.
Downstream utility β the strongest practical result.
| regime | accuracy on real images |
|---|---|
| chance (4 balanced classes) | 0.250 |
| TSTR (train synthetic β test real) | 0.478 |
| TRTR (train real β test real) | 0.621 |
| real + synthetic | 0.641 |
A classifier that has never seen a real radiograph recovers a substantial share of the grading signal, so the synthetic images encode genuine class-discriminative damage structure rather than generic seed-shaped texture β this validates the measurement-derived caption design. At 77 % of TRTR, synthetic data remain clearly inferior to real data, consistent with the coverage result. The +0.020 augmentation gain is not statistically established: error bars overlap substantially (TRTR β 0.591β0.653, combined β 0.625β0.658). Confirming it needs repeated runs across seeds and splits with a paired test on per-fold differences.
Intended use
- Augmenting a canola-seed damage-grading corpus at megapixel resolution.
- Studying caption-conditioned adaptation of a general text-to-image prior to a narrow, non-photographic scientific domain.
Limitations and out-of-scope use
- No spatial control. A caption specifies severity and rough morphology, not location. If you need a fissure at a coordinate, use the DDPM counterpart.
- Narrow coverage. Coverage 0.08β0.19 means the synthetic set occupies a small part of the real distribution, however plausible individual images look. Per-image plausibility is not population-level resemblance.
- Perfectly separable from real images by a two-sample classifier. Do not treat these as distributionally equivalent to real acquisitions.
- Missing detector noise in low-severity grades. If a downstream task keys on grain statistics, this will show.
- Small evaluation samples (22β25 per grade). Large, consistent effects (spectral, coverage, separability) are reliable; small ones are not resolved.
- Not a diagnostic tool. Held-out validation and test partitions for any downstream classifier must consist exclusively of real acquisitions β evaluating on synthetic images generated from the same corpus measures agreement between two models, not agreement with ground truth.
- Detector parameters are fitted, not derived. Feature counts are a consistent relative measure across a corpus, not absolute morphometry, and structure typing is less reliable than structure counting.
Usage sketch
python src/musubi_tuner/qwen_image_generate_image.py \
--dit <path>/qwen_image_bf16.safetensors \
--vae <path>/qwen_image_vae.safetensors \
--text_encoder <path>/qwen_2.5_vl_7b.safetensors \
--lora_weight canolaxray_qwen.safetensors \
--lora_multiplier 1.0 \
--image_size 1024 1024 \
--discrete_flow_shift 2.2 \
--from_file prompts_HD.txt
Sampling prompts should be drawn from realised training captions, not written freehand, so inference operates on the conditioning distribution the adapter was fitted to. Disable block swap at inference; it is a training-memory device and streaming blocks from host RAM during generation is where the earlier empty-output failures came from.
Related artefacts
- Approach 1 (conditional DDPM):
MarwenBellili/canolaxray-conditional-DDPM - Dataset (real + both synthetic corpora):
10.5281/zenodo.22079214 - Code:
Marwenbellili72/CanolaXRay-Damage-Synthesis
Citation
@misc{bellili2026canolaxray,
title = {Generative Synthesis of Canola Seed X-Ray Radiographs: A Comparison of
LoRA Adaptation of a Flow-Matching Transformer and a Conditional
Denoising Diffusion Model},
author = {Bellili, Marwen},
note = {Supervised by Mohammad Nadimi},
year = {2026}
}
Base model: Qwen Team, Qwen-Image Technical Report, 2025. LoRA: Hu et al., arXiv:2106.09685, 2021.
- Downloads last month
- 31
Model tree for MarwenBellili/canolaxray-qwen-lora
Base model
Qwen/Qwen-Image