Image-to-Image
MLX
Diffusers
Safetensors
English
Chinese
apple-silicon
layer-decomposition
rgba
graphic-design
diffusion
ming-image
Instructions to use mlx-community/Ming-Image-0.1-Design-Layer-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Ming-Image-0.1-Design-Layer-bf16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download mlx-community/Ming-Image-0.1-Design-Layer-bf16 --local-dir Ming-Image-0.1-Design-Layer-bf16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Card: measured memory (S5), MLXEngine admission guidance, links to the 8-bit and 4-bit tiers
Browse files
README.md
CHANGED
|
@@ -19,6 +19,10 @@ Ming-Image-0.1-Design-Layer decomposes a flat design (poster, signage, card, UI)
|
|
| 19 |
- the background layer has the covered regions filled in;
|
| 20 |
- the model also returns its own recomposited frame.
|
| 21 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
It shares the architecture of [Ming-Image-0.1-Design](https://huggingface.co/mlx-community/Ming-Image-0.1-Design-bf16),
|
| 23 |
with the DiT trained for multi-frame output: one frame per layer plus the composite. The model conditions on the input
|
| 24 |
twice: as a VAE-encoded reference frame, and through the MLLM's vision tower.
|
|
@@ -72,13 +76,18 @@ Layer 4: Light sky-blue background with clouds, green hills and flowers.
|
|
| 72 |
|
| 73 |
Without a spec, "Decompose this image into N layers." also works, but a precise spec isolates text and objects better.
|
| 74 |
|
| 75 |
-
##
|
| 76 |
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
- 1024 bucket: 14 min, measured while the GPU was shared.
|
| 80 |
|
| 81 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
|
| 83 |
## Use (Swift / MLXEngine)
|
| 84 |
|
|
|
|
| 19 |
- the background layer has the covered regions filled in;
|
| 20 |
- the model also returns its own recomposited frame.
|
| 21 |
|
| 22 |
+
Quantized tiers: [8-bit](https://huggingface.co/mlx-community/Ming-Image-0.1-Design-Layer-8bit) (28.3 GB, within
|
| 23 |
+
0.1 dB of bf16, 48 GB Macs) and [4-bit](https://huggingface.co/mlx-community/Ming-Image-0.1-Design-Layer-4bit) (20.5 GB,
|
| 24 |
+
within 0.2 dB, 36 GB Macs).
|
| 25 |
+
|
| 26 |
It shares the architecture of [Ming-Image-0.1-Design](https://huggingface.co/mlx-community/Ming-Image-0.1-Design-bf16),
|
| 27 |
with the DiT trained for multi-frame output: one frame per layer plus the composite. The model conditions on the input
|
| 28 |
twice: as a VAE-encoded reference frame, and through the MLLM's vision tower.
|
|
|
|
| 76 |
|
| 77 |
Without a spec, "Decompose this image into N layers." also works, but a precise spec isolates text and objects better.
|
| 78 |
|
| 79 |
+
## Memory and speed (M5 Max)
|
| 80 |
|
| 81 |
+
Measured as process `phys_footprint`, with MLX's buffer cache capped at 2 GB (MLXEngine's default). The reference
|
| 82 |
+
profile is 12 steps at CFG 2.0.
|
|
|
|
| 83 |
|
| 84 |
+
- **Post-load resident:** 13.2 GB.
|
| 85 |
+
- **Peak:** 50.1–50.4 GB for 4 to 12 layers at the 1024 bucket, in the conditioning stage (the 34 GB MLLM loads,
|
| 86 |
+
conditions, and is released).
|
| 87 |
+
- **Under MLXEngine:** it declares 13.5 GB resident plus 44.7 GB activation, which needs a 96 GB Mac. Requests are
|
| 88 |
+
capped at 12 layers, the measured envelope. On 48 GB use the 8-bit tier, and on 36 GB the 4-bit tier.
|
| 89 |
+
- **Speed:** 131 s for 4 layers at the 512 bucket from a 2160×3840 still, and about 11 minutes at the 1024 bucket.
|
| 90 |
+
The output keeps the input's aspect ratio at the bucket's size.
|
| 91 |
|
| 92 |
## Use (Swift / MLXEngine)
|
| 93 |
|