xocialize commited on
Commit
7fec4c9
·
verified ·
1 Parent(s): 3f5fd6b

Card: measured memory (S5), MLXEngine admission guidance, links to the 8-bit and 4-bit tiers

Browse files
Files changed (1) hide show
  1. README.md +14 -5
README.md CHANGED
@@ -19,6 +19,10 @@ Ming-Image-0.1-Design-Layer decomposes a flat design (poster, signage, card, UI)
19
  - the background layer has the covered regions filled in;
20
  - the model also returns its own recomposited frame.
21
 
 
 
 
 
22
  It shares the architecture of [Ming-Image-0.1-Design](https://huggingface.co/mlx-community/Ming-Image-0.1-Design-bf16),
23
  with the DiT trained for multi-frame output: one frame per layer plus the composite. The model conditions on the input
24
  twice: as a VAE-encoded reference frame, and through the MLLM's vision tower.
@@ -72,13 +76,18 @@ Layer 4: Light sky-blue background with clouds, green hills and flowers.
72
 
73
  Without a spec, "Decompose this image into N layers." also works, but a precise spec isolates text and objects better.
74
 
75
- ## Performance (M5 Max, bf16)
76
 
77
- The reference profile is 12 steps at CFG 2.0.
78
- - 512 bucket, 4 layers from a 2160×3840 still: about 2.3 min.
79
- - 1024 bucket: 14 min, measured while the GPU was shared.
80
 
81
- The output keeps the input's aspect ratio at the bucket's size.
 
 
 
 
 
 
82
 
83
  ## Use (Swift / MLXEngine)
84
 
 
19
  - the background layer has the covered regions filled in;
20
  - the model also returns its own recomposited frame.
21
 
22
+ Quantized tiers: [8-bit](https://huggingface.co/mlx-community/Ming-Image-0.1-Design-Layer-8bit) (28.3 GB, within
23
+ 0.1 dB of bf16, 48 GB Macs) and [4-bit](https://huggingface.co/mlx-community/Ming-Image-0.1-Design-Layer-4bit) (20.5 GB,
24
+ within 0.2 dB, 36 GB Macs).
25
+
26
  It shares the architecture of [Ming-Image-0.1-Design](https://huggingface.co/mlx-community/Ming-Image-0.1-Design-bf16),
27
  with the DiT trained for multi-frame output: one frame per layer plus the composite. The model conditions on the input
28
  twice: as a VAE-encoded reference frame, and through the MLLM's vision tower.
 
76
 
77
  Without a spec, "Decompose this image into N layers." also works, but a precise spec isolates text and objects better.
78
 
79
+ ## Memory and speed (M5 Max)
80
 
81
+ Measured as process `phys_footprint`, with MLX's buffer cache capped at 2 GB (MLXEngine's default). The reference
82
+ profile is 12 steps at CFG 2.0.
 
83
 
84
+ - **Post-load resident:** 13.2 GB.
85
+ - **Peak:** 50.1–50.4 GB for 4 to 12 layers at the 1024 bucket, in the conditioning stage (the 34 GB MLLM loads,
86
+ conditions, and is released).
87
+ - **Under MLXEngine:** it declares 13.5 GB resident plus 44.7 GB activation, which needs a 96 GB Mac. Requests are
88
+ capped at 12 layers, the measured envelope. On 48 GB use the 8-bit tier, and on 36 GB the 4-bit tier.
89
+ - **Speed:** 131 s for 4 layers at the 512 bucket from a 2160×3840 still, and about 11 minutes at the 1024 bucket.
90
+ The output keeps the input's aspect ratio at the bucket's size.
91
 
92
  ## Use (Swift / MLXEngine)
93