Instructions to use pinecoresystems/Ming-Image-0.1-Design with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use pinecoresystems/Ming-Image-0.1-Design with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("pinecoresystems/Ming-Image-0.1-Design", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Download UPSTREAM-README.md from pinecoresystems/Ming-Image-0.1-Design: direct link, hf CLI and curl.
- Browser
- Download file 3.33 kB
-
https://huggingface.co/pinecoresystems/Ming-Image-0.1-Design/resolve/313f6c19e12406e6dfdcafe18f21f880f4871273/UPSTREAM-README.md
- Command line
-
hf download hf://pinecoresystems/Ming-Image-0.1-Design@313f6c19e12406e6dfdcafe18f21f880f4871273/UPSTREAM-README.md
-
curl -L -o UPSTREAM-README.md https://huggingface.co/pinecoresystems/Ming-Image-0.1-Design/resolve/313f6c19e12406e6dfdcafe18f21f880f4871273/UPSTREAM-README.md
license: mit
library_name: custom
pipeline_tag: text-to-image
inference: false
tags:
- text-to-image
- image-generation
- graphic-design
- text-rendering
- rgba
Ming-Image-0.1-Design
🧩 ModelScope · 🤗 Hugging Face · 📄 Blog · 🖥️ Demo
🎨 Design Skill · 📊 PPT Skill
Ming-Image-0.1-Design is a 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs. It generates complete visual compositions and supports RGBA output with transparent backgrounds.
UI/UX Design leaderboard
Quick Start
Use the companion Ming-Image repository for installation and inference:
git clone https://github.com/inclusionAI/Ming-Image
cd Ming-Image
pip install -r requirements.txt
python infer.py \
--model inclusionAI/Ming-Image-0.1-Design \
--task text-to-image \
--prompt assets/t2i_four_seasons_cabin_prompt.json \
--resolution 2048 \
--output-dir outputs/t2i
Prompt enhancement (PE) can use Ling-3.0-flash-VL or qwen3.8-27B; see
text-to-image prompt rewriting.
Transparent-background generation
For transparent-background generation, prepend exactly one of the recommended RGBA phrases. See the transparent-background generation tip.
Deployment
We recommend the following inference frameworks to serve the model:
- vLLM-Omni: see the recipes and installation guide.
Recommended settings
- Resolution: 2048 x 2048 (recommended), or 1024 x 1024 for faster generation.
- Sampling steps: 12.
- CFG scale: 1.0.
- Precision: BF16.
- Hardware: one CUDA GPU with 80 GiB VRAM (validated configuration).
The public inference code maps text-to-image resolution requests to the supported 1024 or 2048 bucket.
Gallery
Text-to-image
Transparent-background text-to-image
The checkerboard is used only to preview transparency; it is not part of the generated RGBA images.
License
This model is released under the MIT License.