Instructions to use MiniMaxAI/MiniMax-H3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use MiniMaxAI/MiniMax-H3 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
Running on <4gb Vram.
Hi, I found out a way of running this model on <4gb Vram. And I have not used any quantization.
Repo: https://github.com/Jit-Roy/WeeLLM
Relevant measurement for this approach β the wall it hits, and how to tell you've hit it.
I run H3 on 16 GiB of RAM (8 GB VRAM) and my working set is ~34.5 GiB (19.5 GiB model +
15.0 GiB text encoder), so I already live in the regime you're describing β just 4x less
severely. Two things from that may save you time:
1. The slowdown is not linear in time. Same graph, same resolution:
- fresh session: 0.9β6 s/block
- after ~1.7 h of accumulated page pressure: 212 s/block
So if your per-block time drifts upward the longer you run, that is the disk/swap path
doing it β not your swapping logic being wrong. On under 4 GB of RAM I'd expect that wall
much sooner than 1.7 h. The practical consequence on my machine was that restarting the
process between generations mattered more than any tuning β periodic state reset, not a
faster kernel.
2. utilization.gpu cannot tell you whether you are computing or starved β it reads
99β100% in both states, and so does VRAM%. The metric that separates them is board power:
nvidia-smi --query-gpu=power.draw,utilization.gpu --format=csv -l 5
Computing β power swings (roughly 60β100 W on my card). Starved β power sits flat in the
mid-30s W while utilisation still reads 99%. Across a full stalled session my median was
32.8 W over 534 samples.
That's a 5-minute measurement, and it tells you whether what's left is the swapping path or
something else β worth knowing before optimising anything.
Raw telemetry behind both numbers (534 samples, including the power trace):
https://huggingface.co/datasets/FlowForgeLabAi/minimax-h3-8gb-bench
Thanks, the power-vs-utilization point is useful. One question about your setup: does your offloading overlap transfers with compute (double-buffering, with block N+1 copying on a separate CUDA stream while block N computes)? In my approach, two blocks sit in VRAM at once, so if the copy time is below the compute time the GPU should never idle. Your data makes me think the stall is disk and pagefile reads rather than PCIe, since power is flat and mem-util is 3-5%. Did you measure per-block transfer time separately from compute time? And did you try a deeper pipeline where disk to RAM prefetches several blocks ahead of the RAM to VRAM stage? I'm curious whether overlap helps at all once the source is the pagefile.
Thanks β fair question, and the honest answer is that I did not measure transfer vs compute
separately. I'm not running a custom pipeline; it's ComfyUI's stock model management (with
--disable-pinned-memory on Windows), so there's no explicit double-buffering on my side to
describe.
But I have one number that bears directly on whether overlap helps once the source is the
pagefile β same graph, same resolution, one session:
fresh: 0.9-6 s/block
after ~1.7 h: 212 s/block
If what needed hiding were a roughly fixed PCIe transfer, per-block time would be roughly
constant. It isn't β it grows 30-200x within a single run, with no change to the graph. So
what degrades is the rate the bytes arrive at, not the overlap. You can hide a fixed transfer
behind compute; you can't hide a transfer whose cost keeps growing.
Your disk/pagefile read is consistent with that: my working set is ~34.5 GiB against 15.26 GiB
of RAM, so the resident set keeps getting evicted and the same bytes come off disk again.
So the measurement I'd suggest before adding pipeline depth: is your per-block time stable, or
drifting upward over a run? If it drifts, prefetch depth won't save it β the only two things
that helped me were reducing the bytes that must be re-read, and restarting the process between
generations.
For what it's worth, I got to 34.5 GiB using quantization (pruned INT8 ConvRot + an nvfp4 text
encoder). If that still collapsed on 15.26 GiB of RAM, a no-quantization path on under 4 GB is
going to be bandwidth-bound rather than pipeline-bound.