Running on <4gb Vram.

#104
by Jit2024 - opened

Hi, I found out a way of running this model on <4gb Vram. And I have not used any quantization.
Repo: https://github.com/Jit-Roy/WeeLLM

Relevant measurement for this approach β€” the wall it hits, and how to tell you've hit it.

I run H3 on 16 GiB of RAM (8 GB VRAM) and my working set is ~34.5 GiB (19.5 GiB model +
15.0 GiB text encoder), so I already live in the regime you're describing β€” just 4x less
severely. Two things from that may save you time:

1. The slowdown is not linear in time. Same graph, same resolution:

  • fresh session: 0.9–6 s/block
  • after ~1.7 h of accumulated page pressure: 212 s/block

So if your per-block time drifts upward the longer you run, that is the disk/swap path
doing it β€” not your swapping logic being wrong. On under 4 GB of RAM I'd expect that wall
much sooner than 1.7 h. The practical consequence on my machine was that restarting the
process between generations mattered more than any tuning
β€” periodic state reset, not a
faster kernel.

2. utilization.gpu cannot tell you whether you are computing or starved β€” it reads
99–100% in both states, and so does VRAM%. The metric that separates them is board power:

nvidia-smi --query-gpu=power.draw,utilization.gpu --format=csv -l 5

Computing β†’ power swings (roughly 60–100 W on my card). Starved β†’ power sits flat in the
mid-30s W
while utilisation still reads 99%. Across a full stalled session my median was
32.8 W over 534 samples.

That's a 5-minute measurement, and it tells you whether what's left is the swapping path or
something else β€” worth knowing before optimising anything.

Raw telemetry behind both numbers (534 samples, including the power trace):
https://huggingface.co/datasets/FlowForgeLabAi/minimax-h3-8gb-bench

Thanks, the power-vs-utilization point is useful. One question about your setup: does your offloading overlap transfers with compute (double-buffering, with block N+1 copying on a separate CUDA stream while block N computes)? In my approach, two blocks sit in VRAM at once, so if the copy time is below the compute time the GPU should never idle. Your data makes me think the stall is disk and pagefile reads rather than PCIe, since power is flat and mem-util is 3-5%. Did you measure per-block transfer time separately from compute time? And did you try a deeper pipeline where disk to RAM prefetches several blocks ahead of the RAM to VRAM stage? I'm curious whether overlap helps at all once the source is the pagefile.

Thanks β€” fair question, and the honest answer is that I did not measure transfer vs compute
separately. I'm not running a custom pipeline; it's ComfyUI's stock model management (with
--disable-pinned-memory on Windows), so there's no explicit double-buffering on my side to
describe.

But I have one number that bears directly on whether overlap helps once the source is the
pagefile β€” same graph, same resolution, one session:

fresh:            0.9-6 s/block
after ~1.7 h:     212 s/block

If what needed hiding were a roughly fixed PCIe transfer, per-block time would be roughly
constant. It isn't β€” it grows 30-200x within a single run, with no change to the graph. So
what degrades is the rate the bytes arrive at, not the overlap. You can hide a fixed transfer
behind compute; you can't hide a transfer whose cost keeps growing.

Your disk/pagefile read is consistent with that: my working set is ~34.5 GiB against 15.26 GiB
of RAM, so the resident set keeps getting evicted and the same bytes come off disk again.

So the measurement I'd suggest before adding pipeline depth: is your per-block time stable, or
drifting upward over a run? If it drifts, prefetch depth won't save it β€” the only two things
that helped me were reducing the bytes that must be re-read, and restarting the process between
generations.

For what it's worth, I got to 34.5 GiB using quantization (pruned INT8 ConvRot + an nvfp4 text
encoder). If that still collapsed on 15.26 GiB of RAM, a no-quantization path on under 4 GB is
going to be bandwidth-bound rather than pipeline-bound.

Sign up or log in to comment