Running on <4gb Vram.

#104
by Jit2024 - opened

Hi, I found out a way of running this model on <4gb Vram. And I have not used any quantization.
Repo: https://github.com/Jit-Roy/WeeLLM

Relevant measurement for this approach β€” the wall it hits, and how to tell you've hit it.

I run H3 on 16 GiB of RAM (8 GB VRAM) and my working set is ~34.5 GiB (19.5 GiB model +
15.0 GiB text encoder), so I already live in the regime you're describing β€” just 4x less
severely. Two things from that may save you time:

1. The slowdown is not linear in time. Same graph, same resolution:

  • fresh session: 0.9–6 s/block
  • after ~1.7 h of accumulated page pressure: 212 s/block

So if your per-block time drifts upward the longer you run, that is the disk/swap path
doing it β€” not your swapping logic being wrong. On under 4 GB of RAM I'd expect that wall
much sooner than 1.7 h. The practical consequence on my machine was that restarting the
process between generations mattered more than any tuning
β€” periodic state reset, not a
faster kernel.

2. utilization.gpu cannot tell you whether you are computing or starved β€” it reads
99–100% in both states, and so does VRAM%. The metric that separates them is board power:

nvidia-smi --query-gpu=power.draw,utilization.gpu --format=csv -l 5

Computing β†’ power swings (roughly 60–100 W on my card). Starved β†’ power sits flat in the
mid-30s W
while utilisation still reads 99%. Across a full stalled session my median was
32.8 W over 534 samples.

That's a 5-minute measurement, and it tells you whether what's left is the swapping path or
something else β€” worth knowing before optimising anything.

Raw telemetry behind both numbers (534 samples, including the power trace):
https://huggingface.co/datasets/FlowForgeLabAi/minimax-h3-8gb-bench

Thanks, the power-vs-utilization point is useful. One question about your setup: does your offloading overlap transfers with compute (double-buffering, with block N+1 copying on a separate CUDA stream while block N computes)? In my approach, two blocks sit in VRAM at once, so if the copy time is below the compute time the GPU should never idle. Your data makes me think the stall is disk and pagefile reads rather than PCIe, since power is flat and mem-util is 3-5%. Did you measure per-block transfer time separately from compute time? And did you try a deeper pipeline where disk to RAM prefetches several blocks ahead of the RAM to VRAM stage? I'm curious whether overlap helps at all once the source is the pagefile.

Thanks β€” fair question, and the honest answer is that I did not measure transfer vs compute
separately. I'm not running a custom pipeline; it's ComfyUI's stock model management (with
--disable-pinned-memory on Windows), so there's no explicit double-buffering on my side to
describe.

But I have one number that bears directly on whether overlap helps once the source is the
pagefile β€” same graph, same resolution, one session:

fresh:            0.9-6 s/block
after ~1.7 h:     212 s/block

If what needed hiding were a roughly fixed PCIe transfer, per-block time would be roughly
constant. It isn't β€” it grows 30-200x within a single run, with no change to the graph. So
what degrades is the rate the bytes arrive at, not the overlap. You can hide a fixed transfer
behind compute; you can't hide a transfer whose cost keeps growing.

Your disk/pagefile read is consistent with that: my working set is ~34.5 GiB against 15.26 GiB
of RAM, so the resident set keeps getting evicted and the same bytes come off disk again.

So the measurement I'd suggest before adding pipeline depth: is your per-block time stable, or
drifting upward over a run? If it drifts, prefetch depth won't save it β€” the only two things
that helped me were reducing the bytes that must be re-read, and restarting the process between
generations.

For what it's worth, I got to 34.5 GiB using quantization (pruned INT8 ConvRot + an nvfp4 text
encoder). If that still collapsed on 15.26 GiB of RAM, a no-quantization path on under 4 GB is
going to be bandwidth-bound rather than pipeline-bound.

Thanks for the follow-up, and for actually engaging with the numbers.

I can't settle the per-block transfer question from my data, because I never separated transfer
from compute. What I have is block wall time, and it is not stable: same graph, same resolution,
0.9-6 s/block early in a session, 212 s/block after about 1.7 hours.

That is the part I would weigh before adding pipeline depth. Overlap hides a cost that holds
roughly still. Mine does not hold still β€” if the transfer rate itself degrades over a run,
deeper prefetch changes when the bytes move, not how fast they arrive.

One question, and it is the one that would change what I would measure next: how does WeeLLM
decide how many blocks stay resident, and how far ahead the disk-to-RAM stage runs? With two
blocks in VRAM, the overlap argument holds only while transfer time is below compute time, so
the design decision that matters is whether residency depth is fixed or adapts to the measured
transfer time.

it’s capacity-adaptive, not rate-adaptive. Diskβ†’RAM lookahead is min(6, usable_RAM // max_block_bytes), computed once at init from the RAM budget (minus a 2 GB safety margin and an attention-overhead estimate), with block sizes read from shard headers. In VRAM, the in-flight depth is at most one extra block (current + next). Whether that second buffer is allowed is a one-time prediction at the end of block 0 (reserved + block size + margin vs. budget), with an OOM fallback that turns it off. After the first full pass, leftover VRAM is used to pin blocks permanently, greedily in execution order, so those skip transfer entirely.

So you’re right that the overlap only holds while T_transfer < T_compute, and right now nothing measures or enforces that. Per-block disk-wait, H2D and compute times are logged, but depth isn’t adjusted from them. If transfer exceeds compute, the GPU stalls and you see it in the disk-wait numbers. Making depth and pin selection adapt to measured stall time is the obvious next step, and I haven’t done it yet.

WeeLLM logging per-block disk-wait separately is more than I have β€” I only have block wall time,
so I can't decompose it.

The part my data speaks to is exactly the thing you say isn't enforced. My stall time is not a
constant that an init-time calculation can capture. Same graph, same resolution, one session:
0.9-6 s/block early, 212 s/block after about 1.7 hours. Nothing about the graph changes across
that window. So whatever the init-time budget computes, the quantity it is implicitly treating
as stable is not stable β€” it degrades by two orders of magnitude.

That suggests the signal worth acting on is not capacity at init but the trend in the disk-wait
you already log. If per-block disk-wait rises while compute stays flat, the useful response is
to reduce how much gets re-read, or to reset the state that accumulated β€” rather than to change
the depth, which under a capacity rule is already as deep as the budget allows.

Whether your disk-wait series actually trends that way is the test, and you already log it.

Sign up or log in to comment