GLM-5.3-Flash on Single MI300 GPU

#34
by ghostplant - opened

A perfect model whose size exactly fits in a single GPU, with about 100+ tps without MTP, and 200+ tps with MTP.

docker run -e LOCAL_SIZE=1 -p 8000:8000 -it --rm --ipc=host --shm-size=8g \
      --ulimit memlock=-1 --ulimit stack=67108864 -v /:/host -w /host$(pwd) \
      --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --device=/dev/kfd --device=/dev/dri --group-add=video \
      tutelgroup/deepseek-671b:mi300x8-chat-20260929 --serve=core \
        --try_path nvidia/GLM-5.3-Flash-NVFP4 \
        --thinking_effort max

this is not single GPU dude

single MI300 card (not machine) is enough. Do you encounter some troubles in this case (-e LOCAL_SIZE=1)?

single MI300 GPU is of 192 GB VRAM, the model itself is just 328 GB, it wouldn't even load

single MI300 GPU is of 192 GB VRAM, the model itself is just 328 GB, it wouldn't even load

It can be quantized to run using under 176GB.

Sign up or log in to comment