PLE layer in future versions?

#56
by NicSir - opened

Hi, this is my favorite model for several use cases, but with 128gb the max i can run is the q3-xxs from unsloth or similar.
The ple offload would allow us to load a higher quant model, much more capable.
Thank you very much for your hard work, Zai team!

Hi, @NicSir
Thanks for the feedback! Just to clarify, GLM-5.3-Flash does not use a PLE table. Both SGLang and vLLM already have generic weight-offloading support.

I think he meant if you guys are going to add PLE in future models like 5.5 Flash

just let them cook in peace bruh, if you know so much about making good models go get VC funding

Sign up or log in to comment