Model quantization, Efficient inference, LLM inference, Model optimization, NVIDIA Blackwell, NVFP4, vLLM, Multimodal models