Run Qwen3.6-27B on an RTX 5090 with vLLM. See why BF16 does not fit, how NVFP4 changes the memory requirement, and how to serve an API.