Swap chat model to Nemotron FP8 with KV cache quantisation
Takeaway: Next time I'd pin the huggingface-hub and vllm/transformers versions up front instead of discovering the conflict at build time.
Takeaway: Next time I'd pin the huggingface-hub and vllm/transformers versions up front instead of discovering the conflict at build time.