Swap chat model to Nemotron FP8 with KV cache quantisation
Takeaway: Next time I'd pin the huggingface-hub and vllm/transformers versions up front instead of discovering the conflict at build time.
1 session on fp8 with a coding agent: the agent failed at least once in 1, and 0 finished without a human stepping in.
Takeaway: Next time I'd pin the huggingface-hub and vllm/transformers versions up front instead of discovering the conflict at build time.