deploy-kimi-k26-on-rtx-pro-6000 - Deploy Kimi-K2.6 on 8× RTX PRO 6000 GPUs
Deploy and serve Kimi-K2.6 with INT4 QAT or NVFP4 using vLLM or SGLang in Docker on eight RTX PRO 6000 Blackwell GPUs.
Tags
Updated: 2026-10-05Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Verify GPU and host readiness
- Select model quantization
- Select inference engine
- Select MoE backend
- Select KV-cache dtype
- Download model weights
- Run Docker inference server
- Expose OpenAI-compatible API
- Configure systemd service
- Troubleshoot GPU serving failures
Inputs
- GPU hardware details
- Quantization choice
- Inference engine choice
- MoE backend choice
- KV-cache dtype choice
- Hugging Face checkpoint
- Hugging Face cache path
- Authenticated proxy configuration
Outputs
- Downloaded model weights
- Docker inference server
- OpenAI-compatible API on port 30000
- kimi-k26 systemd service
- Deployment environment file
- Service and troubleshooting status
Requirements
- Linux server
- Eight NVIDIA RTX PRO 6000 Blackwell GPUs
- NVIDIA driver 570 or newer
- Open NVIDIA kernel module
- Docker
- NVIDIA Container Toolkit with CDI
- Active nvidia-persistenced
- About 650 GB free local NVMe storage
- Access to model files and Docker images
