llm-inference-scaling - Auto-scale LLM inference clusters on Kubernetes
Scale LLM inference horizontally on Kubernetes with GPU-aware autoscaling, request queuing, and spot instance strategies.
Tags
Updated: 2026-03-26Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Deploy vLLM inference pods
- Configure KEDA autoscaling
- Set up queue-based scaling
- Configure spot instance priorities
- Install NVIDIA GPU Operator
- Monitor GPU metrics
- Scale replicas by queue length
- Scale replicas by GPU usage
- Deploy batch inference jobs
- Configure cluster autoscaler
- Create PodDisruptionBudget
Inputs
- Kubernetes cluster with GPU nodes
- KEDA installed
- Prometheus with GPU metrics
- Helm 3
- Hugging Face Hub token
- vLLM or TGI model image
- Redis queue
- GPU node pools
Outputs
- Scaled vLLM/TGI inference pods
- Prometheus scaling metrics
- Queue-based batch jobs
- Cluster autoscaler configuration
- PodDisruptionBudget policies
Requirements
- Kubernetes cluster with GPU nodes
- KEDA (Kubernetes Event-Driven Autoscaler)
- Prometheus with GPU metrics (dcgm-exporter or gpu-operator)
- Helm 3+
