LogoClawIndex
CasesSkillsAbout
LogoClawIndex

llm-inference-scaling - Auto-scale LLM inference clusters on Kubernetes

Scale LLM inference horizontally on Kubernetes with GPU-aware autoscaling, request queuing, and spot instance strategies.

Tags

Updated: 2026-03-26

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • Deploy vLLM inference pods
  • Configure KEDA autoscaling
  • Set up queue-based scaling
  • Configure spot instance priorities
  • Install NVIDIA GPU Operator
  • Monitor GPU metrics
  • Scale replicas by queue length
  • Scale replicas by GPU usage
  • Deploy batch inference jobs
  • Configure cluster autoscaler
  • Create PodDisruptionBudget

Inputs

  • Kubernetes cluster with GPU nodes
  • KEDA installed
  • Prometheus with GPU metrics
  • Helm 3
  • Hugging Face Hub token
  • vLLM or TGI model image
  • Redis queue
  • GPU node pools

Outputs

  • Scaled vLLM/TGI inference pods
  • Prometheus scaling metrics
  • Queue-based batch jobs
  • Cluster autoscaler configuration
  • PodDisruptionBudget policies

Requirements

  • Kubernetes cluster with GPU nodes
  • KEDA (Kubernetes Event-Driven Autoscaler)
  • Prometheus with GPU metrics (dcgm-exporter or gpu-operator)
  • Helm 3+

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.
kubernetes
autoscaling
LLM
GPU
KEDA
spot-instances
cost-optimization
inference
vllm
prometheus
Deploy vLLM inference pods
Configure KEDA autoscaling
Set up queue-based scaling
Configure spot instance priorities
Kubernetes cluster with GPU nodes
KEDA installed
Prometheus with GPU metrics
Scaled vLLM/TGI inference pods
Prometheus scaling metrics
Queue-based batch jobs