vllm - Operate and troubleshoot vLLM inference servers
Deploy, configure, serve, benchmark, tune, and troubleshoot vLLM inference servers across Docker, Kubernetes, GPUs, APIs, and upgrades.
Tags
Updated: 2026-10-01vLLMinference servingDockerKubernetesOpenAI APIbenchmarkingGPU operationscontinuous batchingquantizationtroubleshooting
Capabilities
What this skill does
- Deploy vLLM with Docker
- Deploy vLLM on Kubernetes
- Configure model serving
- Tune KV cache and batching
- Serve OpenAI-compatible APIs
- Benchmark throughput and latency
- Inspect health and metrics
- Diagnose GPU and startup failures
- Upgrade or roll back deployments
- Verify inference delivery
Inputs
- Deployment target
- vLLM version or image digest
- Model and revision
- Quantization and parallelism settings
- GPU inventory
- Workload specification
- Benchmark dataset
- Server URL
- Human mutation directive
Outputs
- Deployment configuration changes
- OpenAI-compatible API responses
- Health and metrics probe JSON
- Benchmark metrics and run records
- Upgrade or rollback status
- Diagnostic summaries
Requirements
- vLLM 0.26.0 or pinned older release
- Supported NVIDIA CUDA, AMD ROCm, Intel XPU, or CPU environment
- Matching GPU driver when applicable
- Python 3.9+ for the health script
- HTTP(S) access for live probes
- Human confirmation for mutations
