serving-llms-vllm - High-throughput LLM serving framework
High-throughput LLM serving engine supporting OpenAI compatible API, quantization, and tensor parallelism.
Tags
Updated: 2026-09-22Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Deploy production LLM APIs
- Run offline batch inference
- Serve quantized LLM models
- Monitor server performance metrics
Inputs
- LLM models
- Input text prompts
- Server configuration parameters
- GPU hardware resources
Outputs
- OpenAI compatible API server
- Generated text responses
- JSONL result files
- Prometheus metrics endpoint
Requirements
- Linux or macOS
- Python packages vllm, torch, transformers
- Supported GPU hardware
