bench-open-model - Benchmark Open-Weight Models on AWS GPUs
Sizes, provisions, serves, benchmarks, reports on, and tears down AWS GPU infrastructure for vLLM-servable Hugging Face text and multimodal models.
Tags
Updated: 2026-10-02AWSGPU benchmarkingvLLMHugging Faceopen-weight modelsthroughput testinglatency testingmultimodal inference
Capabilities
What this skill does
- Estimate model VRAM needs
- Rank AWS GPU instances
- Provision EC2 GPU infrastructure
- Serve models with vLLM
- Run cache-honest load sweeps
- Measure throughput and latency
- Write benchmark reports
- Teardown benchmark resources
Inputs
- Hugging Face model ID
- AWS credentials
- AWS account and region
- Benchmark workload shape
- Explicit launch approval
- HF access token
- GPU instance preference
- vLLM image version
Outputs
- VRAM sizing results
- GPU instance recommendations
- Throughput and latency results
- Benchmark report file
- Benchmark state file
- vLLM serving endpoint
- Provisioned AWS resources
- Deleted CloudFormation stack
Requirements
- Python 3
- AWS CLI
- AWS launch permissions
- SSH and SCP access
- vLLM architecture support
- Dedicated non-production AWS account
- Available AWS GPU capacity
