llm-benchmark - LLM Benchmarking: Design Suites, Compare Models & Detect Regressions
Design benchmark suites, establish baselines, run A/B model comparisons, detect quality regressions, track production metrics, and generate leaderboards for model selection.
Tags
Updated: 2026-06-30Capabilities
Typical Inputs
Typical Outputs
What this skill does
- design benchmark suites
- establish performance baselines
- compare models with statistical tests
- detect quality regressions
- track production quality metrics
- generate model leaderboards
- run canary tests for model updates
Inputs
- LLM models to benchmark
- Benchmark test cases
- Historical baseline results
- Independent judge model
- Evaluation rubrics
- Cost and latency thresholds
Outputs
- Benchmark comparison reports
- Regression detection alerts
- Model leaderboards
- Baseline data records
- Production monitoring dashboard
Requirements
- Access to LLM API endpoints
- Python runtime environment
- Statistical analysis libraries (numpy, scipy)
- Independent grader/judge model access
- Persistent storage for baseline data
