one-eval - Run end-to-end model evaluations with One-Eval
Evaluates API or local models across text, multimodal, code, function-calling, and agent benchmarks, with metrics and visual reports.
Tags
Updated: 2026-10-01Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Test model connectivity
- Select evaluation benchmarks
- Generate evaluation specs
- Run smoke evaluations
- Run full evaluations
- Resume interrupted evaluations
- Compute evaluation metrics
- Render HTML reports
Inputs
- Model name or path
- API endpoint and key
- Benchmark selection
- Sampling parameters
- Metric selection
- Evaluation repository
Outputs
- Evaluation results JSON
- Metric results JSON
- Per-sample metric records
- HTML evaluation report
- Evaluation run directory
- Benchmark readiness state
Requirements
- One-Eval repository
- Python 3.10–3.11
- Installed project dependencies
- API credentials for API models
- GPU and matching CUDA for local vLLM
