harbor - Agent Evaluation Framework for Testing and Benchmarking Skills
Framework for agent evaluation and skills benchmarking
Tags
Updated: 2026-02-24Capabilities
Typical Inputs
Typical Outputs
What this skill does
- run harbor commands
- validate SkillsBench tasks
- create task templates
- execute agent tests
Inputs
- task configuration files
- agent skills
- API keys
- test cases
Outputs
- execution logs
- test results
- verification files
- benchmark reports
Requirements
- harbor installation
- Docker environment
- API access
- agent skills directories
