LogoClawIndex
CasesSkillsAbout
LogoClawIndex

run-evals - Run Existing Skill Evaluations

Execute and iterate on skill evaluations using existing test cases

Tags

Updated: 2026-03-23
testingevaluationbenchmarkingquality assuranceCI/CDskill development

Capabilities

Run eval commandsRead eval configurationsRead benchmark resultsCreate feedback files

Typical Inputs

evals/evals.jsonskill directoryold skill path

Typical Outputs

benchmark.jsonfeedback.jsonevaluation report

What this skill does

  • Run eval commands
  • Read eval configurations
  • Read benchmark results
  • Create feedback files
  • Modify eval configurations
  • Compare evaluation results
  • Generate improvement suggestions
  • Verify threshold conditions
  • Compare skill versions

Inputs

  • evals/evals.json
  • skill directory
  • old skill path
  • CI threshold value

Outputs

  • benchmark.json
  • feedback.json
  • evaluation report
  • comparison report
  • CI exit code

Requirements

  • Existing evals/evals.json
  • @github/copilot-sdk
  • npx snapeval CLI
  • Copilot CLI authentication
  • Node.js environment

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.