LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

one-eval - Run end-to-end model evaluations with One-Eval

Evaluates API or local models across text, multimodal, code, function-calling, and agent benchmarks, with metrics and visual reports.

Tags

Updated: 2026-10-01

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • Test model connectivity
  • Select evaluation benchmarks
  • Generate evaluation specs
  • Run smoke evaluations
  • Run full evaluations
  • Resume interrupted evaluations
  • Compute evaluation metrics
  • Render HTML reports

Inputs

  • Model name or path
  • API endpoint and key
  • Benchmark selection
  • Sampling parameters
  • Metric selection
  • Evaluation repository

Outputs

  • Evaluation results JSON
  • Metric results JSON
  • Per-sample metric records
  • HTML evaluation report
  • Evaluation run directory
  • Benchmark readiness state

Requirements

  • One-Eval repository
  • Python 3.10–3.11
  • Installed project dependencies
  • API credentials for API models
  • GPU and matching CUDA for local vLLM

Source

  • Spec: SKILL.md
model evaluation
benchmarking
multimodal evaluation
code generation
function calling
agent benchmarks
metrics
HTML reports
Test model connectivity
Select evaluation benchmarks
Generate evaluation specs
Run smoke evaluations
Model name or path
API endpoint and key
Benchmark selection
Evaluation results JSON
Metric results JSON
Per-sample metric records