LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

llm-benchmark - LLM Benchmark Design and Optimization

Guide for designing, running, and interpreting LLM benchmark experiments with statistical analysis and data leakage prevention

Tags

Updated: 2026-05-28

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • design test suites
  • run benchmark experiments
  • compare prompt variants
  • analyze pass rates
  • calculate confidence intervals
  • detect data leakage
  • measure per-turn metrics
  • interpret test results
  • optimize prompts
  • validate test constraints

Inputs

  • test dataset
  • prompt configuration
  • LLM model
  • test suite
  • statistical parameters
  • target effect size

Outputs

  • test results
  • pass rate
  • per-turn metrics
  • statistical analysis
  • performance comparison

Requirements

  • LLM model access
  • statistical analysis tools
  • test environment
  • test data access

Source

  • Spec: SKILL.md
benchmarking
LLM
testing
statistical analysis
prompt optimization
performance
data leakage prevention
design test suites
run benchmark experiments
compare prompt variants
analyze pass rates
test dataset
prompt configuration
LLM model
test results
pass rate
per-turn metrics