LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

make-eval - Generate a deterministic LLM evaluation harness

Builds a small evaluation harness for LLM-backed functions, using exact-match scoring for closed-label outputs and optional LangSmith integration.

Tags

Updated: 2026-10-04

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • Identify the LLM boundary
  • Define the label set
  • Harden output parsing
  • Write evaluation datasets
  • Generate evaluation runners
  • Score exact-match outputs
  • Calculate confusion matrices
  • Configure LangSmith evaluations

Inputs

  • LLM-backed function
  • Closed label set
  • Evaluation examples
  • Project package configuration
  • LANGSMITH_API_KEY environment variable

Outputs

  • Evaluation harness files
  • CSV evaluation dataset
  • Exact-match scores
  • Pass/fail threshold result
  • Confusion matrix
  • LangSmith dataset URL
  • LangSmith experiment URL

Requirements

  • Local project execution environment
  • Callable LLM-backed function
  • LangSmith package for LangSmith mode
  • LANGSMITH_API_KEY for LangSmith mode

Source

  • Spec: SKILL.md
LLM evaluation
Test harness
Classification
Exact-match scoring
Regression testing
Confusion matrix
LangSmith
Identify the LLM boundary
Define the label set
Harden output parsing
Write evaluation datasets
LLM-backed function
Closed label set
Evaluation examples
Evaluation harness files
CSV evaluation dataset
Exact-match scores