LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

llm-as-a-judge - Build LLM Evaluators for Automated Quality Assessment

Build and deploy LLM evaluators for automated Pass/Fail quality assessment of LLM pipeline outputs

Tags

Updated: 2026-05-28

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • write judge prompt
  • split labeled data
  • measure TPR/TNR
  • refine judge prompt
  • estimate success rate
  • calculate bias correction
  • setup CI pipeline

Inputs

  • labeled traces
  • failure mode criteria
  • human ground truth labels
  • production traces
  • reference outputs

Outputs

  • judge judgment
  • TPR/TNR metrics
  • corrected success rate
  • CI evaluation result

Requirements

  • Python runtime
  • labeled training data
  • LLM judge model with version pinning

Source

  • Spec: SKILL.md
LLM evaluation
quality assessment
automated evaluation
TPR
TNR
CI pipeline
write judge prompt
split labeled data
measure TPR/TNR
refine judge prompt
labeled traces
failure mode criteria
human ground truth labels
judge judgment
TPR/TNR metrics
corrected success rate