★ 4 · Updated 2026-10-05
Creates binary Pass/Fail LLM-as-a-Judge evaluators for specific failure modes and validates them against human labels using TPR/TNR.
Browse skills that share this capability.
★ 4 · Updated 2026-10-05
Creates binary Pass/Fail LLM-as-a-Judge evaluators for specific failure modes and validates them against human labels using TPR/TNR.