make-eval - Generate a deterministic LLM evaluation harness
Builds a small evaluation harness for LLM-backed functions, using exact-match scoring for closed-label outputs and optional LangSmith integration.
Tags
Updated: 2026-10-04Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Identify the LLM boundary
- Define the label set
- Harden output parsing
- Write evaluation datasets
- Generate evaluation runners
- Score exact-match outputs
- Calculate confusion matrices
- Configure LangSmith evaluations
Inputs
- LLM-backed function
- Closed label set
- Evaluation examples
- Project package configuration
- LANGSMITH_API_KEY environment variable
Outputs
- Evaluation harness files
- CSV evaluation dataset
- Exact-match scores
- Pass/fail threshold result
- Confusion matrix
- LangSmith dataset URL
- LangSmith experiment URL
Requirements
- Local project execution environment
- Callable LLM-backed function
- LangSmith package for LangSmith mode
- LANGSMITH_API_KEY for LangSmith mode
