LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

Skills tagged: evaluation

Browse skills that share this tag.

  • agent-orchestration-improve-agent - Optimizes existing agent performance and prompt quality
    agent-optimizationprompt-engineeringab-testingperformance-analysis

    ★ 0 · Updated 2026-09-23

    Systematic improvement of existing agents through performance analysis, prompt engineering, and continuous iteration.

    ⚙ Analyze agent performance data⚙ Classify system failure modes⚙ Optimize prompts and few-shot examples
  • math-modeling-skill - Mathematical Modeling Competition Expert
    mathematical-modelingcumcmoptimizationprediction

    ★ 105 · Updated 2026-09-23

    Analyze and solve mathematical modeling competition problems using decision patterns distilled from 139 CUMCM excellent papers.

    ⚙ Decompose modeling problems and dependencies⚙ Classify subproblems by mathematical structure⚙ Retrieve structurally similar historical cases
  • langsmith-observability - Tracing, evaluation, and monitoring for LLM applications
    observabilitytracingevaluationmonitoring

    ★ 602 · Updated 2026-09-21

    Traces LLM calls, evaluates model outputs against datasets, monitors production metrics, and manages evaluation feedback for AI applications.

    ⚙ Trace LLM application calls⚙ Evaluate model outputs against datasets⚙ Monitor production system metrics
  • orchestration-report - Generate evaluation reports on bean orchestration quality.
    orchestrationreporttelemetrymetrics

    ★ 1 · Updated 2026-09-20

    Aggregates Orchestration Telemetry from recent Done beans to generate a report evaluating whether orchestration is effective.

    ⚙ Identify candidate Done beans⚙ Parse Orchestration Telemetry metrics⚙ Compute orchestration quality aggregations
  • workflow-debate - Structured adversarial debate to evaluate ideas and plans
    debatecouncilevaluationagent-teams

    ★ 602 · Updated 2026-09-19

    Evaluates proposals, plans, or decisions through structured adversarial debate and peer cross-examination among specialized AI councillors.

    ⚙ Parse execution flags and input topic⚙ Spawn specialized AI councillor roles⚙ Collect independent opening statements
  • validate-run - Validate agent run checkpoints in parallel.
    testingvalidationcheckpointsub-agent

    ★ 200 · Updated 2026-09-19

    Validates all checkpoints from an agent run directory in parallel by spawning test-validator agents and summarizing verified results.

    ⚙ Discover checkpoints in run directory⚙ Launch test-validator agents in parallel⚙ Collect reports from test validators
  • nemotron-nano3 - Reference desk for Nemotron 3 Nano
    nemotron-3-nanoknowledge-basearchitectureevaluation

    ★ 0 · Updated 2026-09-18

    Answers facts about Nemotron 3 Nano architecture, training data, recipes, evaluation, quantization, and deployment.

    ⚙ Answer model architecture questions⚙ Retrieve training data details⚙ Compare paper claims with recipes
  • nemotron-nano3 - Reference desk for Nemotron 3 Nano model facts and specs
    nemotron-nano3knowledge-basemodel-architectureevaluation

    ★ 2,074 · Updated 2026-09-18

    Provides factual information about Nemotron 3 Nano architecture, training data, recipes, evaluation, quantization, and deployment.

    ⚙ Answer questions about model architecture⚙ Provide details on training data⚙ Compare paper claims with public recipes
  • skill-anything - Generate production-ready skills for software and APIs
    skill-generatorautomationmulti-platformagent-skills

    ★ 470 · Updated 2026-09-18

    SkillAnything automates target analysis, architecture design, implementation, evaluation, description optimization, and multi-platform packaging.

    ⚙ Analyze target applications and APIs⚙ Design skill architecture⚙ Generate skill files and scripts
  • evaluate-presets - Systematically test Ralph hat collection presets using scripts.
    testingevaluationpresetsralph

    ★ 17 · Updated 2026-09-18

    Evaluates Ralph hat collection presets by running test scripts, logging session metrics, and verifying hat routing performance.

    ⚙ Evaluate single preset⚙ Evaluate all presets⚙ Extract session metrics
  • priorart - Check prior art and evaluate project idea uniqueness
    prior-artsearchresearchideation

    ★ 9 · Updated 2026-09-17

    Searches existing projects across code registries, launches, and literature to return a prior-art landscape and actionable build verdict.

    ⚙ Read checked ideas log⚙ Restate project idea⚙ Generate vocabulary framings
  • eval-harness - Formal evaluation framework for Codex sessions based on EDD.
    eval-driven-developmentcodexevaluationtesting

    ★ 0 · Updated 2026-09-17

    Provides a formal evaluation framework for Codex sessions to define, run, and report EDD metrics using code, model, and human graders.

    ⚙ Define capability and regression evals⚙ Execute code-based graders⚙ Execute model-based graders
  • curate-scanner-evaluation-corpus - Curate Scanner Evaluation Corpus
    benchmarkevaluationground-truthscanner

    ★ 0 · Updated 2026-09-16

    Select pinned repositories, define annotation scopes, prepare review packets, maintain ground truth, and adjudicate labels for deterministic scanner evaluation.

    ⚙ Select coverage candidate repositories⚙ Define exhaustive annotation scopes⚙ Prepare human review packets
  • research-task - Execute research via compositional workflow planning
    researchworkflow-planningtask-executionevaluation

    ★ 64 · Updated 2026-09-16

    Executes research tasks by characterizing problems, selecting workflow strategies, executing search, and self-evaluating output quality.

    ⚙ Characterize research tasks⚙ Query workflow library⚙ Select research strategy archetypes
  • ai-assisted-performance-review - Evaluate human performance fairly when work is AI-assisted
    performance-reviewai-assisted-workmanagementcalibration

    ★ 1,361 · Updated 2026-09-16

    Evaluates human performance versus AI tool output, producing rewritten criteria, calibration rules, and review conversation scripts.

    ⚙ Analyze evaluation criteria signals⚙ Rewrite criteria for human capabilities⚙ Establish calibration rules for teams
  • sisyphus-design - High-Quality System and Architecture Design Skill
    designarchitecturedocumentationsystem-design

    ★ 32 · Updated 2026-09-16

    Produces structured design documents for systems and architectures through research, trade-off analysis, and iterative evaluation.

    ⚙ Investigate requirements using research subagents⚙ Generate divergent architectural approaches⚙ Present trade-offs to users
  • sgo - Semantic Gradient Optimization
    optimizationevaluationllmpersona

    ★ 76 · Updated 2026-09-16

    Optimize entities against evaluator populations using LLMs and counterfactual probes.

    ⚙ Download persona dataset⚙ Build entity document⚙ Filter persona data
  • agentic-eval - Agentic Evaluation Patterns
    evaluationreflectionevaluator-optimizerllm-as-judge

    ★ 0 · Updated 2026-09-16

    Patterns and techniques for evaluating and improving AI agent outputs.

    ⚙ Implement self-critique and reflection loops⚙ Build evaluator-optimizer pipelines⚙ Create test-driven code refinement workflows
  • eval-harness - Formal evaluation framework for Claude Code sessions.
    evaluationtestingeddclaude-code

    ★ 602 · Updated 2026-09-15

    Implements eval-driven development to define, run, and report capability and regression evals for Claude Code sessions.

    ⚙ Define capability and regression evals⚙ Execute deterministic code-based graders⚙ Execute model-based output graders
  • arena - Parallel candidate generation and synthesis workflow
    parallel-processingcandidate-synthesismulti-agentevaluation

    ★ 3 · Updated 2026-09-15

    Spawns N parallel candidates for a task, evaluates them against a rubric, selects a base artifact, and grafts the strongest parts into a synthesized result.

    ⚙ Frame task with concrete rubric⚙ Spawn parallel candidate subagents⚙ Run cross-judge evaluation subagent

Scroll to load more