★ 0 · Updated 2026-09-23
Systematic improvement of existing agents through performance analysis, prompt engineering, and continuous iteration.
Browse skills that share this tag.
★ 0 · Updated 2026-09-23
Systematic improvement of existing agents through performance analysis, prompt engineering, and continuous iteration.
★ 105 · Updated 2026-09-23
Analyze and solve mathematical modeling competition problems using decision patterns distilled from 139 CUMCM excellent papers.
★ 602 · Updated 2026-09-21
Traces LLM calls, evaluates model outputs against datasets, monitors production metrics, and manages evaluation feedback for AI applications.
★ 1 · Updated 2026-09-20
Aggregates Orchestration Telemetry from recent Done beans to generate a report evaluating whether orchestration is effective.
★ 602 · Updated 2026-09-19
Evaluates proposals, plans, or decisions through structured adversarial debate and peer cross-examination among specialized AI councillors.
★ 200 · Updated 2026-09-19
Validates all checkpoints from an agent run directory in parallel by spawning test-validator agents and summarizing verified results.
★ 0 · Updated 2026-09-18
Answers facts about Nemotron 3 Nano architecture, training data, recipes, evaluation, quantization, and deployment.
★ 2,074 · Updated 2026-09-18
Provides factual information about Nemotron 3 Nano architecture, training data, recipes, evaluation, quantization, and deployment.
★ 470 · Updated 2026-09-18
SkillAnything automates target analysis, architecture design, implementation, evaluation, description optimization, and multi-platform packaging.
★ 17 · Updated 2026-09-18
Evaluates Ralph hat collection presets by running test scripts, logging session metrics, and verifying hat routing performance.
★ 9 · Updated 2026-09-17
Searches existing projects across code registries, launches, and literature to return a prior-art landscape and actionable build verdict.
★ 0 · Updated 2026-09-17
Provides a formal evaluation framework for Codex sessions to define, run, and report EDD metrics using code, model, and human graders.
★ 0 · Updated 2026-09-16
Select pinned repositories, define annotation scopes, prepare review packets, maintain ground truth, and adjudicate labels for deterministic scanner evaluation.
★ 64 · Updated 2026-09-16
Executes research tasks by characterizing problems, selecting workflow strategies, executing search, and self-evaluating output quality.
★ 1,361 · Updated 2026-09-16
Evaluates human performance versus AI tool output, producing rewritten criteria, calibration rules, and review conversation scripts.
★ 32 · Updated 2026-09-16
Produces structured design documents for systems and architectures through research, trade-off analysis, and iterative evaluation.
★ 76 · Updated 2026-09-16
Optimize entities against evaluator populations using LLMs and counterfactual probes.
★ 0 · Updated 2026-09-16
Patterns and techniques for evaluating and improving AI agent outputs.
★ 602 · Updated 2026-09-15
Implements eval-driven development to define, run, and report capability and regression evals for Claude Code sessions.
★ 3 · Updated 2026-09-15
Spawns N parallel candidates for a task, evaluates them against a rubric, selects a base artifact, and grafts the strongest parts into a synthesized result.
Scroll to load more