★ 26 · Updated 2026-09-17
Produces a structured comparison plan and evidence-linked report across arms without scoring or naming a winner.
Browse skills that share this tag.
★ 26 · Updated 2026-09-17
Produces a structured comparison plan and evidence-linked report across arms without scoring or naming a winner.
★ 0 · Updated 2026-09-16
Select pinned repositories, define annotation scopes, prepare review packets, maintain ground truth, and adjudicate labels for deterministic scanner evaluation.
★ 24 · Updated 2026-09-14
Converts validated benchmark research artifacts into fact-grounded campaign briefs, concept candidates, hook libraries, and test plans for video production.
★ 3 · Updated 2026-09-13
Evaluates text anonymization by measuring span-level masking accuracy and subject-level privacy leakage while assessing text utility on PANORAMA and TAB.
★ 1 · Updated 2026-09-13
Improves AI agents, retrieval systems, benchmark harnesses, or workflows using trace-driven experiments and failure classification.
★ 4 · Updated 2026-09-13
Run Criterion 0.5 benchmarks, compare saved baselines, and analyze estimates.json files to detect performance regressions.
★ 3 · Updated 2026-09-13
Evaluates text-to-image retrieval capabilities of vision-language models on expert-level, ecologically grounded queries.
★ 0 · Updated 2026-09-11
Measures web, API, and build performance metrics, saves baselines, and compares changes before and after PRs.
★ 12 · Updated 2026-09-11
Record, replay, and benchmark real agent sessions against do-memory-cli.
★ 0 · Updated 2026-09-10
Write reproducible PHM benchmark reports with metrics, split manifests, source checksums, validity checks, and deviations.
★ 1 · Updated 2026-09-09
Computes token totals and estimated costs from Codex and Claude SSE logs using Rust binaries or Python scripts.
★ 52 · Updated 2026-06-15
Query TikTok account data and similar account recommendations via RedFox API
★ 3,468 · Updated 2026-06-15
Run benchmarks, capture Tracy profiles, and optimize Isaac Sim workloads
★ 115 · Updated 2026-05-28
Autonomous loop controller for evolutionary code improvement with tournament selection
★ 602 · Updated 2026-05-28
Run WaveCap-SDR test harness with automated parameter sweeps and validation
★ 0 · Updated 2026-05-09
Guidance for querying ML model leaderboards and benchmarks including MTEB and HuggingFace
★ 816 · Updated 2026-03-25
Evaluate scan results against 14 security frameworks, generate SBOMs and compliance reports
★ 0 · Updated 2026-03-22
Automate local CockroachDB performance checks with DDL application and diagnostic reports
★ 57 · Updated 2026-03-09
Computes token totals and estimated costs from Codex and Claude SSE logs
★ 0 · Updated 2026-02-24
Runs harbor commands and manages agent evaluation tasks
Scroll to load more