LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

Skills tagged: benchmark

Browse skills that share this tag.

  • ppa-benchmark - Set up defensible PPA comparison plans across arms
    ppabenchmarkedapnr

    ★ 26 · Updated 2026-09-17

    Produces a structured comparison plan and evidence-linked report across arms without scoring or naming a winner.

    ⚙ Define comparison arms and baseline⚙ Fix baseline design conditions⚙ Declare seed policy and runs
  • curate-scanner-evaluation-corpus - Curate Scanner Evaluation Corpus
    benchmarkevaluationground-truthscanner

    ★ 0 · Updated 2026-09-16

    Select pinned repositories, define annotation scopes, prepare review packets, maintain ground truth, and adjudicate labels for deterministic scanner evaluation.

    ⚙ Select coverage candidate repositories⚙ Define exhaustive annotation scopes⚙ Prepare human review packets
  • benchmark-to-brief - Turn benchmark research into campaign briefs and concepts
    media-productionbenchmarkcampaign-briefshort-form-video

    ★ 24 · Updated 2026-09-14

    Converts validated benchmark research artifacts into fact-grounded campaign briefs, concept candidates, hook libraries, and test plans for video production.

    ⚙ Convert research into campaign briefs⚙ Generate concept candidates and hooks⚙ Prioritize content lanes and matrices
  • subject-level-inference-eval - Subject-level text anonymization evaluation benchmark.
    text-anonymizationprivacy-evaluationbenchmarksubject-level-inference

    ★ 3 · Updated 2026-09-13

    Evaluates text anonymization by measuring span-level masking accuracy and subject-level privacy leakage while assessing text utility on PANORAMA and TAB.

    ⚙ Evaluate text anonymization methods⚙ Measure span-level masking accuracy⚙ Measure subject-level privacy leakage
  • benchmark-hillclimb - Improve AI agents via trace-driven benchmark experiments.
    benchmarkevalsoptimizationtrace-analysis

    ★ 1 · Updated 2026-09-13

    Improves AI agents, retrieval systems, benchmark harnesses, or workflows using trace-driven experiments and failure classification.

    ⚙ Inspect trace artifacts and sources⚙ Classify failure modes⚙ Execute iterative experiment loops
  • criterion-bench - Run Criterion benchmarks and detect performance regressions
    rustcriterionbenchmarkperformance-testing

    ★ 4 · Updated 2026-09-13

    Run Criterion 0.5 benchmarks, compare saved baselines, and analyze estimates.json files to detect performance regressions.

    ⚙ Save benchmark baselines⚙ Compare benchmarks against baselines⚙ Detect performance regressions
  • inquire-eval - Evaluates text-to-image retrieval on INQUIRE benchmark.
    text-to-image-retrievaldataset-evaluationvision-language-modelsbenchmark

    ★ 3 · Updated 2026-09-13

    Evaluates text-to-image retrieval capabilities of vision-language models on expert-level, ecologically grounded queries.

    ⚙ Evaluate text-to-image retrieval performance⚙ Probe fine-grained visual understanding⚙ Measure domain-specific language comprehension
  • benchmark - Performance Baseline & Regression Detection
    performancebenchmarkregression-testingweb-vitals

    ★ 0 · Updated 2026-09-11

    Measures web, API, and build performance metrics, saves baselines, and compares changes before and after PRs.

    ⚙ Measure page performance metrics⚙ Benchmark API endpoint performance⚙ Measure build performance metrics
  • memory-harness - Memory Harness
    testingmemorybenchmarkagent

    ★ 12 · Updated 2026-09-11

    Record, replay, and benchmark real agent sessions against do-memory-cli.

    ⚙ Record live agent session traces⚙ Replay session traces⚙ Benchmark CLI performance
  • phmbs-benchmark-report - Write PHM benchmark reports
    benchmarkreportphmmetrics

    ★ 0 · Updated 2026-09-10

    Write reproducible PHM benchmark reports with metrics, split manifests, source checksums, validity checks, and deviations.

    ⚙ Write reproducible benchmark reports⚙ Generate metrics JSON files⚙ Generate config JSON files
  • compute-token-stats-rust-ultra - Compute Token Stats (Rust Ultra)
    rusttoken-statscost-computationsse-logs

    ★ 1 · Updated 2026-09-09

    Computes token totals and estimated costs from Codex and Claude SSE logs using Rust binaries or Python scripts.

    ⚙ Compute token totals and costs⚙ Filter logs by session ID⚙ Build Rust release binaries
  • douyin-similar-account - TikTok Similar Account Recommendation Tool
    douyinaccount analysissimilar accountsbenchmark

    ★ 52 · Updated 2026-06-15

    Query TikTok account data and similar account recommendations via RedFox API

    ⚙ query account information⚙ match benchmark accounts⚙ recommend top accounts
  • profile-isaac-sim - Profile Isaac Sim Performance
    performanceprofilingbenchmarkIsaac Sim

    ★ 3,468 · Updated 2026-06-15

    Run benchmarks, capture Tracy profiles, and optimize Isaac Sim workloads

    ⚙ run benchmark⚙ capture Tracy profile⚙ export to CSV
  • self-improve - Autonomous evolutionary code improvement engine
    code improvementevolutionarytournament selectionautonomous

    ★ 115 · Updated 2026-05-28

    Autonomous loop controller for evolutionary code improvement with tournament selection

    ⚙ analyze codebase⚙ generate hypotheses⚙ create plans
  • harness-runner - WaveCap-SDR Test Harness Runner
    testingSDRaudioautomation

    ★ 602 · Updated 2026-05-28

    Run WaveCap-SDR test harness with automated parameter sweeps and validation

    ⚙ start server⚙ create capture⚙ add channels
  • mteb-leaderboard - ML Model Leaderboard and Benchmark Query Guide
    machine learningleaderboardbenchmarkMTEB

    ★ 0 · Updated 2026-05-09

    Guidance for querying ML model leaderboards and benchmarks including MTEB and HuggingFace

    ⚙ Find top-performing models on benchmarks⚙ Query current leaderboard standings⚙ Compare model performance across benchmarks
  • agent-bom-compliance - AI Compliance & Policy Engine
    securitycomplianceSBOMpolicy

    ★ 816 · Updated 2026-03-25

    Evaluate scan results against 14 security frameworks, generate SBOMs and compliance reports

    ⚙ evaluate compliance⚙ check policy⚙ run benchmark
  • local-crdb-perf-checker - Local CockroachDB Performance Checker
    cockroachdbperformancedatabasetesting

    ★ 0 · Updated 2026-03-22

    Automate local CockroachDB performance checks with DDL application and diagnostic reports

    ⚙ start database⚙ reset database⚙ apply SQL file
  • compute-token-stats-rust-ultra - Ultra-fast token and cost computation for AI logs
    token-computationcost-analysislog-processingbenchmark

    ★ 57 · Updated 2026-03-09

    Computes token totals and estimated costs from Codex and Claude SSE logs

    ⚙ parse SSE logs⚙ compute token statistics⚙ estimate costs
  • harbor - Harbor Framework for Agent Evaluation
    evaluationframeworkbenchmarktesting

    ★ 0 · Updated 2026-02-24

    Runs harbor commands and manages agent evaluation tasks

    ⚙ run harbor commands⚙ validate task structure⚙ execute agent tasks

Scroll to load more