LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

Skills tagged: benchmarking

Browse skills that share this tag.

  • py-perf - Measure, diagnose, and optimize Python code performance.
    pythonperformanceprofilingbenchmarking

    ★ 1 · Updated 2026-09-22

    Diagnose and improve Python performance using benchmarks, profiling tools, memory analysis, concurrency, and regression tests.

    ⚙ Profile CPU execution time⚙ Track memory allocations⚙ Run reproducible performance benchmarks
  • generating-semantic-game-mutants - Plant semantic game defects to evaluate test or agent power.
    mutation-testinggame-testingbenchmarkingagent-evals

    ★ 16 · Updated 2026-09-22

    Injects controlled game-specific semantic defects into game code or trace reducers and records a manifest to evaluate test suites or agent workflows.

    ⚙ Inject controlled game semantic defects⚙ Record mutant manifest entries⚙ Emit standalone mutant code artifacts
  • dotnet-performance - .NET performance engineering and diagnostics workflow
    dotnetperformanceprofilingbenchmarking

    ★ 36 · Updated 2026-09-21

    Provides an evidence-first performance engineering workflow for .NET applications combining portable diagnostics and profiling.

    ⚙ Investigate .NET application performance⚙ Fix dominant performance causes⚙ Profile managed and native workloads
  • perf-loop - Rust Performance Optimization Loop
    rustperformanceprofilingoptimization

    ★ 221 · Updated 2026-09-21

    Optimizes Rust performance using a strict measure-hypothesize-change-remeasure cycle with specific profiling tools and playbooks.

    ⚙ Measure baseline performance⚙ Profile CPU and allocation bottlenecks⚙ Formulate optimization hypotheses
  • performance-profiling - Analyze CPU, memory, and I/O performance.
    performanceprofilingbenchmarkingflamegraph

    ★ 0 · Updated 2026-09-20

    Profiles CPU, memory, and I/O performance using flame graphs and benchmarks to identify bottlenecks and prevent regressions.

    ⚙ Measure performance baselines⚙ Profile CPU usage⚙ Profile memory allocations
  • gale-v2-refactor - Refactor gale v2 while preserving GPU performance.
    refactoringgpu-optimizationcudarust

    ★ 1 · Updated 2026-09-20

    Consolidate CPU and GPU DG performance optimizations into Sim/Device abstractions while defending measured benchmark baselines.

    ⚙ Consolidate CPU and GPU optimizations⚙ Refactor Sim and Device abstractions⚙ Freeze performance benchmark baselines
  • energize-denver-proposals - Generate Energize Denver proposals and cost estimates.
    energize-denverproposal-generationbuilding-complianceenergy-audit

    ★ 602 · Updated 2026-09-19

    Generates Energize Denver compliance proposals, estimates costs, and verifies requirements against Denver Article XIV for commercial buildings.

    ⚙ Generate customized compliance proposals⚙ Estimate project service costs⚙ Plan compliance project timelines
  • evaluate-presets - Systematically test Ralph hat collection presets using scripts.
    testingevaluationpresetsralph

    ★ 17 · Updated 2026-09-18

    Evaluates Ralph hat collection presets by running test scripts, logging session metrics, and verifying hat routing performance.

    ⚙ Evaluate single preset⚙ Evaluate all presets⚙ Extract session metrics
  • eval-harness - Formal evaluation framework for Codex sessions based on EDD.
    eval-driven-developmentcodexevaluationtesting

    ★ 0 · Updated 2026-09-17

    Provides a formal evaluation framework for Codex sessions to define, run, and report EDD metrics using code, model, and human graders.

    ⚙ Define capability and regression evals⚙ Execute code-based graders⚙ Execute model-based graders
  • ck:autoresearch - Autonomous optimization loop for mechanical metrics
    optimizationgittestingcode-quality

    ★ 1 · Updated 2026-09-16

    Runs iterative code optimization loops against a mechanical metric, learning from git history to automatically keep or discard changes.

    ⚙ Run iterative code optimization loops⚙ Evaluate code with verification commands⚙ Commit or revert code changes
  • profile-performance - Profile and optimize system performance
    performanceprofilingoptimizationmetrics

    ★ 0 · Updated 2026-09-14

    Diagnose and optimize performance using measurement, profiling, runtime metrics, query plans, and before-and-after verification.

    ⚙ Define performance questions and metrics⚙ Establish baseline performance measurements⚙ Characterize workload and resource demands
  • agent-agent-evaluator - AI agent evaluation expert for benchmarking and metrics.
    agent-evaluationbenchmarkinghuman-evalsafety-testing

    ★ 6 · Updated 2026-09-11

    Evaluates AI agents systematically through benchmarking, human evaluation, performance metrics, safety testing, and test suite creation.

    ⚙ Define evaluation criteria⚙ Create test datasets⚙ Execute automated benchmarks
  • java-performance - JVM Performance Tuning and Analysis
    javajvmperformancegc

    ★ 602 · Updated 2026-09-11

    Optimizes JVM performance through garbage collection tuning, memory analysis, CPU profiling, and benchmarking.

    ⚙ Tune garbage collection⚙ Profile CPU hotspots⚙ Analyze memory leaks
  • chicago-tdd-pattern - Chicago TDD State-Based Testing
    state-based testingunit testingintegration testingTDD

    ★ 602 · Updated 2026-06-15

    Create state-based tests verifying observable behavior with real collaborators using AAA pattern

    ⚙ write unit tests⚙ write integration tests⚙ run unit tests
  • competitive-analysis - Competitive Analysis and Intelligence
    competitive analysismarket researchbusiness intelligencestrategic analysis

    ★ 602 · Updated 2026-06-15

    Systematically analyze competitors, compare products and strategies, generate competitive intelligence reports

    ⚙ profile competitors⚙ compare features⚙ analyze pricing
  • performance-profiling - Profile CPU, Memory, Network Performance
    performanceprofilingbenchmarkingmonitoring

    ★ 602 · Updated 2026-06-15

    Profile CPU, memory, and network performance for web apps and APIs

    ⚙ run CPU profiler⚙ create flame graph⚙ capture heap snapshot
  • llm-benchmark - LLM Benchmark Design and Optimization
    benchmarkingLLMtestingstatistical analysis

    ★ 15 · Updated 2026-05-28

    Guide for designing, running, and interpreting LLM benchmark experiments with statistical analysis and data leakage prevention

    ⚙ design test suites⚙ run benchmark experiments⚙ compare prompt variants
  • benchmark-report-creator - Create publication-quality benchmark reports
    benchmarkingreport generationPDF exportdocument automation

    ★ 602 · Updated 2026-05-28

    Orchestrates complete pipeline for benchmark reports with diagrams, PNG captures, and PDF export

    ⚙ orchestrate pipeline⚙ create markdown structure⚙ generate HTML diagrams
  • schelk - Linux Benchmarking CLI for Large Datasets
    linuxbenchmarkingstorageblock-device

    ★ 48 · Updated 2026-05-28

    Linux CLI for repeatable benchmarking on large stateful datasets using block devices and dm-era metadata

    ⚙ install schelk⚙ check prerequisites⚙ initialize new volumes
  • eval-runner - React-Craft Evaluation Test Runner
    testingbenchmarkingcode evaluationpipeline

    ★ 3 · Updated 2026-03-24

    Executes build pipeline against fixtures, applies graders, and produces benchmark reports.

    ⚙ Read fixture metadata⚙ Read fixture input⚙ Sanitize component names

Scroll to load more