LogoClawIndex
CasesSkillsAbout
LogoClawIndex

evaluating-code-models - Evaluate code models with BigCode Evaluation Harness

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics.

Tags

Updated: 2026-09-24
EvaluationCode GenerationHumanEvalMBPPMultiPL-EPass@kBigCodeBenchmarkingCode Models

Capabilities

Evaluate standard code generation benchmarksRun multi-language code evaluationBenchmark instruction-tuned code modelsCompare pass@k metrics across models

Typical Inputs

HuggingFace model identifier or pathBenchmark task namesGeneration configuration parameters

Typical Outputs

JSON files containing pass@k metricsSaved code generation JSON filesModel comparison tables

What this skill does

  • Evaluate standard code generation benchmarks
  • Run multi-language code evaluation
  • Benchmark instruction-tuned code models
  • Compare pass@k metrics across models
  • Execute evaluation safely inside Docker

Inputs

  • HuggingFace model identifier or path
  • Benchmark task names
  • Generation configuration parameters
  • Pre-generated solutions JSON file

Outputs

  • JSON files containing pass@k metrics
  • Saved code generation JSON files
  • Model comparison tables

Requirements

  • bigcode-evaluation-harness
  • transformers>=4.25.1
  • accelerate>=0.13.2
  • datasets>=2.6.1
  • Docker runtime environment
  • Code execution permission

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.