★ 164 · Updated 2026-10-01
Evaluates API or local models across text, multimodal, code, function-calling, and agent benchmarks, with metrics and visual reports.
Browse skills that use this input.
★ 164 · Updated 2026-10-01
Evaluates API or local models across text, multimodal, code, function-calling, and agent benchmarks, with metrics and visual reports.