LogoClawIndex
CasesSkillsAbout
LogoClawIndex

analyze-kernel-bottleneck - Analyze CUDA kernel bottlenecks for optimization

Classifies a CUDA kernel as compute-, memory-, or latency-bound using performance, roofline, occupancy, tile-ratio, and SASS analysis.

Tags

Updated: 2026-10-01

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • Measure baseline kernel performance
  • Classify roofline bottlenecks
  • Calculate GPU occupancy
  • Compute tile load ratios
  • Inspect SASS instruction mix
  • Analyze instruction stall codes
  • Check shared-memory occupancy cliffs
  • Select optimization strategies

Inputs

  • Compiled CUDA kernel
  • CUDA kernel source
  • Kernel build command
  • CUDA benchmark harness
  • Problem dimensions
  • Target GPU architecture
  • Expected peak utilization
  • Prior profiling data

Outputs

  • Baseline timing and throughput metrics
  • Bottleneck classification
  • Occupancy analysis table
  • Compute-to-load ratio
  • cp.async recommendation
  • SASS instruction count table
  • Stall code summary
  • Optimization decision matrix

Requirements

  • CUDA toolkit with nvcc
  • cuobjdump utility
  • CUDA-capable NVIDIA GPU
  • CUDA event timing support
  • Permission to compile and run kernels

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.
gpu
cuda
gpu-optimization
roofline
occupancy
sass
tensor-core
bottleneck-analysis
compute-load-ratio
Measure baseline kernel performance
Classify roofline bottlenecks
Calculate GPU occupancy
Compute tile load ratios
Compiled CUDA kernel
CUDA kernel source
Kernel build command
Baseline timing and throughput metrics
Bottleneck classification
Occupancy analysis table