rocm-kernels - Optimize Triton kernels for AMD ROCm pipelines
Guides development, integration, benchmarking, and debugging of optimized Triton kernels for AMD GPUs on ROCm.
Tags
Updated: 2026-10-07Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Write optimized Triton kernels
- Benchmark kernel performance
- Integrate diffusers pipelines
- Optimize AMD GPU kernels
- Debug ROCm/HIP issues
Inputs
- Kernel type
- Triton or PyTorch source code
- HuggingFace pipeline or model
- Benchmark parameters
- AMD GPU target
Outputs
- Kernel implementation guidance
- Benchmark measurements
- JSON benchmark results
- Pipeline with injected kernels
Requirements
- ROCm environment
- AMD MI355X or R9700 GPU
- Triton
- PyTorch
- HuggingFace diffusers or transformers
