fla-triton-to-gluon - Port FLA Triton Kernels to Gluon
Guides incremental Triton-to-Gluon kernel porting with numerical parity, explicit layouts, asynchronous memory movement, MMA, and scheduling controls.
Tags
Updated: 2026-10-01Capabilities
What this skill does
- Assess porting opportunities
- Translate Triton kernels incrementally
- Configure explicit tensor layouts
- Manage asynchronous memory pipelines
- Integrate Hopper and Blackwell MMA
- Preserve forward-backward numerical parity
- Tune compile-time and autotune settings
Inputs
- Existing Triton kernel
- Operation parity tests
- Installed Triton version
- Target NVIDIA GPU architecture
- Performance measurements
Outputs
- Gluon kernel implementation
- Forward and backward parity results
- Kernel performance measurements
Requirements
- Triton with experimental Gluon support
- NVIDIA GPU environment
- Ampere or newer for cp.async
- Hopper or newer for TMA and WGMMA
- Blackwell for TMEM and tcgen05
