fla-triton-to-gluon - Port Triton kernels to Gluon
Ports Triton kernels in fla/ops/** to Gluon with explicit layouts, shared memory, asynchronous data movement, MMA, and scheduling control.
Tags
Updated: 2026-10-01Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Assess porting feasibility
- Translate Triton APIs to Gluon
- Add explicit tensor layouts
- Implement asynchronous memory pipelines
- Integrate MMA instructions
- Preserve forward and backward parity
- Profile and optimize kernels
Inputs
- Existing Triton kernel
- Frozen pytest contract
- Installed Triton version
- Target NVIDIA GPU architecture
- Kernel profiling data
Outputs
- Gluon kernel source
- Parity test results
- Kernel performance measurements
Requirements
- Triton installation with Gluon support
- NVIDIA GPU environment
- Hardware supporting selected features
- Access to the kernel test suite
