LogoClawIndex
CasesSkillsAbout
LogoClawIndex

fla-triton-to-gluon - Port FLA Triton Kernels to Gluon

Guides incremental Triton-to-Gluon kernel porting with numerical parity, explicit layouts, asynchronous memory movement, MMA, and scheduling controls.

Tags

Updated: 2026-10-01
TritonGluonGPU kernelsNVIDIAkernel portingperformance optimization

Capabilities

Assess porting opportunitiesTranslate Triton kernels incrementallyConfigure explicit tensor layoutsManage asynchronous memory pipelines

Typical Inputs

Existing Triton kernelOperation parity testsInstalled Triton version

Typical Outputs

Gluon kernel implementationForward and backward parity resultsKernel performance measurements

What this skill does

  • Assess porting opportunities
  • Translate Triton kernels incrementally
  • Configure explicit tensor layouts
  • Manage asynchronous memory pipelines
  • Integrate Hopper and Blackwell MMA
  • Preserve forward-backward numerical parity
  • Tune compile-time and autotune settings

Inputs

  • Existing Triton kernel
  • Operation parity tests
  • Installed Triton version
  • Target NVIDIA GPU architecture
  • Performance measurements

Outputs

  • Gluon kernel implementation
  • Forward and backward parity results
  • Kernel performance measurements

Requirements

  • Triton with experimental Gluon support
  • NVIDIA GPU environment
  • Ampere or newer for cp.async
  • Hopper or newer for TMA and WGMMA
  • Blackwell for TMEM and tcgen05

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.