optimizing-attention-flash - Flash Attention Optimization for Transformers
Optimizes transformer attention using Flash Attention via PyTorch SDPA, flash-attn library, H100 FP8, and sliding window attention.
Tags
Updated: 2026-06-30Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Enable Flash Attention backend
- Replace standard attention computation
- Benchmark attention speed and memory
- Verify output accuracy consistency
- Enable sliding window attention
- Enable multi-query attention support
- Convert attention inputs to FP8
- Run FP8 attention on H100
Inputs
- Query, Key, Value tensors
- PyTorch model with attention layers
- NVIDIA GPU device
Outputs
- Optimized attention output tensors
- Performance benchmark results
- Accuracy comparison results
Requirements
- PyTorch 2.2+ for native SDPA support
- NVIDIA Ampere+ GPU (A100, A10, A30) or AMD MI200+
- CUDA 12.0+ (11.8 minimum)
- flash-attn library for advanced features
- float16 or bfloat16 dtype required
- H100 GPU for FP8 support
