★ 0 · Updated 2026-10-06
Analyzes and optimizes KVCache implementations for PyTorch LLM inference on Ascend NPUs, covering continuous, paged, fused-attention, and MLA cache designs.
Browse skills that share this tag.
★ 0 · Updated 2026-10-06
Analyzes and optimizes KVCache implementations for PyTorch LLM inference on Ascend NPUs, covering continuous, paged, fused-attention, and MLA cache designs.