model-infer-kvcache - Optimize KVCache for Ascend NPU inference
Analyzes and optimizes KVCache implementations for PyTorch LLM inference on Ascend NPUs, covering continuous, paged, fused-attention, and MLA cache designs.
Tags
Updated: 2026-10-06Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Analyze KVCache implementations
- Select cache optimization modes
- Construct block tables
- Construct slot mappings
- Integrate fused attention operators
- Optimize MLA compressed caches
- Diagnose memory and performance issues
Inputs
- Model architecture and configuration
- KVCache implementation
- progress.md
- Repository model references
- Memory or performance symptoms
Outputs
- KVCache analysis
- Optimization recommendations
- Implementation guidance
- Cache layout guidance
- Attention operator integration guidance
Requirements
- PyTorch runtime
- Ascend NPU environment
- torch_npu support
- Fused attention operators
- MLA operator support when applicable
