model-infer-parallel-impl - Implement parallel inference for Ascend NPU models
Implements confirmed parallel configurations for PyTorch Ascend NPU models, including parallel layers, communication groups, YAML configuration, and weight handling.
Tags
Updated: 2026-09-30Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Confirm parallel configuration
- Select reference implementation
- Create communication groups
- Replace parallel layers
- Adapt MoE parallelism
- Generate YAML configurations
- Implement weight handling
- Validate parallel inference
Inputs
- Confirmed parallel configuration
- Target model repository
- Single-card model framework
- Reference model implementation
- Model weight path
Outputs
- Parallelized model code
- Communication group definitions
- YAML configuration files
- Weight loading or conversion implementation
- Inference validation results
Requirements
- PyTorch environment
- Ascend NPU environment
- HCCL communication support
- Confirmed parallel_config
- Complete single-card model adaptation
- Access to model source code
