moe-training - Train sparse Mixture of Experts models
Train and optimize sparse Mixture of Experts models with DeepSpeed or HuggingFace, including routing, load balancing, expert parallelism, and inference optimization.
Tags
Updated: 2026-09-29Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Implement sparse MoE architectures
- Configure top-k routing
- Balance expert utilization
- Distribute experts across devices
- Configure DeepSpeed MoE training
- Optimize sparse inference
Inputs
- Training data
- Model architecture configuration
- Training hyperparameters
- Vocabulary file
- Merge file
- DeepSpeed configuration
Outputs
- Trained MoE model checkpoints
- Training metrics
- Evaluation metrics
- Inference results
Requirements
- Python environment
- PyTorch
- DeepSpeed or HuggingFace Transformers
- HuggingFace Accelerate
- Multi-GPU resources for expert parallelism
