simpo - Simple Preference Optimization
Executes reference-free preference optimization for language model post-training without requiring a reference model.
Tags
Updated: 2026-09-16Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Train base models with SimPO
- Fine-tune instruct models
- Optimize reasoning intensive tasks
Inputs
- Preference datasets
- Model checkpoints
- Training configuration files
Outputs
- Trained model checkpoints
- Training logs
Requirements
- Python 3.10
- PyTorch 2.2.2
- alignment-handbook package
- flash-attn package
- NVIDIA A100 or H100 GPU
- Linux, macOS, or Windows
