verl-rl-training - LLM Reinforcement Learning Training with verl
Train LLMs with RLHF, GRPO, and PPO using verl library
Tags
Updated: 2026-06-30Typical Inputs
Typical Outputs
What this skill does
- train LLMs
- configure algorithm
- prepare dataset
- define reward function
- create training config
- launch training
- monitor training
- run evaluation
- manage cluster
- convert model
Inputs
- Dataset
- Base model
- Reward function
- Training config
- GPU cluster
Outputs
- Trained model
- Training metrics
- Evaluation results
Requirements
- verl>=0.3.0
- torch>=2.0.0
- ray>=2.41.0
- vllm>=0.8.2
- transformers>=4.40.0
- GPU cluster
- Python environment
