verl-rl-training - LLM Reinforcement Learning Training with Verl
Train LLMs at scale using verl with PPO, GRPO, and other RL algorithms
Tags
Updated: 2026-06-30Capabilities
Typical Inputs
Typical Outputs
What this skill does
- train models with PPO
- train models with GRPO
- configure RL algorithms
- manage distributed training
- deploy with FSDP backend
- deploy with Megatron-LM backend
- process multi-turn rollouts
- train vision-language models
Inputs
- base model
- training dataset
- GPU cluster
- reward function
- training configuration
- algorithm selection
Outputs
- trained model checkpoints
- training metrics
- evaluation results
- reward logs
Requirements
- verl>=0.3.0
- torch>=2.0.0
- ray>=2.41.0
- vllm>=0.8.2
- transformers>=4.40.0
- 8+ GPUs
- Python environment
