verl-rl-training - Train LLMs with verl RL library
Flexible RL training library for large language models supporting multiple algorithms and backends
Tags
Updated: 2026-06-30Capabilities
Typical Inputs
Typical Outputs
What this skill does
- train LLM with GRPO algorithm
- train LLM with PPO algorithm
- train LLM with RLOO algorithm
- train LLM with DAPO algorithm
- train LLM with REINFORCE++ algorithm
- train LLM with ReMax algorithm
- configure FSDP backend
- configure Megatron-LM backend
- configure vLLM rollout engine
- configure SGLang rollout engine
- run distributed training
- configure LoRA training
- train vision-language models
- enable multi-turn rollout
- compute training rewards
Inputs
- training dataset
- base model
- training configuration
- reward function
- GPU cluster
- model checkpoint
Outputs
- trained model checkpoint
- training metrics
- reward scores
- loss curves
- evaluation results
Requirements
- GPU cluster with 8+ GPUs
- Python 3.x
- PyTorch 2.0.0+
- Ray 2.41.0+
- vLLM 0.8.2+
- verl 0.3.0+
