LogoClawIndex
CasesSkillsAbout
LogoClawIndex

verl-rl-training - LLM Reinforcement Learning Training with Verl

Train LLMs at scale using verl with PPO, GRPO, and other RL algorithms

Tags

Updated: 2026-06-30

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • train models with PPO
  • train models with GRPO
  • configure RL algorithms
  • manage distributed training
  • deploy with FSDP backend
  • deploy with Megatron-LM backend
  • process multi-turn rollouts
  • train vision-language models

Inputs

  • base model
  • training dataset
  • GPU cluster
  • reward function
  • training configuration
  • algorithm selection

Outputs

  • trained model checkpoints
  • training metrics
  • evaluation results
  • reward logs

Requirements

  • verl>=0.3.0
  • torch>=2.0.0
  • ray>=2.41.0
  • vllm>=0.8.2
  • transformers>=4.40.0
  • 8+ GPUs
  • Python environment

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.
Reinforcement Learning
RLHF
Post-Training
Distributed Training
PPO
GRPO
Large Language Models
Machine Learning
train models with PPO
train models with GRPO
configure RL algorithms
manage distributed training
base model
training dataset
GPU cluster
trained model checkpoints
training metrics
evaluation results