LogoClawIndex
CasesSkillsAbout
LogoClawIndex

verl-rl-training - Train LLMs with verl RL library

Flexible RL training library for large language models supporting multiple algorithms and backends

Tags

Updated: 2026-06-30

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • train LLM with GRPO algorithm
  • train LLM with PPO algorithm
  • train LLM with RLOO algorithm
  • train LLM with DAPO algorithm
  • train LLM with REINFORCE++ algorithm
  • train LLM with ReMax algorithm
  • configure FSDP backend
  • configure Megatron-LM backend
  • configure vLLM rollout engine
  • configure SGLang rollout engine
  • run distributed training
  • configure LoRA training
  • train vision-language models
  • enable multi-turn rollout
  • compute training rewards

Inputs

  • training dataset
  • base model
  • training configuration
  • reward function
  • GPU cluster
  • model checkpoint

Outputs

  • trained model checkpoint
  • training metrics
  • reward scores
  • loss curves
  • evaluation results

Requirements

  • GPU cluster with 8+ GPUs
  • Python 3.x
  • PyTorch 2.0.0+
  • Ray 2.41.0+
  • vLLM 0.8.2+
  • verl 0.3.0+

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.
Reinforcement Learning
RLHF
Model Training
Distributed Computing
Large Language Models
PPO
GRPO
train LLM with GRPO algorithm
train LLM with PPO algorithm
train LLM with RLOO algorithm
train LLM with DAPO algorithm
training dataset
base model
training configuration
trained model checkpoint
training metrics
reward scores