LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

verl-rl-training - RL Training Library for LLMs

Provides guidance for training LLMs with RL using verl library

Tags

Updated: 2026-06-30

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • train models with RLHF
  • train models with GRPO
  • train models with PPO
  • train models with RLOO
  • train models with REINFORCE++
  • train models with DAPO
  • train models with ReMax
  • train models with SPIN
  • train models with SPPO
  • switch training backends
  • perform rollout generation
  • compute rewards
  • train vision-language models
  • execute multi-turn rollout
  • call tools during rollout
  • use FSDP backend
  • use Megatron-LM backend
  • use vLLM backend
  • use SGLang backend
  • configure LoRA training

Inputs

  • training dataset
  • base model
  • reward function
  • training configuration
  • GPU cluster
  • model checkpoint

Outputs

  • trained model checkpoint
  • training metrics
  • evaluation results

Requirements

  • Python environment
  • verl>=0.3.0
  • torch>=2.0.0
  • ray>=2.41.0
  • vllm>=0.8.2
  • transformers>=4.40.0
  • GPU cluster with 8+ GPUs

Source

  • Spec: SKILL.md
Reinforcement Learning
RLHF
GRPO
PPO
Post-Training
Distributed Training
Machine Learning
LLM
train models with RLHF
train models with GRPO
train models with PPO
train models with RLOO
training dataset
base model
reward function
trained model checkpoint
training metrics
evaluation results