LogoClawIndex
CasesSkillsAbout
LogoClawIndex

verl-rl-training - LLM Reinforcement Learning Training with verl

Train LLMs with RLHF, GRPO, and PPO using verl library

Tags

Updated: 2026-06-30
Reinforcement LearningRLHFGRPOPPOPost-TrainingDistributed Training

Capabilities

train LLMsconfigure algorithmprepare datasetdefine reward function

Typical Inputs

DatasetBase modelReward function

Typical Outputs

Trained modelTraining metricsEvaluation results

What this skill does

  • train LLMs
  • configure algorithm
  • prepare dataset
  • define reward function
  • create training config
  • launch training
  • monitor training
  • run evaluation
  • manage cluster
  • convert model

Inputs

  • Dataset
  • Base model
  • Reward function
  • Training config
  • GPU cluster

Outputs

  • Trained model
  • Training metrics
  • Evaluation results

Requirements

  • verl>=0.3.0
  • torch>=2.0.0
  • ray>=2.41.0
  • vllm>=0.8.2
  • transformers>=4.40.0
  • GPU cluster
  • Python environment

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.