LogoClawIndex
CasesSkillsAbout
LogoClawIndex

grpo-rl-training - Train language models with GRPO and TRL

Provides guidance for implementing GRPO reinforcement learning fine-tuning with TRL, custom reward functions, and task-specific training workflows.

Tags

Updated: 2026-10-03
Post-trainingReinforcement learningGRPOTRLRLHFReward modelingReasoningDPOPPOStructured output

Capabilities

Prepare chat-format datasetsDefine composite reward functionsConfigure GRPO trainingSet up causal language models

Typical Inputs

Training datasetBase language modelReward functions

Typical Outputs

Trained modelSaved model filesTraining logs

What this skill does

  • Prepare chat-format datasets
  • Define composite reward functions
  • Configure GRPO training
  • Set up causal language models
  • Apply LoRA fine-tuning
  • Save trained models

Inputs

  • Training dataset
  • Base language model
  • Reward functions
  • Training configuration
  • Ground-truth answers

Outputs

  • Trained model
  • Saved model files
  • Training logs
  • Training checkpoints

Requirements

  • Transformers 4.47.0 or later
  • TRL 0.14.0 or later
  • Datasets 3.2.0 or later
  • PEFT 0.14.0 or later
  • PyTorch

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.