★ 0 · Updated 2026-09-24
Provides methods and workflows for LLM post-training, including SFT, DPO, PPO, GRPO, and reward modeling.
Browse skills that share this tag.
★ 0 · Updated 2026-09-24
Provides methods and workflows for LLM post-training, including SFT, DPO, PPO, GRPO, and reward modeling.
★ 0 · Updated 2026-09-24
Provides guidance and workflows for LLM post-training with RL using the slime framework integrating Megatron-LM and SGLang.
★ 3 · Updated 2026-09-15
Replaces hard trust-region masks with a smooth advantage-weighted quadratic regularizer to stabilize LLM reinforcement learning.
★ 602 · Updated 2026-06-30
Provides guidance for training LLMs with RL using verl library
★ 0 · Updated 2026-06-30
Train LLMs at scale using verl with PPO, GRPO, and other RL algorithms
★ 1 · Updated 2026-06-30
Flexible RL training library for large language models supporting multiple algorithms and backends
★ 0 · Updated 2026-06-30
Train LLMs with RLHF, GRPO, and PPO using verl library
★ 1 · Updated 2026-06-30
Flexible RL training library for large language models
★ 0 · Updated 2026-06-30
Train large language models with reinforcement learning algorithms using the verl library
★ 602 · Updated 2026-05-11
Enterprise-grade RL training framework for large MoE models with FP8/INT4 support