★ 0 · Updated 2026-09-24
Provides methods and workflows for LLM post-training, including SFT, DPO, PPO, GRPO, and reward modeling.
Browse skills that use this input.
★ 0 · Updated 2026-09-24
Provides methods and workflows for LLM post-training, including SFT, DPO, PPO, GRPO, and reward modeling.