★ 22 · Updated 2026-10-03
Trains reinforcement learning agents with PufferLib using parallel vectorized environments, native multi-agent support, custom PufferEnv tasks, and integrated environments.
Browse skills that use this input.
★ 22 · Updated 2026-10-03
Trains reinforcement learning agents with PufferLib using parallel vectorized environments, native multi-agent support, custom PufferEnv tasks, and integrated environments.
★ 1 · Updated 2026-10-03
Provides guidance for implementing GRPO reinforcement learning fine-tuning with TRL, custom reward functions, and task-specific training workflows.
★ 85 · Updated 2026-09-28
Fine-tunes and serves Physical Intelligence OpenPI models with JAX or PyTorch for robot policy inference in ALOHA, DROID, and LIBERO environments.
★ 0 · Updated 2026-09-24
Provides methods and workflows for LLM post-training, including SFT, DPO, PPO, GRPO, and reward modeling.
★ 1,182 · Updated 2026-02-18
Add a new model type to the Numerai agents training pipeline for training and evaluation
★ 22 · Updated 2026-02-15
Algorithm and model development skill for dataset design, fine-tuning, evaluation, and deployment