★ 6 · Updated 2026-10-01
Controls LLM reasoning effort with High, Medium, and Low modes using budget-aware fine-tuning and adaptive reward shaping.
Browse skills that use this input.
★ 6 · Updated 2026-10-01
Controls LLM reasoning effort with High, Medium, and Low modes using budget-aware fine-tuning and adaptive reward shaping.
★ 85 · Updated 2026-09-28
Fine-tunes and serves Physical Intelligence OpenPI models with JAX or PyTorch for robot policy inference in ALOHA, DROID, and LIBERO environments.
★ 6 · Updated 2026-09-21
Generate multiple tokens simultaneously by having late transformer layers directly predict multiple outputs after early layer processing.
★ 18 · Updated 2026-09-21
Set up infrastructure for fine-tuning LLMs using QLoRA, LoRA, Hugging Face TRL, Axolotl, DeepSpeed, or FSDP.
★ 1 · Updated 2026-06-30
Flexible RL training library for large language models
★ 650 · Updated 2026-05-09
Train scikit-learn ML models with cross-validation, hyperparameter tuning, and pipelines