Updated 2026-09-15
Replaces hard trust-region masks with a smooth advantage-weighted quadratic regularizer to stabilize LLM reinforcement learning.
Browse skills that use this input.
Updated 2026-09-15
Replaces hard trust-region masks with a smooth advantage-weighted quadratic regularizer to stabilize LLM reinforcement learning.