Updated 2026-09-15
Replaces hard trust-region masks with a smooth advantage-weighted quadratic regularizer to stabilize LLM reinforcement learning.
Browse skills that produce this output.
Updated 2026-09-15
Replaces hard trust-region masks with a smooth advantage-weighted quadratic regularizer to stabilize LLM reinforcement learning.