Updated 2026-09-15
Replaces hard trust-region masks with a smooth advantage-weighted quadratic regularizer to stabilize LLM reinforcement learning.
Browse skills that share this capability.
Updated 2026-09-15
Replaces hard trust-region masks with a smooth advantage-weighted quadratic regularizer to stabilize LLM reinforcement learning.