LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

Skills with capability: Design DRPO quadratic regularizer

Browse skills that share this capability.

  • drpo-llm-rl - DRPO for stable LLM reinforcement learning
    DRPOLLM RLReinforcement LearningPolicy Optimization

    Updated 2026-09-15

    Replaces hard trust-region masks with a smooth advantage-weighted quadratic regularizer to stabilize LLM reinforcement learning.

    ⚙ Assess current RL training setup⚙ Design DRPO quadratic regularizer⚙ Implement smooth policy regularization