thinkdial-reasoning-effort-control - ThinkDial: Control LLM Reasoning Effort
Controls LLM reasoning effort with High, Medium, and Low modes using budget-aware fine-tuning and adaptive reward shaping.
Tags
Updated: 2026-10-01Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Define discrete reasoning modes
- Set token budgets
- Train with budget-aware SFT
- Run two-phase reinforcement learning
- Shape rewards by budget
- Switch reasoning modes at runtime
Inputs
- Trainable LLM model
- Training dataset
- Reasoning mode
- Input prompts
- Ground-truth responses
- Training hyperparameters
Outputs
- Generated responses
- Training losses
- Reward and loss metrics
- Updated model parameters
Requirements
- Python with PyTorch
- Trainable LLM implementation
- Generation API
- Log-probability API
