A general reasoning and code generation model. Trained with process reward models and reinforcement learning from human feedback for reliable multi-step problem solving. Excels at refactoring, debugging, agentic tool use, and structured analysis.
Intended use
- Code generation
- Refactoring
- Debugging
- Agentic workflows
- General reasoning