unsloth-inference - Deploy fine-tuned models for production inference with Unsloth
Deploy fine-tuned models for production inference with optimized kernels and serving engines.
Tags
Updated: 2026-02-14Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Load fine-tuned model
- Load tokenizer
- Enable optimized kernels
- Generate text responses
- Select export method
- Merge LoRA weights
- Start vLLM server
- Create OpenAI-compatible API
Inputs
- Fine-tuned model files
- Tokenizer
- LoRA adapters
- Base model
- vLLM configuration
- SGLang configuration
- OpenAI API specifications
Outputs
- Optimized inference model
- Merged model files
- vLLM server endpoint
- OpenAI-compatible API
- Text generation responses
- Model serving endpoints
Requirements
- Unsloth library installed
- PyTorch installed
- vLLM (optional for production)
- SGLang (optional for production)
- GPU with sufficient VRAM
- Fine-tuned model files available
