LogoClawIndex
CasesSkillsAbout
LogoClawIndex

unsloth-inference - Deploy Fine-Tuned Models for Production Inference

Deploy fine-tuned models for production inference using optimized kernels, vLLM, or SGLang

Tags

Updated: 2026-02-15

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • load fine-tuned model
  • enable optimized kernels
  • merge LoRA weights
  • deploy vLLM server
  • create OpenAI-compatible API
  • generate model responses

Inputs

  • fine-tuned model
  • tokenizer
  • input prompts
  • server configuration
  • LoRA adapters

Outputs

  • inference results
  • OpenAI-compatible endpoint
  • served model files
  • API responses

Requirements

  • unsloth library
  • torch installed
  • Python environment
  • vLLM optional
  • SGLang optional

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.
inference
model deployment
optimization
serving
api
load fine-tuned model
enable optimized kernels
merge LoRA weights
deploy vLLM server
fine-tuned model
tokenizer
input prompts
inference results
OpenAI-compatible endpoint
served model files