gguf-quantization - GGUF Quantization for Efficient Model Inference
Convert and quantize models to GGUF format for efficient CPU/GPU inference
Tags
Updated: 2026-06-30Capabilities
Typical Inputs
Typical Outputs
What this skill does
- convert model to GGUF
- quantize GGUF model
- run model inference
- generate importance matrix
- start inference server
- load GGUF model
Inputs
- HuggingFace model
- quantization type
- calibration text
- GGUF model file
Outputs
- GGUF model file
- quantized model file
- model inference output
- importance matrix file
Requirements
- llama.cpp
- llama-cpp-python
- compatible hardware
- C/C++ compiler
