★ 0 · Updated 2026-06-30
Convert and quantize models to GGUF format for efficient CPU/GPU inference
Browse skills that use this input.
★ 0 · Updated 2026-06-30
Convert and quantize models to GGUF format for efficient CPU/GPU inference
★ 0 · Updated 2026-05-11
Run local GGUF models and discover Hugging Face repositories for llama.cpp
★ 1 · Updated 2026-05-11
Run GGUF models locally and discover models from Hugging Face Hub
★ 18 · Updated 2026-03-10
Secondary local LLM inference engine for running GGUF models directly, loading LoRA adapters, benchmarking, and custom model serving