★ 0 · Updated 2026-09-22
High-throughput LLM serving engine supporting OpenAI compatible API, quantization, and tensor parallelism.
Browse skills that share this tag.
★ 0 · Updated 2026-09-22
High-throughput LLM serving engine supporting OpenAI compatible API, quantization, and tensor parallelism.
★ 0 · Updated 2026-06-30
Convert and quantize models to GGUF format for efficient CPU/GPU inference
★ 0 · Updated 2026-05-11
Run local GGUF models and discover Hugging Face repositories for llama.cpp
★ 17 · Updated 2026-05-11
Local GGUF inference, quant selection, and Hugging Face Hub model discovery for llama.cpp
★ 1 · Updated 2026-05-11
Run GGUF models locally and discover models from Hugging Face Hub