llama-cpp - Local GGUF Inference & Model Discovery
Run local GGUF models and discover Hugging Face repositories for llama.cpp
Tags
Updated: 2026-05-11Capabilities
Typical Inputs
Typical Outputs
What this skill does
- run local models
- search Hugging Face Hub
- discover GGUF files
- generate text
- create chat completions
- generate embeddings
- start llama-server
- run llama-cli
- list available GGUFs
- query tree API
- read repo files
Inputs
- GGUF model file
- Hugging Face repository
- quantization variant
- context length
- GPU layers count
- chat messages
- text prompt
- repository URL
Outputs
- generated text
- chat response
- text embedding
- server endpoint
- GGUF file list
- CLI command
- discovery result
Requirements
- llama.cpp
- llama-cpp-python
- CPU or GPU
- RAM or VRAM
- Linux/macOS/Windows
