llm-integration - Secure Local LLM Integration
Integrates local LLMs with llama.cpp and Ollama, covering secure model loading, inference optimization, prompt handling, and vulnerability protection.
Tags
Updated: 2026-10-07Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Load models securely
- Optimize inference performance
- Sanitize user prompts
- Prevent prompt injection
- Filter generated outputs
- Enforce resource limits
- Stream filtered responses
- Verify model integrity
- Handle model failures
Inputs
- User prompts
- Local model files
- Model selection
- Inference configuration
- Ollama endpoint configuration
Outputs
- Generated responses
- Filtered responses
- Injection warnings
- Inference metrics
- Fallback responses
Requirements
- llama.cpp b2500 or newer
- Ollama 0.1.34 or newer
- llama-cpp-python 0.2.72 or newer
- ollama-python 0.4.0 or newer
- pydantic 2.0 or newer
- jinja2 3.1.3 or newer
- tiktoken 0.5.0 or newer
- structlog 23.0 or newer
- Sandboxed model execution
- Access to approved model directory
