teoria-engine - Production-grade self-hosted LLM inference stack
Starts and manages a local LLM inference stack with OpenAI-compatible API for Linux and macOS
Tags
Updated: 2026-05-09Capabilities
Typical Inputs
Typical Outputs
What this skill does
- start inference engine
- stop inference engine
- check engine status
- switch model profile
- call inference API
- install engine
- configure settings
Inputs
- NVIDIA GPU or Apple Silicon
- Docker environment
- model profile
- API key
- GitHub repository
- HuggingFace token
Outputs
- inference API endpoint
- chat completion text
- streaming responses
- vision-language outputs
- engine status messages
Requirements
- Linux (Ubuntu 22.04+) or macOS Apple Silicon
- NVIDIA GPU with CUDA 12.4+ or Apple Silicon GPU
- Docker with Docker Compose v2+
- Python 3.10+
- Git and Bash
