★ 0 · Updated 2026-09-22
High-throughput LLM serving engine supporting OpenAI compatible API, quantization, and tensor parallelism.
Browse skills that share this tag.
★ 0 · Updated 2026-09-22
High-throughput LLM serving engine supporting OpenAI compatible API, quantization, and tensor parallelism.
★ 602 · Updated 2026-06-15
Validates OpenAI API implementations against official specification including endpoints, parameters, and authentication
★ 0 · Updated 2026-05-09
Starts and manages a local LLM inference stack with OpenAI-compatible API for Linux and macOS