★ 0 · Updated 2026-09-22
High-throughput LLM serving engine supporting OpenAI compatible API, quantization, and tensor parallelism.
Browse skills that share this tag.
★ 0 · Updated 2026-09-22
High-throughput LLM serving engine supporting OpenAI compatible API, quantization, and tensor parallelism.
★ 1 · Updated 2026-06-30
Flexible RL training library for large language models
★ 4 · Updated 2026-05-28
Analyze model code and deployment config to identify CPU/GPU optimization opportunities for KServe inference services
★ 0 · Updated 2026-05-11
Adapt and debug Hugging Face or local models for vLLM on Ascend NPU
★ 0 · Updated 2026-05-11
Adapt and debug models to run vLLM on Ascend NPU
★ 57 · Updated 2026-05-11
Adapt and debug models for vLLM on Ascend NPU.
★ 2,882 · Updated 2026-05-11
Adapt and debug Hugging Face or local models to run on vLLM-Ascend with minimal changes and deterministic validation
★ 602 · Updated 2026-03-26
Fine-tune LLMs with Unsloth using 4-bit quantization and LoRA