★ 174 · Updated 2026-09-22
Review and upgrade MetaX model support against target vLLM revisions and installed MACA components, including model-dependent attention and kernels.
Browse skills that share this tag.
★ 174 · Updated 2026-09-22
Review and upgrade MetaX model support against target vLLM revisions and installed MACA components, including model-dependent attention and kernels.
★ 1 · Updated 2026-09-13
Runs evaluations against Hugging Face Hub models on local hardware using inspect-ai or lighteval.
★ 3 · Updated 2026-09-11
Presents the Axon DSL for compiling shape-safe, framework-agnostic LLM architectures to PyTorch, JAX, MLX, and vLLM.
★ 2 · Updated 2026-05-11
Adapt Hugging Face or local models to run on vllm-ascend with minimal changes
★ 602 · Updated 2026-03-26
Scale LLM inference horizontally on Kubernetes with GPU-aware autoscaling, request queuing, and spot instance strategies.
★ 0 · Updated 2026-02-23
Deploy local or open-source AI models with model selection, quantization, Ollama/vLLM setup, and integration with ModelSelector
★ 602 · Updated 2026-02-14
Deploy fine-tuned models for production inference with optimized kernels and serving engines.