LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

Skills tagged: Inference Serving

Browse skills that share this tag.

  • serving-llms-vllm - High-throughput LLM serving framework
    vLLMInference ServingPagedAttentionContinuous Batching

    ★ 0 · Updated 2026-09-22

    High-throughput LLM serving engine supporting OpenAI compatible API, quantization, and tensor parallelism.

    ⚙ Deploy production LLM APIs⚙ Run offline batch inference⚙ Serve quantized LLM models
  • tensorrt-llm - Optimize LLM inference on NVIDIA GPUs using TensorRT-LLM
    Inference ServingTensorRT-LLMNVIDIAInference Optimization

    ★ 0 · Updated 2026-09-09

    Optimizes LLM inference performance on NVIDIA GPUs with quantization, in-flight batching, and multi-GPU scaling.

    ⚙ Generate text responses from prompts⚙ Serve models via HTTP API⚙ Quantize models to lower precision
  • tensorrt-llm - Optimizing LLM Inference with NVIDIA TensorRT
    Inference ServingTensorRT-LLMNVIDIAInference Optimization

    ★ 0 · Updated 2026-09-09

    Optimizes LLM inference on NVIDIA GPUs using quantization, in-flight batching, and multi-GPU scaling.

    ⚙ Run LLM inference⚙ Serve models via HTTP server⚙ Quantize models with FP8
  • tensorrt-llm - TensorRT-LLM LLM Inference Optimization
    Inference ServingTensorRT-LLMNVIDIAInference Optimization

    ★ 2 · Updated 2026-03-19

    Optimizes LLM inference with NVIDIA TensorRT for high-throughput, low-latency production deployment on NVIDIA GPUs.

    ⚙ optimize LLM inference⚙ compile models with TensorRT⚙ serve models via HTTP API