Vram-GPU-OOM - GPU VRAM Management for Multi-Service Sharing
Manage GPU VRAM sharing across services with OOM retry and auto-unload
Tags
Updated: 2026-03-23Capabilities
Typical Inputs
Typical Outputs
What this skill does
- load model
- catch OOM error
- retry model load
- clear GPU cache
- unload model
- request model unload
Inputs
- model path
- GPU environment
- service configuration
- service URLs
Outputs
- loaded model
- model unloaded
- unload request result
- service status
Requirements
- PyTorch
- GPU
- Python runtime
