Vram-GPU-OOM - GPU VRAM Sharing and OOM Management
Manage GPU VRAM sharing with OOM retry and auto-unload for multiple services
Tags
Updated: 2026-03-22Capabilities
Typical Inputs
Typical Outputs
What this skill does
- catch OOM errors
- retry model loading
- wait between retries
- configure auto-unload
- implement signaling protocol
- request model unload
- check service status
Inputs
- GPU memory
- model files
- service endpoints
Outputs
- loaded model
- unload request response
- service status
Requirements
- CUDA GPU
- PyTorch
- Python runtime
