sglang-diffusion-modelopt-quant - Quantize and validate ModelOpt diffusion checkpoints
Quantizes diffusion DiTs with NVIDIA ModelOpt, converts FP8 or NVFP4 exports for SGLang Diffusion, and validates quality and performance.
Tags
Updated: 2026-10-03Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Quantize diffusion DiTs with ModelOpt
- Convert FP8 exports for SGLang
- Build NVFP4 transformer checkpoints
- Verify trajectory similarity
- Benchmark quantized performance
- Update support matrix documentation
Inputs
- Diffusion model checkpoint
- Quantization format
- Calibration prompt file
- Calibration settings
- Baseline run configuration
- Model component paths
- Output directories
- Benchmark commands
Outputs
- Quantized ModelOpt checkpoint
- SGLang-loadable transformer checkpoint
- BF16 and quantized quality results
- Performance benchmark results
- Updated quantization support matrix
Requirements
- NVIDIA ModelOpt
- SGLang Diffusion environment
- Official ModelOpt quantization script
- Validated SGLang helper tools
- Documented supported model family
