★ 85 · Updated 2026-10-04
Supports image captioning, visual question answering, image-text matching, feature extraction, and multimodal understanding with frozen vision and language models.
Browse skills that share this capability.
★ 85 · Updated 2026-10-04
Supports image captioning, visual question answering, image-text matching, feature extraction, and multimodal understanding with frozen vision and language models.