blip-2-vision-language - BLIP-2 vision-language task framework
Supports image captioning, visual question answering, image-text matching, feature extraction, and multimodal understanding with frozen vision and language models.
Tags
Updated: 2026-10-04Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Generate image captions
- Answer visual questions
- Match images and text
- Extract image features
- Process image batches
- Control text generation
Inputs
- Image files
- Visual questions
- Text prompts
- Model checkpoints
Outputs
- Generated captions
- Visual question answers
- Image-text matching scores
- Image feature embeddings
- Generated text
Requirements
- transformers>=4.30.0
- torch>=1.10.0
- Pillow
- Supported OPT or FlanT5 model
