LogoClawIndex
CasesSkillsAbout
LogoClawIndex

blip-2-vision-language - BLIP-2 vision-language task framework

Supports image captioning, visual question answering, image-text matching, feature extraction, and multimodal understanding with frozen vision and language models.

Tags

Updated: 2026-10-04

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • Generate image captions
  • Answer visual questions
  • Match images and text
  • Extract image features
  • Process image batches
  • Control text generation

Inputs

  • Image files
  • Visual questions
  • Text prompts
  • Model checkpoints

Outputs

  • Generated captions
  • Visual question answers
  • Image-text matching scores
  • Image feature embeddings
  • Generated text

Requirements

  • transformers>=4.30.0
  • torch>=1.10.0
  • Pillow
  • Supported OPT or FlanT5 model

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.
Multimodal
Vision-Language
Image Captioning
VQA
Zero-Shot
Generate image captions
Answer visual questions
Match images and text
Extract image features
Image files
Visual questions
Text prompts
Generated captions
Visual question answers
Image-text matching scores