LogoClawIndex
CasesSkillsAbout
LogoClawIndex

blip-2-vision-language - BLIP-2 vision-language task support

Use BLIP-2 for image captioning, visual question answering, image-text retrieval, feature extraction, and multimodal chat.

Tags

Updated: 2026-10-05
MultimodalVision-languageImage captioningVisual question answeringImage-text retrievalZero-shot

Capabilities

Generate image captionsAnswer visual questionsRetrieve image-text matchesExtract image features

Typical Inputs

Input imagesVisual questionsText prompts

Typical Outputs

Generated captionsQuestion answersMatching scores

What this skill does

  • Generate image captions
  • Answer visual questions
  • Retrieve image-text matches
  • Extract image features
  • Process image batches
  • Control text generation

Inputs

  • Input images
  • Visual questions
  • Text prompts
  • Model checkpoints

Outputs

  • Generated captions
  • Question answers
  • Matching scores
  • Image features
  • Console messages

Requirements

  • Python environment
  • Transformers package
  • PyTorch package
  • Pillow package
  • BLIP-2 model checkpoint

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.