LogoClawIndex
CasesSkillsAbout
LogoClawIndex

gem - Gemini Multimodal AI Content Processing

Process multimodal content using Google Gemini API for analysis and information extraction

Tags

Updated: 2026-02-11
multimodalAIGeminidocument-processingvisioncontent-analysis

Capabilities

analyze PDF documentsprocess image filesanalyze video contentextract YouTube information

Typical Inputs

PDF filesimage filesvideo files

Typical Outputs

analysis reportsdocument summariescomparison results

What this skill does

  • analyze PDF documents
  • process image files
  • analyze video content
  • extract YouTube information
  • compare multiple files
  • perform web searches

Inputs

  • PDF files
  • image files
  • video files
  • YouTube URLs
  • text documents
  • search queries

Outputs

  • analysis reports
  • document summaries
  • comparison results
  • extracted information
  • search results

Requirements

  • GEMINI_API_KEY environment variable
  • hamel Python package
  • Google Gemini API access

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.