LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

Skills tagged: vision

Browse skills that share this tag.

  • draw-your-font - Turn a handwriting photo into an installable TTF font.
    fonthandwritingtypographyttf

    ★ 0 · Updated 2026-09-17

    Converts photos of handwritten letters into installable TTF and web font files.

    ⚙ Generate handwriting font template PDFs⚙ Segment handwriting photos into character crops⚙ Build TTF and web font files
  • ds-vision-skill - Vision extension skill for text reasoning models
    visionocrdocument-parsingrouting

    ★ 166 · Updated 2026-09-15

    Routes image, document, and OCR tasks to specialized execution tools and returns standard JSON output for text models.

    ⚙ Route vision reasoning tasks⚙ Parse document files⚙ Perform OCR text recognition
  • mmx-cli - MiniMax CLI tool for AI content generation and web search
    climinimaxai-generationspeech-synthesis

    ★ 0 · Updated 2026-09-15

    Generate text, images, video, speech, and music, analyze images, and perform web search via MiniMax AI using the mmx terminal CLI.

    ⚙ Generate text chat responses⚙ Generate images from text⚙ Generate videos from text
  • hermes-custom-vision-provider - Configure custom vision providers in Hermes Agent
    hermesvisionconfigurationtroubleshooting

    ★ 0 · Updated 2026-09-11

    Configures custom or OpenAI-compatible vision providers such as Kimi, DeepSeek, or Azure in Hermes Agent.

    ⚙ Configure custom vision provider settings⚙ Debug vision authentication errors⚙ Restart Hermes gateway process
  • claude-vision-skill - Analyze images using the vision helper script
    visionimage-analysisclipboardnode

    ★ 2,284 · Updated 2026-09-10

    Converts local images, image URLs, or clipboard images into text descriptions by running the bundled vision.js script.

    ⚙ Analyze local image files⚙ Analyze remote image URLs⚙ Read clipboard image content
  • mac-use - Control macOS GUI Apps Visually
    macOSGUIautomationOCR

    ★ 0 · Updated 2026-05-28

    Take screenshots, detect UI elements, click, scroll, type, and press keys to control macOS applications through their graphical interface.

    ⚙ capture screenshot⚙ detect text elements⚙ annotate image
  • mac-use - Control macOS GUI Applications Visually
    macOSGUIautomationOCR

    ★ 59 · Updated 2026-05-28

    Control macOS GUI applications through visual interaction using screenshots, clicks, scrolling, and typing

    ⚙ list visible windows⚙ capture screenshot⚙ detect text elements
  • stacks-ai - Stacks AI Integration Framework
    aillmanthropicopenai

    ★ 2 · Updated 2026-05-28

    Integrate AI capabilities into Stacks applications with multiple provider drivers

    ⚙ chat⚙ generate text⚙ generate image
  • stacks-ai - Stacks AI Integration
    AILLMRAGvector search

    ★ 1 · Updated 2026-05-28

    AI/LLM integration with multiple providers, RAG, image generation, and personalization features

    ⚙ configure⚙ chat⚙ stream
  • stacks-ai - Integrate AI into Stacks applications
    AILLMRAGimage generation

    ★ 620 · Updated 2026-05-28

    Integrate AI capabilities into Stacks apps with multiple drivers, image generation, RAG, embeddings, and personalization features

    ⚙ configure driver⚙ send chat message⚙ stream chat response
  • webcam - Webcam Capture for Mac Mini
    webcamvideoWhatsAppvision

    ★ 0 · Updated 2026-03-22

    Captures photos and videos from USB webcam, sends via WhatsApp, optionally describes content with vision model

    ⚙ capture photo⚙ capture video⚙ send via WhatsApp
  • minimax-coding-plan-tool - MiniMax Coding Plan Tool: Web Search & Image Understanding
    web searchimage understandingvisionAPI

    ★ 816 · Updated 2026-03-10

    A lightweight tool that directly calls MiniMax Coding Plan APIs for web search and image understanding using pure JavaScript.

    ⚙ search web⚙ understand image
  • gem - Gemini Multimodal AI Processing Tool
    multimodalaigeminidocument-processing

    ★ 0 · Updated 2026-02-11

    Processes multimodal content using Google Gemini AI

    ⚙ process text queries⚙ analyze documents⚙ analyze images
  • gem - Gemini Multimodal AI Content Processing
    multimodalAIGeminidocument-processing

    ★ 602 · Updated 2026-02-11

    Process multimodal content using Google Gemini API for analysis and information extraction

    ⚙ analyze PDF documents⚙ process image files⚙ analyze video content