LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.

document-ocr-agent - Extract text and fields from documents

Uses Google Document AI to extract OCR text, structured entities, and page-level metadata from PDFs, images, receipts, invoices, and scanned documents.

Tags

Updated: 2026-09-29

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • Process documents with OCR
  • Extract structured entities
  • Return page-level metadata
  • Batch image inputs
  • Select document processors

Inputs

  • Document file URL
  • Cloud file ID
  • Base64-encoded file content
  • Document type
  • MIME type
  • Output limits

Outputs

  • Extracted OCR text
  • Structured entities
  • Per-page summary metadata
  • Raw Document AI response

Requirements

  • AgentPMT-hosted remote tool access
  • process_document action

Source

  • Spec: SKILL.md
OCR
document processing
Google Document AI
receipt parsing
invoice extraction
scanned documents
Process documents with OCR
Extract structured entities
Return page-level metadata
Batch image inputs
Document file URL
Cloud file ID
Base64-encoded file content
Extracted OCR text
Structured entities
Per-page summary metadata