document-ocr-agent - Extract text and fields from documents
Uses Google Document AI to extract OCR text, structured entities, and page-level metadata from PDFs, images, receipts, invoices, and scanned documents.
Tags
Updated: 2026-09-29Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Process documents with OCR
- Extract structured entities
- Return page-level metadata
- Batch image inputs
- Select document processors
Inputs
- Document file URL
- Cloud file ID
- Base64-encoded file content
- Document type
- MIME type
- Output limits
Outputs
- Extracted OCR text
- Structured entities
- Per-page summary metadata
- Raw Document AI response
Requirements
- AgentPMT-hosted remote tool access
- process_document action
