LogoClawIndex
CasesSkillsAbout
LogoClawIndex

ocr_kb - Extract PDF pages and generate DOCX

Reads PDF pages with a multimodal model, extracts text, LaTeX, and figures, and incrementally produces reviewed DOCX files with checkpoint recovery.

Tags

Updated: 2026-09-30

Capabilities

Typical Inputs

Typical Outputs

What this skill does

  • Read PDF pages visually
  • Extract text and LaTeX
  • Crop figures from pages
  • Generate incremental DOCX files
  • Track checkpoints and counters
  • Review every two pages
  • Rerun selected pages

Inputs

  • Source PDF file
  • Formatting specification
  • Optional DOCX template
  • Resume or rerun instruction

Outputs

  • Page PNG files
  • Per-page Markdown files
  • Cropped figure files
  • Compiled Markdown file
  • Checkpoint JSON file
  • Checkpoint DOCX files
  • Final DOCX file
  • Completion report

Requirements

  • Multimodal image-reading model
  • Python runtime
  • Write access to resources/
  • Write access to outputs/

Source

  • Spec: SKILL.md

ClawIndex

OpenClaw Skills & Use Case Index

ClawIndex is an ecosystem-driven index of OpenClaw skills and real-world use cases.

Index

Skills·
Cases

Meta

About·
Disclaimer·
Email·
GitHub
© 2026 ClawIndex All Rights Reserved.
OCR
PDF processing
multimodal extraction
LaTeX extraction
figure cropping
DOCX generation
checkpoint recovery
Read PDF pages visually
Extract text and LaTeX
Crop figures from pages
Generate incremental DOCX files
Source PDF file
Formatting specification
Optional DOCX template
Page PNG files
Per-page Markdown files
Cropped figure files