ocr_kb - Extract PDF pages and generate DOCX
Reads PDF pages with a multimodal model, extracts text, LaTeX, and figures, and incrementally produces reviewed DOCX files with checkpoint recovery.
Tags
Updated: 2026-09-30Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Read PDF pages visually
- Extract text and LaTeX
- Crop figures from pages
- Generate incremental DOCX files
- Track checkpoints and counters
- Review every two pages
- Rerun selected pages
Inputs
- Source PDF file
- Formatting specification
- Optional DOCX template
- Resume or rerun instruction
Outputs
- Page PNG files
- Per-page Markdown files
- Cropped figure files
- Compiled Markdown file
- Checkpoint JSON file
- Checkpoint DOCX files
- Final DOCX file
- Completion report
Requirements
- Multimodal image-reading model
- Python runtime
- Write access to resources/
- Write access to outputs/
