pdf - Extract, create, merge, split, and process PDF documents.
Extract text and tables, create, merge, split, rotate, encrypt, and fill forms in PDF documents using Python libraries and CLI tools.
Tags
Updated: 2026-09-09Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Extract text and tables from PDFs
- Merge multiple PDF files
- Split PDF pages into files
- Create new PDF documents
- Rotate pages in PDF documents
- Extract images from PDF files
- Perform OCR on scanned PDFs
- Encrypt PDF files with passwords
- Generate scientific schematic diagrams
Inputs
- Input PDF documents
- Scanned PDF files
- Schematic diagram descriptions
Outputs
- Extracted text and table data
- Merged or split PDF files
- Generated PDF documents
- Extracted image files
- Generated schematic diagram images
Requirements
- Python 3 environment
- Python libraries (pypdf, pdfplumber, reportlab, pandas)
- Command-line tools (qpdf, poppler-utils, pdftk)
- Tesseract OCR engine for scanned PDFs
