post-ocr-cleanup - Clean and evaluate post-OCR text
Correct OCR errors, apply deterministic cleanup, preserve provenance, and diagnose text quality across languages.
Tags
Updated: 2026-10-05Capabilities
Typical Inputs
Typical Outputs
What this skill does
- Select correction strategies
- Correct OCR text
- Apply deterministic fixes
- Restore language-specific diacritics
- Constrain decoding output
- Trim model overgeneration
- Measure edit distances
- Diagnose page quality
- Flag suspicious pages
- Preserve raw text
Inputs
- OCR text
- Source language metadata
- Document era metadata
- Document type metadata
- Typeface information
- Scan specifications
- OCR engine information
- Document images
- Downstream quality threshold
- Evaluation ground truth
Outputs
- Cleaned OCR text
- Raw text copies
- Provenance records
- Edit-distance metrics
- Change ratios
- Page quality diagnostics
- Manual-review flags
Requirements
- LLM access for model correction
- GPU resources for model inference
- Model logits for token-level CBS
- Reference prompt and schema files
