Scanned document OCR
In developmentExtract searchable text from scanned PDFs using Tesseract. The backend hook is straightforward to wire up; this page needs a review UI so you can check and correct recognized text before exporting.
Planned for this tool:
- Extract text from scanned or image-based PDFs
- English, Hindi, Arabic, and Urdu language support
- Search within OCR'd text
- Export recognized text as TXT, DOCX, or a searchable PDF