feat: OCR scans with tesseract, drop docling (#6) #31

Closed
eSlider wants to merge 1 commits from feat/ocr into main
Owner

Summary

  • pdftotext first; scans pdftoppm + tesseract eng+deu. Docling off the default path.
  • Optional OCR_ENGINE=paddle. Closes #6.

Counterpart: GitHub PR opens with this push.

## Summary - pdftotext first; scans pdftoppm + tesseract eng+deu. Docling off the default path. - Optional OCR_ENGINE=paddle. Closes #6. Counterpart: GitHub PR opens with this push.
eSlider added 1 commit 2026-08-14 11:23:35 +01:00
feat: OCR scans with tesseract, drop docling from the default path.
Tests / Test (push) Skipped
Tests / OCR (tesseract fixture) (push) Skipped
Tests / Release (semver) (push) Skipped
Tests / Test (pull_request) Failing after 6s
Tests / OCR (tesseract fixture) (pull_request) Failing after 4s
Tests / Release (semver) (pull_request) Skipped
0ee4106b99
pdftotext still wins on born-digital PDFs. Empty text layers go through
pdftoppm + tesseract eng+deu. Optional OCR_ENGINE=paddle. Gitea #6.
eSlider closed this pull request 2026-08-14 11:26:07 +01:00

Pull request closed

Please reopen this pull request to perform a merge.
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: eSlider/2dph#31