feat: OCR scans with tesseract, drop docling from the default path. (#33)
Tests / Test (push) Failing after 5s
Tests / OCR (tesseract fixture) (push) Failing after 4s
Tests / Release (semver) (push) Skipped

pdftotext still wins on born-digital PDFs. Empty text layers go through
pdftoppm + tesseract eng+deu. Optional OCR_ENGINE=paddle. Gitea #6.
This commit is contained in:
2026-08-14 11:26:04 +01:00
committed by GitHub
co-authored by GitHub
parent bae1494258
commit f99dfea104
17 changed files with 536 additions and 1743 deletions
+10
View File
@@ -5,6 +5,7 @@
# docker compose --profile picoclaw up brain-mcp
# docker compose --profile reasoner up -d reasoner # CPU Ollama :11435
# docker compose --profile searxng up -d
# OCR_ENGINE=paddle docker compose --profile ocr-paddle run --rm ocr-paddle
#
# Secrets never baked in: search.env + db-profiles.yml from ~/.config/brain.
@@ -131,6 +132,15 @@ services:
- reasoner-ollama:/root/.ollama
restart: unless-stopped
# Optional PP-OCRv5 (not default). Default OCR is tesseract eng+deu.
# OCR_ENGINE=paddle docker compose --profile ocr-paddle run --rm ocr-paddle
ocr-paddle:
profiles: ["ocr-paddle"]
image: python:3.12-slim
environment:
OCR_ENGINE: paddle
command: ["python", "-c", "print('OCR_ENGINE=paddle; install paddleocr on PATH')"]
volumes:
kb-model:
kb-var: