Compare commits

..
5 Commits
Author SHA1 Message Date
eSliderandGitHub 0c4bb87001 feat: parse all Go CLIs with flaggy; dump bash complete. (#36)
Tests / Test (push) Skipped
Tests / OCR (tesseract fixture) (push) Skipped
Tests / Release (semver) (push) Skipped
stdlib flag dropped --hop after the query. One wrapper in internal/cli,
source <(./bin/cli/complete.go bash). Gitea #34.
2026-08-14 12:12:38 +01:00
eSliderandGitHub 0065655e03 feat: D16 adjudication — 2v2 stays hypothesis until a rule fires. (#35)
Tests / Test (push) Skipped
Tests / OCR (tesseract fixture) (push) Skipped
Tests / Release (semver) (push) Skipped
temporal_freshness then authority_pairing; unresolved keeps (not confirmed).
bin/facts/audit contradict. Gitea #29.
2026-08-14 11:43:00 +01:00
eSliderandGitHub a331042488 feat: DuckDB quantiles in-process (gcc CGO), not Ladybug. (#34)
Tests / Test (push) Failing after 6s
Tests / OCR (tesseract fixture) (push) Failing after 4s
Tests / Release (semver) (push) Skipped
OQ3: internal/duckstats + bin/qa/stats.go. Zig stays Ladybug-only.
mikefarah/yq for small structured slices. Gitea #30.
2026-08-14 11:32:23 +01:00
eSliderandGitHub f99dfea104 feat: OCR scans with tesseract, drop docling from the default path. (#33)
Tests / Test (push) Failing after 5s
Tests / OCR (tesseract fixture) (push) Failing after 4s
Tests / Release (semver) (push) Skipped
pdftotext still wins on born-digital PDFs. Empty text layers go through
pdftoppm + tesseract eng+deu. Optional OCR_ENGINE=paddle. Gitea #6.
2026-08-14 11:26:04 +01:00
eSliderandGitHub bae1494258 ci: recall@5 SoT is Zig bin/brain/eval.go (#19) (#32)
Tests / Test (push) Failing after 6s
Tests / Release (semver) (push) Skipped
* ci: recall@5 SoT is Zig bin/brain/eval.go.

GitHub Actions rebuilds the repo corpus and runs the Go eval gate
instead of Python bin/kb/eval. Audit self is a hard fail (Gitea #19).

* ci: point eval at the checkout and keep DevOps in the corpus.

/tmp/brain-eval cannot walk up to var/kb.lbug; KB_ROOT is required.
Recall fragments must exist in README/docs so a repo-only rebuild gates.
2026-08-14 11:03:39 +01:00
75 changed files with 2798 additions and 2220 deletions
+20 -4
View File
@@ -45,10 +45,10 @@ jobs:
run: | run: |
uv run python -m unittest discover -s bin/tools -t . uv run python -m unittest discover -s bin/tools -t .
- name: Go tests (root module, no ladybug cgo) - name: Go tests (root module; duckdb-go CGO via gcc, no ladybug)
run: | run: |
go vet ./... CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= go vet ./...
go test ./... -count=1 CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= go test ./... -count=1
- name: brain ranking tests (no cgo / no ladybug) - name: brain ranking tests (no cgo / no ladybug)
run: go test ./internal/brain/rank -count=1 run: go test ./internal/brain/rank -count=1
@@ -72,10 +72,26 @@ jobs:
uv run python bin/kb/index --rebuild --json uv run python bin/kb/index --rebuild --json
KB_ROOT="$PWD" /tmp/brain-eval --json KB_ROOT="$PWD" /tmp/brain-eval --json
ocr:
name: OCR (tesseract fixture)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
- name: Install tesseract + poppler
run: |
sudo apt-get update
sudo apt-get install -y --no-install-recommends \
tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu poppler-utils
- name: Go OCR tests (synthetic HELLO PNG)
run: go test ./internal/ocr -count=1
release: release:
name: Release (semver) name: Release (semver)
if: github.event_name == 'push' && github.ref == 'refs/heads/main' if: github.event_name == 'push' && github.ref == 'refs/heads/main'
needs: test needs: [test, ocr]
runs-on: ubuntu-latest runs-on: ubuntu-latest
permissions: permissions:
contents: write contents: write
+2
View File
@@ -12,3 +12,5 @@ __pycache__/
lib-ladybug/ lib-ladybug/
go.work.local go.work.local
models/ models/
# Purged from git history. Do not re-add.
docs/crm-associations-proof.md
+13 -7
View File
@@ -42,13 +42,14 @@ skills/ in-project agent skills (vendored, no external links)
bin/ self-describing tools bin/{subject}/{method}.go (shebang) bin/ self-describing tools bin/{subject}/{method}.go (shebang)
bin/brain/ search.go serve.go index.go add.go get.go stats.go eval.go watch.go bin/brain/ search.go serve.go index.go add.go get.go stats.go eval.go watch.go
bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats
bin/mail/ sync.go import.go (index_mail → brain/index.go) bin/mail/ sync.go import.go ocr.go (index_mail → brain/index.go)
bin/markdown/ import.go (H2 leaf split; Python bin/md/import fallback) bin/markdown/ import.go (H2 leaf split; Python bin/md/import fallback)
bin/postgres/ query.go (read-only YAML) bin/postgres/ query.go (read-only YAML)
bin/git/ import.go (go-git history; Python shim execs it) bin/git/ import.go (go-git history; Python shim execs it)
bin/web/ search.go (SearXNG; Python shim execs it) bin/web/ search.go (SearXNG; Python shim execs it)
bin/reasoner/ bakeoff.go (D18 CPU OpenAI tool-call bake-off) bin/reasoner/ bakeoff.go (D18 CPU OpenAI tool-call bake-off)
internal/ shared Go (brain/rank is cgo-free; chats parsers; gitlog; websearch; reasoner) internal/ shared Go (brain/rank is cgo-free; facts D16; cli flaggy D23; chats; gitlog; websearch; reasoner; duckstats)
bin/qa/ stats.go (DuckDB quantiles / JSONL count; gcc CGO, not Zig)
bin/watch/ corpus watcher (used by bin/brain/watch.go) bin/watch/ corpus watcher (used by bin/brain/watch.go)
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch) bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
bin/cgo/ zig zcc zc++ (CGO via zig cc, not gcc) bin/cgo/ zig zcc zc++ (CGO via zig cc, not gcc)
@@ -71,9 +72,9 @@ bin/brain/index.go --rebuild --with-facts --with-chats
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list + - `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
`body.attachmentId` (not partId) for attachments. `body.attachmentId` (not partId) for attachments.
- `import` converts body + attachments to markdown. PDFs use poppler - `import` converts body + attachments to markdown. PDFs use poppler
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to `pdftotext -layout` fast path (~15ms); textless/scanned PDFs use
docling (isolated subprocess — its native onnx can segfault the parent). `pdftoppm` + tesseract `eng+deu` (`bin/mail/ocr.go`). Optional
Conversion never touches the brain DB (crash safety). `OCR_ENGINE=paddle`. Conversion never touches the brain DB (crash safety).
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Bulk - `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Bulk
rebuild still deletes `var/kb.lbug` and creates FTS/HNSW last. Single-leaf rebuild still deletes `var/kb.lbug` and creates FTS/HNSW last. Single-leaf
write is `bin/brain/add.go` (safe while indexes exist; do not DROP INDEX). write is `bin/brain/add.go` (safe while indexes exist; do not DROP INDEX).
@@ -83,11 +84,12 @@ bin/brain/index.go --rebuild --with-facts --with-chats
## Tools ## Tools
```bash ```bash
bin/facts/audit.go ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate bin/facts/audit.go ["self"|"db"|"contradict"] # 2-source + D16 adjudication
bin/facts/crm.go [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT) bin/facts/crm.go [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
bin/brain/search.go "query" --no-web # local graph only bin/brain/search.go "query" --no-web # local graph only
source <(./bin/cli/complete.go bash) # flaggy completions (D23)
eval "$(bin/cgo/zig env)" # Zig cc + liblbug (not gcc) eval "$(bin/cgo/zig env)" # Zig cc + liblbug (not gcc)
bin/brain/index.go --rebuild [--with-mail] [--with-facts] [--with-chats] bin/brain/index.go --rebuild [--with-mail] [--with-facts] [--with-chats]
bin/brain/add.go --text T --root facts --source "a.md x b.md" # incremental write bin/brain/add.go --text T --root facts --source "a.md x b.md" # incremental write
@@ -101,12 +103,16 @@ bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit le
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
bin/reasoner/bakeoff.go [--model ID] [--json] # D18 CPU tool-call bake-off bin/reasoner/bakeoff.go [--model ID] [--json] # D18 CPU tool-call bake-off
bin/postgres/query.go --profile onlyoffice -c 'SELECT 1' bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
bin/qa/stats.go # D22 DuckDB quantiles / JSONL (gcc CGO)
bin/mail/ocr.go <image|pdf> # tesseract eng+deu (scans)
bin/md/tables # what the graph holds → YAML bin/md/tables # what the graph holds → YAML
bin/brain/deduce "question" # thinking wrapper bin/brain/deduce "question" # thinking wrapper
``` ```
Never start a shell command with `cd` — use the tool working-directory Never start a shell command with `cd` — use the tool working-directory
parameter. Search before reading whole files. parameter. Search before reading whole files. For YAML/JSON/XML/CSV/TOML/HCL
prefer mikefarah/yq (`skills/yq/SKILL.md`). For bulk rows and quantiles use
duckdb-go (`internal/duckstats`, `skills/duckdb/SKILL.md`), not Ladybug.
## GitHub safety rules (ABSOLUTE — never violate) ## GitHub safety rules (ABSOLUTE — never violate)
+4
View File
@@ -16,6 +16,10 @@ ENV PYTHONUNBUFFERED=1 \
WORKDIR /app WORKDIR /app
RUN id -u 2dph 2>/dev/null || useradd --create-home --uid 1001 2dph RUN id -u 2dph 2>/dev/null || useradd --create-home --uid 1001 2dph
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
poppler-utils tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu \
&& rm -rf /var/lib/apt/lists/*
COPY requirements.lock.txt /tmp/requirements.lock.txt COPY requirements.lock.txt /tmp/requirements.lock.txt
RUN python -m pip install --no-cache-dir -r /tmp/requirements.lock.txt \ RUN python -m pip install --no-cache-dir -r /tmp/requirements.lock.txt \
+24 -12
View File
@@ -4,8 +4,12 @@ A brain that loves facts and deduction. Evidence-first knowledge graph + hybrid
RAG over the operational Brain/ops/eSlider stack. Built like Sherlock RAG over the operational Brain/ops/eSlider stack. Built like Sherlock
Holmes: nothing is asserted unless it has proof. Holmes: nothing is asserted unless it has proof.
Status: **in progress** — read path + MCP work; v1 goal is [epic #16](https://git.produktor.io/eSlider/2dph/issues/16) Status: **v1 in** (epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed).
(milestone [v1 detective brain](https://git.produktor.io/eSlider/2dph/milestone/12)). v2 board: milestone [v2](https://git.produktor.io/eSlider/2dph/milestone/13)
OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6) in,
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1 in,
[#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3 in,
[#34](https://git.produktor.io/eSlider/2dph/issues/34) D23 in.
Gap: [docs/roadmap.md](docs/roadmap.md). Gap: [docs/roadmap.md](docs/roadmap.md).
## What ## What
@@ -41,12 +45,14 @@ detective method: **a fact needs ≥2 independent sources or it is
| D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. | | D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. |
| D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. | | D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. |
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. | | D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. | | D16 | contradictions | ≥2 yes vs ≥2 no → hypothesis → `(not confirmed)` until a rule fires. Order: **temporal_freshness** (fresh ≥2 vs stale minority), then **authority_pairing** (runtime/config A×B beats narrative C). Store as `a x b vs c x d` on hypothesis leafs. `bin/facts/audit contradict`. [#29](https://git.produktor.io/eSlider/2dph/issues/29). |
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. | | D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
| D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is compose profile `picoclaw`; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. Agent lever/loop: [#15](https://git.produktor.io/eSlider/2dph/issues/15). | | D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is compose profile `picoclaw`; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. Agent lever/loop: [#15](https://git.produktor.io/eSlider/2dph/issues/15). |
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. | | D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
| D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`/`ingest`). | | D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`/`ingest`). |
| D21 | CGO | Ladybug/tokenizers CGO is compiled with **Zig** (`bin/cgo/zcc``zig cc -target …-linux-gnu`), not gcc. `bin/cgo/zig` pins Zig 0.14.1 + liblbug 0.19.1 + libtokenizers 1.27.0. Compose `target: api` has no CPython; write/rebuild is profile `index`. | | D21 | CGO | Ladybug/tokenizers CGO is compiled with **Zig** (`bin/cgo/zcc``zig cc -target …-linux-gnu`), not gcc. `bin/cgo/zig` pins Zig 0.14.1 + liblbug 0.19.1 + libtokenizers 1.27.0. Compose `target: api` has no CPython; write/rebuild is profile `index`. |
| D22 | analytics | **duckdb-go** in-process (`internal/duckstats`, `bin/qa/stats.go`) for quantiles/JSONL. Links with **gcc/g++**, not Zig. Ladybug stays the graph; web-search cache stays modernc sqlite. Slice small structured docs with **mikefarah/yq**, not kislyuk/jq. [#30](https://git.produktor.io/eSlider/2dph/issues/30). |
| D23 | CLI | **flaggy** (`github.com/integrii/flaggy`, 0 deps). Flags at any position. Wrapper `internal/cli`. Bash complete: `source <(./bin/cli/complete.go bash)`. No cobra, no stdlib `flag` in Go tools. Search does not intercept the word `completion`. [#34](https://git.produktor.io/eSlider/2dph/issues/34). |
## Architecture ## Architecture
@@ -63,7 +69,8 @@ detective method: **a fact needs ≥2 independent sources or it is
brain/add.go incremental leaf write (no rebuild) brain/add.go incremental leaf write (no rebuild)
brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback
brain/watch.go brain/watch.go
brain/search.go deduction: facts → info → web-search brain/search.go deduction: facts → info → web
cli/complete.go flaggy bash/zsh/fish complete (D23)
brain/serve.go HTTP API in-process + OpenAPI/MCP (D20); Zig CGO (D21) brain/serve.go HTTP API in-process + OpenAPI/MCP (D20); Zig CGO (D21)
cgo/zig zcc zc++ CGO toolchain (zig cc, not gcc) cgo/zig zcc zc++ CGO toolchain (zig cc, not gcc)
mail/import.go JSON → markdown (no brain write) mail/import.go JSON → markdown (no brain write)
@@ -74,6 +81,7 @@ detective method: **a fact needs ≥2 independent sources or it is
reasoner/bakeoff.go CPU tool-call bake-off (D18; OpenAI tools) reasoner/bakeoff.go CPU tool-call bake-off (D18; OpenAI tools)
chats/sync.go import.go facts.go apply.go chats/sync.go import.go facts.go apply.go
(libs in internal/chats; no chats index) (libs in internal/chats; no chats index)
mail/ocr.go tesseract eng+deu (pdftoppm scans)
md/import (deprecated; bin/markdown/import.go) md/import (deprecated; bin/markdown/import.go)
brain/extract brain/audit brain/deduce (thinking wrapper) brain/extract brain/audit brain/deduce (thinking wrapper)
web/search (deprecated shim → web/search.go) web/search (deprecated shim → web/search.go)
@@ -108,17 +116,20 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
- `bin/{subject}/{method}` — line 2 is a usage comment (mirrors `psql-yq`). - `bin/{subject}/{method}` — line 2 is a usage comment (mirrors `psql-yq`).
- bash + python primary; golang via Go shebang when a compiled helper is right. - bash + python primary; golang via Go shebang when a compiled helper is right.
- YAML default output, `--json` for machines. Slice with `yq`. - YAML default output, `--json` for machines. Slice with mikefarah/yq.
- Everything that touches the network / DB is read-only, throttled, cached. - Everything that touches the network / DB is read-only, throttled, cached.
- Tests (TDD) gate every commit; `gh` + CI/CD on every push. - Tests (TDD) gate every commit; `gh` + CI/CD on every push.
## Open questions (v2) ## Open questions (v2)
- OQ1: mutually-contradicting evidence — how to resolve (authority weighting, - OQ1: **in** — D16 adjudication: `temporal_freshness` then `authority_pairing`.
temporal freshness, audit adjudication). **v2**; does not block epic #16. Unresolved 2v2 stays hypothesis. [#29](https://git.produktor.io/eSlider/2dph/issues/29).
- OQ2: OCR — poppler `pdftotext` fast-path exists; scans still docling. - OQ2: OCR — **in**. `pdftotext -layout` first; scans `pdftoppm` + tesseract
[#6](https://git.produktor.io/eSlider/2dph/issues/6) (v2, does not block #16). `eng+deu` (`bin/mail/ocr.go`, `internal/ocr`). No gocv, no gosseract CGO
- OQ3: optional duckdb-md layer for `SELECT … FORMAT MARKDOWN` export/write-back. (D21 Zig owns Ladybug CGO). Optional `OCR_ENGINE=paddle` / compose profile
`ocr-paddle`. Docling left the default path. [#6](https://git.produktor.io/eSlider/2dph/issues/6).
- OQ3: **in** — duckdb-go (`internal/duckstats`, `bin/qa/stats.go`) for
quantiles / JSONL count. Not a second graph. [#30](https://git.produktor.io/eSlider/2dph/issues/30).
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to - OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
serialize and unambiguous; YAML only where humans edit files. serialize and unambiguous; YAML only where humans edit files.
@@ -127,7 +138,8 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download. 1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
Gmail attachments key off `body.attachmentId`, not MIME `partId`. Gmail attachments key off `body.attachmentId`, not MIME `partId`.
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via 2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
`pdftotext -layout` (~15ms) with docling subprocess fallback; ICS sidecars `pdftotext -layout` (~15ms); textless/scanned PDFs `pdftoppm` + tesseract
`eng+deu`. ICS sidecars
Latin-1→UTF-8 normalized. Latin-1→UTF-8 normalized.
3. `bin/brain/index.go --rebuild` — fresh rebuild (repo corpus + mail) because ladybug 3. `bin/brain/index.go --rebuild` — fresh rebuild (repo corpus + mail) because ladybug
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
@@ -177,4 +189,4 @@ Narrative: [docs/roadmap.md](docs/roadmap.md).
| 4 | [#15](https://git.produktor.io/eSlider/2dph/issues/15) | **in** — lever/loop documented (`search``get``audit`). | | 4 | [#15](https://git.produktor.io/eSlider/2dph/issues/15) | **in** — lever/loop documented (`search``get``audit`). |
| 5 | [#19](https://git.produktor.io/eSlider/2dph/issues/19) | **in** — CI recall SoT is `bin/brain/eval.go` via Zig. Python `bin/kb/eval` stays as an explicit fallback. | | 5 | [#19](https://git.produktor.io/eSlider/2dph/issues/19) | **in** — CI recall SoT is `bin/brain/eval.go` via Zig. Python `bin/kb/eval` stays as an explicit fallback. |
Does **not** block epic close: [#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR, OQ1, OQ3, OQ4. Does **not** block epic close: OQ4. OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6), OQ1 [#29](https://git.produktor.io/eSlider/2dph/issues/29), OQ3 [#30](https://git.produktor.io/eSlider/2dph/issues/30) are **in**.
+91
View File
@@ -0,0 +1,91 @@
//usr/bin/env go run "$0" "$@"; exit
//
// bin/cli/complete.go - dump flaggy shell completions for all Go shebang tools (D23).
//
// source <(./bin/cli/complete.go bash)
// ./bin/cli/complete.go zsh|fish|powershell|nushell
//
// Search does not steal the word "completion"; this binary dumps scripts.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"strings"
mailsync "github.com/eSlider/2dph/bin/mail/sync"
"github.com/eSlider/2dph/internal/brain/rank"
"github.com/eSlider/2dph/internal/chats"
"github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/gitlog"
"github.com/eSlider/2dph/internal/mdleaves"
"github.com/eSlider/2dph/internal/ocr"
"github.com/eSlider/2dph/internal/reasoner"
"github.com/eSlider/2dph/internal/websearch"
"github.com/integrii/flaggy"
)
func tools() []cli.Tool {
return []cli.Tool{
{Path: "bin/brain/search.go", Name: "brain-search", New: rank.Parser},
{Path: "bin/brain/get.go", Name: "brain-get", New: func() *flaggy.Parser {
o := rank.GetOptions{}
return rank.GetParser(&o)
}},
{Path: "bin/brain/stats.go", Name: "brain-stats", New: rank.StatsParser},
{Path: "bin/brain/eval.go", Name: "brain-eval", New: rank.EvalParser},
{Path: "bin/web/search.go", Name: "web-search", New: websearch.Parser},
{Path: "bin/git/import.go", Name: "git-import", New: gitlog.Parser},
{Path: "bin/markdown/import.go", Name: "markdown-import", New: mdleaves.Parser},
{Path: "bin/qa/stats.go", Name: "qa-stats", New: cli.QAParser},
{Path: "bin/reasoner/bakeoff.go", Name: "reasoner-bakeoff", New: reasoner.Parser},
{Path: "bin/mail/ocr.go", Name: "mail-ocr", New: ocr.Parser},
{Path: "bin/mail/sync.go", Name: "mail-sync", New: mailsync.Parser},
{Path: "bin/chats/sync.go", Name: "chats-sync", New: chats.SyncParser},
{Path: "bin/chats/import.go", Name: "chats-import", New: chats.ImportParser},
{Path: "bin/chats/facts.go", Name: "chats-facts", New: chats.FactsParser},
{Path: "bin/chats/apply.go", Name: "chats-apply", New: chats.ApplyParser},
}
}
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
shell := "bash"
if len(args) > 0 {
switch args[0] {
case "bash", "zsh", "fish", "powershell", "nushell":
shell = args[0]
case "-h", "--help", "help":
fmt.Fprintln(os.Stderr, "usage: bin/cli/complete.go [bash|zsh|fish|powershell|nushell]")
return 0
default:
fmt.Fprintf(os.Stderr, "cli/complete: unknown shell %q\n", args[0])
return 2
}
}
if shell == "bash" {
fmt.Print(cli.BashScript(tools()))
return 0
}
var b strings.Builder
for _, t := range tools() {
p := t.New()
p.Name = t.Name
switch shell {
case "zsh":
b.WriteString(flaggy.GenerateZshCompletion(p))
case "fish":
b.WriteString(flaggy.GenerateFishCompletion(p))
case "powershell":
b.WriteString(flaggy.GeneratePowerShellCompletion(p))
case "nushell":
b.WriteString(flaggy.GenerateNushellCompletion(p))
}
}
fmt.Print(b.String())
return 0
}
+45 -18
View File
@@ -1,14 +1,15 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""facts/audit - evidence & lexicon checks for the 2dph brain. """facts/audit - evidence & lexicon checks for the 2dph brain.
bin/facts/audit self # lexicon: every fact in db has >=2 sources bin/facts/audit self # lexicon: docs + two-source rule
bin/facts/audit db # evidence gate: run against var/kb.lbug bin/facts/audit db # evidence gate against var/kb.lbug
bin/facts/audit contradict # D16 adjudication (JSON claim(s) on stdin)
`self` mode checks the repo itself (no network, no runtime deps). It greps `self` mode checks the repo itself (no network, no runtime deps).
for known-good two-source pairings and confirms the docs are consistent. `db` mode loads every Leaf with root=facts. Confirmed facts need ` x `;
`db` mode loads every Leaf with root=facts and asserts each has source_rev hypothesis contradictions need `a x b vs c x d` (both sides ≥2).
and a non-empty `loc` (the "where did you see it" evidence pointer) and that `contradict` applies temporal_freshness then authority_pairing; ≥2 vs ≥2
'confirmed' facts carry a two-source `source` field. with no rule stays hypothesis / `(not confirmed)`.
Exit 0 = all checks pass, 1 = audit failures, 2 = could not evaluate. Exit 0 = all checks pass, 1 = audit failures, 2 = could not evaluate.
""" """
@@ -22,6 +23,8 @@ from pathlib import Path
ROOT = Path(__file__).resolve().parents[2] ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools")) sys.path.insert(0, str(ROOT / "bin" / "tools"))
from contradict import adjudicate, check_fact_row # noqa: E402
def audit_db() -> list[str]: def audit_db() -> list[str]:
from kblib import connect from kblib import connect
@@ -33,14 +36,8 @@ def audit_db() -> list[str]:
r = conn.execute("MATCH (l:Leaf {root:'facts'}) RETURN l.id, l.source, l.loc, l.how, l.confidence") r = conn.execute("MATCH (l:Leaf {root:'facts'}) RETURN l.id, l.source, l.loc, l.how, l.confidence")
problems: list[str] = [] problems: list[str] = []
for lid, source, loc, how, conf in r.get_all(): for lid, source, loc, how, conf in r.get_all():
if conf != "confirmed": problems.extend(check_fact_row(str(lid), str(source or ""), str(loc or ""),
problems.append(f"{lid}: facts require confidence='confirmed', got '{conf}'") str(how or ""), str(conf or "")))
if not source or " x " not in source:
problems.append(f"{lid}: needs 2-source evidence in source, got '{source}'")
if not loc:
problems.append(f"{lid}: missing loc (evidence pointer)")
if not how:
problems.append(f"{lid}: missing how")
conn.close() conn.close()
db.close() db.close()
return problems return problems
@@ -54,20 +51,50 @@ def audit_self() -> list[str]:
problems.append("PLAN.md missing recall@5 gate") problems.append("PLAN.md missing recall@5 gate")
if re.search(r"(?i)facts must have.*2 sources|2.source", plan) is None: if re.search(r"(?i)facts must have.*2 sources|2.source", plan) is None:
problems.append("PLAN.md missing the two-source evidence rule for facts") problems.append("PLAN.md missing the two-source evidence rule for facts")
if "temporal_freshness" not in plan or "authority_pairing" not in plan:
problems.append("PLAN.md missing D16 adjudication rules")
if re.search(r"(?i)HNSW|BM25|deduction", (ROOT / "README.md").read_text()) is None: if re.search(r"(?i)HNSW|BM25|deduction", (ROOT / "README.md").read_text()) is None:
problems.append("README.md missing search/retrieval description") problems.append("README.md missing search/retrieval description")
return problems return problems
def audit_contradict(raw: str) -> tuple[list[str], list[dict]]:
raw = raw.strip()
if not raw:
return ["contradict: empty stdin (JSON claim or {claims:[...]})"], []
try:
payload = json.loads(raw)
except json.JSONDecodeError as e:
return [f"contradict: invalid JSON: {e}"], []
if isinstance(payload, dict) and "claims" in payload:
claims = list(payload.get("claims") or [])
elif isinstance(payload, dict):
claims = [payload]
elif isinstance(payload, list):
claims = payload
else:
return ["contradict: expected object or list"], []
details = [adjudicate(c) for c in claims]
return [], details
def main(argv: list[str]) -> int: def main(argv: list[str]) -> int:
import argparse import argparse
p = argparse.ArgumentParser(description="evidence & lexicon audit") p = argparse.ArgumentParser(description="evidence & lexicon audit")
p.add_argument("mode", choices=("self", "db")) p.add_argument("mode", choices=("self", "db", "contradict"))
p.add_argument("--json", action="store_true") p.add_argument("--json", action="store_true")
a = p.parse_args(argv) a = p.parse_args(argv)
problems = audit_self() if a.mode == "self" else audit_db() details: list[dict] = []
out = {"mode": a.mode, "ok": not problems, "problems": problems} if a.mode == "self":
problems = audit_self()
elif a.mode == "db":
problems = audit_db()
else:
problems, details = audit_contradict(sys.stdin.read())
out: dict = {"mode": a.mode, "ok": not problems, "problems": problems}
if details:
out["contradictions"] = details
if a.json: if a.json:
print(json.dumps(out, indent=2)) print(json.dumps(out, indent=2))
else: else:
+1
View File
@@ -5,6 +5,7 @@
// //
// ./bin/facts/audit.go self // ./bin/facts/audit.go self
// ./bin/facts/audit.go db // ./bin/facts/audit.go db
// ./bin/facts/audit.go contradict --json < claim.json
// //
// Python bin/facts/audit is the implementation (CI runs it directly). // Python bin/facts/audit is the implementation (CI runs it directly).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang. // NOTE: never run `gofmt -w` on this file — it breaks the shebang.
+6 -49
View File
@@ -16,9 +16,8 @@ import (
"fmt" "fmt"
"os" "os"
"path/filepath" "path/filepath"
"strconv"
"time"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/cmdbin" "github.com/eSlider/2dph/internal/cmdbin"
"github.com/eSlider/2dph/internal/gitlog" "github.com/eSlider/2dph/internal/gitlog"
) )
@@ -28,50 +27,17 @@ func main() {
} }
func run(args []string) int { func run(args []string) int {
var repo, root, since string c, err := gitlog.ParseArgs(args)
limit := 0 if err != nil {
jsonOut := false return cliparse.Fail(err)
i := 0
for i < len(args) {
a := args[i]
switch {
case a == "--json":
jsonOut = true
case a == "--limit" && i+1 < len(args):
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 0 {
fmt.Fprintf(os.Stderr, "git/import: --limit must be a non-negative integer\n")
return 2
}
limit = n
case a == "--since" && i+1 < len(args):
i++
since = args[i]
case a == "--root" && i+1 < len(args):
i++
root = args[i]
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, `usage: bin/git/import.go [REPO] [--json] [--limit N] [--since DATE] [--root DIR]`)
return 0
case len(a) > 0 && a[0] != '-':
repo = a
default:
fmt.Fprintf(os.Stderr, "git/import: unknown flag %s\n", a)
return 2
}
i++
} }
repo, root, since, limit, jsonOut := c.Repo, c.Root, c.Since, c.Limit, c.JSONOut
var sinceT time.Time sinceT, err := gitlog.ParseSince(since)
if since != "" {
var err error
sinceT, err = parseSince(since)
if err != nil { if err != nil {
fmt.Fprintf(os.Stderr, "git/import: %v\n", err) fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
return 2 return 2
} }
}
repos := []string{} repos := []string{}
if repo != "" { if repo != "" {
@@ -132,12 +98,3 @@ func run(args []string) int {
} }
return 0 return 0
} }
func parseSince(s string) (time.Time, error) {
for _, layout := range []string{time.RFC3339, "2006-01-02"} {
if t, err := time.Parse(layout, s); err == nil {
return t, nil
}
}
return time.Time{}, fmt.Errorf("cannot parse --since %q", s)
}
+11 -77
View File
@@ -7,7 +7,7 @@
bin/mail/import --since 2026-01-01 only messages after a date bin/mail/import --since 2026-01-01 only messages after a date
bin/mail/import --limit 50 cap messages per run bin/mail/import --limit 50 cap messages per run
bin/mail/import --no-attachments body only, skip attachment conversion bin/mail/import --no-attachments body only, skip attachment conversion
bin/mail/import --ocr OCR scanned PDFs/images via docling bin/mail/import --ocr OCR images (PDFs OCR when textless)
bin/mail/import --dry-run list messages without writing anything bin/mail/import --dry-run list messages without writing anything
Writes one directory per message: var/mail/{folder}/{message_id}/ Writes one directory per message: var/mail/{folder}/{message_id}/
@@ -16,10 +16,11 @@ Writes one directory per message: var/mail/{folder}/{message_id}/
attachments/*.md converted attachment content attachments/*.md converted attachment content
Indexing is a separate step (`bin/brain/index.go --rebuild`): conversion can Indexing is a separate step (`bin/brain/index.go --rebuild`): conversion can
crash in native docling and must not leave the brain DB mid-transaction. crash and must not leave the brain DB mid-transaction.
Requires ONLYOFFICE_URL/USER/PASS in .env (or env). Idempotent: a message Requires ONLYOFFICE_URL/USER/PASS in .env (or env) except `--from-raw`.
already present (message.md exists) is skipped unless --force. Idempotent: a message already present (message.md exists) is skipped unless
--force.
""" """
from __future__ import annotations from __future__ import annotations
@@ -41,10 +42,10 @@ from mailconv import ( # noqa: E402
IMAGE_SUFFIXES, IMAGE_SUFFIXES,
LEGACY_OFFICE_SUFFIXES, LEGACY_OFFICE_SUFFIXES,
TEXT_SUFFIXES, TEXT_SUFFIXES,
convert_pdf,
html_to_markdown, html_to_markdown,
is_convertible,
normalize_markdown, normalize_markdown,
subject_to_filename, ocr_image,
zip_extract_safe, zip_extract_safe,
) )
@@ -147,9 +148,9 @@ def convert_file_to_md(path: Path, ocr: bool) -> str | None:
except Exception as e: except Exception as e:
return f"\n<!-- conversion failed: {e} -->\n" return f"\n<!-- conversion failed: {e} -->\n"
if suffix == ".pdf": if suffix == ".pdf":
return _convert_pdf(path, ocr) return convert_pdf(path, ocr)
if suffix in IMAGE_SUFFIXES and ocr: if suffix in IMAGE_SUFFIXES and ocr:
return _convert_pdf(path, ocr) return ocr_image(path) or "\n<!-- ocr unavailable -->\n"
if suffix in LEGACY_OFFICE_SUFFIXES: if suffix in LEGACY_OFFICE_SUFFIXES:
return _convert_legacy(path) return _convert_legacy(path)
if suffix in ARCHIVE_SUFFIXES: if suffix in ARCHIVE_SUFFIXES:
@@ -157,67 +158,6 @@ def convert_file_to_md(path: Path, ocr: bool) -> str | None:
return None return None
def _convert_pdf(path: Path, ocr: bool) -> str:
"""Convert one PDF to markdown.
Fast path: poppler's pdftotext (-layout) extracts exact text from
born-digital PDFs in ~15ms vs docling's 1-3s. Only textless PDFs (scanned
pages, layout-heavy) fall back to docling, which runs isolated in a
subprocess because its native onnx/RT-DETR has segfaulted the main process.
"""
text = _pdf_fast_text(path)
if ocr or text is None or not text.strip():
return _convert_pdf_docling(path, ocr)
return normalize_markdown(text)
def _pdf_fast_text(path: Path) -> str | None:
"""pdftotext -layout; None when poppler is unavailable (or the PDF has no text layer)."""
try:
proc = subprocess.run(
["pdftotext", "-layout", str(path), "-"],
capture_output=True, timeout=60)
except (OSError, subprocess.TimeoutExpired):
return None
if proc.returncode != 0:
return None
return proc.stdout.decode("utf-8", errors="replace")
def _convert_pdf_docling(path: Path, ocr: bool) -> str:
try:
proc = subprocess.run(
[sys.executable, os.path.abspath(__file__), "--pdf-worker", str(path),
"--ocr" if ocr else "--no-ocr"],
capture_output=True, text=True, timeout=600)
except subprocess.TimeoutExpired:
return "\n<!-- pdf conversion timed out -->\n"
if proc.returncode != 0:
tail = proc.stderr.strip().splitlines()[-3:]
return f"\n<!-- pdf conversion failed: {proc.returncode}: {' | '.join(tail)} -->\n"
return proc.stdout
def _pdf_worker(path: Path, ocr: bool) -> None:
"""docling worker entry: prints converted markdown on stdout, exits non-zero on error."""
try:
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.pipeline_options import PdfPipelineOptions
opts = PdfPipelineOptions()
opts.do_ocr = bool(ocr)
opts.do_table_structure = True
conv = DocumentConverter(format_options={"pdf": PdfFormatOption(pipeline_options=opts)})
res = conv.convert(str(path))
sys.stdout.write(normalize_markdown(res.document.export_to_markdown()))
sys.exit(0)
except Exception as e:
# errors/stacktraces to stderr; the caller only reports a one-liner
print(f"pdf-worker: {e}", file=sys.stderr)
import traceback
traceback.print_exc(file=sys.stderr)
sys.exit(1)
def _convert_legacy(path: Path) -> str: def _convert_legacy(path: Path) -> str:
"""Legacy .doc/.xls/.ppt -> md via pandoc (installed) or a stub.""" """Legacy .doc/.xls/.ppt -> md via pandoc (installed) or a stub."""
try: try:
@@ -356,19 +296,12 @@ def main(argv: list[str]) -> int:
p.add_argument("--from-raw", default="", p.add_argument("--from-raw", default="",
help="convert Go-synced dirs (var/mail/<folder>/<id>/message.json) to markdown") help="convert Go-synced dirs (var/mail/<folder>/<id>/message.json) to markdown")
p.add_argument("--no-attachments", action="store_true", help="skip attachment download+convert") p.add_argument("--no-attachments", action="store_true", help="skip attachment download+convert")
p.add_argument("--ocr", action="store_true", help="OCR scanned PDFs/images via docling") p.add_argument("--ocr", action="store_true", help="OCR images (PDFs OCR when textless)")
p.add_argument("--force", action="store_true", help="re-import even if message.md exists") p.add_argument("--force", action="store_true", help="re-import even if message.md exists")
p.add_argument("--dry-run", action="store_true", help="list messages, write nothing") p.add_argument("--dry-run", action="store_true", help="list messages, write nothing")
p.add_argument("--json", action="store_true") p.add_argument("--json", action="store_true")
p.add_argument("--pdf-worker", default="", help=argparse.SUPPRESS)
p.add_argument("--no-ocr", action="store_true", help=argparse.SUPPRESS)
a = p.parse_args(argv) a = p.parse_args(argv)
if a.pdf_worker:
_pdf_worker(Path(a.pdf_worker), ocr=not a.no_ocr)
return 0
conf = load_env()
fid = folder_id(a.folder) fid = folder_id(a.folder)
out_root = ROOT / "var" / "mail" out_root = ROOT / "var" / "mail"
summary: list[dict] = [] summary: list[dict] = []
@@ -394,6 +327,7 @@ def main(argv: list[str]) -> int:
target_dir=msg_dir.parent)) target_dir=msg_dir.parent))
summary.append(entry) summary.append(entry)
else: else:
conf = load_env()
OOCLIENT = OOClient(conf) OOCLIENT = OOClient(conf)
if a.id: if a.id:
messages = [{"id": i} for i in a.id] messages = [{"id": i} for i in a.id]
+46
View File
@@ -0,0 +1,46 @@
//usr/bin/env go run -tags=mail_ocr "$0" "$@"; exit
//go:build mail_ocr
//
// bin/mail/ocr.go - OCR an image or scanned PDF (tesseract eng+deu).
//
// ./bin/mail/ocr.go scan.png
// ./bin/mail/ocr.go scan.pdf
// OCR_ENGINE=paddle ./bin/mail/ocr.go scan.png
//
// PDFs try pdftotext -layout first; empty text layer uses pdftoppm + tesseract.
// No gocv. Tesseract CGO bindings are not used (D21 Zig owns Ladybug CGO).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"strings"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/ocr"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
c, err := ocr.ParseArgs(args)
if err != nil {
return cliparse.Fail(err)
}
path := c.Path
var text string
if strings.HasSuffix(strings.ToLower(path), ".pdf") {
text, err = ocr.PDFFile(path)
} else {
text, err = ocr.ImageFile(path)
}
if err != nil {
fmt.Fprintf(os.Stderr, "mail/ocr: %v\n", err)
return 1
}
fmt.Println(text)
return 0
}
+68 -40
View File
@@ -5,12 +5,15 @@ package sync
import ( import (
"context" "context"
"flag" "errors"
"fmt" "fmt"
"os" "os"
"path/filepath" "path/filepath"
"strings" "strings"
"time" "time"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
) )
// CLIConfig is a superset of SyncConfig plus flag parsing results. // CLIConfig is a superset of SyncConfig plus flag parsing results.
@@ -21,58 +24,84 @@ type CLIConfig struct {
Help bool Help bool
} }
// ParseCLI reads os.Args into a CLIConfig. Exit codes: 0 ok, 2 usage. type flagVals struct {
env, out, srcs, query string
workers, limit, offset int
force, dryRun bool
}
func Parser() *flaggy.Parser {
v := flagVals{workers: 4, query: "in:inbox", srcs: "onlyoffice"}
return bind(&v)
}
func bind(v *flagVals) *flaggy.Parser {
if v.workers == 0 {
v.workers = 4
}
if v.query == "" {
v.query = "in:inbox"
}
if v.srcs == "" {
v.srcs = "onlyoffice"
}
p := cliparse.New("mail-sync")
p.Description = "download mail to var/mail"
p.String(&v.env, "", "env", ".env file")
p.String(&v.out, "", "out", "var/mail root")
p.Int(&v.workers, "", "workers", "concurrent downloads")
p.Int(&v.limit, "", "limit", "max messages per source (0 = all)")
p.Int(&v.offset, "", "offset", "skip first N messages per source")
p.Bool(&v.force, "", "force", "overwrite existing message.json")
p.Bool(&v.dryRun, "", "dry-run", "list counts without writing")
p.String(&v.query, "", "query", "Gmail search query")
p.String(&v.srcs, "", "source", "comma list: onlyoffice,gmail")
return p
}
// ParseCLI reads args into a CLIConfig. Exit codes: 0 ok, 2 usage.
func ParseCLI(args []string) (CLIConfig, int, error) { func ParseCLI(args []string) (CLIConfig, int, error) {
fs := flag.NewFlagSet("mail/sync", flag.ContinueOnError) v := flagVals{workers: 4, query: "in:inbox", srcs: "onlyoffice"}
var ( p := bind(&v)
env = fs.String("env", "", ".env file (default: <cwd>/.env)") if err := cliparse.Parse(p, args); err != nil {
out = fs.String("out", "", "var/mail root (default: <cwd>/var/mail)") if errors.Is(err, cliparse.ErrHelp) {
workers = fs.Int("workers", 4, "concurrent downloads") return CLIConfig{Help: true}, 0, nil
limit = fs.Int("limit", 0, "max messages per source (0 = all)") }
offset = fs.Int("offset", 0, "skip first N messages per source")
force = fs.Bool("force", false, "overwrite existing message.json + attachments")
dryRun = fs.Bool("dry-run", false, "list message counts without writing")
query = fs.String("query", "in:inbox", "Gmail search query (gmail source only)")
srcs = fs.String("source", "onlyoffice", "comma list: onlyoffice,gmail (default onlyoffice)")
help = fs.Bool("help", false, "usage")
)
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return CLIConfig{}, 2, err return CLIConfig{}, 2, err
} }
if *help || fs.NArg() > 0 { if len(p.TrailingArguments) > 0 {
return CLIConfig{Help: true}, 0, nil return CLIConfig{Help: true}, 0, nil
} }
wd, err := os.Getwd() wd, err := os.Getwd()
if err != nil { if err != nil {
return CLIConfig{}, 2, err return CLIConfig{}, 2, err
} }
if *env == "" { if v.env == "" {
*env = filepath.Join(wd, ".env") v.env = filepath.Join(wd, ".env")
} }
if *out == "" { if v.out == "" {
*out = filepath.Join(wd, "var", "mail") v.out = filepath.Join(wd, "var", "mail")
} }
envVars := readEnv(*env) envVars := readEnv(v.env)
cfg := SyncConfig{ cfg := SyncConfig{
Out: *out, Out: v.out,
Workers: *workers, Workers: v.workers,
Limit: *limit, Limit: v.limit,
Offset: *offset, Offset: v.offset,
Force: *force, Force: v.force,
DryRun: *dryRun, DryRun: v.dryRun,
Query: *query, Query: v.query,
Policy: RetryPolicy{}, Policy: RetryPolicy{},
} }
cli := CLIConfig{Sync: cfg, Env: *env, Sources: *srcs} out := CLIConfig{Sync: cfg, Env: v.env, Sources: v.srcs}
for _, s := range strings.Split(*srcs, ",") { for _, s := range strings.Split(v.srcs, ",") {
switch strings.TrimSpace(s) { switch strings.TrimSpace(s) {
case "onlyoffice": case "onlyoffice":
u := pick(envVars["ONLYOFFICE_URL"], envVars["OO_URL"]) u := pick(envVars["ONLYOFFICE_URL"], envVars["OO_URL"])
user := pick(envVars["ONLYOFFICE_USER"], envVars["OO_USER"]) user := pick(envVars["ONLYOFFICE_USER"], envVars["OO_USER"])
pass := pick(envVars["ONLYOFFICE_PASS"], envVars["OO_PASSWORD"]) pass := pick(envVars["ONLYOFFICE_PASS"], envVars["OO_PASSWORD"])
if u == "" || user == "" || pass == "" { if u == "" || user == "" || pass == "" {
return CLIConfig{}, 2, fmt.Errorf("onlyoffice source needs ONLYOFFICE_URL/USER/PASS in %s", *env) return CLIConfig{}, 2, fmt.Errorf("onlyoffice source needs ONLYOFFICE_URL/USER/PASS in %s", v.env)
} }
cfg.OO = &OOConfig{URL: u, User: user, Password: pass} cfg.OO = &OOConfig{URL: u, User: user, Password: pass}
case "gmail": case "gmail":
@@ -85,30 +114,30 @@ func ParseCLI(args []string) (CLIConfig, int, error) {
return CLIConfig{}, 2, fmt.Errorf("unknown source %q", s) return CLIConfig{}, 2, fmt.Errorf("unknown source %q", s)
} }
} }
cli.Sync = cfg out.Sync = cfg
return cli, 0, nil return out, 0, nil
} }
// Main is the CLI entry: returns process exit code. // Main is the CLI entry: returns process exit code.
func Main(args []string) int { func Main(args []string) int {
cli, code, err := ParseCLI(args) cfg, code, err := ParseCLI(args)
if err != nil { if err != nil {
fmt.Fprintln(os.Stderr, "mail/sync:", err) fmt.Fprintln(os.Stderr, "mail/sync:", err)
return code return code
} }
if cli.Help { if cfg.Help {
fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail] [--query GMAIL_Q] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]") fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail] [--query GMAIL_Q] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]")
return 0 return 0
} }
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour) ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
defer cancel() defer cancel()
start := time.Now() start := time.Now()
stats, err := Run(ctx, cli.Sync) stats, err := Run(ctx, cfg.Sync)
if err != nil { if err != nil {
fmt.Fprintln(os.Stderr, "mail/sync:", err) fmt.Fprintln(os.Stderr, "mail/sync:", err)
return 1 return 1
} }
if cli.Sync.DryRun { if cfg.Sync.DryRun {
fmt.Printf("mail/sync: dry-run checked=%d (no writes)\n", stats.Checked) fmt.Printf("mail/sync: dry-run checked=%d (no writes)\n", stats.Checked)
return 0 return 0
} }
@@ -135,7 +164,6 @@ func readEnv(path string) map[string]string {
k, v, _ := strings.Cut(line, "=") k, v, _ := strings.Cut(line, "=")
out[strings.TrimSpace(k)] = strings.Trim(strings.TrimSpace(v), "\"'") out[strings.TrimSpace(k)] = strings.Trim(strings.TrimSpace(v), "\"'")
} }
// env overrides file
for _, kv := range os.Environ() { for _, kv := range os.Environ() {
k, v, ok := strings.Cut(kv, "=") k, v, ok := strings.Cut(kv, "=")
if !ok { if !ok {
+7 -22
View File
@@ -15,6 +15,7 @@ import (
"os" "os"
"strings" "strings"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/mdleaves" "github.com/eSlider/2dph/internal/mdleaves"
) )
@@ -23,29 +24,13 @@ func main() {
} }
func run(args []string) int { func run(args []string) int {
jsonOut := false c, err := mdleaves.ParseArgs(args)
files := "" if err != nil {
root := "." return cliparse.Fail(err)
for i := 0; i < len(args); i++ {
a := args[i]
switch {
case a == "--json":
jsonOut = true
case a == "--files" && i+1 < len(args):
i++
files = args[i]
case strings.HasPrefix(a, "--files="):
files = strings.TrimPrefix(a, "--files=")
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, "bin/markdown/import.go [dir] [--files a.md,b.md] [--json]")
return 0
case strings.HasPrefix(a, "-"):
fmt.Fprintln(os.Stderr, "unknown arg:", a)
return 2
default:
root = a
}
} }
jsonOut := c.JSONOut
files := c.Files
root := c.Root
var paths []string var paths []string
if files != "" { if files != "" {
+61
View File
@@ -0,0 +1,61 @@
//usr/bin/env go run -tags=qa_stats "$0" "$@"; exit
//go:build qa_stats
//
// bin/qa/stats.go - DuckDB quantiles over a JSON number array or JSONL count.
//
// ./bin/qa/stats.go <<< '[1,2,3,4,5]'
// ./bin/qa/stats.go --jsonl rows.jsonl
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
// DuckDB CGO needs gcc/g++ (not Zig). After eval "$(bin/cgo/zig env)":
// CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= ./bin/qa/stats.go
package main
import (
"encoding/json"
"fmt"
"io"
"os"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/duckstats"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
c, err := cliparse.ParseQAStats(args)
if err != nil {
return cliparse.Fail(err)
}
jsonl := c.JSONL
if jsonl != "" {
n, err := duckstats.CountJSONL(jsonl)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
fmt.Printf("n: %d\n", n)
return 0
}
raw, err := io.ReadAll(os.Stdin)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
var samples []float64
if err := json.Unmarshal(raw, &samples); err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
s, err := duckstats.Quantiles(samples)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
fmt.Printf("n: %d\nmin: %g\np50: %g\np95: %g\nmax: %g\navg: %g\n",
s.N, s.Min, s.P50, s.P95, s.Max, s.Avg)
return 0
}
+17 -35
View File
@@ -7,7 +7,7 @@
// ./bin/reasoner/bakeoff.go --model MichelRosselli/bonsai-27b:Q1_0 --json // ./bin/reasoner/bakeoff.go --model MichelRosselli/bonsai-27b:Q1_0 --json
// //
// Measures OpenAI tool_calls (search/get/audit) and RSS from Ollama /api/ps, not VRAM. // Measures OpenAI tool_calls (search/get/audit) and RSS from Ollama /api/ps, not VRAM.
// PicoClaw is not in this repo; the tool names match internal/httpapi MCP ops. // PicoClaw is compose profile picoclaw; tool names match internal/httpapi MCP ops.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang. // NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main package main
@@ -15,8 +15,9 @@ import (
"encoding/json" "encoding/json"
"fmt" "fmt"
"os" "os"
"strings"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/duckstats"
"github.com/eSlider/2dph/internal/reasoner" "github.com/eSlider/2dph/internal/reasoner"
) )
@@ -25,42 +26,21 @@ func main() {
} }
func run(args []string) int { func run(args []string) int {
base := os.Getenv("REASONER_BASE_URL") c, err := reasoner.ParseArgs(args)
if base == "" { if err != nil {
base = "http://127.0.0.1:11435/v1" return cliparse.Fail(err)
} }
model := os.Getenv("REASONER_MODEL") base, model, jsonOut, device := c.Base, c.Model, c.JSONOut, c.Device
if model == "" { client := reasoner.Client{BaseURL: base, Model: model, Device: device}
model = reasoner.OllamaRAM rep := reasoner.Run(client)
lat := make([]float64, 0, len(rep.Prompts))
for _, p := range rep.Prompts {
lat = append(lat, float64(p.LatencyMS))
} }
jsonOut := false if st, err := duckstats.Quantiles(lat); err == nil {
device := "cpu" rep.LatencyP50MS = st.P50
for i := 0; i < len(args); i++ { rep.LatencyP95MS = st.P95
a := args[i]
switch {
case a == "--json":
jsonOut = true
case a == "--model" && i+1 < len(args):
i++
model = args[i]
case strings.HasPrefix(a, "--model="):
model = strings.TrimPrefix(a, "--model=")
case a == "--base-url" && i+1 < len(args):
i++
base = args[i]
case a == "--device" && i+1 < len(args):
i++
device = args[i]
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, "bin/reasoner/bakeoff.go [--model ID] [--base-url URL] [--device cpu] [--json]")
return 0
default:
fmt.Fprintln(os.Stderr, "unknown arg:", a)
return 2
} }
}
c := reasoner.Client{BaseURL: base, Model: model, Device: device}
rep := reasoner.Run(c)
raw, err := json.MarshalIndent(rep, "", " ") raw, err := json.MarshalIndent(rep, "", " ")
if err != nil { if err != nil {
fmt.Fprintln(os.Stderr, err) fmt.Fprintln(os.Stderr, err)
@@ -76,6 +56,8 @@ func run(args []string) int {
fmt.Printf("xml_leak: %d\n", rep.XMLLeak) fmt.Printf("xml_leak: %d\n", rep.XMLLeak)
fmt.Printf("rss_mb: %d\n", rep.RSSMB) fmt.Printf("rss_mb: %d\n", rep.RSSMB)
fmt.Printf("vram_mb: %d\n", rep.VRAMMB) fmt.Printf("vram_mb: %d\n", rep.VRAMMB)
fmt.Printf("latency_p50_ms: %g\n", rep.LatencyP50MS)
fmt.Printf("latency_p95_ms: %g\n", rep.LatencyP95MS)
for _, p := range rep.Prompts { for _, p := range rep.Prompts {
status := "fail" status := "fail"
if p.OK { if p.OK {
+103
View File
@@ -0,0 +1,103 @@
"""D16 contradiction adjudication (same rules as internal/facts)."""
from __future__ import annotations
from typing import Any
CONF_CONFIRMED = "confirmed"
CONF_HYPOTHESIS = "hypothesis"
RULE_UNRESOLVED = "unresolved"
RULE_TEMPORAL = "temporal_freshness"
RULE_AUTHORITY = "authority_pairing"
RULE_TWO_SOURCE = "two_source"
RULE_SINGLE = "single_source"
KIND_RUNTIME = "runtime"
KIND_CONFIG = "config"
KIND_NARRATIVE = "narrative"
def _independent(sources: list[dict]) -> int:
seen: set[str] = set()
for i, s in enumerate(sources):
sid = str(s.get("id") or "") or f"{s.get('kind', '')}#{i}"
seen.add(sid)
return len(seen)
def _fresh_n(sources: list[dict]) -> int:
return sum(1 for s in sources if not s.get("stale"))
def _strong_n(sources: list[dict]) -> int:
return sum(1 for s in sources if s.get("kind") in (KIND_RUNTIME, KIND_CONFIG))
def adjudicate(claim: dict[str, Any]) -> dict[str, Any]:
yes = list(claim.get("yes") or [])
no = list(claim.get("no") or [])
yes_n, no_n = _independent(yes), _independent(no)
text = str(claim.get("text") or "")
def out(conf: str, rule: str, winner: str = "") -> dict[str, Any]:
return {
"text": text,
"confidence": conf,
"confirmed": conf == CONF_CONFIRMED,
"rule": rule,
"winner": winner,
"yes": yes_n,
"no": no_n,
}
if yes_n < 2 or no_n < 2:
if yes_n >= 2:
return out(CONF_CONFIRMED, RULE_TWO_SOURCE, "yes")
if no_n >= 2:
return out(CONF_CONFIRMED, RULE_TWO_SOURCE, "no")
return out(CONF_HYPOTHESIS, RULE_SINGLE)
yf, nf = _fresh_n(yes), _fresh_n(no)
if yf >= 2 and nf < 2:
return out(CONF_CONFIRMED, RULE_TEMPORAL, "yes")
if nf >= 2 and yf < 2:
return out(CONF_CONFIRMED, RULE_TEMPORAL, "no")
ys, ns = _strong_n(yes), _strong_n(no)
if ys >= 2 and ns < 2:
return out(CONF_CONFIRMED, RULE_AUTHORITY, "yes")
if ns >= 2 and ys < 2:
return out(CONF_CONFIRMED, RULE_AUTHORITY, "no")
return out(CONF_HYPOTHESIS, RULE_UNRESOLVED)
def parse_source_field(source: str) -> tuple[str, str]:
"""Split `a x b vs c x d` into (yes, no). Empty no if no ` vs `."""
if " vs " not in source:
return source, ""
yes, _, no = source.partition(" vs ")
return yes.strip(), no.strip()
def check_fact_row(lid: str, source: str, loc: str, how: str, conf: str) -> list[str]:
"""Lexicon checks for one facts leaf (no Ladybug)."""
problems: list[str] = []
src = source or ""
if conf == CONF_CONFIRMED:
if " vs " in src:
problems.append(f"{lid}: confirmed fact cannot keep a vs-contradiction")
if " x " not in src:
problems.append(f"{lid}: needs 2-source evidence in source, got '{source}'")
elif conf == CONF_HYPOTHESIS:
yes, no = parse_source_field(src)
if not no or " x " not in yes or " x " not in no:
problems.append(
f"{lid}: hypothesis contradiction needs 'a x b vs c x d', got '{source}'"
)
elif conf == "partial":
pass
else:
problems.append(f"{lid}: unknown confidence '{conf}'")
if not loc:
problems.append(f"{lid}: missing loc (evidence pointer)")
if not how:
problems.append(f"{lid}: missing how")
return problems
+84 -1
View File
@@ -7,7 +7,10 @@ offline against fixtures.
from __future__ import annotations from __future__ import annotations
import html import html
import os
import re import re
import subprocess
import tempfile
import zipfile import zipfile
from pathlib import Path from pathlib import Path
@@ -18,9 +21,10 @@ OFFICE_SUFFIXES = {".docx", ".pptx", ".xlsx", ".html", ".htm", ".epub", ".eml",
PDF_SUFFIXES = {".pdf"} PDF_SUFFIXES = {".pdf"}
IMAGE_SUFFIXES = {".png", ".jpg", ".jpeg", ".gif", ".bmp", ".tiff", ".tif", ".webp"} IMAGE_SUFFIXES = {".png", ".jpg", ".jpeg", ".gif", ".bmp", ".tiff", ".tif", ".webp"}
ARCHIVE_SUFFIXES = {".zip"} ARCHIVE_SUFFIXES = {".zip"}
# Legacy binary Office (doc/xls/ppt) — markitdown/docling skip them; we try # Legacy binary Office (doc/xls/ppt) — markitdown skip them; we try
# pandoc first, else leave a stub. # pandoc first, else leave a stub.
LEGACY_OFFICE_SUFFIXES = {".doc", ".xls", ".ppt"} LEGACY_OFFICE_SUFFIXES = {".doc", ".xls", ".ppt"}
TESS_LANG = "eng+deu"
CONVERTIBLE_SUFFIXES = ( CONVERTIBLE_SUFFIXES = (
TEXT_SUFFIXES | OFFICE_SUFFIXES | PDF_SUFFIXES | IMAGE_SUFFIXES | ARCHIVE_SUFFIXES | LEGACY_OFFICE_SUFFIXES TEXT_SUFFIXES | OFFICE_SUFFIXES | PDF_SUFFIXES | IMAGE_SUFFIXES | ARCHIVE_SUFFIXES | LEGACY_OFFICE_SUFFIXES
@@ -146,3 +150,82 @@ def zip_extract_safe(zip_path: Path, dest: Path) -> list[Path]:
def is_convertible(suffix: str) -> bool: def is_convertible(suffix: str) -> bool:
return suffix.lower() in CONVERTIBLE_SUFFIXES return suffix.lower() in CONVERTIBLE_SUFFIXES
def convert_pdf(path: Path, ocr: bool = False) -> str:
"""pdftotext -layout first; empty text layer → pdftoppm + tesseract.
`ocr` is unused for born-digital PDFs (text layer wins). Scans OCR
automatically. This path never execs an ONNX document converter.
"""
del ocr # scans OCR when the text layer is empty; flag is for images
text = pdf_fast_text(path)
if text and text.strip():
return normalize_markdown(text)
scanned = ocr_pdf(path)
if scanned and scanned.strip():
return normalize_markdown(scanned)
if text:
return normalize_markdown(text)
return "\n<!-- pdf has no text layer (ocr unavailable) -->\n"
def pdf_fast_text(path: Path) -> str | None:
"""pdftotext -layout; None when poppler is missing or the command fails."""
try:
proc = subprocess.run(
["pdftotext", "-layout", str(path), "-"],
capture_output=True, timeout=60)
except (OSError, subprocess.TimeoutExpired):
return None
if proc.returncode != 0:
return None
return proc.stdout.decode("utf-8", errors="replace")
def ocr_pdf(path: Path) -> str:
"""Rasterize with pdftoppm and OCR each page (tesseract or paddle)."""
try:
with tempfile.TemporaryDirectory(prefix="2dph-ocr-") as tmp:
prefix = str(Path(tmp) / "page")
proc = subprocess.run(
["pdftoppm", "-png", "-r", "200", str(path), prefix],
capture_output=True, timeout=120)
if proc.returncode != 0:
return ""
pages = sorted(Path(tmp).glob("page*.png"))
parts = [ocr_image(p) for p in pages]
return "\n\n".join(p for p in parts if p and p.strip())
except (OSError, subprocess.TimeoutExpired):
return ""
def ocr_image(path: Path) -> str:
engine = os.environ.get("OCR_ENGINE", "tesseract")
if engine == "paddle":
return _ocr_paddle(path)
return _ocr_tesseract(path)
def _ocr_tesseract(path: Path) -> str:
try:
proc = subprocess.run(
["tesseract", str(path), "stdout", "-l", TESS_LANG, "--psm", "6"],
capture_output=True, timeout=120)
except (OSError, subprocess.TimeoutExpired):
return ""
if proc.returncode != 0:
return ""
return proc.stdout.decode("utf-8", errors="replace").strip()
def _ocr_paddle(path: Path) -> str:
try:
proc = subprocess.run(
["paddleocr", "ocr", "-i", str(path)],
capture_output=True, timeout=180)
except (OSError, subprocess.TimeoutExpired):
return ""
if proc.returncode != 0:
return ""
return proc.stdout.decode("utf-8", errors="replace").strip()
+85
View File
@@ -127,6 +127,38 @@ class BinLayoutTest(unittest.TestCase):
self.assertIn("cmdbin.ExecFile", text) self.assertIn("cmdbin.ExecFile", text)
self.assertIn(f"bin/facts/{method.removesuffix('.go')}", text) self.assertIn(f"bin/facts/{method.removesuffix('.go')}", text)
def test_d16_adjudication_is_cgo_free(self) -> None:
self.assertTrue((ROOT / "internal" / "facts" / "contradict.go").is_file())
go = (ROOT / "internal" / "facts" / "contradict.go").read_text()
py = (ROOT / "bin" / "tools" / "contradict.py").read_text()
audit = (ROOT / "bin" / "facts" / "audit").read_text()
for token in ("temporal_freshness", "authority_pairing", "unresolved"):
self.assertIn(token, go)
self.assertIn(token, py)
self.assertIn("contradict", audit)
self.assertIn(" vs ", py)
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("temporal_freshness", plan)
self.assertIn("authority_pairing", plan)
shebang = (ROOT / "bin" / "facts" / "audit.go").read_text()
self.assertIn("contradict", shebang)
def test_d23_flaggy_cli(self) -> None:
self.assertTrue((ROOT / "internal" / "cli" / "cli.go").is_file())
self.assertIn("github.com/integrii/flaggy", (ROOT / "go.mod").read_text())
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D23", plan)
self.assertIn("flaggy", plan)
complete = (ROOT / "bin" / "cli" / "complete.go").read_text()
first = complete.splitlines()[0]
self.assertTrue(first.startswith("//usr/bin/env go run"), first)
self.assertIn("complete.go bash", complete)
self.assertIn("brain-search", complete)
chats_import = (ROOT / "internal" / "chats" / "import.go").read_text()
self.assertNotIn("flag.NewFlagSet", chats_import)
args = (ROOT / "internal" / "brain" / "rank" / "args.go").read_text()
self.assertIn("internal/cli", args)
def test_mail_import_is_shebang_not_brain_write(self) -> None: def test_mail_import_is_shebang_not_brain_write(self) -> None:
self._assert_shebang("bin/mail/import.go") self._assert_shebang("bin/mail/import.go")
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text() index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
@@ -136,6 +168,33 @@ class BinLayoutTest(unittest.TestCase):
"index_mail must point at bin/brain/index.go", "index_mail must point at bin/brain/index.go",
) )
def test_mail_ocr_is_tesseract_not_docling(self) -> None:
self._assert_shebang("bin/mail/ocr.go")
ocr = (ROOT / "bin" / "mail" / "ocr.go").read_text()
self.assertIn("internal/ocr", ocr)
self.assertIn("mail_ocr", ocr)
self.assertNotIn("github.com/otiai10/gosseract", ocr)
py = (ROOT / "bin" / "mail" / "import").read_text()
self.assertNotIn("from docling", py)
self.assertNotIn("import docling", py)
self.assertIn("convert_pdf", py)
conv = (ROOT / "bin" / "tools" / "mailconv.py").read_text()
self.assertIn("pdftotext", conv)
self.assertIn("pdftoppm", conv)
self.assertIn("tesseract", conv)
self.assertIn("eng+deu", conv)
self.assertNotIn("from docling", conv)
self.assertNotIn("import docling", conv)
self.assertNotIn("gocv", conv.lower())
proj = (ROOT / "pyproject.toml").read_text()
self.assertNotIn("docling", proj)
ci = (ROOT / ".github" / "workflows" / "ci.yml").read_text()
self.assertIn("tesseract-ocr", ci)
self.assertIn("./internal/ocr", ci)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn("ocr-paddle", compose)
self.assertIn("OCR_ENGINE", compose)
def test_markdown_import_is_go_not_python_exec(self) -> None: def test_markdown_import_is_go_not_python_exec(self) -> None:
self._assert_shebang("bin/markdown/import.go") self._assert_shebang("bin/markdown/import.go")
text = (ROOT / "bin" / "markdown" / "import.go").read_text() text = (ROOT / "bin" / "markdown" / "import.go").read_text()
@@ -190,6 +249,32 @@ class BinLayoutTest(unittest.TestCase):
if "go-git/go-git" in line: if "go-git/go-git" in line:
self.assertNotIn("indirect", line) self.assertNotIn("indirect", line)
def test_duckdb_go_is_direct_require(self) -> None:
text = (ROOT / "go.mod").read_text()
first = text.split("require (")[1].split(")")[0]
self.assertRegex(first, r"github.com/duckdb/duckdb-go/v2\s+v")
for line in first.splitlines():
if "duckdb/duckdb-go" in line:
self.assertNotIn("indirect", line)
skill = (ROOT / "skills" / "duckdb" / "SKILL.md").read_text()
self.assertIn("github.com/duckdb/duckdb-go", skill)
self.assertIn("Ladybug", skill)
self.assertIn("sqlite", skill.lower())
self.assertIn("gcc", skill.lower())
self.assertIn("Zig", skill)
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D22", plan)
self.assertIn("duckdb-go", plan)
self._assert_shebang("bin/qa/stats.go")
reasoner = (ROOT / "internal" / "reasoner" / "client.go").read_text()
self.assertNotIn("duckdb", reasoner)
self.assertNotIn("duckstats", reasoner)
bakeoff = (ROOT / "bin" / "reasoner" / "bakeoff.go").read_text()
self.assertIn("internal/duckstats", bakeoff)
webcache = (ROOT / "internal" / "websearch" / "cache.go").read_text()
self.assertNotIn("duckdb", webcache)
self.assertIn("modernc.org/sqlite", webcache)
def test_cgo_uses_zig_not_gcc(self) -> None: def test_cgo_uses_zig_not_gcc(self) -> None:
for rel in ("bin/cgo/zig", "bin/cgo/zcc", "bin/cgo/zc++"): for rel in ("bin/cgo/zig", "bin/cgo/zcc", "bin/cgo/zc++"):
p = ROOT / rel p = ROOT / rel
+104
View File
@@ -0,0 +1,104 @@
import os
import sys
import unittest
sys.path.insert(0, os.path.dirname(__file__))
from contradict import ( # noqa: E402
RULE_AUTHORITY,
RULE_SINGLE,
RULE_TEMPORAL,
RULE_TWO_SOURCE,
RULE_UNRESOLVED,
adjudicate,
check_fact_row,
parse_source_field,
)
def src(i, kind, stale=False):
return {"id": i, "kind": kind, "stale": stale}
class TestContradict(unittest.TestCase):
def test_two_vs_two_stays_hypothesis(self):
r = adjudicate({
"text": "svc listens on 443",
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
"no": [src("docker-old", "runtime"), src("compose-old", "config")],
})
self.assertFalse(r["confirmed"])
self.assertEqual(r["rule"], RULE_UNRESOLVED)
self.assertEqual(r["winner"], "")
def test_temporal_freshness(self):
r = adjudicate({
"text": "svc listens on 443",
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
"no": [src("old-readme", "narrative", True), src("old-wiki", "narrative", True)],
})
self.assertTrue(r["confirmed"])
self.assertEqual(r["rule"], RULE_TEMPORAL)
self.assertEqual(r["winner"], "yes")
def test_authority_pairing(self):
r = adjudicate({
"text": "svc listens on 443",
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
"no": [src("readme", "narrative"), src("wiki", "narrative")],
})
self.assertTrue(r["confirmed"])
self.assertEqual(r["rule"], RULE_AUTHORITY)
self.assertEqual(r["winner"], "yes")
def test_two_source_and_single(self):
two = adjudicate({
"text": "arc-1 runs Matrix",
"yes": [src("compose", "config"), src("docker-ps", "runtime")],
})
self.assertTrue(two["confirmed"])
self.assertEqual(two["rule"], RULE_TWO_SOURCE)
one = adjudicate({"text": "maybe", "yes": [src("readme", "narrative")]})
self.assertFalse(one["confirmed"])
self.assertEqual(one["rule"], RULE_SINGLE)
def test_parse_source_field(self):
yes, no = parse_source_field("docker ps x compose.yml vs old.md x wiki.md")
self.assertIn(" x ", yes)
self.assertIn(" x ", no)
def test_check_fact_row_allows_hypothesis_vs(self):
p = check_fact_row(
"L1", "a.md x b.md vs c.md x d.md", "var/", "audit", "hypothesis",
)
self.assertEqual(p, [])
p = check_fact_row("L2", "a.md x b.md", "var/", "audit", "confirmed")
self.assertEqual(p, [])
p = check_fact_row("L3", "a.md x b.md vs c.md x d.md", "var/", "audit", "confirmed")
self.assertTrue(any("vs-contradiction" in x for x in p))
p = check_fact_row("L4", "only-one.md", "var/", "audit", "hypothesis")
self.assertTrue(any("a x b vs" in x for x in p))
def test_audit_contradict_cli_unresolved(self):
import json
import subprocess
from pathlib import Path
root = Path(__file__).resolve().parents[2]
payload = json.dumps({
"text": "svc 443",
"yes": [src("a", "runtime"), src("b", "config")],
"no": [src("c", "runtime"), src("d", "config")],
})
proc = subprocess.run(
[sys.executable, str(root / "bin" / "facts" / "audit"), "contradict", "--json"],
input=payload, capture_output=True, text=True, check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
out = json.loads(proc.stdout)
self.assertTrue(out["ok"])
self.assertEqual(out["contradictions"][0]["rule"], RULE_UNRESOLVED)
self.assertFalse(out["contradictions"][0]["confirmed"])
if __name__ == "__main__":
unittest.main()
+89
View File
@@ -8,10 +8,13 @@ from pathlib import Path
sys.path.insert(0, os.path.dirname(__file__)) sys.path.insert(0, os.path.dirname(__file__))
from mailconv import ( # noqa: E402 from mailconv import ( # noqa: E402
TESS_LANG,
clean_email_address, clean_email_address,
convert_pdf,
html_to_markdown, html_to_markdown,
is_convertible, is_convertible,
normalize_markdown, normalize_markdown,
ocr_image,
split_zip_members, split_zip_members,
subject_to_filename, subject_to_filename,
zip_extract_safe, zip_extract_safe,
@@ -100,6 +103,92 @@ class TestMailConv(unittest.TestCase):
self.assertFalse(is_convertible(".exe")) self.assertFalse(is_convertible(".exe"))
self.assertFalse(is_convertible(".unknown")) self.assertFalse(is_convertible(".unknown"))
def test_convert_pdf_prefers_pdftotext(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b"Invoice BM25 layout"
stderr = b""
return P()
self._patch_run(mc, fake_run)
out = convert_pdf(Path(self._tmp("born.pdf")))
self.assertIn("BM25", out)
self.assertEqual(calls[0][:2], ["pdftotext", "-layout"])
self.assertFalse(any(c[0] == "tesseract" for c in calls))
self.assertFalse(any(c[0] == "pdftoppm" for c in calls))
def test_convert_pdf_empty_layer_uses_pdftoppm_tesseract(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b""
stderr = b""
if cmd[0] == "pdftotext":
P.stdout = b" \n"
return P()
if cmd[0] == "pdftoppm":
prefix = Path(cmd[-1])
(prefix.parent / "page-1.png").write_bytes(b"fake")
return P()
if cmd[0] == "tesseract":
P.stdout = b"scanned HELLO"
return P()
return P()
self._patch_run(mc, fake_run)
out = convert_pdf(Path(self._tmp("scan.pdf")))
self.assertIn("HELLO", out)
bins = [c[0] for c in calls]
self.assertIn("pdftotext", bins)
self.assertIn("pdftoppm", bins)
self.assertIn("tesseract", bins)
tess = next(c for c in calls if c[0] == "tesseract")
self.assertIn(TESS_LANG, tess)
self.assertNotIn("docling", " ".join(bins))
def test_ocr_image_paddle_engine(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b"paddle text"
stderr = b""
return P()
self._patch_run(mc, fake_run)
os.environ["OCR_ENGINE"] = "paddle"
try:
out = ocr_image(Path(self._tmp("x.png")))
finally:
os.environ.pop("OCR_ENGINE", None)
self.assertEqual(out, "paddle text")
self.assertEqual(calls[0][:2], ["paddleocr", "ocr"])
def _patch_run(self, mod, fn) -> None:
self.addCleanup(setattr, mod.subprocess, "run", mod.subprocess.run)
mod.subprocess.run = fn
def _mk_zip(self, members): def _mk_zip(self, members):
zpath = Path(self._tmp("arc.zip")) zpath = Path(self._tmp("arc.zip"))
with zipfile.ZipFile(zpath, "w") as zf: with zipfile.ZipFile(zpath, "w") as zf:
+1 -1
View File
@@ -121,7 +121,7 @@ class PublishedDocsTest(unittest.TestCase):
self.assertIn("D18", plan) self.assertIn("D18", plan)
self.assertIn("Qwen/Qwen3.5-9B", plan) self.assertIn("Qwen/Qwen3.5-9B", plan)
compose = (ROOT / "compose.yaml").read_text() compose = (ROOT / "compose.yaml").read_text()
self.assertIn('profiles: ["reasoner"]', compose) self.assertIn('"reasoner"', compose)
self.assertIn("OLLAMA_NUM_GPU", compose) self.assertIn("OLLAMA_NUM_GPU", compose)
self.assertIn("127.0.0.1:11435", compose) self.assertIn("127.0.0.1:11435", compose)
dockerfile = (ROOT / "Dockerfile").read_text() dockerfile = (ROOT / "Dockerfile").read_text()
+14
View File
@@ -47,3 +47,17 @@ class SkillsTest(unittest.TestCase):
self.assertIn("throttled", skill.lower()) self.assertIn("throttled", skill.lower())
self.assertIn("not a negative finding", agents) self.assertIn("not a negative finding", agents)
self.assertIn("Fact-check every", agents) self.assertIn("Fact-check every", agents)
def test_yq_is_mikefarah_for_structured_data(self) -> None:
skill = (ROOT / "skills" / "yq" / "SKILL.md").read_text()
self.assertIn("https://github.com/mikefarah/yq", skill)
for fmt in ("YAML", "JSON", "XML", "CSV", "TOML", "HCL"):
self.assertIn(fmt, skill)
self.assertIn("not kislyuk", skill.lower())
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("mikefarah/yq", plan)
agents = (ROOT / "AGENTS.md").read_text()
self.assertIn("mikefarah/yq", agents)
web = (ROOT / "skills" / "web-search" / "SKILL.md").read_text()
self.assertIn("| yq ", web)
self.assertNotIn("| jq ", web)
+36
View File
@@ -0,0 +1,36 @@
"""qa/system_perf.py is an offline-gated system test (no live brain in CI)."""
from __future__ import annotations
import ast
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class SystemPerfScriptTest(unittest.TestCase):
def test_script_compiles_and_is_read_only(self) -> None:
path = ROOT / "qa" / "system_perf.py"
src = path.read_text()
compile(src, str(path), "exec")
self.assertIn("--json", src)
self.assertIn("qwen3.5:9b", src)
self.assertIn("--picoclaw", src)
self.assertIn("BRAIN_URL", src)
self.assertIn("tools/list", src)
self.assertIn("tools/call", src)
self.assertIn("GATE_HEALTH_MS", src)
self.assertIn("GATE_GET_P50_MS", src)
self.assertNotIn("kb.lbug", src)
self.assertNotIn("password", src.lower())
self.assertNotIn("token", src.lower())
def test_script_does_not_write_ladybug(self) -> None:
tree = ast.parse((ROOT / "qa" / "system_perf.py").read_text())
writes = [
n.func.attr
for n in ast.walk(tree)
if isinstance(n, ast.Call) and isinstance(n.func, ast.Attribute)
and n.func.attr in {"write_text", "write_bytes", "dump"}
]
self.assertEqual(writes, [], f"system_perf must not write files: {writes}")
+8 -70
View File
@@ -16,9 +16,9 @@ import (
"fmt" "fmt"
"net/http" "net/http"
"os" "os"
"strconv"
"time" "time"
"github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/websearch" "github.com/eSlider/2dph/internal/websearch"
"golang.org/x/sys/unix" "golang.org/x/sys/unix"
) )
@@ -28,77 +28,15 @@ func main() {
} }
func run(args []string) int { func run(args []string) int {
var ( c, err := websearch.ParseArgs(args)
query, site, lang, fresh, category, engines string
limit = websearch.DefaultLimit
jsonOut, refresh, force bool
ttl = float64(websearch.CacheTTL)
timeout = 25
)
i := 0
for i < len(args) {
a := args[i]
switch {
case a == "--json":
jsonOut = true
case a == "--refresh":
refresh = true
case a == "--force":
force = true
case (a == "-n" || a == "--limit") && i+1 < len(args):
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 0 {
fmt.Fprintln(os.Stderr, "web/search: --limit must be a non-negative integer")
return 2
}
limit = n
case a == "--site" && i+1 < len(args):
i++
site = args[i]
case a == "--lang" && i+1 < len(args):
i++
lang = args[i]
case a == "--fresh" && i+1 < len(args):
i++
fresh = args[i]
case a == "--category" && i+1 < len(args):
i++
category = args[i]
case a == "--engines" && i+1 < len(args):
i++
engines = args[i]
case a == "--ttl" && i+1 < len(args):
i++
v, err := strconv.ParseFloat(args[i], 64)
if err != nil { if err != nil {
fmt.Fprintln(os.Stderr, "web/search: --ttl must be a number") return cli.Fail(err)
return 2
}
ttl = v
case a == "--timeout" && i+1 < len(args):
i++
n, err := strconv.Atoi(args[i])
if err != nil || n <= 0 {
fmt.Fprintln(os.Stderr, "web/search: --timeout must be a positive integer")
return 2
}
timeout = n
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, `usage: bin/web/search.go QUERY [--json] [-n N] [--site HOST] [--lang LANG] [--fresh day|week|month|year] [--category CAT] [--engines LIST] [--refresh] [--force]`)
return 0
case len(a) > 0 && a[0] != '-' && query == "":
query = a
default:
fmt.Fprintf(os.Stderr, "web/search: unknown flag %s\n", a)
return 2
}
i++
}
if query == "" {
fmt.Fprintln(os.Stderr, "web/search: query required")
return 2
} }
query, site, lang, fresh, category, engines := c.Query, c.Site, c.Lang, c.Fresh, c.Category, c.Engines
limit := c.Limit
jsonOut, refresh, force := c.JSONOut, c.Refresh, c.Force
ttl := c.TTL
timeout := c.Timeout
if site != "" { if site != "" {
query = "site:" + site + " " + query query = "site:" + site + " " + query
} }
+43 -4
View File
@@ -2,14 +2,23 @@
# #
# docker compose up -d brain # API (Zig CGO serve) # docker compose up -d brain # API (Zig CGO serve)
# docker compose --profile index run --rm index # Python rebuild # docker compose --profile index run --rm index # Python rebuild
# docker compose --profile picoclaw up brain-mcp # docker compose --profile picoclaw up -d # brain-mcp + CPU reasoner + PicoClaw gateway
# docker compose --profile reasoner up -d reasoner # CPU Ollama :11435 # docker compose --profile reasoner up -d reasoner # CPU Ollama :11435
# docker compose --profile searxng up -d # docker compose --profile searxng up -d
# OCR_ENGINE=paddle docker compose --profile ocr-paddle run --rm ocr-paddle
# #
# Secrets never baked in: search.env + db-profiles.yml from ~/.config/brain. # Secrets never baked in: search.env + db-profiles.yml from ~/.config/brain.
name: 2dph name: 2dph
networks:
default:
name: 2dph_sys
driver: bridge
ipam:
config:
- subnet: 10.23.42.0/24
services: services:
brain: brain:
image: ghcr.io/eslider/2dph:api image: ghcr.io/eslider/2dph:api
@@ -97,8 +106,8 @@ services:
- ./deploy/searxng/limiter.toml:/etc/searxng/limiter.toml:ro - ./deploy/searxng/limiter.toml:/etc/searxng/limiter.toml:ro
restart: unless-stopped restart: unless-stopped
# MCP endpoint for an external agent (PicoClaw is not shipped here). # MCP endpoint for PicoClaw (and any MCP client).
# docker compose --profile picoclaw up brain-mcp # docker compose --profile picoclaw up -d
brain-mcp: brain-mcp:
profiles: ["picoclaw"] profiles: ["picoclaw"]
image: ghcr.io/eslider/2dph:api image: ghcr.io/eslider/2dph:api
@@ -120,7 +129,7 @@ services:
# docker compose --profile reasoner up -d reasoner # docker compose --profile reasoner up -d reasoner
# docker compose --profile reasoner exec reasoner ollama pull qwen3.5:9b # docker compose --profile reasoner exec reasoner ollama pull qwen3.5:9b
reasoner: reasoner:
profiles: ["reasoner"] profiles: ["reasoner", "picoclaw"]
image: docker.io/ollama/ollama:latest image: docker.io/ollama/ollama:latest
environment: environment:
OLLAMA_NUM_GPU: "0" OLLAMA_NUM_GPU: "0"
@@ -131,7 +140,37 @@ services:
- reasoner-ollama:/root/.ollama - reasoner-ollama:/root/.ollama
restart: unless-stopped restart: unless-stopped
# Official PicoClaw gateway. Config has no secrets (Ollama + HTTP MCP).
# Host network: brain/reasoner bind 127.0.0.1 only, so host.docker.internal
# (docker0) cannot reach them. Gateway 127.0.0.1:18790 (not the 18800 launcher).
# If :8630/:11435 are already bound, do not start brain-mcp/reasoner:
# docker compose --profile picoclaw up -d --no-deps picoclaw
picoclaw:
profiles: ["picoclaw"]
image: docker.io/sipeed/picoclaw:v0.3.1
network_mode: host
depends_on:
- brain-mcp
- reasoner
environment:
PICOCLAW_GATEWAY_HOST: "127.0.0.1"
entrypoint: ["picoclaw", "gateway"]
volumes:
- picoclaw-home:/root/.picoclaw
- ./deploy/picoclaw/config.json:/root/.picoclaw/config.json:ro
restart: unless-stopped
# Optional PP-OCRv5 (not default). Default OCR is tesseract eng+deu.
# OCR_ENGINE=paddle docker compose --profile ocr-paddle run --rm ocr-paddle
ocr-paddle:
profiles: ["ocr-paddle"]
image: python:3.12-slim
environment:
OCR_ENGINE: paddle
command: ["python", "-c", "print('OCR_ENGINE=paddle; install paddleocr on PATH')"]
volumes: volumes:
kb-model: kb-model:
kb-var: kb-var:
reasoner-ollama: reasoner-ollama:
picoclaw-home:
+33
View File
@@ -0,0 +1,33 @@
{
"agents": {
"defaults": {
"model_name": "qwen3.5-9b",
"max_tool_iterations": 8,
"max_tokens": 512,
"context_window": 8192
}
},
"model_list": [
{
"model_name": "qwen3.5-9b",
"model": "ollama/qwen3.5:9b",
"api_base": "http://127.0.0.1:11435/v1",
"request_timeout": 600
}
],
"tools": {
"web": {
"enabled": false
},
"mcp": {
"enabled": true,
"servers": {
"2dph": {
"enabled": true,
"type": "http",
"url": "http://127.0.0.1:8630/mcp"
}
}
}
}
}
+1 -1
View File
@@ -20,7 +20,7 @@ Evidence-first knowledge graph. Facts need proof or they are
| explanation | [roadmap](roadmap.md) — gap to v1 (epic #16) | | explanation | [roadmap](roadmap.md) — gap to v1 (epic #16) |
| howto | [picoclaw](picoclaw.md) — MCP agent profile | | howto | [picoclaw](picoclaw.md) — MCP agent profile |
| howto | [reasoner](reasoner.md) — CPU bake-off (D18) | | howto | [reasoner](reasoner.md) — CPU bake-off (D18) |
| reference | [PLAN.md](../PLAN.md) — decisions D1D21 | | reference | [PLAN.md](../PLAN.md) — decisions D1D22 |
Decisions the public face must name: **D3** SearXNG compose, **D6** Go service / Decisions the public face must name: **D3** SearXNG compose, **D6** Go service /
Python write sidecar, **D14** `bin/{subject}/{method}.go`, **D15** Gitea origin, Python write sidecar, **D14** `bin/{subject}/{method}.go`, **D15** Gitea origin,
+7 -1
View File
@@ -38,6 +38,10 @@ bin/brain/search.go "question"
from each hit (1=File, 2=Commit, 3=Person). Rebuild writes FROM_FILE; from each hit (1=File, 2=Commit, 3=Person). Rebuild writes FROM_FILE;
git import writes HAS_VERSION/AUTHORED ([#17](https://git.produktor.io/eSlider/2dph/issues/17)). git import writes HAS_VERSION/AUTHORED ([#17](https://git.produktor.io/eSlider/2dph/issues/17)).
Go CLIs parse with **flaggy** via `internal/cli` (D23). Flags may appear
after positionals (`search q --hop 1`). Completions:
`source <(./bin/cli/complete.go bash)`.
## Who / What / How / Where / When + evidence ## Who / What / How / Where / When + evidence
Every assertion edge carries: Every assertion edge carries:
@@ -71,7 +75,9 @@ corpus HEAD.
- C: narrative — READMEs, AGENTS.md, docs - C: narrative — READMEs, AGENTS.md, docs
Confirmed = A×B or B×C agreement. Single source = hypothesis + `(not confirmed)`. Confirmed = A×B or B×C agreement. Single source = hypothesis + `(not confirmed)`.
Conflicting pairings (≥2 yes vs ≥2 no) = hypothesis (OQ1 → v2 resolution). Conflicting pairings (≥2 yes vs ≥2 no) stay hypothesis until
`temporal_freshness` or `authority_pairing` fires (`bin/facts/audit contradict`,
[#29](https://git.produktor.io/eSlider/2dph/issues/29)).
## Read path ## Read path
+22 -6
View File
@@ -1,18 +1,34 @@
# PicoClaw profile (reference agent) # PicoClaw profile (reference agent)
2dph is the memory/fact gate. PicoClaw (or any MCP client) is the agent loop 2dph is the memory/fact gate. Compose profile `picoclaw` runs the official
and is **not** shipped in this repo. PicoClaw gateway (`docker.io/sipeed/picoclaw:v0.3.1`) plus `brain-mcp` and the
CPU reasoner. Default agent model is `qwen3.5:9b` (RAM path, D18). Weights stay
in the reasoner volume, not in the 2dph image.
No secrets in git: Ollama needs no key; MCP is local HTTP.
```bash ```bash
docker compose --profile picoclaw up brain-mcp docker compose --profile picoclaw up -d
# already serving :8630 / :11435:
docker compose --profile picoclaw up -d --no-deps picoclaw
``` ```
The API listens on `127.0.0.1:8630`. Point the agent at Gateway: `127.0.0.1:18790`. Brain MCP: `http://127.0.0.1:8630/mcp`.
`http://127.0.0.1:8630/mcp` using [deploy/picoclaw/mcp.json.example](../deploy/picoclaw/mcp.json.example). Cursor-style clients can use [deploy/picoclaw/mcp.json.example](../deploy/picoclaw/mcp.json.example).
PicoClaw itself uses [deploy/picoclaw/config.json](../deploy/picoclaw/config.json)
(`127.0.0.1` + host network — loopback publishes are not reachable via docker0).
OpenAPI: `GET http://127.0.0.1:8630/openapi.json`. OpenAPI: `GET http://127.0.0.1:8630/openapi.json`.
Before a factual reply: `search``get``audit`. `throttled` is not a Before a factual reply: `search``get``audit`. `throttled` is not a
negative finding. See `skills/picoclaw/SKILL.md`. negative finding. See `skills/picoclaw/SKILL.md`.
No Cursor required. A live PicoClaw binary/image is an operator choice. System performance (MCP gates + qwen3.5:9b tool_call + PicoClaw gateway):
```bash
./qa/system_perf.py --json | yq '.gates'
REASONER_MODEL=qwen3.5:9b ./qa/system_perf.py --reasoner --picoclaw --json | yq '.reasoner'
```
The default agent model is `qwen3.5:9b`. PicoClaw `context_window` is 8192
(heuristic `max_tokens*4` at 512 is 2048, too small for MCP tool schemas).
`request_timeout` is 600s for a CPU turn (tool_call + MCP search + answer).
+4 -2
View File
@@ -1,8 +1,8 @@
# Reasoner bake-off (D18) # Reasoner bake-off (D18)
Pluggable OpenAI-compatible URL. 2dph does not ship weights. PicoClaw is Pluggable OpenAI-compatible URL. 2dph does not ship weights. PicoClaw is
not in this repo; the bake-off hits the same tool names PicoClaw would compose profile `picoclaw` (`sipeed/picoclaw`); the bake-off hits the same
(`search``get``audit` from `internal/httpapi.Ops`). tool names (`search``get``audit` from `internal/httpapi.Ops`).
```bash ```bash
docker compose --profile reasoner up -d reasoner docker compose --profile reasoner up -d reasoner
@@ -11,6 +11,8 @@ REASONER_BASE_URL=http://127.0.0.1:11435/v1 REASONER_MODEL=qwen3.5:9b \
./bin/reasoner/bakeoff.go --json ./bin/reasoner/bakeoff.go --json
``` ```
JSON includes `latency_p50_ms` / `latency_p95_ms` from DuckDB (`internal/duckstats`, D22).
Host Ollama on `:11434` is left alone. This sidecar binds `127.0.0.1:11435` Host Ollama on `:11434` is left alone. This sidecar binds `127.0.0.1:11435`
with `OLLAMA_NUM_GPU=0` (CPU). Measure RSS (`/api/ps` `size`), not VRAM. with `OLLAMA_NUM_GPU=0` (CPU). Measure RSS (`/api/ps` `size`), not VRAM.
+12 -4
View File
@@ -32,11 +32,19 @@ FROM_FILE / HAS_VERSION / AUTHORED.
`--with-chats` on rebuild (WhatsApp out of v1). `--with-chats` on rebuild (WhatsApp out of v1).
[#19](https://git.produktor.io/eSlider/2dph/issues/19) CI recall SoT = [#19](https://git.produktor.io/eSlider/2dph/issues/19) CI recall SoT =
`bin/brain/eval.go` via Zig. `bin/brain/eval.go` via Zig.
Epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed.
## v2
[#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR — **in**.
[#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3 duckdb-go — **in**.
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1 contradiction
resolution — **in** (`temporal_freshness`, `authority_pairing`).
[#34](https://git.produktor.io/eSlider/2dph/issues/34) D23 flaggy CLI — **in**.
## Blockers ## Blockers
None for epic #16. v2: [#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR, None for epic #16 (closed). Remaining v2: OQ4.
OQ1 contradiction resolution, OQ3 duckdb-md, OQ4 YAML-first leafs.
``` ```
question question
@@ -51,8 +59,8 @@ question
## Not v1 ## Not v1
[#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR (OQ2), OQ1 OQ4 YAML-first leafs.
contradiction resolution, OQ3 duckdb-md export, OQ4 YAML-first leafs. OCR (OQ2), duckdb-go (OQ3/D22), and D16 adjudication (OQ1) are in.
## Close epic #16 when ## Close epic #16 when
+2
View File
@@ -16,6 +16,7 @@ No laptop-absolute paths. Config lives in env files under `$HOME/.config/brain/`
- Go (see `go.mod`) - Go (see `go.mod`)
- Python 3.12 + [uv](https://docs.astral.sh/uv) - Python 3.12 + [uv](https://docs.astral.sh/uv)
- Optional: Docker, Zig CGO via `bin/cgo/zig` (not gcc) - Optional: Docker, Zig CGO via `bin/cgo/zig` (not gcc)
- Optional: poppler (`pdftotext`/`pdftoppm`) + tesseract `eng+deu` for mail OCR
```bash ```bash
uv venv .venv uv venv .venv
@@ -52,6 +53,7 @@ bin/brain/add.go --text "arc-1 runs Matrix" --root facts --source "compose.yml x
bin/brain/index.go --rebuild --with-facts --with-chats bin/brain/index.go --rebuild --with-facts --with-chats
bin/brain/search.go "LadybugDB vector index" # facts → info → web (D17) bin/brain/search.go "LadybugDB vector index" # facts → info → web (D17)
bin/brain/search.go "upstream flag" --no-web bin/brain/search.go "upstream flag" --no-web
source <(./bin/cli/complete.go bash) # D23 flaggy complete
bin/brain/get.go <id> --body bin/brain/get.go <id> --body
bin/brain/stats.go bin/brain/stats.go
``` ```
+9
View File
@@ -7,7 +7,9 @@ require (
github.com/arran4/golang-ical v0.3.5 github.com/arran4/golang-ical v0.3.5
github.com/chewxy/math32 v1.11.2 github.com/chewxy/math32 v1.11.2
github.com/daulet/tokenizers v1.27.0 github.com/daulet/tokenizers v1.27.0
github.com/duckdb/duckdb-go/v2 v2.10505.0
github.com/go-git/go-git/v5 v5.19.2 github.com/go-git/go-git/v5 v5.19.2
github.com/integrii/flaggy v1.8.0
golang.org/x/sys v0.47.0 golang.org/x/sys v0.47.0
golang.org/x/text v0.40.0 golang.org/x/text v0.40.0
modernc.org/sqlite v1.56.0 modernc.org/sqlite v1.56.0
@@ -20,10 +22,17 @@ require (
github.com/apache/arrow-go/v18 v18.6.0 // indirect github.com/apache/arrow-go/v18 v18.6.0 // indirect
github.com/cloudflare/circl v1.6.3 // indirect github.com/cloudflare/circl v1.6.3 // indirect
github.com/cyphar/filepath-securejoin v0.6.1 // indirect github.com/cyphar/filepath-securejoin v0.6.1 // indirect
github.com/duckdb/duckdb-go-bindings v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/darwin-amd64 v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/darwin-arm64 v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/linux-amd64 v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/linux-arm64 v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/windows-amd64 v0.10505.0 // indirect
github.com/dustin/go-humanize v1.0.1 // indirect github.com/dustin/go-humanize v1.0.1 // indirect
github.com/emirpasic/gods v1.18.1 // indirect github.com/emirpasic/gods v1.18.1 // indirect
github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376 // indirect github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376 // indirect
github.com/go-git/go-billy/v5 v5.9.0 // indirect github.com/go-git/go-billy/v5 v5.9.0 // indirect
github.com/go-viper/mapstructure/v2 v2.5.0 // indirect
github.com/goccy/go-json v0.10.6 // indirect github.com/goccy/go-json v0.10.6 // indirect
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 // indirect github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 // indirect
github.com/google/flatbuffers v25.12.19+incompatible // indirect github.com/google/flatbuffers v25.12.19+incompatible // indirect
+18
View File
@@ -31,6 +31,20 @@ github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSs
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM= github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/duckdb/duckdb-go-bindings v0.10505.0 h1:/0pPsTLrcCsTGxT0VrHgJWnOcPe1tQL1vrki1v3jbAI=
github.com/duckdb/duckdb-go-bindings v0.10505.0/go.mod h1:HoD5xePkDj3VZbBnVVfxVVYIljZ9khCprWA7FgwIiC4=
github.com/duckdb/duckdb-go-bindings/lib/darwin-amd64 v0.10505.0 h1:FrMqquFBQlMsi34h2KZgCku54rqA8xEbXZ0NLVDKwYs=
github.com/duckdb/duckdb-go-bindings/lib/darwin-amd64 v0.10505.0/go.mod h1:EnAvZh1kNJHp5yF+M1ZHNEvapnmt6anq1xXHVrAGqMo=
github.com/duckdb/duckdb-go-bindings/lib/darwin-arm64 v0.10505.0 h1:lbRbpQwT1MmUhh/VTwukV9K8bxKByV3UghAP3MvsbBo=
github.com/duckdb/duckdb-go-bindings/lib/darwin-arm64 v0.10505.0/go.mod h1:IGLSeEcFhNeZF16aVjQCULD7TsFZKG5G7SyKJAXKp5c=
github.com/duckdb/duckdb-go-bindings/lib/linux-amd64 v0.10505.0 h1:nrsaVYj3XYCRbS2FpdOMD/KHE7egRMr+/NR1IHmjT84=
github.com/duckdb/duckdb-go-bindings/lib/linux-amd64 v0.10505.0/go.mod h1:KAIynZ0GHCS7X5fRyuFnQMg/SZBPK/bS9OCOVojClxw=
github.com/duckdb/duckdb-go-bindings/lib/linux-arm64 v0.10505.0 h1:qM6oGDgwXBILJGbTY4fCy6QOczLpucUA6yn6g3ORjh4=
github.com/duckdb/duckdb-go-bindings/lib/linux-arm64 v0.10505.0/go.mod h1:81SGOYoEUs8qaAfSk1wRfM5oobrIJ5KI7AzYhK6/bvQ=
github.com/duckdb/duckdb-go-bindings/lib/windows-amd64 v0.10505.0 h1:DjqZl9rYreHkSOqnqLmkrqH5T8UdQNcxZLJVZzGmXXA=
github.com/duckdb/duckdb-go-bindings/lib/windows-amd64 v0.10505.0/go.mod h1:K25pJL26ARblGDeuAkrdblFvUen92+CwksLtPEHRqqQ=
github.com/duckdb/duckdb-go/v2 v2.10505.0 h1:SWwvLn2Qx/RQSnQNupwgIF8VbnJ5A6OQU9lYb/mDETI=
github.com/duckdb/duckdb-go/v2 v2.10505.0/go.mod h1:m0PW4J4FG9hlFlVdXi6Ds9owpyIDaBdE2jyce00fGcE=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY= github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto= github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/elazarl/goproxy v1.7.2 h1:Y2o6urb7Eule09PjlhQRGNsqRfPmYI3KKQLFpCAV3+o= github.com/elazarl/goproxy v1.7.2 h1:Y2o6urb7Eule09PjlhQRGNsqRfPmYI3KKQLFpCAV3+o=
@@ -47,6 +61,8 @@ github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399 h1:eMj
github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399/go.mod h1:1OCfN199q1Jm3HZlxleg+Dw/mwps2Wbk9frAWm+4FII= github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399/go.mod h1:1OCfN199q1Jm3HZlxleg+Dw/mwps2Wbk9frAWm+4FII=
github.com/go-git/go-git/v5 v5.19.2 h1:wkfn7vOlUBu8ivAWKBWisTiwJK4jYHzTF8Ndv1LyGqY= github.com/go-git/go-git/v5 v5.19.2 h1:wkfn7vOlUBu8ivAWKBWisTiwJK4jYHzTF8Ndv1LyGqY=
github.com/go-git/go-git/v5 v5.19.2/go.mod h1:QqCBE1EFN5ddFmrliLQ3/ntRCUjZU3EJuwuB/jWEHjk= github.com/go-git/go-git/v5 v5.19.2/go.mod h1:QqCBE1EFN5ddFmrliLQ3/ntRCUjZU3EJuwuB/jWEHjk=
github.com/go-viper/mapstructure/v2 v2.5.0 h1:vM5IJoUAy3d7zRSVtIwQgBj7BiWtMPfmPEgAXnvj1Ro=
github.com/go-viper/mapstructure/v2 v2.5.0/go.mod h1:oJDH3BJKyqBA2TXFhDsKDGDTlndYOZ6rGS0BRZIxGhM=
github.com/goccy/go-json v0.10.6 h1:p8HrPJzOakx/mn/bQtjgNjdTcN+/S6FcG2CTtQOrHVU= github.com/goccy/go-json v0.10.6 h1:p8HrPJzOakx/mn/bQtjgNjdTcN+/S6FcG2CTtQOrHVU=
github.com/goccy/go-json v0.10.6/go.mod h1:oq7eo15ShAhp70Anwd5lgX2pLfOS3QCiwU/PULtXL6M= github.com/goccy/go-json v0.10.6/go.mod h1:oq7eo15ShAhp70Anwd5lgX2pLfOS3QCiwU/PULtXL6M=
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 h1:f+oWsMOmNPc8JmEHVZIycC7hBoQxHH9pNKQORJNozsQ= github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 h1:f+oWsMOmNPc8JmEHVZIycC7hBoQxHH9pNKQORJNozsQ=
@@ -61,6 +77,8 @@ github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo= github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k= github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM= github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
github.com/integrii/flaggy v1.8.0 h1:tC1qWwg4fhF2Qdaj+MpPK04cxlOSq0+HoMZqAW6Arao=
github.com/integrii/flaggy v1.8.0/go.mod h1:QS4c80m87SXG0pmVUT/Lx2RY5EbkLvLp7IKBD2jwcFA=
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99 h1:BQSFePA1RWJOlocH6Fxy8MmwDt+yVQYULKfN0RoTN8A= github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99 h1:BQSFePA1RWJOlocH6Fxy8MmwDt+yVQYULKfN0RoTN8A=
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99/go.mod h1:1lJo3i6rXxKeerYnT8Nvf0QmHCRC1n8sfWVwXF2Frvo= github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99/go.mod h1:1lJo3i6rXxKeerYnT8Nvf0QmHCRC1n8sfWVwXF2Frvo=
github.com/kevinburke/ssh_config v1.2.0 h1:x584FjTGwHzMwvHx18PXxbBVzfnxogHaAReU4gf13a4= github.com/kevinburke/ssh_config v1.2.0 h1:x584FjTGwHzMwvHx18PXxbBVzfnxogHaAReU4gf13a4=
+43 -45
View File
@@ -3,12 +3,15 @@ package rank
import ( import (
"fmt" "fmt"
"strconv" "strconv"
"strings"
"github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
) )
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--hop N] [--json] [--no-web] const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--hop N] [--json] [--no-web]
bin/brain/search.go serve [port] bin/brain/search.go serve [port]
bin/brain/search.go --list-model` bin/brain/search.go --list-model
source <(./bin/cli/complete.go bash)`
type Options struct { type Options struct {
Query string Query string
@@ -21,62 +24,57 @@ type Options struct {
NoWeb bool NoWeb bool
} }
// NewParser is the flaggy schema for search (also used by bin/cli/complete.go).
func NewParser(opt *Options) *flaggy.Parser {
if opt.Limit == 0 {
opt.Limit = 20
}
p := cli.New("brain-search")
p.Description = "deduction search: facts → info → web"
p.String(&opt.Root, "", "root", "facts or info")
p.String(&opt.Repo, "", "repo", "filter by repo")
p.Int(&opt.Limit, "n", "n", "max hits")
p.Int(&opt.Hop, "", "hop", "walk FROM_FILE depth 1-3")
p.Bool(&opt.JSONOut, "", "json", "JSON output")
p.Bool(&opt.NoWeb, "", "no-web", "stay local")
p.Bool(&opt.ListModel, "", "list-model", "print embedding model")
return p
}
// ParseArgs reads flags. Unknown flags are an error: silently dropping them // ParseArgs reads flags. Unknown flags are an error: silently dropping them
// meant `--hop 1` vanished and its argument `1` was appended to the query. // meant `--hop 1` vanished and its argument `1` was appended to the query.
func ParseArgs(args []string) (Options, error) { func ParseArgs(args []string) (Options, error) {
opt := Options{Limit: 20} opt := Options{Limit: 20}
var queryArgs []string p := NewParser(&opt)
var q string
for i := 0; i < len(args); i++ { p.AddPositionalValue(&q, "query", 1, false, "search query")
arg := args[i] if err := cli.Parse(p, args); err != nil {
wantsValue := arg == "--root" || arg == "--repo" || arg == "-n" || arg == "--hop" return opt, err
if wantsValue && i+1 >= len(args) {
return opt, fmt.Errorf("%s needs a value", arg)
} }
switch arg { opt.Query = cli.Query(q, p.TrailingArguments)
case "--root": if opt.Root != "" && opt.Root != "facts" && opt.Root != "info" {
i++
opt.Root = args[i]
if opt.Root != "facts" && opt.Root != "info" {
return opt, fmt.Errorf("--root must be facts or info, got %q", opt.Root) return opt, fmt.Errorf("--root must be facts or info, got %q", opt.Root)
} }
case "--repo": if opt.Limit < 1 {
i++ return opt, fmt.Errorf("-n must be a positive integer, got %q", strconv.Itoa(opt.Limit))
opt.Repo = args[i]
case "-n":
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 1 {
return opt, fmt.Errorf("-n must be a positive integer, got %q", args[i])
} }
opt.Limit = n if opt.Hop < 0 {
case "--hop": return opt, fmt.Errorf("--hop must be a positive integer, got %q", strconv.Itoa(opt.Hop))
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 1 {
return opt, fmt.Errorf("--hop must be a positive integer, got %q", args[i])
} }
if n > 3 { if opt.Hop > 3 {
return opt, fmt.Errorf("--hop max is 3 (File → Commit → Person)") return opt, fmt.Errorf("--hop max is 3 (File → Commit → Person)")
} }
opt.Hop = n
case "--json":
opt.JSONOut = true
case "--no-web":
opt.NoWeb = true
case "--list-model":
opt.ListModel = true
default:
if strings.HasPrefix(arg, "-") {
return opt, fmt.Errorf("unknown flag %q", arg)
}
queryArgs = append(queryArgs, arg)
}
}
opt.Query = strings.TrimSpace(strings.Join(queryArgs, " "))
if opt.Query == "" && !opt.ListModel { if opt.Query == "" && !opt.ListModel {
return opt, fmt.Errorf("no query given") return opt, fmt.Errorf("no query given")
} }
return opt, nil return opt, nil
} }
// Parser is the search schema for bin/cli/complete.go.
func Parser() *flaggy.Parser {
opt := Options{Limit: 20}
p := NewParser(&opt)
var q string
p.AddPositionalValue(&q, "query", 1, false, "search query")
return p
}
+17 -2
View File
@@ -19,20 +19,35 @@ type SecondSourceHit struct {
type WebFn func(query string) SecondSource type WebFn func(query string) SecondSource
// ShouldEscalate is true when the default deduction path has no facts hit. // ShouldEscalate is true when the default deduction path has no confirmed
// facts hit. Hypothesis/partial facts are `(not confirmed)` (D16).
// `--root facts|info` is a single-root ask: do not mix in the web. // `--root facts|info` is a single-root ask: do not mix in the web.
func ShouldEscalate(hits []Hit, rootFilter string) bool { func ShouldEscalate(hits []Hit, rootFilter string) bool {
if rootFilter != "" { if rootFilter != "" {
return false return false
} }
for _, h := range hits { for _, h := range hits {
if h.Root == "facts" { if ConfirmedFact(h) {
return false return false
} }
} }
return true return true
} }
// ConfirmedFact is a facts-root hit that is not hypothesis/partial.
// Empty confidence is treated as confirmed (legacy leafs).
func ConfirmedFact(h Hit) bool {
if h.Root != "facts" {
return false
}
switch h.Confidence {
case "hypothesis", "partial":
return false
default:
return true
}
}
// Deduce returns the second-source block, or nil when web must not run. // Deduce returns the second-source block, or nil when web must not run.
func Deduce(hits []Hit, query, rootFilter string, noWeb bool, web WebFn) *SecondSource { func Deduce(hits []Hit, query, rootFilter string, noWeb bool, web WebFn) *SecondSource {
if noWeb || web == nil || !ShouldEscalate(hits, rootFilter) { if noWeb || web == nil || !ShouldEscalate(hits, rootFilter) {
+10
View File
@@ -14,6 +14,16 @@ func TestShouldEscalateWhenNoFacts(t *testing.T) {
} }
} }
func TestShouldEscalateWhenHypothesisFacts(t *testing.T) {
hyp := Hit{ID: "c", Root: "facts", Confidence: "hypothesis", Source: "a x b vs c x d"}
if !ShouldEscalate([]Hit{hyp}, "") {
t.Fatal("hypothesis facts are (not confirmed); escalate")
}
if ConfirmedFact(hyp) {
t.Fatal("hypothesis is not confirmed")
}
}
func TestShouldNotEscalateWhenFactsConfirm(t *testing.T) { func TestShouldNotEscalateWhenFactsConfirm(t *testing.T) {
hits := []Hit{h("f", "facts", "docker ps x compose"), h("i", "info", "docs/a.md")} hits := []Hit{h("f", "facts", "docker ps x compose"), h("i", "info", "docs/a.md")}
if ShouldEscalate(hits, "") { if ShouldEscalate(hits, "") {
+2 -2
View File
@@ -3,10 +3,10 @@ package rank
// BM25 ranks best-first, so the top hits are the *highest* scores; cosine // BM25 ranks best-first, so the top hits are the *highest* scores; cosine
// distance ranks best-first ascending. Both mirror kblib.py. // distance ranks best-first ascending. Both mirror kblib.py.
const FTSStmt = "CALL QUERY_FTS_INDEX('Leaf', 'id', $q) " + const FTSStmt = "CALL QUERY_FTS_INDEX('Leaf', 'id', $q) " +
"RETURN node.id, node.text, node.root, node.source, score ORDER BY score DESC LIMIT $n" "RETURN node.id, node.text, node.root, node.source, score, node.confidence ORDER BY score DESC LIMIT $n"
const VecStmt = "CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) " + const VecStmt = "CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) " +
"RETURN node.id, node.text, node.root, node.source, distance ORDER BY distance LIMIT $n" "RETURN node.id, node.text, node.root, node.source, distance, node.confidence ORDER BY distance LIMIT $n"
// HopStmt is the Cypher walk from a search hit. Depth 1 = File, 2 = Commit, 3 = Person. // HopStmt is the Cypher walk from a search hit. Depth 1 = File, 2 = Commit, 3 = Person.
func HopStmt(depth int) string { func HopStmt(depth int) string {
+1
View File
@@ -19,6 +19,7 @@ type Hit struct {
ID string `json:"id"` ID string `json:"id"`
Text string `json:"text"` Text string `json:"text"`
Root string `json:"root"` Root string `json:"root"`
Confidence string `json:"confidence,omitempty"`
Source string `json:"-"` Source string `json:"-"`
Score float64 `json:"score"` Score float64 `json:"score"`
Snippet string `json:"snippet,omitempty"` Snippet string `json:"snippet,omitempty"`
+3
View File
@@ -172,4 +172,7 @@ func TestFTSQueryOrdersByScoreDescending(t *testing.T) {
if !strings.Contains(FTSStmt, "ORDER BY score DESC") { if !strings.Contains(FTSStmt, "ORDER BY score DESC") {
t.Fatalf("FTS query must order by score DESC, got:\n%s", FTSStmt) t.Fatalf("FTS query must order by score DESC, got:\n%s", FTSStmt)
} }
if !strings.Contains(FTSStmt, "node.confidence") {
t.Fatal("FTS must return confidence for D16")
}
} }
+59
View File
@@ -0,0 +1,59 @@
package rank
import (
"fmt"
"github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
)
type GetOptions struct {
ID string
Body bool
JSONOut bool
}
func GetParser(opt *GetOptions) *flaggy.Parser {
p := cli.New("brain-get")
p.Description = "read one leaf"
p.Bool(&opt.Body, "", "body", "full text instead of snippet")
p.Bool(&opt.JSONOut, "", "json", "JSON output")
p.AddPositionalValue(&opt.ID, "id", 1, false, "leaf id")
return p
}
func ParseGet(args []string) (GetOptions, error) {
var opt GetOptions
if err := cli.Parse(GetParser(&opt), args); err != nil {
return opt, err
}
if opt.ID == "" {
return opt, fmt.Errorf("id required")
}
return opt, nil
}
type JSONFlag struct {
JSONOut bool
}
func bindJSON(name string, opt *JSONFlag) *flaggy.Parser {
p := cli.New(name)
p.Bool(&opt.JSONOut, "", "json", "JSON output")
return p
}
func StatsParser() *flaggy.Parser {
opt := JSONFlag{}
return bindJSON("brain-stats", &opt)
}
func EvalParser() *flaggy.Parser {
opt := JSONFlag{}
return bindJSON("brain-eval", &opt)
}
func ParseJSONFlag(name string, args []string) (JSONFlag, error) {
var opt JSONFlag
return opt, cli.Parse(bindJSON(name, &opt), args)
}
+13 -48
View File
@@ -11,30 +11,15 @@ import (
"unicode/utf8" "unicode/utf8"
"github.com/eSlider/2dph/internal/brain/rank" "github.com/eSlider/2dph/internal/brain/rank"
"github.com/eSlider/2dph/internal/cli"
) )
func MainGet(args []string) int { func MainGet(args []string) int {
id, body, jsonOut := "", false, false opt, err := rank.ParseGet(args)
for _, a := range args { if err != nil {
switch { return cli.Fail(err)
case a == "--body":
body = true
case a == "--json":
jsonOut = true
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/get.go <id> [--body] [--json]`)
return 0
case strings.HasPrefix(a, "-"):
fmt.Fprintf(os.Stderr, "brain/get: unknown flag %s\n", a)
return 2
default:
id = a
}
}
if id == "" {
fmt.Fprintln(os.Stderr, "brain/get: id required")
return 2
} }
id, body, jsonOut := opt.ID, opt.Body, opt.JSONOut
if err := openBrain(); err != nil { if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err) fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1 return 1
@@ -72,21 +57,11 @@ func MainGet(args []string) int {
} }
func MainStats(args []string) int { func MainStats(args []string) int {
jsonOut := false opt, err := rank.ParseJSONFlag("brain-stats", args)
for _, a := range args { if err != nil {
switch a { return cli.Fail(err)
case "--json":
jsonOut = true
case "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/stats.go [--json]`)
return 0
default:
if strings.HasPrefix(a, "-") {
fmt.Fprintf(os.Stderr, "brain/stats: unknown flag %s\n", a)
return 2
}
}
} }
jsonOut := opt.JSONOut
if err := openBrain(); err != nil { if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err) fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1 return 1
@@ -124,21 +99,11 @@ func MainStats(args []string) int {
} }
func MainEval(args []string) int { func MainEval(args []string) int {
jsonOut := false opt, err := rank.ParseJSONFlag("brain-eval", args)
for _, a := range args { if err != nil {
switch a { return cli.Fail(err)
case "--json":
jsonOut = true
case "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/eval.go [--json]`)
return 0
default:
if strings.HasPrefix(a, "-") {
fmt.Fprintf(os.Stderr, "brain/eval: unknown flag %s\n", a)
return 2
}
}
} }
jsonOut := opt.JSONOut
if err := openBrain(); err != nil { if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err) fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1 return 1
+14 -1
View File
@@ -20,6 +20,7 @@ import (
lbug "github.com/LadybugDB/go-ladybug" lbug "github.com/LadybugDB/go-ladybug"
"github.com/eSlider/2dph/internal/brain/rank" "github.com/eSlider/2dph/internal/brain/rank"
"github.com/eSlider/2dph/internal/cli"
) )
const defaultPort = 17830 const defaultPort = 17830
@@ -29,6 +30,9 @@ const healthPath = "/health"
func runSearch(args []string) int { func runSearch(args []string) int {
opt, err := rank.ParseArgs(args) opt, err := rank.ParseArgs(args)
if err != nil { if err != nil {
if errors.Is(err, cli.ErrHelp) {
return 0
}
fmt.Fprintf(os.Stderr, "brain/search: %v\n%s\n", err, rank.Usage) fmt.Fprintf(os.Stderr, "brain/search: %v\n%s\n", err, rank.Usage)
return 2 return 2
} }
@@ -212,7 +216,11 @@ func rowsToHits(res *lbug.QueryResult) ([]Hit, error) {
root := fmt.Sprint(vals[2]) root := fmt.Sprint(vals[2])
source := fmt.Sprint(vals[3]) source := fmt.Sprint(vals[3])
score := float64(vals[4].(float64)) score := float64(vals[4].(float64))
hits = append(hits, Hit{ID: id, Text: text, Root: root, Source: source, Score: score}) conf := ""
if len(vals) >= 6 {
conf = fmt.Sprint(vals[5])
}
hits = append(hits, Hit{ID: id, Text: text, Root: root, Source: source, Score: score, Confidence: conf})
} }
return hits, nil return hits, nil
} }
@@ -230,6 +238,7 @@ type jsonHit struct {
ID string `json:"id"` ID string `json:"id"`
Text string `json:"text"` Text string `json:"text"`
Root string `json:"root"` Root string `json:"root"`
Confidence string `json:"confidence,omitempty"`
Score float64 `json:"score"` Score float64 `json:"score"`
Snippet string `json:"snippet,omitempty"` Snippet string `json:"snippet,omitempty"`
Hops []rank.HopNode `json:"hops,omitempty"` Hops []rank.HopNode `json:"hops,omitempty"`
@@ -242,6 +251,7 @@ func toJSONOut(hits []Hit, query, rootFilter string, web *rank.SecondSource) *js
ID: h.ID, ID: h.ID,
Text: h.Text, Text: h.Text,
Root: h.Root, Root: h.Root,
Confidence: h.Confidence,
Score: h.Score, Score: h.Score,
Snippet: h.Snippet, Snippet: h.Snippet,
Hops: h.Hops, Hops: h.Hops,
@@ -265,6 +275,9 @@ func resultsToDicts(hits []Hit) []any {
{"root", h.Root}, {"root", h.Root},
{"score", h.Score}, {"score", h.Score},
} }
if h.Confidence != "" {
d = append(d, KV{"confidence", h.Confidence})
}
if h.Snippet != "" { if h.Snippet != "" {
d = append(d, KV{"snippet", h.Snippet}) d = append(d, KV{"snippet", h.Snippet})
} }
+6 -12
View File
@@ -3,12 +3,13 @@ package chats
import ( import (
"bytes" "bytes"
"encoding/json" "encoding/json"
"flag"
"fmt" "fmt"
"os" "os"
"os/exec" "os/exec"
"path/filepath" "path/filepath"
"strings" "strings"
cliparse "github.com/eSlider/2dph/internal/cli"
) )
type ooContact struct { type ooContact struct {
@@ -25,16 +26,9 @@ type ooContact struct {
} }
func RunApply(args []string) int { func RunApply(args []string) int {
fs := flag.NewFlagSet("chats apply", flag.ContinueOnError) dryRun, err := parseApplyFlags(args)
dryRun := fs.Bool("dry-run", false, "show what would be done without writing") if err != nil {
help := fs.Bool("help", false, "") return cliparse.Fail(err)
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats apply [--dry-run]")
return 0
} }
ooCLI := findOO() ooCLI := findOO()
@@ -128,7 +122,7 @@ func RunApply(args []string) int {
fmt.Printf("\nchats apply: %d actions to apply\n", len(resolved)) fmt.Printf("\nchats apply: %d actions to apply\n", len(resolved))
if *dryRun { if dryRun {
for _, r := range resolved { for _, r := range resolved {
switch r.Action { switch r.Action {
case "info-add": case "info-add":
+74
View File
@@ -0,0 +1,74 @@
package chats
import (
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
)
type syncTelegramFlags struct {
Limit int
Phone string
}
type syncLinkedInFlags struct {
Limit int
Refresh bool
}
func SyncParser() *flaggy.Parser {
p := cliparse.New("chats-sync")
p.Description = "download chats to var/chats"
tg := flaggy.NewSubcommand("telegram")
li := flaggy.NewSubcommand("linkedin")
var limit int
var phone string
var refresh bool
tg.Int(&limit, "", "limit", "max messages per chat")
tg.String(&phone, "", "phone", "phone (default TELEGRAM_PHONE)")
li.Int(&limit, "", "limit", "max messages per conversation")
li.Bool(&refresh, "", "refresh", "refresh webtop session")
p.AttachSubcommand(tg, 1)
p.AttachSubcommand(li, 1)
return p
}
func ImportParser() *flaggy.Parser {
return cliparse.New("chats-import")
}
func FactsParser() *flaggy.Parser {
return cliparse.New("chats-facts")
}
func ApplyParser() *flaggy.Parser {
p := cliparse.New("chats-apply")
dry := false
p.Bool(&dry, "", "dry-run", "show without writing")
return p
}
func parseTelegramFlags(args []string) (syncTelegramFlags, error) {
var f syncTelegramFlags
p := cliparse.New("chats-sync-telegram")
p.Int(&f.Limit, "", "limit", "max messages per chat")
p.String(&f.Phone, "", "phone", "phone (default TELEGRAM_PHONE)")
return f, cliparse.Parse(p, args)
}
func parseLinkedInFlags(args []string) (syncLinkedInFlags, error) {
var f syncLinkedInFlags
p := cliparse.New("chats-sync-linkedin")
p.Int(&f.Limit, "", "limit", "max messages per conversation")
p.Bool(&f.Refresh, "", "refresh", "refresh webtop session")
return f, cliparse.Parse(p, args)
}
func parseApplyFlags(args []string) (dryRun bool, err error) {
p := cliparse.New("chats-apply")
p.Bool(&dryRun, "", "dry-run", "show without writing")
return dryRun, cliparse.Parse(p, args)
}
func parseNoFlags(name string, args []string) error {
return cliparse.Parse(cliparse.New(name), args)
}
+4 -10
View File
@@ -3,12 +3,13 @@ package chats
import ( import (
"bufio" "bufio"
"encoding/json" "encoding/json"
"flag"
"fmt" "fmt"
"os" "os"
"path/filepath" "path/filepath"
"regexp" "regexp"
"strings" "strings"
cliparse "github.com/eSlider/2dph/internal/cli"
) )
var ( var (
@@ -76,15 +77,8 @@ type ExtractedFact struct {
} }
func RunFacts(args []string) int { func RunFacts(args []string) int {
fs := flag.NewFlagSet("chats facts", flag.ContinueOnError) if err := parseNoFlags("chats-facts", args); err != nil {
help := fs.Bool("help", false, "") return cliparse.Fail(err)
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats facts")
return 0
} }
root := Dir() root := Dir()
+4 -10
View File
@@ -4,25 +4,19 @@ import (
"bufio" "bufio"
"bytes" "bytes"
"encoding/json" "encoding/json"
"flag"
"fmt" "fmt"
"html" "html"
"os" "os"
"path/filepath" "path/filepath"
"sort" "sort"
"strings" "strings"
cliparse "github.com/eSlider/2dph/internal/cli"
) )
func RunImport(args []string) int { func RunImport(args []string) int {
fs := flag.NewFlagSet("chats import", flag.ContinueOnError) if err := parseNoFlags("chats-import", args); err != nil {
help := fs.Bool("help", false, "") return cliparse.Fail(err)
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats import")
return 0
} }
root := Dir() root := Dir()
+9 -14
View File
@@ -2,12 +2,13 @@ package chats
import ( import (
"context" "context"
"flag"
"fmt" "fmt"
"os" "os"
"os/exec" "os/exec"
"path/filepath" "path/filepath"
"time" "time"
cliparse "github.com/eSlider/2dph/internal/cli"
) )
func checkLinkedInSession(userDataDir string) (bool, error) { func checkLinkedInSession(userDataDir string) (bool, error) {
@@ -29,18 +30,12 @@ func checkLinkedInSession(userDataDir string) (bool, error) {
} }
func RunSyncLinkedIn(args []string) int { func RunSyncLinkedIn(args []string) int {
fs := flag.NewFlagSet("chats sync linkedin", flag.ContinueOnError) f, err := parseLinkedInFlags(args)
limit := fs.Int("limit", 0, "max messages per conversation (0 = all)") if err != nil {
refresh := fs.Bool("refresh", false, "refresh session from live webtop browser before sync") return cliparse.Fail(err)
help := fs.Bool("help", false, "")
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats sync linkedin [--limit N] [--refresh]")
return 0
} }
limit := f.Limit
refresh := f.Refresh
userDataDir := envVar("LINKEDIN_USER_DATA_DIR", "") userDataDir := envVar("LINKEDIN_USER_DATA_DIR", "")
if userDataDir == "" { if userDataDir == "" {
@@ -48,7 +43,7 @@ func RunSyncLinkedIn(args []string) int {
userDataDir = home + "/.linkedin-mcp/profile" userDataDir = home + "/.linkedin-mcp/profile"
} }
if *refresh { if refresh {
if code := refreshLinkedInSession(userDataDir); code != 0 { if code := refreshLinkedInSession(userDataDir); code != 0 {
return code return code
} }
@@ -72,7 +67,7 @@ func RunSyncLinkedIn(args []string) int {
defer cancel() defer cancel()
start := time.Now() start := time.Now()
if err := src.Sync(ctx, Dir(), *limit); err != nil { if err := src.Sync(ctx, Dir(), limit); err != nil {
fmt.Fprintf(os.Stderr, "chats sync linkedin: %v\n", err) fmt.Fprintf(os.Stderr, "chats sync linkedin: %v\n", err)
return 1 return 1
} }
+9 -14
View File
@@ -2,33 +2,28 @@ package chats
import ( import (
"context" "context"
"flag"
"fmt" "fmt"
"os" "os"
"path/filepath" "path/filepath"
"strconv" "strconv"
"strings" "strings"
"time" "time"
cliparse "github.com/eSlider/2dph/internal/cli"
) )
func RunSyncTelegram(args []string) int { func RunSyncTelegram(args []string) int {
fs := flag.NewFlagSet("chats sync telegram", flag.ContinueOnError) f, err := parseTelegramFlags(args)
limit := fs.Int("limit", 0, "max messages per chat (0 = all)") if err != nil {
phone := fs.String("phone", "", "phone number (default env TELEGRAM_PHONE)") return cliparse.Fail(err)
help := fs.Bool("help", false, "")
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats sync telegram [--limit N] [--phone PHONE]")
return 0
} }
limit := f.Limit
phone := f.Phone
apiIDStr := envVar("TELEGRAM_API_ID", "") apiIDStr := envVar("TELEGRAM_API_ID", "")
apiHash := envVar("TELEGRAM_API_HASH", "") apiHash := envVar("TELEGRAM_API_HASH", "")
sessionStr := envVar("TELEGRAM_SESSION_STRING", "") sessionStr := envVar("TELEGRAM_SESSION_STRING", "")
phoneNum := *phone phoneNum := phone
if phoneNum == "" { if phoneNum == "" {
phoneNum = envVar("TELEGRAM_PHONE", "") phoneNum = envVar("TELEGRAM_PHONE", "")
} }
@@ -78,7 +73,7 @@ func RunSyncTelegram(args []string) int {
defer cancel() defer cancel()
start := time.Now() start := time.Now()
if err := src.Sync(ctx, Dir(), *limit); err != nil { if err := src.Sync(ctx, Dir(), limit); err != nil {
fmt.Fprintf(os.Stderr, "chats sync telegram: %v\n", err) fmt.Fprintf(os.Stderr, "chats sync telegram: %v\n", err)
return 1 return 1
} }
+190
View File
@@ -0,0 +1,190 @@
// Package cli is the shared flaggy wrapper (D23).
//
// flaggy: zero deps, flags at any position, shell completion scripts.
// Individual tools keep ShowCompletion off so a query like "completion" is
// not stolen; dump scripts with bin/cli/complete.go.
package cli
import (
"errors"
"fmt"
"os"
"strings"
"sync"
"github.com/integrii/flaggy"
)
// ErrHelp means -h/--help was requested (exit 0).
var ErrHelp = errors.New("help")
var parseMu sync.Mutex
// New returns a per-call parser. Never reuse: flaggy parses once.
func New(name string) *flaggy.Parser {
p := flaggy.NewParser(name)
p.ShowVersionWithVersionFlag = false
p.ShowCompletion = false
// Extra positionals become TrailingArguments (search "two words --json").
// Unknown dash tokens are rejected in Parse after flaggy returns.
p.ShowHelpOnUnexpected = false
p.ShowHelpWithHFlag = true
return p
}
// Parse runs p.ParseArgs and turns flaggy's os.Exit into an error.
// Not safe to call in parallel (flaggy.PanicInsteadOfExit is process-global).
func Parse(p *flaggy.Parser, args []string) error {
parseMu.Lock()
defer parseMu.Unlock()
prev := flaggy.PanicInsteadOfExit
flaggy.PanicInsteadOfExit = true
defer func() { flaggy.PanicInsteadOfExit = prev }()
var exitMsg string
err := func() error {
defer func() {
if r := recover(); r != nil {
exitMsg = fmt.Sprint(r)
}
}()
return p.ParseArgs(args)
}()
if err != nil {
return err
}
if exitMsg != "" {
if strings.Contains(exitMsg, "code: 0") {
return ErrHelp
}
return errors.New(exitMsg)
}
if u := unknownFlags(p, args); len(u) > 0 {
return fmt.Errorf("unknown flag %q", u[0])
}
return nil
}
func unknownFlags(p *flaggy.Parser, args []string) []string {
flags := collectFlags(&p.Subcommand)
var out []string
skipNext := false
for _, a := range args {
if skipNext {
skipNext = false
continue
}
if a == "--" {
break
}
name, inline := flagName(a)
if name == "" {
continue
}
if name == "h" || name == "help" {
continue
}
f := findFlag(flags, name)
if f == nil {
out = append(out, a)
continue
}
if !inline && !isBoolFlag(f) {
skipNext = true
}
}
return out
}
func flagName(a string) (name string, inline bool) {
if a == "-" || !strings.HasPrefix(a, "-") {
return "", false
}
rest := strings.TrimLeft(a, "-")
name, _, inline = strings.Cut(rest, "=")
return name, inline
}
func collectFlags(sc *flaggy.Subcommand) []*flaggy.Flag {
out := append([]*flaggy.Flag{}, sc.Flags...)
for _, sub := range sc.Subcommands {
out = append(out, collectFlags(sub)...)
}
return out
}
func findFlag(flags []*flaggy.Flag, name string) *flaggy.Flag {
for _, f := range flags {
if f.HasName(name) {
return f
}
}
return nil
}
func isBoolFlag(f *flaggy.Flag) bool {
_, ok := f.AssignmentVar.(*bool)
return ok
}
// Query joins the first positional with leftover trailing words.
func Query(first string, trailing []string) string {
parts := make([]string, 0, 1+len(trailing))
if s := strings.TrimSpace(first); s != "" {
parts = append(parts, s)
}
for _, t := range trailing {
if s := strings.TrimSpace(t); s != "" {
parts = append(parts, s)
}
}
return strings.Join(parts, " ")
}
// Code maps parse errors to process exit codes (0 help, 2 usage).
func Code(err error) int {
if err == nil || errors.Is(err, ErrHelp) {
return 0
}
return 2
}
// Fail prints err unless it is help or a flaggy exit that already wrote stderr.
func Fail(err error) int {
if err == nil || errors.Is(err, ErrHelp) {
return 0
}
if strings.HasPrefix(err.Error(), "Panic instead of exit") {
return 2
}
fmt.Fprintln(os.Stderr, err)
return 2
}
// Tool is one shebang CLI for completion dump.
type Tool struct {
Path string
Name string
New func() *flaggy.Parser
}
// BashScript concatenates flaggy bash complete scripts and binds each
// function to the shebang path (./bin/subject/method.go).
func BashScript(tools []Tool) string {
var b strings.Builder
b.WriteString("# 2dph flaggy completions (D23). source <(./bin/cli/complete.go bash)\n")
for _, t := range tools {
p := t.New()
p.Name = t.Name
script := flaggy.GenerateBashCompletion(p)
b.WriteString(script)
fn := "_" + strings.ReplaceAll(t.Name, "-", "_") + "_complete"
if t.Path != "" && t.Path != t.Name {
fmt.Fprintf(&b, "complete -F %s %s\n", fn, t.Path)
if !strings.HasPrefix(t.Path, "./") {
fmt.Fprintf(&b, "complete -F %s ./%s\n", fn, t.Path)
}
}
}
return b.String()
}
+77
View File
@@ -0,0 +1,77 @@
package cli
import (
"errors"
"strings"
"testing"
"github.com/integrii/flaggy"
)
func TestParseBoolAndIntAnyPosition(t *testing.T) {
p := New("t")
jsonOut := false
n := 20
q := ""
p.Bool(&jsonOut, "", "json", "JSON")
p.Int(&n, "n", "n", "limit")
p.AddPositionalValue(&q, "query", 1, false, "q")
if err := Parse(p, []string{"two", "words", "--json", "-n", "5"}); err != nil {
t.Fatal(err)
}
got := Query(q, p.TrailingArguments)
if got != "two words" || !jsonOut || n != 5 {
t.Fatalf("q=%q json=%v n=%d", got, jsonOut, n)
}
}
func TestParseUnknownFlagIsError(t *testing.T) {
p := New("t")
jsonOut := false
p.Bool(&jsonOut, "", "json", "JSON")
if err := Parse(p, []string{"--nope"}); err == nil {
t.Fatal("unknown flag accepted")
}
}
func TestParseHelpIsErrHelp(t *testing.T) {
p := New("t")
jsonOut := false
p.Bool(&jsonOut, "", "json", "JSON")
err := Parse(p, []string{"--help"})
if !errors.Is(err, ErrHelp) {
t.Fatalf("got %v", err)
}
}
func TestParseMissingFlagValueIsError(t *testing.T) {
p := New("t")
n := 0
p.Int(&n, "", "hop", "hop")
if err := Parse(p, []string{"--hop"}); err == nil {
t.Fatal("expected missing value error")
}
}
func TestBashScriptNamesShebangPath(t *testing.T) {
out := BashScript([]Tool{{
Path: "bin/brain/search.go",
Name: "brain-search",
New: newSearchLike,
}})
if !strings.Contains(out, "--json") || !strings.Contains(out, "--hop") {
t.Fatalf("flags missing:\n%s", out)
}
if !strings.Contains(out, "complete -F") || !strings.Contains(out, "bin/brain/search.go") {
t.Fatalf("shebang complete missing:\n%s", out)
}
}
func newSearchLike() *flaggy.Parser {
p := New("brain-search")
jsonOut := false
hop := 0
p.Bool(&jsonOut, "", "json", "JSON")
p.Int(&hop, "", "hop", "graph hop")
return p
}
+24
View File
@@ -0,0 +1,24 @@
package cli
import "github.com/integrii/flaggy"
type QAStats struct {
JSONL string
}
func QAParser() *flaggy.Parser {
c := QAStats{}
return BindQA(&c)
}
func BindQA(c *QAStats) *flaggy.Parser {
p := New("qa-stats")
p.Description = "DuckDB quantiles / JSONL count"
p.String(&c.JSONL, "", "jsonl", "JSONL file (else stdin JSON [float,…])")
return p
}
func ParseQAStats(args []string) (QAStats, error) {
var c QAStats
return c, Parse(BindQA(&c), args)
}
+47
View File
@@ -0,0 +1,47 @@
// Package duckstats runs in-process DuckDB for columnar aggregates.
// Graph facts stay in Ladybug. Web-search KV cache stays modernc sqlite.
package duckstats
import (
"database/sql"
"fmt"
_ "github.com/duckdb/duckdb-go/v2"
)
type Stats struct {
N int `json:"n"`
Min float64 `json:"min"`
P50 float64 `json:"p50"`
P95 float64 `json:"p95"`
Max float64 `json:"max"`
Avg float64 `json:"avg"`
}
func Quantiles(samples []float64) (Stats, error) {
if len(samples) == 0 {
return Stats{}, fmt.Errorf("duckstats: empty samples")
}
db, err := sql.Open("duckdb", "")
if err != nil {
return Stats{}, err
}
defer db.Close()
var s Stats
err = db.QueryRow(`
SELECT count(v), min(v), quantile_cont(v, 0.5), quantile_cont(v, 0.95), max(v), avg(v)
FROM (SELECT unnest(?) AS v)`, samples).Scan(
&s.N, &s.Min, &s.P50, &s.P95, &s.Max, &s.Avg)
return s, err
}
func CountJSONL(path string) (int64, error) {
db, err := sql.Open("duckdb", "")
if err != nil {
return 0, err
}
defer db.Close()
var n int64
err = db.QueryRow(`SELECT count(*) FROM read_json_auto(?)`, path).Scan(&n)
return n, err
}
+51
View File
@@ -0,0 +1,51 @@
package duckstats
import (
"os"
"testing"
)
func TestQuantilesEmpty(t *testing.T) {
_, err := Quantiles(nil)
if err == nil {
t.Fatal("empty slice must error")
}
}
func TestQuantilesOdd(t *testing.T) {
s, err := Quantiles([]float64{1, 2, 3, 4, 5})
if err != nil {
t.Fatal(err)
}
if s.N != 5 {
t.Fatalf("n=%d", s.N)
}
if s.Min != 1 || s.Max != 5 {
t.Fatalf("min=%v max=%v", s.Min, s.Max)
}
if s.P50 != 3 {
t.Fatalf("p50=%v want 3", s.P50)
}
if s.Avg != 3 {
t.Fatalf("avg=%v want 3", s.Avg)
}
if s.P95 < 4.5 || s.P95 > 5 {
t.Fatalf("p95=%v want in [4.5,5]", s.P95)
}
}
func TestCountJSONL(t *testing.T) {
dir := t.TempDir()
p := dir + "/rows.jsonl"
body := "{\"ms\":1}\n{\"ms\":2}\n{\"ms\":3}\n"
if err := os.WriteFile(p, []byte(body), 0o600); err != nil {
t.Fatal(err)
}
n, err := CountJSONL(p)
if err != nil {
t.Fatal(err)
}
if n != 3 {
t.Fatalf("count=%d want 3", n)
}
}
+120
View File
@@ -0,0 +1,120 @@
// Package facts is cgo-free evidence rules (D16 contradictions).
package facts
import "strconv"
const (
ConfConfirmed = "confirmed"
ConfHypothesis = "hypothesis"
RuleUnresolved = "unresolved"
RuleTemporalFreshness = "temporal_freshness"
RuleAuthorityPairing = "authority_pairing"
RuleTwoSource = "two_source"
RuleSingleSource = "single_source"
KindRuntime = "runtime"
KindConfig = "config"
KindNarrative = "narrative"
)
// Source is one independent pointer on a yes or no side.
type Source struct {
ID string `json:"id"`
Kind string `json:"kind"`
When string `json:"when,omitempty"`
Stale bool `json:"stale,omitempty"`
}
// Claim is one assertion with yes/no evidence lists.
type Claim struct {
Text string `json:"text"`
Yes []Source `json:"yes"`
No []Source `json:"no"`
}
// Result is audit output. Confirmed=false means `(not confirmed)`.
type Result struct {
Text string `json:"text"`
Confidence string `json:"confidence"`
Confirmed bool `json:"confirmed"`
Rule string `json:"rule"`
Winner string `json:"winner,omitempty"`
YesN int `json:"yes"`
NoN int `json:"no"`
}
func independent(ss []Source) int {
seen := map[string]struct{}{}
for i, s := range ss {
id := s.ID
if id == "" {
id = s.Kind + "#" + strconv.Itoa(i)
}
seen[id] = struct{}{}
}
return len(seen)
}
func freshN(ss []Source) int {
n := 0
for _, s := range ss {
if !s.Stale {
n++
}
}
return n
}
func strongN(ss []Source) int {
n := 0
for _, s := range ss {
if s.Kind == KindRuntime || s.Kind == KindConfig {
n++
}
}
return n
}
func out(c Claim, conf, rule, winner string) Result {
return Result{
Text: c.Text,
Confidence: conf,
Confirmed: conf == ConfConfirmed,
Rule: rule,
Winner: winner,
YesN: independent(c.Yes),
NoN: independent(c.No),
}
}
// Adjudicate applies D16: ≥2 yes vs ≥2 no stays hypothesis until a rule fires.
// Order: temporal_freshness, then authority_pairing (A/B beats narrative C).
func Adjudicate(c Claim) Result {
yesN := independent(c.Yes)
noN := independent(c.No)
if yesN < 2 || noN < 2 {
if yesN >= 2 {
return out(c, ConfConfirmed, RuleTwoSource, "yes")
}
if noN >= 2 {
return out(c, ConfConfirmed, RuleTwoSource, "no")
}
return out(c, ConfHypothesis, RuleSingleSource, "")
}
yf, nf := freshN(c.Yes), freshN(c.No)
if yf >= 2 && nf < 2 {
return out(c, ConfConfirmed, RuleTemporalFreshness, "yes")
}
if nf >= 2 && yf < 2 {
return out(c, ConfConfirmed, RuleTemporalFreshness, "no")
}
ys, ns := strongN(c.Yes), strongN(c.No)
if ys >= 2 && ns < 2 {
return out(c, ConfConfirmed, RuleAuthorityPairing, "yes")
}
if ns >= 2 && ys < 2 {
return out(c, ConfConfirmed, RuleAuthorityPairing, "no")
}
return out(c, ConfHypothesis, RuleUnresolved, "")
}
+86
View File
@@ -0,0 +1,86 @@
package facts
import "testing"
func src(id, kind string, stale bool) Source {
return Source{ID: id, Kind: kind, Stale: stale}
}
func TestTwoVsTwoStaysHypothesis(t *testing.T) {
c := Claim{
Text: "svc listens on 443",
Yes: []Source{
src("docker-ps", KindRuntime, false),
src("compose", KindConfig, false),
},
No: []Source{
src("docker-ps-old", KindRuntime, false),
src("compose-old", KindConfig, false),
},
}
r := Adjudicate(c)
if r.Confirmed || r.Confidence != ConfHypothesis || r.Rule != RuleUnresolved {
t.Fatalf("2v2 must stay (not confirmed): %+v", r)
}
if r.Winner != "" {
t.Fatalf("unresolved must not name a winner: %+v", r)
}
}
func TestTemporalFreshnessResolvesStaleSide(t *testing.T) {
c := Claim{
Text: "svc listens on 443",
Yes: []Source{
src("docker-ps", KindRuntime, false),
src("compose", KindConfig, false),
},
No: []Source{
src("old-readme", KindNarrative, true),
src("old-wiki", KindNarrative, true),
},
}
r := Adjudicate(c)
if !r.Confirmed || r.Rule != RuleTemporalFreshness || r.Winner != "yes" {
t.Fatalf("fresh yes vs stale no: %+v", r)
}
}
func TestAuthorityPairingBeatsNarrative(t *testing.T) {
c := Claim{
Text: "svc listens on 443",
Yes: []Source{
src("docker-ps", KindRuntime, false),
src("compose", KindConfig, false),
},
No: []Source{
src("readme", KindNarrative, false),
src("wiki", KindNarrative, false),
},
}
r := Adjudicate(c)
if !r.Confirmed || r.Rule != RuleAuthorityPairing || r.Winner != "yes" {
t.Fatalf("A×B vs C×C: %+v", r)
}
}
func TestTwoSourceYesIsConfirmed(t *testing.T) {
c := Claim{
Text: "arc-1 runs Matrix",
Yes: []Source{
src("compose", KindConfig, false),
src("docker-ps", KindRuntime, false),
},
}
r := Adjudicate(c)
if !r.Confirmed || r.Rule != RuleTwoSource || r.Winner != "yes" {
t.Fatalf("%+v", r)
}
}
func TestSingleSourceIsHypothesis(t *testing.T) {
c := Claim{Text: "maybe", Yes: []Source{src("readme", KindNarrative, false)}}
r := Adjudicate(c)
if r.Confirmed || r.Rule != RuleSingleSource {
t.Fatalf("%+v", r)
}
}
+51
View File
@@ -0,0 +1,51 @@
package gitlog
import (
"fmt"
"time"
"github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
)
type CLI struct {
Repo, Root, Since string
Limit int
JSONOut bool
}
func Parser() *flaggy.Parser {
c := CLI{}
return Bind(&c)
}
func Bind(c *CLI) *flaggy.Parser {
p := cli.New("git-import")
p.Description = "go-git history → commit leafs"
p.Bool(&c.JSONOut, "", "json", "JSON output")
p.Int(&c.Limit, "", "limit", "max commits (0 = all)")
p.String(&c.Since, "", "since", "RFC3339 or YYYY-MM-DD")
p.String(&c.Root, "", "root", "scan dir for git repos")
p.AddPositionalValue(&c.Repo, "repo", 1, false, "git repo path")
return p
}
func ParseArgs(args []string) (CLI, error) {
var c CLI
if err := cli.Parse(Bind(&c), args); err != nil {
return c, err
}
return c, nil
}
func ParseSince(s string) (time.Time, error) {
if s == "" {
return time.Time{}, nil
}
for _, layout := range []string{time.RFC3339, "2006-01-02"} {
if t, err := time.Parse(layout, s); err == nil {
return t, nil
}
}
return time.Time{}, fmt.Errorf("cannot parse --since %q", s)
}
+41
View File
@@ -0,0 +1,41 @@
package mdleaves
import (
"github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
)
type CLI struct {
Root string
Files string
JSONOut bool
}
func Parser() *flaggy.Parser {
c := CLI{Root: "."}
return Bind(&c)
}
func Bind(c *CLI) *flaggy.Parser {
if c.Root == "" {
c.Root = "."
}
p := cli.New("markdown-import")
p.Description = "split markdown H2 leafs"
p.Bool(&c.JSONOut, "", "json", "JSON output")
p.String(&c.Files, "", "files", "comma-separated paths")
p.AddPositionalValue(&c.Root, "dir", 1, false, "markdown root")
return p
}
func ParseArgs(args []string) (CLI, error) {
c := CLI{Root: "."}
p := Bind(&c)
if err := cli.Parse(p, args); err != nil {
return c, err
}
if extra := cli.Query("", p.TrailingArguments); extra != "" && c.Root == "." {
c.Root = extra
}
return c, nil
}
+35
View File
@@ -0,0 +1,35 @@
package ocr
import (
"fmt"
"github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
)
type CLI struct {
Path string
}
func Parser() *flaggy.Parser {
c := CLI{}
return Bind(&c)
}
func Bind(c *CLI) *flaggy.Parser {
p := cli.New("mail-ocr")
p.Description = "tesseract eng+deu on image or scanned PDF"
p.AddPositionalValue(&c.Path, "file", 1, false, "image or pdf")
return p
}
func ParseArgs(args []string) (CLI, error) {
var c CLI
if err := cli.Parse(Bind(&c), args); err != nil {
return c, err
}
if c.Path == "" {
return c, fmt.Errorf("usage: bin/mail/ocr.go <image|pdf>")
}
return c, nil
}
+162
View File
@@ -0,0 +1,162 @@
// Package ocr runs Tesseract (eng+deu) on images and scanned PDFs.
//
// Default engine is the tesseract CLI, not gosseract CGO: Ladybug CGO stays
// Zig-only (D21). Same engine, no gocv. OCR_ENGINE=paddle selects paddleocr
// when that binary is on PATH (compose profile ocr-paddle).
package ocr
import (
"fmt"
"image"
"image/color"
"image/png"
"os"
"os/exec"
"path/filepath"
"strings"
)
const TessLang = "eng+deu"
func ImageFile(path string) (string, error) {
engine := os.Getenv("OCR_ENGINE")
if engine == "paddle" {
return runPaddle(path)
}
return runTesseract(path)
}
func PDFFile(path string) (string, error) {
text, err := pdfToText(path)
if err == nil && strings.TrimSpace(text) != "" {
return strings.TrimSpace(text), nil
}
ocr, oerr := pdfPages(path)
if oerr != nil {
if err != nil {
return "", err
}
return "", oerr
}
if strings.TrimSpace(ocr) != "" {
return strings.TrimSpace(ocr), nil
}
if text != "" {
return strings.TrimSpace(text), nil
}
return "", fmt.Errorf("pdf has no text layer (ocr unavailable)")
}
func pdfToText(path string) (string, error) {
cmd := exec.Command("pdftotext", "-layout", path, "-")
out, err := cmd.Output()
if err != nil {
return "", err
}
return string(out), nil
}
func pdfPages(path string) (string, error) {
dir, err := os.MkdirTemp("", "2dph-ocr-")
if err != nil {
return "", err
}
defer os.RemoveAll(dir)
prefix := filepath.Join(dir, "page")
cmd := exec.Command("pdftoppm", "-png", "-r", "200", path, prefix)
if err := cmd.Run(); err != nil {
return "", err
}
matches, err := filepath.Glob(prefix + "*.png")
if err != nil {
return "", err
}
var parts []string
for _, img := range matches {
t, err := ImageFile(img)
if err != nil {
continue
}
if s := strings.TrimSpace(t); s != "" {
parts = append(parts, s)
}
}
return strings.Join(parts, "\n\n"), nil
}
func runTesseract(path string) (string, error) {
pre, err := preprocessFile(path)
if err != nil {
pre = path
} else {
defer os.Remove(pre)
}
cmd := exec.Command("tesseract", pre, "stdout", "-l", TessLang, "--psm", "6")
out, err := cmd.Output()
if err != nil {
return "", err
}
return strings.TrimSpace(string(out)), nil
}
func runPaddle(path string) (string, error) {
cmd := exec.Command("paddleocr", "ocr", "-i", path)
out, err := cmd.Output()
if err != nil {
return "", err
}
return strings.TrimSpace(string(out)), nil
}
func preprocessFile(path string) (string, error) {
f, err := os.Open(path)
if err != nil {
return "", err
}
defer f.Close()
img, err := png.Decode(f)
if err != nil {
return "", err
}
out := filepath.Join(os.TempDir(), filepath.Base(path)+".gray.png")
w, err := os.Create(out)
if err != nil {
return "", err
}
defer w.Close()
if err := png.Encode(w, GrayContrast(img)); err != nil {
os.Remove(out)
return "", err
}
return out, nil
}
// GrayContrast is a stdlib preprocess (no gocv): grayscale + stretch.
func GrayContrast(src image.Image) image.Image {
b := src.Bounds()
dst := image.NewGray(b)
var minL, maxL uint8 = 255, 0
for y := b.Min.Y; y < b.Max.Y; y++ {
for x := b.Min.X; x < b.Max.X; x++ {
g := color.GrayModel.Convert(src.At(x, y)).(color.Gray)
if g.Y < minL {
minL = g.Y
}
if g.Y > maxL {
maxL = g.Y
}
}
}
span := int(maxL) - int(minL)
if span < 1 {
span = 1
}
for y := b.Min.Y; y < b.Max.Y; y++ {
for x := b.Min.X; x < b.Max.X; x++ {
g := color.GrayModel.Convert(src.At(x, y)).(color.Gray)
v := uint8((int(g.Y) - int(minL)) * 255 / span)
dst.SetGray(x, y, color.Gray{Y: v})
}
}
return dst
}
+54
View File
@@ -0,0 +1,54 @@
package ocr
import (
"image"
"image/color"
"os/exec"
"path/filepath"
"strings"
"testing"
)
func TestGrayContrastStretches(t *testing.T) {
img := image.NewGray(image.Rect(0, 0, 2, 2))
img.SetGray(0, 0, color.Gray{Y: 64})
img.SetGray(0, 1, color.Gray{Y: 64})
img.SetGray(1, 0, color.Gray{Y: 64})
img.SetGray(1, 1, color.Gray{Y: 192})
out := GrayContrast(img).(*image.Gray)
if out.GrayAt(0, 0).Y != 0 {
t.Fatalf("min should map to 0, got %d", out.GrayAt(0, 0).Y)
}
if out.GrayAt(1, 1).Y != 255 {
t.Fatalf("max should map to 255, got %d", out.GrayAt(1, 1).Y)
}
}
func TestHelloPNGFixtureOCR(t *testing.T) {
if _, err := exec.LookPath("tesseract"); err != nil {
t.Skip("tesseract not installed")
}
path := filepath.Join("testdata", "hello.png")
got, err := ImageFile(path)
if err != nil {
t.Fatal(err)
}
up := strings.ToUpper(got)
if !strings.Contains(up, "HELLO") {
t.Fatalf("ocr %q missing HELLO", got)
}
}
func TestPaddleEngineUsesPaddleocrBinary(t *testing.T) {
t.Setenv("OCR_ENGINE", "paddle")
_, err := ImageFile(filepath.Join("testdata", "hello.png"))
if _, look := exec.LookPath("paddleocr"); look != nil {
if err == nil {
t.Fatal("expected error when paddleocr is missing")
}
return
}
if err != nil {
t.Fatal(err)
}
}
BIN
View File
Binary file not shown.

After

Width:  |  Height:  |  Size: 1.7 KiB

+47
View File
@@ -0,0 +1,47 @@
package reasoner
import (
"os"
"github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
)
type CLI struct {
Base string
Model string
Device string
JSONOut bool
}
func Parser() *flaggy.Parser {
c := NewCLI()
return Bind(&c)
}
func NewCLI() CLI {
base := os.Getenv("REASONER_BASE_URL")
if base == "" {
base = "http://127.0.0.1:11435/v1"
}
model := os.Getenv("REASONER_MODEL")
if model == "" {
model = OllamaRAM
}
return CLI{Base: base, Model: model, Device: "cpu"}
}
func Bind(c *CLI) *flaggy.Parser {
p := cli.New("reasoner-bakeoff")
p.Description = "CPU tool-call bake-off"
p.Bool(&c.JSONOut, "", "json", "JSON output")
p.String(&c.Model, "", "model", "Ollama/HF model id")
p.String(&c.Base, "", "base-url", "OpenAI-compatible URL")
p.String(&c.Device, "", "device", "cpu")
return p
}
func ParseArgs(args []string) (CLI, error) {
c := NewCLI()
return c, cli.Parse(Bind(&c), args)
}
+2
View File
@@ -113,6 +113,8 @@ type Report struct {
XMLLeak int `json:"xml_leak"` XMLLeak int `json:"xml_leak"`
RSSMB int `json:"rss_mb"` RSSMB int `json:"rss_mb"`
VRAMMB int `json:"vram_mb"` VRAMMB int `json:"vram_mb"`
LatencyP50MS float64 `json:"latency_p50_ms,omitempty"`
LatencyP95MS float64 `json:"latency_p95_ms,omitempty"`
Prompts []Result `json:"prompts"` Prompts []Result `json:"prompts"`
} }
+63
View File
@@ -0,0 +1,63 @@
package websearch
import (
"fmt"
"github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
)
type CLI struct {
Query, Site, Lang, Fresh, Category, Engines string
Limit int
JSONOut, Refresh, Force bool
TTL float64
Timeout int
}
func NewCLI() CLI {
return CLI{Limit: DefaultLimit, TTL: float64(CacheTTL), Timeout: 25}
}
func Parser() *flaggy.Parser {
c := NewCLI()
return Bind(&c)
}
func Bind(c *CLI) *flaggy.Parser {
p := cli.New("web-search")
p.Description = "SearXNG second source (throttled ≠ absence)"
p.Bool(&c.JSONOut, "", "json", "JSON output")
p.Bool(&c.Refresh, "", "refresh", "bypass cache")
p.Bool(&c.Force, "", "force", "allow PII in query")
p.Int(&c.Limit, "n", "limit", "max hits")
p.String(&c.Site, "", "site", "restrict to host")
p.String(&c.Lang, "", "lang", "language")
p.String(&c.Fresh, "", "fresh", "day|week|month|year")
p.String(&c.Category, "", "category", "searx category")
p.String(&c.Engines, "", "engines", "engine list")
p.Float64(&c.TTL, "", "ttl", "cache ttl seconds")
p.Int(&c.Timeout, "", "timeout", "http timeout seconds")
return p
}
func ParseArgs(args []string) (CLI, error) {
c := NewCLI()
p := Bind(&c)
var q string
p.AddPositionalValue(&q, "query", 1, false, "search query")
if err := cli.Parse(p, args); err != nil {
return c, err
}
c.Query = cli.Query(q, p.TrailingArguments)
if c.Query == "" {
return c, fmt.Errorf("query required")
}
if c.Limit < 0 {
return c, fmt.Errorf("--limit must be a non-negative integer")
}
if c.Timeout <= 0 {
return c, fmt.Errorf("--timeout must be a positive integer")
}
return c, nil
}
-1
View File
@@ -6,7 +6,6 @@ readme = "README.md"
requires-python = ">=3.12" requires-python = ">=3.12"
license = { text = "MIT" } license = { text = "MIT" }
dependencies = [ dependencies = [
"docling>=2.119.0",
"ladybug==0.19.1", "ladybug==0.19.1",
"markitdown[docx,epub,html,image-exif,pdf,pptx,xlsx,zip]>=0.1.7", "markitdown[docx,epub,html,image-exif,pdf,pptx,xlsx,zip]>=0.1.7",
"mistune==3.3.4", "mistune==3.3.4",
+257
View File
@@ -0,0 +1,257 @@
#!/usr/bin/env python3
"""System performance test: PicoClaw surface (brain MCP) + optional reasoner.
BRAIN_URL=http://127.0.0.1:8630 ./qa/system_perf.py --json
REASONER_BASE_URL=http://127.0.0.1:11435/v1 REASONER_MODEL=qwen3.5:9b \\
./qa/system_perf.py --reasoner --picoclaw --json
Does not write Ladybug. Search includes web (D17); expect ~10s+ per search.
Exit 1 if health/get/audit gates fail. Reasoner is measured, not gated.
"""
from __future__ import annotations
import argparse
import json
import os
import statistics
import sys
import time
import urllib.error
import urllib.request
from concurrent.futures import ThreadPoolExecutor
DEFAULT_BRAIN = "http://127.0.0.1:8630"
DEFAULT_REASONER = "http://127.0.0.1:11435/v1"
DEFAULT_MODEL = "qwen3.5:9b"
DEFAULT_PICOCLAW = "http://127.0.0.1:18790"
GATE_HEALTH_MS = 500
GATE_GET_P50_MS = 50
GATE_AUDIT_P50_MS = 50
def _req(url: str, data: bytes | None = None, timeout: float = 90) -> bytes:
headers = {"Content-Type": "application/json"} if data is not None else {}
req = urllib.request.Request(url, data=data, headers=headers)
with urllib.request.urlopen(req, timeout=timeout) as res:
return res.read()
def timed(fn):
t0 = time.perf_counter()
out = fn()
return (time.perf_counter() - t0) * 1000.0, out
def stats(samples: list[float]) -> dict:
s = sorted(samples)
n = len(s)
return {
"n": n,
"min_ms": round(s[0], 1),
"p50_ms": round(s[n // 2], 1),
"p95_ms": round(s[min(n - 1, int(n * 0.95))], 1),
"max_ms": round(s[-1], 1),
"avg_ms": round(statistics.mean(s), 1),
}
def mcp(brain: str, method: str, params=None, timeout: float = 90) -> dict:
payload: dict = {"jsonrpc": "2.0", "id": 1, "method": method}
if params is not None:
payload["params"] = params
raw = _req(brain.rstrip("/") + "/mcp", json.dumps(payload).encode(), timeout=timeout)
return json.loads(raw.decode())
def mcp_call(brain: str, name: str, arguments: dict, timeout: float = 90) -> tuple[bool, str]:
d = mcp(brain, "tools/call", {"name": name, "arguments": arguments}, timeout=timeout)
res = d.get("result") or {}
text = ((res.get("content") or [{}])[0].get("text") or "")
return (not res.get("isError")), text
def reasoner_tool_call(base: str, model: str, user: str) -> str:
payload = {
"model": model,
"messages": [
{"role": "system", "content": "You are PicoClaw. Always call search before answering."},
{"role": "user", "content": user},
],
"tools": [
{
"type": "function",
"function": {
"name": "search",
"description": "deduction search",
"parameters": {
"type": "object",
"properties": {"q": {"type": "string"}},
"required": ["q"],
},
},
}
],
"tool_choice": "required",
}
raw = _req(
base.rstrip("/") + "/chat/completions",
json.dumps(payload).encode(),
timeout=600,
)
chat = json.loads(raw.decode())
tcs = chat["choices"][0]["message"].get("tool_calls") or []
if not tcs:
return ""
return tcs[0]["function"]["name"]
def run(args: argparse.Namespace) -> dict:
brain = args.brain.rstrip("/")
report: dict = {
"brain": brain,
"device": "cpu",
"ok": True,
"gates": {},
"mcp": {},
}
ms, _ = timed(lambda: _req(brain + "/health", timeout=5))
report["mcp"]["health"] = {"n": 1, "avg_ms": round(ms, 1)}
report["gates"]["health"] = ms <= GATE_HEALTH_MS
if ms > GATE_HEALTH_MS:
report["ok"] = False
list_ms = []
for _ in range(args.n):
ms, d = timed(lambda: mcp(brain, "tools/list", timeout=10))
names = [t["name"] for t in ((d.get("result") or {}).get("tools") or [])]
if "search" not in names:
report["ok"] = False
list_ms.append(ms)
report["mcp"]["tools_list"] = stats(list_ms)
audit_ms = []
for _ in range(args.n):
ms, (ok, _) = timed(lambda: mcp_call(brain, "audit", {}))
if not ok:
report["ok"] = False
audit_ms.append(ms)
report["mcp"]["audit"] = stats(audit_ms)
report["gates"]["audit_p50"] = report["mcp"]["audit"]["p50_ms"] <= GATE_AUDIT_P50_MS
if not report["gates"]["audit_p50"]:
report["ok"] = False
ok, text = mcp_call(brain, "search", {"q": "LadybugDB", "n": 2}, timeout=90)
inner = json.loads(text) if ok else {}
hits = inner.get("results") or []
leaf_id = hits[0]["id"] if hits else ""
report["mcp"]["search_seed"] = {
"ok": ok,
"count": inner.get("count"),
"web": (inner.get("web") or {}).get("status"),
}
get_ms = []
if leaf_id:
for _ in range(args.n):
ms, (ok, _) = timed(lambda: mcp_call(brain, "get", {"id": leaf_id, "body": True}))
if not ok:
report["ok"] = False
get_ms.append(ms)
report["mcp"]["get"] = stats(get_ms)
report["gates"]["get_p50"] = report["mcp"]["get"]["p50_ms"] <= GATE_GET_P50_MS
if not report["gates"]["get_p50"]:
report["ok"] = False
def one_get() -> float:
t0 = time.perf_counter()
mcp_call(brain, "get", {"id": leaf_id, "body": True})
return (time.perf_counter() - t0) * 1000.0
t0 = time.perf_counter()
with ThreadPoolExecutor(max_workers=8) as ex:
conc = list(ex.map(lambda _: one_get(), range(8)))
wall = (time.perf_counter() - t0) * 1000.0
report["mcp"]["get_concurrent_8"] = {**stats(conc), "wall_ms": round(wall, 1)}
search_ms = []
for q in ("LadybugDB", "model2vec"):
ms, (ok, text) = timed(lambda q=q: mcp_call(brain, "search", {"q": q, "n": 3}, timeout=90))
inner = json.loads(text) if ok else {}
search_ms.append(ms)
report.setdefault("mcp", {}).setdefault("search_samples", []).append(
{
"q": q,
"ms": round(ms, 1),
"ok": ok,
"count": inner.get("count"),
"web": (inner.get("web") or {}).get("status"),
}
)
if search_ms:
report["mcp"]["search"] = stats(search_ms)
if args.reasoner:
base = args.reasoner_url
model = args.model
report["reasoner"] = {"base_url": base, "model": model, "calls": []}
for user in (
"Use tools. Search the 2dph brain for LadybugDB. Call search.",
"Use tools. Search the 2dph brain for model2vec. Call search.",
):
ms, name = timed(lambda user=user: reasoner_tool_call(base, model, user))
report["reasoner"]["calls"].append({"ms": round(ms, 1), "tool": name})
tools = [c["tool"] for c in report["reasoner"]["calls"]]
report["gates"]["reasoner_tool_call"] = bool(tools) and all(t == "search" for t in tools)
if not report["gates"]["reasoner_tool_call"]:
report["ok"] = False
if args.picoclaw:
gw = args.picoclaw_url.rstrip("/")
ms, raw = timed(lambda: _req(gw + "/health", timeout=5))
body = json.loads(raw.decode())
report["picoclaw"] = {
"url": gw,
"health_ms": round(ms, 1),
"status": body.get("status"),
}
report["gates"]["picoclaw_health"] = body.get("status") == "ok" and ms <= GATE_HEALTH_MS
if not report["gates"]["picoclaw_health"]:
report["ok"] = False
return report
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description="2dph system performance (MCP + optional reasoner)")
p.add_argument("--brain", default=os.environ.get("BRAIN_URL", DEFAULT_BRAIN))
p.add_argument("--n", type=int, default=20)
p.add_argument("--json", action="store_true")
p.add_argument("--reasoner", action="store_true")
p.add_argument("--picoclaw", action="store_true")
p.add_argument("--picoclaw-url", default=os.environ.get("PICOCLAW_URL", DEFAULT_PICOCLAW))
p.add_argument("--reasoner-url", default=os.environ.get("REASONER_BASE_URL", DEFAULT_REASONER))
p.add_argument("--model", default=os.environ.get("REASONER_MODEL", DEFAULT_MODEL))
args = p.parse_args(argv)
try:
report = run(args)
except (urllib.error.URLError, TimeoutError, OSError) as e:
print(f"system_perf: {e}", file=sys.stderr)
return 1
if args.json:
print(json.dumps(report, indent=2))
else:
print(f"ok={report['ok']} brain={report['brain']}")
for name, block in report.get("mcp", {}).items():
if isinstance(block, dict) and "p50_ms" in block:
print(f" {name}: p50={block['p50_ms']} p95={block['p95_ms']} n={block['n']}")
elif name == "health":
print(f" health: {block.get('avg_ms')} ms")
for k, v in report.get("gates", {}).items():
print(f" gate {k}: {v}")
for c in (report.get("reasoner") or {}).get("calls") or []:
print(f" reasoner {c['tool']}: {c['ms']} ms")
return 0 if report["ok"] else 1
if __name__ == "__main__":
raise SystemExit(main(sys.argv[1:]))
+2
View File
@@ -45,6 +45,8 @@ bin/brain/eval.go # recall@5 >= 0.95 gate (
are not evidence of absence. `--root facts|info` and `--no-web` skip the web. are not evidence of absence. `--root facts|info` and `--no-web` skip the web.
- If recall looks wrong, run `bin/brain/eval.go`; it gates control questions and - If recall looks wrong, run `bin/brain/eval.go`; it gates control questions and
should stay at or above 95% recall@5. should stay at or above 95% recall@5.
- Contradictions (≥2 yes vs ≥2 no) stay `(not confirmed)` until
`bin/facts/audit contradict` fires `temporal_freshness` or `authority_pairing`.
- Agents: `GET /openapi.json` and `POST /mcp` on `bin/brain/serve.go` (same - Agents: `GET /openapi.json` and `POST /mcp` on `bin/brain/serve.go` (same
handlers; tool names match paths `search`/`get`/`stats`/`audit`). Generated handlers; tool names match paths `search`/`get`/`stats`/`audit`). Generated
list: [tools.md](tools.md). list: [tools.md](tools.md).
+38
View File
@@ -0,0 +1,38 @@
---
name: duckdb
description: >-
Use https://github.com/duckdb/duckdb-go in-process for columnar analytics
(quantiles, GROUP BY, JSON/CSV/Parquet/JSONL scans) when that is faster than
nested Go loops. Not Ladybug. Not the web-search sqlite cache. Use when
aggregating samples, counting JSONL, or SQL over tabular files.
---
# duckdb-go
Use https://github.com/duckdb/duckdb-go where it makes sense to get better performance in code.
In-process DuckDB (`internal/duckstats`, `database/sql` driver `duckdb`).
Vectorized SQL over tables, JSONL, CSV, Parquet. CGO with bundled libs
(linux/darwin amd64/arm64). Links with **gcc/g++** (libstdc++), not Zig.
D21 Zig (`bin/cgo/zcc`) is Ladybug/tokenizers only. After
`eval "$(bin/cgo/zig env)"`:
```bash
CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= ./bin/qa/stats.go <<< '[1,2,3,4,5]'
CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= go test ./internal/duckstats
```
| Store | Job |
|-------|-----|
| Ladybug | graph + FTS + HNSW (facts/info) |
| modernc sqlite | web-search KV cache + throttle |
| duckdb-go | OLAP: quantiles, counts, scans of many rows/files |
| mikefarah/yq | small YAML/JSON/XML/CSV/TOML/HCL slice, not bulk |
```bash
./bin/qa/stats.go <<< '[1,2,3,4,5]'
./bin/qa/stats.go --jsonl path/to/rows.jsonl
```
Do not open Ladybug through DuckDB. Do not put secrets or client PII into
DuckDB files under the repo.
+5 -5
View File
@@ -1,15 +1,15 @@
--- ---
name: picoclaw name: picoclaw
description: >- description: >-
2dph is the memory/fact gate, not the agent loop. Use when wiring PicoClaw 2dph is the memory/fact gate. Compose runs the official PicoClaw gateway.
or any MCP client: call brain search/get/audit before a factual reply. Use when wiring PicoClaw or any MCP client: call brain search/get/audit
throttled is not a negative finding. before a factual reply. throttled is not a negative finding.
--- ---
# PicoClaw — fact-check before assert # PicoClaw — fact-check before assert
PicoClaw (or any agent) speaks MCP at `POST /mcp` on `bin/brain/serve.go`. PicoClaw speaks MCP at `POST /mcp` on `bin/brain/serve.go`. Compose profile
2dph does not run the agent loop. Compose: `docker compose --profile picoclaw up brain-mcp` `picoclaw` runs the official `sipeed/picoclaw` gateway plus `brain-mcp`
(see [docs/picoclaw.md](../../docs/picoclaw.md)). (see [docs/picoclaw.md](../../docs/picoclaw.md)).
## Tool order (before a factual reply) ## Tool order (before a factual reply)
+1 -1
View File
@@ -9,7 +9,7 @@ description: >-
# postgres # postgres
`bin/postgres/query.go` wraps vendored `bin/db/psql-yq`. Output is YAML `bin/postgres/query.go` wraps vendored `bin/db/psql-yq`. Output is YAML
(cheaper than psql ASCII, easy to slice with `yq`). (cheaper than psql ASCII, easy to slice with mikefarah/yq).
```bash ```bash
bin/postgres/query.go --profile onlyoffice -s document_asset # column list bin/postgres/query.go --profile onlyoffice -s document_asset # column list
+1 -1
View File
@@ -11,7 +11,7 @@ description: >-
```bash ```bash
bin/web/search.go "LadybugDB vector index" bin/web/search.go "LadybugDB vector index"
bin/web/search.go "model2vec multilingual" --category it bin/web/search.go "model2vec multilingual" --category it
bin/web/search.go "hypervisor" --site example.com --json | jq -r '.results[].url' bin/web/search.go "hypervisor" --site example.com --json | yq -r '.results[].url'
bin/web/search.go "postgres partial index" --lang en --fresh year bin/web/search.go "postgres partial index" --lang en --fresh year
``` ```
@@ -43,7 +43,7 @@ found" unless the client refuses to call it absence.
```bash ```bash
for i in $(seq 10); do for i in $(seq 10); do
bin/web/search.go "test $i" -n 1 --refresh --json | jq -r .status bin/web/search.go "test $i" -n 1 --refresh --json | yq -r '.status'
done done
``` ```
+33
View File
@@ -0,0 +1,33 @@
---
name: yq
description: >-
Use https://github.com/mikefarah/yq to work with YAML, JSON, XML, CSV,
TOML, HCL where it's efficient and less code. Use when slicing compose,
config, --json tool output, CSV/TOML/HCL/XML, or converting between those
formats. Not kislyuk Python yq. Not jq when yq already does the job.
---
# yq (mikefarah)
Use https://github.com/mikefarah/yq to work with YAML, JSON, XML, CSV, TOML, HCL where it's efficient and less code.
This is the Go `yq` (`yq --version` contains `mikefarah`). It is not
kislyuk/yq (Python, jq-syntax, YAML-only wrapper). `bin/db/psql-yq` already
calls this binary.
Prefer `yq` over `python3 -c`, `jq`, or ad-hoc parsers when one expression
reads or converts the file. Keep Python/Go for HTTP, binary protocols, and
in-process tests.
```bash
yq '.services.picoclaw.image' compose.yaml
yq -P . deploy/picoclaw/config.json # JSON → YAML
yq -o=json '.gates' # JSON stdin (qa/system_perf.py --json)
yq -p=csv -o=json .
yq -p=xml -o=json .
yq -p=toml '.package.name' file.toml
bin/brain/search.go "LadybugDB" --json | yq '.[].ref'
bin/web/search.go "hypervisor" --json | yq -r '.results[].url'
```
Do not print secrets, PII, or `$HOME/.config/brain/` through `yq`.
Generated
-1648
View File
File diff suppressed because it is too large Load Diff