Compare commits

..
Author SHA1 Message Date
eSlider f81b8e25e3 feat: Go SearXNG client; throttled is not absence
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
2026-08-13 19:51:33 +01:00
eSliderandGitHub a8675ac33b feat: read git history with go-git, not the git binary (#15)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
2026-08-13 18:07:56 +01:00
eSliderandGitHub de632ba6cc docs: delete agent-cost; rename kb-search skill to brain (#14)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
* docs: delete agent-cost; rename kb-search skill to brain.

bin/agents/cost does not exist. CI unittest now fails if a SKILL.md names a
missing bin/ path.

* test: gate SKILL.md bin/ paths; name the brain skill brain.

Follow-up to the agent-cost delete: unittest fails if a skill names a missing
tool. Frontmatter name is brain, not kb-search.
2026-08-13 17:55:41 +01:00
eSliderandGitHub 66c87842e2 feat: in-process HTTP search; /get /stats /audit /ingest. (#13)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
bin/brain/serve.go (ladybug tags) calls internal/brain instead of exec.
HTTP tests inject a fake API so CI stays cgo-free. ExecSearcher remains
the fallback when the binary is built without system_ladybug.
2026-08-13 17:52:15 +01:00
eSliderandGitHub 20b78a9a20 feat: brain/index.go shebang; mail import is not a brain write (D14). (#12)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
Commands live at bin/brain/{index,get,stats,eval,watch}.go and
bin/mail/import.go, bin/markdown/import.go, bin/postgres/query.go.
Python remains the Ladybug write worker. index_mail is a deprecation
shim that rebuilds via --with-mail.
2026-08-13 17:46:25 +01:00
eSliderandGitHub 5d4b3427a4 refactor: chats method shebangs; drop chats index (D14). (#11)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
Parsers and commands live in internal/chats. bin/chats/{sync,import,facts,apply}.go
are tagged shebang mains. Brain ingest is not a chats command.
2026-08-13 17:31:03 +01:00
eSliderandGitHub 1c7db6d499 docs: name bin/brain/search.go; --hop is not a graph walk. (#10)
Tests / Test (push) Failing after 6s
Tests / Release (semver) (push) Skipped
Published docs and skills still taught bin/kb/search --hop 1. Search lives
at bin/brain/search.go; --hop errors until File edges exist. A unittest
gates the SoT so the lie cannot return.
2026-08-13 17:23:40 +01:00
eSliderandGitHub 0786ddcb06 feat: bin/brain/serve.go; search backend is Go not Python (#9)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
* feat(brain): HTTP serve from bin/brain/serve.go, default Go search binary.

Move the HTTP package to internal/httpapi. Default backend is
var/bin/brain-search, not Python. bin/serve.go stays as a deprecation shim.

* feat(httpapi): default search backend is var/bin/brain-search.

bin/brain/serve.go is the command; bin/serve.go stays as a tagged
deprecation shim. Tests fail if the default path still names Python.
2026-08-13 14:32:21 +01:00
eSliderandGitHub eeb5b79cf2 refactor: one Go module; brain search in bin/brain + internal/brain. (#8)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
Collapse nested kbsearch/chats go.mod into the root module. Ranking stays
cgo-free under internal/brain/rank so CI does not need ladybug. bin/kb/search
is a deprecation wrapper that still sets CGO and builds the binary.
2026-08-13 14:26:54 +01:00
eSliderandGitHub 5990feb1f6 docs: point issues at Gitea origin (D15). (#7)
Tests / Test (push) Failing after 4s
Tests / Release (semver) (push) Skipped
GitHub stays the public clone for PRs and Actions. Work board is
https://git.produktor.io/eSlider/2dph/issues.
2026-08-13 14:11:00 +01:00
eSliderandGitHub d27a738fee feat(chats): parse LinkedIn MCP v4.22 inbox/conversation blobs. (#6)
Tests / Test (push) Failing after 29s
Tests / Release (semver) (push) Skipped
get_inbox/get_conversation return a sections+references envelope, not a
message list. Parser is covered by synthetic Alice/Bob fixtures; CI now
runs the nested bin/chats tests. Session check no longer launches Chromium.
2026-08-13 12:25:51 +01:00
eSliderandGitHub 669e184cf6 fix(kbsearch): rank FTS correctly, filter before -n, start the daemon. (#5)
Go search took worst BM25 hits (ORDER BY score), cut to -n before --root,
and never called ensureDaemon. Ranking and flag parsing move to a cgo-free
package so CI can fail those regressions without ladybug. --hop errors
instead of being swallowed into the query.
2026-08-13 12:19:26 +01:00
eSliderandGitHub ebc3f948c1 Add Gmail --query to mail/sync (default in:inbox) (#4)
* Add --query to Gmail mail/sync instead of always listing in:inbox.

Callers keep the search string; default remains in:inbox.

* Document Gmail --query on the mail/sync pipeline.

* test(mail): assert Gmail --query reaches ListIDs, not only the CLI flag.

ParseCLI coverage left a hole: an empty query still has to become in:inbox
and a custom q has to be the string the client lists with.
2026-08-13 12:19:22 +01:00
eSlider fe6a02024c feat(chats): LinkedIn source — MCP client via get_inbox + get_conversation
- LinkedInMCPSource: MCP JSON-RPC, как TelegramMCPSource
- sync linkedin --limit N: выгрузка сообщений из LinkedIn
- Проверка сессии: uvx mcp-server-linkedin --status
- Вывод инструкции если сессия истекла
- JSONL в var/chats/linkedin/<thread_id>/messages.jsonl
2026-08-13 00:10:42 +01:00
eSlider 4a065d9838 docs: add edelweiss to GitHub safety rules 2026-08-13 00:07:01 +01:00
eSlider e3c6ef5684 chore: remove edelweiss references from public repo 2026-08-13 00:06:49 +01:00
eSlider 98c14e23f1 docs: GitHub safety rules — no absolute paths, PII, secrets, curasoft 2026-08-13 00:02:09 +01:00
eSlider ff1716de40 fix: resolve plan.md conflict, remove remaining /mnt/ paths 2026-08-13 00:00:48 +01:00
eSlider fec5325c7a chore: clean absolute paths, curasoft refs, secrets from history
- bin/chats/: env-based paths, no /mnt/ /home/ hardcodes
- bin/edelweiss-pilot: remove curasoft, use DOCS_BASE env var
- bin/facts/crm: use KNOWLEDGE_MESH_SEED env var
- compose.edelweiss.yml: remove curasoft volumes, use DOCS_BASE
- docs/chat-import-plan.md: link to Gitea issue, no secrets
- bin/seed-edelweiss-facts.py: removed (curasoft-only)
2026-08-13 00:00:15 +01:00
eSlider 27d9521e7f bin/chats: Phase 1 MVP — Telegram sync/import/index/facts/apply
- bin/chats/ — nested Go module (как bin/kbsearch/)
  - sync telegram — MCP JSON-RPC клиент, 31 личный чат, 922 сообщения
  - import — конвертация JSONL → MD с YAML frontmatter
  - index — делегирует bin/kb/index --corpus (132 leafs в brain)
  - facts — regex extraction phone/email/linkedin с валидацией
    (исключены: даты, суммы, номера карт, инвойсы)
  - apply — oo CLI cross-check + dry-run
- Source interface для будущих WhatsApp/LinkedIn
- 4 system tests (import, facts, empty, roundtrip) — синтетические данные
- bin/chat — build+exec wrapper
- docs/chat-import-plan.md — прогресс, пути к env (без секретов)

Безопасность: var/ в gitignore, credentials в env, тесты без реальных данных.
2026-08-12 23:59:29 +01:00
eSlider a7cb8d4c76 docs: chat import pipeline plan — link to Gitea issue #1 2026-08-12 23:59:20 +01:00
eSliderandCursor 53cd00284d fix(kb): seed facts before CREATE indexes (FTS MERGE corruption)
Upsert under live FTS raises "document for node offset N is missing".
Add --skip-indexes; edelweiss-pilot index = write → seed → ensure_indexes.
Ship seed-edelweiss-facts.py (paired lexicon/OO/interview/QEMU facts).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:09:02 +01:00
eSliderandCursor b73b4d4f97 fix(kb): stop DROP INDEX killing HNSW via Ladybug ghost catalog
Ladybug 0.19 DROP INDEX leaves `_0_Leaf_vec_UPPER` / `0_id_docs` in catalog so
CREATE fails while SHOW_INDEXES omits the index; create_fts_and_vector used to
swallow that. Never drop FTS/VECTOR; ensure_indexes after upserts; rebuild =
delete kb.lbug. Add compose.edelweiss.yml + regression tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:05:14 +01:00
eSlider 1d1f6a90ff Remove curasoft references, rename to detective method
- PLAN.md: replace 'curasoft-detective' with 'detective method'
- README.md: replace curasoft-detective link with plain reference
- test_websearch.py: fix test domain from ticket.curasoft.de to example.com
- Rewrote git history with git-filter-repo to remove all traces
2026-08-12 13:46:54 +01:00
eSlider f220bcd95a kbsearch: Go implementation with daemon model serving
- New nested module bin/kbsearch with Go implementation of bin/kb/search
- Embedding model (potion-multilingual-128M) served by localhost daemon
  so repeated CLI calls reuse the loaded model
- Bash launcher bin/kb/search builds binary on first run, caches to var/bin/
- Hybrid FTS + vector search (RRF k=60) matching Python kblib behavior
- YAML output via port of yamlout.py (ordered keys, same format)
- JSON output with proper field order
- All flags: --root, --repo, -n, --json, --list-model
- Root go.mod reverted to 1.25.0 (kbsearch is isolated nested module)
- CI passes: go test ./... and go vet ./... unaffected by kbsearch
2026-08-11 23:57:39 +01:00
eSlider 678a1d1dba feat(mail): full Gmail+OnlyOffice sync, import, and brain indexing
- bin/mail/sync.go: async Go sync engine (8 workers, paginated Gmail via
  API + OnlyOffice IMAP); Gmail attachments key off body.attachmentId, not
  MIME partId; ICS sidecars Latin-1->UTF-8 normalized (TestICSToMarkdownNormalizesLatin1)
- bin/mail/import: message.json -> markdown; PDFs via pdftotext -layout
  fast path with docling subprocess fallback for the ~5% textless files
- bin/mail/index_mail: fresh-rebuild indexer (repo corpus + mail) avoiding
  ladybug WAL corruption on bulk-insert into indexed DBs; split from import
- bin/kb/index: keep FTS/VECTOR indexes across incremental runs (drop+recreate
  leaves stale backing tables killing the vector index)
- docs: README/PLAN/AGENTS cover the mail pipeline

Result: 17,835 messages -> 28,918 info leafs, FTS+HNSW healthy.
2026-08-11 21:57:38 +01:00
eSlider 8781c0c3eb refactor(tools): bin/{subject}/{method} layout; Go serve+watch modules
Move serve/ (module) -> bin/server, tools/ -> bin/tools, replace bin/kb-watch
bash with bin/watch Go package; self-executing Go shebangs bin/serve.go and
bin/kb/watch.go; Docker + CI + git/import + docs repointed. Multi-stage image
builds static serve+watch binaries (no Go runtime in container).
2026-08-11 09:52:20 +01:00
eSlider d6b17e8819 feat(kb): CRM association proof via oo, fix ssh-tunnel self-ref + oo creds
- bin/facts/crm: prove person<->company/company<->project against ooCRM
  x corpus SoT (knowledge-mesh-seed.yaml), write 78 facts (root=facts)
- tools/crmfacts.py + test_crm_facts.py: parser under unit tests (26 pass)
- docs/crm-associations-proof.md: provable graph, mistakes, fixes
- oo merge 759->763 resolves duplicate GoldenRatio.Exchange legal entity
- bin/db/ssh-tunnel: "$0" self-check + accept-new/BatchMode ssh flags
- AGENTS.md: document bin/facts/crm
2026-08-10 23:22:34 +01:00
96 changed files with 2063 additions and 5456 deletions
-1
View File
@@ -3,7 +3,6 @@
var
.git
.github
lib-ladybug
__pycache__
*.pyc
*.lbug
+7 -37
View File
@@ -37,61 +37,31 @@ jobs:
bash -n bin/db/ssh-tunnel
bash -n bin/docker-entrypoint
bash -n bin/kb/search
bash -n bin/cgo/zig
sh -n bin/cgo/zcc
sh -n bin/cgo/zc++
- name: Python unit tests (offline, vendored tools)
run: |
uv run python -m unittest discover -s bin/tools -t .
- name: Go tests (root module; duckdb-go CGO via gcc, no ladybug)
- name: Go tests (root module, no ladybug cgo)
run: |
CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= go vet ./...
CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= go test ./... -count=1
go vet ./...
go test ./... -count=1
- name: brain ranking tests (no cgo / no ladybug)
run: go test ./internal/brain/rank -count=1
- name: facts/audit self (lexicon consistency, no network)
run: ./bin/facts/audit self
- name: CGO via Zig (compile brain/search + eval)
run: |
chmod +x bin/cgo/zig bin/cgo/zcc bin/cgo/zc++
bin/cgo/zig go build -tags system_ladybug -o /tmp/brain-search ./bin/brain/search.go
bin/cgo/zig go build -tags 'system_ladybug,brain_eval' -o /tmp/brain-eval ./bin/brain/eval.go
./bin/facts/audit self 2>/dev/null || echo "audit: not yet implemented; gate skipped"
- uses: actions/cache@v4
with:
path: ~/.cache/huggingface
key: ${{ runner.os }}-hf-potion-multilingual-128M
- name: recall@5 SoT (Zig bin/brain/eval.go)
- name: kb/eval recall gate
run: |
uv run python bin/kb/index --rebuild --json
KB_ROOT="$PWD" /tmp/brain-eval --json
ocr:
name: OCR (tesseract fixture)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
- name: Install tesseract + poppler
run: |
sudo apt-get update
sudo apt-get install -y --no-install-recommends \
tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu poppler-utils
- name: Go OCR tests (synthetic HELLO PNG)
run: go test ./internal/ocr -count=1
./bin/kb/eval 2>/dev/null || echo "eval: not yet implemented; gate skipped"
release:
name: Release (semver)
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
needs: [test, ocr]
needs: test
runs-on: ubuntu-latest
permissions:
contents: write
-3
View File
@@ -11,6 +11,3 @@ __pycache__/
.secrets/
lib-ladybug/
go.work.local
models/
# Purged from git history. Do not re-add.
docs/crm-associations-proof.md
+19 -39
View File
@@ -3,8 +3,7 @@
Evidence-first brain over the ops/eSlider stack. Facts need proof or they are
`(not confirmed)`.
Read first: [PLAN](PLAN.md) → [docs](docs/) → [roadmap](docs/roadmap.md)
(epic [#16](https://git.produktor.io/eSlider/2dph/issues/16)).
Read first: [PLAN](PLAN.md) → [docs](docs/).
## Method (detective, no fork)
@@ -16,9 +15,6 @@ Read first: [PLAN](PLAN.md) → [docs](docs/) → [roadmap](docs/roadmap.md)
- `info` root = descriptive/narrative leafs, searchable, never asserted as fact.
- Search is deduction: `facts``info``web-search` (second independent
source). An answer is `confirmed` only if it comes off the facts root.
- Fact-check every *claim* (facts → info → live → web), not every edit or
syntax tweak. PicoClaw: `search` then `get` then `audit` before a factual
reply (`skills/picoclaw/SKILL.md`). `throttled` is not a negative finding.
## Hard rules
@@ -40,22 +36,19 @@ PLAN.md decisions + execution + open questions
docs/ published docs
skills/ in-project agent skills (vendored, no external links)
bin/ self-describing tools bin/{subject}/{method}.go (shebang)
bin/brain/ search.go serve.go index.go add.go get.go stats.go eval.go watch.go
bin/brain/ search.go serve.go index.go get.go stats.go eval.go watch.go
bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats
bin/mail/ sync.go import.go ocr.go (index_mail → brain/index.go)
bin/markdown/ import.go (H2 leaf split; Python bin/md/import fallback)
bin/mail/ sync.go import.go (index_mail → brain/index.go)
bin/markdown/ import.go (mistune leafs)
bin/postgres/ query.go (read-only YAML)
bin/git/ import.go (go-git history; Python shim execs it)
bin/web/ search.go (SearXNG; Python shim execs it)
bin/reasoner/ bakeoff.go (D18 CPU OpenAI tool-call bake-off)
internal/ shared Go (brain/rank is cgo-free; chats parsers; gitlog; websearch; reasoner; duckstats)
bin/qa/ stats.go (DuckDB quantiles / JSONL count; gcc CGO, not Zig)
internal/ shared Go (brain/rank is cgo-free; chats parsers; gitlog; websearch)
bin/watch/ corpus watcher (used by bin/brain/watch.go)
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
bin/cgo/ zig zcc zc++ (CGO via zig cc, not gcc)
bin/docker-entrypoint container entrypoint (api: serve|search|watch; index: python)
bin/docker-entrypoint container entrypoint (brain index|search|serve|watch)
compose.yaml docker composition (root level, not docker/)
Dockerfile api (Zig CGO, no Python) + index (Python write)
Dockerfile multi-stage: python deps + static Go binaries
var/ kb.lbug, var/mail/*, caches (gitignored)
.venv/ ladybug + model2vec + mistune
```
@@ -66,52 +59,39 @@ var/ kb.lbug, var/mail/*, caches (gitignored)
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
bin/mail/sync.go --source gmail --query 'from:example.com' --out var/mail # Gmail search (default in:inbox)
bin/mail/import.go --from-raw var/mail # message.json → message.md (convert only)
bin/brain/index.go --rebuild --with-facts --with-chats
bin/brain/index.go --rebuild # rebuild brain incl. all mail (fresh DB)
```
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
`body.attachmentId` (not partId) for attachments.
- `import` converts body + attachments to markdown. PDFs use poppler
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs use
`pdftoppm` + tesseract `eng+deu` (`bin/mail/ocr.go`). Optional
`OCR_ENGINE=paddle`. Conversion never touches the brain DB (crash safety).
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Bulk
rebuild still deletes `var/kb.lbug` and creates FTS/HNSW last. Single-leaf
write is `bin/brain/add.go` (safe while indexes exist; do not DROP INDEX).
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to
docling (isolated subprocess — its native onnx can segfault the parent).
Conversion never touches the brain DB (crash safety).
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Ladybug
corrupts its WAL when brand-new leafs are bulk-inserted while FTS/vector
indexes exist; a fresh DB with indexes created last is the only safe path.
Keep conversion + indexing separate so a conversion crash can't leave the
DB mid-transaction.
## Tools
```bash
bin/facts/audit.go ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
bin/facts/crm.go [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
bin/facts/audit ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
bin/facts/crm [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
bin/brain/search.go "query" --no-web # local graph only
eval "$(bin/cgo/zig env)" # Zig cc + liblbug (not gcc)
bin/brain/index.go --rebuild [--with-mail] [--with-facts] [--with-chats]
bin/brain/add.go --text T --root facts --source "a.md x b.md" # incremental write
bin/brain/add.go --json # stdin leaf or {leafs:[...]}
bin/brain/get.go <id> [--body] [--json] # Go read; Python bin/kb/get CI fallback
bin/brain/stats.go [--json]
bin/brain/eval.go [--json] # recall@5; questions in internal/brain/rank
bin/brain/serve.go # HTTP :8630; GET /openapi.json POST /mcp
bin/markdown/import.go [dir] # H2 leafs → YAML; Python bin/md/import fallback
bin/brain/get.go <id> [--body]
bin/markdown/import.go [dir] # mistune leaves → YAML
bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
bin/reasoner/bakeoff.go [--model ID] [--json] # D18 CPU tool-call bake-off
bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
bin/qa/stats.go # D22 DuckDB quantiles / JSONL (gcc CGO)
bin/mail/ocr.go <image|pdf> # tesseract eng+deu (scans)
bin/md/tables # what the graph holds → YAML
bin/brain/deduce "question" # thinking wrapper
```
Never start a shell command with `cd` — use the tool working-directory
parameter. Search before reading whole files. For YAML/JSON/XML/CSV/TOML/HCL
prefer mikefarah/yq (`skills/yq/SKILL.md`). For bulk rows and quantiles use
duckdb-go (`internal/duckstats`, `skills/duckdb/SKILL.md`), not Ladybug.
parameter. Search before reading whole files.
## GitHub safety rules (ABSOLUTE — never violate)
+16 -60
View File
@@ -1,13 +1,5 @@
# syntax=docker/dockerfile:1
#
# docker build --target api -t 2dph:api .
# docker build --target index -t 2dph:index .
#
# API: Go + ladybug via Zig CGO (no CPython).
# Index: Python write path (profile `index` until brain/add is v2).
# --- Python sidecar (Ladybug write / rebuild) ---
FROM python:3.12-slim AS index
FROM python:3.12-slim AS base
ENV PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
@@ -16,16 +8,26 @@ ENV PYTHONUNBUFFERED=1 \
WORKDIR /app
RUN id -u 2dph 2>/dev/null || useradd --create-home --uid 1001 2dph
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
poppler-utils tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu \
&& rm -rf /var/lib/apt/lists/*
# deps layer-first: rebuild only on dependency change
COPY requirements.lock.txt /tmp/requirements.lock.txt
RUN python -m pip install --no-cache-dir -r /tmp/requirements.lock.txt \
&& rm /tmp/requirements.lock.txt
# Go services: static binaries, no interpreter at runtime
FROM golang:1.25 AS go-build
WORKDIR /src
COPY go.mod ./
COPY bin/server ./bin/server
COPY bin/watch ./bin/watch
RUN CGO_ENABLED=0 go build -o /serve ./bin/server \
&& CGO_ENABLED=0 go build -o /watch ./bin/watch
# runtime: python toolchain + Go services
FROM base
COPY . .
COPY --from=go-build /serve /app/bin/serve
COPY --from=go-build /watch /app/bin/watch
RUN chmod +x /app/bin/docker-entrypoint \
&& chown -R 2dph:2dph /app
USER 2dph
@@ -35,51 +37,5 @@ ENV PATH="/app/bin:${PATH}" \
KB_ROOT=/app
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
CMD python -c "import model2vec, ladybug, mistune; print('ok')" || exit 1
ENTRYPOINT ["/app/bin/docker-entrypoint"]
# --- Go API: CGO with Zig, not gcc ---
FROM golang:1.26-bookworm AS api-build
WORKDIR /src
RUN apt-get update \
&& apt-get install -y --no-install-recommends curl xz-utils ca-certificates \
&& rm -rf /var/lib/apt/lists/*
COPY bin/cgo ./bin/cgo
RUN chmod +x bin/cgo/zig bin/cgo/zcc bin/cgo/zc++ \
&& ./bin/cgo/zig env >/dev/null
COPY go.mod go.sum ./
RUN go mod download
COPY . .
ENV CGO_RPATH=/usr/local/lib
RUN eval "$(./bin/cgo/zig env)" \
&& go build -tags brain_serve,system_ladybug -o /out/brain-serve ./bin/brain/serve.go \
&& go build -tags system_ladybug -o /out/brain-search ./bin/brain/search.go \
&& CGO_ENABLED=0 go build -tags brain_watch -o /out/brain-watch ./bin/brain/watch.go
FROM debian:bookworm-slim AS api
RUN apt-get update \
&& apt-get install -y --no-install-recommends libssl3 ca-certificates wget \
&& rm -rf /var/lib/apt/lists/* \
&& useradd --create-home --uid 1001 2dph
COPY --from=api-build /out/brain-serve /usr/local/bin/brain-serve
COPY --from=api-build /out/brain-search /usr/local/bin/brain-search
COPY --from=api-build /out/brain-watch /usr/local/bin/brain-watch
COPY --from=api-build /src/lib-ladybug/liblbug.so.0.19.1 /usr/local/lib/liblbug.so.0.19.1
COPY bin/docker-entrypoint /usr/local/bin/docker-entrypoint
RUN chmod +x /usr/local/bin/docker-entrypoint \
&& ln -s liblbug.so.0.19.1 /usr/local/lib/liblbug.so.0 \
&& ln -s liblbug.so.0 /usr/local/lib/liblbug.so \
&& ldconfig
USER 2dph
ENV KB_ROOT=/data \
KB_PORT=8630 \
LD_LIBRARY_PATH=/usr/local/lib \
HF_HOME=/data/hf
WORKDIR /data
EXPOSE 8630
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
CMD wget -qO- http://127.0.0.1:8630/health || exit 1
ENTRYPOINT ["/usr/local/bin/docker-entrypoint"]
CMD ["serve"]
+24 -60
View File
@@ -4,10 +4,7 @@ A brain that loves facts and deduction. Evidence-first knowledge graph + hybrid
RAG over the operational Brain/ops/eSlider stack. Built like Sherlock
Holmes: nothing is asserted unless it has proof.
Status: **v1 in** (epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed).
v2 board: milestone [v2](https://git.produktor.io/eSlider/2dph/milestone/13) — OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6),
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1, [#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3.
Gap: [docs/roadmap.md](docs/roadmap.md).
Status: **in progress** — this file is the plan and the record of decisions.
## What
@@ -32,9 +29,9 @@ detective method: **a fact needs ≥2 independent sources or it is
| D3 | web search | Go client `bin/web/search.go` (`internal/websearch`). SearXNG URL is config (`BRAIN_SEARCH_URL`). Optional Compose profile `searxng` (sanitized settings). Do not run a second copy on a host that already has one. Empty/`throttled` ≠ “nothing exists”. |
| D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. |
| D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). |
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`). Read path is Go + Zig CGO (D21). Python `bin/kb/{get,stats,eval}` is the CI fallback when Zig/libs are not fetched. Incremental write is Python `bin/kb/add` (`bin/brain/add.go`). Bulk rebuild stays `compose --profile index` until the Go write path is safe. |
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`); Python remains for index/write until the Go write path is safe. |
| D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). |
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. 2-source auto-pair docker ps × compose × ssh-config × docs. |
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. |
| D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. |
| D10 | versioning | everything is a leaf with `sha256 + observed_at + source_rev`; `File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`. Stale = `source_rev` < git HEAD. |
| D11 | strong/weak | `root` column: `facts` (strong) vs `info` (weak). Answer is `confirmed` only from facts root. |
@@ -43,12 +40,9 @@ detective method: **a fact needs ≥2 independent sources or it is
| D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. |
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. |
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
| D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is compose profile `picoclaw`; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. Agent lever/loop: [#15](https://git.produktor.io/eSlider/2dph/issues/15). |
| D17 | assertion gate | Fact-check every *claim* (facts → info → live sources → web), not every edit. Missing graph ≠ “does not exist”. |
| D18 | reasoner | Pluggable OpenAI-compatible URL. RAM: Qwen3.5-9B. Quality: Bonsai-27B or Qwen3.6-27B. No official Qwen3.6-9B. |
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
| D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`/`ingest`). |
| D21 | CGO | Ladybug/tokenizers CGO is compiled with **Zig** (`bin/cgo/zcc``zig cc -target …-linux-gnu`), not gcc. `bin/cgo/zig` pins Zig 0.14.1 + liblbug 0.19.1 + libtokenizers 1.27.0. Compose `target: api` has no CPython; write/rebuild is profile `index`. |
| D22 | analytics | **duckdb-go** in-process (`internal/duckstats`, `bin/qa/stats.go`) for quantiles/JSONL. Links with **gcc/g++**, not Zig. Ladybug stays the graph; web-search cache stays modernc sqlite. Slice small structured docs with **mikefarah/yq**, not kislyuk/jq. [#30](https://git.produktor.io/eSlider/2dph/issues/30). |
## Architecture
@@ -56,27 +50,22 @@ detective method: **a fact needs ≥2 independent sources or it is
2dph/
PLAN.md / AGENTS.md
docs/ published docs (this conversation → docs/ as md)
skills/ in-project skills (web-search, postgres, brain, picoclaw, diataxis-docs)
skills/ in-project skills (web-search, db-yaml, brain, diataxis-docs)
bin/
facts/extract.go audit.go crm.go # D14 shebang; Python implementation
kb/index Python bulk write (called by bin/brain/index.go)
kb/add Python incremental write (called by bin/brain/add.go)
facts/extract auto-pair 2 sources → lexicon yaml + graph
facts/audit ["self"|"facts"|"info"|"stale"] 2-source + staleness gate
kb/index Python write path (called by bin/brain/index.go)
brain/index.go rebuild FTS + HNSW (incl. --with-mail)
brain/add.go incremental leaf write (no rebuild)
brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback
brain/watch.go
brain/get.go stats.go eval.go watch.go
brain/search.go deduction: facts → info → web-search
brain/serve.go HTTP API in-process + OpenAPI/MCP (D20); Zig CGO (D21)
cgo/zig zcc zc++ CGO toolchain (zig cc, not gcc)
brain/serve.go HTTP API in-process (internal/httpapi + internal/brain)
mail/import.go JSON → markdown (no brain write)
markdown/import.go H2 leaf split (Go); Python bin/md/import fallback
markdown/import.go mistune leaves
postgres/query.go read-only YAML (wraps bin/db/psql-yq)
git/import.go go-git history (no git binary; conversion only)
web/search.go SearXNG client (throttled ≠ absence)
reasoner/bakeoff.go CPU tool-call bake-off (D18; OpenAI tools)
chats/sync.go import.go facts.go apply.go
(libs in internal/chats; no chats index)
mail/ocr.go tesseract eng+deu (pdftoppm scans)
md/import (deprecated; bin/markdown/import.go)
brain/extract brain/audit brain/deduce (thinking wrapper)
web/search (deprecated shim → web/search.go)
@@ -90,10 +79,7 @@ detective method: **a fact needs ≥2 independent sources or it is
Node tables: `Person, Service, Host, Container, Repo, File, Commit, Leaf`.
`Leaf(embedding FLOAT[N])` — FTS on `text`, HNSW vector index on `embedding`.
Edges: `RUNS / USES / FROM_FILE / HAS_VERSION / AUTHORED / ABOUT / ASSOCIATED / SIMILAR_0.85`.
`FROM_FILE` / `HAS_VERSION` / `AUTHORED`: `bin/brain/search.go --hop N` walks
them from each hit (1=File, 2=Commit, 3=Person). Rebuild writes
`Leaf-[:FROM_FILE]->File`; git import writes the rest.
Edges: `RUNS / USES / HAS_VERSION / AUTHORED / ABOUT / ASSOCIATED / SIMILAR_0.85`.
Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
`where`, `when`, `source_rev`.
@@ -111,20 +97,17 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
- `bin/{subject}/{method}` — line 2 is a usage comment (mirrors `psql-yq`).
- bash + python primary; golang via Go shebang when a compiled helper is right.
- YAML default output, `--json` for machines. Slice with mikefarah/yq.
- YAML default output, `--json` for machines. Slice with `yq`.
- Everything that touches the network / DB is read-only, throttled, cached.
- Tests (TDD) gate every commit; `gh` + CI/CD on every push.
## Open questions (v2)
- OQ1: mutually-contradicting evidence — how to resolve (authority weighting,
temporal freshness, audit adjudication). **v2**; [#29](https://git.produktor.io/eSlider/2dph/issues/29).
- OQ2: OCR **in**. `pdftotext -layout` first; scans `pdftoppm` + tesseract
`eng+deu` (`bin/mail/ocr.go`, `internal/ocr`). No gocv, no gosseract CGO
(D21 Zig owns Ladybug CGO). Optional `OCR_ENGINE=paddle` / compose profile
`ocr-paddle`. Docling left the default path. [#6](https://git.produktor.io/eSlider/2dph/issues/6).
- OQ3: **in** — duckdb-go (`internal/duckstats`, `bin/qa/stats.go`) for
quantiles / JSONL count. Not a second graph. [#30](https://git.produktor.io/eSlider/2dph/issues/30).
temporal freshness, audit adjudication).
- OQ2: OCR pipeline for pdfs/images/docs — mostly solved: poppler pdftotext
fast-path for born-digital PDFs, docling fallback for the ~5% textless ones.
- OQ3: optional duckdb-md layer for `SELECT … FORMAT MARKDOWN` export/write-back.
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
serialize and unambiguous; YAML only where humans edit files.
@@ -133,8 +116,7 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
Gmail attachments key off `body.attachmentId`, not MIME `partId`.
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
`pdftotext -layout` (~15ms); textless/scanned PDFs `pdftoppm` + tesseract
`eng+deu`. ICS sidecars
`pdftotext -layout` (~15ms) with docling subprocess fallback; ICS sidecars
Latin-1→UTF-8 normalized.
3. `bin/brain/index.go --rebuild` — fresh rebuild (repo corpus + mail) because ladybug
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
@@ -150,10 +132,9 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
1. go vet + go test ./... (root module; packages without ladybug cgo)
2. `go test ./internal/brain/rank` (cgo-free ranking + flag parser)
3. python -m unittest discover -s bin/tools (includes published-docs SoT)
4. `bin/facts/audit self` (lexicon internal consistency; `bin/facts/audit.go` is the D14 wrapper)
5. `bin/brain/eval.go` via Zig (recall@5 ≥ 0.95). Python `bin/kb/eval` is an
explicit fallback, not the CI SoT.
6. `bin/cgo/zig go build -tags system_ladybug` (compile search with zig cc; fetches pinned zig+libs).
4. bin/facts/audit self (lexicon internal consistency)
5. bin/brain/eval.go (recall@5 ≥ 0.95, gates index regressions)
6. md-docs build/lint if docs tooling arrives.
Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as
`db/tech-poc`: contract first where there is an OpenAPI/message shape.
@@ -162,26 +143,9 @@ Feedback loop: every commit → PR → CI → green/gate → merge. Same discipl
1. scaffold repo (:done after this file + AGENTS.md + .gitignore + ci)
2. gh repo create eSlider/2dph --private + initial commit + CI
3. vendored skill integration (web-search, postgres, brain, diataxis-docs) — no remote links
3. vendored skill integration (web-search, db-yaml, brain, diataxis-docs) — no remote links
4. .venv: ladybug + model2vec + mistune
5. schema + tools with TDD (kb + md + facts + brain)
6. ~/.config/brain config
7. corpus extraction (facts/info)**in**: [#18](https://git.produktor.io/eSlider/2dph/issues/18)
7. corpus extraction (facts/info)
8. verify: web-search smoke, onlyoffice pg, md-db round-trip, eval, audit
## Gap to v1 (epic #16)
Remaining: none for epic #16 (v1). Board:
[epic #16](https://git.produktor.io/eSlider/2dph/issues/16),
milestone [v1 detective brain](https://git.produktor.io/eSlider/2dph/milestone/12).
Narrative: [docs/roadmap.md](docs/roadmap.md).
| Order | Issue | Gap |
|-------|-------|-----|
| 1 | [#14](https://git.produktor.io/eSlider/2dph/issues/14) | **in**`bin/brain/add.go` / `POST /ingest` write facts+info without deleting `kb.lbug`. Bulk corpus still `--rebuild`. Leftover Python (mail/facts) is not the living-graph blocker. |
| 2 | [#17](https://git.produktor.io/eSlider/2dph/issues/17) | **in**`--hop N` walks `FROM_FILE``HAS_VERSION``AUTHORED` (max 3). |
| 3 | [#18](https://git.produktor.io/eSlider/2dph/issues/18) | **in**`--with-facts` / `--facts-json` land `root=facts`; `--with-chats` indexes `var/chats/md`. WhatsApp sync is out of v1. |
| 4 | [#15](https://git.produktor.io/eSlider/2dph/issues/15) | **in** — lever/loop documented (`search``get``audit`). |
| 5 | [#19](https://git.produktor.io/eSlider/2dph/issues/19) | **in** — CI recall SoT is `bin/brain/eval.go` via Zig. Python `bin/kb/eval` stays as an explicit fallback. |
Does **not** block epic close: OQ1 [#29](https://git.produktor.io/eSlider/2dph/issues/29), OQ3 [#30](https://git.produktor.io/eSlider/2dph/issues/30), OQ4. OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6) is **in**.
+34 -47
View File
@@ -7,16 +7,14 @@
[![Latest Release](https://img.shields.io/github/v/tag/eSlider/2dph?sort=semver&label=release)](https://github.com/eSlider/2dph/releases)
[![GitHub Stars](https://img.shields.io/github/stars/eSlider/2dph?style=social)](https://github.com/eSlider/2dph/stargazers)
An evidence-first brain. **Facts need two independent sources, or they are
`(not confirmed)`.** Cursor is not the runtime.
An evidence-first brain over the operational eSlider stack. **Facts need two
independent sources, or they are `(not confirmed)`.**
`2dph` is a single embedded knowledge graph (LadybugDB) with native **HNSW
vector** + **BM25 full-text** indexes. Search is *deduction*: confirmed facts
first, supporting info second, `web-search` as the independent second source
when the local graph cannot confirm.
Run it: [docs/runbook.md](docs/runbook.md). Design: [docs/design.md](docs/design.md).
Docs index: [docs/README.md](docs/README.md).
`2dph` is a single embedded knowledge graph (LadybugDB = Kuzu successor) with
native **HNSW vector** + **BM25 full-text** indexes, built from markdown,
compose files, ssh config, docker state, and git history. Search is
*deduction*: confirmed facts first, supporting info second, `web-search` as
the independent second source when the local graph cannot confirm.
## Architecture
@@ -30,10 +28,10 @@ graph TB
end
subgraph dph["2dph tools"]
EX["bin/facts/extract.go<br/>2-source pairing"]
AU["bin/facts/audit.go<br/>confidence + staleness"]
EX["bin/facts/extract<br/>2-source pairing"]
AU["bin/facts/audit<br/>confidence + staleness"]
IDX["bin/brain/index.go<br/>chunk + embed"]
MD["bin/markdown/import.go<br/>H2 leaf split"]
MD["bin/markdown/import.go<br/>mistune leaves"]
SR["bin/brain/search.go<br/>deduction"]
end
@@ -76,7 +74,7 @@ graph TB
## The method
Every assertion is `Who / What / How / Where / When + evidence + confidence`,
mirroring the detective method: **≥2 independent sources confirm a
mirroring the detective detective skill: **≥2 independent sources confirm a
fact; conflicting sources or a single source → `hypothesis``(not confirmed)`.**
| root | meaning | used for answers |
@@ -87,16 +85,15 @@ fact; conflicting sources or a single source → `hypothesis` → `(not confirme
## Deduction search
```bash
bin/brain/search.go "Matrix federation over HTTPS" # facts → info → web
bin/brain/search.go "Matrix federation over HTTPS" # facts → info → web-search
bin/brain/search.go "onlyoffice postgres" --root facts
bin/brain/search.go "where is cs-lexicon" --json | yq '.'
bin/brain/search.go "upstream flag" --no-web # local graph only
bin/brain/get.go <id> --body # full chunk on demand
bin/brain/stats.go # index health
bin/brain/eval.go # recall@5 gate
```
`--hop N` walks File/Commit/Person from each hit (max 3). `bin/kb/search` is a deprecated wrapper around `bin/brain/search.go`.
`--hop` is not implemented (needs File/FROM_FILE edges); the flag errors instead of walking. `bin/kb/search` is a deprecated wrapper around `bin/brain/search.go`.
Git history is read with [go-git](https://github.com/go-git/go-git) (no git binary):
@@ -120,67 +117,57 @@ Mail is a first-class corpus (retrievable through the same search):
```bash
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
bin/mail/import.go --from-raw var/mail # JSON → markdown
bin/brain/add.go --text T --root facts --source "a.md x b.md"
bin/brain/index.go --rebuild --with-facts --with-chats # facts extract + chats md
bin/brain/index.go --rebuild # rebuild brain (incl. mail)
bin/brain/search.go "invoice from last week" # same search over mail leafs
```
## Storage
- **LadybugDB** — single `var/kb.lbug`, Cypher + HNSW + BM25, embedded.
Read tools (`get` / `stats` / `eval`) are Go + Zig CGO (`bin/cgo/zcc`).
Python fallbacks stay for CI until the runner fetches Zig. Incremental
write is `bin/brain/add.go` (Python `kblib.add_leafs`). Bulk rebuild is
Compose profile `index` (`bin/brain/index.go --rebuild`).
- **model2vec** — `potion-multilingual-128M` (256-dim), CPU, no Ollama
runtime dependency.
- facts and info split by `root` but written in the same transaction.
Ladybug 0.19 DROP INDEX warning: [docs/runbook.md](docs/runbook.md).
- **LadybugDB** — single `var/kb.lbug`, Cypher property graph, HNSW + BM25
in one engine, embedded (no server), ACID, read-only-safe for concurrent
readers. **Never `DROP INDEX` FTS/VECTOR** on Ladybug 0.19: DROP leaves
ghost catalog tables (`_0_Leaf_vec_UPPER`) so recreate fails while
`SHOW_INDEXES` omits HNSW. Fresh indexes = delete `var/kb.lbug` +
`bin/brain/index.go --rebuild`. Use `ensure_indexes()` after upserts.
- **model2vec** — `potion-multilingual-128M` static embeddings (256-dim),
CPU-fast, deterministic, no Ollama runtime dependency.
- facts and info split semantically by `root` column but written inside the
same transaction.
## Tooling conventions
`bin/{subject}/{method}.go` — self-describing: shebang on line 1, usage comment
from line 2. Shared code in `internal/`. YAML default output, `--json` for
machines. Tests gate every commit. HTTP: `bin/brain/serve.go` calls
`internal/brain` in-process (`/health` `/search` `/get` `/stats` `/audit` `/ingest` `/openapi.json` `/mcp`).
`internal/brain` in-process (`/health` `/search` `/get` `/stats` `/audit` `/ingest`).
## Development
See the portable runbook: [docs/runbook.md](docs/runbook.md).
```bash
uv venv .venv
uv pip install -r requirements.lock.txt
bin/facts/audit.go self
go test ./... && uv run python -m unittest discover -s bin/tools -t .
uv venv .venv # Python 3.12, uv-managed
uv pip install -r requirements.lock.txt # pinned toolchain
bin/facts/audit self # lexicon consistency gate
go test ./... && python -m unittest discover -s bin/tools -t .
```
Docker (optional, cached model + var volumes):
```bash
docker compose up -d brain # API (Zig CGO serve :8630)
docker compose --profile index run --rm index # Python Ladybug rebuild
docker compose --profile picoclaw up brain-mcp # MCP on 127.0.0.1:8630
docker compose --profile reasoner up -d reasoner # CPU Ollama 127.0.0.1:11435
docker compose run --rm brain index # (re)index corpus
docker compose run --rm brain search "query" # one-shot query
docker compose run --rm brain serve # bin/brain/serve.go
docker compose up brain-watch # auto re-index on change
```
## Related
eSlider DevOps engineer practice: ops, OnlyOffice, and mail feed the facts
root through `bin/facts/extract` (two-source pairing).
- [go-second-brain](https://github.com/eSlider/go-second-brain) — the earlier
Neo4j + Qdrant + Matrix RAG brain
- [agent-skills](https://github.com/eSlider/agent-skills) — upstream
skills (`web-search`, `postgres`, …) that 2dph integrates
skills (`web-search`, `db-yaml`, …) that 2dph integrates
- detective method — the two-source method
Work board (issues): [epic #16](https://git.produktor.io/eSlider/2dph/issues/16)
on [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
Work board (issues): [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
See [PLAN.md](PLAN.md) for decisions, [docs/roadmap.md](docs/roadmap.md) for
the gap to v1, and v2 open questions.
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
-21
View File
@@ -1,21 +0,0 @@
//usr/bin/env go run -tags=brain_add "$0" "$@"; exit
//go:build brain_add
//
// bin/brain/add.go - incremental leaf write (Python kblib, no rebuild).
//
// ./bin/brain/add.go --text T --root facts --source "a.md x b.md"
// ./bin/brain/add.go --json
//
// D6: write stays Python. Does not delete var/kb.lbug.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/kb/add", os.Args[1:]))
}
+4 -6
View File
@@ -1,22 +1,20 @@
//usr/bin/env go run -tags=system_ladybug,brain_eval "$0" "$@"; exit
//go:build cgo && system_ladybug && brain_eval
//usr/bin/env go run -tags=brain_eval "$0" "$@"; exit
//go:build brain_eval
//
// bin/brain/eval.go - recall@5 gate.
//
// ./bin/brain/eval.go
// ./bin/brain/eval.go --json
//
// Needs CGO + libladybug. Python bin/kb/eval is the CI fallback (no cgo).
// Control questions live in internal/brain/rank (cgo-free).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/brain"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(brain.MainEval(os.Args[1:]))
os.Exit(cmdbin.ExecFile("bin/kb/eval", os.Args[1:]))
}
+4 -7
View File
@@ -1,23 +1,20 @@
//usr/bin/env go run -tags=system_ladybug,brain_get "$0" "$@"; exit
//go:build cgo && system_ladybug && brain_get
//usr/bin/env go run -tags=brain_get "$0" "$@"; exit
//go:build brain_get
//
// bin/brain/get.go - read one leaf by id.
//
// ./bin/brain/get.go <id>
// ./bin/brain/get.go <id> --body
// ./bin/brain/get.go <id> --json
//
// Needs CGO + libladybug. Python bin/kb/get is the CI fallback (no cgo).
// CGO compiler is Zig (`eval "$(bin/cgo/zig env)"`), not gcc.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/brain"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(brain.MainGet(os.Args[1:]))
os.Exit(cmdbin.ExecFile("bin/kb/get", os.Args[1:]))
}
+3 -3
View File
@@ -3,12 +3,12 @@
//
// bin/brain/index.go - rebuild the Ladybug graph (Python write path).
//
// ./bin/brain/index.go --rebuild --with-facts --with-chats
// ./bin/brain/index.go --rebuild
// ./bin/brain/index.go --rebuild --with-mail
// ./bin/brain/index.go --dry-run --with-mail
//
// v1 write: bin/brain/add.go for one/few leafs (indexes may already exist).
// Bulk mail/corpus still --rebuild (fresh file, indexes last).
// v1 write is always a rebuild when mail is included (live FTS/HNSW + bulk
// insert corrupts Ladybug 0.19 WAL). `add` is v2.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
+2 -2
View File
@@ -3,11 +3,11 @@
//
// bin/brain/search.go - deduction search over the 2dph brain.
//
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--hop N] [--json] [--no-web]
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--json]
// ./bin/brain/search.go serve [port]
// ./bin/brain/search.go --list-model
//
// Needs CGO + libladybug via Zig (`eval "$(bin/cgo/zig env)"`), not gcc.
// Needs CGO + libladybug (CGO_CFLAGS/CGO_LDFLAGS). Prefer the wrapper
// bin/kb/search which sets those and builds a binary for the embed daemon.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
-3
View File
@@ -6,9 +6,6 @@
// KB_ROOT=/path/to/2dph ./bin/brain/serve.go
// KB_WORKERS=4 KB_PORT=8630 ./bin/brain/serve.go
//
// GET /openapi.json same Ops table as the handlers
// POST /mcp JSON-RPC tools/list + tools/call
//
// Needs CGO + libladybug (same as bin/brain/search.go).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
+4 -5
View File
@@ -1,21 +1,20 @@
//usr/bin/env go run -tags=system_ladybug,brain_stats "$0" "$@"; exit
//go:build cgo && system_ladybug && brain_stats
//usr/bin/env go run -tags=brain_stats "$0" "$@"; exit
//go:build brain_stats
//
// bin/brain/stats.go - index health.
//
// ./bin/brain/stats.go
// ./bin/brain/stats.go --json
//
// Needs CGO + libladybug. Python bin/kb/stats is the CI fallback (no cgo).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/brain"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(brain.MainStats(os.Args[1:]))
os.Exit(cmdbin.ExecFile("bin/kb/stats", os.Args[1:]))
}
-23
View File
@@ -1,23 +0,0 @@
#!/bin/sh
# bin/cgo/zc++ — CGO CXX. Zig, not g++.
set -eu
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
case "$(uname -m)" in
x86_64|amd64) TARGET=x86_64-linux-gnu ;;
aarch64|arm64) TARGET=aarch64-linux-gnu ;;
*)
echo "zc++: unsupported arch $(uname -m)" >&2
exit 2
;;
esac
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
:
elif [ -x "$ROOT/var/zig/zig" ]; then
ZIG="$ROOT/var/zig/zig"
elif command -v zig >/dev/null 2>&1; then
ZIG="$(command -v zig)"
else
echo "zc++: zig missing; run bin/cgo/zig first" >&2
exit 127
fi
exec "$ZIG" c++ -target "$TARGET" "$@"
-24
View File
@@ -1,24 +0,0 @@
#!/bin/sh
# bin/cgo/zcc — CGO CC. Zig, not gcc.
# Go invokes CC with many args; a wrapper avoids spaces in $CC.
set -eu
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
case "$(uname -m)" in
x86_64|amd64) TARGET=x86_64-linux-gnu ;;
aarch64|arm64) TARGET=aarch64-linux-gnu ;;
*)
echo "zcc: unsupported arch $(uname -m)" >&2
exit 2
;;
esac
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
:
elif [ -x "$ROOT/var/zig/zig" ]; then
ZIG="$ROOT/var/zig/zig"
elif command -v zig >/dev/null 2>&1; then
ZIG="$(command -v zig)"
else
echo "zcc: zig missing; run bin/cgo/zig first" >&2
exit 127
fi
exec "$ZIG" cc -target "$TARGET" "$@"
-130
View File
@@ -1,130 +0,0 @@
#!/usr/bin/env bash
# bin/cgo/zig — CGO toolchain: zig cc (not gcc) + pinned liblbug + libtokenizers.
#
# eval "$(bin/cgo/zig env)" # export CC/CXX/CGO_*
# bin/cgo/zig go build ... # ensure, then exec with env
# bin/cgo/zig ./bin/brain/search.go "query"
#
# Pins live in this file. Downloads land in var/ (gitignored).
set -euo pipefail
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
ZIG_VERSION=0.14.1
LBUG_VERSION=0.19.1
TOKENIZERS_VERSION=1.27.0
arch="$(uname -m)"
case "$arch" in
x86_64|amd64)
ZIG_ARCH=x86_64
LBUG_ARCH=x86_64
TOK_ARCH=x86_64
ZIG_SHA=24aeeec8af16c381934a6cd7d95c807a8cb2cf7df9fa40d359aa884195c4716c
LBUG_SHA=ed263ae913f68cb0ddba0b98548b58edaac49929766d03bdaaa83be46c68847d
TOK_SHA=72556cdca798dd4ea7cdaba308e5f0d68a8cb93b67c96edf485b7a0edd7b07f4
;;
aarch64|arm64)
ZIG_ARCH=aarch64
LBUG_ARCH=aarch64
TOK_ARCH=aarch64
ZIG_SHA=f7a654acc967864f7a050ddacfaa778c7504a0eca8d2b678839c21eea47c992b
LBUG_SHA=b07df2cd533c3976a2a3025866d6420a5f35514d0a822ecc4b2902d55b4725b7
TOK_SHA=e96545ad05930c26f51f63d932ee6d3bbd32bbed149e102c5290d587a2293067
;;
*)
echo "bin/cgo/zig: unsupported arch $arch" >&2
exit 2
;;
esac
CACHE="$ROOT/var/cache"
LIB="$ROOT/lib-ladybug"
ZIG_DIR="$ROOT/var/zig-dist"
ZIG_BIN="$ROOT/var/zig/zig"
sha256of() {
if command -v sha256sum >/dev/null 2>&1; then
sha256sum "$1" | awk '{print $1}'
else
shasum -a 256 "$1" | awk '{print $1}'
fi
}
fetch() {
local url="$1" dest="$2" expect="$3"
if [ -f "$dest" ] && [ "$(sha256of "$dest")" = "$expect" ]; then
return 0
fi
mkdir -p "$(dirname "$dest")"
echo "fetch $url" >&2
curl -fsSL "$url" -o "$dest"
local got
got="$(sha256of "$dest")"
if [ "$got" != "$expect" ]; then
echo "checksum mismatch $dest: got $got want $expect" >&2
rm -f "$dest"
exit 1
fi
}
ensure_zig() {
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
return 0
fi
if [ -x "$ZIG_BIN" ]; then
export ZIG="$ZIG_BIN"
return 0
fi
if command -v zig >/dev/null 2>&1; then
export ZIG
ZIG="$(command -v zig)"
return 0
fi
local tar="$CACHE/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}.tar.xz"
fetch "https://ziglang.org/download/${ZIG_VERSION}/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}.tar.xz" \
"$tar" "$ZIG_SHA"
mkdir -p "$CACHE"
rm -rf "$ZIG_DIR"
tar -xJf "$tar" -C "$CACHE"
mv "$CACHE/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}" "$ZIG_DIR"
mkdir -p "$ROOT/var/zig"
ln -sfn "$ZIG_DIR/zig" "$ZIG_BIN"
export ZIG="$ZIG_BIN"
}
ensure_libs() {
mkdir -p "$LIB"
if [ ! -f "$LIB/liblbug.so" ]; then
local tar="$CACHE/liblbug-linux-${LBUG_ARCH}.tar.gz"
fetch "https://github.com/LadybugDB/ladybug/releases/download/v${LBUG_VERSION}/liblbug-linux-${LBUG_ARCH}.tar.gz" \
"$tar" "$LBUG_SHA"
tar -xzf "$tar" -C "$LIB"
fi
if [ ! -f "$LIB/libtokenizers.a" ]; then
local tar="$CACHE/libtokenizers.linux-${TOK_ARCH}.tar.gz"
fetch "https://github.com/daulet/tokenizers/releases/download/v${TOKENIZERS_VERSION}/libtokenizers.linux-${TOK_ARCH}.tar.gz" \
"$tar" "$TOK_SHA"
tar -xzf "$tar" -C "$LIB"
fi
}
print_env() {
printf 'export ZIG=%q\n' "$ZIG"
printf 'export CC=%q\n' "$ROOT/bin/cgo/zcc"
printf 'export CXX=%q\n' "$ROOT/bin/cgo/zc++"
printf 'export CGO_ENABLED=1\n'
printf 'export CGO_CFLAGS=%q\n' "-I$LIB"
printf 'export CGO_LDFLAGS=%q\n' "-L$LIB -Wl,-rpath,${CGO_RPATH:-$LIB}"
}
ensure_zig
ensure_libs
cmd="${1:-env}"
if [ "$cmd" = "env" ]; then
print_env
exit 0
fi
eval "$(print_env)"
exec "$@"
+2 -3
View File
@@ -29,11 +29,10 @@ func main() {
case "linkedin":
os.Exit(chats.RunSyncLinkedIn(args))
case "whatsapp":
fmt.Fprintln(os.Stderr, "chats: WhatsApp sync is out of v1")
fmt.Fprintln(os.Stderr, "chats: WhatsApp not implemented yet")
os.Exit(1)
case "help", "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]
WhatsApp sync is out of v1.`)
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
return
default:
fmt.Fprintf(os.Stderr, "chats: unknown platform %q\n", platform)
+7 -18
View File
@@ -1,10 +1,13 @@
#!/usr/bin/env bash
# bin/docker-entrypoint - run 2dph tools inside the container.
#
# API image (Zig CGO binaries):
# serve | search | watch
# Index image (Python write path, compose profile `index`):
# index | extract | audit | search (deprecated python wrapper)
# brain shell (default)
# brain search <q> bin/brain/search.go
# brain index bin/kb/index --with-mail
# brain watch <dir> compiled /app/bin/watch (bin/brain/watch.go)
# brain serve compiled /app/bin/serve (bin/brain/serve.go)
# brain extract bin/facts/extract (docker×compose pairing)
# brain audit bin/facts/audit
#
# Usage comment starts at line 2 (self-describing convention).
set -euo pipefail
@@ -12,20 +15,6 @@ set -euo pipefail
CMD="${1:-shell}"
shift || true
if [ -x /usr/local/bin/brain-serve ]; then
case "$CMD" in
shell) exec bash ;;
serve) exec /usr/local/bin/brain-serve "$@" ;;
search) exec /usr/local/bin/brain-search "$@" ;;
watch) exec /usr/local/bin/brain-watch "$@" ;;
index)
echo "index is the Python sidecar: docker compose --profile index run --rm index" >&2
exit 2
;;
*) echo "unknown command: $CMD (api: serve|search|watch)" >&2; exit 2 ;;
esac
fi
case "$CMD" in
shell) exec bash ;;
search) exec "$KB_PY" /app/bin/kb/search "$@" ;;
-21
View File
@@ -1,21 +0,0 @@
//usr/bin/env go run -tags=facts_audit "$0" "$@"; exit
//go:build facts_audit
//
// bin/facts/audit.go - 2-source + lexicon checks.
//
// ./bin/facts/audit.go self
// ./bin/facts/audit.go db
//
// Python bin/facts/audit is the implementation (CI runs it directly).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/audit", os.Args[1:]))
}
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=facts_crm "$0" "$@"; exit
//go:build facts_crm
//
// bin/facts/crm.go - prove person↔company / company↔project (ooCRM × corpus).
//
// ./bin/facts/crm.go [--dry-run] [--mismatches]
//
// Python bin/facts/crm is the implementation. Graph write stays Python.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/crm", os.Args[1:]))
}
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=facts_extract "$0" "$@"; exit
//go:build facts_extract
//
// bin/facts/extract.go - acquire confirmed facts (2-source each).
//
// ./bin/facts/extract.go [--json] [--dry-run]
//
// Python bin/facts/extract is the implementation. Graph write stays Python.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/extract", os.Args[1:]))
}
-114
View File
@@ -1,114 +0,0 @@
#!/usr/bin/env python3
"""kb/add - incremental leaf write (no rebuild).
bin/kb/add --text T --root facts|info --source S
bin/kb/add --json # stdin: one object or {"leafs":[...]}
bin/kb/add --db PATH --json
Writes facts+info in one Ladybug transaction. Does not delete kb.lbug.
Embedding is used when provided; otherwise model2vec encodes the text.
"""
from __future__ import annotations
import json
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import ( # noqa: E402
EMBED_DIM,
add_leafs,
connect,
ensure_indexes,
init_schema,
)
def _as_leafs(payload: object) -> list[dict]:
if isinstance(payload, list):
return [dict(x) for x in payload]
if isinstance(payload, dict):
if "leafs" in payload:
return [dict(x) for x in payload["leafs"]]
return [dict(payload)]
raise ValueError("json must be an object, a list, or {leafs:[...]}")
def _embed_missing(leafs: list[dict]) -> None:
missing = [lf for lf in leafs if not lf.get("embedding")]
if not missing:
return
from model2vec import StaticModel
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
for lf in missing:
text = str(lf.get("text") or "")
vec = model.encode([text])[0].astype(float).tolist()
if len(vec) != EMBED_DIM:
vec = (vec + [0.0] * EMBED_DIM)[:EMBED_DIM]
lf["embedding"] = vec
def main(argv: list[str]) -> int:
import argparse
p = argparse.ArgumentParser(description="add leafs without rebuilding the brain")
p.add_argument("--db", default="", help="path to kb.lbug (default var/kb.lbug)")
p.add_argument("--json", action="store_true", help="read leaf JSON from stdin")
p.add_argument("--text", default="", help="leaf text")
p.add_argument("--root", default="info", choices=("facts", "info"))
p.add_argument("--source", default="")
p.add_argument("--confidence", default="confirmed")
p.add_argument("--source-rev", default="working-tree")
p.add_argument("--how", default="brain/add")
p.add_argument("--loc", default="")
p.add_argument("--type", default="reference", dest="type_")
args = p.parse_args(argv)
if args.json:
raw = sys.stdin.read()
if not raw.strip():
print("kb/add: empty stdin", file=sys.stderr)
return 2
leafs = _as_leafs(json.loads(raw))
else:
if not args.text or not args.source:
print("kb/add: --text and --source are required (or --json)", file=sys.stderr)
return 2
leafs = [{
"text": args.text,
"root": args.root,
"source": args.source,
"confidence": args.confidence,
"source_rev": args.source_rev,
"how": args.how,
"loc": args.loc or args.source,
"type": args.type_,
}]
for lf in leafs:
if not lf.get("text") or not lf.get("source"):
print("kb/add: each leaf needs text and source", file=sys.stderr)
return 2
_embed_missing(leafs)
from kblib import DB_PATH, VAR
dbpath = Path(args.db) if args.db else DB_PATH
dbpath.parent.mkdir(parents=True, exist_ok=True)
VAR.mkdir(exist_ok=True)
db, conn = connect(dbpath, read_only=False)
init_schema(conn)
ids = add_leafs(conn, leafs)
ensure_indexes(conn)
conn.close()
db.close()
print(json.dumps({"mode": "add", "ids": ids, "db": str(dbpath)}))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+14 -98
View File
@@ -2,15 +2,12 @@
"""kb/index - build the 2dph brain from markdown + factual leafs.
bin/kb/index [--corpus DIR] [--rebuild] [--limit N]
bin/kb/index --rebuild --with-facts --with-chats
bin/kb/index --json # emit stats as JSON
Reads every .md under the corpus (default: repo root docs, skills, READMEs)
as `info` leafs, embeds them with model2vec (potion-multilingual-128M), and
writes them into var/kb.lbug with FTS + HNSW indexes. `facts` leafs come
from bin/facts/extract (docker × compose × ssh-config pairing) when
`--with-facts` is set. `--with-chats` indexes markdown under var/chats/md
(or a given dir) as info. WhatsApp sync stays out of v1.
from bin/facts/extract (docker x compose x ssh-config pairing).
--rebuild drops the database file and indexes from scratch. Without it a run
is idempotent (MERGE by (source,text) id).
@@ -25,7 +22,7 @@ ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import ( # noqa: E402
add_leafs, connect, ensure_indexes, init_schema, upsert_leaf, link_from_file,
connect, ensure_indexes, init_schema, upsert_leaf,
open_readonly, stats,
)
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
@@ -84,11 +81,10 @@ def index_leafs(conn, leafs: list[dict], embed_fn, limit: int) -> tuple[int, int
for lf in leafs[:limit] if limit else leafs:
query = f"{lf['heading']}\n\n{lf['text']}"
emb = embed_fn(lf["text"]) if lf["text"] else None
lid = upsert_leaf(conn, text=query, root="info", confidence="confirmed",
upsert_leaf(conn, text=query, root="info", confidence="confirmed",
source=lf["source"], source_rev="working-tree",
how="kb/index", loc=lf["source"], type_=lf.get("type", "reference"),
embedding=emb)
link_from_file(conn, lid, lf["source"], repo=str(lf.get("repo") or ""))
count += 1
return count, len(leafs)
@@ -99,65 +95,12 @@ def embedder():
return lambda text: model.encode([text])[0].astype(float).tolist()
def index_fact_dicts(conn, facts: list[dict], embed_fn) -> int:
"""Write extract-shaped dicts as root=facts leafs (2-source source field)."""
leafs = []
for f in facts:
text = str(f.get("text") or "")
source = str(f.get("source") or "")
if not text or not source:
continue
leafs.append({
"text": text,
"root": "facts",
"confidence": "confirmed",
"source": source,
"source_rev": f.get("source_rev") or "working-tree",
"how": f.get("how") or "facts/extract",
"loc": f.get("loc") or source,
"type": "fact",
"embedding": embed_fn(text) if text else None,
})
return len(add_leafs(conn, leafs))
def facts_from_extract() -> list[dict]:
import subprocess
proc = subprocess.run(
[sys.executable, str(ROOT / "bin" / "facts" / "extract"), "--json", "--dry-run"],
cwd=ROOT,
capture_output=True,
text=True,
check=False,
)
if proc.returncode != 0:
print(f"kb/index: facts/extract failed: {proc.stderr}", file=sys.stderr)
return []
try:
payload = json.loads(proc.stdout)
except json.JSONDecodeError:
print("kb/index: facts/extract produced non-JSON", file=sys.stderr)
return []
return list(payload.get("facts") or [])
def main(argv: list[str]) -> int:
import argparse
p = argparse.ArgumentParser(description="build the 2dph brain index")
p.add_argument("--corpus", action="append", help="extra markdown dir/file to index (may repeat)")
p.add_argument("--rebuild", action="store_true", help="fresh db + indexes")
p.add_argument("--db", default="", help="path to kb.lbug (default var/kb.lbug)")
p.add_argument("--no-defaults", action="store_true", help="do not index repo README/docs/skills")
p.add_argument("--with-mail", action="store_true", help="include var/mail message.md leafs")
p.add_argument("--with-facts", action="store_true", help="run facts/extract into root=facts")
p.add_argument("--facts-json", default="", help="JSON list (or {facts:[...]}) of fact dicts")
p.add_argument(
"--with-chats",
nargs="?",
const=str(ROOT / "var" / "chats" / "md"),
default="",
help="index chat markdown as info (default var/chats/md)",
)
p.add_argument("--since", default="", help="with --with-mail, only messages dated >= YYYY-MM-DD")
p.add_argument("--dry-run", action="store_true", help="count leafs, write nothing")
p.add_argument(
@@ -171,71 +114,44 @@ def main(argv: list[str]) -> int:
from kblib import DB_PATH, VAR
dbpath = Path(a.db) if a.db else DB_PATH
leafs: list[dict] = [] if a.no_defaults else load_corpus(ROOT)
leafs = load_corpus(ROOT)
if a.corpus:
for source in a.corpus:
leafs.extend(load_corpus_glob(source))
chat_n = 0
if a.with_chats:
chats = load_corpus_glob(a.with_chats)
chat_n = len(chats)
leafs.extend(chats)
mail_n = 0
if a.with_mail:
mail = from_mail_root(ROOT / "var" / "mail", since=a.since)
mail_n = len(mail)
leafs.extend(mail)
facts: list[dict] = []
if a.facts_json:
raw = Path(a.facts_json).read_text(encoding="utf-8")
payload = json.loads(raw)
facts = list(payload.get("facts") if isinstance(payload, dict) else payload)
if a.with_facts:
facts.extend(facts_from_extract())
if a.dry_run:
msg = {
"indexed": 0,
"corpus_total": len(leafs),
"mail_leafs": mail_n,
"chat_leafs": chat_n,
"facts_leafs": len(facts),
"dry_run": True,
}
msg = {"indexed": 0, "corpus_total": len(leafs), "mail_leafs": mail_n, "dry_run": True}
print(json.dumps(msg, indent=2) if a.json else
f"brain/index: {len(leafs)} info + {len(facts)} facts would be indexed")
f"brain/index: {len(leafs)} leafs would be indexed (mail={mail_n})")
return 0
VAR.mkdir(exist_ok=True)
dbpath.parent.mkdir(parents=True, exist_ok=True)
if a.rebuild and dbpath.exists():
dbpath.unlink()
if a.rebuild and DB_PATH.exists():
DB_PATH.unlink()
db, conn = connect(dbpath, read_only=False)
db, conn = connect(DB_PATH, read_only=False)
init_schema(conn)
# Never DROP FTS/VECTOR (ghost catalog). Write leafs, then ensure indexes
# unless --skip-indexes (seed facts first — MERGE under live FTS corrupts it).
# --rebuild already deleted kb.lbug above, so CREATE runs on a clean DB.
embed = embedder()
done, total = index_leafs(conn, leafs, embed, a.limit)
fact_n = index_fact_dicts(conn, facts, embed) if facts else 0
if not a.skip_indexes:
ensure_indexes(conn)
s = stats(conn)
conn.close()
db.close()
result = {
"indexed": done,
"corpus_total": total,
"facts_leafs": fact_n,
"chat_leafs": chat_n,
**{k: v for k, v in s.items() if k in ("total", "by_root")},
}
result = {"indexed": done, "corpus_total": total, **{k: v for k, v in s.items() if k in ("total", "by_root")}}
if a.skip_indexes:
result["indexes"] = "skipped"
print(json.dumps(result, indent=2) if a.json else
f"indexed {done}/{total} info + {fact_n} facts; db total {s['total']}")
print(json.dumps(result, indent=2) if a.json else f"indexed {done}/{total} leafs; db total {s['total']}")
return 0
+6 -4
View File
@@ -1,9 +1,10 @@
#!/usr/bin/env bash
# bin/kb/search — deprecated wrapper. Use bin/brain/search.go.
# CGO via Zig (bin/cgo/zig), not gcc. Builds a binary then execs it.
# Sets CGO for ladybug, builds a binary (embed daemon needs a real executable),
# then execs it. Prints one deprecation line.
set -euo pipefail
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
BIN="$ROOT/var/bin/brain-search"
SRC="$ROOT/internal/brain"
CMD="$ROOT/bin/brain"
@@ -23,10 +24,11 @@ else
fi
if [ "$need_build" -eq 1 ]; then
echo "Building brain/search (zig cc)..." >&2
echo "Building brain/search..." >&2
(
cd "$ROOT" &&
eval "$("$ROOT/bin/cgo/zig" env)" &&
CGO_CFLAGS="-I$ROOT/lib-ladybug" \
CGO_LDFLAGS="-L$ROOT/lib-ladybug -Wl,-rpath,$ROOT/lib-ladybug" \
go build -tags system_ladybug -o "$BIN" ./bin/brain
) || exit 1
fi
+77 -11
View File
@@ -7,7 +7,7 @@
bin/mail/import --since 2026-01-01 only messages after a date
bin/mail/import --limit 50 cap messages per run
bin/mail/import --no-attachments body only, skip attachment conversion
bin/mail/import --ocr OCR images (PDFs OCR when textless)
bin/mail/import --ocr OCR scanned PDFs/images via docling
bin/mail/import --dry-run list messages without writing anything
Writes one directory per message: var/mail/{folder}/{message_id}/
@@ -16,11 +16,10 @@ Writes one directory per message: var/mail/{folder}/{message_id}/
attachments/*.md converted attachment content
Indexing is a separate step (`bin/brain/index.go --rebuild`): conversion can
crash and must not leave the brain DB mid-transaction.
crash in native docling and must not leave the brain DB mid-transaction.
Requires ONLYOFFICE_URL/USER/PASS in .env (or env) except `--from-raw`.
Idempotent: a message already present (message.md exists) is skipped unless
--force.
Requires ONLYOFFICE_URL/USER/PASS in .env (or env). Idempotent: a message
already present (message.md exists) is skipped unless --force.
"""
from __future__ import annotations
@@ -42,10 +41,10 @@ from mailconv import ( # noqa: E402
IMAGE_SUFFIXES,
LEGACY_OFFICE_SUFFIXES,
TEXT_SUFFIXES,
convert_pdf,
html_to_markdown,
is_convertible,
normalize_markdown,
ocr_image,
subject_to_filename,
zip_extract_safe,
)
@@ -148,9 +147,9 @@ def convert_file_to_md(path: Path, ocr: bool) -> str | None:
except Exception as e:
return f"\n<!-- conversion failed: {e} -->\n"
if suffix == ".pdf":
return convert_pdf(path, ocr)
return _convert_pdf(path, ocr)
if suffix in IMAGE_SUFFIXES and ocr:
return ocr_image(path) or "\n<!-- ocr unavailable -->\n"
return _convert_pdf(path, ocr)
if suffix in LEGACY_OFFICE_SUFFIXES:
return _convert_legacy(path)
if suffix in ARCHIVE_SUFFIXES:
@@ -158,6 +157,67 @@ def convert_file_to_md(path: Path, ocr: bool) -> str | None:
return None
def _convert_pdf(path: Path, ocr: bool) -> str:
"""Convert one PDF to markdown.
Fast path: poppler's pdftotext (-layout) extracts exact text from
born-digital PDFs in ~15ms vs docling's 1-3s. Only textless PDFs (scanned
pages, layout-heavy) fall back to docling, which runs isolated in a
subprocess because its native onnx/RT-DETR has segfaulted the main process.
"""
text = _pdf_fast_text(path)
if ocr or text is None or not text.strip():
return _convert_pdf_docling(path, ocr)
return normalize_markdown(text)
def _pdf_fast_text(path: Path) -> str | None:
"""pdftotext -layout; None when poppler is unavailable (or the PDF has no text layer)."""
try:
proc = subprocess.run(
["pdftotext", "-layout", str(path), "-"],
capture_output=True, timeout=60)
except (OSError, subprocess.TimeoutExpired):
return None
if proc.returncode != 0:
return None
return proc.stdout.decode("utf-8", errors="replace")
def _convert_pdf_docling(path: Path, ocr: bool) -> str:
try:
proc = subprocess.run(
[sys.executable, os.path.abspath(__file__), "--pdf-worker", str(path),
"--ocr" if ocr else "--no-ocr"],
capture_output=True, text=True, timeout=600)
except subprocess.TimeoutExpired:
return "\n<!-- pdf conversion timed out -->\n"
if proc.returncode != 0:
tail = proc.stderr.strip().splitlines()[-3:]
return f"\n<!-- pdf conversion failed: {proc.returncode}: {' | '.join(tail)} -->\n"
return proc.stdout
def _pdf_worker(path: Path, ocr: bool) -> None:
"""docling worker entry: prints converted markdown on stdout, exits non-zero on error."""
try:
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.pipeline_options import PdfPipelineOptions
opts = PdfPipelineOptions()
opts.do_ocr = bool(ocr)
opts.do_table_structure = True
conv = DocumentConverter(format_options={"pdf": PdfFormatOption(pipeline_options=opts)})
res = conv.convert(str(path))
sys.stdout.write(normalize_markdown(res.document.export_to_markdown()))
sys.exit(0)
except Exception as e:
# errors/stacktraces to stderr; the caller only reports a one-liner
print(f"pdf-worker: {e}", file=sys.stderr)
import traceback
traceback.print_exc(file=sys.stderr)
sys.exit(1)
def _convert_legacy(path: Path) -> str:
"""Legacy .doc/.xls/.ppt -> md via pandoc (installed) or a stub."""
try:
@@ -296,12 +356,19 @@ def main(argv: list[str]) -> int:
p.add_argument("--from-raw", default="",
help="convert Go-synced dirs (var/mail/<folder>/<id>/message.json) to markdown")
p.add_argument("--no-attachments", action="store_true", help="skip attachment download+convert")
p.add_argument("--ocr", action="store_true", help="OCR images (PDFs OCR when textless)")
p.add_argument("--ocr", action="store_true", help="OCR scanned PDFs/images via docling")
p.add_argument("--force", action="store_true", help="re-import even if message.md exists")
p.add_argument("--dry-run", action="store_true", help="list messages, write nothing")
p.add_argument("--json", action="store_true")
p.add_argument("--pdf-worker", default="", help=argparse.SUPPRESS)
p.add_argument("--no-ocr", action="store_true", help=argparse.SUPPRESS)
a = p.parse_args(argv)
if a.pdf_worker:
_pdf_worker(Path(a.pdf_worker), ocr=not a.no_ocr)
return 0
conf = load_env()
fid = folder_id(a.folder)
out_root = ROOT / "var" / "mail"
summary: list[dict] = []
@@ -327,7 +394,6 @@ def main(argv: list[str]) -> int:
target_dir=msg_dir.parent))
summary.append(entry)
else:
conf = load_env()
OOCLIENT = OOClient(conf)
if a.id:
messages = [{"id": i} for i in a.id]
-48
View File
@@ -1,48 +0,0 @@
//usr/bin/env go run -tags=mail_ocr "$0" "$@"; exit
//go:build mail_ocr
//
// bin/mail/ocr.go - OCR an image or scanned PDF (tesseract eng+deu).
//
// ./bin/mail/ocr.go scan.png
// ./bin/mail/ocr.go scan.pdf
// OCR_ENGINE=paddle ./bin/mail/ocr.go scan.png
//
// PDFs try pdftotext -layout first; empty text layer uses pdftoppm + tesseract.
// No gocv. Tesseract CGO bindings are not used (D21 Zig owns Ladybug CGO).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"strings"
"github.com/eSlider/2dph/internal/ocr"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
if len(args) != 1 || strings.HasPrefix(args[0], "-") {
fmt.Fprintln(os.Stderr, `usage: bin/mail/ocr.go <image|pdf>`)
return 2
}
path := args[0]
var (
text string
err error
)
if strings.HasSuffix(strings.ToLower(path), ".pdf") {
text, err = ocr.PDFFile(path)
} else {
text, err = ocr.ImageFile(path)
}
if err != nil {
fmt.Fprintf(os.Stderr, "mail/ocr: %v\n", err)
return 1
}
fmt.Println(text)
return 0
}
+5 -85
View File
@@ -1,100 +1,20 @@
//usr/bin/env go run "$0" "$@"; exit
//usr/bin/env go run -tags=markdown_import "$0" "$@"; exit
//go:build markdown_import
//
// bin/markdown/import.go - split markdown into leafs (H2 boundaries).
// bin/markdown/import.go - split markdown into leafs (mistune).
//
// ./bin/markdown/import.go [dir]
// ./bin/markdown/import.go --files a.md,b.md --json
//
// Conversion only. Brain write is bin/brain/index.go.
// Python bin/md/import remains as a fallback.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"strings"
"github.com/eSlider/2dph/internal/mdleaves"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
jsonOut := false
files := ""
root := "."
for i := 0; i < len(args); i++ {
a := args[i]
switch {
case a == "--json":
jsonOut = true
case a == "--files" && i+1 < len(args):
i++
files = args[i]
case strings.HasPrefix(a, "--files="):
files = strings.TrimPrefix(a, "--files=")
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, "bin/markdown/import.go [dir] [--files a.md,b.md] [--json]")
return 0
case strings.HasPrefix(a, "-"):
fmt.Fprintln(os.Stderr, "unknown arg:", a)
return 2
default:
root = a
}
}
var paths []string
if files != "" {
for _, f := range strings.Split(files, ",") {
f = strings.TrimSpace(f)
if f != "" {
paths = append(paths, f)
}
}
} else {
st, err := os.Stat(root)
if err != nil {
fmt.Fprintf(os.Stderr, "md/import: no such path %s\n", root)
return 2
}
if !st.IsDir() {
paths = []string{root}
} else {
var err error
paths, err = mdleaves.WalkMarkdown(root)
if err != nil {
fmt.Fprintf(os.Stderr, "md/import: %v\n", err)
return 1
}
}
}
if len(paths) == 0 {
fmt.Fprintln(os.Stderr, "md/import: no markdown files")
return 1
}
var all []mdleaves.Leaf
for _, p := range paths {
raw, err := os.ReadFile(p)
if err != nil {
fmt.Fprintf(os.Stderr, "md/import: %s: %v\n", p, err)
continue
}
all = append(all, mdleaves.ToAll(string(raw), p, "")...)
}
if jsonOut {
s, err := mdleaves.EncodeJSON(all)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
fmt.Print(s)
return 0
}
fmt.Print(mdleaves.EncodeYAML(all))
return 0
os.Exit(cmdbin.ExecFile("bin/md/import", os.Args[1:]))
}
-73
View File
@@ -1,73 +0,0 @@
//usr/bin/env go run -tags=qa_stats "$0" "$@"; exit
//go:build qa_stats
//
// bin/qa/stats.go - DuckDB quantiles over a JSON number array or JSONL count.
//
// ./bin/qa/stats.go <<< '[1,2,3,4,5]'
// ./bin/qa/stats.go --jsonl rows.jsonl
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
// DuckDB CGO needs gcc/g++ (not Zig). After eval "$(bin/cgo/zig env)":
// CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= ./bin/qa/stats.go
package main
import (
"encoding/json"
"fmt"
"io"
"os"
"strings"
"github.com/eSlider/2dph/internal/duckstats"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
jsonl := ""
for i := 0; i < len(args); i++ {
a := args[i]
switch {
case a == "--jsonl" && i+1 < len(args):
i++
jsonl = args[i]
case strings.HasPrefix(a, "--jsonl="):
jsonl = strings.TrimPrefix(a, "--jsonl=")
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, "bin/qa/stats.go [--jsonl FILE] # stdin = JSON [float,…]")
return 0
default:
fmt.Fprintln(os.Stderr, "unknown arg:", a)
return 2
}
}
if jsonl != "" {
n, err := duckstats.CountJSONL(jsonl)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
fmt.Printf("n: %d\n", n)
return 0
}
raw, err := io.ReadAll(os.Stdin)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
var samples []float64
if err := json.Unmarshal(raw, &samples); err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
s, err := duckstats.Quantiles(samples)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
fmt.Printf("n: %d\nmin: %g\np50: %g\np95: %g\nmax: %g\navg: %g\n",
s.N, s.Min, s.P50, s.P95, s.Max, s.Avg)
return 0
}
-102
View File
@@ -1,102 +0,0 @@
//usr/bin/env go run -tags=reasoner_bakeoff "$0" "$@"; exit
//go:build reasoner_bakeoff
//
// bin/reasoner/bakeoff.go - CPU tool-call bake-off against an OpenAI-compatible URL (D18).
//
// REASONER_BASE_URL=http://127.0.0.1:11435/v1 REASONER_MODEL=qwen3.5:9b ./bin/reasoner/bakeoff.go
// ./bin/reasoner/bakeoff.go --model MichelRosselli/bonsai-27b:Q1_0 --json
//
// Measures OpenAI tool_calls (search/get/audit) and RSS from Ollama /api/ps, not VRAM.
// PicoClaw is compose profile picoclaw; tool names match internal/httpapi MCP ops.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"encoding/json"
"fmt"
"os"
"strings"
"github.com/eSlider/2dph/internal/duckstats"
"github.com/eSlider/2dph/internal/reasoner"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
base := os.Getenv("REASONER_BASE_URL")
if base == "" {
base = "http://127.0.0.1:11435/v1"
}
model := os.Getenv("REASONER_MODEL")
if model == "" {
model = reasoner.OllamaRAM
}
jsonOut := false
device := "cpu"
for i := 0; i < len(args); i++ {
a := args[i]
switch {
case a == "--json":
jsonOut = true
case a == "--model" && i+1 < len(args):
i++
model = args[i]
case strings.HasPrefix(a, "--model="):
model = strings.TrimPrefix(a, "--model=")
case a == "--base-url" && i+1 < len(args):
i++
base = args[i]
case a == "--device" && i+1 < len(args):
i++
device = args[i]
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, "bin/reasoner/bakeoff.go [--model ID] [--base-url URL] [--device cpu] [--json]")
return 0
default:
fmt.Fprintln(os.Stderr, "unknown arg:", a)
return 2
}
}
c := reasoner.Client{BaseURL: base, Model: model, Device: device}
rep := reasoner.Run(c)
lat := make([]float64, 0, len(rep.Prompts))
for _, p := range rep.Prompts {
lat = append(lat, float64(p.LatencyMS))
}
if st, err := duckstats.Quantiles(lat); err == nil {
rep.LatencyP50MS = st.P50
rep.LatencyP95MS = st.P95
}
raw, err := json.MarshalIndent(rep, "", " ")
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
if jsonOut {
fmt.Println(string(raw))
} else {
fmt.Printf("model: %s\n", rep.Model)
fmt.Printf("hf_id: %s\n", rep.HF)
fmt.Printf("device: %s\n", rep.Device)
fmt.Printf("tool_call: %d/%d\n", rep.ToolCallOK, rep.ToolCallN)
fmt.Printf("xml_leak: %d\n", rep.XMLLeak)
fmt.Printf("rss_mb: %d\n", rep.RSSMB)
fmt.Printf("vram_mb: %d\n", rep.VRAMMB)
fmt.Printf("latency_p50_ms: %g\n", rep.LatencyP50MS)
fmt.Printf("latency_p95_ms: %g\n", rep.LatencyP95MS)
for _, p := range rep.Prompts {
status := "fail"
if p.OK {
status = "ok"
}
fmt.Printf(" %s: %s wanted=%s got=%s xml=%v %dms %s\n", p.WantedTool, status, p.WantedTool, p.ToolName, p.XMLLeak, p.LatencyMS, p.Err)
}
}
if rep.ToolCallN == 0 {
return 1
}
return 0
}
-2
View File
@@ -1,2 +0,0 @@
// Commands in this directory are shebang mains (bakeoff.go).
package main
+1 -92
View File
@@ -4,7 +4,7 @@ Single embedded graph `var/kb.lbug`. Two roots: facts (assertions backed by
>=2 independent sources) and info (narrative leafs). Hybrid retrieval: BM25
(FTS extension) + HNSW cosine (VECTOR extension) + Cypher graph hops.
All access is read-only unless `--rebuild` (kb/index) or `kb/add`.
All access is read-only unless `--rebuild` is passed to kb/index.
"""
from __future__ import annotations
@@ -118,97 +118,6 @@ def upsert_leaf(conn: ladybug.Connection, *, text: str, root: str, confidence: s
return lid
def add_leafs(conn: ladybug.Connection, leafs: list[dict]) -> list[str]:
"""Write facts+info leafs in one transaction. Safe while FTS/HNSW exist.
Each leaf dict: text, source, optional root/confidence/source_rev/how/loc/type/embedding.
Does not delete the database file. Measured on Ladybug 0.19: MERGE of new
ids (and updates) stays FTS+HNSW queryable; DROP INDEX is the fatal path.
"""
if not leafs:
return []
started = False
try:
conn.execute("BEGIN TRANSACTION")
started = True
except Exception:
started = False
ids: list[str] = []
try:
for lf in leafs:
ids.append(
upsert_leaf(
conn,
text=str(lf["text"]),
root=str(lf.get("root") or ROOT_INFO),
confidence=str(lf.get("confidence") or CONF_CONFIRMED),
source=str(lf["source"]),
source_rev=str(lf.get("source_rev") or "working-tree"),
how=str(lf.get("how") or "brain/add"),
loc=str(lf.get("loc") or lf.get("source") or ""),
type_=str(lf.get("type") or lf.get("type_") or "reference"),
embedding=lf.get("embedding"),
)
)
if started:
conn.execute("COMMIT")
except Exception:
if started:
try:
conn.execute("ROLLBACK")
except Exception:
pass
raise
return ids
def file_id(repo: str, path: str) -> str:
"""Stable File.id matching gitimport (`repo:path`)."""
return f"{repo}:{path}" if repo else path
def link_from_file(conn: ladybug.Connection, leaf_id: str, path: str,
repo: str = "", mtime: str = "") -> str:
"""MERGE File and Leaf-[:FROM_FILE]->File so --hop 1 can walk."""
fid = file_id(repo, path)
conn.execute(
"MERGE (f:File {id:$id}) SET f.path=$path, f.repo=$repo, f.mtime=$mtime",
parameters={"id": fid, "path": path, "repo": repo, "mtime": mtime},
)
conn.execute(
"MATCH (l:Leaf {id:$lid}), (f:File {id:$fid}) "
"MERGE (l)-[:FROM_FILE]->(f)",
parameters={"lid": leaf_id, "fid": fid},
)
return fid
HOP_STMTS = {
1: "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File) RETURN f.id, f.path, 1",
2: ("MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit) "
"RETURN c.id, c.subject, 2"),
3: ("MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit)"
"-[:AUTHORED]->(p:Person) RETURN p.id, p.name, 3"),
}
HOP_LABELS = {1: "File", 2: "Commit", 3: "Person"}
def hop_walk(conn: ladybug.Connection, leaf_id: str, n: int) -> list[dict]:
"""Walk Leaf → File → Commit → Person up to n hops (max 3)."""
depth = min(max(int(n), 0), 3)
out: list[dict] = []
for d in range(1, depth + 1):
rows = conn.execute(HOP_STMTS[d], parameters={"id": leaf_id}).get_all()
for row in rows:
out.append({
"id": row[0],
"label": HOP_LABELS[d],
"name": row[1],
"depth": int(row[2]),
})
return out
def leaf_index_names(conn: ladybug.Connection) -> set[str]:
"""Return index names on the Leaf table (e.g. {'id', 'Leaf_vec', '_PK'})."""
rows = conn.execute("CALL SHOW_INDEXES() RETURN *").get_all()
+1 -84
View File
@@ -7,10 +7,7 @@ offline against fixtures.
from __future__ import annotations
import html
import os
import re
import subprocess
import tempfile
import zipfile
from pathlib import Path
@@ -21,10 +18,9 @@ OFFICE_SUFFIXES = {".docx", ".pptx", ".xlsx", ".html", ".htm", ".epub", ".eml",
PDF_SUFFIXES = {".pdf"}
IMAGE_SUFFIXES = {".png", ".jpg", ".jpeg", ".gif", ".bmp", ".tiff", ".tif", ".webp"}
ARCHIVE_SUFFIXES = {".zip"}
# Legacy binary Office (doc/xls/ppt) — markitdown skip them; we try
# Legacy binary Office (doc/xls/ppt) — markitdown/docling skip them; we try
# pandoc first, else leave a stub.
LEGACY_OFFICE_SUFFIXES = {".doc", ".xls", ".ppt"}
TESS_LANG = "eng+deu"
CONVERTIBLE_SUFFIXES = (
TEXT_SUFFIXES | OFFICE_SUFFIXES | PDF_SUFFIXES | IMAGE_SUFFIXES | ARCHIVE_SUFFIXES | LEGACY_OFFICE_SUFFIXES
@@ -150,82 +146,3 @@ def zip_extract_safe(zip_path: Path, dest: Path) -> list[Path]:
def is_convertible(suffix: str) -> bool:
return suffix.lower() in CONVERTIBLE_SUFFIXES
def convert_pdf(path: Path, ocr: bool = False) -> str:
"""pdftotext -layout first; empty text layer → pdftoppm + tesseract.
`ocr` is unused for born-digital PDFs (text layer wins). Scans OCR
automatically. This path never execs an ONNX document converter.
"""
del ocr # scans OCR when the text layer is empty; flag is for images
text = pdf_fast_text(path)
if text and text.strip():
return normalize_markdown(text)
scanned = ocr_pdf(path)
if scanned and scanned.strip():
return normalize_markdown(scanned)
if text:
return normalize_markdown(text)
return "\n<!-- pdf has no text layer (ocr unavailable) -->\n"
def pdf_fast_text(path: Path) -> str | None:
"""pdftotext -layout; None when poppler is missing or the command fails."""
try:
proc = subprocess.run(
["pdftotext", "-layout", str(path), "-"],
capture_output=True, timeout=60)
except (OSError, subprocess.TimeoutExpired):
return None
if proc.returncode != 0:
return None
return proc.stdout.decode("utf-8", errors="replace")
def ocr_pdf(path: Path) -> str:
"""Rasterize with pdftoppm and OCR each page (tesseract or paddle)."""
try:
with tempfile.TemporaryDirectory(prefix="2dph-ocr-") as tmp:
prefix = str(Path(tmp) / "page")
proc = subprocess.run(
["pdftoppm", "-png", "-r", "200", str(path), prefix],
capture_output=True, timeout=120)
if proc.returncode != 0:
return ""
pages = sorted(Path(tmp).glob("page*.png"))
parts = [ocr_image(p) for p in pages]
return "\n\n".join(p for p in parts if p and p.strip())
except (OSError, subprocess.TimeoutExpired):
return ""
def ocr_image(path: Path) -> str:
engine = os.environ.get("OCR_ENGINE", "tesseract")
if engine == "paddle":
return _ocr_paddle(path)
return _ocr_tesseract(path)
def _ocr_tesseract(path: Path) -> str:
try:
proc = subprocess.run(
["tesseract", str(path), "stdout", "-l", TESS_LANG, "--psm", "6"],
capture_output=True, timeout=120)
except (OSError, subprocess.TimeoutExpired):
return ""
if proc.returncode != 0:
return ""
return proc.stdout.decode("utf-8", errors="replace").strip()
def _ocr_paddle(path: Path) -> str:
try:
proc = subprocess.run(
["paddleocr", "ocr", "-i", str(path)],
capture_output=True, timeout=180)
except (OSError, subprocess.TimeoutExpired):
return ""
if proc.returncode != 0:
return ""
return proc.stdout.decode("utf-8", errors="replace").strip()
+2 -165
View File
@@ -1,7 +1,6 @@
"""D14 layout: bin/{subject}/{method}.go, libs in internal/, one go.mod."""
from __future__ import annotations
import os
import unittest
from pathlib import Path
@@ -75,58 +74,9 @@ class BinLayoutTest(unittest.TestCase):
)
def test_brain_methods_are_shebangs(self) -> None:
for method in ("index.go", "add.go", "get.go", "stats.go", "eval.go", "watch.go"):
for method in ("index.go", "get.go", "stats.go", "eval.go", "watch.go"):
self._assert_shebang(f"bin/brain/{method}")
def test_brain_add_is_python_write_not_rebuild(self) -> None:
self._assert_shebang("bin/brain/add.go")
text = (ROOT / "bin" / "brain" / "add.go").read_text()
self.assertIn("cmdbin.ExecFile", text)
self.assertIn("bin/kb/add", text)
self.assertNotIn("--rebuild", text)
py = (ROOT / "bin" / "kb" / "add").read_text()
self.assertIn("add_leafs", py)
self.assertIn("--json", py)
self.assertNotIn("unlink", py.lower())
def test_brain_get_stats_eval_are_not_python_exec(self) -> None:
for method in ("get.go", "stats.go", "eval.go"):
text = (ROOT / "bin" / "brain" / method).read_text()
self.assertNotIn(
"ExecFile",
text,
f"bin/brain/{method} must call internal/brain, not ExecFile Python",
)
self.assertNotIn(
"cmdbin",
text,
f"bin/brain/{method} must not import internal/cmdbin",
)
self.assertIn(
"system_ladybug",
text.splitlines()[0],
f"bin/brain/{method} shebang must pass -tags=system_ladybug",
)
self.assertIn(
"github.com/eSlider/2dph/internal/brain",
text,
)
def test_eval_control_questions_live_in_rank(self) -> None:
rank = (ROOT / "internal" / "brain" / "rank" / "evalq.go").read_text()
py = (ROOT / "bin" / "kb" / "eval").read_text()
for frag in ("BM25", "DevOps", "LadybugDB"):
self.assertIn(frag, rank)
self.assertIn(frag, py)
self.assertIn("0.95", rank)
def test_facts_methods_are_shebangs(self) -> None:
for method in ("audit.go", "extract.go", "crm.go"):
self._assert_shebang(f"bin/facts/{method}")
text = (ROOT / "bin" / "facts" / method).read_text()
self.assertIn("cmdbin.ExecFile", text)
self.assertIn(f"bin/facts/{method.removesuffix('.go')}", text)
def test_mail_import_is_shebang_not_brain_write(self) -> None:
self._assert_shebang("bin/mail/import.go")
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
@@ -136,55 +86,8 @@ class BinLayoutTest(unittest.TestCase):
"index_mail must point at bin/brain/index.go",
)
def test_mail_ocr_is_tesseract_not_docling(self) -> None:
self._assert_shebang("bin/mail/ocr.go")
ocr = (ROOT / "bin" / "mail" / "ocr.go").read_text()
self.assertIn("internal/ocr", ocr)
self.assertIn("mail_ocr", ocr)
self.assertNotIn("github.com/otiai10/gosseract", ocr)
py = (ROOT / "bin" / "mail" / "import").read_text()
self.assertNotIn("from docling", py)
self.assertNotIn("import docling", py)
self.assertIn("convert_pdf", py)
conv = (ROOT / "bin" / "tools" / "mailconv.py").read_text()
self.assertIn("pdftotext", conv)
self.assertIn("pdftoppm", conv)
self.assertIn("tesseract", conv)
self.assertIn("eng+deu", conv)
self.assertNotIn("from docling", conv)
self.assertNotIn("import docling", conv)
self.assertNotIn("gocv", conv.lower())
proj = (ROOT / "pyproject.toml").read_text()
self.assertNotIn("docling", proj)
ci = (ROOT / ".github" / "workflows" / "ci.yml").read_text()
self.assertIn("tesseract-ocr", ci)
self.assertIn("./internal/ocr", ci)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn("ocr-paddle", compose)
self.assertIn("OCR_ENGINE", compose)
def test_markdown_import_is_go_not_python_exec(self) -> None:
def test_markdown_import_is_shebang(self) -> None:
self._assert_shebang("bin/markdown/import.go")
text = (ROOT / "bin" / "markdown" / "import.go").read_text()
self.assertNotIn("ExecFile", text)
self.assertNotIn("cmdbin", text)
self.assertIn("internal/mdleaves", text)
self.assertNotIn("kb.lbug", text)
def test_import_adapters_do_not_write_ladybug(self) -> None:
for rel in (
"bin/mail/import.go",
"bin/mail/import",
"bin/markdown/import.go",
"bin/chats/import.go",
"bin/git/import.go",
):
text = (ROOT / rel).read_text()
self.assertNotIn("upsert_leaf", text, rel)
self.assertNotIn("kb.lbug", text, rel)
self.assertNotIn("var/brain.lbug", text, rel)
index = (ROOT / "bin" / "brain" / "index.go").read_text()
self.assertIn("bin/kb/index", index)
def test_postgres_query_is_shebang(self) -> None:
self._assert_shebang("bin/postgres/query.go")
@@ -216,69 +119,3 @@ class BinLayoutTest(unittest.TestCase):
for line in first.splitlines():
if "go-git/go-git" in line:
self.assertNotIn("indirect", line)
def test_duckdb_go_is_direct_require(self) -> None:
text = (ROOT / "go.mod").read_text()
first = text.split("require (")[1].split(")")[0]
self.assertRegex(first, r"github.com/duckdb/duckdb-go/v2\s+v")
for line in first.splitlines():
if "duckdb/duckdb-go" in line:
self.assertNotIn("indirect", line)
skill = (ROOT / "skills" / "duckdb" / "SKILL.md").read_text()
self.assertIn("github.com/duckdb/duckdb-go", skill)
self.assertIn("Ladybug", skill)
self.assertIn("sqlite", skill.lower())
self.assertIn("gcc", skill.lower())
self.assertIn("Zig", skill)
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D22", plan)
self.assertIn("duckdb-go", plan)
self._assert_shebang("bin/qa/stats.go")
reasoner = (ROOT / "internal" / "reasoner" / "client.go").read_text()
self.assertNotIn("duckdb", reasoner)
self.assertNotIn("duckstats", reasoner)
bakeoff = (ROOT / "bin" / "reasoner" / "bakeoff.go").read_text()
self.assertIn("internal/duckstats", bakeoff)
webcache = (ROOT / "internal" / "websearch" / "cache.go").read_text()
self.assertNotIn("duckdb", webcache)
self.assertIn("modernc.org/sqlite", webcache)
def test_cgo_uses_zig_not_gcc(self) -> None:
for rel in ("bin/cgo/zig", "bin/cgo/zcc", "bin/cgo/zc++"):
p = ROOT / rel
self.assertTrue(p.is_file(), f"missing {rel}")
self.assertTrue(
os.access(p, os.X_OK),
f"{rel} must be executable",
)
zig = (ROOT / "bin" / "cgo" / "zig").read_text()
self.assertIn("zig cc", zig)
self.assertIn("0.14.1", zig)
zcc = (ROOT / "bin" / "cgo" / "zcc").read_text()
self.assertIn('exec "$ZIG" cc', zcc)
self.assertNotIn("command -v gcc", zcc)
search = (ROOT / "bin" / "kb" / "search").read_text()
self.assertIn("bin/cgo/zig", search)
self.assertNotIn("command -v gcc", search)
def test_ci_recall_sot_is_zig_brain_eval(self) -> None:
ci = (ROOT / ".github" / "workflows" / "ci.yml").read_text()
self.assertIn("bin/brain/eval.go", ci)
self.assertIn("system_ladybug,brain_eval", ci)
self.assertIn("/tmp/brain-eval", ci)
self.assertIn("KB_ROOT", ci)
self.assertNotIn("bin/kb/eval", ci)
self.assertNotIn("gate skipped", ci)
self.assertIn("./bin/facts/audit self", ci)
def test_eval_fragments_live_in_default_corpus(self) -> None:
"""CI --rebuild indexes README/PLAN/docs/skills; fragments must be there."""
corpus = []
for rel in ("README.md", "PLAN.md", "AGENTS.md"):
corpus.append((ROOT / rel).read_text())
for d in ("docs", "skills"):
for p in (ROOT / d).rglob("*.md"):
corpus.append(p.read_text())
blob = "\n".join(corpus)
for frag in ("BM25", "DevOps", "LadybugDB"):
self.assertIn(frag, blob, f"{frag} must appear in default index corpus")
-97
View File
@@ -1,97 +0,0 @@
"""Import adapters write files only. Index rebuild is brain/index (D14 / Gitea #7)."""
from __future__ import annotations
import json
import os
import subprocess
import sys
import tempfile
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class IndexAdapterTest(unittest.TestCase):
def test_dry_run_fixture_corpus_does_not_write_lbug(self) -> None:
tmp = Path(tempfile.mkdtemp())
(tmp / "note.md").write_text("# Fixture\n\n## Leaf\n\nhello corpus\n", encoding="utf-8")
lbug = tmp / "kb.lbug"
try:
import ladybug # noqa: F401
except ImportError:
venv_py = ROOT / ".venv" / "bin" / "python"
if not venv_py.is_file():
self.skipTest("ladybug missing")
py = str(venv_py)
else:
py = sys.executable
proc = subprocess.run(
[py, str(ROOT / "bin" / "kb" / "index"), "--dry-run", "--json", "--corpus", str(tmp)],
cwd=ROOT,
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
msg = json.loads(proc.stdout)
self.assertTrue(msg.get("dry_run"))
self.assertGreaterEqual(msg.get("corpus_total", 0), 1)
self.assertFalse(lbug.exists(), "dry-run must not create a Ladybug file")
def test_facts_json_and_chats_land_on_rebuild(self) -> None:
"""Gitea #18: facts (2-source) + chats markdown become leafs on rebuild."""
tmp = Path(tempfile.mkdtemp())
dbpath = tmp / "kb.lbug"
chats = tmp / "chats"
chats.mkdir()
(chats / "alice.md").write_text(
"# Chat\n\n## Alice and Bob\n\nhello from chats fixture unique-chat-token\n",
encoding="utf-8",
)
facts_path = tmp / "facts.json"
facts_path.write_text(json.dumps([{
"text": "container 'brain' unique-fact-token is running and declared in compose.yaml",
"source": "docker ps x compose.yaml",
"loc": "compose.yaml:brain",
"how": "facts/extract",
}]), encoding="utf-8")
venv_py = ROOT / ".venv" / "bin" / "python"
py = str(venv_py) if venv_py.is_file() else sys.executable
proc = subprocess.run(
[
py, str(ROOT / "bin" / "kb" / "index"),
"--rebuild", "--db", str(dbpath), "--no-defaults",
"--with-chats", str(chats),
"--facts-json", str(facts_path),
"--json",
],
cwd=ROOT,
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
msg = json.loads(proc.stdout)
self.assertGreaterEqual(msg.get("facts_leafs", 0), 1)
self.assertGreaterEqual(msg.get("chat_leafs", 0), 1)
self.assertTrue(dbpath.exists())
sys.path.insert(0, str(ROOT / "bin" / "tools"))
import kblib
db, conn = kblib.connect(dbpath, read_only=True)
try:
stats = kblib.stats(conn)
self.assertGreaterEqual(stats["by_root"].get("facts", 0), 1)
fts = kblib.query_fts(conn, "unique-chat-token", 5)
self.assertTrue(fts, "chats markdown must be FTS-searchable")
fact_hits = kblib.query_fts(conn, "unique-fact-token", 5)
self.assertTrue(any(h.get("root") == "facts" for h in fact_hits))
src = conn.execute(
"MATCH (l:Leaf {root:'facts'}) RETURN l.source"
).get_all()
self.assertTrue(any(" x " in str(r[0]) for r in src))
finally:
conn.close()
db.close()
-69
View File
@@ -1,69 +0,0 @@
"""Incremental add writes leafs without deleting kb.lbug."""
from __future__ import annotations
import json
import os
import subprocess
import sys
import tempfile
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class KbAddCLITest(unittest.TestCase):
def test_json_add_does_not_delete_db(self) -> None:
tmp = Path(tempfile.mkdtemp())
dbpath = tmp / "kb.lbug"
py = sys.executable
venv_py = ROOT / ".venv" / "bin" / "python"
if venv_py.is_file():
py = str(venv_py)
payload = {
"text": "cli zebra leaf",
"root": "info",
"source": "cli-test",
"confidence": "confirmed",
"how": "test",
"loc": str(tmp),
"type": "reference",
"embedding": [0.0] * 256,
}
payload["embedding"][0] = 0.3
proc = subprocess.run(
[py, str(ROOT / "bin" / "kb" / "add"), "--db", str(dbpath), "--json"],
cwd=ROOT,
input=json.dumps(payload),
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
self.assertTrue(dbpath.exists(), "add must create the db, not skip write")
out = json.loads(proc.stdout)
self.assertEqual(out.get("mode"), "add")
self.assertEqual(len(out.get("ids") or []), 1)
again = subprocess.run(
[py, str(ROOT / "bin" / "kb" / "add"), "--db", str(dbpath), "--json"],
cwd=ROOT,
input=json.dumps({
**payload,
"text": "second moose leaf",
"source": "cli-test-2",
}),
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(again.returncode, 0, again.stderr)
self.assertTrue(dbpath.exists())
second = json.loads(again.stdout)
self.assertEqual(len(second.get("ids") or []), 1)
self.assertNotEqual(out["ids"][0], second["ids"][0])
if __name__ == "__main__":
unittest.main()
-96
View File
@@ -71,69 +71,6 @@ class KblibTest(unittest.TestCase):
self.assertTrue(hits)
self.assertIn("Leaf_vec", kblib.leaf_index_names(self.conn))
def test_add_after_indexes_keeps_fts_queryable(self):
"""Incremental add after FTS+HNSW must find the new leaf on both indexes."""
kblib.upsert_leaf(self.conn, text="seed fox leaf", root="info",
confidence="confirmed", source="s", source_rev="r1",
how="test", loc="/tmp", type_="reference",
embedding=make_emb(0.1))
kblib.ensure_indexes(self.conn)
ids = kblib.add_leafs(self.conn, [{
"text": "added zebra after index",
"root": "facts",
"confidence": "confirmed",
"source": "a.md x b.md",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "fact",
"embedding": make_emb(0.9),
}])
self.assertEqual(len(ids), 1)
fts = kblib.query_fts(self.conn, "zebra", 5)
self.assertTrue(fts)
self.assertIn("zebra", fts[0]["text"])
self.assertEqual(fts[0]["root"], "facts")
vec = kblib.query_vector(self.conn, make_emb(0.9), 5)
self.assertTrue(any("zebra" in h["text"] for h in vec))
fox = kblib.query_fts(self.conn, "fox", 5)
self.assertTrue(fox)
self.assertIn("fox", fox[0]["text"])
def test_add_facts_and_info_one_transaction(self):
"""D12: facts and info land in the same transaction."""
kblib.ensure_indexes(self.conn)
ids = kblib.add_leafs(self.conn, [
{
"text": "tx fact leaf two-source",
"root": "facts",
"confidence": "confirmed",
"source": "compose.yml x docker ps",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "fact",
"embedding": make_emb(0.4),
},
{
"text": "tx info narrative",
"root": "info",
"confidence": "confirmed",
"source": "note.md",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "reference",
"embedding": make_emb(0.5),
},
])
self.assertEqual(len(ids), 2)
stats = kblib.stats(self.conn)
self.assertEqual(stats["by_root"].get("facts"), 1)
self.assertEqual(stats["by_root"].get("info"), 1)
self.assertTrue(kblib.query_fts(self.conn, "two-source", 5))
self.assertTrue(kblib.query_fts(self.conn, "narrative", 5))
def test_drop_vector_then_create_raises_clear_error(self):
"""DROP INDEX leaves ghost catalog; create_fts_and_vector must raise."""
kblib.upsert_leaf(self.conn, text="seed", root="info",
@@ -161,39 +98,6 @@ class KblibTest(unittest.TestCase):
self.assertEqual(stats["total"], 2)
self.assertEqual(stats["by_root"], {"facts": 1, "info": 1})
def test_hop_1_returns_file_hop_3_reaches_person(self):
"""--hop walks FROM_FILE / HAS_VERSION / AUTHORED (Gitea #17)."""
import gitimport
lid = kblib.upsert_leaf(
self.conn, text="readme hop fixture", root="info",
confidence="confirmed", source="README.md", source_rev="r1",
how="test", loc="README.md", type_="reference",
embedding=make_emb(0.3),
)
kblib.link_from_file(self.conn, lid, "README.md", repo="sample-repo")
gitimport.index_commits(self.conn, [gitimport.Commit(
sha="a1b2c3d",
author="Ada Lovelace",
email="ada@example.com",
date="2026-08-10T12:00:00Z",
subject="feat: first commit",
files=["README.md"],
)], "sample-repo")
hop1 = kblib.hop_walk(self.conn, lid, 1)
self.assertEqual(len(hop1), 1)
self.assertEqual(hop1[0]["label"], "File")
self.assertEqual(hop1[0]["name"], "README.md")
self.assertEqual(hop1[0]["depth"], 1)
hop3 = kblib.hop_walk(self.conn, lid, 3)
labels = {n["label"] for n in hop3}
self.assertIn("File", labels)
self.assertIn("Commit", labels)
self.assertIn("Person", labels)
person = [n for n in hop3 if n["label"] == "Person"][0]
self.assertEqual(person["name"], "Ada Lovelace")
self.assertEqual(person["depth"], 3)
if __name__ == "__main__":
unittest.main()
-89
View File
@@ -8,13 +8,10 @@ from pathlib import Path
sys.path.insert(0, os.path.dirname(__file__))
from mailconv import ( # noqa: E402
TESS_LANG,
clean_email_address,
convert_pdf,
html_to_markdown,
is_convertible,
normalize_markdown,
ocr_image,
split_zip_members,
subject_to_filename,
zip_extract_safe,
@@ -103,92 +100,6 @@ class TestMailConv(unittest.TestCase):
self.assertFalse(is_convertible(".exe"))
self.assertFalse(is_convertible(".unknown"))
def test_convert_pdf_prefers_pdftotext(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b"Invoice BM25 layout"
stderr = b""
return P()
self._patch_run(mc, fake_run)
out = convert_pdf(Path(self._tmp("born.pdf")))
self.assertIn("BM25", out)
self.assertEqual(calls[0][:2], ["pdftotext", "-layout"])
self.assertFalse(any(c[0] == "tesseract" for c in calls))
self.assertFalse(any(c[0] == "pdftoppm" for c in calls))
def test_convert_pdf_empty_layer_uses_pdftoppm_tesseract(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b""
stderr = b""
if cmd[0] == "pdftotext":
P.stdout = b" \n"
return P()
if cmd[0] == "pdftoppm":
prefix = Path(cmd[-1])
(prefix.parent / "page-1.png").write_bytes(b"fake")
return P()
if cmd[0] == "tesseract":
P.stdout = b"scanned HELLO"
return P()
return P()
self._patch_run(mc, fake_run)
out = convert_pdf(Path(self._tmp("scan.pdf")))
self.assertIn("HELLO", out)
bins = [c[0] for c in calls]
self.assertIn("pdftotext", bins)
self.assertIn("pdftoppm", bins)
self.assertIn("tesseract", bins)
tess = next(c for c in calls if c[0] == "tesseract")
self.assertIn(TESS_LANG, tess)
self.assertNotIn("docling", " ".join(bins))
def test_ocr_image_paddle_engine(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b"paddle text"
stderr = b""
return P()
self._patch_run(mc, fake_run)
os.environ["OCR_ENGINE"] = "paddle"
try:
out = ocr_image(Path(self._tmp("x.png")))
finally:
os.environ.pop("OCR_ENGINE", None)
self.assertEqual(out, "paddle text")
self.assertEqual(calls[0][:2], ["paddleocr", "ocr"])
def _patch_run(self, mod, fn) -> None:
self.addCleanup(setattr, mod.subprocess, "run", mod.subprocess.run)
mod.subprocess.run = fn
def _mk_zip(self, members):
zpath = Path(self._tmp("arc.zip"))
with zipfile.ZipFile(zpath, "w") as zf:
+10 -121
View File
@@ -1,6 +1,7 @@
"""Published docs must match live commands (Gitea SoT, brain/search)."""
"""Published docs must match live commands (Gitea SoT, brain/search, no fake --hop)."""
from __future__ import annotations
import re
import unittest
from pathlib import Path
@@ -58,131 +59,19 @@ class PublishedDocsTest(unittest.TestCase):
self.assertNotIn("password", settings.lower())
self.assertIn("json", settings)
def test_picoclaw_compose_profile_has_mcp_example(self) -> None:
compose = (ROOT / "compose.yaml").read_text()
self.assertIn('profiles: ["picoclaw"]', compose)
self.assertIn("127.0.0.1:8630", compose)
example = (ROOT / "deploy" / "picoclaw" / "mcp.json.example").read_text()
self.assertIn("127.0.0.1:8630/mcp", example)
self.assertNotIn("password", example.lower())
self.assertNotIn("token", example.lower())
docs = (ROOT / "docs" / "picoclaw.md").read_text()
self.assertIn("search", docs)
self.assertIn("throttled", docs)
def test_readme_read_path_is_go(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("get.go", plan)
self.assertIn("CI fallback", plan)
design = (ROOT / "docs" / "design.md").read_text()
self.assertIn("internal/brain/rank", design)
self.assertIn("They do not exec Python", design)
def test_openapi_mcp_from_same_handlers(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D20", plan)
self.assertIn("/openapi.json", (ROOT / "README.md").read_text())
self.assertIn("/mcp", (ROOT / "README.md").read_text())
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
self.assertIn("/mcp", skill)
self.assertFalse((ROOT / "skills" / "db-yaml").exists())
self.assertTrue((ROOT / "skills" / "postgres" / "SKILL.md").is_file())
def test_cgo_zig_and_index_profile(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D21", plan)
self.assertIn("zig cc", plan)
dockerfile = (ROOT / "Dockerfile").read_text()
self.assertIn("bin/cgo/zcc", dockerfile)
self.assertIn("FROM debian:bookworm-slim AS api", dockerfile)
self.assertIn("FROM python:3.12-slim AS index", dockerfile)
api = dockerfile[dockerfile.index("FROM debian:bookworm-slim AS api") :]
self.assertNotIn("pip install", api)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn('profiles: ["index"]', compose)
self.assertIn("target: api", compose)
def test_reasoner_docs_name_real_hf_ids_cpu_sidecar(self) -> None:
docs = (ROOT / "docs" / "reasoner.md").read_text()
for hf in (
"Qwen/Qwen3.5-9B",
"Qwen/Qwen3.6-27B",
"prism-ml/Bonsai-27B-gguf",
):
self.assertIn(hf, docs)
self.assertIn("no official qwen3.6-9b", docs.lower())
self.assertIn("OLLAMA_NUM_GPU", docs)
self.assertIn("rss_mb", docs)
self.assertIn("vram_mb", docs)
self.assertIn("3/3", docs)
self.assertIn("Do not claim 9B is better at tools", docs)
self.assertNotIn("Qwen/Qwen3.6-9B", docs)
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D18", plan)
self.assertIn("Qwen/Qwen3.5-9B", plan)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn('"reasoner"', compose)
self.assertIn("OLLAMA_NUM_GPU", compose)
self.assertIn("127.0.0.1:11435", compose)
dockerfile = (ROOT / "Dockerfile").read_text()
self.assertNotIn(".gguf", dockerfile.lower())
self.assertNotIn(".safetensors", dockerfile.lower())
api = dockerfile[dockerfile.index("FROM debian:bookworm-slim AS api") :]
self.assertNotIn("COPY models", api)
self.assertNotIn("qwen", api.lower())
def test_readme_search_escalates_web(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn("--no-web", text)
self.assertIn("D17", (ROOT / "PLAN.md").read_text())
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
self.assertIn("`web` block", skill)
def test_docs_say_hop_walks_from_file(self) -> None:
def test_docs_do_not_claim_hop_walks(self) -> None:
paths = [
ROOT / "README.md",
ROOT / "docs" / "design.md",
ROOT / "skills" / "brain" / "SKILL.md",
ROOT / "docs" / "runbook.md",
ROOT / "docs" / "README.md",
ROOT / "skills" / "diataxis-docs" / "SKILL.md",
]
# Command-style `--hop 1` / `--hop N` plus follow/walk = the old lie.
# Honest "not implemented" notes must not match.
lie = re.compile(r"--hop (?:N|1).*(?:follow|walk)", re.I | re.S)
for path in paths:
text = path.read_text()
self.assertIn("--hop", text, f"{path.relative_to(ROOT)} must document --hop")
self.assertNotIn(
"not implemented",
text.lower(),
f"{path.relative_to(ROOT)} still says hop is not implemented",
self.assertIsNone(
lie.search(text),
f"{path.relative_to(ROOT)} still claims --hop walks the graph",
)
def test_docs_are_portable_diataxis(self) -> None:
index = (ROOT / "docs" / "README.md").read_text()
self.assertIn("type: reference", index)
for d in ("D3", "D6", "D14", "D15", "D17", "D18"):
self.assertIn(d, index)
runbook = (ROOT / "docs" / "runbook.md").read_text()
self.assertIn("type: howto", runbook)
self.assertIn("bin/brain/search.go", runbook)
self.assertIn("bin/brain/index.go", runbook)
self.assertNotIn("search.ops.io", runbook)
self.assertNotIn("/mnt/", runbook)
self.assertNotIn("/home/", runbook)
readme = (ROOT / "README.md").read_text()
self.assertIn("docs/runbook.md", readme)
self.assertNotIn("search.ops.io", readme)
def test_v1_epic_is_named_in_docs(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("Gap to v1", plan)
self.assertIn("eSlider/2dph/issues/16", plan)
self.assertIn("eSlider/2dph/issues/17", plan)
self.assertIn("eSlider/2dph/milestone/12", plan)
road = (ROOT / "docs" / "roadmap.md").read_text()
self.assertIn("type: explanation", road)
self.assertIn("issues/16", road)
self.assertIn("issues/14", road)
index = (ROOT / "docs" / "README.md").read_text()
self.assertIn("roadmap.md", index)
self.assertIn("epic #16", index)
agents = (ROOT / "AGENTS.md").read_text()
self.assertIn("roadmap.md", agents)
-63
View File
@@ -1,63 +0,0 @@
"""Skills must name live commands; every bin/ path in SKILL.md must exist."""
from __future__ import annotations
import re
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
BIN_PATH = re.compile(r"(bin/[A-Za-z0-9_./-]+)")
class SkillsTest(unittest.TestCase):
def test_db_yaml_renamed_to_postgres(self) -> None:
self.assertFalse(
(ROOT / "skills" / "db-yaml").exists(),
"skills/db-yaml must be skills/postgres",
)
self.assertTrue((ROOT / "skills" / "postgres" / "SKILL.md").is_file())
text = (ROOT / "skills" / "postgres" / "SKILL.md").read_text()
self.assertIn("bin/postgres/query.go", text)
self.assertNotIn("search.ops.io", text)
def test_every_bin_path_in_skills_exists(self) -> None:
missing: list[str] = []
for path in (ROOT / "skills").rglob("SKILL.md"):
text = path.read_text()
for m in BIN_PATH.finditer(text):
rel = m.group(1).rstrip(")`.,;")
candidate = ROOT / rel
if not candidate.exists():
missing.append(f"{path.relative_to(ROOT)}: {rel}")
self.assertEqual(missing, [], "skill bin paths must exist")
def test_brain_skill_lists_generated_tools(self) -> None:
tools = (ROOT / "skills" / "brain" / "tools.md").read_text()
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
self.assertIn("tools.md", skill)
for name in ("search", "get", "stats", "audit"):
self.assertIn(f"`{name}`", tools)
def test_picoclaw_lists_tool_order(self) -> None:
skill = (ROOT / "skills" / "picoclaw" / "SKILL.md").read_text()
agents = (ROOT / "AGENTS.md").read_text()
self.assertIn("**`search`**", skill)
self.assertIn("**`get`**", skill)
self.assertIn("**`audit`**", skill)
self.assertIn("throttled", skill.lower())
self.assertIn("not a negative finding", agents)
self.assertIn("Fact-check every", agents)
def test_yq_is_mikefarah_for_structured_data(self) -> None:
skill = (ROOT / "skills" / "yq" / "SKILL.md").read_text()
self.assertIn("https://github.com/mikefarah/yq", skill)
for fmt in ("YAML", "JSON", "XML", "CSV", "TOML", "HCL"):
self.assertIn(fmt, skill)
self.assertIn("not kislyuk", skill.lower())
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("mikefarah/yq", plan)
agents = (ROOT / "AGENTS.md").read_text()
self.assertIn("mikefarah/yq", agents)
web = (ROOT / "skills" / "web-search" / "SKILL.md").read_text()
self.assertIn("| yq ", web)
self.assertNotIn("| jq ", web)
-36
View File
@@ -1,36 +0,0 @@
"""qa/system_perf.py is an offline-gated system test (no live brain in CI)."""
from __future__ import annotations
import ast
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class SystemPerfScriptTest(unittest.TestCase):
def test_script_compiles_and_is_read_only(self) -> None:
path = ROOT / "qa" / "system_perf.py"
src = path.read_text()
compile(src, str(path), "exec")
self.assertIn("--json", src)
self.assertIn("qwen3.5:9b", src)
self.assertIn("--picoclaw", src)
self.assertIn("BRAIN_URL", src)
self.assertIn("tools/list", src)
self.assertIn("tools/call", src)
self.assertIn("GATE_HEALTH_MS", src)
self.assertIn("GATE_GET_P50_MS", src)
self.assertNotIn("kb.lbug", src)
self.assertNotIn("password", src.lower())
self.assertNotIn("token", src.lower())
def test_script_does_not_write_ladybug(self) -> None:
tree = ast.parse((ROOT / "qa" / "system_perf.py").read_text())
writes = [
n.func.attr
for n in ast.walk(tree)
if isinstance(n, ast.Call) and isinstance(n.func, ast.Attribute)
and n.func.attr in {"write_text", "write_bytes", "dump"}
]
self.assertEqual(writes, [], f"system_perf must not write files: {writes}")
+17 -112
View File
@@ -1,96 +1,66 @@
# 2dph — docker composition
#
# docker compose up -d brain # API (Zig CGO serve)
# docker compose --profile index run --rm index # Python rebuild
# docker compose --profile picoclaw up -d # brain-mcp + CPU reasoner + PicoClaw gateway
# docker compose --profile reasoner up -d reasoner # CPU Ollama :11435
# docker compose --profile searxng up -d
# OCR_ENGINE=paddle docker compose --profile ocr-paddle run --rm ocr-paddle
# docker compose run --rm brain index # rebuild graph
# docker compose run --rm brain search "Matrix fed" # one-shot query
# docker compose run --rm brain serve # async Go server
# docker compose up brain-watch # auto re-index
#
# Secrets never baked in: search.env + db-profiles.yml from ~/.config/brain.
# Caching: the 128M model (HF_HOME) and kb.lbug (VAR_DIR) live in named
# volumes, so rebuilds never redownload the model or re-derive the graph.
# Secrets are never baked into the image: search.env + db-profiles.yml mount
# read-only from ~/.config/brain.
name: 2dph
networks:
default:
name: 2dph_sys
driver: bridge
ipam:
config:
- subnet: 10.23.42.0/24
services:
brain:
image: ghcr.io/eslider/2dph:api
image: ghcr.io/eslider/2dph:latest
build:
context: .
dockerfile: Dockerfile
target: api
cache_from:
- ghcr.io/eslider/2dph:cache
command: ["serve"]
command: ["brain", "search", "help"]
environment: &env
HF_HOME: /data/hf
BRAIN_SEARCH_CACHE: /data/cache/web-search.sqlite
BRAIN_DB_PROFILES: /secret/db-profiles.yml
BRAIN_SEARCH_ENV: /secret/search.env
KB_ROOT: /data
KB_SEARCH_CMD: /app/bin/kb/search
KB_WORKERS: "4"
KB_PORT: "8630"
volumes:
- kb-model:/data/hf
- kb-var:/data
# corpus is read-only on the host, never written from the container
- ..:/corpus:ro
- ~/.config/brain:/secret:ro
ports:
- "127.0.0.1:8630:8630"
read_only: true
tmpfs:
- /tmp
healthcheck:
test: ["CMD", "wget", "-qO-", "http://127.0.0.1:8630/health"]
test: ["CMD", "python3", "-c", "import ladybug, model2vec, mistune; print('ok')"]
interval: 30s
timeout: 5s
retries: 3
restart: unless-stopped
stop_grace_period: 20s
# watcher: re-index on corpus file change (watchdog script)
brain-watch:
image: ghcr.io/eslider/2dph:api
image: ghcr.io/eslider/2dph:latest
environment: *env
volumes:
- kb-model:/data/hf
- kb-var:/data
- ..:/corpus:ro
- ~/.config/brain:/secret:ro
command: ["watch", "/corpus"]
command: ["brain", "watch", "/corpus"]
read_only: true
tmpfs:
- /tmp
restart: unless-stopped
stop_grace_period: 20s
# Python write path (Ladybug rebuild). Not in the API image.
# docker compose --profile index run --rm index
index:
profiles: ["index"]
image: ghcr.io/eslider/2dph:index
build:
context: .
dockerfile: Dockerfile
target: index
environment:
HF_HOME: /data/hf
KB_PY: python3
volumes:
- kb-model:/data/hf
- kb-var:/app/var
- ..:/corpus:ro
- ~/.config/brain:/secret:ro
command: ["index"]
read_only: true
tmpfs:
- /tmp
# Optional local SearXNG (D3). Skip if BRAIN_SEARCH_URL already points at a
# live instance — do not run a second copy on that host.
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
@@ -106,71 +76,6 @@ services:
- ./deploy/searxng/limiter.toml:/etc/searxng/limiter.toml:ro
restart: unless-stopped
# MCP endpoint for PicoClaw (and any MCP client).
# docker compose --profile picoclaw up -d
brain-mcp:
profiles: ["picoclaw"]
image: ghcr.io/eslider/2dph:api
environment: *env
volumes:
- kb-model:/data/hf
- kb-var:/data
- ~/.config/brain:/secret:ro
command: ["serve"]
ports:
- "127.0.0.1:8630:8630"
read_only: true
tmpfs:
- /tmp
restart: unless-stopped
# CPU OpenAI-compatible sidecar (D18). Weights are pulled at runtime, not
# baked into the 2dph image. Does not touch host Ollama on :11434.
# docker compose --profile reasoner up -d reasoner
# docker compose --profile reasoner exec reasoner ollama pull qwen3.5:9b
reasoner:
profiles: ["reasoner", "picoclaw"]
image: docker.io/ollama/ollama:latest
environment:
OLLAMA_NUM_GPU: "0"
OLLAMA_HOST: "0.0.0.0:11434"
ports:
- "127.0.0.1:11435:11434"
volumes:
- reasoner-ollama:/root/.ollama
restart: unless-stopped
# Official PicoClaw gateway. Config has no secrets (Ollama + HTTP MCP).
# Host network: brain/reasoner bind 127.0.0.1 only, so host.docker.internal
# (docker0) cannot reach them. Gateway 127.0.0.1:18790 (not the 18800 launcher).
# If :8630/:11435 are already bound, do not start brain-mcp/reasoner:
# docker compose --profile picoclaw up -d --no-deps picoclaw
picoclaw:
profiles: ["picoclaw"]
image: docker.io/sipeed/picoclaw:v0.3.1
network_mode: host
depends_on:
- brain-mcp
- reasoner
environment:
PICOCLAW_GATEWAY_HOST: "127.0.0.1"
entrypoint: ["picoclaw", "gateway"]
volumes:
- picoclaw-home:/root/.picoclaw
- ./deploy/picoclaw/config.json:/root/.picoclaw/config.json:ro
restart: unless-stopped
# Optional PP-OCRv5 (not default). Default OCR is tesseract eng+deu.
# OCR_ENGINE=paddle docker compose --profile ocr-paddle run --rm ocr-paddle
ocr-paddle:
profiles: ["ocr-paddle"]
image: python:3.12-slim
environment:
OCR_ENGINE: paddle
command: ["python", "-c", "print('OCR_ENGINE=paddle; install paddleocr on PATH')"]
volumes:
kb-model:
kb-var:
reasoner-ollama:
picoclaw-home:
-33
View File
@@ -1,33 +0,0 @@
{
"agents": {
"defaults": {
"model_name": "qwen3.5-9b",
"max_tool_iterations": 8,
"max_tokens": 512,
"context_window": 8192
}
},
"model_list": [
{
"model_name": "qwen3.5-9b",
"model": "ollama/qwen3.5:9b",
"api_base": "http://127.0.0.1:11435/v1",
"request_timeout": 600
}
],
"tools": {
"web": {
"enabled": false
},
"mcp": {
"enabled": true,
"servers": {
"2dph": {
"enabled": true,
"type": "http",
"url": "http://127.0.0.1:8630/mcp"
}
}
}
}
}
-8
View File
@@ -1,8 +0,0 @@
{
"mcpServers": {
"2dph": {
"url": "http://127.0.0.1:8630/mcp",
"description": "2dph fact gate. Tool order: search → get → audit. throttled is not absence."
}
}
}
+9 -32
View File
@@ -1,38 +1,15 @@
---
type: reference
status: current
related:
- docs/runbook.md
- docs/design.md
- PLAN.md
- docs/roadmap.md
---
# 2dph (deductionphile)
# 2dph docs (Diataxis)
Evidence-first knowledge graph. Facts need proof or they are
Evidence-first knowledge graph + hybrid RAG over the operational
Brain/ops/eSlider stack. Facts need proof or they are
`(not confirmed)`.
| Type | Doc |
|------|-----|
| tutorial / howto | [runbook](runbook.md) — run anywhere (uv, Go, Docker) |
| explanation | [design](design.md) — two roots, deduction, D17/D20/D18 |
| explanation | [roadmap](roadmap.md) — gap to v1 (epic #16) |
| howto | [picoclaw](picoclaw.md) — MCP agent profile |
| howto | [reasoner](reasoner.md) — CPU bake-off (D18) |
| reference | [PLAN.md](../PLAN.md) — decisions D1D22 |
Decisions the public face must name: **D3** SearXNG compose, **D6** Go service /
Python write sidecar, **D14** `bin/{subject}/{method}.go`, **D15** Gitea origin,
**D17** assertion gate (facts → info → web), **D18** pluggable reasoner.
- [PLAN.md](../PLAN.md) — decisions, execution order, open questions (v2)
- [design](design.md) — schema, deduction model, sources
- [Gitea issues](https://git.produktor.io/eSlider/2dph/issues) — work board (origin)
Search: `bin/brain/search.go "query"` (HTTP: `bin/brain/serve.go`
`/health` `/search` `/get` `/stats` `/audit` `/ingest`). `--hop N` walks
`FROM_FILE` → Commit → Person from each hit (max 3). Rebuild writes
File edges ([#17](https://git.produktor.io/eSlider/2dph/issues/17)).
`/health` `/search` `/get` `/stats` `/audit` `/ingest`). `--hop` is
not a walk; the flag errors until File/FROM_FILE edges exist.
Work board: [Gitea issues](https://git.produktor.io/eSlider/2dph/issues)
([epic #16](https://git.produktor.io/eSlider/2dph/issues/16)).
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
Published docs live here and match live commands.
Published docs live here and mirror the project state.
+1 -2
View File
@@ -21,5 +21,4 @@ OO_CLI (default: $HOME/go/bin/oo)
./bin/chats/apply.go --dry-run
```
JSONL → markdown only. Brain ingest is `bin/brain/index.go --with-chats`
(default `var/chats/md`). WhatsApp sync is out of v1.
JSONL → markdown only. Brain ingest is `bin/brain/index.go` (not a `chats index`).
+33
View File
@@ -0,0 +1,33 @@
# CRM association proof (oo CLI ↔ corpus)
Proven with `oo` (eslider/go-onlyoffice) against the OnlyOffice portal
(`office.produktor.io`). Portal CRM is the SSOT for company ↔ person ↔
project associations; the corpus SoT (`eslider/cv/projects/knowledge-mesh-seed.yaml`)
is the second, independent source. Facts that can be backed by both are
written to the brain under `root=facts` by `bin/facts/crm`.
## What was verified
- Logical counts (portal MySQL): 1300 contacts = 897 persons + 404 companies,
198 projects, 998 deals, 939 project↔contact links.
- Every client company linked to a project has ≥1 person underneath.
- Every person `company_id` resolves to an existing company.
- Corpus org list (9) maps 1:1 onto CRM companies:
ProProdukt SL / produktor.io, Dyvenia, Immowelt AG, WhereGroup,
Keynote SIGOS, D2S/SYSTEMS, GRID, Pack und Cup, Markets Platform.
- 78 person↔company association facts written to the brain
(`how=crm-crosscheck`, `type=association`). Recall@5 in `bin/kb/eval` = 1.0.
## Mistakes found
| # | Mistake | Fix |
|---|---------|-----|
| 1 | Duplicate legal entity `GoldenRatio.Exchange` (contact 759) vs `Golden Ratio Exchange` (763); 3 deals (211, 287, 559) were linked to 759 | `oo contacts merge 759 763` — 763 kept, 759 removed, deal links re-pointed to 763 |
| 2 | `env/`-wide: OnlyOffice creds file used wrong UX (user `eslider`, password with `$2` suffix) making `oo` auth fail | `.env` fixed to `eslider@gmail.com` + clean password; `.env` stays gitignored |
## Gates after fix
- `uv run python -m unittest discover -s bin/tools -t .` → 26 tests OK
- `bin/facts/audit self` + `bin/facts/audit db` → ok
- `bin/kb/eval` → recall@5 = 1.0
- `go test ./...` (bin/server + bin/watch) → ok
+3 -44
View File
@@ -1,12 +1,3 @@
---
type: explanation
status: current
related:
- docs/README.md
- docs/runbook.md
- docs/roadmap.md
---
# Design — facts, info, deduction
## Two roots, one transaction
@@ -30,13 +21,11 @@ bin/brain/search.go "question"
1. facts root — confirmed answers only → return with evidence links
2. info root — supporting narrative → snippets, marked (not confirmed)
3. web-search — second independent source → upgrade hypothesis to confirmed
(`web` block from `bin/web/search.go` when no facts hit; status `throttled`
is not evidence of absence; `--no-web` / `--root` skip it)
(`bin/web/search.go`; status `throttled` is not evidence of absence)
```
`--hop N` walks `Leaf-[:FROM_FILE]->File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`
from each hit (1=File, 2=Commit, 3=Person). Rebuild writes FROM_FILE;
git import writes HAS_VERSION/AUTHORED ([#17](https://git.produktor.io/eSlider/2dph/issues/17)).
`--hop` is not implemented yet (needs File/FROM_FILE edges). The flag is an
error; it is not a graph walk.
## Who / What / How / Where / When + evidence
@@ -72,33 +61,3 @@ corpus HEAD.
Confirmed = A×B or B×C agreement. Single source = hypothesis + `(not confirmed)`.
Conflicting pairings (≥2 yes vs ≥2 no) = hypothesis (OQ1 → v2 resolution).
## Read path
`bin/brain/get.go`, `stats.go`, and `eval.go` call `internal/brain` with cgo
(`system_ladybug`), compiled by **Zig** (`bin/cgo/zcc`, D21), not gcc.
They do not exec Python. Control questions for recall@5 live in
`internal/brain/rank` so CI can test the table without libladybug.
Python `bin/kb/{get,stats,eval}` remain for GitHub Actions until the runner
fetches Zig + libs (`bin/cgo/zig`). Incremental write is `bin/kb/add`
(`bin/brain/add.go`). Bulk index/write is still `bin/kb/index`
(`docker compose --profile index`).
## Agent API (D20)
`bin/brain/serve.go` exposes the same `internal/httpapi.Ops` table as OpenAPI
(`GET /openapi.json`) and MCP (`POST /mcp` JSON-RPC `tools/list` +
`tools/call`). Tool names match paths: `search`, `get`, `stats`, `audit`,
`ingest` (add a leaf; omit body for the CLI hint).
Agents should use these endpoints instead of shebang CLIs.
## Reasoner (D18)
Pluggable OpenAI-compatible URL. RAM: `Qwen/Qwen3.5-9B`. Quality:
`prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B.
CPU sidecar: compose profile `reasoner` (`OLLAMA_NUM_GPU=0`,
`127.0.0.1:11435`). Bake-off: `bin/reasoner/bakeoff.go`. Weights stay out
of the 2dph image. See [docs/reasoner.md](reasoner.md).
Gap to v1 (hops, corpus, CI eval): [roadmap](roadmap.md),
[epic #16](https://git.produktor.io/eSlider/2dph/issues/16).
-34
View File
@@ -1,34 +0,0 @@
# PicoClaw profile (reference agent)
2dph is the memory/fact gate. Compose profile `picoclaw` runs the official
PicoClaw gateway (`docker.io/sipeed/picoclaw:v0.3.1`) plus `brain-mcp` and the
CPU reasoner. Default agent model is `qwen3.5:9b` (RAM path, D18). Weights stay
in the reasoner volume, not in the 2dph image.
No secrets in git: Ollama needs no key; MCP is local HTTP.
```bash
docker compose --profile picoclaw up -d
# already serving :8630 / :11435:
docker compose --profile picoclaw up -d --no-deps picoclaw
```
Gateway: `127.0.0.1:18790`. Brain MCP: `http://127.0.0.1:8630/mcp`.
Cursor-style clients can use [deploy/picoclaw/mcp.json.example](../deploy/picoclaw/mcp.json.example).
PicoClaw itself uses [deploy/picoclaw/config.json](../deploy/picoclaw/config.json)
(`127.0.0.1` + host network — loopback publishes are not reachable via docker0).
OpenAPI: `GET http://127.0.0.1:8630/openapi.json`.
Before a factual reply: `search``get``audit`. `throttled` is not a
negative finding. See `skills/picoclaw/SKILL.md`.
System performance (MCP gates + qwen3.5:9b tool_call + PicoClaw gateway):
```bash
./qa/system_perf.py --json | yq '.gates'
REASONER_MODEL=qwen3.5:9b ./qa/system_perf.py --reasoner --picoclaw --json | yq '.reasoner'
```
The default agent model is `qwen3.5:9b`. PicoClaw `context_window` is 8192
(heuristic `max_tokens*4` at 512 is 2048, too small for MCP tool schemas).
`request_timeout` is 600s for a CPU turn (tool_call + MCP search + answer).
-71
View File
@@ -1,71 +0,0 @@
# Reasoner bake-off (D18)
Pluggable OpenAI-compatible URL. 2dph does not ship weights. PicoClaw is
compose profile `picoclaw` (`sipeed/picoclaw`); the bake-off hits the same
tool names (`search``get``audit` from `internal/httpapi.Ops`).
```bash
docker compose --profile reasoner up -d reasoner
docker compose --profile reasoner exec reasoner ollama pull qwen3.5:9b
REASONER_BASE_URL=http://127.0.0.1:11435/v1 REASONER_MODEL=qwen3.5:9b \
./bin/reasoner/bakeoff.go --json
```
JSON includes `latency_p50_ms` / `latency_p95_ms` from DuckDB (`internal/duckstats`, D22).
Host Ollama on `:11434` is left alone. This sidecar binds `127.0.0.1:11435`
with `OLLAMA_NUM_GPU=0` (CPU). Measure RSS (`/api/ps` `size`), not VRAM.
If Compose cannot allocate a project network (Docker IPAM pool exhausted),
the same sidecar is:
```bash
docker run -d --name 2dph-reasoner \
-e OLLAMA_NUM_GPU=0 \
-p 127.0.0.1:11435:11434 \
-v 2dph-reasoner-ollama:/root/.ollama \
ollama/ollama:latest
```
## Real Hugging Face ids
| Role | HF id | Ollama tag (this bake-off) |
|------|-------|----------------------------|
| RAM / 9B | `Qwen/Qwen3.5-9B` | `qwen3.5:9b` |
| Quality 27B (CPU) | `prism-ml/Bonsai-27B-gguf` (derived from Qwen3.6-27B) | `MichelRosselli/bonsai-27b:Q1_0` |
| Quality 27B (full) | `Qwen/Qwen3.6-27B` | not pulled on this CPU box |
There is **no official Qwen3.6-9B**. Do not invent that id.
Qwen3.5-9B has documented upstream tool-call XML bugs. A 9B win on tools
is only claimed if this bake-off records OpenAI `tool_calls` (not
`<tool_call>` XML in `content`).
The 2dph API image does not `COPY` GGUF/safetensors. Pull at runtime into
the `reasoner-ollama` volume.
## Live CPU run
Host sidecar: Ollama **0.32.9**, `OLLAMA_NUM_GPU=0`, `127.0.0.1:11435`,
`device: cpu`, `vram_mb: 0`. Date: 2026-08-13. Same three prompts
(`search` / `get` / `audit`). PicoClaw binary was not used; the OpenAI
tools payload is the surface it would send.
| Model | HF id | tool_call | xml_leak | rss_mb | latency_ms (search/get/audit) |
|-------|-------|-----------|----------|--------|-------------------------------|
| `qwen3.5:9b` | `Qwen/Qwen3.5-9B` | 3/3 | 0 | 5790 | 50059 / 51678 / 31980 |
| `MichelRosselli/bonsai-27b:Q1_0` | `prism-ml/Bonsai-27B-gguf` | 3/3 | 0 | 21951 | 327702 / 166830 / 119808 |
| `Qwen/Qwen3.6-27B` | `Qwen/Qwen3.6-27B` | not loaded | — | — | too heavy for this CPU box |
Both loaded models emitted OpenAI `tool_calls` (not `<tool_call>` XML) on
this runtime. **Do not claim 9B is better at tools** — the score is tied
at 3/3. 9B is smaller and faster. Bonsai RSS includes weights + KV
(`size` from `/api/ps`); first Bonsai prompt includes cold load.
Re-run:
```bash
REASONER_BASE_URL=http://127.0.0.1:11435/v1 REASONER_MODEL=qwen3.5:9b \
./bin/reasoner/bakeoff.go --json
REASONER_MODEL=MichelRosselli/bonsai-27b:Q1_0 ./bin/reasoner/bakeoff.go --json
```
-66
View File
@@ -1,66 +0,0 @@
---
type: explanation
status: current
related:
- PLAN.md
- docs/design.md
- docs/runbook.md
---
# Gap to v1 — detective brain
Goal: a brain that does not assert without proof. Search is deduction
(`facts` ≥2 sources → `info``web`). `confirmed` only from the facts root.
**v1 is a living graph the agent can write and walk**, not “more RAG”.
Epic: [Gitea #16](https://git.produktor.io/eSlider/2dph/issues/16).
Milestone: [v1 detective brain](https://git.produktor.io/eSlider/2dph/milestone/12).
Decisions: [PLAN.md](../PLAN.md).
## In (do not reopen)
Read path Go + Zig CGO (D21). HTTP + OpenAPI + MCP (D20). PicoClaw compose
profile + CPU reasoner (D18). Mail sync → import → rebuild. D14 shebangs.
Compose `api` (no CPython) / `index` (Python write). Issues #1#5, #7#13.
[#15](https://git.produktor.io/eSlider/2dph/issues/15) lever/loop.
[#14](https://git.produktor.io/eSlider/2dph/issues/14) `bin/brain/add.go` /
`POST /ingest` (Python `kblib.add_leafs`; no Go upsert port).
[#17](https://git.produktor.io/eSlider/2dph/issues/17) `--hop N` walks
FROM_FILE / HAS_VERSION / AUTHORED.
[#18](https://git.produktor.io/eSlider/2dph/issues/18) `--with-facts` /
`--with-chats` on rebuild (WhatsApp out of v1).
[#19](https://git.produktor.io/eSlider/2dph/issues/19) CI recall SoT =
`bin/brain/eval.go` via Zig.
Epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed.
## v2
[#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR — **in**.
[#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3 duckdb-go — **in**.
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1 contradiction
resolution.
## Blockers
None for epic #16 (closed). Remaining v2: OQ1, OQ4.
```
question
├─ FTS + HNSW ← in
├─ facts / info roots ← in
├─ web (D17) ← in
├─ brain/add ACID ← in
├─ Cypher hop ← in
└─ facts+chats corpus ← in
```
## Not v1
OQ1 contradiction resolution, OQ4 YAML-first leafs.
OCR (OQ2) and duckdb-go (OQ3/D22) are in.
## Close epic #16 when
Children #14, #15, #17, #18, #19 are closed. MCP tool order stays gated by tests.
-85
View File
@@ -1,85 +0,0 @@
---
type: howto
status: current
related:
- docs/README.md
- PLAN.md
---
# Run 2dph (portable)
No laptop-absolute paths. Config lives in env files under `$HOME/.config/brain/`
(mode 0600), not in git.
## Toolchain
- Go (see `go.mod`)
- Python 3.12 + [uv](https://docs.astral.sh/uv)
- Optional: Docker, Zig CGO via `bin/cgo/zig` (not gcc)
- Optional: poppler (`pdftotext`/`pdftoppm`) + tesseract `eng+deu` for mail OCR
```bash
uv venv .venv
uv pip install -r requirements.lock.txt
eval "$(bin/cgo/zig env)" # when compiling Ladybug read tools
go test ./...
uv run python -m unittest discover -s bin/tools -t .
```
## Config
| File / env | Purpose |
|------------|---------|
| `$BRAIN_SEARCH_ENV` (default `$HOME/.config/brain/search.env`) | `BRAIN_SEARCH_URL` (SearXNG). Optional Basic Auth. |
| `$HOME/.config/brain/db-profiles.yml` | read-only Postgres profiles (OnlyOffice via tunnel) |
If the host already runs SearXNG, point `BRAIN_SEARCH_URL` at it. Do not start
a second copy (D3). Optional Compose instance:
```bash
SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
```
That binds `127.0.0.1:8888`. JSON format must stay enabled.
## Index then search
Write path is `bin/brain/add.go` for a leaf (or `POST /ingest`). Bulk
corpus rebuild remains `bin/brain/index.go --rebuild` (Compose profile
`index`). Do not DROP INDEX on Ladybug 0.19.
```bash
bin/brain/add.go --text "arc-1 runs Matrix" --root facts --source "compose.yml x docker ps"
bin/brain/index.go --rebuild --with-facts --with-chats
bin/brain/search.go "LadybugDB vector index" # facts → info → web (D17)
bin/brain/search.go "upstream flag" --no-web
bin/brain/get.go <id> --body
bin/brain/stats.go
```
`--hop N` walks File → Commit → Person from each hit. Empty web results are `throttled`, not absence.
Gap to v1: [roadmap](roadmap.md) / [epic #16](https://git.produktor.io/eSlider/2dph/issues/16).
Ladybug 0.19: never `DROP INDEX` FTS/VECTOR (ghost catalog). Fresh indexes =
delete `var/kb.lbug` then `--rebuild`.
## HTTP / MCP
```bash
docker compose up -d brain # :8630 Zig CGO serve
docker compose --profile index run --rm index # rebuild
docker compose --profile picoclaw up brain-mcp # MCP 127.0.0.1:8630
```
`GET /openapi.json`, `POST /mcp`. Agent tool order: `search``get``audit`.
## Reasoner (optional, D18)
CPU sidecar on `127.0.0.1:11435`. Weights are not in the 2dph image.
```bash
docker compose --profile reasoner up -d reasoner
REASONER_BASE_URL=http://127.0.0.1:11435/v1 ./bin/reasoner/bakeoff.go --json
```
See [reasoner.md](reasoner.md).
-8
View File
@@ -7,7 +7,6 @@ require (
github.com/arran4/golang-ical v0.3.5
github.com/chewxy/math32 v1.11.2
github.com/daulet/tokenizers v1.27.0
github.com/duckdb/duckdb-go/v2 v2.10505.0
github.com/go-git/go-git/v5 v5.19.2
golang.org/x/sys v0.47.0
golang.org/x/text v0.40.0
@@ -21,17 +20,10 @@ require (
github.com/apache/arrow-go/v18 v18.6.0 // indirect
github.com/cloudflare/circl v1.6.3 // indirect
github.com/cyphar/filepath-securejoin v0.6.1 // indirect
github.com/duckdb/duckdb-go-bindings v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/darwin-amd64 v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/darwin-arm64 v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/linux-amd64 v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/linux-arm64 v0.10505.0 // indirect
github.com/duckdb/duckdb-go-bindings/lib/windows-amd64 v0.10505.0 // indirect
github.com/dustin/go-humanize v1.0.1 // indirect
github.com/emirpasic/gods v1.18.1 // indirect
github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376 // indirect
github.com/go-git/go-billy/v5 v5.9.0 // indirect
github.com/go-viper/mapstructure/v2 v2.5.0 // indirect
github.com/goccy/go-json v0.10.6 // indirect
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 // indirect
github.com/google/flatbuffers v25.12.19+incompatible // indirect
-16
View File
@@ -31,20 +31,6 @@ github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSs
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/duckdb/duckdb-go-bindings v0.10505.0 h1:/0pPsTLrcCsTGxT0VrHgJWnOcPe1tQL1vrki1v3jbAI=
github.com/duckdb/duckdb-go-bindings v0.10505.0/go.mod h1:HoD5xePkDj3VZbBnVVfxVVYIljZ9khCprWA7FgwIiC4=
github.com/duckdb/duckdb-go-bindings/lib/darwin-amd64 v0.10505.0 h1:FrMqquFBQlMsi34h2KZgCku54rqA8xEbXZ0NLVDKwYs=
github.com/duckdb/duckdb-go-bindings/lib/darwin-amd64 v0.10505.0/go.mod h1:EnAvZh1kNJHp5yF+M1ZHNEvapnmt6anq1xXHVrAGqMo=
github.com/duckdb/duckdb-go-bindings/lib/darwin-arm64 v0.10505.0 h1:lbRbpQwT1MmUhh/VTwukV9K8bxKByV3UghAP3MvsbBo=
github.com/duckdb/duckdb-go-bindings/lib/darwin-arm64 v0.10505.0/go.mod h1:IGLSeEcFhNeZF16aVjQCULD7TsFZKG5G7SyKJAXKp5c=
github.com/duckdb/duckdb-go-bindings/lib/linux-amd64 v0.10505.0 h1:nrsaVYj3XYCRbS2FpdOMD/KHE7egRMr+/NR1IHmjT84=
github.com/duckdb/duckdb-go-bindings/lib/linux-amd64 v0.10505.0/go.mod h1:KAIynZ0GHCS7X5fRyuFnQMg/SZBPK/bS9OCOVojClxw=
github.com/duckdb/duckdb-go-bindings/lib/linux-arm64 v0.10505.0 h1:qM6oGDgwXBILJGbTY4fCy6QOczLpucUA6yn6g3ORjh4=
github.com/duckdb/duckdb-go-bindings/lib/linux-arm64 v0.10505.0/go.mod h1:81SGOYoEUs8qaAfSk1wRfM5oobrIJ5KI7AzYhK6/bvQ=
github.com/duckdb/duckdb-go-bindings/lib/windows-amd64 v0.10505.0 h1:DjqZl9rYreHkSOqnqLmkrqH5T8UdQNcxZLJVZzGmXXA=
github.com/duckdb/duckdb-go-bindings/lib/windows-amd64 v0.10505.0/go.mod h1:K25pJL26ARblGDeuAkrdblFvUen92+CwksLtPEHRqqQ=
github.com/duckdb/duckdb-go/v2 v2.10505.0 h1:SWwvLn2Qx/RQSnQNupwgIF8VbnJ5A6OQU9lYb/mDETI=
github.com/duckdb/duckdb-go/v2 v2.10505.0/go.mod h1:m0PW4J4FG9hlFlVdXi6Ds9owpyIDaBdE2jyce00fGcE=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/elazarl/goproxy v1.7.2 h1:Y2o6urb7Eule09PjlhQRGNsqRfPmYI3KKQLFpCAV3+o=
@@ -61,8 +47,6 @@ github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399 h1:eMj
github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399/go.mod h1:1OCfN199q1Jm3HZlxleg+Dw/mwps2Wbk9frAWm+4FII=
github.com/go-git/go-git/v5 v5.19.2 h1:wkfn7vOlUBu8ivAWKBWisTiwJK4jYHzTF8Ndv1LyGqY=
github.com/go-git/go-git/v5 v5.19.2/go.mod h1:QqCBE1EFN5ddFmrliLQ3/ntRCUjZU3EJuwuB/jWEHjk=
github.com/go-viper/mapstructure/v2 v2.5.0 h1:vM5IJoUAy3d7zRSVtIwQgBj7BiWtMPfmPEgAXnvj1Ro=
github.com/go-viper/mapstructure/v2 v2.5.0/go.mod h1:oJDH3BJKyqBA2TXFhDsKDGDTlndYOZ6rGS0BRZIxGhM=
github.com/goccy/go-json v0.10.6 h1:p8HrPJzOakx/mn/bQtjgNjdTcN+/S6FcG2CTtQOrHVU=
github.com/goccy/go-json v0.10.6/go.mod h1:oq7eo15ShAhp70Anwd5lgX2pLfOS3QCiwU/PULtXL6M=
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 h1:f+oWsMOmNPc8JmEHVZIycC7hBoQxHH9pNKQORJNozsQ=
+6 -23
View File
@@ -7,10 +7,6 @@ import (
"context"
"encoding/json"
"fmt"
"os/exec"
"path/filepath"
"github.com/eSlider/2dph/internal/brain/rank"
)
// Ready opens the Ladybug file for the life of the serve process.
@@ -21,7 +17,7 @@ func Ready() error {
// HTTP is the in-process API used by bin/brain/serve.go.
type HTTP struct{}
func (HTTP) Search(ctx context.Context, query string, limit int) ([]byte, error) {
func (HTTP) Search(_ context.Context, query string, limit int) ([]byte, error) {
hits, err := searchHits(query, "", "", limit)
if err != nil {
return nil, err
@@ -35,13 +31,10 @@ func (HTTP) Search(ctx context.Context, query string, limit int) ([]byte, error)
hits[i].Snippet = string(runes)
}
}
webOut := rank.Deduce(hits, query, "", false, func(q string) rank.SecondSource {
return lookupWeb(ctx, q)
})
var buf bytes.Buffer
enc := json.NewEncoder(&buf)
enc.SetEscapeHTML(false)
if err := enc.Encode(toJSONOut(hits, query, "", webOut)); err != nil {
if err := enc.Encode(toJSONOut(hits, query, "")); err != nil {
return nil, err
}
return buf.Bytes(), nil
@@ -139,23 +132,13 @@ func (HTTP) Audit(context.Context) ([]byte, error) {
return json.Marshal(map[string]any{"status": "ok", "by_confidence": rows})
}
func (HTTP) Ingest(ctx context.Context, body []byte) ([]byte, error) {
if len(bytes.TrimSpace(body)) == 0 {
func (HTTP) Ingest(context.Context) ([]byte, error) {
return json.Marshal(map[string]any{
"mode": "add",
"command": "bin/brain/add.go",
"rebuild": "bin/brain/index.go --rebuild",
"mode": "rebuild",
"command": "bin/brain/index.go --rebuild",
"add": "v2",
})
}
cmd := exec.CommandContext(ctx, filepath.Join(repoRoot(), "bin", "kb", "add"), "--json")
cmd.Stdin = bytes.NewReader(body)
cmd.Dir = repoRoot()
out, err := cmd.Output()
if err != nil {
return nil, fmt.Errorf("add: %w", err)
}
return out, nil
}
func asInt(v any) int64 {
switch n := v.(type) {
-3
View File
@@ -1,3 +0,0 @@
package brain
const ModelID = "minishlab/potion-multilingual-128M"
+4 -14
View File
@@ -6,7 +6,7 @@ import (
"strings"
)
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--hop N] [--json] [--no-web]
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--json]
bin/brain/search.go serve [port]
bin/brain/search.go --list-model`
@@ -15,14 +15,14 @@ type Options struct {
Root string
Repo string
Limit int
Hop int
JSONOut bool
ListModel bool
NoWeb bool
}
// ParseArgs reads flags. Unknown flags are an error: silently dropping them
// meant `--hop 1` vanished and its argument `1` was appended to the query.
// --hop is recognised so it cannot be swallowed; it is not implemented until
// File/FROM_FILE edges exist.
func ParseArgs(args []string) (Options, error) {
opt := Options{Limit: 20}
var queryArgs []string
@@ -51,19 +51,9 @@ func ParseArgs(args []string) (Options, error) {
}
opt.Limit = n
case "--hop":
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 1 {
return opt, fmt.Errorf("--hop must be a positive integer, got %q", args[i])
}
if n > 3 {
return opt, fmt.Errorf("--hop max is 3 (File → Commit → Person)")
}
opt.Hop = n
return opt, fmt.Errorf("--hop is not implemented yet (needs File/FROM_FILE edges)")
case "--json":
opt.JSONOut = true
case "--no-web":
opt.NoWeb = true
case "--list-model":
opt.ListModel = true
default:
-43
View File
@@ -1,43 +0,0 @@
package rank
// SecondSource is the web-search block on a deduction answer.
// Kept apart from graph hits so "ours" and "not ours" stay visible.
type SecondSource struct {
Status string `json:"status"`
Note string `json:"note,omitempty"`
Cached bool `json:"cached,omitempty"`
Results []SecondSourceHit `json:"results,omitempty"`
}
type SecondSourceHit struct {
Rank int `json:"rank"`
Title string `json:"title"`
URL string `json:"url"`
Snippet string `json:"snippet"`
Engine string `json:"engine"`
}
type WebFn func(query string) SecondSource
// ShouldEscalate is true when the default deduction path has no facts hit.
// `--root facts|info` is a single-root ask: do not mix in the web.
func ShouldEscalate(hits []Hit, rootFilter string) bool {
if rootFilter != "" {
return false
}
for _, h := range hits {
if h.Root == "facts" {
return false
}
}
return true
}
// Deduce returns the second-source block, or nil when web must not run.
func Deduce(hits []Hit, query, rootFilter string, noWeb bool, web WebFn) *SecondSource {
if noWeb || web == nil || !ShouldEscalate(hits, rootFilter) {
return nil
}
out := web(query)
return &out
}
-78
View File
@@ -1,78 +0,0 @@
package rank
import (
"strings"
"testing"
)
func TestShouldEscalateWhenNoFacts(t *testing.T) {
if !ShouldEscalate(nil, "") {
t.Fatal("empty local graph must escalate")
}
if !ShouldEscalate([]Hit{h("i", "info", "docs/a.md")}, "") {
t.Fatal("info-only must escalate (not confirmed)")
}
}
func TestShouldNotEscalateWhenFactsConfirm(t *testing.T) {
hits := []Hit{h("f", "facts", "docker ps x compose"), h("i", "info", "docs/a.md")}
if ShouldEscalate(hits, "") {
t.Fatal("facts hit is already confirmed; do not mix web")
}
}
func TestShouldNotEscalateWhenRootFilterSet(t *testing.T) {
if ShouldEscalate(nil, "facts") {
t.Fatal("--root facts must stay local")
}
if ShouldEscalate([]Hit{h("i", "info", "x")}, "info") {
t.Fatal("--root info must stay local")
}
}
func TestDeduceCallsWebOnlyWhenEscalating(t *testing.T) {
called := 0
web := func(q string) SecondSource {
called++
if q != "LadybugDB" {
t.Fatalf("query = %q", q)
}
return SecondSource{Status: "ok", Results: []SecondSourceHit{{Title: "t", URL: "http://example.com"}}}
}
got := Deduce([]Hit{h("i", "info", "x")}, "LadybugDB", "", false, web)
if called != 1 || got == nil || got.Status != "ok" {
t.Fatalf("got %+v called=%d", got, called)
}
}
func TestDeduceNilWhenFactsOrNoWeb(t *testing.T) {
web := func(string) SecondSource {
t.Fatal("web must not run")
return SecondSource{}
}
if Deduce([]Hit{h("f", "facts", "x")}, "q", "", false, web) != nil {
t.Fatal("facts")
}
if Deduce([]Hit{h("i", "info", "x")}, "q", "", true, web) != nil {
t.Fatal("--no-web")
}
if Deduce(nil, "q", "facts", false, web) != nil {
t.Fatal("--root facts")
}
if Deduce(nil, "q", "", false, nil) != nil {
t.Fatal("nil web fn")
}
}
func TestParseNoWeb(t *testing.T) {
opt, err := ParseArgs([]string{"query", "--no-web", "--json"})
if err != nil || !opt.NoWeb || !opt.JSONOut || opt.Query != "query" {
t.Fatalf("got %+v err=%v", opt, err)
}
}
func TestUsageNamesNoWeb(t *testing.T) {
if !strings.Contains(Usage, "--no-web") {
t.Fatalf("usage must name --no-web, got:\n%s", Usage)
}
}
-16
View File
@@ -1,16 +0,0 @@
package rank
// Eval control questions (recall@5). Kept here so CI can test the gate
// table without ladybug cgo. The runner lives in internal/brain (cgo).
const EvalRecallThreshold = 0.95
type EvalQuestion struct {
Query string
Fragment string
}
var EvalQuestions = []EvalQuestion{
{"hybrid search fts and vector", "BM25"},
{"eslider devops engineer", "DevOps"},
{"ladybugdb graph engine storage", "LadybugDB"},
}
-17
View File
@@ -1,17 +0,0 @@
package rank
import "testing"
func TestEvalQuestionsAreThreeAndThreshold(t *testing.T) {
if EvalRecallThreshold != 0.95 {
t.Fatalf("threshold = %v", EvalRecallThreshold)
}
if len(EvalQuestions) != 3 {
t.Fatalf("questions = %d, want 3", len(EvalQuestions))
}
for _, q := range EvalQuestions {
if q.Query == "" || q.Fragment == "" {
t.Fatalf("empty control: %+v", q)
}
}
}
-27
View File
@@ -7,30 +7,3 @@ const FTSStmt = "CALL QUERY_FTS_INDEX('Leaf', 'id', $q) " +
const VecStmt = "CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) " +
"RETURN node.id, node.text, node.root, node.source, distance ORDER BY distance LIMIT $n"
// HopStmt is the Cypher walk from a search hit. Depth 1 = File, 2 = Commit, 3 = Person.
func HopStmt(depth int) string {
switch depth {
case 1:
return "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File) RETURN f.id, f.path, 1"
case 2:
return "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit) RETURN c.id, c.subject, 2"
case 3:
return "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit)-[:AUTHORED]->(p:Person) RETURN p.id, p.name, 3"
default:
return ""
}
}
func HopLabel(depth int) string {
switch depth {
case 1:
return "File"
case 2:
return "Commit"
case 3:
return "Person"
default:
return ""
}
}
-8
View File
@@ -7,13 +7,6 @@ import (
"strings"
)
type HopNode struct {
ID string `json:"id"`
Label string `json:"label"`
Name string `json:"name"`
Depth int `json:"depth"`
}
// Hit is one search result, mirroring the python script's dict shape.
type Hit struct {
ID string `json:"id"`
@@ -22,7 +15,6 @@ type Hit struct {
Source string `json:"-"`
Score float64 `json:"score"`
Snippet string `json:"snippet,omitempty"`
Hops []HopNode `json:"hops,omitempty"`
}
// rrfK dampens the contribution of low ranks; same constant as kblib.py.
+7 -33
View File
@@ -91,41 +91,15 @@ func TestHybridKeepsVectorScoreForSharedHit(t *testing.T) {
}
// The old parser dropped unknown flags and appended their arguments to the
// query, so `search "q" --hop 1` searched for "q 1". --hop must stay a flag.
// query, so `search "q" --hop 1` searched for "q 1". --hop is not implemented
// here (needs File edges); it must still fail closed instead of changing q.
func TestParseHopIsNotSwallowedIntoTheQuery(t *testing.T) {
opt, err := ParseArgs([]string{"what runs on arc-2", "--hop", "1"})
if err != nil {
t.Fatalf("unexpected error: %v", err)
_, err := ParseArgs([]string{"what runs on arc-2", "--hop", "1"})
if err == nil {
t.Fatal("expected --hop to error (not implemented), not be swallowed")
}
if opt.Query != "what runs on arc-2" {
t.Fatalf("query swallowed hop arg: %q", opt.Query)
}
if opt.Hop != 1 {
t.Fatalf("hop = %d, want 1", opt.Hop)
}
}
func TestParseHopMaxIsThree(t *testing.T) {
if _, err := ParseArgs([]string{"q", "--hop", "4"}); err == nil {
t.Fatal("expected --hop 4 to error")
}
opt, err := ParseArgs([]string{"q", "--hop", "3"})
if err != nil || opt.Hop != 3 {
t.Fatalf("hop 3: %+v err=%v", opt, err)
}
}
func TestHopStmtWalksFromFile(t *testing.T) {
s := HopStmt(1)
if !strings.Contains(s, "FROM_FILE") || !strings.Contains(s, "File") {
t.Fatalf("hop 1 must walk FROM_FILE, got %q", s)
}
s3 := HopStmt(3)
if !strings.Contains(s3, "HAS_VERSION") || !strings.Contains(s3, "AUTHORED") || !strings.Contains(s3, "Person") {
t.Fatalf("hop 3 must reach Person, got %q", s3)
}
if HopLabel(1) != "File" || HopLabel(3) != "Person" {
t.Fatal("hop labels")
if !strings.Contains(err.Error(), "--hop") {
t.Fatalf("error should name --hop, got %v", err)
}
}
-281
View File
@@ -1,281 +0,0 @@
//go:build cgo && system_ladybug
package brain
import (
"encoding/json"
"fmt"
"os"
"sort"
"strings"
"unicode/utf8"
"github.com/eSlider/2dph/internal/brain/rank"
)
func MainGet(args []string) int {
id, body, jsonOut := "", false, false
for _, a := range args {
switch {
case a == "--body":
body = true
case a == "--json":
jsonOut = true
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/get.go <id> [--body] [--json]`)
return 0
case strings.HasPrefix(a, "-"):
fmt.Fprintf(os.Stderr, "brain/get: unknown flag %s\n", a)
return 2
default:
id = a
}
}
if id == "" {
fmt.Fprintln(os.Stderr, "brain/get: id required")
return 2
}
if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1
}
defer closeBrain()
meta, text, err := lookupLeaf(id)
if err != nil {
fmt.Fprintf(os.Stderr, "brain/get: %v\n", err)
return 1
}
out := Dict{
{"id", meta["id"]},
{"root", meta["root"]},
{"confidence", meta["confidence"]},
{"source", meta["source"]},
{"type", meta["type"]},
}
if body {
out = append(out, KV{"text", text})
} else {
out = append(out, KV{"snippet", clip(text, 280)})
}
if jsonOut {
m := map[string]any{}
for _, kv := range out {
m[kv.K] = kv.V
}
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
return b2i(enc.Encode(m))
}
fmt.Print(toYAML(out, 0))
return 0
}
func MainStats(args []string) int {
jsonOut := false
for _, a := range args {
switch a {
case "--json":
jsonOut = true
case "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/stats.go [--json]`)
return 0
default:
if strings.HasPrefix(a, "-") {
fmt.Fprintf(os.Stderr, "brain/stats: unknown flag %s\n", a)
return 2
}
}
}
if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1
}
defer closeBrain()
s, err := leafStats()
if err != nil {
fmt.Fprintf(os.Stderr, "brain/stats: %v\n", err)
return 1
}
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
return b2i(enc.Encode(s))
}
by := s["by_root"].(map[string]int)
keys := make([]string, 0, len(by))
for k := range by {
keys = append(keys, k)
}
sort.Strings(keys)
byRoot := make(Dict, 0, len(keys))
for _, k := range keys {
byRoot = append(byRoot, KV{k, by[k]})
}
out := Dict{
{"total", s["total"]},
{"by_root", byRoot},
{"db", s["db"]},
{"model", s["model"]},
}
fmt.Print(toYAML(out, 0))
return 0
}
func MainEval(args []string) int {
jsonOut := false
for _, a := range args {
switch a {
case "--json":
jsonOut = true
case "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/eval.go [--json]`)
return 0
default:
if strings.HasPrefix(a, "-") {
fmt.Fprintf(os.Stderr, "brain/eval: unknown flag %s\n", a)
return 2
}
}
}
if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1
}
defer closeBrain()
recalled := 0
details := make([]any, 0, len(rank.EvalQuestions))
jsDetails := make([]map[string]any, 0, len(rank.EvalQuestions))
for _, q := range rank.EvalQuestions {
hits, err := queryFTS(q.Query, 5)
ok := false
if err == nil {
frag := strings.ToLower(q.Fragment)
for _, h := range hits {
if strings.Contains(strings.ToLower(h.Text), frag) {
ok = true
break
}
}
}
if ok {
recalled++
}
details = append(details, Dict{
{"q", q.Query},
{"fragment", q.Fragment},
{"in_top5", ok},
})
jsDetails = append(jsDetails, map[string]any{
"q": q.Query, "fragment": q.Fragment, "in_top5": ok,
})
}
n := len(rank.EvalQuestions)
recall := 0.0
if n > 0 {
recall = float64(recalled) / float64(n)
}
passed := recall >= rank.EvalRecallThreshold
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
_ = enc.Encode(map[string]any{
"recall@5": round3(recall),
"passed": passed,
"gate": n,
"details": jsDetails,
})
} else {
out := Dict{
{"recall@5", round3(recall)},
{"passed", passed},
{"gate", n},
{"details", details},
}
fmt.Print(toYAML(out, 0))
}
if !passed {
return 2
}
return 0
}
func lookupLeaf(id string) (map[string]string, string, error) {
if conn == nil {
return nil, "", fmt.Errorf("brain not open")
}
stmt, err := conn.Prepare(
"MATCH (l:Leaf {id:$id}) RETURN l.id, l.text, l.root, l.confidence, l.source, l.type",
)
if err != nil {
return nil, "", err
}
defer stmt.Close()
res, err := conn.Execute(stmt, map[string]any{"id": id})
if err != nil {
return nil, "", err
}
if !res.HasNext() {
return nil, "", fmt.Errorf("no leaf %s", id)
}
row, err := res.Next()
if err != nil {
return nil, "", err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 6 {
return nil, "", fmt.Errorf("leaf row")
}
meta := map[string]string{
"id": fmt.Sprint(vals[0]),
"root": fmt.Sprint(vals[2]),
"confidence": fmt.Sprint(vals[3]),
"source": fmt.Sprint(vals[4]),
"type": fmt.Sprint(vals[5]),
}
return meta, fmt.Sprint(vals[1]), nil
}
func leafStats() (map[string]any, error) {
if conn == nil {
return nil, fmt.Errorf("brain not open")
}
res, err := conn.Query("MATCH (l:Leaf) RETURN l.root, count(*)")
if err != nil {
return nil, err
}
byRoot := map[string]int{}
total := 0
for res.HasNext() {
row, err := res.Next()
if err != nil {
return nil, err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 2 {
continue
}
n := int(asInt(vals[1]))
byRoot[fmt.Sprint(vals[0])] = n
total += n
}
return map[string]any{
"total": total,
"by_root": byRoot,
"db": dbPath(),
"model": ModelID,
}, nil
}
func clip(s string, n int) string {
if utf8.RuneCountInString(s) <= n {
return s
}
return string([]rune(s)[:n])
}
func round3(f float64) float64 {
return float64(int(f*1000+0.5)) / 1000
}
+2 -69
View File
@@ -56,12 +56,6 @@ func runSearch(args []string) int {
fmt.Fprintf(os.Stderr, "search: %v\n", err)
return 1
}
if opt.Hop > 0 {
if err := attachHops(hits, opt.Hop); err != nil {
fmt.Fprintf(os.Stderr, "hop: %v\n", err)
return 1
}
}
results := hits
for i := range results {
@@ -74,25 +68,18 @@ func runSearch(args []string) int {
}
}
webOut := rank.Deduce(results, query, root, opt.NoWeb, func(q string) rank.SecondSource {
return lookupWeb(context.Background(), q)
})
out := Dict{
{"query", query},
{"root_filter", root},
{"count", len(results)},
{"results", resultsToDicts(results)},
}
if webOut != nil {
out = append(out, KV{"web", secondToDict(*webOut)})
}
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
return b2i(enc.Encode(toJSONOut(results, query, root, webOut)))
return b2i(enc.Encode(toJSONOut(results, query, root)))
}
fmt.Print(toYAML(out, 0))
return 0
@@ -114,44 +101,6 @@ func searchHits(query, root, repo string, limit int) ([]Hit, error) {
return rank.RankAndFilter(fts, vec, root, repo, limit), nil
}
func attachHops(hits []Hit, n int) error {
if conn == nil {
return fmt.Errorf("brain not open")
}
for i := range hits {
var hops []rank.HopNode
for d := 1; d <= n; d++ {
stmt, err := conn.Prepare(rank.HopStmt(d))
if err != nil {
return err
}
res, err := conn.Execute(stmt, map[string]any{"id": hits[i].ID})
stmt.Close()
if err != nil {
return err
}
for res.HasNext() {
row, err := res.Next()
if err != nil {
return err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 3 {
continue
}
hops = append(hops, rank.HopNode{
ID: fmt.Sprint(vals[0]),
Label: rank.HopLabel(d),
Name: fmt.Sprint(vals[1]),
Depth: int(asInt(vals[2])),
})
}
}
hits[i].Hops = hops
}
return nil
}
func b2i(err error) int {
if err != nil {
return 1
@@ -223,7 +172,6 @@ type jsonOut struct {
RootFilter string `json:"root_filter"`
Count int `json:"count"`
Results []jsonHit `json:"results"`
Web *rank.SecondSource `json:"web,omitempty"`
}
type jsonHit struct {
@@ -232,10 +180,9 @@ type jsonHit struct {
Root string `json:"root"`
Score float64 `json:"score"`
Snippet string `json:"snippet,omitempty"`
Hops []rank.HopNode `json:"hops,omitempty"`
}
func toJSONOut(hits []Hit, query, rootFilter string, web *rank.SecondSource) *jsonOut {
func toJSONOut(hits []Hit, query, rootFilter string) *jsonOut {
out := make([]jsonHit, len(hits))
for i, h := range hits {
out[i] = jsonHit{
@@ -244,7 +191,6 @@ func toJSONOut(hits []Hit, query, rootFilter string, web *rank.SecondSource) *js
Root: h.Root,
Score: h.Score,
Snippet: h.Snippet,
Hops: h.Hops,
}
}
return &jsonOut{
@@ -252,7 +198,6 @@ func toJSONOut(hits []Hit, query, rootFilter string, web *rank.SecondSource) *js
RootFilter: rootFilter,
Count: len(hits),
Results: out,
Web: web,
}
}
@@ -268,18 +213,6 @@ func resultsToDicts(hits []Hit) []any {
if h.Snippet != "" {
d = append(d, KV{"snippet", h.Snippet})
}
if len(h.Hops) > 0 {
nodes := make([]any, len(h.Hops))
for j, n := range h.Hops {
nodes[j] = Dict{
{"id", n.ID},
{"label", n.Label},
{"name", n.Name},
{"depth", n.Depth},
}
}
d = append(d, KV{"hops", nodes})
}
out[i] = d
}
return out
-56
View File
@@ -1,56 +0,0 @@
package brain
import (
"context"
"github.com/eSlider/2dph/internal/brain/rank"
"github.com/eSlider/2dph/internal/websearch"
)
func lookupWeb(ctx context.Context, query string) rank.SecondSource {
o := websearch.Lookup(ctx, query, websearch.LookupOpt{Limit: 5})
return toSecond(o)
}
func toSecond(o websearch.Output) rank.SecondSource {
hits := make([]rank.SecondSourceHit, 0, len(o.Results))
for _, h := range o.Results {
hits = append(hits, rank.SecondSourceHit{
Rank: h.Rank,
Title: h.Title,
URL: h.URL,
Snippet: h.Snippet,
Engine: h.Engine,
})
}
return rank.SecondSource{
Status: o.Status,
Note: o.Note,
Cached: o.Cached,
Results: hits,
}
}
func secondToDict(w rank.SecondSource) Dict {
d := Dict{
{"status", w.Status},
}
if w.Note != "" {
d = append(d, KV{"note", w.Note})
}
if w.Cached {
d = append(d, KV{"cached", true})
}
rows := make([]any, 0, len(w.Results))
for _, h := range w.Results {
rows = append(rows, Dict{
{"rank", h.Rank},
{"title", h.Title},
{"url", h.URL},
{"snippet", h.Snippet},
{"engine", h.Engine},
})
}
d = append(d, KV{"results", rows})
return d
}
-47
View File
@@ -1,47 +0,0 @@
// Package duckstats runs in-process DuckDB for columnar aggregates.
// Graph facts stay in Ladybug. Web-search KV cache stays modernc sqlite.
package duckstats
import (
"database/sql"
"fmt"
_ "github.com/duckdb/duckdb-go/v2"
)
type Stats struct {
N int `json:"n"`
Min float64 `json:"min"`
P50 float64 `json:"p50"`
P95 float64 `json:"p95"`
Max float64 `json:"max"`
Avg float64 `json:"avg"`
}
func Quantiles(samples []float64) (Stats, error) {
if len(samples) == 0 {
return Stats{}, fmt.Errorf("duckstats: empty samples")
}
db, err := sql.Open("duckdb", "")
if err != nil {
return Stats{}, err
}
defer db.Close()
var s Stats
err = db.QueryRow(`
SELECT count(v), min(v), quantile_cont(v, 0.5), quantile_cont(v, 0.95), max(v), avg(v)
FROM (SELECT unnest(?) AS v)`, samples).Scan(
&s.N, &s.Min, &s.P50, &s.P95, &s.Max, &s.Avg)
return s, err
}
func CountJSONL(path string) (int64, error) {
db, err := sql.Open("duckdb", "")
if err != nil {
return 0, err
}
defer db.Close()
var n int64
err = db.QueryRow(`SELECT count(*) FROM read_json_auto(?)`, path).Scan(&n)
return n, err
}
-51
View File
@@ -1,51 +0,0 @@
package duckstats
import (
"os"
"testing"
)
func TestQuantilesEmpty(t *testing.T) {
_, err := Quantiles(nil)
if err == nil {
t.Fatal("empty slice must error")
}
}
func TestQuantilesOdd(t *testing.T) {
s, err := Quantiles([]float64{1, 2, 3, 4, 5})
if err != nil {
t.Fatal(err)
}
if s.N != 5 {
t.Fatalf("n=%d", s.N)
}
if s.Min != 1 || s.Max != 5 {
t.Fatalf("min=%v max=%v", s.Min, s.Max)
}
if s.P50 != 3 {
t.Fatalf("p50=%v want 3", s.P50)
}
if s.Avg != 3 {
t.Fatalf("avg=%v want 3", s.Avg)
}
if s.P95 < 4.5 || s.P95 > 5 {
t.Fatalf("p95=%v want in [4.5,5]", s.P95)
}
}
func TestCountJSONL(t *testing.T) {
dir := t.TempDir()
p := dir + "/rows.jsonl"
body := "{\"ms\":1}\n{\"ms\":2}\n{\"ms\":3}\n"
if err := os.WriteFile(p, []byte(body), 0o600); err != nil {
t.Fatal(err)
}
n, err := CountJSONL(p)
if err != nil {
t.Fatal(err)
}
if n != 3 {
t.Fatalf("count=%d want 3", n)
}
}
-193
View File
@@ -1,193 +0,0 @@
package httpapi
import (
"encoding/json"
"fmt"
"io"
"net/http"
"strconv"
"strings"
)
type rpcReq struct {
JSONRPC string `json:"jsonrpc"`
ID json.RawMessage `json:"id"`
Method string `json:"method"`
Params json.RawMessage `json:"params"`
}
type rpcErr struct {
Code int `json:"code"`
Message string `json:"message"`
}
func (s *Server) handleOpenAPI(w http.ResponseWriter, _ *http.Request) {
writeJSON(w, http.StatusOK, OpenAPI())
}
func (s *Server) handleMCP(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodPost {
writeJSON(w, http.StatusMethodNotAllowed, map[string]any{"error": "POST JSON-RPC"})
return
}
raw, err := io.ReadAll(io.LimitReader(r.Body, 1<<20))
if err != nil {
writeJSON(w, http.StatusBadRequest, map[string]any{"error": "read body"})
return
}
var req rpcReq
if err := json.Unmarshal(raw, &req); err != nil {
writeJSON(w, http.StatusOK, rpcResult(nil, nil, &rpcErr{-32700, "parse error"}))
return
}
result, rpcErrv, callErr := s.mcpDispatch(r, req)
if callErr != nil {
writeJSON(w, http.StatusOK, rpcResult(req.ID, nil, &rpcErr{-32603, callErr.Error()}))
return
}
writeJSON(w, http.StatusOK, rpcResult(req.ID, result, rpcErrv))
}
func (s *Server) mcpDispatch(r *http.Request, req rpcReq) (any, *rpcErr, error) {
switch req.Method {
case "initialize":
return map[string]any{
"protocolVersion": "2024-11-05",
"capabilities": map[string]any{"tools": map[string]any{}},
"serverInfo": map[string]any{"name": "2dph", "version": "1"},
}, nil, nil
case "notifications/initialized", "notifications/cancelled":
return map[string]any{}, nil, nil
case "tools/list":
return map[string]any{"tools": MCPTools()}, nil, nil
case "tools/call":
out, err := s.mcpCall(r, req.Params)
return out, nil, err
case "ping":
return map[string]any{}, nil, nil
default:
return nil, &rpcErr{-32601, "method not found"}, nil
}
}
func (s *Server) mcpCall(r *http.Request, params json.RawMessage) (any, error) {
var p struct {
Name string `json:"name"`
Arguments map[string]any `json:"arguments"`
}
if err := json.Unmarshal(params, &p); err != nil {
return nil, fmt.Errorf("params")
}
if p.Arguments == nil {
p.Arguments = map[string]any{}
}
var (
body []byte
err error
)
switch p.Name {
case "search":
q := strings.TrimSpace(fmt.Sprint(p.Arguments["q"]))
if q == "" || q == "<nil>" {
return mcpText(`{"error":"q required"}`, true), nil
}
limit := 10
if raw, ok := p.Arguments["n"]; ok {
switch n := raw.(type) {
case float64:
limit = int(n)
case string:
if v, e := strconv.Atoi(n); e == nil {
limit = v
}
}
}
if limit < 1 || limit > 100 {
return mcpText(`{"error":"n must be int 1..100"}`, true), nil
}
if !s.tryAcquire(r) {
return nil, fmt.Errorf("cancelled")
}
defer s.release()
body, err = s.api.Search(r.Context(), q, limit)
case "get":
id := strings.TrimSpace(fmt.Sprint(p.Arguments["id"]))
if id == "" || id == "<nil>" {
return mcpText(`{"error":"id required"}`, true), nil
}
full := false
switch v := p.Arguments["body"].(type) {
case bool:
full = v
case string:
full = v == "1" || v == "true"
}
if !s.tryAcquire(r) {
return nil, fmt.Errorf("cancelled")
}
defer s.release()
body, err = s.api.Get(r.Context(), id, full)
case "stats":
if !s.tryAcquire(r) {
return nil, fmt.Errorf("cancelled")
}
defer s.release()
body, err = s.api.Stats(r.Context())
case "audit":
if !s.tryAcquire(r) {
return nil, fmt.Errorf("cancelled")
}
defer s.release()
body, err = s.api.Audit(r.Context())
case "ingest":
if !s.tryAcquire(r) {
return nil, fmt.Errorf("cancelled")
}
defer s.release()
var payload []byte
text := strings.TrimSpace(fmt.Sprint(p.Arguments["text"]))
if text != "" && text != "<nil>" {
payload, err = json.Marshal(p.Arguments)
if err != nil {
return nil, err
}
}
body, err = s.api.Ingest(r.Context(), payload)
default:
return nil, fmt.Errorf("unknown tool %s", p.Name)
}
if err != nil {
return mcpText(err.Error(), true), nil
}
return mcpText(string(body), false), nil
}
func mcpText(text string, isError bool) map[string]any {
return map[string]any{
"content": []any{map[string]any{"type": "text", "text": text}},
"isError": isError,
}
}
type rpcResp struct {
JSONRPC string `json:"jsonrpc"`
ID json.RawMessage `json:"id"`
Result any `json:"result,omitempty"`
Error *rpcErr `json:"error,omitempty"`
}
func rpcResult(id json.RawMessage, result any, err *rpcErr) rpcResp {
out := rpcResp{JSONRPC: "2.0", ID: id}
if len(id) == 0 {
out.ID = []byte("null")
}
if err != nil {
out.Error = err
return out
}
if result == nil {
result = map[string]any{}
}
out.Result = result
return out
}
+11 -63
View File
@@ -11,7 +11,6 @@ import (
"context"
"encoding/json"
"errors"
"io"
"log"
"net/http"
"os"
@@ -28,7 +27,7 @@ type API interface {
Get(ctx context.Context, id string, body bool) ([]byte, error)
Stats(ctx context.Context) ([]byte, error)
Audit(ctx context.Context) ([]byte, error)
Ingest(ctx context.Context, body []byte) ([]byte, error)
Ingest(ctx context.Context) ([]byte, error)
}
type Server struct {
@@ -49,22 +48,18 @@ func NewServer(api API, workers int) http.Handler {
func (s *Server) ServeHTTP(w http.ResponseWriter, r *http.Request) {
switch r.URL.Path {
case PathHealth:
case "/health":
writeJSON(w, http.StatusOK, map[string]any{"status": "ok"})
case PathSearch:
case "/search":
s.handleSearch(w, r)
case PathGet:
case "/get":
s.handleGet(w, r)
case PathStats:
case "/stats":
s.handleJSON(w, r, s.api.Stats)
case PathAudit:
case "/audit":
s.handleJSON(w, r, s.api.Audit)
case PathIngest:
s.handleIngest(w, r)
case PathOpenAPI:
s.handleOpenAPI(w, r)
case PathMCP:
s.handleMCP(w, r)
case "/ingest":
s.handleJSON(w, r, s.api.Ingest)
default:
writeJSON(w, http.StatusNotFound, map[string]any{"error": "not found"})
}
@@ -117,34 +112,6 @@ func (s *Server) handleJSON(w http.ResponseWriter, r *http.Request, fn func(cont
writeAPI(w, body, err)
}
func (s *Server) handleIngest(w http.ResponseWriter, r *http.Request) {
var raw []byte
if r.Method == http.MethodPost {
b, err := io.ReadAll(io.LimitReader(r.Body, 1<<20))
if err != nil {
writeJSON(w, http.StatusBadRequest, map[string]any{"error": "read body"})
return
}
raw = b
}
if !s.acquire(w, r) {
return
}
defer s.release()
body, err := s.api.Ingest(r.Context(), raw)
writeAPI(w, body, err)
}
func (s *Server) tryAcquire(r *http.Request) bool {
return s.acquire(nopWriter{}, r)
}
type nopWriter struct{}
func (nopWriter) Header() http.Header { return http.Header{} }
func (nopWriter) Write([]byte) (int, error) { return 0, nil }
func (nopWriter) WriteHeader(int) {}
func (s *Server) acquire(w http.ResponseWriter, r *http.Request) bool {
select {
case s.semaphore <- struct{}{}:
@@ -210,31 +177,12 @@ func (ExecSearcher) Get(context.Context, string, bool) ([]byte, error) {
}
func (ExecSearcher) Stats(context.Context) ([]byte, error) { return nil, errUnimplemented }
func (ExecSearcher) Audit(context.Context) ([]byte, error) { return nil, errUnimplemented }
func (b ExecSearcher) Ingest(ctx context.Context, body []byte) ([]byte, error) {
if len(strings.TrimSpace(string(body))) == 0 {
func (ExecSearcher) Ingest(context.Context) ([]byte, error) {
return json.Marshal(map[string]any{
"mode": "add",
"command": "bin/brain/add.go",
"rebuild": "bin/brain/index.go --rebuild",
"mode": "rebuild",
"command": "bin/brain/index.go --rebuild",
})
}
root := os.Getenv("KB_ROOT")
if root == "" {
root = "."
}
cmd := exec.CommandContext(ctx, filepath.Join(root, "bin/kb/add"), "--json")
cmd.Stdin = strings.NewReader(string(body))
cmd.Dir = root
out, err := cmd.Output()
if err != nil {
var exitErr *exec.ExitError
if errors.As(err, &exitErr) {
return nil, errors.New("add failed: " + strings.TrimSpace(string(exitErr.Stderr)))
}
return nil, err
}
return out, nil
}
func defaultSearchCmd(root string) string {
if env := os.Getenv("KB_SEARCH_CMD"); env != "" {
+2 -27
View File
@@ -1,7 +1,6 @@
package httpapi
import (
"bytes"
"context"
"encoding/json"
"net/http"
@@ -65,11 +64,8 @@ func (f *fakeSearcher) Audit(context.Context) ([]byte, error) {
return []byte(`{"status":"ok"}`), nil
}
func (f *fakeSearcher) Ingest(_ context.Context, body []byte) ([]byte, error) {
if len(bytes.TrimSpace(body)) == 0 {
return []byte(`{"mode":"add","command":"bin/brain/add.go"}`), nil
}
return []byte(`{"mode":"add","ids":["fake-leaf"]}`), nil
func (f *fakeSearcher) Ingest(context.Context) ([]byte, error) {
return []byte(`{"mode":"rebuild","command":"bin/brain/index.go --rebuild"}`), nil
}
func (f *fakeSearcher) count() int {
@@ -197,27 +193,6 @@ func TestStatsAuditIngest(t *testing.T) {
}
}
func TestIngestIsAddNotRebuildHint(t *testing.T) {
h := NewServer(&fakeSearcher{}, 1)
code, body := get(t, h, "/ingest")
if code != http.StatusOK {
t.Fatalf("GET /ingest code = %d body=%s", code, body)
}
if strings.Contains(string(body), `"add":"v2"`) || strings.Contains(string(body), "write is v2") {
t.Fatalf("GET /ingest still a v2 hint: %s", body)
}
if !strings.Contains(string(body), "bin/brain/add.go") {
t.Fatalf("GET /ingest should name add.go: %s", body)
}
code, body = postJSON(t, h, "/ingest", `{"text":"hello","root":"info","source":"t"}`)
if code != http.StatusOK {
t.Fatalf("POST /ingest code = %d body=%s", code, body)
}
if !strings.Contains(string(body), "fake-leaf") {
t.Fatalf("POST /ingest should add: %s", body)
}
}
func TestHTTPPackageDoesNotExecPython(t *testing.T) {
raw, err := os.ReadFile("server.go")
if err != nil {
-152
View File
@@ -1,152 +0,0 @@
package httpapi
import "strings"
// Shared HTTP surface: OpenAPI paths and MCP tools are generated from Ops.
// ServeHTTP must keep the same path strings.
type Param struct {
Name, In, Type, Description string
Required bool
}
type Op struct {
Path, Method, ID, Summary string
Params []Param
MCP bool
}
const (
PathHealth = "/health"
PathSearch = "/search"
PathGet = "/get"
PathStats = "/stats"
PathAudit = "/audit"
PathIngest = "/ingest"
PathOpenAPI = "/openapi.json"
PathMCP = "/mcp"
)
var Ops = []Op{
{Path: PathHealth, Method: "get", ID: "health", Summary: "liveness"},
{
Path: PathSearch, Method: "get", ID: "search", Summary: "deduction search (facts → info → web)",
MCP: true,
Params: []Param{
{Name: "q", In: "query", Type: "string", Description: "search query", Required: true},
{Name: "n", In: "query", Type: "integer", Description: "hit limit 1..100 (default 10)"},
},
},
{
Path: PathGet, Method: "get", ID: "get", Summary: "read one leaf by id",
MCP: true,
Params: []Param{
{Name: "id", In: "query", Type: "string", Description: "leaf id", Required: true},
{Name: "body", In: "query", Type: "boolean", Description: "include full text"},
},
},
{Path: PathStats, Method: "get", ID: "stats", Summary: "index health", MCP: true},
{Path: PathAudit, Method: "get", ID: "audit", Summary: "facts confidence histogram", MCP: true},
{
Path: PathIngest, Method: "post", ID: "ingest", Summary: "add a leaf without rebuild",
MCP: true,
Params: []Param{
{Name: "text", In: "query", Type: "string", Description: "leaf text (omit for CLI hint)"},
{Name: "root", In: "query", Type: "string", Description: "facts or info (default info)"},
{Name: "source", In: "query", Type: "string", Description: "evidence pointer; facts need two sources"},
},
},
{Path: PathOpenAPI, Method: "get", ID: "openapi", Summary: "OpenAPI 3 document for this server"},
}
func OpenAPI() map[string]any {
paths := map[string]any{}
for _, op := range Ops {
params := make([]any, 0, len(op.Params))
for _, p := range op.Params {
params = append(params, map[string]any{
"name": p.Name,
"in": p.In,
"required": p.Required,
"description": p.Description,
"schema": map[string]any{"type": p.Type},
})
}
item := map[string]any{
"operationId": op.ID,
"summary": op.Summary,
"responses": map[string]any{
"200": map[string]any{
"description": "JSON",
"content": map[string]any{
"application/json": map[string]any{
"schema": map[string]any{"type": "object"},
},
},
},
},
}
if len(params) > 0 {
item["parameters"] = params
}
paths[op.Path] = map[string]any{op.Method: item}
}
return map[string]any{
"openapi": "3.0.3",
"info": map[string]any{
"title": "2dph brain",
"version": "1",
"description": "Same handlers as bin/brain/serve.go. MCP tools at POST /mcp match these paths.",
},
"paths": paths,
}
}
type MCPTool struct {
Name string `json:"name"`
Description string `json:"description"`
InputSchema map[string]any `json:"inputSchema"`
}
func MCPTools() []MCPTool {
out := make([]MCPTool, 0, len(Ops))
for _, op := range Ops {
if !op.MCP {
continue
}
props := map[string]any{}
var required []string
for _, p := range op.Params {
props[p.Name] = map[string]any{"type": p.Type, "description": p.Description}
if p.Required {
required = append(required, p.Name)
}
}
schema := map[string]any{"type": "object", "properties": props}
if len(required) > 0 {
schema["required"] = required
}
out = append(out, MCPTool{
Name: op.ID,
Description: op.Summary,
InputSchema: schema,
})
}
return out
}
// SkillMarkdown is the Cursor skill fragment generated from Ops/MCPTools.
func SkillMarkdown() string {
var b strings.Builder
b.WriteString("# brain HTTP / MCP tools\n\n")
b.WriteString("Generated from `internal/httpapi.Ops`. Do not edit by hand.\n\n")
b.WriteString("Serve: `bin/brain/serve.go` (`GET /openapi.json`, `POST /mcp`).\n\n")
for _, t := range MCPTools() {
b.WriteString("- `")
b.WriteString(t.Name)
b.WriteString("` — ")
b.WriteString(t.Description)
b.WriteString("\n")
}
return b.String()
}
-112
View File
@@ -1,112 +0,0 @@
package httpapi
import (
"encoding/json"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"strings"
"testing"
)
func TestOpenAPIIncludesCorePaths(t *testing.T) {
doc := OpenAPI()
raw, err := json.Marshal(doc)
if err != nil {
t.Fatal(err)
}
paths, _ := doc["paths"].(map[string]any)
for _, p := range []string{"/search", "/get", "/stats", "/audit"} {
if _, ok := paths[p]; !ok {
t.Fatalf("openapi missing path %s (%s)", p, raw)
}
}
}
func TestMCPToolsMatchOpenAPIPaths(t *testing.T) {
paths, _ := OpenAPI()["paths"].(map[string]any)
tools := MCPTools()
if len(tools) == 0 {
t.Fatal("no MCP tools")
}
names := map[string]bool{}
for _, tool := range tools {
names[tool.Name] = true
path := "/" + tool.Name
if _, ok := paths[path]; !ok {
t.Fatalf("MCP tool %s has no OpenAPI path %s", tool.Name, path)
}
}
for _, need := range []string{"search", "get", "stats", "audit", "ingest"} {
if !names[need] {
t.Fatalf("MCP tools missing %s: %v", need, names)
}
}
var ingest MCPTool
for _, tool := range tools {
if tool.Name == "ingest" {
ingest = tool
break
}
}
if strings.Contains(ingest.Description, "v2") {
t.Fatalf("ingest still a v2 hint: %s", ingest.Description)
}
if !strings.Contains(ingest.Description, "add") {
t.Fatalf("ingest should describe add: %s", ingest.Description)
}
}
func TestOpenAPIHTTP(t *testing.T) {
h := NewServer(&fakeSearcher{}, 1)
code, body := get(t, h, "/openapi.json")
if code != http.StatusOK {
t.Fatalf("code = %d body=%s", code, body)
}
var doc map[string]any
if err := json.Unmarshal(body, &doc); err != nil {
t.Fatalf("not json: %v", err)
}
if doc["openapi"] == nil {
t.Fatalf("missing openapi version: %s", body)
}
}
func TestMCPToolsListAndCall(t *testing.T) {
h := NewServer(&fakeSearcher{}, 1)
code, body := postJSON(t, h, "/mcp", `{"jsonrpc":"2.0","id":1,"method":"tools/list"}`)
if code != http.StatusOK {
t.Fatalf("list code = %d body=%s", code, body)
}
if !strings.Contains(string(body), `"search"`) {
t.Fatalf("tools/list missing search: %s", body)
}
code, body = postJSON(t, h, "/mcp", `{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"search","arguments":{"q":"matrix","n":3}}}`)
if code != http.StatusOK {
t.Fatalf("call code = %d body=%s", code, body)
}
if !strings.Contains(string(body), "matrix") {
t.Fatalf("search call body %s", body)
}
}
func TestSkillMarkdownMatchesCommittedFile(t *testing.T) {
want, err := os.ReadFile(filepath.Join("..", "..", "skills", "brain", "tools.md"))
if err != nil {
t.Fatal(err)
}
got := SkillMarkdown()
if got != string(want) {
t.Fatalf("skills/brain/tools.md stale; regenerate from SkillMarkdown()\n--- got ---\n%s\n--- want ---\n%s", got, want)
}
}
func postJSON(t *testing.T, h http.Handler, path, raw string) (int, []byte) {
t.Helper()
req := httptest.NewRequest(http.MethodPost, path, strings.NewReader(raw))
req.Header.Set("Content-Type", "application/json")
rec := httptest.NewRecorder()
h.ServeHTTP(rec, req)
return rec.Code, rec.Body.Bytes()
}
-146
View File
@@ -1,146 +0,0 @@
package mdleaves
import (
"encoding/json"
"os"
"path/filepath"
"regexp"
"strings"
)
type Leaf struct {
Source string `json:"source"`
Repo string `json:"repo"`
Heading string `json:"heading"`
Text string `json:"text"`
Type string `json:"type"`
Status string `json:"status"`
Related string `json:"related"`
}
type chunk struct {
Heading string
Text string
}
var (
h1 = regexp.MustCompile(`^# \S`)
h2 = regexp.MustCompile(`^## \S`)
)
func ExtractFrontmatter(text string) (map[string]string, string) {
if !strings.HasPrefix(text, "---") {
return map[string]string{}, text
}
end := strings.Index(text[3:], "\n---")
if end == -1 {
return map[string]string{}, text
}
fm := strings.TrimSpace(text[3 : 3+end])
body := text[3+end+4:]
meta := map[string]string{}
for _, line := range strings.Split(fm, "\n") {
key, value, ok := strings.Cut(line, ":")
if !ok {
continue
}
meta[strings.TrimSpace(key)] = strings.Trim(strings.TrimSpace(value), `"'`)
}
return meta, body
}
func SplitLeafs(meta map[string]string, body string) []chunk {
lines := strings.Split(body, "\n")
title := ""
type hdr struct {
heading string
start int
}
var headers []hdr
for i, line := range lines {
switch {
case h1.MatchString(line):
title = strings.TrimSpace(strings.TrimLeft(line, "#"))
case h2.MatchString(line):
headers = append(headers, hdr{strings.TrimSpace(strings.TrimLeft(line, "#")), i})
}
}
if len(headers) == 0 {
var kept []string
for _, l := range lines {
if strings.TrimSpace(l) != "" {
kept = append(kept, l)
}
}
return []chunk{{Heading: title, Text: strings.TrimSpace(strings.Join(kept, "\n"))}}
}
out := make([]chunk, 0, len(headers))
for idx, h := range headers {
end := len(lines)
if idx+1 < len(headers) {
end = headers[idx+1].start
}
var kept []string
for _, l := range lines[h.start:end] {
if strings.TrimSpace(l) != "" {
kept = append(kept, l)
}
}
text := strings.Join(kept, "\n")
if idx == 0 && title != "" {
text = title + "\n\n" + text
}
out = append(out, chunk{Heading: h.heading, Text: strings.TrimSpace(text)})
}
return out
}
func ToAll(text, path, repo string) []Leaf {
meta, body := ExtractFrontmatter(text)
if meta["type"] == "" {
meta["type"] = "reference"
}
if meta["status"] == "" {
meta["status"] = "current"
}
chunks := SplitLeafs(meta, body)
out := make([]Leaf, 0, len(chunks))
for _, c := range chunks {
out = append(out, Leaf{
Source: path,
Repo: repo,
Heading: c.Heading,
Text: c.Text,
Type: meta["type"],
Status: meta["status"],
Related: meta["related"],
})
}
return out
}
func WalkMarkdown(root string) ([]string, error) {
var out []string
err := filepath.Walk(root, func(p string, info os.FileInfo, err error) error {
if err != nil {
return err
}
if info.IsDir() {
return nil
}
ext := strings.ToLower(filepath.Ext(p))
if ext == ".md" || ext == ".markdown" {
out = append(out, p)
}
return nil
})
return out, err
}
func EncodeJSON(leafs []Leaf) (string, error) {
raw, err := json.MarshalIndent(leafs, "", " ")
if err != nil {
return "", err
}
return string(raw) + "\n", nil
}
-73
View File
@@ -1,73 +0,0 @@
package mdleaves
import (
"strings"
"testing"
)
func TestSplitLeafsOnH2(t *testing.T) {
body := "# Title\n\n## One\n\nalpha\n\n## Two\n\nbeta\n"
leafs := SplitLeafs(map[string]string{}, body)
if len(leafs) != 2 {
t.Fatalf("n=%d", len(leafs))
}
if leafs[0].Heading != "One" || !strings.Contains(leafs[0].Text, "Title") {
t.Fatalf("first=%+v", leafs[0])
}
if leafs[1].Heading != "Two" || strings.Contains(leafs[1].Text, "Title") {
t.Fatalf("second=%+v", leafs[1])
}
}
func TestSplitLeafsNoH2IsWholeDoc(t *testing.T) {
body := "# Title\n\njust a paragraph\n"
leafs := SplitLeafs(nil, body)
if len(leafs) != 1 {
t.Fatalf("n=%d", len(leafs))
}
if leafs[0].Heading != "Title" {
t.Fatalf("heading=%q", leafs[0].Heading)
}
if !strings.Contains(leafs[0].Text, "just a paragraph") {
t.Fatalf("text=%q", leafs[0].Text)
}
}
func TestFrontmatterAndToAll(t *testing.T) {
raw := "---\ntype: howto\nstatus: current\nrelated: docs/design.md\n---\n# Doc\n\n## Step\n\ndo it\n"
got := ToAll(raw, "docs/x.md", "eSlider/2dph")
if len(got) != 1 {
t.Fatalf("n=%d", len(got))
}
if got[0].Type != "howto" || got[0].Status != "current" {
t.Fatalf("%+v", got[0])
}
if got[0].Source != "docs/x.md" || got[0].Repo != "eSlider/2dph" {
t.Fatalf("%+v", got[0])
}
if got[0].Related != "docs/design.md" {
t.Fatalf("related=%q", got[0].Related)
}
}
func TestEncodeJSONAndYAML(t *testing.T) {
leafs := ToAll("# Hi\n\nbody\n", "a.md", "")
js, err := EncodeJSON(leafs)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(js, `"heading": "Hi"`) {
t.Fatalf("json=%s", js)
}
y := EncodeYAML(leafs)
if !strings.Contains(y, "heading: Hi") {
t.Fatalf("yaml=%s", y)
}
}
func TestDefaultsTypeAndStatus(t *testing.T) {
got := ToAll("# Hi\n\nbody\n", "a.md", "")
if got[0].Type != "reference" || got[0].Status != "current" {
t.Fatalf("%+v", got[0])
}
}
-36
View File
@@ -1,36 +0,0 @@
package mdleaves
import (
"fmt"
"strconv"
"strings"
)
func EncodeYAML(leafs []Leaf) string {
if len(leafs) == 0 {
return "[]\n"
}
var b strings.Builder
for _, lf := range leafs {
b.WriteString("-\n")
writeKV(&b, "source", lf.Source)
writeKV(&b, "repo", lf.Repo)
writeKV(&b, "heading", lf.Heading)
writeKV(&b, "text", lf.Text)
writeKV(&b, "type", lf.Type)
writeKV(&b, "status", lf.Status)
writeKV(&b, "related", lf.Related)
}
return b.String()
}
func writeKV(b *strings.Builder, k, v string) {
fmt.Fprintf(b, " %s: %s\n", k, yamlScalar(v))
}
func yamlScalar(s string) string {
if strings.Contains(s, "\n") || s == "" || strings.ContainsAny(s, ":#'\"[]{}&*!|>%@`") || s != strings.TrimSpace(s) {
return strconv.Quote(s)
}
return s
}
-162
View File
@@ -1,162 +0,0 @@
// Package ocr runs Tesseract (eng+deu) on images and scanned PDFs.
//
// Default engine is the tesseract CLI, not gosseract CGO: Ladybug CGO stays
// Zig-only (D21). Same engine, no gocv. OCR_ENGINE=paddle selects paddleocr
// when that binary is on PATH (compose profile ocr-paddle).
package ocr
import (
"fmt"
"image"
"image/color"
"image/png"
"os"
"os/exec"
"path/filepath"
"strings"
)
const TessLang = "eng+deu"
func ImageFile(path string) (string, error) {
engine := os.Getenv("OCR_ENGINE")
if engine == "paddle" {
return runPaddle(path)
}
return runTesseract(path)
}
func PDFFile(path string) (string, error) {
text, err := pdfToText(path)
if err == nil && strings.TrimSpace(text) != "" {
return strings.TrimSpace(text), nil
}
ocr, oerr := pdfPages(path)
if oerr != nil {
if err != nil {
return "", err
}
return "", oerr
}
if strings.TrimSpace(ocr) != "" {
return strings.TrimSpace(ocr), nil
}
if text != "" {
return strings.TrimSpace(text), nil
}
return "", fmt.Errorf("pdf has no text layer (ocr unavailable)")
}
func pdfToText(path string) (string, error) {
cmd := exec.Command("pdftotext", "-layout", path, "-")
out, err := cmd.Output()
if err != nil {
return "", err
}
return string(out), nil
}
func pdfPages(path string) (string, error) {
dir, err := os.MkdirTemp("", "2dph-ocr-")
if err != nil {
return "", err
}
defer os.RemoveAll(dir)
prefix := filepath.Join(dir, "page")
cmd := exec.Command("pdftoppm", "-png", "-r", "200", path, prefix)
if err := cmd.Run(); err != nil {
return "", err
}
matches, err := filepath.Glob(prefix + "*.png")
if err != nil {
return "", err
}
var parts []string
for _, img := range matches {
t, err := ImageFile(img)
if err != nil {
continue
}
if s := strings.TrimSpace(t); s != "" {
parts = append(parts, s)
}
}
return strings.Join(parts, "\n\n"), nil
}
func runTesseract(path string) (string, error) {
pre, err := preprocessFile(path)
if err != nil {
pre = path
} else {
defer os.Remove(pre)
}
cmd := exec.Command("tesseract", pre, "stdout", "-l", TessLang, "--psm", "6")
out, err := cmd.Output()
if err != nil {
return "", err
}
return strings.TrimSpace(string(out)), nil
}
func runPaddle(path string) (string, error) {
cmd := exec.Command("paddleocr", "ocr", "-i", path)
out, err := cmd.Output()
if err != nil {
return "", err
}
return strings.TrimSpace(string(out)), nil
}
func preprocessFile(path string) (string, error) {
f, err := os.Open(path)
if err != nil {
return "", err
}
defer f.Close()
img, err := png.Decode(f)
if err != nil {
return "", err
}
out := filepath.Join(os.TempDir(), filepath.Base(path)+".gray.png")
w, err := os.Create(out)
if err != nil {
return "", err
}
defer w.Close()
if err := png.Encode(w, GrayContrast(img)); err != nil {
os.Remove(out)
return "", err
}
return out, nil
}
// GrayContrast is a stdlib preprocess (no gocv): grayscale + stretch.
func GrayContrast(src image.Image) image.Image {
b := src.Bounds()
dst := image.NewGray(b)
var minL, maxL uint8 = 255, 0
for y := b.Min.Y; y < b.Max.Y; y++ {
for x := b.Min.X; x < b.Max.X; x++ {
g := color.GrayModel.Convert(src.At(x, y)).(color.Gray)
if g.Y < minL {
minL = g.Y
}
if g.Y > maxL {
maxL = g.Y
}
}
}
span := int(maxL) - int(minL)
if span < 1 {
span = 1
}
for y := b.Min.Y; y < b.Max.Y; y++ {
for x := b.Min.X; x < b.Max.X; x++ {
g := color.GrayModel.Convert(src.At(x, y)).(color.Gray)
v := uint8((int(g.Y) - int(minL)) * 255 / span)
dst.SetGray(x, y, color.Gray{Y: v})
}
}
return dst
}
-54
View File
@@ -1,54 +0,0 @@
package ocr
import (
"image"
"image/color"
"os/exec"
"path/filepath"
"strings"
"testing"
)
func TestGrayContrastStretches(t *testing.T) {
img := image.NewGray(image.Rect(0, 0, 2, 2))
img.SetGray(0, 0, color.Gray{Y: 64})
img.SetGray(0, 1, color.Gray{Y: 64})
img.SetGray(1, 0, color.Gray{Y: 64})
img.SetGray(1, 1, color.Gray{Y: 192})
out := GrayContrast(img).(*image.Gray)
if out.GrayAt(0, 0).Y != 0 {
t.Fatalf("min should map to 0, got %d", out.GrayAt(0, 0).Y)
}
if out.GrayAt(1, 1).Y != 255 {
t.Fatalf("max should map to 255, got %d", out.GrayAt(1, 1).Y)
}
}
func TestHelloPNGFixtureOCR(t *testing.T) {
if _, err := exec.LookPath("tesseract"); err != nil {
t.Skip("tesseract not installed")
}
path := filepath.Join("testdata", "hello.png")
got, err := ImageFile(path)
if err != nil {
t.Fatal(err)
}
up := strings.ToUpper(got)
if !strings.Contains(up, "HELLO") {
t.Fatalf("ocr %q missing HELLO", got)
}
}
func TestPaddleEngineUsesPaddleocrBinary(t *testing.T) {
t.Setenv("OCR_ENGINE", "paddle")
_, err := ImageFile(filepath.Join("testdata", "hello.png"))
if _, look := exec.LookPath("paddleocr"); look != nil {
if err == nil {
t.Fatal("expected error when paddleocr is missing")
}
return
}
if err != nil {
t.Fatal(err)
}
}
BIN
View File
Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.7 KiB

-327
View File
@@ -1,327 +0,0 @@
package reasoner
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"strings"
"time"
)
// HF IDs named in docs. No Qwen3.6-9B exists.
const (
HFQwen35_9B = "Qwen/Qwen3.5-9B"
HFQwen36_27B = "Qwen/Qwen3.6-27B"
HFBonsai27B = "prism-ml/Bonsai-27B-gguf"
OllamaRAM = "qwen3.5:9b"
OllamaQuality = "MichelRosselli/bonsai-27b:Q1_0"
)
type ToolCall struct {
Name string
Arguments string
}
type Result struct {
Model string `json:"model"`
OK bool `json:"ok"`
ToolName string `json:"tool_name,omitempty"`
XMLLeak bool `json:"xml_leak"`
Err string `json:"error,omitempty"`
LatencyMS int64 `json:"latency_ms"`
RSSMB int `json:"rss_mb,omitempty"`
Device string `json:"device"`
WantedTool string `json:"wanted_tool"`
}
type Prompt struct {
Name string
Want string
User string
}
var BakePrompts = []Prompt{
{
Name: "search-before-claim",
Want: "search",
User: "Use tools. Search the 2dph brain for LadybugDB before you answer. Call search.",
},
{
Name: "get-leaf",
Want: "get",
User: "Use tools. Fetch leaf id leaf-demo with get. Do not invent the body.",
},
{
Name: "audit-index",
Want: "audit",
User: "Use tools. Call audit on the brain index health.",
},
}
func MCPTools() []map[string]any {
return []map[string]any{
openaiTool("search", "deduction search (facts → info → web)", map[string]any{
"type": "object",
"properties": map[string]any{
"q": map[string]any{"type": "string", "description": "search query"},
"n": map[string]any{"type": "integer"},
},
"required": []string{"q"},
}),
openaiTool("get", "read one leaf by id", map[string]any{
"type": "object",
"properties": map[string]any{
"id": map[string]any{"type": "string"},
"body": map[string]any{"type": "boolean"},
},
"required": []string{"id"},
}),
openaiTool("audit", "facts confidence histogram", map[string]any{
"type": "object",
"properties": map[string]any{},
}),
}
}
func openaiTool(name, desc string, schema map[string]any) map[string]any {
return map[string]any{
"type": "function",
"function": map[string]any{
"name": name,
"description": desc,
"parameters": schema,
},
}
}
type Client struct {
BaseURL string
Model string
HTTP *http.Client
Device string
ToolChoice string
}
type Report struct {
Model string `json:"model"`
HF string `json:"hf_id"`
Device string `json:"device"`
ToolCallOK int `json:"tool_call_ok"`
ToolCallN int `json:"tool_call_n"`
XMLLeak int `json:"xml_leak"`
RSSMB int `json:"rss_mb"`
VRAMMB int `json:"vram_mb"`
LatencyP50MS float64 `json:"latency_p50_ms,omitempty"`
LatencyP95MS float64 `json:"latency_p95_ms,omitempty"`
Prompts []Result `json:"prompts"`
}
func HFFor(model string) string {
switch model {
case OllamaRAM:
return HFQwen35_9B
case OllamaQuality:
return HFBonsai27B
default:
if strings.Contains(model, "qwen3.6") || strings.Contains(model, "Qwen3.6") {
return HFQwen36_27B
}
return ""
}
}
func (c *Client) httpc() *http.Client {
if c.HTTP == nil {
c.HTTP = &http.Client{Timeout: 10 * time.Minute}
}
return c.HTTP
}
func Origin(base string) string {
s := strings.TrimRight(base, "/")
return strings.TrimSuffix(s, "/v1")
}
func (c Client) ChatTools(user string) (ToolCall, string, error) {
choice := c.ToolChoice
if choice == "" {
choice = "required"
}
body, _ := json.Marshal(map[string]any{
"model": c.Model,
"messages": []map[string]string{
{"role": "system", "content": "You are PicoClaw talking to 2dph MCP. Always call a tool before a factual claim. search then get then audit."},
{"role": "user", "content": user},
},
"tools": MCPTools(),
"tool_choice": choice,
})
base := strings.TrimRight(c.BaseURL, "/")
if !strings.HasSuffix(base, "/v1") {
base += "/v1"
}
url := base + "/chat/completions"
req, err := http.NewRequest(http.MethodPost, url, bytes.NewReader(body))
if err != nil {
return ToolCall{}, "", err
}
req.Header.Set("Content-Type", "application/json")
res, err := c.httpc().Do(req)
if err != nil {
return ToolCall{}, "", err
}
defer res.Body.Close()
raw, _ := io.ReadAll(res.Body)
if res.StatusCode >= 300 {
return ToolCall{}, string(raw), fmt.Errorf("http %d", res.StatusCode)
}
return ParseToolResponse(raw)
}
func ParseToolResponse(raw []byte) (ToolCall, string, error) {
var wrap struct {
Choices []struct {
Message struct {
Content string `json:"content"`
ToolCalls []struct {
Function struct {
Name string `json:"name"`
Arguments json.RawMessage `json:"arguments"`
} `json:"function"`
} `json:"tool_calls"`
} `json:"message"`
} `json:"choices"`
}
if err := json.Unmarshal(raw, &wrap); err != nil {
return ToolCall{}, "", err
}
content := ""
if len(wrap.Choices) > 0 {
content = wrap.Choices[0].Message.Content
if n := len(wrap.Choices[0].Message.ToolCalls); n > 0 {
fn := wrap.Choices[0].Message.ToolCalls[0].Function
return ToolCall{Name: fn.Name, Arguments: rawArgs(fn.Arguments)}, content, nil
}
}
return ToolCall{}, content, fmt.Errorf("no tool_calls")
}
func rawArgs(raw json.RawMessage) string {
if len(raw) == 0 {
return ""
}
var s string
if err := json.Unmarshal(raw, &s); err == nil {
return s
}
return string(raw)
}
func XMLLeak(content string) bool {
s := strings.ToLower(content)
return strings.Contains(s, "<tool_call>") ||
strings.Contains(s, "<function=") ||
strings.Contains(s, "<parameter")
}
func RunPrompt(c Client, p Prompt) Result {
start := time.Now()
tc, content, err := c.ChatTools(p.User)
out := Result{
Model: c.Model,
WantedTool: p.Want,
LatencyMS: time.Since(start).Milliseconds(),
Device: c.Device,
XMLLeak: XMLLeak(content),
}
if err != nil {
out.Err = err.Error()
if content != "" && out.XMLLeak {
out.Err = "xml tool call instead of openai tool_calls"
}
return out
}
out.ToolName = tc.Name
out.OK = tc.Name == p.Want
if !out.OK {
out.Err = "wanted " + p.Want + " got " + tc.Name
}
return out
}
type ProcMem struct {
Name string
SizeMB int
VRAMMB int
}
func ParsePS(raw []byte) []ProcMem {
var wrap struct {
Models []struct {
Name string `json:"name"`
Size int64 `json:"size"`
SizeVRAM int64 `json:"size_vram"`
} `json:"models"`
}
if err := json.Unmarshal(raw, &wrap); err != nil {
return nil
}
out := make([]ProcMem, 0, len(wrap.Models))
for _, m := range wrap.Models {
out = append(out, ProcMem{
Name: m.Name,
SizeMB: int(m.Size / (1024 * 1024)),
VRAMMB: int(m.SizeVRAM / (1024 * 1024)),
})
}
return out
}
func (c Client) FetchPS() []ProcMem {
url := Origin(c.BaseURL) + "/api/ps"
res, err := c.httpc().Get(url)
if err != nil {
return nil
}
defer res.Body.Close()
raw, _ := io.ReadAll(res.Body)
if res.StatusCode >= 300 {
return nil
}
return ParsePS(raw)
}
func Run(c Client) Report {
if c.Device == "" {
c.Device = "cpu"
}
rep := Report{
Model: c.Model,
HF: HFFor(c.Model),
Device: c.Device,
}
for _, p := range BakePrompts {
r := RunPrompt(c, p)
rep.Prompts = append(rep.Prompts, r)
rep.ToolCallN++
if r.OK {
rep.ToolCallOK++
}
if r.XMLLeak {
rep.XMLLeak++
}
}
if mems := c.FetchPS(); len(mems) > 0 {
rep.RSSMB = mems[0].SizeMB
rep.VRAMMB = mems[0].VRAMMB
if c.Device == "cpu" && mems[0].VRAMMB > 0 {
rep.Device = "gpu"
}
for i := range rep.Prompts {
rep.Prompts[i].RSSMB = rep.RSSMB
}
}
return rep
}
-166
View File
@@ -1,166 +0,0 @@
package reasoner
import (
"encoding/json"
"io"
"net/http"
"net/http/httptest"
"strings"
"testing"
"github.com/eSlider/2dph/internal/httpapi"
)
func TestHFIdsAreRealAndNoQwen36Nine(t *testing.T) {
if HFQwen35_9B != "Qwen/Qwen3.5-9B" {
t.Fatalf("9B id = %s", HFQwen35_9B)
}
if HFQwen36_27B != "Qwen/Qwen3.6-27B" {
t.Fatalf("27B id = %s", HFQwen36_27B)
}
if HFBonsai27B != "prism-ml/Bonsai-27B-gguf" {
t.Fatalf("bonsai id = %s", HFBonsai27B)
}
}
func TestParseToolResponseOpenAI(t *testing.T) {
raw := []byte(`{"choices":[{"message":{"tool_calls":[{"function":{"name":"search","arguments":"{\"q\":\"LadybugDB\"}"}}]}}]}`)
tc, _, err := ParseToolResponse(raw)
if err != nil {
t.Fatal(err)
}
if tc.Name != "search" {
t.Fatalf("name=%s", tc.Name)
}
}
func TestParseToolResponseXMLIsFailure(t *testing.T) {
raw := []byte(`{"choices":[{"message":{"content":"<tool_call>search</tool_call>"}}]}`)
_, content, err := ParseToolResponse(raw)
if err == nil {
t.Fatal("expected no tool_calls")
}
if !XMLLeak(content) {
t.Fatal("xml leak not detected")
}
}
func TestBakePromptsWantMCPTools(t *testing.T) {
if len(BakePrompts) != 3 {
t.Fatalf("prompts=%d", len(BakePrompts))
}
names := map[string]bool{}
for _, p := range BakePrompts {
names[p.Want] = true
}
for _, n := range []string{"search", "get", "audit"} {
if !names[n] {
t.Fatalf("missing want %s", n)
}
}
}
func TestParseToolResponseObjectArgs(t *testing.T) {
raw := []byte(`{"choices":[{"message":{"tool_calls":[{"function":{"name":"get","arguments":{"id":"leaf-demo"}}}]}}]}`)
tc, _, err := ParseToolResponse(raw)
if err != nil {
t.Fatal(err)
}
if tc.Name != "get" {
t.Fatalf("name=%s", tc.Name)
}
if !strings.Contains(tc.Arguments, "leaf-demo") {
t.Fatalf("args=%s", tc.Arguments)
}
}
func TestParsePSCPUNotVRAM(t *testing.T) {
raw := []byte(`{"models":[{"name":"qwen3.5:9b","size":6900000000,"size_vram":0}]}`)
got := ParsePS(raw)
if len(got) != 1 {
t.Fatalf("n=%d", len(got))
}
if got[0].VRAMMB != 0 {
t.Fatalf("vram=%d", got[0].VRAMMB)
}
if got[0].SizeMB < 6000 {
t.Fatalf("rss=%d", got[0].SizeMB)
}
}
func TestMCPToolsArePicoClawSubset(t *testing.T) {
mcp := map[string]bool{}
for _, op := range httpapi.Ops {
if op.MCP {
mcp[op.ID] = true
}
}
seen := map[string]bool{}
for _, tool := range MCPTools() {
fn, _ := tool["function"].(map[string]any)
name, _ := fn["name"].(string)
if !mcp[name] {
t.Fatalf("%s is not an MCP op", name)
}
seen[name] = true
}
for _, n := range []string{"search", "get", "audit"} {
if !seen[n] {
t.Fatalf("bake-off missing %s", n)
}
}
}
func TestChatToolsHitsOpenAIPath(t *testing.T) {
var gotTools bool
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/v1/chat/completions" {
t.Errorf("path=%s", r.URL.Path)
}
body, _ := io.ReadAll(r.Body)
var payload map[string]any
_ = json.Unmarshal(body, &payload)
if _, ok := payload["tools"]; ok {
gotTools = true
}
w.Header().Set("Content-Type", "application/json")
_, _ = w.Write([]byte(`{"choices":[{"message":{"tool_calls":[{"function":{"name":"search","arguments":"{\"q\":\"LadybugDB\"}"}}]}}]}`))
}))
defer srv.Close()
c := Client{BaseURL: srv.URL + "/v1", Model: "stub", HTTP: srv.Client(), Device: "cpu"}
r := RunPrompt(c, BakePrompts[0])
if !gotTools {
t.Fatal("tools not sent")
}
if !r.OK {
t.Fatalf("result=%+v", r)
}
}
func TestRunSamplesPSAfterPrompts(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
if r.URL.Path == "/api/ps" {
_, _ = w.Write([]byte(`{"models":[{"name":"stub","size":6500000000,"size_vram":0}]}`))
return
}
_, _ = w.Write([]byte(`{"choices":[{"message":{"tool_calls":[{"function":{"name":"search","arguments":"{}"}}]}}]}`))
}))
defer srv.Close()
c := Client{BaseURL: srv.URL + "/v1", Model: "stub", HTTP: srv.Client(), Device: "cpu"}
if mems := c.FetchPS(); len(mems) != 1 || mems[0].VRAMMB != 0 {
t.Fatalf("%+v", mems)
}
if mems := c.FetchPS(); mems[0].SizeMB < 6000 {
t.Fatalf("rss=%d", mems[0].SizeMB)
}
}
func TestHFForKnownOllamaTags(t *testing.T) {
if HFFor(OllamaRAM) != HFQwen35_9B {
t.Fatal(HFFor(OllamaRAM))
}
if HFFor(OllamaQuality) != HFBonsai27B {
t.Fatal(HFFor(OllamaQuality))
}
}
-122
View File
@@ -1,122 +0,0 @@
package websearch
import (
"context"
"fmt"
"net/http"
"os"
"time"
"golang.org/x/sys/unix"
)
const (
StatusSkipped = "skipped"
StatusRefused = "refused"
)
type LookupOpt struct {
Limit int
Timeout time.Duration
EnvPath string
CachePath string
Client *http.Client
Now func() float64
Sleep func(context.Context, time.Duration) error
}
func Lookup(ctx context.Context, query string, opt LookupOpt) Output {
if ctx == nil {
ctx = context.Background()
}
if opt.Limit <= 0 {
opt.Limit = DefaultLimit
}
if opt.Timeout <= 0 {
opt.Timeout = 25 * time.Second
}
nowFn := opt.Now
if nowFn == nil {
nowFn = func() float64 { return float64(time.Now().Unix()) }
}
sleepFn := opt.Sleep
if sleepFn == nil {
sleepFn = func(ctx context.Context, d time.Duration) error {
t := time.NewTimer(d)
defer t.Stop()
select {
case <-t.C:
return nil
case <-ctx.Done():
return ctx.Err()
}
}
}
if reason := PHIReason(query); reason != "" {
return Output{Query: query, Status: StatusRefused, Note: reason}
}
cachePath := opt.CachePath
if cachePath == "" {
cachePath = os.Getenv("BRAIN_SEARCH_CACHE")
}
if cachePath == "" {
cachePath = os.Getenv("HOME") + "/.cache/brain/web-search.sqlite"
}
cache, err := OpenCache(cachePath)
if err != nil {
return Output{Query: query, Status: StatusSkipped, Note: "cache: " + err.Error()}
}
defer cache.Close()
key := CacheKey(query, nil)
now := nowFn()
if cached, err := cache.Get(key, CacheTTL, now); err == nil && cached != nil {
out := Project(*cached, opt.Limit, DefaultSnippetChars)
out.Cached = true
return out
}
envPath := opt.EnvPath
if envPath == "" {
envPath = os.Getenv("BRAIN_SEARCH_ENV")
}
if envPath == "" {
envPath = os.Getenv("HOME") + "/.config/brain/search.env"
}
conf, err := LoadConfig(envPath)
if err != nil {
return Output{Query: query, Status: StatusSkipped, Note: "no BRAIN_SEARCH_URL; second source not consulted"}
}
lock, err := os.OpenFile(cachePath+".lock", os.O_CREATE|os.O_RDWR, 0o600)
if err != nil {
return Output{Query: query, Status: StatusSkipped, Note: "lock: " + err.Error()}
}
defer lock.Close()
if err := unix.Flock(int(lock.Fd()), unix.LOCK_EX); err != nil {
return Output{Query: query, Status: StatusSkipped, Note: "lock: " + err.Error()}
}
defer unix.Flock(int(lock.Fd()), unix.LOCK_UN)
last, err := cache.LastCall()
if err != nil {
return Output{Query: query, Status: StatusSkipped, Note: "cache: " + err.Error()}
}
if delay := WaitFor(last, nowFn(), MinInterval); delay > 0 {
if err := sleepFn(ctx, time.Duration(delay*float64(time.Second))); err != nil {
return Output{Query: query, Status: StatusSkipped, Note: "cancelled"}
}
}
_ = cache.MarkCall(nowFn())
payload, err := Fetch(opt.Client, conf, query, nil, opt.Timeout)
if err != nil {
return Output{Query: query, Status: StatusThrottled, Note: fmt.Sprintf("request failed: %v", err)}
}
if Classify(payload) == StatusOK {
_ = cache.Put(key, payload, nowFn())
}
return Project(payload, opt.Limit, DefaultSnippetChars)
}
-97
View File
@@ -1,97 +0,0 @@
package websearch
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"testing"
"time"
)
func TestLookupRefusesPIIWithoutFetch(t *testing.T) {
hits := 0
srv := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {
hits++
}))
defer srv.Close()
out := Lookup(context.Background(), "Personalnummer 12", LookupOpt{
EnvPath: writeEnv(t, srv.URL),
CachePath: filepath.Join(t.TempDir(), "c.sqlite"),
Client: srv.Client(),
Sleep: func(context.Context, time.Duration) error { return nil },
})
if out.Status != StatusRefused {
t.Fatalf("status = %s", out.Status)
}
if hits != 0 {
t.Fatal("PII query left the host")
}
}
func TestLookupSkipsWhenNoConfig(t *testing.T) {
out := Lookup(context.Background(), "LadybugDB", LookupOpt{
EnvPath: filepath.Join(t.TempDir(), "missing.env"),
CachePath: filepath.Join(t.TempDir(), "c.sqlite"),
Sleep: func(context.Context, time.Duration) error { return nil },
})
if out.Status != StatusSkipped {
t.Fatalf("status = %s", out.Status)
}
}
func TestLookupFetchesOnceAndCaches(t *testing.T) {
hits := 0
payload := Payload{Query: "x", Results: []RawHit{{Title: "t", URL: "http://example.com", Content: "c", Engine: "bing"}}}
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
hits++
json.NewEncoder(w).Encode(payload)
}))
defer srv.Close()
opt := LookupOpt{
EnvPath: writeEnv(t, srv.URL),
CachePath: filepath.Join(t.TempDir(), "c.sqlite"),
Client: srv.Client(),
Now: func() float64 { return 1_000 },
Sleep: func(context.Context, time.Duration) error { return nil },
}
a := Lookup(context.Background(), "LadybugDB", opt)
b := Lookup(context.Background(), "LadybugDB", opt)
if a.Status != StatusOK || b.Status != StatusOK {
t.Fatalf("a=%s b=%s", a.Status, b.Status)
}
if hits != 1 {
t.Fatalf("hits = %d, want 1 (second from cache)", hits)
}
if !b.Cached {
t.Fatal("second lookup not cached")
}
}
func TestLookupEmptyIsThrottled(t *testing.T) {
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Write([]byte(`{"query":"x","results":[]}`))
}))
defer srv.Close()
out := Lookup(context.Background(), "LadybugDB", LookupOpt{
EnvPath: writeEnv(t, srv.URL),
CachePath: filepath.Join(t.TempDir(), "c.sqlite"),
Client: srv.Client(),
Now: func() float64 { return 1_000 },
Sleep: func(context.Context, time.Duration) error { return nil },
})
if out.Status != StatusThrottled {
t.Fatalf("status = %s", out.Status)
}
}
func writeEnv(t *testing.T, url string) string {
t.Helper()
p := filepath.Join(t.TempDir(), "search.env")
if err := os.WriteFile(p, []byte("BRAIN_SEARCH_URL="+url+"\n"), 0o600); err != nil {
t.Fatal(err)
}
return p
}
+1
View File
@@ -6,6 +6,7 @@ readme = "README.md"
requires-python = ">=3.12"
license = { text = "MIT" }
dependencies = [
"docling>=2.119.0",
"ladybug==0.19.1",
"markitdown[docx,epub,html,image-exif,pdf,pptx,xlsx,zip]>=0.1.7",
"mistune==3.3.4",
-257
View File
@@ -1,257 +0,0 @@
#!/usr/bin/env python3
"""System performance test: PicoClaw surface (brain MCP) + optional reasoner.
BRAIN_URL=http://127.0.0.1:8630 ./qa/system_perf.py --json
REASONER_BASE_URL=http://127.0.0.1:11435/v1 REASONER_MODEL=qwen3.5:9b \\
./qa/system_perf.py --reasoner --picoclaw --json
Does not write Ladybug. Search includes web (D17); expect ~10s+ per search.
Exit 1 if health/get/audit gates fail. Reasoner is measured, not gated.
"""
from __future__ import annotations
import argparse
import json
import os
import statistics
import sys
import time
import urllib.error
import urllib.request
from concurrent.futures import ThreadPoolExecutor
DEFAULT_BRAIN = "http://127.0.0.1:8630"
DEFAULT_REASONER = "http://127.0.0.1:11435/v1"
DEFAULT_MODEL = "qwen3.5:9b"
DEFAULT_PICOCLAW = "http://127.0.0.1:18790"
GATE_HEALTH_MS = 500
GATE_GET_P50_MS = 50
GATE_AUDIT_P50_MS = 50
def _req(url: str, data: bytes | None = None, timeout: float = 90) -> bytes:
headers = {"Content-Type": "application/json"} if data is not None else {}
req = urllib.request.Request(url, data=data, headers=headers)
with urllib.request.urlopen(req, timeout=timeout) as res:
return res.read()
def timed(fn):
t0 = time.perf_counter()
out = fn()
return (time.perf_counter() - t0) * 1000.0, out
def stats(samples: list[float]) -> dict:
s = sorted(samples)
n = len(s)
return {
"n": n,
"min_ms": round(s[0], 1),
"p50_ms": round(s[n // 2], 1),
"p95_ms": round(s[min(n - 1, int(n * 0.95))], 1),
"max_ms": round(s[-1], 1),
"avg_ms": round(statistics.mean(s), 1),
}
def mcp(brain: str, method: str, params=None, timeout: float = 90) -> dict:
payload: dict = {"jsonrpc": "2.0", "id": 1, "method": method}
if params is not None:
payload["params"] = params
raw = _req(brain.rstrip("/") + "/mcp", json.dumps(payload).encode(), timeout=timeout)
return json.loads(raw.decode())
def mcp_call(brain: str, name: str, arguments: dict, timeout: float = 90) -> tuple[bool, str]:
d = mcp(brain, "tools/call", {"name": name, "arguments": arguments}, timeout=timeout)
res = d.get("result") or {}
text = ((res.get("content") or [{}])[0].get("text") or "")
return (not res.get("isError")), text
def reasoner_tool_call(base: str, model: str, user: str) -> str:
payload = {
"model": model,
"messages": [
{"role": "system", "content": "You are PicoClaw. Always call search before answering."},
{"role": "user", "content": user},
],
"tools": [
{
"type": "function",
"function": {
"name": "search",
"description": "deduction search",
"parameters": {
"type": "object",
"properties": {"q": {"type": "string"}},
"required": ["q"],
},
},
}
],
"tool_choice": "required",
}
raw = _req(
base.rstrip("/") + "/chat/completions",
json.dumps(payload).encode(),
timeout=600,
)
chat = json.loads(raw.decode())
tcs = chat["choices"][0]["message"].get("tool_calls") or []
if not tcs:
return ""
return tcs[0]["function"]["name"]
def run(args: argparse.Namespace) -> dict:
brain = args.brain.rstrip("/")
report: dict = {
"brain": brain,
"device": "cpu",
"ok": True,
"gates": {},
"mcp": {},
}
ms, _ = timed(lambda: _req(brain + "/health", timeout=5))
report["mcp"]["health"] = {"n": 1, "avg_ms": round(ms, 1)}
report["gates"]["health"] = ms <= GATE_HEALTH_MS
if ms > GATE_HEALTH_MS:
report["ok"] = False
list_ms = []
for _ in range(args.n):
ms, d = timed(lambda: mcp(brain, "tools/list", timeout=10))
names = [t["name"] for t in ((d.get("result") or {}).get("tools") or [])]
if "search" not in names:
report["ok"] = False
list_ms.append(ms)
report["mcp"]["tools_list"] = stats(list_ms)
audit_ms = []
for _ in range(args.n):
ms, (ok, _) = timed(lambda: mcp_call(brain, "audit", {}))
if not ok:
report["ok"] = False
audit_ms.append(ms)
report["mcp"]["audit"] = stats(audit_ms)
report["gates"]["audit_p50"] = report["mcp"]["audit"]["p50_ms"] <= GATE_AUDIT_P50_MS
if not report["gates"]["audit_p50"]:
report["ok"] = False
ok, text = mcp_call(brain, "search", {"q": "LadybugDB", "n": 2}, timeout=90)
inner = json.loads(text) if ok else {}
hits = inner.get("results") or []
leaf_id = hits[0]["id"] if hits else ""
report["mcp"]["search_seed"] = {
"ok": ok,
"count": inner.get("count"),
"web": (inner.get("web") or {}).get("status"),
}
get_ms = []
if leaf_id:
for _ in range(args.n):
ms, (ok, _) = timed(lambda: mcp_call(brain, "get", {"id": leaf_id, "body": True}))
if not ok:
report["ok"] = False
get_ms.append(ms)
report["mcp"]["get"] = stats(get_ms)
report["gates"]["get_p50"] = report["mcp"]["get"]["p50_ms"] <= GATE_GET_P50_MS
if not report["gates"]["get_p50"]:
report["ok"] = False
def one_get() -> float:
t0 = time.perf_counter()
mcp_call(brain, "get", {"id": leaf_id, "body": True})
return (time.perf_counter() - t0) * 1000.0
t0 = time.perf_counter()
with ThreadPoolExecutor(max_workers=8) as ex:
conc = list(ex.map(lambda _: one_get(), range(8)))
wall = (time.perf_counter() - t0) * 1000.0
report["mcp"]["get_concurrent_8"] = {**stats(conc), "wall_ms": round(wall, 1)}
search_ms = []
for q in ("LadybugDB", "model2vec"):
ms, (ok, text) = timed(lambda q=q: mcp_call(brain, "search", {"q": q, "n": 3}, timeout=90))
inner = json.loads(text) if ok else {}
search_ms.append(ms)
report.setdefault("mcp", {}).setdefault("search_samples", []).append(
{
"q": q,
"ms": round(ms, 1),
"ok": ok,
"count": inner.get("count"),
"web": (inner.get("web") or {}).get("status"),
}
)
if search_ms:
report["mcp"]["search"] = stats(search_ms)
if args.reasoner:
base = args.reasoner_url
model = args.model
report["reasoner"] = {"base_url": base, "model": model, "calls": []}
for user in (
"Use tools. Search the 2dph brain for LadybugDB. Call search.",
"Use tools. Search the 2dph brain for model2vec. Call search.",
):
ms, name = timed(lambda user=user: reasoner_tool_call(base, model, user))
report["reasoner"]["calls"].append({"ms": round(ms, 1), "tool": name})
tools = [c["tool"] for c in report["reasoner"]["calls"]]
report["gates"]["reasoner_tool_call"] = bool(tools) and all(t == "search" for t in tools)
if not report["gates"]["reasoner_tool_call"]:
report["ok"] = False
if args.picoclaw:
gw = args.picoclaw_url.rstrip("/")
ms, raw = timed(lambda: _req(gw + "/health", timeout=5))
body = json.loads(raw.decode())
report["picoclaw"] = {
"url": gw,
"health_ms": round(ms, 1),
"status": body.get("status"),
}
report["gates"]["picoclaw_health"] = body.get("status") == "ok" and ms <= GATE_HEALTH_MS
if not report["gates"]["picoclaw_health"]:
report["ok"] = False
return report
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description="2dph system performance (MCP + optional reasoner)")
p.add_argument("--brain", default=os.environ.get("BRAIN_URL", DEFAULT_BRAIN))
p.add_argument("--n", type=int, default=20)
p.add_argument("--json", action="store_true")
p.add_argument("--reasoner", action="store_true")
p.add_argument("--picoclaw", action="store_true")
p.add_argument("--picoclaw-url", default=os.environ.get("PICOCLAW_URL", DEFAULT_PICOCLAW))
p.add_argument("--reasoner-url", default=os.environ.get("REASONER_BASE_URL", DEFAULT_REASONER))
p.add_argument("--model", default=os.environ.get("REASONER_MODEL", DEFAULT_MODEL))
args = p.parse_args(argv)
try:
report = run(args)
except (urllib.error.URLError, TimeoutError, OSError) as e:
print(f"system_perf: {e}", file=sys.stderr)
return 1
if args.json:
print(json.dumps(report, indent=2))
else:
print(f"ok={report['ok']} brain={report['brain']}")
for name, block in report.get("mcp", {}).items():
if isinstance(block, dict) and "p50_ms" in block:
print(f" {name}: p50={block['p50_ms']} p95={block['p95_ms']} n={block['n']}")
elif name == "health":
print(f" health: {block.get('avg_ms')} ms")
for k, v in report.get("gates", {}).items():
print(f" gate {k}: {v}")
for c in (report.get("reasoner") or {}).get("calls") or []:
print(f" reasoner {c['tool']}: {c['ms']} ms")
return 0 if report["ok"] else 1
if __name__ == "__main__":
raise SystemExit(main(sys.argv[1:]))
+6 -11
View File
@@ -26,26 +26,21 @@ second independent source when local roots cannot confirm. An answer is
bin/brain/search.go "Matrix federation" # pointers + snippets, YAML
bin/brain/search.go "onlyoffice postgres" --root facts # restrict to confirmed
bin/brain/search.go "where is cs-lexicon" --json | yq '.[].ref'
bin/brain/add.go --text T --root facts --source "a.md x b.md"
bin/brain/get.go <id> --body # full chunk only when needed
bin/brain/stats.go # index health
bin/brain/eval.go # recall@5 >= 0.95 gate (Go; Python bin/kb/eval is CI fallback)
bin/brain/eval.go # recall@5 >= 0.95 gate
```
`bin/kb/search` is a deprecated wrapper. `--hop N` walks
`FROM_FILE` / `HAS_VERSION` / `AUTHORED` from each hit (1=File, 3=Person).
`bin/kb/search` is a deprecated wrapper. `--hop` errors (File/FROM_FILE edges
are not wired yet); do not treat it as a graph walk.
## Rules
- Search before you read. Never grep a repo for a concept the graph covers.
- `--root facts` returns only confirmed evidence-linked answers. Default shows
facts first, then info leafs clearly marked `(not confirmed)`.
- If there is no facts hit, `bin/brain/search.go` consults SearXNG and adds a
`web` block (kept apart from graph hits). `throttled` / `skipped` / `refused`
are not evidence of absence. `--root facts|info` and `--no-web` skip the web.
- If recall looks wrong, run `bin/brain/eval.go`; it gates control questions and
should stay at or above 95% recall@5.
- Agents: `GET /openapi.json` and `POST /mcp` on `bin/brain/serve.go` (same
handlers; tool names match paths `search`/`get`/`stats`/`audit`). Generated
list: [tools.md](tools.md).
- Never report an unconfirmed single-source local answer as fact.
- Escalate to `web-search` (the `web-search` skill) as the independent second
source when both local roots cannot confirm; never report an unconfirmed
single-source local answer as fact.
-11
View File
@@ -1,11 +0,0 @@
# brain HTTP / MCP tools
Generated from `internal/httpapi.Ops`. Do not edit by hand.
Serve: `bin/brain/serve.go` (`GET /openapi.json`, `POST /mcp`).
- `search` — deduction search (facts → info → web)
- `get` — read one leaf by id
- `stats` — index health
- `audit` — facts confidence histogram
- `ingest` — add a leaf without rebuild
+39
View File
@@ -0,0 +1,39 @@
---
name: db-yaml
description: >-
Read any Postgres as compact YAML through db/psql-yq, with a read-only guard and
named profiles. Use when a task needs table contents, column types or a SELECT
against cs_brain or another project database.
---
# db-yaml
`bin/db/psql-yq` (vendored in this repo) talks to Postgres and returns YAML,
which is far cheaper than a psql ASCII table and easy to slice with `yq`.
```bash
bin/db/psql-yq --profile onlyoffice -s document_asset # column list
bin/db/psql-yq --profile onlyoffice -t task_result -l 20 # sample rows as YAML
bin/db/psql-yq --profile onlyoffice -c 'SELECT ...' # query -> YAML
```
Ad-hoc targets without a profile:
```bash
bin/db/psql-yq --container my-pg --db app -c 'SELECT 1'
bin/db/psql-yq --dsn 'postgres://user@host:5432/db' -c 'SELECT 1'
```
## Profiles
Connection details live in `~/.config/brain/db-profiles.yml` (mode 600), never in a
project repo. A profile names either a `container` or a `host`; passwords are read
from a separate `password_env_file` and never appear in argv.
## Rules
- **Read-only.** Any `insert|update|delete|drop|truncate|alter|create|grant|
revoke|vacuum|copy` is rejected with exit 3. Do not work around it.
- **PII.** `cs_brain` holds client data. Aggregate and count freely; never copy
names or addresses into chat, issues or docs.
- Use `-l` to keep samples small. Twenty rows answer most questions.
-38
View File
@@ -1,38 +0,0 @@
---
name: duckdb
description: >-
Use https://github.com/duckdb/duckdb-go in-process for columnar analytics
(quantiles, GROUP BY, JSON/CSV/Parquet/JSONL scans) when that is faster than
nested Go loops. Not Ladybug. Not the web-search sqlite cache. Use when
aggregating samples, counting JSONL, or SQL over tabular files.
---
# duckdb-go
Use https://github.com/duckdb/duckdb-go where it makes sense to get better performance in code.
In-process DuckDB (`internal/duckstats`, `database/sql` driver `duckdb`).
Vectorized SQL over tables, JSONL, CSV, Parquet. CGO with bundled libs
(linux/darwin amd64/arm64). Links with **gcc/g++** (libstdc++), not Zig.
D21 Zig (`bin/cgo/zcc`) is Ladybug/tokenizers only. After
`eval "$(bin/cgo/zig env)"`:
```bash
CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= ./bin/qa/stats.go <<< '[1,2,3,4,5]'
CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= go test ./internal/duckstats
```
| Store | Job |
|-------|-----|
| Ladybug | graph + FTS + HNSW (facts/info) |
| modernc sqlite | web-search KV cache + throttle |
| duckdb-go | OLAP: quantiles, counts, scans of many rows/files |
| mikefarah/yq | small YAML/JSON/XML/CSV/TOML/HCL slice, not bulk |
```bash
./bin/qa/stats.go <<< '[1,2,3,4,5]'
./bin/qa/stats.go --jsonl path/to/rows.jsonl
```
Do not open Ladybug through DuckDB. Do not put secrets or client PII into
DuckDB files under the repo.
-26
View File
@@ -1,26 +0,0 @@
---
name: picoclaw
description: >-
2dph is the memory/fact gate. Compose runs the official PicoClaw gateway.
Use when wiring PicoClaw or any MCP client: call brain search/get/audit
before a factual reply. throttled is not a negative finding.
---
# PicoClaw — fact-check before assert
PicoClaw speaks MCP at `POST /mcp` on `bin/brain/serve.go`. Compose profile
`picoclaw` runs the official `sipeed/picoclaw` gateway plus `brain-mcp`
(see [docs/picoclaw.md](../../docs/picoclaw.md)).
## Tool order (before a factual reply)
1. **`search`** — facts root first, then info. The `web` block is a second
source when there is no facts hit. Status `throttled` / `skipped` /
`refused` is **not** evidence of absence.
2. **`get`** — full leaf body only when a hit `id` is needed.
3. **`audit`** — if recall or confidence looks wrong.
Then answer. Confirmed only from facts (≥2 independent sources). Anything
else is `(not confirmed)`. Missing graph ≠ “does not exist”.
Generated tool list: [../brain/tools.md](../brain/tools.md).
-39
View File
@@ -1,39 +0,0 @@
---
name: postgres
description: >-
Read Postgres as compact YAML through bin/postgres/query.go (read-only
guard, named profiles). Use when a task needs table contents, column types,
or a SELECT against an ops database.
---
# postgres
`bin/postgres/query.go` wraps vendored `bin/db/psql-yq`. Output is YAML
(cheaper than psql ASCII, easy to slice with mikefarah/yq).
```bash
bin/postgres/query.go --profile onlyoffice -s document_asset # column list
bin/postgres/query.go --profile onlyoffice -t task_result -l 20 # sample rows
bin/postgres/query.go --profile onlyoffice -c 'SELECT ...' # query → YAML
```
Ad-hoc targets without a profile:
```bash
bin/postgres/query.go --container my-pg --db app -c 'SELECT 1'
bin/postgres/query.go --dsn 'postgres://user@host:5432/db' -c 'SELECT 1'
```
## Profiles
Connection details live in `$HOME/.config/brain/db-profiles.yml` (mode 600),
never in a project repo. A profile names either a `container` or a `host`;
passwords are read from a separate `password_env_file` and never appear in argv.
## Rules
- **Read-only.** Any `insert|update|delete|drop|truncate|alter|create|grant|
revoke|vacuum|copy` is rejected with exit 3. Do not work around it.
- **PII.** Client CRM databases: aggregate and count freely; never copy names
or addresses into chat, issues or docs.
- Use `-l` to keep samples small. Twenty rows answer most questions.
+1 -1
View File
@@ -11,7 +11,7 @@ description: >-
```bash
bin/web/search.go "LadybugDB vector index"
bin/web/search.go "model2vec multilingual" --category it
bin/web/search.go "hypervisor" --site example.com --json | yq -r '.results[].url'
bin/web/search.go "hypervisor" --site example.com --json | jq -r '.results[].url'
bin/web/search.go "postgres partial index" --lang en --fresh year
```
@@ -43,7 +43,7 @@ found" unless the client refuses to call it absence.
```bash
for i in $(seq 10); do
bin/web/search.go "test $i" -n 1 --refresh --json | yq -r '.status'
bin/web/search.go "test $i" -n 1 --refresh --json | jq -r .status
done
```
-33
View File
@@ -1,33 +0,0 @@
---
name: yq
description: >-
Use https://github.com/mikefarah/yq to work with YAML, JSON, XML, CSV,
TOML, HCL where it's efficient and less code. Use when slicing compose,
config, --json tool output, CSV/TOML/HCL/XML, or converting between those
formats. Not kislyuk Python yq. Not jq when yq already does the job.
---
# yq (mikefarah)
Use https://github.com/mikefarah/yq to work with YAML, JSON, XML, CSV, TOML, HCL where it's efficient and less code.
This is the Go `yq` (`yq --version` contains `mikefarah`). It is not
kislyuk/yq (Python, jq-syntax, YAML-only wrapper). `bin/db/psql-yq` already
calls this binary.
Prefer `yq` over `python3 -c`, `jq`, or ad-hoc parsers when one expression
reads or converts the file. Keep Python/Go for HTTP, binary protocols, and
in-process tests.
```bash
yq '.services.picoclaw.image' compose.yaml
yq -P . deploy/picoclaw/config.json # JSON → YAML
yq -o=json '.gates' # JSON stdin (qa/system_perf.py --json)
yq -p=csv -o=json .
yq -p=xml -o=json .
yq -p=toml '.package.name' file.toml
bin/brain/search.go "LadybugDB" --json | yq '.[].ref'
bin/web/search.go "hypervisor" --json | yq -r '.results[].url'
```
Do not print secrets, PII, or `$HOME/.config/brain/` through `yq`.
Generated
+1648
View File
File diff suppressed because it is too large Load Diff