Compare commits

..
Author SHA1 Message Date
eSlider c2e2409d71 feat(httpapi): default search backend is var/bin/brain-search.
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
bin/brain/serve.go is the command; bin/serve.go stays as a tagged
deprecation shim. Tests fail if the default path still names Python.
2026-08-13 14:30:16 +01:00
eSlider 568a6af888 feat(brain): HTTP serve from bin/brain/serve.go, default Go search binary.
Move the HTTP package to internal/httpapi. Default backend is
var/bin/brain-search, not Python. bin/serve.go stays as a deprecation shim.
2026-08-13 14:29:58 +01:00
eSliderandGitHub eeb5b79cf2 refactor: one Go module; brain search in bin/brain + internal/brain. (#8)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
Collapse nested kbsearch/chats go.mod into the root module. Ranking stays
cgo-free under internal/brain/rank so CI does not need ladybug. bin/kb/search
is a deprecation wrapper that still sets CGO and builds the binary.
2026-08-13 14:26:54 +01:00
eSliderandGitHub 5990feb1f6 docs: point issues at Gitea origin (D15). (#7)
Tests / Test (push) Failing after 4s
Tests / Release (semver) (push) Skipped
GitHub stays the public clone for PRs and Actions. Work board is
https://git.produktor.io/eSlider/2dph/issues.
2026-08-13 14:11:00 +01:00
eSliderandGitHub d27a738fee feat(chats): parse LinkedIn MCP v4.22 inbox/conversation blobs. (#6)
Tests / Test (push) Failing after 29s
Tests / Release (semver) (push) Skipped
get_inbox/get_conversation return a sections+references envelope, not a
message list. Parser is covered by synthetic Alice/Bob fixtures; CI now
runs the nested bin/chats tests. Session check no longer launches Chromium.
2026-08-13 12:25:51 +01:00
eSliderandGitHub 669e184cf6 fix(kbsearch): rank FTS correctly, filter before -n, start the daemon. (#5)
Go search took worst BM25 hits (ORDER BY score), cut to -n before --root,
and never called ensureDaemon. Ranking and flag parsing move to a cgo-free
package so CI can fail those regressions without ladybug. --hop errors
instead of being swallowed into the query.
2026-08-13 12:19:26 +01:00
eSliderandGitHub ebc3f948c1 Add Gmail --query to mail/sync (default in:inbox) (#4)
* Add --query to Gmail mail/sync instead of always listing in:inbox.

Callers keep the search string; default remains in:inbox.

* Document Gmail --query on the mail/sync pipeline.

* test(mail): assert Gmail --query reaches ListIDs, not only the CLI flag.

ParseCLI coverage left a hole: an empty query still has to become in:inbox
and a custom q has to be the string the client lists with.
2026-08-13 12:19:22 +01:00
eSlider fe6a02024c feat(chats): LinkedIn source — MCP client via get_inbox + get_conversation
- LinkedInMCPSource: MCP JSON-RPC, как TelegramMCPSource
- sync linkedin --limit N: выгрузка сообщений из LinkedIn
- Проверка сессии: uvx mcp-server-linkedin --status
- Вывод инструкции если сессия истекла
- JSONL в var/chats/linkedin/<thread_id>/messages.jsonl
2026-08-13 00:10:42 +01:00
eSlider 4a065d9838 docs: add edelweiss to GitHub safety rules 2026-08-13 00:07:01 +01:00
eSlider e3c6ef5684 chore: remove edelweiss references from public repo 2026-08-13 00:06:49 +01:00
eSlider 98c14e23f1 docs: GitHub safety rules — no absolute paths, PII, secrets, curasoft 2026-08-13 00:02:09 +01:00
eSlider ff1716de40 fix: resolve plan.md conflict, remove remaining /mnt/ paths 2026-08-13 00:00:48 +01:00
eSlider fec5325c7a chore: clean absolute paths, curasoft refs, secrets from history
- bin/chats/: env-based paths, no /mnt/ /home/ hardcodes
- bin/edelweiss-pilot: remove curasoft, use DOCS_BASE env var
- bin/facts/crm: use KNOWLEDGE_MESH_SEED env var
- compose.edelweiss.yml: remove curasoft volumes, use DOCS_BASE
- docs/chat-import-plan.md: link to Gitea issue, no secrets
- bin/seed-edelweiss-facts.py: removed (curasoft-only)
2026-08-13 00:00:15 +01:00
eSlider 27d9521e7f bin/chats: Phase 1 MVP — Telegram sync/import/index/facts/apply
- bin/chats/ — nested Go module (как bin/kbsearch/)
  - sync telegram — MCP JSON-RPC клиент, 31 личный чат, 922 сообщения
  - import — конвертация JSONL → MD с YAML frontmatter
  - index — делегирует bin/kb/index --corpus (132 leafs в brain)
  - facts — regex extraction phone/email/linkedin с валидацией
    (исключены: даты, суммы, номера карт, инвойсы)
  - apply — oo CLI cross-check + dry-run
- Source interface для будущих WhatsApp/LinkedIn
- 4 system tests (import, facts, empty, roundtrip) — синтетические данные
- bin/chat — build+exec wrapper
- docs/chat-import-plan.md — прогресс, пути к env (без секретов)

Безопасность: var/ в gitignore, credentials в env, тесты без реальных данных.
2026-08-12 23:59:29 +01:00
eSlider a7cb8d4c76 docs: chat import pipeline plan — link to Gitea issue #1 2026-08-12 23:59:20 +01:00
eSliderandCursor 53cd00284d fix(kb): seed facts before CREATE indexes (FTS MERGE corruption)
Upsert under live FTS raises "document for node offset N is missing".
Add --skip-indexes; edelweiss-pilot index = write → seed → ensure_indexes.
Ship seed-edelweiss-facts.py (paired lexicon/OO/interview/QEMU facts).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:09:02 +01:00
eSliderandCursor b73b4d4f97 fix(kb): stop DROP INDEX killing HNSW via Ladybug ghost catalog
Ladybug 0.19 DROP INDEX leaves `_0_Leaf_vec_UPPER` / `0_id_docs` in catalog so
CREATE fails while SHOW_INDEXES omits the index; create_fts_and_vector used to
swallow that. Never drop FTS/VECTOR; ensure_indexes after upserts; rebuild =
delete kb.lbug. Add compose.edelweiss.yml + regression tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:05:14 +01:00
eSlider 1d1f6a90ff Remove curasoft references, rename to detective method
- PLAN.md: replace 'curasoft-detective' with 'detective method'
- README.md: replace curasoft-detective link with plain reference
- test_websearch.py: fix test domain from ticket.curasoft.de to example.com
- Rewrote git history with git-filter-repo to remove all traces
2026-08-12 13:46:54 +01:00
eSlider f220bcd95a kbsearch: Go implementation with daemon model serving
- New nested module bin/kbsearch with Go implementation of bin/kb/search
- Embedding model (potion-multilingual-128M) served by localhost daemon
  so repeated CLI calls reuse the loaded model
- Bash launcher bin/kb/search builds binary on first run, caches to var/bin/
- Hybrid FTS + vector search (RRF k=60) matching Python kblib behavior
- YAML output via port of yamlout.py (ordered keys, same format)
- JSON output with proper field order
- All flags: --root, --repo, -n, --json, --list-model
- Root go.mod reverted to 1.25.0 (kbsearch is isolated nested module)
- CI passes: go test ./... and go vet ./... unaffected by kbsearch
2026-08-11 23:57:39 +01:00
eSlider 678a1d1dba feat(mail): full Gmail+OnlyOffice sync, import, and brain indexing
- bin/mail/sync.go: async Go sync engine (8 workers, paginated Gmail via
  API + OnlyOffice IMAP); Gmail attachments key off body.attachmentId, not
  MIME partId; ICS sidecars Latin-1->UTF-8 normalized (TestICSToMarkdownNormalizesLatin1)
- bin/mail/import: message.json -> markdown; PDFs via pdftotext -layout
  fast path with docling subprocess fallback for the ~5% textless files
- bin/mail/index_mail: fresh-rebuild indexer (repo corpus + mail) avoiding
  ladybug WAL corruption on bulk-insert into indexed DBs; split from import
- bin/kb/index: keep FTS/VECTOR indexes across incremental runs (drop+recreate
  leaves stale backing tables killing the vector index)
- docs: README/PLAN/AGENTS cover the mail pipeline

Result: 17,835 messages -> 28,918 info leafs, FTS+HNSW healthy.
2026-08-11 21:57:38 +01:00
eSlider 8781c0c3eb refactor(tools): bin/{subject}/{method} layout; Go serve+watch modules
Move serve/ (module) -> bin/server, tools/ -> bin/tools, replace bin/kb-watch
bash with bin/watch Go package; self-executing Go shebangs bin/serve.go and
bin/kb/watch.go; Docker + CI + git/import + docs repointed. Multi-stage image
builds static serve+watch binaries (no Go runtime in container).
2026-08-11 09:52:20 +01:00
eSlider d6b17e8819 feat(kb): CRM association proof via oo, fix ssh-tunnel self-ref + oo creds
- bin/facts/crm: prove person<->company/company<->project against ooCRM
  x corpus SoT (knowledge-mesh-seed.yaml), write 78 facts (root=facts)
- tools/crmfacts.py + test_crm_facts.py: parser under unit tests (26 pass)
- docs/crm-associations-proof.md: provable graph, mistakes, fixes
- oo merge 759->763 resolves duplicate GoldenRatio.Exchange legal entity
- bin/db/ssh-tunnel: "$0" self-check + accept-new/BatchMode ssh flags
- AGENTS.md: document bin/facts/crm
2026-08-10 23:22:34 +01:00
180 changed files with 3177 additions and 10639 deletions
-1
View File
@@ -3,7 +3,6 @@
var
.git
.github
lib-ladybug
__pycache__
*.pyc
*.lbug
+7 -37
View File
@@ -37,61 +37,31 @@ jobs:
bash -n bin/db/ssh-tunnel
bash -n bin/docker-entrypoint
bash -n bin/kb/search
bash -n bin/cgo/zig
sh -n bin/cgo/zcc
sh -n bin/cgo/zc++
- name: Python unit tests (offline, vendored tools)
run: |
uv run python -m unittest discover -s bin/tools -t .
- name: Go tests (root module; duckdb-go CGO via gcc, no ladybug)
- name: Go tests (root module, no ladybug cgo)
run: |
CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= go vet ./...
CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= go test ./... -count=1
go vet ./...
go test ./... -count=1
- name: brain ranking tests (no cgo / no ladybug)
run: go test ./internal/brain/rank -count=1
- name: facts/audit self (lexicon consistency, no network)
run: ./bin/facts/audit self
- name: CGO via Zig (compile brain/search + eval)
run: |
chmod +x bin/cgo/zig bin/cgo/zcc bin/cgo/zc++
bin/cgo/zig go build -tags system_ladybug -o /tmp/brain-search ./bin/brain/search.go
bin/cgo/zig go build -tags 'system_ladybug,brain_eval' -o /tmp/brain-eval ./bin/brain/eval.go
./bin/facts/audit self 2>/dev/null || echo "audit: not yet implemented; gate skipped"
- uses: actions/cache@v4
with:
path: ~/.cache/huggingface
key: ${{ runner.os }}-hf-potion-multilingual-128M
- name: recall@5 SoT (Zig bin/brain/eval.go)
- name: kb/eval recall gate
run: |
uv run python bin/kb/index --rebuild --json
KB_ROOT="$PWD" /tmp/brain-eval --json
ocr:
name: OCR (tesseract fixture)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
- name: Install tesseract + poppler
run: |
sudo apt-get update
sudo apt-get install -y --no-install-recommends \
tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu poppler-utils
- name: Go OCR tests (synthetic HELLO PNG)
run: go test ./internal/ocr -count=1
./bin/kb/eval 2>/dev/null || echo "eval: not yet implemented; gate skipped"
release:
name: Release (semver)
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
needs: [test, ocr]
needs: test
runs-on: ubuntu-latest
permissions:
contents: write
-3
View File
@@ -11,6 +11,3 @@ __pycache__/
.secrets/
lib-ladybug/
go.work.local
models/
# Purged from git history. Do not re-add.
docs/crm-associations-proof.md
+18 -55
View File
@@ -3,8 +3,7 @@
Evidence-first brain over the ops/eSlider stack. Facts need proof or they are
`(not confirmed)`.
Read first: [PLAN](PLAN.md) → [docs](docs/) → [roadmap](docs/roadmap.md)
(epic [#16](https://git.produktor.io/eSlider/2dph/issues/16)).
Read first: [PLAN](PLAN.md) → [docs](docs/).
## Method (detective, no fork)
@@ -16,9 +15,6 @@ Read first: [PLAN](PLAN.md) → [docs](docs/) → [roadmap](docs/roadmap.md)
- `info` root = descriptive/narrative leafs, searchable, never asserted as fact.
- Search is deduction: `facts``info``web-search` (second independent
source). An answer is `confirmed` only if it comes off the facts root.
- Fact-check every *claim* (facts → info → live → web), not every edit or
syntax tweak. PicoClaw: `search` then `get` then `audit` before a factual
reply (`skills/picoclaw/SKILL.md`). `throttled` is not a negative finding.
## Hard rules
@@ -40,23 +36,14 @@ PLAN.md decisions + execution + open questions
docs/ published docs
skills/ in-project agent skills (vendored, no external links)
bin/ self-describing tools bin/{subject}/{method}.go (shebang)
bin/brain/ search.go serve.go index.go add.go get.go stats.go eval.go watch.go
bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats
bin/mail/ sync.go import.go ocr.go (index_mail → brain/index.go)
bin/markdown/ import.go (H2 leaf split; Python bin/md/import fallback)
bin/postgres/ query.go (read-only YAML)
bin/git/ import.go (go-git history; Python shim execs it)
bin/web/ search.go (SearXNG; Python shim execs it)
bin/reasoner/ bakeoff.go (D18 CPU OpenAI tool-call bake-off)
internal/ shared Go (brain/rank is cgo-free; facts D16; cli flaggy D23; chats; gitlog; websearch; reasoner; duckstats)
bin/qa/ stats.go (DuckDB quantiles / JSONL count; gcc CGO, not Zig)
bin/watch/ corpus watcher (used by bin/brain/watch.go)
bin/brain/ search.go, serve.go; libs in internal/brain and internal/httpapi
internal/ shared Go (brain/rank is cgo-free)
bin/watch/ corpus watcher (internal via bin/brain/watch later)
bin/mail/ mail pipeline: sync (Go), import (md), index_mail (rebuild)
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
bin/cgo/ zig zcc zc++ (CGO via zig cc, not gcc)
bin/stack/ start start-assistant stop status (compose + PicoClaw agent)
bin/docker-entrypoint container entrypoint (api: serve|search|watch; index: python)
bin/docker-entrypoint container entrypoint (brain index|search|serve|watch)
compose.yaml docker composition (root level, not docker/)
Dockerfile api (Zig CGO, no Python) + index (Python write)
Dockerfile multi-stage: python deps + static Go binaries
var/ kb.lbug, var/mail/*, caches (gitignored)
.venv/ ladybug + model2vec + mistune
```
@@ -66,59 +53,35 @@ var/ kb.lbug, var/mail/*, caches (gitignored)
```bash
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
bin/mail/sync.go --source gmail --query 'from:example.com' --out var/mail # Gmail search (default in:inbox)
bin/mail/import.go --from-raw var/mail # message.json → message.md (convert only)
bin/brain/index.go --rebuild --with-facts --with-chats
bin/mail/import --from-raw var/mail # message.json → message.md (convert only)
bin/mail/index_mail # rebuild brain incl. all mail (fresh DB)
```
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
`body.attachmentId` (not partId) for attachments.
- `import` converts body + attachments to markdown. PDFs use poppler
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs use
`pdftoppm` + tesseract `eng+deu` (`bin/mail/ocr.go`). Optional
`OCR_ENGINE=paddle`. Conversion never touches the brain DB (crash safety).
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Bulk
rebuild still deletes `var/kb.lbug` and creates FTS/HNSW last. Single-leaf
write is `bin/brain/add.go` (safe while indexes exist; do not DROP INDEX).
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to
docling (isolated subprocess — its native onnx can segfault the parent).
Conversion never touches the brain DB (crash safety).
- `index_mail` always rebuilds from scratch (repo corpus + mail). Ladybug
corrupts its WAL when brand-new leafs are bulk-inserted while FTS/vector
indexes exist; a fresh DB with indexes created last is the only safe path.
Keep conversion + indexing separate so a conversion crash can't leave the
DB mid-transaction.
## Tools
```bash
bin/facts/audit.go ["self"|"db"|"contradict"] # 2-source + D16 adjudication
bin/facts/crm.go [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
bin/facts/audit ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
bin/facts/crm [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
bin/brain/search.go "query" --as-of 2025-01-01 # D24 fact intervals
bin/brain/search.go "query" --no-web # local graph only
source <(./bin/cli/complete.go bash) # flaggy completions (D23)
eval "$(bin/cgo/zig env)" # Zig cc + liblbug (not gcc)
bin/brain/index.go --rebuild [--with-mail] [--with-facts] [--with-chats]
bin/brain/add.go --text T --root facts --source "a.md x b.md" # incremental write
bin/brain/add.go --json # stdin leaf or {leafs:[...]}
bin/brain/get.go <id> [--body] [--json] # Go read; Python bin/kb/get CI fallback
bin/brain/stats.go [--json]
bin/brain/eval.go [--json] # recall@5; questions in internal/brain/rank
bin/brain/serve.go # HTTP :8630; GET /openapi.json POST /mcp
bin/stack/start # brain HTTP/MCP (reuse healthy :8630)
bin/stack/start-assistant # + reasoner + PicoClaw agent
bin/stack/status # YAML health
bin/stack/stop # compose stop; volumes kept
bin/markdown/import.go [dir] # H2 leafs → YAML; Python bin/md/import fallback
bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
bin/reasoner/bakeoff.go [--model ID] [--json] # D18 CPU tool-call bake-off
bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
bin/qa/stats.go # D22 DuckDB quantiles / JSONL (gcc CGO)
bin/mail/ocr.go <image|pdf> # tesseract eng+deu (scans)
bin/md/tables # what the graph holds → YAML
bin/brain/deduce "question" # thinking wrapper
```
Never start a shell command with `cd` — use the tool working-directory
parameter. Search before reading whole files. For YAML/JSON/XML/CSV/TOML/HCL
prefer mikefarah/yq (`skills/yq/SKILL.md`). For bulk rows and quantiles use
duckdb-go (`internal/duckstats`, `skills/duckdb/SKILL.md`), not Ladybug.
parameter. Search before reading whole files.
## GitHub safety rules (ABSOLUTE — never violate)
+16 -60
View File
@@ -1,13 +1,5 @@
# syntax=docker/dockerfile:1
#
# docker build --target api -t 2dph:api .
# docker build --target index -t 2dph:index .
#
# API: Go + ladybug via Zig CGO (no CPython).
# Index: Python write path (profile `index` until brain/add is v2).
# --- Python sidecar (Ladybug write / rebuild) ---
FROM python:3.12-slim AS index
FROM python:3.12-slim AS base
ENV PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
@@ -16,16 +8,26 @@ ENV PYTHONUNBUFFERED=1 \
WORKDIR /app
RUN id -u 2dph 2>/dev/null || useradd --create-home --uid 1001 2dph
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
poppler-utils tesseract-ocr tesseract-ocr-eng tesseract-ocr-deu \
&& rm -rf /var/lib/apt/lists/*
# deps layer-first: rebuild only on dependency change
COPY requirements.lock.txt /tmp/requirements.lock.txt
RUN python -m pip install --no-cache-dir -r /tmp/requirements.lock.txt \
&& rm /tmp/requirements.lock.txt
# Go services: static binaries, no interpreter at runtime
FROM golang:1.25 AS go-build
WORKDIR /src
COPY go.mod ./
COPY bin/server ./bin/server
COPY bin/watch ./bin/watch
RUN CGO_ENABLED=0 go build -o /serve ./bin/server \
&& CGO_ENABLED=0 go build -o /watch ./bin/watch
# runtime: python toolchain + Go services
FROM base
COPY . .
COPY --from=go-build /serve /app/bin/serve
COPY --from=go-build /watch /app/bin/watch
RUN chmod +x /app/bin/docker-entrypoint \
&& chown -R 2dph:2dph /app
USER 2dph
@@ -35,51 +37,5 @@ ENV PATH="/app/bin:${PATH}" \
KB_ROOT=/app
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
CMD python -c "import model2vec, ladybug, mistune; print('ok')" || exit 1
ENTRYPOINT ["/app/bin/docker-entrypoint"]
# --- Go API: CGO with Zig, not gcc ---
FROM golang:1.26-bookworm AS api-build
WORKDIR /src
RUN apt-get update \
&& apt-get install -y --no-install-recommends curl xz-utils ca-certificates \
&& rm -rf /var/lib/apt/lists/*
COPY bin/cgo ./bin/cgo
RUN chmod +x bin/cgo/zig bin/cgo/zcc bin/cgo/zc++ \
&& ./bin/cgo/zig env >/dev/null
COPY go.mod go.sum ./
RUN go mod download
COPY . .
ENV CGO_RPATH=/usr/local/lib
RUN eval "$(./bin/cgo/zig env)" \
&& go build -tags brain_serve,system_ladybug -o /out/brain-serve ./bin/brain/serve.go \
&& go build -tags system_ladybug -o /out/brain-search ./bin/brain/search.go \
&& CGO_ENABLED=0 go build -tags brain_watch -o /out/brain-watch ./bin/brain/watch.go
FROM debian:bookworm-slim AS api
RUN apt-get update \
&& apt-get install -y --no-install-recommends libssl3 ca-certificates wget \
&& rm -rf /var/lib/apt/lists/* \
&& useradd --create-home --uid 1001 2dph
COPY --from=api-build /out/brain-serve /usr/local/bin/brain-serve
COPY --from=api-build /out/brain-search /usr/local/bin/brain-search
COPY --from=api-build /out/brain-watch /usr/local/bin/brain-watch
COPY --from=api-build /src/lib-ladybug/liblbug.so.0.19.1 /usr/local/lib/liblbug.so.0.19.1
COPY bin/docker-entrypoint /usr/local/bin/docker-entrypoint
RUN chmod +x /usr/local/bin/docker-entrypoint \
&& ln -s liblbug.so.0.19.1 /usr/local/lib/liblbug.so.0 \
&& ln -s liblbug.so.0 /usr/local/lib/liblbug.so \
&& ldconfig
USER 2dph
ENV KB_ROOT=/data \
KB_PORT=8630 \
LD_LIBRARY_PATH=/usr/local/lib \
HF_HOME=/data/hf
WORKDIR /data
EXPOSE 8630
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
CMD wget -qO- http://127.0.0.1:8630/health || exit 1
ENTRYPOINT ["/usr/local/bin/docker-entrypoint"]
CMD ["serve"]
+35 -94
View File
@@ -4,14 +4,7 @@ A brain that loves facts and deduction. Evidence-first knowledge graph + hybrid
RAG over the operational Brain/ops/eSlider stack. Built like Sherlock
Holmes: nothing is asserted unless it has proof.
Status: **v1 in** (epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed).
v2 board: milestone [v2](https://git.produktor.io/eSlider/2dph/milestone/13) —
OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6) in,
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1 in,
[#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3 in,
[#34](https://git.produktor.io/eSlider/2dph/issues/34) D23 in,
[#36](https://git.produktor.io/eSlider/2dph/issues/36) OQ5/D24 in.
Gap: [docs/roadmap.md](docs/roadmap.md).
Status: **in progress** — this file is the plan and the record of decisions.
## What
@@ -33,12 +26,12 @@ detective method: **a fact needs ≥2 independent sources or it is
|---|----------|--------|
| D1 | RAG corpus | ops stack (chat, onlyoffice, gitea/NPM, searchxng, observability, ai-bot, mcp-servers, `~/.ssh/config`) + portfolio. Exclude `office.dev` + jobs/applications. |
| D2 | skill merging | integrate skills **in this project** `skills/`; skip gitea / brain-dependent skills. |
| D3 | web search | Go client `bin/web/search.go` (`internal/websearch`). SearXNG URL is config (`BRAIN_SEARCH_URL`). Optional Compose profile `searxng` (sanitized settings). Do not run a second copy on a host that already has one. Empty/`throttled` ≠ “nothing exists”. |
| D3 | web search | import `web-search`, retire local `searxng-ops`. Vendored here, no remote link. |
| D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. |
| D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). |
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`). Read path is Go + Zig CGO (D21). Python `bin/kb/{get,stats,eval}` is the CI fallback when Zig/libs are not fetched. Incremental write is Python `bin/kb/add` (`bin/brain/add.go`). Bulk rebuild stays `compose --profile index` until the Go write path is safe. |
| D6 | graph engine | **LadybugDB** (Kuzu successor, MIT, embedded, native FTS+vector+Cypher). Python binding for `bin/*`; Go shebang for golang tools. |
| D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). |
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. 2-source auto-pair docker ps × compose × ssh-config × docs. |
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. |
| D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. |
| D10 | versioning | everything is a leaf with `sha256 + observed_at + source_rev`; `File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`. Stale = `source_rev` < git HEAD. |
| D11 | strong/weak | `root` column: `facts` (strong) vs `info` (weak). Answer is `confirmed` only from facts root. |
@@ -46,15 +39,7 @@ detective method: **a fact needs ≥2 independent sources or it is
| D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. |
| D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. |
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
| D16 | contradictions | ≥2 yes vs ≥2 no → hypothesis → `(not confirmed)` until a rule fires. Order: **temporal_freshness** (fresh ≥2 vs stale minority), then **authority_pairing** (runtime/config A×B beats narrative C). Store as `a x b vs c x d` on hypothesis leafs. `bin/facts/audit contradict`. [#29](https://git.produktor.io/eSlider/2dph/issues/29). |
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
| D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is compose profile `picoclaw`; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. Agent lever/loop: [#15](https://git.produktor.io/eSlider/2dph/issues/15). |
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
| D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`/`ingest`). |
| D21 | CGO | Ladybug/tokenizers CGO is compiled with **Zig** (`bin/cgo/zcc``zig cc -target …-linux-gnu`), not gcc. `bin/cgo/zig` pins Zig 0.14.1 + liblbug 0.19.1 + libtokenizers 1.27.0. Compose `target: api` has no CPython; write/rebuild is profile `index`. |
| D22 | analytics | **duckdb-go** in-process (`internal/duckstats`, `bin/qa/stats.go`) for quantiles/JSONL. Links with **gcc/g++**, not Zig. Ladybug stays the graph; web-search cache stays modernc sqlite. Slice small structured docs with **mikefarah/yq**, not kislyuk/jq. [#30](https://git.produktor.io/eSlider/2dph/issues/30). |
| D23 | CLI | **flaggy** (`github.com/integrii/flaggy`, 0 deps). Flags at any position. Wrapper `internal/cli`. Bash complete: `source <(./bin/cli/complete.go bash)`. No cobra, no stdlib `flag` in Go tools. Search does not intercept the word `completion`. [#34](https://git.produktor.io/eSlider/2dph/issues/34). |
| D24 | fact intervals | Leaf `valid_from` / `valid_to` (YYYY-MM-DD, inclusive; empty = open/legacy). Search `--as-of` / MCP `as_of` keeps facts active that day. Not D16 `temporal_freshness` (source stale vs HEAD). Empty interval = always visible. [#36](https://git.produktor.io/eSlider/2dph/issues/36). |
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. |
## Architecture
@@ -62,32 +47,16 @@ detective method: **a fact needs ≥2 independent sources or it is
2dph/
PLAN.md / AGENTS.md
docs/ published docs (this conversation → docs/ as md)
skills/ in-project skills (web-search, postgres, brain, picoclaw, diataxis-docs)
skills/ in-project skills (web-search, db-yaml, kb-search, agent-cost, diataxis-docs, …)
bin/
facts/extract.go audit.go crm.go # D14 shebang; Python implementation
kb/index Python bulk write (called by bin/brain/index.go)
kb/add Python incremental write (called by bin/brain/add.go)
brain/index.go rebuild FTS + HNSW (incl. --with-mail)
brain/add.go incremental leaf write (no rebuild)
brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback
brain/watch.go
brain/search.go deduction: facts → info → web
cli/complete.go flaggy bash/zsh/fish complete (D23)
brain/serve.go HTTP API in-process + OpenAPI/MCP (D20); Zig CGO (D21)
cgo/zig zcc zc++ CGO toolchain (zig cc, not gcc)
mail/import.go JSON → markdown (no brain write)
markdown/import.go H2 leaf split (Go); Python bin/md/import fallback
postgres/query.go read-only YAML (wraps bin/db/psql-yq)
git/import.go go-git history (no git binary; conversion only)
web/search.go SearXNG client (throttled ≠ absence)
reasoner/bakeoff.go CPU tool-call bake-off (D18; OpenAI tools)
chats/sync.go import.go facts.go apply.go
(libs in internal/chats; no chats index)
mail/ocr.go tesseract eng+deu (pdftoppm scans)
md/import (deprecated; bin/markdown/import.go)
facts/extract auto-pair 2 sources → lexicon yaml + graph
facts/audit ["self"|"facts"|"info"|"stale"] 2-source + staleness gate
kb/index build FTS + HNSW from corpus
kb/search deduction: facts → info → web-search; --hop N
kb/get kb/stats kb/eval
md/import md/select md/tables md/gaps (mistune)
brain/extract brain/audit brain/deduce (thinking wrapper)
stack/start start-assistant stop status
web/search (deprecated shim → web/search.go)
web/search (vendored)
db/psql-yq (vendored)
ssh-tunnel onlyoffice pg tunnel 5433
var/kb.lbug single embedded store (gitignored)
@@ -98,14 +67,10 @@ detective method: **a fact needs ≥2 independent sources or it is
Node tables: `Person, Service, Host, Container, Repo, File, Commit, Leaf`.
`Leaf(embedding FLOAT[N])` — FTS on `text`, HNSW vector index on `embedding`.
Edges: `RUNS / USES / FROM_FILE / HAS_VERSION / AUTHORED / ABOUT / ASSOCIATED / SIMILAR_0.85`.
`FROM_FILE` / `HAS_VERSION` / `AUTHORED`: `bin/brain/search.go --hop N` walks
them from each hit (1=File, 2=Commit, 3=Person). Rebuild writes
`Leaf-[:FROM_FILE]->File`; git import writes the rest.
Edges: `RUNS / USES / HAS_VERSION / AUTHORED / ABOUT / ASSOCIATED / SIMILAR_0.85`.
Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
`where`, `when`, `source_rev`. Leaf interval of truth (D24): `valid_from`,
`valid_to`.
`where`, `when`, `source_rev`.
## Config
@@ -120,51 +85,44 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
- `bin/{subject}/{method}` — line 2 is a usage comment (mirrors `psql-yq`).
- bash + python primary; golang via Go shebang when a compiled helper is right.
- YAML default output, `--json` for machines. Slice with mikefarah/yq.
- YAML default output, `--json` for machines. Slice with `yq`.
- Everything that touches the network / DB is read-only, throttled, cached.
- Tests (TDD) gate every commit; `gh` + CI/CD on every push.
## Open questions (v2)
- OQ1: **in** — D16 adjudication: `temporal_freshness` then `authority_pairing`.
Unresolved 2v2 stays hypothesis. [#29](https://git.produktor.io/eSlider/2dph/issues/29).
- OQ2: OCR **in**. `pdftotext -layout` first; scans `pdftoppm` + tesseract
`eng+deu` (`bin/mail/ocr.go`, `internal/ocr`). No gocv, no gosseract CGO
(D21 Zig owns Ladybug CGO). Optional `OCR_ENGINE=paddle` / compose profile
`ocr-paddle`. Docling left the default path. [#6](https://git.produktor.io/eSlider/2dph/issues/6).
- OQ3: **in** — duckdb-go (`internal/duckstats`, `bin/qa/stats.go`) for
quantiles / JSONL count. Not a second graph. [#30](https://git.produktor.io/eSlider/2dph/issues/30).
- OQ1: mutually-contradicting evidence — how to resolve (authority weighting,
temporal freshness, audit adjudication).
- OQ2: OCR pipeline for pdfs/images/docs — mostly solved: poppler pdftotext
fast-path for born-digital PDFs, docling fallback for the ~5% textless ones.
- OQ3: optional duckdb-md layer for `SELECT … FORMAT MARKDOWN` export/write-back.
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
serialize and unambiguous; YAML only where humans edit files.
- OQ5: **in** — fact `valid_from` / `valid_to` + `--as-of` / MCP `as_of` (D24).
Not D16 `temporal_freshness`. [#36](https://git.produktor.io/eSlider/2dph/issues/36).
## Mail pipeline (done)
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
Gmail attachments key off `body.attachmentId`, not MIME `partId`.
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
`pdftotext -layout` (~15ms); textless/scanned PDFs `pdftoppm` + tesseract
`eng+deu`. ICS sidecars
2. `bin/mail/import --from-raw` — message.json → message.md; PDFs via
`pdftotext -layout` (~15ms) with docling subprocess fallback; ICS sidecars
Latin-1→UTF-8 normalized.
3. `bin/brain/index.go --rebuild` — fresh rebuild (repo corpus + mail) because ladybug
3. `bin/mail/index_mail` — fresh rebuild (repo corpus + mail) because ladybug
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
indexing stay separate for crash safety. `bin/mail/index_mail` is a
deprecation shim.
indexing stay separate for crash safety.
4. Result: 17,835 messages → 28,918 info leafs, FTS + HNSW healthy, searchable
via `bin/brain/search.go`.
via `bin/kb/search`.
## CI/CD pipeline (D15)
`.github/workflows/ci.yml`:
1. go vet + go test ./... (root module; packages without ladybug cgo)
2. `go test ./internal/brain/rank` (cgo-free ranking + flag parser)
3. python -m unittest discover -s bin/tools (includes published-docs SoT)
4. `bin/facts/audit self` (lexicon internal consistency; `bin/facts/audit.go` is the D14 wrapper)
5. `bin/brain/eval.go` via Zig (recall@5 ≥ 0.95). Python `bin/kb/eval` is an
explicit fallback, not the CI SoT.
6. `bin/cgo/zig go build -tags system_ladybug` (compile search with zig cc; fetches pinned zig+libs).
1. go vet + go test ./... (Go tools; root module)
2. `go test ./rank` in `bin/kbsearch` (cgo-free ranking + flag parser; nested module still needs ladybug for the rest)
3. `go test ./...` in `bin/chats` (Telegram + LinkedIn parsers; nested module)
4. python -m unittest discover (Py tools)
5. bin/facts/audit self (lexicon internal consistency)
6. bin/kb/eval (recall@5 ≥ 0.95, gates index regressions)
7. md-docs build/lint if docs tooling arrives.
Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as
`db/tech-poc`: contract first where there is an OpenAPI/message shape.
@@ -173,26 +131,9 @@ Feedback loop: every commit → PR → CI → green/gate → merge. Same discipl
1. scaffold repo (:done after this file + AGENTS.md + .gitignore + ci)
2. gh repo create eSlider/2dph --private + initial commit + CI
3. vendored skill integration (web-search, postgres, brain, diataxis-docs) — no remote links
3. vendored skill integration (web-search, db-yaml, kb-search, agent-cost, diataxis-docs) — no remote links
4. .venv: ladybug + model2vec + mistune
5. schema + tools with TDD (kb + md + facts + brain)
6. ~/.config/brain config
7. corpus extraction (facts/info)**in**: [#18](https://git.produktor.io/eSlider/2dph/issues/18)
7. corpus extraction (facts/info)
8. verify: web-search smoke, onlyoffice pg, md-db round-trip, eval, audit
## Gap to v1 (epic #16)
Remaining: none for epic #16 (v1). Board:
[epic #16](https://git.produktor.io/eSlider/2dph/issues/16),
milestone [v1 detective brain](https://git.produktor.io/eSlider/2dph/milestone/12).
Narrative: [docs/roadmap.md](docs/roadmap.md).
| Order | Issue | Gap |
|-------|-------|-----|
| 1 | [#14](https://git.produktor.io/eSlider/2dph/issues/14) | **in**`bin/brain/add.go` / `POST /ingest` write facts+info without deleting `kb.lbug`. Bulk corpus still `--rebuild`. Leftover Python (mail/facts) is not the living-graph blocker. |
| 2 | [#17](https://git.produktor.io/eSlider/2dph/issues/17) | **in**`--hop N` walks `FROM_FILE``HAS_VERSION``AUTHORED` (max 3). |
| 3 | [#18](https://git.produktor.io/eSlider/2dph/issues/18) | **in**`--with-facts` / `--facts-json` land `root=facts`; `--with-chats` indexes `var/chats/md`. WhatsApp sync is out of v1. |
| 4 | [#15](https://git.produktor.io/eSlider/2dph/issues/15) | **in** — lever/loop documented (`search``get``audit`). |
| 5 | [#19](https://git.produktor.io/eSlider/2dph/issues/19) | **in** — CI recall SoT is `bin/brain/eval.go` via Zig. Python `bin/kb/eval` stays as an explicit fallback. |
Does **not** block epic close: OQ4. OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6), OQ1 [#29](https://git.produktor.io/eSlider/2dph/issues/29), OQ3 [#30](https://git.produktor.io/eSlider/2dph/issues/30) are **in**.
+46 -82
View File
@@ -7,16 +7,14 @@
[![Latest Release](https://img.shields.io/github/v/tag/eSlider/2dph?sort=semver&label=release)](https://github.com/eSlider/2dph/releases)
[![GitHub Stars](https://img.shields.io/github/stars/eSlider/2dph?style=social)](https://github.com/eSlider/2dph/stargazers)
An evidence-first brain. **Facts need two independent sources, or they are
`(not confirmed)`.** Cursor is not the runtime.
An evidence-first brain over the operational eSlider stack. **Facts need two
independent sources, or they are `(not confirmed)`.**
`2dph` is a single embedded knowledge graph (LadybugDB) with native **HNSW
vector** + **BM25 full-text** indexes. Search is *deduction*: confirmed facts
first, supporting info second, `web-search` as the independent second source
when the local graph cannot confirm.
Run it: [docs/runbook.md](docs/runbook.md). Design: [docs/design.md](docs/design.md).
Docs index: [docs/README.md](docs/README.md).
`2dph` is a single embedded knowledge graph (LadybugDB = Kuzu successor) with
native **HNSW vector** + **BM25 full-text** indexes, built from markdown,
compose files, ssh config, docker state, and git history. Search is
*deduction*: confirmed facts first, supporting info second, `web-search` as
the independent second source when the local graph cannot confirm.
## Architecture
@@ -30,11 +28,11 @@ graph TB
end
subgraph dph["2dph tools"]
EX["bin/facts/extract.go<br/>2-source pairing"]
AU["bin/facts/audit.go<br/>confidence + staleness"]
IDX["bin/brain/index.go<br/>chunk + embed"]
MD["bin/markdown/import.go<br/>H2 leaf split"]
SR["bin/brain/search.go<br/>deduction"]
EX["bin/facts/extract<br/>2-source pairing"]
AU["bin/facts/audit<br/>confidence + staleness"]
IDX["bin/kb/index<br/>chunk + embed"]
MD["bin/md/import<br/>mistune leaves"]
SR["bin/kb/search<br/>deduction + --hop"]
end
subgraph store["Ladybug var/kb.lbug"]
@@ -76,7 +74,7 @@ graph TB
## The method
Every assertion is `Who / What / How / Where / When + evidence + confidence`,
mirroring the detective method: **≥2 independent sources confirm a
mirroring the detective detective skill: **≥2 independent sources confirm a
fact; conflicting sources or a single source → `hypothesis``(not confirmed)`.**
| root | meaning | used for answers |
@@ -87,104 +85,70 @@ fact; conflicting sources or a single source → `hypothesis` → `(not confirme
## Deduction search
```bash
bin/brain/search.go "Matrix federation over HTTPS" # facts → info → web
bin/brain/search.go "onlyoffice postgres" --root facts
bin/brain/search.go "where is cs-lexicon" --json | yq '.'
bin/brain/search.go "upstream flag" --no-web # local graph only
bin/brain/get.go <id> --body # full chunk on demand
bin/brain/stats.go # index health
bin/brain/eval.go # recall@5 gate
```
`--hop N` walks File/Commit/Person from each hit (max 3). `bin/kb/search` is a deprecated wrapper around `bin/brain/search.go`.
Git history is read with [go-git](https://github.com/go-git/go-git) (no git binary):
```bash
bin/git/import.go --json --limit 100 # commit leafs for this repo
bin/git/import.go --root "$PROJECTS_ROOT" --json # one pass per .git under root
```
Conversion only. Graph write (`File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`) stays with `bin/brain/index.go`.
Web search (second independent source) goes through SearXNG. Empty results mean **throttled**, not “nothing exists”:
```bash
bin/web/search.go "LadybugDB vector index" --json
# Optional local instance (skip if BRAIN_SEARCH_URL already points at one):
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
bin/kb/search "Matrix federation over HTTPS" # facts → info → web-search
bin/kb/search "what runs on arc-2" --hop 1 # walk graph edges
bin/kb/search "where is cs-lexicon" --json | yq '.' # YAML by default
bin/kb/get <id> --body # full chunk on demand
bin/kb/stats # index health
bin/kb/eval # recall@5 gate
```
Mail is a first-class corpus (retrievable through the same search):
```bash
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
bin/mail/import.go --from-raw var/mail # JSON → markdown
bin/brain/add.go --text T --root facts --source "a.md x b.md"
bin/brain/index.go --rebuild --with-facts --with-chats # facts extract + chats md
bin/brain/index.go --rebuild # rebuild brain (incl. mail)
bin/brain/search.go "invoice from last week" # same search over mail leafs
bin/mail/import --from-raw var/mail # JSON → markdown
bin/mail/index_mail # rebuild brain incl. mail
bin/kb/search "Mietwagen Nürnberg invoice" # now answers from mail
```
## Storage
- **LadybugDB** — single `var/kb.lbug`, Cypher + HNSW + BM25, embedded.
Read tools (`get` / `stats` / `eval`) are Go + Zig CGO (`bin/cgo/zcc`).
Python fallbacks stay for CI until the runner fetches Zig. Incremental
write is `bin/brain/add.go` (Python `kblib.add_leafs`). Bulk rebuild is
Compose profile `index` (`bin/brain/index.go --rebuild`).
- **model2vec** — `potion-multilingual-128M` (256-dim), CPU, no Ollama
runtime dependency.
- facts and info split by `root` but written in the same transaction.
Ladybug 0.19 DROP INDEX warning: [docs/runbook.md](docs/runbook.md).
- **LadybugDB** — single `var/kb.lbug`, Cypher property graph, HNSW + BM25
in one engine, embedded (no server), ACID, read-only-safe for concurrent
readers. **Never `DROP INDEX` FTS/VECTOR** on Ladybug 0.19: DROP leaves
ghost catalog tables (`_0_Leaf_vec_UPPER`) so recreate fails while
`SHOW_INDEXES` omits HNSW. Fresh indexes = delete `var/kb.lbug` +
`bin/kb/index --rebuild`. Use `ensure_indexes()` after upserts.
- **model2vec** — `potion-multilingual-128M` static embeddings (256-dim),
CPU-fast, deterministic, no Ollama runtime dependency.
- facts and info split semantically by `root` column but written inside the
same transaction.
## Tooling conventions
`bin/{subject}/{method}.go` — self-describing: shebang on line 1, usage comment
from line 2. Shared code in `internal/`. YAML default output, `--json` for
machines. Tests gate every commit. HTTP: `bin/brain/serve.go` calls
`internal/brain` in-process (`/health` `/search` `/get` `/stats` `/audit` `/ingest` `/openapi.json` `/mcp`).
`bin/{subject}/{method}` — self-describing: shebang on line 1, usage comment
from line 2. bash + python primary; golang via the Go shebang when a compiled
helper is right. YAML default output, `--json` for machines. Everything that
touches network/db is read-only, throttled, cached. Tests gate every commit.
## Development
See the portable runbook: [docs/runbook.md](docs/runbook.md).
```bash
uv venv .venv
uv pip install -r requirements.lock.txt
bin/facts/audit.go self
go test ./... && uv run python -m unittest discover -s bin/tools -t .
uv venv .venv # Python 3.12, uv-managed
uv pip install -r requirements.lock.txt # pinned toolchain
bin/facts/audit self # lexicon consistency gate
go test ./... && python -m unittest discover -s bin/tools -t .
```
Docker (optional, cached model + var volumes):
```bash
bin/stack/start # brain HTTP/MCP :8630
bin/stack/start-assistant # + qwen3.5:9b + PicoClaw agent
bin/stack/status
bin/stack/stop
docker compose up -d brain # API (Zig CGO serve :8630)
docker compose --profile index run --rm index # Python Ladybug rebuild
docker compose --profile picoclaw up brain-mcp # MCP on 127.0.0.1:8630
docker compose --profile reasoner up -d reasoner # CPU Ollama 127.0.0.1:11435
docker compose run --rm brain index # (re)index corpus
docker compose run --rm brain search "query" # one-shot query
docker compose run --rm brain serve # async Go HTTP server
docker compose up brain-watch # auto re-index on change
```
## Related
eSlider DevOps engineer practice: ops, OnlyOffice, and mail feed the facts
root through `bin/facts/extract` (two-source pairing).
- [go-second-brain](https://github.com/eSlider/go-second-brain) — the earlier
Neo4j + Qdrant + Matrix RAG brain
- [agent-skills](https://github.com/eSlider/agent-skills) — upstream
skills (`web-search`, `postgres`, …) that 2dph integrates
skills (`web-search`, `db-yaml`, …) that 2dph integrates
- detective method — the two-source method
Work board (issues): [epic #16](https://git.produktor.io/eSlider/2dph/issues/16)
on [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
Work board (issues): [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
See [PLAN.md](PLAN.md) for decisions, [docs/roadmap.md](docs/roadmap.md) for
the gap to v1, and v2 open questions.
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
-21
View File
@@ -1,21 +0,0 @@
//usr/bin/env go run -tags=brain_add "$0" "$@"; exit
//go:build brain_add
//
// bin/brain/add.go - incremental leaf write (Python kblib, no rebuild).
//
// ./bin/brain/add.go --text T --root facts --source "a.md x b.md"
// ./bin/brain/add.go --json
//
// D6: write stays Python. Does not delete var/kb.lbug.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/kb/add", os.Args[1:]))
}
+2 -2
View File
@@ -1,3 +1,3 @@
// Commands in this directory are shebang mains (search.go, serve.go, index.go,
// get.go, stats.go, eval.go, watch.go), each behind an exclusive build tag.
// Commands in this directory are shebang mains (search.go).
// search.go is behind the system_ladybug build tag (cgo).
package main
-22
View File
@@ -1,22 +0,0 @@
//usr/bin/env go run -tags=system_ladybug,brain_eval "$0" "$@"; exit
//go:build cgo && system_ladybug && brain_eval
//
// bin/brain/eval.go - recall@5 gate.
//
// ./bin/brain/eval.go
// ./bin/brain/eval.go --json
//
// Needs CGO + libladybug. Python bin/kb/eval is the CI fallback (no cgo).
// Control questions live in internal/brain/rank (cgo-free).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/brain"
)
func main() {
os.Exit(brain.MainEval(os.Args[1:]))
}
-23
View File
@@ -1,23 +0,0 @@
//usr/bin/env go run -tags=system_ladybug,brain_get "$0" "$@"; exit
//go:build cgo && system_ladybug && brain_get
//
// bin/brain/get.go - read one leaf by id.
//
// ./bin/brain/get.go <id>
// ./bin/brain/get.go <id> --body
// ./bin/brain/get.go <id> --json
//
// Needs CGO + libladybug. Python bin/kb/get is the CI fallback (no cgo).
// CGO compiler is Zig (`eval "$(bin/cgo/zig env)"`), not gcc.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/brain"
)
func main() {
os.Exit(brain.MainGet(os.Args[1:]))
}
-24
View File
@@ -1,24 +0,0 @@
//usr/bin/env go run -tags=brain_index "$0" "$@"; exit
//go:build brain_index
//
// bin/brain/index.go - rebuild the Ladybug graph (Python write path).
//
// ./bin/brain/index.go --rebuild --with-facts --with-chats
// ./bin/brain/index.go --rebuild --with-mail
// ./bin/brain/index.go --dry-run --with-mail
//
// v1 write: bin/brain/add.go for one/few leafs (indexes may already exist).
// Bulk mail/corpus still --rebuild (fresh file, indexes last).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
args := append([]string{"--with-mail"}, os.Args[1:]...)
os.Exit(cmdbin.ExecFile("bin/kb/index", args))
}
+2 -2
View File
@@ -3,11 +3,11 @@
//
// bin/brain/search.go - deduction search over the 2dph brain.
//
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--hop N] [--json] [--no-web]
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--json]
// ./bin/brain/search.go serve [port]
// ./bin/brain/search.go --list-model
//
// Needs CGO + libladybug via Zig (`eval "$(bin/cgo/zig env)"`), not gcc.
// Needs CGO + libladybug (CGO_CFLAGS/CGO_LDFLAGS). Prefer the wrapper
// bin/kb/search which sets those and builds a binary for the embed daemon.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
+6 -14
View File
@@ -1,23 +1,18 @@
//usr/bin/env go run -tags=brain_serve,system_ladybug "$0" "$@"; exit
//go:build brain_serve && cgo && system_ladybug
//usr/bin/env go run -tags=brain_serve "$0" "$@"; exit
//go:build brain_serve
//
// bin/brain/serve.go - HTTP API (in-process ladybug search).
// bin/brain/serve.go - HTTP API for the 2dph brain.
//
// KB_ROOT=/path/to/2dph ./bin/brain/serve.go
// KB_WORKERS=4 KB_PORT=8630 ./bin/brain/serve.go
// KB_SEARCH_CMD=... KB_WORKERS=4 KB_PORT=8630 ./bin/brain/serve.go
//
// GET /openapi.json same Ops table as the handlers
// POST /mcp JSON-RPC tools/list + tools/call
//
// Needs CGO + libladybug (same as bin/brain/search.go).
// Default search backend is var/bin/brain-search (Go), not Python.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"log"
"os"
"github.com/eSlider/2dph/internal/brain"
"github.com/eSlider/2dph/internal/httpapi"
)
@@ -27,8 +22,5 @@ func main() {
os.Setenv("KB_ROOT", wd)
}
}
if err := brain.Ready(); err != nil {
log.Fatal(err)
}
httpapi.Run(brain.HTTP{})
httpapi.Run()
}
-20
View File
@@ -1,20 +0,0 @@
//go:build brain_serve && !system_ladybug
//
// Fallback serve when ladybug cgo is not in the build (CI / tags=brain_serve).
// Production shebang is serve.go (in-process).
package main
import (
"os"
"github.com/eSlider/2dph/internal/httpapi"
)
func main() {
if os.Getenv("KB_ROOT") == "" {
if wd, err := os.Getwd(); err == nil {
os.Setenv("KB_ROOT", wd)
}
}
httpapi.Run(nil)
}
-21
View File
@@ -1,21 +0,0 @@
//usr/bin/env go run -tags=system_ladybug,brain_stats "$0" "$@"; exit
//go:build cgo && system_ladybug && brain_stats
//
// bin/brain/stats.go - index health.
//
// ./bin/brain/stats.go
// ./bin/brain/stats.go --json
//
// Needs CGO + libladybug. Python bin/kb/stats is the CI fallback (no cgo).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/brain"
)
func main() {
os.Exit(brain.MainStats(os.Args[1:]))
}
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=brain_watch "$0" "$@"; exit
//go:build brain_watch
//
// bin/brain/watch.go - re-index when corpus files change.
//
// ./bin/brain/watch.go [dir...]
// KB_WATCH_INTERVAL=15 ./bin/brain/watch.go
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/bin/watch"
)
func main() {
watch.Run(os.Args[1:])
}
-23
View File
@@ -1,23 +0,0 @@
#!/bin/sh
# bin/cgo/zc++ — CGO CXX. Zig, not g++.
set -eu
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
case "$(uname -m)" in
x86_64|amd64) TARGET=x86_64-linux-gnu ;;
aarch64|arm64) TARGET=aarch64-linux-gnu ;;
*)
echo "zc++: unsupported arch $(uname -m)" >&2
exit 2
;;
esac
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
:
elif [ -x "$ROOT/var/zig/zig" ]; then
ZIG="$ROOT/var/zig/zig"
elif command -v zig >/dev/null 2>&1; then
ZIG="$(command -v zig)"
else
echo "zc++: zig missing; run bin/cgo/zig first" >&2
exit 127
fi
exec "$ZIG" c++ -target "$TARGET" "$@"
-24
View File
@@ -1,24 +0,0 @@
#!/bin/sh
# bin/cgo/zcc — CGO CC. Zig, not gcc.
# Go invokes CC with many args; a wrapper avoids spaces in $CC.
set -eu
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
case "$(uname -m)" in
x86_64|amd64) TARGET=x86_64-linux-gnu ;;
aarch64|arm64) TARGET=aarch64-linux-gnu ;;
*)
echo "zcc: unsupported arch $(uname -m)" >&2
exit 2
;;
esac
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
:
elif [ -x "$ROOT/var/zig/zig" ]; then
ZIG="$ROOT/var/zig/zig"
elif command -v zig >/dev/null 2>&1; then
ZIG="$(command -v zig)"
else
echo "zcc: zig missing; run bin/cgo/zig first" >&2
exit 127
fi
exec "$ZIG" cc -target "$TARGET" "$@"
-130
View File
@@ -1,130 +0,0 @@
#!/usr/bin/env bash
# bin/cgo/zig — CGO toolchain: zig cc (not gcc) + pinned liblbug + libtokenizers.
#
# eval "$(bin/cgo/zig env)" # export CC/CXX/CGO_*
# bin/cgo/zig go build ... # ensure, then exec with env
# bin/cgo/zig ./bin/brain/search.go "query"
#
# Pins live in this file. Downloads land in var/ (gitignored).
set -euo pipefail
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
ZIG_VERSION=0.14.1
LBUG_VERSION=0.19.1
TOKENIZERS_VERSION=1.27.0
arch="$(uname -m)"
case "$arch" in
x86_64|amd64)
ZIG_ARCH=x86_64
LBUG_ARCH=x86_64
TOK_ARCH=x86_64
ZIG_SHA=24aeeec8af16c381934a6cd7d95c807a8cb2cf7df9fa40d359aa884195c4716c
LBUG_SHA=ed263ae913f68cb0ddba0b98548b58edaac49929766d03bdaaa83be46c68847d
TOK_SHA=72556cdca798dd4ea7cdaba308e5f0d68a8cb93b67c96edf485b7a0edd7b07f4
;;
aarch64|arm64)
ZIG_ARCH=aarch64
LBUG_ARCH=aarch64
TOK_ARCH=aarch64
ZIG_SHA=f7a654acc967864f7a050ddacfaa778c7504a0eca8d2b678839c21eea47c992b
LBUG_SHA=b07df2cd533c3976a2a3025866d6420a5f35514d0a822ecc4b2902d55b4725b7
TOK_SHA=e96545ad05930c26f51f63d932ee6d3bbd32bbed149e102c5290d587a2293067
;;
*)
echo "bin/cgo/zig: unsupported arch $arch" >&2
exit 2
;;
esac
CACHE="$ROOT/var/cache"
LIB="$ROOT/lib-ladybug"
ZIG_DIR="$ROOT/var/zig-dist"
ZIG_BIN="$ROOT/var/zig/zig"
sha256of() {
if command -v sha256sum >/dev/null 2>&1; then
sha256sum "$1" | awk '{print $1}'
else
shasum -a 256 "$1" | awk '{print $1}'
fi
}
fetch() {
local url="$1" dest="$2" expect="$3"
if [ -f "$dest" ] && [ "$(sha256of "$dest")" = "$expect" ]; then
return 0
fi
mkdir -p "$(dirname "$dest")"
echo "fetch $url" >&2
curl -fsSL "$url" -o "$dest"
local got
got="$(sha256of "$dest")"
if [ "$got" != "$expect" ]; then
echo "checksum mismatch $dest: got $got want $expect" >&2
rm -f "$dest"
exit 1
fi
}
ensure_zig() {
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
return 0
fi
if [ -x "$ZIG_BIN" ]; then
export ZIG="$ZIG_BIN"
return 0
fi
if command -v zig >/dev/null 2>&1; then
export ZIG
ZIG="$(command -v zig)"
return 0
fi
local tar="$CACHE/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}.tar.xz"
fetch "https://ziglang.org/download/${ZIG_VERSION}/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}.tar.xz" \
"$tar" "$ZIG_SHA"
mkdir -p "$CACHE"
rm -rf "$ZIG_DIR"
tar -xJf "$tar" -C "$CACHE"
mv "$CACHE/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}" "$ZIG_DIR"
mkdir -p "$ROOT/var/zig"
ln -sfn "$ZIG_DIR/zig" "$ZIG_BIN"
export ZIG="$ZIG_BIN"
}
ensure_libs() {
mkdir -p "$LIB"
if [ ! -f "$LIB/liblbug.so" ]; then
local tar="$CACHE/liblbug-linux-${LBUG_ARCH}.tar.gz"
fetch "https://github.com/LadybugDB/ladybug/releases/download/v${LBUG_VERSION}/liblbug-linux-${LBUG_ARCH}.tar.gz" \
"$tar" "$LBUG_SHA"
tar -xzf "$tar" -C "$LIB"
fi
if [ ! -f "$LIB/libtokenizers.a" ]; then
local tar="$CACHE/libtokenizers.linux-${TOK_ARCH}.tar.gz"
fetch "https://github.com/daulet/tokenizers/releases/download/v${TOKENIZERS_VERSION}/libtokenizers.linux-${TOK_ARCH}.tar.gz" \
"$tar" "$TOK_SHA"
tar -xzf "$tar" -C "$LIB"
fi
}
print_env() {
printf 'export ZIG=%q\n' "$ZIG"
printf 'export CC=%q\n' "$ROOT/bin/cgo/zcc"
printf 'export CXX=%q\n' "$ROOT/bin/cgo/zc++"
printf 'export CGO_ENABLED=1\n'
printf 'export CGO_CFLAGS=%q\n' "-I$LIB"
printf 'export CGO_LDFLAGS=%q\n' "-L$LIB -Wl,-rpath,${CGO_RPATH:-$LIB}"
}
ensure_zig
ensure_libs
cmd="${1:-env}"
if [ "$cmd" = "env" ]; then
print_env
exit 0
fi
eval "$(print_env)"
exec "$@"
-19
View File
@@ -1,19 +0,0 @@
//usr/bin/env go run -tags=chats_apply "$0" "$@"; exit
//go:build chats_apply
//
// bin/chats/apply.go - push extracted chat facts to OnlyOffice CRM.
//
// ./bin/chats/apply.go [--dry-run]
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/chats"
)
func main() {
os.Exit(chats.RunApply(os.Args[1:]))
}
@@ -1,15 +1,14 @@
package chats
package main
import (
"bytes"
"encoding/json"
"flag"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
cliparse "github.com/eSlider/2dph/internal/cli"
)
type ooContact struct {
@@ -25,10 +24,17 @@ type ooContact struct {
} `json:"commonData"`
}
func RunApply(args []string) int {
dryRun, err := parseApplyFlags(args)
if err != nil {
return cliparse.Fail(err)
func runApply(args []string) int {
fs := flag.NewFlagSet("chats apply", flag.ContinueOnError)
dryRun := fs.Bool("dry-run", false, "show what would be done without writing")
help := fs.Bool("help", false, "")
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats apply [--dry-run]")
return 0
}
ooCLI := findOO()
@@ -122,7 +128,7 @@ func RunApply(args []string) int {
fmt.Printf("\nchats apply: %d actions to apply\n", len(resolved))
if dryRun {
if *dryRun {
for _, r := range resolved {
switch r.Action {
case "info-add":
@@ -170,7 +176,7 @@ func RunApply(args []string) int {
}
func loadFacts() ([]ExtractedFact, error) {
factsPath := filepath.Join(Dir(), "facts", "chat-facts.json")
factsPath := filepath.Join(chatsDir(), "facts", "chat-facts.json")
data, err := os.ReadFile(factsPath)
if err != nil {
if os.IsNotExist(err) {
@@ -3,7 +3,7 @@
// These are integration tests using real data and real Telegram API (when
// credentials are available). They follow the TDD workflow pattern:
// sync → import → facts → verify.
package chats
package main
import (
"encoding/json"
@@ -48,7 +48,7 @@ func TestChatsImport(t *testing.T) {
t.Cleanup(func() { os.Chdir(cwd) })
t.Setenv("KB_ROOT", dir)
exitCode := RunImport([]string{})
exitCode := runImport([]string{})
if exitCode != 0 {
t.Fatalf("import exit code %d", exitCode)
}
@@ -140,7 +140,7 @@ func TestChatsImportEmpty(t *testing.T) {
t.Cleanup(func() { os.Chdir(cwd) })
t.Setenv("KB_ROOT", dir)
exitCode := RunImport([]string{})
exitCode := runImport([]string{})
if exitCode == 0 {
t.Fatal("expected non-zero exit for empty data dir")
}
@@ -171,7 +171,7 @@ func TestChatsRoundTrip(t *testing.T) {
t.Cleanup(func() { os.Chdir(cwd) })
t.Setenv("KB_ROOT", dir)
if code := RunImport([]string{}); code != 0 {
if code := runImport([]string{}); code != 0 {
t.Fatalf("import exit %d", code)
}
-4
View File
@@ -1,4 +0,0 @@
// Commands in this directory are shebang mains (sync.go, import.go, facts.go,
// apply.go), each behind an exclusive build tag so `go build ./bin/chats`
// does not see two mains. Shared code lives in internal/chats.
package main
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=chats_facts "$0" "$@"; exit
//go:build chats_facts
//
// bin/chats/facts.go - extract phone/email/linkedin facts from JSONL.
//
// ./bin/chats/facts.go
//
// Writes var/chats/facts/. Does not index the brain.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/chats"
)
func main() {
os.Exit(chats.RunFacts(os.Args[1:]))
}
@@ -1,15 +1,16 @@
package chats
package main
import (
"bufio"
"bytes"
"encoding/json"
"flag"
"fmt"
"os"
"os/exec"
"path/filepath"
"regexp"
"strings"
cliparse "github.com/eSlider/2dph/internal/cli"
)
var (
@@ -76,12 +77,19 @@ type ExtractedFact struct {
MessageID string `json:"message_id"`
}
func RunFacts(args []string) int {
if err := parseNoFlags("chats-facts", args); err != nil {
return cliparse.Fail(err)
func runFacts(args []string) int {
fs := flag.NewFlagSet("chats facts", flag.ContinueOnError)
help := fs.Bool("help", false, "")
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats facts")
return 0
}
root := Dir()
root := chatsDir()
telegramDir := filepath.Join(root, "telegram")
entries, err := os.ReadDir(telegramDir)
@@ -141,7 +149,7 @@ func RunFacts(args []string) int {
}
fmt.Printf("chats facts: saved to %s\n", factsPath)
writeFactsMarkdown(allFacts)
writeFactsToBrain(root, allFacts)
return 0
}
@@ -265,10 +273,14 @@ func filterFacts(facts []ExtractedFact, factType string) []ExtractedFact {
return result
}
// writeFactsMarkdown stores a sidecar for humans. Brain ingest is
// bin/brain/index.go (not this subject).
func writeFactsMarkdown(facts []ExtractedFact) {
mdDir := filepath.Join(Dir(), "facts")
func writeFactsToBrain(root string, facts []ExtractedFact) {
indexScript := filepath.Join(root, "bin", "kb", "index")
if _, err := os.Stat(indexScript); os.IsNotExist(err) {
fmt.Fprintf(os.Stderr, "chats facts: kb/index not found, skipping brain write\n")
return
}
mdDir := filepath.Join(chatsDir(), "facts")
if err := os.MkdirAll(mdDir, 0755); err != nil {
fmt.Fprintf(os.Stderr, "chats facts: mkdir %s: %v\n", mdDir, err)
return
@@ -276,7 +288,7 @@ func writeFactsMarkdown(facts []ExtractedFact) {
var sb strings.Builder
sb.WriteString("---\n")
sb.WriteString("root: info\n")
sb.WriteString("root: facts\n")
sb.WriteString("---\n\n")
sb.WriteString("# Chat-Derived Facts\n\n")
for _, f := range facts {
@@ -290,5 +302,15 @@ func writeFactsMarkdown(facts []ExtractedFact) {
fmt.Fprintf(os.Stderr, "chats facts: write %s: %v\n", factsMD, err)
return
}
fmt.Printf("chats facts: markdown %s (index via brain, not chats)\n", factsMD)
cmd := exec.Command(indexScript, "--corpus", mdDir, "--skip-indexes")
var outBuf, errBuf bytes.Buffer
cmd.Stdout = &outBuf
cmd.Stderr = &errBuf
cmd.Dir = root
if err := cmd.Run(); err != nil {
fmt.Fprintf(os.Stderr, "chats facts: brain index: %v\n%s", err, errBuf.String())
return
}
fmt.Printf("chats facts: written to brain (%s)\n", strings.TrimSpace(outBuf.String()))
}
View File
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=chats_import "$0" "$@"; exit
//go:build chats_import
//
// bin/chats/import.go - JSONL → markdown under var/chats/md/.
//
// ./bin/chats/import.go
//
// Conversion only. Brain ingest is bin/brain/index.go, not this command.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/chats"
)
func main() {
os.Exit(chats.RunImport(os.Args[1:]))
}
@@ -1,25 +1,31 @@
package chats
package main
import (
"bufio"
"bytes"
"encoding/json"
"flag"
"fmt"
"html"
"os"
"path/filepath"
"sort"
"strings"
cliparse "github.com/eSlider/2dph/internal/cli"
)
func RunImport(args []string) int {
if err := parseNoFlags("chats-import", args); err != nil {
return cliparse.Fail(err)
func runImport(args []string) int {
fs := flag.NewFlagSet("chats import", flag.ContinueOnError)
help := fs.Bool("help", false, "")
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats import")
return 0
}
root := Dir()
root := chatsDir()
mdRoot := filepath.Join(root, "md")
glob := filepath.Join(root, "telegram", "*", "messages.jsonl")
+56
View File
@@ -0,0 +1,56 @@
package main
import (
"bytes"
"flag"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
)
func runIndex(args []string) int {
fs := flag.NewFlagSet("chats index", flag.ContinueOnError)
help := fs.Bool("help", false, "")
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats index")
return 0
}
root := repoRoot()
mdDir := filepath.Join(chatsDir(), "md")
_, err := os.Stat(mdDir)
if os.IsNotExist(err) {
fmt.Fprintf(os.Stderr, "chats index: no chat markdown at %s; run 'chats import' first\n", mdDir)
return 1
}
indexScript := filepath.Join(root, "bin", "kb", "index")
if _, err := os.Stat(indexScript); os.IsNotExist(err) {
fmt.Fprintf(os.Stderr, "chats index: %s not found\n", indexScript)
return 1
}
cmd := exec.Command(indexScript, "--corpus", mdDir)
var outBuf, errBuf bytes.Buffer
cmd.Stdout = &outBuf
cmd.Stderr = &errBuf
cmd.Dir = root
if err := cmd.Run(); err != nil {
fmt.Fprintf(os.Stderr, "chats index: %v\n%s", err, errBuf.String())
return 1
}
result := strings.TrimSpace(outBuf.String())
if result == "" {
result = strings.TrimSpace(errBuf.String())
}
fmt.Printf("chats index: %s\n", result)
return 0
}
@@ -1,4 +1,4 @@
package chats
package main
import (
"bufio"
@@ -1,4 +1,4 @@
package chats
package main
import (
"errors"
+115
View File
@@ -0,0 +1,115 @@
// bin/chats - sync, import, index, extract facts, and apply chat data
// from Telegram, WhatsApp, LinkedIn into the brain and OnlyOffice CRM.
//
// Usage:
//
// chats sync telegram [--limit N] [--since DATE] [--phone PHONE]
// chats sync whatsapp [--qr] [--limit N]
// chats sync linkedin [--limit N]
// chats import # JSONL → MD (all sources)
// chats index # rebuild var/kb.lbug with chats
// chats facts # extract + cross-check
// chats apply [--dry-run] # push to OnlyOffice CRM
package main
import (
"fmt"
"os"
"strings"
)
func main() {
if len(os.Args) < 2 {
usage()
os.Exit(2)
}
cmd := os.Args[1]
args := os.Args[2:]
switch cmd {
case "sync":
if len(args) < 1 {
usage()
os.Exit(2)
}
platform := args[0]
platformArgs := args[1:]
switch platform {
case "telegram":
os.Exit(runSyncTelegram(platformArgs))
case "whatsapp":
fmt.Fprintf(os.Stderr, "chats: WhatsApp not implemented yet\n")
os.Exit(1)
case "linkedin":
os.Exit(runSyncLinkedIn(platformArgs))
default:
fmt.Fprintf(os.Stderr, "chats: unknown platform %q\n", platform)
os.Exit(2)
}
case "import":
os.Exit(runImport(args))
case "index":
os.Exit(runIndex(args))
case "facts":
os.Exit(runFacts(args))
case "apply":
os.Exit(runApply(args))
case "help", "-h", "--help":
usage()
return
default:
fmt.Fprintf(os.Stderr, "chats: unknown command %q\n", cmd)
usage()
os.Exit(2)
}
}
func usage() {
w := os.Stderr
fmt.Fprintln(w, `Usage: chats <command> [args]
Commands:
sync telegram [--limit N] [--since DATE] [--phone PHONE]
sync whatsapp [--qr] [--limit N]
sync linkedin [--limit N]
import JSONL → MD (all sources)
index rebuild var/kb.lbug with chats
facts extract + cross-check facts
apply [--dry-run] push to OnlyOffice CRM
Output layout:
var/chats/<platform>/<chat_id>/messages.jsonl
var/chats/md/<platform>/<chat_name>/messages.md`)
}
// repoRoot locates the 2dph project root by walking up from the binary.
func repoRoot() string {
if v := os.Getenv("KB_ROOT"); v != "" {
return v
}
wd, err := os.Getwd()
if err != nil {
return "."
}
for i := 0; i < 10; i++ {
if _, err := os.Stat(wd + "/var"); err == nil {
return wd
}
if _, err := os.Stat(wd + "/.git"); err == nil {
return wd
}
parent := wd
if idx := strings.LastIndex(wd, "/"); idx >= 0 {
parent = wd[:idx]
}
if parent == wd {
break
}
wd = parent
}
return "."
}
// chatsDir returns var/chats under the repo root.
func chatsDir() string {
return repoRoot() + "/var/chats"
}
@@ -1,4 +1,4 @@
package chats
package main
import (
"bufio"
@@ -1,4 +1,4 @@
package chats
package main
import (
"context"
-42
View File
@@ -1,42 +0,0 @@
//usr/bin/env go run -tags=chats_sync "$0" "$@"; exit
//go:build chats_sync
//
// bin/chats/sync.go - download chat messages to var/chats/<platform>/.
//
// ./bin/chats/sync.go telegram [--limit N] [--phone PHONE]
// ./bin/chats/sync.go linkedin [--limit N] [--refresh]
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"github.com/eSlider/2dph/internal/chats"
)
func main() {
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
os.Exit(2)
}
platform := os.Args[1]
args := os.Args[2:]
switch platform {
case "telegram":
os.Exit(chats.RunSyncTelegram(args))
case "linkedin":
os.Exit(chats.RunSyncLinkedIn(args))
case "whatsapp":
fmt.Fprintln(os.Stderr, "chats: WhatsApp sync is out of v1")
os.Exit(1)
case "help", "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]
WhatsApp sync is out of v1.`)
return
default:
fmt.Fprintf(os.Stderr, "chats: unknown platform %q\n", platform)
os.Exit(2)
}
}
@@ -1,29 +1,34 @@
package chats
package main
import (
"context"
"flag"
"fmt"
"os"
"path/filepath"
"strconv"
"strings"
"time"
cliparse "github.com/eSlider/2dph/internal/cli"
)
func RunSyncTelegram(args []string) int {
f, err := parseTelegramFlags(args)
if err != nil {
return cliparse.Fail(err)
func runSyncTelegram(args []string) int {
fs := flag.NewFlagSet("chats sync telegram", flag.ContinueOnError)
limit := fs.Int("limit", 0, "max messages per chat (0 = all)")
phone := fs.String("phone", "", "phone number (default env TELEGRAM_PHONE)")
help := fs.Bool("help", false, "")
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats sync telegram [--limit N] [--phone PHONE]")
return 0
}
limit := f.Limit
phone := f.Phone
apiIDStr := envVar("TELEGRAM_API_ID", "")
apiHash := envVar("TELEGRAM_API_HASH", "")
sessionStr := envVar("TELEGRAM_SESSION_STRING", "")
phoneNum := phone
phoneNum := *phone
if phoneNum == "" {
phoneNum = envVar("TELEGRAM_PHONE", "")
}
@@ -73,7 +78,7 @@ func RunSyncTelegram(args []string) int {
defer cancel()
start := time.Now()
if err := src.Sync(ctx, Dir(), limit); err != nil {
if err := src.Sync(ctx, chatsDir(), *limit); err != nil {
fmt.Fprintf(os.Stderr, "chats sync telegram: %v\n", err)
return 1
}
@@ -1,14 +1,13 @@
package chats
package main
import (
"context"
"flag"
"fmt"
"os"
"os/exec"
"path/filepath"
"time"
cliparse "github.com/eSlider/2dph/internal/cli"
)
func checkLinkedInSession(userDataDir string) (bool, error) {
@@ -29,13 +28,19 @@ func checkLinkedInSession(userDataDir string) (bool, error) {
return false, nil
}
func RunSyncLinkedIn(args []string) int {
f, err := parseLinkedInFlags(args)
if err != nil {
return cliparse.Fail(err)
func runSyncLinkedIn(args []string) int {
fs := flag.NewFlagSet("chats sync linkedin", flag.ContinueOnError)
limit := fs.Int("limit", 0, "max messages per conversation (0 = all)")
refresh := fs.Bool("refresh", false, "refresh session from live webtop browser before sync")
help := fs.Bool("help", false, "")
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return 2
}
if *help {
fmt.Fprintln(os.Stderr, "usage: chats sync linkedin [--limit N] [--refresh]")
return 0
}
limit := f.Limit
refresh := f.Refresh
userDataDir := envVar("LINKEDIN_USER_DATA_DIR", "")
if userDataDir == "" {
@@ -43,7 +48,7 @@ func RunSyncLinkedIn(args []string) int {
userDataDir = home + "/.linkedin-mcp/profile"
}
if refresh {
if *refresh {
if code := refreshLinkedInSession(userDataDir); code != 0 {
return code
}
@@ -67,7 +72,7 @@ func RunSyncLinkedIn(args []string) int {
defer cancel()
start := time.Now()
if err := src.Sync(ctx, Dir(), limit); err != nil {
if err := src.Sync(ctx, chatsDir(), *limit); err != nil {
fmt.Fprintf(os.Stderr, "chats sync linkedin: %v\n", err)
return 1
}
-91
View File
@@ -1,91 +0,0 @@
//usr/bin/env go run "$0" "$@"; exit
//
// bin/cli/complete.go - dump flaggy shell completions for all Go shebang tools (D23).
//
// source <(./bin/cli/complete.go bash)
// ./bin/cli/complete.go zsh|fish|powershell|nushell
//
// Search does not steal the word "completion"; this binary dumps scripts.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"strings"
mailsync "github.com/eSlider/2dph/bin/mail/sync"
"github.com/eSlider/2dph/internal/brain/rank"
"github.com/eSlider/2dph/internal/chats"
"github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/gitlog"
"github.com/eSlider/2dph/internal/mdleaves"
"github.com/eSlider/2dph/internal/ocr"
"github.com/eSlider/2dph/internal/reasoner"
"github.com/eSlider/2dph/internal/websearch"
"github.com/integrii/flaggy"
)
func tools() []cli.Tool {
return []cli.Tool{
{Path: "bin/brain/search.go", Name: "brain-search", New: rank.Parser},
{Path: "bin/brain/get.go", Name: "brain-get", New: func() *flaggy.Parser {
o := rank.GetOptions{}
return rank.GetParser(&o)
}},
{Path: "bin/brain/stats.go", Name: "brain-stats", New: rank.StatsParser},
{Path: "bin/brain/eval.go", Name: "brain-eval", New: rank.EvalParser},
{Path: "bin/web/search.go", Name: "web-search", New: websearch.Parser},
{Path: "bin/git/import.go", Name: "git-import", New: gitlog.Parser},
{Path: "bin/markdown/import.go", Name: "markdown-import", New: mdleaves.Parser},
{Path: "bin/qa/stats.go", Name: "qa-stats", New: cli.QAParser},
{Path: "bin/reasoner/bakeoff.go", Name: "reasoner-bakeoff", New: reasoner.Parser},
{Path: "bin/mail/ocr.go", Name: "mail-ocr", New: ocr.Parser},
{Path: "bin/mail/sync.go", Name: "mail-sync", New: mailsync.Parser},
{Path: "bin/chats/sync.go", Name: "chats-sync", New: chats.SyncParser},
{Path: "bin/chats/import.go", Name: "chats-import", New: chats.ImportParser},
{Path: "bin/chats/facts.go", Name: "chats-facts", New: chats.FactsParser},
{Path: "bin/chats/apply.go", Name: "chats-apply", New: chats.ApplyParser},
}
}
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
shell := "bash"
if len(args) > 0 {
switch args[0] {
case "bash", "zsh", "fish", "powershell", "nushell":
shell = args[0]
case "-h", "--help", "help":
fmt.Fprintln(os.Stderr, "usage: bin/cli/complete.go [bash|zsh|fish|powershell|nushell]")
return 0
default:
fmt.Fprintf(os.Stderr, "cli/complete: unknown shell %q\n", args[0])
return 2
}
}
if shell == "bash" {
fmt.Print(cli.BashScript(tools()))
return 0
}
var b strings.Builder
for _, t := range tools() {
p := t.New()
p.Name = t.Name
switch shell {
case "zsh":
b.WriteString(flaggy.GenerateZshCompletion(p))
case "fish":
b.WriteString(flaggy.GenerateFishCompletion(p))
case "powershell":
b.WriteString(flaggy.GeneratePowerShellCompletion(p))
case "nushell":
b.WriteString(flaggy.GenerateNushellCompletion(p))
}
}
fmt.Print(b.String())
return 0
}
+8 -19
View File
@@ -1,10 +1,13 @@
#!/usr/bin/env bash
# bin/docker-entrypoint - run 2dph tools inside the container.
#
# API image (Zig CGO binaries):
# serve | search | watch
# Index image (Python write path, compose profile `index`):
# index | extract | audit | search (deprecated python wrapper)
# brain shell (default)
# brain search <q> bin/kb/search
# brain index bin/kb/index
# brain watch <dir> watchdog re-indexer (bin/kb/watch)
# brain serve async Go HTTP server (bin/serve)
# brain extract bin/facts/extract (docker×compose pairing)
# brain audit bin/facts/audit
#
# Usage comment starts at line 2 (self-describing convention).
set -euo pipefail
@@ -12,24 +15,10 @@ set -euo pipefail
CMD="${1:-shell}"
shift || true
if [ -x /usr/local/bin/brain-serve ]; then
case "$CMD" in
shell) exec bash ;;
serve) exec /usr/local/bin/brain-serve "$@" ;;
search) exec /usr/local/bin/brain-search "$@" ;;
watch) exec /usr/local/bin/brain-watch "$@" ;;
index)
echo "index is the Python sidecar: docker compose --profile index run --rm index" >&2
exit 2
;;
*) echo "unknown command: $CMD (api: serve|search|watch)" >&2; exit 2 ;;
esac
fi
case "$CMD" in
shell) exec bash ;;
search) exec "$KB_PY" /app/bin/kb/search "$@" ;;
index) exec "$KB_PY" /app/bin/kb/index --with-mail "$@" ;;
index) exec "$KB_PY" /app/bin/kb/index "$@" ;;
watch) exec /app/bin/watch "$@" ;;
serve) exec /app/bin/serve "$@" ;;
extract) exec "$KB_PY" /app/bin/facts/extract "$@" ;;
+18 -45
View File
@@ -1,15 +1,14 @@
#!/usr/bin/env python3
"""facts/audit - evidence & lexicon checks for the 2dph brain.
bin/facts/audit self # lexicon: docs + two-source rule
bin/facts/audit db # evidence gate against var/kb.lbug
bin/facts/audit contradict # D16 adjudication (JSON claim(s) on stdin)
bin/facts/audit self # lexicon: every fact in db has >=2 sources
bin/facts/audit db # evidence gate: run against var/kb.lbug
`self` mode checks the repo itself (no network, no runtime deps).
`db` mode loads every Leaf with root=facts. Confirmed facts need ` x `;
hypothesis contradictions need `a x b vs c x d` (both sides ≥2).
`contradict` applies temporal_freshness then authority_pairing; ≥2 vs ≥2
with no rule stays hypothesis / `(not confirmed)`.
`self` mode checks the repo itself (no network, no runtime deps). It greps
for known-good two-source pairings and confirms the docs are consistent.
`db` mode loads every Leaf with root=facts and asserts each has source_rev
and a non-empty `loc` (the "where did you see it" evidence pointer) and that
'confirmed' facts carry a two-source `source` field.
Exit 0 = all checks pass, 1 = audit failures, 2 = could not evaluate.
"""
@@ -23,8 +22,6 @@ from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from contradict import adjudicate, check_fact_row # noqa: E402
def audit_db() -> list[str]:
from kblib import connect
@@ -36,8 +33,14 @@ def audit_db() -> list[str]:
r = conn.execute("MATCH (l:Leaf {root:'facts'}) RETURN l.id, l.source, l.loc, l.how, l.confidence")
problems: list[str] = []
for lid, source, loc, how, conf in r.get_all():
problems.extend(check_fact_row(str(lid), str(source or ""), str(loc or ""),
str(how or ""), str(conf or "")))
if conf != "confirmed":
problems.append(f"{lid}: facts require confidence='confirmed', got '{conf}'")
if not source or " x " not in source:
problems.append(f"{lid}: needs 2-source evidence in source, got '{source}'")
if not loc:
problems.append(f"{lid}: missing loc (evidence pointer)")
if not how:
problems.append(f"{lid}: missing how")
conn.close()
db.close()
return problems
@@ -51,50 +54,20 @@ def audit_self() -> list[str]:
problems.append("PLAN.md missing recall@5 gate")
if re.search(r"(?i)facts must have.*2 sources|2.source", plan) is None:
problems.append("PLAN.md missing the two-source evidence rule for facts")
if "temporal_freshness" not in plan or "authority_pairing" not in plan:
problems.append("PLAN.md missing D16 adjudication rules")
if re.search(r"(?i)HNSW|BM25|deduction", (ROOT / "README.md").read_text()) is None:
problems.append("README.md missing search/retrieval description")
return problems
def audit_contradict(raw: str) -> tuple[list[str], list[dict]]:
raw = raw.strip()
if not raw:
return ["contradict: empty stdin (JSON claim or {claims:[...]})"], []
try:
payload = json.loads(raw)
except json.JSONDecodeError as e:
return [f"contradict: invalid JSON: {e}"], []
if isinstance(payload, dict) and "claims" in payload:
claims = list(payload.get("claims") or [])
elif isinstance(payload, dict):
claims = [payload]
elif isinstance(payload, list):
claims = payload
else:
return ["contradict: expected object or list"], []
details = [adjudicate(c) for c in claims]
return [], details
def main(argv: list[str]) -> int:
import argparse
p = argparse.ArgumentParser(description="evidence & lexicon audit")
p.add_argument("mode", choices=("self", "db", "contradict"))
p.add_argument("mode", choices=("self", "db"))
p.add_argument("--json", action="store_true")
a = p.parse_args(argv)
details: list[dict] = []
if a.mode == "self":
problems = audit_self()
elif a.mode == "db":
problems = audit_db()
else:
problems, details = audit_contradict(sys.stdin.read())
out: dict = {"mode": a.mode, "ok": not problems, "problems": problems}
if details:
out["contradictions"] = details
problems = audit_self() if a.mode == "self" else audit_db()
out = {"mode": a.mode, "ok": not problems, "problems": problems}
if a.json:
print(json.dumps(out, indent=2))
else:
-22
View File
@@ -1,22 +0,0 @@
//usr/bin/env go run -tags=facts_audit "$0" "$@"; exit
//go:build facts_audit
//
// bin/facts/audit.go - 2-source + lexicon checks.
//
// ./bin/facts/audit.go self
// ./bin/facts/audit.go db
// ./bin/facts/audit.go contradict --json < claim.json
//
// Python bin/facts/audit is the implementation (CI runs it directly).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/audit", os.Args[1:]))
}
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=facts_crm "$0" "$@"; exit
//go:build facts_crm
//
// bin/facts/crm.go - prove person↔company / company↔project (ooCRM × corpus).
//
// ./bin/facts/crm.go [--dry-run] [--mismatches]
//
// Python bin/facts/crm is the implementation. Graph write stays Python.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/crm", os.Args[1:]))
}
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=facts_extract "$0" "$@"; exit
//go:build facts_extract
//
// bin/facts/extract.go - acquire confirmed facts (2-source each).
//
// ./bin/facts/extract.go [--json] [--dry-run]
//
// Python bin/facts/extract is the implementation. Graph write stays Python.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/extract", os.Args[1:]))
}
+138 -10
View File
@@ -1,25 +1,153 @@
#!/usr/bin/env python3
"""git/import — deprecated. Use bin/git/import.go (go-git, no git binary).
"""git/import - import git history (commits, authors, files) into the brain.
bin/git/import.go [REPO] [--json] [--limit N] [--since DATE]
bin/git/import [REPO] import all commits -> leafs + graph
bin/git/import --json emit import leafs as JSON, no write
bin/git/import --limit 100 cap commits processed
bin/git/import --since 2026-01-01 only recent commits
bin/git/import --root DIR run per repo dir under DIR
bin/git/import --no-env never read .env anywhere (default: true)
Reads `git log --no-merges --name-only` from the repo, maps commits to
`info` leafs (root=info, type=commit) and writes the version graph
`File -[:HAS_VERSION]-> Commit -[:AUTHORED]-> Person` into var/kb.lbug.
Idempotent: leaf MERGE by (source,text via leaf_id), graph MERGE by sha.
"""
from __future__ import annotations
import os
import json
import subprocess
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import ( # noqa: E402
connect, ensure_indexes, init_schema, upsert_leaf,
)
from gitimport import commits_to_leafs, ensure_git_schema, index_commits, parse_log # noqa: E402
LOG_FMT = "--format=%x1e%H%x1f%an%x1f%ae%x1f%aI%x1f%s"
def git_log(repo: Path, limit: int = 0, since: str = "") -> str:
cmd = ["git", "-C", str(repo), "log", "--no-merges", "--name-only", LOG_FMT]
if since:
cmd += ["--since", since]
if limit:
cmd += ["-n", str(limit)]
try:
out = subprocess.run(cmd, capture_output=True, text=True, timeout=120)
except (FileNotFoundError, subprocess.TimeoutExpired):
return ""
if out.returncode != 0:
print(f"git/import: {repo}: {out.stderr.strip()}", file=sys.stderr)
return ""
return out.stdout
def repo_name(repo: Path) -> str:
try:
out = subprocess.run(
["git", "-C", str(repo), "remote", "get-url", "origin"],
capture_output=True, text=True, timeout=20)
url = out.stdout.strip()
return url.rstrip("/").split("/")[-1].removesuffix(".git") if url else repo.name
except (FileNotFoundError, subprocess.TimeoutExpired):
return repo.name
def embedder():
from model2vec import StaticModel
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
return lambda text: model.encode([text])[0].astype(float).tolist()
def import_repo(conn, repo: Path, embed, limit: int, since: str,
no_write: bool = False) -> tuple[int, int]:
raw = git_log(repo, limit, since)
commits = parse_log(raw)
leafs = commits_to_leafs(commits, repo_name(repo))
if no_write:
return len(commits), 0
written = 0
for lf in leafs:
query = f"{lf['heading']}\n\n{lf['text']}"
emb = embed(lf["text"]) if lf["text"] else None
upsert_leaf(conn, text=query, root="info", confidence="confirmed",
source=lf["source"], source_rev="git", how="git/import",
loc=lf["source"], type_=lf.get("type", "commit"),
embedding=emb)
written += 1
index_commits(conn, commits, repo_name(repo))
return len(commits), written
def main(argv: list[str]) -> int:
print(
"bin/git/import is deprecated; use bin/git/import.go (go-git)",
file=sys.stderr,
)
target = ROOT / "bin" / "git" / "import.go"
os.execvp("go", ["go", "run", str(target), *argv])
return 1
import argparse
p = argparse.ArgumentParser(description="import git history into the brain")
p.add_argument("repo", nargs="?", default=None)
p.add_argument("--root", default=None, help="directory of repos to import (each git dir separately)")
p.add_argument("--limit", type=int, default=0)
p.add_argument("--since", default="")
p.add_argument("--json", action="store_true")
p.add_argument("--dry-run", action="store_true", help="parse + report, no db write")
a = p.parse_args(argv)
repos: list[Path] = []
if a.repo:
repos = [Path(a.repo)]
elif a.root:
root = Path(a.root)
if root.is_file():
repos = [root]
else:
repos = [dp for dp in sorted(root.iterdir()) if (dp / ".git").exists() or dp.is_file()]
else:
repos = [ROOT]
total_commits = 0
results: list[dict] = []
if a.dry_run:
for repo in repos:
if not repo.exists():
continue
commits = parse_log(git_log(repo, a.limit, a.since))
name = repo_name(repo)
total_commits += len(commits)
results.append({"repo": name, "commits": len(commits),
"leafs": len(commits_to_leafs(commits, name)), "path": str(repo)})
if a.json:
print(json.dumps(results, indent=2))
else:
for r in results:
print(f"{r['repo']:<24} {r['commits']:>5} commits -> {r['leafs']} leafs {r['path']}")
return 0
# Never DROP FTS/VECTOR (ghost catalog). Upsert while indexes exist is OK;
# ensure_indexes only CREATEs when missing.
db, conn = connect(ROOT / "var" / "kb.lbug", read_only=False)
init_schema(conn)
embed = embedder()
rows: list[dict] = []
for repo in repos:
if not repo.exists():
continue
reached, written = import_repo(conn, repo, embed, a.limit, a.since)
total_commits += reached
rows.append({"repo": repo_name(repo), "commits": reached, "written": written})
ensure_indexes(conn)
conn.close()
db.close()
if a.json:
print(json.dumps(rows, indent=2))
else:
for r in rows:
print(f"imported {r['commits']:>5} commits -> {r['written']} leafs {r['repo']}")
print(f"total: {total_commits} commits")
return 0
if __name__ == "__main__":
-100
View File
@@ -1,100 +0,0 @@
//usr/bin/env go run "$0" "$@"; exit
//
// bin/git/import.go - read git history with go-git (no git binary).
//
// ./bin/git/import.go [REPO]
// ./bin/git/import.go --json
// ./bin/git/import.go --limit 100 --since 2026-01-01
// ./bin/git/import.go --root DIR
//
// Conversion only: prints commit leafs. Brain write is bin/brain/index.go.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"encoding/json"
"fmt"
"os"
"path/filepath"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/cmdbin"
"github.com/eSlider/2dph/internal/gitlog"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
c, err := gitlog.ParseArgs(args)
if err != nil {
return cliparse.Fail(err)
}
repo, root, since, limit, jsonOut := c.Repo, c.Root, c.Since, c.Limit, c.JSONOut
sinceT, err := gitlog.ParseSince(since)
if err != nil {
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
return 2
}
repos := []string{}
if repo != "" {
repos = []string{repo}
} else if root != "" {
entries, err := os.ReadDir(root)
if err != nil {
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
return 1
}
for _, e := range entries {
p := filepath.Join(root, e.Name())
if _, err := os.Stat(filepath.Join(p, ".git")); err == nil {
repos = append(repos, p)
}
}
} else {
repos = []string{cmdbin.Root()}
}
opt := gitlog.Options{Limit: limit, Since: sinceT}
type row struct {
Repo string `json:"repo"`
Path string `json:"path"`
Commits int `json:"commits"`
Leafs []gitlog.Leaf `json:"leafs,omitempty"`
}
var rows []row
for _, p := range repos {
name, err := gitlog.RepoName(p)
if err != nil && name == "" {
fmt.Fprintf(os.Stderr, "git/import: %s: %v\n", p, err)
continue
}
cs, err := gitlog.Log(p, opt)
if err != nil {
fmt.Fprintf(os.Stderr, "git/import: %s: %v\n", p, err)
return 1
}
leafs := make([]gitlog.Leaf, 0, len(cs))
for _, c := range cs {
leafs = append(leafs, gitlog.ToLeaf(c, name))
}
rows = append(rows, row{Repo: name, Path: p, Commits: len(cs), Leafs: leafs})
}
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
if err := enc.Encode(rows); err != nil {
return 1
}
return 0
}
for _, r := range rows {
fmt.Printf("%-24s %5d commits %s\n", r.Repo, r.Commits, r.Path)
}
return 0
}
-120
View File
@@ -1,120 +0,0 @@
#!/usr/bin/env python3
"""kb/add - incremental leaf write (no rebuild).
bin/kb/add --text T --root facts|info --source S
bin/kb/add --json # stdin: one object or {"leafs":[...]}
bin/kb/add --db PATH --json
Writes facts+info in one Ladybug transaction. Does not delete kb.lbug.
Embedding is used when provided; otherwise model2vec encodes the text.
"""
from __future__ import annotations
import json
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import ( # noqa: E402
EMBED_DIM,
add_leafs,
connect,
ensure_indexes,
init_schema,
)
def _as_leafs(payload: object) -> list[dict]:
if isinstance(payload, list):
return [dict(x) for x in payload]
if isinstance(payload, dict):
if "leafs" in payload:
return [dict(x) for x in payload["leafs"]]
return [dict(payload)]
raise ValueError("json must be an object, a list, or {leafs:[...]}")
def _embed_missing(leafs: list[dict]) -> None:
missing = [lf for lf in leafs if not lf.get("embedding")]
if not missing:
return
from model2vec import StaticModel
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
for lf in missing:
text = str(lf.get("text") or "")
vec = model.encode([text])[0].astype(float).tolist()
if len(vec) != EMBED_DIM:
vec = (vec + [0.0] * EMBED_DIM)[:EMBED_DIM]
lf["embedding"] = vec
def main(argv: list[str]) -> int:
import argparse
p = argparse.ArgumentParser(description="add leafs without rebuilding the brain")
p.add_argument("--db", default="", help="path to kb.lbug (default var/kb.lbug)")
p.add_argument("--json", action="store_true", help="read leaf JSON from stdin")
p.add_argument("--text", default="", help="leaf text")
p.add_argument("--root", default="info", choices=("facts", "info"))
p.add_argument("--source", default="")
p.add_argument("--confidence", default="confirmed")
p.add_argument("--source-rev", default="working-tree")
p.add_argument("--how", default="brain/add")
p.add_argument("--loc", default="")
p.add_argument("--type", default="reference", dest="type_")
p.add_argument("--valid-from", default="", dest="valid_from",
help="fact interval start YYYY-MM-DD (D24)")
p.add_argument("--valid-to", default="", dest="valid_to",
help="fact interval end YYYY-MM-DD inclusive; empty=open (D24)")
args = p.parse_args(argv)
if args.json:
raw = sys.stdin.read()
if not raw.strip():
print("kb/add: empty stdin", file=sys.stderr)
return 2
leafs = _as_leafs(json.loads(raw))
else:
if not args.text or not args.source:
print("kb/add: --text and --source are required (or --json)", file=sys.stderr)
return 2
leafs = [{
"text": args.text,
"root": args.root,
"source": args.source,
"confidence": args.confidence,
"source_rev": args.source_rev,
"how": args.how,
"loc": args.loc or args.source,
"type": args.type_,
"valid_from": args.valid_from,
"valid_to": args.valid_to,
}]
for lf in leafs:
if not lf.get("text") or not lf.get("source"):
print("kb/add: each leaf needs text and source", file=sys.stderr)
return 2
_embed_missing(leafs)
from kblib import DB_PATH, VAR
dbpath = Path(args.db) if args.db else DB_PATH
dbpath.parent.mkdir(parents=True, exist_ok=True)
VAR.mkdir(exist_ok=True)
db, conn = connect(dbpath, read_only=False)
init_schema(conn)
ids = add_leafs(conn, leafs)
ensure_indexes(conn)
conn.close()
db.close()
print(json.dumps({"mode": "add", "ids": ids, "db": str(dbpath)}))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+13 -113
View File
@@ -2,15 +2,12 @@
"""kb/index - build the 2dph brain from markdown + factual leafs.
bin/kb/index [--corpus DIR] [--rebuild] [--limit N]
bin/kb/index --rebuild --with-facts --with-chats
bin/kb/index --json # emit stats as JSON
Reads every .md under the corpus (default: repo root docs, skills, READMEs)
as `info` leafs, embeds them with model2vec (potion-multilingual-128M), and
writes them into var/kb.lbug with FTS + HNSW indexes. `facts` leafs come
from bin/facts/extract (docker × compose × ssh-config pairing) when
`--with-facts` is set. `--with-chats` indexes markdown under var/chats/md
(or a given dir) as info. WhatsApp sync stays out of v1.
from bin/facts/extract (docker x compose x ssh-config pairing).
--rebuild drops the database file and indexes from scratch. Without it a run
is idempotent (MERGE by (source,text) id).
@@ -25,11 +22,10 @@ ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import ( # noqa: E402
add_leafs, connect, ensure_indexes, init_schema, upsert_leaf, link_from_file,
connect, ensure_indexes, init_schema, upsert_leaf,
open_readonly, stats,
)
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
from mailleafs import from_mail_root # noqa: E402
CORPUS_DEFAULTS = ["README.md", "PLAN.md", "AGENTS.md", "docs", "skills"]
@@ -84,11 +80,10 @@ def index_leafs(conn, leafs: list[dict], embed_fn, limit: int) -> tuple[int, int
for lf in leafs[:limit] if limit else leafs:
query = f"{lf['heading']}\n\n{lf['text']}"
emb = embed_fn(lf["text"]) if lf["text"] else None
lid = upsert_leaf(conn, text=query, root="info", confidence="confirmed",
upsert_leaf(conn, text=query, root="info", confidence="confirmed",
source=lf["source"], source_rev="working-tree",
how="kb/index", loc=lf["source"], type_=lf.get("type", "reference"),
embedding=emb)
link_from_file(conn, lid, lf["source"], repo=str(lf.get("repo") or ""))
count += 1
return count, len(leafs)
@@ -99,67 +94,11 @@ def embedder():
return lambda text: model.encode([text])[0].astype(float).tolist()
def index_fact_dicts(conn, facts: list[dict], embed_fn) -> int:
"""Write extract-shaped dicts as root=facts leafs (2-source source field)."""
leafs = []
for f in facts:
text = str(f.get("text") or "")
source = str(f.get("source") or "")
if not text or not source:
continue
leafs.append({
"text": text,
"root": "facts",
"confidence": "confirmed",
"source": source,
"source_rev": f.get("source_rev") or "working-tree",
"how": f.get("how") or "facts/extract",
"loc": f.get("loc") or source,
"type": "fact",
"embedding": embed_fn(text) if text else None,
})
return len(add_leafs(conn, leafs))
def facts_from_extract() -> list[dict]:
import subprocess
proc = subprocess.run(
[sys.executable, str(ROOT / "bin" / "facts" / "extract"), "--json", "--dry-run"],
cwd=ROOT,
capture_output=True,
text=True,
check=False,
)
if proc.returncode != 0:
print(f"kb/index: facts/extract failed: {proc.stderr}", file=sys.stderr)
return []
try:
payload = json.loads(proc.stdout)
except json.JSONDecodeError:
print("kb/index: facts/extract produced non-JSON", file=sys.stderr)
return []
return list(payload.get("facts") or [])
def main(argv: list[str]) -> int:
import argparse
p = argparse.ArgumentParser(description="build the 2dph brain index")
p.add_argument("--corpus", action="append", help="extra markdown dir/file to index (may repeat)")
p.add_argument("--rebuild", action="store_true", help="fresh db + indexes")
p.add_argument("--db", default="", help="path to kb.lbug (default var/kb.lbug)")
p.add_argument("--no-defaults", action="store_true", help="do not index repo README/docs/skills")
p.add_argument("--with-mail", action="store_true", help="include var/mail message.md leafs")
p.add_argument("--with-facts", action="store_true", help="run facts/extract into root=facts")
p.add_argument("--facts-json", default="", help="JSON list (or {facts:[...]}) of fact dicts")
p.add_argument(
"--with-chats",
nargs="?",
const=str(ROOT / "var" / "chats" / "md"),
default="",
help="index chat markdown as info (default var/chats/md)",
)
p.add_argument("--since", default="", help="with --with-mail, only messages dated >= YYYY-MM-DD")
p.add_argument("--dry-run", action="store_true", help="count leafs, write nothing")
p.add_argument(
"--skip-indexes",
action="store_true",
@@ -170,72 +109,33 @@ def main(argv: list[str]) -> int:
a = p.parse_args(argv)
from kblib import DB_PATH, VAR
VAR.mkdir(exist_ok=True)
if a.rebuild and DB_PATH.exists():
DB_PATH.unlink()
dbpath = Path(a.db) if a.db else DB_PATH
leafs: list[dict] = [] if a.no_defaults else load_corpus(ROOT)
leafs = load_corpus(ROOT)
if a.corpus:
for source in a.corpus:
leafs.extend(load_corpus_glob(source))
chat_n = 0
if a.with_chats:
chats = load_corpus_glob(a.with_chats)
chat_n = len(chats)
leafs.extend(chats)
mail_n = 0
if a.with_mail:
mail = from_mail_root(ROOT / "var" / "mail", since=a.since)
mail_n = len(mail)
leafs.extend(mail)
facts: list[dict] = []
if a.facts_json:
raw = Path(a.facts_json).read_text(encoding="utf-8")
payload = json.loads(raw)
facts = list(payload.get("facts") if isinstance(payload, dict) else payload)
if a.with_facts:
facts.extend(facts_from_extract())
if a.dry_run:
msg = {
"indexed": 0,
"corpus_total": len(leafs),
"mail_leafs": mail_n,
"chat_leafs": chat_n,
"facts_leafs": len(facts),
"dry_run": True,
}
print(json.dumps(msg, indent=2) if a.json else
f"brain/index: {len(leafs)} info + {len(facts)} facts would be indexed")
return 0
VAR.mkdir(exist_ok=True)
dbpath.parent.mkdir(parents=True, exist_ok=True)
if a.rebuild and dbpath.exists():
dbpath.unlink()
db, conn = connect(dbpath, read_only=False)
db, conn = connect(DB_PATH, read_only=False)
init_schema(conn)
# Never DROP FTS/VECTOR (ghost catalog). Write leafs, then ensure indexes
# unless --skip-indexes (seed facts first — MERGE under live FTS corrupts it).
# --rebuild already deleted kb.lbug above, so CREATE runs on a clean DB.
embed = embedder()
done, total = index_leafs(conn, leafs, embed, a.limit)
fact_n = index_fact_dicts(conn, facts, embed) if facts else 0
if not a.skip_indexes:
ensure_indexes(conn)
s = stats(conn)
conn.close()
db.close()
result = {
"indexed": done,
"corpus_total": total,
"facts_leafs": fact_n,
"chat_leafs": chat_n,
**{k: v for k, v in s.items() if k in ("total", "by_root")},
}
result = {"indexed": done, "corpus_total": total, **{k: v for k, v in s.items() if k in ("total", "by_root")}}
if a.skip_indexes:
result["indexes"] = "skipped"
print(json.dumps(result, indent=2) if a.json else
f"indexed {done}/{total} info + {fact_n} facts; db total {s['total']}")
print(json.dumps(result, indent=2) if a.json else f"indexed {done}/{total} leafs; db total {s['total']}")
return 0
+6 -4
View File
@@ -1,9 +1,10 @@
#!/usr/bin/env bash
# bin/kb/search — deprecated wrapper. Use bin/brain/search.go.
# CGO via Zig (bin/cgo/zig), not gcc. Builds a binary then execs it.
# Sets CGO for ladybug, builds a binary (embed daemon needs a real executable),
# then execs it. Prints one deprecation line.
set -euo pipefail
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
BIN="$ROOT/var/bin/brain-search"
SRC="$ROOT/internal/brain"
CMD="$ROOT/bin/brain"
@@ -23,10 +24,11 @@ else
fi
if [ "$need_build" -eq 1 ]; then
echo "Building brain/search (zig cc)..." >&2
echo "Building brain/search..." >&2
(
cd "$ROOT" &&
eval "$("$ROOT/bin/cgo/zig" env)" &&
CGO_CFLAGS="-I$ROOT/lib-ladybug" \
CGO_LDFLAGS="-L$ROOT/lib-ladybug -Wl,-rpath,$ROOT/lib-ladybug" \
go build -tags system_ladybug -o "$BIN" ./bin/brain
) || exit 1
fi
+1 -1
View File
@@ -1,5 +1,5 @@
//usr/bin/env go run "$0" "$@"; exit
// bin/kb/watch.go — deprecated. Use bin/brain/watch.go.
// bin/kb/watch.go - re-index the 2dph brain when corpus files change.
//
// Usage:
//
+78 -12
View File
@@ -7,7 +7,7 @@
bin/mail/import --since 2026-01-01 only messages after a date
bin/mail/import --limit 50 cap messages per run
bin/mail/import --no-attachments body only, skip attachment conversion
bin/mail/import --ocr OCR images (PDFs OCR when textless)
bin/mail/import --ocr OCR scanned PDFs/images via docling
bin/mail/import --dry-run list messages without writing anything
Writes one directory per message: var/mail/{folder}/{message_id}/
@@ -15,12 +15,11 @@ Writes one directory per message: var/mail/{folder}/{message_id}/
attachments/ raw attachment files (zips unpacked to _unpacked/)
attachments/*.md converted attachment content
Indexing is a separate step (`bin/brain/index.go --rebuild`): conversion can
crash and must not leave the brain DB mid-transaction.
Indexing is a separate step (bin/mail/index_mail): conversion can crash in
native docling and must not leave the brain DB mid-transaction.
Requires ONLYOFFICE_URL/USER/PASS in .env (or env) except `--from-raw`.
Idempotent: a message already present (message.md exists) is skipped unless
--force.
Requires ONLYOFFICE_URL/USER/PASS in .env (or env). Idempotent: a message
already present (message.md exists) is skipped unless --force.
"""
from __future__ import annotations
@@ -42,10 +41,10 @@ from mailconv import ( # noqa: E402
IMAGE_SUFFIXES,
LEGACY_OFFICE_SUFFIXES,
TEXT_SUFFIXES,
convert_pdf,
html_to_markdown,
is_convertible,
normalize_markdown,
ocr_image,
subject_to_filename,
zip_extract_safe,
)
@@ -148,9 +147,9 @@ def convert_file_to_md(path: Path, ocr: bool) -> str | None:
except Exception as e:
return f"\n<!-- conversion failed: {e} -->\n"
if suffix == ".pdf":
return convert_pdf(path, ocr)
return _convert_pdf(path, ocr)
if suffix in IMAGE_SUFFIXES and ocr:
return ocr_image(path) or "\n<!-- ocr unavailable -->\n"
return _convert_pdf(path, ocr)
if suffix in LEGACY_OFFICE_SUFFIXES:
return _convert_legacy(path)
if suffix in ARCHIVE_SUFFIXES:
@@ -158,6 +157,67 @@ def convert_file_to_md(path: Path, ocr: bool) -> str | None:
return None
def _convert_pdf(path: Path, ocr: bool) -> str:
"""Convert one PDF to markdown.
Fast path: poppler's pdftotext (-layout) extracts exact text from
born-digital PDFs in ~15ms vs docling's 1-3s. Only textless PDFs (scanned
pages, layout-heavy) fall back to docling, which runs isolated in a
subprocess because its native onnx/RT-DETR has segfaulted the main process.
"""
text = _pdf_fast_text(path)
if ocr or text is None or not text.strip():
return _convert_pdf_docling(path, ocr)
return normalize_markdown(text)
def _pdf_fast_text(path: Path) -> str | None:
"""pdftotext -layout; None when poppler is unavailable (or the PDF has no text layer)."""
try:
proc = subprocess.run(
["pdftotext", "-layout", str(path), "-"],
capture_output=True, timeout=60)
except (OSError, subprocess.TimeoutExpired):
return None
if proc.returncode != 0:
return None
return proc.stdout.decode("utf-8", errors="replace")
def _convert_pdf_docling(path: Path, ocr: bool) -> str:
try:
proc = subprocess.run(
[sys.executable, os.path.abspath(__file__), "--pdf-worker", str(path),
"--ocr" if ocr else "--no-ocr"],
capture_output=True, text=True, timeout=600)
except subprocess.TimeoutExpired:
return "\n<!-- pdf conversion timed out -->\n"
if proc.returncode != 0:
tail = proc.stderr.strip().splitlines()[-3:]
return f"\n<!-- pdf conversion failed: {proc.returncode}: {' | '.join(tail)} -->\n"
return proc.stdout
def _pdf_worker(path: Path, ocr: bool) -> None:
"""docling worker entry: prints converted markdown on stdout, exits non-zero on error."""
try:
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.pipeline_options import PdfPipelineOptions
opts = PdfPipelineOptions()
opts.do_ocr = bool(ocr)
opts.do_table_structure = True
conv = DocumentConverter(format_options={"pdf": PdfFormatOption(pipeline_options=opts)})
res = conv.convert(str(path))
sys.stdout.write(normalize_markdown(res.document.export_to_markdown()))
sys.exit(0)
except Exception as e:
# errors/stacktraces to stderr; the caller only reports a one-liner
print(f"pdf-worker: {e}", file=sys.stderr)
import traceback
traceback.print_exc(file=sys.stderr)
sys.exit(1)
def _convert_legacy(path: Path) -> str:
"""Legacy .doc/.xls/.ppt -> md via pandoc (installed) or a stub."""
try:
@@ -296,12 +356,19 @@ def main(argv: list[str]) -> int:
p.add_argument("--from-raw", default="",
help="convert Go-synced dirs (var/mail/<folder>/<id>/message.json) to markdown")
p.add_argument("--no-attachments", action="store_true", help="skip attachment download+convert")
p.add_argument("--ocr", action="store_true", help="OCR images (PDFs OCR when textless)")
p.add_argument("--ocr", action="store_true", help="OCR scanned PDFs/images via docling")
p.add_argument("--force", action="store_true", help="re-import even if message.md exists")
p.add_argument("--dry-run", action="store_true", help="list messages, write nothing")
p.add_argument("--json", action="store_true")
p.add_argument("--pdf-worker", default="", help=argparse.SUPPRESS)
p.add_argument("--no-ocr", action="store_true", help=argparse.SUPPRESS)
a = p.parse_args(argv)
if a.pdf_worker:
_pdf_worker(Path(a.pdf_worker), ocr=not a.no_ocr)
return 0
conf = load_env()
fid = folder_id(a.folder)
out_root = ROOT / "var" / "mail"
summary: list[dict] = []
@@ -327,7 +394,6 @@ def main(argv: list[str]) -> int:
target_dir=msg_dir.parent))
summary.append(entry)
else:
conf = load_env()
OOCLIENT = OOClient(conf)
if a.id:
messages = [{"id": i} for i in a.id]
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=mail_import "$0" "$@"; exit
//go:build mail_import
//
// bin/mail/import.go - message.json → markdown (no brain write).
//
// ./bin/mail/import.go --from-raw var/mail
//
// Indexing is bin/brain/index.go --rebuild, not this command.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/mail/import", os.Args[1:]))
}
+120 -11
View File
@@ -1,26 +1,135 @@
#!/usr/bin/env python3
"""mail/index_mail — deprecated. Use bin/brain/index.go --rebuild --with-mail.
"""mail/index_mail - rebuild the brain with every markdown under var/mail.
Ladybug corrupts its WAL on bulk-insert into an already-indexed DB, so this
shim always rebuilds (repo corpus + var/mail). Conversion stays in mail/import.
Ladybug corrupts its WAL when brand-new leafs are bulk-inserted while the
FTS/VECTOR indexes already exist, so indexing ALWAYS runs as a fresh rebuild
(repo corpus + var/mail), matching the proven-safe `kb/index --rebuild` path.
Conversion and indexing stay separate: conversion can crash in native docling
and must not leave the brain DB mid-transaction.
bin/mail/index_mail rebuild the index incl. all mail
bin/mail/index_mail --dry-run count without writing
bin/mail/index_mail --limit N cap messages included
bin/mail/index_mail --since D only messages dated >= D (YYYY-MM-DD)
"""
from __future__ import annotations
import os
import argparse
import json
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import DB_PATH, VAR, connect, ensure_indexes, init_schema, stats, upsert_leaf # noqa: E402
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
def msg_date(md: Path) -> str:
j = md.parent / "message.json"
try:
d = json.loads(j.read_text(encoding="utf-8"))
return (d.get("receivedDate") or d.get("receivedAt") or "")[:10]
except Exception:
return ""
def mail_leafs(limit: int, since: str, repo: str = "ooMail") -> list[dict]:
root = ROOT / "var" / "mail"
mds = sorted(root.rglob("message.md"))
if since:
mds = [m for m in mds if msg_date(m) >= since]
if limit:
mds = mds[:limit]
leafs: list[dict] = []
for md in mds:
files = [md] + sorted((md.parent / "attachments").glob("*.md"))
for f in files:
if not f.exists():
continue
for lf in to_all(read_markdown(f), f, repo=repo):
lf["source"] = f"ooMail:{md.parent.name}:{f.name}"
lf["how"] = "mail/import"
leafs.append(lf)
return leafs
def main(argv: list[str]) -> int:
print(
"bin/mail/index_mail is deprecated; use bin/brain/index.go --rebuild --with-mail",
file=sys.stderr,
)
index = ROOT / "bin" / "kb" / "index"
os.execv(sys.executable, [sys.executable, str(index), "--rebuild", "--with-mail", *argv])
return 1
p = argparse.ArgumentParser(description="rebuild the brain incl. all mail")
p.add_argument("--dry-run", action="store_true", help="count only, write nothing")
p.add_argument("--limit", type=int, default=0, help="cap messages included")
p.add_argument("--since", default="", help="only messages dated >= YYYY-MM-DD")
p.add_argument("--json", action="store_true")
a = p.parse_args(argv)
mail = mail_leafs(a.limit, a.since)
if a.dry_run:
print(f"mail/index_mail: {len(mail)} mail leafs would be indexed")
return 0
# Fresh rebuild: delete DB, index repo corpus + mail, create indexes once
# at the end. Never insert into an already-indexed DB (WAL corruption).
VAR.mkdir(exist_ok=True)
if DB_PATH.exists():
DB_PATH.unlink()
corpus = _load_corpus()
leafs = corpus + mail
db, conn = connect(DB_PATH, read_only=False)
init_schema(conn)
embed = _embedder()
done, total = _index_leafs(conn, leafs, embed)
ensure_indexes(conn)
s = stats(conn)
conn.close()
db.close()
result = {"indexed": done, "corpus_total": total, "mail_leafs": len(mail),
**{k: v for k, v in s.items() if k in ("total", "by_root")}}
print(json.dumps(result, indent=2) if a.json else
f"mail/index_mail: indexed {done}/{total} leafs (mail={len(mail)}); db total {s['total']}")
return 0
CORPUS_DEFAULTS = ["README.md", "PLAN.md", "AGENTS.md", "docs", "skills"]
def _load_corpus() -> list[dict]:
files: list[Path] = []
for entry in CORPUS_DEFAULTS:
p = ROOT / entry
if p.is_file():
files.append(p)
elif p.is_dir():
files.extend(walk_markdown(p))
leafs: list[dict] = []
for path in files:
try:
leafs.extend(to_all(read_markdown(path), path, repo="eSlider/2dph"))
except OSError as e:
print(f"mail/index_mail: skip {path}: {e}", file=sys.stderr)
return leafs
def _index_leafs(conn, leafs: list[dict], embed_fn) -> tuple[int, int]:
count = 0
for lf in leafs:
query = f"{lf['heading']}\n\n{lf['text']}"
emb = embed_fn(lf["text"]) if lf["text"] else None
upsert_leaf(conn, text=query, root="info", confidence="confirmed",
source=lf["source"], source_rev="mail" if lf.get("how") == "mail/import" else "working-tree",
how=lf.get("how", "kb/index"), loc=lf["source"], type_=lf.get("type", "reference"),
embedding=emb)
count += 1
return count, len(leafs)
def _embedder():
from model2vec import StaticModel
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
return lambda text: model.encode([text])[0].astype(float).tolist()
if __name__ == "__main__":
-46
View File
@@ -1,46 +0,0 @@
//usr/bin/env go run -tags=mail_ocr "$0" "$@"; exit
//go:build mail_ocr
//
// bin/mail/ocr.go - OCR an image or scanned PDF (tesseract eng+deu).
//
// ./bin/mail/ocr.go scan.png
// ./bin/mail/ocr.go scan.pdf
// OCR_ENGINE=paddle ./bin/mail/ocr.go scan.png
//
// PDFs try pdftotext -layout first; empty text layer uses pdftoppm + tesseract.
// No gocv. Tesseract CGO bindings are not used (D21 Zig owns Ladybug CGO).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"strings"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/ocr"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
c, err := ocr.ParseArgs(args)
if err != nil {
return cliparse.Fail(err)
}
path := c.Path
var text string
if strings.HasSuffix(strings.ToLower(path), ".pdf") {
text, err = ocr.PDFFile(path)
} else {
text, err = ocr.ImageFile(path)
}
if err != nil {
fmt.Fprintf(os.Stderr, "mail/ocr: %v\n", err)
return 1
}
fmt.Println(text)
return 0
}
+1 -1
View File
@@ -6,7 +6,7 @@
// ./bin/mail/sync.go --dry-run
//
// Writes raw message.json + attachments under var/mail/<folder>/<id>/; run
// bin/mail/import.go --from-raw afterwards to convert everything to markdown.
// bin/mail/import --from-raw afterwards to convert everything to markdown.
//
// Shebang trick: first line is a Go `//` comment; the real code lives in the
// importable package (module path, never a relative import).
+40 -68
View File
@@ -5,15 +5,12 @@ package sync
import (
"context"
"errors"
"flag"
"fmt"
"os"
"path/filepath"
"strings"
"time"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/integrii/flaggy"
)
// CLIConfig is a superset of SyncConfig plus flag parsing results.
@@ -24,84 +21,58 @@ type CLIConfig struct {
Help bool
}
type flagVals struct {
env, out, srcs, query string
workers, limit, offset int
force, dryRun bool
}
func Parser() *flaggy.Parser {
v := flagVals{workers: 4, query: "in:inbox", srcs: "onlyoffice"}
return bind(&v)
}
func bind(v *flagVals) *flaggy.Parser {
if v.workers == 0 {
v.workers = 4
}
if v.query == "" {
v.query = "in:inbox"
}
if v.srcs == "" {
v.srcs = "onlyoffice"
}
p := cliparse.New("mail-sync")
p.Description = "download mail to var/mail"
p.String(&v.env, "", "env", ".env file")
p.String(&v.out, "", "out", "var/mail root")
p.Int(&v.workers, "", "workers", "concurrent downloads")
p.Int(&v.limit, "", "limit", "max messages per source (0 = all)")
p.Int(&v.offset, "", "offset", "skip first N messages per source")
p.Bool(&v.force, "", "force", "overwrite existing message.json")
p.Bool(&v.dryRun, "", "dry-run", "list counts without writing")
p.String(&v.query, "", "query", "Gmail search query")
p.String(&v.srcs, "", "source", "comma list: onlyoffice,gmail")
return p
}
// ParseCLI reads args into a CLIConfig. Exit codes: 0 ok, 2 usage.
// ParseCLI reads os.Args into a CLIConfig. Exit codes: 0 ok, 2 usage.
func ParseCLI(args []string) (CLIConfig, int, error) {
v := flagVals{workers: 4, query: "in:inbox", srcs: "onlyoffice"}
p := bind(&v)
if err := cliparse.Parse(p, args); err != nil {
if errors.Is(err, cliparse.ErrHelp) {
return CLIConfig{Help: true}, 0, nil
}
fs := flag.NewFlagSet("mail/sync", flag.ContinueOnError)
var (
env = fs.String("env", "", ".env file (default: <cwd>/.env)")
out = fs.String("out", "", "var/mail root (default: <cwd>/var/mail)")
workers = fs.Int("workers", 4, "concurrent downloads")
limit = fs.Int("limit", 0, "max messages per source (0 = all)")
offset = fs.Int("offset", 0, "skip first N messages per source")
force = fs.Bool("force", false, "overwrite existing message.json + attachments")
dryRun = fs.Bool("dry-run", false, "list message counts without writing")
query = fs.String("query", "in:inbox", "Gmail search query (gmail source only)")
srcs = fs.String("source", "onlyoffice", "comma list: onlyoffice,gmail (default onlyoffice)")
help = fs.Bool("help", false, "usage")
)
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return CLIConfig{}, 2, err
}
if len(p.TrailingArguments) > 0 {
if *help || fs.NArg() > 0 {
return CLIConfig{Help: true}, 0, nil
}
wd, err := os.Getwd()
if err != nil {
return CLIConfig{}, 2, err
}
if v.env == "" {
v.env = filepath.Join(wd, ".env")
if *env == "" {
*env = filepath.Join(wd, ".env")
}
if v.out == "" {
v.out = filepath.Join(wd, "var", "mail")
if *out == "" {
*out = filepath.Join(wd, "var", "mail")
}
envVars := readEnv(v.env)
envVars := readEnv(*env)
cfg := SyncConfig{
Out: v.out,
Workers: v.workers,
Limit: v.limit,
Offset: v.offset,
Force: v.force,
DryRun: v.dryRun,
Query: v.query,
Out: *out,
Workers: *workers,
Limit: *limit,
Offset: *offset,
Force: *force,
DryRun: *dryRun,
Query: *query,
Policy: RetryPolicy{},
}
out := CLIConfig{Sync: cfg, Env: v.env, Sources: v.srcs}
for _, s := range strings.Split(v.srcs, ",") {
cli := CLIConfig{Sync: cfg, Env: *env, Sources: *srcs}
for _, s := range strings.Split(*srcs, ",") {
switch strings.TrimSpace(s) {
case "onlyoffice":
u := pick(envVars["ONLYOFFICE_URL"], envVars["OO_URL"])
user := pick(envVars["ONLYOFFICE_USER"], envVars["OO_USER"])
pass := pick(envVars["ONLYOFFICE_PASS"], envVars["OO_PASSWORD"])
if u == "" || user == "" || pass == "" {
return CLIConfig{}, 2, fmt.Errorf("onlyoffice source needs ONLYOFFICE_URL/USER/PASS in %s", v.env)
return CLIConfig{}, 2, fmt.Errorf("onlyoffice source needs ONLYOFFICE_URL/USER/PASS in %s", *env)
}
cfg.OO = &OOConfig{URL: u, User: user, Password: pass}
case "gmail":
@@ -114,30 +85,30 @@ func ParseCLI(args []string) (CLIConfig, int, error) {
return CLIConfig{}, 2, fmt.Errorf("unknown source %q", s)
}
}
out.Sync = cfg
return out, 0, nil
cli.Sync = cfg
return cli, 0, nil
}
// Main is the CLI entry: returns process exit code.
func Main(args []string) int {
cfg, code, err := ParseCLI(args)
cli, code, err := ParseCLI(args)
if err != nil {
fmt.Fprintln(os.Stderr, "mail/sync:", err)
return code
}
if cfg.Help {
if cli.Help {
fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail] [--query GMAIL_Q] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]")
return 0
}
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
defer cancel()
start := time.Now()
stats, err := Run(ctx, cfg.Sync)
stats, err := Run(ctx, cli.Sync)
if err != nil {
fmt.Fprintln(os.Stderr, "mail/sync:", err)
return 1
}
if cfg.Sync.DryRun {
if cli.Sync.DryRun {
fmt.Printf("mail/sync: dry-run checked=%d (no writes)\n", stats.Checked)
return 0
}
@@ -164,6 +135,7 @@ func readEnv(path string) map[string]string {
k, v, _ := strings.Cut(line, "=")
out[strings.TrimSpace(k)] = strings.Trim(strings.TrimSpace(v), "\"'")
}
// env overrides file
for _, kv := range os.Environ() {
k, v, ok := strings.Cut(kv, "=")
if !ok {
-3
View File
@@ -1,3 +0,0 @@
// Commands in this directory are shebang mains (import.go), tagged so
// `go build ./bin/markdown` does not see two mains.
package main
-85
View File
@@ -1,85 +0,0 @@
//usr/bin/env go run "$0" "$@"; exit
//
// bin/markdown/import.go - split markdown into leafs (H2 boundaries).
//
// ./bin/markdown/import.go [dir]
// ./bin/markdown/import.go --files a.md,b.md --json
//
// Conversion only. Brain write is bin/brain/index.go.
// Python bin/md/import remains as a fallback.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"strings"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/mdleaves"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
c, err := mdleaves.ParseArgs(args)
if err != nil {
return cliparse.Fail(err)
}
jsonOut := c.JSONOut
files := c.Files
root := c.Root
var paths []string
if files != "" {
for _, f := range strings.Split(files, ",") {
f = strings.TrimSpace(f)
if f != "" {
paths = append(paths, f)
}
}
} else {
st, err := os.Stat(root)
if err != nil {
fmt.Fprintf(os.Stderr, "md/import: no such path %s\n", root)
return 2
}
if !st.IsDir() {
paths = []string{root}
} else {
var err error
paths, err = mdleaves.WalkMarkdown(root)
if err != nil {
fmt.Fprintf(os.Stderr, "md/import: %v\n", err)
return 1
}
}
}
if len(paths) == 0 {
fmt.Fprintln(os.Stderr, "md/import: no markdown files")
return 1
}
var all []mdleaves.Leaf
for _, p := range paths {
raw, err := os.ReadFile(p)
if err != nil {
fmt.Fprintf(os.Stderr, "md/import: %s: %v\n", p, err)
continue
}
all = append(all, mdleaves.ToAll(string(raw), p, "")...)
}
if jsonOut {
s, err := mdleaves.EncodeJSON(all)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
fmt.Print(s)
return 0
}
fmt.Print(mdleaves.EncodeYAML(all))
return 0
}
-2
View File
@@ -1,2 +0,0 @@
// Commands in this directory are shebang mains (query.go).
package main
-20
View File
@@ -1,20 +0,0 @@
//usr/bin/env go run -tags=postgres_query "$0" "$@"; exit
//go:build postgres_query
//
// bin/postgres/query.go - read-only Postgres as YAML.
//
// ./bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
//
// Profiles: $HOME/.config/brain/db-profiles.yml (credentials stay out of git).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/db/psql-yq", os.Args[1:]))
}
-61
View File
@@ -1,61 +0,0 @@
//usr/bin/env go run -tags=qa_stats "$0" "$@"; exit
//go:build qa_stats
//
// bin/qa/stats.go - DuckDB quantiles over a JSON number array or JSONL count.
//
// ./bin/qa/stats.go <<< '[1,2,3,4,5]'
// ./bin/qa/stats.go --jsonl rows.jsonl
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
// DuckDB CGO needs gcc/g++ (not Zig). After eval "$(bin/cgo/zig env)":
// CC=gcc CXX=g++ CGO_CFLAGS= CGO_LDFLAGS= ./bin/qa/stats.go
package main
import (
"encoding/json"
"fmt"
"io"
"os"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/duckstats"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
c, err := cliparse.ParseQAStats(args)
if err != nil {
return cliparse.Fail(err)
}
jsonl := c.JSONL
if jsonl != "" {
n, err := duckstats.CountJSONL(jsonl)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
fmt.Printf("n: %d\n", n)
return 0
}
raw, err := io.ReadAll(os.Stdin)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
var samples []float64
if err := json.Unmarshal(raw, &samples); err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
s, err := duckstats.Quantiles(samples)
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
fmt.Printf("n: %d\nmin: %g\np50: %g\np95: %g\nmax: %g\navg: %g\n",
s.N, s.Min, s.P50, s.P95, s.Max, s.Avg)
return 0
}
-73
View File
@@ -1,73 +0,0 @@
//usr/bin/env go run -tags=reasoner_bakeoff "$0" "$@"; exit
//go:build reasoner_bakeoff
//
// bin/reasoner/bakeoff.go - CPU tool-call bake-off against an OpenAI-compatible URL (D18).
//
// REASONER_BASE_URL=http://127.0.0.1:11435/v1 REASONER_MODEL=qwen3.5:9b ./bin/reasoner/bakeoff.go
// ./bin/reasoner/bakeoff.go --model MichelRosselli/bonsai-27b:Q1_0 --json
//
// Measures OpenAI tool_calls (search/get/audit) and RSS from Ollama /api/ps, not VRAM.
// PicoClaw is compose profile picoclaw; tool names match internal/httpapi MCP ops.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"encoding/json"
"fmt"
"os"
cliparse "github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/duckstats"
"github.com/eSlider/2dph/internal/reasoner"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
c, err := reasoner.ParseArgs(args)
if err != nil {
return cliparse.Fail(err)
}
base, model, jsonOut, device := c.Base, c.Model, c.JSONOut, c.Device
client := reasoner.Client{BaseURL: base, Model: model, Device: device}
rep := reasoner.Run(client)
lat := make([]float64, 0, len(rep.Prompts))
for _, p := range rep.Prompts {
lat = append(lat, float64(p.LatencyMS))
}
if st, err := duckstats.Quantiles(lat); err == nil {
rep.LatencyP50MS = st.P50
rep.LatencyP95MS = st.P95
}
raw, err := json.MarshalIndent(rep, "", " ")
if err != nil {
fmt.Fprintln(os.Stderr, err)
return 1
}
if jsonOut {
fmt.Println(string(raw))
} else {
fmt.Printf("model: %s\n", rep.Model)
fmt.Printf("hf_id: %s\n", rep.HF)
fmt.Printf("device: %s\n", rep.Device)
fmt.Printf("tool_call: %d/%d\n", rep.ToolCallOK, rep.ToolCallN)
fmt.Printf("xml_leak: %d\n", rep.XMLLeak)
fmt.Printf("rss_mb: %d\n", rep.RSSMB)
fmt.Printf("vram_mb: %d\n", rep.VRAMMB)
fmt.Printf("latency_p50_ms: %g\n", rep.LatencyP50MS)
fmt.Printf("latency_p95_ms: %g\n", rep.LatencyP95MS)
for _, p := range rep.Prompts {
status := "fail"
if p.OK {
status = "ok"
}
fmt.Printf(" %s: %s wanted=%s got=%s xml=%v %dms %s\n", p.WantedTool, status, p.WantedTool, p.ToolName, p.XMLLeak, p.LatencyMS, p.Err)
}
}
if rep.ToolCallN == 0 {
return 1
}
return 0
}
-2
View File
@@ -1,2 +0,0 @@
// Commands in this directory are shebang mains (bakeoff.go).
package main
+1 -1
View File
@@ -18,5 +18,5 @@ func main() {
os.Setenv("KB_ROOT", wd)
}
}
httpapi.Run(nil)
httpapi.Run()
}
-235
View File
@@ -1,235 +0,0 @@
# bin/stack/lib.sh — compose helpers for start / start-assistant / stop / status.
# Sourced, not executed. No secrets. No host-absolute paths.
BRAIN_URL="${BRAIN_URL:-http://127.0.0.1:8630}"
REASONER_URL="${REASONER_URL:-http://127.0.0.1:11435}"
PICOCLAW_URL="${PICOCLAW_URL:-http://127.0.0.1:18790}"
REASONER_MODEL="${REASONER_MODEL:-qwen3.5:9b}"
STACK_WAIT_SECS="${STACK_WAIT_SECS:-90}"
STACK_WAIT_INTERVAL="${STACK_WAIT_INTERVAL:-2}"
STACK_PULL_SECS="${STACK_PULL_SECS:-600}"
if [[ -z "${ROOT:-}" ]]; then
STACK_DIR="$(CDPATH= cd -- "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
ROOT="$(CDPATH= cd -- "$STACK_DIR/../.." && pwd)"
fi
stack_usage() {
awk 'NR == 1 { next } /^#/ { sub(/^# ?/, ""); print; next } { exit }' "$1"
}
stack_die() {
echo "bin/stack: $*" >&2
return 1
}
compose() {
docker compose -f "$ROOT/compose.yaml" --project-directory "$ROOT" "$@"
}
http_get() {
local url=$1
local timeout=${2:-5}
curl -sS --max-time "$timeout" "$url" 2>/dev/null || return 1
}
health_ok() {
local url=$1
local timeout=${2:-5}
local body
body=$(http_get "$url" "$timeout") || return 1
printf '%s' "$body" | grep -q '"status":"ok"'
}
wait_health() {
local url=$1
local n=0
while ((n <= STACK_WAIT_SECS)); do
if health_ok "$url"; then
return 0
fi
n=$((n + 1))
if ((n <= STACK_WAIT_SECS)); then
sleep "$STACK_WAIT_INTERVAL"
fi
done
return 1
}
wait_http() {
local url=$1
local n=0
while ((n <= STACK_WAIT_SECS)); do
if http_get "$url" 5 >/dev/null; then
return 0
fi
n=$((n + 1))
if ((n <= STACK_WAIT_SECS)); then
sleep "$STACK_WAIT_INTERVAL"
fi
done
return 1
}
mcp_body() {
curl -sS --max-time 10 \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' \
"$BRAIN_URL/mcp" 2>/dev/null || return 1
}
mcp_ok() {
local body
body=$(mcp_body) || return 1
printf '%s' "$body" | grep -Eq '"name": ?"search"' || return 1
printf '%s' "$body" | grep -Eq '"name": ?"get"' || return 1
printf '%s' "$body" | grep -Eq '"name": ?"audit"' || return 1
return 0
}
reasoner_tags() {
http_get "$REASONER_URL/api/tags" 5
}
reasoner_has_model() {
local body
body=$(reasoner_tags) || return 1
printf '%s' "$body" | grep -Fq "$REASONER_MODEL"
}
ensure_mcp() {
mcp_ok || stack_die "MCP tools/list missing search/get/audit at $BRAIN_URL/mcp"
}
ensure_brain() {
if health_ok "$BRAIN_URL/health"; then
echo "brain: reuse $BRAIN_URL" >&2
else
echo "brain: compose up" >&2
compose up -d brain
wait_health "$BRAIN_URL/health" || stack_die "brain health failed at $BRAIN_URL/health"
fi
ensure_mcp
}
pull_reasoner_model() {
echo "reasoner: pulling $REASONER_MODEL (CPU, may take minutes)" >&2
curl -sS --max-time "$STACK_PULL_SECS" \
-H 'Content-Type: application/json' \
-d "{\"name\":\"$REASONER_MODEL\"}" \
"$REASONER_URL/api/pull" >/dev/null
}
ensure_reasoner() {
if reasoner_has_model; then
echo "reasoner: reuse $REASONER_URL model $REASONER_MODEL" >&2
return 0
fi
if ! reasoner_tags >/dev/null; then
echo "reasoner: compose up" >&2
compose --profile reasoner up -d reasoner
wait_http "$REASONER_URL/api/tags" || stack_die "reasoner not listening at $REASONER_URL"
fi
if reasoner_has_model; then
return 0
fi
pull_reasoner_model
reasoner_has_model || stack_die "reasoner missing model $REASONER_MODEL"
}
ensure_picoclaw() {
echo "picoclaw: compose up --no-deps (reuse healthy :8630/:11435)" >&2
compose --profile picoclaw up -d --no-deps picoclaw
wait_health "$PICOCLAW_URL/health" || stack_die "picoclaw health failed at $PICOCLAW_URL/health"
}
stack_status() {
local bh=down mcp=down ph=down present=false
health_ok "$BRAIN_URL/health" && bh=ok
mcp_ok && mcp=ok
reasoner_has_model && present=true
health_ok "$PICOCLAW_URL/health" && ph=ok
cat <<EOF
brain:
url: $BRAIN_URL
health: $bh
mcp: $mcp
reasoner:
url: $REASONER_URL
model: $REASONER_MODEL
present: $present
picoclaw:
url: $PICOCLAW_URL
health: $ph
EOF
}
stack_start() {
ensure_brain
}
stack_attach_agent() {
local opts=()
if [[ -t 0 && -t 1 ]]; then
opts+=(-it)
else
opts+=(-T)
fi
if [[ (! -t 0 || ! -t 1) && $# -eq 0 ]]; then
echo "picoclaw: no TTY. Attach with:" >&2
echo " $ROOT/bin/stack/start-assistant" >&2
echo " docker compose --profile picoclaw exec -it picoclaw picoclaw agent" >&2
return 0
fi
echo "picoclaw: agent (search → get → audit before a factual reply)" >&2
exec docker compose -f "$ROOT/compose.yaml" --project-directory "$ROOT" \
--profile picoclaw exec "${opts[@]}" picoclaw picoclaw agent "$@"
}
stack_start_assistant() {
local attach=1
local agent_args=()
while (($#)); do
case "$1" in
-h | --help)
stack_usage "$ROOT/bin/stack/start-assistant"
return 0
;;
--no-attach)
attach=0
shift
;;
--)
shift
agent_args+=("$@")
break
;;
*)
agent_args+=("$1")
shift
;;
esac
done
stack_start
ensure_reasoner
ensure_picoclaw
stack_status
if ((attach == 0)); then
echo "picoclaw: gateway $PICOCLAW_URL (agent not attached)" >&2
echo "ask the brain: $ROOT/bin/stack/start-assistant" >&2
echo "one-shot: $ROOT/bin/stack/start-assistant -- -m \"search the 2dph brain for LadybugDB\"" >&2
return 0
fi
stack_attach_agent "${agent_args[@]}"
}
stack_stop() {
case "${1:-}" in
-h | --help)
stack_usage "$ROOT/bin/stack/stop"
return 0
;;
esac
echo "stack: stop brain brain-mcp reasoner picoclaw (volumes kept)" >&2
compose --profile picoclaw --profile reasoner stop picoclaw brain-mcp reasoner brain
}
-21
View File
@@ -1,21 +0,0 @@
#!/usr/bin/env bash
# bin/stack/start - bring up brain HTTP/MCP and wait until search/get/audit respond.
#
# bin/stack/start
# bin/stack/status
#
# Reuses a healthy process on :8630 (host serve or compose). Does not start
# PicoClaw. Does not rebuild Ladybug.
set -euo pipefail
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
# shellcheck source=lib.sh
source "$STACK_DIR/lib.sh"
case "${1:-}" in
-h | --help)
stack_usage "$0"
exit 0
;;
esac
stack_start "$@"
stack_status
-15
View File
@@ -1,15 +0,0 @@
#!/usr/bin/env bash
# bin/stack/start-assistant - start + CPU reasoner + PicoClaw, then attach agent.
#
# bin/stack/start-assistant
# bin/stack/start-assistant --no-attach
# bin/stack/start-assistant -- -m "search the 2dph brain for LadybugDB"
#
# Pulls qwen3.5:9b if missing. Gateway :18790. Agent uses MCP search → get → audit.
# --no-attach leaves the gateway up without exec.
set -euo pipefail
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
# shellcheck source=lib.sh
source "$STACK_DIR/lib.sh"
stack_start_assistant "$@"
-17
View File
@@ -1,17 +0,0 @@
#!/usr/bin/env bash
# bin/stack/status - YAML health for brain MCP, reasoner model, PicoClaw gateway.
#
# bin/stack/status
# bin/stack/status | yq '.picoclaw'
set -euo pipefail
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
# shellcheck source=lib.sh
source "$STACK_DIR/lib.sh"
case "${1:-}" in
-h | --help)
stack_usage "$0"
exit 0
;;
esac
stack_status
-13
View File
@@ -1,13 +0,0 @@
#!/usr/bin/env bash
# bin/stack/stop - stop compose brain / brain-mcp / reasoner / picoclaw.
#
# bin/stack/stop
#
# Volumes kept (kb, reasoner weights, picoclaw-home). Does not kill a host
# bin/brain/serve.go that is not a compose service.
set -euo pipefail
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
# shellcheck source=lib.sh
source "$STACK_DIR/lib.sh"
stack_stop "$@"
-103
View File
@@ -1,103 +0,0 @@
"""D16 contradiction adjudication (same rules as internal/facts)."""
from __future__ import annotations
from typing import Any
CONF_CONFIRMED = "confirmed"
CONF_HYPOTHESIS = "hypothesis"
RULE_UNRESOLVED = "unresolved"
RULE_TEMPORAL = "temporal_freshness"
RULE_AUTHORITY = "authority_pairing"
RULE_TWO_SOURCE = "two_source"
RULE_SINGLE = "single_source"
KIND_RUNTIME = "runtime"
KIND_CONFIG = "config"
KIND_NARRATIVE = "narrative"
def _independent(sources: list[dict]) -> int:
seen: set[str] = set()
for i, s in enumerate(sources):
sid = str(s.get("id") or "") or f"{s.get('kind', '')}#{i}"
seen.add(sid)
return len(seen)
def _fresh_n(sources: list[dict]) -> int:
return sum(1 for s in sources if not s.get("stale"))
def _strong_n(sources: list[dict]) -> int:
return sum(1 for s in sources if s.get("kind") in (KIND_RUNTIME, KIND_CONFIG))
def adjudicate(claim: dict[str, Any]) -> dict[str, Any]:
yes = list(claim.get("yes") or [])
no = list(claim.get("no") or [])
yes_n, no_n = _independent(yes), _independent(no)
text = str(claim.get("text") or "")
def out(conf: str, rule: str, winner: str = "") -> dict[str, Any]:
return {
"text": text,
"confidence": conf,
"confirmed": conf == CONF_CONFIRMED,
"rule": rule,
"winner": winner,
"yes": yes_n,
"no": no_n,
}
if yes_n < 2 or no_n < 2:
if yes_n >= 2:
return out(CONF_CONFIRMED, RULE_TWO_SOURCE, "yes")
if no_n >= 2:
return out(CONF_CONFIRMED, RULE_TWO_SOURCE, "no")
return out(CONF_HYPOTHESIS, RULE_SINGLE)
yf, nf = _fresh_n(yes), _fresh_n(no)
if yf >= 2 and nf < 2:
return out(CONF_CONFIRMED, RULE_TEMPORAL, "yes")
if nf >= 2 and yf < 2:
return out(CONF_CONFIRMED, RULE_TEMPORAL, "no")
ys, ns = _strong_n(yes), _strong_n(no)
if ys >= 2 and ns < 2:
return out(CONF_CONFIRMED, RULE_AUTHORITY, "yes")
if ns >= 2 and ys < 2:
return out(CONF_CONFIRMED, RULE_AUTHORITY, "no")
return out(CONF_HYPOTHESIS, RULE_UNRESOLVED)
def parse_source_field(source: str) -> tuple[str, str]:
"""Split `a x b vs c x d` into (yes, no). Empty no if no ` vs `."""
if " vs " not in source:
return source, ""
yes, _, no = source.partition(" vs ")
return yes.strip(), no.strip()
def check_fact_row(lid: str, source: str, loc: str, how: str, conf: str) -> list[str]:
"""Lexicon checks for one facts leaf (no Ladybug)."""
problems: list[str] = []
src = source or ""
if conf == CONF_CONFIRMED:
if " vs " in src:
problems.append(f"{lid}: confirmed fact cannot keep a vs-contradiction")
if " x " not in src:
problems.append(f"{lid}: needs 2-source evidence in source, got '{source}'")
elif conf == CONF_HYPOTHESIS:
yes, no = parse_source_field(src)
if not no or " x " not in yes or " x " not in no:
problems.append(
f"{lid}: hypothesis contradiction needs 'a x b vs c x d', got '{source}'"
)
elif conf == "partial":
pass
else:
problems.append(f"{lid}: unknown confidence '{conf}'")
if not loc:
problems.append(f"{lid}: missing loc (evidence pointer)")
if not how:
problems.append(f"{lid}: missing how")
return problems
+60 -3
View File
@@ -1,12 +1,21 @@
"""gitimport - Ladybug graph writes for Commit/File/Person (no git binary).
"""gitimport - parse `git log` output and turn commits into brain leafs.
Commit records come from bin/git/import.go (go-git). This module only MERGEs
the version graph File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person.
Pure, testable functions. Field grammar (see bin/git/import):
git log --no-merges --name-only \
--format='%x1e%H%x1f%an%x1f%ae%x1f%aI%x1f%s'
0x1e = record separator, 0x1f = field separator.
Files: newline-separated lines following each record's subject.
"""
from __future__ import annotations
from dataclasses import dataclass, field
REC_SEP = "\x1e"
FIELD_SEP = "\x1f"
@dataclass
class Commit:
@@ -17,6 +26,54 @@ class Commit:
subject: str
files: list[str] = field(default_factory=list)
def leaf_text(self, repo: str) -> str:
head = f"commit {self.sha[:12]} in {repo}{self.subject}"
body = [head, f"Author: {self.author} <{self.email}>", f"Date: {self.date}"]
if self.files:
body.append("Changing: " + ", ".join(self.files))
return "\n".join(body)
def parse_log(text: str) -> list[Commit]:
"""Parse `git log` output into Commit records.
Records are separated by 0x1e. A record is fields joined by 0x1f,
followed by optional newline-separated file paths inside the next
segment (git emits blank line + files after each record).
"""
commits: list[Commit] = []
# field records and file lists alternate; simpler: split on REC_SEP,
# each chunk = header line, possibly followed by newline + files.
for chunk in text.split(REC_SEP):
chunk = chunk.strip("\n")
if not chunk:
continue
lines = chunk.split("\n", 1)
header = lines[0].split(FIELD_SEP)
if len(header) < 5:
continue
sha, author, email, date, subject = header[:5]
files = [ln.strip() for ln in lines[1].splitlines() if ln.strip()] if len(lines) > 1 else []
commits.append(Commit(sha=sha, author=author, email=email,
date=date, subject=subject, files=files))
return commits
def commits_to_leafs(commits: list[Commit], repo: str) -> list[dict]:
"""Map commits to the leaf shape bin/kb/index expects (source/repo/...)."""
out: list[dict] = []
for c in commits:
out.append({
"source": f"{repo}@{c.sha}",
"repo": repo,
"heading": f"commit {c.sha[:12]}{c.subject}",
"text": c.leaf_text(repo),
"type": "commit",
"status": "current",
"related": ",".join(c.files),
})
return out
GIT_SCHEMA = (
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
+16 -169
View File
@@ -4,7 +4,7 @@ Single embedded graph `var/kb.lbug`. Two roots: facts (assertions backed by
>=2 independent sources) and info (narrative leafs). Hybrid retrieval: BM25
(FTS extension) + HNSW cosine (VECTOR extension) + Cypher graph hops.
All access is read-only unless `--rebuild` (kb/index) or `kb/add`.
All access is read-only unless `--rebuild` is passed to kb/index.
"""
from __future__ import annotations
@@ -62,9 +62,7 @@ def init_schema(conn: ladybug.Connection) -> None:
"CREATE NODE TABLE IF NOT EXISTS Leaf ("
" id STRING, text STRING, root STRING, confidence STRING, "
" sha256 STRING, source STRING, source_rev STRING, observed_at STRING, "
" how STRING, loc STRING, type STRING, "
" valid_from STRING, valid_to STRING, "
" embedding FLOAT[256], "
" how STRING, loc STRING, type STRING, embedding FLOAT[256], "
" PRIMARY KEY(id))"
)
conn.execute(
@@ -93,46 +91,6 @@ def init_schema(conn: ladybug.Connection) -> None:
conn.execute(
"CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
)
ensure_interval_columns(conn)
def ensure_interval_columns(conn: ladybug.Connection) -> None:
"""D24: add valid_from/valid_to on older Leaf tables (idempotent ALTER)."""
for col in ("valid_from", "valid_to"):
try:
conn.execute(f"ALTER TABLE Leaf ADD {col} STRING")
except Exception:
pass
def normalize_day(s: str) -> str:
s = (s or "").strip()
if len(s) >= 10 and s[4] == "-" and s[7] == "-":
return s[:10]
return s
def active_at(valid_from: str, valid_to: str, as_of: str) -> bool:
"""D24: fact interval of truth. Empty ends = always; empty as_of = no filter."""
as_of = normalize_day(as_of)
if not as_of:
return True
fro = normalize_day(valid_from)
to = normalize_day(valid_to)
if fro and as_of < fro:
return False
if to and as_of > to:
return False
return True
def filter_as_of(hits: list[dict], as_of: str) -> list[dict]:
if not as_of:
return hits
return [
h for h in hits
if active_at(str(h.get("valid_from") or ""), str(h.get("valid_to") or ""), as_of)
]
def leaf_id(text: str, source: str) -> str:
@@ -141,124 +99,25 @@ def leaf_id(text: str, source: str) -> str:
def upsert_leaf(conn: ladybug.Connection, *, text: str, root: str, confidence: str,
source: str, source_rev: str, how: str, loc: str, type_: str,
embedding: list[float] | None,
valid_from: str = "", valid_to: str = "") -> str:
embedding: list[float] | None) -> str:
lid = leaf_id(text, source)
obs = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
vf = normalize_day(valid_from)
vt = normalize_day(valid_to)
conn.execute(
"MERGE (l:Leaf {id:$id}) "
"SET l.text=$text, l.root=$root, l.confidence=$confidence, "
" l.sha256=$sha, l.source=$source, l.source_rev=$rev, l.observed_at=$obs, "
" l.how=$how, l.loc=$location, l.type=$type, "
" l.valid_from=$vf, l.valid_to=$vt"
" l.how=$how, l.loc=$location, l.type=$type"
+ (", l.embedding=$emb" if embedding else ""),
parameters={
"id": lid, "text": text, "root": root, "confidence": confidence,
"sha": sha256_b64(text), "source": source, "rev": source_rev,
"obs": obs, "how": how, "location": loc, "type": type_,
"vf": vf, "vt": vt,
"emb": (embedding if embedding else None),
},
)
return lid
def add_leafs(conn: ladybug.Connection, leafs: list[dict]) -> list[str]:
"""Write facts+info leafs in one transaction. Safe while FTS/HNSW exist.
Each leaf dict: text, source, optional root/confidence/source_rev/how/loc/type/
embedding/valid_from/valid_to. Does not delete the database file. Measured on
Ladybug 0.19: MERGE of new ids (and updates) stays FTS+HNSW queryable; DROP
INDEX is the fatal path.
"""
if not leafs:
return []
started = False
try:
conn.execute("BEGIN TRANSACTION")
started = True
except Exception:
started = False
ids: list[str] = []
try:
for lf in leafs:
ids.append(
upsert_leaf(
conn,
text=str(lf["text"]),
root=str(lf.get("root") or ROOT_INFO),
confidence=str(lf.get("confidence") or CONF_CONFIRMED),
source=str(lf["source"]),
source_rev=str(lf.get("source_rev") or "working-tree"),
how=str(lf.get("how") or "brain/add"),
loc=str(lf.get("loc") or lf.get("source") or ""),
type_=str(lf.get("type") or lf.get("type_") or "reference"),
embedding=lf.get("embedding"),
valid_from=str(lf.get("valid_from") or ""),
valid_to=str(lf.get("valid_to") or ""),
)
)
if started:
conn.execute("COMMIT")
except Exception:
if started:
try:
conn.execute("ROLLBACK")
except Exception:
pass
raise
return ids
def file_id(repo: str, path: str) -> str:
"""Stable File.id matching gitimport (`repo:path`)."""
return f"{repo}:{path}" if repo else path
def link_from_file(conn: ladybug.Connection, leaf_id: str, path: str,
repo: str = "", mtime: str = "") -> str:
"""MERGE File and Leaf-[:FROM_FILE]->File so --hop 1 can walk."""
fid = file_id(repo, path)
conn.execute(
"MERGE (f:File {id:$id}) SET f.path=$path, f.repo=$repo, f.mtime=$mtime",
parameters={"id": fid, "path": path, "repo": repo, "mtime": mtime},
)
conn.execute(
"MATCH (l:Leaf {id:$lid}), (f:File {id:$fid}) "
"MERGE (l)-[:FROM_FILE]->(f)",
parameters={"lid": leaf_id, "fid": fid},
)
return fid
HOP_STMTS = {
1: "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File) RETURN f.id, f.path, 1",
2: ("MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit) "
"RETURN c.id, c.subject, 2"),
3: ("MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit)"
"-[:AUTHORED]->(p:Person) RETURN p.id, p.name, 3"),
}
HOP_LABELS = {1: "File", 2: "Commit", 3: "Person"}
def hop_walk(conn: ladybug.Connection, leaf_id: str, n: int) -> list[dict]:
"""Walk Leaf → File → Commit → Person up to n hops (max 3)."""
depth = min(max(int(n), 0), 3)
out: list[dict] = []
for d in range(1, depth + 1):
rows = conn.execute(HOP_STMTS[d], parameters={"id": leaf_id}).get_all()
for row in rows:
out.append({
"id": row[0],
"label": HOP_LABELS[d],
"name": row[1],
"depth": int(row[2]),
})
return out
def leaf_index_names(conn: ladybug.Connection) -> set[str]:
"""Return index names on the Leaf table (e.g. {'id', 'Leaf_vec', '_PK'})."""
rows = conn.execute("CALL SHOW_INDEXES() RETURN *").get_all()
@@ -276,7 +135,7 @@ def create_fts_and_vector(conn: ladybug.Connection, force: bool = False) -> None
`force=True` is accepted for API compatibility but does **not** drop.
Fresh indexes require deleting `var/kb.lbug` and rebuilding
(`bin/brain/index.go --rebuild`).
(`bin/kb/index --rebuild`).
"""
del force # API compat; DROP is unsafe — see docstring
names = leaf_index_names(conn)
@@ -286,7 +145,7 @@ def create_fts_and_vector(conn: ladybug.Connection, force: bool = False) -> None
except Exception as e:
raise RuntimeError(
"CREATE_FTS_INDEX failed (often ghost catalog after DROP INDEX). "
"Delete var/kb.lbug and run bin/brain/index.go --rebuild. "
"Delete var/kb.lbug and run bin/kb/index --rebuild. "
f"Cause: {e}"
) from e
if "Leaf_vec" not in names:
@@ -299,7 +158,7 @@ def create_fts_and_vector(conn: ladybug.Connection, force: bool = False) -> None
raise RuntimeError(
"CREATE_VECTOR_INDEX failed (often ghost catalog after DROP INDEX "
"Leaf.Leaf_vec → `_0_Leaf_vec_UPPER already exists in catalog`). "
"Delete var/kb.lbug and run bin/brain/index.go --rebuild. "
"Delete var/kb.lbug and run bin/kb/index --rebuild. "
f"Cause: {e}"
) from e
names = leaf_index_names(conn)
@@ -335,39 +194,29 @@ def drop_indexes(conn: ladybug.Connection) -> None:
def query_fts(conn: ladybug.Connection, text: str, limit: int = 10) -> list[dict]:
r = conn.execute(
"CALL QUERY_FTS_INDEX('Leaf', 'id', $q) "
"RETURN node.id, node.text, node.root, score, node.valid_from, node.valid_to "
"ORDER BY score DESC LIMIT $n",
"RETURN node.id, node.text, node.root, score ORDER BY score DESC LIMIT $n",
parameters={"q": text, "n": limit},
)
return [
{
"id": row[0], "text": row[1], "root": row[2], "score": row[3],
"valid_from": row[4] or "", "valid_to": row[5] or "",
}
for row in r.get_all()
]
return [{"id": row[0], "text": row[1], "root": row[2], "score": row[3]} for row in r.get_all()]
def query_vector(conn: ladybug.Connection, embedding: list[float], limit: int = 10) -> list[dict]:
r = conn.execute(
"CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) "
"RETURN node.id, node.text, node.root, distance, node.valid_from, node.valid_to "
"ORDER BY distance LIMIT $n",
"RETURN node.id, node.text, node.root, distance ORDER BY distance LIMIT $n",
parameters={"q": embedding, "n": limit},
)
out = []
for row in r.get_all():
# distance -> similarity reasonable for cosine
score = 1.0 - row[3] if row[3] is not None else 0.0
out.append({
"id": row[0], "text": row[1], "root": row[2], "score": score,
"valid_from": row[4] or "", "valid_to": row[5] or "",
})
out.append({"id": row[0], "text": row[1], "root": row[2], "score": score})
return out
def hybrid_search(conn: ladybug.Connection, embedding: list[float], fts_hits: list[dict],
limit: int = 10, as_of: str = "") -> list[dict]:
"""Merge FTS + vector by reciprocal rank fusion; optional D24 as-of filter."""
limit: int = 10) -> list[dict]:
"""Merge FTS + vector by reciprocal rank fusion."""
fused: dict[str, dict] = {}
for rank, hit in enumerate(fts_hits):
fused.setdefault(hit["id"], {**hit, "rrf": 0.0})["rrf"] = 1.0 / (60 + rank + 1)
@@ -375,10 +224,8 @@ def hybrid_search(conn: ladybug.Connection, embedding: list[float], fts_hits: li
entry = fused.setdefault(hit["id"], {**hit, "rrf": 0.0})
entry["rrf"] += 1.0 / (60 + rank + 1)
entry.setdefault("score", hit.get("score", 0.0))
entry.setdefault("valid_from", hit.get("valid_from") or "")
entry.setdefault("valid_to", hit.get("valid_to") or "")
ranked = sorted(fused.values(), key=lambda h: h.get("rrf", 0.0), reverse=True)
return filter_as_of(ranked, as_of)[:limit]
return ranked[:limit]
def stats(conn: ladybug.Connection) -> dict:
@@ -390,6 +237,6 @@ def stats(conn: ladybug.Connection) -> dict:
def open_readonly() -> tuple[ladybug.Database, ladybug.Connection]:
if not DB_PATH.exists():
raise FileNotFoundError(f"{DB_PATH} missing - run bin/brain/index.go --rebuild first")
raise FileNotFoundError(f"{DB_PATH} missing - run bin/kb/index first")
db, conn = connect(read_only=True)
return db, conn
+1 -84
View File
@@ -7,10 +7,7 @@ offline against fixtures.
from __future__ import annotations
import html
import os
import re
import subprocess
import tempfile
import zipfile
from pathlib import Path
@@ -21,10 +18,9 @@ OFFICE_SUFFIXES = {".docx", ".pptx", ".xlsx", ".html", ".htm", ".epub", ".eml",
PDF_SUFFIXES = {".pdf"}
IMAGE_SUFFIXES = {".png", ".jpg", ".jpeg", ".gif", ".bmp", ".tiff", ".tif", ".webp"}
ARCHIVE_SUFFIXES = {".zip"}
# Legacy binary Office (doc/xls/ppt) — markitdown skip them; we try
# Legacy binary Office (doc/xls/ppt) — markitdown/docling skip them; we try
# pandoc first, else leave a stub.
LEGACY_OFFICE_SUFFIXES = {".doc", ".xls", ".ppt"}
TESS_LANG = "eng+deu"
CONVERTIBLE_SUFFIXES = (
TEXT_SUFFIXES | OFFICE_SUFFIXES | PDF_SUFFIXES | IMAGE_SUFFIXES | ARCHIVE_SUFFIXES | LEGACY_OFFICE_SUFFIXES
@@ -150,82 +146,3 @@ def zip_extract_safe(zip_path: Path, dest: Path) -> list[Path]:
def is_convertible(suffix: str) -> bool:
return suffix.lower() in CONVERTIBLE_SUFFIXES
def convert_pdf(path: Path, ocr: bool = False) -> str:
"""pdftotext -layout first; empty text layer → pdftoppm + tesseract.
`ocr` is unused for born-digital PDFs (text layer wins). Scans OCR
automatically. This path never execs an ONNX document converter.
"""
del ocr # scans OCR when the text layer is empty; flag is for images
text = pdf_fast_text(path)
if text and text.strip():
return normalize_markdown(text)
scanned = ocr_pdf(path)
if scanned and scanned.strip():
return normalize_markdown(scanned)
if text:
return normalize_markdown(text)
return "\n<!-- pdf has no text layer (ocr unavailable) -->\n"
def pdf_fast_text(path: Path) -> str | None:
"""pdftotext -layout; None when poppler is missing or the command fails."""
try:
proc = subprocess.run(
["pdftotext", "-layout", str(path), "-"],
capture_output=True, timeout=60)
except (OSError, subprocess.TimeoutExpired):
return None
if proc.returncode != 0:
return None
return proc.stdout.decode("utf-8", errors="replace")
def ocr_pdf(path: Path) -> str:
"""Rasterize with pdftoppm and OCR each page (tesseract or paddle)."""
try:
with tempfile.TemporaryDirectory(prefix="2dph-ocr-") as tmp:
prefix = str(Path(tmp) / "page")
proc = subprocess.run(
["pdftoppm", "-png", "-r", "200", str(path), prefix],
capture_output=True, timeout=120)
if proc.returncode != 0:
return ""
pages = sorted(Path(tmp).glob("page*.png"))
parts = [ocr_image(p) for p in pages]
return "\n\n".join(p for p in parts if p and p.strip())
except (OSError, subprocess.TimeoutExpired):
return ""
def ocr_image(path: Path) -> str:
engine = os.environ.get("OCR_ENGINE", "tesseract")
if engine == "paddle":
return _ocr_paddle(path)
return _ocr_tesseract(path)
def _ocr_tesseract(path: Path) -> str:
try:
proc = subprocess.run(
["tesseract", str(path), "stdout", "-l", TESS_LANG, "--psm", "6"],
capture_output=True, timeout=120)
except (OSError, subprocess.TimeoutExpired):
return ""
if proc.returncode != 0:
return ""
return proc.stdout.decode("utf-8", errors="replace").strip()
def _ocr_paddle(path: Path) -> str:
try:
proc = subprocess.run(
["paddleocr", "ocr", "-i", str(path)],
capture_output=True, timeout=180)
except (OSError, subprocess.TimeoutExpired):
return ""
if proc.returncode != 0:
return ""
return proc.stdout.decode("utf-8", errors="replace").strip()
-37
View File
@@ -1,37 +0,0 @@
"""Mail markdown under var/mail → info leafs. Conversion stays off the brain DB."""
from __future__ import annotations
import json
from pathlib import Path
from mdleaves import read_markdown, to_all
def msg_date(md: Path) -> str:
j = md.parent / "message.json"
try:
d = json.loads(j.read_text(encoding="utf-8"))
return (d.get("receivedDate") or d.get("receivedAt") or "")[:10]
except (OSError, json.JSONDecodeError, TypeError):
return ""
def from_mail_root(root: Path, limit: int = 0, since: str = "", repo: str = "ooMail") -> list[dict]:
if not root.is_dir():
return []
mds = sorted(root.rglob("message.md"))
if since:
mds = [m for m in mds if msg_date(m) >= since]
if limit:
mds = mds[:limit]
leafs: list[dict] = []
for md in mds:
files = [md] + sorted((md.parent / "attachments").glob("*.md"))
for f in files:
if not f.exists():
continue
for lf in to_all(read_markdown(f), f, repo=repo):
lf["source"] = f"ooMail:{md.parent.name}:{f.name}"
lf["how"] = "mail/import"
leafs.append(lf)
return leafs
-280
View File
@@ -1,7 +1,6 @@
"""D14 layout: bin/{subject}/{method}.go, libs in internal/, one go.mod."""
from __future__ import annotations
import os
import unittest
from pathlib import Path
@@ -35,282 +34,3 @@ class BinLayoutTest(unittest.TestCase):
def test_no_main_go_under_bin_brain(self) -> None:
main = ROOT / "bin" / "brain" / "main.go"
self.assertFalse(main.exists(), "bin/brain/main.go is not a method")
def test_chats_methods_are_shebangs_not_main(self) -> None:
chats = ROOT / "bin" / "chats"
self.assertFalse(
(chats / "main.go").exists(),
"bin/chats/main.go is a dispatcher, not a method",
)
self.assertFalse(
(chats / "index_cmd.go").exists(),
"chats index is a brain write hiding under the wrong subject",
)
for method in ("sync.go", "import.go", "facts.go", "apply.go"):
p = chats / method
self.assertTrue(p.is_file(), f"missing bin/chats/{method}")
first = p.read_text().splitlines()[0]
self.assertTrue(
first.startswith("//usr/bin/env go run"),
f"{method} shebang, got {first!r}",
)
def test_chats_lib_lives_in_internal(self) -> None:
self.assertTrue(
(ROOT / "internal" / "chats" / "linkedin.go").is_file(),
"LinkedIn parser must live in internal/chats",
)
self.assertFalse(
(ROOT / "bin" / "chats" / "linkedin.go").exists(),
"parser must not stay under bin/chats as a second main",
)
def _assert_shebang(self, rel: str) -> None:
p = ROOT / rel
self.assertTrue(p.is_file(), f"missing {rel}")
first = p.read_text().splitlines()[0]
self.assertTrue(
first.startswith("//usr/bin/env go run"),
f"{rel} shebang, got {first!r}",
)
def test_brain_methods_are_shebangs(self) -> None:
for method in ("index.go", "add.go", "get.go", "stats.go", "eval.go", "watch.go"):
self._assert_shebang(f"bin/brain/{method}")
def test_brain_add_is_python_write_not_rebuild(self) -> None:
self._assert_shebang("bin/brain/add.go")
text = (ROOT / "bin" / "brain" / "add.go").read_text()
self.assertIn("cmdbin.ExecFile", text)
self.assertIn("bin/kb/add", text)
self.assertNotIn("--rebuild", text)
py = (ROOT / "bin" / "kb" / "add").read_text()
self.assertIn("add_leafs", py)
self.assertIn("--json", py)
self.assertNotIn("unlink", py.lower())
def test_brain_get_stats_eval_are_not_python_exec(self) -> None:
for method in ("get.go", "stats.go", "eval.go"):
text = (ROOT / "bin" / "brain" / method).read_text()
self.assertNotIn(
"ExecFile",
text,
f"bin/brain/{method} must call internal/brain, not ExecFile Python",
)
self.assertNotIn(
"cmdbin",
text,
f"bin/brain/{method} must not import internal/cmdbin",
)
self.assertIn(
"system_ladybug",
text.splitlines()[0],
f"bin/brain/{method} shebang must pass -tags=system_ladybug",
)
self.assertIn(
"github.com/eSlider/2dph/internal/brain",
text,
)
def test_eval_control_questions_live_in_rank(self) -> None:
rank = (ROOT / "internal" / "brain" / "rank" / "evalq.go").read_text()
py = (ROOT / "bin" / "kb" / "eval").read_text()
for frag in ("BM25", "DevOps", "LadybugDB"):
self.assertIn(frag, rank)
self.assertIn(frag, py)
self.assertIn("0.95", rank)
def test_facts_methods_are_shebangs(self) -> None:
for method in ("audit.go", "extract.go", "crm.go"):
self._assert_shebang(f"bin/facts/{method}")
text = (ROOT / "bin" / "facts" / method).read_text()
self.assertIn("cmdbin.ExecFile", text)
self.assertIn(f"bin/facts/{method.removesuffix('.go')}", text)
def test_d16_adjudication_is_cgo_free(self) -> None:
self.assertTrue((ROOT / "internal" / "facts" / "contradict.go").is_file())
go = (ROOT / "internal" / "facts" / "contradict.go").read_text()
py = (ROOT / "bin" / "tools" / "contradict.py").read_text()
audit = (ROOT / "bin" / "facts" / "audit").read_text()
for token in ("temporal_freshness", "authority_pairing", "unresolved"):
self.assertIn(token, go)
self.assertIn(token, py)
self.assertIn("contradict", audit)
self.assertIn(" vs ", py)
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("temporal_freshness", plan)
self.assertIn("authority_pairing", plan)
shebang = (ROOT / "bin" / "facts" / "audit.go").read_text()
self.assertIn("contradict", shebang)
def test_d23_flaggy_cli(self) -> None:
self.assertTrue((ROOT / "internal" / "cli" / "cli.go").is_file())
self.assertIn("github.com/integrii/flaggy", (ROOT / "go.mod").read_text())
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D23", plan)
self.assertIn("flaggy", plan)
complete = (ROOT / "bin" / "cli" / "complete.go").read_text()
first = complete.splitlines()[0]
self.assertTrue(first.startswith("//usr/bin/env go run"), first)
self.assertIn("complete.go bash", complete)
self.assertIn("brain-search", complete)
chats_import = (ROOT / "internal" / "chats" / "import.go").read_text()
self.assertNotIn("flag.NewFlagSet", chats_import)
args = (ROOT / "internal" / "brain" / "rank" / "args.go").read_text()
self.assertIn("internal/cli", args)
def test_mail_import_is_shebang_not_brain_write(self) -> None:
self._assert_shebang("bin/mail/import.go")
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
self.assertIn(
"bin/brain/index.go",
index_mail,
"index_mail must point at bin/brain/index.go",
)
def test_mail_ocr_is_tesseract_not_docling(self) -> None:
self._assert_shebang("bin/mail/ocr.go")
ocr = (ROOT / "bin" / "mail" / "ocr.go").read_text()
self.assertIn("internal/ocr", ocr)
self.assertIn("mail_ocr", ocr)
self.assertNotIn("github.com/otiai10/gosseract", ocr)
py = (ROOT / "bin" / "mail" / "import").read_text()
self.assertNotIn("from docling", py)
self.assertNotIn("import docling", py)
self.assertIn("convert_pdf", py)
conv = (ROOT / "bin" / "tools" / "mailconv.py").read_text()
self.assertIn("pdftotext", conv)
self.assertIn("pdftoppm", conv)
self.assertIn("tesseract", conv)
self.assertIn("eng+deu", conv)
self.assertNotIn("from docling", conv)
self.assertNotIn("import docling", conv)
self.assertNotIn("gocv", conv.lower())
proj = (ROOT / "pyproject.toml").read_text()
self.assertNotIn("docling", proj)
ci = (ROOT / ".github" / "workflows" / "ci.yml").read_text()
self.assertIn("tesseract-ocr", ci)
self.assertIn("./internal/ocr", ci)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn("ocr-paddle", compose)
self.assertIn("OCR_ENGINE", compose)
def test_markdown_import_is_go_not_python_exec(self) -> None:
self._assert_shebang("bin/markdown/import.go")
text = (ROOT / "bin" / "markdown" / "import.go").read_text()
self.assertNotIn("ExecFile", text)
self.assertNotIn("cmdbin", text)
self.assertIn("internal/mdleaves", text)
self.assertNotIn("kb.lbug", text)
def test_import_adapters_do_not_write_ladybug(self) -> None:
for rel in (
"bin/mail/import.go",
"bin/mail/import",
"bin/markdown/import.go",
"bin/chats/import.go",
"bin/git/import.go",
):
text = (ROOT / rel).read_text()
self.assertNotIn("upsert_leaf", text, rel)
self.assertNotIn("kb.lbug", text, rel)
self.assertNotIn("var/brain.lbug", text, rel)
index = (ROOT / "bin" / "brain" / "index.go").read_text()
self.assertIn("bin/kb/index", index)
def test_postgres_query_is_shebang(self) -> None:
self._assert_shebang("bin/postgres/query.go")
def test_git_import_is_gogit_shebang(self) -> None:
self._assert_shebang("bin/git/import.go")
py = (ROOT / "bin" / "git" / "import").read_text()
self.assertNotIn(
'["git"',
py,
"Python git/import must not subprocess the git binary",
)
self.assertIn("bin/git/import.go", py)
def test_web_search_is_shebang(self) -> None:
self._assert_shebang("bin/web/search.go")
py = (ROOT / "bin" / "web" / "search").read_text()
self.assertIn("bin/web/search.go", py)
def test_gitimport_py_has_no_git_binary(self) -> None:
py = (ROOT / "bin" / "tools" / "gitimport.py").read_text()
self.assertNotIn("subprocess", py)
self.assertNotIn("git log", py)
def test_gogit_is_direct_go_mod_require(self) -> None:
text = (ROOT / "go.mod").read_text()
first = text.split("require (")[1].split(")")[0]
self.assertRegex(first, r"github.com/go-git/go-git/v5\s+v")
for line in first.splitlines():
if "go-git/go-git" in line:
self.assertNotIn("indirect", line)
def test_duckdb_go_is_direct_require(self) -> None:
text = (ROOT / "go.mod").read_text()
first = text.split("require (")[1].split(")")[0]
self.assertRegex(first, r"github.com/duckdb/duckdb-go/v2\s+v")
for line in first.splitlines():
if "duckdb/duckdb-go" in line:
self.assertNotIn("indirect", line)
skill = (ROOT / "skills" / "duckdb" / "SKILL.md").read_text()
self.assertIn("github.com/duckdb/duckdb-go", skill)
self.assertIn("Ladybug", skill)
self.assertIn("sqlite", skill.lower())
self.assertIn("gcc", skill.lower())
self.assertIn("Zig", skill)
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D22", plan)
self.assertIn("duckdb-go", plan)
self._assert_shebang("bin/qa/stats.go")
reasoner = (ROOT / "internal" / "reasoner" / "client.go").read_text()
self.assertNotIn("duckdb", reasoner)
self.assertNotIn("duckstats", reasoner)
bakeoff = (ROOT / "bin" / "reasoner" / "bakeoff.go").read_text()
self.assertIn("internal/duckstats", bakeoff)
webcache = (ROOT / "internal" / "websearch" / "cache.go").read_text()
self.assertNotIn("duckdb", webcache)
self.assertIn("modernc.org/sqlite", webcache)
def test_cgo_uses_zig_not_gcc(self) -> None:
for rel in ("bin/cgo/zig", "bin/cgo/zcc", "bin/cgo/zc++"):
p = ROOT / rel
self.assertTrue(p.is_file(), f"missing {rel}")
self.assertTrue(
os.access(p, os.X_OK),
f"{rel} must be executable",
)
zig = (ROOT / "bin" / "cgo" / "zig").read_text()
self.assertIn("zig cc", zig)
self.assertIn("0.14.1", zig)
zcc = (ROOT / "bin" / "cgo" / "zcc").read_text()
self.assertIn('exec "$ZIG" cc', zcc)
self.assertNotIn("command -v gcc", zcc)
search = (ROOT / "bin" / "kb" / "search").read_text()
self.assertIn("bin/cgo/zig", search)
self.assertNotIn("command -v gcc", search)
def test_ci_recall_sot_is_zig_brain_eval(self) -> None:
ci = (ROOT / ".github" / "workflows" / "ci.yml").read_text()
self.assertIn("bin/brain/eval.go", ci)
self.assertIn("system_ladybug,brain_eval", ci)
self.assertIn("/tmp/brain-eval", ci)
self.assertIn("KB_ROOT", ci)
self.assertNotIn("bin/kb/eval", ci)
self.assertNotIn("gate skipped", ci)
self.assertIn("./bin/facts/audit self", ci)
def test_eval_fragments_live_in_default_corpus(self) -> None:
"""CI --rebuild indexes README/PLAN/docs/skills; fragments must be there."""
corpus = []
for rel in ("README.md", "PLAN.md", "AGENTS.md"):
corpus.append((ROOT / rel).read_text())
for d in ("docs", "skills"):
for p in (ROOT / d).rglob("*.md"):
corpus.append(p.read_text())
blob = "\n".join(corpus)
for frag in ("BM25", "DevOps", "LadybugDB"):
self.assertIn(frag, blob, f"{frag} must appear in default index corpus")
-104
View File
@@ -1,104 +0,0 @@
import os
import sys
import unittest
sys.path.insert(0, os.path.dirname(__file__))
from contradict import ( # noqa: E402
RULE_AUTHORITY,
RULE_SINGLE,
RULE_TEMPORAL,
RULE_TWO_SOURCE,
RULE_UNRESOLVED,
adjudicate,
check_fact_row,
parse_source_field,
)
def src(i, kind, stale=False):
return {"id": i, "kind": kind, "stale": stale}
class TestContradict(unittest.TestCase):
def test_two_vs_two_stays_hypothesis(self):
r = adjudicate({
"text": "svc listens on 443",
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
"no": [src("docker-old", "runtime"), src("compose-old", "config")],
})
self.assertFalse(r["confirmed"])
self.assertEqual(r["rule"], RULE_UNRESOLVED)
self.assertEqual(r["winner"], "")
def test_temporal_freshness(self):
r = adjudicate({
"text": "svc listens on 443",
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
"no": [src("old-readme", "narrative", True), src("old-wiki", "narrative", True)],
})
self.assertTrue(r["confirmed"])
self.assertEqual(r["rule"], RULE_TEMPORAL)
self.assertEqual(r["winner"], "yes")
def test_authority_pairing(self):
r = adjudicate({
"text": "svc listens on 443",
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
"no": [src("readme", "narrative"), src("wiki", "narrative")],
})
self.assertTrue(r["confirmed"])
self.assertEqual(r["rule"], RULE_AUTHORITY)
self.assertEqual(r["winner"], "yes")
def test_two_source_and_single(self):
two = adjudicate({
"text": "arc-1 runs Matrix",
"yes": [src("compose", "config"), src("docker-ps", "runtime")],
})
self.assertTrue(two["confirmed"])
self.assertEqual(two["rule"], RULE_TWO_SOURCE)
one = adjudicate({"text": "maybe", "yes": [src("readme", "narrative")]})
self.assertFalse(one["confirmed"])
self.assertEqual(one["rule"], RULE_SINGLE)
def test_parse_source_field(self):
yes, no = parse_source_field("docker ps x compose.yml vs old.md x wiki.md")
self.assertIn(" x ", yes)
self.assertIn(" x ", no)
def test_check_fact_row_allows_hypothesis_vs(self):
p = check_fact_row(
"L1", "a.md x b.md vs c.md x d.md", "var/", "audit", "hypothesis",
)
self.assertEqual(p, [])
p = check_fact_row("L2", "a.md x b.md", "var/", "audit", "confirmed")
self.assertEqual(p, [])
p = check_fact_row("L3", "a.md x b.md vs c.md x d.md", "var/", "audit", "confirmed")
self.assertTrue(any("vs-contradiction" in x for x in p))
p = check_fact_row("L4", "only-one.md", "var/", "audit", "hypothesis")
self.assertTrue(any("a x b vs" in x for x in p))
def test_audit_contradict_cli_unresolved(self):
import json
import subprocess
from pathlib import Path
root = Path(__file__).resolve().parents[2]
payload = json.dumps({
"text": "svc 443",
"yes": [src("a", "runtime"), src("b", "config")],
"no": [src("c", "runtime"), src("d", "config")],
})
proc = subprocess.run(
[sys.executable, str(root / "bin" / "facts" / "audit"), "contradict", "--json"],
input=payload, capture_output=True, text=True, check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
out = json.loads(proc.stdout)
self.assertTrue(out["ok"])
self.assertEqual(out["contradictions"][0]["rule"], RULE_UNRESOLVED)
self.assertFalse(out["contradictions"][0]["confirmed"])
if __name__ == "__main__":
unittest.main()
+10 -13
View File
@@ -9,6 +9,12 @@ sys.path.insert(0, str(Path(__file__).resolve().parent))
import kblib # noqa: E402
import gitimport # noqa: E402
SAMPLE = (
"\x1e" + "a1b2c3d" + "\x1f" + "Ada Lovelace" + "\x1f" + "ada@example.com"
+ "\x1f" + "2026-08-10T12:00:00+01:00" + "\x1f" + "feat: first commit"
+ "\n\nREADME.md\nsrc/main.c\n"
)
COMMIT_PERSON_SCHEMA = (
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
"author STRING, email STRING, date STRING, PRIMARY KEY(id))"
@@ -20,17 +26,6 @@ HAS_VERSION_SCHEMA = "CREATE REL TABLE IF NOT EXISTS HAS_VERSION (FROM File TO C
AUTHORED_SCHEMA = "CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
def sample_commit() -> gitimport.Commit:
return gitimport.Commit(
sha="a1b2c3d",
author="Ada Lovelace",
email="ada@example.com",
date="2026-08-10T12:00:00+01:00",
subject="feat: first commit",
files=["README.md", "src/main.c"],
)
class GitGraphTest(unittest.TestCase):
def setUp(self):
self.dir = tempfile.mkdtemp()
@@ -47,12 +42,14 @@ class GitGraphTest(unittest.TestCase):
self.db.close()
def test_index_commits_creates_nodes_and_edges(self):
gitimport.index_commits(self.conn, [sample_commit()], "sample-repo")
cs = gitimport.parse_log(SAMPLE)
gitimport.index_commits(self.conn, cs, "sample-repo")
rp = self.conn.execute("MATCH (p:Person) RETURN p.name, p.email").get_all()
self.assertEqual([tuple(r) for r in rp], [("Ada Lovelace", "ada@example.com")])
rc = self.conn.execute("MATCH (c:Commit) RETURN c.id, c.repo").get_all()
self.assertEqual(len(rc), 1)
self.assertEqual(rc[0][1], "sample-repo")
# File -[:HAS_VERSION]-> Commit -[:AUTHORED]-> Person
rf = self.conn.execute(
"MATCH (f:File)-[:HAS_VERSION]->(c:Commit)-[:AUTHORED]->(p:Person) "
"RETURN f.path, c.id, p.email").get_all()
@@ -61,7 +58,7 @@ class GitGraphTest(unittest.TestCase):
self.assertTrue(all(r[2] == "ada@example.com" for r in rf))
def test_index_commits_idempotent(self):
cs = [sample_commit()]
cs = gitimport.parse_log(SAMPLE)
gitimport.index_commits(self.conn, cs, "sample-repo")
gitimport.index_commits(self.conn, cs, "sample-repo")
n = self.conn.execute("MATCH (c:Commit) RETURN count(*)").get_all()[0][0]
+57
View File
@@ -0,0 +1,57 @@
import sys
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import gitimport # noqa: E402
SAMPLE = (
"\x1e" + "a1b2c3d" + "\x1f" + "Ada Lovelace" + "\x1f" + "ada@example.com"
+ "\x1f" + "2026-08-10T12:00:00+01:00" + "\x1f" + "feat: first commit"
+ "\n\nREADME.md\nsrc/main.c\n"
+ "\x1e" + "e4f5a6b" + "\x1f" + "Bob Babbage" + "\x1f" + "bob@example.com"
+ "\x1f" + "2026-08-11T09:30:00+01:00" + "\x1f" + "fix: typo"
+ "\n\ndocs/notes.md"
)
class GitparseTest(unittest.TestCase):
def test_parses_records(self):
cs = gitimport.parse_log(SAMPLE)
self.assertEqual(len(cs), 2)
def test_parses_commit_fields(self):
cs = gitimport.parse_log(SAMPLE)
c = cs[0]
self.assertEqual(c.sha, "a1b2c3d")
self.assertEqual(c.author, "Ada Lovelace")
self.assertEqual(c.email, "ada@example.com")
self.assertEqual(c.date, "2026-08-10T12:00:00+01:00")
self.assertEqual(c.subject, "feat: first commit")
def test_parses_changed_files(self):
cs = gitimport.parse_log(SAMPLE)
self.assertEqual(cs[0].files, ["README.md", "src/main.c"])
self.assertEqual(cs[1].files, ["docs/notes.md"])
def test_ignores_empty(self):
self.assertEqual(gitimport.parse_log(""), [])
def test_skip_malformed_record(self):
self.assertEqual(gitimport.parse_log("\x1eweird\x1e"), [])
def test_commit_leaf_shape(self):
leafs = gitimport.commits_to_leafs(gitimport.parse_log(SAMPLE), "sample-repo")
self.assertEqual(len(leafs), 2)
lf = leafs[0]
self.assertEqual(lf["type"], "commit")
self.assertEqual(lf["repo"], "sample-repo")
self.assertEqual(lf["source"], "sample-repo@a1b2c3d")
self.assertIn("Ada Lovelace", lf["text"])
self.assertIn("README.md", lf["related"])
self.assertIn("feat: first commit", lf["heading"])
if __name__ == "__main__":
unittest.main()
-97
View File
@@ -1,97 +0,0 @@
"""Import adapters write files only. Index rebuild is brain/index (D14 / Gitea #7)."""
from __future__ import annotations
import json
import os
import subprocess
import sys
import tempfile
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class IndexAdapterTest(unittest.TestCase):
def test_dry_run_fixture_corpus_does_not_write_lbug(self) -> None:
tmp = Path(tempfile.mkdtemp())
(tmp / "note.md").write_text("# Fixture\n\n## Leaf\n\nhello corpus\n", encoding="utf-8")
lbug = tmp / "kb.lbug"
try:
import ladybug # noqa: F401
except ImportError:
venv_py = ROOT / ".venv" / "bin" / "python"
if not venv_py.is_file():
self.skipTest("ladybug missing")
py = str(venv_py)
else:
py = sys.executable
proc = subprocess.run(
[py, str(ROOT / "bin" / "kb" / "index"), "--dry-run", "--json", "--corpus", str(tmp)],
cwd=ROOT,
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
msg = json.loads(proc.stdout)
self.assertTrue(msg.get("dry_run"))
self.assertGreaterEqual(msg.get("corpus_total", 0), 1)
self.assertFalse(lbug.exists(), "dry-run must not create a Ladybug file")
def test_facts_json_and_chats_land_on_rebuild(self) -> None:
"""Gitea #18: facts (2-source) + chats markdown become leafs on rebuild."""
tmp = Path(tempfile.mkdtemp())
dbpath = tmp / "kb.lbug"
chats = tmp / "chats"
chats.mkdir()
(chats / "alice.md").write_text(
"# Chat\n\n## Alice and Bob\n\nhello from chats fixture unique-chat-token\n",
encoding="utf-8",
)
facts_path = tmp / "facts.json"
facts_path.write_text(json.dumps([{
"text": "container 'brain' unique-fact-token is running and declared in compose.yaml",
"source": "docker ps x compose.yaml",
"loc": "compose.yaml:brain",
"how": "facts/extract",
}]), encoding="utf-8")
venv_py = ROOT / ".venv" / "bin" / "python"
py = str(venv_py) if venv_py.is_file() else sys.executable
proc = subprocess.run(
[
py, str(ROOT / "bin" / "kb" / "index"),
"--rebuild", "--db", str(dbpath), "--no-defaults",
"--with-chats", str(chats),
"--facts-json", str(facts_path),
"--json",
],
cwd=ROOT,
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
msg = json.loads(proc.stdout)
self.assertGreaterEqual(msg.get("facts_leafs", 0), 1)
self.assertGreaterEqual(msg.get("chat_leafs", 0), 1)
self.assertTrue(dbpath.exists())
sys.path.insert(0, str(ROOT / "bin" / "tools"))
import kblib
db, conn = kblib.connect(dbpath, read_only=True)
try:
stats = kblib.stats(conn)
self.assertGreaterEqual(stats["by_root"].get("facts", 0), 1)
fts = kblib.query_fts(conn, "unique-chat-token", 5)
self.assertTrue(fts, "chats markdown must be FTS-searchable")
fact_hits = kblib.query_fts(conn, "unique-fact-token", 5)
self.assertTrue(any(h.get("root") == "facts" for h in fact_hits))
src = conn.execute(
"MATCH (l:Leaf {root:'facts'}) RETURN l.source"
).get_all()
self.assertTrue(any(" x " in str(r[0]) for r in src))
finally:
conn.close()
db.close()
-69
View File
@@ -1,69 +0,0 @@
"""Incremental add writes leafs without deleting kb.lbug."""
from __future__ import annotations
import json
import os
import subprocess
import sys
import tempfile
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class KbAddCLITest(unittest.TestCase):
def test_json_add_does_not_delete_db(self) -> None:
tmp = Path(tempfile.mkdtemp())
dbpath = tmp / "kb.lbug"
py = sys.executable
venv_py = ROOT / ".venv" / "bin" / "python"
if venv_py.is_file():
py = str(venv_py)
payload = {
"text": "cli zebra leaf",
"root": "info",
"source": "cli-test",
"confidence": "confirmed",
"how": "test",
"loc": str(tmp),
"type": "reference",
"embedding": [0.0] * 256,
}
payload["embedding"][0] = 0.3
proc = subprocess.run(
[py, str(ROOT / "bin" / "kb" / "add"), "--db", str(dbpath), "--json"],
cwd=ROOT,
input=json.dumps(payload),
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
self.assertTrue(dbpath.exists(), "add must create the db, not skip write")
out = json.loads(proc.stdout)
self.assertEqual(out.get("mode"), "add")
self.assertEqual(len(out.get("ids") or []), 1)
again = subprocess.run(
[py, str(ROOT / "bin" / "kb" / "add"), "--db", str(dbpath), "--json"],
cwd=ROOT,
input=json.dumps({
**payload,
"text": "second moose leaf",
"source": "cli-test-2",
}),
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(again.returncode, 0, again.stderr)
self.assertTrue(dbpath.exists())
second = json.loads(again.stdout)
self.assertEqual(len(second.get("ids") or []), 1)
self.assertNotEqual(out["ids"][0], second["ids"][0])
if __name__ == "__main__":
unittest.main()
-128
View File
@@ -71,69 +71,6 @@ class KblibTest(unittest.TestCase):
self.assertTrue(hits)
self.assertIn("Leaf_vec", kblib.leaf_index_names(self.conn))
def test_add_after_indexes_keeps_fts_queryable(self):
"""Incremental add after FTS+HNSW must find the new leaf on both indexes."""
kblib.upsert_leaf(self.conn, text="seed fox leaf", root="info",
confidence="confirmed", source="s", source_rev="r1",
how="test", loc="/tmp", type_="reference",
embedding=make_emb(0.1))
kblib.ensure_indexes(self.conn)
ids = kblib.add_leafs(self.conn, [{
"text": "added zebra after index",
"root": "facts",
"confidence": "confirmed",
"source": "a.md x b.md",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "fact",
"embedding": make_emb(0.9),
}])
self.assertEqual(len(ids), 1)
fts = kblib.query_fts(self.conn, "zebra", 5)
self.assertTrue(fts)
self.assertIn("zebra", fts[0]["text"])
self.assertEqual(fts[0]["root"], "facts")
vec = kblib.query_vector(self.conn, make_emb(0.9), 5)
self.assertTrue(any("zebra" in h["text"] for h in vec))
fox = kblib.query_fts(self.conn, "fox", 5)
self.assertTrue(fox)
self.assertIn("fox", fox[0]["text"])
def test_add_facts_and_info_one_transaction(self):
"""D12: facts and info land in the same transaction."""
kblib.ensure_indexes(self.conn)
ids = kblib.add_leafs(self.conn, [
{
"text": "tx fact leaf two-source",
"root": "facts",
"confidence": "confirmed",
"source": "compose.yml x docker ps",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "fact",
"embedding": make_emb(0.4),
},
{
"text": "tx info narrative",
"root": "info",
"confidence": "confirmed",
"source": "note.md",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "reference",
"embedding": make_emb(0.5),
},
])
self.assertEqual(len(ids), 2)
stats = kblib.stats(self.conn)
self.assertEqual(stats["by_root"].get("facts"), 1)
self.assertEqual(stats["by_root"].get("info"), 1)
self.assertTrue(kblib.query_fts(self.conn, "two-source", 5))
self.assertTrue(kblib.query_fts(self.conn, "narrative", 5))
def test_drop_vector_then_create_raises_clear_error(self):
"""DROP INDEX leaves ghost catalog; create_fts_and_vector must raise."""
kblib.upsert_leaf(self.conn, text="seed", root="info",
@@ -161,71 +98,6 @@ class KblibTest(unittest.TestCase):
self.assertEqual(stats["total"], 2)
self.assertEqual(stats["by_root"], {"facts": 1, "info": 1})
def test_hop_1_returns_file_hop_3_reaches_person(self):
"""--hop walks FROM_FILE / HAS_VERSION / AUTHORED (Gitea #17)."""
import gitimport
lid = kblib.upsert_leaf(
self.conn, text="readme hop fixture", root="info",
confidence="confirmed", source="README.md", source_rev="r1",
how="test", loc="README.md", type_="reference",
embedding=make_emb(0.3),
)
kblib.link_from_file(self.conn, lid, "README.md", repo="sample-repo")
gitimport.index_commits(self.conn, [gitimport.Commit(
sha="a1b2c3d",
author="Ada Lovelace",
email="ada@example.com",
date="2026-08-10T12:00:00Z",
subject="feat: first commit",
files=["README.md"],
)], "sample-repo")
hop1 = kblib.hop_walk(self.conn, lid, 1)
self.assertEqual(len(hop1), 1)
self.assertEqual(hop1[0]["label"], "File")
self.assertEqual(hop1[0]["name"], "README.md")
self.assertEqual(hop1[0]["depth"], 1)
hop3 = kblib.hop_walk(self.conn, lid, 3)
labels = {n["label"] for n in hop3}
self.assertIn("File", labels)
self.assertIn("Commit", labels)
self.assertIn("Person", labels)
person = [n for n in hop3 if n["label"] == "Person"][0]
self.assertEqual(person["name"], "Ada Lovelace")
self.assertEqual(person["depth"], 3)
def test_as_of_keeps_x_drops_y(self) -> None:
"""OQ5/#36: as of 2025-01-01 → works-at-X, not works-at-Y."""
kblib.upsert_leaf(
self.conn, text="Andrey works at X", root="facts",
confidence="confirmed", source="crm.md x contract.md",
source_rev="r1", how="test", loc="/tmp", type_="fact",
embedding=make_emb(0.5),
valid_from="2024-03-01", valid_to="2025-07-15",
)
kblib.upsert_leaf(
self.conn, text="Andrey works at Y", root="facts",
confidence="confirmed", source="offer.md x payroll.md",
source_rev="r1", how="test", loc="/tmp", type_="fact",
embedding=make_emb(0.6),
valid_from="2025-07-16", valid_to="",
)
kblib.ensure_indexes(self.conn)
hits = kblib.query_fts(self.conn, "Andrey works", 10)
kept = kblib.filter_as_of(hits, "2025-01-01")
texts = [h["text"] for h in kept]
self.assertTrue(any("works at X" in t for t in texts), texts)
self.assertFalse(any("works at Y" in t for t in texts), texts)
later = kblib.filter_as_of(hits, "2025-08-01")
later_texts = [h["text"] for h in later]
self.assertTrue(any("works at Y" in t for t in later_texts), later_texts)
self.assertFalse(any("works at X" in t for t in later_texts), later_texts)
def test_active_at_pure(self) -> None:
self.assertTrue(kblib.active_at("2024-03-01", "2025-07-15", "2025-01-01"))
self.assertFalse(kblib.active_at("2025-07-16", "", "2025-01-01"))
self.assertTrue(kblib.active_at("", "", "2025-01-01"))
if __name__ == "__main__":
unittest.main()
-89
View File
@@ -8,13 +8,10 @@ from pathlib import Path
sys.path.insert(0, os.path.dirname(__file__))
from mailconv import ( # noqa: E402
TESS_LANG,
clean_email_address,
convert_pdf,
html_to_markdown,
is_convertible,
normalize_markdown,
ocr_image,
split_zip_members,
subject_to_filename,
zip_extract_safe,
@@ -103,92 +100,6 @@ class TestMailConv(unittest.TestCase):
self.assertFalse(is_convertible(".exe"))
self.assertFalse(is_convertible(".unknown"))
def test_convert_pdf_prefers_pdftotext(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b"Invoice BM25 layout"
stderr = b""
return P()
self._patch_run(mc, fake_run)
out = convert_pdf(Path(self._tmp("born.pdf")))
self.assertIn("BM25", out)
self.assertEqual(calls[0][:2], ["pdftotext", "-layout"])
self.assertFalse(any(c[0] == "tesseract" for c in calls))
self.assertFalse(any(c[0] == "pdftoppm" for c in calls))
def test_convert_pdf_empty_layer_uses_pdftoppm_tesseract(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b""
stderr = b""
if cmd[0] == "pdftotext":
P.stdout = b" \n"
return P()
if cmd[0] == "pdftoppm":
prefix = Path(cmd[-1])
(prefix.parent / "page-1.png").write_bytes(b"fake")
return P()
if cmd[0] == "tesseract":
P.stdout = b"scanned HELLO"
return P()
return P()
self._patch_run(mc, fake_run)
out = convert_pdf(Path(self._tmp("scan.pdf")))
self.assertIn("HELLO", out)
bins = [c[0] for c in calls]
self.assertIn("pdftotext", bins)
self.assertIn("pdftoppm", bins)
self.assertIn("tesseract", bins)
tess = next(c for c in calls if c[0] == "tesseract")
self.assertIn(TESS_LANG, tess)
self.assertNotIn("docling", " ".join(bins))
def test_ocr_image_paddle_engine(self):
import mailconv as mc
calls: list[list[str]] = []
def fake_run(cmd, **kwargs):
calls.append(list(cmd))
class P:
returncode = 0
stdout = b"paddle text"
stderr = b""
return P()
self._patch_run(mc, fake_run)
os.environ["OCR_ENGINE"] = "paddle"
try:
out = ocr_image(Path(self._tmp("x.png")))
finally:
os.environ.pop("OCR_ENGINE", None)
self.assertEqual(out, "paddle text")
self.assertEqual(calls[0][:2], ["paddleocr", "ocr"])
def _patch_run(self, mod, fn) -> None:
self.addCleanup(setattr, mod.subprocess, "run", mod.subprocess.run)
mod.subprocess.run = fn
def _mk_zip(self, members):
zpath = Path(self._tmp("arc.zip"))
with zipfile.ZipFile(zpath, "w") as zf:
-46
View File
@@ -1,46 +0,0 @@
"""Mail markdown → leafs (no Ladybug). Brain index --with-mail uses this."""
from __future__ import annotations
import json
import sys
import tempfile
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import mailleafs # noqa: E402
class MailLeafsTest(unittest.TestCase):
def test_message_md_becomes_info_leaf(self) -> None:
root = Path(tempfile.mkdtemp())
msg = root / "inbox" / "alice-1"
msg.mkdir(parents=True)
(msg / "message.json").write_text(
json.dumps({"receivedDate": "2026-01-15T10:00:00Z", "subject": "Hello"}),
encoding="utf-8",
)
(msg / "message.md").write_text(
"---\nroot: info\n---\n\n# Hello\n\nFrom Alice to Bob.\n",
encoding="utf-8",
)
leafs = mailleafs.from_mail_root(root)
self.assertEqual(len(leafs), 1)
self.assertIn("Alice", leafs[0]["text"])
self.assertTrue(leafs[0]["source"].startswith("ooMail:"))
self.assertEqual(leafs[0]["how"], "mail/import")
def test_since_filters_by_message_json_date(self) -> None:
root = Path(tempfile.mkdtemp())
for name, day in (("old", "2025-01-01"), ("new", "2026-06-01")):
d = root / "inbox" / name
d.mkdir(parents=True)
(d / "message.json").write_text(
json.dumps({"receivedDate": f"{day}T00:00:00Z"}),
encoding="utf-8",
)
(d / "message.md").write_text(f"# {name}\n\nbody\n", encoding="utf-8")
leafs = mailleafs.from_mail_root(root, since="2026-01-01")
self.assertEqual(len(leafs), 1)
self.assertIn("new", leafs[0]["text"])
-206
View File
@@ -1,206 +0,0 @@
"""Published docs must match live commands (Gitea SoT, brain/search)."""
from __future__ import annotations
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class PublishedDocsTest(unittest.TestCase):
def test_readme_points_issues_at_gitea(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn(
"https://git.produktor.io/eSlider/2dph/issues",
text,
"README must point issues at Gitea",
)
def test_plan_d15_names_gitea_origin(self) -> None:
text = (ROOT / "PLAN.md").read_text()
self.assertIn("D15", text)
self.assertIn("git.produktor.io/eSlider/2dph", text)
def test_readme_primary_search_is_brain(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn(
"bin/brain/search.go",
text,
"README deduction search must name bin/brain/search.go",
)
def test_readme_index_is_brain_not_index_mail(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn("bin/brain/index.go", text)
self.assertNotIn(
"bin/mail/index_mail",
text,
"mail index is a brain write; README must name bin/brain/index.go",
)
def test_readme_git_import_is_gogit(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn("bin/git/import.go", text)
self.assertIn("go-git", text)
self.assertIn("D19", (ROOT / "PLAN.md").read_text())
def test_web_search_is_go_not_ops_host(self) -> None:
readme = (ROOT / "README.md").read_text()
self.assertIn("bin/web/search.go", readme)
skill = (ROOT / "skills" / "web-search" / "SKILL.md").read_text()
self.assertIn("bin/web/search.go", skill)
self.assertNotIn("search.ops.io", skill)
self.assertNotIn("search.ops.io", readme)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn("searxng", compose)
self.assertNotIn("search.ops.io", compose)
settings = (ROOT / "deploy" / "searxng" / "settings.yml").read_text()
self.assertNotIn("password", settings.lower())
self.assertIn("json", settings)
def test_picoclaw_compose_profile_has_mcp_example(self) -> None:
compose = (ROOT / "compose.yaml").read_text()
self.assertIn('profiles: ["picoclaw"]', compose)
self.assertIn("127.0.0.1:8630", compose)
example = (ROOT / "deploy" / "picoclaw" / "mcp.json.example").read_text()
self.assertIn("127.0.0.1:8630/mcp", example)
self.assertNotIn("password", example.lower())
self.assertNotIn("token", example.lower())
docs = (ROOT / "docs" / "picoclaw.md").read_text()
self.assertIn("search", docs)
self.assertIn("throttled", docs)
def test_readme_read_path_is_go(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("get.go", plan)
self.assertIn("CI fallback", plan)
design = (ROOT / "docs" / "design.md").read_text()
self.assertIn("internal/brain/rank", design)
self.assertIn("They do not exec Python", design)
def test_openapi_mcp_from_same_handlers(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D20", plan)
self.assertIn("/openapi.json", (ROOT / "README.md").read_text())
self.assertIn("/mcp", (ROOT / "README.md").read_text())
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
self.assertIn("/mcp", skill)
self.assertFalse((ROOT / "skills" / "db-yaml").exists())
self.assertTrue((ROOT / "skills" / "postgres" / "SKILL.md").is_file())
def test_cgo_zig_and_index_profile(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D21", plan)
self.assertIn("zig cc", plan)
dockerfile = (ROOT / "Dockerfile").read_text()
self.assertIn("bin/cgo/zcc", dockerfile)
self.assertIn("FROM debian:bookworm-slim AS api", dockerfile)
self.assertIn("FROM python:3.12-slim AS index", dockerfile)
api = dockerfile[dockerfile.index("FROM debian:bookworm-slim AS api") :]
self.assertNotIn("pip install", api)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn('profiles: ["index"]', compose)
self.assertIn("target: api", compose)
def test_reasoner_docs_name_real_hf_ids_cpu_sidecar(self) -> None:
docs = (ROOT / "docs" / "reasoner.md").read_text()
for hf in (
"Qwen/Qwen3.5-9B",
"Qwen/Qwen3.6-27B",
"prism-ml/Bonsai-27B-gguf",
):
self.assertIn(hf, docs)
self.assertIn("no official qwen3.6-9b", docs.lower())
self.assertIn("OLLAMA_NUM_GPU", docs)
self.assertIn("rss_mb", docs)
self.assertIn("vram_mb", docs)
self.assertIn("3/3", docs)
self.assertIn("Do not claim 9B is better at tools", docs)
self.assertNotIn("Qwen/Qwen3.6-9B", docs)
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("D18", plan)
self.assertIn("Qwen/Qwen3.5-9B", plan)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn('"reasoner"', compose)
self.assertIn("OLLAMA_NUM_GPU", compose)
self.assertIn("127.0.0.1:11435", compose)
dockerfile = (ROOT / "Dockerfile").read_text()
self.assertNotIn(".gguf", dockerfile.lower())
self.assertNotIn(".safetensors", dockerfile.lower())
api = dockerfile[dockerfile.index("FROM debian:bookworm-slim AS api") :]
self.assertNotIn("COPY models", api)
self.assertNotIn("qwen", api.lower())
def test_readme_search_escalates_web(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn("--no-web", text)
self.assertIn("D17", (ROOT / "PLAN.md").read_text())
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
self.assertIn("`web` block", skill)
def test_docs_say_hop_walks_from_file(self) -> None:
paths = [
ROOT / "README.md",
ROOT / "docs" / "design.md",
ROOT / "skills" / "brain" / "SKILL.md",
ROOT / "docs" / "runbook.md",
ROOT / "docs" / "README.md",
]
for path in paths:
text = path.read_text()
self.assertIn("--hop", text, f"{path.relative_to(ROOT)} must document --hop")
self.assertNotIn(
"not implemented",
text.lower(),
f"{path.relative_to(ROOT)} still says hop is not implemented",
)
def test_docs_are_portable_diataxis(self) -> None:
index = (ROOT / "docs" / "README.md").read_text()
self.assertIn("type: reference", index)
for d in ("D3", "D6", "D14", "D15", "D17", "D18"):
self.assertIn(d, index)
runbook = (ROOT / "docs" / "runbook.md").read_text()
self.assertIn("type: howto", runbook)
self.assertIn("bin/brain/search.go", runbook)
self.assertIn("bin/brain/index.go", runbook)
self.assertNotIn("search.ops.io", runbook)
self.assertNotIn("/mnt/", runbook)
self.assertNotIn("/home/", runbook)
readme = (ROOT / "README.md").read_text()
self.assertIn("docs/runbook.md", readme)
self.assertNotIn("search.ops.io", readme)
def test_v1_epic_is_named_in_docs(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("Gap to v1", plan)
self.assertIn("eSlider/2dph/issues/16", plan)
self.assertIn("eSlider/2dph/issues/17", plan)
self.assertIn("eSlider/2dph/milestone/12", plan)
road = (ROOT / "docs" / "roadmap.md").read_text()
self.assertIn("type: explanation", road)
self.assertIn("issues/16", road)
self.assertIn("issues/14", road)
index = (ROOT / "docs" / "README.md").read_text()
self.assertIn("roadmap.md", index)
self.assertIn("epic #16", index)
agents = (ROOT / "AGENTS.md").read_text()
self.assertIn("roadmap.md", agents)
def test_oq5_intervals_are_not_d16_stale(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
design = (ROOT / "docs" / "design.md").read_text()
road = (ROOT / "docs" / "roadmap.md").read_text()
contradict = (ROOT / "internal" / "facts" / "contradict.go").read_text()
interval = (ROOT / "internal" / "facts" / "interval.go").read_text()
self.assertIn("OQ5", plan)
self.assertIn("D24", plan)
self.assertIn("issues/36", plan)
self.assertIn("valid_from", plan)
self.assertIn("--as-of", design)
self.assertIn("issues/36", road)
self.assertIn("**in**", road[road.index("OQ5"):road.index("OQ5") + 80])
self.assertIn("temporal_freshness", contradict)
self.assertNotIn("valid_from", contradict)
self.assertIn("ActiveAt", interval)
self.assertIn("NormalizeDay", interval)
-63
View File
@@ -1,63 +0,0 @@
"""Skills must name live commands; every bin/ path in SKILL.md must exist."""
from __future__ import annotations
import re
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
BIN_PATH = re.compile(r"(bin/[A-Za-z0-9_./-]+)")
class SkillsTest(unittest.TestCase):
def test_db_yaml_renamed_to_postgres(self) -> None:
self.assertFalse(
(ROOT / "skills" / "db-yaml").exists(),
"skills/db-yaml must be skills/postgres",
)
self.assertTrue((ROOT / "skills" / "postgres" / "SKILL.md").is_file())
text = (ROOT / "skills" / "postgres" / "SKILL.md").read_text()
self.assertIn("bin/postgres/query.go", text)
self.assertNotIn("search.ops.io", text)
def test_every_bin_path_in_skills_exists(self) -> None:
missing: list[str] = []
for path in (ROOT / "skills").rglob("SKILL.md"):
text = path.read_text()
for m in BIN_PATH.finditer(text):
rel = m.group(1).rstrip(")`.,;")
candidate = ROOT / rel
if not candidate.exists():
missing.append(f"{path.relative_to(ROOT)}: {rel}")
self.assertEqual(missing, [], "skill bin paths must exist")
def test_brain_skill_lists_generated_tools(self) -> None:
tools = (ROOT / "skills" / "brain" / "tools.md").read_text()
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
self.assertIn("tools.md", skill)
for name in ("search", "get", "stats", "audit"):
self.assertIn(f"`{name}`", tools)
def test_picoclaw_lists_tool_order(self) -> None:
skill = (ROOT / "skills" / "picoclaw" / "SKILL.md").read_text()
agents = (ROOT / "AGENTS.md").read_text()
self.assertIn("**`search`**", skill)
self.assertIn("**`get`**", skill)
self.assertIn("**`audit`**", skill)
self.assertIn("throttled", skill.lower())
self.assertIn("not a negative finding", agents)
self.assertIn("Fact-check every", agents)
def test_yq_is_mikefarah_for_structured_data(self) -> None:
skill = (ROOT / "skills" / "yq" / "SKILL.md").read_text()
self.assertIn("https://github.com/mikefarah/yq", skill)
for fmt in ("YAML", "JSON", "XML", "CSV", "TOML", "HCL"):
self.assertIn(fmt, skill)
self.assertIn("not kislyuk", skill.lower())
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("mikefarah/yq", plan)
agents = (ROOT / "AGENTS.md").read_text()
self.assertIn("mikefarah/yq", agents)
web = (ROOT / "skills" / "web-search" / "SKILL.md").read_text()
self.assertIn("| yq ", web)
self.assertNotIn("| jq ", web)
-33
View File
@@ -1,33 +0,0 @@
"""Every bin/ path named in skills/ must exist on disk."""
from __future__ import annotations
import re
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
BIN_PATH = re.compile(r"\b(bin/[A-Za-z0-9_./-]+)")
class SkillsBinPathsTest(unittest.TestCase):
def test_agent_cost_skill_is_gone(self) -> None:
self.assertFalse(
(ROOT / "skills" / "agent-cost").exists(),
"skills/agent-cost documents bin/agents/cost which does not exist",
)
def test_brain_skill_replaces_kb_search(self) -> None:
self.assertTrue((ROOT / "skills" / "brain" / "SKILL.md").is_file())
self.assertFalse((ROOT / "skills" / "kb-search").exists())
def test_skill_bin_paths_exist(self) -> None:
missing: list[str] = []
for skill in sorted((ROOT / "skills").rglob("SKILL.md")):
text = skill.read_text()
for match in BIN_PATH.findall(text):
rel = match.rstrip("`'.,")
if rel.endswith(".go") or Path(rel).suffix == "" or Path(rel).suffix in {".go", ".py"}:
p = ROOT / rel
if not p.exists():
missing.append(f"{skill.relative_to(ROOT)}: {rel}")
self.assertEqual(missing, [], "SKILL.md names bin/ paths that do not exist")
-213
View File
@@ -1,213 +0,0 @@
"""bin/stack/{start,start-assistant,stop,status} — offline contract + fake PATH."""
from __future__ import annotations
import os
import stat
import subprocess
import tempfile
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
METHODS = ("start", "start-assistant", "stop", "status")
class StackLayoutTest(unittest.TestCase):
def test_methods_are_bash_with_usage_comment(self) -> None:
lib = ROOT / "bin" / "stack" / "lib.sh"
self.assertTrue(lib.is_file(), "missing bin/stack/lib.sh")
for name in METHODS:
p = ROOT / "bin" / "stack" / name
self.assertTrue(p.is_file(), f"missing bin/stack/{name}")
self.assertTrue(os.access(p, os.X_OK), f"bin/stack/{name} must be executable")
lines = p.read_text().splitlines()
self.assertEqual(lines[0], "#!/usr/bin/env bash", name)
self.assertTrue(lines[1].startswith("# bin/stack/"), name)
text = "\n".join(lines)
self.assertIn("lib.sh", text, name)
self.assertNotIn("/mnt/", text, name)
self.assertNotIn("/home/", text, name)
def test_lib_has_no_host_paths_or_secrets(self) -> None:
lib = (ROOT / "bin" / "stack" / "lib.sh").read_text()
self.assertIn("stack_start", lib)
self.assertIn("stack_start_assistant", lib)
self.assertIn("stack_stop", lib)
self.assertIn("stack_status", lib)
self.assertIn("qwen3.5:9b", lib)
self.assertIn("picoclaw agent", lib)
self.assertIn("--no-deps", lib)
self.assertIn("tools/list", lib)
self.assertNotIn("/mnt/", lib)
self.assertNotIn("/home/", lib)
self.assertNotIn("password", lib.lower())
self.assertNotIn("GITEA_TOKEN", lib)
def test_start_does_not_launch_picoclaw(self) -> None:
start = (ROOT / "bin" / "stack" / "start").read_text()
self.assertIn("stack_start", start)
self.assertNotIn("stack_start_assistant", start)
self.assertNotIn("picoclaw agent", start)
def test_start_assistant_attaches_agent(self) -> None:
src = (ROOT / "bin" / "stack" / "start-assistant").read_text()
self.assertIn("stack_start_assistant", src)
self.assertIn("--no-attach", src)
def test_stop_does_not_down_volumes(self) -> None:
lib = (ROOT / "bin" / "stack" / "lib.sh").read_text()
self.assertIn(" compose ", lib)
self.assertRegex(lib, r"\bstop\b")
self.assertNotIn(" compose down", lib)
self.assertNotIn("compose down", lib)
def test_docs_name_stack_commands(self) -> None:
runbook = (ROOT / "docs" / "runbook.md").read_text()
pico = (ROOT / "docs" / "picoclaw.md").read_text()
agents = (ROOT / "AGENTS.md").read_text()
for text in (runbook, pico, agents):
self.assertIn("bin/stack/start", text)
self.assertIn("bin/stack/start-assistant", text)
self.assertIn("bin/stack/status", text)
self.assertIn("bin/stack/stop", text)
def test_help_prints_comments_not_source(self) -> None:
r = subprocess.run(
[str(ROOT / "bin" / "stack" / "start"), "--help"],
cwd=str(ROOT),
capture_output=True,
text=True,
check=False,
)
self.assertEqual(r.returncode, 0, r.stderr)
self.assertIn("bin/stack/start", r.stdout)
self.assertNotIn("set -euo pipefail", r.stdout)
self.assertNotIn("source ", r.stdout)
class StackFakePathTest(unittest.TestCase):
def _fake_bin(self, tmp: Path, *, health_ok: bool) -> Path:
bindir = tmp / "bin"
bindir.mkdir()
curl = bindir / "curl"
docker = bindir / "docker"
log = tmp / "docker.log"
curl.write_text(
f"""#!/usr/bin/env bash
url=""
for a in "$@"; do
case "$a" in http*) url=$a ;;
esac
done
if [ "{int(health_ok)}" = "0" ] && [[ "$url" == */health ]]; then
echo '{{"status":"down"}}'
exit 7
fi
case "$url" in
*/mcp)
echo '{{"jsonrpc":"2.0","id":1,"result":{{"tools":[{{"name":"search"}},{{"name":"get"}},{{"name":"audit"}}]}}}}'
;;
*/api/tags)
echo '{{"models":[{{"name":"qwen3.5:9b"}}]}}'
;;
*/api/pull)
echo '{{"status":"success"}}'
;;
*)
echo '{{"status":"ok"}}'
;;
esac
"""
)
docker.write_text(
f"""#!/usr/bin/env bash
echo "$*" >> "{log}"
exit 0
"""
)
curl.chmod(curl.stat().st_mode | stat.S_IEXEC)
docker.chmod(docker.stat().st_mode | stat.S_IEXEC)
return bindir
def _env(self, bindir: Path) -> dict[str, str]:
env = os.environ.copy()
env["PATH"] = f"{bindir}:{env.get('PATH', '')}"
env["STACK_WAIT_SECS"] = "1"
env["STACK_WAIT_INTERVAL"] = "0"
return env
def test_start_skips_compose_when_brain_healthy(self) -> None:
with tempfile.TemporaryDirectory() as raw:
tmp = Path(raw)
bindir = self._fake_bin(tmp, health_ok=True)
log = tmp / "docker.log"
r = subprocess.run(
[str(ROOT / "bin" / "stack" / "start")],
cwd=str(ROOT),
env=self._env(bindir),
capture_output=True,
text=True,
check=False,
)
self.assertEqual(r.returncode, 0, r.stderr)
self.assertFalse(log.exists(), "healthy brain must not docker compose up")
def test_start_ups_brain_when_unhealthy(self) -> None:
with tempfile.TemporaryDirectory() as raw:
tmp = Path(raw)
bindir = self._fake_bin(tmp, health_ok=False)
log = tmp / "docker.log"
r = subprocess.run(
[str(ROOT / "bin" / "stack" / "start")],
cwd=str(ROOT),
env=self._env(bindir),
capture_output=True,
text=True,
check=False,
)
self.assertNotEqual(r.returncode, 0, "unhealthy brain without recovering compose must fail")
self.assertTrue(log.exists(), r.stderr)
logged = log.read_text()
self.assertIn("up -d", logged)
self.assertIn("brain", logged)
self.assertNotIn("picoclaw", logged)
def test_stop_stops_named_services(self) -> None:
with tempfile.TemporaryDirectory() as raw:
tmp = Path(raw)
bindir = self._fake_bin(tmp, health_ok=True)
log = tmp / "docker.log"
r = subprocess.run(
[str(ROOT / "bin" / "stack" / "stop")],
cwd=str(ROOT),
env=self._env(bindir),
capture_output=True,
text=True,
check=False,
)
self.assertEqual(r.returncode, 0, r.stderr)
logged = log.read_text()
self.assertIn("stop", logged)
self.assertNotIn(" down", logged)
for svc in ("brain", "brain-mcp", "reasoner", "picoclaw"):
self.assertIn(svc, logged)
def test_start_assistant_no_attach_starts_picoclaw(self) -> None:
with tempfile.TemporaryDirectory() as raw:
tmp = Path(raw)
bindir = self._fake_bin(tmp, health_ok=True)
log = tmp / "docker.log"
r = subprocess.run(
[str(ROOT / "bin" / "stack" / "start-assistant"), "--no-attach"],
cwd=str(ROOT),
env=self._env(bindir),
capture_output=True,
text=True,
check=False,
)
self.assertEqual(r.returncode, 0, r.stderr + r.stdout)
logged = log.read_text() if log.exists() else ""
self.assertIn("picoclaw", logged)
self.assertIn("--no-deps", logged)
self.assertNotIn("picoclaw agent", logged)
-36
View File
@@ -1,36 +0,0 @@
"""qa/system_perf.py is an offline-gated system test (no live brain in CI)."""
from __future__ import annotations
import ast
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class SystemPerfScriptTest(unittest.TestCase):
def test_script_compiles_and_is_read_only(self) -> None:
path = ROOT / "qa" / "system_perf.py"
src = path.read_text()
compile(src, str(path), "exec")
self.assertIn("--json", src)
self.assertIn("qwen3.5:9b", src)
self.assertIn("--picoclaw", src)
self.assertIn("BRAIN_URL", src)
self.assertIn("tools/list", src)
self.assertIn("tools/call", src)
self.assertIn("GATE_HEALTH_MS", src)
self.assertIn("GATE_GET_P50_MS", src)
self.assertNotIn("kb.lbug", src)
self.assertNotIn("password", src.lower())
self.assertNotIn("token", src.lower())
def test_script_does_not_write_ladybug(self) -> None:
tree = ast.parse((ROOT / "qa" / "system_perf.py").read_text())
writes = [
n.func.attr
for n in ast.walk(tree)
if isinstance(n, ast.Call) and isinstance(n.func, ast.Attribute)
and n.func.attr in {"write_text", "write_bytes", "dump"}
]
self.assertEqual(writes, [], f"system_perf must not write files: {writes}")
+4 -4
View File
@@ -1,4 +1,4 @@
// Package watch polls corpus directories for changes and re-runs brain/index.
// Package watch polls corpus directories for changes and re-runs bin/kb/index.
//
// Port of the former bin/kb-watch bash script to an importable, testable Go
// package. Polls file mtimes (no inotify deps); cheap and reliable.
@@ -18,8 +18,8 @@ import (
type Options struct {
Dirs []string
Interval time.Duration
// IndexCmd is the index command template. %s is replaced by the repo
// root (from KB_ROOT). Defaults to `python3 <root>/bin/kb/index --with-mail`.
// IndexCmd is the kb/index command template. %s is replaced by the repo
// root (from KB_ROOT). Defaults to `python3 <root>/bin/kb/index`.
IndexCmd string
}
@@ -67,7 +67,7 @@ func fromEnv(args []string) Options {
if pys == "" {
pys = "python3"
}
opts.IndexCmd = pys + " <root>/bin/kb/index --with-mail"
opts.IndexCmd = pys + " <root>/bin/kb/index"
return opts
}
+2 -6
View File
@@ -3,7 +3,6 @@ package watch
import (
"os"
"path/filepath"
"strings"
"testing"
"time"
)
@@ -44,10 +43,7 @@ func TestFromEnvDefaults(t *testing.T) {
if opts.Interval != 30*time.Second {
t.Fatalf("default interval = %s, want 30s", opts.Interval)
}
if !strings.Contains(opts.IndexCmd, "kb/index") {
t.Fatalf("default index cmd = %q, want kb/index", opts.IndexCmd)
}
if !strings.Contains(opts.IndexCmd, "--with-mail") {
t.Fatalf("default index cmd must include --with-mail, got %q", opts.IndexCmd)
if opts.IndexCmd == "" {
t.Fatal("default index cmd is empty")
}
}
+136 -12
View File
@@ -1,26 +1,150 @@
#!/usr/bin/env python3
"""web/search — deprecated. Use bin/web/search.go (SearXNG, no Python client).
"""web/search - web search through the self-hosted SearXNG at search.ops.io.
bin/web/search.go QUERY [--json] [-n N] [--site HOST]
bin/web/search "LadybugDB vector search"
bin/web/search "model2vec multilingual" --site github.com
bin/web/search "uclancy" --category it -n 3 --json | jq -r '.results[].url'
bin/web/search "sqlite-vec" --refresh # ignore the cached answer
This complements bin/kb/search: the knowledge base holds our own facts, this
reaches the public web. Use it as the second, independent source that the
detective method asks for.
Exit codes: 0 results, 2 refused as possible PII, 3 throttled (not "nothing
found" - the instance answers 200 with an empty list when it throttles).
"""
from __future__ import annotations
import argparse
import fcntl
import json
import os
import sys
import time
import urllib.parse
import urllib.request
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
TOOLS = Path(__file__).resolve().parents[1] / "tools"
sys.path.insert(0, str(TOOLS))
sys.path.insert(0, str(TOOLS / "web-search"))
import websearch as ws # noqa: E402
from yamlout import to_yaml # noqa: E402
CONFIG = Path(os.environ.get("BRAIN_SEARCH_ENV", Path.home() / ".config/brain/search.env"))
CACHE = Path(os.environ.get("BRAIN_SEARCH_CACHE", Path.home() / ".cache/brain/web-search.sqlite"))
LOCK = CACHE.with_suffix(".lock")
def main(argv: list[str]) -> int:
print(
"bin/web/search is deprecated; use bin/web/search.go",
file=sys.stderr,
)
target = ROOT / "bin" / "web" / "search.go"
os.execvp("go", ["go", "run", str(target), *argv])
return 1
def load_config() -> dict:
if not CONFIG.exists():
sys.exit(f"no credentials at {CONFIG} (mode 600, BRAIN_SEARCH_URL/USER/PASS)")
conf = {}
for line in CONFIG.read_text().splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
conf[key.strip()] = value.strip().strip("\"'")
missing = {"BRAIN_SEARCH_URL", "BRAIN_SEARCH_USER", "BRAIN_SEARCH_PASS"} - conf.keys()
if missing:
sys.exit(f"{CONFIG} is missing {', '.join(sorted(missing))}")
return conf
def fetch(conf: dict, query: str, params: dict, timeout: int) -> dict:
args = {"q": query, "format": "json", **params}
url = f"{conf['BRAIN_SEARCH_URL'].rstrip('/')}/search?{urllib.parse.urlencode(args)}"
request = urllib.request.Request(url)
token = f"{conf['BRAIN_SEARCH_USER']}:{conf['BRAIN_SEARCH_PASS']}".encode()
import base64
request.add_header("Authorization", "Basic " + base64.b64encode(token).decode())
with urllib.request.urlopen(request, timeout=timeout) as response:
return json.loads(response.read().decode())
def main() -> int:
parser = argparse.ArgumentParser(description="web search via SearXNG")
parser.add_argument("query")
parser.add_argument("-n", "--limit", type=int, default=ws.DEFAULT_LIMIT)
parser.add_argument("--site", help="restrict to one domain")
parser.add_argument("--lang", help="language code, e.g. de")
parser.add_argument("--fresh", choices=["day", "week", "month", "year"],
help="time range")
parser.add_argument("--category", help="SearXNG category, e.g. it, science, news")
parser.add_argument("--engines", help="comma separated engine list")
parser.add_argument("--json", action="store_true")
parser.add_argument("--refresh", action="store_true", help="bypass the cache")
parser.add_argument("--ttl", type=float, default=ws.CACHE_TTL)
parser.add_argument("--timeout", type=int, default=25)
parser.add_argument("--force", action="store_true",
help="send even if the query looks like PII")
args = parser.parse_args()
query = f"site:{args.site} {args.query}" if args.site else args.query
reason = ws.phi_reason(query)
if reason and not args.force:
print(f"refused: {reason}. This query would leave the host.", file=sys.stderr)
print("Rephrase without identifiers, or pass --force if it is genuinely public.",
file=sys.stderr)
return 2
params = {}
if args.lang:
params["language"] = args.lang
if args.fresh:
params["time_range"] = args.fresh
if args.category:
params["categories"] = args.category
if args.engines:
params["engines"] = args.engines
key = ws.cache_key(query, params)
conn = ws.open_cache(CACHE)
if not args.refresh:
cached = ws.cache_get(conn, key, ttl=args.ttl)
if cached is not None:
out = ws.project(cached, limit=args.limit)
out["cached"] = True
sys.stdout.write(json.dumps(out, indent=2, ensure_ascii=False) + "\n"
if args.json else to_yaml(out))
return 0
conf = load_config()
LOCK.parent.mkdir(parents=True, exist_ok=True)
# One request at a time across every agent on this host: the instance
# suspends engines for minutes when several of us ask at once.
with open(LOCK, "w") as lock:
fcntl.flock(lock, fcntl.LOCK_EX)
payload = None
for attempt in range(1 + len(ws.RETRY_BACKOFF)):
delay = ws.wait_for(ws.last_call(conn), time.time())
if delay:
time.sleep(delay)
ws.mark_call(conn)
try:
payload = fetch(conf, query, params, args.timeout)
except Exception as error: # noqa: BLE001 - report, do not crash
print(f"request failed: {error}", file=sys.stderr)
return 3
if ws.classify(payload) == "ok":
break
if attempt < len(ws.RETRY_BACKOFF):
time.sleep(ws.RETRY_BACKOFF[attempt])
if ws.classify(payload) == "ok":
ws.cache_put(conn, key, payload)
out = ws.project(payload, limit=args.limit)
sys.stdout.write(json.dumps(out, indent=2, ensure_ascii=False) + "\n"
if args.json else to_yaml(out))
return 0 if out["status"] == "ok" else 3
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
sys.exit(main())
-170
View File
@@ -1,170 +0,0 @@
//usr/bin/env go run "$0" "$@"; exit
//
// bin/web/search.go - SearXNG as the second independent source (D3).
//
// ./bin/web/search.go "LadybugDB vector search"
// ./bin/web/search.go "model2vec" --category it --json
// ./bin/web/search.go "postgres" --site github.com --fresh year
//
// Empty results mean throttled, not "nothing exists". Exit 2 = PII refuse, 3 = throttled.
// Config: $BRAIN_SEARCH_ENV (default $HOME/.config/brain/search.env).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"encoding/json"
"fmt"
"net/http"
"os"
"time"
"github.com/eSlider/2dph/internal/cli"
"github.com/eSlider/2dph/internal/websearch"
"golang.org/x/sys/unix"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
c, err := websearch.ParseArgs(args)
if err != nil {
return cli.Fail(err)
}
query, site, lang, fresh, category, engines := c.Query, c.Site, c.Lang, c.Fresh, c.Category, c.Engines
limit := c.Limit
jsonOut, refresh, force := c.JSONOut, c.Refresh, c.Force
ttl := c.TTL
timeout := c.Timeout
if site != "" {
query = "site:" + site + " " + query
}
if reason := websearch.PHIReason(query); reason != "" && !force {
fmt.Fprintf(os.Stderr, "refused: %s. This query would leave the host.\n", reason)
fmt.Fprintln(os.Stderr, "Rephrase without identifiers, or pass --force if it is genuinely public.")
return 2
}
params := map[string]string{}
if lang != "" {
params["language"] = lang
}
if fresh != "" {
params["time_range"] = fresh
}
if category != "" {
params["categories"] = category
}
if engines != "" {
params["engines"] = engines
}
cachePath := os.Getenv("BRAIN_SEARCH_CACHE")
if cachePath == "" {
cachePath = os.Getenv("HOME") + "/.cache/brain/web-search.sqlite"
}
cache, err := websearch.OpenCache(cachePath)
if err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
return 1
}
defer cache.Close()
key := websearch.CacheKey(query, params)
now := float64(time.Now().Unix())
if !refresh {
if cached, err := cache.Get(key, ttl, now); err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
return 1
} else if cached != nil {
out := websearch.Project(*cached, limit, websearch.DefaultSnippetChars)
out.Cached = true
return writeOut(out, jsonOut)
}
}
envPath := os.Getenv("BRAIN_SEARCH_ENV")
if envPath == "" {
envPath = os.Getenv("HOME") + "/.config/brain/search.env"
}
conf, err := websearch.LoadConfig(envPath)
if err != nil {
fmt.Fprintf(os.Stderr, "web/search: %v\n", err)
return 1
}
lockPath := cachePath + ".lock"
lock, err := os.OpenFile(lockPath, os.O_CREATE|os.O_RDWR, 0o600)
if err != nil {
fmt.Fprintf(os.Stderr, "web/search: lock: %v\n", err)
return 1
}
defer lock.Close()
if err := unix.Flock(int(lock.Fd()), unix.LOCK_EX); err != nil {
fmt.Fprintf(os.Stderr, "web/search: lock: %v\n", err)
return 1
}
defer unix.Flock(int(lock.Fd()), unix.LOCK_UN)
var payload websearch.Payload
attempts := 1 + len(websearch.RetryBackoff)
client := &http.Client{}
for attempt := 0; attempt < attempts; attempt++ {
last, err := cache.LastCall()
if err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
return 1
}
if delay := websearch.WaitFor(last, float64(time.Now().Unix()), websearch.MinInterval); delay > 0 {
time.Sleep(time.Duration(delay * float64(time.Second)))
}
if err := cache.MarkCall(float64(time.Now().Unix())); err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
return 1
}
payload, err = websearch.Fetch(client, conf, query, params, time.Duration(timeout)*time.Second)
if err != nil {
fmt.Fprintf(os.Stderr, "request failed: %v\n", err)
return 3
}
if websearch.Classify(payload) == websearch.StatusOK {
break
}
if attempt < len(websearch.RetryBackoff) {
time.Sleep(time.Duration(websearch.RetryBackoff[attempt] * float64(time.Second)))
}
}
if websearch.Classify(payload) == websearch.StatusOK {
if err := cache.Put(key, payload, float64(time.Now().Unix())); err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
}
}
out := websearch.Project(payload, limit, websearch.DefaultSnippetChars)
code := writeOut(out, jsonOut)
if out.Status != websearch.StatusOK && code == 0 {
return 3
}
return code
}
func writeOut(out websearch.Output, jsonOut bool) int {
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
if err := enc.Encode(out); err != nil {
return 1
}
if out.Status != websearch.StatusOK {
return 3
}
return 0
}
fmt.Print(out.YAML())
if out.Status != websearch.StatusOK {
return 3
}
return 0
}
+17 -128
View File
@@ -1,177 +1,66 @@
# 2dph — docker composition
#
# bin/stack/start / start-assistant / status / stop
# docker compose up -d brain # API (Zig CGO serve)
# docker compose --profile index run --rm index # Python rebuild
# docker compose --profile picoclaw up -d # brain-mcp + CPU reasoner + PicoClaw gateway
# docker compose --profile reasoner up -d reasoner # CPU Ollama :11435
# docker compose --profile searxng up -d
# OCR_ENGINE=paddle docker compose --profile ocr-paddle run --rm ocr-paddle
# docker compose run --rm brain index # rebuild graph
# docker compose run --rm brain search "Matrix fed" # one-shot query
# docker compose run --rm brain serve # async Go server
# docker compose up brain-watch # auto re-index
#
# Secrets never baked in: search.env + db-profiles.yml from ~/.config/brain.
# Caching: the 128M model (HF_HOME) and kb.lbug (VAR_DIR) live in named
# volumes, so rebuilds never redownload the model or re-derive the graph.
# Secrets are never baked into the image: search.env + db-profiles.yml mount
# read-only from ~/.config/brain.
name: 2dph
networks:
default:
name: 2dph_sys
driver: bridge
ipam:
config:
- subnet: 10.23.42.0/24
services:
brain:
image: ghcr.io/eslider/2dph:api
image: ghcr.io/eslider/2dph:latest
build:
context: .
dockerfile: Dockerfile
target: api
cache_from:
- ghcr.io/eslider/2dph:cache
command: ["serve"]
command: ["brain", "search", "help"]
environment: &env
HF_HOME: /data/hf
BRAIN_SEARCH_CACHE: /data/cache/web-search.sqlite
BRAIN_DB_PROFILES: /secret/db-profiles.yml
BRAIN_SEARCH_ENV: /secret/search.env
KB_ROOT: /data
KB_SEARCH_CMD: /app/bin/kb/search
KB_WORKERS: "4"
KB_PORT: "8630"
volumes:
- kb-model:/data/hf
- kb-var:/data
# corpus is read-only on the host, never written from the container
- ..:/corpus:ro
- ~/.config/brain:/secret:ro
ports:
- "127.0.0.1:8630:8630"
read_only: true
tmpfs:
- /tmp
healthcheck:
test: ["CMD", "wget", "-qO-", "http://127.0.0.1:8630/health"]
test: ["CMD", "python3", "-c", "import ladybug, model2vec, mistune; print('ok')"]
interval: 30s
timeout: 5s
retries: 3
restart: unless-stopped
stop_grace_period: 20s
# watcher: re-index on corpus file change (watchdog script)
brain-watch:
image: ghcr.io/eslider/2dph:api
image: ghcr.io/eslider/2dph:latest
environment: *env
volumes:
- kb-model:/data/hf
- kb-var:/data
- ..:/corpus:ro
- ~/.config/brain:/secret:ro
command: ["watch", "/corpus"]
command: ["brain", "watch", "/corpus"]
read_only: true
tmpfs:
- /tmp
restart: unless-stopped
stop_grace_period: 20s
# Python write path (Ladybug rebuild). Not in the API image.
# docker compose --profile index run --rm index
index:
profiles: ["index"]
image: ghcr.io/eslider/2dph:index
build:
context: .
dockerfile: Dockerfile
target: index
environment:
HF_HOME: /data/hf
KB_PY: python3
volumes:
- kb-model:/data/hf
- kb-var:/app/var
- ..:/corpus:ro
- ~/.config/brain:/secret:ro
command: ["index"]
read_only: true
tmpfs:
- /tmp
# Optional local SearXNG (D3). Skip if BRAIN_SEARCH_URL already points at a
# live instance — do not run a second copy on that host.
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
searxng:
profiles: ["searxng"]
image: docker.io/searxng/searxng:2026.8.10-0a118066d
ports:
- "127.0.0.1:8888:8080"
environment:
SEARXNG_SECRET: ${SEARXNG_SECRET:-}
volumes:
- ./deploy/searxng/settings.yml:/etc/searxng/settings.yml:ro
- ./deploy/searxng/limiter.toml:/etc/searxng/limiter.toml:ro
restart: unless-stopped
# MCP endpoint for PicoClaw (and any MCP client).
# docker compose --profile picoclaw up -d
brain-mcp:
profiles: ["picoclaw"]
image: ghcr.io/eslider/2dph:api
environment: *env
volumes:
- kb-model:/data/hf
- kb-var:/data
- ~/.config/brain:/secret:ro
command: ["serve"]
ports:
- "127.0.0.1:8630:8630"
read_only: true
tmpfs:
- /tmp
restart: unless-stopped
# CPU OpenAI-compatible sidecar (D18). Weights are pulled at runtime, not
# baked into the 2dph image. Does not touch host Ollama on :11434.
# docker compose --profile reasoner up -d reasoner
# docker compose --profile reasoner exec reasoner ollama pull qwen3.5:9b
reasoner:
profiles: ["reasoner", "picoclaw"]
image: docker.io/ollama/ollama:latest
environment:
OLLAMA_NUM_GPU: "0"
OLLAMA_HOST: "0.0.0.0:11434"
ports:
- "127.0.0.1:11435:11434"
volumes:
- reasoner-ollama:/root/.ollama
restart: unless-stopped
# Official PicoClaw gateway. Config has no secrets (Ollama + HTTP MCP).
# Host network: brain/reasoner bind 127.0.0.1 only, so host.docker.internal
# (docker0) cannot reach them. Gateway 127.0.0.1:18790 (not the 18800 launcher).
# If :8630/:11435 are already bound, do not start brain-mcp/reasoner:
# docker compose --profile picoclaw up -d --no-deps picoclaw
picoclaw:
profiles: ["picoclaw"]
image: docker.io/sipeed/picoclaw:v0.3.1
network_mode: host
depends_on:
- brain-mcp
- reasoner
environment:
PICOCLAW_GATEWAY_HOST: "127.0.0.1"
entrypoint: ["picoclaw", "gateway"]
volumes:
- picoclaw-home:/root/.picoclaw
- ./deploy/picoclaw/config.json:/root/.picoclaw/config.json:ro
restart: unless-stopped
# Optional PP-OCRv5 (not default). Default OCR is tesseract eng+deu.
# OCR_ENGINE=paddle docker compose --profile ocr-paddle run --rm ocr-paddle
ocr-paddle:
profiles: ["ocr-paddle"]
image: python:3.12-slim
environment:
OCR_ENGINE: paddle
command: ["python", "-c", "print('OCR_ENGINE=paddle; install paddleocr on PATH')"]
volumes:
kb-model:
kb-var:
reasoner-ollama:
picoclaw-home:
-33
View File
@@ -1,33 +0,0 @@
{
"agents": {
"defaults": {
"model_name": "qwen3.5-9b",
"max_tool_iterations": 8,
"max_tokens": 512,
"context_window": 8192
}
},
"model_list": [
{
"model_name": "qwen3.5-9b",
"model": "ollama/qwen3.5:9b",
"api_base": "http://127.0.0.1:11435/v1",
"request_timeout": 600
}
],
"tools": {
"web": {
"enabled": false
},
"mcp": {
"enabled": true,
"servers": {
"2dph": {
"enabled": true,
"type": "http",
"url": "http://127.0.0.1:8630/mcp"
}
}
}
}
}
-8
View File
@@ -1,8 +0,0 @@
{
"mcpServers": {
"2dph": {
"url": "http://127.0.0.1:8630/mcp",
"description": "2dph fact gate. Tool order: search → get → audit. throttled is not absence."
}
}
}
-7
View File
@@ -1,7 +0,0 @@
[botdetection.ip_lists]
# RFC1918 only. Do not copy a live instance egress IP into git.
pass_ip = [
"10.0.0.0/8",
"172.16.0.0/12",
"192.168.0.0/16",
]
-28
View File
@@ -1,28 +0,0 @@
use_default_settings: true
general:
instance_name: "2dph"
search:
formats:
- html
- json
suspended_times:
SearxEngineCaptcha: 300
SearxEngineTooManyRequests: 120
SearxEngineAccessDenied: 300
server:
limiter: true
image_proxy: false
# secret_key comes from SEARXNG_SECRET (never commit a real secret)
engines:
- name: bing
disabled: false
- name: google
disabled: false
- name: duckduckgo
disabled: false
- name: wikipedia
disabled: false
+7 -34
View File
@@ -1,38 +1,11 @@
---
type: reference
status: current
related:
- docs/runbook.md
- docs/design.md
- PLAN.md
- docs/roadmap.md
---
# 2dph (deductionphile)
# 2dph docs (Diataxis)
Evidence-first knowledge graph. Facts need proof or they are
Evidence-first knowledge graph + hybrid RAG over the operational
Brain/ops/eSlider stack. Facts need proof or they are
`(not confirmed)`.
| Type | Doc |
|------|-----|
| tutorial / howto | [runbook](runbook.md) — run anywhere (uv, Go, Docker) |
| explanation | [design](design.md) — two roots, deduction, D17/D20/D18 |
| explanation | [roadmap](roadmap.md) — gap to v1 (epic #16) |
| howto | [picoclaw](picoclaw.md) — MCP agent (`bin/stack/start-assistant`) |
| howto | [reasoner](reasoner.md) — CPU bake-off (D18) |
| reference | [PLAN.md](../PLAN.md) — decisions D1D24 |
- [PLAN.md](../PLAN.md) — decisions, execution order, open questions (v2)
- [design](design.md) — schema, deduction model, sources
- [Gitea issues](https://git.produktor.io/eSlider/2dph/issues) — work board (origin)
Decisions the public face must name: **D3** SearXNG compose, **D6** Go service /
Python write sidecar, **D14** `bin/{subject}/{method}.go`, **D15** Gitea origin,
**D17** assertion gate (facts → info → web), **D18** pluggable reasoner.
Search: `bin/brain/search.go "query"` (HTTP: `bin/brain/serve.go`
`/health` `/search` `/get` `/stats` `/audit` `/ingest`). `--hop N` walks
`FROM_FILE` → Commit → Person from each hit (max 3). Rebuild writes
File edges ([#17](https://git.produktor.io/eSlider/2dph/issues/17)).
Work board: [Gitea issues](https://git.produktor.io/eSlider/2dph/issues)
([epic #16](https://git.produktor.io/eSlider/2dph/issues/16)).
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
Published docs live here and match live commands.
Published docs live here and mirror the project state.

Some files were not shown because too many files have changed in this diff Show More