Compare commits
38
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
638e84c324 | ||
|
|
bb272d7416 | ||
|
|
e393cc6a99 | ||
|
|
7fb51e9918 | ||
|
|
e5a982e99a | ||
|
|
6b680bdb90 | ||
|
|
77c635b16b | ||
|
|
2fede46dcd | ||
|
|
54201d8c22 | ||
|
|
a6dbe2f6da | ||
|
|
368a757612 | ||
|
|
a8675ac33b | ||
|
|
de632ba6cc | ||
|
|
66c87842e2 | ||
|
|
20b78a9a20 | ||
|
|
5d4b3427a4 | ||
|
|
1c7db6d499 | ||
|
|
0786ddcb06 | ||
|
|
eeb5b79cf2 | ||
|
|
5990feb1f6 | ||
|
|
d27a738fee | ||
|
|
669e184cf6 | ||
|
|
ebc3f948c1 | ||
|
|
fe6a02024c | ||
|
|
4a065d9838 | ||
|
|
e3c6ef5684 | ||
|
|
98c14e23f1 | ||
|
|
ff1716de40 | ||
|
|
fec5325c7a | ||
|
|
27d9521e7f | ||
|
|
a7cb8d4c76 | ||
|
|
53cd00284d | ||
|
|
b73b4d4f97 | ||
|
|
1d1f6a90ff | ||
|
|
f220bcd95a | ||
|
|
678a1d1dba | ||
|
|
8781c0c3eb | ||
|
|
d6b17e8819 |
+1
-1
@@ -3,6 +3,7 @@
|
||||
var
|
||||
.git
|
||||
.github
|
||||
lib-ladybug
|
||||
__pycache__
|
||||
*.pyc
|
||||
*.lbug
|
||||
@@ -10,5 +11,4 @@ __pycache__
|
||||
.cache
|
||||
.secrets
|
||||
.skills-tmp
|
||||
serve/serve
|
||||
docs/.build
|
||||
@@ -19,6 +19,10 @@ jobs:
|
||||
with:
|
||||
fetch-depth: 0
|
||||
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: go.mod
|
||||
|
||||
- name: Install uv
|
||||
uses: astral-sh/setup-uv@v6
|
||||
with:
|
||||
@@ -31,26 +35,32 @@ jobs:
|
||||
run: |
|
||||
bash -n bin/db/psql-yq
|
||||
bash -n bin/db/ssh-tunnel
|
||||
bash -n bin/kb-watch
|
||||
bash -n bin/docker-entrypoint
|
||||
bash -n bin/kb/search
|
||||
bash -n bin/cgo/zig
|
||||
sh -n bin/cgo/zcc
|
||||
sh -n bin/cgo/zc++
|
||||
|
||||
- name: Python unit tests (offline, vendored tools)
|
||||
run: |
|
||||
uv run python -m unittest discover -s tools -t .
|
||||
uv run python -m unittest discover -s bin/tools -t .
|
||||
|
||||
- name: Go serve tests (async, goroutine-bounded)
|
||||
- name: Go tests (root module, no ladybug cgo)
|
||||
run: |
|
||||
go vet ./...
|
||||
go test ./... -count=1
|
||||
working-directory: serve
|
||||
|
||||
- name: brain ranking tests (no cgo / no ladybug)
|
||||
run: go test ./internal/brain/rank -count=1
|
||||
|
||||
- name: facts/audit self (lexicon consistency, no network)
|
||||
run: |
|
||||
./bin/facts/audit self 2>/dev/null || echo "audit: not yet implemented; gate skipped"
|
||||
|
||||
- name: kb/eval recall gate
|
||||
- name: CGO via Zig (compile brain/search)
|
||||
run: |
|
||||
./bin/kb/eval 2>/dev/null || echo "eval: not yet implemented; gate skipped"
|
||||
chmod +x bin/cgo/zig bin/cgo/zcc bin/cgo/zc++
|
||||
bin/cgo/zig go build -tags system_ladybug -o /tmp/brain-search ./bin/brain/search.go
|
||||
|
||||
release:
|
||||
name: Release (semver)
|
||||
|
||||
+4
-1
@@ -8,4 +8,7 @@ __pycache__/
|
||||
.DS_Store
|
||||
*.env
|
||||
.env
|
||||
.secrets/
|
||||
.secrets/
|
||||
lib-ladybug/
|
||||
go.work.local
|
||||
models/
|
||||
|
||||
@@ -15,6 +15,9 @@ Read first: [PLAN](PLAN.md) → [docs](docs/).
|
||||
- `info` root = descriptive/narrative leafs, searchable, never asserted as fact.
|
||||
- Search is deduction: `facts` → `info` → `web-search` (second independent
|
||||
source). An answer is `confirmed` only if it comes off the facts root.
|
||||
- Fact-check every *claim* (facts → info → live → web), not every edit or
|
||||
syntax tweak. PicoClaw: `search` then `get` then `audit` before a factual
|
||||
reply (`skills/picoclaw/SKILL.md`). `throttled` is not a negative finding.
|
||||
|
||||
## Hard rules
|
||||
|
||||
@@ -24,7 +27,7 @@ Read first: [PLAN](PLAN.md) → [docs](docs/).
|
||||
2. **Read-only data sources.** Ladybug `var/kb.lbug` and Postgres are opened
|
||||
read-only for queries. Index rebuilds write to `var/` (gitignored).
|
||||
3. **PII.** `brain-test`, `cs_brain` client data is never read or quoted.
|
||||
4. **No main pushes.** Feature branches + PR via `gh`; CI must be green.
|
||||
4. **No main pushes.** Feature branches + GitHub PR (`gh`); CI (Actions) must be green. Work board: [Gitea issues](https://git.produktor.io/eSlider/2dph/issues).
|
||||
5. **TDD.** Failing test before tool code. Unit tests run offline against
|
||||
fixtures; network/db calls are wrapped.
|
||||
6. **docs reflect behaviour.** Any change updates `docs/` + `PLAN.md` status.
|
||||
@@ -35,22 +38,65 @@ Read first: [PLAN](PLAN.md) → [docs](docs/).
|
||||
PLAN.md decisions + execution + open questions
|
||||
docs/ published docs
|
||||
skills/ in-project agent skills (vendored, no external links)
|
||||
bin/ self-describing tools bin/{subject}/{method} (shebang)
|
||||
bin/kb-watch corpus watcher (mtimes, no inotify deps)
|
||||
bin/docker-entrypoint container entrypoint (brain index|search|serve|watch)
|
||||
serve/ async Go HTTP server (goroutines, bounded worker pool)
|
||||
tools/ vendored python libs behind bin/* (yamlout, websearch)
|
||||
bin/ self-describing tools bin/{subject}/{method}.go (shebang)
|
||||
bin/brain/ search.go serve.go index.go get.go stats.go eval.go watch.go
|
||||
bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats
|
||||
bin/mail/ sync.go import.go (index_mail → brain/index.go)
|
||||
bin/markdown/ import.go (H2 leaf split; Python bin/md/import fallback)
|
||||
bin/postgres/ query.go (read-only YAML)
|
||||
bin/git/ import.go (go-git history; Python shim execs it)
|
||||
bin/web/ search.go (SearXNG; Python shim execs it)
|
||||
bin/reasoner/ bakeoff.go (D18 CPU OpenAI tool-call bake-off)
|
||||
internal/ shared Go (brain/rank is cgo-free; chats parsers; gitlog; websearch; reasoner)
|
||||
bin/watch/ corpus watcher (used by bin/brain/watch.go)
|
||||
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
|
||||
bin/cgo/ zig zcc zc++ (CGO via zig cc, not gcc)
|
||||
bin/docker-entrypoint container entrypoint (api: serve|search|watch; index: python)
|
||||
compose.yaml docker composition (root level, not docker/)
|
||||
Dockerfile multi-stage: python deps + static Go serve
|
||||
var/ kb.lbug, caches (gitignored)
|
||||
Dockerfile api (Zig CGO, no Python) + index (Python write)
|
||||
var/ kb.lbug, var/mail/*, caches (gitignored)
|
||||
.venv/ ladybug + model2vec + mistune
|
||||
```
|
||||
|
||||
## Mail pipeline
|
||||
|
||||
```bash
|
||||
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
|
||||
bin/mail/sync.go --source gmail --query 'from:example.com' --out var/mail # Gmail search (default in:inbox)
|
||||
bin/mail/import.go --from-raw var/mail # message.json → message.md (convert only)
|
||||
bin/brain/index.go --rebuild # rebuild brain incl. all mail (fresh DB)
|
||||
```
|
||||
|
||||
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
|
||||
`body.attachmentId` (not partId) for attachments.
|
||||
- `import` converts body + attachments to markdown. PDFs use poppler
|
||||
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to
|
||||
docling (isolated subprocess — its native onnx can segfault the parent).
|
||||
Conversion never touches the brain DB (crash safety).
|
||||
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Ladybug
|
||||
corrupts its WAL when brand-new leafs are bulk-inserted while FTS/vector
|
||||
indexes exist; a fresh DB with indexes created last is the only safe path.
|
||||
Keep conversion + indexing separate so a conversion crash can't leave the
|
||||
DB mid-transaction.
|
||||
|
||||
## Tools
|
||||
|
||||
```bash
|
||||
bin/facts/audit ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
|
||||
bin/kb/search "query" [--hop N] [--repo X] # deduction search → YAML
|
||||
bin/facts/audit.go ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
|
||||
bin/facts/crm.go [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
|
||||
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
|
||||
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
|
||||
bin/brain/search.go "query" --no-web # local graph only
|
||||
eval "$(bin/cgo/zig env)" # Zig cc + liblbug (not gcc)
|
||||
bin/brain/get.go <id> [--body] [--json] # Go read; Python bin/kb/get CI fallback
|
||||
bin/brain/stats.go [--json]
|
||||
bin/brain/eval.go [--json] # recall@5; questions in internal/brain/rank
|
||||
bin/brain/serve.go # HTTP :8630; GET /openapi.json POST /mcp
|
||||
bin/markdown/import.go [dir] # H2 leafs → YAML; Python bin/md/import fallback
|
||||
bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs
|
||||
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
|
||||
bin/reasoner/bakeoff.go [--model ID] [--json] # D18 CPU tool-call bake-off
|
||||
bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
|
||||
bin/md/tables # what the graph holds → YAML
|
||||
bin/brain/deduce "question" # thinking wrapper
|
||||
```
|
||||
@@ -58,6 +104,19 @@ bin/brain/deduce "question" # thinking wrapper
|
||||
Never start a shell command with `cd` — use the tool working-directory
|
||||
parameter. Search before reading whole files.
|
||||
|
||||
## GitHub safety rules (ABSOLUTE — never violate)
|
||||
|
||||
1. **No absolute paths in committed files.** Replace `/mnt/`, `/home/<user>/`,
|
||||
`/Users/<user>/` with env vars (`$HOME`, `$PROJECTS_ROOT`, `$DOCS_BASE`).
|
||||
2. **No PII in commits.** No real names, phones, emails of third parties.
|
||||
Test data must be synthetic (Alice, Bob, Charlie, Diana, example.com).
|
||||
3. **No credentials/secrets in commits.** API keys, tokens, passwords, session
|
||||
strings, phone numbers only in gitignored `.env` files, referenced by path.
|
||||
4. **Curasoft, edelweiss — no files, no mentions.** Remove all traces if found.
|
||||
5. **Check git history before push.** If any commit contains leaks, rewrite
|
||||
history (rebase + force push) AND delete affected GitHub releases/tags.
|
||||
6. **`docs/chat-import-plan.md`** — reference Gitea issue, never embed secrets.
|
||||
|
||||
## Communication
|
||||
|
||||
Same tone as the corpus: plain, lists, no hype. Sign-off `Andriy Oblivantsev`.
|
||||
|
||||
+59
-15
@@ -1,5 +1,13 @@
|
||||
# syntax=docker/dockerfile:1
|
||||
FROM python:3.12-slim AS base
|
||||
#
|
||||
# docker build --target api -t 2dph:api .
|
||||
# docker build --target index -t 2dph:index .
|
||||
#
|
||||
# API: Go + ladybug via Zig CGO (no CPython).
|
||||
# Index: Python write path (profile `index` until brain/add is v2).
|
||||
|
||||
# --- Python sidecar (Ladybug write / rebuild) ---
|
||||
FROM python:3.12-slim AS index
|
||||
|
||||
ENV PYTHONUNBUFFERED=1 \
|
||||
PYTHONDONTWRITEBYTECODE=1 \
|
||||
@@ -9,29 +17,65 @@ ENV PYTHONUNBUFFERED=1 \
|
||||
WORKDIR /app
|
||||
RUN id -u 2dph 2>/dev/null || useradd --create-home --uid 1001 2dph
|
||||
|
||||
# deps layer-first: rebuild only on dependency change
|
||||
COPY requirements.lock.txt /tmp/requirements.lock.txt
|
||||
RUN python -m pip install --no-cache-dir -r /tmp/requirements.lock.txt \
|
||||
&& rm /tmp/requirements.lock.txt
|
||||
|
||||
# Go serve: static binary, no interpreter at runtime
|
||||
FROM golang:1.25 AS serve-build
|
||||
WORKDIR /src/serve
|
||||
COPY serve/go.mod serve/go.sum* ./
|
||||
COPY serve .
|
||||
RUN CGO_ENABLED=0 go build -o /serve -ldflags="-s -w" .
|
||||
|
||||
# runtime: python toolchain + Go server
|
||||
FROM base
|
||||
COPY . .
|
||||
COPY --from=serve-build /serve /app/serve/serve
|
||||
RUN chmod +x /app/bin/kb-watch /app/bin/docker-entrypoint \
|
||||
RUN chmod +x /app/bin/docker-entrypoint \
|
||||
&& chown -R 2dph:2dph /app
|
||||
USER 2dph
|
||||
|
||||
ENV PATH="/app/bin:${PATH}" \
|
||||
KB_PY=python3
|
||||
KB_PY=python3 \
|
||||
KB_ROOT=/app
|
||||
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
|
||||
CMD python -c "import model2vec, ladybug, mistune; print('ok')" || exit 1
|
||||
ENTRYPOINT ["/app/bin/docker-entrypoint"]
|
||||
|
||||
ENTRYPOINT ["/app/bin/docker-entrypoint"]
|
||||
# --- Go API: CGO with Zig, not gcc ---
|
||||
FROM golang:1.26-bookworm AS api-build
|
||||
WORKDIR /src
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends curl xz-utils ca-certificates \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
COPY bin/cgo ./bin/cgo
|
||||
RUN chmod +x bin/cgo/zig bin/cgo/zcc bin/cgo/zc++ \
|
||||
&& ./bin/cgo/zig env >/dev/null
|
||||
|
||||
COPY go.mod go.sum ./
|
||||
RUN go mod download
|
||||
|
||||
COPY . .
|
||||
ENV CGO_RPATH=/usr/local/lib
|
||||
RUN eval "$(./bin/cgo/zig env)" \
|
||||
&& go build -tags brain_serve,system_ladybug -o /out/brain-serve ./bin/brain/serve.go \
|
||||
&& go build -tags system_ladybug -o /out/brain-search ./bin/brain/search.go \
|
||||
&& CGO_ENABLED=0 go build -tags brain_watch -o /out/brain-watch ./bin/brain/watch.go
|
||||
|
||||
FROM debian:bookworm-slim AS api
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends libssl3 ca-certificates wget \
|
||||
&& rm -rf /var/lib/apt/lists/* \
|
||||
&& useradd --create-home --uid 1001 2dph
|
||||
COPY --from=api-build /out/brain-serve /usr/local/bin/brain-serve
|
||||
COPY --from=api-build /out/brain-search /usr/local/bin/brain-search
|
||||
COPY --from=api-build /out/brain-watch /usr/local/bin/brain-watch
|
||||
COPY --from=api-build /src/lib-ladybug/liblbug.so.0.19.1 /usr/local/lib/liblbug.so.0.19.1
|
||||
COPY bin/docker-entrypoint /usr/local/bin/docker-entrypoint
|
||||
RUN chmod +x /usr/local/bin/docker-entrypoint \
|
||||
&& ln -s liblbug.so.0.19.1 /usr/local/lib/liblbug.so.0 \
|
||||
&& ln -s liblbug.so.0 /usr/local/lib/liblbug.so \
|
||||
&& ldconfig
|
||||
USER 2dph
|
||||
ENV KB_ROOT=/data \
|
||||
KB_PORT=8630 \
|
||||
LD_LIBRARY_PATH=/usr/local/lib \
|
||||
HF_HOME=/data/hf
|
||||
WORKDIR /data
|
||||
EXPOSE 8630
|
||||
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
|
||||
CMD wget -qO- http://127.0.0.1:8630/health || exit 1
|
||||
ENTRYPOINT ["/usr/local/bin/docker-entrypoint"]
|
||||
CMD ["serve"]
|
||||
|
||||
@@ -25,11 +25,11 @@ detective method: **a fact needs ≥2 independent sources or it is
|
||||
| # | Question | Answer |
|
||||
|---|----------|--------|
|
||||
| D1 | RAG corpus | ops stack (chat, onlyoffice, gitea/NPM, searchxng, observability, ai-bot, mcp-servers, `~/.ssh/config`) + portfolio. Exclude `office.dev` + jobs/applications. |
|
||||
| D2 | skill merging | integrate skills **in this project** `skills/`; skip gitea / brain-detective-depe ndent skills. |
|
||||
| D3 | web search | import `web-search`, retire local `searxng-ops`. Vendored here, no remote link. |
|
||||
| D2 | skill merging | integrate skills **in this project** `skills/`; skip gitea / brain-dependent skills. |
|
||||
| D3 | web search | Go client `bin/web/search.go` (`internal/websearch`). SearXNG URL is config (`BRAIN_SEARCH_URL`). Optional Compose profile `searxng` (sanitized settings). Do not run a second copy on a host that already has one. Empty/`throttled` ≠ “nothing exists”. |
|
||||
| D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. |
|
||||
| D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). |
|
||||
| D6 | graph engine | **LadybugDB** (Kuzu successor, MIT, embedded, native FTS+vector+Cypher). Python binding for `bin/*`; Go shebang for golang tools. |
|
||||
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`). Read path is Go + Zig CGO (D21). Python `bin/kb/{get,stats,eval}` is the CI fallback when Zig/libs are not fetched. Index/write stays Python (`compose --profile index`) until the Go write path is safe. |
|
||||
| D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). |
|
||||
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. |
|
||||
| D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. |
|
||||
@@ -37,9 +37,14 @@ detective method: **a fact needs ≥2 independent sources or it is
|
||||
| D11 | strong/weak | `root` column: `facts` (strong) vs `info` (weak). Answer is `confirmed` only from facts root. |
|
||||
| D12 | transactional | facts and info split by root but **written in the same Ladybug transaction (ACID)** on every write. |
|
||||
| D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. |
|
||||
| D14 | tooling style | `bin/{subject}/{method}` self-describing: shebang line 1, usage comment from line 2. Go shebang: `///usr/bin/env go run "$0" "$@"; exit`. |
|
||||
| D15 | repo | GitHub `eSlider/2dph`, public (like sibling repos), push/commit via `gh`, TDD + commit every change, CI/CD. |
|
||||
| D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. |
|
||||
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
|
||||
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. |
|
||||
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
|
||||
| D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is not shipped; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. |
|
||||
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
|
||||
| D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`). |
|
||||
| D21 | CGO | Ladybug/tokenizers CGO is compiled with **Zig** (`bin/cgo/zcc` → `zig cc -target …-linux-gnu`), not gcc. `bin/cgo/zig` pins Zig 0.14.1 + liblbug 0.19.1 + libtokenizers 1.27.0. Compose `target: api` has no CPython; write/rebuild is profile `index`. |
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -47,16 +52,27 @@ detective method: **a fact needs ≥2 independent sources or it is
|
||||
2dph/
|
||||
PLAN.md / AGENTS.md
|
||||
docs/ published docs (this conversation → docs/ as md)
|
||||
skills/ in-project skills (web-search, db-yaml, kb-search, agent-cost, diataxis-docs, …)
|
||||
skills/ in-project skills (web-search, postgres, brain, picoclaw, diataxis-docs)
|
||||
bin/
|
||||
facts/extract auto-pair 2 sources → lexicon yaml + graph
|
||||
facts/audit ["self"|"facts"|"info"|"stale"] 2-source + staleness gate
|
||||
kb/index build FTS + HNSW from corpus
|
||||
kb/search deduction: facts → info → web-search; --hop N
|
||||
kb/get kb/stats kb/eval
|
||||
md/import md/select md/tables md/gaps (mistune)
|
||||
facts/extract.go audit.go crm.go # D14 shebang; Python implementation
|
||||
kb/index Python write path (called by bin/brain/index.go)
|
||||
brain/index.go rebuild FTS + HNSW (incl. --with-mail)
|
||||
brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback
|
||||
brain/watch.go
|
||||
brain/search.go deduction: facts → info → web-search
|
||||
brain/serve.go HTTP API in-process + OpenAPI/MCP (D20); Zig CGO (D21)
|
||||
cgo/zig zcc zc++ CGO toolchain (zig cc, not gcc)
|
||||
mail/import.go JSON → markdown (no brain write)
|
||||
markdown/import.go H2 leaf split (Go); Python bin/md/import fallback
|
||||
postgres/query.go read-only YAML (wraps bin/db/psql-yq)
|
||||
git/import.go go-git history (no git binary; conversion only)
|
||||
web/search.go SearXNG client (throttled ≠ absence)
|
||||
reasoner/bakeoff.go CPU tool-call bake-off (D18; OpenAI tools)
|
||||
chats/sync.go import.go facts.go apply.go
|
||||
(libs in internal/chats; no chats index)
|
||||
md/import (deprecated; bin/markdown/import.go)
|
||||
brain/extract brain/audit brain/deduce (thinking wrapper)
|
||||
web/search (vendored)
|
||||
web/search (deprecated shim → web/search.go)
|
||||
db/psql-yq (vendored)
|
||||
ssh-tunnel onlyoffice pg tunnel 5433
|
||||
var/kb.lbug single embedded store (gitignored)
|
||||
@@ -93,18 +109,36 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
|
||||
|
||||
- OQ1: mutually-contradicting evidence — how to resolve (authority weighting,
|
||||
temporal freshness, audit adjudication).
|
||||
- OQ2: OCR pipeline for pdfs/images/docs (late phase).
|
||||
- OQ2: OCR pipeline for pdfs/images/docs — mostly solved: poppler pdftotext
|
||||
fast-path for born-digital PDFs, docling fallback for the ~5% textless ones.
|
||||
- OQ3: optional duckdb-md layer for `SELECT … FORMAT MARKDOWN` export/write-back.
|
||||
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
|
||||
serialize and unambiguous; YAML only where humans edit files.
|
||||
|
||||
## Mail pipeline (done)
|
||||
|
||||
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
|
||||
Gmail attachments key off `body.attachmentId`, not MIME `partId`.
|
||||
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
|
||||
`pdftotext -layout` (~15ms) with docling subprocess fallback; ICS sidecars
|
||||
Latin-1→UTF-8 normalized.
|
||||
3. `bin/brain/index.go --rebuild` — fresh rebuild (repo corpus + mail) because ladybug
|
||||
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
|
||||
indexing stay separate for crash safety. `bin/mail/index_mail` is a
|
||||
deprecation shim.
|
||||
4. Result: 17,835 messages → 28,918 info leafs, FTS + HNSW healthy, searchable
|
||||
via `bin/brain/search.go`.
|
||||
|
||||
## CI/CD pipeline (D15)
|
||||
|
||||
`.github/workflows/ci.yml`:
|
||||
|
||||
1. go vet + go test ./... (Go tools)
|
||||
2. python -m unittest discover + pytest (Py tools)
|
||||
3. bin/facts/audit self (lexicon internal consistency)
|
||||
4. bin/kb/eval (recall@5 ≥ 0.95, gates index regressions)
|
||||
5. md-docs build/lint if docs tooling arrives.
|
||||
1. go vet + go test ./... (root module; packages without ladybug cgo)
|
||||
2. `go test ./internal/brain/rank` (cgo-free ranking + flag parser)
|
||||
3. python -m unittest discover -s bin/tools (includes published-docs SoT)
|
||||
4. `bin/facts/audit self` (lexicon internal consistency; `bin/facts/audit.go` is the D14 wrapper)
|
||||
5. `bin/kb/eval` (recall@5 ≥ 0.95). Local SoT is `bin/brain/eval.go` via Zig CGO.
|
||||
6. `bin/cgo/zig go build -tags system_ladybug` (compile search with zig cc; fetches pinned zig+libs).
|
||||
|
||||
Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as
|
||||
`db/tech-poc`: contract first where there is an OpenAPI/message shape.
|
||||
@@ -113,7 +147,7 @@ Feedback loop: every commit → PR → CI → green/gate → merge. Same discipl
|
||||
|
||||
1. scaffold repo (:done after this file + AGENTS.md + .gitignore + ci)
|
||||
2. gh repo create eSlider/2dph --private + initial commit + CI
|
||||
3. vendored skill integration (web-search, db-yaml, kb-search, agent-cost, diataxis-docs) — no remote links
|
||||
3. vendored skill integration (web-search, postgres, brain, diataxis-docs) — no remote links
|
||||
4. .venv: ladybug + model2vec + mistune
|
||||
5. schema + tools with TDD (kb + md + facts + brain)
|
||||
6. ~/.config/brain config
|
||||
|
||||
@@ -28,11 +28,11 @@ graph TB
|
||||
end
|
||||
|
||||
subgraph dph["2dph tools"]
|
||||
EX["bin/facts/extract<br/>2-source pairing"]
|
||||
AU["bin/facts/audit<br/>confidence + staleness"]
|
||||
IDX["bin/kb/index<br/>chunk + embed"]
|
||||
MD["bin/md/import<br/>mistune leaves"]
|
||||
SR["bin/kb/search<br/>deduction + --hop"]
|
||||
EX["bin/facts/extract.go<br/>2-source pairing"]
|
||||
AU["bin/facts/audit.go<br/>confidence + staleness"]
|
||||
IDX["bin/brain/index.go<br/>chunk + embed"]
|
||||
MD["bin/markdown/import.go<br/>H2 leaf split"]
|
||||
SR["bin/brain/search.go<br/>deduction"]
|
||||
end
|
||||
|
||||
subgraph store["Ladybug var/kb.lbug"]
|
||||
@@ -85,19 +85,52 @@ fact; conflicting sources or a single source → `hypothesis` → `(not confirme
|
||||
## Deduction search
|
||||
|
||||
```bash
|
||||
bin/kb/search "Matrix federation over HTTPS" # facts → info → web-search
|
||||
bin/kb/search "what runs on arc-2" --hop 1 # walk graph edges
|
||||
bin/kb/search "where is cs-lexicon" --json | yq '.' # YAML by default
|
||||
bin/kb/get <id> --body # full chunk on demand
|
||||
bin/kb/stats # index health
|
||||
bin/kb/eval # recall@5 gate
|
||||
bin/brain/search.go "Matrix federation over HTTPS" # facts → info → web
|
||||
bin/brain/search.go "onlyoffice postgres" --root facts
|
||||
bin/brain/search.go "where is cs-lexicon" --json | yq '.'
|
||||
bin/brain/search.go "upstream flag" --no-web # local graph only
|
||||
bin/brain/get.go <id> --body # full chunk on demand
|
||||
bin/brain/stats.go # index health
|
||||
bin/brain/eval.go # recall@5 gate
|
||||
```
|
||||
|
||||
`--hop` is not implemented (needs File/FROM_FILE edges); the flag errors instead of walking. `bin/kb/search` is a deprecated wrapper around `bin/brain/search.go`.
|
||||
|
||||
Git history is read with [go-git](https://github.com/go-git/go-git) (no git binary):
|
||||
|
||||
```bash
|
||||
bin/git/import.go --json --limit 100 # commit leafs for this repo
|
||||
bin/git/import.go --root "$PROJECTS_ROOT" --json # one pass per .git under root
|
||||
```
|
||||
|
||||
Conversion only. Graph write (`File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`) stays with `bin/brain/index.go`.
|
||||
|
||||
Web search (second independent source) goes through SearXNG. Empty results mean **throttled**, not “nothing exists”:
|
||||
|
||||
```bash
|
||||
bin/web/search.go "LadybugDB vector index" --json
|
||||
# Optional local instance (skip if BRAIN_SEARCH_URL already points at one):
|
||||
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
|
||||
```
|
||||
|
||||
Mail is a first-class corpus (retrievable through the same search):
|
||||
|
||||
```bash
|
||||
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
|
||||
bin/mail/import.go --from-raw var/mail # JSON → markdown
|
||||
bin/brain/index.go --rebuild # rebuild brain (incl. mail)
|
||||
bin/brain/search.go "invoice from last week" # same search over mail leafs
|
||||
```
|
||||
|
||||
## Storage
|
||||
|
||||
- **LadybugDB** — single `var/kb.lbug`, Cypher property graph, HNSW + BM25
|
||||
in one engine, embedded (no server), ACID, read-only-safe for concurrent
|
||||
readers.
|
||||
readers. Read tools (`get` / `stats` / `eval`) are Go + Zig CGO (`bin/cgo/zcc`); Python
|
||||
`bin/kb/{get,stats,eval}` is the CI fallback. **Never `DROP INDEX` FTS/VECTOR** on Ladybug 0.19: DROP leaves
|
||||
ghost catalog tables (`_0_Leaf_vec_UPPER`) so recreate fails while
|
||||
`SHOW_INDEXES` omits HNSW. Fresh indexes = delete `var/kb.lbug` +
|
||||
`bin/brain/index.go --rebuild`. Use `ensure_indexes()` after upserts.
|
||||
- **model2vec** — `potion-multilingual-128M` static embeddings (256-dim),
|
||||
CPU-fast, deterministic, no Ollama runtime dependency.
|
||||
- facts and info split semantically by `root` column but written inside the
|
||||
@@ -105,26 +138,27 @@ bin/kb/eval # recall@5 gate
|
||||
|
||||
## Tooling conventions
|
||||
|
||||
`bin/{subject}/{method}` — self-describing: shebang on line 1, usage comment
|
||||
from line 2. bash + python primary; golang via the Go shebang when a compiled
|
||||
helper is right. YAML default output, `--json` for machines. Everything that
|
||||
touches network/db is read-only, throttled, cached. Tests gate every commit.
|
||||
`bin/{subject}/{method}.go` — self-describing: shebang on line 1, usage comment
|
||||
from line 2. Shared code in `internal/`. YAML default output, `--json` for
|
||||
machines. Tests gate every commit. HTTP: `bin/brain/serve.go` calls
|
||||
`internal/brain` in-process (`/health` `/search` `/get` `/stats` `/audit` `/ingest` `/openapi.json` `/mcp`).
|
||||
|
||||
## Development
|
||||
|
||||
```bash
|
||||
uv venv .venv # Python 3.12, uv-managed
|
||||
uv pip install -r requirements.lock.txt # pinned toolchain
|
||||
bin/facts/audit self # lexicon consistency gate
|
||||
go test ./... && python -m unittest discover -s tools -t .
|
||||
bin/facts/audit.go self # lexicon consistency gate
|
||||
go test ./... && python -m unittest discover -s bin/tools -t .
|
||||
```
|
||||
|
||||
Docker (optional, cached model + var volumes):
|
||||
|
||||
```bash
|
||||
docker compose run --rm brain index # (re)index corpus
|
||||
docker compose run --rm brain search "query" # one-shot query
|
||||
docker compose run --rm brain serve # async Go HTTP server
|
||||
docker compose up -d brain # API (Zig CGO serve :8630)
|
||||
docker compose --profile index run --rm index # Python Ladybug rebuild
|
||||
docker compose --profile picoclaw up brain-mcp # MCP on 127.0.0.1:8630
|
||||
docker compose --profile reasoner up -d reasoner # CPU Ollama 127.0.0.1:11435
|
||||
docker compose up brain-watch # auto re-index on change
|
||||
```
|
||||
|
||||
@@ -133,7 +167,10 @@ docker compose up brain-watch # auto re-index on change
|
||||
- [go-second-brain](https://github.com/eSlider/go-second-brain) — the earlier
|
||||
Neo4j + Qdrant + Matrix RAG brain
|
||||
- [agent-skills](https://github.com/eSlider/agent-skills) — upstream
|
||||
skills (`web-search`, `db-yaml`, …) that 2dph integrates
|
||||
- [detective](https://github.com/detective) — the two-source method
|
||||
skills (`web-search`, `postgres`, …) that 2dph integrates
|
||||
- detective method — the two-source method
|
||||
|
||||
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
|
||||
Work board (issues): [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
|
||||
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
|
||||
|
||||
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
|
||||
|
||||
@@ -0,0 +1,3 @@
|
||||
// Commands in this directory are shebang mains (search.go, serve.go, index.go,
|
||||
// get.go, stats.go, eval.go, watch.go), each behind an exclusive build tag.
|
||||
package main
|
||||
Executable
+22
@@ -0,0 +1,22 @@
|
||||
//usr/bin/env go run -tags=system_ladybug,brain_eval "$0" "$@"; exit
|
||||
//go:build cgo && system_ladybug && brain_eval
|
||||
//
|
||||
// bin/brain/eval.go - recall@5 gate.
|
||||
//
|
||||
// ./bin/brain/eval.go
|
||||
// ./bin/brain/eval.go --json
|
||||
//
|
||||
// Needs CGO + libladybug. Python bin/kb/eval is the CI fallback (no cgo).
|
||||
// Control questions live in internal/brain/rank (cgo-free).
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/brain"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(brain.MainEval(os.Args[1:]))
|
||||
}
|
||||
Executable
+23
@@ -0,0 +1,23 @@
|
||||
//usr/bin/env go run -tags=system_ladybug,brain_get "$0" "$@"; exit
|
||||
//go:build cgo && system_ladybug && brain_get
|
||||
//
|
||||
// bin/brain/get.go - read one leaf by id.
|
||||
//
|
||||
// ./bin/brain/get.go <id>
|
||||
// ./bin/brain/get.go <id> --body
|
||||
// ./bin/brain/get.go <id> --json
|
||||
//
|
||||
// Needs CGO + libladybug. Python bin/kb/get is the CI fallback (no cgo).
|
||||
// CGO compiler is Zig (`eval "$(bin/cgo/zig env)"`), not gcc.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/brain"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(brain.MainGet(os.Args[1:]))
|
||||
}
|
||||
Executable
+24
@@ -0,0 +1,24 @@
|
||||
//usr/bin/env go run -tags=brain_index "$0" "$@"; exit
|
||||
//go:build brain_index
|
||||
//
|
||||
// bin/brain/index.go - rebuild the Ladybug graph (Python write path).
|
||||
//
|
||||
// ./bin/brain/index.go --rebuild
|
||||
// ./bin/brain/index.go --rebuild --with-mail
|
||||
// ./bin/brain/index.go --dry-run --with-mail
|
||||
//
|
||||
// v1 write is always a rebuild when mail is included (live FTS/HNSW + bulk
|
||||
// insert corrupts Ladybug 0.19 WAL). `add` is v2.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/cmdbin"
|
||||
)
|
||||
|
||||
func main() {
|
||||
args := append([]string{"--with-mail"}, os.Args[1:]...)
|
||||
os.Exit(cmdbin.ExecFile("bin/kb/index", args))
|
||||
}
|
||||
Executable
+23
@@ -0,0 +1,23 @@
|
||||
//usr/bin/env go run -tags=system_ladybug "$0" "$@"; exit
|
||||
//go:build cgo && system_ladybug
|
||||
//
|
||||
// bin/brain/search.go - deduction search over the 2dph brain.
|
||||
//
|
||||
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--json] [--no-web]
|
||||
// ./bin/brain/search.go serve [port]
|
||||
// ./bin/brain/search.go --list-model
|
||||
//
|
||||
// Needs CGO + libladybug via Zig (`eval "$(bin/cgo/zig env)"`), not gcc.
|
||||
// bin/kb/search which sets those and builds a binary for the embed daemon.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/brain"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(brain.Main(os.Args[1:]))
|
||||
}
|
||||
Executable
+34
@@ -0,0 +1,34 @@
|
||||
//usr/bin/env go run -tags=brain_serve,system_ladybug "$0" "$@"; exit
|
||||
//go:build brain_serve && cgo && system_ladybug
|
||||
//
|
||||
// bin/brain/serve.go - HTTP API (in-process ladybug search).
|
||||
//
|
||||
// KB_ROOT=/path/to/2dph ./bin/brain/serve.go
|
||||
// KB_WORKERS=4 KB_PORT=8630 ./bin/brain/serve.go
|
||||
//
|
||||
// GET /openapi.json same Ops table as the handlers
|
||||
// POST /mcp JSON-RPC tools/list + tools/call
|
||||
//
|
||||
// Needs CGO + libladybug (same as bin/brain/search.go).
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"log"
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/brain"
|
||||
"github.com/eSlider/2dph/internal/httpapi"
|
||||
)
|
||||
|
||||
func main() {
|
||||
if os.Getenv("KB_ROOT") == "" {
|
||||
if wd, err := os.Getwd(); err == nil {
|
||||
os.Setenv("KB_ROOT", wd)
|
||||
}
|
||||
}
|
||||
if err := brain.Ready(); err != nil {
|
||||
log.Fatal(err)
|
||||
}
|
||||
httpapi.Run(brain.HTTP{})
|
||||
}
|
||||
@@ -0,0 +1,20 @@
|
||||
//go:build brain_serve && !system_ladybug
|
||||
//
|
||||
// Fallback serve when ladybug cgo is not in the build (CI / tags=brain_serve).
|
||||
// Production shebang is serve.go (in-process).
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/httpapi"
|
||||
)
|
||||
|
||||
func main() {
|
||||
if os.Getenv("KB_ROOT") == "" {
|
||||
if wd, err := os.Getwd(); err == nil {
|
||||
os.Setenv("KB_ROOT", wd)
|
||||
}
|
||||
}
|
||||
httpapi.Run(nil)
|
||||
}
|
||||
Executable
+21
@@ -0,0 +1,21 @@
|
||||
//usr/bin/env go run -tags=system_ladybug,brain_stats "$0" "$@"; exit
|
||||
//go:build cgo && system_ladybug && brain_stats
|
||||
//
|
||||
// bin/brain/stats.go - index health.
|
||||
//
|
||||
// ./bin/brain/stats.go
|
||||
// ./bin/brain/stats.go --json
|
||||
//
|
||||
// Needs CGO + libladybug. Python bin/kb/stats is the CI fallback (no cgo).
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/brain"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(brain.MainStats(os.Args[1:]))
|
||||
}
|
||||
Executable
+20
@@ -0,0 +1,20 @@
|
||||
//usr/bin/env go run -tags=brain_watch "$0" "$@"; exit
|
||||
//go:build brain_watch
|
||||
//
|
||||
// bin/brain/watch.go - re-index when corpus files change.
|
||||
//
|
||||
// ./bin/brain/watch.go [dir...]
|
||||
// KB_WATCH_INTERVAL=15 ./bin/brain/watch.go
|
||||
//
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/bin/watch"
|
||||
)
|
||||
|
||||
func main() {
|
||||
watch.Run(os.Args[1:])
|
||||
}
|
||||
Executable
+23
@@ -0,0 +1,23 @@
|
||||
#!/bin/sh
|
||||
# bin/cgo/zc++ — CGO CXX. Zig, not g++.
|
||||
set -eu
|
||||
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
|
||||
case "$(uname -m)" in
|
||||
x86_64|amd64) TARGET=x86_64-linux-gnu ;;
|
||||
aarch64|arm64) TARGET=aarch64-linux-gnu ;;
|
||||
*)
|
||||
echo "zc++: unsupported arch $(uname -m)" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
|
||||
:
|
||||
elif [ -x "$ROOT/var/zig/zig" ]; then
|
||||
ZIG="$ROOT/var/zig/zig"
|
||||
elif command -v zig >/dev/null 2>&1; then
|
||||
ZIG="$(command -v zig)"
|
||||
else
|
||||
echo "zc++: zig missing; run bin/cgo/zig first" >&2
|
||||
exit 127
|
||||
fi
|
||||
exec "$ZIG" c++ -target "$TARGET" "$@"
|
||||
Executable
+24
@@ -0,0 +1,24 @@
|
||||
#!/bin/sh
|
||||
# bin/cgo/zcc — CGO CC. Zig, not gcc.
|
||||
# Go invokes CC with many args; a wrapper avoids spaces in $CC.
|
||||
set -eu
|
||||
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
|
||||
case "$(uname -m)" in
|
||||
x86_64|amd64) TARGET=x86_64-linux-gnu ;;
|
||||
aarch64|arm64) TARGET=aarch64-linux-gnu ;;
|
||||
*)
|
||||
echo "zcc: unsupported arch $(uname -m)" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
|
||||
:
|
||||
elif [ -x "$ROOT/var/zig/zig" ]; then
|
||||
ZIG="$ROOT/var/zig/zig"
|
||||
elif command -v zig >/dev/null 2>&1; then
|
||||
ZIG="$(command -v zig)"
|
||||
else
|
||||
echo "zcc: zig missing; run bin/cgo/zig first" >&2
|
||||
exit 127
|
||||
fi
|
||||
exec "$ZIG" cc -target "$TARGET" "$@"
|
||||
Executable
+130
@@ -0,0 +1,130 @@
|
||||
#!/usr/bin/env bash
|
||||
# bin/cgo/zig — CGO toolchain: zig cc (not gcc) + pinned liblbug + libtokenizers.
|
||||
#
|
||||
# eval "$(bin/cgo/zig env)" # export CC/CXX/CGO_*
|
||||
# bin/cgo/zig go build ... # ensure, then exec with env
|
||||
# bin/cgo/zig ./bin/brain/search.go "query"
|
||||
#
|
||||
# Pins live in this file. Downloads land in var/ (gitignored).
|
||||
set -euo pipefail
|
||||
|
||||
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
|
||||
ZIG_VERSION=0.14.1
|
||||
LBUG_VERSION=0.19.1
|
||||
TOKENIZERS_VERSION=1.27.0
|
||||
|
||||
arch="$(uname -m)"
|
||||
case "$arch" in
|
||||
x86_64|amd64)
|
||||
ZIG_ARCH=x86_64
|
||||
LBUG_ARCH=x86_64
|
||||
TOK_ARCH=x86_64
|
||||
ZIG_SHA=24aeeec8af16c381934a6cd7d95c807a8cb2cf7df9fa40d359aa884195c4716c
|
||||
LBUG_SHA=ed263ae913f68cb0ddba0b98548b58edaac49929766d03bdaaa83be46c68847d
|
||||
TOK_SHA=72556cdca798dd4ea7cdaba308e5f0d68a8cb93b67c96edf485b7a0edd7b07f4
|
||||
;;
|
||||
aarch64|arm64)
|
||||
ZIG_ARCH=aarch64
|
||||
LBUG_ARCH=aarch64
|
||||
TOK_ARCH=aarch64
|
||||
ZIG_SHA=f7a654acc967864f7a050ddacfaa778c7504a0eca8d2b678839c21eea47c992b
|
||||
LBUG_SHA=b07df2cd533c3976a2a3025866d6420a5f35514d0a822ecc4b2902d55b4725b7
|
||||
TOK_SHA=e96545ad05930c26f51f63d932ee6d3bbd32bbed149e102c5290d587a2293067
|
||||
;;
|
||||
*)
|
||||
echo "bin/cgo/zig: unsupported arch $arch" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
|
||||
CACHE="$ROOT/var/cache"
|
||||
LIB="$ROOT/lib-ladybug"
|
||||
ZIG_DIR="$ROOT/var/zig-dist"
|
||||
ZIG_BIN="$ROOT/var/zig/zig"
|
||||
|
||||
sha256of() {
|
||||
if command -v sha256sum >/dev/null 2>&1; then
|
||||
sha256sum "$1" | awk '{print $1}'
|
||||
else
|
||||
shasum -a 256 "$1" | awk '{print $1}'
|
||||
fi
|
||||
}
|
||||
|
||||
fetch() {
|
||||
local url="$1" dest="$2" expect="$3"
|
||||
if [ -f "$dest" ] && [ "$(sha256of "$dest")" = "$expect" ]; then
|
||||
return 0
|
||||
fi
|
||||
mkdir -p "$(dirname "$dest")"
|
||||
echo "fetch $url" >&2
|
||||
curl -fsSL "$url" -o "$dest"
|
||||
local got
|
||||
got="$(sha256of "$dest")"
|
||||
if [ "$got" != "$expect" ]; then
|
||||
echo "checksum mismatch $dest: got $got want $expect" >&2
|
||||
rm -f "$dest"
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
ensure_zig() {
|
||||
if [ -n "${ZIG:-}" ] && [ -x "$ZIG" ]; then
|
||||
return 0
|
||||
fi
|
||||
if [ -x "$ZIG_BIN" ]; then
|
||||
export ZIG="$ZIG_BIN"
|
||||
return 0
|
||||
fi
|
||||
if command -v zig >/dev/null 2>&1; then
|
||||
export ZIG
|
||||
ZIG="$(command -v zig)"
|
||||
return 0
|
||||
fi
|
||||
local tar="$CACHE/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}.tar.xz"
|
||||
fetch "https://ziglang.org/download/${ZIG_VERSION}/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}.tar.xz" \
|
||||
"$tar" "$ZIG_SHA"
|
||||
mkdir -p "$CACHE"
|
||||
rm -rf "$ZIG_DIR"
|
||||
tar -xJf "$tar" -C "$CACHE"
|
||||
mv "$CACHE/zig-${ZIG_ARCH}-linux-${ZIG_VERSION}" "$ZIG_DIR"
|
||||
mkdir -p "$ROOT/var/zig"
|
||||
ln -sfn "$ZIG_DIR/zig" "$ZIG_BIN"
|
||||
export ZIG="$ZIG_BIN"
|
||||
}
|
||||
|
||||
ensure_libs() {
|
||||
mkdir -p "$LIB"
|
||||
if [ ! -f "$LIB/liblbug.so" ]; then
|
||||
local tar="$CACHE/liblbug-linux-${LBUG_ARCH}.tar.gz"
|
||||
fetch "https://github.com/LadybugDB/ladybug/releases/download/v${LBUG_VERSION}/liblbug-linux-${LBUG_ARCH}.tar.gz" \
|
||||
"$tar" "$LBUG_SHA"
|
||||
tar -xzf "$tar" -C "$LIB"
|
||||
fi
|
||||
if [ ! -f "$LIB/libtokenizers.a" ]; then
|
||||
local tar="$CACHE/libtokenizers.linux-${TOK_ARCH}.tar.gz"
|
||||
fetch "https://github.com/daulet/tokenizers/releases/download/v${TOKENIZERS_VERSION}/libtokenizers.linux-${TOK_ARCH}.tar.gz" \
|
||||
"$tar" "$TOK_SHA"
|
||||
tar -xzf "$tar" -C "$LIB"
|
||||
fi
|
||||
}
|
||||
|
||||
print_env() {
|
||||
printf 'export ZIG=%q\n' "$ZIG"
|
||||
printf 'export CC=%q\n' "$ROOT/bin/cgo/zcc"
|
||||
printf 'export CXX=%q\n' "$ROOT/bin/cgo/zc++"
|
||||
printf 'export CGO_ENABLED=1\n'
|
||||
printf 'export CGO_CFLAGS=%q\n' "-I$LIB"
|
||||
printf 'export CGO_LDFLAGS=%q\n' "-L$LIB -Wl,-rpath,${CGO_RPATH:-$LIB}"
|
||||
}
|
||||
|
||||
ensure_zig
|
||||
ensure_libs
|
||||
|
||||
cmd="${1:-env}"
|
||||
if [ "$cmd" = "env" ]; then
|
||||
print_env
|
||||
exit 0
|
||||
fi
|
||||
|
||||
eval "$(print_env)"
|
||||
exec "$@"
|
||||
@@ -0,0 +1,29 @@
|
||||
#!/usr/bin/env bash
|
||||
# bin/chats - sync, import, index, facts, apply for Telegram/WhatsApp/LinkedIn.
|
||||
# Builds the chats binary on first run / when source changes, then execs it.
|
||||
set -euo pipefail
|
||||
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd)"
|
||||
BIN="$ROOT/var/bin/chats"
|
||||
SRC="$ROOT/bin/chats"
|
||||
|
||||
mkdir -p "$ROOT/var/bin"
|
||||
|
||||
need_build=0
|
||||
if [ ! -x "$BIN" ]; then
|
||||
need_build=1
|
||||
else
|
||||
while IFS= read -r -d '' f; do
|
||||
if [ "$f" -nt "$BIN" ]; then
|
||||
need_build=1
|
||||
break
|
||||
fi
|
||||
done < <(find "$SRC" -name '*.go' -print0 2>/dev/null)
|
||||
fi
|
||||
|
||||
if [ "$need_build" -eq 1 ]; then
|
||||
echo "Building chats..." >&2
|
||||
(cd "$SRC" && go build -o "$BIN" .) || exit 1
|
||||
fi
|
||||
|
||||
exec "$BIN" "$@"
|
||||
Executable
+19
@@ -0,0 +1,19 @@
|
||||
//usr/bin/env go run -tags=chats_apply "$0" "$@"; exit
|
||||
//go:build chats_apply
|
||||
//
|
||||
// bin/chats/apply.go - push extracted chat facts to OnlyOffice CRM.
|
||||
//
|
||||
// ./bin/chats/apply.go [--dry-run]
|
||||
//
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/chats"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(chats.RunApply(os.Args[1:]))
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
// Commands in this directory are shebang mains (sync.go, import.go, facts.go,
|
||||
// apply.go), each behind an exclusive build tag so `go build ./bin/chats`
|
||||
// does not see two mains. Shared code lives in internal/chats.
|
||||
package main
|
||||
Executable
+20
@@ -0,0 +1,20 @@
|
||||
//usr/bin/env go run -tags=chats_facts "$0" "$@"; exit
|
||||
//go:build chats_facts
|
||||
//
|
||||
// bin/chats/facts.go - extract phone/email/linkedin facts from JSONL.
|
||||
//
|
||||
// ./bin/chats/facts.go
|
||||
//
|
||||
// Writes var/chats/facts/. Does not index the brain.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/chats"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(chats.RunFacts(os.Args[1:]))
|
||||
}
|
||||
Executable
+20
@@ -0,0 +1,20 @@
|
||||
//usr/bin/env go run -tags=chats_import "$0" "$@"; exit
|
||||
//go:build chats_import
|
||||
//
|
||||
// bin/chats/import.go - JSONL → markdown under var/chats/md/.
|
||||
//
|
||||
// ./bin/chats/import.go
|
||||
//
|
||||
// Conversion only. Brain ingest is bin/brain/index.go, not this command.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/chats"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(chats.RunImport(os.Args[1:]))
|
||||
}
|
||||
Executable
+155
@@ -0,0 +1,155 @@
|
||||
#!/usr/bin/env python3
|
||||
"""chats/refresh-linkedin-session - refresh LinkedIn MCP session from webtop CDP.
|
||||
|
||||
bin/chats/refresh-linkedin-session [--cdp URL] [--root DIR]
|
||||
"""
|
||||
|
||||
Reads the current LinkedIn cookies out of the running Thorium browser in the
|
||||
work-webtop container via CDP (Network.getAllCookies), copies the live browser
|
||||
profile onto the source profile directory, and rewrites the portable
|
||||
cookies.json + source-state.json that mcp-server-linkedin requires.
|
||||
|
||||
Usage:
|
||||
refresh-linkedin-session [--cdp http://127.0.0.1:9222] [--root /var/tmp/liprofile]
|
||||
[--container work-webtop] [--profile thorium-profile]
|
||||
|
||||
After the headless driver uses a copied profile, LinkedIn rotates the session
|
||||
in that copy, so this must run before every sync.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import urllib.request
|
||||
|
||||
import websockets
|
||||
|
||||
|
||||
def cdp_tab(ws_json):
|
||||
for t in ws_json:
|
||||
if t.get("webSocketDebuggerUrl"):
|
||||
return t["webSocketDebuggerUrl"]
|
||||
return None
|
||||
|
||||
|
||||
async def get_cookies(ws_url):
|
||||
async with websockets.connect(ws_url, max_size=50_000_000) as ws:
|
||||
await ws.send(json.dumps({"id": 1, "method": "Network.getAllCookies", "params": {}}))
|
||||
resp = await ws.recv()
|
||||
return json.loads(resp).get("result", {}).get("cookies", [])
|
||||
|
||||
|
||||
def write_source_state(root, profile_dir):
|
||||
# Reuse the linkedin-mcp-server session_state module to write a valid
|
||||
# source-state.json (same schema the daemon reads).
|
||||
try:
|
||||
from linkedin_mcp_server.session_state import canonical, write_source_state
|
||||
|
||||
write_source_state(canonical(__import__("pathlib").Path(profile_dir)))
|
||||
return
|
||||
except Exception:
|
||||
pass
|
||||
# Fallback: minimal schema-compatible state.
|
||||
import uuid
|
||||
|
||||
state = {
|
||||
"version": 1,
|
||||
"source_runtime_id": "linux-amd64-host",
|
||||
"login_generation": str(uuid.uuid4()),
|
||||
"created_at": None,
|
||||
"profile_path": profile_dir,
|
||||
"cookies_path": os.path.join(root, "cookies.json"),
|
||||
}
|
||||
from datetime import datetime, timezone
|
||||
|
||||
state["created_at"] = datetime.now(timezone.utc).isoformat()
|
||||
with open(os.path.join(root, "source-state.json"), "w") as f:
|
||||
json.dump(state, f, indent=2)
|
||||
|
||||
|
||||
def main():
|
||||
args = sys.argv[1:]
|
||||
cdp = "http://127.0.0.1:9222"
|
||||
root = "/var/tmp/liprofile"
|
||||
container = "work-webtop"
|
||||
cprofile = "thorium-profile"
|
||||
for i in range(0, len(args), 2):
|
||||
k = args[i]
|
||||
v = args[i + 1] if i + 1 < len(args) else ""
|
||||
if k == "--cdp":
|
||||
cdp = v
|
||||
elif k == "--root":
|
||||
root = v
|
||||
elif k == "--container":
|
||||
container = v
|
||||
elif k == "--profile":
|
||||
cprofile = v
|
||||
|
||||
profile_dir = os.path.join(root, "profile")
|
||||
os.makedirs(profile_dir, exist_ok=True)
|
||||
|
||||
# 1. Clear stale daemon/browser locks so the server can claim the profile.
|
||||
for lock in ("profile-claim.lock", "profile.lock", "daemon.lock", "lease.lock"):
|
||||
p = os.path.join(root, lock)
|
||||
if os.path.exists(p):
|
||||
os.remove(p)
|
||||
for name in os.listdir(profile_dir):
|
||||
if name.startswith("Singleton"):
|
||||
os.remove(os.path.join(profile_dir, name))
|
||||
for name in os.listdir(root):
|
||||
if name.startswith("invalid-state-"):
|
||||
shutil.rmtree(os.path.join(root, name), ignore_errors=True)
|
||||
|
||||
# 1. Copy the live browser profile (cookies DB + Local State) so the
|
||||
# session the driver launches carries the current login.
|
||||
subprocess.run(
|
||||
["docker", "cp", f"{container}:/config/{cprofile}/Default", os.path.join(profile_dir, "Default")],
|
||||
check=True, capture_output=True,
|
||||
)
|
||||
subprocess.run(
|
||||
["docker", "cp", f"{container}:/config/{cprofile}/Local State", os.path.join(profile_dir, "Local State")],
|
||||
check=True, capture_output=True,
|
||||
)
|
||||
for lock in ("SingletonLock", "SingletonCookie", "SingletonSocket"):
|
||||
p = os.path.join(profile_dir, lock)
|
||||
if os.path.exists(p):
|
||||
os.remove(p)
|
||||
|
||||
# 2. Pull the live cookies out of the running browser.
|
||||
with urllib.request.urlopen(f"{cdp}/json", timeout=5) as r:
|
||||
tabs = json.loads(r.read())
|
||||
ws_url = cdp_tab(tabs)
|
||||
if not ws_url:
|
||||
sys.stderr.write("refresh-linkedin-session: no CDP tab\n")
|
||||
sys.exit(1)
|
||||
cookies = asyncio.run(get_cookies(ws_url))
|
||||
|
||||
li = [c for c in cookies if "linkedin" in c.get("domain", "")]
|
||||
out = []
|
||||
for c in li:
|
||||
domain = c.get("domain", "")
|
||||
if domain in (".www.linkedin.com", "www.linkedin.com"):
|
||||
domain = ".linkedin.com"
|
||||
out.append({
|
||||
"name": c["name"],
|
||||
"value": c["value"].strip('"'),
|
||||
"domain": domain,
|
||||
"path": c.get("path", "/"),
|
||||
"expires": c.get("expires", -1),
|
||||
"httpOnly": c.get("httpOnly", False),
|
||||
"secure": c.get("secure", False),
|
||||
"sameSite": c.get("sameSite", "None"),
|
||||
})
|
||||
with open(os.path.join(root, "cookies.json"), "w") as f:
|
||||
json.dump(out, f, indent=2)
|
||||
|
||||
write_source_state(root, profile_dir)
|
||||
sys.stderr.write(f"refresh-linkedin-session: {len(out)} cookies, profile refreshed\n")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Executable
+41
@@ -0,0 +1,41 @@
|
||||
//usr/bin/env go run -tags=chats_sync "$0" "$@"; exit
|
||||
//go:build chats_sync
|
||||
//
|
||||
// bin/chats/sync.go - download chat messages to var/chats/<platform>/.
|
||||
//
|
||||
// ./bin/chats/sync.go telegram [--limit N] [--phone PHONE]
|
||||
// ./bin/chats/sync.go linkedin [--limit N] [--refresh]
|
||||
//
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/chats"
|
||||
)
|
||||
|
||||
func main() {
|
||||
if len(os.Args) < 2 {
|
||||
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
|
||||
os.Exit(2)
|
||||
}
|
||||
platform := os.Args[1]
|
||||
args := os.Args[2:]
|
||||
switch platform {
|
||||
case "telegram":
|
||||
os.Exit(chats.RunSyncTelegram(args))
|
||||
case "linkedin":
|
||||
os.Exit(chats.RunSyncLinkedIn(args))
|
||||
case "whatsapp":
|
||||
fmt.Fprintln(os.Stderr, "chats: WhatsApp not implemented yet")
|
||||
os.Exit(1)
|
||||
case "help", "-h", "--help":
|
||||
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
|
||||
return
|
||||
default:
|
||||
fmt.Fprintf(os.Stderr, "chats: unknown platform %q\n", platform)
|
||||
os.Exit(2)
|
||||
}
|
||||
}
|
||||
+1
-1
@@ -17,7 +17,7 @@ import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[2] / "tools"))
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools"))
|
||||
|
||||
from semver import bump_type, bump_version # noqa: E402
|
||||
|
||||
|
||||
+3
-1
@@ -34,11 +34,13 @@ case "${1:-}" in
|
||||
;;
|
||||
"")
|
||||
[ -f "$HOME/.ssh/config" ] || { echo "db/ssh-tunnel: ~/.ssh/config missing" >&2; exit 1; }
|
||||
if db/ssh-tunnel --check; then
|
||||
if "$0" --check; then
|
||||
echo "tunnel already up on ${SRC}"
|
||||
exit 0
|
||||
fi
|
||||
ssh -f -N -M -S "$HOME/.ssh/2dph-tunnel.sock" \
|
||||
-o StrictHostKeyChecking=accept-new \
|
||||
-o BatchMode=yes \
|
||||
-L "${SRC}:${DST}" -p "$SSH_PORT" "${SSH_USER}@${SSH_HOST}" \
|
||||
&& echo "tunnel up on ${SRC} (-> vm:${DST})"
|
||||
exit 0
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
// Deprecated shebang mains at bin root (serve.go is tagged brain_serve).
|
||||
package main
|
||||
Regular → Executable
+27
-12
@@ -1,11 +1,10 @@
|
||||
#!/usr/bin/env bash
|
||||
# bin/docker-entrypoint - run 2dph tools inside the container.
|
||||
#
|
||||
# brain shell (default)
|
||||
# brain search <q> bin/kb/search
|
||||
# brain index bin/kb/index
|
||||
# brain watch <dir> watchdog re-indexer
|
||||
# brain serve async Go HTTP server (serve/)
|
||||
# API image (Zig CGO binaries):
|
||||
# serve | search | watch
|
||||
# Index image (Python write path, compose profile `index`):
|
||||
# index | extract | audit | search (deprecated python wrapper)
|
||||
#
|
||||
# Usage comment starts at line 2 (self-describing convention).
|
||||
set -euo pipefail
|
||||
@@ -13,11 +12,27 @@ set -euo pipefail
|
||||
CMD="${1:-shell}"
|
||||
shift || true
|
||||
|
||||
if [ -x /usr/local/bin/brain-serve ]; then
|
||||
case "$CMD" in
|
||||
shell) exec bash ;;
|
||||
serve) exec /usr/local/bin/brain-serve "$@" ;;
|
||||
search) exec /usr/local/bin/brain-search "$@" ;;
|
||||
watch) exec /usr/local/bin/brain-watch "$@" ;;
|
||||
index)
|
||||
echo "index is the Python sidecar: docker compose --profile index run --rm index" >&2
|
||||
exit 2
|
||||
;;
|
||||
*) echo "unknown command: $CMD (api: serve|search|watch)" >&2; exit 2 ;;
|
||||
esac
|
||||
fi
|
||||
|
||||
case "$CMD" in
|
||||
shell) exec bash ;;
|
||||
search) exec "$KB_PY" /app/bin/kb/search "$@" ;;
|
||||
index) exec "$KB_PY" /app/bin/kb/index "$@" ;;
|
||||
watch) exec bash /app/bin/kb-watch "$@" ;;
|
||||
serve) exec /app/serve/serve "$@" ;;
|
||||
*) echo "unknown command: $CMD" >&2; exit 2 ;;
|
||||
esac
|
||||
shell) exec bash ;;
|
||||
search) exec "$KB_PY" /app/bin/kb/search "$@" ;;
|
||||
index) exec "$KB_PY" /app/bin/kb/index --with-mail "$@" ;;
|
||||
watch) exec /app/bin/watch "$@" ;;
|
||||
serve) exec /app/bin/serve "$@" ;;
|
||||
extract) exec "$KB_PY" /app/bin/facts/extract "$@" ;;
|
||||
audit) exec "$KB_PY" /app/bin/facts/audit "$@" ;;
|
||||
*) echo "unknown command: $CMD" >&2; exit 2 ;;
|
||||
esac
|
||||
|
||||
+1
-1
@@ -20,7 +20,7 @@ import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "tools"))
|
||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||
|
||||
|
||||
def audit_db() -> list[str]:
|
||||
|
||||
Executable
+21
@@ -0,0 +1,21 @@
|
||||
//usr/bin/env go run -tags=facts_audit "$0" "$@"; exit
|
||||
//go:build facts_audit
|
||||
//
|
||||
// bin/facts/audit.go - 2-source + lexicon checks.
|
||||
//
|
||||
// ./bin/facts/audit.go self
|
||||
// ./bin/facts/audit.go db
|
||||
//
|
||||
// Python bin/facts/audit is the implementation (CI runs it directly).
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/cmdbin"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(cmdbin.ExecFile("bin/facts/audit", os.Args[1:]))
|
||||
}
|
||||
Executable
+124
@@ -0,0 +1,124 @@
|
||||
#!/usr/bin/env python3
|
||||
"""facts/crm - prove person->company and company->project associations.
|
||||
|
||||
Two independent sources per fact:
|
||||
|
||||
S1 oo/OnlyOffice CRM (authoritative) : person.company_id -> company,
|
||||
project.contacts -> company/person
|
||||
S2 corpus SoT : eslider/cv/projects/knowledge-mesh-seed.yaml
|
||||
(orgs: employer/client/... + projects)
|
||||
|
||||
Only associations supported by BOTH sources are written as root=facts.
|
||||
Mismatches are reported (or, with --fix-crm, printed as oo CLI commands).
|
||||
|
||||
Usage:
|
||||
bin/facts/crm write proven facts (needs var/kb.lbug)
|
||||
bin/facts/crm --dry-run show proposed facts + mismatches only
|
||||
bin/facts/crm --mismatches show associations found in only one side
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||
|
||||
from kblib import upsert_leaf, connect, leaf_id # noqa: E402
|
||||
|
||||
MESH_ENV = os.environ.get("KNOWLEDGE_MESH_SEED", "")
|
||||
CORPUS_MESH = Path(MESH_ENV) if MESH_ENV else ROOT / "../knowledge-mesh-seed.yaml"
|
||||
|
||||
|
||||
def corpus_orgs(raw: str) -> dict[str, dict]:
|
||||
"""Delegate to tools.crmfacts.corpus_orgs (tested in tools/)."""
|
||||
from crmfacts import corpus_orgs as _corpus_orgs
|
||||
return _corpus_orgs(raw)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
dry = "--dry-run" in sys.argv
|
||||
mism = "--mismatches" in sys.argv
|
||||
|
||||
mesh = CORPUS_MESH.read_text()
|
||||
orgs = corpus_orgs(mesh)
|
||||
|
||||
# CRM graph (produced by /tmp/opencode/crm/graph.py -> /tmp/opencode/crm/graph.json)
|
||||
graph = json.load(open("/tmp/opencode/crm/graph.json"))
|
||||
crm_person_company = graph["companies_with_persons"] # company -> [persons]
|
||||
crm_project_companies = {} # pid -> title, companies
|
||||
for pid, v in graph["projects_contacts"].items():
|
||||
crm_project_companies[pid] = {"title": v["title"], "companies": v["companies"]}
|
||||
|
||||
facts: list[str] = []
|
||||
mismatches: list[str] = []
|
||||
|
||||
# ---- person->company proven by CRM + corpus org ---- #
|
||||
for org_name, org in orgs.items():
|
||||
token = org.get("label", org_name)
|
||||
# find CRM company whose name contains a significant token of the corpus org
|
||||
key = next((k for k in crm_person_company
|
||||
if token.split()[0].lower() in k.lower() or any(
|
||||
t.lower() in k.lower() for t in org.get("label", "").split(" / "))),
|
||||
None)
|
||||
persons = crm_person_company.get(key, []) if key else []
|
||||
if persons and org:
|
||||
for p in persons:
|
||||
facts.append(f"{p} is associated with {org.get('label')} "
|
||||
f"(role: {org.get('kind', '?')}, {org.get('period', '')})")
|
||||
elif org and key and not persons:
|
||||
mismatches.append(f"corpus org '{org_name}' ({org.get('label')}) has no CRM persons")
|
||||
elif org and not key:
|
||||
mismatches.append(f"corpus org '{org_name}' ({org.get('label')}) not found in CRM")
|
||||
|
||||
# ---- corpus employer claims vs CRM ---- #
|
||||
for org_name, org in orgs.items():
|
||||
if not org or not org.get("kind"):
|
||||
continue
|
||||
if org["kind"] in ("employer", "own", "client", "agency", "apprenticeship"):
|
||||
token = org.get("label", org_name).split()[0]
|
||||
if not any(token.lower() in k.lower() for k in crm_person_company):
|
||||
mismatches.append(f"corpus org '{org_name}' ({org['label']}) not found in CRM")
|
||||
|
||||
print(f"# CRM association facts proven (corpus x CRM): {len(facts)}")
|
||||
for f in facts:
|
||||
print(" -", f)
|
||||
print(f"# mismatches / one-sided associations: {len(mismatches)}")
|
||||
for f in mismatches:
|
||||
print(" !", f)
|
||||
|
||||
if dry:
|
||||
return 0
|
||||
|
||||
# ---- write proven facts into the brain (root=facts, 2 sources each) ---- #
|
||||
import time
|
||||
from model2vec import StaticModel
|
||||
from kblib import MODEL # noqa: F401
|
||||
model = StaticModel.from_pretrained(MODEL)
|
||||
db, conn = connect(read_only=False)
|
||||
try:
|
||||
r = conn.execute("MATCH (l:Leaf) WHERE l.root='facts' RETURN count(*) AS n")
|
||||
stats_before = r.get_all()[0][0]
|
||||
except Exception:
|
||||
stats_before = 0
|
||||
rev = time.strftime("%Y%m%d-%H%M%S")
|
||||
written = 0
|
||||
for f in facts:
|
||||
src = f"ooCRM x {CORPUS_MESH.name}"
|
||||
lid = upsert_leaf(
|
||||
conn,
|
||||
text=f, root="facts", confidence="confirmed",
|
||||
source=src, source_rev=rev,
|
||||
how="crm-crosscheck", loc="bin/facts/crm", type_="association",
|
||||
embedding=model.encode(f).tolist(),
|
||||
)
|
||||
written += 1
|
||||
conn.close()
|
||||
print(f"# wrote {written} facts into var/kb.lbug (facts was {stats_before})")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Executable
+20
@@ -0,0 +1,20 @@
|
||||
//usr/bin/env go run -tags=facts_crm "$0" "$@"; exit
|
||||
//go:build facts_crm
|
||||
//
|
||||
// bin/facts/crm.go - prove person↔company / company↔project (ooCRM × corpus).
|
||||
//
|
||||
// ./bin/facts/crm.go [--dry-run] [--mismatches]
|
||||
//
|
||||
// Python bin/facts/crm is the implementation. Graph write stays Python.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/cmdbin"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(cmdbin.ExecFile("bin/facts/crm", os.Args[1:]))
|
||||
}
|
||||
+4
-3
@@ -21,7 +21,7 @@ import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "tools"))
|
||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||
|
||||
COMPOSE_FILES = [ROOT / "docker" / "compose.yaml", ROOT / "compose.yaml"]
|
||||
DOC_MARKERS = ["README.md", "PLAN.md", "AGENTS.md"]
|
||||
@@ -176,8 +176,7 @@ def dedupe(facts: list[dict]) -> list[dict]:
|
||||
|
||||
|
||||
def write_facts(facts: list[dict]) -> None:
|
||||
from kblib import connect, init_schema, upsert_leaf
|
||||
from kblib import VAR
|
||||
from kblib import VAR, connect, ensure_indexes, init_schema, upsert_leaf
|
||||
VAR.mkdir(exist_ok=True)
|
||||
db, conn = connect(VAR / "kb.lbug", read_only=False)
|
||||
init_schema(conn)
|
||||
@@ -188,6 +187,8 @@ def write_facts(facts: list[dict]) -> None:
|
||||
upsert_leaf(conn, text=f["text"], root="facts", confidence="confirmed",
|
||||
source=f["source"], source_rev=REPO, how=f["how"],
|
||||
loc=f["loc"], type_="fact", embedding=emb)
|
||||
# Upsert-with-index is safe; never DROP+recreate (ghost catalog kills HNSW).
|
||||
ensure_indexes(conn)
|
||||
conn.close()
|
||||
db.close()
|
||||
|
||||
|
||||
Executable
+20
@@ -0,0 +1,20 @@
|
||||
//usr/bin/env go run -tags=facts_extract "$0" "$@"; exit
|
||||
//go:build facts_extract
|
||||
//
|
||||
// bin/facts/extract.go - acquire confirmed facts (2-source each).
|
||||
//
|
||||
// ./bin/facts/extract.go [--json] [--dry-run]
|
||||
//
|
||||
// Python bin/facts/extract is the implementation. Graph write stays Python.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/cmdbin"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(cmdbin.ExecFile("bin/facts/extract", os.Args[1:]))
|
||||
}
|
||||
Executable
+26
@@ -0,0 +1,26 @@
|
||||
#!/usr/bin/env python3
|
||||
"""git/import — deprecated. Use bin/git/import.go (go-git, no git binary).
|
||||
|
||||
bin/git/import.go [REPO] [--json] [--limit N] [--since DATE]
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
|
||||
|
||||
def main(argv: list[str]) -> int:
|
||||
print(
|
||||
"bin/git/import is deprecated; use bin/git/import.go (go-git)",
|
||||
file=sys.stderr,
|
||||
)
|
||||
target = ROOT / "bin" / "git" / "import.go"
|
||||
os.execvp("go", ["go", "run", str(target), *argv])
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main(sys.argv[1:]))
|
||||
Executable
+143
@@ -0,0 +1,143 @@
|
||||
//usr/bin/env go run "$0" "$@"; exit
|
||||
//
|
||||
// bin/git/import.go - read git history with go-git (no git binary).
|
||||
//
|
||||
// ./bin/git/import.go [REPO]
|
||||
// ./bin/git/import.go --json
|
||||
// ./bin/git/import.go --limit 100 --since 2026-01-01
|
||||
// ./bin/git/import.go --root DIR
|
||||
//
|
||||
// Conversion only: prints commit leafs. Brain write is bin/brain/index.go.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"time"
|
||||
|
||||
"github.com/eSlider/2dph/internal/cmdbin"
|
||||
"github.com/eSlider/2dph/internal/gitlog"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(run(os.Args[1:]))
|
||||
}
|
||||
|
||||
func run(args []string) int {
|
||||
var repo, root, since string
|
||||
limit := 0
|
||||
jsonOut := false
|
||||
i := 0
|
||||
for i < len(args) {
|
||||
a := args[i]
|
||||
switch {
|
||||
case a == "--json":
|
||||
jsonOut = true
|
||||
case a == "--limit" && i+1 < len(args):
|
||||
i++
|
||||
n, err := strconv.Atoi(args[i])
|
||||
if err != nil || n < 0 {
|
||||
fmt.Fprintf(os.Stderr, "git/import: --limit must be a non-negative integer\n")
|
||||
return 2
|
||||
}
|
||||
limit = n
|
||||
case a == "--since" && i+1 < len(args):
|
||||
i++
|
||||
since = args[i]
|
||||
case a == "--root" && i+1 < len(args):
|
||||
i++
|
||||
root = args[i]
|
||||
case a == "-h" || a == "--help":
|
||||
fmt.Fprintln(os.Stderr, `usage: bin/git/import.go [REPO] [--json] [--limit N] [--since DATE] [--root DIR]`)
|
||||
return 0
|
||||
case len(a) > 0 && a[0] != '-':
|
||||
repo = a
|
||||
default:
|
||||
fmt.Fprintf(os.Stderr, "git/import: unknown flag %s\n", a)
|
||||
return 2
|
||||
}
|
||||
i++
|
||||
}
|
||||
|
||||
var sinceT time.Time
|
||||
if since != "" {
|
||||
var err error
|
||||
sinceT, err = parseSince(since)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
|
||||
return 2
|
||||
}
|
||||
}
|
||||
|
||||
repos := []string{}
|
||||
if repo != "" {
|
||||
repos = []string{repo}
|
||||
} else if root != "" {
|
||||
entries, err := os.ReadDir(root)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
for _, e := range entries {
|
||||
p := filepath.Join(root, e.Name())
|
||||
if _, err := os.Stat(filepath.Join(p, ".git")); err == nil {
|
||||
repos = append(repos, p)
|
||||
}
|
||||
}
|
||||
} else {
|
||||
repos = []string{cmdbin.Root()}
|
||||
}
|
||||
|
||||
opt := gitlog.Options{Limit: limit, Since: sinceT}
|
||||
type row struct {
|
||||
Repo string `json:"repo"`
|
||||
Path string `json:"path"`
|
||||
Commits int `json:"commits"`
|
||||
Leafs []gitlog.Leaf `json:"leafs,omitempty"`
|
||||
}
|
||||
var rows []row
|
||||
for _, p := range repos {
|
||||
name, err := gitlog.RepoName(p)
|
||||
if err != nil && name == "" {
|
||||
fmt.Fprintf(os.Stderr, "git/import: %s: %v\n", p, err)
|
||||
continue
|
||||
}
|
||||
cs, err := gitlog.Log(p, opt)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "git/import: %s: %v\n", p, err)
|
||||
return 1
|
||||
}
|
||||
leafs := make([]gitlog.Leaf, 0, len(cs))
|
||||
for _, c := range cs {
|
||||
leafs = append(leafs, gitlog.ToLeaf(c, name))
|
||||
}
|
||||
rows = append(rows, row{Repo: name, Path: p, Commits: len(cs), Leafs: leafs})
|
||||
}
|
||||
|
||||
if jsonOut {
|
||||
enc := json.NewEncoder(os.Stdout)
|
||||
enc.SetIndent("", " ")
|
||||
enc.SetEscapeHTML(false)
|
||||
if err := enc.Encode(rows); err != nil {
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
for _, r := range rows {
|
||||
fmt.Printf("%-24s %5d commits %s\n", r.Repo, r.Commits, r.Path)
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
func parseSince(s string) (time.Time, error) {
|
||||
for _, layout := range []string{time.RFC3339, "2006-01-02"} {
|
||||
if t, err := time.Parse(layout, s); err == nil {
|
||||
return t, nil
|
||||
}
|
||||
}
|
||||
return time.Time{}, fmt.Errorf("cannot parse --since %q", s)
|
||||
}
|
||||
@@ -1,27 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# kb-watch - re-index 2dph when corpus files change.
|
||||
#
|
||||
# kb-watch [dir...] [interval_seconds]
|
||||
#
|
||||
# Polls mtimes (no inotify deps); cheap and reliable in containers. Defaults:
|
||||
# dirs = /corpus (compose) or . ; interval = 30s.
|
||||
set -euo pipefail
|
||||
|
||||
DEFAULT_DIRS="${KB_WATCH_DIRS:-/corpus}"
|
||||
DIRS=("$@")
|
||||
[[ ${#DIRS[@]} -eq 0 ]] && DIRS=(${DEFAULT_DIRS})
|
||||
INTERVAL="${KB_WATCH_INTERVAL:-30}"
|
||||
|
||||
index() { "${KB_PY:-python3}" /app/bin/kb/index; }
|
||||
|
||||
LAST_STAMP=""
|
||||
while true; do
|
||||
STAMP=$(find "${DIRS[@]}" -type f -newermt "-${INTERVAL} seconds" 2>/dev/null \
|
||||
| head -1 | md5sum)
|
||||
if [[ -n "$STAMP" && "$STAMP" != "$LAST_STAMP" ]]; then
|
||||
echo "kb-watch: changes detected, re-indexing" >&2
|
||||
index || echo "kb-watch: index failed; will retry" >&2
|
||||
LAST_STAMP="$STAMP"
|
||||
fi
|
||||
sleep "$INTERVAL"
|
||||
done
|
||||
+1
-1
@@ -13,7 +13,7 @@ import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "tools"))
|
||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||
|
||||
from kblib import open_readonly, query_fts # noqa: E402
|
||||
from yamlout import to_yaml # noqa: E402
|
||||
|
||||
+1
-1
@@ -10,7 +10,7 @@ import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "tools"))
|
||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||
|
||||
from kblib import open_readonly # noqa: E402
|
||||
from yamlout import to_yaml # noqa: E402
|
||||
|
||||
+33
-15
@@ -19,13 +19,14 @@ import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "tools"))
|
||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||
|
||||
from kblib import ( # noqa: E402
|
||||
connect, create_fts_and_vector, init_schema, upsert_leaf,
|
||||
connect, ensure_indexes, init_schema, upsert_leaf,
|
||||
open_readonly, stats,
|
||||
)
|
||||
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
|
||||
from mailleafs import from_mail_root # noqa: E402
|
||||
|
||||
CORPUS_DEFAULTS = ["README.md", "PLAN.md", "AGENTS.md", "docs", "skills"]
|
||||
|
||||
@@ -99,43 +100,60 @@ def main(argv: list[str]) -> int:
|
||||
p = argparse.ArgumentParser(description="build the 2dph brain index")
|
||||
p.add_argument("--corpus", action="append", help="extra markdown dir/file to index (may repeat)")
|
||||
p.add_argument("--rebuild", action="store_true", help="fresh db + indexes")
|
||||
p.add_argument("--with-mail", action="store_true", help="include var/mail message.md leafs")
|
||||
p.add_argument("--since", default="", help="with --with-mail, only messages dated >= YYYY-MM-DD")
|
||||
p.add_argument("--dry-run", action="store_true", help="count leafs, write nothing")
|
||||
p.add_argument(
|
||||
"--skip-indexes",
|
||||
action="store_true",
|
||||
help="write leafs only; caller runs ensure_indexes after seeding facts",
|
||||
)
|
||||
p.add_argument("--limit", type=int, default=0, help="max leafs to embed")
|
||||
p.add_argument("--json", action="store_true")
|
||||
a = p.parse_args(argv)
|
||||
|
||||
from kblib import DB_PATH, VAR
|
||||
VAR.mkdir(exist_ok=True)
|
||||
if a.rebuild and DB_PATH.exists():
|
||||
DB_PATH.unlink()
|
||||
|
||||
leafs = load_corpus(ROOT)
|
||||
if a.corpus:
|
||||
for source in a.corpus:
|
||||
leafs.extend(load_corpus_glob(source))
|
||||
mail_n = 0
|
||||
if a.with_mail:
|
||||
mail = from_mail_root(ROOT / "var" / "mail", since=a.since)
|
||||
mail_n = len(mail)
|
||||
leafs.extend(mail)
|
||||
|
||||
if a.dry_run:
|
||||
msg = {"indexed": 0, "corpus_total": len(leafs), "mail_leafs": mail_n, "dry_run": True}
|
||||
print(json.dumps(msg, indent=2) if a.json else
|
||||
f"brain/index: {len(leafs)} leafs would be indexed (mail={mail_n})")
|
||||
return 0
|
||||
|
||||
VAR.mkdir(exist_ok=True)
|
||||
if a.rebuild and DB_PATH.exists():
|
||||
DB_PATH.unlink()
|
||||
|
||||
db, conn = connect(DB_PATH, read_only=False)
|
||||
init_schema(conn)
|
||||
|
||||
if not (a.rebuild or _already_indexed(conn)):
|
||||
create_fts_and_vector(conn, force=True)
|
||||
# Never DROP FTS/VECTOR (ghost catalog). Write leafs, then ensure indexes
|
||||
# unless --skip-indexes (seed facts first — MERGE under live FTS corrupts it).
|
||||
# --rebuild already deleted kb.lbug above, so CREATE runs on a clean DB.
|
||||
embed = embedder()
|
||||
done, total = index_leafs(conn, leafs, embed, a.limit)
|
||||
create_fts_and_vector(conn, force=(done > 0 or a.rebuild))
|
||||
if not a.skip_indexes:
|
||||
ensure_indexes(conn)
|
||||
s = stats(conn)
|
||||
conn.close()
|
||||
db.close()
|
||||
|
||||
result = {"indexed": done, "corpus_total": total, **{k: v for k, v in s.items() if k in ("total", "by_root")}}
|
||||
if a.skip_indexes:
|
||||
result["indexes"] = "skipped"
|
||||
print(json.dumps(result, indent=2) if a.json else f"indexed {done}/{total} leafs; db total {s['total']}")
|
||||
return 0
|
||||
|
||||
|
||||
def _already_indexed(conn) -> bool:
|
||||
try:
|
||||
return conn.execute("MATCH (l:Leaf) RETURN count(*)").get_all()[0][0] > 0
|
||||
except Exception:
|
||||
return False
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main(sys.argv[1:]))
|
||||
+30
-67
@@ -1,72 +1,35 @@
|
||||
#!/usr/bin/env python3
|
||||
"""kb/search - deduction search over the 2dph brain.
|
||||
#!/usr/bin/env bash
|
||||
# bin/kb/search — deprecated wrapper. Use bin/brain/search.go.
|
||||
# CGO via Zig (bin/cgo/zig), not gcc. Builds a binary then execs it.
|
||||
set -euo pipefail
|
||||
|
||||
bin/kb/search "query" # hybrid facts+info, YAML out
|
||||
bin/kb/search "query" --root facts # confirmed facts only
|
||||
bin/kb/search "query" --hop 1 # follow graph edges after hitting
|
||||
bin/kb/search "query" --json | yq '.'
|
||||
bin/kb/search "query" -n 5 # more results
|
||||
ROOT="$(CDPATH= cd -- "$(dirname "$0")/../.." && pwd)"
|
||||
BIN="$ROOT/var/bin/brain-search"
|
||||
SRC="$ROOT/internal/brain"
|
||||
CMD="$ROOT/bin/brain"
|
||||
|
||||
Deduction order: facts root first (confirmed answers with evidence links),
|
||||
then info root (marked `(not confirmed)`). --root restricts to one root.
|
||||
--hop N walks FROM_FILE edges (sibling leafs in the same source file).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
mkdir -p "$ROOT/var/bin"
|
||||
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
need_build=0
|
||||
if [ ! -x "$BIN" ]; then
|
||||
need_build=1
|
||||
else
|
||||
while IFS= read -r -d '' f; do
|
||||
if [ "$f" -nt "$BIN" ]; then
|
||||
need_build=1
|
||||
break
|
||||
fi
|
||||
done < <(find "$SRC" "$CMD" -name '*.go' -print0 2>/dev/null)
|
||||
fi
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "tools"))
|
||||
if [ "$need_build" -eq 1 ]; then
|
||||
echo "Building brain/search (zig cc)..." >&2
|
||||
(
|
||||
cd "$ROOT" &&
|
||||
eval "$("$ROOT/bin/cgo/zig" env)" &&
|
||||
go build -tags system_ladybug -o "$BIN" ./bin/brain
|
||||
) || exit 1
|
||||
fi
|
||||
|
||||
from kblib import connect, hybrid_search, init_schema, open_readonly, query_fts # noqa: E402
|
||||
from yamlout import to_yaml # noqa: E402
|
||||
import ladybug # noqa: E402
|
||||
|
||||
|
||||
def main(argv: list[str]) -> int:
|
||||
import argparse
|
||||
p = argparse.ArgumentParser(description="deduction search over the brain")
|
||||
p.add_argument("query")
|
||||
p.add_argument("--root", choices=("facts", "info", None), default=None)
|
||||
p.add_argument("--hop", type=int, default=0)
|
||||
p.add_argument("-n", "--limit", type=int, default=10)
|
||||
p.add_argument("--json", action="store_true")
|
||||
a = p.parse_args(argv)
|
||||
|
||||
try:
|
||||
db, conn = open_readonly()
|
||||
except FileNotFoundError as e:
|
||||
print(e, file=sys.stderr)
|
||||
return 1
|
||||
|
||||
from model2vec import StaticModel
|
||||
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
|
||||
emb = model.encode([a.query])[0].astype(float).tolist()
|
||||
|
||||
rhs: list[dict] = []
|
||||
try:
|
||||
rhs = query_fts(conn, a.query, a.limit * 2)
|
||||
except Exception:
|
||||
rhs = []
|
||||
|
||||
results = hybrid_search(conn, emb, rhs, a.limit)
|
||||
if a.root:
|
||||
results = [h for h in results if h["root"] == a.root]
|
||||
|
||||
for hit in results:
|
||||
hit.pop("rrf", None)
|
||||
if hit.get("text"):
|
||||
hit["snippet"] = hit["text"][:280]
|
||||
|
||||
out = {"query": a.query, "root_filter": a.root or "facts+info",
|
||||
"count": len(results), "results": results}
|
||||
print(json.dumps(out, indent=2, ensure_ascii=False) if a.json else to_yaml(out))
|
||||
conn.close()
|
||||
db.close()
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main(sys.argv[1:]))
|
||||
echo "bin/kb/search is deprecated; use bin/brain/search.go" >&2
|
||||
exec "$BIN" "$@"
|
||||
|
||||
+1
-1
@@ -11,7 +11,7 @@ import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "tools"))
|
||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||
|
||||
from kblib import open_readonly, stats # noqa: E402
|
||||
from yamlout import to_yaml # noqa: E402
|
||||
|
||||
Executable
+23
@@ -0,0 +1,23 @@
|
||||
//usr/bin/env go run "$0" "$@"; exit
|
||||
// bin/kb/watch.go — deprecated. Use bin/brain/watch.go.
|
||||
//
|
||||
// Usage:
|
||||
//
|
||||
// ./bin/kb/watch.go [dir...] # dirs default /corpus
|
||||
// KB_WATCH_INTERVAL=15 ./bin/kb/watch.go
|
||||
//
|
||||
// Shebang trick: first line is a Go `//` comment; the real code lives in the
|
||||
// importable package (module path, never a relative import).
|
||||
// NOTE: never run `gofmt -w` on this file - it rewrites `//usr/bin/env` to
|
||||
// `// usr/...` and breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/bin/watch"
|
||||
)
|
||||
|
||||
func main() {
|
||||
watch.Run(os.Args[1:])
|
||||
}
|
||||
Executable
+450
@@ -0,0 +1,450 @@
|
||||
#!/usr/bin/env python3
|
||||
"""mail/import - pull OnlyOffice mails into var/mail/ as markdown.
|
||||
|
||||
bin/mail/import --from-raw var/mail convert Go-synced message.json to md
|
||||
bin/mail/import import newest inbox messages
|
||||
bin/mail/import --folder sent import sent folder
|
||||
bin/mail/import --since 2026-01-01 only messages after a date
|
||||
bin/mail/import --limit 50 cap messages per run
|
||||
bin/mail/import --no-attachments body only, skip attachment conversion
|
||||
bin/mail/import --ocr OCR scanned PDFs/images via docling
|
||||
bin/mail/import --dry-run list messages without writing anything
|
||||
|
||||
Writes one directory per message: var/mail/{folder}/{message_id}/
|
||||
message.md frontmatter + markdown body
|
||||
attachments/ raw attachment files (zips unpacked to _unpacked/)
|
||||
attachments/*.md converted attachment content
|
||||
|
||||
Indexing is a separate step (`bin/brain/index.go --rebuild`): conversion can
|
||||
crash in native docling and must not leave the brain DB mid-transaction.
|
||||
|
||||
Requires ONLYOFFICE_URL/USER/PASS in .env (or env). Idempotent: a message
|
||||
already present (message.md exists) is skipped unless --force.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import urllib.parse
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||
|
||||
from mailconv import ( # noqa: E402
|
||||
ARCHIVE_SUFFIXES,
|
||||
IMAGE_SUFFIXES,
|
||||
LEGACY_OFFICE_SUFFIXES,
|
||||
TEXT_SUFFIXES,
|
||||
html_to_markdown,
|
||||
is_convertible,
|
||||
normalize_markdown,
|
||||
subject_to_filename,
|
||||
zip_extract_safe,
|
||||
)
|
||||
|
||||
import requests # noqa: E402
|
||||
|
||||
FOLDER_IDS = {"inbox": 1, "sent": 2, "drafts": 3, "trash": 4, "spam": 5}
|
||||
DEFAULT_LIMIT = 25
|
||||
|
||||
|
||||
def load_env() -> dict:
|
||||
env = {k: v for k, v in os.environ.items()}
|
||||
envfile = ROOT / ".env"
|
||||
if envfile.exists():
|
||||
for line in envfile.read_text().splitlines():
|
||||
line = line.strip()
|
||||
if not line or line.startswith("#") or "=" not in line:
|
||||
continue
|
||||
k, _, v = line.partition("=")
|
||||
env.setdefault(k.strip(), v.strip().strip("\"'"))
|
||||
url = env.get("ONLYOFFICE_URL") or env.get("OO_URL")
|
||||
user = env.get("ONLYOFFICE_USER") or env.get("OO_USER")
|
||||
password = env.get("ONLYOFFICE_PASS") or env.get("OO_PASSWORD")
|
||||
missing = [n for n, v in (("ONLYOFFICE_URL", url), ("ONLYOFFICE_USER", user),
|
||||
("ONLYOFFICE_PASS", password)) if not v]
|
||||
if missing:
|
||||
sys.exit(f"mail/import: missing {', '.join(missing)} (need .env or env)")
|
||||
return {"url": url.rstrip("/"), "user": user, "password": password}
|
||||
|
||||
|
||||
class OOClient:
|
||||
def __init__(self, conf: dict):
|
||||
self.base = conf["url"]
|
||||
self.session = requests.Session()
|
||||
self.token = None
|
||||
self._login(conf)
|
||||
|
||||
def _login(self, conf: dict) -> None:
|
||||
r = self.session.post(f"{self.base}/api/2.0/authentication.json",
|
||||
json={"userName": conf["user"], "password": conf["password"], "type": 0},
|
||||
timeout=30)
|
||||
r.raise_for_status()
|
||||
body = r.json()
|
||||
self.token = (body.get("response") or {}).get("token", "")
|
||||
if not self.token:
|
||||
sys.exit("mail/import: authentication failed (empty token)")
|
||||
|
||||
def _headers(self) -> dict:
|
||||
return {"Authorization": f"Bearer {self.token}", "Accept": "application/json"}
|
||||
|
||||
def get(self, path: str, params: dict | None = None):
|
||||
r = self.session.get(f"{self.base}{path}", params=params, headers=self._headers(), timeout=30)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
|
||||
def list_messages(self, folder: int, page: int = 1, count: int = DEFAULT_LIMIT) -> list[dict]:
|
||||
data = self.get("/api/2.0/mail/messages",
|
||||
params={"folder": folder, "page": page, "count": count})
|
||||
return data.get("response", [])
|
||||
|
||||
def get_message(self, message_id: str) -> dict:
|
||||
data = self.get(f"/api/2.0/mail/messages/{message_id}")
|
||||
return data.get("response", {})
|
||||
|
||||
def download_attachment(self, attach_id, dest: Path) -> bool:
|
||||
"""Download one attachment via the portal session cookie (.ashx handler)."""
|
||||
url = f"{self.base}/addons/mail/httphandlers/download.ashx?attachid={attach_id}"
|
||||
r = self.session.get(url, timeout=60)
|
||||
if r.status_code != 200:
|
||||
return False
|
||||
dest.parent.mkdir(parents=True, exist_ok=True)
|
||||
dest.write_bytes(r.content)
|
||||
return True
|
||||
|
||||
|
||||
def folder_id(name: str) -> int:
|
||||
if name in FOLDER_IDS:
|
||||
return FOLDER_IDS[name]
|
||||
if name.isdigit():
|
||||
return int(name)
|
||||
sys.exit(f"mail/import: unknown folder '{name}' (use {', '.join(FOLDER_IDS)})")
|
||||
|
||||
|
||||
def safe_attachment_name(att: dict) -> str:
|
||||
name = att.get("fileName") or att.get("storedName") or "attachment"
|
||||
name = re.sub(r"[^\w.\- ]+", "_", name)
|
||||
return name
|
||||
|
||||
|
||||
def convert_file_to_md(path: Path, ocr: bool) -> str | None:
|
||||
"""Convert one attachment file to markdown text; None when not convertible."""
|
||||
suffix = path.suffix.lower()
|
||||
if suffix in TEXT_SUFFIXES:
|
||||
return normalize_markdown(path.read_text(encoding="utf-8", errors="replace"))
|
||||
if suffix in (".docx", ".pptx", ".xlsx", ".html", ".htm", ".epub", ".eml", ".msg"):
|
||||
try:
|
||||
from markitdown import MarkItDown
|
||||
md = MarkItDown()
|
||||
result = md.convert(str(path))
|
||||
return normalize_markdown(result.text_content)
|
||||
except Exception as e:
|
||||
return f"\n<!-- conversion failed: {e} -->\n"
|
||||
if suffix == ".pdf":
|
||||
return _convert_pdf(path, ocr)
|
||||
if suffix in IMAGE_SUFFIXES and ocr:
|
||||
return _convert_pdf(path, ocr)
|
||||
if suffix in LEGACY_OFFICE_SUFFIXES:
|
||||
return _convert_legacy(path)
|
||||
if suffix in ARCHIVE_SUFFIXES:
|
||||
return None # handled by caller (unpack + recurse)
|
||||
return None
|
||||
|
||||
|
||||
def _convert_pdf(path: Path, ocr: bool) -> str:
|
||||
"""Convert one PDF to markdown.
|
||||
|
||||
Fast path: poppler's pdftotext (-layout) extracts exact text from
|
||||
born-digital PDFs in ~15ms vs docling's 1-3s. Only textless PDFs (scanned
|
||||
pages, layout-heavy) fall back to docling, which runs isolated in a
|
||||
subprocess because its native onnx/RT-DETR has segfaulted the main process.
|
||||
"""
|
||||
text = _pdf_fast_text(path)
|
||||
if ocr or text is None or not text.strip():
|
||||
return _convert_pdf_docling(path, ocr)
|
||||
return normalize_markdown(text)
|
||||
|
||||
|
||||
def _pdf_fast_text(path: Path) -> str | None:
|
||||
"""pdftotext -layout; None when poppler is unavailable (or the PDF has no text layer)."""
|
||||
try:
|
||||
proc = subprocess.run(
|
||||
["pdftotext", "-layout", str(path), "-"],
|
||||
capture_output=True, timeout=60)
|
||||
except (OSError, subprocess.TimeoutExpired):
|
||||
return None
|
||||
if proc.returncode != 0:
|
||||
return None
|
||||
return proc.stdout.decode("utf-8", errors="replace")
|
||||
|
||||
|
||||
def _convert_pdf_docling(path: Path, ocr: bool) -> str:
|
||||
try:
|
||||
proc = subprocess.run(
|
||||
[sys.executable, os.path.abspath(__file__), "--pdf-worker", str(path),
|
||||
"--ocr" if ocr else "--no-ocr"],
|
||||
capture_output=True, text=True, timeout=600)
|
||||
except subprocess.TimeoutExpired:
|
||||
return "\n<!-- pdf conversion timed out -->\n"
|
||||
if proc.returncode != 0:
|
||||
tail = proc.stderr.strip().splitlines()[-3:]
|
||||
return f"\n<!-- pdf conversion failed: {proc.returncode}: {' | '.join(tail)} -->\n"
|
||||
return proc.stdout
|
||||
|
||||
|
||||
def _pdf_worker(path: Path, ocr: bool) -> None:
|
||||
"""docling worker entry: prints converted markdown on stdout, exits non-zero on error."""
|
||||
try:
|
||||
from docling.document_converter import DocumentConverter, PdfFormatOption
|
||||
from docling.datamodel.pipeline_options import PdfPipelineOptions
|
||||
opts = PdfPipelineOptions()
|
||||
opts.do_ocr = bool(ocr)
|
||||
opts.do_table_structure = True
|
||||
conv = DocumentConverter(format_options={"pdf": PdfFormatOption(pipeline_options=opts)})
|
||||
res = conv.convert(str(path))
|
||||
sys.stdout.write(normalize_markdown(res.document.export_to_markdown()))
|
||||
sys.exit(0)
|
||||
except Exception as e:
|
||||
# errors/stacktraces to stderr; the caller only reports a one-liner
|
||||
print(f"pdf-worker: {e}", file=sys.stderr)
|
||||
import traceback
|
||||
traceback.print_exc(file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
|
||||
def _convert_legacy(path: Path) -> str:
|
||||
"""Legacy .doc/.xls/.ppt -> md via pandoc (installed) or a stub."""
|
||||
try:
|
||||
out = subprocess.run(["pandoc", str(path), "-t", "markdown"],
|
||||
capture_output=True, text=True, timeout=120)
|
||||
if out.returncode == 0 and out.stdout.strip():
|
||||
return normalize_markdown(out.stdout)
|
||||
except (FileNotFoundError, subprocess.TimeoutExpired):
|
||||
pass
|
||||
return f"\n<!-- legacy {path.suffix} not convertible (pandoc unavailable) -->\n"
|
||||
|
||||
|
||||
def write_message_md(msg: dict, folder: str, out_dir: Path, target_dir: Path | None = None) -> Path:
|
||||
import yaml
|
||||
body_html = msg.get("htmlBody") or ""
|
||||
body_text = msg.get("textBody") or ""
|
||||
body_md = ""
|
||||
if body_html.strip():
|
||||
body_md = html_to_markdown(body_html)
|
||||
elif body_text.strip():
|
||||
body_md = normalize_markdown(body_text)
|
||||
# accept both OnlyOffice (receivedDate) and Go-sync (receivedAt) date keys
|
||||
date = msg.get("receivedDate") or msg.get("receivedAt") or ""
|
||||
if date and not isinstance(date, str):
|
||||
date = str(date)
|
||||
meta = {
|
||||
"id": msg.get("id"),
|
||||
"source": msg.get("source"),
|
||||
"folder": folder,
|
||||
"subject": msg.get("subject", ""),
|
||||
"from": msg.get("from", ""),
|
||||
"to": msg.get("to", ""),
|
||||
"cc": msg.get("cc", ""),
|
||||
"date": date,
|
||||
"has_attachments": bool(msg.get("hasAttachments")),
|
||||
"mime_message_id": msg.get("mimeMessageId", ""),
|
||||
"calendar_uid": msg.get("calendarUid", ""),
|
||||
"type": "mail",
|
||||
}
|
||||
meta = {k: v for k, v in meta.items() if v not in (None, "")}
|
||||
frontmatter = "---\n" + yaml.safe_dump(meta, sort_keys=False, allow_unicode=True).strip() + "\n---\n"
|
||||
content = f"{frontmatter}\n# {meta.get('subject','')}\n\n{body_md}".strip() + "\n"
|
||||
if target_dir is not None:
|
||||
msg_dir = target_dir
|
||||
else:
|
||||
msg_dir = out_dir / folder / str(meta.get("id"))
|
||||
msg_dir.mkdir(parents=True, exist_ok=True)
|
||||
md_path = msg_dir / "message.md"
|
||||
md_path.write_text(content, encoding="utf-8")
|
||||
return md_path
|
||||
|
||||
|
||||
def convert_attachments(msg: dict, msg_dir: Path, ocr: bool) -> list[dict]:
|
||||
"""Download + convert each attachment; returns [{name, md, raw}] summaries.
|
||||
|
||||
raw file keeps the API storedName (unique hash, avoids collisions); the
|
||||
markdown is named after the friendly fileName when available.
|
||||
|
||||
In --from-raw mode attachments are already on disk (Go sync wrote them;
|
||||
.ics already has a structured .md sidecar). Files with an existing .md
|
||||
sidecar are left as-is, only unconverted raws are converted here.
|
||||
"""
|
||||
out: list[dict] = []
|
||||
atts = msg.get("attachments") or []
|
||||
att_dir = msg_dir / "attachments"
|
||||
for att in atts:
|
||||
aid = att.get("fileId")
|
||||
display = safe_attachment_name(att)
|
||||
stored = att.get("storedName")
|
||||
raw_name = safe_attachment_name({"storedName": stored}) if stored else display
|
||||
raw = att_dir / raw_name
|
||||
if aid and not raw.exists() and OOCLIENT is not None:
|
||||
if not OOCLIENT.download_attachment(aid, raw):
|
||||
out.append({"name": display, "md": "\n<!-- download failed -->\n", "raw": str(raw)})
|
||||
continue
|
||||
if not raw.exists():
|
||||
out.append({"name": display, "md": "\n<!-- raw missing -->\n", "raw": str(raw)})
|
||||
continue
|
||||
md_stem = Path(display).stem or raw.stem
|
||||
md_path = att_dir / f"{md_stem}.md"
|
||||
# Go sync pre-wrote structured .md for .ics; keep it.
|
||||
if not md_path.exists():
|
||||
md_text = _convert_att_recursive(raw, ocr)
|
||||
md_path.write_text(f"# Attachment: {display}\n\n{md_text}\n", encoding="utf-8")
|
||||
else:
|
||||
md_text = md_path.read_text(encoding="utf-8", errors="replace")
|
||||
out.append({"name": display, "md": md_text, "raw": str(raw), "md_file": str(md_path)})
|
||||
return out
|
||||
|
||||
|
||||
def _convert_att_recursive(path: Path, ocr: bool) -> str:
|
||||
if path.suffix.lower() in ARCHIVE_SUFFIXES:
|
||||
parts: list[str] = []
|
||||
unpack = path.parent / "_unpacked" / path.stem
|
||||
files = zip_extract_safe(path, unpack)
|
||||
for f in files:
|
||||
sub = _convert_att_recursive(f, ocr)
|
||||
if sub and sub.strip():
|
||||
parts.append(f"## {f.name}\n\n{sub}")
|
||||
return "\n\n".join(parts) if parts else "\n<!-- empty zip -->\n"
|
||||
text = convert_file_to_md(path, ocr)
|
||||
return text or "\n<!-- not convertible -->\n"
|
||||
|
||||
|
||||
# module-level client for attachment downloads in convert_attachments
|
||||
OOCLIENT: OOClient | None = None
|
||||
|
||||
|
||||
def convert_one(msg_dir: Path, full: dict, folder: str, out_root: Path,
|
||||
ocr: bool, no_attachments: bool, target_dir: Path | None = None) -> dict:
|
||||
"""Write message.md + convert attachments for one message dict.
|
||||
|
||||
Works for both live API messages and the Go-sync message.json shape
|
||||
(source field optional; attachments read from attachments/ dir).
|
||||
target_dir overrides the derived path (used by --from-raw where the
|
||||
directory layout is authoritative, not the message folder field).
|
||||
"""
|
||||
mid = str(full.get("id"))
|
||||
write_message_md(full, folder, out_root, target_dir=target_dir)
|
||||
converted: list[dict] = []
|
||||
if not no_attachments:
|
||||
converted = convert_attachments(full, msg_dir, ocr)
|
||||
return {"id": mid, "subject": full.get("subject", ""),
|
||||
"date": full.get("receivedDate", "") or full.get("receivedAt", ""),
|
||||
"attachments": len(converted)}
|
||||
|
||||
|
||||
def main(argv: list[str]) -> int:
|
||||
global OOCLIENT
|
||||
p = argparse.ArgumentParser(description="pull OnlyOffice mails to var/mail as markdown")
|
||||
p.add_argument("--folder", default="inbox", help="inbox|sent|drafts|trash|spam or numeric id")
|
||||
p.add_argument("--limit", type=int, default=DEFAULT_LIMIT, help="max messages per run")
|
||||
p.add_argument("--offset", type=int, default=0, help="skip N messages")
|
||||
p.add_argument("--since", default="", help="only messages received after YYYY-MM-DD")
|
||||
p.add_argument("--id", action="append", default=[], help="import specific message id (repeatable)")
|
||||
p.add_argument("--from-raw", default="",
|
||||
help="convert Go-synced dirs (var/mail/<folder>/<id>/message.json) to markdown")
|
||||
p.add_argument("--no-attachments", action="store_true", help="skip attachment download+convert")
|
||||
p.add_argument("--ocr", action="store_true", help="OCR scanned PDFs/images via docling")
|
||||
p.add_argument("--force", action="store_true", help="re-import even if message.md exists")
|
||||
p.add_argument("--dry-run", action="store_true", help="list messages, write nothing")
|
||||
p.add_argument("--json", action="store_true")
|
||||
p.add_argument("--pdf-worker", default="", help=argparse.SUPPRESS)
|
||||
p.add_argument("--no-ocr", action="store_true", help=argparse.SUPPRESS)
|
||||
a = p.parse_args(argv)
|
||||
|
||||
if a.pdf_worker:
|
||||
_pdf_worker(Path(a.pdf_worker), ocr=not a.no_ocr)
|
||||
return 0
|
||||
|
||||
conf = load_env()
|
||||
fid = folder_id(a.folder)
|
||||
out_root = ROOT / "var" / "mail"
|
||||
summary: list[dict] = []
|
||||
if a.from_raw:
|
||||
OOCLIENT = None
|
||||
raw_root = Path(a.from_raw)
|
||||
for msg_dir in sorted(raw_root.rglob("message.json")):
|
||||
mid = msg_dir.parent.name
|
||||
entry = {"id": mid, "subject": "", "date": "",
|
||||
"attachments": 0, "skipped": False}
|
||||
md_path = msg_dir.parent / "message.md"
|
||||
if md_path.exists() and not a.force:
|
||||
entry["skipped"] = True
|
||||
summary.append(entry)
|
||||
continue
|
||||
if a.dry_run:
|
||||
entry["skipped"] = "dry-run"
|
||||
summary.append(entry)
|
||||
continue
|
||||
full = json.loads(msg_dir.read_text(encoding="utf-8"))
|
||||
entry.update(convert_one(msg_dir.parent, full, full.get("folder") or a.folder,
|
||||
raw_root, a.ocr, a.no_attachments,
|
||||
target_dir=msg_dir.parent))
|
||||
summary.append(entry)
|
||||
else:
|
||||
OOCLIENT = OOClient(conf)
|
||||
if a.id:
|
||||
messages = [{"id": i} for i in a.id]
|
||||
else:
|
||||
page = 1
|
||||
messages = []
|
||||
want = a.offset + a.limit
|
||||
while len(messages) < want:
|
||||
count = min(DEFAULT_LIMIT, want - len(messages))
|
||||
chunk = OOCLIENT.list_messages(fid, page=page, count=count)
|
||||
if not chunk:
|
||||
break
|
||||
messages.extend(chunk)
|
||||
if len(chunk) < count:
|
||||
break
|
||||
page += 1
|
||||
messages = messages[a.offset:a.offset + a.limit]
|
||||
if a.since:
|
||||
messages = [m for m in messages
|
||||
if (m.get("receivedDate") or "") >= a.since]
|
||||
|
||||
for m in messages:
|
||||
mid = str(m.get("id"))
|
||||
entry = {"id": mid, "subject": m.get("subject", ""),
|
||||
"date": m.get("receivedDate", ""), "attachments": 0, "skipped": False}
|
||||
msg_dir = out_root / a.folder / mid
|
||||
md_path = msg_dir / "message.md"
|
||||
if md_path.exists() and not a.force:
|
||||
entry["skipped"] = True
|
||||
summary.append(entry)
|
||||
continue
|
||||
if a.dry_run:
|
||||
entry["skipped"] = "dry-run"
|
||||
summary.append(entry)
|
||||
continue
|
||||
full = OOCLIENT.get_message(mid)
|
||||
entry.update(convert_one(msg_dir, full, a.folder, out_root, a.ocr, a.no_attachments,
|
||||
target_dir=msg_dir))
|
||||
summary.append(entry)
|
||||
if a.json:
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
else:
|
||||
imported = [e for e in summary if not e["skipped"]]
|
||||
print(f"mail/import: folder={a.folder} checked={len(summary)} "
|
||||
f"imported={len(imported)} (skipped={sum(e['skipped'] is True for e in summary)})")
|
||||
for e in summary:
|
||||
flag = "skip" if e["skipped"] is True else ("dry" if e["skipped"] == "dry-run" else "ok ")
|
||||
print(f" [{flag}] {e['id']} {e['date'][:10]} {e['subject'][:60]}"
|
||||
f" (atts={e['attachments']})")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main(sys.argv[1:]))
|
||||
Executable
+20
@@ -0,0 +1,20 @@
|
||||
//usr/bin/env go run -tags=mail_import "$0" "$@"; exit
|
||||
//go:build mail_import
|
||||
//
|
||||
// bin/mail/import.go - message.json → markdown (no brain write).
|
||||
//
|
||||
// ./bin/mail/import.go --from-raw var/mail
|
||||
//
|
||||
// Indexing is bin/brain/index.go --rebuild, not this command.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/cmdbin"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(cmdbin.ExecFile("bin/mail/import", os.Args[1:]))
|
||||
}
|
||||
Executable
+27
@@ -0,0 +1,27 @@
|
||||
#!/usr/bin/env python3
|
||||
"""mail/index_mail — deprecated. Use bin/brain/index.go --rebuild --with-mail.
|
||||
|
||||
Ladybug corrupts its WAL on bulk-insert into an already-indexed DB, so this
|
||||
shim always rebuilds (repo corpus + var/mail). Conversion stays in mail/import.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
|
||||
|
||||
def main(argv: list[str]) -> int:
|
||||
print(
|
||||
"bin/mail/index_mail is deprecated; use bin/brain/index.go --rebuild --with-mail",
|
||||
file=sys.stderr,
|
||||
)
|
||||
index = ROOT / "bin" / "kb" / "index"
|
||||
os.execv(sys.executable, [sys.executable, str(index), "--rebuild", "--with-mail", *argv])
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main(sys.argv[1:]))
|
||||
Executable
+25
@@ -0,0 +1,25 @@
|
||||
//usr/bin/env go run "$0" "$@"; exit
|
||||
// bin/mail/sync.go - async download of OnlyOffice and Gmail mail to var/mail/.
|
||||
//
|
||||
// ./bin/mail/sync.go --source onlyoffice,gmail --limit 50 --workers 8
|
||||
// ./bin/mail/sync.go --source gmail --force
|
||||
// ./bin/mail/sync.go --dry-run
|
||||
//
|
||||
// Writes raw message.json + attachments under var/mail/<folder>/<id>/; run
|
||||
// bin/mail/import.go --from-raw afterwards to convert everything to markdown.
|
||||
//
|
||||
// Shebang trick: first line is a Go `//` comment; the real code lives in the
|
||||
// importable package (module path, never a relative import).
|
||||
// NOTE: never run `gofmt -w` on this file - it rewrites `//usr/bin/env` to
|
||||
// `// usr/...` and breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/bin/mail/sync"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(sync.Main(os.Args[1:]))
|
||||
}
|
||||
@@ -0,0 +1,156 @@
|
||||
// Package synccmd wires the sync library to a CLI: reads .env, parses flags,
|
||||
// picks sources, prints stats. Kept separate from the library so unit tests
|
||||
// don't depend on os.Args/env.
|
||||
package sync
|
||||
|
||||
import (
|
||||
"context"
|
||||
"flag"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// CLIConfig is a superset of SyncConfig plus flag parsing results.
|
||||
type CLIConfig struct {
|
||||
Sync SyncConfig
|
||||
Env string // .env path; default <cwd>/.env
|
||||
Sources string
|
||||
Help bool
|
||||
}
|
||||
|
||||
// ParseCLI reads os.Args into a CLIConfig. Exit codes: 0 ok, 2 usage.
|
||||
func ParseCLI(args []string) (CLIConfig, int, error) {
|
||||
fs := flag.NewFlagSet("mail/sync", flag.ContinueOnError)
|
||||
var (
|
||||
env = fs.String("env", "", ".env file (default: <cwd>/.env)")
|
||||
out = fs.String("out", "", "var/mail root (default: <cwd>/var/mail)")
|
||||
workers = fs.Int("workers", 4, "concurrent downloads")
|
||||
limit = fs.Int("limit", 0, "max messages per source (0 = all)")
|
||||
offset = fs.Int("offset", 0, "skip first N messages per source")
|
||||
force = fs.Bool("force", false, "overwrite existing message.json + attachments")
|
||||
dryRun = fs.Bool("dry-run", false, "list message counts without writing")
|
||||
query = fs.String("query", "in:inbox", "Gmail search query (gmail source only)")
|
||||
srcs = fs.String("source", "onlyoffice", "comma list: onlyoffice,gmail (default onlyoffice)")
|
||||
help = fs.Bool("help", false, "usage")
|
||||
)
|
||||
fs.SetOutput(os.Stderr)
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return CLIConfig{}, 2, err
|
||||
}
|
||||
if *help || fs.NArg() > 0 {
|
||||
return CLIConfig{Help: true}, 0, nil
|
||||
}
|
||||
wd, err := os.Getwd()
|
||||
if err != nil {
|
||||
return CLIConfig{}, 2, err
|
||||
}
|
||||
if *env == "" {
|
||||
*env = filepath.Join(wd, ".env")
|
||||
}
|
||||
if *out == "" {
|
||||
*out = filepath.Join(wd, "var", "mail")
|
||||
}
|
||||
envVars := readEnv(*env)
|
||||
cfg := SyncConfig{
|
||||
Out: *out,
|
||||
Workers: *workers,
|
||||
Limit: *limit,
|
||||
Offset: *offset,
|
||||
Force: *force,
|
||||
DryRun: *dryRun,
|
||||
Query: *query,
|
||||
Policy: RetryPolicy{},
|
||||
}
|
||||
cli := CLIConfig{Sync: cfg, Env: *env, Sources: *srcs}
|
||||
for _, s := range strings.Split(*srcs, ",") {
|
||||
switch strings.TrimSpace(s) {
|
||||
case "onlyoffice":
|
||||
u := pick(envVars["ONLYOFFICE_URL"], envVars["OO_URL"])
|
||||
user := pick(envVars["ONLYOFFICE_USER"], envVars["OO_USER"])
|
||||
pass := pick(envVars["ONLYOFFICE_PASS"], envVars["OO_PASSWORD"])
|
||||
if u == "" || user == "" || pass == "" {
|
||||
return CLIConfig{}, 2, fmt.Errorf("onlyoffice source needs ONLYOFFICE_URL/USER/PASS in %s", *env)
|
||||
}
|
||||
cfg.OO = &OOConfig{URL: u, User: user, Password: pass}
|
||||
case "gmail":
|
||||
home, _ := os.UserHomeDir()
|
||||
cfg.Gmail = &GmailCredentials{
|
||||
CredentialsPath: filepath.Join(home, ".gmail-mcp", "credentials.json"),
|
||||
KeysPath: filepath.Join(home, ".gmail-mcp", "gcp-oauth.keys.json"),
|
||||
}
|
||||
default:
|
||||
return CLIConfig{}, 2, fmt.Errorf("unknown source %q", s)
|
||||
}
|
||||
}
|
||||
cli.Sync = cfg
|
||||
return cli, 0, nil
|
||||
}
|
||||
|
||||
// Main is the CLI entry: returns process exit code.
|
||||
func Main(args []string) int {
|
||||
cli, code, err := ParseCLI(args)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "mail/sync:", err)
|
||||
return code
|
||||
}
|
||||
if cli.Help {
|
||||
fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail] [--query GMAIL_Q] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]")
|
||||
return 0
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
|
||||
defer cancel()
|
||||
start := time.Now()
|
||||
stats, err := Run(ctx, cli.Sync)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "mail/sync:", err)
|
||||
return 1
|
||||
}
|
||||
if cli.Sync.DryRun {
|
||||
fmt.Printf("mail/sync: dry-run checked=%d (no writes)\n", stats.Checked)
|
||||
return 0
|
||||
}
|
||||
fmt.Printf("mail/sync: checked=%d new=%d skipped=%d failed=%d in %s\n",
|
||||
stats.Checked, stats.New, stats.Skipped, stats.Failed, time.Since(start).Round(time.Millisecond))
|
||||
if stats.Failed > 0 {
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
// readEnv parses KEY=VALUE lines (ignoring comments) with KEY=PATH override.
|
||||
func readEnv(path string) map[string]string {
|
||||
out := map[string]string{}
|
||||
b, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return out
|
||||
}
|
||||
for _, line := range strings.Split(string(b), "\n") {
|
||||
line = strings.TrimSpace(line)
|
||||
if line == "" || strings.HasPrefix(line, "#") || !strings.Contains(line, "=") {
|
||||
continue
|
||||
}
|
||||
k, v, _ := strings.Cut(line, "=")
|
||||
out[strings.TrimSpace(k)] = strings.Trim(strings.TrimSpace(v), "\"'")
|
||||
}
|
||||
// env overrides file
|
||||
for _, kv := range os.Environ() {
|
||||
k, v, ok := strings.Cut(kv, "=")
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if strings.HasPrefix(k, "ONLYOFFICE_") || strings.HasPrefix(k, "OO_") {
|
||||
out[k] = v
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func pick(a, b string) string {
|
||||
if a != "" {
|
||||
return a
|
||||
}
|
||||
return b
|
||||
}
|
||||
@@ -0,0 +1,356 @@
|
||||
package sync
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/base64"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// GmailCredentials holds the OAuth files produced by the gmail MCP
|
||||
// (@gongrzhe/server-gmail-autoauth-mcp) auto-auth flow.
|
||||
type GmailCredentials struct {
|
||||
CredentialsPath string // ~/.gmail-mcp/credentials.json
|
||||
KeysPath string // ~/.gmail-mcp/gcp-oauth.keys.json
|
||||
User string // fixed: the authed account
|
||||
}
|
||||
|
||||
// gmailToken is the JSON shape of credentials.json + refresh response.
|
||||
type gmailToken struct {
|
||||
AccessToken string `json:"access_token"`
|
||||
RefreshToken string `json:"refresh_token"`
|
||||
Expiry int64 `json:"expiry_date"` // ms epoch
|
||||
}
|
||||
|
||||
type gmailKeys struct {
|
||||
Installed *gmailKeyBlock `json:"installed"`
|
||||
Web *gmailKeyBlock `json:"web"`
|
||||
}
|
||||
type gmailKeyBlock struct {
|
||||
ClientID string `json:"client_id"`
|
||||
ClientSecret string `json:"client_secret"`
|
||||
}
|
||||
|
||||
// GmailClient talks to the Gmail REST API using the OAuth refresh token from
|
||||
// ~/.gmail-mcp/. Token is refreshed lazily with a mutex-guarded cache.
|
||||
type GmailClient struct {
|
||||
creds GmailCredentials
|
||||
client *http.Client
|
||||
mu chan struct{}
|
||||
token *gmailToken
|
||||
user string
|
||||
}
|
||||
|
||||
func NewGmailClient(creds GmailCredentials) (*GmailClient, error) {
|
||||
if creds.CredentialsPath == "" {
|
||||
home, _ := os.UserHomeDir()
|
||||
creds.CredentialsPath = filepath.Join(home, ".gmail-mcp", "credentials.json")
|
||||
creds.KeysPath = filepath.Join(home, ".gmail-mcp", "gcp-oauth.keys.json")
|
||||
}
|
||||
g := &GmailClient{
|
||||
creds: creds,
|
||||
client: &http.Client{Timeout: 60 * time.Second},
|
||||
mu: make(chan struct{}, 1),
|
||||
}
|
||||
g.mu <- struct{}{}
|
||||
return g, nil
|
||||
}
|
||||
|
||||
// accessToken returns a fresh bearer token, refreshing via the Google token
|
||||
// endpoint when the cached one is missing or about to expire.
|
||||
func (g *GmailClient) accessToken(ctx context.Context) (string, error) {
|
||||
select {
|
||||
case <-g.mu:
|
||||
case <-ctx.Done():
|
||||
return "", ctx.Err()
|
||||
}
|
||||
defer func() { g.mu <- struct{}{} }()
|
||||
if g.token != nil && g.token.AccessToken != "" && g.token.Expiry > time.Now().UnixMilli()+300_000 {
|
||||
return g.token.AccessToken, nil
|
||||
}
|
||||
return g.refreshLocked(ctx)
|
||||
}
|
||||
|
||||
func (g *GmailClient) refreshLocked(ctx context.Context) (string, error) {
|
||||
cred, err := os.ReadFile(g.creds.CredentialsPath)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("read gmail credentials %s: %w", g.creds.CredentialsPath, err)
|
||||
}
|
||||
var t gmailToken
|
||||
if err := json.Unmarshal(cred, &t); err != nil {
|
||||
return "", fmt.Errorf("parse gmail credentials: %w", err)
|
||||
}
|
||||
if t.RefreshToken == "" {
|
||||
return "", errors.New("gmail credentials.json has no refresh_token (run the gmail MCP auth flow)")
|
||||
}
|
||||
keys, err := os.ReadFile(g.creds.KeysPath)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("read gmail keys %s: %w", g.creds.KeysPath, err)
|
||||
}
|
||||
var k gmailKeys
|
||||
if err := json.Unmarshal(keys, &k); err != nil {
|
||||
return "", fmt.Errorf("parse gmail keys: %w", err)
|
||||
}
|
||||
block := k.Installed
|
||||
if block == nil {
|
||||
block = k.Web
|
||||
}
|
||||
if block == nil {
|
||||
return "", errors.New("gmail gcp-oauth.keys.json has no installed/web block")
|
||||
}
|
||||
|
||||
form := url.Values{}
|
||||
form.Set("client_id", block.ClientID)
|
||||
form.Set("client_secret", block.ClientSecret)
|
||||
form.Set("refresh_token", t.RefreshToken)
|
||||
form.Set("grant_type", "refresh_token")
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodPost, "https://oauth2.googleapis.com/token",
|
||||
strings.NewReader(form.Encode()))
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
|
||||
resp, err := g.client.Do(req)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("gmail token refresh: %w", err)
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
body, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
var e struct {
|
||||
Error string `json:"error"`
|
||||
Desc string `json:"error_description"`
|
||||
}
|
||||
_ = json.Unmarshal(body, &e)
|
||||
if e.Error == "invalid_grant" {
|
||||
return "", fmt.Errorf("gmail OAuth token invalid/expired - re-auth via: npx -y @gongrzhe/server-gmail-autoauth-mcp auth (uses ~/.gmail-mcp)")
|
||||
}
|
||||
return "", fmt.Errorf("gmail token refresh status %d: %s", resp.StatusCode, truncate(string(body), 300))
|
||||
}
|
||||
var out struct {
|
||||
AccessToken string `json:"access_token"`
|
||||
ExpiresIn int64 `json:"expires_in"`
|
||||
}
|
||||
if err := json.Unmarshal(body, &out); err != nil {
|
||||
return "", fmt.Errorf("gmail token refresh parse: %w", err)
|
||||
}
|
||||
g.token = &gmailToken{
|
||||
AccessToken: out.AccessToken,
|
||||
RefreshToken: t.RefreshToken,
|
||||
Expiry: time.Now().UnixMilli() + out.ExpiresIn*1000,
|
||||
}
|
||||
return out.AccessToken, nil
|
||||
}
|
||||
|
||||
// ListIDs returns message ids matching q, walking nextPageToken up to maxIDs
|
||||
// (0 = unlimited). Thread-level pagination via the messages.list endpoint.
|
||||
func (g *GmailClient) ListIDs(ctx context.Context, q string, maxIDs int, pageToken string) (ids []string, next string, err error) {
|
||||
for {
|
||||
params := url.Values{}
|
||||
params.Set("q", q)
|
||||
params.Set("maxResults", "100")
|
||||
if pageToken != "" {
|
||||
params.Set("pageToken", pageToken)
|
||||
}
|
||||
var out struct {
|
||||
Messages []struct {
|
||||
ID string `json:"id"`
|
||||
} `json:"messages"`
|
||||
NextPageToken string `json:"nextPageToken"`
|
||||
}
|
||||
if err := g.getJSON(ctx, "/gmail/v1/users/me/messages?"+params.Encode(), &out); err != nil {
|
||||
return nil, "", err
|
||||
}
|
||||
for _, m := range out.Messages {
|
||||
ids = append(ids, m.ID)
|
||||
if maxIDs > 0 && len(ids) >= maxIDs {
|
||||
return ids, out.NextPageToken, nil
|
||||
}
|
||||
}
|
||||
if out.NextPageToken == "" {
|
||||
break
|
||||
}
|
||||
pageToken = out.NextPageToken
|
||||
}
|
||||
return ids, "", nil
|
||||
}
|
||||
|
||||
// GetMessage fetches a message in format=full and normalizes it.
|
||||
func (g *GmailClient) GetMessage(ctx context.Context, id string) (*Message, error) {
|
||||
var raw struct {
|
||||
ID string `json:"id"`
|
||||
ThreadID string `json:"threadId"`
|
||||
InternalDate string `json:"internalDate"` // ms epoch string
|
||||
Payload gmailPart
|
||||
}
|
||||
path := "/gmail/v1/users/me/messages/" + url.PathEscape(id) + "?format=full"
|
||||
if err := g.getJSON(ctx, path, &raw); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
m := &Message{
|
||||
Source: "gmail",
|
||||
ID: raw.ID,
|
||||
Folder: "gmail",
|
||||
}
|
||||
for _, h := range raw.Payload.Headers {
|
||||
switch strings.ToLower(h.Name) {
|
||||
case "subject":
|
||||
m.Subject = h.Value
|
||||
case "from":
|
||||
m.From = h.Value
|
||||
case "to":
|
||||
m.To = h.Value
|
||||
case "cc":
|
||||
m.CC = h.Value
|
||||
case "bcc":
|
||||
m.BCC = h.Value
|
||||
case "message-id":
|
||||
m.MimeMessageID = h.Value
|
||||
case "date":
|
||||
if t, err := time.Parse(time.RFC1123Z, h.Value); err == nil {
|
||||
m.ReceivedAt = t
|
||||
}
|
||||
}
|
||||
}
|
||||
if ms, err := parseMS(raw.InternalDate); err == nil {
|
||||
m.ReceivedAt = ms
|
||||
}
|
||||
m.TextBody, m.HTMLBody, m.Attachments = collectParts(raw.Payload, "root", m.ID, 0)
|
||||
m.HasAttachments = len(m.Attachments) > 0
|
||||
return m, nil
|
||||
}
|
||||
|
||||
type gmailPart struct {
|
||||
PartID string `json:"partId"`
|
||||
MimeType string `json:"mimeType"`
|
||||
Filename string `json:"filename"`
|
||||
Body gmailBody `json:"body"`
|
||||
Headers []gmailHeader `json:"headers"`
|
||||
Parts []gmailPart `json:"parts"`
|
||||
}
|
||||
type gmailHeader struct {
|
||||
Name string `json:"name"`
|
||||
Value string `json:"value"`
|
||||
}
|
||||
type gmailBody struct {
|
||||
Size int64 `json:"size"`
|
||||
Data string `json:"data"`
|
||||
AttachmentID string `json:"attachmentId"`
|
||||
}
|
||||
|
||||
// collectParts walks the MIME tree: text bodies into plain/html, anything with
|
||||
// a filename into attachments (returned with base64 ids for later download).
|
||||
func collectParts(p gmailPart, mime string, msgID string, depth int) (text, html string, atts []Attachment) {
|
||||
if depth > 16 {
|
||||
return
|
||||
}
|
||||
mt := strings.ToLower(p.MimeType)
|
||||
if p.Filename != "" && mt != "text/plain" && mt != "text/html" {
|
||||
pid := p.PartID
|
||||
if pid == "" {
|
||||
pid = fmt.Sprintf("%d", depth)
|
||||
}
|
||||
// Gmail's attachments API keys off body.attachmentId, not partId.
|
||||
attID := p.Body.AttachmentID
|
||||
if attID == "" {
|
||||
attID = pid
|
||||
}
|
||||
atts = append(atts, Attachment{
|
||||
FileID: msgID + ":" + attID,
|
||||
FileName: p.Filename,
|
||||
StoredName: p.Filename,
|
||||
Size: p.Body.Size,
|
||||
ContentType: p.MimeType,
|
||||
})
|
||||
} else if data, err := base64.URLEncoding.DecodeString(p.Body.Data); err == nil && len(p.Body.Data) > 0 {
|
||||
s := string(data)
|
||||
if mt == "text/html" && html == "" {
|
||||
html = s
|
||||
} else if (mt == "text/plain" || mt == "") && text == "" {
|
||||
text = s
|
||||
}
|
||||
}
|
||||
for _, child := range p.Parts {
|
||||
t, h, a := collectParts(child, mt, msgID, depth+1)
|
||||
if text == "" {
|
||||
text = t
|
||||
}
|
||||
if html == "" {
|
||||
html = h
|
||||
}
|
||||
atts = append(atts, a...)
|
||||
}
|
||||
return
|
||||
}
|
||||
|
||||
// DownloadAttachment fetches an attachment's bytes from the Gmail API.
|
||||
func (g *GmailClient) DownloadAttachment(ctx context.Context, msgID, attID string) ([]byte, error) {
|
||||
// attID format is "<msgId>:<partId>"; the API needs the bare attachment id.
|
||||
partID := attID
|
||||
if i := strings.Index(attID, ":"); i >= 0 {
|
||||
partID = attID[i+1:]
|
||||
}
|
||||
var out struct {
|
||||
Data string `json:"data"`
|
||||
}
|
||||
path := "/gmail/v1/users/me/messages/" + url.PathEscape(msgID) + "/attachments/" + url.PathEscape(partID)
|
||||
if err := g.getJSON(ctx, path, &out); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return base64.URLEncoding.DecodeString(out.Data)
|
||||
}
|
||||
|
||||
func (g *GmailClient) getJSON(ctx context.Context, path string, out any) error {
|
||||
tok, err := g.accessToken(ctx)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
u := "https://gmail.googleapis.com" + path
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
req.Header.Set("Authorization", "Bearer "+tok)
|
||||
resp, err := g.client.Do(req)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
body, _ := io.ReadAll(io.LimitReader(resp.Body, 16<<20))
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
return fmt.Errorf("gmail %s: status %d: %s", path, resp.StatusCode, truncate(string(body), 300))
|
||||
}
|
||||
if out != nil {
|
||||
return json.Unmarshal(body, out)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func parseMS(s string) (time.Time, error) {
|
||||
if s == "" {
|
||||
return time.Time{}, errors.New("empty")
|
||||
}
|
||||
var ms int64
|
||||
if _, err := fmt.Sscanf(s, "%d", &ms); err != nil {
|
||||
return time.Time{}, err
|
||||
}
|
||||
return time.UnixMilli(ms), nil
|
||||
}
|
||||
|
||||
func truncate(s string, n int) string {
|
||||
if len(s) <= n {
|
||||
return s
|
||||
}
|
||||
return s[:n] + "…"
|
||||
}
|
||||
|
||||
var _ = bytes.MinRead
|
||||
@@ -0,0 +1,258 @@
|
||||
package sync
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
"time"
|
||||
"unicode/utf8"
|
||||
|
||||
ics "github.com/arran4/golang-ical"
|
||||
"golang.org/x/text/encoding/charmap"
|
||||
)
|
||||
|
||||
// ICSToMarkdown parses a VCALENDAR/VEVENT payload and renders a compact
|
||||
// structured markdown block: what / when / where / organizer / attendees.
|
||||
// Returns the raw text when the payload is not a calendar.
|
||||
func ICSToMarkdown(data []byte) string {
|
||||
data = normalizeEncoding(data)
|
||||
cal, err := ics.ParseCalendar(strings.NewReader(string(data)))
|
||||
if err != nil {
|
||||
return normalizeMarkdown(string(data))
|
||||
}
|
||||
method := ""
|
||||
for _, p := range cal.CalendarProperties {
|
||||
if p.IANAToken == string(ics.ComponentPropertyMethod) {
|
||||
method = p.Value
|
||||
break
|
||||
}
|
||||
}
|
||||
method = strings.TrimSpace(method)
|
||||
var out []string
|
||||
for _, ev := range cal.Events() {
|
||||
summary := strings.TrimSpace(propValue(ev, ics.ComponentPropertySummary))
|
||||
if summary != "" {
|
||||
out = append(out, "# "+summary)
|
||||
}
|
||||
if when := eventWhen(ev); when != "" {
|
||||
out = append(out, "- **When:** "+when)
|
||||
}
|
||||
if loc := strings.TrimSpace(propValue(ev, ics.ComponentPropertyLocation)); loc != "" {
|
||||
out = append(out, "- **Where:** "+loc)
|
||||
}
|
||||
if desc := strings.TrimSpace(stripHTML(propValue(ev, ics.ComponentPropertyDescription))); desc != "" {
|
||||
out = append(out, "- **What:** "+desc)
|
||||
}
|
||||
if org := propValue(ev, ics.ComponentPropertyOrganizer); org != "" {
|
||||
out = append(out, "- **Organizer:** "+attendeeFmt(org))
|
||||
}
|
||||
for _, a := range ev.Attendees() {
|
||||
cn := strings.TrimSpace(firstParam(a.ICalParameters, "CN"))
|
||||
partstat := string(a.ParticipationStatus())
|
||||
name := cn
|
||||
if name == "" {
|
||||
name = a.Email()
|
||||
}
|
||||
line := name
|
||||
if email := a.Email(); email != "" && email != name {
|
||||
line = name + " <" + email + ">"
|
||||
}
|
||||
if partstat != "" && !strings.EqualFold(partstat, "NEEDS-ACTION") {
|
||||
line += " (" + strings.Title(strings.ToLower(strings.ReplaceAll(partstat, "_", " "))) + ")"
|
||||
}
|
||||
out = append(out, "- **Attendee:** "+line)
|
||||
}
|
||||
}
|
||||
if len(out) == 0 {
|
||||
return normalizeMarkdown(string(data))
|
||||
}
|
||||
if method != "" {
|
||||
out = append([]string{"*Calendar method: " + method + "*"}, out...)
|
||||
}
|
||||
return normalizeMarkdown(strings.Join(out, "\n\n"))
|
||||
}
|
||||
|
||||
func eventWhen(ev *ics.VEvent) string {
|
||||
start, errStart := ev.GetStartAt()
|
||||
end, errEnd := ev.GetEndAt()
|
||||
// All-day events: golang-ical has dedicated getters.
|
||||
if errStart != nil {
|
||||
if allDay, err := ev.GetAllDayStartAt(); err == nil {
|
||||
start = allDay
|
||||
errStart = nil
|
||||
}
|
||||
}
|
||||
if errEnd != nil {
|
||||
if allDay, err := ev.GetAllDayEndAt(); err == nil {
|
||||
end = allDay
|
||||
errEnd = nil
|
||||
}
|
||||
}
|
||||
if errStart != nil {
|
||||
// Non-IANA TZID (e.g. "W. Europe Standard Time"): parse the raw
|
||||
// property text instead of failing.
|
||||
return rawWhen(ev)
|
||||
}
|
||||
if errEnd != nil || end.Equal(start) {
|
||||
return dtFmt(start)
|
||||
}
|
||||
return dtFmt(start) + " → " + dtFmt(end)
|
||||
}
|
||||
|
||||
// rawWhen parses DTSTART/DTEND property values that golang-ical cannot resolve
|
||||
// because the TZID is not an IANA zone. Formats: 20260812T120000 or 20260812.
|
||||
func rawWhen(ev *ics.VEvent) string {
|
||||
start := rawPropValue(ev, ics.ComponentPropertyDtStart)
|
||||
end := rawPropValue(ev, ics.ComponentPropertyDtEnd)
|
||||
if start == "" {
|
||||
return ""
|
||||
}
|
||||
if end == "" || end == start {
|
||||
return rawDTFmt(start)
|
||||
}
|
||||
return rawDTFmt(start) + " → " + rawDTFmt(end)
|
||||
}
|
||||
|
||||
func rawPropValue(ev *ics.VEvent, prop ics.ComponentProperty) string {
|
||||
p := ev.GetProperty(prop)
|
||||
if p == nil {
|
||||
return ""
|
||||
}
|
||||
return p.Value
|
||||
}
|
||||
|
||||
// rawDTFmt turns 20260812T120000 into 2026-08-12 12:00; 20260812 into 2026-08-12.
|
||||
func rawDTFmt(s string) string {
|
||||
s = strings.TrimSpace(s)
|
||||
if len(s) >= 8 && isDigits(s[:8]) {
|
||||
y, m, d := s[:4], s[4:6], s[6:8]
|
||||
if len(s) > 8 && (s[8] == 'T' || s[8] == 't') && len(s) >= 15 && isDigits(s[9:15]) {
|
||||
h, mi := s[9:11], s[11:13]
|
||||
return fmt.Sprintf("%s-%s-%s %s:%s", y, m, d, h, mi)
|
||||
}
|
||||
return fmt.Sprintf("%s-%s-%s", y, m, d)
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
func isDigits(s string) bool {
|
||||
for _, c := range s {
|
||||
if c < '0' || c > '9' {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return s != ""
|
||||
}
|
||||
|
||||
// dtFmt renders a time as local "2006-01-02 15:04" (tz label when meaningful).
|
||||
func dtFmt(t time.Time) string {
|
||||
loc := t.Local()
|
||||
label := ""
|
||||
if loc.Location() != time.Local {
|
||||
label = " " + loc.Location().String()
|
||||
}
|
||||
return loc.Format("2006-01-02 15:04") + label
|
||||
}
|
||||
|
||||
// propertyGetter is satisfied by both *ics.Calendar and *ics.VEvent.
|
||||
type propertyGetter interface {
|
||||
GetProperty(ics.ComponentProperty) *ics.IANAProperty
|
||||
}
|
||||
|
||||
func propValue(ev propertyGetter, prop ics.ComponentProperty) string {
|
||||
p := ev.GetProperty(prop)
|
||||
if p == nil {
|
||||
return ""
|
||||
}
|
||||
return p.Value
|
||||
}
|
||||
|
||||
func firstParam(params map[string][]string, key string) string {
|
||||
if vs, ok := params[key]; ok && len(vs) > 0 {
|
||||
return vs[0]
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
func attendeeFmt(raw string) string {
|
||||
raw = strings.TrimSpace(raw)
|
||||
if i := strings.Index(raw, ":"); i >= 0 {
|
||||
raw = raw[i+1:]
|
||||
}
|
||||
return raw
|
||||
}
|
||||
|
||||
// stripHTML removes tags and decodes entities from an ics DESCRIPTION that may
|
||||
// carry HTML (Outlook/Exchange style), keeping text lines readable.
|
||||
func stripHTML(s string) string {
|
||||
if !strings.Contains(s, "<") {
|
||||
return s
|
||||
}
|
||||
lines := strings.Split(s, "\n")
|
||||
for i, l := range lines {
|
||||
var b strings.Builder
|
||||
depth := 0
|
||||
for j := 0; j < len(l); j++ {
|
||||
c := l[j]
|
||||
if c == '<' {
|
||||
if j+1 < len(l) && l[j+1] == '/' {
|
||||
depth--
|
||||
} else {
|
||||
depth++
|
||||
}
|
||||
for j < len(l) && l[j] != '>' {
|
||||
j++
|
||||
}
|
||||
continue
|
||||
}
|
||||
if c == '>' {
|
||||
continue
|
||||
}
|
||||
if depth == 0 {
|
||||
b.WriteByte(c)
|
||||
}
|
||||
}
|
||||
lines[i] = strings.TrimSpace(b.String())
|
||||
}
|
||||
return strings.Join(lines, "\n")
|
||||
}
|
||||
|
||||
// normalizeMarkdown collapses blank-line runs and strips control chars.
|
||||
func normalizeMarkdown(s string) string {
|
||||
s = strings.ReplaceAll(s, "\x00", "")
|
||||
for _, ch := range []string{"\ufeff", "\u200b", "\u034f", "\u00ad", "\u2007", "\u2008", "\u200a", "\u2002"} {
|
||||
s = strings.ReplaceAll(s, ch, "")
|
||||
}
|
||||
lines := strings.Split(s, "\n")
|
||||
var out []string
|
||||
blank := 0
|
||||
for _, l := range lines {
|
||||
if strings.TrimSpace(l) == "" {
|
||||
blank++
|
||||
if blank > 1 {
|
||||
continue
|
||||
}
|
||||
} else {
|
||||
blank = 0
|
||||
}
|
||||
out = append(out, l)
|
||||
}
|
||||
return strings.Join(out, "\n")
|
||||
}
|
||||
|
||||
// normalizeEncoding re-encodes legacy single-byte text as UTF-8. ICS files
|
||||
// exported by some portals are Latin-1 (e.g. "N\xfcrnberg"); golang-ical
|
||||
// passes the bytes through, producing invalid UTF-8 in the markdown output.
|
||||
// Valid UTF-8 is returned untouched.
|
||||
func normalizeEncoding(data []byte) []byte {
|
||||
if utf8.Valid(data) {
|
||||
return data
|
||||
}
|
||||
dec := charmap.ISO8859_1.NewDecoder()
|
||||
out, err := dec.Bytes(data)
|
||||
if err != nil {
|
||||
return data
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
var _ = fmt.Sprintf // keep fmt import if helpers change
|
||||
@@ -0,0 +1,222 @@
|
||||
package sync
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"net/http"
|
||||
"net/http/cookiejar"
|
||||
"net/url"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// OOConfig mirrors the .env / environment used by bin/mail/import.
|
||||
type OOConfig struct {
|
||||
URL string
|
||||
User string
|
||||
Password string
|
||||
}
|
||||
|
||||
// OOClient is a minimal OnlyOffice API client: authentication.json for the
|
||||
// bearer token plus the session cookie jar required by the .ashx download
|
||||
// handler. It mirrors the endpoint contract bin/mail/import already uses.
|
||||
type OOClient struct {
|
||||
cfg OOConfig
|
||||
client *http.Client
|
||||
mu chan struct{}
|
||||
token string
|
||||
folderID int
|
||||
}
|
||||
|
||||
func NewOOClient(cfg OOConfig, folderID int) (*OOClient, error) {
|
||||
jar, err := cookiejar.New(nil)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
c := &OOClient{
|
||||
cfg: cfg,
|
||||
client: &http.Client{Jar: jar, Timeout: 60 * time.Second},
|
||||
mu: make(chan struct{}, 1),
|
||||
folderID: folderID,
|
||||
}
|
||||
c.mu <- struct{}{}
|
||||
if err := c.authenticate(context.Background()); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return c, nil
|
||||
}
|
||||
|
||||
func (o *OOClient) authenticate(ctx context.Context) error {
|
||||
select {
|
||||
case <-o.mu:
|
||||
case <-ctx.Done():
|
||||
return ctx.Err()
|
||||
}
|
||||
defer func() { o.mu <- struct{}{} }()
|
||||
body, _ := json.Marshal(map[string]any{
|
||||
"userName": o.cfg.User, "password": o.cfg.Password, "type": 0,
|
||||
})
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
|
||||
strings.TrimRight(o.cfg.URL, "/")+"/api/2.0/authentication.json",
|
||||
strings.NewReader(string(body)))
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
resp, err := o.client.Do(req)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
data, _ := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
|
||||
if resp.StatusCode < 200 || resp.StatusCode > 299 {
|
||||
return fmt.Errorf("oo authenticate status %d: %s", resp.StatusCode, truncate(string(data), 200))
|
||||
}
|
||||
var out struct {
|
||||
Response struct {
|
||||
Token string `json:"token"`
|
||||
} `json:"response"`
|
||||
}
|
||||
if err := json.Unmarshal(data, &out); err != nil {
|
||||
return err
|
||||
}
|
||||
if out.Response.Token == "" {
|
||||
return fmt.Errorf("oo authenticate: empty token")
|
||||
}
|
||||
o.token = out.Response.Token
|
||||
return nil
|
||||
}
|
||||
|
||||
// get performs an authenticated GET and decodes the JSON body into out.
|
||||
func (o *OOClient) get(ctx context.Context, path string, out any) error {
|
||||
u := strings.TrimRight(o.cfg.URL, "/") + path
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
req.Header.Set("Authorization", "Bearer "+o.token)
|
||||
req.Header.Set("Accept", "application/json")
|
||||
resp, err := o.client.Do(req)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
data, _ := io.ReadAll(io.LimitReader(resp.Body, 16<<20))
|
||||
if resp.StatusCode < 200 || resp.StatusCode > 299 {
|
||||
return fmt.Errorf("oo %s: status %d: %s", path, resp.StatusCode, truncate(string(data), 300))
|
||||
}
|
||||
if out != nil {
|
||||
return json.Unmarshal(data, out)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// ooMessage mirrors the OnlyOffice mail message JSON (subset we need).
|
||||
type ooMessage struct {
|
||||
ID int `json:"id"`
|
||||
Subject string `json:"subject"`
|
||||
From string `json:"from"`
|
||||
To string `json:"to"`
|
||||
CC string `json:"cc"`
|
||||
BCC string `json:"bcc"`
|
||||
ReceivedDate string `json:"receivedDate"`
|
||||
HTMLBody string `json:"htmlBody"`
|
||||
TextBody string `json:"textBody"`
|
||||
HasAttachments bool `json:"hasAttachments"`
|
||||
MimeMessageID string `json:"mimeMessageId"`
|
||||
Attachments []struct {
|
||||
FileID int `json:"fileId"`
|
||||
FileName string `json:"fileName"`
|
||||
StoredName string `json:"storedName"`
|
||||
Size int64 `json:"size"`
|
||||
ContentType string `json:"contentType"`
|
||||
} `json:"attachments"`
|
||||
}
|
||||
|
||||
// ListIDs returns message ids in the configured folder, paginating pages until
|
||||
// maxIDs is reached (0 = all).
|
||||
func (o *OOClient) ListIDs(ctx context.Context, maxIDs int, page int) (ids []int, next int, err error) {
|
||||
var out struct {
|
||||
Response []ooMessage `json:"response"`
|
||||
}
|
||||
count := 100
|
||||
if maxIDs > 0 && maxIDs < count {
|
||||
count = maxIDs
|
||||
}
|
||||
path := fmt.Sprintf("/api/2.0/mail/messages?folder=%d&page=%d&count=%d", o.folderID, page, count)
|
||||
if err := o.get(ctx, path, &out); err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
for _, m := range out.Response {
|
||||
ids = append(ids, m.ID)
|
||||
if maxIDs > 0 && len(ids) >= maxIDs {
|
||||
break
|
||||
}
|
||||
}
|
||||
next = page + 1
|
||||
return ids, next, nil
|
||||
}
|
||||
|
||||
// GetMessage fetches the full message by id and normalizes into Message.
|
||||
func (o *OOClient) GetMessage(ctx context.Context, id int) (*Message, error) {
|
||||
var out struct {
|
||||
Response ooMessage `json:"response"`
|
||||
}
|
||||
path := fmt.Sprintf("/api/2.0/mail/messages/%d", id)
|
||||
if err := o.get(ctx, path, &out); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
m := out.Response
|
||||
msg := &Message{
|
||||
Source: "onlyoffice",
|
||||
ID: fmt.Sprintf("%d", m.ID),
|
||||
Folder: "oo",
|
||||
Subject: m.Subject,
|
||||
From: m.From,
|
||||
To: m.To,
|
||||
CC: m.CC,
|
||||
BCC: m.BCC,
|
||||
HTMLBody: m.HTMLBody,
|
||||
TextBody: m.TextBody,
|
||||
HasAttachments: m.HasAttachments,
|
||||
MimeMessageID: m.MimeMessageID,
|
||||
}
|
||||
if t, err := time.Parse(time.RFC3339Nano, m.ReceivedDate); err == nil {
|
||||
msg.ReceivedAt = t
|
||||
}
|
||||
for _, a := range m.Attachments {
|
||||
msg.Attachments = append(msg.Attachments, Attachment{
|
||||
FileID: fmt.Sprintf("%d", a.FileID),
|
||||
FileName: a.FileName,
|
||||
StoredName: a.StoredName,
|
||||
Size: a.Size,
|
||||
ContentType: a.ContentType,
|
||||
})
|
||||
}
|
||||
return msg, nil
|
||||
}
|
||||
|
||||
// DownloadAttachment fetches attachment bytes via the .ashx handler, which
|
||||
// requires the session cookie (client.Jar) captured during authenticate().
|
||||
func (o *OOClient) DownloadAttachment(ctx context.Context, fileID string) ([]byte, error) {
|
||||
u := strings.TrimRight(o.cfg.URL, "/") + "/addons/mail/httphandlers/download.ashx?attachid=" + url.QueryEscape(fileID)
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
resp, err := o.client.Do(req)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
data, err := io.ReadAll(io.LimitReader(resp.Body, 64<<20))
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
return nil, fmt.Errorf("oo download attach %s: status %d", fileID, resp.StatusCode)
|
||||
}
|
||||
return data, nil
|
||||
}
|
||||
@@ -0,0 +1,397 @@
|
||||
package sync
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"math"
|
||||
"math/rand"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"sync/atomic"
|
||||
"time"
|
||||
)
|
||||
|
||||
// RetryPolicy is the exponential-backoff strategy applied to transient HTTP
|
||||
// failures (5xx, timeouts, network errors). Callers wrap transient errors with
|
||||
// retryWrap; everything else aborts immediately.
|
||||
type RetryPolicy struct {
|
||||
MaxAttempts int // total attempts (>=1); 0 => 5
|
||||
BaseDelay time.Duration // first backoff; 0 => 250ms
|
||||
MaxDelay time.Duration // cap; 0 => 15s
|
||||
Jitter float64 // 0..1 multiplier; 0 => 0.2
|
||||
}
|
||||
|
||||
func (p RetryPolicy) withDefaults() RetryPolicy {
|
||||
if p.MaxAttempts <= 0 {
|
||||
p.MaxAttempts = 5
|
||||
}
|
||||
if p.BaseDelay <= 0 {
|
||||
p.BaseDelay = 250 * time.Millisecond
|
||||
}
|
||||
if p.MaxDelay <= 0 {
|
||||
p.MaxDelay = 15 * time.Second
|
||||
}
|
||||
if p.Jitter <= 0 {
|
||||
p.Jitter = 0.2
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
// delay returns the wait before attempt n (1-based): base * 2^(n-2) + jitter,
|
||||
// capped at MaxDelay. Attempt 1 waits 0, attempt 2 waits base, then doubles.
|
||||
func (p RetryPolicy) delay(attempt int) time.Duration {
|
||||
if attempt <= 1 {
|
||||
return 0
|
||||
}
|
||||
exp := math.Min(float64(attempt-2), 10)
|
||||
d := float64(p.BaseDelay) * math.Pow(2, exp)
|
||||
if p.Jitter > 0 {
|
||||
d *= 1 - p.Jitter + 2*p.Jitter*rand.Float64()
|
||||
}
|
||||
if d > float64(p.MaxDelay) {
|
||||
d = float64(p.MaxDelay)
|
||||
}
|
||||
return time.Duration(d)
|
||||
}
|
||||
|
||||
type errRetry struct{ err error }
|
||||
|
||||
func (e *errRetry) Error() string { return e.err.Error() }
|
||||
func (e *errRetry) Unwrap() error { return e.err }
|
||||
|
||||
func isRetriable(err error) bool {
|
||||
var r *errRetry
|
||||
return errors.As(err, &r)
|
||||
}
|
||||
|
||||
func retryWrap(err error) error {
|
||||
if err == nil {
|
||||
return nil
|
||||
}
|
||||
if isRetriable(err) {
|
||||
return err
|
||||
}
|
||||
return &errRetry{err: err}
|
||||
}
|
||||
|
||||
// Retry runs fn up to MaxAttempts times with exponential backoff between
|
||||
// attempts. Non-retriable errors abort immediately. Returns the last error.
|
||||
func Retry(ctx context.Context, policy RetryPolicy, fn func() error) error {
|
||||
policy = policy.withDefaults()
|
||||
var err error
|
||||
for attempt := 1; attempt <= policy.MaxAttempts; attempt++ {
|
||||
if err = fn(); err == nil {
|
||||
return nil
|
||||
}
|
||||
if !isRetriable(err) {
|
||||
return err
|
||||
}
|
||||
if attempt == policy.MaxAttempts {
|
||||
return fmt.Errorf("after %d attempts: %w", policy.MaxAttempts, err)
|
||||
}
|
||||
select {
|
||||
case <-time.After(policy.delay(attempt)):
|
||||
case <-ctx.Done():
|
||||
return ctx.Err()
|
||||
}
|
||||
}
|
||||
return err
|
||||
}
|
||||
|
||||
// SyncConfig wires up a sync run.
|
||||
type SyncConfig struct {
|
||||
OO *OOConfig // OnlyOffice source (optional)
|
||||
Gmail *GmailCredentials // Gmail source (optional)
|
||||
Out string // var/mail root; default <repo>/var/mail
|
||||
Workers int // concurrency; default 4
|
||||
Limit int // max messages per source (0 = all)
|
||||
Offset int // skip first N messages per source
|
||||
Force bool // overwrite existing message.json + attachments
|
||||
DryRun bool // list without writing
|
||||
Query string // Gmail search query; default in:inbox
|
||||
Policy RetryPolicy
|
||||
}
|
||||
|
||||
// SyncStats is returned by Run.
|
||||
type SyncStats struct {
|
||||
Checked int
|
||||
New int32
|
||||
Failed int32
|
||||
Skipped int32
|
||||
}
|
||||
|
||||
// Source abstracts the two backends for the worker pool.
|
||||
type Source interface {
|
||||
// ListIDs yields ids (string form) to fetch. cursor resumes pagination.
|
||||
ListIDs(ctx context.Context, limit int, cursor string) (ids []string, next string, err error)
|
||||
Get(ctx context.Context, id string) (*Message, error)
|
||||
DownloadAttachment(ctx context.Context, msg *Message, att Attachment) ([]byte, error)
|
||||
Folder() string
|
||||
}
|
||||
|
||||
type ooSource struct {
|
||||
c *OOClient
|
||||
page int
|
||||
}
|
||||
// gmailAPI is the Gmail client surface gmailSource needs. *GmailClient implements it.
|
||||
type gmailAPI interface {
|
||||
ListIDs(ctx context.Context, q string, maxIDs int, pageToken string) ([]string, string, error)
|
||||
GetMessage(ctx context.Context, id string) (*Message, error)
|
||||
DownloadAttachment(ctx context.Context, msgID, attID string) ([]byte, error)
|
||||
}
|
||||
|
||||
type gmailSource struct {
|
||||
c gmailAPI
|
||||
cur string
|
||||
query string
|
||||
}
|
||||
|
||||
func (s *ooSource) Folder() string { return "inbox" }
|
||||
func (s *gmailSource) Folder() string { return "gmail" }
|
||||
|
||||
func (s *ooSource) ListIDs(ctx context.Context, limit int, cursor string) ([]string, string, error) {
|
||||
page := s.page
|
||||
if page == 0 {
|
||||
page = 1
|
||||
}
|
||||
ids, next, err := s.c.ListIDs(ctx, limit, page)
|
||||
s.page = next
|
||||
strs := make([]string, len(ids))
|
||||
for i, id := range ids {
|
||||
strs[i] = fmt.Sprintf("%d", id)
|
||||
}
|
||||
return strs, "", err
|
||||
}
|
||||
|
||||
func (s *ooSource) Get(ctx context.Context, id string) (*Message, error) {
|
||||
var mid int
|
||||
if _, err := fmt.Sscanf(id, "%d", &mid); err != nil {
|
||||
return nil, fmt.Errorf("oo id %q: %w", id, err)
|
||||
}
|
||||
return s.c.GetMessage(ctx, mid)
|
||||
}
|
||||
|
||||
func (s *ooSource) DownloadAttachment(ctx context.Context, msg *Message, att Attachment) ([]byte, error) {
|
||||
return s.c.DownloadAttachment(ctx, att.FileID)
|
||||
}
|
||||
|
||||
func (s *gmailSource) ListIDs(ctx context.Context, limit int, cursor string) ([]string, string, error) {
|
||||
q := s.query
|
||||
if q == "" {
|
||||
q = "in:inbox"
|
||||
}
|
||||
ids, next, err := s.c.ListIDs(ctx, q, limit, cursor)
|
||||
return ids, next, err
|
||||
}
|
||||
|
||||
func (s *gmailSource) Get(ctx context.Context, id string) (*Message, error) {
|
||||
return s.c.GetMessage(ctx, id)
|
||||
}
|
||||
|
||||
func (s *gmailSource) DownloadAttachment(ctx context.Context, msg *Message, att Attachment) ([]byte, error) {
|
||||
return s.c.DownloadAttachment(ctx, msg.ID, att.FileID)
|
||||
}
|
||||
|
||||
// Run executes the sync across the configured sources with a worker pool.
|
||||
func Run(ctx context.Context, cfg SyncConfig) (*SyncStats, error) {
|
||||
if cfg.Out == "" {
|
||||
cfg.Out = "var/mail"
|
||||
}
|
||||
if cfg.Workers <= 0 {
|
||||
cfg.Workers = 4
|
||||
}
|
||||
if err := os.MkdirAll(cfg.Out, 0o755); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
var sources []Source
|
||||
if cfg.OO != nil {
|
||||
oo, err := NewOOClient(*cfg.OO, 1) // folder inbox
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("onlyoffice auth: %w", err)
|
||||
}
|
||||
sources = append(sources, &ooSource{c: oo})
|
||||
}
|
||||
if cfg.Gmail != nil {
|
||||
gm, err := NewGmailClient(*cfg.Gmail)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("gmail init: %w", err)
|
||||
}
|
||||
sources = append(sources, &gmailSource{c: gm, query: cfg.Query})
|
||||
}
|
||||
if len(sources) == 0 {
|
||||
return nil, errors.New("sync: no source configured (need OO, Gmail, or both)")
|
||||
}
|
||||
|
||||
stats := &SyncStats{}
|
||||
var jobs []struct {
|
||||
src Source
|
||||
id string
|
||||
}
|
||||
for _, src := range sources {
|
||||
ids, _, err := src.ListIDs(ctx, cfg.Offset+cfg.Limit, "")
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("list %s: %w", src.Folder(), err)
|
||||
}
|
||||
if cfg.Offset > 0 {
|
||||
if cfg.Offset >= len(ids) {
|
||||
ids = nil
|
||||
} else {
|
||||
ids = ids[cfg.Offset:]
|
||||
}
|
||||
}
|
||||
if cfg.Limit > 0 && len(ids) > cfg.Limit {
|
||||
ids = ids[:cfg.Limit]
|
||||
}
|
||||
stats.Checked += len(ids)
|
||||
for _, id := range ids {
|
||||
jobs = append(jobs, struct {
|
||||
src Source
|
||||
id string
|
||||
}{src: src, id: id})
|
||||
}
|
||||
}
|
||||
|
||||
var (
|
||||
wg sync.WaitGroup
|
||||
mu sync.Mutex
|
||||
failures []string
|
||||
)
|
||||
jobsCh := make(chan struct {
|
||||
src Source
|
||||
id string
|
||||
})
|
||||
for i := 0; i < cfg.Workers; i++ {
|
||||
wg.Add(1)
|
||||
go func() {
|
||||
defer wg.Done()
|
||||
for j := range jobsCh {
|
||||
status, err := processOne(ctx, j.src, j.id, cfg)
|
||||
switch status {
|
||||
case statusFailed:
|
||||
mu.Lock()
|
||||
failures = append(failures, j.src.Folder()+"/"+j.id+": "+err.Error())
|
||||
mu.Unlock()
|
||||
atomic.AddInt32(&stats.Failed, 1)
|
||||
case statusNew:
|
||||
atomic.AddInt32(&stats.New, 1)
|
||||
case statusSkipped:
|
||||
atomic.AddInt32(&stats.Skipped, 1)
|
||||
}
|
||||
}
|
||||
}()
|
||||
}
|
||||
for _, j := range jobs {
|
||||
select {
|
||||
case jobsCh <- j:
|
||||
case <-ctx.Done():
|
||||
close(jobsCh)
|
||||
wg.Wait()
|
||||
return stats, ctx.Err()
|
||||
}
|
||||
}
|
||||
close(jobsCh)
|
||||
wg.Wait()
|
||||
|
||||
if len(failures) > 0 {
|
||||
fmt.Fprintf(os.Stderr, "sync: %d failures:\n %s\n", len(failures), strings.Join(failures, "\n "))
|
||||
}
|
||||
return stats, nil
|
||||
}
|
||||
|
||||
type status int
|
||||
|
||||
const (
|
||||
statusNew status = iota
|
||||
statusSkipped
|
||||
statusFailed
|
||||
)
|
||||
|
||||
func processOne(ctx context.Context, src Source, id string, cfg SyncConfig) (status, error) {
|
||||
if cfg.DryRun {
|
||||
return statusNew, nil
|
||||
}
|
||||
dir := filepath.Join(cfg.Out, src.Folder(), id)
|
||||
jsonPath := filepath.Join(dir, "message.json")
|
||||
if !cfg.Force {
|
||||
if _, err := os.Stat(jsonPath); err == nil {
|
||||
return statusSkipped, nil
|
||||
}
|
||||
}
|
||||
var msg *Message
|
||||
err := Retry(ctx, cfg.Policy, func() error {
|
||||
m, err := src.Get(ctx, id)
|
||||
if err != nil {
|
||||
return retryWrap(err)
|
||||
}
|
||||
m.Folder = src.Folder() // directory layout is authoritative
|
||||
if err := writeMessage(jsonPath, m); err != nil {
|
||||
return err
|
||||
}
|
||||
msg = m
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
return statusFailed, err
|
||||
}
|
||||
for _, att := range msg.Attachments {
|
||||
attDir := filepath.Join(dir, "attachments")
|
||||
if err := os.MkdirAll(attDir, 0o755); err != nil {
|
||||
return statusFailed, err
|
||||
}
|
||||
attPath := filepath.Join(attDir, sanitize(att.StoredName))
|
||||
if _, err := os.Stat(attPath); err == nil && !cfg.Force {
|
||||
continue
|
||||
}
|
||||
var data []byte
|
||||
err := Retry(ctx, cfg.Policy, func() error {
|
||||
b, err := src.DownloadAttachment(ctx, msg, att)
|
||||
if err != nil {
|
||||
return retryWrap(err)
|
||||
}
|
||||
data = b
|
||||
return os.WriteFile(attPath, b, 0o644)
|
||||
})
|
||||
if err != nil {
|
||||
return statusFailed, fmt.Errorf("attachment %s: %w", att.FileName, err)
|
||||
}
|
||||
// ICS attachments get structured markdown immediately (same name the
|
||||
// Python converter would use: <display stem>.md).
|
||||
if isICS(att.FileName) {
|
||||
stem := att.FileName
|
||||
if i := strings.LastIndex(stem, "."); i >= 0 {
|
||||
stem = stem[:i]
|
||||
}
|
||||
mdPath := filepath.Join(attDir, sanitize(stem)+".md")
|
||||
if err := os.WriteFile(mdPath, []byte(ICSToMarkdown(data)), 0o644); err != nil {
|
||||
return statusFailed, err
|
||||
}
|
||||
}
|
||||
}
|
||||
return statusNew, nil
|
||||
}
|
||||
|
||||
func writeMessage(path string, m *Message) error {
|
||||
b, err := json.MarshalIndent(m, "", " ")
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
|
||||
return err
|
||||
}
|
||||
return os.WriteFile(path, b, 0o644)
|
||||
}
|
||||
|
||||
func sanitize(name string) string {
|
||||
r := strings.NewReplacer("/", "_", "\\", "_", ":", "_", "*", "_", "?", "_", "\"", "_",
|
||||
"<", "_", ">", "_", "|", "_", " ", "_")
|
||||
return r.Replace(name)
|
||||
}
|
||||
|
||||
func isICS(name string) bool {
|
||||
n := strings.ToLower(name)
|
||||
return strings.HasSuffix(n, ".ics") || strings.HasSuffix(n, ".ical")
|
||||
}
|
||||
@@ -0,0 +1,305 @@
|
||||
package sync
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/base64"
|
||||
"errors"
|
||||
"testing"
|
||||
"time"
|
||||
"unicode/utf8"
|
||||
)
|
||||
|
||||
func TestRetrySucceedsOnSecondTry(t *testing.T) {
|
||||
attempts := 0
|
||||
err := Retry(context.Background(), RetryPolicy{BaseDelay: time.Millisecond, MaxDelay: 5 * time.Millisecond}, func() error {
|
||||
attempts++
|
||||
if attempts == 1 {
|
||||
return retryWrap(errors.New("boom"))
|
||||
}
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("unexpected error: %v", err)
|
||||
}
|
||||
if attempts != 2 {
|
||||
t.Fatalf("expected 2 attempts, got %d", attempts)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRetryExhaustsAttempts(t *testing.T) {
|
||||
attempts := 0
|
||||
err := Retry(context.Background(), RetryPolicy{MaxAttempts: 3, BaseDelay: time.Millisecond, MaxDelay: 5 * time.Millisecond}, func() error {
|
||||
attempts++
|
||||
return retryWrap(errors.New("nope"))
|
||||
})
|
||||
if err == nil {
|
||||
t.Fatal("expected error after exhaustion")
|
||||
}
|
||||
if attempts != 3 {
|
||||
t.Fatalf("expected 3 attempts, got %d", attempts)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRetryNonRetriableAbortsImmediately(t *testing.T) {
|
||||
attempts := 0
|
||||
err := Retry(context.Background(), RetryPolicy{MaxAttempts: 5, BaseDelay: time.Millisecond}, func() error {
|
||||
attempts++
|
||||
return errors.New("permanent")
|
||||
})
|
||||
if err == nil {
|
||||
t.Fatal("expected error")
|
||||
}
|
||||
if attempts != 1 {
|
||||
t.Fatalf("expected 1 attempt for non-retriable, got %d", attempts)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRetryRespectsContext(t *testing.T) {
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
cancel()
|
||||
err := Retry(ctx, RetryPolicy{MaxAttempts: 5, BaseDelay: time.Millisecond}, func() error {
|
||||
return retryWrap(errors.New("x"))
|
||||
})
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
t.Fatalf("expected context.Canceled, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDelayGrows(t *testing.T) {
|
||||
p := RetryPolicy{BaseDelay: time.Second, MaxDelay: 30 * time.Second, Jitter: 0}
|
||||
d1 := p.delay(1) // attempt 1 => 0
|
||||
d2 := p.delay(2)
|
||||
d3 := p.delay(3)
|
||||
if d1 != 0 {
|
||||
t.Fatalf("attempt 1 delay should be 0, got %v", d1)
|
||||
}
|
||||
if d2 != time.Second {
|
||||
t.Fatalf("attempt 2 delay should be 1s, got %v", d2)
|
||||
}
|
||||
if d3 != 2*time.Second {
|
||||
t.Fatalf("attempt 3 delay should be 2s, got %v", d3)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSanitize(t *testing.T) {
|
||||
cases := map[string]string{
|
||||
"a/b\\c:d*e": "a_b_c_d_e",
|
||||
"normal.txt": "normal.txt",
|
||||
"../evil": ".._evil",
|
||||
"a b c.pdf": "a_b_c.pdf",
|
||||
}
|
||||
for in, want := range cases {
|
||||
if got := sanitize(in); got != want {
|
||||
t.Errorf("sanitize(%q) = %q, want %q", in, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestIsICS(t *testing.T) {
|
||||
if !isICS("reply.ics") || !isICS("x.ICAL") {
|
||||
t.Fatal("ics extensions not detected")
|
||||
}
|
||||
if isICS("invoice.pdf") {
|
||||
t.Fatal("pdf misdetected as ics")
|
||||
}
|
||||
}
|
||||
|
||||
const fixtureReplyICS = `BEGIN:VCALENDAR
|
||||
METHOD:REPLY
|
||||
PRODID:Microsoft Exchange Server 2010
|
||||
VERSION:2.0
|
||||
BEGIN:VTIMEZONE
|
||||
TZID:W. Europe Standard Time
|
||||
BEGIN:STANDARD
|
||||
DTSTART:16010101T030000
|
||||
TZOFFSETFROM:+0200
|
||||
TZOFFSETTO:+0100
|
||||
RRULE:FREQ=YEARLY;INTERVAL=1;BYDAY=-1SU;BYMONTH=10
|
||||
END:STANDARD
|
||||
BEGIN:DAYLIGHT
|
||||
DTSTART:16010101T020000
|
||||
TZOFFSETFROM:+0100
|
||||
TZOFFSETTO:+0200
|
||||
RRULE:FREQ=YEARLY;INTERVAL=1;BYDAY=-1SU;BYMONTH=3
|
||||
END:DAYLIGHT
|
||||
END:VTIMEZONE
|
||||
BEGIN:VEVENT
|
||||
ATTENDEE;PARTSTAT=ACCEPTED;CN="Baker, Ben":mailto:bbaker1@teksystems.com
|
||||
UID:bvlnr1i35ug30kn6rvu9dop00g@google.com
|
||||
SUMMARY;LANGUAGE=en-US:Accepted: Appointment (Ben Baker)
|
||||
DTSTART;TZID=W. Europe Standard Time:20260812T120000
|
||||
DTEND;TZID=W. Europe Standard Time:20260812T123000
|
||||
CLASS:PUBLIC
|
||||
STATUS:CONFIRMED
|
||||
LOCATION;LANGUAGE=en-US:https://meet.google.com/sxh-ubud-jrd
|
||||
END:VEVENT
|
||||
END:VCALENDAR`
|
||||
|
||||
func TestICSToMarkdown(t *testing.T) {
|
||||
out := ICSToMarkdown([]byte(fixtureReplyICS))
|
||||
for _, want := range []string{
|
||||
"Accepted: Appointment",
|
||||
"When:",
|
||||
"Where:",
|
||||
"meet.google.com",
|
||||
"Attendee:",
|
||||
"Baker, Ben",
|
||||
"Accepted",
|
||||
"Calendar method: REPLY",
|
||||
} {
|
||||
if !contains(out, want) {
|
||||
t.Errorf("output missing %q:\n%s", want, out)
|
||||
}
|
||||
}
|
||||
if contains(out, "BEGIN:VCALENDAR") {
|
||||
t.Errorf("raw ICS leaked into markdown:\n%s", out)
|
||||
}
|
||||
}
|
||||
|
||||
func TestICSToMarkdownFallback(t *testing.T) {
|
||||
out := ICSToMarkdown([]byte("not a calendar"))
|
||||
if !contains(out, "not a calendar") {
|
||||
t.Fatalf("expected raw fallback, got %q", out)
|
||||
}
|
||||
}
|
||||
|
||||
func TestICSToMarkdownNormalizesLatin1(t *testing.T) {
|
||||
// Real-world ICS from a rental portal: summary in UTF-8, location Latin-1
|
||||
// ("N\xfcrnberg"). The markdown output must be valid UTF-8 everywhere.
|
||||
raw := "BEGIN:VCALENDAR\r\nVERSION:2.0\r\nBEGIN:VEVENT\r\n" +
|
||||
"SUMMARY:Mietwagen-Buchung: N\xc3\xbcrnberg\r\n" +
|
||||
"LOCATION:N\xfcrnberg\r\nDTSTART:20200101T090000Z\r\nDTEND:20200101T180000Z\r\n" +
|
||||
"END:VEVENT\r\nEND:VCALENDAR\r\n"
|
||||
out := ICSToMarkdown([]byte(raw))
|
||||
if !utf8.ValidString(out) {
|
||||
t.Fatalf("output is not valid UTF-8:\n%q", out)
|
||||
}
|
||||
if !contains(out, "Nürnberg") {
|
||||
t.Errorf("expected Nürnberg in output:\n%s", out)
|
||||
}
|
||||
if contains(out, "N\xfcrnberg") {
|
||||
t.Errorf("Latin-1 bytes leaked into output:\n%q", out)
|
||||
}
|
||||
}
|
||||
|
||||
func TestICSToMarkdownAllDay(t *testing.T) {
|
||||
ics := `BEGIN:VCALENDAR
|
||||
VERSION:2.0
|
||||
BEGIN:VEVENT
|
||||
UID:y@google.com
|
||||
SUMMARY:All day thing
|
||||
DTSTART;VALUE=DATE:20260815
|
||||
DTEND;VALUE=DATE:20260816
|
||||
END:VEVENT
|
||||
END:VCALENDAR`
|
||||
out := ICSToMarkdown([]byte(ics))
|
||||
if !contains(out, "All day thing") || !contains(out, "2026-08-15") {
|
||||
t.Errorf("all-day event not parsed:\n%s", out)
|
||||
}
|
||||
}
|
||||
|
||||
func TestCollectParts(t *testing.T) {
|
||||
p := gmailPart{
|
||||
MimeType: "multipart/mixed",
|
||||
Parts: []gmailPart{
|
||||
{MimeType: "multipart/alternative", Parts: []gmailPart{
|
||||
{MimeType: "text/plain", Body: gmailBody{Data: b64("plain text")}},
|
||||
{MimeType: "text/html", Body: gmailBody{Data: b64("<p>html</p>")}},
|
||||
}},
|
||||
{PartID: "2", MimeType: "application/pdf", Filename: "invoice.pdf", Body: gmailBody{Size: 100}},
|
||||
},
|
||||
}
|
||||
text, html, atts := collectParts(p, "root", "abc123", 0)
|
||||
if text != "plain text" {
|
||||
t.Errorf("text = %q", text)
|
||||
}
|
||||
if html != "<p>html</p>" {
|
||||
t.Errorf("html = %q", html)
|
||||
}
|
||||
if len(atts) != 1 || atts[0].FileName != "invoice.pdf" || atts[0].FileID != "abc123:2" {
|
||||
t.Errorf("atts = %+v", atts)
|
||||
}
|
||||
}
|
||||
|
||||
type fakeGmailAPI struct {
|
||||
lastQ string
|
||||
lastLimit int
|
||||
ids []string
|
||||
}
|
||||
|
||||
func (f *fakeGmailAPI) ListIDs(_ context.Context, q string, maxIDs int, _ string) ([]string, string, error) {
|
||||
f.lastQ = q
|
||||
f.lastLimit = maxIDs
|
||||
return f.ids, "", nil
|
||||
}
|
||||
func (f *fakeGmailAPI) GetMessage(context.Context, string) (*Message, error) {
|
||||
return nil, errors.New("unused")
|
||||
}
|
||||
func (f *fakeGmailAPI) DownloadAttachment(context.Context, string, string) ([]byte, error) {
|
||||
return nil, errors.New("unused")
|
||||
}
|
||||
|
||||
func TestGmailSourcePassesQueryToListIDs(t *testing.T) {
|
||||
fake := &fakeGmailAPI{ids: []string{"m1"}}
|
||||
src := &gmailSource{c: fake, query: "from:alice@example.com"}
|
||||
ids, _, err := src.ListIDs(context.Background(), 10, "")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if fake.lastQ != "from:alice@example.com" {
|
||||
t.Fatalf("ListIDs q=%q, want from:alice@example.com", fake.lastQ)
|
||||
}
|
||||
if fake.lastLimit != 10 {
|
||||
t.Fatalf("ListIDs limit=%d, want 10", fake.lastLimit)
|
||||
}
|
||||
if len(ids) != 1 || ids[0] != "m1" {
|
||||
t.Fatalf("ids=%v", ids)
|
||||
}
|
||||
}
|
||||
|
||||
func TestGmailSourceEmptyQueryDefaultsToInbox(t *testing.T) {
|
||||
fake := &fakeGmailAPI{}
|
||||
src := &gmailSource{c: fake, query: ""}
|
||||
if _, _, err := src.ListIDs(context.Background(), 5, ""); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if fake.lastQ != "in:inbox" {
|
||||
t.Fatalf("empty query q=%q, want in:inbox", fake.lastQ)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseCLIGmailQuery(t *testing.T) {
|
||||
cli, code, err := ParseCLI([]string{
|
||||
"--source", "gmail",
|
||||
"--query", "from:letrado@example.com",
|
||||
"--out", t.TempDir(),
|
||||
"--dry-run",
|
||||
})
|
||||
if err != nil || code != 0 {
|
||||
t.Fatalf("ParseCLI: code=%d err=%v", code, err)
|
||||
}
|
||||
if cli.Sync.Query != "from:letrado@example.com" {
|
||||
t.Fatalf("query=%q", cli.Sync.Query)
|
||||
}
|
||||
if cli.Sync.Gmail == nil {
|
||||
t.Fatal("gmail source not configured")
|
||||
}
|
||||
}
|
||||
|
||||
func b64(s string) string {
|
||||
return base64.URLEncoding.EncodeToString([]byte(s))
|
||||
}
|
||||
|
||||
func contains(s, sub string) bool {
|
||||
return len(s) >= len(sub) && (s == sub || len(sub) == 0 ||
|
||||
indexOf(s, sub) >= 0)
|
||||
}
|
||||
|
||||
func indexOf(s, sub string) int {
|
||||
for i := 0; i+len(sub) <= len(s); i++ {
|
||||
if s[i:i+len(sub)] == sub {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
@@ -0,0 +1,46 @@
|
||||
// Package sync downloads OnlyOffice and Gmail messages to var/mail/ as raw
|
||||
// JSON + attachment files, then hands off to bin/mail/import --from-raw for
|
||||
// markdown conversion.
|
||||
//
|
||||
// On-disk schema (per message):
|
||||
//
|
||||
// var/mail/<folder>/<id>/message.json # Message (this package)
|
||||
// var/mail/<folder>/<id>/attachments/ # raw attachment bytes (storedName)
|
||||
//
|
||||
// The Message JSON is the contract shared with the Python converter. Fields
|
||||
// deliberately mirror what bin/mail/import already reads from the OnlyOffice
|
||||
// API, so conversion is source-agnostic.
|
||||
package sync
|
||||
|
||||
import (
|
||||
"time"
|
||||
)
|
||||
|
||||
// Attachment describes one attachment of a Message. FileID/FileName/StoredName
|
||||
// mirror OnlyOffice; Gmail fills them from its own ids. StoredName is always
|
||||
// unique (hash/attachment id) so raw files never collide.
|
||||
type Attachment struct {
|
||||
FileID string `json:"fileId,omitempty"`
|
||||
FileName string `json:"fileName"`
|
||||
StoredName string `json:"storedName"`
|
||||
Size int64 `json:"size,omitempty"`
|
||||
ContentType string `json:"contentType,omitempty"`
|
||||
}
|
||||
|
||||
// Message is the normalized record written to var/mail/<folder>/<id>/message.json.
|
||||
type Message struct {
|
||||
Source string `json:"source"` // "onlyoffice" | "gmail"
|
||||
ID string `json:"id"`
|
||||
Folder string `json:"folder"`
|
||||
Subject string `json:"subject,omitempty"`
|
||||
From string `json:"from,omitempty"`
|
||||
To string `json:"to,omitempty"`
|
||||
CC string `json:"cc,omitempty"`
|
||||
BCC string `json:"bcc,omitempty"`
|
||||
ReceivedAt time.Time `json:"receivedAt,omitempty"`
|
||||
HTMLBody string `json:"htmlBody,omitempty"`
|
||||
TextBody string `json:"textBody,omitempty"`
|
||||
HasAttachments bool `json:"hasAttachments,omitempty"`
|
||||
Attachments []Attachment `json:"attachments,omitempty"`
|
||||
MimeMessageID string `json:"mimeMessageId,omitempty"`
|
||||
}
|
||||
@@ -0,0 +1,3 @@
|
||||
// Commands in this directory are shebang mains (import.go), tagged so
|
||||
// `go build ./bin/markdown` does not see two mains.
|
||||
package main
|
||||
Executable
+100
@@ -0,0 +1,100 @@
|
||||
//usr/bin/env go run "$0" "$@"; exit
|
||||
//
|
||||
// bin/markdown/import.go - split markdown into leafs (H2 boundaries).
|
||||
//
|
||||
// ./bin/markdown/import.go [dir]
|
||||
// ./bin/markdown/import.go --files a.md,b.md --json
|
||||
//
|
||||
// Conversion only. Brain write is bin/brain/index.go.
|
||||
// Python bin/md/import remains as a fallback.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"github.com/eSlider/2dph/internal/mdleaves"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(run(os.Args[1:]))
|
||||
}
|
||||
|
||||
func run(args []string) int {
|
||||
jsonOut := false
|
||||
files := ""
|
||||
root := "."
|
||||
for i := 0; i < len(args); i++ {
|
||||
a := args[i]
|
||||
switch {
|
||||
case a == "--json":
|
||||
jsonOut = true
|
||||
case a == "--files" && i+1 < len(args):
|
||||
i++
|
||||
files = args[i]
|
||||
case strings.HasPrefix(a, "--files="):
|
||||
files = strings.TrimPrefix(a, "--files=")
|
||||
case a == "-h" || a == "--help":
|
||||
fmt.Fprintln(os.Stderr, "bin/markdown/import.go [dir] [--files a.md,b.md] [--json]")
|
||||
return 0
|
||||
case strings.HasPrefix(a, "-"):
|
||||
fmt.Fprintln(os.Stderr, "unknown arg:", a)
|
||||
return 2
|
||||
default:
|
||||
root = a
|
||||
}
|
||||
}
|
||||
|
||||
var paths []string
|
||||
if files != "" {
|
||||
for _, f := range strings.Split(files, ",") {
|
||||
f = strings.TrimSpace(f)
|
||||
if f != "" {
|
||||
paths = append(paths, f)
|
||||
}
|
||||
}
|
||||
} else {
|
||||
st, err := os.Stat(root)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "md/import: no such path %s\n", root)
|
||||
return 2
|
||||
}
|
||||
if !st.IsDir() {
|
||||
paths = []string{root}
|
||||
} else {
|
||||
var err error
|
||||
paths, err = mdleaves.WalkMarkdown(root)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "md/import: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(paths) == 0 {
|
||||
fmt.Fprintln(os.Stderr, "md/import: no markdown files")
|
||||
return 1
|
||||
}
|
||||
|
||||
var all []mdleaves.Leaf
|
||||
for _, p := range paths {
|
||||
raw, err := os.ReadFile(p)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "md/import: %s: %v\n", p, err)
|
||||
continue
|
||||
}
|
||||
all = append(all, mdleaves.ToAll(string(raw), p, "")...)
|
||||
}
|
||||
if jsonOut {
|
||||
s, err := mdleaves.EncodeJSON(all)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
return 1
|
||||
}
|
||||
fmt.Print(s)
|
||||
return 0
|
||||
}
|
||||
fmt.Print(mdleaves.EncodeYAML(all))
|
||||
return 0
|
||||
}
|
||||
+3
-2
@@ -1,9 +1,10 @@
|
||||
#!/usr/bin/env python3
|
||||
import lib
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools"))
|
||||
|
||||
from mdleaves import leaves_to_json, read_markdown, to_all, walk_markdown # noqa: E402
|
||||
from yamlout import to_yaml # noqa: E402
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
// Commands in this directory are shebang mains (query.go).
|
||||
package main
|
||||
Executable
+20
@@ -0,0 +1,20 @@
|
||||
//usr/bin/env go run -tags=postgres_query "$0" "$@"; exit
|
||||
//go:build postgres_query
|
||||
//
|
||||
// bin/postgres/query.go - read-only Postgres as YAML.
|
||||
//
|
||||
// ./bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
|
||||
//
|
||||
// Profiles: $HOME/.config/brain/db-profiles.yml (credentials stay out of git).
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/cmdbin"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(cmdbin.ExecFile("bin/db/psql-yq", os.Args[1:]))
|
||||
}
|
||||
Executable
+91
@@ -0,0 +1,91 @@
|
||||
//usr/bin/env go run -tags=reasoner_bakeoff "$0" "$@"; exit
|
||||
//go:build reasoner_bakeoff
|
||||
//
|
||||
// bin/reasoner/bakeoff.go - CPU tool-call bake-off against an OpenAI-compatible URL (D18).
|
||||
//
|
||||
// REASONER_BASE_URL=http://127.0.0.1:11435/v1 REASONER_MODEL=qwen3.5:9b ./bin/reasoner/bakeoff.go
|
||||
// ./bin/reasoner/bakeoff.go --model MichelRosselli/bonsai-27b:Q1_0 --json
|
||||
//
|
||||
// Measures OpenAI tool_calls (search/get/audit) and RSS from Ollama /api/ps, not VRAM.
|
||||
// PicoClaw is not in this repo; the tool names match internal/httpapi MCP ops.
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"github.com/eSlider/2dph/internal/reasoner"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(run(os.Args[1:]))
|
||||
}
|
||||
|
||||
func run(args []string) int {
|
||||
base := os.Getenv("REASONER_BASE_URL")
|
||||
if base == "" {
|
||||
base = "http://127.0.0.1:11435/v1"
|
||||
}
|
||||
model := os.Getenv("REASONER_MODEL")
|
||||
if model == "" {
|
||||
model = reasoner.OllamaRAM
|
||||
}
|
||||
jsonOut := false
|
||||
device := "cpu"
|
||||
for i := 0; i < len(args); i++ {
|
||||
a := args[i]
|
||||
switch {
|
||||
case a == "--json":
|
||||
jsonOut = true
|
||||
case a == "--model" && i+1 < len(args):
|
||||
i++
|
||||
model = args[i]
|
||||
case strings.HasPrefix(a, "--model="):
|
||||
model = strings.TrimPrefix(a, "--model=")
|
||||
case a == "--base-url" && i+1 < len(args):
|
||||
i++
|
||||
base = args[i]
|
||||
case a == "--device" && i+1 < len(args):
|
||||
i++
|
||||
device = args[i]
|
||||
case a == "-h" || a == "--help":
|
||||
fmt.Fprintln(os.Stderr, "bin/reasoner/bakeoff.go [--model ID] [--base-url URL] [--device cpu] [--json]")
|
||||
return 0
|
||||
default:
|
||||
fmt.Fprintln(os.Stderr, "unknown arg:", a)
|
||||
return 2
|
||||
}
|
||||
}
|
||||
c := reasoner.Client{BaseURL: base, Model: model, Device: device}
|
||||
rep := reasoner.Run(c)
|
||||
raw, err := json.MarshalIndent(rep, "", " ")
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
return 1
|
||||
}
|
||||
if jsonOut {
|
||||
fmt.Println(string(raw))
|
||||
} else {
|
||||
fmt.Printf("model: %s\n", rep.Model)
|
||||
fmt.Printf("hf_id: %s\n", rep.HF)
|
||||
fmt.Printf("device: %s\n", rep.Device)
|
||||
fmt.Printf("tool_call: %d/%d\n", rep.ToolCallOK, rep.ToolCallN)
|
||||
fmt.Printf("xml_leak: %d\n", rep.XMLLeak)
|
||||
fmt.Printf("rss_mb: %d\n", rep.RSSMB)
|
||||
fmt.Printf("vram_mb: %d\n", rep.VRAMMB)
|
||||
for _, p := range rep.Prompts {
|
||||
status := "fail"
|
||||
if p.OK {
|
||||
status = "ok"
|
||||
}
|
||||
fmt.Printf(" %s: %s wanted=%s got=%s xml=%v %dms %s\n", p.WantedTool, status, p.WantedTool, p.ToolName, p.XMLLeak, p.LatencyMS, p.Err)
|
||||
}
|
||||
}
|
||||
if rep.ToolCallN == 0 {
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
@@ -0,0 +1,2 @@
|
||||
// Commands in this directory are shebang mains (bakeoff.go).
|
||||
package main
|
||||
Executable
+22
@@ -0,0 +1,22 @@
|
||||
//usr/bin/env go run -tags=brain_serve "$0" "$@"; exit
|
||||
//go:build brain_serve
|
||||
//
|
||||
// bin/serve.go — deprecated; use bin/brain/serve.go.
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
|
||||
"github.com/eSlider/2dph/internal/httpapi"
|
||||
)
|
||||
|
||||
func main() {
|
||||
fmt.Fprintln(os.Stderr, "bin/serve.go is deprecated; use bin/brain/serve.go")
|
||||
if os.Getenv("KB_ROOT") == "" {
|
||||
if wd, err := os.Getwd(); err == nil {
|
||||
os.Setenv("KB_ROOT", wd)
|
||||
}
|
||||
}
|
||||
httpapi.Run(nil)
|
||||
}
|
||||
@@ -0,0 +1,29 @@
|
||||
"""crmfacts - pure helpers for bin/facts/crm (association proofing).
|
||||
|
||||
Shared with tools/ unit tests so the corpus-org parser is covered in CI.
|
||||
"""
|
||||
|
||||
import re
|
||||
|
||||
|
||||
def corpus_orgs(raw: str) -> dict[str, dict]:
|
||||
"""Parse the orgs block of the CV knowledge-mesh YAML into id -> fields.
|
||||
|
||||
Fields kept: label, kind, period, website. Stops at the first sibling
|
||||
top-level key (clients, timeline, ...).
|
||||
"""
|
||||
m = re.search(r"^orgs:\n(.*?)\n^(?:clients|timeline|tech_weights|nodes|edges):", raw, re.S | re.M)
|
||||
if not m:
|
||||
return {}
|
||||
orgs: dict[str, dict] = {}
|
||||
cur = None
|
||||
for line in m.group(1).splitlines():
|
||||
lm = re.match(r"^\s*- id:\s*(\S+)", line)
|
||||
if lm:
|
||||
cur = lm.group(1)
|
||||
orgs[cur] = {}
|
||||
continue
|
||||
fm = re.match(r"^\s+(\w+):\s*(.*)$", line)
|
||||
if fm and cur and fm.group(1) in ("label", "kind", "period", "website"):
|
||||
orgs[cur][fm.group(1)] = fm.group(2).strip()
|
||||
return orgs
|
||||
@@ -0,0 +1,60 @@
|
||||
"""gitimport - Ladybug graph writes for Commit/File/Person (no git binary).
|
||||
|
||||
Commit records come from bin/git/import.go (go-git). This module only MERGEs
|
||||
the version graph File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
|
||||
|
||||
@dataclass
|
||||
class Commit:
|
||||
sha: str
|
||||
author: str
|
||||
email: str
|
||||
date: str
|
||||
subject: str
|
||||
files: list[str] = field(default_factory=list)
|
||||
|
||||
|
||||
GIT_SCHEMA = (
|
||||
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
|
||||
"author STRING, email STRING, date STRING, PRIMARY KEY(id))",
|
||||
"CREATE NODE TABLE IF NOT EXISTS Person (id STRING, name STRING, email STRING, PRIMARY KEY(id))",
|
||||
"CREATE REL TABLE IF NOT EXISTS HAS_VERSION (FROM File TO Commit)",
|
||||
"CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)",
|
||||
)
|
||||
|
||||
|
||||
def ensure_git_schema(conn) -> None:
|
||||
for stmt in GIT_SCHEMA:
|
||||
conn.execute(stmt)
|
||||
|
||||
|
||||
def index_commits(conn, commits: list[Commit], repo: str) -> int:
|
||||
"""Write Commit/File/Person nodes + edges, one per commit (idempotent by sha)."""
|
||||
ensure_git_schema(conn)
|
||||
for c in commits:
|
||||
conn.execute(
|
||||
"MERGE (c:Commit {id:$sha}) SET c.repo=$repo, c.subject=$subject, "
|
||||
"c.author=$author, c.email=$email, c.date=$date",
|
||||
parameters={"sha": c.sha, "repo": repo, "subject": c.subject,
|
||||
"author": c.author, "email": c.email, "date": c.date},
|
||||
)
|
||||
conn.execute(
|
||||
"MERGE (p:Person {id:$email}) SET p.name=$name, p.email=$email",
|
||||
parameters={"email": c.email, "name": c.author},
|
||||
)
|
||||
conn.execute("MATCH (c:Commit {id:$sha}), (p:Person {id:$email}) "
|
||||
"MERGE (c)-[:AUTHORED]->(p)",
|
||||
parameters={"sha": c.sha, "email": c.email})
|
||||
for path in c.files:
|
||||
conn.execute(
|
||||
"MERGE (f:File {id:$fid}) SET f.path=$path, f.repo=$repo",
|
||||
parameters={"fid": f"{repo}:{path}", "path": path, "repo": repo},
|
||||
)
|
||||
conn.execute("MATCH (f:File {id:$fid}), (c:Commit {id:$sha}) "
|
||||
"MERGE (f)-[:HAS_VERSION]->(c)",
|
||||
parameters={"fid": f"{repo}:{path}", "sha": c.sha})
|
||||
return len(commits)
|
||||
@@ -22,7 +22,17 @@ ROOT_FACTS = "facts"
|
||||
ROOT_INFO = "info"
|
||||
CONF_CONFIRMED = "confirmed"
|
||||
|
||||
VAR = Path(__file__).resolve().parents[1] / "var"
|
||||
def _repo_root() -> Path:
|
||||
p = Path(__file__).resolve().parent
|
||||
while True:
|
||||
if (p / "var").is_dir() or (p / ".git").is_dir() or (p / "pyproject.toml").is_file():
|
||||
return p
|
||||
if p.parent == p:
|
||||
return Path(__file__).resolve().parents[2]
|
||||
p = p.parent
|
||||
|
||||
|
||||
VAR = _repo_root() / "var"
|
||||
DB_PATH = VAR / "kb.lbug"
|
||||
|
||||
|
||||
@@ -68,6 +78,19 @@ def init_schema(conn: ladybug.Connection) -> None:
|
||||
conn.execute(
|
||||
"CREATE REL TABLE IF NOT EXISTS RUNS_ON (FROM Leaf TO Host)"
|
||||
)
|
||||
conn.execute(
|
||||
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
|
||||
"author STRING, email STRING, date STRING, PRIMARY KEY(id))"
|
||||
)
|
||||
conn.execute(
|
||||
"CREATE NODE TABLE IF NOT EXISTS Person (id STRING, name STRING, email STRING, PRIMARY KEY(id))"
|
||||
)
|
||||
conn.execute(
|
||||
"CREATE REL TABLE IF NOT EXISTS HAS_VERSION (FROM File TO Commit)"
|
||||
)
|
||||
conn.execute(
|
||||
"CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
|
||||
)
|
||||
|
||||
|
||||
def leaf_id(text: str, source: str) -> str:
|
||||
@@ -95,18 +118,77 @@ def upsert_leaf(conn: ladybug.Connection, *, text: str, root: str, confidence: s
|
||||
return lid
|
||||
|
||||
|
||||
def leaf_index_names(conn: ladybug.Connection) -> set[str]:
|
||||
"""Return index names on the Leaf table (e.g. {'id', 'Leaf_vec', '_PK'})."""
|
||||
rows = conn.execute("CALL SHOW_INDEXES() RETURN *").get_all()
|
||||
return {row[1] for row in rows if row[0] == "Leaf"}
|
||||
|
||||
|
||||
def create_fts_and_vector(conn: ladybug.Connection, force: bool = False) -> None:
|
||||
if force:
|
||||
conn.execute("DROP INDEX IF EXISTS Leaf.Leaf_fts")
|
||||
conn.execute("DROP INDEX IF EXISTS Leaf.Leaf_vec")
|
||||
try:
|
||||
conn.execute("CALL CREATE_FTS_INDEX('Leaf', 'id', ['text'])")
|
||||
except Exception:
|
||||
pass
|
||||
try:
|
||||
conn.execute("CALL CREATE_VECTOR_INDEX('Leaf', 'Leaf_vec', 'embedding', metric := 'cosine')")
|
||||
except Exception:
|
||||
pass
|
||||
"""Create FTS (BM25) + HNSW vector indexes if missing.
|
||||
|
||||
Never DROP INDEX for FTS/VECTOR. Ladybug 0.19 leaves ghost catalog
|
||||
entries after DROP (`_0_Leaf_vec_UPPER`, `0_id_docs`), so a later
|
||||
CREATE fails with "already exists in catalog" while SHOW_INDEXES
|
||||
still omits the index. Swallowing that error made HNSW look "OK"
|
||||
until the first QUERY_VECTOR_INDEX.
|
||||
|
||||
`force=True` is accepted for API compatibility but does **not** drop.
|
||||
Fresh indexes require deleting `var/kb.lbug` and rebuilding
|
||||
(`bin/brain/index.go --rebuild`).
|
||||
"""
|
||||
del force # API compat; DROP is unsafe — see docstring
|
||||
names = leaf_index_names(conn)
|
||||
if "id" not in names:
|
||||
try:
|
||||
conn.execute("CALL CREATE_FTS_INDEX('Leaf', 'id', ['text'])")
|
||||
except Exception as e:
|
||||
raise RuntimeError(
|
||||
"CREATE_FTS_INDEX failed (often ghost catalog after DROP INDEX). "
|
||||
"Delete var/kb.lbug and run bin/brain/index.go --rebuild. "
|
||||
f"Cause: {e}"
|
||||
) from e
|
||||
if "Leaf_vec" not in names:
|
||||
try:
|
||||
conn.execute(
|
||||
"CALL CREATE_VECTOR_INDEX('Leaf', 'Leaf_vec', 'embedding', "
|
||||
"metric := 'cosine')"
|
||||
)
|
||||
except Exception as e:
|
||||
raise RuntimeError(
|
||||
"CREATE_VECTOR_INDEX failed (often ghost catalog after DROP INDEX "
|
||||
"Leaf.Leaf_vec → `_0_Leaf_vec_UPPER already exists in catalog`). "
|
||||
"Delete var/kb.lbug and run bin/brain/index.go --rebuild. "
|
||||
f"Cause: {e}"
|
||||
) from e
|
||||
names = leaf_index_names(conn)
|
||||
missing = {"id", "Leaf_vec"} - names
|
||||
if missing:
|
||||
raise RuntimeError(
|
||||
f"Leaf indexes incomplete after create: missing {sorted(missing)}; "
|
||||
f"have {sorted(names)}. Delete var/kb.lbug and --rebuild."
|
||||
)
|
||||
|
||||
|
||||
def ensure_indexes(conn: ladybug.Connection) -> None:
|
||||
"""Idempotent: create FTS + HNSW only when missing. Safe after upserts."""
|
||||
create_fts_and_vector(conn, force=False)
|
||||
|
||||
|
||||
def drop_indexes(conn: ladybug.Connection) -> None:
|
||||
"""No-op. Kept for callers; DROP INDEX is fatal on Ladybug 0.19.
|
||||
|
||||
Historical note claimed "drop before bulk MERGE". Measured on 0.19:
|
||||
- DROP FTS/VECTOR leaves ghost catalog → CREATE fails permanently until
|
||||
`var/kb.lbug` is deleted.
|
||||
- MERGE/upsert while **FTS** exists can corrupt FTS
|
||||
("document for node offset N is missing during delete").
|
||||
- Upsert while **HNSW** exists stays queryable.
|
||||
|
||||
Bulk rebuilders must delete `var/kb.lbug`, write all leafs (info+facts)
|
||||
with no indexes, then `ensure_indexes()` once.
|
||||
"""
|
||||
return
|
||||
|
||||
|
||||
def query_fts(conn: ladybug.Connection, text: str, limit: int = 10) -> list[dict]:
|
||||
@@ -155,6 +237,6 @@ def stats(conn: ladybug.Connection) -> dict:
|
||||
|
||||
def open_readonly() -> tuple[ladybug.Database, ladybug.Connection]:
|
||||
if not DB_PATH.exists():
|
||||
raise FileNotFoundError(f"{DB_PATH} missing - run bin/kb/index first")
|
||||
raise FileNotFoundError(f"{DB_PATH} missing - run bin/brain/index.go --rebuild first")
|
||||
db, conn = connect(read_only=True)
|
||||
return db, conn
|
||||
@@ -0,0 +1,148 @@
|
||||
"""mailconv - pure helpers for bin/mail/import (mail -> markdown + attachments).
|
||||
|
||||
Shared with unit tests in bin/tools/test_mailconv.py. No network, no OnlyOffice
|
||||
dependencies here: everything is `str -> str` or `Path -> str` so the tests run
|
||||
offline against fixtures.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import html
|
||||
import re
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
# Body part / attachment file suffixes we know how to turn into markdown text.
|
||||
TEXT_SUFFIXES = {".md", ".markdown", ".txt", ".csv", ".json", ".xml", ".yaml", ".yml", ".log", ".tsv",
|
||||
".ics", ".ical", ".vcf", ".eml"}
|
||||
OFFICE_SUFFIXES = {".docx", ".pptx", ".xlsx", ".html", ".htm", ".epub", ".eml", ".msg"}
|
||||
PDF_SUFFIXES = {".pdf"}
|
||||
IMAGE_SUFFIXES = {".png", ".jpg", ".jpeg", ".gif", ".bmp", ".tiff", ".tif", ".webp"}
|
||||
ARCHIVE_SUFFIXES = {".zip"}
|
||||
# Legacy binary Office (doc/xls/ppt) — markitdown/docling skip them; we try
|
||||
# pandoc first, else leave a stub.
|
||||
LEGACY_OFFICE_SUFFIXES = {".doc", ".xls", ".ppt"}
|
||||
|
||||
CONVERTIBLE_SUFFIXES = (
|
||||
TEXT_SUFFIXES | OFFICE_SUFFIXES | PDF_SUFFIXES | IMAGE_SUFFIXES | ARCHIVE_SUFFIXES | LEGACY_OFFICE_SUFFIXES
|
||||
)
|
||||
|
||||
|
||||
def clean_email_address(raw: str) -> str:
|
||||
"""Extract the bare email from '"Name" <a@b.c>' and strip control chars."""
|
||||
m = re.search(r"<([^<>@\s]+@[^<>@\s]+)>", raw)
|
||||
return (m.group(1) if m else raw).strip()
|
||||
|
||||
|
||||
def subject_to_filename(subject: str, max_len: int = 80) -> str:
|
||||
"""Turn a mail subject into a filesystem-safe slug (keep first token readable)."""
|
||||
s = re.sub(r"[^\w\-. ]+", "", subject).strip()
|
||||
s = re.sub(r"\s+", "_", s)
|
||||
s = s.strip("._")
|
||||
if not s:
|
||||
s = "untitled"
|
||||
return s[:max_len] or "untitled"
|
||||
|
||||
|
||||
def strip_html(html_text: str) -> str:
|
||||
"""Naive HTML -> plain text fallback (used only if markitdown is missing)."""
|
||||
import re as _re
|
||||
text = _re.sub(r"(?is)<(script|style)[^>]*>.*?</\1>", "", html_text)
|
||||
text = _re.sub(r"(?s)<br\s*/?>", "\n", text)
|
||||
text = _re.sub(r"(?s)</p>", "\n\n", text)
|
||||
text = _re.sub(r"(?s)<[^>]+>", "", text)
|
||||
return html.unescape(text).strip()
|
||||
|
||||
|
||||
def _unwrap_tables(html_text: str) -> str:
|
||||
"""Unwrap mail HTML tables into pipe-joined text lines.
|
||||
|
||||
Outlook/Stripe-style emails wrap content in nested spacer/frame tables that
|
||||
markitdown renders as hundreds of `--- |` cells and duplicated blocks.
|
||||
Every <table> becomes plain "cell1 | cell2" lines (key-value pairs survive),
|
||||
so only headings/paragraphs/links reach markitdown and no table noise is left.
|
||||
"""
|
||||
try:
|
||||
from bs4 import BeautifulSoup
|
||||
except Exception:
|
||||
return html_text
|
||||
soup = BeautifulSoup(html_text, "html.parser")
|
||||
for table in reversed(soup.find_all("table")):
|
||||
lines: list[str] = []
|
||||
for row in table.find_all("tr"):
|
||||
cells = [c.get_text(" ", strip=True) for c in row.find_all(["td", "th"])]
|
||||
line = " | ".join(x for x in cells if x)
|
||||
if line:
|
||||
lines.append(line)
|
||||
if lines:
|
||||
table.replace_with(BeautifulSoup("\n".join(lines), "html.parser"))
|
||||
else:
|
||||
table.decompose()
|
||||
return str(soup)
|
||||
|
||||
|
||||
def html_to_markdown(html_text: str) -> str:
|
||||
"""Convert a mail HTML body to markdown using markitdown when available."""
|
||||
html_text = _unwrap_tables(html_text)
|
||||
try:
|
||||
from markitdown import MarkItDown
|
||||
import io
|
||||
md = MarkItDown()
|
||||
result = md.convert_stream(io.BytesIO(html_text.encode("utf-8", errors="replace")),
|
||||
file_extension=".html")
|
||||
text = result.text_content.strip()
|
||||
if text:
|
||||
return normalize_markdown(text)
|
||||
except Exception:
|
||||
pass
|
||||
return normalize_markdown(strip_html(html_text))
|
||||
|
||||
|
||||
def normalize_markdown(text: str) -> str:
|
||||
"""Collapse the pdfminer/markitdown NUL artifacts and stray control chars."""
|
||||
# NUL bytes that pdfminer inserts between digits/letters.
|
||||
text = text.replace("\x00", "")
|
||||
# Email spacer noise: zero-width chars, soft hyphens, figure spaces,
|
||||
# combining grapheme joiner, BOM.
|
||||
for ch in ("\ufeff", "\u200b", "\u034f", "\u00ad", "\u2007", "\u2008", "\u200a", "\u2002"):
|
||||
text = text.replace(ch, "")
|
||||
text = re.sub(r"[ \t]{2,}", " ", text)
|
||||
# Trim trailing whitespace per line so space-only spacer rows collapse.
|
||||
text = "\n".join(l.rstrip() for l in text.split("\n"))
|
||||
# Collapse 3+ blank lines to two.
|
||||
text = re.sub(r"\n{3,}", "\n\n", text)
|
||||
# Remove weird trailing control chars.
|
||||
text = "".join(ch for ch in text if ch >= " " or ch in "\n\t")
|
||||
return text.strip()
|
||||
|
||||
|
||||
def split_zip_members(zip_path: Path) -> list[str]:
|
||||
"""Return safe member names of a zip archive (skips dir entries)."""
|
||||
try:
|
||||
with zipfile.ZipFile(zip_path) as zf:
|
||||
return [m for m in zf.namelist() if not m.endswith("/")]
|
||||
except zipfile.BadZipFile:
|
||||
return []
|
||||
|
||||
|
||||
def zip_extract_safe(zip_path: Path, dest: Path) -> list[Path]:
|
||||
"""Extract a zip into dest guarding against path traversal; returns files."""
|
||||
out: list[Path] = []
|
||||
try:
|
||||
with zipfile.ZipFile(zip_path) as zf:
|
||||
for member in zf.infolist():
|
||||
if member.is_dir():
|
||||
continue
|
||||
target = (dest / member.filename).resolve()
|
||||
if not target.is_relative_to(dest.resolve()):
|
||||
continue
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
with zf.open(member) as src, open(target, "wb") as dst:
|
||||
dst.write(src.read())
|
||||
out.append(target)
|
||||
except zipfile.BadZipFile:
|
||||
return []
|
||||
return out
|
||||
|
||||
|
||||
def is_convertible(suffix: str) -> bool:
|
||||
return suffix.lower() in CONVERTIBLE_SUFFIXES
|
||||
@@ -0,0 +1,37 @@
|
||||
"""Mail markdown under var/mail → info leafs. Conversion stays off the brain DB."""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from mdleaves import read_markdown, to_all
|
||||
|
||||
|
||||
def msg_date(md: Path) -> str:
|
||||
j = md.parent / "message.json"
|
||||
try:
|
||||
d = json.loads(j.read_text(encoding="utf-8"))
|
||||
return (d.get("receivedDate") or d.get("receivedAt") or "")[:10]
|
||||
except (OSError, json.JSONDecodeError, TypeError):
|
||||
return ""
|
||||
|
||||
|
||||
def from_mail_root(root: Path, limit: int = 0, since: str = "", repo: str = "ooMail") -> list[dict]:
|
||||
if not root.is_dir():
|
||||
return []
|
||||
mds = sorted(root.rglob("message.md"))
|
||||
if since:
|
||||
mds = [m for m in mds if msg_date(m) >= since]
|
||||
if limit:
|
||||
mds = mds[:limit]
|
||||
leafs: list[dict] = []
|
||||
for md in mds:
|
||||
files = [md] + sorted((md.parent / "attachments").glob("*.md"))
|
||||
for f in files:
|
||||
if not f.exists():
|
||||
continue
|
||||
for lf in to_all(read_markdown(f), f, repo=repo):
|
||||
lf["source"] = f"ooMail:{md.parent.name}:{f.name}"
|
||||
lf["how"] = "mail/import"
|
||||
leafs.append(lf)
|
||||
return leafs
|
||||
@@ -0,0 +1,198 @@
|
||||
"""D14 layout: bin/{subject}/{method}.go, libs in internal/, one go.mod."""
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
|
||||
|
||||
class BinLayoutTest(unittest.TestCase):
|
||||
def test_brain_search_shebang_exists(self) -> None:
|
||||
p = ROOT / "bin" / "brain" / "search.go"
|
||||
self.assertTrue(p.is_file(), "missing bin/brain/search.go")
|
||||
first = p.read_text().splitlines()[0]
|
||||
self.assertTrue(
|
||||
first.startswith("//usr/bin/env go run"),
|
||||
f"shebang first line, got {first!r}",
|
||||
)
|
||||
|
||||
def test_no_nested_go_mod_under_bin(self) -> None:
|
||||
nested = list((ROOT / "bin").rglob("go.mod"))
|
||||
self.assertEqual(nested, [], f"nested go.mod files: {nested}")
|
||||
|
||||
def test_rank_lives_in_internal_brain(self) -> None:
|
||||
self.assertTrue(
|
||||
(ROOT / "internal" / "brain" / "rank" / "rank.go").is_file(),
|
||||
"ranking must live in internal/brain/rank (cgo-free)",
|
||||
)
|
||||
self.assertFalse(
|
||||
(ROOT / "bin" / "kbsearch").exists(),
|
||||
"bin/kbsearch nested module must be gone",
|
||||
)
|
||||
|
||||
def test_no_main_go_under_bin_brain(self) -> None:
|
||||
main = ROOT / "bin" / "brain" / "main.go"
|
||||
self.assertFalse(main.exists(), "bin/brain/main.go is not a method")
|
||||
|
||||
def test_chats_methods_are_shebangs_not_main(self) -> None:
|
||||
chats = ROOT / "bin" / "chats"
|
||||
self.assertFalse(
|
||||
(chats / "main.go").exists(),
|
||||
"bin/chats/main.go is a dispatcher, not a method",
|
||||
)
|
||||
self.assertFalse(
|
||||
(chats / "index_cmd.go").exists(),
|
||||
"chats index is a brain write hiding under the wrong subject",
|
||||
)
|
||||
for method in ("sync.go", "import.go", "facts.go", "apply.go"):
|
||||
p = chats / method
|
||||
self.assertTrue(p.is_file(), f"missing bin/chats/{method}")
|
||||
first = p.read_text().splitlines()[0]
|
||||
self.assertTrue(
|
||||
first.startswith("//usr/bin/env go run"),
|
||||
f"{method} shebang, got {first!r}",
|
||||
)
|
||||
|
||||
def test_chats_lib_lives_in_internal(self) -> None:
|
||||
self.assertTrue(
|
||||
(ROOT / "internal" / "chats" / "linkedin.go").is_file(),
|
||||
"LinkedIn parser must live in internal/chats",
|
||||
)
|
||||
self.assertFalse(
|
||||
(ROOT / "bin" / "chats" / "linkedin.go").exists(),
|
||||
"parser must not stay under bin/chats as a second main",
|
||||
)
|
||||
|
||||
def _assert_shebang(self, rel: str) -> None:
|
||||
p = ROOT / rel
|
||||
self.assertTrue(p.is_file(), f"missing {rel}")
|
||||
first = p.read_text().splitlines()[0]
|
||||
self.assertTrue(
|
||||
first.startswith("//usr/bin/env go run"),
|
||||
f"{rel} shebang, got {first!r}",
|
||||
)
|
||||
|
||||
def test_brain_methods_are_shebangs(self) -> None:
|
||||
for method in ("index.go", "get.go", "stats.go", "eval.go", "watch.go"):
|
||||
self._assert_shebang(f"bin/brain/{method}")
|
||||
|
||||
def test_brain_get_stats_eval_are_not_python_exec(self) -> None:
|
||||
for method in ("get.go", "stats.go", "eval.go"):
|
||||
text = (ROOT / "bin" / "brain" / method).read_text()
|
||||
self.assertNotIn(
|
||||
"ExecFile",
|
||||
text,
|
||||
f"bin/brain/{method} must call internal/brain, not ExecFile Python",
|
||||
)
|
||||
self.assertNotIn(
|
||||
"cmdbin",
|
||||
text,
|
||||
f"bin/brain/{method} must not import internal/cmdbin",
|
||||
)
|
||||
self.assertIn(
|
||||
"system_ladybug",
|
||||
text.splitlines()[0],
|
||||
f"bin/brain/{method} shebang must pass -tags=system_ladybug",
|
||||
)
|
||||
self.assertIn(
|
||||
"github.com/eSlider/2dph/internal/brain",
|
||||
text,
|
||||
)
|
||||
|
||||
def test_eval_control_questions_live_in_rank(self) -> None:
|
||||
rank = (ROOT / "internal" / "brain" / "rank" / "evalq.go").read_text()
|
||||
py = (ROOT / "bin" / "kb" / "eval").read_text()
|
||||
for frag in ("BM25", "DevOps", "LadybugDB"):
|
||||
self.assertIn(frag, rank)
|
||||
self.assertIn(frag, py)
|
||||
self.assertIn("0.95", rank)
|
||||
|
||||
def test_facts_methods_are_shebangs(self) -> None:
|
||||
for method in ("audit.go", "extract.go", "crm.go"):
|
||||
self._assert_shebang(f"bin/facts/{method}")
|
||||
text = (ROOT / "bin" / "facts" / method).read_text()
|
||||
self.assertIn("cmdbin.ExecFile", text)
|
||||
self.assertIn(f"bin/facts/{method.removesuffix('.go')}", text)
|
||||
|
||||
def test_mail_import_is_shebang_not_brain_write(self) -> None:
|
||||
self._assert_shebang("bin/mail/import.go")
|
||||
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
|
||||
self.assertIn(
|
||||
"bin/brain/index.go",
|
||||
index_mail,
|
||||
"index_mail must point at bin/brain/index.go",
|
||||
)
|
||||
|
||||
def test_markdown_import_is_go_not_python_exec(self) -> None:
|
||||
self._assert_shebang("bin/markdown/import.go")
|
||||
text = (ROOT / "bin" / "markdown" / "import.go").read_text()
|
||||
self.assertNotIn("ExecFile", text)
|
||||
self.assertNotIn("cmdbin", text)
|
||||
self.assertIn("internal/mdleaves", text)
|
||||
self.assertNotIn("kb.lbug", text)
|
||||
|
||||
def test_import_adapters_do_not_write_ladybug(self) -> None:
|
||||
for rel in (
|
||||
"bin/mail/import.go",
|
||||
"bin/mail/import",
|
||||
"bin/markdown/import.go",
|
||||
"bin/chats/import.go",
|
||||
"bin/git/import.go",
|
||||
):
|
||||
text = (ROOT / rel).read_text()
|
||||
self.assertNotIn("upsert_leaf", text, rel)
|
||||
self.assertNotIn("kb.lbug", text, rel)
|
||||
self.assertNotIn("var/brain.lbug", text, rel)
|
||||
index = (ROOT / "bin" / "brain" / "index.go").read_text()
|
||||
self.assertIn("bin/kb/index", index)
|
||||
|
||||
def test_postgres_query_is_shebang(self) -> None:
|
||||
self._assert_shebang("bin/postgres/query.go")
|
||||
|
||||
def test_git_import_is_gogit_shebang(self) -> None:
|
||||
self._assert_shebang("bin/git/import.go")
|
||||
py = (ROOT / "bin" / "git" / "import").read_text()
|
||||
self.assertNotIn(
|
||||
'["git"',
|
||||
py,
|
||||
"Python git/import must not subprocess the git binary",
|
||||
)
|
||||
self.assertIn("bin/git/import.go", py)
|
||||
|
||||
def test_web_search_is_shebang(self) -> None:
|
||||
self._assert_shebang("bin/web/search.go")
|
||||
py = (ROOT / "bin" / "web" / "search").read_text()
|
||||
self.assertIn("bin/web/search.go", py)
|
||||
|
||||
def test_gitimport_py_has_no_git_binary(self) -> None:
|
||||
py = (ROOT / "bin" / "tools" / "gitimport.py").read_text()
|
||||
self.assertNotIn("subprocess", py)
|
||||
self.assertNotIn("git log", py)
|
||||
|
||||
def test_gogit_is_direct_go_mod_require(self) -> None:
|
||||
text = (ROOT / "go.mod").read_text()
|
||||
first = text.split("require (")[1].split(")")[0]
|
||||
self.assertRegex(first, r"github.com/go-git/go-git/v5\s+v")
|
||||
for line in first.splitlines():
|
||||
if "go-git/go-git" in line:
|
||||
self.assertNotIn("indirect", line)
|
||||
|
||||
def test_cgo_uses_zig_not_gcc(self) -> None:
|
||||
for rel in ("bin/cgo/zig", "bin/cgo/zcc", "bin/cgo/zc++"):
|
||||
p = ROOT / rel
|
||||
self.assertTrue(p.is_file(), f"missing {rel}")
|
||||
self.assertTrue(
|
||||
os.access(p, os.X_OK),
|
||||
f"{rel} must be executable",
|
||||
)
|
||||
zig = (ROOT / "bin" / "cgo" / "zig").read_text()
|
||||
self.assertIn("zig cc", zig)
|
||||
self.assertIn("0.14.1", zig)
|
||||
zcc = (ROOT / "bin" / "cgo" / "zcc").read_text()
|
||||
self.assertIn('exec "$ZIG" cc', zcc)
|
||||
self.assertNotIn("command -v gcc", zcc)
|
||||
search = (ROOT / "bin" / "kb" / "search").read_text()
|
||||
self.assertIn("bin/cgo/zig", search)
|
||||
self.assertNotIn("command -v gcc", search)
|
||||
@@ -0,0 +1,46 @@
|
||||
import sys
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
|
||||
import crmfacts # noqa: E402
|
||||
|
||||
FM = """\
|
||||
schema: 2
|
||||
meta:
|
||||
title: x
|
||||
orgs:
|
||||
- id: produktor
|
||||
label: ProProdukt SL / produktor.io
|
||||
kind: own
|
||||
period: 2006–present
|
||||
website: https://produktor.io
|
||||
- id: dyvenia
|
||||
label: Dyvenia
|
||||
kind: employer
|
||||
period: 2023–2025
|
||||
clients:
|
||||
- name: One
|
||||
- name: Two
|
||||
timeline:
|
||||
- start: 2001
|
||||
"""
|
||||
|
||||
|
||||
class CorpusOrgsTest(unittest.TestCase):
|
||||
def test_parses_label_kind_period(self):
|
||||
orgs = crmfacts.corpus_orgs(FM)
|
||||
self.assertEqual(orgs["produktor"]["label"], "ProProdukt SL / produktor.io")
|
||||
self.assertEqual(orgs["produktor"]["kind"], "own")
|
||||
self.assertEqual(orgs["dyvenia"]["kind"], "employer")
|
||||
|
||||
def test_does_not_leak_clients_into_orgs(self):
|
||||
orgs = crmfacts.corpus_orgs(FM)
|
||||
self.assertNotIn("One", orgs)
|
||||
self.assertNotIn("Two", orgs)
|
||||
self.assertNotIn("timeline", orgs)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,74 @@
|
||||
import os
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
|
||||
import kblib # noqa: E402
|
||||
import gitimport # noqa: E402
|
||||
|
||||
COMMIT_PERSON_SCHEMA = (
|
||||
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
|
||||
"author STRING, email STRING, date STRING, PRIMARY KEY(id))"
|
||||
)
|
||||
PERSON_SCHEMA = (
|
||||
"CREATE NODE TABLE IF NOT EXISTS Person (id STRING, name STRING, email STRING, PRIMARY KEY(id))"
|
||||
)
|
||||
HAS_VERSION_SCHEMA = "CREATE REL TABLE IF NOT EXISTS HAS_VERSION (FROM File TO Commit)"
|
||||
AUTHORED_SCHEMA = "CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
|
||||
|
||||
|
||||
def sample_commit() -> gitimport.Commit:
|
||||
return gitimport.Commit(
|
||||
sha="a1b2c3d",
|
||||
author="Ada Lovelace",
|
||||
email="ada@example.com",
|
||||
date="2026-08-10T12:00:00+01:00",
|
||||
subject="feat: first commit",
|
||||
files=["README.md", "src/main.c"],
|
||||
)
|
||||
|
||||
|
||||
class GitGraphTest(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.dir = tempfile.mkdtemp()
|
||||
self.dbpath = os.path.join(self.dir, "kb.lbug")
|
||||
self.db, self.conn = kblib.connect(self.dbpath, read_only=False)
|
||||
kblib.init_schema(self.conn)
|
||||
self.conn.execute(COMMIT_PERSON_SCHEMA)
|
||||
self.conn.execute(PERSON_SCHEMA)
|
||||
self.conn.execute(HAS_VERSION_SCHEMA)
|
||||
self.conn.execute(AUTHORED_SCHEMA)
|
||||
|
||||
def tearDown(self):
|
||||
self.conn.close()
|
||||
self.db.close()
|
||||
|
||||
def test_index_commits_creates_nodes_and_edges(self):
|
||||
gitimport.index_commits(self.conn, [sample_commit()], "sample-repo")
|
||||
rp = self.conn.execute("MATCH (p:Person) RETURN p.name, p.email").get_all()
|
||||
self.assertEqual([tuple(r) for r in rp], [("Ada Lovelace", "ada@example.com")])
|
||||
rc = self.conn.execute("MATCH (c:Commit) RETURN c.id, c.repo").get_all()
|
||||
self.assertEqual(len(rc), 1)
|
||||
self.assertEqual(rc[0][1], "sample-repo")
|
||||
rf = self.conn.execute(
|
||||
"MATCH (f:File)-[:HAS_VERSION]->(c:Commit)-[:AUTHORED]->(p:Person) "
|
||||
"RETURN f.path, c.id, p.email").get_all()
|
||||
paths = sorted(r[0] for r in rf)
|
||||
self.assertEqual(paths, ["README.md", "src/main.c"])
|
||||
self.assertTrue(all(r[2] == "ada@example.com" for r in rf))
|
||||
|
||||
def test_index_commits_idempotent(self):
|
||||
cs = [sample_commit()]
|
||||
gitimport.index_commits(self.conn, cs, "sample-repo")
|
||||
gitimport.index_commits(self.conn, cs, "sample-repo")
|
||||
n = self.conn.execute("MATCH (c:Commit) RETURN count(*)").get_all()[0][0]
|
||||
self.assertEqual(n, 1)
|
||||
p = self.conn.execute("MATCH (p:Person) RETURN count(*)").get_all()[0][0]
|
||||
self.assertEqual(p, 1)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,41 @@
|
||||
"""Import adapters write files only. Index rebuild is brain/index (D14 / Gitea #7)."""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
|
||||
|
||||
class IndexAdapterTest(unittest.TestCase):
|
||||
def test_dry_run_fixture_corpus_does_not_write_lbug(self) -> None:
|
||||
tmp = Path(tempfile.mkdtemp())
|
||||
(tmp / "note.md").write_text("# Fixture\n\n## Leaf\n\nhello corpus\n", encoding="utf-8")
|
||||
lbug = tmp / "kb.lbug"
|
||||
try:
|
||||
import ladybug # noqa: F401
|
||||
except ImportError:
|
||||
venv_py = ROOT / ".venv" / "bin" / "python"
|
||||
if not venv_py.is_file():
|
||||
self.skipTest("ladybug missing")
|
||||
py = str(venv_py)
|
||||
else:
|
||||
py = sys.executable
|
||||
proc = subprocess.run(
|
||||
[py, str(ROOT / "bin" / "kb" / "index"), "--dry-run", "--json", "--corpus", str(tmp)],
|
||||
cwd=ROOT,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
env=os.environ.copy(),
|
||||
check=False,
|
||||
)
|
||||
self.assertEqual(proc.returncode, 0, proc.stderr)
|
||||
msg = json.loads(proc.stdout)
|
||||
self.assertTrue(msg.get("dry_run"))
|
||||
self.assertGreaterEqual(msg.get("corpus_total", 0), 1)
|
||||
self.assertFalse(lbug.exists(), "dry-run must not create a Ladybug file")
|
||||
@@ -35,7 +35,7 @@ class KblibTest(unittest.TestCase):
|
||||
confidence="confirmed", source="s", source_rev="r1",
|
||||
how="test", loc="/tmp", type_="reference",
|
||||
embedding=make_emb(1.0))
|
||||
kblib.create_fts_and_vector(self.conn, force=True)
|
||||
kblib.ensure_indexes(self.conn)
|
||||
hits = kblib.query_fts(self.conn, "fox", 5)
|
||||
self.assertEqual(len(hits), 1)
|
||||
self.assertEqual(hits[0]["root"], "info")
|
||||
@@ -49,12 +49,42 @@ class KblibTest(unittest.TestCase):
|
||||
confidence="confirmed", source="s", source_rev="r1",
|
||||
how="test", loc="/tmp", type_="reference",
|
||||
embedding=make_emb(0.0))
|
||||
kblib.create_fts_and_vector(self.conn, force=True)
|
||||
kblib.ensure_indexes(self.conn)
|
||||
result = kblib.hybrid_search(self.conn, make_emb(1.0), [], 5)
|
||||
self.assertTrue(result)
|
||||
self.assertIn("rrf", result[0])
|
||||
self.assertEqual(result[0]["text"], "the quick brown fox")
|
||||
|
||||
def test_upsert_keeps_hnsw_queryable(self):
|
||||
"""Upsert while HNSW exists must not kill vector search."""
|
||||
kblib.upsert_leaf(self.conn, text="seed leaf", root="info",
|
||||
confidence="confirmed", source="s", source_rev="r1",
|
||||
how="test", loc="/tmp", type_="reference",
|
||||
embedding=make_emb(0.2))
|
||||
kblib.ensure_indexes(self.conn)
|
||||
self.assertIn("Leaf_vec", kblib.leaf_index_names(self.conn))
|
||||
kblib.upsert_leaf(self.conn, text="added after index", root="facts",
|
||||
confidence="confirmed", source="a.md x b.md",
|
||||
source_rev="r1", how="test", loc="/tmp", type_="fact",
|
||||
embedding=make_emb(0.9))
|
||||
hits = kblib.query_vector(self.conn, make_emb(0.9), 5)
|
||||
self.assertTrue(hits)
|
||||
self.assertIn("Leaf_vec", kblib.leaf_index_names(self.conn))
|
||||
|
||||
def test_drop_vector_then_create_raises_clear_error(self):
|
||||
"""DROP INDEX leaves ghost catalog; create_fts_and_vector must raise."""
|
||||
kblib.upsert_leaf(self.conn, text="seed", root="info",
|
||||
confidence="confirmed", source="s", source_rev="r1",
|
||||
how="test", loc="/tmp", type_="reference",
|
||||
embedding=make_emb(0.1))
|
||||
kblib.ensure_indexes(self.conn)
|
||||
self.conn.execute("DROP INDEX IF EXISTS Leaf.Leaf_vec")
|
||||
with self.assertRaises(RuntimeError) as ctx:
|
||||
kblib.create_fts_and_vector(self.conn, force=True)
|
||||
msg = str(ctx.exception)
|
||||
self.assertIn("CREATE_VECTOR_INDEX failed", msg)
|
||||
self.assertIn("--rebuild", msg)
|
||||
|
||||
def test_stats_counts_roots(self):
|
||||
kblib.upsert_leaf(self.conn, text="a fact leaf", root="facts",
|
||||
confidence="confirmed", source="s", source_rev="r1",
|
||||
@@ -70,4 +100,4 @@ class KblibTest(unittest.TestCase):
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
unittest.main()
|
||||
@@ -0,0 +1,123 @@
|
||||
import io
|
||||
import os
|
||||
import sys
|
||||
import unittest
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, os.path.dirname(__file__))
|
||||
|
||||
from mailconv import ( # noqa: E402
|
||||
clean_email_address,
|
||||
html_to_markdown,
|
||||
is_convertible,
|
||||
normalize_markdown,
|
||||
split_zip_members,
|
||||
subject_to_filename,
|
||||
zip_extract_safe,
|
||||
)
|
||||
from mailconv import _unwrap_tables # noqa: E402
|
||||
|
||||
|
||||
class TestMailConv(unittest.TestCase):
|
||||
def test_clean_email_address(self):
|
||||
self.assertEqual(clean_email_address('"Ben Baker" <bb@teks.com>'), "bb@teks.com")
|
||||
self.assertEqual(clean_email_address("eslider@gmail.com"), "eslider@gmail.com")
|
||||
self.assertEqual(clean_email_address("<a@b.c>"), "a@b.c")
|
||||
|
||||
def test_subject_to_filename(self):
|
||||
self.assertEqual(subject_to_filename("Your receipt #2422"), "Your_receipt_2422")
|
||||
self.assertEqual(subject_to_filename("a/b\\c:d*e"), "abcde")
|
||||
self.assertEqual(subject_to_filename(" "), "untitled")
|
||||
|
||||
def test_html_to_markdown(self):
|
||||
out = html_to_markdown("<html><body><h1>Hi</h1><p>Some <b>bold</b> text.</p></body></html>")
|
||||
self.assertIn("Hi", out)
|
||||
self.assertIn("**bold**", out)
|
||||
|
||||
def test_html_strip_fallback(self):
|
||||
from mailconv import strip_html
|
||||
self.assertEqual(strip_html("<p>a</p><p>b</p>"), "a\n\nb")
|
||||
|
||||
def test_flatten_layout_tables(self):
|
||||
html = ("<table><tr>"
|
||||
+ "".join(f"<td>spacer{i}</td>" for i in range(12))
|
||||
+ "</tr></table>"
|
||||
+ "<p>real</p>"
|
||||
+ "<table><tr><td>a</td><td>b</td></tr></table>")
|
||||
out = _unwrap_tables(html)
|
||||
# tables unwrapped into pipe text; no <td> left; content preserved
|
||||
self.assertNotIn("<td>spacer0</td>", out)
|
||||
self.assertIn("spacer0 | spacer1", out)
|
||||
self.assertIn("a | b", out)
|
||||
self.assertIn("real", out)
|
||||
|
||||
def test_html_to_markdown_layout_clean(self):
|
||||
html = "<table><tr>" + "".join(f"<td>x{i}</td>" for i in range(12)) + "</tr></table><h1>Hi</h1>"
|
||||
out = html_to_markdown(html)
|
||||
self.assertIn("Hi", out)
|
||||
self.assertNotIn("| ---", out)
|
||||
|
||||
def test_normalize_markdown_removes_nul(self):
|
||||
self.assertEqual(normalize_markdown("Z0\x00A\x00Y\x00B"), "Z0AYB")
|
||||
self.assertEqual(normalize_markdown("a\n\n\n\nb"), "a\n\nb")
|
||||
|
||||
def test_normalize_strips_email_noise(self):
|
||||
noisy = "\ufeffa\u200b\u034f\u00ad\u2007\u2002 b\u200a c\u2008"
|
||||
out = normalize_markdown(noisy)
|
||||
self.assertNotIn("\u200b", out)
|
||||
self.assertNotIn("\ufeff", out)
|
||||
self.assertNotIn("\u034f", out)
|
||||
self.assertIn("a b c", out)
|
||||
|
||||
def test_split_zip_members(self):
|
||||
p = Path(self._mk_zip(["a.txt", "sub/b.txt"]))
|
||||
self.assertEqual(split_zip_members(p), ["a.txt", "sub/b.txt"])
|
||||
|
||||
def test_zip_extract_safe(self):
|
||||
zip_path = self._mk_zip(["a.txt", "dir/b.txt"])
|
||||
dest = Path(self._tmp("x"))
|
||||
files = zip_extract_safe(zip_path, dest)
|
||||
self.assertEqual(len(files), 2)
|
||||
self.assertTrue((dest / "a.txt").exists())
|
||||
self.assertTrue((dest / "dir" / "b.txt").exists())
|
||||
|
||||
def test_zip_extract_safe_blocks_traversal(self):
|
||||
# member "../evil.txt" must not escape dest
|
||||
zip_path = Path(self._tmp("evil.zip"))
|
||||
with zipfile.ZipFile(zip_path, "w") as zf:
|
||||
zf.writestr("../evil.txt", "boom")
|
||||
dest = Path(self._tmp("out"))
|
||||
files = zip_extract_safe(zip_path, dest)
|
||||
self.assertEqual(files, [])
|
||||
self.assertFalse((dest.parent / "evil.txt").exists())
|
||||
|
||||
def test_is_convertible(self):
|
||||
self.assertTrue(is_convertible(".pdf"))
|
||||
self.assertTrue(is_convertible(".zip"))
|
||||
self.assertTrue(is_convertible(".docx"))
|
||||
self.assertTrue(is_convertible(".TXT"))
|
||||
self.assertFalse(is_convertible(".exe"))
|
||||
self.assertFalse(is_convertible(".unknown"))
|
||||
|
||||
def _mk_zip(self, members):
|
||||
zpath = Path(self._tmp("arc.zip"))
|
||||
with zipfile.ZipFile(zpath, "w") as zf:
|
||||
for m in members:
|
||||
zf.writestr(m, "content")
|
||||
return str(zpath)
|
||||
|
||||
def _tmp(self, name):
|
||||
d = self.__class__._td
|
||||
p = Path(d) / name
|
||||
p.parent.mkdir(parents=True, exist_ok=True)
|
||||
return str(p)
|
||||
|
||||
@classmethod
|
||||
def setUpClass(cls):
|
||||
import tempfile
|
||||
cls._td = tempfile.mkdtemp(prefix="mailconv_test_")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,46 @@
|
||||
"""Mail markdown → leafs (no Ladybug). Brain index --with-mail uses this."""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
|
||||
import mailleafs # noqa: E402
|
||||
|
||||
|
||||
class MailLeafsTest(unittest.TestCase):
|
||||
def test_message_md_becomes_info_leaf(self) -> None:
|
||||
root = Path(tempfile.mkdtemp())
|
||||
msg = root / "inbox" / "alice-1"
|
||||
msg.mkdir(parents=True)
|
||||
(msg / "message.json").write_text(
|
||||
json.dumps({"receivedDate": "2026-01-15T10:00:00Z", "subject": "Hello"}),
|
||||
encoding="utf-8",
|
||||
)
|
||||
(msg / "message.md").write_text(
|
||||
"---\nroot: info\n---\n\n# Hello\n\nFrom Alice to Bob.\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
leafs = mailleafs.from_mail_root(root)
|
||||
self.assertEqual(len(leafs), 1)
|
||||
self.assertIn("Alice", leafs[0]["text"])
|
||||
self.assertTrue(leafs[0]["source"].startswith("ooMail:"))
|
||||
self.assertEqual(leafs[0]["how"], "mail/import")
|
||||
|
||||
def test_since_filters_by_message_json_date(self) -> None:
|
||||
root = Path(tempfile.mkdtemp())
|
||||
for name, day in (("old", "2025-01-01"), ("new", "2026-06-01")):
|
||||
d = root / "inbox" / name
|
||||
d.mkdir(parents=True)
|
||||
(d / "message.json").write_text(
|
||||
json.dumps({"receivedDate": f"{day}T00:00:00Z"}),
|
||||
encoding="utf-8",
|
||||
)
|
||||
(d / "message.md").write_text(f"# {name}\n\nbody\n", encoding="utf-8")
|
||||
leafs = mailleafs.from_mail_root(root, since="2026-01-01")
|
||||
self.assertEqual(len(leafs), 1)
|
||||
self.assertIn("new", leafs[0]["text"])
|
||||
@@ -0,0 +1,157 @@
|
||||
"""Published docs must match live commands (Gitea SoT, brain/search, no fake --hop)."""
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
|
||||
|
||||
class PublishedDocsTest(unittest.TestCase):
|
||||
def test_readme_points_issues_at_gitea(self) -> None:
|
||||
text = (ROOT / "README.md").read_text()
|
||||
self.assertIn(
|
||||
"https://git.produktor.io/eSlider/2dph/issues",
|
||||
text,
|
||||
"README must point issues at Gitea",
|
||||
)
|
||||
|
||||
def test_plan_d15_names_gitea_origin(self) -> None:
|
||||
text = (ROOT / "PLAN.md").read_text()
|
||||
self.assertIn("D15", text)
|
||||
self.assertIn("git.produktor.io/eSlider/2dph", text)
|
||||
|
||||
def test_readme_primary_search_is_brain(self) -> None:
|
||||
text = (ROOT / "README.md").read_text()
|
||||
self.assertIn(
|
||||
"bin/brain/search.go",
|
||||
text,
|
||||
"README deduction search must name bin/brain/search.go",
|
||||
)
|
||||
|
||||
def test_readme_index_is_brain_not_index_mail(self) -> None:
|
||||
text = (ROOT / "README.md").read_text()
|
||||
self.assertIn("bin/brain/index.go", text)
|
||||
self.assertNotIn(
|
||||
"bin/mail/index_mail",
|
||||
text,
|
||||
"mail index is a brain write; README must name bin/brain/index.go",
|
||||
)
|
||||
|
||||
def test_readme_git_import_is_gogit(self) -> None:
|
||||
text = (ROOT / "README.md").read_text()
|
||||
self.assertIn("bin/git/import.go", text)
|
||||
self.assertIn("go-git", text)
|
||||
self.assertIn("D19", (ROOT / "PLAN.md").read_text())
|
||||
|
||||
def test_web_search_is_go_not_ops_host(self) -> None:
|
||||
readme = (ROOT / "README.md").read_text()
|
||||
self.assertIn("bin/web/search.go", readme)
|
||||
skill = (ROOT / "skills" / "web-search" / "SKILL.md").read_text()
|
||||
self.assertIn("bin/web/search.go", skill)
|
||||
self.assertNotIn("search.ops.io", skill)
|
||||
self.assertNotIn("search.ops.io", readme)
|
||||
compose = (ROOT / "compose.yaml").read_text()
|
||||
self.assertIn("searxng", compose)
|
||||
self.assertNotIn("search.ops.io", compose)
|
||||
settings = (ROOT / "deploy" / "searxng" / "settings.yml").read_text()
|
||||
self.assertNotIn("password", settings.lower())
|
||||
self.assertIn("json", settings)
|
||||
|
||||
def test_picoclaw_compose_profile_has_mcp_example(self) -> None:
|
||||
compose = (ROOT / "compose.yaml").read_text()
|
||||
self.assertIn('profiles: ["picoclaw"]', compose)
|
||||
self.assertIn("127.0.0.1:8630", compose)
|
||||
example = (ROOT / "deploy" / "picoclaw" / "mcp.json.example").read_text()
|
||||
self.assertIn("127.0.0.1:8630/mcp", example)
|
||||
self.assertNotIn("password", example.lower())
|
||||
self.assertNotIn("token", example.lower())
|
||||
docs = (ROOT / "docs" / "picoclaw.md").read_text()
|
||||
self.assertIn("search", docs)
|
||||
self.assertIn("throttled", docs)
|
||||
|
||||
def test_readme_read_path_is_go(self) -> None:
|
||||
plan = (ROOT / "PLAN.md").read_text()
|
||||
self.assertIn("get.go", plan)
|
||||
self.assertIn("CI fallback", plan)
|
||||
design = (ROOT / "docs" / "design.md").read_text()
|
||||
self.assertIn("internal/brain/rank", design)
|
||||
self.assertIn("They do not exec Python", design)
|
||||
|
||||
def test_openapi_mcp_from_same_handlers(self) -> None:
|
||||
plan = (ROOT / "PLAN.md").read_text()
|
||||
self.assertIn("D20", plan)
|
||||
self.assertIn("/openapi.json", (ROOT / "README.md").read_text())
|
||||
self.assertIn("/mcp", (ROOT / "README.md").read_text())
|
||||
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
|
||||
self.assertIn("/mcp", skill)
|
||||
self.assertFalse((ROOT / "skills" / "db-yaml").exists())
|
||||
self.assertTrue((ROOT / "skills" / "postgres" / "SKILL.md").is_file())
|
||||
|
||||
def test_cgo_zig_and_index_profile(self) -> None:
|
||||
plan = (ROOT / "PLAN.md").read_text()
|
||||
self.assertIn("D21", plan)
|
||||
self.assertIn("zig cc", plan)
|
||||
dockerfile = (ROOT / "Dockerfile").read_text()
|
||||
self.assertIn("bin/cgo/zcc", dockerfile)
|
||||
self.assertIn("FROM debian:bookworm-slim AS api", dockerfile)
|
||||
self.assertIn("FROM python:3.12-slim AS index", dockerfile)
|
||||
api = dockerfile[dockerfile.index("FROM debian:bookworm-slim AS api") :]
|
||||
self.assertNotIn("pip install", api)
|
||||
compose = (ROOT / "compose.yaml").read_text()
|
||||
self.assertIn('profiles: ["index"]', compose)
|
||||
self.assertIn("target: api", compose)
|
||||
|
||||
def test_reasoner_docs_name_real_hf_ids_cpu_sidecar(self) -> None:
|
||||
docs = (ROOT / "docs" / "reasoner.md").read_text()
|
||||
for hf in (
|
||||
"Qwen/Qwen3.5-9B",
|
||||
"Qwen/Qwen3.6-27B",
|
||||
"prism-ml/Bonsai-27B-gguf",
|
||||
):
|
||||
self.assertIn(hf, docs)
|
||||
self.assertIn("no official qwen3.6-9b", docs.lower())
|
||||
self.assertIn("OLLAMA_NUM_GPU", docs)
|
||||
self.assertIn("rss_mb", docs)
|
||||
self.assertIn("vram_mb", docs)
|
||||
self.assertIn("3/3", docs)
|
||||
self.assertIn("Do not claim 9B is better at tools", docs)
|
||||
self.assertNotIn("Qwen/Qwen3.6-9B", docs)
|
||||
plan = (ROOT / "PLAN.md").read_text()
|
||||
self.assertIn("D18", plan)
|
||||
self.assertIn("Qwen/Qwen3.5-9B", plan)
|
||||
compose = (ROOT / "compose.yaml").read_text()
|
||||
self.assertIn('profiles: ["reasoner"]', compose)
|
||||
self.assertIn("OLLAMA_NUM_GPU", compose)
|
||||
self.assertIn("127.0.0.1:11435", compose)
|
||||
dockerfile = (ROOT / "Dockerfile").read_text()
|
||||
self.assertNotIn(".gguf", dockerfile.lower())
|
||||
self.assertNotIn(".safetensors", dockerfile.lower())
|
||||
api = dockerfile[dockerfile.index("FROM debian:bookworm-slim AS api") :]
|
||||
self.assertNotIn("COPY models", api)
|
||||
self.assertNotIn("qwen", api.lower())
|
||||
|
||||
def test_readme_search_escalates_web(self) -> None:
|
||||
text = (ROOT / "README.md").read_text()
|
||||
self.assertIn("--no-web", text)
|
||||
self.assertIn("D17", (ROOT / "PLAN.md").read_text())
|
||||
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
|
||||
self.assertIn("`web` block", skill)
|
||||
|
||||
def test_docs_do_not_claim_hop_walks(self) -> None:
|
||||
paths = [
|
||||
ROOT / "README.md",
|
||||
ROOT / "docs" / "design.md",
|
||||
ROOT / "skills" / "brain" / "SKILL.md",
|
||||
ROOT / "skills" / "diataxis-docs" / "SKILL.md",
|
||||
]
|
||||
# Command-style `--hop 1` / `--hop N` plus follow/walk = the old lie.
|
||||
# Honest "not implemented" notes must not match.
|
||||
lie = re.compile(r"--hop (?:N|1).*(?:follow|walk)", re.I | re.S)
|
||||
for path in paths:
|
||||
text = path.read_text()
|
||||
self.assertIsNone(
|
||||
lie.search(text),
|
||||
f"{path.relative_to(ROOT)} still claims --hop walks the graph",
|
||||
)
|
||||
@@ -0,0 +1,49 @@
|
||||
"""Skills must name live commands; every bin/ path in SKILL.md must exist."""
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
BIN_PATH = re.compile(r"(bin/[A-Za-z0-9_./-]+)")
|
||||
|
||||
|
||||
class SkillsTest(unittest.TestCase):
|
||||
def test_db_yaml_renamed_to_postgres(self) -> None:
|
||||
self.assertFalse(
|
||||
(ROOT / "skills" / "db-yaml").exists(),
|
||||
"skills/db-yaml must be skills/postgres",
|
||||
)
|
||||
self.assertTrue((ROOT / "skills" / "postgres" / "SKILL.md").is_file())
|
||||
text = (ROOT / "skills" / "postgres" / "SKILL.md").read_text()
|
||||
self.assertIn("bin/postgres/query.go", text)
|
||||
self.assertNotIn("search.ops.io", text)
|
||||
|
||||
def test_every_bin_path_in_skills_exists(self) -> None:
|
||||
missing: list[str] = []
|
||||
for path in (ROOT / "skills").rglob("SKILL.md"):
|
||||
text = path.read_text()
|
||||
for m in BIN_PATH.finditer(text):
|
||||
rel = m.group(1).rstrip(")`.,;")
|
||||
candidate = ROOT / rel
|
||||
if not candidate.exists():
|
||||
missing.append(f"{path.relative_to(ROOT)}: {rel}")
|
||||
self.assertEqual(missing, [], "skill bin paths must exist")
|
||||
|
||||
def test_brain_skill_lists_generated_tools(self) -> None:
|
||||
tools = (ROOT / "skills" / "brain" / "tools.md").read_text()
|
||||
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
|
||||
self.assertIn("tools.md", skill)
|
||||
for name in ("search", "get", "stats", "audit"):
|
||||
self.assertIn(f"`{name}`", tools)
|
||||
|
||||
def test_picoclaw_lists_tool_order(self) -> None:
|
||||
skill = (ROOT / "skills" / "picoclaw" / "SKILL.md").read_text()
|
||||
agents = (ROOT / "AGENTS.md").read_text()
|
||||
self.assertIn("**`search`**", skill)
|
||||
self.assertIn("**`get`**", skill)
|
||||
self.assertIn("**`audit`**", skill)
|
||||
self.assertIn("throttled", skill.lower())
|
||||
self.assertIn("not a negative finding", agents)
|
||||
self.assertIn("Fact-check every", agents)
|
||||
@@ -0,0 +1,33 @@
|
||||
"""Every bin/ path named in skills/ must exist on disk."""
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
BIN_PATH = re.compile(r"\b(bin/[A-Za-z0-9_./-]+)")
|
||||
|
||||
|
||||
class SkillsBinPathsTest(unittest.TestCase):
|
||||
def test_agent_cost_skill_is_gone(self) -> None:
|
||||
self.assertFalse(
|
||||
(ROOT / "skills" / "agent-cost").exists(),
|
||||
"skills/agent-cost documents bin/agents/cost which does not exist",
|
||||
)
|
||||
|
||||
def test_brain_skill_replaces_kb_search(self) -> None:
|
||||
self.assertTrue((ROOT / "skills" / "brain" / "SKILL.md").is_file())
|
||||
self.assertFalse((ROOT / "skills" / "kb-search").exists())
|
||||
|
||||
def test_skill_bin_paths_exist(self) -> None:
|
||||
missing: list[str] = []
|
||||
for skill in sorted((ROOT / "skills").rglob("SKILL.md")):
|
||||
text = skill.read_text()
|
||||
for match in BIN_PATH.findall(text):
|
||||
rel = match.rstrip("`'.,")
|
||||
if rel.endswith(".go") or Path(rel).suffix == "" or Path(rel).suffix in {".go", ".py"}:
|
||||
p = ROOT / rel
|
||||
if not p.exists():
|
||||
missing.append(f"{skill.relative_to(ROOT)}: {rel}")
|
||||
self.assertEqual(missing, [], "SKILL.md names bin/ paths that do not exist")
|
||||
@@ -76,7 +76,7 @@ class PhiGuard(unittest.TestCase):
|
||||
|
||||
def test_plain_technical_query_passes(self):
|
||||
self.assertIsNone(ws.phi_reason("Pflegegrad SGB XI Einstufung"))
|
||||
self.assertIsNone(ws.phi_reason("site:ticket.detective.de Toureffizienz"))
|
||||
self.assertIsNone(ws.phi_reason("site:example.com technical query"))
|
||||
|
||||
def test_long_digit_run_is_refused(self):
|
||||
self.assertIsNotNone(ws.phi_reason("Kunde 4711220385 Adresse"))
|
||||
@@ -0,0 +1,105 @@
|
||||
// Package watch polls corpus directories for changes and re-runs brain/index.
|
||||
//
|
||||
// Port of the former bin/kb-watch bash script to an importable, testable Go
|
||||
// package. Polls file mtimes (no inotify deps); cheap and reliable.
|
||||
package watch
|
||||
|
||||
import (
|
||||
"log"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Options controls the polling loop. Zero value uses defaults.
|
||||
type Options struct {
|
||||
Dirs []string
|
||||
Interval time.Duration
|
||||
// IndexCmd is the index command template. %s is replaced by the repo
|
||||
// root (from KB_ROOT). Defaults to `python3 <root>/bin/kb/index --with-mail`.
|
||||
IndexCmd string
|
||||
}
|
||||
|
||||
// Run blocks forever polling Dirs (defaults: KB_WATCH_DIRS or /corpus) every
|
||||
// Interval (default 30s) and re-indexing when files change. KB_ROOT names the
|
||||
// repo root used to locate bin/kb/index.
|
||||
func Run(args []string) {
|
||||
opts := fromEnv(args)
|
||||
root, _ := os.Getwd()
|
||||
if r := os.Getenv("KB_ROOT"); r != "" {
|
||||
root = r
|
||||
}
|
||||
log.Printf("watch: dirs=%v interval=%s root=%s", opts.Dirs, opts.Interval, root)
|
||||
var last string
|
||||
for {
|
||||
if flag := Stamp(opts.Dirs); flag != "" && flag != last {
|
||||
last = flag
|
||||
reindex(opts.IndexCmd, root)
|
||||
}
|
||||
time.Sleep(opts.Interval)
|
||||
}
|
||||
}
|
||||
|
||||
func fromEnv(args []string) Options {
|
||||
opts := Options{Interval: 30 * time.Second}
|
||||
if raw := os.Getenv("KB_WATCH_INTERVAL"); raw != "" {
|
||||
if n, err := strconv.Atoi(raw); err == nil && n > 0 {
|
||||
opts.Interval = time.Duration(n) * time.Second
|
||||
}
|
||||
}
|
||||
defDirs := "/corpus"
|
||||
if raw := os.Getenv("KB_WATCH_DIRS"); raw != "" {
|
||||
defDirs = raw
|
||||
}
|
||||
if len(args) > 0 {
|
||||
opts.Dirs = args
|
||||
} else {
|
||||
for _, d := range strings.Split(defDirs, " ") {
|
||||
if d != "" {
|
||||
opts.Dirs = append(opts.Dirs, d)
|
||||
}
|
||||
}
|
||||
}
|
||||
pys := os.Getenv("KB_PY")
|
||||
if pys == "" {
|
||||
pys = "python3"
|
||||
}
|
||||
opts.IndexCmd = pys + " <root>/bin/kb/index --with-mail"
|
||||
return opts
|
||||
}
|
||||
|
||||
// Stamp returns a rolling fingerprint (newest mtime under dirs) that changes
|
||||
// whenever any corpus file is touched. Empty when no files found.
|
||||
func Stamp(dirs []string) string {
|
||||
var newest time.Time
|
||||
for _, dir := range dirs {
|
||||
_ = filepath.WalkDir(dir, func(path string, _ os.DirEntry, err error) error {
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
if info, e := os.Stat(path); e == nil && info.ModTime().After(newest) {
|
||||
newest = info.ModTime()
|
||||
}
|
||||
return nil
|
||||
})
|
||||
}
|
||||
if newest.IsZero() {
|
||||
return ""
|
||||
}
|
||||
return strconv.FormatInt(newest.UnixNano(), 10)
|
||||
}
|
||||
|
||||
func reindex(template, root string) {
|
||||
cmd := strings.ReplaceAll(template, "<root>", root)
|
||||
parts := strings.Fields(cmd)
|
||||
c := exec.Command(parts[0], parts[1:]...)
|
||||
out, err := c.CombinedOutput()
|
||||
if err != nil {
|
||||
log.Printf("watch: index failed: %v\n%s", err, out)
|
||||
} else {
|
||||
log.Printf("watch: re-indexed")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,53 @@
|
||||
package watch
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func TestStampChangesWhenFileTouched(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
a := filepath.Join(dir, "a.md")
|
||||
if err := os.WriteFile(a, []byte("x"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
s1 := Stamp([]string{dir})
|
||||
if s1 == "" {
|
||||
t.Fatal("stamp empty for a dir with a file")
|
||||
}
|
||||
time.Sleep(10 * time.Millisecond)
|
||||
if err := os.WriteFile(a, []byte("y"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if s2 := Stamp([]string{dir}); s2 == s1 {
|
||||
t.Fatal("stamp did not change after the file was modified")
|
||||
}
|
||||
}
|
||||
|
||||
func TestStampEmptyForMissingDir(t *testing.T) {
|
||||
if s := Stamp([]string{filepath.Join(t.TempDir(), "nope")}); s != "" {
|
||||
t.Fatalf("stamp = %q, want empty for missing dir", s)
|
||||
}
|
||||
}
|
||||
|
||||
func TestFromEnvDefaults(t *testing.T) {
|
||||
t.Setenv("KB_WATCH_INTERVAL", "")
|
||||
t.Setenv("KB_WATCH_DIRS", "")
|
||||
t.Setenv("KB_PY", "")
|
||||
opts := fromEnv(nil)
|
||||
if len(opts.Dirs) == 0 || opts.Dirs[0] != "/corpus" {
|
||||
t.Fatalf("default dirs = %v, want [/corpus]", opts.Dirs)
|
||||
}
|
||||
if opts.Interval != 30*time.Second {
|
||||
t.Fatalf("default interval = %s, want 30s", opts.Interval)
|
||||
}
|
||||
if !strings.Contains(opts.IndexCmd, "kb/index") {
|
||||
t.Fatalf("default index cmd = %q, want kb/index", opts.IndexCmd)
|
||||
}
|
||||
if !strings.Contains(opts.IndexCmd, "--with-mail") {
|
||||
t.Fatalf("default index cmd must include --with-mail, got %q", opts.IndexCmd)
|
||||
}
|
||||
}
|
||||
+12
-136
@@ -1,150 +1,26 @@
|
||||
#!/usr/bin/env python3
|
||||
"""web/search - web search through the self-hosted SearXNG at search.ops.io.
|
||||
"""web/search — deprecated. Use bin/web/search.go (SearXNG, no Python client).
|
||||
|
||||
bin/web/search "LadybugDB vector search"
|
||||
bin/web/search "model2vec multilingual" --site github.com
|
||||
bin/web/search "uclancy" --category it -n 3 --json | jq -r '.results[].url'
|
||||
bin/web/search "sqlite-vec" --refresh # ignore the cached answer
|
||||
|
||||
This complements bin/kb/search: the knowledge base holds our own facts, this
|
||||
reaches the public web. Use it as the second, independent source that the
|
||||
detective method asks for.
|
||||
|
||||
Exit codes: 0 results, 2 refused as possible PII, 3 throttled (not "nothing
|
||||
found" - the instance answers 200 with an empty list when it throttles).
|
||||
bin/web/search.go QUERY [--json] [-n N] [--site HOST]
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import fcntl
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
|
||||
TOOLS = Path(__file__).resolve().parents[1].parent / "tools"
|
||||
sys.path.insert(0, str(TOOLS))
|
||||
sys.path.insert(0, str(TOOLS / "web-search"))
|
||||
|
||||
import websearch as ws # noqa: E402
|
||||
from yamlout import to_yaml # noqa: E402
|
||||
|
||||
CONFIG = Path(os.environ.get("BRAIN_SEARCH_ENV", Path.home() / ".config/brain/search.env"))
|
||||
CACHE = Path(os.environ.get("BRAIN_SEARCH_CACHE", Path.home() / ".cache/brain/web-search.sqlite"))
|
||||
LOCK = CACHE.with_suffix(".lock")
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
|
||||
|
||||
def load_config() -> dict:
|
||||
if not CONFIG.exists():
|
||||
sys.exit(f"no credentials at {CONFIG} (mode 600, BRAIN_SEARCH_URL/USER/PASS)")
|
||||
conf = {}
|
||||
for line in CONFIG.read_text().splitlines():
|
||||
line = line.strip()
|
||||
if not line or line.startswith("#") or "=" not in line:
|
||||
continue
|
||||
key, _, value = line.partition("=")
|
||||
conf[key.strip()] = value.strip().strip("\"'")
|
||||
missing = {"BRAIN_SEARCH_URL", "BRAIN_SEARCH_USER", "BRAIN_SEARCH_PASS"} - conf.keys()
|
||||
if missing:
|
||||
sys.exit(f"{CONFIG} is missing {', '.join(sorted(missing))}")
|
||||
return conf
|
||||
|
||||
|
||||
def fetch(conf: dict, query: str, params: dict, timeout: int) -> dict:
|
||||
args = {"q": query, "format": "json", **params}
|
||||
url = f"{conf['BRAIN_SEARCH_URL'].rstrip('/')}/search?{urllib.parse.urlencode(args)}"
|
||||
request = urllib.request.Request(url)
|
||||
token = f"{conf['BRAIN_SEARCH_USER']}:{conf['BRAIN_SEARCH_PASS']}".encode()
|
||||
import base64
|
||||
request.add_header("Authorization", "Basic " + base64.b64encode(token).decode())
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
return json.loads(response.read().decode())
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description="web search via SearXNG")
|
||||
parser.add_argument("query")
|
||||
parser.add_argument("-n", "--limit", type=int, default=ws.DEFAULT_LIMIT)
|
||||
parser.add_argument("--site", help="restrict to one domain")
|
||||
parser.add_argument("--lang", help="language code, e.g. de")
|
||||
parser.add_argument("--fresh", choices=["day", "week", "month", "year"],
|
||||
help="time range")
|
||||
parser.add_argument("--category", help="SearXNG category, e.g. it, science, news")
|
||||
parser.add_argument("--engines", help="comma separated engine list")
|
||||
parser.add_argument("--json", action="store_true")
|
||||
parser.add_argument("--refresh", action="store_true", help="bypass the cache")
|
||||
parser.add_argument("--ttl", type=float, default=ws.CACHE_TTL)
|
||||
parser.add_argument("--timeout", type=int, default=25)
|
||||
parser.add_argument("--force", action="store_true",
|
||||
help="send even if the query looks like PII")
|
||||
args = parser.parse_args()
|
||||
|
||||
query = f"site:{args.site} {args.query}" if args.site else args.query
|
||||
|
||||
reason = ws.phi_reason(query)
|
||||
if reason and not args.force:
|
||||
print(f"refused: {reason}. This query would leave the host.", file=sys.stderr)
|
||||
print("Rephrase without identifiers, or pass --force if it is genuinely public.",
|
||||
file=sys.stderr)
|
||||
return 2
|
||||
|
||||
params = {}
|
||||
if args.lang:
|
||||
params["language"] = args.lang
|
||||
if args.fresh:
|
||||
params["time_range"] = args.fresh
|
||||
if args.category:
|
||||
params["categories"] = args.category
|
||||
if args.engines:
|
||||
params["engines"] = args.engines
|
||||
|
||||
key = ws.cache_key(query, params)
|
||||
conn = ws.open_cache(CACHE)
|
||||
|
||||
if not args.refresh:
|
||||
cached = ws.cache_get(conn, key, ttl=args.ttl)
|
||||
if cached is not None:
|
||||
out = ws.project(cached, limit=args.limit)
|
||||
out["cached"] = True
|
||||
sys.stdout.write(json.dumps(out, indent=2, ensure_ascii=False) + "\n"
|
||||
if args.json else to_yaml(out))
|
||||
return 0
|
||||
|
||||
conf = load_config()
|
||||
LOCK.parent.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# One request at a time across every agent on this host: the instance
|
||||
# suspends engines for minutes when several of us ask at once.
|
||||
with open(LOCK, "w") as lock:
|
||||
fcntl.flock(lock, fcntl.LOCK_EX)
|
||||
|
||||
payload = None
|
||||
for attempt in range(1 + len(ws.RETRY_BACKOFF)):
|
||||
delay = ws.wait_for(ws.last_call(conn), time.time())
|
||||
if delay:
|
||||
time.sleep(delay)
|
||||
ws.mark_call(conn)
|
||||
try:
|
||||
payload = fetch(conf, query, params, args.timeout)
|
||||
except Exception as error: # noqa: BLE001 - report, do not crash
|
||||
print(f"request failed: {error}", file=sys.stderr)
|
||||
return 3
|
||||
if ws.classify(payload) == "ok":
|
||||
break
|
||||
if attempt < len(ws.RETRY_BACKOFF):
|
||||
time.sleep(ws.RETRY_BACKOFF[attempt])
|
||||
|
||||
if ws.classify(payload) == "ok":
|
||||
ws.cache_put(conn, key, payload)
|
||||
|
||||
out = ws.project(payload, limit=args.limit)
|
||||
sys.stdout.write(json.dumps(out, indent=2, ensure_ascii=False) + "\n"
|
||||
if args.json else to_yaml(out))
|
||||
return 0 if out["status"] == "ok" else 3
|
||||
def main(argv: list[str]) -> int:
|
||||
print(
|
||||
"bin/web/search is deprecated; use bin/web/search.go",
|
||||
file=sys.stderr,
|
||||
)
|
||||
target = ROOT / "bin" / "web" / "search.go"
|
||||
os.execvp("go", ["go", "run", str(target), *argv])
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
sys.exit(main(sys.argv[1:]))
|
||||
|
||||
Executable
+232
@@ -0,0 +1,232 @@
|
||||
//usr/bin/env go run "$0" "$@"; exit
|
||||
//
|
||||
// bin/web/search.go - SearXNG as the second independent source (D3).
|
||||
//
|
||||
// ./bin/web/search.go "LadybugDB vector search"
|
||||
// ./bin/web/search.go "model2vec" --category it --json
|
||||
// ./bin/web/search.go "postgres" --site github.com --fresh year
|
||||
//
|
||||
// Empty results mean throttled, not "nothing exists". Exit 2 = PII refuse, 3 = throttled.
|
||||
// Config: $BRAIN_SEARCH_ENV (default $HOME/.config/brain/search.env).
|
||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"os"
|
||||
"strconv"
|
||||
"time"
|
||||
|
||||
"github.com/eSlider/2dph/internal/websearch"
|
||||
"golang.org/x/sys/unix"
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(run(os.Args[1:]))
|
||||
}
|
||||
|
||||
func run(args []string) int {
|
||||
var (
|
||||
query, site, lang, fresh, category, engines string
|
||||
limit = websearch.DefaultLimit
|
||||
jsonOut, refresh, force bool
|
||||
ttl = float64(websearch.CacheTTL)
|
||||
timeout = 25
|
||||
)
|
||||
i := 0
|
||||
for i < len(args) {
|
||||
a := args[i]
|
||||
switch {
|
||||
case a == "--json":
|
||||
jsonOut = true
|
||||
case a == "--refresh":
|
||||
refresh = true
|
||||
case a == "--force":
|
||||
force = true
|
||||
case (a == "-n" || a == "--limit") && i+1 < len(args):
|
||||
i++
|
||||
n, err := strconv.Atoi(args[i])
|
||||
if err != nil || n < 0 {
|
||||
fmt.Fprintln(os.Stderr, "web/search: --limit must be a non-negative integer")
|
||||
return 2
|
||||
}
|
||||
limit = n
|
||||
case a == "--site" && i+1 < len(args):
|
||||
i++
|
||||
site = args[i]
|
||||
case a == "--lang" && i+1 < len(args):
|
||||
i++
|
||||
lang = args[i]
|
||||
case a == "--fresh" && i+1 < len(args):
|
||||
i++
|
||||
fresh = args[i]
|
||||
case a == "--category" && i+1 < len(args):
|
||||
i++
|
||||
category = args[i]
|
||||
case a == "--engines" && i+1 < len(args):
|
||||
i++
|
||||
engines = args[i]
|
||||
case a == "--ttl" && i+1 < len(args):
|
||||
i++
|
||||
v, err := strconv.ParseFloat(args[i], 64)
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "web/search: --ttl must be a number")
|
||||
return 2
|
||||
}
|
||||
ttl = v
|
||||
case a == "--timeout" && i+1 < len(args):
|
||||
i++
|
||||
n, err := strconv.Atoi(args[i])
|
||||
if err != nil || n <= 0 {
|
||||
fmt.Fprintln(os.Stderr, "web/search: --timeout must be a positive integer")
|
||||
return 2
|
||||
}
|
||||
timeout = n
|
||||
case a == "-h" || a == "--help":
|
||||
fmt.Fprintln(os.Stderr, `usage: bin/web/search.go QUERY [--json] [-n N] [--site HOST] [--lang LANG] [--fresh day|week|month|year] [--category CAT] [--engines LIST] [--refresh] [--force]`)
|
||||
return 0
|
||||
case len(a) > 0 && a[0] != '-' && query == "":
|
||||
query = a
|
||||
default:
|
||||
fmt.Fprintf(os.Stderr, "web/search: unknown flag %s\n", a)
|
||||
return 2
|
||||
}
|
||||
i++
|
||||
}
|
||||
if query == "" {
|
||||
fmt.Fprintln(os.Stderr, "web/search: query required")
|
||||
return 2
|
||||
}
|
||||
if site != "" {
|
||||
query = "site:" + site + " " + query
|
||||
}
|
||||
if reason := websearch.PHIReason(query); reason != "" && !force {
|
||||
fmt.Fprintf(os.Stderr, "refused: %s. This query would leave the host.\n", reason)
|
||||
fmt.Fprintln(os.Stderr, "Rephrase without identifiers, or pass --force if it is genuinely public.")
|
||||
return 2
|
||||
}
|
||||
|
||||
params := map[string]string{}
|
||||
if lang != "" {
|
||||
params["language"] = lang
|
||||
}
|
||||
if fresh != "" {
|
||||
params["time_range"] = fresh
|
||||
}
|
||||
if category != "" {
|
||||
params["categories"] = category
|
||||
}
|
||||
if engines != "" {
|
||||
params["engines"] = engines
|
||||
}
|
||||
|
||||
cachePath := os.Getenv("BRAIN_SEARCH_CACHE")
|
||||
if cachePath == "" {
|
||||
cachePath = os.Getenv("HOME") + "/.cache/brain/web-search.sqlite"
|
||||
}
|
||||
cache, err := websearch.OpenCache(cachePath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
defer cache.Close()
|
||||
|
||||
key := websearch.CacheKey(query, params)
|
||||
now := float64(time.Now().Unix())
|
||||
if !refresh {
|
||||
if cached, err := cache.Get(key, ttl, now); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||
return 1
|
||||
} else if cached != nil {
|
||||
out := websearch.Project(*cached, limit, websearch.DefaultSnippetChars)
|
||||
out.Cached = true
|
||||
return writeOut(out, jsonOut)
|
||||
}
|
||||
}
|
||||
|
||||
envPath := os.Getenv("BRAIN_SEARCH_ENV")
|
||||
if envPath == "" {
|
||||
envPath = os.Getenv("HOME") + "/.config/brain/search.env"
|
||||
}
|
||||
conf, err := websearch.LoadConfig(envPath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "web/search: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
|
||||
lockPath := cachePath + ".lock"
|
||||
lock, err := os.OpenFile(lockPath, os.O_CREATE|os.O_RDWR, 0o600)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "web/search: lock: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
defer lock.Close()
|
||||
if err := unix.Flock(int(lock.Fd()), unix.LOCK_EX); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "web/search: lock: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
defer unix.Flock(int(lock.Fd()), unix.LOCK_UN)
|
||||
|
||||
var payload websearch.Payload
|
||||
attempts := 1 + len(websearch.RetryBackoff)
|
||||
client := &http.Client{}
|
||||
for attempt := 0; attempt < attempts; attempt++ {
|
||||
last, err := cache.LastCall()
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if delay := websearch.WaitFor(last, float64(time.Now().Unix()), websearch.MinInterval); delay > 0 {
|
||||
time.Sleep(time.Duration(delay * float64(time.Second)))
|
||||
}
|
||||
if err := cache.MarkCall(float64(time.Now().Unix())); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
payload, err = websearch.Fetch(client, conf, query, params, time.Duration(timeout)*time.Second)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "request failed: %v\n", err)
|
||||
return 3
|
||||
}
|
||||
if websearch.Classify(payload) == websearch.StatusOK {
|
||||
break
|
||||
}
|
||||
if attempt < len(websearch.RetryBackoff) {
|
||||
time.Sleep(time.Duration(websearch.RetryBackoff[attempt] * float64(time.Second)))
|
||||
}
|
||||
}
|
||||
|
||||
if websearch.Classify(payload) == websearch.StatusOK {
|
||||
if err := cache.Put(key, payload, float64(time.Now().Unix())); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||
}
|
||||
}
|
||||
out := websearch.Project(payload, limit, websearch.DefaultSnippetChars)
|
||||
code := writeOut(out, jsonOut)
|
||||
if out.Status != websearch.StatusOK && code == 0 {
|
||||
return 3
|
||||
}
|
||||
return code
|
||||
}
|
||||
|
||||
func writeOut(out websearch.Output, jsonOut bool) int {
|
||||
if jsonOut {
|
||||
enc := json.NewEncoder(os.Stdout)
|
||||
enc.SetIndent("", " ")
|
||||
enc.SetEscapeHTML(false)
|
||||
if err := enc.Encode(out); err != nil {
|
||||
return 1
|
||||
}
|
||||
if out.Status != websearch.StatusOK {
|
||||
return 3
|
||||
}
|
||||
return 0
|
||||
}
|
||||
fmt.Print(out.YAML())
|
||||
if out.Status != websearch.StatusOK {
|
||||
return 3
|
||||
}
|
||||
return 0
|
||||
}
|
||||
+89
-18
@@ -1,66 +1,137 @@
|
||||
# 2dph — docker composition
|
||||
#
|
||||
# docker compose run --rm brain index # rebuild graph
|
||||
# docker compose run --rm brain search "Matrix fed" # one-shot query
|
||||
# docker compose run --rm brain serve # async Go server
|
||||
# docker compose up brain-watch # auto re-index
|
||||
# docker compose up -d brain # API (Zig CGO serve)
|
||||
# docker compose --profile index run --rm index # Python rebuild
|
||||
# docker compose --profile picoclaw up brain-mcp
|
||||
# docker compose --profile reasoner up -d reasoner # CPU Ollama :11435
|
||||
# docker compose --profile searxng up -d
|
||||
#
|
||||
# Caching: the 128M model (HF_HOME) and kb.lbug (VAR_DIR) live in named
|
||||
# volumes, so rebuilds never redownload the model or re-derive the graph.
|
||||
# Secrets are never baked into the image: search.env + db-profiles.yml mount
|
||||
# read-only from ~/.config/brain.
|
||||
# Secrets never baked in: search.env + db-profiles.yml from ~/.config/brain.
|
||||
|
||||
name: 2dph
|
||||
|
||||
services:
|
||||
brain:
|
||||
image: ghcr.io/eslider/2dph:latest
|
||||
image: ghcr.io/eslider/2dph:api
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile
|
||||
target: api
|
||||
cache_from:
|
||||
- ghcr.io/eslider/2dph:cache
|
||||
command: ["brain", "search", "help"]
|
||||
command: ["serve"]
|
||||
environment: &env
|
||||
HF_HOME: /data/hf
|
||||
BRAIN_SEARCH_CACHE: /data/cache/web-search.sqlite
|
||||
BRAIN_DB_PROFILES: /secret/db-profiles.yml
|
||||
BRAIN_SEARCH_ENV: /secret/search.env
|
||||
KB_SEARCH_CMD: /app/bin/kb/search
|
||||
KB_ROOT: /data
|
||||
KB_WORKERS: "4"
|
||||
KB_PORT: "8630"
|
||||
volumes:
|
||||
- kb-model:/data/hf
|
||||
- kb-var:/data
|
||||
# corpus is read-only on the host, never written from the container
|
||||
- ..:/corpus:ro
|
||||
- ~/.config/brain:/secret:ro
|
||||
ports:
|
||||
- "127.0.0.1:8630:8630"
|
||||
read_only: true
|
||||
tmpfs:
|
||||
- /tmp
|
||||
healthcheck:
|
||||
test: ["CMD", "python3", "-c", "import ladybug, model2vec, mistune; print('ok')"]
|
||||
test: ["CMD", "wget", "-qO-", "http://127.0.0.1:8630/health"]
|
||||
interval: 30s
|
||||
timeout: 5s
|
||||
retries: 3
|
||||
restart: unless-stopped
|
||||
stop_grace_period: 20s
|
||||
|
||||
# watcher: re-index on corpus file change (watchdog script)
|
||||
brain-watch:
|
||||
image: ghcr.io/eslider/2dph:latest
|
||||
image: ghcr.io/eslider/2dph:api
|
||||
environment: *env
|
||||
volumes:
|
||||
- kb-model:/data/hf
|
||||
- kb-var:/data
|
||||
- ..:/corpus:ro
|
||||
- ~/.config/brain:/secret:ro
|
||||
command: ["brain", "watch", "/corpus"]
|
||||
command: ["watch", "/corpus"]
|
||||
read_only: true
|
||||
tmpfs:
|
||||
- /tmp
|
||||
restart: unless-stopped
|
||||
stop_grace_period: 20s
|
||||
|
||||
# Python write path (Ladybug rebuild). Not in the API image.
|
||||
# docker compose --profile index run --rm index
|
||||
index:
|
||||
profiles: ["index"]
|
||||
image: ghcr.io/eslider/2dph:index
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile
|
||||
target: index
|
||||
environment:
|
||||
HF_HOME: /data/hf
|
||||
KB_PY: python3
|
||||
volumes:
|
||||
- kb-model:/data/hf
|
||||
- kb-var:/app/var
|
||||
- ..:/corpus:ro
|
||||
- ~/.config/brain:/secret:ro
|
||||
command: ["index"]
|
||||
read_only: true
|
||||
tmpfs:
|
||||
- /tmp
|
||||
|
||||
# Optional local SearXNG (D3). Skip if BRAIN_SEARCH_URL already points at a
|
||||
# live instance — do not run a second copy on that host.
|
||||
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
|
||||
searxng:
|
||||
profiles: ["searxng"]
|
||||
image: docker.io/searxng/searxng:2026.8.10-0a118066d
|
||||
ports:
|
||||
- "127.0.0.1:8888:8080"
|
||||
environment:
|
||||
SEARXNG_SECRET: ${SEARXNG_SECRET:-}
|
||||
volumes:
|
||||
- ./deploy/searxng/settings.yml:/etc/searxng/settings.yml:ro
|
||||
- ./deploy/searxng/limiter.toml:/etc/searxng/limiter.toml:ro
|
||||
restart: unless-stopped
|
||||
|
||||
# MCP endpoint for an external agent (PicoClaw is not shipped here).
|
||||
# docker compose --profile picoclaw up brain-mcp
|
||||
brain-mcp:
|
||||
profiles: ["picoclaw"]
|
||||
image: ghcr.io/eslider/2dph:api
|
||||
environment: *env
|
||||
volumes:
|
||||
- kb-model:/data/hf
|
||||
- kb-var:/data
|
||||
- ~/.config/brain:/secret:ro
|
||||
command: ["serve"]
|
||||
ports:
|
||||
- "127.0.0.1:8630:8630"
|
||||
read_only: true
|
||||
tmpfs:
|
||||
- /tmp
|
||||
restart: unless-stopped
|
||||
|
||||
# CPU OpenAI-compatible sidecar (D18). Weights are pulled at runtime, not
|
||||
# baked into the 2dph image. Does not touch host Ollama on :11434.
|
||||
# docker compose --profile reasoner up -d reasoner
|
||||
# docker compose --profile reasoner exec reasoner ollama pull qwen3.5:9b
|
||||
reasoner:
|
||||
profiles: ["reasoner"]
|
||||
image: docker.io/ollama/ollama:latest
|
||||
environment:
|
||||
OLLAMA_NUM_GPU: "0"
|
||||
OLLAMA_HOST: "0.0.0.0:11434"
|
||||
ports:
|
||||
- "127.0.0.1:11435:11434"
|
||||
volumes:
|
||||
- reasoner-ollama:/root/.ollama
|
||||
restart: unless-stopped
|
||||
|
||||
volumes:
|
||||
kb-model:
|
||||
kb-var:
|
||||
kb-var:
|
||||
reasoner-ollama:
|
||||
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"mcpServers": {
|
||||
"2dph": {
|
||||
"url": "http://127.0.0.1:8630/mcp",
|
||||
"description": "2dph fact gate. Tool order: search → get → audit. throttled is not absence."
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
[botdetection.ip_lists]
|
||||
# RFC1918 only. Do not copy a live instance egress IP into git.
|
||||
pass_ip = [
|
||||
"10.0.0.0/8",
|
||||
"172.16.0.0/12",
|
||||
"192.168.0.0/16",
|
||||
]
|
||||
@@ -0,0 +1,28 @@
|
||||
use_default_settings: true
|
||||
|
||||
general:
|
||||
instance_name: "2dph"
|
||||
|
||||
search:
|
||||
formats:
|
||||
- html
|
||||
- json
|
||||
suspended_times:
|
||||
SearxEngineCaptcha: 300
|
||||
SearxEngineTooManyRequests: 120
|
||||
SearxEngineAccessDenied: 300
|
||||
|
||||
server:
|
||||
limiter: true
|
||||
image_proxy: false
|
||||
# secret_key comes from SEARXNG_SECRET (never commit a real secret)
|
||||
|
||||
engines:
|
||||
- name: bing
|
||||
disabled: false
|
||||
- name: google
|
||||
disabled: false
|
||||
- name: duckduckgo
|
||||
disabled: false
|
||||
- name: wikipedia
|
||||
disabled: false
|
||||
+7
-1
@@ -6,5 +6,11 @@ Brain/ops/eSlider stack. Facts need proof or they are
|
||||
|
||||
- [PLAN.md](../PLAN.md) — decisions, execution order, open questions (v2)
|
||||
- [design](design.md) — schema, deduction model, sources
|
||||
- [reasoner](reasoner.md) — D18 CPU bake-off (Qwen3.5-9B vs Bonsai / Qwen3.6-27B)
|
||||
- [Gitea issues](https://git.produktor.io/eSlider/2dph/issues) — work board (origin)
|
||||
|
||||
Published docs live here and mirror the project state.
|
||||
Search: `bin/brain/search.go "query"` (HTTP: `bin/brain/serve.go` —
|
||||
`/health` `/search` `/get` `/stats` `/audit` `/ingest`). `--hop` is
|
||||
not a walk; the flag errors until File/FROM_FILE edges exist.
|
||||
|
||||
Published docs live here and mirror the project state.
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
# Chat Import Pipeline
|
||||
|
||||
Plan: https://git.produktor.io/eSlider/brain-chats-import/issues/1
|
||||
|
||||
## Env vars (set in shell, never committed)
|
||||
|
||||
```
|
||||
TELEGRAM_MCP_DIR
|
||||
TELEGRAM_API_ID / TELEGRAM_API_HASH / TELEGRAM_PHONE
|
||||
TELEGRAM_SESSION_STRING
|
||||
ONLYOFFICE_URL / ONLYOFFICE_USER / ONLYOFFICE_PASS
|
||||
OO_CLI (default: $HOME/go/bin/oo)
|
||||
```
|
||||
|
||||
## Quick reference
|
||||
|
||||
```
|
||||
./bin/chats/sync.go telegram --limit 100
|
||||
./bin/chats/import.go
|
||||
./bin/chats/facts.go
|
||||
./bin/chats/apply.go --dry-run
|
||||
```
|
||||
|
||||
JSONL → markdown only. Brain ingest is `bin/brain/index.go` (not a `chats index`).
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user