Compare commits

...
Author SHA1 Message Date
eSlider 515088f4e1 feat: escalate brain search to web when facts cannot confirm
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
2026-08-13 20:23:07 +01:00
eSliderandGitHub 368a757612 feat: Go SearXNG client; throttled is not absence (#16)
Tests / Test (push) Failing after 4s
Tests / Release (semver) (push) Skipped
2026-08-13 19:53:31 +01:00
eSliderandGitHub a8675ac33b feat: read git history with go-git, not the git binary (#15)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
2026-08-13 18:07:56 +01:00
eSliderandGitHub de632ba6cc docs: delete agent-cost; rename kb-search skill to brain (#14)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
* docs: delete agent-cost; rename kb-search skill to brain.

bin/agents/cost does not exist. CI unittest now fails if a SKILL.md names a
missing bin/ path.

* test: gate SKILL.md bin/ paths; name the brain skill brain.

Follow-up to the agent-cost delete: unittest fails if a skill names a missing
tool. Frontmatter name is brain, not kb-search.
2026-08-13 17:55:41 +01:00
eSliderandGitHub 66c87842e2 feat: in-process HTTP search; /get /stats /audit /ingest. (#13)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
bin/brain/serve.go (ladybug tags) calls internal/brain instead of exec.
HTTP tests inject a fake API so CI stays cgo-free. ExecSearcher remains
the fallback when the binary is built without system_ladybug.
2026-08-13 17:52:15 +01:00
eSliderandGitHub 20b78a9a20 feat: brain/index.go shebang; mail import is not a brain write (D14). (#12)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
Commands live at bin/brain/{index,get,stats,eval,watch}.go and
bin/mail/import.go, bin/markdown/import.go, bin/postgres/query.go.
Python remains the Ladybug write worker. index_mail is a deprecation
shim that rebuilds via --with-mail.
2026-08-13 17:46:25 +01:00
eSliderandGitHub 5d4b3427a4 refactor: chats method shebangs; drop chats index (D14). (#11)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
Parsers and commands live in internal/chats. bin/chats/{sync,import,facts,apply}.go
are tagged shebang mains. Brain ingest is not a chats command.
2026-08-13 17:31:03 +01:00
eSliderandGitHub 1c7db6d499 docs: name bin/brain/search.go; --hop is not a graph walk. (#10)
Tests / Test (push) Failing after 6s
Tests / Release (semver) (push) Skipped
Published docs and skills still taught bin/kb/search --hop 1. Search lives
at bin/brain/search.go; --hop errors until File edges exist. A unittest
gates the SoT so the lie cannot return.
2026-08-13 17:23:40 +01:00
eSliderandGitHub 0786ddcb06 feat: bin/brain/serve.go; search backend is Go not Python (#9)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
* feat(brain): HTTP serve from bin/brain/serve.go, default Go search binary.

Move the HTTP package to internal/httpapi. Default backend is
var/bin/brain-search, not Python. bin/serve.go stays as a deprecation shim.

* feat(httpapi): default search backend is var/bin/brain-search.

bin/brain/serve.go is the command; bin/serve.go stays as a tagged
deprecation shim. Tests fail if the default path still names Python.
2026-08-13 14:32:21 +01:00
eSliderandGitHub eeb5b79cf2 refactor: one Go module; brain search in bin/brain + internal/brain. (#8)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
Collapse nested kbsearch/chats go.mod into the root module. Ranking stays
cgo-free under internal/brain/rank so CI does not need ladybug. bin/kb/search
is a deprecation wrapper that still sets CGO and builds the binary.
2026-08-13 14:26:54 +01:00
eSliderandGitHub 5990feb1f6 docs: point issues at Gitea origin (D15). (#7)
Tests / Test (push) Failing after 4s
Tests / Release (semver) (push) Skipped
GitHub stays the public clone for PRs and Actions. Work board is
https://git.produktor.io/eSlider/2dph/issues.
2026-08-13 14:11:00 +01:00
eSliderandGitHub d27a738fee feat(chats): parse LinkedIn MCP v4.22 inbox/conversation blobs. (#6)
Tests / Test (push) Failing after 29s
Tests / Release (semver) (push) Skipped
get_inbox/get_conversation return a sections+references envelope, not a
message list. Parser is covered by synthetic Alice/Bob fixtures; CI now
runs the nested bin/chats tests. Session check no longer launches Chromium.
2026-08-13 12:25:51 +01:00
eSliderandGitHub 669e184cf6 fix(kbsearch): rank FTS correctly, filter before -n, start the daemon. (#5)
Go search took worst BM25 hits (ORDER BY score), cut to -n before --root,
and never called ensureDaemon. Ranking and flag parsing move to a cgo-free
package so CI can fail those regressions without ladybug. --hop errors
instead of being swallowed into the query.
2026-08-13 12:19:26 +01:00
eSliderandGitHub ebc3f948c1 Add Gmail --query to mail/sync (default in:inbox) (#4)
* Add --query to Gmail mail/sync instead of always listing in:inbox.

Callers keep the search string; default remains in:inbox.

* Document Gmail --query on the mail/sync pipeline.

* test(mail): assert Gmail --query reaches ListIDs, not only the CLI flag.

ParseCLI coverage left a hole: an empty query still has to become in:inbox
and a custom q has to be the string the client lists with.
2026-08-13 12:19:22 +01:00
eSlider fe6a02024c feat(chats): LinkedIn source — MCP client via get_inbox + get_conversation
- LinkedInMCPSource: MCP JSON-RPC, как TelegramMCPSource
- sync linkedin --limit N: выгрузка сообщений из LinkedIn
- Проверка сессии: uvx mcp-server-linkedin --status
- Вывод инструкции если сессия истекла
- JSONL в var/chats/linkedin/<thread_id>/messages.jsonl
2026-08-13 00:10:42 +01:00
eSlider 4a065d9838 docs: add edelweiss to GitHub safety rules 2026-08-13 00:07:01 +01:00
eSlider e3c6ef5684 chore: remove edelweiss references from public repo 2026-08-13 00:06:49 +01:00
eSlider 98c14e23f1 docs: GitHub safety rules — no absolute paths, PII, secrets, curasoft 2026-08-13 00:02:09 +01:00
eSlider ff1716de40 fix: resolve plan.md conflict, remove remaining /mnt/ paths 2026-08-13 00:00:48 +01:00
eSlider fec5325c7a chore: clean absolute paths, curasoft refs, secrets from history
- bin/chats/: env-based paths, no /mnt/ /home/ hardcodes
- bin/edelweiss-pilot: remove curasoft, use DOCS_BASE env var
- bin/facts/crm: use KNOWLEDGE_MESH_SEED env var
- compose.edelweiss.yml: remove curasoft volumes, use DOCS_BASE
- docs/chat-import-plan.md: link to Gitea issue, no secrets
- bin/seed-edelweiss-facts.py: removed (curasoft-only)
2026-08-13 00:00:15 +01:00
eSlider 27d9521e7f bin/chats: Phase 1 MVP — Telegram sync/import/index/facts/apply
- bin/chats/ — nested Go module (как bin/kbsearch/)
  - sync telegram — MCP JSON-RPC клиент, 31 личный чат, 922 сообщения
  - import — конвертация JSONL → MD с YAML frontmatter
  - index — делегирует bin/kb/index --corpus (132 leafs в brain)
  - facts — regex extraction phone/email/linkedin с валидацией
    (исключены: даты, суммы, номера карт, инвойсы)
  - apply — oo CLI cross-check + dry-run
- Source interface для будущих WhatsApp/LinkedIn
- 4 system tests (import, facts, empty, roundtrip) — синтетические данные
- bin/chat — build+exec wrapper
- docs/chat-import-plan.md — прогресс, пути к env (без секретов)

Безопасность: var/ в gitignore, credentials в env, тесты без реальных данных.
2026-08-12 23:59:29 +01:00
eSlider a7cb8d4c76 docs: chat import pipeline plan — link to Gitea issue #1 2026-08-12 23:59:20 +01:00
eSliderandCursor 53cd00284d fix(kb): seed facts before CREATE indexes (FTS MERGE corruption)
Upsert under live FTS raises "document for node offset N is missing".
Add --skip-indexes; edelweiss-pilot index = write → seed → ensure_indexes.
Ship seed-edelweiss-facts.py (paired lexicon/OO/interview/QEMU facts).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:09:02 +01:00
eSliderandCursor b73b4d4f97 fix(kb): stop DROP INDEX killing HNSW via Ladybug ghost catalog
Ladybug 0.19 DROP INDEX leaves `_0_Leaf_vec_UPPER` / `0_id_docs` in catalog so
CREATE fails while SHOW_INDEXES omits the index; create_fts_and_vector used to
swallow that. Never drop FTS/VECTOR; ensure_indexes after upserts; rebuild =
delete kb.lbug. Add compose.edelweiss.yml + regression tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:05:14 +01:00
eSlider 1d1f6a90ff Remove curasoft references, rename to detective method
- PLAN.md: replace 'curasoft-detective' with 'detective method'
- README.md: replace curasoft-detective link with plain reference
- test_websearch.py: fix test domain from ticket.curasoft.de to example.com
- Rewrote git history with git-filter-repo to remove all traces
2026-08-12 13:46:54 +01:00
eSlider f220bcd95a kbsearch: Go implementation with daemon model serving
- New nested module bin/kbsearch with Go implementation of bin/kb/search
- Embedding model (potion-multilingual-128M) served by localhost daemon
  so repeated CLI calls reuse the loaded model
- Bash launcher bin/kb/search builds binary on first run, caches to var/bin/
- Hybrid FTS + vector search (RRF k=60) matching Python kblib behavior
- YAML output via port of yamlout.py (ordered keys, same format)
- JSON output with proper field order
- All flags: --root, --repo, -n, --json, --list-model
- Root go.mod reverted to 1.25.0 (kbsearch is isolated nested module)
- CI passes: go test ./... and go vet ./... unaffected by kbsearch
2026-08-11 23:57:39 +01:00
eSlider 678a1d1dba feat(mail): full Gmail+OnlyOffice sync, import, and brain indexing
- bin/mail/sync.go: async Go sync engine (8 workers, paginated Gmail via
  API + OnlyOffice IMAP); Gmail attachments key off body.attachmentId, not
  MIME partId; ICS sidecars Latin-1->UTF-8 normalized (TestICSToMarkdownNormalizesLatin1)
- bin/mail/import: message.json -> markdown; PDFs via pdftotext -layout
  fast path with docling subprocess fallback for the ~5% textless files
- bin/mail/index_mail: fresh-rebuild indexer (repo corpus + mail) avoiding
  ladybug WAL corruption on bulk-insert into indexed DBs; split from import
- bin/kb/index: keep FTS/VECTOR indexes across incremental runs (drop+recreate
  leaves stale backing tables killing the vector index)
- docs: README/PLAN/AGENTS cover the mail pipeline

Result: 17,835 messages -> 28,918 info leafs, FTS+HNSW healthy.
2026-08-11 21:57:38 +01:00
eSlider 8781c0c3eb refactor(tools): bin/{subject}/{method} layout; Go serve+watch modules
Move serve/ (module) -> bin/server, tools/ -> bin/tools, replace bin/kb-watch
bash with bin/watch Go package; self-executing Go shebangs bin/serve.go and
bin/kb/watch.go; Docker + CI + git/import + docs repointed. Multi-stage image
builds static serve+watch binaries (no Go runtime in container).
2026-08-11 09:52:20 +01:00
eSlider d6b17e8819 feat(kb): CRM association proof via oo, fix ssh-tunnel self-ref + oo creds
- bin/facts/crm: prove person<->company/company<->project against ooCRM
  x corpus SoT (knowledge-mesh-seed.yaml), write 78 facts (root=facts)
- tools/crmfacts.py + test_crm_facts.py: parser under unit tests (26 pass)
- docs/crm-associations-proof.md: provable graph, mistakes, fixes
- oo merge 759->763 resolves duplicate GoldenRatio.Exchange legal entity
- bin/db/ssh-tunnel: "$0" self-check + accept-new/BatchMode ssh flags
- AGENTS.md: document bin/facts/crm
2026-08-10 23:22:34 +01:00
147 changed files with 13701 additions and 667 deletions
-1
View File
@@ -10,5 +10,4 @@ __pycache__
.cache
.secrets
.skills-tmp
serve/serve
docs/.build
+10 -4
View File
@@ -19,6 +19,10 @@ jobs:
with:
fetch-depth: 0
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
@@ -31,18 +35,20 @@ jobs:
run: |
bash -n bin/db/psql-yq
bash -n bin/db/ssh-tunnel
bash -n bin/kb-watch
bash -n bin/docker-entrypoint
bash -n bin/kb/search
- name: Python unit tests (offline, vendored tools)
run: |
uv run python -m unittest discover -s tools -t .
uv run python -m unittest discover -s bin/tools -t .
- name: Go serve tests (async, goroutine-bounded)
- name: Go tests (root module, no ladybug cgo)
run: |
go vet ./...
go test ./... -count=1
working-directory: serve
- name: brain ranking tests (no cgo / no ladybug)
run: go test ./internal/brain/rank -count=1
- name: facts/audit self (lexicon consistency, no network)
run: |
+3 -1
View File
@@ -8,4 +8,6 @@ __pycache__/
.DS_Store
*.env
.env
.secrets/
.secrets/
lib-ladybug/
go.work.local
+57 -8
View File
@@ -24,7 +24,7 @@ Read first: [PLAN](PLAN.md) → [docs](docs/).
2. **Read-only data sources.** Ladybug `var/kb.lbug` and Postgres are opened
read-only for queries. Index rebuilds write to `var/` (gitignored).
3. **PII.** `brain-test`, `cs_brain` client data is never read or quoted.
4. **No main pushes.** Feature branches + PR via `gh`; CI must be green.
4. **No main pushes.** Feature branches + GitHub PR (`gh`); CI (Actions) must be green. Work board: [Gitea issues](https://git.produktor.io/eSlider/2dph/issues).
5. **TDD.** Failing test before tool code. Unit tests run offline against
fixtures; network/db calls are wrapped.
6. **docs reflect behaviour.** Any change updates `docs/` + `PLAN.md` status.
@@ -35,22 +35,58 @@ Read first: [PLAN](PLAN.md) → [docs](docs/).
PLAN.md decisions + execution + open questions
docs/ published docs
skills/ in-project agent skills (vendored, no external links)
bin/ self-describing tools bin/{subject}/{method} (shebang)
bin/kb-watch corpus watcher (mtimes, no inotify deps)
bin/ self-describing tools bin/{subject}/{method}.go (shebang)
bin/brain/ search.go serve.go index.go get.go stats.go eval.go watch.go
bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats
bin/mail/ sync.go import.go (index_mail → brain/index.go)
bin/markdown/ import.go (mistune leafs)
bin/postgres/ query.go (read-only YAML)
bin/git/ import.go (go-git history; Python shim execs it)
bin/web/ search.go (SearXNG; Python shim execs it)
internal/ shared Go (brain/rank is cgo-free; chats parsers; gitlog; websearch)
bin/watch/ corpus watcher (used by bin/brain/watch.go)
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
bin/docker-entrypoint container entrypoint (brain index|search|serve|watch)
serve/ async Go HTTP server (goroutines, bounded worker pool)
tools/ vendored python libs behind bin/* (yamlout, websearch)
compose.yaml docker composition (root level, not docker/)
Dockerfile multi-stage: python deps + static Go serve
var/ kb.lbug, caches (gitignored)
Dockerfile multi-stage: python deps + static Go binaries
var/ kb.lbug, var/mail/*, caches (gitignored)
.venv/ ladybug + model2vec + mistune
```
## Mail pipeline
```bash
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
bin/mail/sync.go --source gmail --query 'from:example.com' --out var/mail # Gmail search (default in:inbox)
bin/mail/import.go --from-raw var/mail # message.json → message.md (convert only)
bin/brain/index.go --rebuild # rebuild brain incl. all mail (fresh DB)
```
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
`body.attachmentId` (not partId) for attachments.
- `import` converts body + attachments to markdown. PDFs use poppler
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to
docling (isolated subprocess — its native onnx can segfault the parent).
Conversion never touches the brain DB (crash safety).
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Ladybug
corrupts its WAL when brand-new leafs are bulk-inserted while FTS/vector
indexes exist; a fresh DB with indexes created last is the only safe path.
Keep conversion + indexing separate so a conversion crash can't leave the
DB mid-transaction.
## Tools
```bash
bin/facts/audit ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
bin/kb/search "query" [--hop N] [--repo X] # deduction search → YAML
bin/facts/crm [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
bin/brain/search.go "query" --no-web # local graph only
bin/brain/get.go <id> [--body]
bin/markdown/import.go [dir] # mistune leaves → YAML
bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
bin/md/tables # what the graph holds → YAML
bin/brain/deduce "question" # thinking wrapper
```
@@ -58,6 +94,19 @@ bin/brain/deduce "question" # thinking wrapper
Never start a shell command with `cd` — use the tool working-directory
parameter. Search before reading whole files.
## GitHub safety rules (ABSOLUTE — never violate)
1. **No absolute paths in committed files.** Replace `/mnt/`, `/home/<user>/`,
`/Users/<user>/` with env vars (`$HOME`, `$PROJECTS_ROOT`, `$DOCS_BASE`).
2. **No PII in commits.** No real names, phones, emails of third parties.
Test data must be synthetic (Alice, Bob, Charlie, Diana, example.com).
3. **No credentials/secrets in commits.** API keys, tokens, passwords, session
strings, phone numbers only in gitignored `.env` files, referenced by path.
4. **Curasoft, edelweiss — no files, no mentions.** Remove all traces if found.
5. **Check git history before push.** If any commit contains leaks, rewrite
history (rebase + force push) AND delete affected GitHub releases/tags.
6. **`docs/chat-import-plan.md`** — reference Gitea issue, never embed secrets.
## Communication
Same tone as the corpus: plain, lists, no hype. Sign-off `Andriy Oblivantsev`.
+14 -10
View File
@@ -14,23 +14,27 @@ COPY requirements.lock.txt /tmp/requirements.lock.txt
RUN python -m pip install --no-cache-dir -r /tmp/requirements.lock.txt \
&& rm /tmp/requirements.lock.txt
# Go serve: static binary, no interpreter at runtime
FROM golang:1.25 AS serve-build
WORKDIR /src/serve
COPY serve/go.mod serve/go.sum* ./
COPY serve .
RUN CGO_ENABLED=0 go build -o /serve -ldflags="-s -w" .
# Go services: static binaries, no interpreter at runtime
FROM golang:1.25 AS go-build
WORKDIR /src
COPY go.mod ./
COPY bin/server ./bin/server
COPY bin/watch ./bin/watch
RUN CGO_ENABLED=0 go build -o /serve ./bin/server \
&& CGO_ENABLED=0 go build -o /watch ./bin/watch
# runtime: python toolchain + Go server
# runtime: python toolchain + Go services
FROM base
COPY . .
COPY --from=serve-build /serve /app/serve/serve
RUN chmod +x /app/bin/kb-watch /app/bin/docker-entrypoint \
COPY --from=go-build /serve /app/bin/serve
COPY --from=go-build /watch /app/bin/watch
RUN chmod +x /app/bin/docker-entrypoint \
&& chown -R 2dph:2dph /app
USER 2dph
ENV PATH="/app/bin:${PATH}" \
KB_PY=python3
KB_PY=python3 \
KB_ROOT=/app
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
CMD python -c "import model2vec, ladybug, mistune; print('ok')" || exit 1
+48 -18
View File
@@ -25,11 +25,11 @@ detective method: **a fact needs ≥2 independent sources or it is
| # | Question | Answer |
|---|----------|--------|
| D1 | RAG corpus | ops stack (chat, onlyoffice, gitea/NPM, searchxng, observability, ai-bot, mcp-servers, `~/.ssh/config`) + portfolio. Exclude `office.dev` + jobs/applications. |
| D2 | skill merging | integrate skills **in this project** `skills/`; skip gitea / brain-detective-depe ndent skills. |
| D3 | web search | import `web-search`, retire local `searxng-ops`. Vendored here, no remote link. |
| D2 | skill merging | integrate skills **in this project** `skills/`; skip gitea / brain-dependent skills. |
| D3 | web search | Go client `bin/web/search.go` (`internal/websearch`). SearXNG URL is config (`BRAIN_SEARCH_URL`). Optional Compose profile `searxng` (sanitized settings). Do not run a second copy on a host that already has one. Empty/`throttled` ≠ “nothing exists”. |
| D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. |
| D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). |
| D6 | graph engine | **LadybugDB** (Kuzu successor, MIT, embedded, native FTS+vector+Cypher). Python binding for `bin/*`; Go shebang for golang tools. |
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`); Python remains for index/write until the Go write path is safe. |
| D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). |
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. |
| D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. |
@@ -37,9 +37,12 @@ detective method: **a fact needs ≥2 independent sources or it is
| D11 | strong/weak | `root` column: `facts` (strong) vs `info` (weak). Answer is `confirmed` only from facts root. |
| D12 | transactional | facts and info split by root but **written in the same Ladybug transaction (ACID)** on every write. |
| D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. |
| D14 | tooling style | `bin/{subject}/{method}` self-describing: shebang line 1, usage comment from line 2. Go shebang: `///usr/bin/env go run "$0" "$@"; exit`. |
| D15 | repo | GitHub `eSlider/2dph`, public (like sibling repos), push/commit via `gh`, TDD + commit every change, CI/CD. |
| D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. |
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. |
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
| D18 | reasoner | Pluggable OpenAI-compatible URL. RAM: Qwen3.5-9B. Quality: Bonsai-27B or Qwen3.6-27B. No official Qwen3.6-9B. |
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
## Architecture
@@ -47,16 +50,25 @@ detective method: **a fact needs ≥2 independent sources or it is
2dph/
PLAN.md / AGENTS.md
docs/ published docs (this conversation → docs/ as md)
skills/ in-project skills (web-search, db-yaml, kb-search, agent-cost, diataxis-docs, …)
skills/ in-project skills (web-search, db-yaml, brain, diataxis-docs)
bin/
facts/extract auto-pair 2 sources → lexicon yaml + graph
facts/audit ["self"|"facts"|"info"|"stale"] 2-source + staleness gate
kb/index build FTS + HNSW from corpus
kb/search deduction: facts → info → web-search; --hop N
kb/get kb/stats kb/eval
md/import md/select md/tables md/gaps (mistune)
kb/index Python write path (called by bin/brain/index.go)
brain/index.go rebuild FTS + HNSW (incl. --with-mail)
brain/get.go stats.go eval.go watch.go
brain/search.go deduction: facts → info → web-search
brain/serve.go HTTP API in-process (internal/httpapi + internal/brain)
mail/import.go JSON → markdown (no brain write)
markdown/import.go mistune leaves
postgres/query.go read-only YAML (wraps bin/db/psql-yq)
git/import.go go-git history (no git binary; conversion only)
web/search.go SearXNG client (throttled ≠ absence)
chats/sync.go import.go facts.go apply.go
(libs in internal/chats; no chats index)
md/import (deprecated; bin/markdown/import.go)
brain/extract brain/audit brain/deduce (thinking wrapper)
web/search (vendored)
web/search (deprecated shim → web/search.go)
db/psql-yq (vendored)
ssh-tunnel onlyoffice pg tunnel 5433
var/kb.lbug single embedded store (gitignored)
@@ -93,18 +105,36 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
- OQ1: mutually-contradicting evidence — how to resolve (authority weighting,
temporal freshness, audit adjudication).
- OQ2: OCR pipeline for pdfs/images/docs (late phase).
- OQ2: OCR pipeline for pdfs/images/docs — mostly solved: poppler pdftotext
fast-path for born-digital PDFs, docling fallback for the ~5% textless ones.
- OQ3: optional duckdb-md layer for `SELECT … FORMAT MARKDOWN` export/write-back.
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
serialize and unambiguous; YAML only where humans edit files.
## Mail pipeline (done)
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
Gmail attachments key off `body.attachmentId`, not MIME `partId`.
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
`pdftotext -layout` (~15ms) with docling subprocess fallback; ICS sidecars
Latin-1→UTF-8 normalized.
3. `bin/brain/index.go --rebuild` — fresh rebuild (repo corpus + mail) because ladybug
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
indexing stay separate for crash safety. `bin/mail/index_mail` is a
deprecation shim.
4. Result: 17,835 messages → 28,918 info leafs, FTS + HNSW healthy, searchable
via `bin/brain/search.go`.
## CI/CD pipeline (D15)
`.github/workflows/ci.yml`:
1. go vet + go test ./... (Go tools)
2. python -m unittest discover + pytest (Py tools)
3. bin/facts/audit self (lexicon internal consistency)
4. bin/kb/eval (recall@5 ≥ 0.95, gates index regressions)
5. md-docs build/lint if docs tooling arrives.
1. go vet + go test ./... (root module; packages without ladybug cgo)
2. `go test ./internal/brain/rank` (cgo-free ranking + flag parser)
3. python -m unittest discover -s bin/tools (includes published-docs SoT)
4. bin/facts/audit self (lexicon internal consistency)
5. bin/brain/eval.go (recall@5 ≥ 0.95, gates index regressions)
6. md-docs build/lint if docs tooling arrives.
Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as
`db/tech-poc`: contract first where there is an OpenAPI/message shape.
@@ -113,7 +143,7 @@ Feedback loop: every commit → PR → CI → green/gate → merge. Same discipl
1. scaffold repo (:done after this file + AGENTS.md + .gitignore + ci)
2. gh repo create eSlider/2dph --private + initial commit + CI
3. vendored skill integration (web-search, db-yaml, kb-search, agent-cost, diataxis-docs) — no remote links
3. vendored skill integration (web-search, db-yaml, brain, diataxis-docs) — no remote links
4. .venv: ladybug + model2vec + mistune
5. schema + tools with TDD (kb + md + facts + brain)
6. ~/.config/brain config
+53 -18
View File
@@ -30,9 +30,9 @@ graph TB
subgraph dph["2dph tools"]
EX["bin/facts/extract<br/>2-source pairing"]
AU["bin/facts/audit<br/>confidence + staleness"]
IDX["bin/kb/index<br/>chunk + embed"]
MD["bin/md/import<br/>mistune leaves"]
SR["bin/kb/search<br/>deduction + --hop"]
IDX["bin/brain/index.go<br/>chunk + embed"]
MD["bin/markdown/import.go<br/>mistune leaves"]
SR["bin/brain/search.go<br/>deduction"]
end
subgraph store["Ladybug var/kb.lbug"]
@@ -85,19 +85,51 @@ fact; conflicting sources or a single source → `hypothesis` → `(not confirme
## Deduction search
```bash
bin/kb/search "Matrix federation over HTTPS" # facts → info → web-search
bin/kb/search "what runs on arc-2" --hop 1 # walk graph edges
bin/kb/search "where is cs-lexicon" --json | yq '.' # YAML by default
bin/kb/get <id> --body # full chunk on demand
bin/kb/stats # index health
bin/kb/eval # recall@5 gate
bin/brain/search.go "Matrix federation over HTTPS" # facts → info → web
bin/brain/search.go "onlyoffice postgres" --root facts
bin/brain/search.go "where is cs-lexicon" --json | yq '.'
bin/brain/search.go "upstream flag" --no-web # local graph only
bin/brain/get.go <id> --body # full chunk on demand
bin/brain/stats.go # index health
bin/brain/eval.go # recall@5 gate
```
`--hop` is not implemented (needs File/FROM_FILE edges); the flag errors instead of walking. `bin/kb/search` is a deprecated wrapper around `bin/brain/search.go`.
Git history is read with [go-git](https://github.com/go-git/go-git) (no git binary):
```bash
bin/git/import.go --json --limit 100 # commit leafs for this repo
bin/git/import.go --root "$PROJECTS_ROOT" --json # one pass per .git under root
```
Conversion only. Graph write (`File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`) stays with `bin/brain/index.go`.
Web search (second independent source) goes through SearXNG. Empty results mean **throttled**, not “nothing exists”:
```bash
bin/web/search.go "LadybugDB vector index" --json
# Optional local instance (skip if BRAIN_SEARCH_URL already points at one):
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
```
Mail is a first-class corpus (retrievable through the same search):
```bash
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
bin/mail/import.go --from-raw var/mail # JSON → markdown
bin/brain/index.go --rebuild # rebuild brain (incl. mail)
bin/brain/search.go "invoice from last week" # same search over mail leafs
```
## Storage
- **LadybugDB** — single `var/kb.lbug`, Cypher property graph, HNSW + BM25
in one engine, embedded (no server), ACID, read-only-safe for concurrent
readers.
readers. **Never `DROP INDEX` FTS/VECTOR** on Ladybug 0.19: DROP leaves
ghost catalog tables (`_0_Leaf_vec_UPPER`) so recreate fails while
`SHOW_INDEXES` omits HNSW. Fresh indexes = delete `var/kb.lbug` +
`bin/brain/index.go --rebuild`. Use `ensure_indexes()` after upserts.
- **model2vec** — `potion-multilingual-128M` static embeddings (256-dim),
CPU-fast, deterministic, no Ollama runtime dependency.
- facts and info split semantically by `root` column but written inside the
@@ -105,10 +137,10 @@ bin/kb/eval # recall@5 gate
## Tooling conventions
`bin/{subject}/{method}` — self-describing: shebang on line 1, usage comment
from line 2. bash + python primary; golang via the Go shebang when a compiled
helper is right. YAML default output, `--json` for machines. Everything that
touches network/db is read-only, throttled, cached. Tests gate every commit.
`bin/{subject}/{method}.go` — self-describing: shebang on line 1, usage comment
from line 2. Shared code in `internal/`. YAML default output, `--json` for
machines. Tests gate every commit. HTTP: `bin/brain/serve.go` calls
`internal/brain` in-process (`/health` `/search` `/get` `/stats` `/audit` `/ingest`).
## Development
@@ -116,7 +148,7 @@ touches network/db is read-only, throttled, cached. Tests gate every commit.
uv venv .venv # Python 3.12, uv-managed
uv pip install -r requirements.lock.txt # pinned toolchain
bin/facts/audit self # lexicon consistency gate
go test ./... && python -m unittest discover -s tools -t .
go test ./... && python -m unittest discover -s bin/tools -t .
```
Docker (optional, cached model + var volumes):
@@ -124,7 +156,7 @@ Docker (optional, cached model + var volumes):
```bash
docker compose run --rm brain index # (re)index corpus
docker compose run --rm brain search "query" # one-shot query
docker compose run --rm brain serve # async Go HTTP server
docker compose run --rm brain serve # bin/brain/serve.go
docker compose up brain-watch # auto re-index on change
```
@@ -134,6 +166,9 @@ docker compose up brain-watch # auto re-index on change
Neo4j + Qdrant + Matrix RAG brain
- [agent-skills](https://github.com/eSlider/agent-skills) — upstream
skills (`web-search`, `db-yaml`, …) that 2dph integrates
- [detective](https://github.com/detective) — the two-source method
- detective method — the two-source method
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
Work board (issues): [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
+3
View File
@@ -0,0 +1,3 @@
// Commands in this directory are shebang mains (search.go, serve.go, index.go,
// get.go, stats.go, eval.go, watch.go), each behind an exclusive build tag.
package main
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=brain_eval "$0" "$@"; exit
//go:build brain_eval
//
// bin/brain/eval.go - recall@5 gate.
//
// ./bin/brain/eval.go
// ./bin/brain/eval.go --json
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/kb/eval", os.Args[1:]))
}
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=brain_get "$0" "$@"; exit
//go:build brain_get
//
// bin/brain/get.go - read one leaf by id.
//
// ./bin/brain/get.go <id>
// ./bin/brain/get.go <id> --body
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/kb/get", os.Args[1:]))
}
+24
View File
@@ -0,0 +1,24 @@
//usr/bin/env go run -tags=brain_index "$0" "$@"; exit
//go:build brain_index
//
// bin/brain/index.go - rebuild the Ladybug graph (Python write path).
//
// ./bin/brain/index.go --rebuild
// ./bin/brain/index.go --rebuild --with-mail
// ./bin/brain/index.go --dry-run --with-mail
//
// v1 write is always a rebuild when mail is included (live FTS/HNSW + bulk
// insert corrupts Ladybug 0.19 WAL). `add` is v2.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
args := append([]string{"--with-mail"}, os.Args[1:]...)
os.Exit(cmdbin.ExecFile("bin/kb/index", args))
}
+23
View File
@@ -0,0 +1,23 @@
//usr/bin/env go run -tags=system_ladybug "$0" "$@"; exit
//go:build cgo && system_ladybug
//
// bin/brain/search.go - deduction search over the 2dph brain.
//
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--json] [--no-web]
// ./bin/brain/search.go serve [port]
// ./bin/brain/search.go --list-model
//
// Needs CGO + libladybug (CGO_CFLAGS/CGO_LDFLAGS). Prefer the wrapper
// bin/kb/search which sets those and builds a binary for the embed daemon.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/brain"
)
func main() {
os.Exit(brain.Main(os.Args[1:]))
}
+31
View File
@@ -0,0 +1,31 @@
//usr/bin/env go run -tags=brain_serve,system_ladybug "$0" "$@"; exit
//go:build brain_serve && cgo && system_ladybug
//
// bin/brain/serve.go - HTTP API (in-process ladybug search).
//
// KB_ROOT=/path/to/2dph ./bin/brain/serve.go
// KB_WORKERS=4 KB_PORT=8630 ./bin/brain/serve.go
//
// Needs CGO + libladybug (same as bin/brain/search.go).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"log"
"os"
"github.com/eSlider/2dph/internal/brain"
"github.com/eSlider/2dph/internal/httpapi"
)
func main() {
if os.Getenv("KB_ROOT") == "" {
if wd, err := os.Getwd(); err == nil {
os.Setenv("KB_ROOT", wd)
}
}
if err := brain.Ready(); err != nil {
log.Fatal(err)
}
httpapi.Run(brain.HTTP{})
}
+20
View File
@@ -0,0 +1,20 @@
//go:build brain_serve && !system_ladybug
//
// Fallback serve when ladybug cgo is not in the build (CI / tags=brain_serve).
// Production shebang is serve.go (in-process).
package main
import (
"os"
"github.com/eSlider/2dph/internal/httpapi"
)
func main() {
if os.Getenv("KB_ROOT") == "" {
if wd, err := os.Getwd(); err == nil {
os.Setenv("KB_ROOT", wd)
}
}
httpapi.Run(nil)
}
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=brain_stats "$0" "$@"; exit
//go:build brain_stats
//
// bin/brain/stats.go - index health.
//
// ./bin/brain/stats.go
// ./bin/brain/stats.go --json
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/kb/stats", os.Args[1:]))
}
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=brain_watch "$0" "$@"; exit
//go:build brain_watch
//
// bin/brain/watch.go - re-index when corpus files change.
//
// ./bin/brain/watch.go [dir...]
// KB_WATCH_INTERVAL=15 ./bin/brain/watch.go
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/bin/watch"
)
func main() {
watch.Run(os.Args[1:])
}
Executable
+29
View File
@@ -0,0 +1,29 @@
#!/usr/bin/env bash
# bin/chats - sync, import, index, facts, apply for Telegram/WhatsApp/LinkedIn.
# Builds the chats binary on first run / when source changes, then execs it.
set -euo pipefail
ROOT="$(cd "$(dirname "$0")/.." && pwd)"
BIN="$ROOT/var/bin/chats"
SRC="$ROOT/bin/chats"
mkdir -p "$ROOT/var/bin"
need_build=0
if [ ! -x "$BIN" ]; then
need_build=1
else
while IFS= read -r -d '' f; do
if [ "$f" -nt "$BIN" ]; then
need_build=1
break
fi
done < <(find "$SRC" -name '*.go' -print0 2>/dev/null)
fi
if [ "$need_build" -eq 1 ]; then
echo "Building chats..." >&2
(cd "$SRC" && go build -o "$BIN" .) || exit 1
fi
exec "$BIN" "$@"
+19
View File
@@ -0,0 +1,19 @@
//usr/bin/env go run -tags=chats_apply "$0" "$@"; exit
//go:build chats_apply
//
// bin/chats/apply.go - push extracted chat facts to OnlyOffice CRM.
//
// ./bin/chats/apply.go [--dry-run]
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/chats"
)
func main() {
os.Exit(chats.RunApply(os.Args[1:]))
}
+4
View File
@@ -0,0 +1,4 @@
// Commands in this directory are shebang mains (sync.go, import.go, facts.go,
// apply.go), each behind an exclusive build tag so `go build ./bin/chats`
// does not see two mains. Shared code lives in internal/chats.
package main
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=chats_facts "$0" "$@"; exit
//go:build chats_facts
//
// bin/chats/facts.go - extract phone/email/linkedin facts from JSONL.
//
// ./bin/chats/facts.go
//
// Writes var/chats/facts/. Does not index the brain.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/chats"
)
func main() {
os.Exit(chats.RunFacts(os.Args[1:]))
}
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=chats_import "$0" "$@"; exit
//go:build chats_import
//
// bin/chats/import.go - JSONL → markdown under var/chats/md/.
//
// ./bin/chats/import.go
//
// Conversion only. Brain ingest is bin/brain/index.go, not this command.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/chats"
)
func main() {
os.Exit(chats.RunImport(os.Args[1:]))
}
+155
View File
@@ -0,0 +1,155 @@
#!/usr/bin/env python3
"""chats/refresh-linkedin-session - refresh LinkedIn MCP session from webtop CDP.
bin/chats/refresh-linkedin-session [--cdp URL] [--root DIR]
"""
Reads the current LinkedIn cookies out of the running Thorium browser in the
work-webtop container via CDP (Network.getAllCookies), copies the live browser
profile onto the source profile directory, and rewrites the portable
cookies.json + source-state.json that mcp-server-linkedin requires.
Usage:
refresh-linkedin-session [--cdp http://127.0.0.1:9222] [--root /var/tmp/liprofile]
[--container work-webtop] [--profile thorium-profile]
After the headless driver uses a copied profile, LinkedIn rotates the session
in that copy, so this must run before every sync.
"""
import asyncio
import json
import os
import shutil
import subprocess
import sys
import tempfile
import urllib.request
import websockets
def cdp_tab(ws_json):
for t in ws_json:
if t.get("webSocketDebuggerUrl"):
return t["webSocketDebuggerUrl"]
return None
async def get_cookies(ws_url):
async with websockets.connect(ws_url, max_size=50_000_000) as ws:
await ws.send(json.dumps({"id": 1, "method": "Network.getAllCookies", "params": {}}))
resp = await ws.recv()
return json.loads(resp).get("result", {}).get("cookies", [])
def write_source_state(root, profile_dir):
# Reuse the linkedin-mcp-server session_state module to write a valid
# source-state.json (same schema the daemon reads).
try:
from linkedin_mcp_server.session_state import canonical, write_source_state
write_source_state(canonical(__import__("pathlib").Path(profile_dir)))
return
except Exception:
pass
# Fallback: minimal schema-compatible state.
import uuid
state = {
"version": 1,
"source_runtime_id": "linux-amd64-host",
"login_generation": str(uuid.uuid4()),
"created_at": None,
"profile_path": profile_dir,
"cookies_path": os.path.join(root, "cookies.json"),
}
from datetime import datetime, timezone
state["created_at"] = datetime.now(timezone.utc).isoformat()
with open(os.path.join(root, "source-state.json"), "w") as f:
json.dump(state, f, indent=2)
def main():
args = sys.argv[1:]
cdp = "http://127.0.0.1:9222"
root = "/var/tmp/liprofile"
container = "work-webtop"
cprofile = "thorium-profile"
for i in range(0, len(args), 2):
k = args[i]
v = args[i + 1] if i + 1 < len(args) else ""
if k == "--cdp":
cdp = v
elif k == "--root":
root = v
elif k == "--container":
container = v
elif k == "--profile":
cprofile = v
profile_dir = os.path.join(root, "profile")
os.makedirs(profile_dir, exist_ok=True)
# 1. Clear stale daemon/browser locks so the server can claim the profile.
for lock in ("profile-claim.lock", "profile.lock", "daemon.lock", "lease.lock"):
p = os.path.join(root, lock)
if os.path.exists(p):
os.remove(p)
for name in os.listdir(profile_dir):
if name.startswith("Singleton"):
os.remove(os.path.join(profile_dir, name))
for name in os.listdir(root):
if name.startswith("invalid-state-"):
shutil.rmtree(os.path.join(root, name), ignore_errors=True)
# 1. Copy the live browser profile (cookies DB + Local State) so the
# session the driver launches carries the current login.
subprocess.run(
["docker", "cp", f"{container}:/config/{cprofile}/Default", os.path.join(profile_dir, "Default")],
check=True, capture_output=True,
)
subprocess.run(
["docker", "cp", f"{container}:/config/{cprofile}/Local State", os.path.join(profile_dir, "Local State")],
check=True, capture_output=True,
)
for lock in ("SingletonLock", "SingletonCookie", "SingletonSocket"):
p = os.path.join(profile_dir, lock)
if os.path.exists(p):
os.remove(p)
# 2. Pull the live cookies out of the running browser.
with urllib.request.urlopen(f"{cdp}/json", timeout=5) as r:
tabs = json.loads(r.read())
ws_url = cdp_tab(tabs)
if not ws_url:
sys.stderr.write("refresh-linkedin-session: no CDP tab\n")
sys.exit(1)
cookies = asyncio.run(get_cookies(ws_url))
li = [c for c in cookies if "linkedin" in c.get("domain", "")]
out = []
for c in li:
domain = c.get("domain", "")
if domain in (".www.linkedin.com", "www.linkedin.com"):
domain = ".linkedin.com"
out.append({
"name": c["name"],
"value": c["value"].strip('"'),
"domain": domain,
"path": c.get("path", "/"),
"expires": c.get("expires", -1),
"httpOnly": c.get("httpOnly", False),
"secure": c.get("secure", False),
"sameSite": c.get("sameSite", "None"),
})
with open(os.path.join(root, "cookies.json"), "w") as f:
json.dump(out, f, indent=2)
write_source_state(root, profile_dir)
sys.stderr.write(f"refresh-linkedin-session: {len(out)} cookies, profile refreshed\n")
if __name__ == "__main__":
main()
+41
View File
@@ -0,0 +1,41 @@
//usr/bin/env go run -tags=chats_sync "$0" "$@"; exit
//go:build chats_sync
//
// bin/chats/sync.go - download chat messages to var/chats/<platform>/.
//
// ./bin/chats/sync.go telegram [--limit N] [--phone PHONE]
// ./bin/chats/sync.go linkedin [--limit N] [--refresh]
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"fmt"
"os"
"github.com/eSlider/2dph/internal/chats"
)
func main() {
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
os.Exit(2)
}
platform := os.Args[1]
args := os.Args[2:]
switch platform {
case "telegram":
os.Exit(chats.RunSyncTelegram(args))
case "linkedin":
os.Exit(chats.RunSyncLinkedIn(args))
case "whatsapp":
fmt.Fprintln(os.Stderr, "chats: WhatsApp not implemented yet")
os.Exit(1)
case "help", "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
return
default:
fmt.Fprintf(os.Stderr, "chats: unknown platform %q\n", platform)
os.Exit(2)
}
}
+1 -1
View File
@@ -17,7 +17,7 @@ import subprocess
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[2] / "tools"))
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools"))
from semver import bump_type, bump_version # noqa: E402
+3 -1
View File
@@ -34,11 +34,13 @@ case "${1:-}" in
;;
"")
[ -f "$HOME/.ssh/config" ] || { echo "db/ssh-tunnel: ~/.ssh/config missing" >&2; exit 1; }
if db/ssh-tunnel --check; then
if "$0" --check; then
echo "tunnel already up on ${SRC}"
exit 0
fi
ssh -f -N -M -S "$HOME/.ssh/2dph-tunnel.sock" \
-o StrictHostKeyChecking=accept-new \
-o BatchMode=yes \
-L "${SRC}:${DST}" -p "$SSH_PORT" "${SSH_USER}@${SSH_HOST}" \
&& echo "tunnel up on ${SRC} (-> vm:${DST})"
exit 0
+2
View File
@@ -0,0 +1,2 @@
// Deprecated shebang mains at bin root (serve.go is tagged brain_serve).
package main
Regular → Executable
+12 -8
View File
@@ -2,10 +2,12 @@
# bin/docker-entrypoint - run 2dph tools inside the container.
#
# brain shell (default)
# brain search <q> bin/kb/search
# brain index bin/kb/index
# brain watch <dir> watchdog re-indexer
# brain serve async Go HTTP server (serve/)
# brain search <q> bin/brain/search.go
# brain index bin/kb/index --with-mail
# brain watch <dir> compiled /app/bin/watch (bin/brain/watch.go)
# brain serve compiled /app/bin/serve (bin/brain/serve.go)
# brain extract bin/facts/extract (docker×compose pairing)
# brain audit bin/facts/audit
#
# Usage comment starts at line 2 (self-describing convention).
set -euo pipefail
@@ -16,8 +18,10 @@ shift || true
case "$CMD" in
shell) exec bash ;;
search) exec "$KB_PY" /app/bin/kb/search "$@" ;;
index) exec "$KB_PY" /app/bin/kb/index "$@" ;;
watch) exec bash /app/bin/kb-watch "$@" ;;
serve) exec /app/serve/serve "$@" ;;
index) exec "$KB_PY" /app/bin/kb/index --with-mail "$@" ;;
watch) exec /app/bin/watch "$@" ;;
serve) exec /app/bin/serve "$@" ;;
extract) exec "$KB_PY" /app/bin/facts/extract "$@" ;;
audit) exec "$KB_PY" /app/bin/facts/audit "$@" ;;
*) echo "unknown command: $CMD" >&2; exit 2 ;;
esac
esac
+1 -1
View File
@@ -20,7 +20,7 @@ import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "tools"))
sys.path.insert(0, str(ROOT / "bin" / "tools"))
def audit_db() -> list[str]:
Executable
+124
View File
@@ -0,0 +1,124 @@
#!/usr/bin/env python3
"""facts/crm - prove person->company and company->project associations.
Two independent sources per fact:
S1 oo/OnlyOffice CRM (authoritative) : person.company_id -> company,
project.contacts -> company/person
S2 corpus SoT : eslider/cv/projects/knowledge-mesh-seed.yaml
(orgs: employer/client/... + projects)
Only associations supported by BOTH sources are written as root=facts.
Mismatches are reported (or, with --fix-crm, printed as oo CLI commands).
Usage:
bin/facts/crm write proven facts (needs var/kb.lbug)
bin/facts/crm --dry-run show proposed facts + mismatches only
bin/facts/crm --mismatches show associations found in only one side
"""
from __future__ import annotations
import json
import os
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import upsert_leaf, connect, leaf_id # noqa: E402
MESH_ENV = os.environ.get("KNOWLEDGE_MESH_SEED", "")
CORPUS_MESH = Path(MESH_ENV) if MESH_ENV else ROOT / "../knowledge-mesh-seed.yaml"
def corpus_orgs(raw: str) -> dict[str, dict]:
"""Delegate to tools.crmfacts.corpus_orgs (tested in tools/)."""
from crmfacts import corpus_orgs as _corpus_orgs
return _corpus_orgs(raw)
def main() -> int:
dry = "--dry-run" in sys.argv
mism = "--mismatches" in sys.argv
mesh = CORPUS_MESH.read_text()
orgs = corpus_orgs(mesh)
# CRM graph (produced by /tmp/opencode/crm/graph.py -> /tmp/opencode/crm/graph.json)
graph = json.load(open("/tmp/opencode/crm/graph.json"))
crm_person_company = graph["companies_with_persons"] # company -> [persons]
crm_project_companies = {} # pid -> title, companies
for pid, v in graph["projects_contacts"].items():
crm_project_companies[pid] = {"title": v["title"], "companies": v["companies"]}
facts: list[str] = []
mismatches: list[str] = []
# ---- person->company proven by CRM + corpus org ---- #
for org_name, org in orgs.items():
token = org.get("label", org_name)
# find CRM company whose name contains a significant token of the corpus org
key = next((k for k in crm_person_company
if token.split()[0].lower() in k.lower() or any(
t.lower() in k.lower() for t in org.get("label", "").split(" / "))),
None)
persons = crm_person_company.get(key, []) if key else []
if persons and org:
for p in persons:
facts.append(f"{p} is associated with {org.get('label')} "
f"(role: {org.get('kind', '?')}, {org.get('period', '')})")
elif org and key and not persons:
mismatches.append(f"corpus org '{org_name}' ({org.get('label')}) has no CRM persons")
elif org and not key:
mismatches.append(f"corpus org '{org_name}' ({org.get('label')}) not found in CRM")
# ---- corpus employer claims vs CRM ---- #
for org_name, org in orgs.items():
if not org or not org.get("kind"):
continue
if org["kind"] in ("employer", "own", "client", "agency", "apprenticeship"):
token = org.get("label", org_name).split()[0]
if not any(token.lower() in k.lower() for k in crm_person_company):
mismatches.append(f"corpus org '{org_name}' ({org['label']}) not found in CRM")
print(f"# CRM association facts proven (corpus x CRM): {len(facts)}")
for f in facts:
print(" -", f)
print(f"# mismatches / one-sided associations: {len(mismatches)}")
for f in mismatches:
print(" !", f)
if dry:
return 0
# ---- write proven facts into the brain (root=facts, 2 sources each) ---- #
import time
from model2vec import StaticModel
from kblib import MODEL # noqa: F401
model = StaticModel.from_pretrained(MODEL)
db, conn = connect(read_only=False)
try:
r = conn.execute("MATCH (l:Leaf) WHERE l.root='facts' RETURN count(*) AS n")
stats_before = r.get_all()[0][0]
except Exception:
stats_before = 0
rev = time.strftime("%Y%m%d-%H%M%S")
written = 0
for f in facts:
src = f"ooCRM x {CORPUS_MESH.name}"
lid = upsert_leaf(
conn,
text=f, root="facts", confidence="confirmed",
source=src, source_rev=rev,
how="crm-crosscheck", loc="bin/facts/crm", type_="association",
embedding=model.encode(f).tolist(),
)
written += 1
conn.close()
print(f"# wrote {written} facts into var/kb.lbug (facts was {stats_before})")
return 0
if __name__ == "__main__":
sys.exit(main())
+4 -3
View File
@@ -21,7 +21,7 @@ import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "tools"))
sys.path.insert(0, str(ROOT / "bin" / "tools"))
COMPOSE_FILES = [ROOT / "docker" / "compose.yaml", ROOT / "compose.yaml"]
DOC_MARKERS = ["README.md", "PLAN.md", "AGENTS.md"]
@@ -176,8 +176,7 @@ def dedupe(facts: list[dict]) -> list[dict]:
def write_facts(facts: list[dict]) -> None:
from kblib import connect, init_schema, upsert_leaf
from kblib import VAR
from kblib import VAR, connect, ensure_indexes, init_schema, upsert_leaf
VAR.mkdir(exist_ok=True)
db, conn = connect(VAR / "kb.lbug", read_only=False)
init_schema(conn)
@@ -188,6 +187,8 @@ def write_facts(facts: list[dict]) -> None:
upsert_leaf(conn, text=f["text"], root="facts", confidence="confirmed",
source=f["source"], source_rev=REPO, how=f["how"],
loc=f["loc"], type_="fact", embedding=emb)
# Upsert-with-index is safe; never DROP+recreate (ghost catalog kills HNSW).
ensure_indexes(conn)
conn.close()
db.close()
Executable
+26
View File
@@ -0,0 +1,26 @@
#!/usr/bin/env python3
"""git/import — deprecated. Use bin/git/import.go (go-git, no git binary).
bin/git/import.go [REPO] [--json] [--limit N] [--since DATE]
"""
from __future__ import annotations
import os
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
def main(argv: list[str]) -> int:
print(
"bin/git/import is deprecated; use bin/git/import.go (go-git)",
file=sys.stderr,
)
target = ROOT / "bin" / "git" / "import.go"
os.execvp("go", ["go", "run", str(target), *argv])
return 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+143
View File
@@ -0,0 +1,143 @@
//usr/bin/env go run "$0" "$@"; exit
//
// bin/git/import.go - read git history with go-git (no git binary).
//
// ./bin/git/import.go [REPO]
// ./bin/git/import.go --json
// ./bin/git/import.go --limit 100 --since 2026-01-01
// ./bin/git/import.go --root DIR
//
// Conversion only: prints commit leafs. Brain write is bin/brain/index.go.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"encoding/json"
"fmt"
"os"
"path/filepath"
"strconv"
"time"
"github.com/eSlider/2dph/internal/cmdbin"
"github.com/eSlider/2dph/internal/gitlog"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
var repo, root, since string
limit := 0
jsonOut := false
i := 0
for i < len(args) {
a := args[i]
switch {
case a == "--json":
jsonOut = true
case a == "--limit" && i+1 < len(args):
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 0 {
fmt.Fprintf(os.Stderr, "git/import: --limit must be a non-negative integer\n")
return 2
}
limit = n
case a == "--since" && i+1 < len(args):
i++
since = args[i]
case a == "--root" && i+1 < len(args):
i++
root = args[i]
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, `usage: bin/git/import.go [REPO] [--json] [--limit N] [--since DATE] [--root DIR]`)
return 0
case len(a) > 0 && a[0] != '-':
repo = a
default:
fmt.Fprintf(os.Stderr, "git/import: unknown flag %s\n", a)
return 2
}
i++
}
var sinceT time.Time
if since != "" {
var err error
sinceT, err = parseSince(since)
if err != nil {
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
return 2
}
}
repos := []string{}
if repo != "" {
repos = []string{repo}
} else if root != "" {
entries, err := os.ReadDir(root)
if err != nil {
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
return 1
}
for _, e := range entries {
p := filepath.Join(root, e.Name())
if _, err := os.Stat(filepath.Join(p, ".git")); err == nil {
repos = append(repos, p)
}
}
} else {
repos = []string{cmdbin.Root()}
}
opt := gitlog.Options{Limit: limit, Since: sinceT}
type row struct {
Repo string `json:"repo"`
Path string `json:"path"`
Commits int `json:"commits"`
Leafs []gitlog.Leaf `json:"leafs,omitempty"`
}
var rows []row
for _, p := range repos {
name, err := gitlog.RepoName(p)
if err != nil && name == "" {
fmt.Fprintf(os.Stderr, "git/import: %s: %v\n", p, err)
continue
}
cs, err := gitlog.Log(p, opt)
if err != nil {
fmt.Fprintf(os.Stderr, "git/import: %s: %v\n", p, err)
return 1
}
leafs := make([]gitlog.Leaf, 0, len(cs))
for _, c := range cs {
leafs = append(leafs, gitlog.ToLeaf(c, name))
}
rows = append(rows, row{Repo: name, Path: p, Commits: len(cs), Leafs: leafs})
}
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
if err := enc.Encode(rows); err != nil {
return 1
}
return 0
}
for _, r := range rows {
fmt.Printf("%-24s %5d commits %s\n", r.Repo, r.Commits, r.Path)
}
return 0
}
func parseSince(s string) (time.Time, error) {
for _, layout := range []string{time.RFC3339, "2006-01-02"} {
if t, err := time.Parse(layout, s); err == nil {
return t, nil
}
}
return time.Time{}, fmt.Errorf("cannot parse --since %q", s)
}
-27
View File
@@ -1,27 +0,0 @@
#!/usr/bin/env bash
# kb-watch - re-index 2dph when corpus files change.
#
# kb-watch [dir...] [interval_seconds]
#
# Polls mtimes (no inotify deps); cheap and reliable in containers. Defaults:
# dirs = /corpus (compose) or . ; interval = 30s.
set -euo pipefail
DEFAULT_DIRS="${KB_WATCH_DIRS:-/corpus}"
DIRS=("$@")
[[ ${#DIRS[@]} -eq 0 ]] && DIRS=(${DEFAULT_DIRS})
INTERVAL="${KB_WATCH_INTERVAL:-30}"
index() { "${KB_PY:-python3}" /app/bin/kb/index; }
LAST_STAMP=""
while true; do
STAMP=$(find "${DIRS[@]}" -type f -newermt "-${INTERVAL} seconds" 2>/dev/null \
| head -1 | md5sum)
if [[ -n "$STAMP" && "$STAMP" != "$LAST_STAMP" ]]; then
echo "kb-watch: changes detected, re-indexing" >&2
index || echo "kb-watch: index failed; will retry" >&2
LAST_STAMP="$STAMP"
fi
sleep "$INTERVAL"
done
+1 -1
View File
@@ -13,7 +13,7 @@ import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "tools"))
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import open_readonly, query_fts # noqa: E402
from yamlout import to_yaml # noqa: E402
+1 -1
View File
@@ -10,7 +10,7 @@ import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "tools"))
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import open_readonly # noqa: E402
from yamlout import to_yaml # noqa: E402
+33 -15
View File
@@ -19,13 +19,14 @@ import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "tools"))
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import ( # noqa: E402
connect, create_fts_and_vector, init_schema, upsert_leaf,
connect, ensure_indexes, init_schema, upsert_leaf,
open_readonly, stats,
)
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
from mailleafs import from_mail_root # noqa: E402
CORPUS_DEFAULTS = ["README.md", "PLAN.md", "AGENTS.md", "docs", "skills"]
@@ -99,43 +100,60 @@ def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description="build the 2dph brain index")
p.add_argument("--corpus", action="append", help="extra markdown dir/file to index (may repeat)")
p.add_argument("--rebuild", action="store_true", help="fresh db + indexes")
p.add_argument("--with-mail", action="store_true", help="include var/mail message.md leafs")
p.add_argument("--since", default="", help="with --with-mail, only messages dated >= YYYY-MM-DD")
p.add_argument("--dry-run", action="store_true", help="count leafs, write nothing")
p.add_argument(
"--skip-indexes",
action="store_true",
help="write leafs only; caller runs ensure_indexes after seeding facts",
)
p.add_argument("--limit", type=int, default=0, help="max leafs to embed")
p.add_argument("--json", action="store_true")
a = p.parse_args(argv)
from kblib import DB_PATH, VAR
VAR.mkdir(exist_ok=True)
if a.rebuild and DB_PATH.exists():
DB_PATH.unlink()
leafs = load_corpus(ROOT)
if a.corpus:
for source in a.corpus:
leafs.extend(load_corpus_glob(source))
mail_n = 0
if a.with_mail:
mail = from_mail_root(ROOT / "var" / "mail", since=a.since)
mail_n = len(mail)
leafs.extend(mail)
if a.dry_run:
msg = {"indexed": 0, "corpus_total": len(leafs), "mail_leafs": mail_n, "dry_run": True}
print(json.dumps(msg, indent=2) if a.json else
f"brain/index: {len(leafs)} leafs would be indexed (mail={mail_n})")
return 0
VAR.mkdir(exist_ok=True)
if a.rebuild and DB_PATH.exists():
DB_PATH.unlink()
db, conn = connect(DB_PATH, read_only=False)
init_schema(conn)
if not (a.rebuild or _already_indexed(conn)):
create_fts_and_vector(conn, force=True)
# Never DROP FTS/VECTOR (ghost catalog). Write leafs, then ensure indexes
# unless --skip-indexes (seed facts first — MERGE under live FTS corrupts it).
# --rebuild already deleted kb.lbug above, so CREATE runs on a clean DB.
embed = embedder()
done, total = index_leafs(conn, leafs, embed, a.limit)
create_fts_and_vector(conn, force=(done > 0 or a.rebuild))
if not a.skip_indexes:
ensure_indexes(conn)
s = stats(conn)
conn.close()
db.close()
result = {"indexed": done, "corpus_total": total, **{k: v for k, v in s.items() if k in ("total", "by_root")}}
if a.skip_indexes:
result["indexes"] = "skipped"
print(json.dumps(result, indent=2) if a.json else f"indexed {done}/{total} leafs; db total {s['total']}")
return 0
def _already_indexed(conn) -> bool:
try:
return conn.execute("MATCH (l:Leaf) RETURN count(*)").get_all()[0][0] > 0
except Exception:
return False
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+32 -67
View File
@@ -1,72 +1,37 @@
#!/usr/bin/env python3
"""kb/search - deduction search over the 2dph brain.
#!/usr/bin/env bash
# bin/kb/search deprecated wrapper. Use bin/brain/search.go.
# Sets CGO for ladybug, builds a binary (embed daemon needs a real executable),
# then execs it. Prints one deprecation line.
set -euo pipefail
bin/kb/search "query" # hybrid facts+info, YAML out
bin/kb/search "query" --root facts # confirmed facts only
bin/kb/search "query" --hop 1 # follow graph edges after hitting
bin/kb/search "query" --json | yq '.'
bin/kb/search "query" -n 5 # more results
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
BIN="$ROOT/var/bin/brain-search"
SRC="$ROOT/internal/brain"
CMD="$ROOT/bin/brain"
Deduction order: facts root first (confirmed answers with evidence links),
then info root (marked `(not confirmed)`). --root restricts to one root.
--hop N walks FROM_FILE edges (sibling leafs in the same source file).
"""
from __future__ import annotations
mkdir -p "$ROOT/var/bin"
import json
import sys
from pathlib import Path
need_build=0
if [ ! -x "$BIN" ]; then
need_build=1
else
while IFS= read -r -d '' f; do
if [ "$f" -nt "$BIN" ]; then
need_build=1
break
fi
done < <(find "$SRC" "$CMD" -name '*.go' -print0 2>/dev/null)
fi
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "tools"))
if [ "$need_build" -eq 1 ]; then
echo "Building brain/search..." >&2
(
cd "$ROOT" &&
CGO_CFLAGS="-I$ROOT/lib-ladybug" \
CGO_LDFLAGS="-L$ROOT/lib-ladybug -Wl,-rpath,$ROOT/lib-ladybug" \
go build -tags system_ladybug -o "$BIN" ./bin/brain
) || exit 1
fi
from kblib import connect, hybrid_search, init_schema, open_readonly, query_fts # noqa: E402
from yamlout import to_yaml # noqa: E402
import ladybug # noqa: E402
def main(argv: list[str]) -> int:
import argparse
p = argparse.ArgumentParser(description="deduction search over the brain")
p.add_argument("query")
p.add_argument("--root", choices=("facts", "info", None), default=None)
p.add_argument("--hop", type=int, default=0)
p.add_argument("-n", "--limit", type=int, default=10)
p.add_argument("--json", action="store_true")
a = p.parse_args(argv)
try:
db, conn = open_readonly()
except FileNotFoundError as e:
print(e, file=sys.stderr)
return 1
from model2vec import StaticModel
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
emb = model.encode([a.query])[0].astype(float).tolist()
rhs: list[dict] = []
try:
rhs = query_fts(conn, a.query, a.limit * 2)
except Exception:
rhs = []
results = hybrid_search(conn, emb, rhs, a.limit)
if a.root:
results = [h for h in results if h["root"] == a.root]
for hit in results:
hit.pop("rrf", None)
if hit.get("text"):
hit["snippet"] = hit["text"][:280]
out = {"query": a.query, "root_filter": a.root or "facts+info",
"count": len(results), "results": results}
print(json.dumps(out, indent=2, ensure_ascii=False) if a.json else to_yaml(out))
conn.close()
db.close()
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
echo "bin/kb/search is deprecated; use bin/brain/search.go" >&2
exec "$BIN" "$@"
+1 -1
View File
@@ -11,7 +11,7 @@ import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "tools"))
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import open_readonly, stats # noqa: E402
from yamlout import to_yaml # noqa: E402
+23
View File
@@ -0,0 +1,23 @@
//usr/bin/env go run "$0" "$@"; exit
// bin/kb/watch.go — deprecated. Use bin/brain/watch.go.
//
// Usage:
//
// ./bin/kb/watch.go [dir...] # dirs default /corpus
// KB_WATCH_INTERVAL=15 ./bin/kb/watch.go
//
// Shebang trick: first line is a Go `//` comment; the real code lives in the
// importable package (module path, never a relative import).
// NOTE: never run `gofmt -w` on this file - it rewrites `//usr/bin/env` to
// `// usr/...` and breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/bin/watch"
)
func main() {
watch.Run(os.Args[1:])
}
+450
View File
@@ -0,0 +1,450 @@
#!/usr/bin/env python3
"""mail/import - pull OnlyOffice mails into var/mail/ as markdown.
bin/mail/import --from-raw var/mail convert Go-synced message.json to md
bin/mail/import import newest inbox messages
bin/mail/import --folder sent import sent folder
bin/mail/import --since 2026-01-01 only messages after a date
bin/mail/import --limit 50 cap messages per run
bin/mail/import --no-attachments body only, skip attachment conversion
bin/mail/import --ocr OCR scanned PDFs/images via docling
bin/mail/import --dry-run list messages without writing anything
Writes one directory per message: var/mail/{folder}/{message_id}/
message.md frontmatter + markdown body
attachments/ raw attachment files (zips unpacked to _unpacked/)
attachments/*.md converted attachment content
Indexing is a separate step (`bin/brain/index.go --rebuild`): conversion can
crash in native docling and must not leave the brain DB mid-transaction.
Requires ONLYOFFICE_URL/USER/PASS in .env (or env). Idempotent: a message
already present (message.md exists) is skipped unless --force.
"""
from __future__ import annotations
import argparse
import json
import os
import re
import subprocess
import sys
import time
import urllib.parse
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from mailconv import ( # noqa: E402
ARCHIVE_SUFFIXES,
IMAGE_SUFFIXES,
LEGACY_OFFICE_SUFFIXES,
TEXT_SUFFIXES,
html_to_markdown,
is_convertible,
normalize_markdown,
subject_to_filename,
zip_extract_safe,
)
import requests # noqa: E402
FOLDER_IDS = {"inbox": 1, "sent": 2, "drafts": 3, "trash": 4, "spam": 5}
DEFAULT_LIMIT = 25
def load_env() -> dict:
env = {k: v for k, v in os.environ.items()}
envfile = ROOT / ".env"
if envfile.exists():
for line in envfile.read_text().splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
k, _, v = line.partition("=")
env.setdefault(k.strip(), v.strip().strip("\"'"))
url = env.get("ONLYOFFICE_URL") or env.get("OO_URL")
user = env.get("ONLYOFFICE_USER") or env.get("OO_USER")
password = env.get("ONLYOFFICE_PASS") or env.get("OO_PASSWORD")
missing = [n for n, v in (("ONLYOFFICE_URL", url), ("ONLYOFFICE_USER", user),
("ONLYOFFICE_PASS", password)) if not v]
if missing:
sys.exit(f"mail/import: missing {', '.join(missing)} (need .env or env)")
return {"url": url.rstrip("/"), "user": user, "password": password}
class OOClient:
def __init__(self, conf: dict):
self.base = conf["url"]
self.session = requests.Session()
self.token = None
self._login(conf)
def _login(self, conf: dict) -> None:
r = self.session.post(f"{self.base}/api/2.0/authentication.json",
json={"userName": conf["user"], "password": conf["password"], "type": 0},
timeout=30)
r.raise_for_status()
body = r.json()
self.token = (body.get("response") or {}).get("token", "")
if not self.token:
sys.exit("mail/import: authentication failed (empty token)")
def _headers(self) -> dict:
return {"Authorization": f"Bearer {self.token}", "Accept": "application/json"}
def get(self, path: str, params: dict | None = None):
r = self.session.get(f"{self.base}{path}", params=params, headers=self._headers(), timeout=30)
r.raise_for_status()
return r.json()
def list_messages(self, folder: int, page: int = 1, count: int = DEFAULT_LIMIT) -> list[dict]:
data = self.get("/api/2.0/mail/messages",
params={"folder": folder, "page": page, "count": count})
return data.get("response", [])
def get_message(self, message_id: str) -> dict:
data = self.get(f"/api/2.0/mail/messages/{message_id}")
return data.get("response", {})
def download_attachment(self, attach_id, dest: Path) -> bool:
"""Download one attachment via the portal session cookie (.ashx handler)."""
url = f"{self.base}/addons/mail/httphandlers/download.ashx?attachid={attach_id}"
r = self.session.get(url, timeout=60)
if r.status_code != 200:
return False
dest.parent.mkdir(parents=True, exist_ok=True)
dest.write_bytes(r.content)
return True
def folder_id(name: str) -> int:
if name in FOLDER_IDS:
return FOLDER_IDS[name]
if name.isdigit():
return int(name)
sys.exit(f"mail/import: unknown folder '{name}' (use {', '.join(FOLDER_IDS)})")
def safe_attachment_name(att: dict) -> str:
name = att.get("fileName") or att.get("storedName") or "attachment"
name = re.sub(r"[^\w.\- ]+", "_", name)
return name
def convert_file_to_md(path: Path, ocr: bool) -> str | None:
"""Convert one attachment file to markdown text; None when not convertible."""
suffix = path.suffix.lower()
if suffix in TEXT_SUFFIXES:
return normalize_markdown(path.read_text(encoding="utf-8", errors="replace"))
if suffix in (".docx", ".pptx", ".xlsx", ".html", ".htm", ".epub", ".eml", ".msg"):
try:
from markitdown import MarkItDown
md = MarkItDown()
result = md.convert(str(path))
return normalize_markdown(result.text_content)
except Exception as e:
return f"\n<!-- conversion failed: {e} -->\n"
if suffix == ".pdf":
return _convert_pdf(path, ocr)
if suffix in IMAGE_SUFFIXES and ocr:
return _convert_pdf(path, ocr)
if suffix in LEGACY_OFFICE_SUFFIXES:
return _convert_legacy(path)
if suffix in ARCHIVE_SUFFIXES:
return None # handled by caller (unpack + recurse)
return None
def _convert_pdf(path: Path, ocr: bool) -> str:
"""Convert one PDF to markdown.
Fast path: poppler's pdftotext (-layout) extracts exact text from
born-digital PDFs in ~15ms vs docling's 1-3s. Only textless PDFs (scanned
pages, layout-heavy) fall back to docling, which runs isolated in a
subprocess because its native onnx/RT-DETR has segfaulted the main process.
"""
text = _pdf_fast_text(path)
if ocr or text is None or not text.strip():
return _convert_pdf_docling(path, ocr)
return normalize_markdown(text)
def _pdf_fast_text(path: Path) -> str | None:
"""pdftotext -layout; None when poppler is unavailable (or the PDF has no text layer)."""
try:
proc = subprocess.run(
["pdftotext", "-layout", str(path), "-"],
capture_output=True, timeout=60)
except (OSError, subprocess.TimeoutExpired):
return None
if proc.returncode != 0:
return None
return proc.stdout.decode("utf-8", errors="replace")
def _convert_pdf_docling(path: Path, ocr: bool) -> str:
try:
proc = subprocess.run(
[sys.executable, os.path.abspath(__file__), "--pdf-worker", str(path),
"--ocr" if ocr else "--no-ocr"],
capture_output=True, text=True, timeout=600)
except subprocess.TimeoutExpired:
return "\n<!-- pdf conversion timed out -->\n"
if proc.returncode != 0:
tail = proc.stderr.strip().splitlines()[-3:]
return f"\n<!-- pdf conversion failed: {proc.returncode}: {' | '.join(tail)} -->\n"
return proc.stdout
def _pdf_worker(path: Path, ocr: bool) -> None:
"""docling worker entry: prints converted markdown on stdout, exits non-zero on error."""
try:
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.pipeline_options import PdfPipelineOptions
opts = PdfPipelineOptions()
opts.do_ocr = bool(ocr)
opts.do_table_structure = True
conv = DocumentConverter(format_options={"pdf": PdfFormatOption(pipeline_options=opts)})
res = conv.convert(str(path))
sys.stdout.write(normalize_markdown(res.document.export_to_markdown()))
sys.exit(0)
except Exception as e:
# errors/stacktraces to stderr; the caller only reports a one-liner
print(f"pdf-worker: {e}", file=sys.stderr)
import traceback
traceback.print_exc(file=sys.stderr)
sys.exit(1)
def _convert_legacy(path: Path) -> str:
"""Legacy .doc/.xls/.ppt -> md via pandoc (installed) or a stub."""
try:
out = subprocess.run(["pandoc", str(path), "-t", "markdown"],
capture_output=True, text=True, timeout=120)
if out.returncode == 0 and out.stdout.strip():
return normalize_markdown(out.stdout)
except (FileNotFoundError, subprocess.TimeoutExpired):
pass
return f"\n<!-- legacy {path.suffix} not convertible (pandoc unavailable) -->\n"
def write_message_md(msg: dict, folder: str, out_dir: Path, target_dir: Path | None = None) -> Path:
import yaml
body_html = msg.get("htmlBody") or ""
body_text = msg.get("textBody") or ""
body_md = ""
if body_html.strip():
body_md = html_to_markdown(body_html)
elif body_text.strip():
body_md = normalize_markdown(body_text)
# accept both OnlyOffice (receivedDate) and Go-sync (receivedAt) date keys
date = msg.get("receivedDate") or msg.get("receivedAt") or ""
if date and not isinstance(date, str):
date = str(date)
meta = {
"id": msg.get("id"),
"source": msg.get("source"),
"folder": folder,
"subject": msg.get("subject", ""),
"from": msg.get("from", ""),
"to": msg.get("to", ""),
"cc": msg.get("cc", ""),
"date": date,
"has_attachments": bool(msg.get("hasAttachments")),
"mime_message_id": msg.get("mimeMessageId", ""),
"calendar_uid": msg.get("calendarUid", ""),
"type": "mail",
}
meta = {k: v for k, v in meta.items() if v not in (None, "")}
frontmatter = "---\n" + yaml.safe_dump(meta, sort_keys=False, allow_unicode=True).strip() + "\n---\n"
content = f"{frontmatter}\n# {meta.get('subject','')}\n\n{body_md}".strip() + "\n"
if target_dir is not None:
msg_dir = target_dir
else:
msg_dir = out_dir / folder / str(meta.get("id"))
msg_dir.mkdir(parents=True, exist_ok=True)
md_path = msg_dir / "message.md"
md_path.write_text(content, encoding="utf-8")
return md_path
def convert_attachments(msg: dict, msg_dir: Path, ocr: bool) -> list[dict]:
"""Download + convert each attachment; returns [{name, md, raw}] summaries.
raw file keeps the API storedName (unique hash, avoids collisions); the
markdown is named after the friendly fileName when available.
In --from-raw mode attachments are already on disk (Go sync wrote them;
.ics already has a structured .md sidecar). Files with an existing .md
sidecar are left as-is, only unconverted raws are converted here.
"""
out: list[dict] = []
atts = msg.get("attachments") or []
att_dir = msg_dir / "attachments"
for att in atts:
aid = att.get("fileId")
display = safe_attachment_name(att)
stored = att.get("storedName")
raw_name = safe_attachment_name({"storedName": stored}) if stored else display
raw = att_dir / raw_name
if aid and not raw.exists() and OOCLIENT is not None:
if not OOCLIENT.download_attachment(aid, raw):
out.append({"name": display, "md": "\n<!-- download failed -->\n", "raw": str(raw)})
continue
if not raw.exists():
out.append({"name": display, "md": "\n<!-- raw missing -->\n", "raw": str(raw)})
continue
md_stem = Path(display).stem or raw.stem
md_path = att_dir / f"{md_stem}.md"
# Go sync pre-wrote structured .md for .ics; keep it.
if not md_path.exists():
md_text = _convert_att_recursive(raw, ocr)
md_path.write_text(f"# Attachment: {display}\n\n{md_text}\n", encoding="utf-8")
else:
md_text = md_path.read_text(encoding="utf-8", errors="replace")
out.append({"name": display, "md": md_text, "raw": str(raw), "md_file": str(md_path)})
return out
def _convert_att_recursive(path: Path, ocr: bool) -> str:
if path.suffix.lower() in ARCHIVE_SUFFIXES:
parts: list[str] = []
unpack = path.parent / "_unpacked" / path.stem
files = zip_extract_safe(path, unpack)
for f in files:
sub = _convert_att_recursive(f, ocr)
if sub and sub.strip():
parts.append(f"## {f.name}\n\n{sub}")
return "\n\n".join(parts) if parts else "\n<!-- empty zip -->\n"
text = convert_file_to_md(path, ocr)
return text or "\n<!-- not convertible -->\n"
# module-level client for attachment downloads in convert_attachments
OOCLIENT: OOClient | None = None
def convert_one(msg_dir: Path, full: dict, folder: str, out_root: Path,
ocr: bool, no_attachments: bool, target_dir: Path | None = None) -> dict:
"""Write message.md + convert attachments for one message dict.
Works for both live API messages and the Go-sync message.json shape
(source field optional; attachments read from attachments/ dir).
target_dir overrides the derived path (used by --from-raw where the
directory layout is authoritative, not the message folder field).
"""
mid = str(full.get("id"))
write_message_md(full, folder, out_root, target_dir=target_dir)
converted: list[dict] = []
if not no_attachments:
converted = convert_attachments(full, msg_dir, ocr)
return {"id": mid, "subject": full.get("subject", ""),
"date": full.get("receivedDate", "") or full.get("receivedAt", ""),
"attachments": len(converted)}
def main(argv: list[str]) -> int:
global OOCLIENT
p = argparse.ArgumentParser(description="pull OnlyOffice mails to var/mail as markdown")
p.add_argument("--folder", default="inbox", help="inbox|sent|drafts|trash|spam or numeric id")
p.add_argument("--limit", type=int, default=DEFAULT_LIMIT, help="max messages per run")
p.add_argument("--offset", type=int, default=0, help="skip N messages")
p.add_argument("--since", default="", help="only messages received after YYYY-MM-DD")
p.add_argument("--id", action="append", default=[], help="import specific message id (repeatable)")
p.add_argument("--from-raw", default="",
help="convert Go-synced dirs (var/mail/<folder>/<id>/message.json) to markdown")
p.add_argument("--no-attachments", action="store_true", help="skip attachment download+convert")
p.add_argument("--ocr", action="store_true", help="OCR scanned PDFs/images via docling")
p.add_argument("--force", action="store_true", help="re-import even if message.md exists")
p.add_argument("--dry-run", action="store_true", help="list messages, write nothing")
p.add_argument("--json", action="store_true")
p.add_argument("--pdf-worker", default="", help=argparse.SUPPRESS)
p.add_argument("--no-ocr", action="store_true", help=argparse.SUPPRESS)
a = p.parse_args(argv)
if a.pdf_worker:
_pdf_worker(Path(a.pdf_worker), ocr=not a.no_ocr)
return 0
conf = load_env()
fid = folder_id(a.folder)
out_root = ROOT / "var" / "mail"
summary: list[dict] = []
if a.from_raw:
OOCLIENT = None
raw_root = Path(a.from_raw)
for msg_dir in sorted(raw_root.rglob("message.json")):
mid = msg_dir.parent.name
entry = {"id": mid, "subject": "", "date": "",
"attachments": 0, "skipped": False}
md_path = msg_dir.parent / "message.md"
if md_path.exists() and not a.force:
entry["skipped"] = True
summary.append(entry)
continue
if a.dry_run:
entry["skipped"] = "dry-run"
summary.append(entry)
continue
full = json.loads(msg_dir.read_text(encoding="utf-8"))
entry.update(convert_one(msg_dir.parent, full, full.get("folder") or a.folder,
raw_root, a.ocr, a.no_attachments,
target_dir=msg_dir.parent))
summary.append(entry)
else:
OOCLIENT = OOClient(conf)
if a.id:
messages = [{"id": i} for i in a.id]
else:
page = 1
messages = []
want = a.offset + a.limit
while len(messages) < want:
count = min(DEFAULT_LIMIT, want - len(messages))
chunk = OOCLIENT.list_messages(fid, page=page, count=count)
if not chunk:
break
messages.extend(chunk)
if len(chunk) < count:
break
page += 1
messages = messages[a.offset:a.offset + a.limit]
if a.since:
messages = [m for m in messages
if (m.get("receivedDate") or "") >= a.since]
for m in messages:
mid = str(m.get("id"))
entry = {"id": mid, "subject": m.get("subject", ""),
"date": m.get("receivedDate", ""), "attachments": 0, "skipped": False}
msg_dir = out_root / a.folder / mid
md_path = msg_dir / "message.md"
if md_path.exists() and not a.force:
entry["skipped"] = True
summary.append(entry)
continue
if a.dry_run:
entry["skipped"] = "dry-run"
summary.append(entry)
continue
full = OOCLIENT.get_message(mid)
entry.update(convert_one(msg_dir, full, a.folder, out_root, a.ocr, a.no_attachments,
target_dir=msg_dir))
summary.append(entry)
if a.json:
print(json.dumps(summary, ensure_ascii=False, indent=2))
else:
imported = [e for e in summary if not e["skipped"]]
print(f"mail/import: folder={a.folder} checked={len(summary)} "
f"imported={len(imported)} (skipped={sum(e['skipped'] is True for e in summary)})")
for e in summary:
flag = "skip" if e["skipped"] is True else ("dry" if e["skipped"] == "dry-run" else "ok ")
print(f" [{flag}] {e['id']} {e['date'][:10]} {e['subject'][:60]}"
f" (atts={e['attachments']})")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=mail_import "$0" "$@"; exit
//go:build mail_import
//
// bin/mail/import.go - message.json → markdown (no brain write).
//
// ./bin/mail/import.go --from-raw var/mail
//
// Indexing is bin/brain/index.go --rebuild, not this command.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/mail/import", os.Args[1:]))
}
+27
View File
@@ -0,0 +1,27 @@
#!/usr/bin/env python3
"""mail/index_mail — deprecated. Use bin/brain/index.go --rebuild --with-mail.
Ladybug corrupts its WAL on bulk-insert into an already-indexed DB, so this
shim always rebuilds (repo corpus + var/mail). Conversion stays in mail/import.
"""
from __future__ import annotations
import os
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
def main(argv: list[str]) -> int:
print(
"bin/mail/index_mail is deprecated; use bin/brain/index.go --rebuild --with-mail",
file=sys.stderr,
)
index = ROOT / "bin" / "kb" / "index"
os.execv(sys.executable, [sys.executable, str(index), "--rebuild", "--with-mail", *argv])
return 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+25
View File
@@ -0,0 +1,25 @@
//usr/bin/env go run "$0" "$@"; exit
// bin/mail/sync.go - async download of OnlyOffice and Gmail mail to var/mail/.
//
// ./bin/mail/sync.go --source onlyoffice,gmail --limit 50 --workers 8
// ./bin/mail/sync.go --source gmail --force
// ./bin/mail/sync.go --dry-run
//
// Writes raw message.json + attachments under var/mail/<folder>/<id>/; run
// bin/mail/import.go --from-raw afterwards to convert everything to markdown.
//
// Shebang trick: first line is a Go `//` comment; the real code lives in the
// importable package (module path, never a relative import).
// NOTE: never run `gofmt -w` on this file - it rewrites `//usr/bin/env` to
// `// usr/...` and breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/bin/mail/sync"
)
func main() {
os.Exit(sync.Main(os.Args[1:]))
}
+156
View File
@@ -0,0 +1,156 @@
// Package synccmd wires the sync library to a CLI: reads .env, parses flags,
// picks sources, prints stats. Kept separate from the library so unit tests
// don't depend on os.Args/env.
package sync
import (
"context"
"flag"
"fmt"
"os"
"path/filepath"
"strings"
"time"
)
// CLIConfig is a superset of SyncConfig plus flag parsing results.
type CLIConfig struct {
Sync SyncConfig
Env string // .env path; default <cwd>/.env
Sources string
Help bool
}
// ParseCLI reads os.Args into a CLIConfig. Exit codes: 0 ok, 2 usage.
func ParseCLI(args []string) (CLIConfig, int, error) {
fs := flag.NewFlagSet("mail/sync", flag.ContinueOnError)
var (
env = fs.String("env", "", ".env file (default: <cwd>/.env)")
out = fs.String("out", "", "var/mail root (default: <cwd>/var/mail)")
workers = fs.Int("workers", 4, "concurrent downloads")
limit = fs.Int("limit", 0, "max messages per source (0 = all)")
offset = fs.Int("offset", 0, "skip first N messages per source")
force = fs.Bool("force", false, "overwrite existing message.json + attachments")
dryRun = fs.Bool("dry-run", false, "list message counts without writing")
query = fs.String("query", "in:inbox", "Gmail search query (gmail source only)")
srcs = fs.String("source", "onlyoffice", "comma list: onlyoffice,gmail (default onlyoffice)")
help = fs.Bool("help", false, "usage")
)
fs.SetOutput(os.Stderr)
if err := fs.Parse(args); err != nil {
return CLIConfig{}, 2, err
}
if *help || fs.NArg() > 0 {
return CLIConfig{Help: true}, 0, nil
}
wd, err := os.Getwd()
if err != nil {
return CLIConfig{}, 2, err
}
if *env == "" {
*env = filepath.Join(wd, ".env")
}
if *out == "" {
*out = filepath.Join(wd, "var", "mail")
}
envVars := readEnv(*env)
cfg := SyncConfig{
Out: *out,
Workers: *workers,
Limit: *limit,
Offset: *offset,
Force: *force,
DryRun: *dryRun,
Query: *query,
Policy: RetryPolicy{},
}
cli := CLIConfig{Sync: cfg, Env: *env, Sources: *srcs}
for _, s := range strings.Split(*srcs, ",") {
switch strings.TrimSpace(s) {
case "onlyoffice":
u := pick(envVars["ONLYOFFICE_URL"], envVars["OO_URL"])
user := pick(envVars["ONLYOFFICE_USER"], envVars["OO_USER"])
pass := pick(envVars["ONLYOFFICE_PASS"], envVars["OO_PASSWORD"])
if u == "" || user == "" || pass == "" {
return CLIConfig{}, 2, fmt.Errorf("onlyoffice source needs ONLYOFFICE_URL/USER/PASS in %s", *env)
}
cfg.OO = &OOConfig{URL: u, User: user, Password: pass}
case "gmail":
home, _ := os.UserHomeDir()
cfg.Gmail = &GmailCredentials{
CredentialsPath: filepath.Join(home, ".gmail-mcp", "credentials.json"),
KeysPath: filepath.Join(home, ".gmail-mcp", "gcp-oauth.keys.json"),
}
default:
return CLIConfig{}, 2, fmt.Errorf("unknown source %q", s)
}
}
cli.Sync = cfg
return cli, 0, nil
}
// Main is the CLI entry: returns process exit code.
func Main(args []string) int {
cli, code, err := ParseCLI(args)
if err != nil {
fmt.Fprintln(os.Stderr, "mail/sync:", err)
return code
}
if cli.Help {
fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail] [--query GMAIL_Q] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]")
return 0
}
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
defer cancel()
start := time.Now()
stats, err := Run(ctx, cli.Sync)
if err != nil {
fmt.Fprintln(os.Stderr, "mail/sync:", err)
return 1
}
if cli.Sync.DryRun {
fmt.Printf("mail/sync: dry-run checked=%d (no writes)\n", stats.Checked)
return 0
}
fmt.Printf("mail/sync: checked=%d new=%d skipped=%d failed=%d in %s\n",
stats.Checked, stats.New, stats.Skipped, stats.Failed, time.Since(start).Round(time.Millisecond))
if stats.Failed > 0 {
return 1
}
return 0
}
// readEnv parses KEY=VALUE lines (ignoring comments) with KEY=PATH override.
func readEnv(path string) map[string]string {
out := map[string]string{}
b, err := os.ReadFile(path)
if err != nil {
return out
}
for _, line := range strings.Split(string(b), "\n") {
line = strings.TrimSpace(line)
if line == "" || strings.HasPrefix(line, "#") || !strings.Contains(line, "=") {
continue
}
k, v, _ := strings.Cut(line, "=")
out[strings.TrimSpace(k)] = strings.Trim(strings.TrimSpace(v), "\"'")
}
// env overrides file
for _, kv := range os.Environ() {
k, v, ok := strings.Cut(kv, "=")
if !ok {
continue
}
if strings.HasPrefix(k, "ONLYOFFICE_") || strings.HasPrefix(k, "OO_") {
out[k] = v
}
}
return out
}
func pick(a, b string) string {
if a != "" {
return a
}
return b
}
+356
View File
@@ -0,0 +1,356 @@
package sync
import (
"bytes"
"context"
"encoding/base64"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"path/filepath"
"strings"
"time"
)
// GmailCredentials holds the OAuth files produced by the gmail MCP
// (@gongrzhe/server-gmail-autoauth-mcp) auto-auth flow.
type GmailCredentials struct {
CredentialsPath string // ~/.gmail-mcp/credentials.json
KeysPath string // ~/.gmail-mcp/gcp-oauth.keys.json
User string // fixed: the authed account
}
// gmailToken is the JSON shape of credentials.json + refresh response.
type gmailToken struct {
AccessToken string `json:"access_token"`
RefreshToken string `json:"refresh_token"`
Expiry int64 `json:"expiry_date"` // ms epoch
}
type gmailKeys struct {
Installed *gmailKeyBlock `json:"installed"`
Web *gmailKeyBlock `json:"web"`
}
type gmailKeyBlock struct {
ClientID string `json:"client_id"`
ClientSecret string `json:"client_secret"`
}
// GmailClient talks to the Gmail REST API using the OAuth refresh token from
// ~/.gmail-mcp/. Token is refreshed lazily with a mutex-guarded cache.
type GmailClient struct {
creds GmailCredentials
client *http.Client
mu chan struct{}
token *gmailToken
user string
}
func NewGmailClient(creds GmailCredentials) (*GmailClient, error) {
if creds.CredentialsPath == "" {
home, _ := os.UserHomeDir()
creds.CredentialsPath = filepath.Join(home, ".gmail-mcp", "credentials.json")
creds.KeysPath = filepath.Join(home, ".gmail-mcp", "gcp-oauth.keys.json")
}
g := &GmailClient{
creds: creds,
client: &http.Client{Timeout: 60 * time.Second},
mu: make(chan struct{}, 1),
}
g.mu <- struct{}{}
return g, nil
}
// accessToken returns a fresh bearer token, refreshing via the Google token
// endpoint when the cached one is missing or about to expire.
func (g *GmailClient) accessToken(ctx context.Context) (string, error) {
select {
case <-g.mu:
case <-ctx.Done():
return "", ctx.Err()
}
defer func() { g.mu <- struct{}{} }()
if g.token != nil && g.token.AccessToken != "" && g.token.Expiry > time.Now().UnixMilli()+300_000 {
return g.token.AccessToken, nil
}
return g.refreshLocked(ctx)
}
func (g *GmailClient) refreshLocked(ctx context.Context) (string, error) {
cred, err := os.ReadFile(g.creds.CredentialsPath)
if err != nil {
return "", fmt.Errorf("read gmail credentials %s: %w", g.creds.CredentialsPath, err)
}
var t gmailToken
if err := json.Unmarshal(cred, &t); err != nil {
return "", fmt.Errorf("parse gmail credentials: %w", err)
}
if t.RefreshToken == "" {
return "", errors.New("gmail credentials.json has no refresh_token (run the gmail MCP auth flow)")
}
keys, err := os.ReadFile(g.creds.KeysPath)
if err != nil {
return "", fmt.Errorf("read gmail keys %s: %w", g.creds.KeysPath, err)
}
var k gmailKeys
if err := json.Unmarshal(keys, &k); err != nil {
return "", fmt.Errorf("parse gmail keys: %w", err)
}
block := k.Installed
if block == nil {
block = k.Web
}
if block == nil {
return "", errors.New("gmail gcp-oauth.keys.json has no installed/web block")
}
form := url.Values{}
form.Set("client_id", block.ClientID)
form.Set("client_secret", block.ClientSecret)
form.Set("refresh_token", t.RefreshToken)
form.Set("grant_type", "refresh_token")
req, err := http.NewRequestWithContext(ctx, http.MethodPost, "https://oauth2.googleapis.com/token",
strings.NewReader(form.Encode()))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
resp, err := g.client.Do(req)
if err != nil {
return "", fmt.Errorf("gmail token refresh: %w", err)
}
defer resp.Body.Close()
body, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
if resp.StatusCode != http.StatusOK {
var e struct {
Error string `json:"error"`
Desc string `json:"error_description"`
}
_ = json.Unmarshal(body, &e)
if e.Error == "invalid_grant" {
return "", fmt.Errorf("gmail OAuth token invalid/expired - re-auth via: npx -y @gongrzhe/server-gmail-autoauth-mcp auth (uses ~/.gmail-mcp)")
}
return "", fmt.Errorf("gmail token refresh status %d: %s", resp.StatusCode, truncate(string(body), 300))
}
var out struct {
AccessToken string `json:"access_token"`
ExpiresIn int64 `json:"expires_in"`
}
if err := json.Unmarshal(body, &out); err != nil {
return "", fmt.Errorf("gmail token refresh parse: %w", err)
}
g.token = &gmailToken{
AccessToken: out.AccessToken,
RefreshToken: t.RefreshToken,
Expiry: time.Now().UnixMilli() + out.ExpiresIn*1000,
}
return out.AccessToken, nil
}
// ListIDs returns message ids matching q, walking nextPageToken up to maxIDs
// (0 = unlimited). Thread-level pagination via the messages.list endpoint.
func (g *GmailClient) ListIDs(ctx context.Context, q string, maxIDs int, pageToken string) (ids []string, next string, err error) {
for {
params := url.Values{}
params.Set("q", q)
params.Set("maxResults", "100")
if pageToken != "" {
params.Set("pageToken", pageToken)
}
var out struct {
Messages []struct {
ID string `json:"id"`
} `json:"messages"`
NextPageToken string `json:"nextPageToken"`
}
if err := g.getJSON(ctx, "/gmail/v1/users/me/messages?"+params.Encode(), &out); err != nil {
return nil, "", err
}
for _, m := range out.Messages {
ids = append(ids, m.ID)
if maxIDs > 0 && len(ids) >= maxIDs {
return ids, out.NextPageToken, nil
}
}
if out.NextPageToken == "" {
break
}
pageToken = out.NextPageToken
}
return ids, "", nil
}
// GetMessage fetches a message in format=full and normalizes it.
func (g *GmailClient) GetMessage(ctx context.Context, id string) (*Message, error) {
var raw struct {
ID string `json:"id"`
ThreadID string `json:"threadId"`
InternalDate string `json:"internalDate"` // ms epoch string
Payload gmailPart
}
path := "/gmail/v1/users/me/messages/" + url.PathEscape(id) + "?format=full"
if err := g.getJSON(ctx, path, &raw); err != nil {
return nil, err
}
m := &Message{
Source: "gmail",
ID: raw.ID,
Folder: "gmail",
}
for _, h := range raw.Payload.Headers {
switch strings.ToLower(h.Name) {
case "subject":
m.Subject = h.Value
case "from":
m.From = h.Value
case "to":
m.To = h.Value
case "cc":
m.CC = h.Value
case "bcc":
m.BCC = h.Value
case "message-id":
m.MimeMessageID = h.Value
case "date":
if t, err := time.Parse(time.RFC1123Z, h.Value); err == nil {
m.ReceivedAt = t
}
}
}
if ms, err := parseMS(raw.InternalDate); err == nil {
m.ReceivedAt = ms
}
m.TextBody, m.HTMLBody, m.Attachments = collectParts(raw.Payload, "root", m.ID, 0)
m.HasAttachments = len(m.Attachments) > 0
return m, nil
}
type gmailPart struct {
PartID string `json:"partId"`
MimeType string `json:"mimeType"`
Filename string `json:"filename"`
Body gmailBody `json:"body"`
Headers []gmailHeader `json:"headers"`
Parts []gmailPart `json:"parts"`
}
type gmailHeader struct {
Name string `json:"name"`
Value string `json:"value"`
}
type gmailBody struct {
Size int64 `json:"size"`
Data string `json:"data"`
AttachmentID string `json:"attachmentId"`
}
// collectParts walks the MIME tree: text bodies into plain/html, anything with
// a filename into attachments (returned with base64 ids for later download).
func collectParts(p gmailPart, mime string, msgID string, depth int) (text, html string, atts []Attachment) {
if depth > 16 {
return
}
mt := strings.ToLower(p.MimeType)
if p.Filename != "" && mt != "text/plain" && mt != "text/html" {
pid := p.PartID
if pid == "" {
pid = fmt.Sprintf("%d", depth)
}
// Gmail's attachments API keys off body.attachmentId, not partId.
attID := p.Body.AttachmentID
if attID == "" {
attID = pid
}
atts = append(atts, Attachment{
FileID: msgID + ":" + attID,
FileName: p.Filename,
StoredName: p.Filename,
Size: p.Body.Size,
ContentType: p.MimeType,
})
} else if data, err := base64.URLEncoding.DecodeString(p.Body.Data); err == nil && len(p.Body.Data) > 0 {
s := string(data)
if mt == "text/html" && html == "" {
html = s
} else if (mt == "text/plain" || mt == "") && text == "" {
text = s
}
}
for _, child := range p.Parts {
t, h, a := collectParts(child, mt, msgID, depth+1)
if text == "" {
text = t
}
if html == "" {
html = h
}
atts = append(atts, a...)
}
return
}
// DownloadAttachment fetches an attachment's bytes from the Gmail API.
func (g *GmailClient) DownloadAttachment(ctx context.Context, msgID, attID string) ([]byte, error) {
// attID format is "<msgId>:<partId>"; the API needs the bare attachment id.
partID := attID
if i := strings.Index(attID, ":"); i >= 0 {
partID = attID[i+1:]
}
var out struct {
Data string `json:"data"`
}
path := "/gmail/v1/users/me/messages/" + url.PathEscape(msgID) + "/attachments/" + url.PathEscape(partID)
if err := g.getJSON(ctx, path, &out); err != nil {
return nil, err
}
return base64.URLEncoding.DecodeString(out.Data)
}
func (g *GmailClient) getJSON(ctx context.Context, path string, out any) error {
tok, err := g.accessToken(ctx)
if err != nil {
return err
}
u := "https://gmail.googleapis.com" + path
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+tok)
resp, err := g.client.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
body, _ := io.ReadAll(io.LimitReader(resp.Body, 16<<20))
if resp.StatusCode != http.StatusOK {
return fmt.Errorf("gmail %s: status %d: %s", path, resp.StatusCode, truncate(string(body), 300))
}
if out != nil {
return json.Unmarshal(body, out)
}
return nil
}
func parseMS(s string) (time.Time, error) {
if s == "" {
return time.Time{}, errors.New("empty")
}
var ms int64
if _, err := fmt.Sscanf(s, "%d", &ms); err != nil {
return time.Time{}, err
}
return time.UnixMilli(ms), nil
}
func truncate(s string, n int) string {
if len(s) <= n {
return s
}
return s[:n] + "…"
}
var _ = bytes.MinRead
+258
View File
@@ -0,0 +1,258 @@
package sync
import (
"fmt"
"strings"
"time"
"unicode/utf8"
ics "github.com/arran4/golang-ical"
"golang.org/x/text/encoding/charmap"
)
// ICSToMarkdown parses a VCALENDAR/VEVENT payload and renders a compact
// structured markdown block: what / when / where / organizer / attendees.
// Returns the raw text when the payload is not a calendar.
func ICSToMarkdown(data []byte) string {
data = normalizeEncoding(data)
cal, err := ics.ParseCalendar(strings.NewReader(string(data)))
if err != nil {
return normalizeMarkdown(string(data))
}
method := ""
for _, p := range cal.CalendarProperties {
if p.IANAToken == string(ics.ComponentPropertyMethod) {
method = p.Value
break
}
}
method = strings.TrimSpace(method)
var out []string
for _, ev := range cal.Events() {
summary := strings.TrimSpace(propValue(ev, ics.ComponentPropertySummary))
if summary != "" {
out = append(out, "# "+summary)
}
if when := eventWhen(ev); when != "" {
out = append(out, "- **When:** "+when)
}
if loc := strings.TrimSpace(propValue(ev, ics.ComponentPropertyLocation)); loc != "" {
out = append(out, "- **Where:** "+loc)
}
if desc := strings.TrimSpace(stripHTML(propValue(ev, ics.ComponentPropertyDescription))); desc != "" {
out = append(out, "- **What:** "+desc)
}
if org := propValue(ev, ics.ComponentPropertyOrganizer); org != "" {
out = append(out, "- **Organizer:** "+attendeeFmt(org))
}
for _, a := range ev.Attendees() {
cn := strings.TrimSpace(firstParam(a.ICalParameters, "CN"))
partstat := string(a.ParticipationStatus())
name := cn
if name == "" {
name = a.Email()
}
line := name
if email := a.Email(); email != "" && email != name {
line = name + " <" + email + ">"
}
if partstat != "" && !strings.EqualFold(partstat, "NEEDS-ACTION") {
line += " (" + strings.Title(strings.ToLower(strings.ReplaceAll(partstat, "_", " "))) + ")"
}
out = append(out, "- **Attendee:** "+line)
}
}
if len(out) == 0 {
return normalizeMarkdown(string(data))
}
if method != "" {
out = append([]string{"*Calendar method: " + method + "*"}, out...)
}
return normalizeMarkdown(strings.Join(out, "\n\n"))
}
func eventWhen(ev *ics.VEvent) string {
start, errStart := ev.GetStartAt()
end, errEnd := ev.GetEndAt()
// All-day events: golang-ical has dedicated getters.
if errStart != nil {
if allDay, err := ev.GetAllDayStartAt(); err == nil {
start = allDay
errStart = nil
}
}
if errEnd != nil {
if allDay, err := ev.GetAllDayEndAt(); err == nil {
end = allDay
errEnd = nil
}
}
if errStart != nil {
// Non-IANA TZID (e.g. "W. Europe Standard Time"): parse the raw
// property text instead of failing.
return rawWhen(ev)
}
if errEnd != nil || end.Equal(start) {
return dtFmt(start)
}
return dtFmt(start) + " → " + dtFmt(end)
}
// rawWhen parses DTSTART/DTEND property values that golang-ical cannot resolve
// because the TZID is not an IANA zone. Formats: 20260812T120000 or 20260812.
func rawWhen(ev *ics.VEvent) string {
start := rawPropValue(ev, ics.ComponentPropertyDtStart)
end := rawPropValue(ev, ics.ComponentPropertyDtEnd)
if start == "" {
return ""
}
if end == "" || end == start {
return rawDTFmt(start)
}
return rawDTFmt(start) + " → " + rawDTFmt(end)
}
func rawPropValue(ev *ics.VEvent, prop ics.ComponentProperty) string {
p := ev.GetProperty(prop)
if p == nil {
return ""
}
return p.Value
}
// rawDTFmt turns 20260812T120000 into 2026-08-12 12:00; 20260812 into 2026-08-12.
func rawDTFmt(s string) string {
s = strings.TrimSpace(s)
if len(s) >= 8 && isDigits(s[:8]) {
y, m, d := s[:4], s[4:6], s[6:8]
if len(s) > 8 && (s[8] == 'T' || s[8] == 't') && len(s) >= 15 && isDigits(s[9:15]) {
h, mi := s[9:11], s[11:13]
return fmt.Sprintf("%s-%s-%s %s:%s", y, m, d, h, mi)
}
return fmt.Sprintf("%s-%s-%s", y, m, d)
}
return s
}
func isDigits(s string) bool {
for _, c := range s {
if c < '0' || c > '9' {
return false
}
}
return s != ""
}
// dtFmt renders a time as local "2006-01-02 15:04" (tz label when meaningful).
func dtFmt(t time.Time) string {
loc := t.Local()
label := ""
if loc.Location() != time.Local {
label = " " + loc.Location().String()
}
return loc.Format("2006-01-02 15:04") + label
}
// propertyGetter is satisfied by both *ics.Calendar and *ics.VEvent.
type propertyGetter interface {
GetProperty(ics.ComponentProperty) *ics.IANAProperty
}
func propValue(ev propertyGetter, prop ics.ComponentProperty) string {
p := ev.GetProperty(prop)
if p == nil {
return ""
}
return p.Value
}
func firstParam(params map[string][]string, key string) string {
if vs, ok := params[key]; ok && len(vs) > 0 {
return vs[0]
}
return ""
}
func attendeeFmt(raw string) string {
raw = strings.TrimSpace(raw)
if i := strings.Index(raw, ":"); i >= 0 {
raw = raw[i+1:]
}
return raw
}
// stripHTML removes tags and decodes entities from an ics DESCRIPTION that may
// carry HTML (Outlook/Exchange style), keeping text lines readable.
func stripHTML(s string) string {
if !strings.Contains(s, "<") {
return s
}
lines := strings.Split(s, "\n")
for i, l := range lines {
var b strings.Builder
depth := 0
for j := 0; j < len(l); j++ {
c := l[j]
if c == '<' {
if j+1 < len(l) && l[j+1] == '/' {
depth--
} else {
depth++
}
for j < len(l) && l[j] != '>' {
j++
}
continue
}
if c == '>' {
continue
}
if depth == 0 {
b.WriteByte(c)
}
}
lines[i] = strings.TrimSpace(b.String())
}
return strings.Join(lines, "\n")
}
// normalizeMarkdown collapses blank-line runs and strips control chars.
func normalizeMarkdown(s string) string {
s = strings.ReplaceAll(s, "\x00", "")
for _, ch := range []string{"\ufeff", "\u200b", "\u034f", "\u00ad", "\u2007", "\u2008", "\u200a", "\u2002"} {
s = strings.ReplaceAll(s, ch, "")
}
lines := strings.Split(s, "\n")
var out []string
blank := 0
for _, l := range lines {
if strings.TrimSpace(l) == "" {
blank++
if blank > 1 {
continue
}
} else {
blank = 0
}
out = append(out, l)
}
return strings.Join(out, "\n")
}
// normalizeEncoding re-encodes legacy single-byte text as UTF-8. ICS files
// exported by some portals are Latin-1 (e.g. "N\xfcrnberg"); golang-ical
// passes the bytes through, producing invalid UTF-8 in the markdown output.
// Valid UTF-8 is returned untouched.
func normalizeEncoding(data []byte) []byte {
if utf8.Valid(data) {
return data
}
dec := charmap.ISO8859_1.NewDecoder()
out, err := dec.Bytes(data)
if err != nil {
return data
}
return out
}
var _ = fmt.Sprintf // keep fmt import if helpers change
+222
View File
@@ -0,0 +1,222 @@
package sync
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"net/http/cookiejar"
"net/url"
"strings"
"time"
)
// OOConfig mirrors the .env / environment used by bin/mail/import.
type OOConfig struct {
URL string
User string
Password string
}
// OOClient is a minimal OnlyOffice API client: authentication.json for the
// bearer token plus the session cookie jar required by the .ashx download
// handler. It mirrors the endpoint contract bin/mail/import already uses.
type OOClient struct {
cfg OOConfig
client *http.Client
mu chan struct{}
token string
folderID int
}
func NewOOClient(cfg OOConfig, folderID int) (*OOClient, error) {
jar, err := cookiejar.New(nil)
if err != nil {
return nil, err
}
c := &OOClient{
cfg: cfg,
client: &http.Client{Jar: jar, Timeout: 60 * time.Second},
mu: make(chan struct{}, 1),
folderID: folderID,
}
c.mu <- struct{}{}
if err := c.authenticate(context.Background()); err != nil {
return nil, err
}
return c, nil
}
func (o *OOClient) authenticate(ctx context.Context) error {
select {
case <-o.mu:
case <-ctx.Done():
return ctx.Err()
}
defer func() { o.mu <- struct{}{} }()
body, _ := json.Marshal(map[string]any{
"userName": o.cfg.User, "password": o.cfg.Password, "type": 0,
})
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
strings.TrimRight(o.cfg.URL, "/")+"/api/2.0/authentication.json",
strings.NewReader(string(body)))
if err != nil {
return err
}
req.Header.Set("Content-Type", "application/json")
resp, err := o.client.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
data, _ := io.ReadAll(io.LimitReader(resp.Body, 4<<20))
if resp.StatusCode < 200 || resp.StatusCode > 299 {
return fmt.Errorf("oo authenticate status %d: %s", resp.StatusCode, truncate(string(data), 200))
}
var out struct {
Response struct {
Token string `json:"token"`
} `json:"response"`
}
if err := json.Unmarshal(data, &out); err != nil {
return err
}
if out.Response.Token == "" {
return fmt.Errorf("oo authenticate: empty token")
}
o.token = out.Response.Token
return nil
}
// get performs an authenticated GET and decodes the JSON body into out.
func (o *OOClient) get(ctx context.Context, path string, out any) error {
u := strings.TrimRight(o.cfg.URL, "/") + path
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+o.token)
req.Header.Set("Accept", "application/json")
resp, err := o.client.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
data, _ := io.ReadAll(io.LimitReader(resp.Body, 16<<20))
if resp.StatusCode < 200 || resp.StatusCode > 299 {
return fmt.Errorf("oo %s: status %d: %s", path, resp.StatusCode, truncate(string(data), 300))
}
if out != nil {
return json.Unmarshal(data, out)
}
return nil
}
// ooMessage mirrors the OnlyOffice mail message JSON (subset we need).
type ooMessage struct {
ID int `json:"id"`
Subject string `json:"subject"`
From string `json:"from"`
To string `json:"to"`
CC string `json:"cc"`
BCC string `json:"bcc"`
ReceivedDate string `json:"receivedDate"`
HTMLBody string `json:"htmlBody"`
TextBody string `json:"textBody"`
HasAttachments bool `json:"hasAttachments"`
MimeMessageID string `json:"mimeMessageId"`
Attachments []struct {
FileID int `json:"fileId"`
FileName string `json:"fileName"`
StoredName string `json:"storedName"`
Size int64 `json:"size"`
ContentType string `json:"contentType"`
} `json:"attachments"`
}
// ListIDs returns message ids in the configured folder, paginating pages until
// maxIDs is reached (0 = all).
func (o *OOClient) ListIDs(ctx context.Context, maxIDs int, page int) (ids []int, next int, err error) {
var out struct {
Response []ooMessage `json:"response"`
}
count := 100
if maxIDs > 0 && maxIDs < count {
count = maxIDs
}
path := fmt.Sprintf("/api/2.0/mail/messages?folder=%d&page=%d&count=%d", o.folderID, page, count)
if err := o.get(ctx, path, &out); err != nil {
return nil, 0, err
}
for _, m := range out.Response {
ids = append(ids, m.ID)
if maxIDs > 0 && len(ids) >= maxIDs {
break
}
}
next = page + 1
return ids, next, nil
}
// GetMessage fetches the full message by id and normalizes into Message.
func (o *OOClient) GetMessage(ctx context.Context, id int) (*Message, error) {
var out struct {
Response ooMessage `json:"response"`
}
path := fmt.Sprintf("/api/2.0/mail/messages/%d", id)
if err := o.get(ctx, path, &out); err != nil {
return nil, err
}
m := out.Response
msg := &Message{
Source: "onlyoffice",
ID: fmt.Sprintf("%d", m.ID),
Folder: "oo",
Subject: m.Subject,
From: m.From,
To: m.To,
CC: m.CC,
BCC: m.BCC,
HTMLBody: m.HTMLBody,
TextBody: m.TextBody,
HasAttachments: m.HasAttachments,
MimeMessageID: m.MimeMessageID,
}
if t, err := time.Parse(time.RFC3339Nano, m.ReceivedDate); err == nil {
msg.ReceivedAt = t
}
for _, a := range m.Attachments {
msg.Attachments = append(msg.Attachments, Attachment{
FileID: fmt.Sprintf("%d", a.FileID),
FileName: a.FileName,
StoredName: a.StoredName,
Size: a.Size,
ContentType: a.ContentType,
})
}
return msg, nil
}
// DownloadAttachment fetches attachment bytes via the .ashx handler, which
// requires the session cookie (client.Jar) captured during authenticate().
func (o *OOClient) DownloadAttachment(ctx context.Context, fileID string) ([]byte, error) {
u := strings.TrimRight(o.cfg.URL, "/") + "/addons/mail/httphandlers/download.ashx?attachid=" + url.QueryEscape(fileID)
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
if err != nil {
return nil, err
}
resp, err := o.client.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
data, err := io.ReadAll(io.LimitReader(resp.Body, 64<<20))
if err != nil {
return nil, err
}
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("oo download attach %s: status %d", fileID, resp.StatusCode)
}
return data, nil
}
+397
View File
@@ -0,0 +1,397 @@
package sync
import (
"context"
"encoding/json"
"errors"
"fmt"
"math"
"math/rand"
"os"
"path/filepath"
"strings"
"sync"
"sync/atomic"
"time"
)
// RetryPolicy is the exponential-backoff strategy applied to transient HTTP
// failures (5xx, timeouts, network errors). Callers wrap transient errors with
// retryWrap; everything else aborts immediately.
type RetryPolicy struct {
MaxAttempts int // total attempts (>=1); 0 => 5
BaseDelay time.Duration // first backoff; 0 => 250ms
MaxDelay time.Duration // cap; 0 => 15s
Jitter float64 // 0..1 multiplier; 0 => 0.2
}
func (p RetryPolicy) withDefaults() RetryPolicy {
if p.MaxAttempts <= 0 {
p.MaxAttempts = 5
}
if p.BaseDelay <= 0 {
p.BaseDelay = 250 * time.Millisecond
}
if p.MaxDelay <= 0 {
p.MaxDelay = 15 * time.Second
}
if p.Jitter <= 0 {
p.Jitter = 0.2
}
return p
}
// delay returns the wait before attempt n (1-based): base * 2^(n-2) + jitter,
// capped at MaxDelay. Attempt 1 waits 0, attempt 2 waits base, then doubles.
func (p RetryPolicy) delay(attempt int) time.Duration {
if attempt <= 1 {
return 0
}
exp := math.Min(float64(attempt-2), 10)
d := float64(p.BaseDelay) * math.Pow(2, exp)
if p.Jitter > 0 {
d *= 1 - p.Jitter + 2*p.Jitter*rand.Float64()
}
if d > float64(p.MaxDelay) {
d = float64(p.MaxDelay)
}
return time.Duration(d)
}
type errRetry struct{ err error }
func (e *errRetry) Error() string { return e.err.Error() }
func (e *errRetry) Unwrap() error { return e.err }
func isRetriable(err error) bool {
var r *errRetry
return errors.As(err, &r)
}
func retryWrap(err error) error {
if err == nil {
return nil
}
if isRetriable(err) {
return err
}
return &errRetry{err: err}
}
// Retry runs fn up to MaxAttempts times with exponential backoff between
// attempts. Non-retriable errors abort immediately. Returns the last error.
func Retry(ctx context.Context, policy RetryPolicy, fn func() error) error {
policy = policy.withDefaults()
var err error
for attempt := 1; attempt <= policy.MaxAttempts; attempt++ {
if err = fn(); err == nil {
return nil
}
if !isRetriable(err) {
return err
}
if attempt == policy.MaxAttempts {
return fmt.Errorf("after %d attempts: %w", policy.MaxAttempts, err)
}
select {
case <-time.After(policy.delay(attempt)):
case <-ctx.Done():
return ctx.Err()
}
}
return err
}
// SyncConfig wires up a sync run.
type SyncConfig struct {
OO *OOConfig // OnlyOffice source (optional)
Gmail *GmailCredentials // Gmail source (optional)
Out string // var/mail root; default <repo>/var/mail
Workers int // concurrency; default 4
Limit int // max messages per source (0 = all)
Offset int // skip first N messages per source
Force bool // overwrite existing message.json + attachments
DryRun bool // list without writing
Query string // Gmail search query; default in:inbox
Policy RetryPolicy
}
// SyncStats is returned by Run.
type SyncStats struct {
Checked int
New int32
Failed int32
Skipped int32
}
// Source abstracts the two backends for the worker pool.
type Source interface {
// ListIDs yields ids (string form) to fetch. cursor resumes pagination.
ListIDs(ctx context.Context, limit int, cursor string) (ids []string, next string, err error)
Get(ctx context.Context, id string) (*Message, error)
DownloadAttachment(ctx context.Context, msg *Message, att Attachment) ([]byte, error)
Folder() string
}
type ooSource struct {
c *OOClient
page int
}
// gmailAPI is the Gmail client surface gmailSource needs. *GmailClient implements it.
type gmailAPI interface {
ListIDs(ctx context.Context, q string, maxIDs int, pageToken string) ([]string, string, error)
GetMessage(ctx context.Context, id string) (*Message, error)
DownloadAttachment(ctx context.Context, msgID, attID string) ([]byte, error)
}
type gmailSource struct {
c gmailAPI
cur string
query string
}
func (s *ooSource) Folder() string { return "inbox" }
func (s *gmailSource) Folder() string { return "gmail" }
func (s *ooSource) ListIDs(ctx context.Context, limit int, cursor string) ([]string, string, error) {
page := s.page
if page == 0 {
page = 1
}
ids, next, err := s.c.ListIDs(ctx, limit, page)
s.page = next
strs := make([]string, len(ids))
for i, id := range ids {
strs[i] = fmt.Sprintf("%d", id)
}
return strs, "", err
}
func (s *ooSource) Get(ctx context.Context, id string) (*Message, error) {
var mid int
if _, err := fmt.Sscanf(id, "%d", &mid); err != nil {
return nil, fmt.Errorf("oo id %q: %w", id, err)
}
return s.c.GetMessage(ctx, mid)
}
func (s *ooSource) DownloadAttachment(ctx context.Context, msg *Message, att Attachment) ([]byte, error) {
return s.c.DownloadAttachment(ctx, att.FileID)
}
func (s *gmailSource) ListIDs(ctx context.Context, limit int, cursor string) ([]string, string, error) {
q := s.query
if q == "" {
q = "in:inbox"
}
ids, next, err := s.c.ListIDs(ctx, q, limit, cursor)
return ids, next, err
}
func (s *gmailSource) Get(ctx context.Context, id string) (*Message, error) {
return s.c.GetMessage(ctx, id)
}
func (s *gmailSource) DownloadAttachment(ctx context.Context, msg *Message, att Attachment) ([]byte, error) {
return s.c.DownloadAttachment(ctx, msg.ID, att.FileID)
}
// Run executes the sync across the configured sources with a worker pool.
func Run(ctx context.Context, cfg SyncConfig) (*SyncStats, error) {
if cfg.Out == "" {
cfg.Out = "var/mail"
}
if cfg.Workers <= 0 {
cfg.Workers = 4
}
if err := os.MkdirAll(cfg.Out, 0o755); err != nil {
return nil, err
}
var sources []Source
if cfg.OO != nil {
oo, err := NewOOClient(*cfg.OO, 1) // folder inbox
if err != nil {
return nil, fmt.Errorf("onlyoffice auth: %w", err)
}
sources = append(sources, &ooSource{c: oo})
}
if cfg.Gmail != nil {
gm, err := NewGmailClient(*cfg.Gmail)
if err != nil {
return nil, fmt.Errorf("gmail init: %w", err)
}
sources = append(sources, &gmailSource{c: gm, query: cfg.Query})
}
if len(sources) == 0 {
return nil, errors.New("sync: no source configured (need OO, Gmail, or both)")
}
stats := &SyncStats{}
var jobs []struct {
src Source
id string
}
for _, src := range sources {
ids, _, err := src.ListIDs(ctx, cfg.Offset+cfg.Limit, "")
if err != nil {
return nil, fmt.Errorf("list %s: %w", src.Folder(), err)
}
if cfg.Offset > 0 {
if cfg.Offset >= len(ids) {
ids = nil
} else {
ids = ids[cfg.Offset:]
}
}
if cfg.Limit > 0 && len(ids) > cfg.Limit {
ids = ids[:cfg.Limit]
}
stats.Checked += len(ids)
for _, id := range ids {
jobs = append(jobs, struct {
src Source
id string
}{src: src, id: id})
}
}
var (
wg sync.WaitGroup
mu sync.Mutex
failures []string
)
jobsCh := make(chan struct {
src Source
id string
})
for i := 0; i < cfg.Workers; i++ {
wg.Add(1)
go func() {
defer wg.Done()
for j := range jobsCh {
status, err := processOne(ctx, j.src, j.id, cfg)
switch status {
case statusFailed:
mu.Lock()
failures = append(failures, j.src.Folder()+"/"+j.id+": "+err.Error())
mu.Unlock()
atomic.AddInt32(&stats.Failed, 1)
case statusNew:
atomic.AddInt32(&stats.New, 1)
case statusSkipped:
atomic.AddInt32(&stats.Skipped, 1)
}
}
}()
}
for _, j := range jobs {
select {
case jobsCh <- j:
case <-ctx.Done():
close(jobsCh)
wg.Wait()
return stats, ctx.Err()
}
}
close(jobsCh)
wg.Wait()
if len(failures) > 0 {
fmt.Fprintf(os.Stderr, "sync: %d failures:\n %s\n", len(failures), strings.Join(failures, "\n "))
}
return stats, nil
}
type status int
const (
statusNew status = iota
statusSkipped
statusFailed
)
func processOne(ctx context.Context, src Source, id string, cfg SyncConfig) (status, error) {
if cfg.DryRun {
return statusNew, nil
}
dir := filepath.Join(cfg.Out, src.Folder(), id)
jsonPath := filepath.Join(dir, "message.json")
if !cfg.Force {
if _, err := os.Stat(jsonPath); err == nil {
return statusSkipped, nil
}
}
var msg *Message
err := Retry(ctx, cfg.Policy, func() error {
m, err := src.Get(ctx, id)
if err != nil {
return retryWrap(err)
}
m.Folder = src.Folder() // directory layout is authoritative
if err := writeMessage(jsonPath, m); err != nil {
return err
}
msg = m
return nil
})
if err != nil {
return statusFailed, err
}
for _, att := range msg.Attachments {
attDir := filepath.Join(dir, "attachments")
if err := os.MkdirAll(attDir, 0o755); err != nil {
return statusFailed, err
}
attPath := filepath.Join(attDir, sanitize(att.StoredName))
if _, err := os.Stat(attPath); err == nil && !cfg.Force {
continue
}
var data []byte
err := Retry(ctx, cfg.Policy, func() error {
b, err := src.DownloadAttachment(ctx, msg, att)
if err != nil {
return retryWrap(err)
}
data = b
return os.WriteFile(attPath, b, 0o644)
})
if err != nil {
return statusFailed, fmt.Errorf("attachment %s: %w", att.FileName, err)
}
// ICS attachments get structured markdown immediately (same name the
// Python converter would use: <display stem>.md).
if isICS(att.FileName) {
stem := att.FileName
if i := strings.LastIndex(stem, "."); i >= 0 {
stem = stem[:i]
}
mdPath := filepath.Join(attDir, sanitize(stem)+".md")
if err := os.WriteFile(mdPath, []byte(ICSToMarkdown(data)), 0o644); err != nil {
return statusFailed, err
}
}
}
return statusNew, nil
}
func writeMessage(path string, m *Message) error {
b, err := json.MarshalIndent(m, "", " ")
if err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
return err
}
return os.WriteFile(path, b, 0o644)
}
func sanitize(name string) string {
r := strings.NewReplacer("/", "_", "\\", "_", ":", "_", "*", "_", "?", "_", "\"", "_",
"<", "_", ">", "_", "|", "_", " ", "_")
return r.Replace(name)
}
func isICS(name string) bool {
n := strings.ToLower(name)
return strings.HasSuffix(n, ".ics") || strings.HasSuffix(n, ".ical")
}
+305
View File
@@ -0,0 +1,305 @@
package sync
import (
"context"
"encoding/base64"
"errors"
"testing"
"time"
"unicode/utf8"
)
func TestRetrySucceedsOnSecondTry(t *testing.T) {
attempts := 0
err := Retry(context.Background(), RetryPolicy{BaseDelay: time.Millisecond, MaxDelay: 5 * time.Millisecond}, func() error {
attempts++
if attempts == 1 {
return retryWrap(errors.New("boom"))
}
return nil
})
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if attempts != 2 {
t.Fatalf("expected 2 attempts, got %d", attempts)
}
}
func TestRetryExhaustsAttempts(t *testing.T) {
attempts := 0
err := Retry(context.Background(), RetryPolicy{MaxAttempts: 3, BaseDelay: time.Millisecond, MaxDelay: 5 * time.Millisecond}, func() error {
attempts++
return retryWrap(errors.New("nope"))
})
if err == nil {
t.Fatal("expected error after exhaustion")
}
if attempts != 3 {
t.Fatalf("expected 3 attempts, got %d", attempts)
}
}
func TestRetryNonRetriableAbortsImmediately(t *testing.T) {
attempts := 0
err := Retry(context.Background(), RetryPolicy{MaxAttempts: 5, BaseDelay: time.Millisecond}, func() error {
attempts++
return errors.New("permanent")
})
if err == nil {
t.Fatal("expected error")
}
if attempts != 1 {
t.Fatalf("expected 1 attempt for non-retriable, got %d", attempts)
}
}
func TestRetryRespectsContext(t *testing.T) {
ctx, cancel := context.WithCancel(context.Background())
cancel()
err := Retry(ctx, RetryPolicy{MaxAttempts: 5, BaseDelay: time.Millisecond}, func() error {
return retryWrap(errors.New("x"))
})
if !errors.Is(err, context.Canceled) {
t.Fatalf("expected context.Canceled, got %v", err)
}
}
func TestDelayGrows(t *testing.T) {
p := RetryPolicy{BaseDelay: time.Second, MaxDelay: 30 * time.Second, Jitter: 0}
d1 := p.delay(1) // attempt 1 => 0
d2 := p.delay(2)
d3 := p.delay(3)
if d1 != 0 {
t.Fatalf("attempt 1 delay should be 0, got %v", d1)
}
if d2 != time.Second {
t.Fatalf("attempt 2 delay should be 1s, got %v", d2)
}
if d3 != 2*time.Second {
t.Fatalf("attempt 3 delay should be 2s, got %v", d3)
}
}
func TestSanitize(t *testing.T) {
cases := map[string]string{
"a/b\\c:d*e": "a_b_c_d_e",
"normal.txt": "normal.txt",
"../evil": ".._evil",
"a b c.pdf": "a_b_c.pdf",
}
for in, want := range cases {
if got := sanitize(in); got != want {
t.Errorf("sanitize(%q) = %q, want %q", in, got, want)
}
}
}
func TestIsICS(t *testing.T) {
if !isICS("reply.ics") || !isICS("x.ICAL") {
t.Fatal("ics extensions not detected")
}
if isICS("invoice.pdf") {
t.Fatal("pdf misdetected as ics")
}
}
const fixtureReplyICS = `BEGIN:VCALENDAR
METHOD:REPLY
PRODID:Microsoft Exchange Server 2010
VERSION:2.0
BEGIN:VTIMEZONE
TZID:W. Europe Standard Time
BEGIN:STANDARD
DTSTART:16010101T030000
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
RRULE:FREQ=YEARLY;INTERVAL=1;BYDAY=-1SU;BYMONTH=10
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:16010101T020000
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
RRULE:FREQ=YEARLY;INTERVAL=1;BYDAY=-1SU;BYMONTH=3
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
ATTENDEE;PARTSTAT=ACCEPTED;CN="Baker, Ben":mailto:bbaker1@teksystems.com
UID:bvlnr1i35ug30kn6rvu9dop00g@google.com
SUMMARY;LANGUAGE=en-US:Accepted: Appointment (Ben Baker)
DTSTART;TZID=W. Europe Standard Time:20260812T120000
DTEND;TZID=W. Europe Standard Time:20260812T123000
CLASS:PUBLIC
STATUS:CONFIRMED
LOCATION;LANGUAGE=en-US:https://meet.google.com/sxh-ubud-jrd
END:VEVENT
END:VCALENDAR`
func TestICSToMarkdown(t *testing.T) {
out := ICSToMarkdown([]byte(fixtureReplyICS))
for _, want := range []string{
"Accepted: Appointment",
"When:",
"Where:",
"meet.google.com",
"Attendee:",
"Baker, Ben",
"Accepted",
"Calendar method: REPLY",
} {
if !contains(out, want) {
t.Errorf("output missing %q:\n%s", want, out)
}
}
if contains(out, "BEGIN:VCALENDAR") {
t.Errorf("raw ICS leaked into markdown:\n%s", out)
}
}
func TestICSToMarkdownFallback(t *testing.T) {
out := ICSToMarkdown([]byte("not a calendar"))
if !contains(out, "not a calendar") {
t.Fatalf("expected raw fallback, got %q", out)
}
}
func TestICSToMarkdownNormalizesLatin1(t *testing.T) {
// Real-world ICS from a rental portal: summary in UTF-8, location Latin-1
// ("N\xfcrnberg"). The markdown output must be valid UTF-8 everywhere.
raw := "BEGIN:VCALENDAR\r\nVERSION:2.0\r\nBEGIN:VEVENT\r\n" +
"SUMMARY:Mietwagen-Buchung: N\xc3\xbcrnberg\r\n" +
"LOCATION:N\xfcrnberg\r\nDTSTART:20200101T090000Z\r\nDTEND:20200101T180000Z\r\n" +
"END:VEVENT\r\nEND:VCALENDAR\r\n"
out := ICSToMarkdown([]byte(raw))
if !utf8.ValidString(out) {
t.Fatalf("output is not valid UTF-8:\n%q", out)
}
if !contains(out, "Nürnberg") {
t.Errorf("expected Nürnberg in output:\n%s", out)
}
if contains(out, "N\xfcrnberg") {
t.Errorf("Latin-1 bytes leaked into output:\n%q", out)
}
}
func TestICSToMarkdownAllDay(t *testing.T) {
ics := `BEGIN:VCALENDAR
VERSION:2.0
BEGIN:VEVENT
UID:y@google.com
SUMMARY:All day thing
DTSTART;VALUE=DATE:20260815
DTEND;VALUE=DATE:20260816
END:VEVENT
END:VCALENDAR`
out := ICSToMarkdown([]byte(ics))
if !contains(out, "All day thing") || !contains(out, "2026-08-15") {
t.Errorf("all-day event not parsed:\n%s", out)
}
}
func TestCollectParts(t *testing.T) {
p := gmailPart{
MimeType: "multipart/mixed",
Parts: []gmailPart{
{MimeType: "multipart/alternative", Parts: []gmailPart{
{MimeType: "text/plain", Body: gmailBody{Data: b64("plain text")}},
{MimeType: "text/html", Body: gmailBody{Data: b64("<p>html</p>")}},
}},
{PartID: "2", MimeType: "application/pdf", Filename: "invoice.pdf", Body: gmailBody{Size: 100}},
},
}
text, html, atts := collectParts(p, "root", "abc123", 0)
if text != "plain text" {
t.Errorf("text = %q", text)
}
if html != "<p>html</p>" {
t.Errorf("html = %q", html)
}
if len(atts) != 1 || atts[0].FileName != "invoice.pdf" || atts[0].FileID != "abc123:2" {
t.Errorf("atts = %+v", atts)
}
}
type fakeGmailAPI struct {
lastQ string
lastLimit int
ids []string
}
func (f *fakeGmailAPI) ListIDs(_ context.Context, q string, maxIDs int, _ string) ([]string, string, error) {
f.lastQ = q
f.lastLimit = maxIDs
return f.ids, "", nil
}
func (f *fakeGmailAPI) GetMessage(context.Context, string) (*Message, error) {
return nil, errors.New("unused")
}
func (f *fakeGmailAPI) DownloadAttachment(context.Context, string, string) ([]byte, error) {
return nil, errors.New("unused")
}
func TestGmailSourcePassesQueryToListIDs(t *testing.T) {
fake := &fakeGmailAPI{ids: []string{"m1"}}
src := &gmailSource{c: fake, query: "from:alice@example.com"}
ids, _, err := src.ListIDs(context.Background(), 10, "")
if err != nil {
t.Fatal(err)
}
if fake.lastQ != "from:alice@example.com" {
t.Fatalf("ListIDs q=%q, want from:alice@example.com", fake.lastQ)
}
if fake.lastLimit != 10 {
t.Fatalf("ListIDs limit=%d, want 10", fake.lastLimit)
}
if len(ids) != 1 || ids[0] != "m1" {
t.Fatalf("ids=%v", ids)
}
}
func TestGmailSourceEmptyQueryDefaultsToInbox(t *testing.T) {
fake := &fakeGmailAPI{}
src := &gmailSource{c: fake, query: ""}
if _, _, err := src.ListIDs(context.Background(), 5, ""); err != nil {
t.Fatal(err)
}
if fake.lastQ != "in:inbox" {
t.Fatalf("empty query q=%q, want in:inbox", fake.lastQ)
}
}
func TestParseCLIGmailQuery(t *testing.T) {
cli, code, err := ParseCLI([]string{
"--source", "gmail",
"--query", "from:letrado@example.com",
"--out", t.TempDir(),
"--dry-run",
})
if err != nil || code != 0 {
t.Fatalf("ParseCLI: code=%d err=%v", code, err)
}
if cli.Sync.Query != "from:letrado@example.com" {
t.Fatalf("query=%q", cli.Sync.Query)
}
if cli.Sync.Gmail == nil {
t.Fatal("gmail source not configured")
}
}
func b64(s string) string {
return base64.URLEncoding.EncodeToString([]byte(s))
}
func contains(s, sub string) bool {
return len(s) >= len(sub) && (s == sub || len(sub) == 0 ||
indexOf(s, sub) >= 0)
}
func indexOf(s, sub string) int {
for i := 0; i+len(sub) <= len(s); i++ {
if s[i:i+len(sub)] == sub {
return i
}
}
return -1
}
+46
View File
@@ -0,0 +1,46 @@
// Package sync downloads OnlyOffice and Gmail messages to var/mail/ as raw
// JSON + attachment files, then hands off to bin/mail/import --from-raw for
// markdown conversion.
//
// On-disk schema (per message):
//
// var/mail/<folder>/<id>/message.json # Message (this package)
// var/mail/<folder>/<id>/attachments/ # raw attachment bytes (storedName)
//
// The Message JSON is the contract shared with the Python converter. Fields
// deliberately mirror what bin/mail/import already reads from the OnlyOffice
// API, so conversion is source-agnostic.
package sync
import (
"time"
)
// Attachment describes one attachment of a Message. FileID/FileName/StoredName
// mirror OnlyOffice; Gmail fills them from its own ids. StoredName is always
// unique (hash/attachment id) so raw files never collide.
type Attachment struct {
FileID string `json:"fileId,omitempty"`
FileName string `json:"fileName"`
StoredName string `json:"storedName"`
Size int64 `json:"size,omitempty"`
ContentType string `json:"contentType,omitempty"`
}
// Message is the normalized record written to var/mail/<folder>/<id>/message.json.
type Message struct {
Source string `json:"source"` // "onlyoffice" | "gmail"
ID string `json:"id"`
Folder string `json:"folder"`
Subject string `json:"subject,omitempty"`
From string `json:"from,omitempty"`
To string `json:"to,omitempty"`
CC string `json:"cc,omitempty"`
BCC string `json:"bcc,omitempty"`
ReceivedAt time.Time `json:"receivedAt,omitempty"`
HTMLBody string `json:"htmlBody,omitempty"`
TextBody string `json:"textBody,omitempty"`
HasAttachments bool `json:"hasAttachments,omitempty"`
Attachments []Attachment `json:"attachments,omitempty"`
MimeMessageID string `json:"mimeMessageId,omitempty"`
}
+3
View File
@@ -0,0 +1,3 @@
// Commands in this directory are shebang mains (import.go), tagged so
// `go build ./bin/markdown` does not see two mains.
package main
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=markdown_import "$0" "$@"; exit
//go:build markdown_import
//
// bin/markdown/import.go - split markdown into leafs (mistune).
//
// ./bin/markdown/import.go [dir]
// ./bin/markdown/import.go --files a.md,b.md --json
//
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/md/import", os.Args[1:]))
}
+3 -2
View File
@@ -1,9 +1,10 @@
#!/usr/bin/env python3
import lib
from __future__ import annotations
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools"))
from mdleaves import leaves_to_json, read_markdown, to_all, walk_markdown # noqa: E402
from yamlout import to_yaml # noqa: E402
+2
View File
@@ -0,0 +1,2 @@
// Commands in this directory are shebang mains (query.go).
package main
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=postgres_query "$0" "$@"; exit
//go:build postgres_query
//
// bin/postgres/query.go - read-only Postgres as YAML.
//
// ./bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
//
// Profiles: $HOME/.config/brain/db-profiles.yml (credentials stay out of git).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/db/psql-yq", os.Args[1:]))
}
Executable
+22
View File
@@ -0,0 +1,22 @@
//usr/bin/env go run -tags=brain_serve "$0" "$@"; exit
//go:build brain_serve
//
// bin/serve.go — deprecated; use bin/brain/serve.go.
package main
import (
"fmt"
"os"
"github.com/eSlider/2dph/internal/httpapi"
)
func main() {
fmt.Fprintln(os.Stderr, "bin/serve.go is deprecated; use bin/brain/serve.go")
if os.Getenv("KB_ROOT") == "" {
if wd, err := os.Getwd(); err == nil {
os.Setenv("KB_ROOT", wd)
}
}
httpapi.Run(nil)
}
+29
View File
@@ -0,0 +1,29 @@
"""crmfacts - pure helpers for bin/facts/crm (association proofing).
Shared with tools/ unit tests so the corpus-org parser is covered in CI.
"""
import re
def corpus_orgs(raw: str) -> dict[str, dict]:
"""Parse the orgs block of the CV knowledge-mesh YAML into id -> fields.
Fields kept: label, kind, period, website. Stops at the first sibling
top-level key (clients, timeline, ...).
"""
m = re.search(r"^orgs:\n(.*?)\n^(?:clients|timeline|tech_weights|nodes|edges):", raw, re.S | re.M)
if not m:
return {}
orgs: dict[str, dict] = {}
cur = None
for line in m.group(1).splitlines():
lm = re.match(r"^\s*- id:\s*(\S+)", line)
if lm:
cur = lm.group(1)
orgs[cur] = {}
continue
fm = re.match(r"^\s+(\w+):\s*(.*)$", line)
if fm and cur and fm.group(1) in ("label", "kind", "period", "website"):
orgs[cur][fm.group(1)] = fm.group(2).strip()
return orgs
+60
View File
@@ -0,0 +1,60 @@
"""gitimport - Ladybug graph writes for Commit/File/Person (no git binary).
Commit records come from bin/git/import.go (go-git). This module only MERGEs
the version graph File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person.
"""
from __future__ import annotations
from dataclasses import dataclass, field
@dataclass
class Commit:
sha: str
author: str
email: str
date: str
subject: str
files: list[str] = field(default_factory=list)
GIT_SCHEMA = (
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
"author STRING, email STRING, date STRING, PRIMARY KEY(id))",
"CREATE NODE TABLE IF NOT EXISTS Person (id STRING, name STRING, email STRING, PRIMARY KEY(id))",
"CREATE REL TABLE IF NOT EXISTS HAS_VERSION (FROM File TO Commit)",
"CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)",
)
def ensure_git_schema(conn) -> None:
for stmt in GIT_SCHEMA:
conn.execute(stmt)
def index_commits(conn, commits: list[Commit], repo: str) -> int:
"""Write Commit/File/Person nodes + edges, one per commit (idempotent by sha)."""
ensure_git_schema(conn)
for c in commits:
conn.execute(
"MERGE (c:Commit {id:$sha}) SET c.repo=$repo, c.subject=$subject, "
"c.author=$author, c.email=$email, c.date=$date",
parameters={"sha": c.sha, "repo": repo, "subject": c.subject,
"author": c.author, "email": c.email, "date": c.date},
)
conn.execute(
"MERGE (p:Person {id:$email}) SET p.name=$name, p.email=$email",
parameters={"email": c.email, "name": c.author},
)
conn.execute("MATCH (c:Commit {id:$sha}), (p:Person {id:$email}) "
"MERGE (c)-[:AUTHORED]->(p)",
parameters={"sha": c.sha, "email": c.email})
for path in c.files:
conn.execute(
"MERGE (f:File {id:$fid}) SET f.path=$path, f.repo=$repo",
parameters={"fid": f"{repo}:{path}", "path": path, "repo": repo},
)
conn.execute("MATCH (f:File {id:$fid}), (c:Commit {id:$sha}) "
"MERGE (f)-[:HAS_VERSION]->(c)",
parameters={"fid": f"{repo}:{path}", "sha": c.sha})
return len(commits)
+95 -13
View File
@@ -22,7 +22,17 @@ ROOT_FACTS = "facts"
ROOT_INFO = "info"
CONF_CONFIRMED = "confirmed"
VAR = Path(__file__).resolve().parents[1] / "var"
def _repo_root() -> Path:
p = Path(__file__).resolve().parent
while True:
if (p / "var").is_dir() or (p / ".git").is_dir() or (p / "pyproject.toml").is_file():
return p
if p.parent == p:
return Path(__file__).resolve().parents[2]
p = p.parent
VAR = _repo_root() / "var"
DB_PATH = VAR / "kb.lbug"
@@ -68,6 +78,19 @@ def init_schema(conn: ladybug.Connection) -> None:
conn.execute(
"CREATE REL TABLE IF NOT EXISTS RUNS_ON (FROM Leaf TO Host)"
)
conn.execute(
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
"author STRING, email STRING, date STRING, PRIMARY KEY(id))"
)
conn.execute(
"CREATE NODE TABLE IF NOT EXISTS Person (id STRING, name STRING, email STRING, PRIMARY KEY(id))"
)
conn.execute(
"CREATE REL TABLE IF NOT EXISTS HAS_VERSION (FROM File TO Commit)"
)
conn.execute(
"CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
)
def leaf_id(text: str, source: str) -> str:
@@ -95,18 +118,77 @@ def upsert_leaf(conn: ladybug.Connection, *, text: str, root: str, confidence: s
return lid
def leaf_index_names(conn: ladybug.Connection) -> set[str]:
"""Return index names on the Leaf table (e.g. {'id', 'Leaf_vec', '_PK'})."""
rows = conn.execute("CALL SHOW_INDEXES() RETURN *").get_all()
return {row[1] for row in rows if row[0] == "Leaf"}
def create_fts_and_vector(conn: ladybug.Connection, force: bool = False) -> None:
if force:
conn.execute("DROP INDEX IF EXISTS Leaf.Leaf_fts")
conn.execute("DROP INDEX IF EXISTS Leaf.Leaf_vec")
try:
conn.execute("CALL CREATE_FTS_INDEX('Leaf', 'id', ['text'])")
except Exception:
pass
try:
conn.execute("CALL CREATE_VECTOR_INDEX('Leaf', 'Leaf_vec', 'embedding', metric := 'cosine')")
except Exception:
pass
"""Create FTS (BM25) + HNSW vector indexes if missing.
Never DROP INDEX for FTS/VECTOR. Ladybug 0.19 leaves ghost catalog
entries after DROP (`_0_Leaf_vec_UPPER`, `0_id_docs`), so a later
CREATE fails with "already exists in catalog" while SHOW_INDEXES
still omits the index. Swallowing that error made HNSW look "OK"
until the first QUERY_VECTOR_INDEX.
`force=True` is accepted for API compatibility but does **not** drop.
Fresh indexes require deleting `var/kb.lbug` and rebuilding
(`bin/brain/index.go --rebuild`).
"""
del force # API compat; DROP is unsafe — see docstring
names = leaf_index_names(conn)
if "id" not in names:
try:
conn.execute("CALL CREATE_FTS_INDEX('Leaf', 'id', ['text'])")
except Exception as e:
raise RuntimeError(
"CREATE_FTS_INDEX failed (often ghost catalog after DROP INDEX). "
"Delete var/kb.lbug and run bin/brain/index.go --rebuild. "
f"Cause: {e}"
) from e
if "Leaf_vec" not in names:
try:
conn.execute(
"CALL CREATE_VECTOR_INDEX('Leaf', 'Leaf_vec', 'embedding', "
"metric := 'cosine')"
)
except Exception as e:
raise RuntimeError(
"CREATE_VECTOR_INDEX failed (often ghost catalog after DROP INDEX "
"Leaf.Leaf_vec → `_0_Leaf_vec_UPPER already exists in catalog`). "
"Delete var/kb.lbug and run bin/brain/index.go --rebuild. "
f"Cause: {e}"
) from e
names = leaf_index_names(conn)
missing = {"id", "Leaf_vec"} - names
if missing:
raise RuntimeError(
f"Leaf indexes incomplete after create: missing {sorted(missing)}; "
f"have {sorted(names)}. Delete var/kb.lbug and --rebuild."
)
def ensure_indexes(conn: ladybug.Connection) -> None:
"""Idempotent: create FTS + HNSW only when missing. Safe after upserts."""
create_fts_and_vector(conn, force=False)
def drop_indexes(conn: ladybug.Connection) -> None:
"""No-op. Kept for callers; DROP INDEX is fatal on Ladybug 0.19.
Historical note claimed "drop before bulk MERGE". Measured on 0.19:
- DROP FTS/VECTOR leaves ghost catalog CREATE fails permanently until
`var/kb.lbug` is deleted.
- MERGE/upsert while **FTS** exists can corrupt FTS
("document for node offset N is missing during delete").
- Upsert while **HNSW** exists stays queryable.
Bulk rebuilders must delete `var/kb.lbug`, write all leafs (info+facts)
with no indexes, then `ensure_indexes()` once.
"""
return
def query_fts(conn: ladybug.Connection, text: str, limit: int = 10) -> list[dict]:
@@ -155,6 +237,6 @@ def stats(conn: ladybug.Connection) -> dict:
def open_readonly() -> tuple[ladybug.Database, ladybug.Connection]:
if not DB_PATH.exists():
raise FileNotFoundError(f"{DB_PATH} missing - run bin/kb/index first")
raise FileNotFoundError(f"{DB_PATH} missing - run bin/brain/index.go --rebuild first")
db, conn = connect(read_only=True)
return db, conn
+148
View File
@@ -0,0 +1,148 @@
"""mailconv - pure helpers for bin/mail/import (mail -> markdown + attachments).
Shared with unit tests in bin/tools/test_mailconv.py. No network, no OnlyOffice
dependencies here: everything is `str -> str` or `Path -> str` so the tests run
offline against fixtures.
"""
from __future__ import annotations
import html
import re
import zipfile
from pathlib import Path
# Body part / attachment file suffixes we know how to turn into markdown text.
TEXT_SUFFIXES = {".md", ".markdown", ".txt", ".csv", ".json", ".xml", ".yaml", ".yml", ".log", ".tsv",
".ics", ".ical", ".vcf", ".eml"}
OFFICE_SUFFIXES = {".docx", ".pptx", ".xlsx", ".html", ".htm", ".epub", ".eml", ".msg"}
PDF_SUFFIXES = {".pdf"}
IMAGE_SUFFIXES = {".png", ".jpg", ".jpeg", ".gif", ".bmp", ".tiff", ".tif", ".webp"}
ARCHIVE_SUFFIXES = {".zip"}
# Legacy binary Office (doc/xls/ppt) — markitdown/docling skip them; we try
# pandoc first, else leave a stub.
LEGACY_OFFICE_SUFFIXES = {".doc", ".xls", ".ppt"}
CONVERTIBLE_SUFFIXES = (
TEXT_SUFFIXES | OFFICE_SUFFIXES | PDF_SUFFIXES | IMAGE_SUFFIXES | ARCHIVE_SUFFIXES | LEGACY_OFFICE_SUFFIXES
)
def clean_email_address(raw: str) -> str:
"""Extract the bare email from '"Name" <a@b.c>' and strip control chars."""
m = re.search(r"<([^<>@\s]+@[^<>@\s]+)>", raw)
return (m.group(1) if m else raw).strip()
def subject_to_filename(subject: str, max_len: int = 80) -> str:
"""Turn a mail subject into a filesystem-safe slug (keep first token readable)."""
s = re.sub(r"[^\w\-. ]+", "", subject).strip()
s = re.sub(r"\s+", "_", s)
s = s.strip("._")
if not s:
s = "untitled"
return s[:max_len] or "untitled"
def strip_html(html_text: str) -> str:
"""Naive HTML -> plain text fallback (used only if markitdown is missing)."""
import re as _re
text = _re.sub(r"(?is)<(script|style)[^>]*>.*?</\1>", "", html_text)
text = _re.sub(r"(?s)<br\s*/?>", "\n", text)
text = _re.sub(r"(?s)</p>", "\n\n", text)
text = _re.sub(r"(?s)<[^>]+>", "", text)
return html.unescape(text).strip()
def _unwrap_tables(html_text: str) -> str:
"""Unwrap mail HTML tables into pipe-joined text lines.
Outlook/Stripe-style emails wrap content in nested spacer/frame tables that
markitdown renders as hundreds of `--- |` cells and duplicated blocks.
Every <table> becomes plain "cell1 | cell2" lines (key-value pairs survive),
so only headings/paragraphs/links reach markitdown and no table noise is left.
"""
try:
from bs4 import BeautifulSoup
except Exception:
return html_text
soup = BeautifulSoup(html_text, "html.parser")
for table in reversed(soup.find_all("table")):
lines: list[str] = []
for row in table.find_all("tr"):
cells = [c.get_text(" ", strip=True) for c in row.find_all(["td", "th"])]
line = " | ".join(x for x in cells if x)
if line:
lines.append(line)
if lines:
table.replace_with(BeautifulSoup("\n".join(lines), "html.parser"))
else:
table.decompose()
return str(soup)
def html_to_markdown(html_text: str) -> str:
"""Convert a mail HTML body to markdown using markitdown when available."""
html_text = _unwrap_tables(html_text)
try:
from markitdown import MarkItDown
import io
md = MarkItDown()
result = md.convert_stream(io.BytesIO(html_text.encode("utf-8", errors="replace")),
file_extension=".html")
text = result.text_content.strip()
if text:
return normalize_markdown(text)
except Exception:
pass
return normalize_markdown(strip_html(html_text))
def normalize_markdown(text: str) -> str:
"""Collapse the pdfminer/markitdown NUL artifacts and stray control chars."""
# NUL bytes that pdfminer inserts between digits/letters.
text = text.replace("\x00", "")
# Email spacer noise: zero-width chars, soft hyphens, figure spaces,
# combining grapheme joiner, BOM.
for ch in ("\ufeff", "\u200b", "\u034f", "\u00ad", "\u2007", "\u2008", "\u200a", "\u2002"):
text = text.replace(ch, "")
text = re.sub(r"[ \t]{2,}", " ", text)
# Trim trailing whitespace per line so space-only spacer rows collapse.
text = "\n".join(l.rstrip() for l in text.split("\n"))
# Collapse 3+ blank lines to two.
text = re.sub(r"\n{3,}", "\n\n", text)
# Remove weird trailing control chars.
text = "".join(ch for ch in text if ch >= " " or ch in "\n\t")
return text.strip()
def split_zip_members(zip_path: Path) -> list[str]:
"""Return safe member names of a zip archive (skips dir entries)."""
try:
with zipfile.ZipFile(zip_path) as zf:
return [m for m in zf.namelist() if not m.endswith("/")]
except zipfile.BadZipFile:
return []
def zip_extract_safe(zip_path: Path, dest: Path) -> list[Path]:
"""Extract a zip into dest guarding against path traversal; returns files."""
out: list[Path] = []
try:
with zipfile.ZipFile(zip_path) as zf:
for member in zf.infolist():
if member.is_dir():
continue
target = (dest / member.filename).resolve()
if not target.is_relative_to(dest.resolve()):
continue
target.parent.mkdir(parents=True, exist_ok=True)
with zf.open(member) as src, open(target, "wb") as dst:
dst.write(src.read())
out.append(target)
except zipfile.BadZipFile:
return []
return out
def is_convertible(suffix: str) -> bool:
return suffix.lower() in CONVERTIBLE_SUFFIXES
+37
View File
@@ -0,0 +1,37 @@
"""Mail markdown under var/mail → info leafs. Conversion stays off the brain DB."""
from __future__ import annotations
import json
from pathlib import Path
from mdleaves import read_markdown, to_all
def msg_date(md: Path) -> str:
j = md.parent / "message.json"
try:
d = json.loads(j.read_text(encoding="utf-8"))
return (d.get("receivedDate") or d.get("receivedAt") or "")[:10]
except (OSError, json.JSONDecodeError, TypeError):
return ""
def from_mail_root(root: Path, limit: int = 0, since: str = "", repo: str = "ooMail") -> list[dict]:
if not root.is_dir():
return []
mds = sorted(root.rglob("message.md"))
if since:
mds = [m for m in mds if msg_date(m) >= since]
if limit:
mds = mds[:limit]
leafs: list[dict] = []
for md in mds:
files = [md] + sorted((md.parent / "attachments").glob("*.md"))
for f in files:
if not f.exists():
continue
for lf in to_all(read_markdown(f), f, repo=repo):
lf["source"] = f"ooMail:{md.parent.name}:{f.name}"
lf["how"] = "mail/import"
leafs.append(lf)
return leafs
+121
View File
@@ -0,0 +1,121 @@
"""D14 layout: bin/{subject}/{method}.go, libs in internal/, one go.mod."""
from __future__ import annotations
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class BinLayoutTest(unittest.TestCase):
def test_brain_search_shebang_exists(self) -> None:
p = ROOT / "bin" / "brain" / "search.go"
self.assertTrue(p.is_file(), "missing bin/brain/search.go")
first = p.read_text().splitlines()[0]
self.assertTrue(
first.startswith("//usr/bin/env go run"),
f"shebang first line, got {first!r}",
)
def test_no_nested_go_mod_under_bin(self) -> None:
nested = list((ROOT / "bin").rglob("go.mod"))
self.assertEqual(nested, [], f"nested go.mod files: {nested}")
def test_rank_lives_in_internal_brain(self) -> None:
self.assertTrue(
(ROOT / "internal" / "brain" / "rank" / "rank.go").is_file(),
"ranking must live in internal/brain/rank (cgo-free)",
)
self.assertFalse(
(ROOT / "bin" / "kbsearch").exists(),
"bin/kbsearch nested module must be gone",
)
def test_no_main_go_under_bin_brain(self) -> None:
main = ROOT / "bin" / "brain" / "main.go"
self.assertFalse(main.exists(), "bin/brain/main.go is not a method")
def test_chats_methods_are_shebangs_not_main(self) -> None:
chats = ROOT / "bin" / "chats"
self.assertFalse(
(chats / "main.go").exists(),
"bin/chats/main.go is a dispatcher, not a method",
)
self.assertFalse(
(chats / "index_cmd.go").exists(),
"chats index is a brain write hiding under the wrong subject",
)
for method in ("sync.go", "import.go", "facts.go", "apply.go"):
p = chats / method
self.assertTrue(p.is_file(), f"missing bin/chats/{method}")
first = p.read_text().splitlines()[0]
self.assertTrue(
first.startswith("//usr/bin/env go run"),
f"{method} shebang, got {first!r}",
)
def test_chats_lib_lives_in_internal(self) -> None:
self.assertTrue(
(ROOT / "internal" / "chats" / "linkedin.go").is_file(),
"LinkedIn parser must live in internal/chats",
)
self.assertFalse(
(ROOT / "bin" / "chats" / "linkedin.go").exists(),
"parser must not stay under bin/chats as a second main",
)
def _assert_shebang(self, rel: str) -> None:
p = ROOT / rel
self.assertTrue(p.is_file(), f"missing {rel}")
first = p.read_text().splitlines()[0]
self.assertTrue(
first.startswith("//usr/bin/env go run"),
f"{rel} shebang, got {first!r}",
)
def test_brain_methods_are_shebangs(self) -> None:
for method in ("index.go", "get.go", "stats.go", "eval.go", "watch.go"):
self._assert_shebang(f"bin/brain/{method}")
def test_mail_import_is_shebang_not_brain_write(self) -> None:
self._assert_shebang("bin/mail/import.go")
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
self.assertIn(
"bin/brain/index.go",
index_mail,
"index_mail must point at bin/brain/index.go",
)
def test_markdown_import_is_shebang(self) -> None:
self._assert_shebang("bin/markdown/import.go")
def test_postgres_query_is_shebang(self) -> None:
self._assert_shebang("bin/postgres/query.go")
def test_git_import_is_gogit_shebang(self) -> None:
self._assert_shebang("bin/git/import.go")
py = (ROOT / "bin" / "git" / "import").read_text()
self.assertNotIn(
'["git"',
py,
"Python git/import must not subprocess the git binary",
)
self.assertIn("bin/git/import.go", py)
def test_web_search_is_shebang(self) -> None:
self._assert_shebang("bin/web/search.go")
py = (ROOT / "bin" / "web" / "search").read_text()
self.assertIn("bin/web/search.go", py)
def test_gitimport_py_has_no_git_binary(self) -> None:
py = (ROOT / "bin" / "tools" / "gitimport.py").read_text()
self.assertNotIn("subprocess", py)
self.assertNotIn("git log", py)
def test_gogit_is_direct_go_mod_require(self) -> None:
text = (ROOT / "go.mod").read_text()
first = text.split("require (")[1].split(")")[0]
self.assertRegex(first, r"github.com/go-git/go-git/v5\s+v")
for line in first.splitlines():
if "go-git/go-git" in line:
self.assertNotIn("indirect", line)
+46
View File
@@ -0,0 +1,46 @@
import sys
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import crmfacts # noqa: E402
FM = """\
schema: 2
meta:
title: x
orgs:
- id: produktor
label: ProProdukt SL / produktor.io
kind: own
period: 2006present
website: https://produktor.io
- id: dyvenia
label: Dyvenia
kind: employer
period: 20232025
clients:
- name: One
- name: Two
timeline:
- start: 2001
"""
class CorpusOrgsTest(unittest.TestCase):
def test_parses_label_kind_period(self):
orgs = crmfacts.corpus_orgs(FM)
self.assertEqual(orgs["produktor"]["label"], "ProProdukt SL / produktor.io")
self.assertEqual(orgs["produktor"]["kind"], "own")
self.assertEqual(orgs["dyvenia"]["kind"], "employer")
def test_does_not_leak_clients_into_orgs(self):
orgs = crmfacts.corpus_orgs(FM)
self.assertNotIn("One", orgs)
self.assertNotIn("Two", orgs)
self.assertNotIn("timeline", orgs)
if __name__ == "__main__":
unittest.main()
+74
View File
@@ -0,0 +1,74 @@
import os
import sys
import tempfile
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import kblib # noqa: E402
import gitimport # noqa: E402
COMMIT_PERSON_SCHEMA = (
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
"author STRING, email STRING, date STRING, PRIMARY KEY(id))"
)
PERSON_SCHEMA = (
"CREATE NODE TABLE IF NOT EXISTS Person (id STRING, name STRING, email STRING, PRIMARY KEY(id))"
)
HAS_VERSION_SCHEMA = "CREATE REL TABLE IF NOT EXISTS HAS_VERSION (FROM File TO Commit)"
AUTHORED_SCHEMA = "CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
def sample_commit() -> gitimport.Commit:
return gitimport.Commit(
sha="a1b2c3d",
author="Ada Lovelace",
email="ada@example.com",
date="2026-08-10T12:00:00+01:00",
subject="feat: first commit",
files=["README.md", "src/main.c"],
)
class GitGraphTest(unittest.TestCase):
def setUp(self):
self.dir = tempfile.mkdtemp()
self.dbpath = os.path.join(self.dir, "kb.lbug")
self.db, self.conn = kblib.connect(self.dbpath, read_only=False)
kblib.init_schema(self.conn)
self.conn.execute(COMMIT_PERSON_SCHEMA)
self.conn.execute(PERSON_SCHEMA)
self.conn.execute(HAS_VERSION_SCHEMA)
self.conn.execute(AUTHORED_SCHEMA)
def tearDown(self):
self.conn.close()
self.db.close()
def test_index_commits_creates_nodes_and_edges(self):
gitimport.index_commits(self.conn, [sample_commit()], "sample-repo")
rp = self.conn.execute("MATCH (p:Person) RETURN p.name, p.email").get_all()
self.assertEqual([tuple(r) for r in rp], [("Ada Lovelace", "ada@example.com")])
rc = self.conn.execute("MATCH (c:Commit) RETURN c.id, c.repo").get_all()
self.assertEqual(len(rc), 1)
self.assertEqual(rc[0][1], "sample-repo")
rf = self.conn.execute(
"MATCH (f:File)-[:HAS_VERSION]->(c:Commit)-[:AUTHORED]->(p:Person) "
"RETURN f.path, c.id, p.email").get_all()
paths = sorted(r[0] for r in rf)
self.assertEqual(paths, ["README.md", "src/main.c"])
self.assertTrue(all(r[2] == "ada@example.com" for r in rf))
def test_index_commits_idempotent(self):
cs = [sample_commit()]
gitimport.index_commits(self.conn, cs, "sample-repo")
gitimport.index_commits(self.conn, cs, "sample-repo")
n = self.conn.execute("MATCH (c:Commit) RETURN count(*)").get_all()[0][0]
self.assertEqual(n, 1)
p = self.conn.execute("MATCH (p:Person) RETURN count(*)").get_all()[0][0]
self.assertEqual(p, 1)
if __name__ == "__main__":
unittest.main()
@@ -35,7 +35,7 @@ class KblibTest(unittest.TestCase):
confidence="confirmed", source="s", source_rev="r1",
how="test", loc="/tmp", type_="reference",
embedding=make_emb(1.0))
kblib.create_fts_and_vector(self.conn, force=True)
kblib.ensure_indexes(self.conn)
hits = kblib.query_fts(self.conn, "fox", 5)
self.assertEqual(len(hits), 1)
self.assertEqual(hits[0]["root"], "info")
@@ -49,12 +49,42 @@ class KblibTest(unittest.TestCase):
confidence="confirmed", source="s", source_rev="r1",
how="test", loc="/tmp", type_="reference",
embedding=make_emb(0.0))
kblib.create_fts_and_vector(self.conn, force=True)
kblib.ensure_indexes(self.conn)
result = kblib.hybrid_search(self.conn, make_emb(1.0), [], 5)
self.assertTrue(result)
self.assertIn("rrf", result[0])
self.assertEqual(result[0]["text"], "the quick brown fox")
def test_upsert_keeps_hnsw_queryable(self):
"""Upsert while HNSW exists must not kill vector search."""
kblib.upsert_leaf(self.conn, text="seed leaf", root="info",
confidence="confirmed", source="s", source_rev="r1",
how="test", loc="/tmp", type_="reference",
embedding=make_emb(0.2))
kblib.ensure_indexes(self.conn)
self.assertIn("Leaf_vec", kblib.leaf_index_names(self.conn))
kblib.upsert_leaf(self.conn, text="added after index", root="facts",
confidence="confirmed", source="a.md x b.md",
source_rev="r1", how="test", loc="/tmp", type_="fact",
embedding=make_emb(0.9))
hits = kblib.query_vector(self.conn, make_emb(0.9), 5)
self.assertTrue(hits)
self.assertIn("Leaf_vec", kblib.leaf_index_names(self.conn))
def test_drop_vector_then_create_raises_clear_error(self):
"""DROP INDEX leaves ghost catalog; create_fts_and_vector must raise."""
kblib.upsert_leaf(self.conn, text="seed", root="info",
confidence="confirmed", source="s", source_rev="r1",
how="test", loc="/tmp", type_="reference",
embedding=make_emb(0.1))
kblib.ensure_indexes(self.conn)
self.conn.execute("DROP INDEX IF EXISTS Leaf.Leaf_vec")
with self.assertRaises(RuntimeError) as ctx:
kblib.create_fts_and_vector(self.conn, force=True)
msg = str(ctx.exception)
self.assertIn("CREATE_VECTOR_INDEX failed", msg)
self.assertIn("--rebuild", msg)
def test_stats_counts_roots(self):
kblib.upsert_leaf(self.conn, text="a fact leaf", root="facts",
confidence="confirmed", source="s", source_rev="r1",
@@ -70,4 +100,4 @@ class KblibTest(unittest.TestCase):
if __name__ == "__main__":
unittest.main()
unittest.main()
+123
View File
@@ -0,0 +1,123 @@
import io
import os
import sys
import unittest
import zipfile
from pathlib import Path
sys.path.insert(0, os.path.dirname(__file__))
from mailconv import ( # noqa: E402
clean_email_address,
html_to_markdown,
is_convertible,
normalize_markdown,
split_zip_members,
subject_to_filename,
zip_extract_safe,
)
from mailconv import _unwrap_tables # noqa: E402
class TestMailConv(unittest.TestCase):
def test_clean_email_address(self):
self.assertEqual(clean_email_address('"Ben Baker" <bb@teks.com>'), "bb@teks.com")
self.assertEqual(clean_email_address("eslider@gmail.com"), "eslider@gmail.com")
self.assertEqual(clean_email_address("<a@b.c>"), "a@b.c")
def test_subject_to_filename(self):
self.assertEqual(subject_to_filename("Your receipt #2422"), "Your_receipt_2422")
self.assertEqual(subject_to_filename("a/b\\c:d*e"), "abcde")
self.assertEqual(subject_to_filename(" "), "untitled")
def test_html_to_markdown(self):
out = html_to_markdown("<html><body><h1>Hi</h1><p>Some <b>bold</b> text.</p></body></html>")
self.assertIn("Hi", out)
self.assertIn("**bold**", out)
def test_html_strip_fallback(self):
from mailconv import strip_html
self.assertEqual(strip_html("<p>a</p><p>b</p>"), "a\n\nb")
def test_flatten_layout_tables(self):
html = ("<table><tr>"
+ "".join(f"<td>spacer{i}</td>" for i in range(12))
+ "</tr></table>"
+ "<p>real</p>"
+ "<table><tr><td>a</td><td>b</td></tr></table>")
out = _unwrap_tables(html)
# tables unwrapped into pipe text; no <td> left; content preserved
self.assertNotIn("<td>spacer0</td>", out)
self.assertIn("spacer0 | spacer1", out)
self.assertIn("a | b", out)
self.assertIn("real", out)
def test_html_to_markdown_layout_clean(self):
html = "<table><tr>" + "".join(f"<td>x{i}</td>" for i in range(12)) + "</tr></table><h1>Hi</h1>"
out = html_to_markdown(html)
self.assertIn("Hi", out)
self.assertNotIn("| ---", out)
def test_normalize_markdown_removes_nul(self):
self.assertEqual(normalize_markdown("Z0\x00A\x00Y\x00B"), "Z0AYB")
self.assertEqual(normalize_markdown("a\n\n\n\nb"), "a\n\nb")
def test_normalize_strips_email_noise(self):
noisy = "\ufeffa\u200b\u034f\u00ad\u2007\u2002 b\u200a c\u2008"
out = normalize_markdown(noisy)
self.assertNotIn("\u200b", out)
self.assertNotIn("\ufeff", out)
self.assertNotIn("\u034f", out)
self.assertIn("a b c", out)
def test_split_zip_members(self):
p = Path(self._mk_zip(["a.txt", "sub/b.txt"]))
self.assertEqual(split_zip_members(p), ["a.txt", "sub/b.txt"])
def test_zip_extract_safe(self):
zip_path = self._mk_zip(["a.txt", "dir/b.txt"])
dest = Path(self._tmp("x"))
files = zip_extract_safe(zip_path, dest)
self.assertEqual(len(files), 2)
self.assertTrue((dest / "a.txt").exists())
self.assertTrue((dest / "dir" / "b.txt").exists())
def test_zip_extract_safe_blocks_traversal(self):
# member "../evil.txt" must not escape dest
zip_path = Path(self._tmp("evil.zip"))
with zipfile.ZipFile(zip_path, "w") as zf:
zf.writestr("../evil.txt", "boom")
dest = Path(self._tmp("out"))
files = zip_extract_safe(zip_path, dest)
self.assertEqual(files, [])
self.assertFalse((dest.parent / "evil.txt").exists())
def test_is_convertible(self):
self.assertTrue(is_convertible(".pdf"))
self.assertTrue(is_convertible(".zip"))
self.assertTrue(is_convertible(".docx"))
self.assertTrue(is_convertible(".TXT"))
self.assertFalse(is_convertible(".exe"))
self.assertFalse(is_convertible(".unknown"))
def _mk_zip(self, members):
zpath = Path(self._tmp("arc.zip"))
with zipfile.ZipFile(zpath, "w") as zf:
for m in members:
zf.writestr(m, "content")
return str(zpath)
def _tmp(self, name):
d = self.__class__._td
p = Path(d) / name
p.parent.mkdir(parents=True, exist_ok=True)
return str(p)
@classmethod
def setUpClass(cls):
import tempfile
cls._td = tempfile.mkdtemp(prefix="mailconv_test_")
if __name__ == "__main__":
unittest.main()
+46
View File
@@ -0,0 +1,46 @@
"""Mail markdown → leafs (no Ladybug). Brain index --with-mail uses this."""
from __future__ import annotations
import json
import sys
import tempfile
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
import mailleafs # noqa: E402
class MailLeafsTest(unittest.TestCase):
def test_message_md_becomes_info_leaf(self) -> None:
root = Path(tempfile.mkdtemp())
msg = root / "inbox" / "alice-1"
msg.mkdir(parents=True)
(msg / "message.json").write_text(
json.dumps({"receivedDate": "2026-01-15T10:00:00Z", "subject": "Hello"}),
encoding="utf-8",
)
(msg / "message.md").write_text(
"---\nroot: info\n---\n\n# Hello\n\nFrom Alice to Bob.\n",
encoding="utf-8",
)
leafs = mailleafs.from_mail_root(root)
self.assertEqual(len(leafs), 1)
self.assertIn("Alice", leafs[0]["text"])
self.assertTrue(leafs[0]["source"].startswith("ooMail:"))
self.assertEqual(leafs[0]["how"], "mail/import")
def test_since_filters_by_message_json_date(self) -> None:
root = Path(tempfile.mkdtemp())
for name, day in (("old", "2025-01-01"), ("new", "2026-06-01")):
d = root / "inbox" / name
d.mkdir(parents=True)
(d / "message.json").write_text(
json.dumps({"receivedDate": f"{day}T00:00:00Z"}),
encoding="utf-8",
)
(d / "message.md").write_text(f"# {name}\n\nbody\n", encoding="utf-8")
leafs = mailleafs.from_mail_root(root, since="2026-01-01")
self.assertEqual(len(leafs), 1)
self.assertIn("new", leafs[0]["text"])
+84
View File
@@ -0,0 +1,84 @@
"""Published docs must match live commands (Gitea SoT, brain/search, no fake --hop)."""
from __future__ import annotations
import re
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class PublishedDocsTest(unittest.TestCase):
def test_readme_points_issues_at_gitea(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn(
"https://git.produktor.io/eSlider/2dph/issues",
text,
"README must point issues at Gitea",
)
def test_plan_d15_names_gitea_origin(self) -> None:
text = (ROOT / "PLAN.md").read_text()
self.assertIn("D15", text)
self.assertIn("git.produktor.io/eSlider/2dph", text)
def test_readme_primary_search_is_brain(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn(
"bin/brain/search.go",
text,
"README deduction search must name bin/brain/search.go",
)
def test_readme_index_is_brain_not_index_mail(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn("bin/brain/index.go", text)
self.assertNotIn(
"bin/mail/index_mail",
text,
"mail index is a brain write; README must name bin/brain/index.go",
)
def test_readme_git_import_is_gogit(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn("bin/git/import.go", text)
self.assertIn("go-git", text)
self.assertIn("D19", (ROOT / "PLAN.md").read_text())
def test_web_search_is_go_not_ops_host(self) -> None:
readme = (ROOT / "README.md").read_text()
self.assertIn("bin/web/search.go", readme)
skill = (ROOT / "skills" / "web-search" / "SKILL.md").read_text()
self.assertIn("bin/web/search.go", skill)
self.assertNotIn("search.ops.io", skill)
self.assertNotIn("search.ops.io", readme)
compose = (ROOT / "compose.yaml").read_text()
self.assertIn("searxng", compose)
self.assertNotIn("search.ops.io", compose)
settings = (ROOT / "deploy" / "searxng" / "settings.yml").read_text()
self.assertNotIn("password", settings.lower())
self.assertIn("json", settings)
def test_readme_search_escalates_web(self) -> None:
text = (ROOT / "README.md").read_text()
self.assertIn("--no-web", text)
self.assertIn("D17", (ROOT / "PLAN.md").read_text())
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
self.assertIn("`web` block", skill)
def test_docs_do_not_claim_hop_walks(self) -> None:
paths = [
ROOT / "README.md",
ROOT / "docs" / "design.md",
ROOT / "skills" / "brain" / "SKILL.md",
ROOT / "skills" / "diataxis-docs" / "SKILL.md",
]
# Command-style `--hop 1` / `--hop N` plus follow/walk = the old lie.
# Honest "not implemented" notes must not match.
lie = re.compile(r"--hop (?:N|1).*(?:follow|walk)", re.I | re.S)
for path in paths:
text = path.read_text()
self.assertIsNone(
lie.search(text),
f"{path.relative_to(ROOT)} still claims --hop walks the graph",
)
+33
View File
@@ -0,0 +1,33 @@
"""Every bin/ path named in skills/ must exist on disk."""
from __future__ import annotations
import re
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
BIN_PATH = re.compile(r"\b(bin/[A-Za-z0-9_./-]+)")
class SkillsBinPathsTest(unittest.TestCase):
def test_agent_cost_skill_is_gone(self) -> None:
self.assertFalse(
(ROOT / "skills" / "agent-cost").exists(),
"skills/agent-cost documents bin/agents/cost which does not exist",
)
def test_brain_skill_replaces_kb_search(self) -> None:
self.assertTrue((ROOT / "skills" / "brain" / "SKILL.md").is_file())
self.assertFalse((ROOT / "skills" / "kb-search").exists())
def test_skill_bin_paths_exist(self) -> None:
missing: list[str] = []
for skill in sorted((ROOT / "skills").rglob("SKILL.md")):
text = skill.read_text()
for match in BIN_PATH.findall(text):
rel = match.rstrip("`'.,")
if rel.endswith(".go") or Path(rel).suffix == "" or Path(rel).suffix in {".go", ".py"}:
p = ROOT / rel
if not p.exists():
missing.append(f"{skill.relative_to(ROOT)}: {rel}")
self.assertEqual(missing, [], "SKILL.md names bin/ paths that do not exist")
View File
@@ -76,7 +76,7 @@ class PhiGuard(unittest.TestCase):
def test_plain_technical_query_passes(self):
self.assertIsNone(ws.phi_reason("Pflegegrad SGB XI Einstufung"))
self.assertIsNone(ws.phi_reason("site:ticket.detective.de Toureffizienz"))
self.assertIsNone(ws.phi_reason("site:example.com technical query"))
def test_long_digit_run_is_refused(self):
self.assertIsNotNone(ws.phi_reason("Kunde 4711220385 Adresse"))
+105
View File
@@ -0,0 +1,105 @@
// Package watch polls corpus directories for changes and re-runs brain/index.
//
// Port of the former bin/kb-watch bash script to an importable, testable Go
// package. Polls file mtimes (no inotify deps); cheap and reliable.
package watch
import (
"log"
"os"
"os/exec"
"path/filepath"
"strconv"
"strings"
"time"
)
// Options controls the polling loop. Zero value uses defaults.
type Options struct {
Dirs []string
Interval time.Duration
// IndexCmd is the index command template. %s is replaced by the repo
// root (from KB_ROOT). Defaults to `python3 <root>/bin/kb/index --with-mail`.
IndexCmd string
}
// Run blocks forever polling Dirs (defaults: KB_WATCH_DIRS or /corpus) every
// Interval (default 30s) and re-indexing when files change. KB_ROOT names the
// repo root used to locate bin/kb/index.
func Run(args []string) {
opts := fromEnv(args)
root, _ := os.Getwd()
if r := os.Getenv("KB_ROOT"); r != "" {
root = r
}
log.Printf("watch: dirs=%v interval=%s root=%s", opts.Dirs, opts.Interval, root)
var last string
for {
if flag := Stamp(opts.Dirs); flag != "" && flag != last {
last = flag
reindex(opts.IndexCmd, root)
}
time.Sleep(opts.Interval)
}
}
func fromEnv(args []string) Options {
opts := Options{Interval: 30 * time.Second}
if raw := os.Getenv("KB_WATCH_INTERVAL"); raw != "" {
if n, err := strconv.Atoi(raw); err == nil && n > 0 {
opts.Interval = time.Duration(n) * time.Second
}
}
defDirs := "/corpus"
if raw := os.Getenv("KB_WATCH_DIRS"); raw != "" {
defDirs = raw
}
if len(args) > 0 {
opts.Dirs = args
} else {
for _, d := range strings.Split(defDirs, " ") {
if d != "" {
opts.Dirs = append(opts.Dirs, d)
}
}
}
pys := os.Getenv("KB_PY")
if pys == "" {
pys = "python3"
}
opts.IndexCmd = pys + " <root>/bin/kb/index --with-mail"
return opts
}
// Stamp returns a rolling fingerprint (newest mtime under dirs) that changes
// whenever any corpus file is touched. Empty when no files found.
func Stamp(dirs []string) string {
var newest time.Time
for _, dir := range dirs {
_ = filepath.WalkDir(dir, func(path string, _ os.DirEntry, err error) error {
if err != nil {
return nil
}
if info, e := os.Stat(path); e == nil && info.ModTime().After(newest) {
newest = info.ModTime()
}
return nil
})
}
if newest.IsZero() {
return ""
}
return strconv.FormatInt(newest.UnixNano(), 10)
}
func reindex(template, root string) {
cmd := strings.ReplaceAll(template, "<root>", root)
parts := strings.Fields(cmd)
c := exec.Command(parts[0], parts[1:]...)
out, err := c.CombinedOutput()
if err != nil {
log.Printf("watch: index failed: %v\n%s", err, out)
} else {
log.Printf("watch: re-indexed")
}
}
+53
View File
@@ -0,0 +1,53 @@
package watch
import (
"os"
"path/filepath"
"strings"
"testing"
"time"
)
func TestStampChangesWhenFileTouched(t *testing.T) {
dir := t.TempDir()
a := filepath.Join(dir, "a.md")
if err := os.WriteFile(a, []byte("x"), 0o644); err != nil {
t.Fatal(err)
}
s1 := Stamp([]string{dir})
if s1 == "" {
t.Fatal("stamp empty for a dir with a file")
}
time.Sleep(10 * time.Millisecond)
if err := os.WriteFile(a, []byte("y"), 0o644); err != nil {
t.Fatal(err)
}
if s2 := Stamp([]string{dir}); s2 == s1 {
t.Fatal("stamp did not change after the file was modified")
}
}
func TestStampEmptyForMissingDir(t *testing.T) {
if s := Stamp([]string{filepath.Join(t.TempDir(), "nope")}); s != "" {
t.Fatalf("stamp = %q, want empty for missing dir", s)
}
}
func TestFromEnvDefaults(t *testing.T) {
t.Setenv("KB_WATCH_INTERVAL", "")
t.Setenv("KB_WATCH_DIRS", "")
t.Setenv("KB_PY", "")
opts := fromEnv(nil)
if len(opts.Dirs) == 0 || opts.Dirs[0] != "/corpus" {
t.Fatalf("default dirs = %v, want [/corpus]", opts.Dirs)
}
if opts.Interval != 30*time.Second {
t.Fatalf("default interval = %s, want 30s", opts.Interval)
}
if !strings.Contains(opts.IndexCmd, "kb/index") {
t.Fatalf("default index cmd = %q, want kb/index", opts.IndexCmd)
}
if !strings.Contains(opts.IndexCmd, "--with-mail") {
t.Fatalf("default index cmd must include --with-mail, got %q", opts.IndexCmd)
}
}
+12 -136
View File
@@ -1,150 +1,26 @@
#!/usr/bin/env python3
"""web/search - web search through the self-hosted SearXNG at search.ops.io.
"""web/search — deprecated. Use bin/web/search.go (SearXNG, no Python client).
bin/web/search "LadybugDB vector search"
bin/web/search "model2vec multilingual" --site github.com
bin/web/search "uclancy" --category it -n 3 --json | jq -r '.results[].url'
bin/web/search "sqlite-vec" --refresh # ignore the cached answer
This complements bin/kb/search: the knowledge base holds our own facts, this
reaches the public web. Use it as the second, independent source that the
detective method asks for.
Exit codes: 0 results, 2 refused as possible PII, 3 throttled (not "nothing
found" - the instance answers 200 with an empty list when it throttles).
bin/web/search.go QUERY [--json] [-n N] [--site HOST]
"""
from __future__ import annotations
import argparse
import fcntl
import json
import os
import sys
import time
import urllib.parse
import urllib.request
from pathlib import Path
TOOLS = Path(__file__).resolve().parents[1].parent / "tools"
sys.path.insert(0, str(TOOLS))
sys.path.insert(0, str(TOOLS / "web-search"))
import websearch as ws # noqa: E402
from yamlout import to_yaml # noqa: E402
CONFIG = Path(os.environ.get("BRAIN_SEARCH_ENV", Path.home() / ".config/brain/search.env"))
CACHE = Path(os.environ.get("BRAIN_SEARCH_CACHE", Path.home() / ".cache/brain/web-search.sqlite"))
LOCK = CACHE.with_suffix(".lock")
ROOT = Path(__file__).resolve().parents[2]
def load_config() -> dict:
if not CONFIG.exists():
sys.exit(f"no credentials at {CONFIG} (mode 600, BRAIN_SEARCH_URL/USER/PASS)")
conf = {}
for line in CONFIG.read_text().splitlines():
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
key, _, value = line.partition("=")
conf[key.strip()] = value.strip().strip("\"'")
missing = {"BRAIN_SEARCH_URL", "BRAIN_SEARCH_USER", "BRAIN_SEARCH_PASS"} - conf.keys()
if missing:
sys.exit(f"{CONFIG} is missing {', '.join(sorted(missing))}")
return conf
def fetch(conf: dict, query: str, params: dict, timeout: int) -> dict:
args = {"q": query, "format": "json", **params}
url = f"{conf['BRAIN_SEARCH_URL'].rstrip('/')}/search?{urllib.parse.urlencode(args)}"
request = urllib.request.Request(url)
token = f"{conf['BRAIN_SEARCH_USER']}:{conf['BRAIN_SEARCH_PASS']}".encode()
import base64
request.add_header("Authorization", "Basic " + base64.b64encode(token).decode())
with urllib.request.urlopen(request, timeout=timeout) as response:
return json.loads(response.read().decode())
def main() -> int:
parser = argparse.ArgumentParser(description="web search via SearXNG")
parser.add_argument("query")
parser.add_argument("-n", "--limit", type=int, default=ws.DEFAULT_LIMIT)
parser.add_argument("--site", help="restrict to one domain")
parser.add_argument("--lang", help="language code, e.g. de")
parser.add_argument("--fresh", choices=["day", "week", "month", "year"],
help="time range")
parser.add_argument("--category", help="SearXNG category, e.g. it, science, news")
parser.add_argument("--engines", help="comma separated engine list")
parser.add_argument("--json", action="store_true")
parser.add_argument("--refresh", action="store_true", help="bypass the cache")
parser.add_argument("--ttl", type=float, default=ws.CACHE_TTL)
parser.add_argument("--timeout", type=int, default=25)
parser.add_argument("--force", action="store_true",
help="send even if the query looks like PII")
args = parser.parse_args()
query = f"site:{args.site} {args.query}" if args.site else args.query
reason = ws.phi_reason(query)
if reason and not args.force:
print(f"refused: {reason}. This query would leave the host.", file=sys.stderr)
print("Rephrase without identifiers, or pass --force if it is genuinely public.",
file=sys.stderr)
return 2
params = {}
if args.lang:
params["language"] = args.lang
if args.fresh:
params["time_range"] = args.fresh
if args.category:
params["categories"] = args.category
if args.engines:
params["engines"] = args.engines
key = ws.cache_key(query, params)
conn = ws.open_cache(CACHE)
if not args.refresh:
cached = ws.cache_get(conn, key, ttl=args.ttl)
if cached is not None:
out = ws.project(cached, limit=args.limit)
out["cached"] = True
sys.stdout.write(json.dumps(out, indent=2, ensure_ascii=False) + "\n"
if args.json else to_yaml(out))
return 0
conf = load_config()
LOCK.parent.mkdir(parents=True, exist_ok=True)
# One request at a time across every agent on this host: the instance
# suspends engines for minutes when several of us ask at once.
with open(LOCK, "w") as lock:
fcntl.flock(lock, fcntl.LOCK_EX)
payload = None
for attempt in range(1 + len(ws.RETRY_BACKOFF)):
delay = ws.wait_for(ws.last_call(conn), time.time())
if delay:
time.sleep(delay)
ws.mark_call(conn)
try:
payload = fetch(conf, query, params, args.timeout)
except Exception as error: # noqa: BLE001 - report, do not crash
print(f"request failed: {error}", file=sys.stderr)
return 3
if ws.classify(payload) == "ok":
break
if attempt < len(ws.RETRY_BACKOFF):
time.sleep(ws.RETRY_BACKOFF[attempt])
if ws.classify(payload) == "ok":
ws.cache_put(conn, key, payload)
out = ws.project(payload, limit=args.limit)
sys.stdout.write(json.dumps(out, indent=2, ensure_ascii=False) + "\n"
if args.json else to_yaml(out))
return 0 if out["status"] == "ok" else 3
def main(argv: list[str]) -> int:
print(
"bin/web/search is deprecated; use bin/web/search.go",
file=sys.stderr,
)
target = ROOT / "bin" / "web" / "search.go"
os.execvp("go", ["go", "run", str(target), *argv])
return 1
if __name__ == "__main__":
sys.exit(main())
sys.exit(main(sys.argv[1:]))
+232
View File
@@ -0,0 +1,232 @@
//usr/bin/env go run "$0" "$@"; exit
//
// bin/web/search.go - SearXNG as the second independent source (D3).
//
// ./bin/web/search.go "LadybugDB vector search"
// ./bin/web/search.go "model2vec" --category it --json
// ./bin/web/search.go "postgres" --site github.com --fresh year
//
// Empty results mean throttled, not "nothing exists". Exit 2 = PII refuse, 3 = throttled.
// Config: $BRAIN_SEARCH_ENV (default $HOME/.config/brain/search.env).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"encoding/json"
"fmt"
"net/http"
"os"
"strconv"
"time"
"github.com/eSlider/2dph/internal/websearch"
"golang.org/x/sys/unix"
)
func main() {
os.Exit(run(os.Args[1:]))
}
func run(args []string) int {
var (
query, site, lang, fresh, category, engines string
limit = websearch.DefaultLimit
jsonOut, refresh, force bool
ttl = float64(websearch.CacheTTL)
timeout = 25
)
i := 0
for i < len(args) {
a := args[i]
switch {
case a == "--json":
jsonOut = true
case a == "--refresh":
refresh = true
case a == "--force":
force = true
case (a == "-n" || a == "--limit") && i+1 < len(args):
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 0 {
fmt.Fprintln(os.Stderr, "web/search: --limit must be a non-negative integer")
return 2
}
limit = n
case a == "--site" && i+1 < len(args):
i++
site = args[i]
case a == "--lang" && i+1 < len(args):
i++
lang = args[i]
case a == "--fresh" && i+1 < len(args):
i++
fresh = args[i]
case a == "--category" && i+1 < len(args):
i++
category = args[i]
case a == "--engines" && i+1 < len(args):
i++
engines = args[i]
case a == "--ttl" && i+1 < len(args):
i++
v, err := strconv.ParseFloat(args[i], 64)
if err != nil {
fmt.Fprintln(os.Stderr, "web/search: --ttl must be a number")
return 2
}
ttl = v
case a == "--timeout" && i+1 < len(args):
i++
n, err := strconv.Atoi(args[i])
if err != nil || n <= 0 {
fmt.Fprintln(os.Stderr, "web/search: --timeout must be a positive integer")
return 2
}
timeout = n
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, `usage: bin/web/search.go QUERY [--json] [-n N] [--site HOST] [--lang LANG] [--fresh day|week|month|year] [--category CAT] [--engines LIST] [--refresh] [--force]`)
return 0
case len(a) > 0 && a[0] != '-' && query == "":
query = a
default:
fmt.Fprintf(os.Stderr, "web/search: unknown flag %s\n", a)
return 2
}
i++
}
if query == "" {
fmt.Fprintln(os.Stderr, "web/search: query required")
return 2
}
if site != "" {
query = "site:" + site + " " + query
}
if reason := websearch.PHIReason(query); reason != "" && !force {
fmt.Fprintf(os.Stderr, "refused: %s. This query would leave the host.\n", reason)
fmt.Fprintln(os.Stderr, "Rephrase without identifiers, or pass --force if it is genuinely public.")
return 2
}
params := map[string]string{}
if lang != "" {
params["language"] = lang
}
if fresh != "" {
params["time_range"] = fresh
}
if category != "" {
params["categories"] = category
}
if engines != "" {
params["engines"] = engines
}
cachePath := os.Getenv("BRAIN_SEARCH_CACHE")
if cachePath == "" {
cachePath = os.Getenv("HOME") + "/.cache/brain/web-search.sqlite"
}
cache, err := websearch.OpenCache(cachePath)
if err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
return 1
}
defer cache.Close()
key := websearch.CacheKey(query, params)
now := float64(time.Now().Unix())
if !refresh {
if cached, err := cache.Get(key, ttl, now); err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
return 1
} else if cached != nil {
out := websearch.Project(*cached, limit, websearch.DefaultSnippetChars)
out.Cached = true
return writeOut(out, jsonOut)
}
}
envPath := os.Getenv("BRAIN_SEARCH_ENV")
if envPath == "" {
envPath = os.Getenv("HOME") + "/.config/brain/search.env"
}
conf, err := websearch.LoadConfig(envPath)
if err != nil {
fmt.Fprintf(os.Stderr, "web/search: %v\n", err)
return 1
}
lockPath := cachePath + ".lock"
lock, err := os.OpenFile(lockPath, os.O_CREATE|os.O_RDWR, 0o600)
if err != nil {
fmt.Fprintf(os.Stderr, "web/search: lock: %v\n", err)
return 1
}
defer lock.Close()
if err := unix.Flock(int(lock.Fd()), unix.LOCK_EX); err != nil {
fmt.Fprintf(os.Stderr, "web/search: lock: %v\n", err)
return 1
}
defer unix.Flock(int(lock.Fd()), unix.LOCK_UN)
var payload websearch.Payload
attempts := 1 + len(websearch.RetryBackoff)
client := &http.Client{}
for attempt := 0; attempt < attempts; attempt++ {
last, err := cache.LastCall()
if err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
return 1
}
if delay := websearch.WaitFor(last, float64(time.Now().Unix()), websearch.MinInterval); delay > 0 {
time.Sleep(time.Duration(delay * float64(time.Second)))
}
if err := cache.MarkCall(float64(time.Now().Unix())); err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
return 1
}
payload, err = websearch.Fetch(client, conf, query, params, time.Duration(timeout)*time.Second)
if err != nil {
fmt.Fprintf(os.Stderr, "request failed: %v\n", err)
return 3
}
if websearch.Classify(payload) == websearch.StatusOK {
break
}
if attempt < len(websearch.RetryBackoff) {
time.Sleep(time.Duration(websearch.RetryBackoff[attempt] * float64(time.Second)))
}
}
if websearch.Classify(payload) == websearch.StatusOK {
if err := cache.Put(key, payload, float64(time.Now().Unix())); err != nil {
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
}
}
out := websearch.Project(payload, limit, websearch.DefaultSnippetChars)
code := writeOut(out, jsonOut)
if out.Status != websearch.StatusOK && code == 0 {
return 3
}
return code
}
func writeOut(out websearch.Output, jsonOut bool) int {
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
if err := enc.Encode(out); err != nil {
return 1
}
if out.Status != websearch.StatusOK {
return 3
}
return 0
}
fmt.Print(out.YAML())
if out.Status != websearch.StatusOK {
return 3
}
return 0
}
+15
View File
@@ -61,6 +61,21 @@ services:
restart: unless-stopped
stop_grace_period: 20s
# Optional local SearXNG (D3). Skip if BRAIN_SEARCH_URL already points at a
# live instance — do not run a second copy on that host.
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
searxng:
profiles: ["searxng"]
image: docker.io/searxng/searxng:2026.8.10-0a118066d
ports:
- "127.0.0.1:8888:8080"
environment:
SEARXNG_SECRET: ${SEARXNG_SECRET:-}
volumes:
- ./deploy/searxng/settings.yml:/etc/searxng/settings.yml:ro
- ./deploy/searxng/limiter.toml:/etc/searxng/limiter.toml:ro
restart: unless-stopped
volumes:
kb-model:
kb-var:
+7
View File
@@ -0,0 +1,7 @@
[botdetection.ip_lists]
# RFC1918 only. Do not copy a live instance egress IP into git.
pass_ip = [
"10.0.0.0/8",
"172.16.0.0/12",
"192.168.0.0/16",
]
+28
View File
@@ -0,0 +1,28 @@
use_default_settings: true
general:
instance_name: "2dph"
search:
formats:
- html
- json
suspended_times:
SearxEngineCaptcha: 300
SearxEngineTooManyRequests: 120
SearxEngineAccessDenied: 300
server:
limiter: true
image_proxy: false
# secret_key comes from SEARXNG_SECRET (never commit a real secret)
engines:
- name: bing
disabled: false
- name: google
disabled: false
- name: duckduckgo
disabled: false
- name: wikipedia
disabled: false
+6 -1
View File
@@ -6,5 +6,10 @@ Brain/ops/eSlider stack. Facts need proof or they are
- [PLAN.md](../PLAN.md) — decisions, execution order, open questions (v2)
- [design](design.md) — schema, deduction model, sources
- [Gitea issues](https://git.produktor.io/eSlider/2dph/issues) — work board (origin)
Published docs live here and mirror the project state.
Search: `bin/brain/search.go "query"` (HTTP: `bin/brain/serve.go`
`/health` `/search` `/get` `/stats` `/audit` `/ingest`). `--hop` is
not a walk; the flag errors until File/FROM_FILE edges exist.
Published docs live here and mirror the project state.
+24
View File
@@ -0,0 +1,24 @@
# Chat Import Pipeline
Plan: https://git.produktor.io/eSlider/brain-chats-import/issues/1
## Env vars (set in shell, never committed)
```
TELEGRAM_MCP_DIR
TELEGRAM_API_ID / TELEGRAM_API_HASH / TELEGRAM_PHONE
TELEGRAM_SESSION_STRING
ONLYOFFICE_URL / ONLYOFFICE_USER / ONLYOFFICE_PASS
OO_CLI (default: $HOME/go/bin/oo)
```
## Quick reference
```
./bin/chats/sync.go telegram --limit 100
./bin/chats/import.go
./bin/chats/facts.go
./bin/chats/apply.go --dry-run
```
JSONL → markdown only. Brain ingest is `bin/brain/index.go` (not a `chats index`).
+33
View File
@@ -0,0 +1,33 @@
# CRM association proof (oo CLI ↔ corpus)
Proven with `oo` (eslider/go-onlyoffice) against the OnlyOffice portal
(`office.produktor.io`). Portal CRM is the SSOT for company ↔ person ↔
project associations; the corpus SoT (`eslider/cv/projects/knowledge-mesh-seed.yaml`)
is the second, independent source. Facts that can be backed by both are
written to the brain under `root=facts` by `bin/facts/crm`.
## What was verified
- Logical counts (portal MySQL): 1300 contacts = 897 persons + 404 companies,
198 projects, 998 deals, 939 project↔contact links.
- Every client company linked to a project has ≥1 person underneath.
- Every person `company_id` resolves to an existing company.
- Corpus org list (9) maps 1:1 onto CRM companies:
ProProdukt SL / produktor.io, Dyvenia, Immowelt AG, WhereGroup,
Keynote SIGOS, D2S/SYSTEMS, GRID, Pack und Cup, Markets Platform.
- 78 person↔company association facts written to the brain
(`how=crm-crosscheck`, `type=association`). Recall@5 in `bin/kb/eval` = 1.0.
## Mistakes found
| # | Mistake | Fix |
|---|---------|-----|
| 1 | Duplicate legal entity `GoldenRatio.Exchange` (contact 759) vs `Golden Ratio Exchange` (763); 3 deals (211, 287, 559) were linked to 759 | `oo contacts merge 759 763` — 763 kept, 759 removed, deal links re-pointed to 763 |
| 2 | `env/`-wide: OnlyOffice creds file used wrong UX (user `eslider`, password with `$2` suffix) making `oo` auth fail | `.env` fixed to `eslider@gmail.com` + clean password; `.env` stays gitignored |
## Gates after fix
- `uv run python -m unittest discover -s bin/tools -t .` → 26 tests OK
- `bin/facts/audit self` + `bin/facts/audit db` → ok
- `bin/kb/eval` → recall@5 = 1.0
- `go test ./...` (bin/server + bin/watch) → ok
+7 -4
View File
@@ -17,14 +17,16 @@ as one consistent state.
## Deduction search
```
bin/kb/search "question"
bin/brain/search.go "question"
1. facts root — confirmed answers only → return with evidence links
2. info root — supporting narrative → snippets, marked (not confirmed)
3. web-search — second independent source → upgrade hypothesis to confirmed
(`web` block from `bin/web/search.go` when no facts hit; status `throttled`
is not evidence of absence; `--no-web` / `--root` skip it)
```
`--hop N` follows graph edges (sibling leaves under a heading, owning file,
`related:` files, vector-neighbour leaves) — the deduction walk.
`--hop` is not implemented yet (needs File/FROM_FILE edges). The flag is an
error; it is not a graph walk.
## Who / What / How / Where / When + evidence
@@ -46,7 +48,8 @@ Every assertion edge carries:
Content leafs: `sha256`, `observed_at`, `source_rev`, `confidence`. Stale = a
file changed on disk (git HEAD/mtime) after its last observed `source_rev`.
`File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person` records the history of every
content leaf.
content leaf. Commit records come from `bin/git/import.go` (go-git, no git
binary); conversion prints leafs, brain write is `bin/brain/index.go`.
`bin/facts/audit stale` flags leafs whose observed revision is behind the
corpus HEAD.
+52
View File
@@ -0,0 +1,52 @@
module github.com/eSlider/2dph
go 1.26
require (
github.com/LadybugDB/go-ladybug v0.17.0
github.com/arran4/golang-ical v0.3.5
github.com/chewxy/math32 v1.11.2
github.com/daulet/tokenizers v1.27.0
github.com/go-git/go-git/v5 v5.19.2
golang.org/x/sys v0.47.0
golang.org/x/text v0.40.0
modernc.org/sqlite v1.56.0
)
require (
dario.cat/mergo v1.0.0 // indirect
github.com/Microsoft/go-winio v0.6.2 // indirect
github.com/ProtonMail/go-crypto v1.1.6 // indirect
github.com/apache/arrow-go/v18 v18.6.0 // indirect
github.com/cloudflare/circl v1.6.3 // indirect
github.com/cyphar/filepath-securejoin v0.6.1 // indirect
github.com/dustin/go-humanize v1.0.1 // indirect
github.com/emirpasic/gods v1.18.1 // indirect
github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376 // indirect
github.com/go-git/go-billy/v5 v5.9.0 // indirect
github.com/goccy/go-json v0.10.6 // indirect
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 // indirect
github.com/google/flatbuffers v25.12.19+incompatible // indirect
github.com/google/uuid v1.6.0 // indirect
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99 // indirect
github.com/kevinburke/ssh_config v1.2.0 // indirect
github.com/klauspost/compress v1.18.5 // indirect
github.com/klauspost/cpuid/v2 v2.3.0 // indirect
github.com/mattn/go-isatty v0.0.24 // indirect
github.com/ncruces/go-strftime v1.0.0 // indirect
github.com/pierrec/lz4/v4 v4.1.26 // indirect
github.com/pjbgf/sha1cd v0.6.0 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/sergi/go-diff v1.3.2-0.20230802210424-5b0b94c5c0d3 // indirect
github.com/shopspring/decimal v1.4.0 // indirect
github.com/skeema/knownhosts v1.3.1 // indirect
github.com/xanzy/ssh-agent v0.3.3 // indirect
github.com/zeebo/xxh3 v1.1.0 // indirect
golang.org/x/crypto v0.53.0 // indirect
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f // indirect
golang.org/x/net v0.56.0 // indirect
gopkg.in/warnings.v0 v0.1.2 // indirect
modernc.org/libc v1.74.4 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
)
+184
View File
@@ -0,0 +1,184 @@
dario.cat/mergo v1.0.0 h1:AGCNq9Evsj31mOgNPcLyXc+4PNABt905YmuqPYYpBWk=
dario.cat/mergo v1.0.0/go.mod h1:uNxQE+84aUszobStD9th8a29P2fMDhsBdgRYvZOxGmk=
github.com/LadybugDB/go-ladybug v0.17.0 h1:RXDbkBjrbRmLdEbhGl4CLOIEzSt09gbP0n9UbKDEfwI=
github.com/LadybugDB/go-ladybug v0.17.0/go.mod h1:GeIXmE8XyF5TFS94NAuTag7vgCC+no/HTBMRA6Rd5Cs=
github.com/Microsoft/go-winio v0.5.2/go.mod h1:WpS1mjBmmwHBEWmogvA2mj8546UReBk4v8QkMxJ6pZY=
github.com/Microsoft/go-winio v0.6.2 h1:F2VQgta7ecxGYO8k3ZZz3RS8fVIXVxONVUPlNERoyfY=
github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU=
github.com/ProtonMail/go-crypto v1.1.6 h1:ZcV+Ropw6Qn0AX9brlQLAUXfqLBc7Bl+f/DmNxpLfdw=
github.com/ProtonMail/go-crypto v1.1.6/go.mod h1:rA3QumHc/FZ8pAHreoekgiAbzpNsfQAosU5td4SnOrE=
github.com/andybalholm/brotli v1.2.1 h1:R+f5xP285VArJDRgowrfb9DqL18yVK0gKAW/F+eTWro=
github.com/andybalholm/brotli v1.2.1/go.mod h1:rzTDkvFWvIrjDXZHkuS16NPggd91W3kUSvPlQ1pLaKY=
github.com/anmitsu/go-shlex v0.0.0-20200514113438-38f4b401e2be h1:9AeTilPcZAjCFIImctFaOjnTIavg87rW78vTPkQqLI8=
github.com/anmitsu/go-shlex v0.0.0-20200514113438-38f4b401e2be/go.mod h1:ySMOLuWl6zY27l47sB3qLNK6tF2fkHG55UZxx8oIVo4=
github.com/apache/arrow-go/v18 v18.6.0 h1:GX/Jyd3R7mCLiECAwY9FWbbaYblie2WXBSz4Sw8fNpM=
github.com/apache/arrow-go/v18 v18.6.0/go.mod h1:gm3MiPpY82fLYK5VKPB3WoJbsiLVDfT7flD5/vHReKw=
github.com/apache/thrift v0.22.0 h1:r7mTJdj51TMDe6RtcmNdQxgn9XcyfGDOzegMDRg47uc=
github.com/apache/thrift v0.22.0/go.mod h1:1e7J/O1Ae6ZQMTYdy9xa3w9k+XHWPfRvdPyJeynQ+/g=
github.com/armon/go-socks5 v0.0.0-20160902184237-e75332964ef5 h1:0CwZNZbxp69SHPdPJAN/hZIm0C4OItdklCFmMRWYpio=
github.com/armon/go-socks5 v0.0.0-20160902184237-e75332964ef5/go.mod h1:wHh0iHkYZB8zMSxRWpUBQtwG5a7fFgvEO+odwuTv2gs=
github.com/arran4/golang-ical v0.3.5 h1:bbz6ld4dC+MmCKiFfOd6SkmIGnhNMBACZ485ULh7p9A=
github.com/arran4/golang-ical v0.3.5/go.mod h1:OnguFgjN0Hmx8jzpmWcC+AkHio94ujmLHKoaef7xQh8=
github.com/chewxy/math32 v1.11.2 h1:IufN08Zwr1NKuWfY+4Tz55BcwKmyKKNdOP7KtumehnM=
github.com/chewxy/math32 v1.11.2/go.mod h1:dOB2rcuFrCn6UHrze36WSLVPKtzPMRAQvBvUwkSsLqs=
github.com/cloudflare/circl v1.6.3 h1:9GPOhQGF9MCYUeXyMYlqTR6a5gTrgR/fBLXvUgtVcg8=
github.com/cloudflare/circl v1.6.3/go.mod h1:2eXP6Qfat4O/Yhh8BznvKnJ+uzEoTQ6jVKJRn81BiS4=
github.com/cyphar/filepath-securejoin v0.6.1 h1:5CeZ1jPXEiYt3+Z6zqprSAgSWiggmpVyciv8syjIpVE=
github.com/cyphar/filepath-securejoin v0.6.1/go.mod h1:A8hd4EnAeyujCJRrICiOWqjS1AX0a9kM5XL+NwKoYSc=
github.com/daulet/tokenizers v1.27.0 h1:MmFYAEDFz69s/nNQfHg59DWqHz3v94m99kEZ/JbL+s4=
github.com/daulet/tokenizers v1.27.0/go.mod h1:YjFY1o1HGMyWkQgbXJDghhvke/yFDp2vGdIO2hYs4MQ=
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/elazarl/goproxy v1.7.2 h1:Y2o6urb7Eule09PjlhQRGNsqRfPmYI3KKQLFpCAV3+o=
github.com/elazarl/goproxy v1.7.2/go.mod h1:82vkLNir0ALaW14Rc399OTTjyNREgmdL2cVoIbS6XaE=
github.com/emirpasic/gods v1.18.1 h1:FXtiHYKDGKCW2KzwZKx0iC0PQmdlorYgdFG9jPXJ1Bc=
github.com/emirpasic/gods v1.18.1/go.mod h1:8tpGGwCnJ5H4r6BWwaV6OrWmMoPhUl5jm/FMNAnJvWQ=
github.com/gliderlabs/ssh v0.3.8 h1:a4YXD1V7xMF9g5nTkdfnja3Sxy1PVDCj1Zg4Wb8vY6c=
github.com/gliderlabs/ssh v0.3.8/go.mod h1:xYoytBv1sV0aL3CavoDuJIQNURXkkfPA/wxQ1pL1fAU=
github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376 h1:+zs/tPmkDkHx3U66DAb0lQFJrpS6731Oaa12ikc+DiI=
github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376/go.mod h1:an3vInlBmSxCcxctByoQdvwPiA7DTK7jaaFDBTtu0ic=
github.com/go-git/go-billy/v5 v5.9.0 h1:jItGXszUDRtR/AlferWPTMN4j38BQ88XnXKbilmmBPA=
github.com/go-git/go-billy/v5 v5.9.0/go.mod h1:jCnQMLj9eUgGU7+ludSTYoZL/GGmii14RxKFj7ROgHw=
github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399 h1:eMje31YglSBqCdIqdhKBW8lokaMrL3uTkpGYlE2OOT4=
github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399/go.mod h1:1OCfN199q1Jm3HZlxleg+Dw/mwps2Wbk9frAWm+4FII=
github.com/go-git/go-git/v5 v5.19.2 h1:wkfn7vOlUBu8ivAWKBWisTiwJK4jYHzTF8Ndv1LyGqY=
github.com/go-git/go-git/v5 v5.19.2/go.mod h1:QqCBE1EFN5ddFmrliLQ3/ntRCUjZU3EJuwuB/jWEHjk=
github.com/goccy/go-json v0.10.6 h1:p8HrPJzOakx/mn/bQtjgNjdTcN+/S6FcG2CTtQOrHVU=
github.com/goccy/go-json v0.10.6/go.mod h1:oq7eo15ShAhp70Anwd5lgX2pLfOS3QCiwU/PULtXL6M=
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 h1:f+oWsMOmNPc8JmEHVZIycC7hBoQxHH9pNKQORJNozsQ=
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8/go.mod h1:wcDNUvekVysuuOpQKo3191zZyTpiI6se1N1ULghS0sw=
github.com/google/flatbuffers v25.12.19+incompatible h1:haMV2JRRJCe1998HeW/p0X9UaMTK6SDo0ffLn2+DbLs=
github.com/google/flatbuffers v25.12.19+incompatible/go.mod h1:1AeVuKshWv4vARoZatz6mlQ0JxURH0Kv5+zNeJKJCa8=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3 h1:LMLX+LgTNWpfvCBdFebv6EsYotImrt/Ppc5cXIriCSo=
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3/go.mod h1:jl5iWTm0/hd5PjEYEOuwAJ57L/CibdZfrqZ5XA5GrCk=
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99 h1:BQSFePA1RWJOlocH6Fxy8MmwDt+yVQYULKfN0RoTN8A=
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99/go.mod h1:1lJo3i6rXxKeerYnT8Nvf0QmHCRC1n8sfWVwXF2Frvo=
github.com/kevinburke/ssh_config v1.2.0 h1:x584FjTGwHzMwvHx18PXxbBVzfnxogHaAReU4gf13a4=
github.com/kevinburke/ssh_config v1.2.0/go.mod h1:CT57kijsi8u/K/BOFA39wgDQJ9CxiF4nAY/ojJ6r6mM=
github.com/klauspost/compress v1.18.5 h1:/h1gH5Ce+VWNLSWqPzOVn6XBO+vJbCNGvjoaGBFW2IE=
github.com/klauspost/compress v1.18.5/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
github.com/klauspost/cpuid/v2 v2.3.0 h1:S4CRMLnYUhGeDFDqkGriYKdfoFlDnMtqTiI/sFzhA9Y=
github.com/klauspost/cpuid/v2 v2.3.0/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0=
github.com/kr/pretty v0.1.0/go.mod h1:dAy3ld7l9f0ibDNOQOHHMYYIIbhfbHSm3C4ZsoJORNo=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/pty v1.1.1/go.mod h1:pFQYn66WHrOpPYNljwOMqo10TkYh1fy3cYio2l3bCsQ=
github.com/kr/text v0.1.0/go.mod h1:4Jbv+DJW3UT/LiOwJeYQe1efqtUx/iVham/4vfdArNI=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
github.com/mattn/go-isatty v0.0.24/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
github.com/ncruces/go-strftime v1.0.0 h1:HMFp8mLCTPp341M/ZnA4qaf7ZlsbTc+miZjCLOFAw7w=
github.com/ncruces/go-strftime v1.0.0/go.mod h1:Fwc5htZGVVkseilnfgOVb9mKy6w1naJmn9CehxcKcls=
github.com/onsi/gomega v1.34.1 h1:EUMJIKUjM8sKjYbtxQI9A4z2o+rruxnzNvpknOXie6k=
github.com/onsi/gomega v1.34.1/go.mod h1:kU1QgUvBDLXBJq618Xvm2LUX6rSAfRaFRTcdOeDLwwY=
github.com/pierrec/lz4/v4 v4.1.26 h1:GrpZw1gZttORinvzBdXPUXATeqlJjqUG/D87TKMnhjY=
github.com/pierrec/lz4/v4 v4.1.26/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pjbgf/sha1cd v0.6.0 h1:3WJ8Wz8gvDz29quX1OcEmkAlUg9diU4GxJHqs0/XiwU=
github.com/pjbgf/sha1cd v0.6.0/go.mod h1:lhpGlyHLpQZoxMv8HcgXvZEhcGs0PG/vsZnEJ7H0iCM=
github.com/pkg/errors v0.9.1 h1:FEBLx1zS214owpjy7qsBeixbURkuhQAwrK5UwLGTwt4=
github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo=
github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ=
github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc=
github.com/sergi/go-diff v1.3.2-0.20230802210424-5b0b94c5c0d3 h1:n661drycOFuPLCN3Uc8sB6B/s6Z4t2xvBgU1htSHuq8=
github.com/sergi/go-diff v1.3.2-0.20230802210424-5b0b94c5c0d3/go.mod h1:A0bzQcvG0E7Rwjx0REVgAGH58e96+X0MeOfepqsbeW4=
github.com/shopspring/decimal v1.4.0 h1:bxl37RwXBklmTi0C79JfXCEBD1cqqHt0bbgBAGFp81k=
github.com/shopspring/decimal v1.4.0/go.mod h1:gawqmDU56v4yIKSwfBSFip1HdCCXN8/+DMd9qYNcwME=
github.com/sirupsen/logrus v1.7.0/go.mod h1:yWOB1SBYBC5VeMP7gHvWumXLIWorT60ONWic61uBYv0=
github.com/skeema/knownhosts v1.3.1 h1:X2osQ+RAjK76shCbvhHHHVl3ZlgDm8apHEHFqRjnBY8=
github.com/skeema/knownhosts v1.3.1/go.mod h1:r7KTdC8l4uxWRyK2TpQZ/1o5HaSzh06ePQNxPwTcfiY=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
github.com/stretchr/testify v1.2.2/go.mod h1:a8OnRcib4nhh0OaRAV+Yts87kKdq0PP7pXfy6kDkUVs=
github.com/stretchr/testify v1.4.0/go.mod h1:j7eGeouHqKxXV5pUuKE4zz7dFj8WfuZ+81PSLYec5m4=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/xanzy/ssh-agent v0.3.3 h1:+/15pJfg/RsTxqYcX6fHqOXZwwMP+2VyYWJeWM2qQFM=
github.com/xanzy/ssh-agent v0.3.3/go.mod h1:6dzNDKs0J9rVPHPhaGCukekBHKqfl+L3KghI1Bc68Uw=
github.com/zeebo/assert v1.3.0 h1:g7C04CbJuIDKNPFHmsk4hwZDO5O+kntRxzaUoNXj+IQ=
github.com/zeebo/assert v1.3.0/go.mod h1:Pq9JiuJQpG8JLJdtkwrJESF0Foym2/D9XMU5ciN/wJ0=
github.com/zeebo/xxh3 v1.1.0 h1:s7DLGDK45Dyfg7++yxI0khrfwq9661w9EN78eP/UZVs=
github.com/zeebo/xxh3 v1.1.0/go.mod h1:IisAie1LELR4xhVinxWS5+zf1lA4p0MW4T+w+W07F5s=
golang.org/x/crypto v0.0.0-20220622213112-05595931fe9d/go.mod h1:IxCIyHEi3zRg3s0A5j5BB6A9Jmi73HwBIUl50j+osU4=
golang.org/x/crypto v0.53.0 h1:QZ4Muo8THX6CizN2vPPd5fBGHyogrdK9fG4wLPFUsto=
golang.org/x/crypto v0.53.0/go.mod h1:DNLU434OwVakk9PzuwV8w62mAJpRJL3vsgcfp4Qnsio=
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f h1:W3F4c+6OLc6H2lb//N1q4WpJkhzJCK5J6kUi1NTVXfM=
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f/go.mod h1:J1xhfL/vlindoeF/aINzNzt2Bket5bjo9sdOYzOsU80=
golang.org/x/mod v0.37.0 h1:vF1DjpVEshcIqoEaauuHebaLk1O1forxjxBaVn884JQ=
golang.org/x/mod v0.37.0/go.mod h1:m8S8VeM9r4dzDwjrKO0a1sZP3YjeMamRRlD+fmR2Q/0=
golang.org/x/net v0.0.0-20211112202133-69e39bad7dc2/go.mod h1:9nx3DQGgdP8bBQD5qxJ1jj9UTztislL4KSBs9R2vV5Y=
golang.org/x/net v0.56.0 h1:Rw8j/hFzGvJUZwNBXnAtf5sVDVt+65SK2C7IxCxZt5o=
golang.org/x/net v0.56.0/go.mod h1:D3Ku6r+V6JROoZK144D2XfMHFcMq/0zSfLelVTCFKec=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.0.0-20191026070338-33540a1f6037/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20201119102817-f84b799fce68/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210124154548-22da62e12c0c/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210423082822-04245dca01da/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210615035016-665e8c7367d1/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.0.0-20220715151400-c0bba94af5f8/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.44.0 h1:0rLvDRCtNj0gZkyIXhCyOb2OAzEhLVqc4B+hrsBhrmc=
golang.org/x/term v0.44.0/go.mod h1:7ze4MdzUzLXpSAoFP1H0bOI9aXDqveSvatT5vKcFh2Y=
golang.org/x/text v0.3.6/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q=
golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA=
gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4=
gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
gopkg.in/warnings.v0 v0.1.2 h1:wFXVbFY8DY5/xOe1ECiWdKCzZlxgshcYVNkBHstARME=
gopkg.in/warnings.v0 v0.1.2/go.mod h1:jksf8JmL6Qr/oQM2OXTHunEvvTAsrWBLb6OOjuVWRNI=
gopkg.in/yaml.v2 v2.2.2/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
gopkg.in/yaml.v2 v2.4.0/go.mod h1:RDklbk79AGWmwhnvt/jBztapEOGDOx6ZbXqjP6csGnQ=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI=
modernc.org/cc/v4 v4.29.1/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
modernc.org/ccgo/v4 v4.34.6 h1:sBgfIwyN0TQ9C5hwIeuqyeAKyMWnbvj2fvpF4L11uzU=
modernc.org/ccgo/v4 v4.34.6/go.mod h1:SZ8YcN9NG7XVsQYdm6jYBvi8PQP1qi+kqB6OhjqI3Fk=
modernc.org/fileutil v1.4.0 h1:j6ZzNTftVS054gi281TyLjHPp6CPHr2KCxEXjEbD6SM=
modernc.org/fileutil v1.4.0/go.mod h1:EqdKFDxiByqxLk8ozOxObDSfcVOv/54xDs/DUHdvCUU=
modernc.org/gc/v2 v2.6.5 h1:nyqdV8q46KvTpZlsw66kWqwXRHdjIlJOhG6kxiV/9xI=
modernc.org/gc/v2 v2.6.5/go.mod h1:YgIahr1ypgfe7chRuJi2gD7DBQiKSLMPgBQe9oIiito=
modernc.org/gc/v3 v3.1.4 h1:2g65LGVSmFQrXeITAw97x7hCRvZFcyE1uDP+7Vng7JI=
modernc.org/gc/v3 v3.1.4/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
modernc.org/goabi0 v0.2.0 h1:HvEowk7LxcPd0eq6mVOAEMai46V+i7Jrj13t4AzuNks=
modernc.org/goabi0 v0.2.0/go.mod h1:CEFRnnJhKvWT1c1JTI3Avm+tgOWbkOu5oPA8eH8LnMI=
modernc.org/libc v1.74.4 h1:fX1Omw4o2/1C2iRkkIsrQTasJQldLhRmuPreXLoWs9k=
modernc.org/libc v1.74.4/go.mod h1:eeQAS9W3sZeKYMFubydxJpII9ybHWshk+7or7bLG9co=
modernc.org/mathutil v1.7.1 h1:GCZVGXdaN8gTqB1Mf/usp1Y/hSqgI2vAGGP4jZMCxOU=
modernc.org/mathutil v1.7.1/go.mod h1:4p5IwJITfppl0G4sUEDtCr4DthTaT47/N3aT6MhfgJg=
modernc.org/memory v1.11.0 h1:o4QC8aMQzmcwCK3t3Ux/ZHmwFPzE6hf2Y5LbkRs+hbI=
modernc.org/memory v1.11.0/go.mod h1:/JP4VbVC+K5sU2wZi9bHoq2MAkCnrt2r98UGeSK7Mjw=
modernc.org/opt v0.2.0 h1:tGyef5ApycA7FSEOMraay9SaTk5zmbx7Tu+cJs4QKZg=
modernc.org/opt v0.2.0/go.mod h1:03fq9lsNfvkYSfxrfUhZCWPk1lm4cq4N+Bh//bEtgns=
modernc.org/sortutil v1.2.1 h1:+xyoGf15mM3NMlPDnFqrteY07klSFxLElE2PVuWIJ7w=
modernc.org/sortutil v1.2.1/go.mod h1:7ZI3a3REbai7gzCLcotuw9AC4VZVpYMjDzETGsSMqJE=
modernc.org/sqlite v1.56.0 h1:/D8e2RfFqoy/Zc6PuC76U28zFwmI/sYx1Kjm4yEn9e0=
modernc.org/sqlite v1.56.0/go.mod h1:yCJ2cmAaIkHQ25oXWrF8H4O1lIfPYPR26yCEDj2P3pQ=
modernc.org/strutil v1.2.1 h1:UneZBkQA+DX2Rp35KcM69cSsNES9ly8mQWD71HKlOA0=
modernc.org/strutil v1.2.1/go.mod h1:EHkiggD70koQxjVdSBM3JKM7k6L0FbGE5eymy9i3B9A=
modernc.org/token v1.1.0 h1:Xl7Ap9dKaEs5kLoOQeQmPWevfnk/DM5qcLcYlA8ys6Y=
modernc.org/token v1.1.0/go.mod h1:UGzOrNV1mAFSEB63lOFHIpNRUVMvYTc6yu1SMY/XTDM=
+3
View File
@@ -0,0 +1,3 @@
go 1.26
use .
+100
View File
@@ -0,0 +1,100 @@
//go:build cgo && system_ladybug
package brain
import (
"fmt"
"os"
"path/filepath"
"strings"
lbug "github.com/LadybugDB/go-ladybug"
)
var (
db *lbug.Database
conn *lbug.Connection
)
func repoRoot() string {
// Try KB_ROOT env, then walk up from binary
if v := os.Getenv("KB_ROOT"); v != "" {
return v
}
self, err := os.Executable()
if err == nil {
dir := filepath.Dir(self)
for i := 0; i < 5; i++ {
if _, err := os.Stat(filepath.Join(dir, "var")); err == nil {
return dir
}
if _, err := os.Stat(filepath.Join(dir, ".git")); err == nil {
return dir
}
parent := filepath.Dir(dir)
if parent == dir {
break
}
dir = parent
}
}
return "."
}
func dbPath() string {
return filepath.Join(repoRoot(), "var", "kb.lbug")
}
func openBrain() error {
return openWithSandbox(eps())
}
func openWithSandbox(epsv string) error {
cfg := lbug.DefaultSystemConfig()
cfg.MaxNumThreads = 8
cfg.BufferPoolSize = 1 << 30 // 1GB
var err error
db, err = lbug.OpenDatabase(dbPath(), cfg)
if err != nil {
return fmt.Errorf("OpenDatabase: %w", err)
}
conn, err = lbug.OpenConnection(db)
if err != nil {
closeBrain()
return fmt.Errorf("OpenConnection: %w", err)
}
// Session settings need a live connection; running this before
// OpenConnection dereferenced a nil *Connection.
if epsv != "" {
if strings.ContainsAny(epsv, "'\\") {
closeBrain()
return fmt.Errorf("SET STREAM_SANDBOX: invalid value")
}
if _, err := conn.Query("SET STREAM_SANDBOX = '" + epsv + "'"); err != nil {
closeBrain()
return fmt.Errorf("SET STREAM_SANDBOX: %w", err)
}
}
if _, err := conn.Query("LOAD EXTENSION FTS"); err != nil {
closeBrain()
return fmt.Errorf("LOAD EXTENSION FTS: %w", err)
}
if _, err := conn.Query("LOAD EXTENSION VECTOR"); err != nil {
closeBrain()
return fmt.Errorf("LOAD EXTENSION VECTOR: %w", err)
}
return nil
}
func closeBrain() {
if conn != nil {
conn.Close()
conn = nil
}
if db != nil {
db.Close()
db = nil
}
}
+3
View File
@@ -0,0 +1,3 @@
// Package brain is deduction search over Ladybug (FTS + HNSW).
// Query/embed code that needs cgo lives behind the system_ladybug tag.
package brain
+159
View File
@@ -0,0 +1,159 @@
//go:build cgo && system_ladybug
package brain
import (
"bytes"
"context"
"encoding/json"
"fmt"
"github.com/eSlider/2dph/internal/brain/rank"
)
// Ready opens the Ladybug file for the life of the serve process.
func Ready() error {
return openBrain()
}
// HTTP is the in-process API used by bin/brain/serve.go.
type HTTP struct{}
func (HTTP) Search(ctx context.Context, query string, limit int) ([]byte, error) {
hits, err := searchHits(query, "", "", limit)
if err != nil {
return nil, err
}
for i := range hits {
if hits[i].Text != "" {
runes := []rune(hits[i].Text)
if len(runes) > 280 {
runes = runes[:280]
}
hits[i].Snippet = string(runes)
}
}
webOut := rank.Deduce(hits, query, "", false, func(q string) rank.SecondSource {
return lookupWeb(ctx, q)
})
var buf bytes.Buffer
enc := json.NewEncoder(&buf)
enc.SetEscapeHTML(false)
if err := enc.Encode(toJSONOut(hits, query, "", webOut)); err != nil {
return nil, err
}
return buf.Bytes(), nil
}
func (HTTP) Get(_ context.Context, id string, body bool) ([]byte, error) {
if conn == nil {
return nil, fmt.Errorf("brain not open")
}
stmt, err := conn.Prepare(
"MATCH (l:Leaf {id:$id}) RETURN l.id, l.text, l.root, l.confidence, l.source, l.type",
)
if err != nil {
return nil, err
}
defer stmt.Close()
res, err := conn.Execute(stmt, map[string]any{"id": id})
if err != nil {
return nil, err
}
if !res.HasNext() {
return nil, fmt.Errorf("no leaf %s", id)
}
row, err := res.Next()
if err != nil {
return nil, err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 6 {
return nil, fmt.Errorf("leaf row")
}
out := map[string]any{
"id": fmt.Sprint(vals[0]),
"root": fmt.Sprint(vals[2]),
"confidence": fmt.Sprint(vals[3]),
"source": fmt.Sprint(vals[4]),
"type": fmt.Sprint(vals[5]),
}
if body {
out["text"] = fmt.Sprint(vals[1])
}
return json.Marshal(out)
}
func (HTTP) Stats(context.Context) ([]byte, error) {
if conn == nil {
return nil, fmt.Errorf("brain not open")
}
res, err := conn.Query("MATCH (l:Leaf) RETURN l.root, count(*)")
if err != nil {
return nil, err
}
byRoot := map[string]int{}
total := 0
for res.HasNext() {
row, err := res.Next()
if err != nil {
return nil, err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 2 {
continue
}
n := int(asInt(vals[1]))
byRoot[fmt.Sprint(vals[0])] = n
total += n
}
return json.Marshal(map[string]any{"total": total, "by_root": byRoot, "db": dbPath()})
}
func (HTTP) Audit(context.Context) ([]byte, error) {
if conn == nil {
return nil, fmt.Errorf("brain not open")
}
res, err := conn.Query("MATCH (l:Leaf) RETURN l.root, l.confidence, count(*)")
if err != nil {
return nil, err
}
var rows []map[string]any
for res.HasNext() {
row, err := res.Next()
if err != nil {
return nil, err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 3 {
continue
}
rows = append(rows, map[string]any{
"root": fmt.Sprint(vals[0]),
"confidence": fmt.Sprint(vals[1]),
"count": asInt(vals[2]),
})
}
return json.Marshal(map[string]any{"status": "ok", "by_confidence": rows})
}
func (HTTP) Ingest(context.Context) ([]byte, error) {
return json.Marshal(map[string]any{
"mode": "rebuild",
"command": "bin/brain/index.go --rebuild",
"add": "v2",
})
}
func asInt(v any) int64 {
switch n := v.(type) {
case int64:
return n
case int:
return int64(n)
case float64:
return int64(n)
default:
return 0
}
}
+238
View File
@@ -0,0 +1,238 @@
//go:build cgo && system_ladybug
// StaticModel wraps the potion-multilingual-128m embedding model.
//
// Mirrors model2vec.StaticModel: tokenizer (daulet Unigram) + safetensors matrix.
// Embed(text) applies the same preprocessing: median_token_length pre-truncation,
// add_special_tokens=false, drop unk (id=1), truncate to 512, mean pool, L2 normalize +1e-32.
package brain
import (
"encoding/json"
"fmt"
"io"
"math"
"os"
"path/filepath"
"sort"
"github.com/chewxy/math32"
"github.com/daulet/tokenizers"
)
type StaticModel struct {
tok *tokenizers.Tokenizer
mat []float32 // row-major: vocab_size x 128
medianLen int
vocabSize int
dim int
}
func loadModel() (*StaticModel, error) {
dir, err := modelDir()
if err != nil {
return nil, err
}
tok, err := tokenizers.FromFile(filepath.Join(dir, "tokenizer.json"))
if err != nil {
return nil, fmt.Errorf("tokenizer: %w", err)
}
mat, vocabSize, dim, err := loadMatrix(filepath.Join(dir, "model.safetensors"))
if err != nil {
return nil, fmt.Errorf("safetensors: %w", err)
}
median := medianTokenLength(filepath.Join(dir, "tokenizer.json"))
return &StaticModel{
tok: tok,
mat: mat,
vocabSize: vocabSize,
dim: dim,
medianLen: median,
}, nil
}
func (m *StaticModel) Close() error {
if m.tok != nil {
m.tok.Close()
m.tok = nil
}
return nil
}
func (m *StaticModel) Embed(text string) ([]float64, error) {
const maxLen = 512
if m.medianLen > 0 {
maxChars := maxLen * m.medianLen
runes := []rune(text)
if len(runes) > maxChars {
text = string(runes[:maxChars])
}
}
ids, _, err := m.tok.EncodeErr(text, false)
if err != nil {
return nil, fmt.Errorf("encode: %w", err)
}
filtered := make([]uint32, 0, len(ids))
for _, id := range ids {
if id != 1 {
filtered = append(filtered, id)
}
if len(filtered) >= maxLen {
break
}
}
if len(filtered) == 0 {
return make([]float64, m.dim), nil
}
acc := make([]float32, m.dim)
for _, id := range filtered {
if int(id) >= m.vocabSize {
continue
}
off := int(id) * m.dim
for d := 0; d < m.dim; d++ {
acc[d] += m.mat[off+d]
}
}
inv := 1.0 / float32(len(filtered))
for d := 0; d < m.dim; d++ {
acc[d] *= inv
}
var norm float32
for d := 0; d < m.dim; d++ {
norm += acc[d] * acc[d]
}
norm = math32.Sqrt(norm) + 1e-32
for d := 0; d < m.dim; d++ {
acc[d] /= norm
}
out := make([]float64, m.dim)
for d := 0; d < m.dim; d++ {
out[d] = float64(acc[d])
}
return out, nil
}
func medianTokenLength(tokenizerPath string) int {
data, err := os.ReadFile(tokenizerPath)
if err != nil {
return 0
}
var parsed struct {
Model struct {
Vocab [][]json.RawMessage `json:"vocab"`
} `json:"model"`
}
if err := json.Unmarshal(data, &parsed); err != nil {
return 0
}
vocab := parsed.Model.Vocab
if len(vocab) == 0 {
return 0
}
lengths := make([]int, 0, len(vocab))
for _, pair := range vocab {
if len(pair) < 1 {
continue
}
var tok string
if err := json.Unmarshal(pair[0], &tok); err != nil {
continue
}
lengths = append(lengths, len([]rune(tok)))
}
if len(lengths) == 0 {
return 0
}
sort.Ints(lengths)
return lengths[len(lengths)/2]
}
func loadMatrix(path string) ([]float32, int, int, error) {
f, err := os.Open(path)
if err != nil {
return nil, 0, 0, err
}
defer f.Close()
var hdrLen uint64
if err := binaryRead(f, &hdrLen); err != nil {
return nil, 0, 0, err
}
hdrBytes := make([]byte, hdrLen)
if _, err := io.ReadFull(f, hdrBytes); err != nil {
return nil, 0, 0, err
}
var hdr struct {
Embeddings struct {
Dtype string `json:"dtype"`
Shape []int `json:"shape"`
Offset []uint64 `json:"data_offsets"`
} `json:"embeddings"`
}
if err := json.Unmarshal(hdrBytes, &hdr); err != nil {
return nil, 0, 0, err
}
if hdr.Embeddings.Dtype != "F32" {
return nil, 0, 0, fmt.Errorf("unsupported dtype %s", hdr.Embeddings.Dtype)
}
if len(hdr.Embeddings.Shape) != 2 {
return nil, 0, 0, fmt.Errorf("expected 2D shape, got %v", hdr.Embeddings.Shape)
}
vocabSize := hdr.Embeddings.Shape[0]
dim := hdr.Embeddings.Shape[1]
if len(hdr.Embeddings.Offset) != 2 {
return nil, 0, 0, fmt.Errorf("bad offsets")
}
start := hdr.Embeddings.Offset[0]
end := hdr.Embeddings.Offset[1]
size := end - start
if size != uint64(vocabSize*dim*4) {
return nil, 0, 0, fmt.Errorf("size mismatch")
}
if _, err := f.Seek(int64(8+hdrLen+start), io.SeekStart); err != nil {
return nil, 0, 0, err
}
buf := make([]byte, size)
if _, err := io.ReadFull(f, buf); err != nil {
return nil, 0, 0, err
}
mat := make([]float32, vocabSize*dim)
for i := 0; i < len(mat); i++ {
off := i * 4
mat[i] = math.Float32frombits(
uint32(buf[off]) |
uint32(buf[off+1])<<8 |
uint32(buf[off+2])<<16 |
uint32(buf[off+3])<<24,
)
}
return mat, vocabSize, dim, nil
}
func binaryRead(r io.Reader, v any) error {
switch p := v.(type) {
case *uint64:
var b [8]byte
if _, err := io.ReadFull(r, b[:]); err != nil {
return err
}
*p = uint64(b[0]) | uint64(b[1])<<8 | uint64(b[2])<<16 | uint64(b[3])<<24 |
uint64(b[4])<<32 | uint64(b[5])<<40 | uint64(b[6])<<48 | uint64(b[7])<<56
}
return nil
}
+59
View File
@@ -0,0 +1,59 @@
package brain
import (
"fmt"
"os"
"path/filepath"
"strings"
)
func modelDir() (string, error) {
// 1. Explicit env
if v := os.Getenv("KBSEARCH_MODEL"); v != "" {
return v, nil
}
// 2. Next to the binary (dev or installed)
self, err := os.Executable()
if err == nil {
if dir, err := filepath.EvalSymlinks(filepath.Dir(self)); err == nil {
cand := filepath.Join(dir, "potion-multilingual-128m")
if st, err := os.Stat(cand); err == nil && st.IsDir() {
return cand, nil
}
}
}
// 3. Repo root lib/ (where other scripts expect it)
if v := os.Getenv("KB_ROOT"); v != "" {
cand := filepath.Join(v, "lib", "potion-multilingual-128m")
if st, err := os.Stat(cand); err == nil && st.IsDir() {
return cand, nil
}
}
// 4. HF cache (new layout: models--*/snapshots/*)
if v := os.Getenv("HF_HOME"); v != "" {
base := filepath.Join(v, "hub")
if entries, err := os.ReadDir(base); err == nil {
for _, e := range entries {
if strings.HasPrefix(e.Name(), "models--") {
snapDir := filepath.Join(base, e.Name(), "snapshots")
if snaps, err := os.ReadDir(snapDir); err == nil {
for _, s := range snaps {
cand := filepath.Join(snapDir, s.Name())
if st, _ := os.Stat(cand); st != nil && st.IsDir() {
return cand, nil
}
}
}
}
}
}
}
// 5. Legacy HF cache (symlinked model dir)
if v := os.Getenv("HF_HOME"); v != "" {
cand := filepath.Join(v, "potion-multilingual-128m")
if st, err := os.Stat(cand); err == nil && st.IsDir() {
return cand, nil
}
}
return "", fmt.Errorf("model not found (set KBSEARCH_MODEL or KB_ROOT, or download to HF cache)")
}
+75
View File
@@ -0,0 +1,75 @@
package rank
import (
"fmt"
"strconv"
"strings"
)
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--json] [--no-web]
bin/brain/search.go serve [port]
bin/brain/search.go --list-model`
type Options struct {
Query string
Root string
Repo string
Limit int
JSONOut bool
ListModel bool
NoWeb bool
}
// ParseArgs reads flags. Unknown flags are an error: silently dropping them
// meant `--hop 1` vanished and its argument `1` was appended to the query.
// --hop is recognised so it cannot be swallowed; it is not implemented until
// File/FROM_FILE edges exist.
func ParseArgs(args []string) (Options, error) {
opt := Options{Limit: 20}
var queryArgs []string
for i := 0; i < len(args); i++ {
arg := args[i]
wantsValue := arg == "--root" || arg == "--repo" || arg == "-n" || arg == "--hop"
if wantsValue && i+1 >= len(args) {
return opt, fmt.Errorf("%s needs a value", arg)
}
switch arg {
case "--root":
i++
opt.Root = args[i]
if opt.Root != "facts" && opt.Root != "info" {
return opt, fmt.Errorf("--root must be facts or info, got %q", opt.Root)
}
case "--repo":
i++
opt.Repo = args[i]
case "-n":
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 1 {
return opt, fmt.Errorf("-n must be a positive integer, got %q", args[i])
}
opt.Limit = n
case "--hop":
return opt, fmt.Errorf("--hop is not implemented yet (needs File/FROM_FILE edges)")
case "--json":
opt.JSONOut = true
case "--no-web":
opt.NoWeb = true
case "--list-model":
opt.ListModel = true
default:
if strings.HasPrefix(arg, "-") {
return opt, fmt.Errorf("unknown flag %q", arg)
}
queryArgs = append(queryArgs, arg)
}
}
opt.Query = strings.TrimSpace(strings.Join(queryArgs, " "))
if opt.Query == "" && !opt.ListModel {
return opt, fmt.Errorf("no query given")
}
return opt, nil
}

Some files were not shown because too many files have changed in this diff Show More