Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
aca05626bd | ||
|
|
39ae2abe8d | ||
|
|
ba5cc3a6e2 | ||
|
|
15d59054ff | ||
|
|
3f30052ea8 | ||
|
|
c1ee920b0a | ||
|
|
cec0161ff6 | ||
|
|
46310f8773 | ||
|
|
7e0f3c9e06 | ||
|
|
f14025304e | ||
|
|
dd6d7e9395 | ||
|
|
68d478224f | ||
|
|
117f3c2cfd | ||
|
|
140d86a4b9 |
@@ -19,6 +19,10 @@ jobs:
|
|||||||
with:
|
with:
|
||||||
fetch-depth: 0
|
fetch-depth: 0
|
||||||
|
|
||||||
|
- uses: actions/setup-go@v5
|
||||||
|
with:
|
||||||
|
go-version-file: go.mod
|
||||||
|
|
||||||
- name: Install uv
|
- name: Install uv
|
||||||
uses: astral-sh/setup-uv@v6
|
uses: astral-sh/setup-uv@v6
|
||||||
with:
|
with:
|
||||||
@@ -32,16 +36,20 @@ jobs:
|
|||||||
bash -n bin/db/psql-yq
|
bash -n bin/db/psql-yq
|
||||||
bash -n bin/db/ssh-tunnel
|
bash -n bin/db/ssh-tunnel
|
||||||
bash -n bin/docker-entrypoint
|
bash -n bin/docker-entrypoint
|
||||||
|
bash -n bin/kb/search
|
||||||
|
|
||||||
- name: Python unit tests (offline, vendored tools)
|
- name: Python unit tests (offline, vendored tools)
|
||||||
run: |
|
run: |
|
||||||
uv run python -m unittest discover -s bin/tools -t .
|
uv run python -m unittest discover -s bin/tools -t .
|
||||||
|
|
||||||
- name: Go tests (server + watch packages)
|
- name: Go tests (root module, no ladybug cgo)
|
||||||
run: |
|
run: |
|
||||||
go vet ./...
|
go vet ./...
|
||||||
go test ./... -count=1
|
go test ./... -count=1
|
||||||
|
|
||||||
|
- name: brain ranking tests (no cgo / no ladybug)
|
||||||
|
run: go test ./internal/brain/rank -count=1
|
||||||
|
|
||||||
- name: facts/audit self (lexicon consistency, no network)
|
- name: facts/audit self (lexicon consistency, no network)
|
||||||
run: |
|
run: |
|
||||||
./bin/facts/audit self 2>/dev/null || echo "audit: not yet implemented; gate skipped"
|
./bin/facts/audit self 2>/dev/null || echo "audit: not yet implemented; gate skipped"
|
||||||
|
|||||||
@@ -10,3 +10,4 @@ __pycache__/
|
|||||||
.env
|
.env
|
||||||
.secrets/
|
.secrets/
|
||||||
lib-ladybug/
|
lib-ladybug/
|
||||||
|
go.work.local
|
||||||
|
|||||||
@@ -24,7 +24,7 @@ Read first: [PLAN](PLAN.md) → [docs](docs/).
|
|||||||
2. **Read-only data sources.** Ladybug `var/kb.lbug` and Postgres are opened
|
2. **Read-only data sources.** Ladybug `var/kb.lbug` and Postgres are opened
|
||||||
read-only for queries. Index rebuilds write to `var/` (gitignored).
|
read-only for queries. Index rebuilds write to `var/` (gitignored).
|
||||||
3. **PII.** `brain-test`, `cs_brain` client data is never read or quoted.
|
3. **PII.** `brain-test`, `cs_brain` client data is never read or quoted.
|
||||||
4. **No main pushes.** Feature branches + PR via `gh`; CI must be green.
|
4. **No main pushes.** Feature branches + GitHub PR (`gh`); CI (Actions) must be green. Work board: [Gitea issues](https://git.produktor.io/eSlider/2dph/issues).
|
||||||
5. **TDD.** Failing test before tool code. Unit tests run offline against
|
5. **TDD.** Failing test before tool code. Unit tests run offline against
|
||||||
fixtures; network/db calls are wrapped.
|
fixtures; network/db calls are wrapped.
|
||||||
6. **docs reflect behaviour.** Any change updates `docs/` + `PLAN.md` status.
|
6. **docs reflect behaviour.** Any change updates `docs/` + `PLAN.md` status.
|
||||||
@@ -35,11 +35,16 @@ Read first: [PLAN](PLAN.md) → [docs](docs/).
|
|||||||
PLAN.md decisions + execution + open questions
|
PLAN.md decisions + execution + open questions
|
||||||
docs/ published docs
|
docs/ published docs
|
||||||
skills/ in-project agent skills (vendored, no external links)
|
skills/ in-project agent skills (vendored, no external links)
|
||||||
bin/ self-describing tools bin/{subject}/{method} (shebang)
|
bin/ self-describing tools bin/{subject}/{method}.go (shebang)
|
||||||
bin/serve.go async Go HTTP server entry (self-executing go run shebang)
|
bin/brain/ search.go serve.go index.go get.go stats.go eval.go watch.go
|
||||||
bin/watch/ corpus watcher Go package (mtimes, no inotify deps)
|
bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats
|
||||||
bin/server/ async Go HTTP server (goroutines, bounded worker pool)
|
bin/mail/ sync.go import.go (index_mail → brain/index.go)
|
||||||
bin/mail/ mail pipeline: sync (Go), import (md), index_mail (rebuild)
|
bin/markdown/ import.go (mistune leafs)
|
||||||
|
bin/postgres/ query.go (read-only YAML)
|
||||||
|
bin/git/ import.go (go-git history; Python shim execs it)
|
||||||
|
bin/web/ search.go (SearXNG; Python shim execs it)
|
||||||
|
internal/ shared Go (brain/rank is cgo-free; chats parsers; gitlog; websearch)
|
||||||
|
bin/watch/ corpus watcher (used by bin/brain/watch.go)
|
||||||
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
|
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
|
||||||
bin/docker-entrypoint container entrypoint (brain index|search|serve|watch)
|
bin/docker-entrypoint container entrypoint (brain index|search|serve|watch)
|
||||||
compose.yaml docker composition (root level, not docker/)
|
compose.yaml docker composition (root level, not docker/)
|
||||||
@@ -52,8 +57,9 @@ var/ kb.lbug, var/mail/*, caches (gitignored)
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
|
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
|
||||||
bin/mail/import --from-raw var/mail # message.json → message.md (convert only)
|
bin/mail/sync.go --source gmail --query 'from:example.com' --out var/mail # Gmail search (default in:inbox)
|
||||||
bin/mail/index_mail # rebuild brain incl. all mail (fresh DB)
|
bin/mail/import.go --from-raw var/mail # message.json → message.md (convert only)
|
||||||
|
bin/brain/index.go --rebuild # rebuild brain incl. all mail (fresh DB)
|
||||||
```
|
```
|
||||||
|
|
||||||
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
|
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
|
||||||
@@ -62,7 +68,7 @@ bin/mail/index_mail # rebuil
|
|||||||
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to
|
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to
|
||||||
docling (isolated subprocess — its native onnx can segfault the parent).
|
docling (isolated subprocess — its native onnx can segfault the parent).
|
||||||
Conversion never touches the brain DB (crash safety).
|
Conversion never touches the brain DB (crash safety).
|
||||||
- `index_mail` always rebuilds from scratch (repo corpus + mail). Ladybug
|
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Ladybug
|
||||||
corrupts its WAL when brand-new leafs are bulk-inserted while FTS/vector
|
corrupts its WAL when brand-new leafs are bulk-inserted while FTS/vector
|
||||||
indexes exist; a fresh DB with indexes created last is the only safe path.
|
indexes exist; a fresh DB with indexes created last is the only safe path.
|
||||||
Keep conversion + indexing separate so a conversion crash can't leave the
|
Keep conversion + indexing separate so a conversion crash can't leave the
|
||||||
@@ -73,7 +79,14 @@ bin/mail/index_mail # rebuil
|
|||||||
```bash
|
```bash
|
||||||
bin/facts/audit ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
|
bin/facts/audit ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
|
||||||
bin/facts/crm [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
|
bin/facts/crm [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
|
||||||
bin/kb/search "query" [--hop N] [--repo X] # deduction search → YAML
|
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
|
||||||
|
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
|
||||||
|
bin/brain/search.go "query" --no-web # local graph only
|
||||||
|
bin/brain/get.go <id> [--body]
|
||||||
|
bin/markdown/import.go [dir] # mistune leaves → YAML
|
||||||
|
bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs
|
||||||
|
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
|
||||||
|
bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
|
||||||
bin/md/tables # what the graph holds → YAML
|
bin/md/tables # what the graph holds → YAML
|
||||||
bin/brain/deduce "question" # thinking wrapper
|
bin/brain/deduce "question" # thinking wrapper
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -26,10 +26,10 @@ detective method: **a fact needs ≥2 independent sources or it is
|
|||||||
|---|----------|--------|
|
|---|----------|--------|
|
||||||
| D1 | RAG corpus | ops stack (chat, onlyoffice, gitea/NPM, searchxng, observability, ai-bot, mcp-servers, `~/.ssh/config`) + portfolio. Exclude `office.dev` + jobs/applications. |
|
| D1 | RAG corpus | ops stack (chat, onlyoffice, gitea/NPM, searchxng, observability, ai-bot, mcp-servers, `~/.ssh/config`) + portfolio. Exclude `office.dev` + jobs/applications. |
|
||||||
| D2 | skill merging | integrate skills **in this project** `skills/`; skip gitea / brain-dependent skills. |
|
| D2 | skill merging | integrate skills **in this project** `skills/`; skip gitea / brain-dependent skills. |
|
||||||
| D3 | web search | import `web-search`, retire local `searxng-ops`. Vendored here, no remote link. |
|
| D3 | web search | Go client `bin/web/search.go` (`internal/websearch`). SearXNG URL is config (`BRAIN_SEARCH_URL`). Optional Compose profile `searxng` (sanitized settings). Do not run a second copy on a host that already has one. Empty/`throttled` ≠ “nothing exists”. |
|
||||||
| D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. |
|
| D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. |
|
||||||
| D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). |
|
| D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). |
|
||||||
| D6 | graph engine | **LadybugDB** (Kuzu successor, MIT, embedded, native FTS+vector+Cypher). Python binding for `bin/*`; Go shebang for golang tools. |
|
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`); Python remains for index/write until the Go write path is safe. |
|
||||||
| D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). |
|
| D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). |
|
||||||
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. |
|
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. |
|
||||||
| D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. |
|
| D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. |
|
||||||
@@ -37,9 +37,12 @@ detective method: **a fact needs ≥2 independent sources or it is
|
|||||||
| D11 | strong/weak | `root` column: `facts` (strong) vs `info` (weak). Answer is `confirmed` only from facts root. |
|
| D11 | strong/weak | `root` column: `facts` (strong) vs `info` (weak). Answer is `confirmed` only from facts root. |
|
||||||
| D12 | transactional | facts and info split by root but **written in the same Ladybug transaction (ACID)** on every write. |
|
| D12 | transactional | facts and info split by root but **written in the same Ladybug transaction (ACID)** on every write. |
|
||||||
| D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. |
|
| D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. |
|
||||||
| D14 | tooling style | `bin/{subject}/{method}` self-describing: shebang line 1, usage comment from line 2. Go shebang: `///usr/bin/env go run "$0" "$@"; exit`. |
|
| D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. |
|
||||||
| D15 | repo | GitHub `eSlider/2dph`, public (like sibling repos), push/commit via `gh`, TDD + commit every change, CI/CD. |
|
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
|
||||||
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. |
|
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. |
|
||||||
|
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
|
||||||
|
| D18 | reasoner | Pluggable OpenAI-compatible URL. RAM: Qwen3.5-9B. Quality: Bonsai-27B or Qwen3.6-27B. No official Qwen3.6-9B. |
|
||||||
|
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
|
|
||||||
@@ -47,16 +50,25 @@ detective method: **a fact needs ≥2 independent sources or it is
|
|||||||
2dph/
|
2dph/
|
||||||
PLAN.md / AGENTS.md
|
PLAN.md / AGENTS.md
|
||||||
docs/ published docs (this conversation → docs/ as md)
|
docs/ published docs (this conversation → docs/ as md)
|
||||||
skills/ in-project skills (web-search, db-yaml, kb-search, agent-cost, diataxis-docs, …)
|
skills/ in-project skills (web-search, db-yaml, brain, diataxis-docs)
|
||||||
bin/
|
bin/
|
||||||
facts/extract auto-pair 2 sources → lexicon yaml + graph
|
facts/extract auto-pair 2 sources → lexicon yaml + graph
|
||||||
facts/audit ["self"|"facts"|"info"|"stale"] 2-source + staleness gate
|
facts/audit ["self"|"facts"|"info"|"stale"] 2-source + staleness gate
|
||||||
kb/index build FTS + HNSW from corpus
|
kb/index Python write path (called by bin/brain/index.go)
|
||||||
kb/search deduction: facts → info → web-search; --hop N
|
brain/index.go rebuild FTS + HNSW (incl. --with-mail)
|
||||||
kb/get kb/stats kb/eval
|
brain/get.go stats.go eval.go watch.go
|
||||||
md/import md/select md/tables md/gaps (mistune)
|
brain/search.go deduction: facts → info → web-search
|
||||||
|
brain/serve.go HTTP API in-process (internal/httpapi + internal/brain)
|
||||||
|
mail/import.go JSON → markdown (no brain write)
|
||||||
|
markdown/import.go mistune leaves
|
||||||
|
postgres/query.go read-only YAML (wraps bin/db/psql-yq)
|
||||||
|
git/import.go go-git history (no git binary; conversion only)
|
||||||
|
web/search.go SearXNG client (throttled ≠ absence)
|
||||||
|
chats/sync.go import.go facts.go apply.go
|
||||||
|
(libs in internal/chats; no chats index)
|
||||||
|
md/import (deprecated; bin/markdown/import.go)
|
||||||
brain/extract brain/audit brain/deduce (thinking wrapper)
|
brain/extract brain/audit brain/deduce (thinking wrapper)
|
||||||
web/search (vendored)
|
web/search (deprecated shim → web/search.go)
|
||||||
db/psql-yq (vendored)
|
db/psql-yq (vendored)
|
||||||
ssh-tunnel onlyoffice pg tunnel 5433
|
ssh-tunnel onlyoffice pg tunnel 5433
|
||||||
var/kb.lbug single embedded store (gitignored)
|
var/kb.lbug single embedded store (gitignored)
|
||||||
@@ -103,24 +115,26 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
|
|||||||
|
|
||||||
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
|
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
|
||||||
Gmail attachments key off `body.attachmentId`, not MIME `partId`.
|
Gmail attachments key off `body.attachmentId`, not MIME `partId`.
|
||||||
2. `bin/mail/import --from-raw` — message.json → message.md; PDFs via
|
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
|
||||||
`pdftotext -layout` (~15ms) with docling subprocess fallback; ICS sidecars
|
`pdftotext -layout` (~15ms) with docling subprocess fallback; ICS sidecars
|
||||||
Latin-1→UTF-8 normalized.
|
Latin-1→UTF-8 normalized.
|
||||||
3. `bin/mail/index_mail` — fresh rebuild (repo corpus + mail) because ladybug
|
3. `bin/brain/index.go --rebuild` — fresh rebuild (repo corpus + mail) because ladybug
|
||||||
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
|
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
|
||||||
indexing stay separate for crash safety.
|
indexing stay separate for crash safety. `bin/mail/index_mail` is a
|
||||||
|
deprecation shim.
|
||||||
4. Result: 17,835 messages → 28,918 info leafs, FTS + HNSW healthy, searchable
|
4. Result: 17,835 messages → 28,918 info leafs, FTS + HNSW healthy, searchable
|
||||||
via `bin/kb/search`.
|
via `bin/brain/search.go`.
|
||||||
|
|
||||||
## CI/CD pipeline (D15)
|
## CI/CD pipeline (D15)
|
||||||
|
|
||||||
`.github/workflows/ci.yml`:
|
`.github/workflows/ci.yml`:
|
||||||
|
|
||||||
1. go vet + go test ./... (Go tools)
|
1. go vet + go test ./... (root module; packages without ladybug cgo)
|
||||||
2. python -m unittest discover + pytest (Py tools)
|
2. `go test ./internal/brain/rank` (cgo-free ranking + flag parser)
|
||||||
3. bin/facts/audit self (lexicon internal consistency)
|
3. python -m unittest discover -s bin/tools (includes published-docs SoT)
|
||||||
4. bin/kb/eval (recall@5 ≥ 0.95, gates index regressions)
|
4. bin/facts/audit self (lexicon internal consistency)
|
||||||
5. md-docs build/lint if docs tooling arrives.
|
5. bin/brain/eval.go (recall@5 ≥ 0.95, gates index regressions)
|
||||||
|
6. md-docs build/lint if docs tooling arrives.
|
||||||
|
|
||||||
Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as
|
Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as
|
||||||
`db/tech-poc`: contract first where there is an OpenAPI/message shape.
|
`db/tech-poc`: contract first where there is an OpenAPI/message shape.
|
||||||
@@ -129,7 +143,7 @@ Feedback loop: every commit → PR → CI → green/gate → merge. Same discipl
|
|||||||
|
|
||||||
1. scaffold repo (:done after this file + AGENTS.md + .gitignore + ci)
|
1. scaffold repo (:done after this file + AGENTS.md + .gitignore + ci)
|
||||||
2. gh repo create eSlider/2dph --private + initial commit + CI
|
2. gh repo create eSlider/2dph --private + initial commit + CI
|
||||||
3. vendored skill integration (web-search, db-yaml, kb-search, agent-cost, diataxis-docs) — no remote links
|
3. vendored skill integration (web-search, db-yaml, brain, diataxis-docs) — no remote links
|
||||||
4. .venv: ladybug + model2vec + mistune
|
4. .venv: ladybug + model2vec + mistune
|
||||||
5. schema + tools with TDD (kb + md + facts + brain)
|
5. schema + tools with TDD (kb + md + facts + brain)
|
||||||
6. ~/.config/brain config
|
6. ~/.config/brain config
|
||||||
|
|||||||
@@ -30,9 +30,9 @@ graph TB
|
|||||||
subgraph dph["2dph tools"]
|
subgraph dph["2dph tools"]
|
||||||
EX["bin/facts/extract<br/>2-source pairing"]
|
EX["bin/facts/extract<br/>2-source pairing"]
|
||||||
AU["bin/facts/audit<br/>confidence + staleness"]
|
AU["bin/facts/audit<br/>confidence + staleness"]
|
||||||
IDX["bin/kb/index<br/>chunk + embed"]
|
IDX["bin/brain/index.go<br/>chunk + embed"]
|
||||||
MD["bin/md/import<br/>mistune leaves"]
|
MD["bin/markdown/import.go<br/>mistune leaves"]
|
||||||
SR["bin/kb/search<br/>deduction + --hop"]
|
SR["bin/brain/search.go<br/>deduction"]
|
||||||
end
|
end
|
||||||
|
|
||||||
subgraph store["Ladybug var/kb.lbug"]
|
subgraph store["Ladybug var/kb.lbug"]
|
||||||
@@ -85,21 +85,41 @@ fact; conflicting sources or a single source → `hypothesis` → `(not confirme
|
|||||||
## Deduction search
|
## Deduction search
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
bin/kb/search "Matrix federation over HTTPS" # facts → info → web-search
|
bin/brain/search.go "Matrix federation over HTTPS" # facts → info → web
|
||||||
bin/kb/search "what runs on arc-2" --hop 1 # walk graph edges
|
bin/brain/search.go "onlyoffice postgres" --root facts
|
||||||
bin/kb/search "where is cs-lexicon" --json | yq '.' # YAML by default
|
bin/brain/search.go "where is cs-lexicon" --json | yq '.'
|
||||||
bin/kb/get <id> --body # full chunk on demand
|
bin/brain/search.go "upstream flag" --no-web # local graph only
|
||||||
bin/kb/stats # index health
|
bin/brain/get.go <id> --body # full chunk on demand
|
||||||
bin/kb/eval # recall@5 gate
|
bin/brain/stats.go # index health
|
||||||
|
bin/brain/eval.go # recall@5 gate
|
||||||
|
```
|
||||||
|
|
||||||
|
`--hop` is not implemented (needs File/FROM_FILE edges); the flag errors instead of walking. `bin/kb/search` is a deprecated wrapper around `bin/brain/search.go`.
|
||||||
|
|
||||||
|
Git history is read with [go-git](https://github.com/go-git/go-git) (no git binary):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/git/import.go --json --limit 100 # commit leafs for this repo
|
||||||
|
bin/git/import.go --root "$PROJECTS_ROOT" --json # one pass per .git under root
|
||||||
|
```
|
||||||
|
|
||||||
|
Conversion only. Graph write (`File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`) stays with `bin/brain/index.go`.
|
||||||
|
|
||||||
|
Web search (second independent source) goes through SearXNG. Empty results mean **throttled**, not “nothing exists”:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/web/search.go "LadybugDB vector index" --json
|
||||||
|
# Optional local instance (skip if BRAIN_SEARCH_URL already points at one):
|
||||||
|
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
|
||||||
```
|
```
|
||||||
|
|
||||||
Mail is a first-class corpus (retrievable through the same search):
|
Mail is a first-class corpus (retrievable through the same search):
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
|
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
|
||||||
bin/mail/import --from-raw var/mail # JSON → markdown
|
bin/mail/import.go --from-raw var/mail # JSON → markdown
|
||||||
bin/mail/index_mail # rebuild brain incl. mail
|
bin/brain/index.go --rebuild # rebuild brain (incl. mail)
|
||||||
bin/kb/search "Mietwagen Nürnberg invoice" # now answers from mail
|
bin/brain/search.go "invoice from last week" # same search over mail leafs
|
||||||
```
|
```
|
||||||
|
|
||||||
## Storage
|
## Storage
|
||||||
@@ -109,7 +129,7 @@ bin/kb/search "Mietwagen Nürnberg invoice" # now a
|
|||||||
readers. **Never `DROP INDEX` FTS/VECTOR** on Ladybug 0.19: DROP leaves
|
readers. **Never `DROP INDEX` FTS/VECTOR** on Ladybug 0.19: DROP leaves
|
||||||
ghost catalog tables (`_0_Leaf_vec_UPPER`) so recreate fails while
|
ghost catalog tables (`_0_Leaf_vec_UPPER`) so recreate fails while
|
||||||
`SHOW_INDEXES` omits HNSW. Fresh indexes = delete `var/kb.lbug` +
|
`SHOW_INDEXES` omits HNSW. Fresh indexes = delete `var/kb.lbug` +
|
||||||
`bin/kb/index --rebuild`. Use `ensure_indexes()` after upserts.
|
`bin/brain/index.go --rebuild`. Use `ensure_indexes()` after upserts.
|
||||||
- **model2vec** — `potion-multilingual-128M` static embeddings (256-dim),
|
- **model2vec** — `potion-multilingual-128M` static embeddings (256-dim),
|
||||||
CPU-fast, deterministic, no Ollama runtime dependency.
|
CPU-fast, deterministic, no Ollama runtime dependency.
|
||||||
- facts and info split semantically by `root` column but written inside the
|
- facts and info split semantically by `root` column but written inside the
|
||||||
@@ -117,10 +137,10 @@ bin/kb/search "Mietwagen Nürnberg invoice" # now a
|
|||||||
|
|
||||||
## Tooling conventions
|
## Tooling conventions
|
||||||
|
|
||||||
`bin/{subject}/{method}` — self-describing: shebang on line 1, usage comment
|
`bin/{subject}/{method}.go` — self-describing: shebang on line 1, usage comment
|
||||||
from line 2. bash + python primary; golang via the Go shebang when a compiled
|
from line 2. Shared code in `internal/`. YAML default output, `--json` for
|
||||||
helper is right. YAML default output, `--json` for machines. Everything that
|
machines. Tests gate every commit. HTTP: `bin/brain/serve.go` calls
|
||||||
touches network/db is read-only, throttled, cached. Tests gate every commit.
|
`internal/brain` in-process (`/health` `/search` `/get` `/stats` `/audit` `/ingest`).
|
||||||
|
|
||||||
## Development
|
## Development
|
||||||
|
|
||||||
@@ -136,7 +156,7 @@ Docker (optional, cached model + var volumes):
|
|||||||
```bash
|
```bash
|
||||||
docker compose run --rm brain index # (re)index corpus
|
docker compose run --rm brain index # (re)index corpus
|
||||||
docker compose run --rm brain search "query" # one-shot query
|
docker compose run --rm brain search "query" # one-shot query
|
||||||
docker compose run --rm brain serve # async Go HTTP server
|
docker compose run --rm brain serve # bin/brain/serve.go
|
||||||
docker compose up brain-watch # auto re-index on change
|
docker compose up brain-watch # auto re-index on change
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -148,4 +168,7 @@ docker compose up brain-watch # auto re-index on change
|
|||||||
skills (`web-search`, `db-yaml`, …) that 2dph integrates
|
skills (`web-search`, `db-yaml`, …) that 2dph integrates
|
||||||
- detective method — the two-source method
|
- detective method — the two-source method
|
||||||
|
|
||||||
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
|
Work board (issues): [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
|
||||||
|
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
|
||||||
|
|
||||||
|
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
|
||||||
|
|||||||
@@ -0,0 +1,3 @@
|
|||||||
|
// Commands in this directory are shebang mains (search.go, serve.go, index.go,
|
||||||
|
// get.go, stats.go, eval.go, watch.go), each behind an exclusive build tag.
|
||||||
|
package main
|
||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=brain_eval "$0" "$@"; exit
|
||||||
|
//go:build brain_eval
|
||||||
|
//
|
||||||
|
// bin/brain/eval.go - recall@5 gate.
|
||||||
|
//
|
||||||
|
// ./bin/brain/eval.go
|
||||||
|
// ./bin/brain/eval.go --json
|
||||||
|
//
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(cmdbin.ExecFile("bin/kb/eval", os.Args[1:]))
|
||||||
|
}
|
||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=brain_get "$0" "$@"; exit
|
||||||
|
//go:build brain_get
|
||||||
|
//
|
||||||
|
// bin/brain/get.go - read one leaf by id.
|
||||||
|
//
|
||||||
|
// ./bin/brain/get.go <id>
|
||||||
|
// ./bin/brain/get.go <id> --body
|
||||||
|
//
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(cmdbin.ExecFile("bin/kb/get", os.Args[1:]))
|
||||||
|
}
|
||||||
Executable
+24
@@ -0,0 +1,24 @@
|
|||||||
|
//usr/bin/env go run -tags=brain_index "$0" "$@"; exit
|
||||||
|
//go:build brain_index
|
||||||
|
//
|
||||||
|
// bin/brain/index.go - rebuild the Ladybug graph (Python write path).
|
||||||
|
//
|
||||||
|
// ./bin/brain/index.go --rebuild
|
||||||
|
// ./bin/brain/index.go --rebuild --with-mail
|
||||||
|
// ./bin/brain/index.go --dry-run --with-mail
|
||||||
|
//
|
||||||
|
// v1 write is always a rebuild when mail is included (live FTS/HNSW + bulk
|
||||||
|
// insert corrupts Ladybug 0.19 WAL). `add` is v2.
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
args := append([]string{"--with-mail"}, os.Args[1:]...)
|
||||||
|
os.Exit(cmdbin.ExecFile("bin/kb/index", args))
|
||||||
|
}
|
||||||
Executable
+23
@@ -0,0 +1,23 @@
|
|||||||
|
//usr/bin/env go run -tags=system_ladybug "$0" "$@"; exit
|
||||||
|
//go:build cgo && system_ladybug
|
||||||
|
//
|
||||||
|
// bin/brain/search.go - deduction search over the 2dph brain.
|
||||||
|
//
|
||||||
|
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--json] [--no-web]
|
||||||
|
// ./bin/brain/search.go serve [port]
|
||||||
|
// ./bin/brain/search.go --list-model
|
||||||
|
//
|
||||||
|
// Needs CGO + libladybug (CGO_CFLAGS/CGO_LDFLAGS). Prefer the wrapper
|
||||||
|
// bin/kb/search which sets those and builds a binary for the embed daemon.
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/brain"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(brain.Main(os.Args[1:]))
|
||||||
|
}
|
||||||
Executable
+31
@@ -0,0 +1,31 @@
|
|||||||
|
//usr/bin/env go run -tags=brain_serve,system_ladybug "$0" "$@"; exit
|
||||||
|
//go:build brain_serve && cgo && system_ladybug
|
||||||
|
//
|
||||||
|
// bin/brain/serve.go - HTTP API (in-process ladybug search).
|
||||||
|
//
|
||||||
|
// KB_ROOT=/path/to/2dph ./bin/brain/serve.go
|
||||||
|
// KB_WORKERS=4 KB_PORT=8630 ./bin/brain/serve.go
|
||||||
|
//
|
||||||
|
// Needs CGO + libladybug (same as bin/brain/search.go).
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"log"
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/brain"
|
||||||
|
"github.com/eSlider/2dph/internal/httpapi"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
if os.Getenv("KB_ROOT") == "" {
|
||||||
|
if wd, err := os.Getwd(); err == nil {
|
||||||
|
os.Setenv("KB_ROOT", wd)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if err := brain.Ready(); err != nil {
|
||||||
|
log.Fatal(err)
|
||||||
|
}
|
||||||
|
httpapi.Run(brain.HTTP{})
|
||||||
|
}
|
||||||
@@ -0,0 +1,20 @@
|
|||||||
|
//go:build brain_serve && !system_ladybug
|
||||||
|
//
|
||||||
|
// Fallback serve when ladybug cgo is not in the build (CI / tags=brain_serve).
|
||||||
|
// Production shebang is serve.go (in-process).
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/httpapi"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
if os.Getenv("KB_ROOT") == "" {
|
||||||
|
if wd, err := os.Getwd(); err == nil {
|
||||||
|
os.Setenv("KB_ROOT", wd)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
httpapi.Run(nil)
|
||||||
|
}
|
||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=brain_stats "$0" "$@"; exit
|
||||||
|
//go:build brain_stats
|
||||||
|
//
|
||||||
|
// bin/brain/stats.go - index health.
|
||||||
|
//
|
||||||
|
// ./bin/brain/stats.go
|
||||||
|
// ./bin/brain/stats.go --json
|
||||||
|
//
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(cmdbin.ExecFile("bin/kb/stats", os.Args[1:]))
|
||||||
|
}
|
||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=brain_watch "$0" "$@"; exit
|
||||||
|
//go:build brain_watch
|
||||||
|
//
|
||||||
|
// bin/brain/watch.go - re-index when corpus files change.
|
||||||
|
//
|
||||||
|
// ./bin/brain/watch.go [dir...]
|
||||||
|
// KB_WATCH_INTERVAL=15 ./bin/brain/watch.go
|
||||||
|
//
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/bin/watch"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
watch.Run(os.Args[1:])
|
||||||
|
}
|
||||||
Executable
+19
@@ -0,0 +1,19 @@
|
|||||||
|
//usr/bin/env go run -tags=chats_apply "$0" "$@"; exit
|
||||||
|
//go:build chats_apply
|
||||||
|
//
|
||||||
|
// bin/chats/apply.go - push extracted chat facts to OnlyOffice CRM.
|
||||||
|
//
|
||||||
|
// ./bin/chats/apply.go [--dry-run]
|
||||||
|
//
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/chats"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(chats.RunApply(os.Args[1:]))
|
||||||
|
}
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
// Commands in this directory are shebang mains (sync.go, import.go, facts.go,
|
||||||
|
// apply.go), each behind an exclusive build tag so `go build ./bin/chats`
|
||||||
|
// does not see two mains. Shared code lives in internal/chats.
|
||||||
|
package main
|
||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=chats_facts "$0" "$@"; exit
|
||||||
|
//go:build chats_facts
|
||||||
|
//
|
||||||
|
// bin/chats/facts.go - extract phone/email/linkedin facts from JSONL.
|
||||||
|
//
|
||||||
|
// ./bin/chats/facts.go
|
||||||
|
//
|
||||||
|
// Writes var/chats/facts/. Does not index the brain.
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/chats"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(chats.RunFacts(os.Args[1:]))
|
||||||
|
}
|
||||||
@@ -1,3 +0,0 @@
|
|||||||
module github.com/eSlider/2dph/bin/chats
|
|
||||||
|
|
||||||
go 1.25.0
|
|
||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=chats_import "$0" "$@"; exit
|
||||||
|
//go:build chats_import
|
||||||
|
//
|
||||||
|
// bin/chats/import.go - JSONL → markdown under var/chats/md/.
|
||||||
|
//
|
||||||
|
// ./bin/chats/import.go
|
||||||
|
//
|
||||||
|
// Conversion only. Brain ingest is bin/brain/index.go, not this command.
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/chats"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(chats.RunImport(os.Args[1:]))
|
||||||
|
}
|
||||||
@@ -1,56 +0,0 @@
|
|||||||
package main
|
|
||||||
|
|
||||||
import (
|
|
||||||
"bytes"
|
|
||||||
"flag"
|
|
||||||
"fmt"
|
|
||||||
"os"
|
|
||||||
"os/exec"
|
|
||||||
"path/filepath"
|
|
||||||
"strings"
|
|
||||||
)
|
|
||||||
|
|
||||||
func runIndex(args []string) int {
|
|
||||||
fs := flag.NewFlagSet("chats index", flag.ContinueOnError)
|
|
||||||
help := fs.Bool("help", false, "")
|
|
||||||
fs.SetOutput(os.Stderr)
|
|
||||||
if err := fs.Parse(args); err != nil {
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
if *help {
|
|
||||||
fmt.Fprintln(os.Stderr, "usage: chats index")
|
|
||||||
return 0
|
|
||||||
}
|
|
||||||
|
|
||||||
root := repoRoot()
|
|
||||||
mdDir := filepath.Join(chatsDir(), "md")
|
|
||||||
|
|
||||||
_, err := os.Stat(mdDir)
|
|
||||||
if os.IsNotExist(err) {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats index: no chat markdown at %s; run 'chats import' first\n", mdDir)
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
|
|
||||||
indexScript := filepath.Join(root, "bin", "kb", "index")
|
|
||||||
if _, err := os.Stat(indexScript); os.IsNotExist(err) {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats index: %s not found\n", indexScript)
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
|
|
||||||
cmd := exec.Command(indexScript, "--corpus", mdDir)
|
|
||||||
var outBuf, errBuf bytes.Buffer
|
|
||||||
cmd.Stdout = &outBuf
|
|
||||||
cmd.Stderr = &errBuf
|
|
||||||
cmd.Dir = root
|
|
||||||
|
|
||||||
if err := cmd.Run(); err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats index: %v\n%s", err, errBuf.String())
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
result := strings.TrimSpace(outBuf.String())
|
|
||||||
if result == "" {
|
|
||||||
result = strings.TrimSpace(errBuf.String())
|
|
||||||
}
|
|
||||||
fmt.Printf("chats index: %s\n", result)
|
|
||||||
return 0
|
|
||||||
}
|
|
||||||
@@ -1,356 +0,0 @@
|
|||||||
package main
|
|
||||||
|
|
||||||
import (
|
|
||||||
"bufio"
|
|
||||||
"context"
|
|
||||||
"encoding/json"
|
|
||||||
"fmt"
|
|
||||||
"os"
|
|
||||||
"os/exec"
|
|
||||||
"path/filepath"
|
|
||||||
"strings"
|
|
||||||
"time"
|
|
||||||
)
|
|
||||||
|
|
||||||
type LinkedInMCPSource struct {
|
|
||||||
userDataDir string
|
|
||||||
limit int
|
|
||||||
}
|
|
||||||
|
|
||||||
type lnInboxItem struct {
|
|
||||||
ThreadID string `json:"thread_id"`
|
|
||||||
Participants string `json:"participants"`
|
|
||||||
LastMessage string `json:"last_message"`
|
|
||||||
LastMessageDate string `json:"last_message_date"`
|
|
||||||
Unread bool `json:"unread"`
|
|
||||||
}
|
|
||||||
|
|
||||||
type lnInboxEnvelope struct {
|
|
||||||
Results []lnInboxItem `json:"results"`
|
|
||||||
HasMore bool `json:"hasMore"`
|
|
||||||
}
|
|
||||||
|
|
||||||
type lnMessage struct {
|
|
||||||
From string `json:"from"`
|
|
||||||
Date string `json:"date"`
|
|
||||||
Text string `json:"text"`
|
|
||||||
}
|
|
||||||
|
|
||||||
type lnConvEnvelope struct {
|
|
||||||
Results []lnMessage `json:"results"`
|
|
||||||
HasMore bool `json:"hasMore"`
|
|
||||||
TotalCount int `json:"total_count"`
|
|
||||||
}
|
|
||||||
|
|
||||||
func NewLinkedInMCPSource(userDataDir string) *LinkedInMCPSource {
|
|
||||||
return &LinkedInMCPSource{userDataDir: userDataDir}
|
|
||||||
}
|
|
||||||
|
|
||||||
func (s *LinkedInMCPSource) Name() string { return "linkedin" }
|
|
||||||
|
|
||||||
func (s *LinkedInMCPSource) Sync(ctx context.Context, outDir string, limit int) error {
|
|
||||||
if limit > 0 {
|
|
||||||
s.limit = limit
|
|
||||||
}
|
|
||||||
|
|
||||||
client, err := newLinkedInMCP(ctx, s.userDataDir)
|
|
||||||
if err != nil {
|
|
||||||
return fmt.Errorf("linkedin mcp: %w", err)
|
|
||||||
}
|
|
||||||
defer client.Close()
|
|
||||||
|
|
||||||
inbox, err := client.GetInbox(ctx, 50)
|
|
||||||
if err != nil {
|
|
||||||
return fmt.Errorf("get_inbox: %w", err)
|
|
||||||
}
|
|
||||||
if len(inbox) == 0 {
|
|
||||||
fmt.Println("chats: no LinkedIn conversations found")
|
|
||||||
return nil
|
|
||||||
}
|
|
||||||
|
|
||||||
fmt.Printf("chats: found %d LinkedIn conversations\n", len(inbox))
|
|
||||||
|
|
||||||
for _, conv := range inbox {
|
|
||||||
convID := sanitizeDir(conv.ThreadID)
|
|
||||||
if convID == "" {
|
|
||||||
convID = fmt.Sprintf("conv_%d", time.Now().UnixNano())
|
|
||||||
}
|
|
||||||
|
|
||||||
parts := strings.SplitN(conv.Participants, ",", 2)
|
|
||||||
chatName := strings.TrimSpace(parts[0])
|
|
||||||
if chatName == "" {
|
|
||||||
chatName = convID
|
|
||||||
}
|
|
||||||
|
|
||||||
chatDir := filepath.Join(outDir, "linkedin", convID)
|
|
||||||
if err := os.MkdirAll(chatDir, 0755); err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: mkdir %s: %v\n", chatDir, err)
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
|
|
||||||
msgLimit := 100
|
|
||||||
if s.limit > 0 {
|
|
||||||
msgLimit = s.limit
|
|
||||||
}
|
|
||||||
|
|
||||||
msgs, err := client.GetConversation(ctx, "", conv.ThreadID, msgLimit)
|
|
||||||
if err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: get_conversation %s: %v\n", convID, err)
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
|
|
||||||
jsonlPath := filepath.Join(chatDir, "messages.jsonl")
|
|
||||||
f, err := os.Create(jsonlPath)
|
|
||||||
if err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: create %s: %v\n", jsonlPath, err)
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
|
|
||||||
enc := json.NewEncoder(f)
|
|
||||||
written := 0
|
|
||||||
for i, m := range msgs {
|
|
||||||
text := m.Text
|
|
||||||
if text == "" {
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
|
|
||||||
ts := m.Date
|
|
||||||
if t, err := time.Parse("2006-01-02T15:04:05Z07:00", m.Date); err == nil {
|
|
||||||
ts = t.UTC().Format(time.RFC3339)
|
|
||||||
} else if t, err := time.Parse(time.RFC3339, m.Date); err == nil {
|
|
||||||
ts = t.UTC().Format(time.RFC3339)
|
|
||||||
}
|
|
||||||
|
|
||||||
chatMsg := Message{
|
|
||||||
ID: fmt.Sprintf("li_%s_%d", convID, i),
|
|
||||||
Timestamp: ts,
|
|
||||||
From: m.From,
|
|
||||||
Text: text,
|
|
||||||
Platform: "linkedin",
|
|
||||||
}
|
|
||||||
if err := enc.Encode(chatMsg); err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: encode: %v\n", err)
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
written++
|
|
||||||
}
|
|
||||||
f.Close()
|
|
||||||
|
|
||||||
fmt.Printf("chats: synced %s (%s) — %d messages\n", chatName, convID, written)
|
|
||||||
}
|
|
||||||
|
|
||||||
return nil
|
|
||||||
}
|
|
||||||
|
|
||||||
type linkedInMCPClient struct {
|
|
||||||
cmd *exec.Cmd
|
|
||||||
stdin *bufio.Writer
|
|
||||||
stdout *bufio.Scanner
|
|
||||||
msgID int
|
|
||||||
}
|
|
||||||
|
|
||||||
func newLinkedInMCP(ctx context.Context, userDataDir string) (*linkedInMCPClient, error) {
|
|
||||||
args := []string{
|
|
||||||
"mcp-server-linkedin@latest",
|
|
||||||
"--user-data-dir", userDataDir,
|
|
||||||
"--no-auto-import",
|
|
||||||
"--transport", "stdio",
|
|
||||||
"--login-timeout", "10",
|
|
||||||
"--browser-wait", "1",
|
|
||||||
"--browser-idle-timeout", "10",
|
|
||||||
"--log-level", "ERROR",
|
|
||||||
}
|
|
||||||
|
|
||||||
cmd := exec.CommandContext(ctx, "uvx", args...)
|
|
||||||
cmd.Env = os.Environ()
|
|
||||||
|
|
||||||
stdin, err := cmd.StdinPipe()
|
|
||||||
if err != nil {
|
|
||||||
return nil, fmt.Errorf("stdin pipe: %w", err)
|
|
||||||
}
|
|
||||||
stdout, err := cmd.StdoutPipe()
|
|
||||||
if err != nil {
|
|
||||||
return nil, fmt.Errorf("stdout pipe: %w", err)
|
|
||||||
}
|
|
||||||
cmd.Stderr = os.Stderr
|
|
||||||
|
|
||||||
if err := cmd.Start(); err != nil {
|
|
||||||
return nil, fmt.Errorf("start: %w", err)
|
|
||||||
}
|
|
||||||
|
|
||||||
c := &linkedInMCPClient{
|
|
||||||
cmd: cmd,
|
|
||||||
stdin: bufio.NewWriter(stdin),
|
|
||||||
stdout: bufio.NewScanner(stdout),
|
|
||||||
msgID: 0,
|
|
||||||
}
|
|
||||||
c.stdout.Buffer(make([]byte, 1<<20), 1<<20)
|
|
||||||
|
|
||||||
if err := c.initialize(ctx); err != nil {
|
|
||||||
c.Close()
|
|
||||||
return nil, fmt.Errorf("init: %w", err)
|
|
||||||
}
|
|
||||||
return c, nil
|
|
||||||
}
|
|
||||||
|
|
||||||
func (c *linkedInMCPClient) nextID() int {
|
|
||||||
c.msgID++
|
|
||||||
return c.msgID
|
|
||||||
}
|
|
||||||
|
|
||||||
func (c *linkedInMCPClient) initialize(ctx context.Context) error {
|
|
||||||
params := map[string]interface{}{
|
|
||||||
"protocolVersion": "2024-11-05",
|
|
||||||
"capabilities": map[string]interface{}{},
|
|
||||||
"clientInfo": map[string]string{
|
|
||||||
"name": "chats-sync",
|
|
||||||
"version": "0.1.0",
|
|
||||||
},
|
|
||||||
}
|
|
||||||
_, err := c.send(ctx, "initialize", params)
|
|
||||||
return err
|
|
||||||
}
|
|
||||||
|
|
||||||
func (c *linkedInMCPClient) send(ctx context.Context, method string, params interface{}) (json.RawMessage, error) {
|
|
||||||
id := c.nextID()
|
|
||||||
req := map[string]interface{}{
|
|
||||||
"jsonrpc": "2.0",
|
|
||||||
"id": id,
|
|
||||||
"method": method,
|
|
||||||
}
|
|
||||||
if params != nil {
|
|
||||||
req["params"] = params
|
|
||||||
}
|
|
||||||
body, err := json.Marshal(req)
|
|
||||||
if err != nil {
|
|
||||||
return nil, fmt.Errorf("marshal: %w", err)
|
|
||||||
}
|
|
||||||
|
|
||||||
if _, err := c.stdin.Write(body); err != nil {
|
|
||||||
return nil, fmt.Errorf("write: %w", err)
|
|
||||||
}
|
|
||||||
if err := c.stdin.WriteByte('\n'); err != nil {
|
|
||||||
return nil, fmt.Errorf("newline: %w", err)
|
|
||||||
}
|
|
||||||
if err := c.stdin.Flush(); err != nil {
|
|
||||||
return nil, fmt.Errorf("flush: %w", err)
|
|
||||||
}
|
|
||||||
|
|
||||||
for c.stdout.Scan() {
|
|
||||||
line := c.stdout.Text()
|
|
||||||
if line == "" {
|
|
||||||
continue
|
|
||||||
}
|
|
||||||
var resp struct {
|
|
||||||
JSONRPC string `json:"jsonrpc"`
|
|
||||||
ID int `json:"id"`
|
|
||||||
Result json.RawMessage `json:"result,omitempty"`
|
|
||||||
Error *struct {
|
|
||||||
Code int `json:"code"`
|
|
||||||
Message string `json:"message"`
|
|
||||||
} `json:"error,omitempty"`
|
|
||||||
}
|
|
||||||
if err := json.Unmarshal([]byte(line), &resp); err != nil {
|
|
||||||
return nil, fmt.Errorf("unmarshal: %w\nline: %s", err, line[:min(len(line), 500)])
|
|
||||||
}
|
|
||||||
if resp.Error != nil {
|
|
||||||
return nil, fmt.Errorf("rpc error %d: %s", resp.Error.Code, resp.Error.Message)
|
|
||||||
}
|
|
||||||
return resp.Result, nil
|
|
||||||
}
|
|
||||||
return nil, fmt.Errorf("no response: %w", c.stdout.Err())
|
|
||||||
}
|
|
||||||
|
|
||||||
func (c *linkedInMCPClient) GetInbox(ctx context.Context, limit int) ([]lnInboxItem, error) {
|
|
||||||
params := map[string]interface{}{
|
|
||||||
"limit": limit,
|
|
||||||
}
|
|
||||||
result, err := c.send(ctx, "tools/call", map[string]interface{}{
|
|
||||||
"name": "get_inbox",
|
|
||||||
"arguments": params,
|
|
||||||
})
|
|
||||||
if err != nil {
|
|
||||||
return nil, err
|
|
||||||
}
|
|
||||||
|
|
||||||
var toolRes struct {
|
|
||||||
Content []struct {
|
|
||||||
Type string `json:"type"`
|
|
||||||
Text string `json:"text"`
|
|
||||||
} `json:"content"`
|
|
||||||
IsError bool `json:"isError"`
|
|
||||||
}
|
|
||||||
if err := json.Unmarshal(result, &toolRes); err != nil {
|
|
||||||
return nil, fmt.Errorf("unmarshal tool: %w", err)
|
|
||||||
}
|
|
||||||
if toolRes.IsError {
|
|
||||||
return nil, fmt.Errorf("get_inbox error")
|
|
||||||
}
|
|
||||||
if len(toolRes.Content) == 0 {
|
|
||||||
return nil, nil
|
|
||||||
}
|
|
||||||
|
|
||||||
text := toolRes.Content[0].Text
|
|
||||||
var env lnInboxEnvelope
|
|
||||||
if err := json.Unmarshal([]byte(text), &env); err != nil {
|
|
||||||
var arr []lnInboxItem
|
|
||||||
if err2 := json.Unmarshal([]byte(text), &arr); err2 == nil {
|
|
||||||
return arr, nil
|
|
||||||
}
|
|
||||||
return nil, fmt.Errorf("parse inbox: %w", err)
|
|
||||||
}
|
|
||||||
return env.Results, nil
|
|
||||||
}
|
|
||||||
|
|
||||||
func (c *linkedInMCPClient) GetConversation(ctx context.Context, username, threadID string, limit int) ([]lnMessage, error) {
|
|
||||||
params := map[string]interface{}{
|
|
||||||
"linkedin_username": username,
|
|
||||||
"thread_id": threadID,
|
|
||||||
"index": limit,
|
|
||||||
}
|
|
||||||
result, err := c.send(ctx, "tools/call", map[string]interface{}{
|
|
||||||
"name": "get_conversation",
|
|
||||||
"arguments": params,
|
|
||||||
})
|
|
||||||
if err != nil {
|
|
||||||
return nil, err
|
|
||||||
}
|
|
||||||
|
|
||||||
var toolRes struct {
|
|
||||||
Content []struct {
|
|
||||||
Type string `json:"type"`
|
|
||||||
Text string `json:"text"`
|
|
||||||
} `json:"content"`
|
|
||||||
IsError bool `json:"isError"`
|
|
||||||
}
|
|
||||||
if err := json.Unmarshal(result, &toolRes); err != nil {
|
|
||||||
return nil, fmt.Errorf("unmarshal tool: %w", err)
|
|
||||||
}
|
|
||||||
if toolRes.IsError {
|
|
||||||
return nil, nil
|
|
||||||
}
|
|
||||||
if len(toolRes.Content) == 0 {
|
|
||||||
return nil, nil
|
|
||||||
}
|
|
||||||
|
|
||||||
text := toolRes.Content[0].Text
|
|
||||||
var env lnConvEnvelope
|
|
||||||
if err := json.Unmarshal([]byte(text), &env); err != nil {
|
|
||||||
var arr []lnMessage
|
|
||||||
if err2 := json.Unmarshal([]byte(text), &arr); err2 == nil {
|
|
||||||
return arr, nil
|
|
||||||
}
|
|
||||||
return nil, fmt.Errorf("parse conv: %w", err)
|
|
||||||
}
|
|
||||||
return env.Results, nil
|
|
||||||
}
|
|
||||||
|
|
||||||
func (c *linkedInMCPClient) Close() error {
|
|
||||||
if c.stdin != nil {
|
|
||||||
c.stdin.Flush()
|
|
||||||
}
|
|
||||||
if c.cmd != nil && c.cmd.Process != nil {
|
|
||||||
c.cmd.Process.Kill()
|
|
||||||
}
|
|
||||||
return nil
|
|
||||||
}
|
|
||||||
@@ -1,115 +0,0 @@
|
|||||||
// bin/chats - sync, import, index, extract facts, and apply chat data
|
|
||||||
// from Telegram, WhatsApp, LinkedIn into the brain and OnlyOffice CRM.
|
|
||||||
//
|
|
||||||
// Usage:
|
|
||||||
//
|
|
||||||
// chats sync telegram [--limit N] [--since DATE] [--phone PHONE]
|
|
||||||
// chats sync whatsapp [--qr] [--limit N]
|
|
||||||
// chats sync linkedin [--limit N]
|
|
||||||
// chats import # JSONL → MD (all sources)
|
|
||||||
// chats index # rebuild var/kb.lbug with chats
|
|
||||||
// chats facts # extract + cross-check
|
|
||||||
// chats apply [--dry-run] # push to OnlyOffice CRM
|
|
||||||
package main
|
|
||||||
|
|
||||||
import (
|
|
||||||
"fmt"
|
|
||||||
"os"
|
|
||||||
"strings"
|
|
||||||
)
|
|
||||||
|
|
||||||
func main() {
|
|
||||||
if len(os.Args) < 2 {
|
|
||||||
usage()
|
|
||||||
os.Exit(2)
|
|
||||||
}
|
|
||||||
cmd := os.Args[1]
|
|
||||||
args := os.Args[2:]
|
|
||||||
switch cmd {
|
|
||||||
case "sync":
|
|
||||||
if len(args) < 1 {
|
|
||||||
usage()
|
|
||||||
os.Exit(2)
|
|
||||||
}
|
|
||||||
platform := args[0]
|
|
||||||
platformArgs := args[1:]
|
|
||||||
switch platform {
|
|
||||||
case "telegram":
|
|
||||||
os.Exit(runSyncTelegram(platformArgs))
|
|
||||||
case "whatsapp":
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: WhatsApp not implemented yet\n")
|
|
||||||
os.Exit(1)
|
|
||||||
case "linkedin":
|
|
||||||
os.Exit(runSyncLinkedIn(platformArgs))
|
|
||||||
default:
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: unknown platform %q\n", platform)
|
|
||||||
os.Exit(2)
|
|
||||||
}
|
|
||||||
case "import":
|
|
||||||
os.Exit(runImport(args))
|
|
||||||
case "index":
|
|
||||||
os.Exit(runIndex(args))
|
|
||||||
case "facts":
|
|
||||||
os.Exit(runFacts(args))
|
|
||||||
case "apply":
|
|
||||||
os.Exit(runApply(args))
|
|
||||||
case "help", "-h", "--help":
|
|
||||||
usage()
|
|
||||||
return
|
|
||||||
default:
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: unknown command %q\n", cmd)
|
|
||||||
usage()
|
|
||||||
os.Exit(2)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
func usage() {
|
|
||||||
w := os.Stderr
|
|
||||||
fmt.Fprintln(w, `Usage: chats <command> [args]
|
|
||||||
|
|
||||||
Commands:
|
|
||||||
sync telegram [--limit N] [--since DATE] [--phone PHONE]
|
|
||||||
sync whatsapp [--qr] [--limit N]
|
|
||||||
sync linkedin [--limit N]
|
|
||||||
import JSONL → MD (all sources)
|
|
||||||
index rebuild var/kb.lbug with chats
|
|
||||||
facts extract + cross-check facts
|
|
||||||
apply [--dry-run] push to OnlyOffice CRM
|
|
||||||
|
|
||||||
Output layout:
|
|
||||||
var/chats/<platform>/<chat_id>/messages.jsonl
|
|
||||||
var/chats/md/<platform>/<chat_name>/messages.md`)
|
|
||||||
}
|
|
||||||
|
|
||||||
// repoRoot locates the 2dph project root by walking up from the binary.
|
|
||||||
func repoRoot() string {
|
|
||||||
if v := os.Getenv("KB_ROOT"); v != "" {
|
|
||||||
return v
|
|
||||||
}
|
|
||||||
wd, err := os.Getwd()
|
|
||||||
if err != nil {
|
|
||||||
return "."
|
|
||||||
}
|
|
||||||
for i := 0; i < 10; i++ {
|
|
||||||
if _, err := os.Stat(wd + "/var"); err == nil {
|
|
||||||
return wd
|
|
||||||
}
|
|
||||||
if _, err := os.Stat(wd + "/.git"); err == nil {
|
|
||||||
return wd
|
|
||||||
}
|
|
||||||
parent := wd
|
|
||||||
if idx := strings.LastIndex(wd, "/"); idx >= 0 {
|
|
||||||
parent = wd[:idx]
|
|
||||||
}
|
|
||||||
if parent == wd {
|
|
||||||
break
|
|
||||||
}
|
|
||||||
wd = parent
|
|
||||||
}
|
|
||||||
return "."
|
|
||||||
}
|
|
||||||
|
|
||||||
// chatsDir returns var/chats under the repo root.
|
|
||||||
func chatsDir() string {
|
|
||||||
return repoRoot() + "/var/chats"
|
|
||||||
}
|
|
||||||
Executable
+155
@@ -0,0 +1,155 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""chats/refresh-linkedin-session - refresh LinkedIn MCP session from webtop CDP.
|
||||||
|
|
||||||
|
bin/chats/refresh-linkedin-session [--cdp URL] [--root DIR]
|
||||||
|
"""
|
||||||
|
|
||||||
|
Reads the current LinkedIn cookies out of the running Thorium browser in the
|
||||||
|
work-webtop container via CDP (Network.getAllCookies), copies the live browser
|
||||||
|
profile onto the source profile directory, and rewrites the portable
|
||||||
|
cookies.json + source-state.json that mcp-server-linkedin requires.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
refresh-linkedin-session [--cdp http://127.0.0.1:9222] [--root /var/tmp/liprofile]
|
||||||
|
[--container work-webtop] [--profile thorium-profile]
|
||||||
|
|
||||||
|
After the headless driver uses a copied profile, LinkedIn rotates the session
|
||||||
|
in that copy, so this must run before every sync.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import shutil
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import urllib.request
|
||||||
|
|
||||||
|
import websockets
|
||||||
|
|
||||||
|
|
||||||
|
def cdp_tab(ws_json):
|
||||||
|
for t in ws_json:
|
||||||
|
if t.get("webSocketDebuggerUrl"):
|
||||||
|
return t["webSocketDebuggerUrl"]
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
async def get_cookies(ws_url):
|
||||||
|
async with websockets.connect(ws_url, max_size=50_000_000) as ws:
|
||||||
|
await ws.send(json.dumps({"id": 1, "method": "Network.getAllCookies", "params": {}}))
|
||||||
|
resp = await ws.recv()
|
||||||
|
return json.loads(resp).get("result", {}).get("cookies", [])
|
||||||
|
|
||||||
|
|
||||||
|
def write_source_state(root, profile_dir):
|
||||||
|
# Reuse the linkedin-mcp-server session_state module to write a valid
|
||||||
|
# source-state.json (same schema the daemon reads).
|
||||||
|
try:
|
||||||
|
from linkedin_mcp_server.session_state import canonical, write_source_state
|
||||||
|
|
||||||
|
write_source_state(canonical(__import__("pathlib").Path(profile_dir)))
|
||||||
|
return
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
# Fallback: minimal schema-compatible state.
|
||||||
|
import uuid
|
||||||
|
|
||||||
|
state = {
|
||||||
|
"version": 1,
|
||||||
|
"source_runtime_id": "linux-amd64-host",
|
||||||
|
"login_generation": str(uuid.uuid4()),
|
||||||
|
"created_at": None,
|
||||||
|
"profile_path": profile_dir,
|
||||||
|
"cookies_path": os.path.join(root, "cookies.json"),
|
||||||
|
}
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
state["created_at"] = datetime.now(timezone.utc).isoformat()
|
||||||
|
with open(os.path.join(root, "source-state.json"), "w") as f:
|
||||||
|
json.dump(state, f, indent=2)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
args = sys.argv[1:]
|
||||||
|
cdp = "http://127.0.0.1:9222"
|
||||||
|
root = "/var/tmp/liprofile"
|
||||||
|
container = "work-webtop"
|
||||||
|
cprofile = "thorium-profile"
|
||||||
|
for i in range(0, len(args), 2):
|
||||||
|
k = args[i]
|
||||||
|
v = args[i + 1] if i + 1 < len(args) else ""
|
||||||
|
if k == "--cdp":
|
||||||
|
cdp = v
|
||||||
|
elif k == "--root":
|
||||||
|
root = v
|
||||||
|
elif k == "--container":
|
||||||
|
container = v
|
||||||
|
elif k == "--profile":
|
||||||
|
cprofile = v
|
||||||
|
|
||||||
|
profile_dir = os.path.join(root, "profile")
|
||||||
|
os.makedirs(profile_dir, exist_ok=True)
|
||||||
|
|
||||||
|
# 1. Clear stale daemon/browser locks so the server can claim the profile.
|
||||||
|
for lock in ("profile-claim.lock", "profile.lock", "daemon.lock", "lease.lock"):
|
||||||
|
p = os.path.join(root, lock)
|
||||||
|
if os.path.exists(p):
|
||||||
|
os.remove(p)
|
||||||
|
for name in os.listdir(profile_dir):
|
||||||
|
if name.startswith("Singleton"):
|
||||||
|
os.remove(os.path.join(profile_dir, name))
|
||||||
|
for name in os.listdir(root):
|
||||||
|
if name.startswith("invalid-state-"):
|
||||||
|
shutil.rmtree(os.path.join(root, name), ignore_errors=True)
|
||||||
|
|
||||||
|
# 1. Copy the live browser profile (cookies DB + Local State) so the
|
||||||
|
# session the driver launches carries the current login.
|
||||||
|
subprocess.run(
|
||||||
|
["docker", "cp", f"{container}:/config/{cprofile}/Default", os.path.join(profile_dir, "Default")],
|
||||||
|
check=True, capture_output=True,
|
||||||
|
)
|
||||||
|
subprocess.run(
|
||||||
|
["docker", "cp", f"{container}:/config/{cprofile}/Local State", os.path.join(profile_dir, "Local State")],
|
||||||
|
check=True, capture_output=True,
|
||||||
|
)
|
||||||
|
for lock in ("SingletonLock", "SingletonCookie", "SingletonSocket"):
|
||||||
|
p = os.path.join(profile_dir, lock)
|
||||||
|
if os.path.exists(p):
|
||||||
|
os.remove(p)
|
||||||
|
|
||||||
|
# 2. Pull the live cookies out of the running browser.
|
||||||
|
with urllib.request.urlopen(f"{cdp}/json", timeout=5) as r:
|
||||||
|
tabs = json.loads(r.read())
|
||||||
|
ws_url = cdp_tab(tabs)
|
||||||
|
if not ws_url:
|
||||||
|
sys.stderr.write("refresh-linkedin-session: no CDP tab\n")
|
||||||
|
sys.exit(1)
|
||||||
|
cookies = asyncio.run(get_cookies(ws_url))
|
||||||
|
|
||||||
|
li = [c for c in cookies if "linkedin" in c.get("domain", "")]
|
||||||
|
out = []
|
||||||
|
for c in li:
|
||||||
|
domain = c.get("domain", "")
|
||||||
|
if domain in (".www.linkedin.com", "www.linkedin.com"):
|
||||||
|
domain = ".linkedin.com"
|
||||||
|
out.append({
|
||||||
|
"name": c["name"],
|
||||||
|
"value": c["value"].strip('"'),
|
||||||
|
"domain": domain,
|
||||||
|
"path": c.get("path", "/"),
|
||||||
|
"expires": c.get("expires", -1),
|
||||||
|
"httpOnly": c.get("httpOnly", False),
|
||||||
|
"secure": c.get("secure", False),
|
||||||
|
"sameSite": c.get("sameSite", "None"),
|
||||||
|
})
|
||||||
|
with open(os.path.join(root, "cookies.json"), "w") as f:
|
||||||
|
json.dump(out, f, indent=2)
|
||||||
|
|
||||||
|
write_source_state(root, profile_dir)
|
||||||
|
sys.stderr.write(f"refresh-linkedin-session: {len(out)} cookies, profile refreshed\n")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
Executable
+41
@@ -0,0 +1,41 @@
|
|||||||
|
//usr/bin/env go run -tags=chats_sync "$0" "$@"; exit
|
||||||
|
//go:build chats_sync
|
||||||
|
//
|
||||||
|
// bin/chats/sync.go - download chat messages to var/chats/<platform>/.
|
||||||
|
//
|
||||||
|
// ./bin/chats/sync.go telegram [--limit N] [--phone PHONE]
|
||||||
|
// ./bin/chats/sync.go linkedin [--limit N] [--refresh]
|
||||||
|
//
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/chats"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
if len(os.Args) < 2 {
|
||||||
|
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
|
||||||
|
os.Exit(2)
|
||||||
|
}
|
||||||
|
platform := os.Args[1]
|
||||||
|
args := os.Args[2:]
|
||||||
|
switch platform {
|
||||||
|
case "telegram":
|
||||||
|
os.Exit(chats.RunSyncTelegram(args))
|
||||||
|
case "linkedin":
|
||||||
|
os.Exit(chats.RunSyncLinkedIn(args))
|
||||||
|
case "whatsapp":
|
||||||
|
fmt.Fprintln(os.Stderr, "chats: WhatsApp not implemented yet")
|
||||||
|
os.Exit(1)
|
||||||
|
case "help", "-h", "--help":
|
||||||
|
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
|
||||||
|
return
|
||||||
|
default:
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: unknown platform %q\n", platform)
|
||||||
|
os.Exit(2)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,69 +0,0 @@
|
|||||||
package main
|
|
||||||
|
|
||||||
import (
|
|
||||||
"context"
|
|
||||||
"flag"
|
|
||||||
"fmt"
|
|
||||||
"os"
|
|
||||||
"os/exec"
|
|
||||||
"strings"
|
|
||||||
"time"
|
|
||||||
)
|
|
||||||
|
|
||||||
func checkLinkedInSession(userDataDir string) (bool, error) {
|
|
||||||
cmd := exec.Command("uvx", "mcp-server-linkedin@latest",
|
|
||||||
"--user-data-dir", userDataDir,
|
|
||||||
"--no-auto-import",
|
|
||||||
"--status",
|
|
||||||
)
|
|
||||||
out, err := cmd.CombinedOutput()
|
|
||||||
if err != nil {
|
|
||||||
return true, fmt.Errorf("status check: %w\n%s", err, string(out))
|
|
||||||
}
|
|
||||||
return !strings.Contains(string(out), "✅"), nil
|
|
||||||
}
|
|
||||||
|
|
||||||
func runSyncLinkedIn(args []string) int {
|
|
||||||
fs := flag.NewFlagSet("chats sync linkedin", flag.ContinueOnError)
|
|
||||||
limit := fs.Int("limit", 0, "max messages per conversation (0 = all)")
|
|
||||||
help := fs.Bool("help", false, "")
|
|
||||||
fs.SetOutput(os.Stderr)
|
|
||||||
if err := fs.Parse(args); err != nil {
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
if *help {
|
|
||||||
fmt.Fprintln(os.Stderr, "usage: chats sync linkedin [--limit N]")
|
|
||||||
return 0
|
|
||||||
}
|
|
||||||
|
|
||||||
userDataDir := envVar("LINKEDIN_USER_DATA_DIR", "")
|
|
||||||
if userDataDir == "" {
|
|
||||||
home, _ := os.UserHomeDir()
|
|
||||||
userDataDir = home + "/.linkedin-mcp/profile"
|
|
||||||
}
|
|
||||||
|
|
||||||
// Check session first
|
|
||||||
loginNeeded, err := checkLinkedInSession(userDataDir)
|
|
||||||
if err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: linkedin status check: %v\n", err)
|
|
||||||
}
|
|
||||||
if loginNeeded {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats: LinkedIn session expired. Run:\n")
|
|
||||||
fmt.Fprintf(os.Stderr, " uvx mcp-server-linkedin@latest --user-data-dir %s --login\n", userDataDir)
|
|
||||||
fmt.Fprintf(os.Stderr, "Then retry 'chats sync linkedin'\n")
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
|
|
||||||
src := NewLinkedInMCPSource(userDataDir)
|
|
||||||
|
|
||||||
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute)
|
|
||||||
defer cancel()
|
|
||||||
|
|
||||||
start := time.Now()
|
|
||||||
if err := src.Sync(ctx, chatsDir(), *limit); err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats sync linkedin: %v\n", err)
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
fmt.Printf("chats sync linkedin: completed in %s\n", time.Since(start).Round(time.Millisecond))
|
|
||||||
return 0
|
|
||||||
}
|
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
// Deprecated shebang mains at bin root (serve.go is tagged brain_serve).
|
||||||
|
package main
|
||||||
@@ -2,10 +2,10 @@
|
|||||||
# bin/docker-entrypoint - run 2dph tools inside the container.
|
# bin/docker-entrypoint - run 2dph tools inside the container.
|
||||||
#
|
#
|
||||||
# brain shell (default)
|
# brain shell (default)
|
||||||
# brain search <q> bin/kb/search
|
# brain search <q> bin/brain/search.go
|
||||||
# brain index bin/kb/index
|
# brain index bin/kb/index --with-mail
|
||||||
# brain watch <dir> watchdog re-indexer (bin/kb/watch)
|
# brain watch <dir> compiled /app/bin/watch (bin/brain/watch.go)
|
||||||
# brain serve async Go HTTP server (bin/serve)
|
# brain serve compiled /app/bin/serve (bin/brain/serve.go)
|
||||||
# brain extract bin/facts/extract (docker×compose pairing)
|
# brain extract bin/facts/extract (docker×compose pairing)
|
||||||
# brain audit bin/facts/audit
|
# brain audit bin/facts/audit
|
||||||
#
|
#
|
||||||
@@ -18,7 +18,7 @@ shift || true
|
|||||||
case "$CMD" in
|
case "$CMD" in
|
||||||
shell) exec bash ;;
|
shell) exec bash ;;
|
||||||
search) exec "$KB_PY" /app/bin/kb/search "$@" ;;
|
search) exec "$KB_PY" /app/bin/kb/search "$@" ;;
|
||||||
index) exec "$KB_PY" /app/bin/kb/index "$@" ;;
|
index) exec "$KB_PY" /app/bin/kb/index --with-mail "$@" ;;
|
||||||
watch) exec /app/bin/watch "$@" ;;
|
watch) exec /app/bin/watch "$@" ;;
|
||||||
serve) exec /app/bin/serve "$@" ;;
|
serve) exec /app/bin/serve "$@" ;;
|
||||||
extract) exec "$KB_PY" /app/bin/facts/extract "$@" ;;
|
extract) exec "$KB_PY" /app/bin/facts/extract "$@" ;;
|
||||||
|
|||||||
+11
-139
@@ -1,154 +1,26 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""git/import - import git history (commits, authors, files) into the brain.
|
"""git/import — deprecated. Use bin/git/import.go (go-git, no git binary).
|
||||||
|
|
||||||
bin/git/import [REPO] import all commits -> leafs + graph
|
bin/git/import.go [REPO] [--json] [--limit N] [--since DATE]
|
||||||
bin/git/import --json emit import leafs as JSON, no write
|
|
||||||
bin/git/import --limit 100 cap commits processed
|
|
||||||
bin/git/import --since 2026-01-01 only recent commits
|
|
||||||
bin/git/import --root DIR run per repo dir under DIR
|
|
||||||
bin/git/import --no-env never read .env anywhere (default: true)
|
|
||||||
|
|
||||||
Reads `git log --no-merges --name-only` from the repo, maps commits to
|
|
||||||
`info` leafs (root=info, type=commit) and writes the version graph
|
|
||||||
`File -[:HAS_VERSION]-> Commit -[:AUTHORED]-> Person` into var/kb.lbug.
|
|
||||||
Idempotent: leaf MERGE by (source,text via leaf_id), graph MERGE by sha.
|
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import json
|
import os
|
||||||
import subprocess
|
|
||||||
import sys
|
import sys
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
ROOT = Path(__file__).resolve().parents[2]
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
|
||||||
|
|
||||||
from kblib import ( # noqa: E402
|
|
||||||
connect, ensure_indexes, init_schema, upsert_leaf,
|
|
||||||
)
|
|
||||||
from gitimport import commits_to_leafs, ensure_git_schema, index_commits, parse_log # noqa: E402
|
|
||||||
|
|
||||||
LOG_FMT = "--format=%x1e%H%x1f%an%x1f%ae%x1f%aI%x1f%s"
|
|
||||||
|
|
||||||
|
|
||||||
def git_log(repo: Path, limit: int = 0, since: str = "") -> str:
|
|
||||||
cmd = ["git", "-C", str(repo), "log", "--no-merges", "--name-only", LOG_FMT]
|
|
||||||
if since:
|
|
||||||
cmd += ["--since", since]
|
|
||||||
if limit:
|
|
||||||
cmd += ["-n", str(limit)]
|
|
||||||
try:
|
|
||||||
out = subprocess.run(cmd, capture_output=True, text=True, timeout=120)
|
|
||||||
except (FileNotFoundError, subprocess.TimeoutExpired):
|
|
||||||
return ""
|
|
||||||
if out.returncode != 0:
|
|
||||||
print(f"git/import: {repo}: {out.stderr.strip()}", file=sys.stderr)
|
|
||||||
return ""
|
|
||||||
return out.stdout
|
|
||||||
|
|
||||||
|
|
||||||
def repo_name(repo: Path) -> str:
|
|
||||||
try:
|
|
||||||
out = subprocess.run(
|
|
||||||
["git", "-C", str(repo), "remote", "get-url", "origin"],
|
|
||||||
capture_output=True, text=True, timeout=20)
|
|
||||||
url = out.stdout.strip()
|
|
||||||
return url.rstrip("/").split("/")[-1].removesuffix(".git") if url else repo.name
|
|
||||||
except (FileNotFoundError, subprocess.TimeoutExpired):
|
|
||||||
return repo.name
|
|
||||||
|
|
||||||
|
|
||||||
def embedder():
|
|
||||||
from model2vec import StaticModel
|
|
||||||
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
|
|
||||||
return lambda text: model.encode([text])[0].astype(float).tolist()
|
|
||||||
|
|
||||||
|
|
||||||
def import_repo(conn, repo: Path, embed, limit: int, since: str,
|
|
||||||
no_write: bool = False) -> tuple[int, int]:
|
|
||||||
raw = git_log(repo, limit, since)
|
|
||||||
commits = parse_log(raw)
|
|
||||||
leafs = commits_to_leafs(commits, repo_name(repo))
|
|
||||||
if no_write:
|
|
||||||
return len(commits), 0
|
|
||||||
written = 0
|
|
||||||
for lf in leafs:
|
|
||||||
query = f"{lf['heading']}\n\n{lf['text']}"
|
|
||||||
emb = embed(lf["text"]) if lf["text"] else None
|
|
||||||
upsert_leaf(conn, text=query, root="info", confidence="confirmed",
|
|
||||||
source=lf["source"], source_rev="git", how="git/import",
|
|
||||||
loc=lf["source"], type_=lf.get("type", "commit"),
|
|
||||||
embedding=emb)
|
|
||||||
written += 1
|
|
||||||
index_commits(conn, commits, repo_name(repo))
|
|
||||||
return len(commits), written
|
|
||||||
|
|
||||||
|
|
||||||
def main(argv: list[str]) -> int:
|
def main(argv: list[str]) -> int:
|
||||||
import argparse
|
print(
|
||||||
p = argparse.ArgumentParser(description="import git history into the brain")
|
"bin/git/import is deprecated; use bin/git/import.go (go-git)",
|
||||||
p.add_argument("repo", nargs="?", default=None)
|
file=sys.stderr,
|
||||||
p.add_argument("--root", default=None, help="directory of repos to import (each git dir separately)")
|
)
|
||||||
p.add_argument("--limit", type=int, default=0)
|
target = ROOT / "bin" / "git" / "import.go"
|
||||||
p.add_argument("--since", default="")
|
os.execvp("go", ["go", "run", str(target), *argv])
|
||||||
p.add_argument("--json", action="store_true")
|
return 1
|
||||||
p.add_argument("--dry-run", action="store_true", help="parse + report, no db write")
|
|
||||||
a = p.parse_args(argv)
|
|
||||||
|
|
||||||
repos: list[Path] = []
|
|
||||||
if a.repo:
|
|
||||||
repos = [Path(a.repo)]
|
|
||||||
elif a.root:
|
|
||||||
root = Path(a.root)
|
|
||||||
if root.is_file():
|
|
||||||
repos = [root]
|
|
||||||
else:
|
|
||||||
repos = [dp for dp in sorted(root.iterdir()) if (dp / ".git").exists() or dp.is_file()]
|
|
||||||
else:
|
|
||||||
repos = [ROOT]
|
|
||||||
|
|
||||||
total_commits = 0
|
|
||||||
results: list[dict] = []
|
|
||||||
if a.dry_run:
|
|
||||||
for repo in repos:
|
|
||||||
if not repo.exists():
|
|
||||||
continue
|
|
||||||
commits = parse_log(git_log(repo, a.limit, a.since))
|
|
||||||
name = repo_name(repo)
|
|
||||||
total_commits += len(commits)
|
|
||||||
results.append({"repo": name, "commits": len(commits),
|
|
||||||
"leafs": len(commits_to_leafs(commits, name)), "path": str(repo)})
|
|
||||||
if a.json:
|
|
||||||
print(json.dumps(results, indent=2))
|
|
||||||
else:
|
|
||||||
for r in results:
|
|
||||||
print(f"{r['repo']:<24} {r['commits']:>5} commits -> {r['leafs']} leafs {r['path']}")
|
|
||||||
return 0
|
|
||||||
|
|
||||||
# Never DROP FTS/VECTOR (ghost catalog). Upsert while indexes exist is OK;
|
|
||||||
# ensure_indexes only CREATEs when missing.
|
|
||||||
db, conn = connect(ROOT / "var" / "kb.lbug", read_only=False)
|
|
||||||
init_schema(conn)
|
|
||||||
embed = embedder()
|
|
||||||
rows: list[dict] = []
|
|
||||||
for repo in repos:
|
|
||||||
if not repo.exists():
|
|
||||||
continue
|
|
||||||
reached, written = import_repo(conn, repo, embed, a.limit, a.since)
|
|
||||||
total_commits += reached
|
|
||||||
rows.append({"repo": repo_name(repo), "commits": reached, "written": written})
|
|
||||||
ensure_indexes(conn)
|
|
||||||
conn.close()
|
|
||||||
db.close()
|
|
||||||
|
|
||||||
if a.json:
|
|
||||||
print(json.dumps(rows, indent=2))
|
|
||||||
else:
|
|
||||||
for r in rows:
|
|
||||||
print(f"imported {r['commits']:>5} commits -> {r['written']} leafs {r['repo']}")
|
|
||||||
print(f"total: {total_commits} commits")
|
|
||||||
return 0
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
sys.exit(main(sys.argv[1:]))
|
sys.exit(main(sys.argv[1:]))
|
||||||
|
|||||||
Executable
+143
@@ -0,0 +1,143 @@
|
|||||||
|
//usr/bin/env go run "$0" "$@"; exit
|
||||||
|
//
|
||||||
|
// bin/git/import.go - read git history with go-git (no git binary).
|
||||||
|
//
|
||||||
|
// ./bin/git/import.go [REPO]
|
||||||
|
// ./bin/git/import.go --json
|
||||||
|
// ./bin/git/import.go --limit 100 --since 2026-01-01
|
||||||
|
// ./bin/git/import.go --root DIR
|
||||||
|
//
|
||||||
|
// Conversion only: prints commit leafs. Brain write is bin/brain/index.go.
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strconv"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
|
"github.com/eSlider/2dph/internal/gitlog"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(run(os.Args[1:]))
|
||||||
|
}
|
||||||
|
|
||||||
|
func run(args []string) int {
|
||||||
|
var repo, root, since string
|
||||||
|
limit := 0
|
||||||
|
jsonOut := false
|
||||||
|
i := 0
|
||||||
|
for i < len(args) {
|
||||||
|
a := args[i]
|
||||||
|
switch {
|
||||||
|
case a == "--json":
|
||||||
|
jsonOut = true
|
||||||
|
case a == "--limit" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
n, err := strconv.Atoi(args[i])
|
||||||
|
if err != nil || n < 0 {
|
||||||
|
fmt.Fprintf(os.Stderr, "git/import: --limit must be a non-negative integer\n")
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
limit = n
|
||||||
|
case a == "--since" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
since = args[i]
|
||||||
|
case a == "--root" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
root = args[i]
|
||||||
|
case a == "-h" || a == "--help":
|
||||||
|
fmt.Fprintln(os.Stderr, `usage: bin/git/import.go [REPO] [--json] [--limit N] [--since DATE] [--root DIR]`)
|
||||||
|
return 0
|
||||||
|
case len(a) > 0 && a[0] != '-':
|
||||||
|
repo = a
|
||||||
|
default:
|
||||||
|
fmt.Fprintf(os.Stderr, "git/import: unknown flag %s\n", a)
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
i++
|
||||||
|
}
|
||||||
|
|
||||||
|
var sinceT time.Time
|
||||||
|
if since != "" {
|
||||||
|
var err error
|
||||||
|
sinceT, err = parseSince(since)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
repos := []string{}
|
||||||
|
if repo != "" {
|
||||||
|
repos = []string{repo}
|
||||||
|
} else if root != "" {
|
||||||
|
entries, err := os.ReadDir(root)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
for _, e := range entries {
|
||||||
|
p := filepath.Join(root, e.Name())
|
||||||
|
if _, err := os.Stat(filepath.Join(p, ".git")); err == nil {
|
||||||
|
repos = append(repos, p)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
repos = []string{cmdbin.Root()}
|
||||||
|
}
|
||||||
|
|
||||||
|
opt := gitlog.Options{Limit: limit, Since: sinceT}
|
||||||
|
type row struct {
|
||||||
|
Repo string `json:"repo"`
|
||||||
|
Path string `json:"path"`
|
||||||
|
Commits int `json:"commits"`
|
||||||
|
Leafs []gitlog.Leaf `json:"leafs,omitempty"`
|
||||||
|
}
|
||||||
|
var rows []row
|
||||||
|
for _, p := range repos {
|
||||||
|
name, err := gitlog.RepoName(p)
|
||||||
|
if err != nil && name == "" {
|
||||||
|
fmt.Fprintf(os.Stderr, "git/import: %s: %v\n", p, err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
cs, err := gitlog.Log(p, opt)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "git/import: %s: %v\n", p, err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
leafs := make([]gitlog.Leaf, 0, len(cs))
|
||||||
|
for _, c := range cs {
|
||||||
|
leafs = append(leafs, gitlog.ToLeaf(c, name))
|
||||||
|
}
|
||||||
|
rows = append(rows, row{Repo: name, Path: p, Commits: len(cs), Leafs: leafs})
|
||||||
|
}
|
||||||
|
|
||||||
|
if jsonOut {
|
||||||
|
enc := json.NewEncoder(os.Stdout)
|
||||||
|
enc.SetIndent("", " ")
|
||||||
|
enc.SetEscapeHTML(false)
|
||||||
|
if err := enc.Encode(rows); err != nil {
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
for _, r := range rows {
|
||||||
|
fmt.Printf("%-24s %5d commits %s\n", r.Repo, r.Commits, r.Path)
|
||||||
|
}
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
func parseSince(s string) (time.Time, error) {
|
||||||
|
for _, layout := range []string{time.RFC3339, "2006-01-02"} {
|
||||||
|
if t, err := time.Parse(layout, s); err == nil {
|
||||||
|
return t, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return time.Time{}, fmt.Errorf("cannot parse --since %q", s)
|
||||||
|
}
|
||||||
+19
-3
@@ -26,6 +26,7 @@ from kblib import ( # noqa: E402
|
|||||||
open_readonly, stats,
|
open_readonly, stats,
|
||||||
)
|
)
|
||||||
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
|
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
|
||||||
|
from mailleafs import from_mail_root # noqa: E402
|
||||||
|
|
||||||
CORPUS_DEFAULTS = ["README.md", "PLAN.md", "AGENTS.md", "docs", "skills"]
|
CORPUS_DEFAULTS = ["README.md", "PLAN.md", "AGENTS.md", "docs", "skills"]
|
||||||
|
|
||||||
@@ -99,6 +100,9 @@ def main(argv: list[str]) -> int:
|
|||||||
p = argparse.ArgumentParser(description="build the 2dph brain index")
|
p = argparse.ArgumentParser(description="build the 2dph brain index")
|
||||||
p.add_argument("--corpus", action="append", help="extra markdown dir/file to index (may repeat)")
|
p.add_argument("--corpus", action="append", help="extra markdown dir/file to index (may repeat)")
|
||||||
p.add_argument("--rebuild", action="store_true", help="fresh db + indexes")
|
p.add_argument("--rebuild", action="store_true", help="fresh db + indexes")
|
||||||
|
p.add_argument("--with-mail", action="store_true", help="include var/mail message.md leafs")
|
||||||
|
p.add_argument("--since", default="", help="with --with-mail, only messages dated >= YYYY-MM-DD")
|
||||||
|
p.add_argument("--dry-run", action="store_true", help="count leafs, write nothing")
|
||||||
p.add_argument(
|
p.add_argument(
|
||||||
"--skip-indexes",
|
"--skip-indexes",
|
||||||
action="store_true",
|
action="store_true",
|
||||||
@@ -109,14 +113,26 @@ def main(argv: list[str]) -> int:
|
|||||||
a = p.parse_args(argv)
|
a = p.parse_args(argv)
|
||||||
|
|
||||||
from kblib import DB_PATH, VAR
|
from kblib import DB_PATH, VAR
|
||||||
VAR.mkdir(exist_ok=True)
|
|
||||||
if a.rebuild and DB_PATH.exists():
|
|
||||||
DB_PATH.unlink()
|
|
||||||
|
|
||||||
leafs = load_corpus(ROOT)
|
leafs = load_corpus(ROOT)
|
||||||
if a.corpus:
|
if a.corpus:
|
||||||
for source in a.corpus:
|
for source in a.corpus:
|
||||||
leafs.extend(load_corpus_glob(source))
|
leafs.extend(load_corpus_glob(source))
|
||||||
|
mail_n = 0
|
||||||
|
if a.with_mail:
|
||||||
|
mail = from_mail_root(ROOT / "var" / "mail", since=a.since)
|
||||||
|
mail_n = len(mail)
|
||||||
|
leafs.extend(mail)
|
||||||
|
|
||||||
|
if a.dry_run:
|
||||||
|
msg = {"indexed": 0, "corpus_total": len(leafs), "mail_leafs": mail_n, "dry_run": True}
|
||||||
|
print(json.dumps(msg, indent=2) if a.json else
|
||||||
|
f"brain/index: {len(leafs)} leafs would be indexed (mail={mail_n})")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
VAR.mkdir(exist_ok=True)
|
||||||
|
if a.rebuild and DB_PATH.exists():
|
||||||
|
DB_PATH.unlink()
|
||||||
|
|
||||||
db, conn = connect(DB_PATH, read_only=False)
|
db, conn = connect(DB_PATH, read_only=False)
|
||||||
init_schema(conn)
|
init_schema(conn)
|
||||||
|
|||||||
+24
-21
@@ -1,34 +1,37 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
# bin/kb/search - Go deduction search over the brain (model served by daemon).
|
# bin/kb/search — deprecated wrapper. Use bin/brain/search.go.
|
||||||
# Builds the kbsearch binary on first run / when source changes, then execs it.
|
# Sets CGO for ladybug, builds a binary (embed daemon needs a real executable),
|
||||||
|
# then execs it. Prints one deprecation line.
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
KB="$(cd "$(dirname "$0")/../.." && pwd)"
|
ROOT="$(cd "$(dirname "$0")/../.." && pwd)"
|
||||||
BIN="$KB/var/bin/kbsearch"
|
BIN="$ROOT/var/bin/brain-search"
|
||||||
SRC="$KB/bin/kbsearch"
|
SRC="$ROOT/internal/brain"
|
||||||
|
CMD="$ROOT/bin/brain"
|
||||||
|
|
||||||
mkdir -p "$KB/var/bin"
|
mkdir -p "$ROOT/var/bin"
|
||||||
|
|
||||||
# Rebuild if binary missing or any .go source newer
|
|
||||||
need_build=0
|
need_build=0
|
||||||
if [ ! -x "$BIN" ]; then
|
if [ ! -x "$BIN" ]; then
|
||||||
need_build=1
|
need_build=1
|
||||||
else
|
else
|
||||||
# Check if any .go in kbsearch is newer than binary
|
while IFS= read -r -d '' f; do
|
||||||
while IFS= read -r -d '' f; do
|
if [ "$f" -nt "$BIN" ]; then
|
||||||
if [ "$f" -nt "$BIN" ]; then
|
need_build=1
|
||||||
need_build=1
|
break
|
||||||
break
|
fi
|
||||||
fi
|
done < <(find "$SRC" "$CMD" -name '*.go' -print0 2>/dev/null)
|
||||||
done < <(find "$SRC" -name '*.go' -print0 2>/dev/null)
|
|
||||||
fi
|
fi
|
||||||
|
|
||||||
if [ "$need_build" -eq 1 ]; then
|
if [ "$need_build" -eq 1 ]; then
|
||||||
echo "Building kbsearch..." >&2
|
echo "Building brain/search..." >&2
|
||||||
(cd "$SRC" && \
|
(
|
||||||
CGO_CFLAGS="-I$KB/lib-ladybug" \
|
cd "$ROOT" &&
|
||||||
CGO_LDFLAGS="-L$KB/lib-ladybug -Wl,-rpath,$KB/lib-ladybug" \
|
CGO_CFLAGS="-I$ROOT/lib-ladybug" \
|
||||||
go build -tags system_ladybug -o "$BIN" .) || exit 1
|
CGO_LDFLAGS="-L$ROOT/lib-ladybug -Wl,-rpath,$ROOT/lib-ladybug" \
|
||||||
|
go build -tags system_ladybug -o "$BIN" ./bin/brain
|
||||||
|
) || exit 1
|
||||||
fi
|
fi
|
||||||
|
|
||||||
exec "$BIN" "$@"
|
echo "bin/kb/search is deprecated; use bin/brain/search.go" >&2
|
||||||
|
exec "$BIN" "$@"
|
||||||
|
|||||||
+1
-1
@@ -1,5 +1,5 @@
|
|||||||
//usr/bin/env go run "$0" "$@"; exit
|
//usr/bin/env go run "$0" "$@"; exit
|
||||||
// bin/kb/watch.go - re-index the 2dph brain when corpus files change.
|
// bin/kb/watch.go — deprecated. Use bin/brain/watch.go.
|
||||||
//
|
//
|
||||||
// Usage:
|
// Usage:
|
||||||
//
|
//
|
||||||
|
|||||||
@@ -1,23 +0,0 @@
|
|||||||
module github.com/eSlider/2dph/bin/kbsearch
|
|
||||||
|
|
||||||
go 1.26.0
|
|
||||||
|
|
||||||
require (
|
|
||||||
github.com/LadybugDB/go-ladybug v0.17.0
|
|
||||||
github.com/chewxy/math32 v1.11.2
|
|
||||||
github.com/daulet/tokenizers v1.27.0
|
|
||||||
)
|
|
||||||
|
|
||||||
require (
|
|
||||||
github.com/apache/arrow-go/v18 v18.6.0 // indirect
|
|
||||||
github.com/goccy/go-json v0.10.6 // indirect
|
|
||||||
github.com/google/flatbuffers v25.12.19+incompatible // indirect
|
|
||||||
github.com/google/uuid v1.6.0 // indirect
|
|
||||||
github.com/klauspost/compress v1.18.5 // indirect
|
|
||||||
github.com/klauspost/cpuid/v2 v2.3.0 // indirect
|
|
||||||
github.com/pierrec/lz4/v4 v4.1.26 // indirect
|
|
||||||
github.com/shopspring/decimal v1.4.0 // indirect
|
|
||||||
github.com/zeebo/xxh3 v1.1.0 // indirect
|
|
||||||
golang.org/x/exp v0.0.0-20260112195511-716be5621a96 // indirect
|
|
||||||
golang.org/x/sys v0.43.0 // indirect
|
|
||||||
)
|
|
||||||
@@ -1,44 +0,0 @@
|
|||||||
github.com/LadybugDB/go-ladybug v0.17.0 h1:RXDbkBjrbRmLdEbhGl4CLOIEzSt09gbP0n9UbKDEfwI=
|
|
||||||
github.com/LadybugDB/go-ladybug v0.17.0/go.mod h1:GeIXmE8XyF5TFS94NAuTag7vgCC+no/HTBMRA6Rd5Cs=
|
|
||||||
github.com/andybalholm/brotli v1.2.1 h1:R+f5xP285VArJDRgowrfb9DqL18yVK0gKAW/F+eTWro=
|
|
||||||
github.com/andybalholm/brotli v1.2.1/go.mod h1:rzTDkvFWvIrjDXZHkuS16NPggd91W3kUSvPlQ1pLaKY=
|
|
||||||
github.com/apache/arrow-go/v18 v18.6.0 h1:GX/Jyd3R7mCLiECAwY9FWbbaYblie2WXBSz4Sw8fNpM=
|
|
||||||
github.com/apache/arrow-go/v18 v18.6.0/go.mod h1:gm3MiPpY82fLYK5VKPB3WoJbsiLVDfT7flD5/vHReKw=
|
|
||||||
github.com/apache/thrift v0.22.0 h1:r7mTJdj51TMDe6RtcmNdQxgn9XcyfGDOzegMDRg47uc=
|
|
||||||
github.com/apache/thrift v0.22.0/go.mod h1:1e7J/O1Ae6ZQMTYdy9xa3w9k+XHWPfRvdPyJeynQ+/g=
|
|
||||||
github.com/chewxy/math32 v1.11.2 h1:IufN08Zwr1NKuWfY+4Tz55BcwKmyKKNdOP7KtumehnM=
|
|
||||||
github.com/chewxy/math32 v1.11.2/go.mod h1:dOB2rcuFrCn6UHrze36WSLVPKtzPMRAQvBvUwkSsLqs=
|
|
||||||
github.com/daulet/tokenizers v1.27.0 h1:MmFYAEDFz69s/nNQfHg59DWqHz3v94m99kEZ/JbL+s4=
|
|
||||||
github.com/daulet/tokenizers v1.27.0/go.mod h1:YjFY1o1HGMyWkQgbXJDghhvke/yFDp2vGdIO2hYs4MQ=
|
|
||||||
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
|
|
||||||
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
|
||||||
github.com/goccy/go-json v0.10.6 h1:p8HrPJzOakx/mn/bQtjgNjdTcN+/S6FcG2CTtQOrHVU=
|
|
||||||
github.com/goccy/go-json v0.10.6/go.mod h1:oq7eo15ShAhp70Anwd5lgX2pLfOS3QCiwU/PULtXL6M=
|
|
||||||
github.com/google/flatbuffers v25.12.19+incompatible h1:haMV2JRRJCe1998HeW/p0X9UaMTK6SDo0ffLn2+DbLs=
|
|
||||||
github.com/google/flatbuffers v25.12.19+incompatible/go.mod h1:1AeVuKshWv4vARoZatz6mlQ0JxURH0Kv5+zNeJKJCa8=
|
|
||||||
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
|
|
||||||
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
|
|
||||||
github.com/klauspost/compress v1.18.5 h1:/h1gH5Ce+VWNLSWqPzOVn6XBO+vJbCNGvjoaGBFW2IE=
|
|
||||||
github.com/klauspost/compress v1.18.5/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
|
|
||||||
github.com/klauspost/cpuid/v2 v2.3.0 h1:S4CRMLnYUhGeDFDqkGriYKdfoFlDnMtqTiI/sFzhA9Y=
|
|
||||||
github.com/klauspost/cpuid/v2 v2.3.0/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0=
|
|
||||||
github.com/pierrec/lz4/v4 v4.1.26 h1:GrpZw1gZttORinvzBdXPUXATeqlJjqUG/D87TKMnhjY=
|
|
||||||
github.com/pierrec/lz4/v4 v4.1.26/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
|
|
||||||
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
|
|
||||||
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
|
|
||||||
github.com/shopspring/decimal v1.4.0 h1:bxl37RwXBklmTi0C79JfXCEBD1cqqHt0bbgBAGFp81k=
|
|
||||||
github.com/shopspring/decimal v1.4.0/go.mod h1:gawqmDU56v4yIKSwfBSFip1HdCCXN8/+DMd9qYNcwME=
|
|
||||||
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
|
|
||||||
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
|
|
||||||
github.com/zeebo/assert v1.3.0 h1:g7C04CbJuIDKNPFHmsk4hwZDO5O+kntRxzaUoNXj+IQ=
|
|
||||||
github.com/zeebo/assert v1.3.0/go.mod h1:Pq9JiuJQpG8JLJdtkwrJESF0Foym2/D9XMU5ciN/wJ0=
|
|
||||||
github.com/zeebo/xxh3 v1.1.0 h1:s7DLGDK45Dyfg7++yxI0khrfwq9661w9EN78eP/UZVs=
|
|
||||||
github.com/zeebo/xxh3 v1.1.0/go.mod h1:IisAie1LELR4xhVinxWS5+zf1lA4p0MW4T+w+W07F5s=
|
|
||||||
golang.org/x/exp v0.0.0-20260112195511-716be5621a96 h1:Z/6YuSHTLOHfNFdb8zVZomZr7cqNgTJvA8+Qz75D8gU=
|
|
||||||
golang.org/x/exp v0.0.0-20260112195511-716be5621a96/go.mod h1:nzimsREAkjBCIEFtHiYkrJyT+2uy9YZJB7H1k68CXZU=
|
|
||||||
golang.org/x/sys v0.43.0 h1:Rlag2XtaFTxp19wS8MXlJwTvoh8ArU6ezoyFsMyCTNI=
|
|
||||||
golang.org/x/sys v0.43.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
|
|
||||||
gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4=
|
|
||||||
gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E=
|
|
||||||
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
|
|
||||||
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
|
|
||||||
@@ -1,44 +0,0 @@
|
|||||||
// bin/kbsearch - the Go implementation of bin/kb/search (nested module so the
|
|
||||||
// root `go test ./...` and CI never compile it against native ladyships).
|
|
||||||
//
|
|
||||||
// Usage (built/run by ./bin/kb/search):
|
|
||||||
//
|
|
||||||
// kbsearch "query" [--root facts|info] [--repo P] [-n N] [--json]
|
|
||||||
// kbsearch serve [port] start the embedding daemon
|
|
||||||
// kbsearch --list-model print the resolved model dir
|
|
||||||
//
|
|
||||||
// The potion-multilingual model is loaded only in `serve`; a CLI reuses the
|
|
||||||
// daemon over localhost HTTP (falling back to in-process embedding).
|
|
||||||
package main
|
|
||||||
|
|
||||||
import (
|
|
||||||
"fmt"
|
|
||||||
"log"
|
|
||||||
"os"
|
|
||||||
"strconv"
|
|
||||||
)
|
|
||||||
|
|
||||||
func main() {
|
|
||||||
if len(os.Args) > 1 && os.Args[1] == "serve" {
|
|
||||||
port := 17830
|
|
||||||
if len(os.Args) > 2 {
|
|
||||||
if p, err := strconv.Atoi(os.Args[2]); err == nil {
|
|
||||||
port = p
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if err := serve(port); err != nil {
|
|
||||||
log.Fatalf("kbsearch serve: %v", err)
|
|
||||||
}
|
|
||||||
return
|
|
||||||
}
|
|
||||||
if len(os.Args) > 1 && os.Args[1] == "--list-model" {
|
|
||||||
dir, err := modelDir()
|
|
||||||
if err != nil {
|
|
||||||
fmt.Fprintln(os.Stderr, err)
|
|
||||||
os.Exit(1)
|
|
||||||
}
|
|
||||||
fmt.Println(dir)
|
|
||||||
return
|
|
||||||
}
|
|
||||||
os.Exit(runSearch(os.Args[1:]))
|
|
||||||
}
|
|
||||||
@@ -1,16 +0,0 @@
|
|||||||
// Common types and helpers for kbsearch.
|
|
||||||
package main
|
|
||||||
|
|
||||||
import "os"
|
|
||||||
|
|
||||||
func eps() string { return os.Getenv("KBTEST_EPS") }
|
|
||||||
|
|
||||||
// Hit is one search result, mirroring the python script's dict shape.
|
|
||||||
type Hit struct {
|
|
||||||
ID string `json:"id"`
|
|
||||||
Text string `json:"text"`
|
|
||||||
Root string `json:"root"`
|
|
||||||
Source string `json:"-"` // for repo filtering, not in output
|
|
||||||
Score float64 `json:"score"`
|
|
||||||
Snippet string `json:"snippet,omitempty"`
|
|
||||||
}
|
|
||||||
+2
-2
@@ -15,8 +15,8 @@ Writes one directory per message: var/mail/{folder}/{message_id}/
|
|||||||
attachments/ raw attachment files (zips unpacked to _unpacked/)
|
attachments/ raw attachment files (zips unpacked to _unpacked/)
|
||||||
attachments/*.md converted attachment content
|
attachments/*.md converted attachment content
|
||||||
|
|
||||||
Indexing is a separate step (bin/mail/index_mail): conversion can crash in
|
Indexing is a separate step (`bin/brain/index.go --rebuild`): conversion can
|
||||||
native docling and must not leave the brain DB mid-transaction.
|
crash in native docling and must not leave the brain DB mid-transaction.
|
||||||
|
|
||||||
Requires ONLYOFFICE_URL/USER/PASS in .env (or env). Idempotent: a message
|
Requires ONLYOFFICE_URL/USER/PASS in .env (or env). Idempotent: a message
|
||||||
already present (message.md exists) is skipped unless --force.
|
already present (message.md exists) is skipped unless --force.
|
||||||
|
|||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=mail_import "$0" "$@"; exit
|
||||||
|
//go:build mail_import
|
||||||
|
//
|
||||||
|
// bin/mail/import.go - message.json → markdown (no brain write).
|
||||||
|
//
|
||||||
|
// ./bin/mail/import.go --from-raw var/mail
|
||||||
|
//
|
||||||
|
// Indexing is bin/brain/index.go --rebuild, not this command.
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(cmdbin.ExecFile("bin/mail/import", os.Args[1:]))
|
||||||
|
}
|
||||||
+11
-120
@@ -1,135 +1,26 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""mail/index_mail - rebuild the brain with every markdown under var/mail.
|
"""mail/index_mail — deprecated. Use bin/brain/index.go --rebuild --with-mail.
|
||||||
|
|
||||||
Ladybug corrupts its WAL when brand-new leafs are bulk-inserted while the
|
Ladybug corrupts its WAL on bulk-insert into an already-indexed DB, so this
|
||||||
FTS/VECTOR indexes already exist, so indexing ALWAYS runs as a fresh rebuild
|
shim always rebuilds (repo corpus + var/mail). Conversion stays in mail/import.
|
||||||
(repo corpus + var/mail), matching the proven-safe `kb/index --rebuild` path.
|
|
||||||
Conversion and indexing stay separate: conversion can crash in native docling
|
|
||||||
and must not leave the brain DB mid-transaction.
|
|
||||||
|
|
||||||
bin/mail/index_mail rebuild the index incl. all mail
|
|
||||||
bin/mail/index_mail --dry-run count without writing
|
|
||||||
bin/mail/index_mail --limit N cap messages included
|
|
||||||
bin/mail/index_mail --since D only messages dated >= D (YYYY-MM-DD)
|
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import argparse
|
import os
|
||||||
import json
|
|
||||||
import sys
|
import sys
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
ROOT = Path(__file__).resolve().parents[2]
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
|
||||||
|
|
||||||
from kblib import DB_PATH, VAR, connect, ensure_indexes, init_schema, stats, upsert_leaf # noqa: E402
|
|
||||||
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
|
|
||||||
|
|
||||||
|
|
||||||
def msg_date(md: Path) -> str:
|
|
||||||
j = md.parent / "message.json"
|
|
||||||
try:
|
|
||||||
d = json.loads(j.read_text(encoding="utf-8"))
|
|
||||||
return (d.get("receivedDate") or d.get("receivedAt") or "")[:10]
|
|
||||||
except Exception:
|
|
||||||
return ""
|
|
||||||
|
|
||||||
|
|
||||||
def mail_leafs(limit: int, since: str, repo: str = "ooMail") -> list[dict]:
|
|
||||||
root = ROOT / "var" / "mail"
|
|
||||||
mds = sorted(root.rglob("message.md"))
|
|
||||||
if since:
|
|
||||||
mds = [m for m in mds if msg_date(m) >= since]
|
|
||||||
if limit:
|
|
||||||
mds = mds[:limit]
|
|
||||||
leafs: list[dict] = []
|
|
||||||
for md in mds:
|
|
||||||
files = [md] + sorted((md.parent / "attachments").glob("*.md"))
|
|
||||||
for f in files:
|
|
||||||
if not f.exists():
|
|
||||||
continue
|
|
||||||
for lf in to_all(read_markdown(f), f, repo=repo):
|
|
||||||
lf["source"] = f"ooMail:{md.parent.name}:{f.name}"
|
|
||||||
lf["how"] = "mail/import"
|
|
||||||
leafs.append(lf)
|
|
||||||
return leafs
|
|
||||||
|
|
||||||
|
|
||||||
def main(argv: list[str]) -> int:
|
def main(argv: list[str]) -> int:
|
||||||
p = argparse.ArgumentParser(description="rebuild the brain incl. all mail")
|
print(
|
||||||
p.add_argument("--dry-run", action="store_true", help="count only, write nothing")
|
"bin/mail/index_mail is deprecated; use bin/brain/index.go --rebuild --with-mail",
|
||||||
p.add_argument("--limit", type=int, default=0, help="cap messages included")
|
file=sys.stderr,
|
||||||
p.add_argument("--since", default="", help="only messages dated >= YYYY-MM-DD")
|
)
|
||||||
p.add_argument("--json", action="store_true")
|
index = ROOT / "bin" / "kb" / "index"
|
||||||
a = p.parse_args(argv)
|
os.execv(sys.executable, [sys.executable, str(index), "--rebuild", "--with-mail", *argv])
|
||||||
|
return 1
|
||||||
mail = mail_leafs(a.limit, a.since)
|
|
||||||
if a.dry_run:
|
|
||||||
print(f"mail/index_mail: {len(mail)} mail leafs would be indexed")
|
|
||||||
return 0
|
|
||||||
|
|
||||||
# Fresh rebuild: delete DB, index repo corpus + mail, create indexes once
|
|
||||||
# at the end. Never insert into an already-indexed DB (WAL corruption).
|
|
||||||
VAR.mkdir(exist_ok=True)
|
|
||||||
if DB_PATH.exists():
|
|
||||||
DB_PATH.unlink()
|
|
||||||
|
|
||||||
corpus = _load_corpus()
|
|
||||||
leafs = corpus + mail
|
|
||||||
|
|
||||||
db, conn = connect(DB_PATH, read_only=False)
|
|
||||||
init_schema(conn)
|
|
||||||
embed = _embedder()
|
|
||||||
done, total = _index_leafs(conn, leafs, embed)
|
|
||||||
ensure_indexes(conn)
|
|
||||||
s = stats(conn)
|
|
||||||
conn.close()
|
|
||||||
db.close()
|
|
||||||
|
|
||||||
result = {"indexed": done, "corpus_total": total, "mail_leafs": len(mail),
|
|
||||||
**{k: v for k, v in s.items() if k in ("total", "by_root")}}
|
|
||||||
print(json.dumps(result, indent=2) if a.json else
|
|
||||||
f"mail/index_mail: indexed {done}/{total} leafs (mail={len(mail)}); db total {s['total']}")
|
|
||||||
return 0
|
|
||||||
|
|
||||||
|
|
||||||
CORPUS_DEFAULTS = ["README.md", "PLAN.md", "AGENTS.md", "docs", "skills"]
|
|
||||||
|
|
||||||
|
|
||||||
def _load_corpus() -> list[dict]:
|
|
||||||
files: list[Path] = []
|
|
||||||
for entry in CORPUS_DEFAULTS:
|
|
||||||
p = ROOT / entry
|
|
||||||
if p.is_file():
|
|
||||||
files.append(p)
|
|
||||||
elif p.is_dir():
|
|
||||||
files.extend(walk_markdown(p))
|
|
||||||
leafs: list[dict] = []
|
|
||||||
for path in files:
|
|
||||||
try:
|
|
||||||
leafs.extend(to_all(read_markdown(path), path, repo="eSlider/2dph"))
|
|
||||||
except OSError as e:
|
|
||||||
print(f"mail/index_mail: skip {path}: {e}", file=sys.stderr)
|
|
||||||
return leafs
|
|
||||||
|
|
||||||
|
|
||||||
def _index_leafs(conn, leafs: list[dict], embed_fn) -> tuple[int, int]:
|
|
||||||
count = 0
|
|
||||||
for lf in leafs:
|
|
||||||
query = f"{lf['heading']}\n\n{lf['text']}"
|
|
||||||
emb = embed_fn(lf["text"]) if lf["text"] else None
|
|
||||||
upsert_leaf(conn, text=query, root="info", confidence="confirmed",
|
|
||||||
source=lf["source"], source_rev="mail" if lf.get("how") == "mail/import" else "working-tree",
|
|
||||||
how=lf.get("how", "kb/index"), loc=lf["source"], type_=lf.get("type", "reference"),
|
|
||||||
embedding=emb)
|
|
||||||
count += 1
|
|
||||||
return count, len(leafs)
|
|
||||||
|
|
||||||
|
|
||||||
def _embedder():
|
|
||||||
from model2vec import StaticModel
|
|
||||||
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
|
|
||||||
return lambda text: model.encode([text])[0].astype(float).tolist()
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|||||||
+1
-1
@@ -6,7 +6,7 @@
|
|||||||
// ./bin/mail/sync.go --dry-run
|
// ./bin/mail/sync.go --dry-run
|
||||||
//
|
//
|
||||||
// Writes raw message.json + attachments under var/mail/<folder>/<id>/; run
|
// Writes raw message.json + attachments under var/mail/<folder>/<id>/; run
|
||||||
// bin/mail/import --from-raw afterwards to convert everything to markdown.
|
// bin/mail/import.go --from-raw afterwards to convert everything to markdown.
|
||||||
//
|
//
|
||||||
// Shebang trick: first line is a Go `//` comment; the real code lives in the
|
// Shebang trick: first line is a Go `//` comment; the real code lives in the
|
||||||
// importable package (module path, never a relative import).
|
// importable package (module path, never a relative import).
|
||||||
|
|||||||
@@ -32,6 +32,7 @@ func ParseCLI(args []string) (CLIConfig, int, error) {
|
|||||||
offset = fs.Int("offset", 0, "skip first N messages per source")
|
offset = fs.Int("offset", 0, "skip first N messages per source")
|
||||||
force = fs.Bool("force", false, "overwrite existing message.json + attachments")
|
force = fs.Bool("force", false, "overwrite existing message.json + attachments")
|
||||||
dryRun = fs.Bool("dry-run", false, "list message counts without writing")
|
dryRun = fs.Bool("dry-run", false, "list message counts without writing")
|
||||||
|
query = fs.String("query", "in:inbox", "Gmail search query (gmail source only)")
|
||||||
srcs = fs.String("source", "onlyoffice", "comma list: onlyoffice,gmail (default onlyoffice)")
|
srcs = fs.String("source", "onlyoffice", "comma list: onlyoffice,gmail (default onlyoffice)")
|
||||||
help = fs.Bool("help", false, "usage")
|
help = fs.Bool("help", false, "usage")
|
||||||
)
|
)
|
||||||
@@ -60,6 +61,7 @@ func ParseCLI(args []string) (CLIConfig, int, error) {
|
|||||||
Offset: *offset,
|
Offset: *offset,
|
||||||
Force: *force,
|
Force: *force,
|
||||||
DryRun: *dryRun,
|
DryRun: *dryRun,
|
||||||
|
Query: *query,
|
||||||
Policy: RetryPolicy{},
|
Policy: RetryPolicy{},
|
||||||
}
|
}
|
||||||
cli := CLIConfig{Sync: cfg, Env: *env, Sources: *srcs}
|
cli := CLIConfig{Sync: cfg, Env: *env, Sources: *srcs}
|
||||||
@@ -95,7 +97,7 @@ func Main(args []string) int {
|
|||||||
return code
|
return code
|
||||||
}
|
}
|
||||||
if cli.Help {
|
if cli.Help {
|
||||||
fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]")
|
fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail] [--query GMAIL_Q] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]")
|
||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
|
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
|
||||||
|
|||||||
+17
-4
@@ -112,6 +112,7 @@ type SyncConfig struct {
|
|||||||
Offset int // skip first N messages per source
|
Offset int // skip first N messages per source
|
||||||
Force bool // overwrite existing message.json + attachments
|
Force bool // overwrite existing message.json + attachments
|
||||||
DryRun bool // list without writing
|
DryRun bool // list without writing
|
||||||
|
Query string // Gmail search query; default in:inbox
|
||||||
Policy RetryPolicy
|
Policy RetryPolicy
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -136,9 +137,17 @@ type ooSource struct {
|
|||||||
c *OOClient
|
c *OOClient
|
||||||
page int
|
page int
|
||||||
}
|
}
|
||||||
|
// gmailAPI is the Gmail client surface gmailSource needs. *GmailClient implements it.
|
||||||
|
type gmailAPI interface {
|
||||||
|
ListIDs(ctx context.Context, q string, maxIDs int, pageToken string) ([]string, string, error)
|
||||||
|
GetMessage(ctx context.Context, id string) (*Message, error)
|
||||||
|
DownloadAttachment(ctx context.Context, msgID, attID string) ([]byte, error)
|
||||||
|
}
|
||||||
|
|
||||||
type gmailSource struct {
|
type gmailSource struct {
|
||||||
c *GmailClient
|
c gmailAPI
|
||||||
cur string
|
cur string
|
||||||
|
query string
|
||||||
}
|
}
|
||||||
|
|
||||||
func (s *ooSource) Folder() string { return "inbox" }
|
func (s *ooSource) Folder() string { return "inbox" }
|
||||||
@@ -171,7 +180,11 @@ func (s *ooSource) DownloadAttachment(ctx context.Context, msg *Message, att Att
|
|||||||
}
|
}
|
||||||
|
|
||||||
func (s *gmailSource) ListIDs(ctx context.Context, limit int, cursor string) ([]string, string, error) {
|
func (s *gmailSource) ListIDs(ctx context.Context, limit int, cursor string) ([]string, string, error) {
|
||||||
ids, next, err := s.c.ListIDs(ctx, "in:inbox", limit, cursor)
|
q := s.query
|
||||||
|
if q == "" {
|
||||||
|
q = "in:inbox"
|
||||||
|
}
|
||||||
|
ids, next, err := s.c.ListIDs(ctx, q, limit, cursor)
|
||||||
return ids, next, err
|
return ids, next, err
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -207,7 +220,7 @@ func Run(ctx context.Context, cfg SyncConfig) (*SyncStats, error) {
|
|||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, fmt.Errorf("gmail init: %w", err)
|
return nil, fmt.Errorf("gmail init: %w", err)
|
||||||
}
|
}
|
||||||
sources = append(sources, &gmailSource{c: gm})
|
sources = append(sources, &gmailSource{c: gm, query: cfg.Query})
|
||||||
}
|
}
|
||||||
if len(sources) == 0 {
|
if len(sources) == 0 {
|
||||||
return nil, errors.New("sync: no source configured (need OO, Gmail, or both)")
|
return nil, errors.New("sync: no source configured (need OO, Gmail, or both)")
|
||||||
|
|||||||
@@ -221,6 +221,71 @@ func TestCollectParts(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
type fakeGmailAPI struct {
|
||||||
|
lastQ string
|
||||||
|
lastLimit int
|
||||||
|
ids []string
|
||||||
|
}
|
||||||
|
|
||||||
|
func (f *fakeGmailAPI) ListIDs(_ context.Context, q string, maxIDs int, _ string) ([]string, string, error) {
|
||||||
|
f.lastQ = q
|
||||||
|
f.lastLimit = maxIDs
|
||||||
|
return f.ids, "", nil
|
||||||
|
}
|
||||||
|
func (f *fakeGmailAPI) GetMessage(context.Context, string) (*Message, error) {
|
||||||
|
return nil, errors.New("unused")
|
||||||
|
}
|
||||||
|
func (f *fakeGmailAPI) DownloadAttachment(context.Context, string, string) ([]byte, error) {
|
||||||
|
return nil, errors.New("unused")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestGmailSourcePassesQueryToListIDs(t *testing.T) {
|
||||||
|
fake := &fakeGmailAPI{ids: []string{"m1"}}
|
||||||
|
src := &gmailSource{c: fake, query: "from:alice@example.com"}
|
||||||
|
ids, _, err := src.ListIDs(context.Background(), 10, "")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if fake.lastQ != "from:alice@example.com" {
|
||||||
|
t.Fatalf("ListIDs q=%q, want from:alice@example.com", fake.lastQ)
|
||||||
|
}
|
||||||
|
if fake.lastLimit != 10 {
|
||||||
|
t.Fatalf("ListIDs limit=%d, want 10", fake.lastLimit)
|
||||||
|
}
|
||||||
|
if len(ids) != 1 || ids[0] != "m1" {
|
||||||
|
t.Fatalf("ids=%v", ids)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestGmailSourceEmptyQueryDefaultsToInbox(t *testing.T) {
|
||||||
|
fake := &fakeGmailAPI{}
|
||||||
|
src := &gmailSource{c: fake, query: ""}
|
||||||
|
if _, _, err := src.ListIDs(context.Background(), 5, ""); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if fake.lastQ != "in:inbox" {
|
||||||
|
t.Fatalf("empty query q=%q, want in:inbox", fake.lastQ)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseCLIGmailQuery(t *testing.T) {
|
||||||
|
cli, code, err := ParseCLI([]string{
|
||||||
|
"--source", "gmail",
|
||||||
|
"--query", "from:letrado@example.com",
|
||||||
|
"--out", t.TempDir(),
|
||||||
|
"--dry-run",
|
||||||
|
})
|
||||||
|
if err != nil || code != 0 {
|
||||||
|
t.Fatalf("ParseCLI: code=%d err=%v", code, err)
|
||||||
|
}
|
||||||
|
if cli.Sync.Query != "from:letrado@example.com" {
|
||||||
|
t.Fatalf("query=%q", cli.Sync.Query)
|
||||||
|
}
|
||||||
|
if cli.Sync.Gmail == nil {
|
||||||
|
t.Fatal("gmail source not configured")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func b64(s string) string {
|
func b64(s string) string {
|
||||||
return base64.URLEncoding.EncodeToString([]byte(s))
|
return base64.URLEncoding.EncodeToString([]byte(s))
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,3 @@
|
|||||||
|
// Commands in this directory are shebang mains (import.go), tagged so
|
||||||
|
// `go build ./bin/markdown` does not see two mains.
|
||||||
|
package main
|
||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=markdown_import "$0" "$@"; exit
|
||||||
|
//go:build markdown_import
|
||||||
|
//
|
||||||
|
// bin/markdown/import.go - split markdown into leafs (mistune).
|
||||||
|
//
|
||||||
|
// ./bin/markdown/import.go [dir]
|
||||||
|
// ./bin/markdown/import.go --files a.md,b.md --json
|
||||||
|
//
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(cmdbin.ExecFile("bin/md/import", os.Args[1:]))
|
||||||
|
}
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
// Commands in this directory are shebang mains (query.go).
|
||||||
|
package main
|
||||||
Executable
+20
@@ -0,0 +1,20 @@
|
|||||||
|
//usr/bin/env go run -tags=postgres_query "$0" "$@"; exit
|
||||||
|
//go:build postgres_query
|
||||||
|
//
|
||||||
|
// bin/postgres/query.go - read-only Postgres as YAML.
|
||||||
|
//
|
||||||
|
// ./bin/postgres/query.go --profile onlyoffice -c 'SELECT 1'
|
||||||
|
//
|
||||||
|
// Profiles: $HOME/.config/brain/db-profiles.yml (credentials stay out of git).
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(cmdbin.ExecFile("bin/db/psql-yq", os.Args[1:]))
|
||||||
|
}
|
||||||
+8
-13
@@ -1,27 +1,22 @@
|
|||||||
//usr/bin/env go run "$0" "$@"; exit
|
//usr/bin/env go run -tags=brain_serve "$0" "$@"; exit
|
||||||
// bin/serve.go - async Go HTTP server for the 2dph brain (see bin/server).
|
//go:build brain_serve
|
||||||
//
|
//
|
||||||
// KB_ROOT=/path/to/2dph ./bin/serve.go # serve the brain
|
// bin/serve.go — deprecated; use bin/brain/serve.go.
|
||||||
// KB_SEARCH_CMD=... KB_WORKERS=4 KB_PORT=8630 ./bin/serve.go
|
|
||||||
//
|
|
||||||
// Shebang trick: the first line is a Go `//` comment; when executed, env runs
|
|
||||||
// `go run "$0"` so this file doubles as an executable script. The real code
|
|
||||||
// lives in the importable package (module path, never a relative import).
|
|
||||||
// NOTE: never run `gofmt -w` on this file - it rewrites `//usr/bin/env` to
|
|
||||||
// `// usr/...` and breaks the shebang.
|
|
||||||
package main
|
package main
|
||||||
|
|
||||||
import (
|
import (
|
||||||
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
|
|
||||||
"github.com/eSlider/2dph/bin/server"
|
"github.com/eSlider/2dph/internal/httpapi"
|
||||||
)
|
)
|
||||||
|
|
||||||
func main() {
|
func main() {
|
||||||
if env := os.Getenv("KB_ROOT"); env == "" {
|
fmt.Fprintln(os.Stderr, "bin/serve.go is deprecated; use bin/brain/serve.go")
|
||||||
|
if os.Getenv("KB_ROOT") == "" {
|
||||||
if wd, err := os.Getwd(); err == nil {
|
if wd, err := os.Getwd(); err == nil {
|
||||||
os.Setenv("KB_ROOT", wd)
|
os.Setenv("KB_ROOT", wd)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
server.Run()
|
httpapi.Run(nil)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,154 +0,0 @@
|
|||||||
// Package server serves the 2dph brain over HTTP.
|
|
||||||
//
|
|
||||||
// Async by design: every request runs on its own goroutine, and CPU-heavy
|
|
||||||
// searches are serialized through a bounded worker pool (a counting
|
|
||||||
// semaphore) so N requests can't spawn N Python interpreters at once.
|
|
||||||
//
|
|
||||||
// Used by bin/serve.go which is a self-executing shebang script:
|
|
||||||
//
|
|
||||||
// ///usr/bin/env go run "$0" "$@"; exit
|
|
||||||
// package main
|
|
||||||
// import "github.com/eSlider/2dph/bin/server"
|
|
||||||
// func main() { server.Run() }
|
|
||||||
package server
|
|
||||||
|
|
||||||
import (
|
|
||||||
"context"
|
|
||||||
"encoding/json"
|
|
||||||
"errors"
|
|
||||||
"log"
|
|
||||||
"net/http"
|
|
||||||
"os"
|
|
||||||
"os/exec"
|
|
||||||
"path/filepath"
|
|
||||||
"strconv"
|
|
||||||
"strings"
|
|
||||||
"time"
|
|
||||||
)
|
|
||||||
|
|
||||||
type Searcher interface {
|
|
||||||
Search(ctx context.Context, query string, limit int) ([]byte, error)
|
|
||||||
}
|
|
||||||
|
|
||||||
type Server struct {
|
|
||||||
searcher Searcher
|
|
||||||
semaphore chan struct{}
|
|
||||||
}
|
|
||||||
|
|
||||||
const defaultPort = 8630
|
|
||||||
|
|
||||||
func NewServer(searcher Searcher, workers int) http.Handler {
|
|
||||||
return &Server{
|
|
||||||
searcher: searcher,
|
|
||||||
semaphore: make(chan struct{}, workers),
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
func (s *Server) ServeHTTP(w http.ResponseWriter, r *http.Request) {
|
|
||||||
switch {
|
|
||||||
case r.URL.Path == "/health":
|
|
||||||
writeJSON(w, http.StatusOK, map[string]any{"status": "ok"})
|
|
||||||
case r.URL.Path == "/search":
|
|
||||||
s.handleSearch(w, r)
|
|
||||||
default:
|
|
||||||
writeJSON(w, http.StatusNotFound, map[string]any{"error": "not found"})
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
func (s *Server) handleSearch(w http.ResponseWriter, r *http.Request) {
|
|
||||||
q := strings.TrimSpace(r.URL.Query().Get("q"))
|
|
||||||
if q == "" {
|
|
||||||
writeJSON(w, http.StatusBadRequest, map[string]any{"error": "q required"})
|
|
||||||
return
|
|
||||||
}
|
|
||||||
limit := 10
|
|
||||||
if raw := r.URL.Query().Get("n"); raw != "" {
|
|
||||||
n, err := strconv.Atoi(raw)
|
|
||||||
if err != nil || n < 1 || n > 100 {
|
|
||||||
writeJSON(w, http.StatusBadRequest, map[string]any{"error": "n must be int 1..100"})
|
|
||||||
return
|
|
||||||
}
|
|
||||||
limit = n
|
|
||||||
}
|
|
||||||
|
|
||||||
// Worker pool: block until a slot frees, so burst concurrency still
|
|
||||||
// bounds memory (no unbounded python processes).
|
|
||||||
select {
|
|
||||||
case s.semaphore <- struct{}{}:
|
|
||||||
defer func() { <-s.semaphore }()
|
|
||||||
case <-r.Context().Done():
|
|
||||||
return
|
|
||||||
}
|
|
||||||
|
|
||||||
body, err := s.searcher.Search(r.Context(), q, limit)
|
|
||||||
if err != nil {
|
|
||||||
writeJSON(w, http.StatusGatewayTimeout, map[string]any{"error": err.Error()})
|
|
||||||
return
|
|
||||||
}
|
|
||||||
writeRaw(w, http.StatusOK, body)
|
|
||||||
}
|
|
||||||
|
|
||||||
func writeJSON(w http.ResponseWriter, code int, obj any) {
|
|
||||||
body, _ := json.Marshal(obj)
|
|
||||||
writeRaw(w, code, body)
|
|
||||||
}
|
|
||||||
|
|
||||||
func writeRaw(w http.ResponseWriter, code int, body []byte) {
|
|
||||||
w.Header().Set("Content-Type", "application/json")
|
|
||||||
w.Header().Set("Content-Length", strconv.Itoa(len(body)))
|
|
||||||
w.WriteHeader(code)
|
|
||||||
w.Write(body)
|
|
||||||
}
|
|
||||||
|
|
||||||
// brainSearcher shells out to bin/kb/search --json. A single python search
|
|
||||||
// is bounded and short-lived; the worker pool keeps at most N live.
|
|
||||||
type brainSearcher struct {
|
|
||||||
cmdPath string
|
|
||||||
timeout time.Duration
|
|
||||||
}
|
|
||||||
|
|
||||||
func (b *brainSearcher) Search(ctx context.Context, query string, limit int) ([]byte, error) {
|
|
||||||
ctx, cancel := context.WithTimeout(ctx, b.timeout)
|
|
||||||
defer cancel()
|
|
||||||
cmd := exec.CommandContext(ctx, b.cmdPath, "--json", "-n", strconv.Itoa(limit), query)
|
|
||||||
out, err := cmd.Output()
|
|
||||||
if err != nil {
|
|
||||||
var exitErr *exec.ExitError
|
|
||||||
if errors.As(err, &exitErr) {
|
|
||||||
return nil, errors.New("search backend failed: " + strings.TrimSpace(string(exitErr.Stderr)))
|
|
||||||
}
|
|
||||||
return nil, err
|
|
||||||
}
|
|
||||||
return out, nil
|
|
||||||
}
|
|
||||||
|
|
||||||
// Run starts the HTTP server. Reads env: KB_SEARCH_CMD (default bin/kb/search,
|
|
||||||
// relative to the repo root given by KB_ROOT), KB_WORKERS (default 4), KB_PORT
|
|
||||||
// (default 8630).
|
|
||||||
func Run() {
|
|
||||||
root := os.Getenv("KB_ROOT")
|
|
||||||
searchPath := os.Getenv("KB_SEARCH_CMD")
|
|
||||||
if searchPath == "" {
|
|
||||||
searchPath = filepath.Join(root, "bin", "kb", "search")
|
|
||||||
}
|
|
||||||
workers := 4
|
|
||||||
if raw := os.Getenv("KB_WORKERS"); raw != "" {
|
|
||||||
if n, err := strconv.Atoi(raw); err == nil && n > 0 {
|
|
||||||
workers = n
|
|
||||||
}
|
|
||||||
}
|
|
||||||
port := defaultPort
|
|
||||||
if raw := os.Getenv("KB_PORT"); raw != "" {
|
|
||||||
if n, err := strconv.Atoi(raw); err == nil && n > 0 {
|
|
||||||
port = n
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
searcher := &brainSearcher{cmdPath: searchPath, timeout: 60 * time.Second}
|
|
||||||
handler := NewServer(searcher, workers)
|
|
||||||
addr := "127.0.0.1:" + strconv.Itoa(port)
|
|
||||||
log.Printf("serve: %s (workers=%d)", addr, workers)
|
|
||||||
if err := http.ListenAndServe(addr, handler); err != nil {
|
|
||||||
log.Fatal(err)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
+4
-61
@@ -1,21 +1,12 @@
|
|||||||
"""gitimport - parse `git log` output and turn commits into brain leafs.
|
"""gitimport - Ladybug graph writes for Commit/File/Person (no git binary).
|
||||||
|
|
||||||
Pure, testable functions. Field grammar (see bin/git/import):
|
Commit records come from bin/git/import.go (go-git). This module only MERGEs
|
||||||
|
the version graph File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person.
|
||||||
git log --no-merges --name-only \
|
|
||||||
--format='%x1e%H%x1f%an%x1f%ae%x1f%aI%x1f%s'
|
|
||||||
|
|
||||||
0x1e = record separator, 0x1f = field separator.
|
|
||||||
Files: newline-separated lines following each record's subject.
|
|
||||||
"""
|
"""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
from dataclasses import dataclass, field
|
from dataclasses import dataclass, field
|
||||||
|
|
||||||
REC_SEP = "\x1e"
|
|
||||||
FIELD_SEP = "\x1f"
|
|
||||||
|
|
||||||
|
|
||||||
@dataclass
|
@dataclass
|
||||||
class Commit:
|
class Commit:
|
||||||
@@ -26,54 +17,6 @@ class Commit:
|
|||||||
subject: str
|
subject: str
|
||||||
files: list[str] = field(default_factory=list)
|
files: list[str] = field(default_factory=list)
|
||||||
|
|
||||||
def leaf_text(self, repo: str) -> str:
|
|
||||||
head = f"commit {self.sha[:12]} in {repo} — {self.subject}"
|
|
||||||
body = [head, f"Author: {self.author} <{self.email}>", f"Date: {self.date}"]
|
|
||||||
if self.files:
|
|
||||||
body.append("Changing: " + ", ".join(self.files))
|
|
||||||
return "\n".join(body)
|
|
||||||
|
|
||||||
|
|
||||||
def parse_log(text: str) -> list[Commit]:
|
|
||||||
"""Parse `git log` output into Commit records.
|
|
||||||
|
|
||||||
Records are separated by 0x1e. A record is fields joined by 0x1f,
|
|
||||||
followed by optional newline-separated file paths inside the next
|
|
||||||
segment (git emits blank line + files after each record).
|
|
||||||
"""
|
|
||||||
commits: list[Commit] = []
|
|
||||||
# field records and file lists alternate; simpler: split on REC_SEP,
|
|
||||||
# each chunk = header line, possibly followed by newline + files.
|
|
||||||
for chunk in text.split(REC_SEP):
|
|
||||||
chunk = chunk.strip("\n")
|
|
||||||
if not chunk:
|
|
||||||
continue
|
|
||||||
lines = chunk.split("\n", 1)
|
|
||||||
header = lines[0].split(FIELD_SEP)
|
|
||||||
if len(header) < 5:
|
|
||||||
continue
|
|
||||||
sha, author, email, date, subject = header[:5]
|
|
||||||
files = [ln.strip() for ln in lines[1].splitlines() if ln.strip()] if len(lines) > 1 else []
|
|
||||||
commits.append(Commit(sha=sha, author=author, email=email,
|
|
||||||
date=date, subject=subject, files=files))
|
|
||||||
return commits
|
|
||||||
|
|
||||||
|
|
||||||
def commits_to_leafs(commits: list[Commit], repo: str) -> list[dict]:
|
|
||||||
"""Map commits to the leaf shape bin/kb/index expects (source/repo/...)."""
|
|
||||||
out: list[dict] = []
|
|
||||||
for c in commits:
|
|
||||||
out.append({
|
|
||||||
"source": f"{repo}@{c.sha}",
|
|
||||||
"repo": repo,
|
|
||||||
"heading": f"commit {c.sha[:12]} — {c.subject}",
|
|
||||||
"text": c.leaf_text(repo),
|
|
||||||
"type": "commit",
|
|
||||||
"status": "current",
|
|
||||||
"related": ",".join(c.files),
|
|
||||||
})
|
|
||||||
return out
|
|
||||||
|
|
||||||
|
|
||||||
GIT_SCHEMA = (
|
GIT_SCHEMA = (
|
||||||
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
|
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
|
||||||
@@ -114,4 +57,4 @@ def index_commits(conn, commits: list[Commit], repo: str) -> int:
|
|||||||
conn.execute("MATCH (f:File {id:$fid}), (c:Commit {id:$sha}) "
|
conn.execute("MATCH (f:File {id:$fid}), (c:Commit {id:$sha}) "
|
||||||
"MERGE (f)-[:HAS_VERSION]->(c)",
|
"MERGE (f)-[:HAS_VERSION]->(c)",
|
||||||
parameters={"fid": f"{repo}:{path}", "sha": c.sha})
|
parameters={"fid": f"{repo}:{path}", "sha": c.sha})
|
||||||
return len(commits)
|
return len(commits)
|
||||||
|
|||||||
+4
-4
@@ -135,7 +135,7 @@ def create_fts_and_vector(conn: ladybug.Connection, force: bool = False) -> None
|
|||||||
|
|
||||||
`force=True` is accepted for API compatibility but does **not** drop.
|
`force=True` is accepted for API compatibility but does **not** drop.
|
||||||
Fresh indexes require deleting `var/kb.lbug` and rebuilding
|
Fresh indexes require deleting `var/kb.lbug` and rebuilding
|
||||||
(`bin/kb/index --rebuild`).
|
(`bin/brain/index.go --rebuild`).
|
||||||
"""
|
"""
|
||||||
del force # API compat; DROP is unsafe — see docstring
|
del force # API compat; DROP is unsafe — see docstring
|
||||||
names = leaf_index_names(conn)
|
names = leaf_index_names(conn)
|
||||||
@@ -145,7 +145,7 @@ def create_fts_and_vector(conn: ladybug.Connection, force: bool = False) -> None
|
|||||||
except Exception as e:
|
except Exception as e:
|
||||||
raise RuntimeError(
|
raise RuntimeError(
|
||||||
"CREATE_FTS_INDEX failed (often ghost catalog after DROP INDEX). "
|
"CREATE_FTS_INDEX failed (often ghost catalog after DROP INDEX). "
|
||||||
"Delete var/kb.lbug and run bin/kb/index --rebuild. "
|
"Delete var/kb.lbug and run bin/brain/index.go --rebuild. "
|
||||||
f"Cause: {e}"
|
f"Cause: {e}"
|
||||||
) from e
|
) from e
|
||||||
if "Leaf_vec" not in names:
|
if "Leaf_vec" not in names:
|
||||||
@@ -158,7 +158,7 @@ def create_fts_and_vector(conn: ladybug.Connection, force: bool = False) -> None
|
|||||||
raise RuntimeError(
|
raise RuntimeError(
|
||||||
"CREATE_VECTOR_INDEX failed (often ghost catalog after DROP INDEX "
|
"CREATE_VECTOR_INDEX failed (often ghost catalog after DROP INDEX "
|
||||||
"Leaf.Leaf_vec → `_0_Leaf_vec_UPPER already exists in catalog`). "
|
"Leaf.Leaf_vec → `_0_Leaf_vec_UPPER already exists in catalog`). "
|
||||||
"Delete var/kb.lbug and run bin/kb/index --rebuild. "
|
"Delete var/kb.lbug and run bin/brain/index.go --rebuild. "
|
||||||
f"Cause: {e}"
|
f"Cause: {e}"
|
||||||
) from e
|
) from e
|
||||||
names = leaf_index_names(conn)
|
names = leaf_index_names(conn)
|
||||||
@@ -237,6 +237,6 @@ def stats(conn: ladybug.Connection) -> dict:
|
|||||||
|
|
||||||
def open_readonly() -> tuple[ladybug.Database, ladybug.Connection]:
|
def open_readonly() -> tuple[ladybug.Database, ladybug.Connection]:
|
||||||
if not DB_PATH.exists():
|
if not DB_PATH.exists():
|
||||||
raise FileNotFoundError(f"{DB_PATH} missing - run bin/kb/index first")
|
raise FileNotFoundError(f"{DB_PATH} missing - run bin/brain/index.go --rebuild first")
|
||||||
db, conn = connect(read_only=True)
|
db, conn = connect(read_only=True)
|
||||||
return db, conn
|
return db, conn
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
"""Mail markdown under var/mail → info leafs. Conversion stays off the brain DB."""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from mdleaves import read_markdown, to_all
|
||||||
|
|
||||||
|
|
||||||
|
def msg_date(md: Path) -> str:
|
||||||
|
j = md.parent / "message.json"
|
||||||
|
try:
|
||||||
|
d = json.loads(j.read_text(encoding="utf-8"))
|
||||||
|
return (d.get("receivedDate") or d.get("receivedAt") or "")[:10]
|
||||||
|
except (OSError, json.JSONDecodeError, TypeError):
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def from_mail_root(root: Path, limit: int = 0, since: str = "", repo: str = "ooMail") -> list[dict]:
|
||||||
|
if not root.is_dir():
|
||||||
|
return []
|
||||||
|
mds = sorted(root.rglob("message.md"))
|
||||||
|
if since:
|
||||||
|
mds = [m for m in mds if msg_date(m) >= since]
|
||||||
|
if limit:
|
||||||
|
mds = mds[:limit]
|
||||||
|
leafs: list[dict] = []
|
||||||
|
for md in mds:
|
||||||
|
files = [md] + sorted((md.parent / "attachments").glob("*.md"))
|
||||||
|
for f in files:
|
||||||
|
if not f.exists():
|
||||||
|
continue
|
||||||
|
for lf in to_all(read_markdown(f), f, repo=repo):
|
||||||
|
lf["source"] = f"ooMail:{md.parent.name}:{f.name}"
|
||||||
|
lf["how"] = "mail/import"
|
||||||
|
leafs.append(lf)
|
||||||
|
return leafs
|
||||||
@@ -0,0 +1,121 @@
|
|||||||
|
"""D14 layout: bin/{subject}/{method}.go, libs in internal/, one go.mod."""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
|
|
||||||
|
|
||||||
|
class BinLayoutTest(unittest.TestCase):
|
||||||
|
def test_brain_search_shebang_exists(self) -> None:
|
||||||
|
p = ROOT / "bin" / "brain" / "search.go"
|
||||||
|
self.assertTrue(p.is_file(), "missing bin/brain/search.go")
|
||||||
|
first = p.read_text().splitlines()[0]
|
||||||
|
self.assertTrue(
|
||||||
|
first.startswith("//usr/bin/env go run"),
|
||||||
|
f"shebang first line, got {first!r}",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_no_nested_go_mod_under_bin(self) -> None:
|
||||||
|
nested = list((ROOT / "bin").rglob("go.mod"))
|
||||||
|
self.assertEqual(nested, [], f"nested go.mod files: {nested}")
|
||||||
|
|
||||||
|
def test_rank_lives_in_internal_brain(self) -> None:
|
||||||
|
self.assertTrue(
|
||||||
|
(ROOT / "internal" / "brain" / "rank" / "rank.go").is_file(),
|
||||||
|
"ranking must live in internal/brain/rank (cgo-free)",
|
||||||
|
)
|
||||||
|
self.assertFalse(
|
||||||
|
(ROOT / "bin" / "kbsearch").exists(),
|
||||||
|
"bin/kbsearch nested module must be gone",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_no_main_go_under_bin_brain(self) -> None:
|
||||||
|
main = ROOT / "bin" / "brain" / "main.go"
|
||||||
|
self.assertFalse(main.exists(), "bin/brain/main.go is not a method")
|
||||||
|
|
||||||
|
def test_chats_methods_are_shebangs_not_main(self) -> None:
|
||||||
|
chats = ROOT / "bin" / "chats"
|
||||||
|
self.assertFalse(
|
||||||
|
(chats / "main.go").exists(),
|
||||||
|
"bin/chats/main.go is a dispatcher, not a method",
|
||||||
|
)
|
||||||
|
self.assertFalse(
|
||||||
|
(chats / "index_cmd.go").exists(),
|
||||||
|
"chats index is a brain write hiding under the wrong subject",
|
||||||
|
)
|
||||||
|
for method in ("sync.go", "import.go", "facts.go", "apply.go"):
|
||||||
|
p = chats / method
|
||||||
|
self.assertTrue(p.is_file(), f"missing bin/chats/{method}")
|
||||||
|
first = p.read_text().splitlines()[0]
|
||||||
|
self.assertTrue(
|
||||||
|
first.startswith("//usr/bin/env go run"),
|
||||||
|
f"{method} shebang, got {first!r}",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_chats_lib_lives_in_internal(self) -> None:
|
||||||
|
self.assertTrue(
|
||||||
|
(ROOT / "internal" / "chats" / "linkedin.go").is_file(),
|
||||||
|
"LinkedIn parser must live in internal/chats",
|
||||||
|
)
|
||||||
|
self.assertFalse(
|
||||||
|
(ROOT / "bin" / "chats" / "linkedin.go").exists(),
|
||||||
|
"parser must not stay under bin/chats as a second main",
|
||||||
|
)
|
||||||
|
|
||||||
|
def _assert_shebang(self, rel: str) -> None:
|
||||||
|
p = ROOT / rel
|
||||||
|
self.assertTrue(p.is_file(), f"missing {rel}")
|
||||||
|
first = p.read_text().splitlines()[0]
|
||||||
|
self.assertTrue(
|
||||||
|
first.startswith("//usr/bin/env go run"),
|
||||||
|
f"{rel} shebang, got {first!r}",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_brain_methods_are_shebangs(self) -> None:
|
||||||
|
for method in ("index.go", "get.go", "stats.go", "eval.go", "watch.go"):
|
||||||
|
self._assert_shebang(f"bin/brain/{method}")
|
||||||
|
|
||||||
|
def test_mail_import_is_shebang_not_brain_write(self) -> None:
|
||||||
|
self._assert_shebang("bin/mail/import.go")
|
||||||
|
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
|
||||||
|
self.assertIn(
|
||||||
|
"bin/brain/index.go",
|
||||||
|
index_mail,
|
||||||
|
"index_mail must point at bin/brain/index.go",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_markdown_import_is_shebang(self) -> None:
|
||||||
|
self._assert_shebang("bin/markdown/import.go")
|
||||||
|
|
||||||
|
def test_postgres_query_is_shebang(self) -> None:
|
||||||
|
self._assert_shebang("bin/postgres/query.go")
|
||||||
|
|
||||||
|
def test_git_import_is_gogit_shebang(self) -> None:
|
||||||
|
self._assert_shebang("bin/git/import.go")
|
||||||
|
py = (ROOT / "bin" / "git" / "import").read_text()
|
||||||
|
self.assertNotIn(
|
||||||
|
'["git"',
|
||||||
|
py,
|
||||||
|
"Python git/import must not subprocess the git binary",
|
||||||
|
)
|
||||||
|
self.assertIn("bin/git/import.go", py)
|
||||||
|
|
||||||
|
def test_web_search_is_shebang(self) -> None:
|
||||||
|
self._assert_shebang("bin/web/search.go")
|
||||||
|
py = (ROOT / "bin" / "web" / "search").read_text()
|
||||||
|
self.assertIn("bin/web/search.go", py)
|
||||||
|
|
||||||
|
def test_gitimport_py_has_no_git_binary(self) -> None:
|
||||||
|
py = (ROOT / "bin" / "tools" / "gitimport.py").read_text()
|
||||||
|
self.assertNotIn("subprocess", py)
|
||||||
|
self.assertNotIn("git log", py)
|
||||||
|
|
||||||
|
def test_gogit_is_direct_go_mod_require(self) -> None:
|
||||||
|
text = (ROOT / "go.mod").read_text()
|
||||||
|
first = text.split("require (")[1].split(")")[0]
|
||||||
|
self.assertRegex(first, r"github.com/go-git/go-git/v5\s+v")
|
||||||
|
for line in first.splitlines():
|
||||||
|
if "go-git/go-git" in line:
|
||||||
|
self.assertNotIn("indirect", line)
|
||||||
+14
-11
@@ -9,12 +9,6 @@ sys.path.insert(0, str(Path(__file__).resolve().parent))
|
|||||||
import kblib # noqa: E402
|
import kblib # noqa: E402
|
||||||
import gitimport # noqa: E402
|
import gitimport # noqa: E402
|
||||||
|
|
||||||
SAMPLE = (
|
|
||||||
"\x1e" + "a1b2c3d" + "\x1f" + "Ada Lovelace" + "\x1f" + "ada@example.com"
|
|
||||||
+ "\x1f" + "2026-08-10T12:00:00+01:00" + "\x1f" + "feat: first commit"
|
|
||||||
+ "\n\nREADME.md\nsrc/main.c\n"
|
|
||||||
)
|
|
||||||
|
|
||||||
COMMIT_PERSON_SCHEMA = (
|
COMMIT_PERSON_SCHEMA = (
|
||||||
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
|
"CREATE NODE TABLE IF NOT EXISTS Commit (id STRING, repo STRING, subject STRING, "
|
||||||
"author STRING, email STRING, date STRING, PRIMARY KEY(id))"
|
"author STRING, email STRING, date STRING, PRIMARY KEY(id))"
|
||||||
@@ -26,6 +20,17 @@ HAS_VERSION_SCHEMA = "CREATE REL TABLE IF NOT EXISTS HAS_VERSION (FROM File TO C
|
|||||||
AUTHORED_SCHEMA = "CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
|
AUTHORED_SCHEMA = "CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
|
||||||
|
|
||||||
|
|
||||||
|
def sample_commit() -> gitimport.Commit:
|
||||||
|
return gitimport.Commit(
|
||||||
|
sha="a1b2c3d",
|
||||||
|
author="Ada Lovelace",
|
||||||
|
email="ada@example.com",
|
||||||
|
date="2026-08-10T12:00:00+01:00",
|
||||||
|
subject="feat: first commit",
|
||||||
|
files=["README.md", "src/main.c"],
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
class GitGraphTest(unittest.TestCase):
|
class GitGraphTest(unittest.TestCase):
|
||||||
def setUp(self):
|
def setUp(self):
|
||||||
self.dir = tempfile.mkdtemp()
|
self.dir = tempfile.mkdtemp()
|
||||||
@@ -42,14 +47,12 @@ class GitGraphTest(unittest.TestCase):
|
|||||||
self.db.close()
|
self.db.close()
|
||||||
|
|
||||||
def test_index_commits_creates_nodes_and_edges(self):
|
def test_index_commits_creates_nodes_and_edges(self):
|
||||||
cs = gitimport.parse_log(SAMPLE)
|
gitimport.index_commits(self.conn, [sample_commit()], "sample-repo")
|
||||||
gitimport.index_commits(self.conn, cs, "sample-repo")
|
|
||||||
rp = self.conn.execute("MATCH (p:Person) RETURN p.name, p.email").get_all()
|
rp = self.conn.execute("MATCH (p:Person) RETURN p.name, p.email").get_all()
|
||||||
self.assertEqual([tuple(r) for r in rp], [("Ada Lovelace", "ada@example.com")])
|
self.assertEqual([tuple(r) for r in rp], [("Ada Lovelace", "ada@example.com")])
|
||||||
rc = self.conn.execute("MATCH (c:Commit) RETURN c.id, c.repo").get_all()
|
rc = self.conn.execute("MATCH (c:Commit) RETURN c.id, c.repo").get_all()
|
||||||
self.assertEqual(len(rc), 1)
|
self.assertEqual(len(rc), 1)
|
||||||
self.assertEqual(rc[0][1], "sample-repo")
|
self.assertEqual(rc[0][1], "sample-repo")
|
||||||
# File -[:HAS_VERSION]-> Commit -[:AUTHORED]-> Person
|
|
||||||
rf = self.conn.execute(
|
rf = self.conn.execute(
|
||||||
"MATCH (f:File)-[:HAS_VERSION]->(c:Commit)-[:AUTHORED]->(p:Person) "
|
"MATCH (f:File)-[:HAS_VERSION]->(c:Commit)-[:AUTHORED]->(p:Person) "
|
||||||
"RETURN f.path, c.id, p.email").get_all()
|
"RETURN f.path, c.id, p.email").get_all()
|
||||||
@@ -58,7 +61,7 @@ class GitGraphTest(unittest.TestCase):
|
|||||||
self.assertTrue(all(r[2] == "ada@example.com" for r in rf))
|
self.assertTrue(all(r[2] == "ada@example.com" for r in rf))
|
||||||
|
|
||||||
def test_index_commits_idempotent(self):
|
def test_index_commits_idempotent(self):
|
||||||
cs = gitimport.parse_log(SAMPLE)
|
cs = [sample_commit()]
|
||||||
gitimport.index_commits(self.conn, cs, "sample-repo")
|
gitimport.index_commits(self.conn, cs, "sample-repo")
|
||||||
gitimport.index_commits(self.conn, cs, "sample-repo")
|
gitimport.index_commits(self.conn, cs, "sample-repo")
|
||||||
n = self.conn.execute("MATCH (c:Commit) RETURN count(*)").get_all()[0][0]
|
n = self.conn.execute("MATCH (c:Commit) RETURN count(*)").get_all()[0][0]
|
||||||
@@ -68,4 +71,4 @@ class GitGraphTest(unittest.TestCase):
|
|||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
unittest.main()
|
unittest.main()
|
||||||
|
|||||||
@@ -1,57 +0,0 @@
|
|||||||
import sys
|
|
||||||
import unittest
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
|
||||||
|
|
||||||
import gitimport # noqa: E402
|
|
||||||
|
|
||||||
SAMPLE = (
|
|
||||||
"\x1e" + "a1b2c3d" + "\x1f" + "Ada Lovelace" + "\x1f" + "ada@example.com"
|
|
||||||
+ "\x1f" + "2026-08-10T12:00:00+01:00" + "\x1f" + "feat: first commit"
|
|
||||||
+ "\n\nREADME.md\nsrc/main.c\n"
|
|
||||||
+ "\x1e" + "e4f5a6b" + "\x1f" + "Bob Babbage" + "\x1f" + "bob@example.com"
|
|
||||||
+ "\x1f" + "2026-08-11T09:30:00+01:00" + "\x1f" + "fix: typo"
|
|
||||||
+ "\n\ndocs/notes.md"
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
class GitparseTest(unittest.TestCase):
|
|
||||||
def test_parses_records(self):
|
|
||||||
cs = gitimport.parse_log(SAMPLE)
|
|
||||||
self.assertEqual(len(cs), 2)
|
|
||||||
|
|
||||||
def test_parses_commit_fields(self):
|
|
||||||
cs = gitimport.parse_log(SAMPLE)
|
|
||||||
c = cs[0]
|
|
||||||
self.assertEqual(c.sha, "a1b2c3d")
|
|
||||||
self.assertEqual(c.author, "Ada Lovelace")
|
|
||||||
self.assertEqual(c.email, "ada@example.com")
|
|
||||||
self.assertEqual(c.date, "2026-08-10T12:00:00+01:00")
|
|
||||||
self.assertEqual(c.subject, "feat: first commit")
|
|
||||||
|
|
||||||
def test_parses_changed_files(self):
|
|
||||||
cs = gitimport.parse_log(SAMPLE)
|
|
||||||
self.assertEqual(cs[0].files, ["README.md", "src/main.c"])
|
|
||||||
self.assertEqual(cs[1].files, ["docs/notes.md"])
|
|
||||||
|
|
||||||
def test_ignores_empty(self):
|
|
||||||
self.assertEqual(gitimport.parse_log(""), [])
|
|
||||||
|
|
||||||
def test_skip_malformed_record(self):
|
|
||||||
self.assertEqual(gitimport.parse_log("\x1eweird\x1e"), [])
|
|
||||||
|
|
||||||
def test_commit_leaf_shape(self):
|
|
||||||
leafs = gitimport.commits_to_leafs(gitimport.parse_log(SAMPLE), "sample-repo")
|
|
||||||
self.assertEqual(len(leafs), 2)
|
|
||||||
lf = leafs[0]
|
|
||||||
self.assertEqual(lf["type"], "commit")
|
|
||||||
self.assertEqual(lf["repo"], "sample-repo")
|
|
||||||
self.assertEqual(lf["source"], "sample-repo@a1b2c3d")
|
|
||||||
self.assertIn("Ada Lovelace", lf["text"])
|
|
||||||
self.assertIn("README.md", lf["related"])
|
|
||||||
self.assertIn("feat: first commit", lf["heading"])
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
unittest.main()
|
|
||||||
@@ -0,0 +1,46 @@
|
|||||||
|
"""Mail markdown → leafs (no Ladybug). Brain index --with-mail uses this."""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||||
|
|
||||||
|
import mailleafs # noqa: E402
|
||||||
|
|
||||||
|
|
||||||
|
class MailLeafsTest(unittest.TestCase):
|
||||||
|
def test_message_md_becomes_info_leaf(self) -> None:
|
||||||
|
root = Path(tempfile.mkdtemp())
|
||||||
|
msg = root / "inbox" / "alice-1"
|
||||||
|
msg.mkdir(parents=True)
|
||||||
|
(msg / "message.json").write_text(
|
||||||
|
json.dumps({"receivedDate": "2026-01-15T10:00:00Z", "subject": "Hello"}),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
(msg / "message.md").write_text(
|
||||||
|
"---\nroot: info\n---\n\n# Hello\n\nFrom Alice to Bob.\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
leafs = mailleafs.from_mail_root(root)
|
||||||
|
self.assertEqual(len(leafs), 1)
|
||||||
|
self.assertIn("Alice", leafs[0]["text"])
|
||||||
|
self.assertTrue(leafs[0]["source"].startswith("ooMail:"))
|
||||||
|
self.assertEqual(leafs[0]["how"], "mail/import")
|
||||||
|
|
||||||
|
def test_since_filters_by_message_json_date(self) -> None:
|
||||||
|
root = Path(tempfile.mkdtemp())
|
||||||
|
for name, day in (("old", "2025-01-01"), ("new", "2026-06-01")):
|
||||||
|
d = root / "inbox" / name
|
||||||
|
d.mkdir(parents=True)
|
||||||
|
(d / "message.json").write_text(
|
||||||
|
json.dumps({"receivedDate": f"{day}T00:00:00Z"}),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
(d / "message.md").write_text(f"# {name}\n\nbody\n", encoding="utf-8")
|
||||||
|
leafs = mailleafs.from_mail_root(root, since="2026-01-01")
|
||||||
|
self.assertEqual(len(leafs), 1)
|
||||||
|
self.assertIn("new", leafs[0]["text"])
|
||||||
@@ -0,0 +1,84 @@
|
|||||||
|
"""Published docs must match live commands (Gitea SoT, brain/search, no fake --hop)."""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import re
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
|
|
||||||
|
|
||||||
|
class PublishedDocsTest(unittest.TestCase):
|
||||||
|
def test_readme_points_issues_at_gitea(self) -> None:
|
||||||
|
text = (ROOT / "README.md").read_text()
|
||||||
|
self.assertIn(
|
||||||
|
"https://git.produktor.io/eSlider/2dph/issues",
|
||||||
|
text,
|
||||||
|
"README must point issues at Gitea",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_plan_d15_names_gitea_origin(self) -> None:
|
||||||
|
text = (ROOT / "PLAN.md").read_text()
|
||||||
|
self.assertIn("D15", text)
|
||||||
|
self.assertIn("git.produktor.io/eSlider/2dph", text)
|
||||||
|
|
||||||
|
def test_readme_primary_search_is_brain(self) -> None:
|
||||||
|
text = (ROOT / "README.md").read_text()
|
||||||
|
self.assertIn(
|
||||||
|
"bin/brain/search.go",
|
||||||
|
text,
|
||||||
|
"README deduction search must name bin/brain/search.go",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_readme_index_is_brain_not_index_mail(self) -> None:
|
||||||
|
text = (ROOT / "README.md").read_text()
|
||||||
|
self.assertIn("bin/brain/index.go", text)
|
||||||
|
self.assertNotIn(
|
||||||
|
"bin/mail/index_mail",
|
||||||
|
text,
|
||||||
|
"mail index is a brain write; README must name bin/brain/index.go",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_readme_git_import_is_gogit(self) -> None:
|
||||||
|
text = (ROOT / "README.md").read_text()
|
||||||
|
self.assertIn("bin/git/import.go", text)
|
||||||
|
self.assertIn("go-git", text)
|
||||||
|
self.assertIn("D19", (ROOT / "PLAN.md").read_text())
|
||||||
|
|
||||||
|
def test_web_search_is_go_not_ops_host(self) -> None:
|
||||||
|
readme = (ROOT / "README.md").read_text()
|
||||||
|
self.assertIn("bin/web/search.go", readme)
|
||||||
|
skill = (ROOT / "skills" / "web-search" / "SKILL.md").read_text()
|
||||||
|
self.assertIn("bin/web/search.go", skill)
|
||||||
|
self.assertNotIn("search.ops.io", skill)
|
||||||
|
self.assertNotIn("search.ops.io", readme)
|
||||||
|
compose = (ROOT / "compose.yaml").read_text()
|
||||||
|
self.assertIn("searxng", compose)
|
||||||
|
self.assertNotIn("search.ops.io", compose)
|
||||||
|
settings = (ROOT / "deploy" / "searxng" / "settings.yml").read_text()
|
||||||
|
self.assertNotIn("password", settings.lower())
|
||||||
|
self.assertIn("json", settings)
|
||||||
|
|
||||||
|
def test_readme_search_escalates_web(self) -> None:
|
||||||
|
text = (ROOT / "README.md").read_text()
|
||||||
|
self.assertIn("--no-web", text)
|
||||||
|
self.assertIn("D17", (ROOT / "PLAN.md").read_text())
|
||||||
|
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
|
||||||
|
self.assertIn("`web` block", skill)
|
||||||
|
|
||||||
|
def test_docs_do_not_claim_hop_walks(self) -> None:
|
||||||
|
paths = [
|
||||||
|
ROOT / "README.md",
|
||||||
|
ROOT / "docs" / "design.md",
|
||||||
|
ROOT / "skills" / "brain" / "SKILL.md",
|
||||||
|
ROOT / "skills" / "diataxis-docs" / "SKILL.md",
|
||||||
|
]
|
||||||
|
# Command-style `--hop 1` / `--hop N` plus follow/walk = the old lie.
|
||||||
|
# Honest "not implemented" notes must not match.
|
||||||
|
lie = re.compile(r"--hop (?:N|1).*(?:follow|walk)", re.I | re.S)
|
||||||
|
for path in paths:
|
||||||
|
text = path.read_text()
|
||||||
|
self.assertIsNone(
|
||||||
|
lie.search(text),
|
||||||
|
f"{path.relative_to(ROOT)} still claims --hop walks the graph",
|
||||||
|
)
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
"""Every bin/ path named in skills/ must exist on disk."""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import re
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
|
BIN_PATH = re.compile(r"\b(bin/[A-Za-z0-9_./-]+)")
|
||||||
|
|
||||||
|
|
||||||
|
class SkillsBinPathsTest(unittest.TestCase):
|
||||||
|
def test_agent_cost_skill_is_gone(self) -> None:
|
||||||
|
self.assertFalse(
|
||||||
|
(ROOT / "skills" / "agent-cost").exists(),
|
||||||
|
"skills/agent-cost documents bin/agents/cost which does not exist",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_brain_skill_replaces_kb_search(self) -> None:
|
||||||
|
self.assertTrue((ROOT / "skills" / "brain" / "SKILL.md").is_file())
|
||||||
|
self.assertFalse((ROOT / "skills" / "kb-search").exists())
|
||||||
|
|
||||||
|
def test_skill_bin_paths_exist(self) -> None:
|
||||||
|
missing: list[str] = []
|
||||||
|
for skill in sorted((ROOT / "skills").rglob("SKILL.md")):
|
||||||
|
text = skill.read_text()
|
||||||
|
for match in BIN_PATH.findall(text):
|
||||||
|
rel = match.rstrip("`'.,")
|
||||||
|
if rel.endswith(".go") or Path(rel).suffix == "" or Path(rel).suffix in {".go", ".py"}:
|
||||||
|
p = ROOT / rel
|
||||||
|
if not p.exists():
|
||||||
|
missing.append(f"{skill.relative_to(ROOT)}: {rel}")
|
||||||
|
self.assertEqual(missing, [], "SKILL.md names bin/ paths that do not exist")
|
||||||
+4
-4
@@ -1,4 +1,4 @@
|
|||||||
// Package watch polls corpus directories for changes and re-runs bin/kb/index.
|
// Package watch polls corpus directories for changes and re-runs brain/index.
|
||||||
//
|
//
|
||||||
// Port of the former bin/kb-watch bash script to an importable, testable Go
|
// Port of the former bin/kb-watch bash script to an importable, testable Go
|
||||||
// package. Polls file mtimes (no inotify deps); cheap and reliable.
|
// package. Polls file mtimes (no inotify deps); cheap and reliable.
|
||||||
@@ -18,8 +18,8 @@ import (
|
|||||||
type Options struct {
|
type Options struct {
|
||||||
Dirs []string
|
Dirs []string
|
||||||
Interval time.Duration
|
Interval time.Duration
|
||||||
// IndexCmd is the kb/index command template. %s is replaced by the repo
|
// IndexCmd is the index command template. %s is replaced by the repo
|
||||||
// root (from KB_ROOT). Defaults to `python3 <root>/bin/kb/index`.
|
// root (from KB_ROOT). Defaults to `python3 <root>/bin/kb/index --with-mail`.
|
||||||
IndexCmd string
|
IndexCmd string
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -67,7 +67,7 @@ func fromEnv(args []string) Options {
|
|||||||
if pys == "" {
|
if pys == "" {
|
||||||
pys = "python3"
|
pys = "python3"
|
||||||
}
|
}
|
||||||
opts.IndexCmd = pys + " <root>/bin/kb/index"
|
opts.IndexCmd = pys + " <root>/bin/kb/index --with-mail"
|
||||||
return opts
|
return opts
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -3,6 +3,7 @@ package watch
|
|||||||
import (
|
import (
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
"testing"
|
"testing"
|
||||||
"time"
|
"time"
|
||||||
)
|
)
|
||||||
@@ -43,7 +44,10 @@ func TestFromEnvDefaults(t *testing.T) {
|
|||||||
if opts.Interval != 30*time.Second {
|
if opts.Interval != 30*time.Second {
|
||||||
t.Fatalf("default interval = %s, want 30s", opts.Interval)
|
t.Fatalf("default interval = %s, want 30s", opts.Interval)
|
||||||
}
|
}
|
||||||
if opts.IndexCmd == "" {
|
if !strings.Contains(opts.IndexCmd, "kb/index") {
|
||||||
t.Fatal("default index cmd is empty")
|
t.Fatalf("default index cmd = %q, want kb/index", opts.IndexCmd)
|
||||||
|
}
|
||||||
|
if !strings.Contains(opts.IndexCmd, "--with-mail") {
|
||||||
|
t.Fatalf("default index cmd must include --with-mail, got %q", opts.IndexCmd)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
+12
-136
@@ -1,150 +1,26 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""web/search - web search through the self-hosted SearXNG at search.ops.io.
|
"""web/search — deprecated. Use bin/web/search.go (SearXNG, no Python client).
|
||||||
|
|
||||||
bin/web/search "LadybugDB vector search"
|
bin/web/search.go QUERY [--json] [-n N] [--site HOST]
|
||||||
bin/web/search "model2vec multilingual" --site github.com
|
|
||||||
bin/web/search "uclancy" --category it -n 3 --json | jq -r '.results[].url'
|
|
||||||
bin/web/search "sqlite-vec" --refresh # ignore the cached answer
|
|
||||||
|
|
||||||
This complements bin/kb/search: the knowledge base holds our own facts, this
|
|
||||||
reaches the public web. Use it as the second, independent source that the
|
|
||||||
detective method asks for.
|
|
||||||
|
|
||||||
Exit codes: 0 results, 2 refused as possible PII, 3 throttled (not "nothing
|
|
||||||
found" - the instance answers 200 with an empty list when it throttles).
|
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import argparse
|
|
||||||
import fcntl
|
|
||||||
import json
|
|
||||||
import os
|
import os
|
||||||
import sys
|
import sys
|
||||||
import time
|
|
||||||
import urllib.parse
|
|
||||||
import urllib.request
|
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
TOOLS = Path(__file__).resolve().parents[1] / "tools"
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
sys.path.insert(0, str(TOOLS))
|
|
||||||
sys.path.insert(0, str(TOOLS / "web-search"))
|
|
||||||
|
|
||||||
import websearch as ws # noqa: E402
|
|
||||||
from yamlout import to_yaml # noqa: E402
|
|
||||||
|
|
||||||
CONFIG = Path(os.environ.get("BRAIN_SEARCH_ENV", Path.home() / ".config/brain/search.env"))
|
|
||||||
CACHE = Path(os.environ.get("BRAIN_SEARCH_CACHE", Path.home() / ".cache/brain/web-search.sqlite"))
|
|
||||||
LOCK = CACHE.with_suffix(".lock")
|
|
||||||
|
|
||||||
|
|
||||||
def load_config() -> dict:
|
def main(argv: list[str]) -> int:
|
||||||
if not CONFIG.exists():
|
print(
|
||||||
sys.exit(f"no credentials at {CONFIG} (mode 600, BRAIN_SEARCH_URL/USER/PASS)")
|
"bin/web/search is deprecated; use bin/web/search.go",
|
||||||
conf = {}
|
file=sys.stderr,
|
||||||
for line in CONFIG.read_text().splitlines():
|
)
|
||||||
line = line.strip()
|
target = ROOT / "bin" / "web" / "search.go"
|
||||||
if not line or line.startswith("#") or "=" not in line:
|
os.execvp("go", ["go", "run", str(target), *argv])
|
||||||
continue
|
return 1
|
||||||
key, _, value = line.partition("=")
|
|
||||||
conf[key.strip()] = value.strip().strip("\"'")
|
|
||||||
missing = {"BRAIN_SEARCH_URL", "BRAIN_SEARCH_USER", "BRAIN_SEARCH_PASS"} - conf.keys()
|
|
||||||
if missing:
|
|
||||||
sys.exit(f"{CONFIG} is missing {', '.join(sorted(missing))}")
|
|
||||||
return conf
|
|
||||||
|
|
||||||
|
|
||||||
def fetch(conf: dict, query: str, params: dict, timeout: int) -> dict:
|
|
||||||
args = {"q": query, "format": "json", **params}
|
|
||||||
url = f"{conf['BRAIN_SEARCH_URL'].rstrip('/')}/search?{urllib.parse.urlencode(args)}"
|
|
||||||
request = urllib.request.Request(url)
|
|
||||||
token = f"{conf['BRAIN_SEARCH_USER']}:{conf['BRAIN_SEARCH_PASS']}".encode()
|
|
||||||
import base64
|
|
||||||
request.add_header("Authorization", "Basic " + base64.b64encode(token).decode())
|
|
||||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
|
||||||
return json.loads(response.read().decode())
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> int:
|
|
||||||
parser = argparse.ArgumentParser(description="web search via SearXNG")
|
|
||||||
parser.add_argument("query")
|
|
||||||
parser.add_argument("-n", "--limit", type=int, default=ws.DEFAULT_LIMIT)
|
|
||||||
parser.add_argument("--site", help="restrict to one domain")
|
|
||||||
parser.add_argument("--lang", help="language code, e.g. de")
|
|
||||||
parser.add_argument("--fresh", choices=["day", "week", "month", "year"],
|
|
||||||
help="time range")
|
|
||||||
parser.add_argument("--category", help="SearXNG category, e.g. it, science, news")
|
|
||||||
parser.add_argument("--engines", help="comma separated engine list")
|
|
||||||
parser.add_argument("--json", action="store_true")
|
|
||||||
parser.add_argument("--refresh", action="store_true", help="bypass the cache")
|
|
||||||
parser.add_argument("--ttl", type=float, default=ws.CACHE_TTL)
|
|
||||||
parser.add_argument("--timeout", type=int, default=25)
|
|
||||||
parser.add_argument("--force", action="store_true",
|
|
||||||
help="send even if the query looks like PII")
|
|
||||||
args = parser.parse_args()
|
|
||||||
|
|
||||||
query = f"site:{args.site} {args.query}" if args.site else args.query
|
|
||||||
|
|
||||||
reason = ws.phi_reason(query)
|
|
||||||
if reason and not args.force:
|
|
||||||
print(f"refused: {reason}. This query would leave the host.", file=sys.stderr)
|
|
||||||
print("Rephrase without identifiers, or pass --force if it is genuinely public.",
|
|
||||||
file=sys.stderr)
|
|
||||||
return 2
|
|
||||||
|
|
||||||
params = {}
|
|
||||||
if args.lang:
|
|
||||||
params["language"] = args.lang
|
|
||||||
if args.fresh:
|
|
||||||
params["time_range"] = args.fresh
|
|
||||||
if args.category:
|
|
||||||
params["categories"] = args.category
|
|
||||||
if args.engines:
|
|
||||||
params["engines"] = args.engines
|
|
||||||
|
|
||||||
key = ws.cache_key(query, params)
|
|
||||||
conn = ws.open_cache(CACHE)
|
|
||||||
|
|
||||||
if not args.refresh:
|
|
||||||
cached = ws.cache_get(conn, key, ttl=args.ttl)
|
|
||||||
if cached is not None:
|
|
||||||
out = ws.project(cached, limit=args.limit)
|
|
||||||
out["cached"] = True
|
|
||||||
sys.stdout.write(json.dumps(out, indent=2, ensure_ascii=False) + "\n"
|
|
||||||
if args.json else to_yaml(out))
|
|
||||||
return 0
|
|
||||||
|
|
||||||
conf = load_config()
|
|
||||||
LOCK.parent.mkdir(parents=True, exist_ok=True)
|
|
||||||
|
|
||||||
# One request at a time across every agent on this host: the instance
|
|
||||||
# suspends engines for minutes when several of us ask at once.
|
|
||||||
with open(LOCK, "w") as lock:
|
|
||||||
fcntl.flock(lock, fcntl.LOCK_EX)
|
|
||||||
|
|
||||||
payload = None
|
|
||||||
for attempt in range(1 + len(ws.RETRY_BACKOFF)):
|
|
||||||
delay = ws.wait_for(ws.last_call(conn), time.time())
|
|
||||||
if delay:
|
|
||||||
time.sleep(delay)
|
|
||||||
ws.mark_call(conn)
|
|
||||||
try:
|
|
||||||
payload = fetch(conf, query, params, args.timeout)
|
|
||||||
except Exception as error: # noqa: BLE001 - report, do not crash
|
|
||||||
print(f"request failed: {error}", file=sys.stderr)
|
|
||||||
return 3
|
|
||||||
if ws.classify(payload) == "ok":
|
|
||||||
break
|
|
||||||
if attempt < len(ws.RETRY_BACKOFF):
|
|
||||||
time.sleep(ws.RETRY_BACKOFF[attempt])
|
|
||||||
|
|
||||||
if ws.classify(payload) == "ok":
|
|
||||||
ws.cache_put(conn, key, payload)
|
|
||||||
|
|
||||||
out = ws.project(payload, limit=args.limit)
|
|
||||||
sys.stdout.write(json.dumps(out, indent=2, ensure_ascii=False) + "\n"
|
|
||||||
if args.json else to_yaml(out))
|
|
||||||
return 0 if out["status"] == "ok" else 3
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
sys.exit(main())
|
sys.exit(main(sys.argv[1:]))
|
||||||
|
|||||||
Executable
+232
@@ -0,0 +1,232 @@
|
|||||||
|
//usr/bin/env go run "$0" "$@"; exit
|
||||||
|
//
|
||||||
|
// bin/web/search.go - SearXNG as the second independent source (D3).
|
||||||
|
//
|
||||||
|
// ./bin/web/search.go "LadybugDB vector search"
|
||||||
|
// ./bin/web/search.go "model2vec" --category it --json
|
||||||
|
// ./bin/web/search.go "postgres" --site github.com --fresh year
|
||||||
|
//
|
||||||
|
// Empty results mean throttled, not "nothing exists". Exit 2 = PII refuse, 3 = throttled.
|
||||||
|
// Config: $BRAIN_SEARCH_ENV (default $HOME/.config/brain/search.env).
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"net/http"
|
||||||
|
"os"
|
||||||
|
"strconv"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/websearch"
|
||||||
|
"golang.org/x/sys/unix"
|
||||||
|
)
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(run(os.Args[1:]))
|
||||||
|
}
|
||||||
|
|
||||||
|
func run(args []string) int {
|
||||||
|
var (
|
||||||
|
query, site, lang, fresh, category, engines string
|
||||||
|
limit = websearch.DefaultLimit
|
||||||
|
jsonOut, refresh, force bool
|
||||||
|
ttl = float64(websearch.CacheTTL)
|
||||||
|
timeout = 25
|
||||||
|
)
|
||||||
|
i := 0
|
||||||
|
for i < len(args) {
|
||||||
|
a := args[i]
|
||||||
|
switch {
|
||||||
|
case a == "--json":
|
||||||
|
jsonOut = true
|
||||||
|
case a == "--refresh":
|
||||||
|
refresh = true
|
||||||
|
case a == "--force":
|
||||||
|
force = true
|
||||||
|
case (a == "-n" || a == "--limit") && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
n, err := strconv.Atoi(args[i])
|
||||||
|
if err != nil || n < 0 {
|
||||||
|
fmt.Fprintln(os.Stderr, "web/search: --limit must be a non-negative integer")
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
limit = n
|
||||||
|
case a == "--site" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
site = args[i]
|
||||||
|
case a == "--lang" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
lang = args[i]
|
||||||
|
case a == "--fresh" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
fresh = args[i]
|
||||||
|
case a == "--category" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
category = args[i]
|
||||||
|
case a == "--engines" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
engines = args[i]
|
||||||
|
case a == "--ttl" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
v, err := strconv.ParseFloat(args[i], 64)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintln(os.Stderr, "web/search: --ttl must be a number")
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
ttl = v
|
||||||
|
case a == "--timeout" && i+1 < len(args):
|
||||||
|
i++
|
||||||
|
n, err := strconv.Atoi(args[i])
|
||||||
|
if err != nil || n <= 0 {
|
||||||
|
fmt.Fprintln(os.Stderr, "web/search: --timeout must be a positive integer")
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
timeout = n
|
||||||
|
case a == "-h" || a == "--help":
|
||||||
|
fmt.Fprintln(os.Stderr, `usage: bin/web/search.go QUERY [--json] [-n N] [--site HOST] [--lang LANG] [--fresh day|week|month|year] [--category CAT] [--engines LIST] [--refresh] [--force]`)
|
||||||
|
return 0
|
||||||
|
case len(a) > 0 && a[0] != '-' && query == "":
|
||||||
|
query = a
|
||||||
|
default:
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: unknown flag %s\n", a)
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
i++
|
||||||
|
}
|
||||||
|
if query == "" {
|
||||||
|
fmt.Fprintln(os.Stderr, "web/search: query required")
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
if site != "" {
|
||||||
|
query = "site:" + site + " " + query
|
||||||
|
}
|
||||||
|
if reason := websearch.PHIReason(query); reason != "" && !force {
|
||||||
|
fmt.Fprintf(os.Stderr, "refused: %s. This query would leave the host.\n", reason)
|
||||||
|
fmt.Fprintln(os.Stderr, "Rephrase without identifiers, or pass --force if it is genuinely public.")
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
|
||||||
|
params := map[string]string{}
|
||||||
|
if lang != "" {
|
||||||
|
params["language"] = lang
|
||||||
|
}
|
||||||
|
if fresh != "" {
|
||||||
|
params["time_range"] = fresh
|
||||||
|
}
|
||||||
|
if category != "" {
|
||||||
|
params["categories"] = category
|
||||||
|
}
|
||||||
|
if engines != "" {
|
||||||
|
params["engines"] = engines
|
||||||
|
}
|
||||||
|
|
||||||
|
cachePath := os.Getenv("BRAIN_SEARCH_CACHE")
|
||||||
|
if cachePath == "" {
|
||||||
|
cachePath = os.Getenv("HOME") + "/.cache/brain/web-search.sqlite"
|
||||||
|
}
|
||||||
|
cache, err := websearch.OpenCache(cachePath)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
defer cache.Close()
|
||||||
|
|
||||||
|
key := websearch.CacheKey(query, params)
|
||||||
|
now := float64(time.Now().Unix())
|
||||||
|
if !refresh {
|
||||||
|
if cached, err := cache.Get(key, ttl, now); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||||
|
return 1
|
||||||
|
} else if cached != nil {
|
||||||
|
out := websearch.Project(*cached, limit, websearch.DefaultSnippetChars)
|
||||||
|
out.Cached = true
|
||||||
|
return writeOut(out, jsonOut)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
envPath := os.Getenv("BRAIN_SEARCH_ENV")
|
||||||
|
if envPath == "" {
|
||||||
|
envPath = os.Getenv("HOME") + "/.config/brain/search.env"
|
||||||
|
}
|
||||||
|
conf, err := websearch.LoadConfig(envPath)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
lockPath := cachePath + ".lock"
|
||||||
|
lock, err := os.OpenFile(lockPath, os.O_CREATE|os.O_RDWR, 0o600)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: lock: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
defer lock.Close()
|
||||||
|
if err := unix.Flock(int(lock.Fd()), unix.LOCK_EX); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: lock: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
defer unix.Flock(int(lock.Fd()), unix.LOCK_UN)
|
||||||
|
|
||||||
|
var payload websearch.Payload
|
||||||
|
attempts := 1 + len(websearch.RetryBackoff)
|
||||||
|
client := &http.Client{}
|
||||||
|
for attempt := 0; attempt < attempts; attempt++ {
|
||||||
|
last, err := cache.LastCall()
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
if delay := websearch.WaitFor(last, float64(time.Now().Unix()), websearch.MinInterval); delay > 0 {
|
||||||
|
time.Sleep(time.Duration(delay * float64(time.Second)))
|
||||||
|
}
|
||||||
|
if err := cache.MarkCall(float64(time.Now().Unix())); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
payload, err = websearch.Fetch(client, conf, query, params, time.Duration(timeout)*time.Second)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "request failed: %v\n", err)
|
||||||
|
return 3
|
||||||
|
}
|
||||||
|
if websearch.Classify(payload) == websearch.StatusOK {
|
||||||
|
break
|
||||||
|
}
|
||||||
|
if attempt < len(websearch.RetryBackoff) {
|
||||||
|
time.Sleep(time.Duration(websearch.RetryBackoff[attempt] * float64(time.Second)))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if websearch.Classify(payload) == websearch.StatusOK {
|
||||||
|
if err := cache.Put(key, payload, float64(time.Now().Unix())); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "web/search: cache: %v\n", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
out := websearch.Project(payload, limit, websearch.DefaultSnippetChars)
|
||||||
|
code := writeOut(out, jsonOut)
|
||||||
|
if out.Status != websearch.StatusOK && code == 0 {
|
||||||
|
return 3
|
||||||
|
}
|
||||||
|
return code
|
||||||
|
}
|
||||||
|
|
||||||
|
func writeOut(out websearch.Output, jsonOut bool) int {
|
||||||
|
if jsonOut {
|
||||||
|
enc := json.NewEncoder(os.Stdout)
|
||||||
|
enc.SetIndent("", " ")
|
||||||
|
enc.SetEscapeHTML(false)
|
||||||
|
if err := enc.Encode(out); err != nil {
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
if out.Status != websearch.StatusOK {
|
||||||
|
return 3
|
||||||
|
}
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
fmt.Print(out.YAML())
|
||||||
|
if out.Status != websearch.StatusOK {
|
||||||
|
return 3
|
||||||
|
}
|
||||||
|
return 0
|
||||||
|
}
|
||||||
@@ -61,6 +61,21 @@ services:
|
|||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
stop_grace_period: 20s
|
stop_grace_period: 20s
|
||||||
|
|
||||||
|
# Optional local SearXNG (D3). Skip if BRAIN_SEARCH_URL already points at a
|
||||||
|
# live instance — do not run a second copy on that host.
|
||||||
|
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
|
||||||
|
searxng:
|
||||||
|
profiles: ["searxng"]
|
||||||
|
image: docker.io/searxng/searxng:2026.8.10-0a118066d
|
||||||
|
ports:
|
||||||
|
- "127.0.0.1:8888:8080"
|
||||||
|
environment:
|
||||||
|
SEARXNG_SECRET: ${SEARXNG_SECRET:-}
|
||||||
|
volumes:
|
||||||
|
- ./deploy/searxng/settings.yml:/etc/searxng/settings.yml:ro
|
||||||
|
- ./deploy/searxng/limiter.toml:/etc/searxng/limiter.toml:ro
|
||||||
|
restart: unless-stopped
|
||||||
|
|
||||||
volumes:
|
volumes:
|
||||||
kb-model:
|
kb-model:
|
||||||
kb-var:
|
kb-var:
|
||||||
@@ -0,0 +1,7 @@
|
|||||||
|
[botdetection.ip_lists]
|
||||||
|
# RFC1918 only. Do not copy a live instance egress IP into git.
|
||||||
|
pass_ip = [
|
||||||
|
"10.0.0.0/8",
|
||||||
|
"172.16.0.0/12",
|
||||||
|
"192.168.0.0/16",
|
||||||
|
]
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
use_default_settings: true
|
||||||
|
|
||||||
|
general:
|
||||||
|
instance_name: "2dph"
|
||||||
|
|
||||||
|
search:
|
||||||
|
formats:
|
||||||
|
- html
|
||||||
|
- json
|
||||||
|
suspended_times:
|
||||||
|
SearxEngineCaptcha: 300
|
||||||
|
SearxEngineTooManyRequests: 120
|
||||||
|
SearxEngineAccessDenied: 300
|
||||||
|
|
||||||
|
server:
|
||||||
|
limiter: true
|
||||||
|
image_proxy: false
|
||||||
|
# secret_key comes from SEARXNG_SECRET (never commit a real secret)
|
||||||
|
|
||||||
|
engines:
|
||||||
|
- name: bing
|
||||||
|
disabled: false
|
||||||
|
- name: google
|
||||||
|
disabled: false
|
||||||
|
- name: duckduckgo
|
||||||
|
disabled: false
|
||||||
|
- name: wikipedia
|
||||||
|
disabled: false
|
||||||
+6
-1
@@ -6,5 +6,10 @@ Brain/ops/eSlider stack. Facts need proof or they are
|
|||||||
|
|
||||||
- [PLAN.md](../PLAN.md) — decisions, execution order, open questions (v2)
|
- [PLAN.md](../PLAN.md) — decisions, execution order, open questions (v2)
|
||||||
- [design](design.md) — schema, deduction model, sources
|
- [design](design.md) — schema, deduction model, sources
|
||||||
|
- [Gitea issues](https://git.produktor.io/eSlider/2dph/issues) — work board (origin)
|
||||||
|
|
||||||
Published docs live here and mirror the project state.
|
Search: `bin/brain/search.go "query"` (HTTP: `bin/brain/serve.go` —
|
||||||
|
`/health` `/search` `/get` `/stats` `/audit` `/ingest`). `--hop` is
|
||||||
|
not a walk; the flag errors until File/FROM_FILE edges exist.
|
||||||
|
|
||||||
|
Published docs live here and mirror the project state.
|
||||||
|
|||||||
@@ -15,9 +15,10 @@ OO_CLI (default: $HOME/go/bin/oo)
|
|||||||
## Quick reference
|
## Quick reference
|
||||||
|
|
||||||
```
|
```
|
||||||
./bin/chat sync telegram --limit 100
|
./bin/chats/sync.go telegram --limit 100
|
||||||
./bin/chat import
|
./bin/chats/import.go
|
||||||
./bin/chat index
|
./bin/chats/facts.go
|
||||||
./bin/chat facts
|
./bin/chats/apply.go --dry-run
|
||||||
./bin/chat apply --dry-run
|
|
||||||
```
|
```
|
||||||
|
|
||||||
|
JSONL → markdown only. Brain ingest is `bin/brain/index.go` (not a `chats index`).
|
||||||
|
|||||||
+7
-4
@@ -17,14 +17,16 @@ as one consistent state.
|
|||||||
## Deduction search
|
## Deduction search
|
||||||
|
|
||||||
```
|
```
|
||||||
bin/kb/search "question"
|
bin/brain/search.go "question"
|
||||||
1. facts root — confirmed answers only → return with evidence links
|
1. facts root — confirmed answers only → return with evidence links
|
||||||
2. info root — supporting narrative → snippets, marked (not confirmed)
|
2. info root — supporting narrative → snippets, marked (not confirmed)
|
||||||
3. web-search — second independent source → upgrade hypothesis to confirmed
|
3. web-search — second independent source → upgrade hypothesis to confirmed
|
||||||
|
(`web` block from `bin/web/search.go` when no facts hit; status `throttled`
|
||||||
|
is not evidence of absence; `--no-web` / `--root` skip it)
|
||||||
```
|
```
|
||||||
|
|
||||||
`--hop N` follows graph edges (sibling leaves under a heading, owning file,
|
`--hop` is not implemented yet (needs File/FROM_FILE edges). The flag is an
|
||||||
`related:` files, vector-neighbour leaves) — the deduction walk.
|
error; it is not a graph walk.
|
||||||
|
|
||||||
## Who / What / How / Where / When + evidence
|
## Who / What / How / Where / When + evidence
|
||||||
|
|
||||||
@@ -46,7 +48,8 @@ Every assertion edge carries:
|
|||||||
Content leafs: `sha256`, `observed_at`, `source_rev`, `confidence`. Stale = a
|
Content leafs: `sha256`, `observed_at`, `source_rev`, `confidence`. Stale = a
|
||||||
file changed on disk (git HEAD/mtime) after its last observed `source_rev`.
|
file changed on disk (git HEAD/mtime) after its last observed `source_rev`.
|
||||||
`File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person` records the history of every
|
`File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person` records the history of every
|
||||||
content leaf.
|
content leaf. Commit records come from `bin/git/import.go` (go-git, no git
|
||||||
|
binary); conversion prints leafs, brain write is `bin/brain/index.go`.
|
||||||
|
|
||||||
`bin/facts/audit stale` flags leafs whose observed revision is behind the
|
`bin/facts/audit stale` flags leafs whose observed revision is behind the
|
||||||
corpus HEAD.
|
corpus HEAD.
|
||||||
|
|||||||
@@ -1,8 +1,52 @@
|
|||||||
module github.com/eSlider/2dph
|
module github.com/eSlider/2dph
|
||||||
|
|
||||||
go 1.25.0
|
go 1.26
|
||||||
|
|
||||||
require (
|
require (
|
||||||
|
github.com/LadybugDB/go-ladybug v0.17.0
|
||||||
github.com/arran4/golang-ical v0.3.5
|
github.com/arran4/golang-ical v0.3.5
|
||||||
|
github.com/chewxy/math32 v1.11.2
|
||||||
|
github.com/daulet/tokenizers v1.27.0
|
||||||
|
github.com/go-git/go-git/v5 v5.19.2
|
||||||
|
golang.org/x/sys v0.47.0
|
||||||
golang.org/x/text v0.40.0
|
golang.org/x/text v0.40.0
|
||||||
|
modernc.org/sqlite v1.56.0
|
||||||
|
)
|
||||||
|
|
||||||
|
require (
|
||||||
|
dario.cat/mergo v1.0.0 // indirect
|
||||||
|
github.com/Microsoft/go-winio v0.6.2 // indirect
|
||||||
|
github.com/ProtonMail/go-crypto v1.1.6 // indirect
|
||||||
|
github.com/apache/arrow-go/v18 v18.6.0 // indirect
|
||||||
|
github.com/cloudflare/circl v1.6.3 // indirect
|
||||||
|
github.com/cyphar/filepath-securejoin v0.6.1 // indirect
|
||||||
|
github.com/dustin/go-humanize v1.0.1 // indirect
|
||||||
|
github.com/emirpasic/gods v1.18.1 // indirect
|
||||||
|
github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376 // indirect
|
||||||
|
github.com/go-git/go-billy/v5 v5.9.0 // indirect
|
||||||
|
github.com/goccy/go-json v0.10.6 // indirect
|
||||||
|
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 // indirect
|
||||||
|
github.com/google/flatbuffers v25.12.19+incompatible // indirect
|
||||||
|
github.com/google/uuid v1.6.0 // indirect
|
||||||
|
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99 // indirect
|
||||||
|
github.com/kevinburke/ssh_config v1.2.0 // indirect
|
||||||
|
github.com/klauspost/compress v1.18.5 // indirect
|
||||||
|
github.com/klauspost/cpuid/v2 v2.3.0 // indirect
|
||||||
|
github.com/mattn/go-isatty v0.0.24 // indirect
|
||||||
|
github.com/ncruces/go-strftime v1.0.0 // indirect
|
||||||
|
github.com/pierrec/lz4/v4 v4.1.26 // indirect
|
||||||
|
github.com/pjbgf/sha1cd v0.6.0 // indirect
|
||||||
|
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
|
||||||
|
github.com/sergi/go-diff v1.3.2-0.20230802210424-5b0b94c5c0d3 // indirect
|
||||||
|
github.com/shopspring/decimal v1.4.0 // indirect
|
||||||
|
github.com/skeema/knownhosts v1.3.1 // indirect
|
||||||
|
github.com/xanzy/ssh-agent v0.3.3 // indirect
|
||||||
|
github.com/zeebo/xxh3 v1.1.0 // indirect
|
||||||
|
golang.org/x/crypto v0.53.0 // indirect
|
||||||
|
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f // indirect
|
||||||
|
golang.org/x/net v0.56.0 // indirect
|
||||||
|
gopkg.in/warnings.v0 v0.1.2 // indirect
|
||||||
|
modernc.org/libc v1.74.4 // indirect
|
||||||
|
modernc.org/mathutil v1.7.1 // indirect
|
||||||
|
modernc.org/memory v1.11.0 // indirect
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -1,14 +1,184 @@
|
|||||||
|
dario.cat/mergo v1.0.0 h1:AGCNq9Evsj31mOgNPcLyXc+4PNABt905YmuqPYYpBWk=
|
||||||
|
dario.cat/mergo v1.0.0/go.mod h1:uNxQE+84aUszobStD9th8a29P2fMDhsBdgRYvZOxGmk=
|
||||||
|
github.com/LadybugDB/go-ladybug v0.17.0 h1:RXDbkBjrbRmLdEbhGl4CLOIEzSt09gbP0n9UbKDEfwI=
|
||||||
|
github.com/LadybugDB/go-ladybug v0.17.0/go.mod h1:GeIXmE8XyF5TFS94NAuTag7vgCC+no/HTBMRA6Rd5Cs=
|
||||||
|
github.com/Microsoft/go-winio v0.5.2/go.mod h1:WpS1mjBmmwHBEWmogvA2mj8546UReBk4v8QkMxJ6pZY=
|
||||||
|
github.com/Microsoft/go-winio v0.6.2 h1:F2VQgta7ecxGYO8k3ZZz3RS8fVIXVxONVUPlNERoyfY=
|
||||||
|
github.com/Microsoft/go-winio v0.6.2/go.mod h1:yd8OoFMLzJbo9gZq8j5qaps8bJ9aShtEA8Ipt1oGCvU=
|
||||||
|
github.com/ProtonMail/go-crypto v1.1.6 h1:ZcV+Ropw6Qn0AX9brlQLAUXfqLBc7Bl+f/DmNxpLfdw=
|
||||||
|
github.com/ProtonMail/go-crypto v1.1.6/go.mod h1:rA3QumHc/FZ8pAHreoekgiAbzpNsfQAosU5td4SnOrE=
|
||||||
|
github.com/andybalholm/brotli v1.2.1 h1:R+f5xP285VArJDRgowrfb9DqL18yVK0gKAW/F+eTWro=
|
||||||
|
github.com/andybalholm/brotli v1.2.1/go.mod h1:rzTDkvFWvIrjDXZHkuS16NPggd91W3kUSvPlQ1pLaKY=
|
||||||
|
github.com/anmitsu/go-shlex v0.0.0-20200514113438-38f4b401e2be h1:9AeTilPcZAjCFIImctFaOjnTIavg87rW78vTPkQqLI8=
|
||||||
|
github.com/anmitsu/go-shlex v0.0.0-20200514113438-38f4b401e2be/go.mod h1:ySMOLuWl6zY27l47sB3qLNK6tF2fkHG55UZxx8oIVo4=
|
||||||
|
github.com/apache/arrow-go/v18 v18.6.0 h1:GX/Jyd3R7mCLiECAwY9FWbbaYblie2WXBSz4Sw8fNpM=
|
||||||
|
github.com/apache/arrow-go/v18 v18.6.0/go.mod h1:gm3MiPpY82fLYK5VKPB3WoJbsiLVDfT7flD5/vHReKw=
|
||||||
|
github.com/apache/thrift v0.22.0 h1:r7mTJdj51TMDe6RtcmNdQxgn9XcyfGDOzegMDRg47uc=
|
||||||
|
github.com/apache/thrift v0.22.0/go.mod h1:1e7J/O1Ae6ZQMTYdy9xa3w9k+XHWPfRvdPyJeynQ+/g=
|
||||||
|
github.com/armon/go-socks5 v0.0.0-20160902184237-e75332964ef5 h1:0CwZNZbxp69SHPdPJAN/hZIm0C4OItdklCFmMRWYpio=
|
||||||
|
github.com/armon/go-socks5 v0.0.0-20160902184237-e75332964ef5/go.mod h1:wHh0iHkYZB8zMSxRWpUBQtwG5a7fFgvEO+odwuTv2gs=
|
||||||
github.com/arran4/golang-ical v0.3.5 h1:bbz6ld4dC+MmCKiFfOd6SkmIGnhNMBACZ485ULh7p9A=
|
github.com/arran4/golang-ical v0.3.5 h1:bbz6ld4dC+MmCKiFfOd6SkmIGnhNMBACZ485ULh7p9A=
|
||||||
github.com/arran4/golang-ical v0.3.5/go.mod h1:OnguFgjN0Hmx8jzpmWcC+AkHio94ujmLHKoaef7xQh8=
|
github.com/arran4/golang-ical v0.3.5/go.mod h1:OnguFgjN0Hmx8jzpmWcC+AkHio94ujmLHKoaef7xQh8=
|
||||||
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
|
github.com/chewxy/math32 v1.11.2 h1:IufN08Zwr1NKuWfY+4Tz55BcwKmyKKNdOP7KtumehnM=
|
||||||
|
github.com/chewxy/math32 v1.11.2/go.mod h1:dOB2rcuFrCn6UHrze36WSLVPKtzPMRAQvBvUwkSsLqs=
|
||||||
|
github.com/cloudflare/circl v1.6.3 h1:9GPOhQGF9MCYUeXyMYlqTR6a5gTrgR/fBLXvUgtVcg8=
|
||||||
|
github.com/cloudflare/circl v1.6.3/go.mod h1:2eXP6Qfat4O/Yhh8BznvKnJ+uzEoTQ6jVKJRn81BiS4=
|
||||||
|
github.com/cyphar/filepath-securejoin v0.6.1 h1:5CeZ1jPXEiYt3+Z6zqprSAgSWiggmpVyciv8syjIpVE=
|
||||||
|
github.com/cyphar/filepath-securejoin v0.6.1/go.mod h1:A8hd4EnAeyujCJRrICiOWqjS1AX0a9kM5XL+NwKoYSc=
|
||||||
|
github.com/daulet/tokenizers v1.27.0 h1:MmFYAEDFz69s/nNQfHg59DWqHz3v94m99kEZ/JbL+s4=
|
||||||
|
github.com/daulet/tokenizers v1.27.0/go.mod h1:YjFY1o1HGMyWkQgbXJDghhvke/yFDp2vGdIO2hYs4MQ=
|
||||||
|
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
||||||
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
||||||
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI=
|
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM=
|
||||||
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY=
|
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
||||||
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
|
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
|
||||||
|
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
|
||||||
|
github.com/elazarl/goproxy v1.7.2 h1:Y2o6urb7Eule09PjlhQRGNsqRfPmYI3KKQLFpCAV3+o=
|
||||||
|
github.com/elazarl/goproxy v1.7.2/go.mod h1:82vkLNir0ALaW14Rc399OTTjyNREgmdL2cVoIbS6XaE=
|
||||||
|
github.com/emirpasic/gods v1.18.1 h1:FXtiHYKDGKCW2KzwZKx0iC0PQmdlorYgdFG9jPXJ1Bc=
|
||||||
|
github.com/emirpasic/gods v1.18.1/go.mod h1:8tpGGwCnJ5H4r6BWwaV6OrWmMoPhUl5jm/FMNAnJvWQ=
|
||||||
|
github.com/gliderlabs/ssh v0.3.8 h1:a4YXD1V7xMF9g5nTkdfnja3Sxy1PVDCj1Zg4Wb8vY6c=
|
||||||
|
github.com/gliderlabs/ssh v0.3.8/go.mod h1:xYoytBv1sV0aL3CavoDuJIQNURXkkfPA/wxQ1pL1fAU=
|
||||||
|
github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376 h1:+zs/tPmkDkHx3U66DAb0lQFJrpS6731Oaa12ikc+DiI=
|
||||||
|
github.com/go-git/gcfg v1.5.1-0.20230307220236-3a3c6141e376/go.mod h1:an3vInlBmSxCcxctByoQdvwPiA7DTK7jaaFDBTtu0ic=
|
||||||
|
github.com/go-git/go-billy/v5 v5.9.0 h1:jItGXszUDRtR/AlferWPTMN4j38BQ88XnXKbilmmBPA=
|
||||||
|
github.com/go-git/go-billy/v5 v5.9.0/go.mod h1:jCnQMLj9eUgGU7+ludSTYoZL/GGmii14RxKFj7ROgHw=
|
||||||
|
github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399 h1:eMje31YglSBqCdIqdhKBW8lokaMrL3uTkpGYlE2OOT4=
|
||||||
|
github.com/go-git/go-git-fixtures/v4 v4.3.2-0.20231010084843-55a94097c399/go.mod h1:1OCfN199q1Jm3HZlxleg+Dw/mwps2Wbk9frAWm+4FII=
|
||||||
|
github.com/go-git/go-git/v5 v5.19.2 h1:wkfn7vOlUBu8ivAWKBWisTiwJK4jYHzTF8Ndv1LyGqY=
|
||||||
|
github.com/go-git/go-git/v5 v5.19.2/go.mod h1:QqCBE1EFN5ddFmrliLQ3/ntRCUjZU3EJuwuB/jWEHjk=
|
||||||
|
github.com/goccy/go-json v0.10.6 h1:p8HrPJzOakx/mn/bQtjgNjdTcN+/S6FcG2CTtQOrHVU=
|
||||||
|
github.com/goccy/go-json v0.10.6/go.mod h1:oq7eo15ShAhp70Anwd5lgX2pLfOS3QCiwU/PULtXL6M=
|
||||||
|
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8 h1:f+oWsMOmNPc8JmEHVZIycC7hBoQxHH9pNKQORJNozsQ=
|
||||||
|
github.com/golang/groupcache v0.0.0-20241129210726-2c02b8208cf8/go.mod h1:wcDNUvekVysuuOpQKo3191zZyTpiI6se1N1ULghS0sw=
|
||||||
|
github.com/google/flatbuffers v25.12.19+incompatible h1:haMV2JRRJCe1998HeW/p0X9UaMTK6SDo0ffLn2+DbLs=
|
||||||
|
github.com/google/flatbuffers v25.12.19+incompatible/go.mod h1:1AeVuKshWv4vARoZatz6mlQ0JxURH0Kv5+zNeJKJCa8=
|
||||||
|
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
|
||||||
|
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
|
||||||
|
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3 h1:LMLX+LgTNWpfvCBdFebv6EsYotImrt/Ppc5cXIriCSo=
|
||||||
|
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3/go.mod h1:jl5iWTm0/hd5PjEYEOuwAJ57L/CibdZfrqZ5XA5GrCk=
|
||||||
|
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
|
||||||
|
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
|
||||||
|
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
|
||||||
|
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
|
||||||
|
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99 h1:BQSFePA1RWJOlocH6Fxy8MmwDt+yVQYULKfN0RoTN8A=
|
||||||
|
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99/go.mod h1:1lJo3i6rXxKeerYnT8Nvf0QmHCRC1n8sfWVwXF2Frvo=
|
||||||
|
github.com/kevinburke/ssh_config v1.2.0 h1:x584FjTGwHzMwvHx18PXxbBVzfnxogHaAReU4gf13a4=
|
||||||
|
github.com/kevinburke/ssh_config v1.2.0/go.mod h1:CT57kijsi8u/K/BOFA39wgDQJ9CxiF4nAY/ojJ6r6mM=
|
||||||
|
github.com/klauspost/compress v1.18.5 h1:/h1gH5Ce+VWNLSWqPzOVn6XBO+vJbCNGvjoaGBFW2IE=
|
||||||
|
github.com/klauspost/compress v1.18.5/go.mod h1:cwPg85FWrGar70rWktvGQj8/hthj3wpl0PGDogxkrSQ=
|
||||||
|
github.com/klauspost/cpuid/v2 v2.3.0 h1:S4CRMLnYUhGeDFDqkGriYKdfoFlDnMtqTiI/sFzhA9Y=
|
||||||
|
github.com/klauspost/cpuid/v2 v2.3.0/go.mod h1:hqwkgyIinND0mEev00jJYCxPNVRVXFQeu1XKlok6oO0=
|
||||||
|
github.com/kr/pretty v0.1.0/go.mod h1:dAy3ld7l9f0ibDNOQOHHMYYIIbhfbHSm3C4ZsoJORNo=
|
||||||
|
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
|
||||||
|
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
|
||||||
|
github.com/kr/pty v1.1.1/go.mod h1:pFQYn66WHrOpPYNljwOMqo10TkYh1fy3cYio2l3bCsQ=
|
||||||
|
github.com/kr/text v0.1.0/go.mod h1:4Jbv+DJW3UT/LiOwJeYQe1efqtUx/iVham/4vfdArNI=
|
||||||
|
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
|
||||||
|
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
|
||||||
|
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
|
||||||
|
github.com/mattn/go-isatty v0.0.24/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
|
||||||
|
github.com/ncruces/go-strftime v1.0.0 h1:HMFp8mLCTPp341M/ZnA4qaf7ZlsbTc+miZjCLOFAw7w=
|
||||||
|
github.com/ncruces/go-strftime v1.0.0/go.mod h1:Fwc5htZGVVkseilnfgOVb9mKy6w1naJmn9CehxcKcls=
|
||||||
|
github.com/onsi/gomega v1.34.1 h1:EUMJIKUjM8sKjYbtxQI9A4z2o+rruxnzNvpknOXie6k=
|
||||||
|
github.com/onsi/gomega v1.34.1/go.mod h1:kU1QgUvBDLXBJq618Xvm2LUX6rSAfRaFRTcdOeDLwwY=
|
||||||
|
github.com/pierrec/lz4/v4 v4.1.26 h1:GrpZw1gZttORinvzBdXPUXATeqlJjqUG/D87TKMnhjY=
|
||||||
|
github.com/pierrec/lz4/v4 v4.1.26/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
|
||||||
|
github.com/pjbgf/sha1cd v0.6.0 h1:3WJ8Wz8gvDz29quX1OcEmkAlUg9diU4GxJHqs0/XiwU=
|
||||||
|
github.com/pjbgf/sha1cd v0.6.0/go.mod h1:lhpGlyHLpQZoxMv8HcgXvZEhcGs0PG/vsZnEJ7H0iCM=
|
||||||
|
github.com/pkg/errors v0.9.1 h1:FEBLx1zS214owpjy7qsBeixbURkuhQAwrK5UwLGTwt4=
|
||||||
|
github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
|
||||||
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
|
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
|
||||||
github.com/stretchr/testify v1.7.0 h1:nwc3DEeHmmLAfoZucVR881uASk0Mfjw8xYJ99tb5CcY=
|
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 h1:Jamvg5psRIccs7FGNTlIRMkT8wgtp5eCXdBlqhYGL6U=
|
||||||
github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
|
github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
|
||||||
|
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE=
|
||||||
|
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo=
|
||||||
|
github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ=
|
||||||
|
github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc=
|
||||||
|
github.com/sergi/go-diff v1.3.2-0.20230802210424-5b0b94c5c0d3 h1:n661drycOFuPLCN3Uc8sB6B/s6Z4t2xvBgU1htSHuq8=
|
||||||
|
github.com/sergi/go-diff v1.3.2-0.20230802210424-5b0b94c5c0d3/go.mod h1:A0bzQcvG0E7Rwjx0REVgAGH58e96+X0MeOfepqsbeW4=
|
||||||
|
github.com/shopspring/decimal v1.4.0 h1:bxl37RwXBklmTi0C79JfXCEBD1cqqHt0bbgBAGFp81k=
|
||||||
|
github.com/shopspring/decimal v1.4.0/go.mod h1:gawqmDU56v4yIKSwfBSFip1HdCCXN8/+DMd9qYNcwME=
|
||||||
|
github.com/sirupsen/logrus v1.7.0/go.mod h1:yWOB1SBYBC5VeMP7gHvWumXLIWorT60ONWic61uBYv0=
|
||||||
|
github.com/skeema/knownhosts v1.3.1 h1:X2osQ+RAjK76shCbvhHHHVl3ZlgDm8apHEHFqRjnBY8=
|
||||||
|
github.com/skeema/knownhosts v1.3.1/go.mod h1:r7KTdC8l4uxWRyK2TpQZ/1o5HaSzh06ePQNxPwTcfiY=
|
||||||
|
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
|
||||||
|
github.com/stretchr/testify v1.2.2/go.mod h1:a8OnRcib4nhh0OaRAV+Yts87kKdq0PP7pXfy6kDkUVs=
|
||||||
|
github.com/stretchr/testify v1.4.0/go.mod h1:j7eGeouHqKxXV5pUuKE4zz7dFj8WfuZ+81PSLYec5m4=
|
||||||
|
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
|
||||||
|
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
|
||||||
|
github.com/xanzy/ssh-agent v0.3.3 h1:+/15pJfg/RsTxqYcX6fHqOXZwwMP+2VyYWJeWM2qQFM=
|
||||||
|
github.com/xanzy/ssh-agent v0.3.3/go.mod h1:6dzNDKs0J9rVPHPhaGCukekBHKqfl+L3KghI1Bc68Uw=
|
||||||
|
github.com/zeebo/assert v1.3.0 h1:g7C04CbJuIDKNPFHmsk4hwZDO5O+kntRxzaUoNXj+IQ=
|
||||||
|
github.com/zeebo/assert v1.3.0/go.mod h1:Pq9JiuJQpG8JLJdtkwrJESF0Foym2/D9XMU5ciN/wJ0=
|
||||||
|
github.com/zeebo/xxh3 v1.1.0 h1:s7DLGDK45Dyfg7++yxI0khrfwq9661w9EN78eP/UZVs=
|
||||||
|
github.com/zeebo/xxh3 v1.1.0/go.mod h1:IisAie1LELR4xhVinxWS5+zf1lA4p0MW4T+w+W07F5s=
|
||||||
|
golang.org/x/crypto v0.0.0-20220622213112-05595931fe9d/go.mod h1:IxCIyHEi3zRg3s0A5j5BB6A9Jmi73HwBIUl50j+osU4=
|
||||||
|
golang.org/x/crypto v0.53.0 h1:QZ4Muo8THX6CizN2vPPd5fBGHyogrdK9fG4wLPFUsto=
|
||||||
|
golang.org/x/crypto v0.53.0/go.mod h1:DNLU434OwVakk9PzuwV8w62mAJpRJL3vsgcfp4Qnsio=
|
||||||
|
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f h1:W3F4c+6OLc6H2lb//N1q4WpJkhzJCK5J6kUi1NTVXfM=
|
||||||
|
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f/go.mod h1:J1xhfL/vlindoeF/aINzNzt2Bket5bjo9sdOYzOsU80=
|
||||||
|
golang.org/x/mod v0.37.0 h1:vF1DjpVEshcIqoEaauuHebaLk1O1forxjxBaVn884JQ=
|
||||||
|
golang.org/x/mod v0.37.0/go.mod h1:m8S8VeM9r4dzDwjrKO0a1sZP3YjeMamRRlD+fmR2Q/0=
|
||||||
|
golang.org/x/net v0.0.0-20211112202133-69e39bad7dc2/go.mod h1:9nx3DQGgdP8bBQD5qxJ1jj9UTztislL4KSBs9R2vV5Y=
|
||||||
|
golang.org/x/net v0.56.0 h1:Rw8j/hFzGvJUZwNBXnAtf5sVDVt+65SK2C7IxCxZt5o=
|
||||||
|
golang.org/x/net v0.56.0/go.mod h1:D3Ku6r+V6JROoZK144D2XfMHFcMq/0zSfLelVTCFKec=
|
||||||
|
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
|
||||||
|
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
|
||||||
|
golang.org/x/sys v0.0.0-20191026070338-33540a1f6037/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
|
||||||
|
golang.org/x/sys v0.0.0-20201119102817-f84b799fce68/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
|
||||||
|
golang.org/x/sys v0.0.0-20210124154548-22da62e12c0c/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
|
||||||
|
golang.org/x/sys v0.0.0-20210423082822-04245dca01da/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
|
||||||
|
golang.org/x/sys v0.0.0-20210615035016-665e8c7367d1/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
|
||||||
|
golang.org/x/sys v0.0.0-20220715151400-c0bba94af5f8/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
|
||||||
|
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
|
||||||
|
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
|
||||||
|
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
|
||||||
|
golang.org/x/term v0.44.0 h1:0rLvDRCtNj0gZkyIXhCyOb2OAzEhLVqc4B+hrsBhrmc=
|
||||||
|
golang.org/x/term v0.44.0/go.mod h1:7ze4MdzUzLXpSAoFP1H0bOI9aXDqveSvatT5vKcFh2Y=
|
||||||
|
golang.org/x/text v0.3.6/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ=
|
||||||
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
|
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
|
||||||
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
|
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
|
||||||
|
golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ=
|
||||||
|
golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q=
|
||||||
|
golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA=
|
||||||
|
gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4=
|
||||||
|
gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E=
|
||||||
|
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
|
||||||
|
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
|
||||||
|
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
|
||||||
|
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
|
||||||
|
gopkg.in/warnings.v0 v0.1.2 h1:wFXVbFY8DY5/xOe1ECiWdKCzZlxgshcYVNkBHstARME=
|
||||||
|
gopkg.in/warnings.v0 v0.1.2/go.mod h1:jksf8JmL6Qr/oQM2OXTHunEvvTAsrWBLb6OOjuVWRNI=
|
||||||
|
gopkg.in/yaml.v2 v2.2.2/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
|
||||||
|
gopkg.in/yaml.v2 v2.4.0/go.mod h1:RDklbk79AGWmwhnvt/jBztapEOGDOx6ZbXqjP6csGnQ=
|
||||||
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
|
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
|
||||||
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
|
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
|
||||||
|
modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI=
|
||||||
|
modernc.org/cc/v4 v4.29.1/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
|
||||||
|
modernc.org/ccgo/v4 v4.34.6 h1:sBgfIwyN0TQ9C5hwIeuqyeAKyMWnbvj2fvpF4L11uzU=
|
||||||
|
modernc.org/ccgo/v4 v4.34.6/go.mod h1:SZ8YcN9NG7XVsQYdm6jYBvi8PQP1qi+kqB6OhjqI3Fk=
|
||||||
|
modernc.org/fileutil v1.4.0 h1:j6ZzNTftVS054gi281TyLjHPp6CPHr2KCxEXjEbD6SM=
|
||||||
|
modernc.org/fileutil v1.4.0/go.mod h1:EqdKFDxiByqxLk8ozOxObDSfcVOv/54xDs/DUHdvCUU=
|
||||||
|
modernc.org/gc/v2 v2.6.5 h1:nyqdV8q46KvTpZlsw66kWqwXRHdjIlJOhG6kxiV/9xI=
|
||||||
|
modernc.org/gc/v2 v2.6.5/go.mod h1:YgIahr1ypgfe7chRuJi2gD7DBQiKSLMPgBQe9oIiito=
|
||||||
|
modernc.org/gc/v3 v3.1.4 h1:2g65LGVSmFQrXeITAw97x7hCRvZFcyE1uDP+7Vng7JI=
|
||||||
|
modernc.org/gc/v3 v3.1.4/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
|
||||||
|
modernc.org/goabi0 v0.2.0 h1:HvEowk7LxcPd0eq6mVOAEMai46V+i7Jrj13t4AzuNks=
|
||||||
|
modernc.org/goabi0 v0.2.0/go.mod h1:CEFRnnJhKvWT1c1JTI3Avm+tgOWbkOu5oPA8eH8LnMI=
|
||||||
|
modernc.org/libc v1.74.4 h1:fX1Omw4o2/1C2iRkkIsrQTasJQldLhRmuPreXLoWs9k=
|
||||||
|
modernc.org/libc v1.74.4/go.mod h1:eeQAS9W3sZeKYMFubydxJpII9ybHWshk+7or7bLG9co=
|
||||||
|
modernc.org/mathutil v1.7.1 h1:GCZVGXdaN8gTqB1Mf/usp1Y/hSqgI2vAGGP4jZMCxOU=
|
||||||
|
modernc.org/mathutil v1.7.1/go.mod h1:4p5IwJITfppl0G4sUEDtCr4DthTaT47/N3aT6MhfgJg=
|
||||||
|
modernc.org/memory v1.11.0 h1:o4QC8aMQzmcwCK3t3Ux/ZHmwFPzE6hf2Y5LbkRs+hbI=
|
||||||
|
modernc.org/memory v1.11.0/go.mod h1:/JP4VbVC+K5sU2wZi9bHoq2MAkCnrt2r98UGeSK7Mjw=
|
||||||
|
modernc.org/opt v0.2.0 h1:tGyef5ApycA7FSEOMraay9SaTk5zmbx7Tu+cJs4QKZg=
|
||||||
|
modernc.org/opt v0.2.0/go.mod h1:03fq9lsNfvkYSfxrfUhZCWPk1lm4cq4N+Bh//bEtgns=
|
||||||
|
modernc.org/sortutil v1.2.1 h1:+xyoGf15mM3NMlPDnFqrteY07klSFxLElE2PVuWIJ7w=
|
||||||
|
modernc.org/sortutil v1.2.1/go.mod h1:7ZI3a3REbai7gzCLcotuw9AC4VZVpYMjDzETGsSMqJE=
|
||||||
|
modernc.org/sqlite v1.56.0 h1:/D8e2RfFqoy/Zc6PuC76U28zFwmI/sYx1Kjm4yEn9e0=
|
||||||
|
modernc.org/sqlite v1.56.0/go.mod h1:yCJ2cmAaIkHQ25oXWrF8H4O1lIfPYPR26yCEDj2P3pQ=
|
||||||
|
modernc.org/strutil v1.2.1 h1:UneZBkQA+DX2Rp35KcM69cSsNES9ly8mQWD71HKlOA0=
|
||||||
|
modernc.org/strutil v1.2.1/go.mod h1:EHkiggD70koQxjVdSBM3JKM7k6L0FbGE5eymy9i3B9A=
|
||||||
|
modernc.org/token v1.1.0 h1:Xl7Ap9dKaEs5kLoOQeQmPWevfnk/DM5qcLcYlA8ys6Y=
|
||||||
|
modernc.org/token v1.1.0/go.mod h1:UGzOrNV1mAFSEB63lOFHIpNRUVMvYTc6yu1SMY/XTDM=
|
||||||
|
|||||||
@@ -1,10 +1,12 @@
|
|||||||
// Brain connection management using go-ladybug.
|
//go:build cgo && system_ladybug
|
||||||
package main
|
|
||||||
|
package brain
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
|
||||||
lbug "github.com/LadybugDB/go-ladybug"
|
lbug "github.com/LadybugDB/go-ladybug"
|
||||||
)
|
)
|
||||||
@@ -44,10 +46,10 @@ func dbPath() string {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func openBrain() error {
|
func openBrain() error {
|
||||||
return openWithOpts(2, eps())
|
return openWithSandbox(eps())
|
||||||
}
|
}
|
||||||
|
|
||||||
func openWithOpts(allow int, epsv string) error {
|
func openWithSandbox(epsv string) error {
|
||||||
cfg := lbug.DefaultSystemConfig()
|
cfg := lbug.DefaultSystemConfig()
|
||||||
cfg.MaxNumThreads = 8
|
cfg.MaxNumThreads = 8
|
||||||
cfg.BufferPoolSize = 1 << 30 // 1GB
|
cfg.BufferPoolSize = 1 << 30 // 1GB
|
||||||
@@ -57,20 +59,30 @@ func openWithOpts(allow int, epsv string) error {
|
|||||||
if err != nil {
|
if err != nil {
|
||||||
return fmt.Errorf("OpenDatabase: %w", err)
|
return fmt.Errorf("OpenDatabase: %w", err)
|
||||||
}
|
}
|
||||||
if epsv != "" {
|
|
||||||
if _, err := conn.Query("SET STREAM_SANDBOX = '" + epsv + "'"); err != nil {
|
|
||||||
return err
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
conn, err = lbug.OpenConnection(db)
|
conn, err = lbug.OpenConnection(db)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
|
closeBrain()
|
||||||
return fmt.Errorf("OpenConnection: %w", err)
|
return fmt.Errorf("OpenConnection: %w", err)
|
||||||
}
|
}
|
||||||
|
// Session settings need a live connection; running this before
|
||||||
|
// OpenConnection dereferenced a nil *Connection.
|
||||||
|
if epsv != "" {
|
||||||
|
if strings.ContainsAny(epsv, "'\\") {
|
||||||
|
closeBrain()
|
||||||
|
return fmt.Errorf("SET STREAM_SANDBOX: invalid value")
|
||||||
|
}
|
||||||
|
if _, err := conn.Query("SET STREAM_SANDBOX = '" + epsv + "'"); err != nil {
|
||||||
|
closeBrain()
|
||||||
|
return fmt.Errorf("SET STREAM_SANDBOX: %w", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
if _, err := conn.Query("LOAD EXTENSION FTS"); err != nil {
|
if _, err := conn.Query("LOAD EXTENSION FTS"); err != nil {
|
||||||
|
closeBrain()
|
||||||
return fmt.Errorf("LOAD EXTENSION FTS: %w", err)
|
return fmt.Errorf("LOAD EXTENSION FTS: %w", err)
|
||||||
}
|
}
|
||||||
if _, err := conn.Query("LOAD EXTENSION VECTOR"); err != nil {
|
if _, err := conn.Query("LOAD EXTENSION VECTOR"); err != nil {
|
||||||
|
closeBrain()
|
||||||
return fmt.Errorf("LOAD EXTENSION VECTOR: %w", err)
|
return fmt.Errorf("LOAD EXTENSION VECTOR: %w", err)
|
||||||
}
|
}
|
||||||
return nil
|
return nil
|
||||||
@@ -85,4 +97,4 @@ func closeBrain() {
|
|||||||
db.Close()
|
db.Close()
|
||||||
db = nil
|
db = nil
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
// Package brain is deduction search over Ladybug (FTS + HNSW).
|
||||||
|
// Query/embed code that needs cgo lives behind the system_ladybug tag.
|
||||||
|
package brain
|
||||||
@@ -0,0 +1,159 @@
|
|||||||
|
//go:build cgo && system_ladybug
|
||||||
|
|
||||||
|
package brain
|
||||||
|
|
||||||
|
import (
|
||||||
|
"bytes"
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/brain/rank"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Ready opens the Ladybug file for the life of the serve process.
|
||||||
|
func Ready() error {
|
||||||
|
return openBrain()
|
||||||
|
}
|
||||||
|
|
||||||
|
// HTTP is the in-process API used by bin/brain/serve.go.
|
||||||
|
type HTTP struct{}
|
||||||
|
|
||||||
|
func (HTTP) Search(ctx context.Context, query string, limit int) ([]byte, error) {
|
||||||
|
hits, err := searchHits(query, "", "", limit)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
for i := range hits {
|
||||||
|
if hits[i].Text != "" {
|
||||||
|
runes := []rune(hits[i].Text)
|
||||||
|
if len(runes) > 280 {
|
||||||
|
runes = runes[:280]
|
||||||
|
}
|
||||||
|
hits[i].Snippet = string(runes)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
webOut := rank.Deduce(hits, query, "", false, func(q string) rank.SecondSource {
|
||||||
|
return lookupWeb(ctx, q)
|
||||||
|
})
|
||||||
|
var buf bytes.Buffer
|
||||||
|
enc := json.NewEncoder(&buf)
|
||||||
|
enc.SetEscapeHTML(false)
|
||||||
|
if err := enc.Encode(toJSONOut(hits, query, "", webOut)); err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
return buf.Bytes(), nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (HTTP) Get(_ context.Context, id string, body bool) ([]byte, error) {
|
||||||
|
if conn == nil {
|
||||||
|
return nil, fmt.Errorf("brain not open")
|
||||||
|
}
|
||||||
|
stmt, err := conn.Prepare(
|
||||||
|
"MATCH (l:Leaf {id:$id}) RETURN l.id, l.text, l.root, l.confidence, l.source, l.type",
|
||||||
|
)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
defer stmt.Close()
|
||||||
|
res, err := conn.Execute(stmt, map[string]any{"id": id})
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
if !res.HasNext() {
|
||||||
|
return nil, fmt.Errorf("no leaf %s", id)
|
||||||
|
}
|
||||||
|
row, err := res.Next()
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
vals, err := row.GetAsSlice()
|
||||||
|
if err != nil || len(vals) < 6 {
|
||||||
|
return nil, fmt.Errorf("leaf row")
|
||||||
|
}
|
||||||
|
out := map[string]any{
|
||||||
|
"id": fmt.Sprint(vals[0]),
|
||||||
|
"root": fmt.Sprint(vals[2]),
|
||||||
|
"confidence": fmt.Sprint(vals[3]),
|
||||||
|
"source": fmt.Sprint(vals[4]),
|
||||||
|
"type": fmt.Sprint(vals[5]),
|
||||||
|
}
|
||||||
|
if body {
|
||||||
|
out["text"] = fmt.Sprint(vals[1])
|
||||||
|
}
|
||||||
|
return json.Marshal(out)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (HTTP) Stats(context.Context) ([]byte, error) {
|
||||||
|
if conn == nil {
|
||||||
|
return nil, fmt.Errorf("brain not open")
|
||||||
|
}
|
||||||
|
res, err := conn.Query("MATCH (l:Leaf) RETURN l.root, count(*)")
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
byRoot := map[string]int{}
|
||||||
|
total := 0
|
||||||
|
for res.HasNext() {
|
||||||
|
row, err := res.Next()
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
vals, err := row.GetAsSlice()
|
||||||
|
if err != nil || len(vals) < 2 {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
n := int(asInt(vals[1]))
|
||||||
|
byRoot[fmt.Sprint(vals[0])] = n
|
||||||
|
total += n
|
||||||
|
}
|
||||||
|
return json.Marshal(map[string]any{"total": total, "by_root": byRoot, "db": dbPath()})
|
||||||
|
}
|
||||||
|
|
||||||
|
func (HTTP) Audit(context.Context) ([]byte, error) {
|
||||||
|
if conn == nil {
|
||||||
|
return nil, fmt.Errorf("brain not open")
|
||||||
|
}
|
||||||
|
res, err := conn.Query("MATCH (l:Leaf) RETURN l.root, l.confidence, count(*)")
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
var rows []map[string]any
|
||||||
|
for res.HasNext() {
|
||||||
|
row, err := res.Next()
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
vals, err := row.GetAsSlice()
|
||||||
|
if err != nil || len(vals) < 3 {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
rows = append(rows, map[string]any{
|
||||||
|
"root": fmt.Sprint(vals[0]),
|
||||||
|
"confidence": fmt.Sprint(vals[1]),
|
||||||
|
"count": asInt(vals[2]),
|
||||||
|
})
|
||||||
|
}
|
||||||
|
return json.Marshal(map[string]any{"status": "ok", "by_confidence": rows})
|
||||||
|
}
|
||||||
|
|
||||||
|
func (HTTP) Ingest(context.Context) ([]byte, error) {
|
||||||
|
return json.Marshal(map[string]any{
|
||||||
|
"mode": "rebuild",
|
||||||
|
"command": "bin/brain/index.go --rebuild",
|
||||||
|
"add": "v2",
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
func asInt(v any) int64 {
|
||||||
|
switch n := v.(type) {
|
||||||
|
case int64:
|
||||||
|
return n
|
||||||
|
case int:
|
||||||
|
return int64(n)
|
||||||
|
case float64:
|
||||||
|
return int64(n)
|
||||||
|
default:
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,9 +1,11 @@
|
|||||||
|
//go:build cgo && system_ladybug
|
||||||
|
|
||||||
// StaticModel wraps the potion-multilingual-128m embedding model.
|
// StaticModel wraps the potion-multilingual-128m embedding model.
|
||||||
//
|
//
|
||||||
// Mirrors model2vec.StaticModel: tokenizer (daulet Unigram) + safetensors matrix.
|
// Mirrors model2vec.StaticModel: tokenizer (daulet Unigram) + safetensors matrix.
|
||||||
// Embed(text) applies the same preprocessing: median_token_length pre-truncation,
|
// Embed(text) applies the same preprocessing: median_token_length pre-truncation,
|
||||||
// add_special_tokens=false, drop unk (id=1), truncate to 512, mean pool, L2 normalize +1e-32.
|
// add_special_tokens=false, drop unk (id=1), truncate to 512, mean pool, L2 normalize +1e-32.
|
||||||
package main
|
package brain
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"encoding/json"
|
"encoding/json"
|
||||||
@@ -1,5 +1,4 @@
|
|||||||
// modelDir returns the resolved potion-multilingual-128m model directory.
|
package brain
|
||||||
package main
|
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"fmt"
|
"fmt"
|
||||||
@@ -0,0 +1,75 @@
|
|||||||
|
package rank
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
"strconv"
|
||||||
|
"strings"
|
||||||
|
)
|
||||||
|
|
||||||
|
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--json] [--no-web]
|
||||||
|
bin/brain/search.go serve [port]
|
||||||
|
bin/brain/search.go --list-model`
|
||||||
|
|
||||||
|
type Options struct {
|
||||||
|
Query string
|
||||||
|
Root string
|
||||||
|
Repo string
|
||||||
|
Limit int
|
||||||
|
JSONOut bool
|
||||||
|
ListModel bool
|
||||||
|
NoWeb bool
|
||||||
|
}
|
||||||
|
|
||||||
|
// ParseArgs reads flags. Unknown flags are an error: silently dropping them
|
||||||
|
// meant `--hop 1` vanished and its argument `1` was appended to the query.
|
||||||
|
// --hop is recognised so it cannot be swallowed; it is not implemented until
|
||||||
|
// File/FROM_FILE edges exist.
|
||||||
|
func ParseArgs(args []string) (Options, error) {
|
||||||
|
opt := Options{Limit: 20}
|
||||||
|
var queryArgs []string
|
||||||
|
|
||||||
|
for i := 0; i < len(args); i++ {
|
||||||
|
arg := args[i]
|
||||||
|
wantsValue := arg == "--root" || arg == "--repo" || arg == "-n" || arg == "--hop"
|
||||||
|
if wantsValue && i+1 >= len(args) {
|
||||||
|
return opt, fmt.Errorf("%s needs a value", arg)
|
||||||
|
}
|
||||||
|
switch arg {
|
||||||
|
case "--root":
|
||||||
|
i++
|
||||||
|
opt.Root = args[i]
|
||||||
|
if opt.Root != "facts" && opt.Root != "info" {
|
||||||
|
return opt, fmt.Errorf("--root must be facts or info, got %q", opt.Root)
|
||||||
|
}
|
||||||
|
case "--repo":
|
||||||
|
i++
|
||||||
|
opt.Repo = args[i]
|
||||||
|
case "-n":
|
||||||
|
i++
|
||||||
|
n, err := strconv.Atoi(args[i])
|
||||||
|
if err != nil || n < 1 {
|
||||||
|
return opt, fmt.Errorf("-n must be a positive integer, got %q", args[i])
|
||||||
|
}
|
||||||
|
opt.Limit = n
|
||||||
|
case "--hop":
|
||||||
|
return opt, fmt.Errorf("--hop is not implemented yet (needs File/FROM_FILE edges)")
|
||||||
|
case "--json":
|
||||||
|
opt.JSONOut = true
|
||||||
|
case "--no-web":
|
||||||
|
opt.NoWeb = true
|
||||||
|
case "--list-model":
|
||||||
|
opt.ListModel = true
|
||||||
|
default:
|
||||||
|
if strings.HasPrefix(arg, "-") {
|
||||||
|
return opt, fmt.Errorf("unknown flag %q", arg)
|
||||||
|
}
|
||||||
|
queryArgs = append(queryArgs, arg)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
opt.Query = strings.TrimSpace(strings.Join(queryArgs, " "))
|
||||||
|
if opt.Query == "" && !opt.ListModel {
|
||||||
|
return opt, fmt.Errorf("no query given")
|
||||||
|
}
|
||||||
|
return opt, nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,43 @@
|
|||||||
|
package rank
|
||||||
|
|
||||||
|
// SecondSource is the web-search block on a deduction answer.
|
||||||
|
// Kept apart from graph hits so "ours" and "not ours" stay visible.
|
||||||
|
type SecondSource struct {
|
||||||
|
Status string `json:"status"`
|
||||||
|
Note string `json:"note,omitempty"`
|
||||||
|
Cached bool `json:"cached,omitempty"`
|
||||||
|
Results []SecondSourceHit `json:"results,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
type SecondSourceHit struct {
|
||||||
|
Rank int `json:"rank"`
|
||||||
|
Title string `json:"title"`
|
||||||
|
URL string `json:"url"`
|
||||||
|
Snippet string `json:"snippet"`
|
||||||
|
Engine string `json:"engine"`
|
||||||
|
}
|
||||||
|
|
||||||
|
type WebFn func(query string) SecondSource
|
||||||
|
|
||||||
|
// ShouldEscalate is true when the default deduction path has no facts hit.
|
||||||
|
// `--root facts|info` is a single-root ask: do not mix in the web.
|
||||||
|
func ShouldEscalate(hits []Hit, rootFilter string) bool {
|
||||||
|
if rootFilter != "" {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
for _, h := range hits {
|
||||||
|
if h.Root == "facts" {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
|
||||||
|
// Deduce returns the second-source block, or nil when web must not run.
|
||||||
|
func Deduce(hits []Hit, query, rootFilter string, noWeb bool, web WebFn) *SecondSource {
|
||||||
|
if noWeb || web == nil || !ShouldEscalate(hits, rootFilter) {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
out := web(query)
|
||||||
|
return &out
|
||||||
|
}
|
||||||
@@ -0,0 +1,78 @@
|
|||||||
|
package rank
|
||||||
|
|
||||||
|
import (
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestShouldEscalateWhenNoFacts(t *testing.T) {
|
||||||
|
if !ShouldEscalate(nil, "") {
|
||||||
|
t.Fatal("empty local graph must escalate")
|
||||||
|
}
|
||||||
|
if !ShouldEscalate([]Hit{h("i", "info", "docs/a.md")}, "") {
|
||||||
|
t.Fatal("info-only must escalate (not confirmed)")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestShouldNotEscalateWhenFactsConfirm(t *testing.T) {
|
||||||
|
hits := []Hit{h("f", "facts", "docker ps x compose"), h("i", "info", "docs/a.md")}
|
||||||
|
if ShouldEscalate(hits, "") {
|
||||||
|
t.Fatal("facts hit is already confirmed; do not mix web")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestShouldNotEscalateWhenRootFilterSet(t *testing.T) {
|
||||||
|
if ShouldEscalate(nil, "facts") {
|
||||||
|
t.Fatal("--root facts must stay local")
|
||||||
|
}
|
||||||
|
if ShouldEscalate([]Hit{h("i", "info", "x")}, "info") {
|
||||||
|
t.Fatal("--root info must stay local")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestDeduceCallsWebOnlyWhenEscalating(t *testing.T) {
|
||||||
|
called := 0
|
||||||
|
web := func(q string) SecondSource {
|
||||||
|
called++
|
||||||
|
if q != "LadybugDB" {
|
||||||
|
t.Fatalf("query = %q", q)
|
||||||
|
}
|
||||||
|
return SecondSource{Status: "ok", Results: []SecondSourceHit{{Title: "t", URL: "http://example.com"}}}
|
||||||
|
}
|
||||||
|
got := Deduce([]Hit{h("i", "info", "x")}, "LadybugDB", "", false, web)
|
||||||
|
if called != 1 || got == nil || got.Status != "ok" {
|
||||||
|
t.Fatalf("got %+v called=%d", got, called)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestDeduceNilWhenFactsOrNoWeb(t *testing.T) {
|
||||||
|
web := func(string) SecondSource {
|
||||||
|
t.Fatal("web must not run")
|
||||||
|
return SecondSource{}
|
||||||
|
}
|
||||||
|
if Deduce([]Hit{h("f", "facts", "x")}, "q", "", false, web) != nil {
|
||||||
|
t.Fatal("facts")
|
||||||
|
}
|
||||||
|
if Deduce([]Hit{h("i", "info", "x")}, "q", "", true, web) != nil {
|
||||||
|
t.Fatal("--no-web")
|
||||||
|
}
|
||||||
|
if Deduce(nil, "q", "facts", false, web) != nil {
|
||||||
|
t.Fatal("--root facts")
|
||||||
|
}
|
||||||
|
if Deduce(nil, "q", "", false, nil) != nil {
|
||||||
|
t.Fatal("nil web fn")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseNoWeb(t *testing.T) {
|
||||||
|
opt, err := ParseArgs([]string{"query", "--no-web", "--json"})
|
||||||
|
if err != nil || !opt.NoWeb || !opt.JSONOut || opt.Query != "query" {
|
||||||
|
t.Fatalf("got %+v err=%v", opt, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUsageNamesNoWeb(t *testing.T) {
|
||||||
|
if !strings.Contains(Usage, "--no-web") {
|
||||||
|
t.Fatalf("usage must name --no-web, got:\n%s", Usage)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
package rank
|
||||||
|
|
||||||
|
// BM25 ranks best-first, so the top hits are the *highest* scores; cosine
|
||||||
|
// distance ranks best-first ascending. Both mirror kblib.py.
|
||||||
|
const FTSStmt = "CALL QUERY_FTS_INDEX('Leaf', 'id', $q) " +
|
||||||
|
"RETURN node.id, node.text, node.root, node.source, score ORDER BY score DESC LIMIT $n"
|
||||||
|
|
||||||
|
const VecStmt = "CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) " +
|
||||||
|
"RETURN node.id, node.text, node.root, node.source, distance ORDER BY distance LIMIT $n"
|
||||||
@@ -0,0 +1,100 @@
|
|||||||
|
// Package rank is the cgo-free ranking and CLI parsing for brain search.
|
||||||
|
// CI can `go test ./rank` without the native ladybug library.
|
||||||
|
package rank
|
||||||
|
|
||||||
|
import (
|
||||||
|
"sort"
|
||||||
|
"strings"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Hit is one search result, mirroring the python script's dict shape.
|
||||||
|
type Hit struct {
|
||||||
|
ID string `json:"id"`
|
||||||
|
Text string `json:"text"`
|
||||||
|
Root string `json:"root"`
|
||||||
|
Source string `json:"-"`
|
||||||
|
Score float64 `json:"score"`
|
||||||
|
Snippet string `json:"snippet,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// rrfK dampens the contribution of low ranks; same constant as kblib.py.
|
||||||
|
const rrfK = 60
|
||||||
|
|
||||||
|
// RankAndFilter fuses the two hit lists, applies --root/--repo, then cuts to
|
||||||
|
// limit. Cutting first dropped every matching leaf ranked below the cut, so
|
||||||
|
// `--root facts` came back empty whenever info leafs filled the top N.
|
||||||
|
// limit <= 0 keeps everything.
|
||||||
|
func RankAndFilter(fts, vec []Hit, root, repo string, limit int) []Hit {
|
||||||
|
out := Hybrid(fts, vec, 0)
|
||||||
|
if root != "" {
|
||||||
|
out = FilterRoot(out, root)
|
||||||
|
}
|
||||||
|
if repo != "" {
|
||||||
|
out = FilterRepo(out, repo)
|
||||||
|
}
|
||||||
|
if limit > 0 && len(out) > limit {
|
||||||
|
out = out[:limit]
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
// Hybrid merges FTS and vector hits by reciprocal rank fusion.
|
||||||
|
// limit <= 0 returns the full fused list.
|
||||||
|
func Hybrid(fts, vec []Hit, limit int) []Hit {
|
||||||
|
byID := make(map[string]Hit, len(fts)+len(vec))
|
||||||
|
rrf := make(map[string]float64, len(fts)+len(vec))
|
||||||
|
|
||||||
|
for i, h := range fts {
|
||||||
|
byID[h.ID] = h
|
||||||
|
rrf[h.ID] += 1.0 / (rrfK + float64(i+1))
|
||||||
|
}
|
||||||
|
for i, h := range vec {
|
||||||
|
if existing, ok := byID[h.ID]; !ok {
|
||||||
|
byID[h.ID] = h
|
||||||
|
} else if existing.Score == 0 {
|
||||||
|
existing.Score = h.Score
|
||||||
|
byID[h.ID] = existing
|
||||||
|
}
|
||||||
|
rrf[h.ID] += 1.0 / (rrfK + float64(i+1))
|
||||||
|
}
|
||||||
|
|
||||||
|
ids := make([]string, 0, len(rrf))
|
||||||
|
for id := range rrf {
|
||||||
|
ids = append(ids, id)
|
||||||
|
}
|
||||||
|
sort.Slice(ids, func(i, j int) bool {
|
||||||
|
if rrf[ids[i]] != rrf[ids[j]] {
|
||||||
|
return rrf[ids[i]] > rrf[ids[j]]
|
||||||
|
}
|
||||||
|
return ids[i] < ids[j]
|
||||||
|
})
|
||||||
|
if limit > 0 && len(ids) > limit {
|
||||||
|
ids = ids[:limit]
|
||||||
|
}
|
||||||
|
|
||||||
|
out := make([]Hit, 0, len(ids))
|
||||||
|
for _, id := range ids {
|
||||||
|
out = append(out, byID[id])
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
func FilterRoot(hits []Hit, root string) []Hit {
|
||||||
|
var out []Hit
|
||||||
|
for _, h := range hits {
|
||||||
|
if h.Root == root {
|
||||||
|
out = append(out, h)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
func FilterRepo(hits []Hit, repo string) []Hit {
|
||||||
|
var out []Hit
|
||||||
|
for _, h := range hits {
|
||||||
|
if strings.Contains(h.Source, repo) {
|
||||||
|
out = append(out, h)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
@@ -0,0 +1,149 @@
|
|||||||
|
// Unit tests for ranking/filtering and CLI parsing (no db, no model, offline).
|
||||||
|
package rank
|
||||||
|
|
||||||
|
import (
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
func h(id, root, source string) Hit {
|
||||||
|
return Hit{ID: id, Text: id, Root: root, Source: source}
|
||||||
|
}
|
||||||
|
|
||||||
|
func ids(hits []Hit) []string {
|
||||||
|
out := make([]string, len(hits))
|
||||||
|
for i, hit := range hits {
|
||||||
|
out[i] = hit.ID
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
func eq(t *testing.T, got []Hit, want ...string) {
|
||||||
|
t.Helper()
|
||||||
|
g := ids(got)
|
||||||
|
if len(g) != len(want) {
|
||||||
|
t.Fatalf("got %v, want %v", g, want)
|
||||||
|
}
|
||||||
|
for i := range want {
|
||||||
|
if g[i] != want[i] {
|
||||||
|
t.Fatalf("got %v, want %v", g, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// A facts leaf that ranks below the limit in the unfiltered list must still
|
||||||
|
// be returned for --root facts. Filtering after truncation loses it.
|
||||||
|
func TestRankAndFilterFiltersBeforeLimit(t *testing.T) {
|
||||||
|
fts := []Hit{
|
||||||
|
h("i1", "info", "docs/a.md"),
|
||||||
|
h("i2", "info", "docs/b.md"),
|
||||||
|
h("i3", "info", "docs/c.md"),
|
||||||
|
h("f1", "facts", "docker ps x compose"),
|
||||||
|
}
|
||||||
|
eq(t, RankAndFilter(fts, nil, "facts", "", 2), "f1")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRankAndFilterRepoFiltersBeforeLimit(t *testing.T) {
|
||||||
|
fts := []Hit{
|
||||||
|
h("a", "info", "eSlider/2dph:README.md"),
|
||||||
|
h("b", "info", "eSlider/2dph:PLAN.md"),
|
||||||
|
h("c", "info", "eSlider/ops:compose.yaml"),
|
||||||
|
}
|
||||||
|
eq(t, RankAndFilter(fts, nil, "", "ops", 2), "c")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRankAndFilterTruncatesToLimit(t *testing.T) {
|
||||||
|
fts := []Hit{h("a", "info", "x"), h("b", "info", "x"), h("c", "info", "x")}
|
||||||
|
eq(t, RankAndFilter(fts, nil, "", "", 2), "a", "b")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRankAndFilterLimitZeroKeepsAll(t *testing.T) {
|
||||||
|
fts := []Hit{h("a", "info", "x"), h("b", "info", "x")}
|
||||||
|
eq(t, RankAndFilter(fts, nil, "", "", 0), "a", "b")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestHybridFusesBothRetrievers(t *testing.T) {
|
||||||
|
fts := []Hit{h("only-fts", "info", "x"), h("both", "info", "x")}
|
||||||
|
vec := []Hit{h("only-vec", "info", "x"), h("both", "info", "x")}
|
||||||
|
eq(t, Hybrid(fts, vec, 0), "both", "only-fts", "only-vec")
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestHybridTiesAreDeterministic(t *testing.T) {
|
||||||
|
fts := []Hit{h("b", "info", "x"), h("a", "info", "x")}
|
||||||
|
first := ids(Hybrid(fts, nil, 0))
|
||||||
|
for i := 0; i < 50; i++ {
|
||||||
|
got := ids(Hybrid(fts, nil, 0))
|
||||||
|
for j := range first {
|
||||||
|
if got[j] != first[j] {
|
||||||
|
t.Fatalf("unstable order: %v then %v", first, got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestHybridKeepsVectorScoreForSharedHit(t *testing.T) {
|
||||||
|
fts := []Hit{{ID: "x", Root: "info", Score: 0}}
|
||||||
|
vec := []Hit{{ID: "x", Root: "info", Score: 0.87}}
|
||||||
|
got := Hybrid(fts, vec, 0)
|
||||||
|
if len(got) != 1 || got[0].Score != 0.87 {
|
||||||
|
t.Fatalf("got %+v, want score 0.87", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// The old parser dropped unknown flags and appended their arguments to the
|
||||||
|
// query, so `search "q" --hop 1` searched for "q 1". --hop is not implemented
|
||||||
|
// here (needs File edges); it must still fail closed instead of changing q.
|
||||||
|
func TestParseHopIsNotSwallowedIntoTheQuery(t *testing.T) {
|
||||||
|
_, err := ParseArgs([]string{"what runs on arc-2", "--hop", "1"})
|
||||||
|
if err == nil {
|
||||||
|
t.Fatal("expected --hop to error (not implemented), not be swallowed")
|
||||||
|
}
|
||||||
|
if !strings.Contains(err.Error(), "--hop") {
|
||||||
|
t.Fatalf("error should name --hop, got %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseRejectsUnknownFlags(t *testing.T) {
|
||||||
|
if _, err := ParseArgs([]string{"query", "--nope"}); err == nil {
|
||||||
|
t.Fatal("unknown flag accepted")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseRejectsBadValues(t *testing.T) {
|
||||||
|
for _, args := range [][]string{
|
||||||
|
{"q", "-n", "zero"},
|
||||||
|
{"q", "-n", "0"},
|
||||||
|
{"q", "--root", "nonsense"},
|
||||||
|
{"q", "--hop"},
|
||||||
|
{"--json"},
|
||||||
|
} {
|
||||||
|
if _, err := ParseArgs(args); err == nil {
|
||||||
|
t.Errorf("accepted %v", args)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseDefaults(t *testing.T) {
|
||||||
|
opt, err := ParseArgs([]string{"two", "words", "--json"})
|
||||||
|
if err != nil || opt.Query != "two words" || opt.Limit != 20 || !opt.JSONOut {
|
||||||
|
t.Fatalf("got %+v err=%v", opt, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestListModelNeedsNoQuery(t *testing.T) {
|
||||||
|
if _, err := ParseArgs([]string{"--list-model"}); err != nil {
|
||||||
|
t.Fatalf("unexpected error: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestUsageNamesBrainSearch(t *testing.T) {
|
||||||
|
if !strings.Contains(Usage, "bin/brain/search.go") {
|
||||||
|
t.Fatalf("usage must name bin/brain/search.go, got:\n%s", Usage)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestFTSQueryOrdersByScoreDescending(t *testing.T) {
|
||||||
|
if !strings.Contains(FTSStmt, "ORDER BY score DESC") {
|
||||||
|
t.Fatalf("FTS query must order by score DESC, got:\n%s", FTSStmt)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,5 +1,6 @@
|
|||||||
// Hybrid FTS + vector search implementation, plus daemon client/server.
|
//go:build cgo && system_ladybug
|
||||||
package main
|
|
||||||
|
package brain
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"bytes"
|
"bytes"
|
||||||
@@ -13,12 +14,12 @@ import (
|
|||||||
"os"
|
"os"
|
||||||
"os/exec"
|
"os/exec"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"sort"
|
|
||||||
"strconv"
|
"strconv"
|
||||||
"strings"
|
"syscall"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
lbug "github.com/LadybugDB/go-ladybug"
|
lbug "github.com/LadybugDB/go-ladybug"
|
||||||
|
"github.com/eSlider/2dph/internal/brain/rank"
|
||||||
)
|
)
|
||||||
|
|
||||||
const defaultPort = 17830
|
const defaultPort = 17830
|
||||||
@@ -26,45 +27,15 @@ const daemonPath = "/embed"
|
|||||||
const healthPath = "/health"
|
const healthPath = "/health"
|
||||||
|
|
||||||
func runSearch(args []string) int {
|
func runSearch(args []string) int {
|
||||||
// Manual flag parsing to allow flags after query (like Python argparse)
|
opt, err := rank.ParseArgs(args)
|
||||||
root := ""
|
if err != nil {
|
||||||
repo := ""
|
fmt.Fprintf(os.Stderr, "brain/search: %v\n%s\n", err, rank.Usage)
|
||||||
limit := 20
|
return 2
|
||||||
jsonOut := false
|
|
||||||
listModel := false
|
|
||||||
|
|
||||||
var queryArgs []string
|
|
||||||
for i := 0; i < len(args); i++ {
|
|
||||||
switch args[i] {
|
|
||||||
case "--root":
|
|
||||||
if i+1 < len(args) {
|
|
||||||
root = args[i+1]
|
|
||||||
i++
|
|
||||||
}
|
|
||||||
case "--repo":
|
|
||||||
if i+1 < len(args) {
|
|
||||||
repo = args[i+1]
|
|
||||||
i++
|
|
||||||
}
|
|
||||||
case "-n":
|
|
||||||
if i+1 < len(args) {
|
|
||||||
if n, err := strconv.Atoi(args[i+1]); err == nil {
|
|
||||||
limit = n
|
|
||||||
}
|
|
||||||
i++
|
|
||||||
}
|
|
||||||
case "--json":
|
|
||||||
jsonOut = true
|
|
||||||
case "--list-model":
|
|
||||||
listModel = true
|
|
||||||
default:
|
|
||||||
if !strings.HasPrefix(args[i], "-") {
|
|
||||||
queryArgs = append(queryArgs, args[i])
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
root, repo, limit, query := opt.Root, opt.Repo, opt.Limit, opt.Query
|
||||||
|
jsonOut := opt.JSONOut
|
||||||
|
|
||||||
if listModel {
|
if opt.ListModel {
|
||||||
dir, err := modelDir()
|
dir, err := modelDir()
|
||||||
if err != nil {
|
if err != nil {
|
||||||
fmt.Fprintln(os.Stderr, err)
|
fmt.Fprintln(os.Stderr, err)
|
||||||
@@ -74,47 +45,19 @@ func runSearch(args []string) int {
|
|||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
|
|
||||||
query := strings.TrimSpace(strings.Join(queryArgs, " "))
|
|
||||||
if query == "" {
|
|
||||||
fmt.Fprintln(os.Stderr, "usage: kbsearch \"query\" [--root facts|info] [--repo REPO] [-n N] [--json]")
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
|
|
||||||
if err := openBrain(); err != nil {
|
if err := openBrain(); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
|
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
}
|
}
|
||||||
defer closeBrain()
|
defer closeBrain()
|
||||||
|
|
||||||
emb, err := embedQuery(query)
|
hits, err := searchHits(query, root, repo, limit)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "embed: %v\n", err)
|
fmt.Fprintf(os.Stderr, "search: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
}
|
}
|
||||||
|
|
||||||
fts, err := queryFTS(query, limit*3)
|
results := hits
|
||||||
if err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "fts: %v\n", err)
|
|
||||||
return 1
|
|
||||||
}
|
|
||||||
|
|
||||||
var vec []Hit
|
|
||||||
if vec, err = queryVector(emb, limit*3); err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "vec: %v\n", err)
|
|
||||||
}
|
|
||||||
|
|
||||||
results := hybrid(fts, vec, limit)
|
|
||||||
|
|
||||||
if root != "" {
|
|
||||||
results = filterRoot(results, root)
|
|
||||||
}
|
|
||||||
if repo != "" {
|
|
||||||
results = filterRepo(results, repo)
|
|
||||||
}
|
|
||||||
if len(results) > limit {
|
|
||||||
results = results[:limit]
|
|
||||||
}
|
|
||||||
|
|
||||||
for i := range results {
|
for i := range results {
|
||||||
if results[i].Text != "" {
|
if results[i].Text != "" {
|
||||||
runes := []rune(results[i].Text)
|
runes := []rune(results[i].Text)
|
||||||
@@ -125,23 +68,46 @@ func runSearch(args []string) int {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
webOut := rank.Deduce(results, query, root, opt.NoWeb, func(q string) rank.SecondSource {
|
||||||
|
return lookupWeb(context.Background(), q)
|
||||||
|
})
|
||||||
|
|
||||||
out := Dict{
|
out := Dict{
|
||||||
{"query", query},
|
{"query", query},
|
||||||
{"root_filter", root},
|
{"root_filter", root},
|
||||||
{"count", len(results)},
|
{"count", len(results)},
|
||||||
{"results", resultsToDicts(results)},
|
{"results", resultsToDicts(results)},
|
||||||
}
|
}
|
||||||
|
if webOut != nil {
|
||||||
|
out = append(out, KV{"web", secondToDict(*webOut)})
|
||||||
|
}
|
||||||
|
|
||||||
if jsonOut {
|
if jsonOut {
|
||||||
enc := json.NewEncoder(os.Stdout)
|
enc := json.NewEncoder(os.Stdout)
|
||||||
enc.SetIndent("", " ")
|
enc.SetIndent("", " ")
|
||||||
enc.SetEscapeHTML(false)
|
enc.SetEscapeHTML(false)
|
||||||
return b2i(enc.Encode(toJSONOut(results, query, root)))
|
return b2i(enc.Encode(toJSONOut(results, query, root, webOut)))
|
||||||
}
|
}
|
||||||
fmt.Print(toYAML(out, 0))
|
fmt.Print(toYAML(out, 0))
|
||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func searchHits(query, root, repo string, limit int) ([]Hit, error) {
|
||||||
|
emb, err := embedQuery(query)
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("embed: %w", err)
|
||||||
|
}
|
||||||
|
fts, err := queryFTS(query, limit*3)
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("fts: %w", err)
|
||||||
|
}
|
||||||
|
var vec []Hit
|
||||||
|
if vec, err = queryVector(emb, limit*3); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "vec: %v\n", err)
|
||||||
|
}
|
||||||
|
return rank.RankAndFilter(fts, vec, root, repo, limit), nil
|
||||||
|
}
|
||||||
|
|
||||||
func b2i(err error) int {
|
func b2i(err error) int {
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return 1
|
return 1
|
||||||
@@ -150,10 +116,7 @@ func b2i(err error) int {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func queryFTS(text string, limit int) ([]Hit, error) {
|
func queryFTS(text string, limit int) ([]Hit, error) {
|
||||||
stmt, err := conn.Prepare(
|
stmt, err := conn.Prepare(rank.FTSStmt)
|
||||||
"CALL QUERY_FTS_INDEX('Leaf', 'id', $q) " +
|
|
||||||
"RETURN node.id, node.text, node.root, node.source, score ORDER BY score LIMIT $n",
|
|
||||||
)
|
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
@@ -170,10 +133,7 @@ func queryVector(emb []float64, limit int) ([]Hit, error) {
|
|||||||
for i, v := range emb {
|
for i, v := range emb {
|
||||||
embList[i] = v
|
embList[i] = v
|
||||||
}
|
}
|
||||||
stmt, err := conn.Prepare(
|
stmt, err := conn.Prepare(rank.VecStmt)
|
||||||
"CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) " +
|
|
||||||
"RETURN node.id, node.text, node.root, node.source, distance ORDER BY distance LIMIT $n",
|
|
||||||
)
|
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
@@ -215,10 +175,11 @@ func rowsToHits(res *lbug.QueryResult) ([]Hit, error) {
|
|||||||
|
|
||||||
// JSON output types
|
// JSON output types
|
||||||
type jsonOut struct {
|
type jsonOut struct {
|
||||||
Query string `json:"query"`
|
Query string `json:"query"`
|
||||||
RootFilter string `json:"root_filter"`
|
RootFilter string `json:"root_filter"`
|
||||||
Count int `json:"count"`
|
Count int `json:"count"`
|
||||||
Results []jsonHit `json:"results"`
|
Results []jsonHit `json:"results"`
|
||||||
|
Web *rank.SecondSource `json:"web,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
type jsonHit struct {
|
type jsonHit struct {
|
||||||
@@ -229,7 +190,7 @@ type jsonHit struct {
|
|||||||
Snippet string `json:"snippet,omitempty"`
|
Snippet string `json:"snippet,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
func toJSONOut(hits []Hit, query, rootFilter string) *jsonOut {
|
func toJSONOut(hits []Hit, query, rootFilter string, web *rank.SecondSource) *jsonOut {
|
||||||
out := make([]jsonHit, len(hits))
|
out := make([]jsonHit, len(hits))
|
||||||
for i, h := range hits {
|
for i, h := range hits {
|
||||||
out[i] = jsonHit{
|
out[i] = jsonHit{
|
||||||
@@ -245,73 +206,10 @@ func toJSONOut(hits []Hit, query, rootFilter string) *jsonOut {
|
|||||||
RootFilter: rootFilter,
|
RootFilter: rootFilter,
|
||||||
Count: len(hits),
|
Count: len(hits),
|
||||||
Results: out,
|
Results: out,
|
||||||
|
Web: web,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
func hybrid(fts, vec []Hit, limit int) []Hit {
|
|
||||||
byID := make(map[string]Hit)
|
|
||||||
rrf := make(map[string]float64)
|
|
||||||
|
|
||||||
for rank, h := range fts {
|
|
||||||
byID[h.ID] = h
|
|
||||||
rrf[h.ID] += 1.0 / (60 + float64(rank+1))
|
|
||||||
}
|
|
||||||
for rank, h := range vec {
|
|
||||||
if _, ok := byID[h.ID]; !ok {
|
|
||||||
byID[h.ID] = h
|
|
||||||
} else {
|
|
||||||
existing := byID[h.ID]
|
|
||||||
if existing.Score == 0 {
|
|
||||||
existing.Score = h.Score
|
|
||||||
byID[h.ID] = existing
|
|
||||||
}
|
|
||||||
}
|
|
||||||
rrf[h.ID] += 1.0 / (60 + float64(rank+1))
|
|
||||||
}
|
|
||||||
|
|
||||||
type scored struct {
|
|
||||||
id string
|
|
||||||
rrf float64
|
|
||||||
}
|
|
||||||
var scoredList []scored
|
|
||||||
for id, v := range rrf {
|
|
||||||
scoredList = append(scoredList, scored{id, v})
|
|
||||||
}
|
|
||||||
sort.Slice(scoredList, func(i, j int) bool {
|
|
||||||
return scoredList[i].rrf > scoredList[j].rrf
|
|
||||||
})
|
|
||||||
|
|
||||||
var out []Hit
|
|
||||||
for i, s := range scoredList {
|
|
||||||
if i >= limit {
|
|
||||||
break
|
|
||||||
}
|
|
||||||
h := byID[s.id]
|
|
||||||
out = append(out, h)
|
|
||||||
}
|
|
||||||
return out
|
|
||||||
}
|
|
||||||
|
|
||||||
func filterRoot(hits []Hit, root string) []Hit {
|
|
||||||
var out []Hit
|
|
||||||
for _, h := range hits {
|
|
||||||
if h.Root == root {
|
|
||||||
out = append(out, h)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return out
|
|
||||||
}
|
|
||||||
|
|
||||||
func filterRepo(hits []Hit, repo string) []Hit {
|
|
||||||
var out []Hit
|
|
||||||
for _, h := range hits {
|
|
||||||
if strings.Contains(h.Source, repo) {
|
|
||||||
out = append(out, h)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return out
|
|
||||||
}
|
|
||||||
|
|
||||||
func resultsToDicts(hits []Hit) []any {
|
func resultsToDicts(hits []Hit) []any {
|
||||||
out := make([]any, len(hits))
|
out := make([]any, len(hits))
|
||||||
for i, h := range hits {
|
for i, h := range hits {
|
||||||
@@ -363,7 +261,7 @@ func serve(port int) error {
|
|||||||
})
|
})
|
||||||
|
|
||||||
addr := fmt.Sprintf("127.0.0.1:%d", port)
|
addr := fmt.Sprintf("127.0.0.1:%d", port)
|
||||||
log.Printf("kbsearch daemon listening on %s", addr)
|
log.Printf("brain search daemon listening on %s", addr)
|
||||||
return http.ListenAndServe(addr, mux)
|
return http.ListenAndServe(addr, mux)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -382,10 +280,16 @@ func embedQuery(text string) ([]float64, error) {
|
|||||||
port = p
|
port = p
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
emb, err := tryDaemon(text, port)
|
if emb, err := tryDaemon(text, port); err == nil {
|
||||||
if err == nil {
|
|
||||||
return emb, nil
|
return emb, nil
|
||||||
}
|
}
|
||||||
|
if os.Getenv("KBSEARCH_NO_DAEMON") == "" {
|
||||||
|
if err := ensureDaemon(port); err == nil {
|
||||||
|
if emb, err := tryDaemon(text, port); err == nil {
|
||||||
|
return emb, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
model, err := loadModel()
|
model, err := loadModel()
|
||||||
if err != nil {
|
if err != nil {
|
||||||
@@ -447,9 +351,13 @@ func ensureDaemon(port int) error {
|
|||||||
cmd.Dir, _ = filepath.Split(self)
|
cmd.Dir, _ = filepath.Split(self)
|
||||||
cmd.Stdout = nil
|
cmd.Stdout = nil
|
||||||
cmd.Stderr = nil
|
cmd.Stderr = nil
|
||||||
|
cmd.SysProcAttr = &syscall.SysProcAttr{Setsid: true}
|
||||||
if err := cmd.Start(); err != nil {
|
if err := cmd.Start(); err != nil {
|
||||||
return err
|
return err
|
||||||
}
|
}
|
||||||
|
if cmd.Process != nil {
|
||||||
|
_ = cmd.Process.Release()
|
||||||
|
}
|
||||||
|
|
||||||
for i := 0; i < 40; i++ {
|
for i := 0; i < 40; i++ {
|
||||||
time.Sleep(250 * time.Millisecond)
|
time.Sleep(250 * time.Millisecond)
|
||||||
@@ -462,4 +370,22 @@ func ensureDaemon(port int) error {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
return fmt.Errorf("daemon failed to start on port %d", port)
|
return fmt.Errorf("daemon failed to start on port %d", port)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Main is the bin/brain/search.go entry: search, serve, or --list-model.
|
||||||
|
func Main(args []string) int {
|
||||||
|
if len(args) > 0 && args[0] == "serve" {
|
||||||
|
port := defaultPort
|
||||||
|
if len(args) > 1 {
|
||||||
|
if p, err := strconv.Atoi(args[1]); err == nil {
|
||||||
|
port = p
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if err := serve(port); err != nil {
|
||||||
|
log.Printf("brain/search serve: %v", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
return runSearch(args)
|
||||||
|
}
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
package brain
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/brain/rank"
|
||||||
|
)
|
||||||
|
|
||||||
|
func eps() string { return os.Getenv("KBTEST_EPS") }
|
||||||
|
|
||||||
|
// Hit is the search hit type; ranking lives in package rank so CI can test
|
||||||
|
// it without the native ladybug library.
|
||||||
|
type Hit = rank.Hit
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
package brain
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/brain/rank"
|
||||||
|
"github.com/eSlider/2dph/internal/websearch"
|
||||||
|
)
|
||||||
|
|
||||||
|
func lookupWeb(ctx context.Context, query string) rank.SecondSource {
|
||||||
|
o := websearch.Lookup(ctx, query, websearch.LookupOpt{Limit: 5})
|
||||||
|
return toSecond(o)
|
||||||
|
}
|
||||||
|
|
||||||
|
func toSecond(o websearch.Output) rank.SecondSource {
|
||||||
|
hits := make([]rank.SecondSourceHit, 0, len(o.Results))
|
||||||
|
for _, h := range o.Results {
|
||||||
|
hits = append(hits, rank.SecondSourceHit{
|
||||||
|
Rank: h.Rank,
|
||||||
|
Title: h.Title,
|
||||||
|
URL: h.URL,
|
||||||
|
Snippet: h.Snippet,
|
||||||
|
Engine: h.Engine,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
return rank.SecondSource{
|
||||||
|
Status: o.Status,
|
||||||
|
Note: o.Note,
|
||||||
|
Cached: o.Cached,
|
||||||
|
Results: hits,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func secondToDict(w rank.SecondSource) Dict {
|
||||||
|
d := Dict{
|
||||||
|
{"status", w.Status},
|
||||||
|
}
|
||||||
|
if w.Note != "" {
|
||||||
|
d = append(d, KV{"note", w.Note})
|
||||||
|
}
|
||||||
|
if w.Cached {
|
||||||
|
d = append(d, KV{"cached", true})
|
||||||
|
}
|
||||||
|
rows := make([]any, 0, len(w.Results))
|
||||||
|
for _, h := range w.Results {
|
||||||
|
rows = append(rows, Dict{
|
||||||
|
{"rank", h.Rank},
|
||||||
|
{"title", h.Title},
|
||||||
|
{"url", h.URL},
|
||||||
|
{"snippet", h.Snippet},
|
||||||
|
{"engine", h.Engine},
|
||||||
|
})
|
||||||
|
}
|
||||||
|
d = append(d, KV{"results", rows})
|
||||||
|
return d
|
||||||
|
}
|
||||||
@@ -1,5 +1,5 @@
|
|||||||
// YAML emitter ported from bin/kb/yamlout.py — preserves insertion order.
|
// YAML emitter ported from bin/kb/yamlout.py — preserves insertion order.
|
||||||
package main
|
package brain
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"fmt"
|
"fmt"
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
package main
|
package chats
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"bytes"
|
"bytes"
|
||||||
@@ -24,7 +24,7 @@ type ooContact struct {
|
|||||||
} `json:"commonData"`
|
} `json:"commonData"`
|
||||||
}
|
}
|
||||||
|
|
||||||
func runApply(args []string) int {
|
func RunApply(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats apply", flag.ContinueOnError)
|
fs := flag.NewFlagSet("chats apply", flag.ContinueOnError)
|
||||||
dryRun := fs.Bool("dry-run", false, "show what would be done without writing")
|
dryRun := fs.Bool("dry-run", false, "show what would be done without writing")
|
||||||
help := fs.Bool("help", false, "")
|
help := fs.Bool("help", false, "")
|
||||||
@@ -176,7 +176,7 @@ func runApply(args []string) int {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func loadFacts() ([]ExtractedFact, error) {
|
func loadFacts() ([]ExtractedFact, error) {
|
||||||
factsPath := filepath.Join(chatsDir(), "facts", "chat-facts.json")
|
factsPath := filepath.Join(Dir(), "facts", "chat-facts.json")
|
||||||
data, err := os.ReadFile(factsPath)
|
data, err := os.ReadFile(factsPath)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
if os.IsNotExist(err) {
|
if os.IsNotExist(err) {
|
||||||
@@ -3,7 +3,7 @@
|
|||||||
// These are integration tests using real data and real Telegram API (when
|
// These are integration tests using real data and real Telegram API (when
|
||||||
// credentials are available). They follow the TDD workflow pattern:
|
// credentials are available). They follow the TDD workflow pattern:
|
||||||
// sync → import → facts → verify.
|
// sync → import → facts → verify.
|
||||||
package main
|
package chats
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"encoding/json"
|
"encoding/json"
|
||||||
@@ -48,7 +48,7 @@ func TestChatsImport(t *testing.T) {
|
|||||||
t.Cleanup(func() { os.Chdir(cwd) })
|
t.Cleanup(func() { os.Chdir(cwd) })
|
||||||
t.Setenv("KB_ROOT", dir)
|
t.Setenv("KB_ROOT", dir)
|
||||||
|
|
||||||
exitCode := runImport([]string{})
|
exitCode := RunImport([]string{})
|
||||||
if exitCode != 0 {
|
if exitCode != 0 {
|
||||||
t.Fatalf("import exit code %d", exitCode)
|
t.Fatalf("import exit code %d", exitCode)
|
||||||
}
|
}
|
||||||
@@ -140,7 +140,7 @@ func TestChatsImportEmpty(t *testing.T) {
|
|||||||
t.Cleanup(func() { os.Chdir(cwd) })
|
t.Cleanup(func() { os.Chdir(cwd) })
|
||||||
t.Setenv("KB_ROOT", dir)
|
t.Setenv("KB_ROOT", dir)
|
||||||
|
|
||||||
exitCode := runImport([]string{})
|
exitCode := RunImport([]string{})
|
||||||
if exitCode == 0 {
|
if exitCode == 0 {
|
||||||
t.Fatal("expected non-zero exit for empty data dir")
|
t.Fatal("expected non-zero exit for empty data dir")
|
||||||
}
|
}
|
||||||
@@ -171,7 +171,7 @@ func TestChatsRoundTrip(t *testing.T) {
|
|||||||
t.Cleanup(func() { os.Chdir(cwd) })
|
t.Cleanup(func() { os.Chdir(cwd) })
|
||||||
t.Setenv("KB_ROOT", dir)
|
t.Setenv("KB_ROOT", dir)
|
||||||
|
|
||||||
if code := runImport([]string{}); code != 0 {
|
if code := RunImport([]string{}); code != 0 {
|
||||||
t.Fatalf("import exit %d", code)
|
t.Fatalf("import exit %d", code)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1,13 +1,11 @@
|
|||||||
package main
|
package chats
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"bufio"
|
"bufio"
|
||||||
"bytes"
|
|
||||||
"encoding/json"
|
"encoding/json"
|
||||||
"flag"
|
"flag"
|
||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"os/exec"
|
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"regexp"
|
"regexp"
|
||||||
"strings"
|
"strings"
|
||||||
@@ -77,7 +75,7 @@ type ExtractedFact struct {
|
|||||||
MessageID string `json:"message_id"`
|
MessageID string `json:"message_id"`
|
||||||
}
|
}
|
||||||
|
|
||||||
func runFacts(args []string) int {
|
func RunFacts(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats facts", flag.ContinueOnError)
|
fs := flag.NewFlagSet("chats facts", flag.ContinueOnError)
|
||||||
help := fs.Bool("help", false, "")
|
help := fs.Bool("help", false, "")
|
||||||
fs.SetOutput(os.Stderr)
|
fs.SetOutput(os.Stderr)
|
||||||
@@ -89,7 +87,7 @@ func runFacts(args []string) int {
|
|||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
|
|
||||||
root := chatsDir()
|
root := Dir()
|
||||||
telegramDir := filepath.Join(root, "telegram")
|
telegramDir := filepath.Join(root, "telegram")
|
||||||
|
|
||||||
entries, err := os.ReadDir(telegramDir)
|
entries, err := os.ReadDir(telegramDir)
|
||||||
@@ -149,7 +147,7 @@ func runFacts(args []string) int {
|
|||||||
}
|
}
|
||||||
fmt.Printf("chats facts: saved to %s\n", factsPath)
|
fmt.Printf("chats facts: saved to %s\n", factsPath)
|
||||||
|
|
||||||
writeFactsToBrain(root, allFacts)
|
writeFactsMarkdown(allFacts)
|
||||||
|
|
||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
@@ -273,14 +271,10 @@ func filterFacts(facts []ExtractedFact, factType string) []ExtractedFact {
|
|||||||
return result
|
return result
|
||||||
}
|
}
|
||||||
|
|
||||||
func writeFactsToBrain(root string, facts []ExtractedFact) {
|
// writeFactsMarkdown stores a sidecar for humans. Brain ingest is
|
||||||
indexScript := filepath.Join(root, "bin", "kb", "index")
|
// bin/brain/index.go (not this subject).
|
||||||
if _, err := os.Stat(indexScript); os.IsNotExist(err) {
|
func writeFactsMarkdown(facts []ExtractedFact) {
|
||||||
fmt.Fprintf(os.Stderr, "chats facts: kb/index not found, skipping brain write\n")
|
mdDir := filepath.Join(Dir(), "facts")
|
||||||
return
|
|
||||||
}
|
|
||||||
|
|
||||||
mdDir := filepath.Join(chatsDir(), "facts")
|
|
||||||
if err := os.MkdirAll(mdDir, 0755); err != nil {
|
if err := os.MkdirAll(mdDir, 0755); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "chats facts: mkdir %s: %v\n", mdDir, err)
|
fmt.Fprintf(os.Stderr, "chats facts: mkdir %s: %v\n", mdDir, err)
|
||||||
return
|
return
|
||||||
@@ -288,7 +282,7 @@ func writeFactsToBrain(root string, facts []ExtractedFact) {
|
|||||||
|
|
||||||
var sb strings.Builder
|
var sb strings.Builder
|
||||||
sb.WriteString("---\n")
|
sb.WriteString("---\n")
|
||||||
sb.WriteString("root: facts\n")
|
sb.WriteString("root: info\n")
|
||||||
sb.WriteString("---\n\n")
|
sb.WriteString("---\n\n")
|
||||||
sb.WriteString("# Chat-Derived Facts\n\n")
|
sb.WriteString("# Chat-Derived Facts\n\n")
|
||||||
for _, f := range facts {
|
for _, f := range facts {
|
||||||
@@ -302,15 +296,5 @@ func writeFactsToBrain(root string, facts []ExtractedFact) {
|
|||||||
fmt.Fprintf(os.Stderr, "chats facts: write %s: %v\n", factsMD, err)
|
fmt.Fprintf(os.Stderr, "chats facts: write %s: %v\n", factsMD, err)
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
|
fmt.Printf("chats facts: markdown %s (index via brain, not chats)\n", factsMD)
|
||||||
cmd := exec.Command(indexScript, "--corpus", mdDir, "--skip-indexes")
|
|
||||||
var outBuf, errBuf bytes.Buffer
|
|
||||||
cmd.Stdout = &outBuf
|
|
||||||
cmd.Stderr = &errBuf
|
|
||||||
cmd.Dir = root
|
|
||||||
if err := cmd.Run(); err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "chats facts: brain index: %v\n%s", err, errBuf.String())
|
|
||||||
return
|
|
||||||
}
|
|
||||||
fmt.Printf("chats facts: written to brain (%s)\n", strings.TrimSpace(outBuf.String()))
|
|
||||||
}
|
}
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
package main
|
package chats
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"bufio"
|
"bufio"
|
||||||
@@ -13,7 +13,7 @@ import (
|
|||||||
"strings"
|
"strings"
|
||||||
)
|
)
|
||||||
|
|
||||||
func runImport(args []string) int {
|
func RunImport(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats import", flag.ContinueOnError)
|
fs := flag.NewFlagSet("chats import", flag.ContinueOnError)
|
||||||
help := fs.Bool("help", false, "")
|
help := fs.Bool("help", false, "")
|
||||||
fs.SetOutput(os.Stderr)
|
fs.SetOutput(os.Stderr)
|
||||||
@@ -25,7 +25,7 @@ func runImport(args []string) int {
|
|||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
|
|
||||||
root := chatsDir()
|
root := Dir()
|
||||||
mdRoot := filepath.Join(root, "md")
|
mdRoot := filepath.Join(root, "md")
|
||||||
glob := filepath.Join(root, "telegram", "*", "messages.jsonl")
|
glob := filepath.Join(root, "telegram", "*", "messages.jsonl")
|
||||||
|
|
||||||
@@ -0,0 +1,669 @@
|
|||||||
|
package chats
|
||||||
|
|
||||||
|
import (
|
||||||
|
"bufio"
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"fmt"
|
||||||
|
"os"
|
||||||
|
"os/exec"
|
||||||
|
"path/filepath"
|
||||||
|
"regexp"
|
||||||
|
"strings"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
type LinkedInMCPSource struct {
|
||||||
|
userDataDir string
|
||||||
|
limit int
|
||||||
|
}
|
||||||
|
|
||||||
|
type lnInboxItem struct {
|
||||||
|
ThreadID string `json:"thread_id"`
|
||||||
|
Participants string `json:"participants"`
|
||||||
|
LastMessage string `json:"last_message"`
|
||||||
|
LastMessageDate string `json:"last_message_date"`
|
||||||
|
Unread bool `json:"unread"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// mcp-server-linkedin v4.22 returns get_inbox / get_conversation as
|
||||||
|
// {url, sections:{inbox|conversation: textblob}, references:{...}}.
|
||||||
|
// The conversation list lives in references (kind=conversation); messages live
|
||||||
|
// in the sections text blob, delimited by "<From> sent the following message
|
||||||
|
// at <time>" markers. See testdata/linkedin_*.json for the wire shape.
|
||||||
|
|
||||||
|
type lnEnvelope struct {
|
||||||
|
URL string `json:"url"`
|
||||||
|
Sections map[string]any `json:"sections"`
|
||||||
|
References map[string]any `json:"references"`
|
||||||
|
}
|
||||||
|
|
||||||
|
type lnReference struct {
|
||||||
|
Kind string `json:"kind"`
|
||||||
|
URL string `json:"url"`
|
||||||
|
Text string `json:"text"`
|
||||||
|
Context string `json:"context"`
|
||||||
|
}
|
||||||
|
|
||||||
|
type lnMessage struct {
|
||||||
|
From string `json:"from"`
|
||||||
|
Date string `json:"date"`
|
||||||
|
Text string `json:"text"`
|
||||||
|
}
|
||||||
|
|
||||||
|
var (
|
||||||
|
lnWeekdays = map[string]time.Weekday{
|
||||||
|
"SUNDAY": time.Sunday, "MONDAY": time.Monday, "TUESDAY": time.Tuesday,
|
||||||
|
"WEDNESDAY": time.Wednesday, "THURSDAY": time.Thursday,
|
||||||
|
"FRIDAY": time.Friday, "SATURDAY": time.Saturday,
|
||||||
|
}
|
||||||
|
lnMsgStartRe = regexp.MustCompile(`^(.+?) sent the following messages? at (.+)$`)
|
||||||
|
lnTimeRe = regexp.MustCompile(`\d{1,2}:\d{2}\s*[AP]M`)
|
||||||
|
)
|
||||||
|
|
||||||
|
func isWeekdayLine(s string) bool {
|
||||||
|
if _, ok := lnWeekdays[s]; ok {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
switch s {
|
||||||
|
case "TODAY", "YESTERDAY", "THIS WEEK", "LAST WEEK":
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
return lnMonthDayRe.MatchString(s)
|
||||||
|
}
|
||||||
|
|
||||||
|
var lnMonthDayRe = regexp.MustCompile(`^[A-Z]{3}\s+\d{1,2}$`)
|
||||||
|
|
||||||
|
// parseLinkedInInbox extracts conversations from a get_inbox response.
|
||||||
|
func parseLinkedInInbox(text string) []lnInboxItem {
|
||||||
|
var env lnEnvelope
|
||||||
|
if err := json.Unmarshal([]byte(text), &env); err != nil {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
refs, _ := env.References["inbox"].([]any)
|
||||||
|
var items []lnInboxItem
|
||||||
|
for _, r := range refs {
|
||||||
|
rr, ok := r.(map[string]any)
|
||||||
|
if !ok {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if rr["kind"] != "conversation" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
u, _ := rr["url"].(string)
|
||||||
|
tid := threadIDFromURL(u)
|
||||||
|
if !validThreadID(tid) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
name, _ := rr["text"].(string)
|
||||||
|
items = append(items, lnInboxItem{
|
||||||
|
ThreadID: tid,
|
||||||
|
Participants: name,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
return items
|
||||||
|
}
|
||||||
|
|
||||||
|
// parseLinkedInConversation parses the sections.conversation text blob into
|
||||||
|
// messages. Messages are delimited by "<From> sent the following message(s) at
|
||||||
|
// <time>" lines; each message body runs until the next marker. Day headers
|
||||||
|
// (all-caps weekdays) provide date context; times are mapped to the most
|
||||||
|
// recent matching weekday.
|
||||||
|
func parseLinkedInConversation(text string) []lnMessage {
|
||||||
|
var env lnEnvelope
|
||||||
|
if err := json.Unmarshal([]byte(text), &env); err != nil {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
blob, _ := env.Sections["conversation"].(string)
|
||||||
|
if blob == "" {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
var msgs []lnMessage
|
||||||
|
var cur *lnMessage
|
||||||
|
var body []string
|
||||||
|
day := ""
|
||||||
|
|
||||||
|
flush := func() {
|
||||||
|
if cur == nil {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
cur.Text = strings.TrimSpace(strings.Join(body, "\n"))
|
||||||
|
if ts := linkedInTimestamp(day, cur.Date); ts != "" {
|
||||||
|
cur.Date = ts
|
||||||
|
}
|
||||||
|
if cur.Text != "" {
|
||||||
|
msgs = append(msgs, *cur)
|
||||||
|
}
|
||||||
|
cur = nil
|
||||||
|
body = nil
|
||||||
|
}
|
||||||
|
|
||||||
|
for _, raw := range strings.Split(blob, "\n") {
|
||||||
|
line := strings.TrimSpace(raw)
|
||||||
|
if line == "" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if isWeekdayLine(line) {
|
||||||
|
if line != day {
|
||||||
|
// A new day header terminates the previous message,
|
||||||
|
// which must keep the earlier date context.
|
||||||
|
flush()
|
||||||
|
}
|
||||||
|
day = line
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if m := lnMsgStartRe.FindStringSubmatch(line); m != nil {
|
||||||
|
flush()
|
||||||
|
cur = &lnMessage{From: strings.TrimSpace(m[1]), Date: strings.TrimSpace(m[2])}
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if cur == nil {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
// Skip "View X's profile" and the "<From> (pronouns) <time>" header.
|
||||||
|
if strings.HasPrefix(line, "View ") && strings.HasSuffix(line, "'s profile") {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if strings.HasPrefix(line, cur.From) && lnTimeRe.MatchString(line) {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
body = append(body, line)
|
||||||
|
}
|
||||||
|
flush()
|
||||||
|
return msgs
|
||||||
|
}
|
||||||
|
|
||||||
|
// linkedInTimestamp maps a weekday, relative, or MON DD date header + clock
|
||||||
|
// string to a timestamp, or returns "" when the clock cannot be parsed.
|
||||||
|
func linkedInTimestamp(day, clock string) string {
|
||||||
|
t, err := time.Parse("3:04 PM", clock)
|
||||||
|
if err != nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
now := time.Now()
|
||||||
|
var d time.Time
|
||||||
|
if wd, ok := lnWeekdays[day]; ok {
|
||||||
|
diff := (int(now.Weekday()) - int(wd) + 7) % 7
|
||||||
|
d = now.AddDate(0, 0, -diff)
|
||||||
|
} else {
|
||||||
|
switch day {
|
||||||
|
case "TODAY":
|
||||||
|
d = now
|
||||||
|
case "YESTERDAY":
|
||||||
|
d = now.AddDate(0, 0, -1)
|
||||||
|
case "THIS WEEK":
|
||||||
|
diff := int(now.Weekday())
|
||||||
|
d = now.AddDate(0, 0, -diff)
|
||||||
|
case "LAST WEEK":
|
||||||
|
diff := int(now.Weekday()) + 7
|
||||||
|
d = now.AddDate(0, 0, -diff)
|
||||||
|
default:
|
||||||
|
if m := lnMonthDayRe.FindStringSubmatch(day); m != nil {
|
||||||
|
// MON DD without a year: resolve to the most recent
|
||||||
|
// occurrence that is not in the future.
|
||||||
|
d = monthDayDate(day, now)
|
||||||
|
if d.IsZero() {
|
||||||
|
return t.Format("15:04")
|
||||||
|
}
|
||||||
|
} else {
|
||||||
|
// No date context; keep bare clock time.
|
||||||
|
return t.Format("15:04")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
res := time.Date(d.Year(), d.Month(), d.Day(), t.Hour(), t.Minute(), 0, 0, time.UTC)
|
||||||
|
return res.UTC().Format(time.RFC3339)
|
||||||
|
}
|
||||||
|
|
||||||
|
var lnMonths = map[string]time.Month{
|
||||||
|
"JAN": time.January, "FEB": time.February, "MAR": time.March,
|
||||||
|
"APR": time.April, "MAY": time.May, "JUN": time.June,
|
||||||
|
"JUL": time.July, "AUG": time.August, "SEP": time.September,
|
||||||
|
"OCT": time.October, "NOV": time.November, "DEC": time.December,
|
||||||
|
}
|
||||||
|
|
||||||
|
// monthDayDate resolves "MON DD" to the most recent occurrence of that date,
|
||||||
|
// preferring the current year and falling back to the previous year when the
|
||||||
|
// date is in the future. Returns zero time when unresolvable.
|
||||||
|
func monthDayDate(day string, now time.Time) time.Time {
|
||||||
|
parts := strings.Fields(day)
|
||||||
|
if len(parts) != 2 {
|
||||||
|
return time.Time{}
|
||||||
|
}
|
||||||
|
mo, ok := lnMonths[parts[0]]
|
||||||
|
if !ok {
|
||||||
|
return time.Time{}
|
||||||
|
}
|
||||||
|
var dd int
|
||||||
|
if _, err := fmt.Sscanf(parts[1], "%d", &dd); err != nil {
|
||||||
|
return time.Time{}
|
||||||
|
}
|
||||||
|
if dd < 1 || dd > 31 {
|
||||||
|
return time.Time{}
|
||||||
|
}
|
||||||
|
d := time.Date(now.Year(), mo, dd, 0, 0, 0, 0, time.UTC)
|
||||||
|
if d.After(now) {
|
||||||
|
d = d.AddDate(-1, 0, 0)
|
||||||
|
}
|
||||||
|
if d.After(now) {
|
||||||
|
return time.Time{}
|
||||||
|
}
|
||||||
|
return d
|
||||||
|
}
|
||||||
|
|
||||||
|
func threadIDFromURL(u string) string {
|
||||||
|
u = strings.TrimSuffix(u, "/")
|
||||||
|
idx := strings.LastIndex(u, "/")
|
||||||
|
if idx < 0 {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return u[idx+1:]
|
||||||
|
}
|
||||||
|
|
||||||
|
// validThreadID rejects path segments that are not real thread ids (e.g. the
|
||||||
|
// literal "thread" or an empty trailing segment).
|
||||||
|
func validThreadID(id string) bool {
|
||||||
|
if id == "" || id == "thread" {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
|
||||||
|
func NewLinkedInMCPSource(userDataDir string) *LinkedInMCPSource {
|
||||||
|
return &LinkedInMCPSource{userDataDir: userDataDir}
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *LinkedInMCPSource) Name() string { return "linkedin" }
|
||||||
|
|
||||||
|
func (s *LinkedInMCPSource) Sync(ctx context.Context, outDir string, limit int) error {
|
||||||
|
if limit > 0 {
|
||||||
|
s.limit = limit
|
||||||
|
}
|
||||||
|
|
||||||
|
// getConversation fetches one thread, recreating the MCP server when it
|
||||||
|
// wedges. A single 429 makes mcp-server-linkedin close its browser and
|
||||||
|
// refuse every later call ("still has a browser open"), so a broken server
|
||||||
|
// must be restarted rather than hammered.
|
||||||
|
getConversation := func(threadID string) ([]lnMessage, error) {
|
||||||
|
client, err := newLinkedInMCP(ctx, s.userDataDir)
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("linkedin mcp: %w", err)
|
||||||
|
}
|
||||||
|
defer client.Close()
|
||||||
|
msgs, err := client.GetConversation(ctx, "", threadID, msgLimitFor(s.limit))
|
||||||
|
if err != nil && wedged(err) {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: %s: server wedged, restarting broker\n", threadID)
|
||||||
|
time.Sleep(5 * time.Second)
|
||||||
|
client2, cerr := newLinkedInMCP(ctx, s.userDataDir)
|
||||||
|
if cerr == nil {
|
||||||
|
defer client2.Close()
|
||||||
|
msgs, err = client2.GetConversation(ctx, "", threadID, msgLimitFor(s.limit))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return msgs, err
|
||||||
|
}
|
||||||
|
|
||||||
|
client, err := newLinkedInMCP(ctx, s.userDataDir)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("linkedin mcp: %w", err)
|
||||||
|
}
|
||||||
|
inbox, err := client.GetInbox(ctx, 50)
|
||||||
|
if err != nil {
|
||||||
|
client.Close()
|
||||||
|
return fmt.Errorf("get_inbox: %w", err)
|
||||||
|
}
|
||||||
|
client.Close()
|
||||||
|
if len(inbox) == 0 {
|
||||||
|
fmt.Println("chats: no LinkedIn conversations found")
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
fmt.Printf("chats: found %d LinkedIn conversations\n", len(inbox))
|
||||||
|
|
||||||
|
for _, conv := range inbox {
|
||||||
|
convID := sanitizeDir(conv.ThreadID)
|
||||||
|
if convID == "" {
|
||||||
|
convID = fmt.Sprintf("conv_%d", time.Now().UnixNano())
|
||||||
|
}
|
||||||
|
|
||||||
|
parts := strings.SplitN(conv.Participants, ",", 2)
|
||||||
|
chatName := strings.TrimSpace(parts[0])
|
||||||
|
if chatName == "" {
|
||||||
|
chatName = convID
|
||||||
|
}
|
||||||
|
|
||||||
|
chatDir := filepath.Join(outDir, "linkedin", convID)
|
||||||
|
if err := os.MkdirAll(chatDir, 0755); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: mkdir %s: %v\n", chatDir, err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
|
||||||
|
msgs, err := getConversation(conv.ThreadID)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: get_conversation %s: %v\n", convID, err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
|
||||||
|
jsonlPath := filepath.Join(chatDir, "messages.jsonl")
|
||||||
|
|
||||||
|
// A rate-limited response can parse to zero messages. Never clobber
|
||||||
|
// previously synced data with an empty file.
|
||||||
|
if len(msgs) == 0 {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: %s (%s): 0 messages parsed, keeping existing file\n", chatName, convID)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
|
||||||
|
f, err := os.Create(jsonlPath)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: create %s: %v\n", jsonlPath, err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
|
||||||
|
enc := json.NewEncoder(f)
|
||||||
|
written := 0
|
||||||
|
for i, m := range msgs {
|
||||||
|
text := m.Text
|
||||||
|
if text == "" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
|
||||||
|
ts := m.Date
|
||||||
|
if t, err := time.Parse("2006-01-02T15:04:05Z07:00", m.Date); err == nil {
|
||||||
|
ts = t.UTC().Format(time.RFC3339)
|
||||||
|
} else if t, err := time.Parse(time.RFC3339, m.Date); err == nil {
|
||||||
|
ts = t.UTC().Format(time.RFC3339)
|
||||||
|
}
|
||||||
|
|
||||||
|
chatMsg := Message{
|
||||||
|
ID: fmt.Sprintf("li_%s_%d", convID, i),
|
||||||
|
Timestamp: ts,
|
||||||
|
From: m.From,
|
||||||
|
Text: text,
|
||||||
|
Platform: "linkedin",
|
||||||
|
}
|
||||||
|
if err := enc.Encode(chatMsg); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: encode: %v\n", err)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
written++
|
||||||
|
}
|
||||||
|
f.Close()
|
||||||
|
|
||||||
|
fmt.Printf("chats: synced %s (%s) — %d messages\n", chatName, convID, written)
|
||||||
|
|
||||||
|
// Pause between conversations to reduce LinkedIn rate limiting.
|
||||||
|
select {
|
||||||
|
case <-ctx.Done():
|
||||||
|
return ctx.Err()
|
||||||
|
case <-time.After(2 * time.Second):
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
type linkedInMCPClient struct {
|
||||||
|
cmd *exec.Cmd
|
||||||
|
stdin *bufio.Writer
|
||||||
|
stdout *bufio.Scanner
|
||||||
|
msgID int
|
||||||
|
}
|
||||||
|
|
||||||
|
func newLinkedInMCP(ctx context.Context, userDataDir string) (*linkedInMCPClient, error) {
|
||||||
|
args := []string{
|
||||||
|
"mcp-server-linkedin@latest",
|
||||||
|
"--user-data-dir", userDataDir,
|
||||||
|
"--no-auto-import",
|
||||||
|
"--no-daemon",
|
||||||
|
"--transport", "stdio",
|
||||||
|
"--login-timeout", "10",
|
||||||
|
"--browser-wait", "1",
|
||||||
|
"--browser-idle-timeout", "10",
|
||||||
|
"--log-level", "ERROR",
|
||||||
|
}
|
||||||
|
|
||||||
|
cmd := exec.CommandContext(ctx, "uvx", args...)
|
||||||
|
cmd.Env = os.Environ()
|
||||||
|
|
||||||
|
stdin, err := cmd.StdinPipe()
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("stdin pipe: %w", err)
|
||||||
|
}
|
||||||
|
stdout, err := cmd.StdoutPipe()
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("stdout pipe: %w", err)
|
||||||
|
}
|
||||||
|
cmd.Stderr = os.Stderr
|
||||||
|
|
||||||
|
if err := cmd.Start(); err != nil {
|
||||||
|
return nil, fmt.Errorf("start: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
c := &linkedInMCPClient{
|
||||||
|
cmd: cmd,
|
||||||
|
stdin: bufio.NewWriter(stdin),
|
||||||
|
stdout: bufio.NewScanner(stdout),
|
||||||
|
msgID: 0,
|
||||||
|
}
|
||||||
|
c.stdout.Buffer(make([]byte, 1<<20), 1<<20)
|
||||||
|
|
||||||
|
if err := c.initialize(ctx); err != nil {
|
||||||
|
c.Close()
|
||||||
|
return nil, fmt.Errorf("init: %w", err)
|
||||||
|
}
|
||||||
|
return c, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *linkedInMCPClient) nextID() int {
|
||||||
|
c.msgID++
|
||||||
|
return c.msgID
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *linkedInMCPClient) initialize(ctx context.Context) error {
|
||||||
|
params := map[string]interface{}{
|
||||||
|
"protocolVersion": "2024-11-05",
|
||||||
|
"capabilities": map[string]interface{}{},
|
||||||
|
"clientInfo": map[string]string{
|
||||||
|
"name": "chats-sync",
|
||||||
|
"version": "0.1.0",
|
||||||
|
},
|
||||||
|
}
|
||||||
|
_, err := c.send(ctx, "initialize", params)
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *linkedInMCPClient) send(ctx context.Context, method string, params interface{}) (json.RawMessage, error) {
|
||||||
|
id := c.nextID()
|
||||||
|
req := map[string]interface{}{
|
||||||
|
"jsonrpc": "2.0",
|
||||||
|
"id": id,
|
||||||
|
"method": method,
|
||||||
|
}
|
||||||
|
if params != nil {
|
||||||
|
req["params"] = params
|
||||||
|
}
|
||||||
|
body, err := json.Marshal(req)
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("marshal: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
if _, err := c.stdin.Write(body); err != nil {
|
||||||
|
return nil, fmt.Errorf("write: %w", err)
|
||||||
|
}
|
||||||
|
if err := c.stdin.WriteByte('\n'); err != nil {
|
||||||
|
return nil, fmt.Errorf("newline: %w", err)
|
||||||
|
}
|
||||||
|
if err := c.stdin.Flush(); err != nil {
|
||||||
|
return nil, fmt.Errorf("flush: %w", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
for c.stdout.Scan() {
|
||||||
|
line := c.stdout.Text()
|
||||||
|
if line == "" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
var resp struct {
|
||||||
|
JSONRPC string `json:"jsonrpc"`
|
||||||
|
ID int `json:"id"`
|
||||||
|
Result json.RawMessage `json:"result,omitempty"`
|
||||||
|
Error *struct {
|
||||||
|
Code int `json:"code"`
|
||||||
|
Message string `json:"message"`
|
||||||
|
} `json:"error,omitempty"`
|
||||||
|
}
|
||||||
|
if err := json.Unmarshal([]byte(line), &resp); err != nil {
|
||||||
|
return nil, fmt.Errorf("unmarshal: %w\nline: %s", err, line[:min(len(line), 500)])
|
||||||
|
}
|
||||||
|
if resp.Error != nil {
|
||||||
|
return nil, fmt.Errorf("rpc error %d: %s", resp.Error.Code, resp.Error.Message)
|
||||||
|
}
|
||||||
|
return resp.Result, nil
|
||||||
|
}
|
||||||
|
return nil, fmt.Errorf("no response: %w", c.stdout.Err())
|
||||||
|
}
|
||||||
|
|
||||||
|
// msgLimitFor returns the per-conversation message cap for a sync.
|
||||||
|
func msgLimitFor(limit int) int {
|
||||||
|
if limit > 0 {
|
||||||
|
return limit
|
||||||
|
}
|
||||||
|
return 100
|
||||||
|
}
|
||||||
|
|
||||||
|
// wedged reports whether a conversation fetch failure means the MCP server
|
||||||
|
// closed its browser and will refuse every later call.
|
||||||
|
func wedged(err error) bool {
|
||||||
|
return strings.Contains(err.Error(), "still has a browser open")
|
||||||
|
}
|
||||||
|
|
||||||
|
// callTool invokes an MCP tool, retrying transient (rate-limit) failures.
|
||||||
|
func (c *linkedInMCPClient) callTool(ctx context.Context, name string, params map[string]interface{}) (json.RawMessage, error) {
|
||||||
|
var lastErr error
|
||||||
|
for attempt := 0; attempt < 3; attempt++ {
|
||||||
|
if attempt > 0 {
|
||||||
|
delay := time.Duration(1<<uint(attempt)) * 5 * time.Second
|
||||||
|
select {
|
||||||
|
case <-ctx.Done():
|
||||||
|
return nil, ctx.Err()
|
||||||
|
case <-time.After(delay):
|
||||||
|
}
|
||||||
|
}
|
||||||
|
result, err := c.send(ctx, "tools/call", map[string]interface{}{
|
||||||
|
"name": name,
|
||||||
|
"arguments": params,
|
||||||
|
})
|
||||||
|
if err == nil {
|
||||||
|
// Tool-level errors surface as a successful RPC with an
|
||||||
|
// isError=true content entry.
|
||||||
|
if hint := toolErrorHint(result); hint != "" {
|
||||||
|
lastErr = fmt.Errorf("%s error: %s", name, hint)
|
||||||
|
if !isTransientLinkedInError(lastErr.Error()) {
|
||||||
|
return nil, lastErr
|
||||||
|
}
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
return result, nil
|
||||||
|
}
|
||||||
|
lastErr = err
|
||||||
|
if !isTransientLinkedInError(err.Error()) {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return nil, fmt.Errorf("%s: %w", name, lastErr)
|
||||||
|
}
|
||||||
|
|
||||||
|
// toolErrorHint returns the tool's error text when the result has isError set.
|
||||||
|
func toolErrorHint(result json.RawMessage) string {
|
||||||
|
var toolRes struct {
|
||||||
|
Content []struct {
|
||||||
|
Type string `json:"type"`
|
||||||
|
Text string `json:"text"`
|
||||||
|
} `json:"content"`
|
||||||
|
IsError bool `json:"isError"`
|
||||||
|
}
|
||||||
|
if err := json.Unmarshal(result, &toolRes); err != nil || !toolRes.IsError {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
if len(toolRes.Content) > 0 {
|
||||||
|
return toolRes.Content[0].Text
|
||||||
|
}
|
||||||
|
return "unknown tool error"
|
||||||
|
}
|
||||||
|
|
||||||
|
// isTransientLinkedInError reports whether a fetch failed due to rate limiting
|
||||||
|
// or a transient server error, which may succeed on retry.
|
||||||
|
func isTransientLinkedInError(msg string) bool {
|
||||||
|
return strings.Contains(msg, "503") || strings.Contains(msg, "429") ||
|
||||||
|
strings.Contains(msg, "ERR_HTTP_RESPONSE_CODE_FAILURE") ||
|
||||||
|
strings.Contains(msg, "ERR_ABORTED") ||
|
||||||
|
strings.Contains(msg, "Error calling tool") ||
|
||||||
|
strings.Contains(msg, "Unexpected error") ||
|
||||||
|
strings.Contains(msg, "still has a browser open")
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *linkedInMCPClient) GetInbox(ctx context.Context, limit int) ([]lnInboxItem, error) {
|
||||||
|
params := map[string]interface{}{
|
||||||
|
"limit": limit,
|
||||||
|
}
|
||||||
|
result, err := c.callTool(ctx, "get_inbox", params)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
|
||||||
|
var toolRes struct {
|
||||||
|
Content []struct {
|
||||||
|
Type string `json:"type"`
|
||||||
|
Text string `json:"text"`
|
||||||
|
} `json:"content"`
|
||||||
|
IsError bool `json:"isError"`
|
||||||
|
}
|
||||||
|
if err := json.Unmarshal(result, &toolRes); err != nil {
|
||||||
|
return nil, fmt.Errorf("unmarshal tool: %w", err)
|
||||||
|
}
|
||||||
|
if len(toolRes.Content) == 0 {
|
||||||
|
return nil, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
text := toolRes.Content[0].Text
|
||||||
|
return parseLinkedInInbox(text), nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *linkedInMCPClient) GetConversation(ctx context.Context, username, threadID string, limit int) ([]lnMessage, error) {
|
||||||
|
params := map[string]interface{}{
|
||||||
|
"linkedin_username": username,
|
||||||
|
"thread_id": threadID,
|
||||||
|
"index": limit,
|
||||||
|
}
|
||||||
|
result, err := c.callTool(ctx, "get_conversation", params)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
|
||||||
|
var toolRes struct {
|
||||||
|
Content []struct {
|
||||||
|
Type string `json:"type"`
|
||||||
|
Text string `json:"text"`
|
||||||
|
} `json:"content"`
|
||||||
|
IsError bool `json:"isError"`
|
||||||
|
}
|
||||||
|
if err := json.Unmarshal(result, &toolRes); err != nil {
|
||||||
|
return nil, fmt.Errorf("unmarshal tool: %w", err)
|
||||||
|
}
|
||||||
|
if len(toolRes.Content) == 0 {
|
||||||
|
return nil, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
text := toolRes.Content[0].Text
|
||||||
|
return parseLinkedInConversation(text), nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *linkedInMCPClient) Close() error {
|
||||||
|
if c.stdin != nil {
|
||||||
|
c.stdin.Flush()
|
||||||
|
}
|
||||||
|
if c.cmd != nil && c.cmd.Process != nil {
|
||||||
|
c.cmd.Process.Kill()
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,215 @@
|
|||||||
|
package chats
|
||||||
|
|
||||||
|
import (
|
||||||
|
"errors"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
func readFixture(t *testing.T, name string) string {
|
||||||
|
t.Helper()
|
||||||
|
data, err := os.ReadFile(filepath.Join("testdata", name))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
return string(data)
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestParseLinkedInInbox verifies get_inbox parsing against the v4.22 wire
|
||||||
|
// format (testdata/linkedin_inbox.json — synthetic Alice/Bob/Charlie).
|
||||||
|
func TestParseLinkedInInbox(t *testing.T) {
|
||||||
|
text := readFixture(t, "linkedin_inbox.json")
|
||||||
|
items := parseLinkedInInbox(text)
|
||||||
|
if len(items) != 2 {
|
||||||
|
t.Fatalf("expected 2 conversations (empty thread url skipped), got %d", len(items))
|
||||||
|
}
|
||||||
|
|
||||||
|
first := items[0]
|
||||||
|
if first.ThreadID == "" {
|
||||||
|
t.Error("expected thread id extracted from reference url")
|
||||||
|
}
|
||||||
|
if first.Participants != "Alice Example" {
|
||||||
|
t.Errorf("participants=%q, want Alice Example", first.Participants)
|
||||||
|
}
|
||||||
|
if !strings.HasPrefix(first.ThreadID, "2-") {
|
||||||
|
t.Errorf("unexpected thread id format %q", first.ThreadID)
|
||||||
|
}
|
||||||
|
if items[1].Participants != "Charlie Example" {
|
||||||
|
t.Errorf("second participant=%q, want Charlie Example", items[1].Participants)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestParseLinkedInInboxBadJSON verifies a non-JSON response yields no items
|
||||||
|
// rather than a panic or error.
|
||||||
|
func TestParseLinkedInInboxBadJSON(t *testing.T) {
|
||||||
|
if got := parseLinkedInInbox("Session expired"); len(got) != 0 {
|
||||||
|
t.Fatalf("expected no items for non-JSON, got %d", len(got))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestParseLinkedInConversation verifies message extraction from the sections
|
||||||
|
// blob (testdata/linkedin_conversation.json — synthetic Alice/Bob).
|
||||||
|
func TestParseLinkedInConversation(t *testing.T) {
|
||||||
|
text := readFixture(t, "linkedin_conversation.json")
|
||||||
|
msgs := parseLinkedInConversation(text)
|
||||||
|
if len(msgs) != 2 {
|
||||||
|
t.Fatalf("expected 2 messages, got %d", len(msgs))
|
||||||
|
}
|
||||||
|
|
||||||
|
if msgs[0].From != "Alice Example" {
|
||||||
|
t.Errorf("from=%q, want Alice Example", msgs[0].From)
|
||||||
|
}
|
||||||
|
if !strings.Contains(msgs[0].Text, "Senior Software Engineer") {
|
||||||
|
t.Errorf("alice text missing role, got %q", msgs[0].Text)
|
||||||
|
}
|
||||||
|
if msgs[0].Date == "" {
|
||||||
|
t.Error("expected message date")
|
||||||
|
}
|
||||||
|
if msgs[1].From != "Bob Example" {
|
||||||
|
t.Errorf("from=%q, want Bob Example", msgs[1].From)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestParseLinkedInConversationEmpty verifies empty/non-JSON blobs parse to
|
||||||
|
// zero messages.
|
||||||
|
func TestParseLinkedInConversationEmpty(t *testing.T) {
|
||||||
|
if got := parseLinkedInConversation("no data here"); len(got) != 0 {
|
||||||
|
t.Fatalf("expected 0 messages, got %d", len(got))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestLinkedInTimestamp verifies weekday+clock resolution to a recent UTC date.
|
||||||
|
func TestLinkedInTimestamp(t *testing.T) {
|
||||||
|
// The most recent Wednesday before/equal to "now".
|
||||||
|
ts := linkedInTimestamp("WEDNESDAY", "10:02 AM")
|
||||||
|
parsed, err := time.Parse(time.RFC3339, ts)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("unparseable timestamp %q: %v", ts, err)
|
||||||
|
}
|
||||||
|
if parsed.Weekday() != time.Wednesday {
|
||||||
|
t.Errorf("expected Wednesday, got %s", parsed.Weekday())
|
||||||
|
}
|
||||||
|
if parsed.Hour() != 10 || parsed.Minute() != 2 {
|
||||||
|
t.Errorf("expected 10:02, got %02d:%02d", parsed.Hour(), parsed.Minute())
|
||||||
|
}
|
||||||
|
|
||||||
|
now := time.Now()
|
||||||
|
diff := now.Sub(parsed)
|
||||||
|
if diff < 0 || diff > 7*24*time.Hour {
|
||||||
|
t.Errorf("timestamp %s is not within the last week of %s", parsed, now)
|
||||||
|
}
|
||||||
|
|
||||||
|
if got := linkedInTimestamp("MONDAY", "garbage"); got != "" {
|
||||||
|
t.Errorf("expected empty for bad clock, got %q", got)
|
||||||
|
}
|
||||||
|
if got := linkedInTimestamp("", "1:22 PM"); got != "13:22" {
|
||||||
|
t.Errorf("expected bare 13:22 for missing weekday, got %q", got)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Relative day headers must resolve to full dates, not bare clocks.
|
||||||
|
today := linkedInTimestamp("TODAY", "9:42 AM")
|
||||||
|
yp, err := time.Parse(time.RFC3339, today)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("TODAY unparseable %q: %v", today, err)
|
||||||
|
}
|
||||||
|
if yp.Year() != now.Year() || yp.Month() != now.Month() || yp.Day() != now.Day() {
|
||||||
|
t.Errorf("TODAY expected %v, got %v", now, yp)
|
||||||
|
}
|
||||||
|
yest := linkedInTimestamp("YESTERDAY", "3:00 PM")
|
||||||
|
yp, err = time.Parse(time.RFC3339, yest)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("YESTERDAY unparseable %q: %v", yest, err)
|
||||||
|
}
|
||||||
|
if yp.Day() != now.AddDate(0, 0, -1).Day() {
|
||||||
|
t.Errorf("YESTERDAY expected day %d, got %d", now.AddDate(0, 0, -1).Day(), yp.Day())
|
||||||
|
}
|
||||||
|
|
||||||
|
// MON DD header (e.g. "JUN 25"): must resolve to a full date. The
|
||||||
|
// timestamp should fall within the current year (falling back to the
|
||||||
|
// prior year if the date would be in the future).
|
||||||
|
md := linkedInTimestamp("JUN 25", "10:48 AM")
|
||||||
|
mp, err := time.Parse(time.RFC3339, md)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("MON DD unparseable %q: %v", md, err)
|
||||||
|
}
|
||||||
|
if mp.Year() != now.Year() && mp.Year() != now.Year()-1 {
|
||||||
|
t.Errorf("JUN 25 expected year %d or %d, got %d", now.Year(), now.Year()-1, mp.Year())
|
||||||
|
}
|
||||||
|
if mp.Month() != time.June || mp.Day() != 25 {
|
||||||
|
t.Errorf("JUN 25 expected Jun 25, got %s %d", mp.Month(), mp.Day())
|
||||||
|
}
|
||||||
|
if mp.After(now) {
|
||||||
|
t.Errorf("JUN 25 resolved to the future: %s > %s", mp, now)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestTransientLinkedInError verifies rate-limit errors are retryable but
|
||||||
|
// genuine failures are not.
|
||||||
|
func TestTransientLinkedInError(t *testing.T) {
|
||||||
|
retryable := []string{
|
||||||
|
"get_conversation error: Error calling tool 'get_conversation'",
|
||||||
|
"get_conversation error: Unexpected error in get_conversation: net::ERR_HTTP_RESPONSE_CODE_FAILURE",
|
||||||
|
"rpc error 503: rate limited",
|
||||||
|
"rpc error 429: too many requests",
|
||||||
|
"get_conversation: get_conversation error: This server still has a browser open on the profile.",
|
||||||
|
}
|
||||||
|
for _, msg := range retryable {
|
||||||
|
if !isTransientLinkedInError(msg) {
|
||||||
|
t.Errorf("expected %q to be transient", msg)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
permanent := []string{
|
||||||
|
"get_inbox error: bad credentials",
|
||||||
|
"rpc error -32602: Invalid request parameters",
|
||||||
|
"unmarshal: unexpected end of JSON input",
|
||||||
|
}
|
||||||
|
for _, msg := range permanent {
|
||||||
|
if isTransientLinkedInError(msg) {
|
||||||
|
t.Errorf("expected %q to be permanent", msg)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestMsgLimitFor verifies the per-conversation message cap resolution.
|
||||||
|
func TestMsgLimitFor(t *testing.T) {
|
||||||
|
if got := msgLimitFor(0); got != 100 {
|
||||||
|
t.Errorf("expected default 100, got %d", got)
|
||||||
|
}
|
||||||
|
if got := msgLimitFor(5); got != 5 {
|
||||||
|
t.Errorf("expected 5, got %d", got)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestWedged verifies the browser-open failure is recognized as a wedge.
|
||||||
|
func TestWedged(t *testing.T) {
|
||||||
|
if !wedged(errors.New("get_conversation error: This server still has a browser open on the profile")) {
|
||||||
|
t.Error("expected wedged error to be recognized")
|
||||||
|
}
|
||||||
|
if wedged(errors.New("get_conversation error: bad thing")) {
|
||||||
|
t.Error("unexpected wedge detection")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestThreadIDFromURL verifies thread id extraction.
|
||||||
|
func TestThreadIDFromURL(t *testing.T) {
|
||||||
|
cases := []struct {
|
||||||
|
url, want string
|
||||||
|
}{
|
||||||
|
{"/messaging/thread/2-abc123/", "2-abc123"},
|
||||||
|
{"/messaging/thread/2-abc123", "2-abc123"},
|
||||||
|
{"", ""},
|
||||||
|
{"/messaging/thread/", ""},
|
||||||
|
}
|
||||||
|
for _, c := range cases {
|
||||||
|
got := threadIDFromURL(c.url)
|
||||||
|
if c.want == "" && validThreadID(got) {
|
||||||
|
t.Errorf("threadIDFromURL(%q) = %q, want empty", c.url, got)
|
||||||
|
}
|
||||||
|
if c.want != "" && got != c.want {
|
||||||
|
t.Errorf("threadIDFromURL(%q) = %q, want %q", c.url, got, c.want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
package main
|
package chats
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"bufio"
|
"bufio"
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
package chats
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
"strings"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Root locates the 2dph project root (KB_ROOT, or walk up for var/ or .git).
|
||||||
|
func Root() string {
|
||||||
|
if v := os.Getenv("KB_ROOT"); v != "" {
|
||||||
|
return v
|
||||||
|
}
|
||||||
|
wd, err := os.Getwd()
|
||||||
|
if err != nil {
|
||||||
|
return "."
|
||||||
|
}
|
||||||
|
for i := 0; i < 10; i++ {
|
||||||
|
if _, err := os.Stat(wd + "/var"); err == nil {
|
||||||
|
return wd
|
||||||
|
}
|
||||||
|
if _, err := os.Stat(wd + "/.git"); err == nil {
|
||||||
|
return wd
|
||||||
|
}
|
||||||
|
parent := wd
|
||||||
|
if idx := strings.LastIndex(wd, "/"); idx >= 0 {
|
||||||
|
parent = wd[:idx]
|
||||||
|
}
|
||||||
|
if parent == wd {
|
||||||
|
break
|
||||||
|
}
|
||||||
|
wd = parent
|
||||||
|
}
|
||||||
|
return "."
|
||||||
|
}
|
||||||
|
|
||||||
|
// Dir is var/chats under the project root.
|
||||||
|
func Dir() string {
|
||||||
|
return Root() + "/var/chats"
|
||||||
|
}
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
package main
|
package chats
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
@@ -0,0 +1,105 @@
|
|||||||
|
package chats
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"flag"
|
||||||
|
"fmt"
|
||||||
|
"os"
|
||||||
|
"os/exec"
|
||||||
|
"path/filepath"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
func checkLinkedInSession(userDataDir string) (bool, error) {
|
||||||
|
// Validate the source-session files without launching a browser. A full
|
||||||
|
// `--status` run spawns Chromium and loads /feed/, doubling the automation
|
||||||
|
// exposed to LinkedIn (429 rate limits) before the sync even starts.
|
||||||
|
root := filepath.Dir(userDataDir)
|
||||||
|
sessionFiles := []string{
|
||||||
|
filepath.Join(root, "source-state.json"),
|
||||||
|
filepath.Join(root, "cookies.json"),
|
||||||
|
filepath.Join(userDataDir, "Default", "Cookies"),
|
||||||
|
}
|
||||||
|
for _, f := range sessionFiles {
|
||||||
|
if _, err := os.Stat(f); err != nil {
|
||||||
|
return true, fmt.Errorf("missing session file %s", f)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return false, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func RunSyncLinkedIn(args []string) int {
|
||||||
|
fs := flag.NewFlagSet("chats sync linkedin", flag.ContinueOnError)
|
||||||
|
limit := fs.Int("limit", 0, "max messages per conversation (0 = all)")
|
||||||
|
refresh := fs.Bool("refresh", false, "refresh session from live webtop browser before sync")
|
||||||
|
help := fs.Bool("help", false, "")
|
||||||
|
fs.SetOutput(os.Stderr)
|
||||||
|
if err := fs.Parse(args); err != nil {
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
if *help {
|
||||||
|
fmt.Fprintln(os.Stderr, "usage: chats sync linkedin [--limit N] [--refresh]")
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
userDataDir := envVar("LINKEDIN_USER_DATA_DIR", "")
|
||||||
|
if userDataDir == "" {
|
||||||
|
home, _ := os.UserHomeDir()
|
||||||
|
userDataDir = home + "/.linkedin-mcp/profile"
|
||||||
|
}
|
||||||
|
|
||||||
|
if *refresh {
|
||||||
|
if code := refreshLinkedInSession(userDataDir); code != 0 {
|
||||||
|
return code
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Check session files first (no browser launch).
|
||||||
|
loginNeeded, err := checkLinkedInSession(userDataDir)
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: linkedin status check: %v\n", err)
|
||||||
|
}
|
||||||
|
if loginNeeded {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: LinkedIn session missing. Run:\n")
|
||||||
|
fmt.Fprintf(os.Stderr, " chats sync linkedin --refresh\n")
|
||||||
|
fmt.Fprintf(os.Stderr, "or point LINKEDIN_USER_DATA_DIR at a valid session\n")
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
src := NewLinkedInMCPSource(userDataDir)
|
||||||
|
|
||||||
|
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute)
|
||||||
|
defer cancel()
|
||||||
|
|
||||||
|
start := time.Now()
|
||||||
|
if err := src.Sync(ctx, Dir(), *limit); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats sync linkedin: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
fmt.Printf("chats sync linkedin: completed in %s\n", time.Since(start).Round(time.Millisecond))
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
// refreshLinkedInSession re-syncs the LinkedIn source session from the live
|
||||||
|
// webtop browser via the vendored refresh-linkedin-session helper.
|
||||||
|
func refreshLinkedInSession(userDataDir string) int {
|
||||||
|
exe, err := os.Executable()
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: resolve executable: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
helper := filepath.Join(filepath.Dir(exe), "refresh-linkedin-session")
|
||||||
|
if _, err := os.Stat(helper); err != nil {
|
||||||
|
// Fall back to the source tree helper next to this command file.
|
||||||
|
helper = "bin/chats/refresh-linkedin-session"
|
||||||
|
}
|
||||||
|
root := filepath.Dir(userDataDir)
|
||||||
|
cmd := exec.Command(helper, "--root", root)
|
||||||
|
cmd.Stdout = os.Stderr
|
||||||
|
cmd.Stderr = os.Stderr
|
||||||
|
if err := cmd.Run(); err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "chats: linkedin session refresh: %v\n", err)
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
return 0
|
||||||
|
}
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
package main
|
package chats
|
||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
@@ -11,7 +11,7 @@ import (
|
|||||||
"time"
|
"time"
|
||||||
)
|
)
|
||||||
|
|
||||||
func runSyncTelegram(args []string) int {
|
func RunSyncTelegram(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats sync telegram", flag.ContinueOnError)
|
fs := flag.NewFlagSet("chats sync telegram", flag.ContinueOnError)
|
||||||
limit := fs.Int("limit", 0, "max messages per chat (0 = all)")
|
limit := fs.Int("limit", 0, "max messages per chat (0 = all)")
|
||||||
phone := fs.String("phone", "", "phone number (default env TELEGRAM_PHONE)")
|
phone := fs.String("phone", "", "phone number (default env TELEGRAM_PHONE)")
|
||||||
@@ -78,7 +78,7 @@ func runSyncTelegram(args []string) int {
|
|||||||
defer cancel()
|
defer cancel()
|
||||||
|
|
||||||
start := time.Now()
|
start := time.Now()
|
||||||
if err := src.Sync(ctx, chatsDir(), *limit); err != nil {
|
if err := src.Sync(ctx, Dir(), *limit); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "chats sync telegram: %v\n", err)
|
fmt.Fprintf(os.Stderr, "chats sync telegram: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
}
|
}
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
{
|
||||||
|
"url": "https://www.linkedin.com/messaging/thread/2-YWxpY2UtYm9iLXRocmVhZC0xMjM=/",
|
||||||
|
"sections": {
|
||||||
|
"conversation": "WEDNESDAY\nAlice Example sent the following message at 10:02 AM\nView Alice Example's profile\nAlice Example (She/Her) 10:02 AM\nHi Bob, we have a Senior Software Engineer role that matches your Go and Python background. Happy to share more if you are open to a chat.\n\nBob Example sent the following messages at 1:22 PM\nView Bob Example's profile\nBob Example 1:22 PM\nHi Alice, thanks for reaching out — yes, I am open to exploring a Senior Software Engineer role. Happy to do a short video call.\n"
|
||||||
|
},
|
||||||
|
"references": {
|
||||||
|
"conversation": [
|
||||||
|
{"kind": "person", "url": "/in/alice-example/", "text": "Alice Example"},
|
||||||
|
{"kind": "person", "url": "/in/bob-example/", "text": "Bob Example"}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
+34
@@ -0,0 +1,34 @@
|
|||||||
|
{
|
||||||
|
"url": "https://www.linkedin.com/messaging/",
|
||||||
|
"sections": {
|
||||||
|
"inbox": "Messaging\nInbox\nConversation List\nAlice Example\nExciting opportunity for a senior software engineer\n"
|
||||||
|
},
|
||||||
|
"references": {
|
||||||
|
"inbox": [
|
||||||
|
{
|
||||||
|
"kind": "conversation",
|
||||||
|
"url": "/messaging/thread/2-YWxpY2UtYm9iLXRocmVhZC0xMjM=/",
|
||||||
|
"context": "inbox",
|
||||||
|
"text": "Alice Example"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"kind": "person",
|
||||||
|
"url": "/in/alice-example/",
|
||||||
|
"text": "Alice Example",
|
||||||
|
"context": "inbox"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"kind": "conversation",
|
||||||
|
"url": "/messaging/thread/2-Y2hhcmxpZS1ib2ItdGhyZWFkLTQ1Ng==/",
|
||||||
|
"context": "inbox",
|
||||||
|
"text": "Charlie Example"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"kind": "conversation",
|
||||||
|
"url": "/messaging/thread/",
|
||||||
|
"context": "inbox",
|
||||||
|
"text": "should-be-skipped"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user