Compare commits

..
Author SHA1 Message Date
eSliderandGitHub 36976d9b53 feat: facts audit/extract/crm shebang wrappers (D14) (#19)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
Python stays the implementation. Commands are bin/facts/{audit,extract,crm}.go
like postgres/query.go.
2026-08-13 21:20:28 +01:00
eSliderandGitHub 8e6f67cc97 feat: brain get/stats/eval call internal/brain, not Python (#18)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
Read path is cgo like search. Control questions live in rank so CI can
test them without ladybug. Python bin/kb/{get,stats,eval} stays the
runner fallback.
2026-08-13 21:16:48 +01:00
eSliderandGitHub aca05626bd feat: escalate brain search to web when facts cannot confirm (#17)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
2026-08-13 20:24:30 +01:00
eSliderandGitHub 39ae2abe8d feat: Go SearXNG client; throttled is not absence (#16)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
2026-08-13 19:53:31 +01:00
eSliderandGitHub ba5cc3a6e2 feat: read git history with go-git, not the git binary (#15)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
2026-08-13 18:07:56 +01:00
eSliderandGitHub 15d59054ff docs: delete agent-cost; rename kb-search skill to brain (#14)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
* docs: delete agent-cost; rename kb-search skill to brain.

bin/agents/cost does not exist. CI unittest now fails if a SKILL.md names a
missing bin/ path.

* test: gate SKILL.md bin/ paths; name the brain skill brain.

Follow-up to the agent-cost delete: unittest fails if a skill names a missing
tool. Frontmatter name is brain, not kb-search.
2026-08-13 17:55:41 +01:00
eSliderandGitHub 3f30052ea8 feat: in-process HTTP search; /get /stats /audit /ingest. (#13)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
bin/brain/serve.go (ladybug tags) calls internal/brain instead of exec.
HTTP tests inject a fake API so CI stays cgo-free. ExecSearcher remains
the fallback when the binary is built without system_ladybug.
2026-08-13 17:52:15 +01:00
eSliderandGitHub c1ee920b0a feat: brain/index.go shebang; mail import is not a brain write (D14). (#12)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
Commands live at bin/brain/{index,get,stats,eval,watch}.go and
bin/mail/import.go, bin/markdown/import.go, bin/postgres/query.go.
Python remains the Ladybug write worker. index_mail is a deprecation
shim that rebuilds via --with-mail.
2026-08-13 17:46:25 +01:00
eSliderandGitHub cec0161ff6 refactor: chats method shebangs; drop chats index (D14). (#11)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
Parsers and commands live in internal/chats. bin/chats/{sync,import,facts,apply}.go
are tagged shebang mains. Brain ingest is not a chats command.
2026-08-13 17:31:03 +01:00
eSliderandGitHub 46310f8773 docs: name bin/brain/search.go; --hop is not a graph walk. (#10)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
Published docs and skills still taught bin/kb/search --hop 1. Search lives
at bin/brain/search.go; --hop errors until File edges exist. A unittest
gates the SoT so the lie cannot return.
2026-08-13 17:23:40 +01:00
eSliderandGitHub 7e0f3c9e06 feat: bin/brain/serve.go; search backend is Go not Python (#9)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
* feat(brain): HTTP serve from bin/brain/serve.go, default Go search binary.

Move the HTTP package to internal/httpapi. Default backend is
var/bin/brain-search, not Python. bin/serve.go stays as a deprecation shim.

* feat(httpapi): default search backend is var/bin/brain-search.

bin/brain/serve.go is the command; bin/serve.go stays as a tagged
deprecation shim. Tests fail if the default path still names Python.
2026-08-13 14:32:21 +01:00
eSliderandGitHub f14025304e refactor: one Go module; brain search in bin/brain + internal/brain. (#8)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
Collapse nested kbsearch/chats go.mod into the root module. Ranking stays
cgo-free under internal/brain/rank so CI does not need ladybug. bin/kb/search
is a deprecation wrapper that still sets CGO and builds the binary.
2026-08-13 14:26:54 +01:00
eSliderandGitHub dd6d7e9395 docs: point issues at Gitea origin (D15). (#7)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
GitHub stays the public clone for PRs and Actions. Work board is
https://git.produktor.io/eSlider/2dph/issues.
2026-08-13 14:11:00 +01:00
eSliderandGitHub 68d478224f feat(chats): parse LinkedIn MCP v4.22 inbox/conversation blobs. (#6)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
get_inbox/get_conversation return a sections+references envelope, not a
message list. Parser is covered by synthetic Alice/Bob fixtures; CI now
runs the nested bin/chats tests. Session check no longer launches Chromium.
2026-08-13 12:25:51 +01:00
eSliderandGitHub 117f3c2cfd fix(kbsearch): rank FTS correctly, filter before -n, start the daemon. (#5)
Go search took worst BM25 hits (ORDER BY score), cut to -n before --root,
and never called ensureDaemon. Ranking and flag parsing move to a cgo-free
package so CI can fail those regressions without ladybug. --hop errors
instead of being swallowed into the query.
2026-08-13 12:19:26 +01:00
eSliderandGitHub 140d86a4b9 Add Gmail --query to mail/sync (default in:inbox) (#4)
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
* Add --query to Gmail mail/sync instead of always listing in:inbox.

Callers keep the search string; default remains in:inbox.

* Document Gmail --query on the mail/sync pipeline.

* test(mail): assert Gmail --query reaches ListIDs, not only the CLI flag.

ParseCLI coverage left a hole: an empty query still has to become in:inbox
and a custom q has to be the string the client lists with.
2026-08-13 12:19:22 +01:00
eSlider c96c393a4a feat(chats): LinkedIn source — MCP client via get_inbox + get_conversation
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
- LinkedInMCPSource: MCP JSON-RPC, как TelegramMCPSource
- sync linkedin --limit N: выгрузка сообщений из LinkedIn
- Проверка сессии: uvx mcp-server-linkedin --status
- Вывод инструкции если сессия истекла
- JSONL в var/chats/linkedin/<thread_id>/messages.jsonl
2026-08-13 00:10:42 +01:00
eSlider 3d0d95cf00 docs: add edelweiss to GitHub safety rules
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
2026-08-13 00:07:01 +01:00
eSlider ed28fdbd2a chore: remove edelweiss references from public repo
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
2026-08-13 00:06:49 +01:00
eSlider 7e511d5b78 docs: GitHub safety rules — no absolute paths, PII, secrets, curasoft 2026-08-13 00:02:09 +01:00
eSlider 0d26519fab fix: resolve plan.md conflict, remove remaining /mnt/ paths 2026-08-13 00:00:48 +01:00
eSlider a429b823e5 chore: clean absolute paths, curasoft refs, secrets from history
- bin/chats/: env-based paths, no /mnt/ /home/ hardcodes
- bin/edelweiss-pilot: remove curasoft, use DOCS_BASE env var
- bin/facts/crm: use KNOWLEDGE_MESH_SEED env var
- compose.edelweiss.yml: remove curasoft volumes, use DOCS_BASE
- docs/chat-import-plan.md: link to Gitea issue, no secrets
- bin/seed-edelweiss-facts.py: removed (curasoft-only)
2026-08-13 00:00:15 +01:00
eSlider 6847233183 bin/chats: Phase 1 MVP — Telegram sync/import/index/facts/apply
- bin/chats/ — nested Go module (как bin/kbsearch/)
  - sync telegram — MCP JSON-RPC клиент, 31 личный чат, 922 сообщения
  - import — конвертация JSONL → MD с YAML frontmatter
  - index — делегирует bin/kb/index --corpus (132 leafs в brain)
  - facts — regex extraction phone/email/linkedin с валидацией
    (исключены: даты, суммы, номера карт, инвойсы)
  - apply — oo CLI cross-check + dry-run
- Source interface для будущих WhatsApp/LinkedIn
- 4 system tests (import, facts, empty, roundtrip) — синтетические данные
- bin/chat — build+exec wrapper
- docs/chat-import-plan.md — прогресс, пути к env (без секретов)

Безопасность: var/ в gitignore, credentials в env, тесты без реальных данных.
2026-08-12 23:59:29 +01:00
eSlider 6d7638ab73 docs: chat import pipeline plan — link to Gitea issue #1 2026-08-12 23:59:20 +01:00
eSliderandCursor bd1a91dab7 fix(kb): seed facts before CREATE indexes (FTS MERGE corruption)
Upsert under live FTS raises "document for node offset N is missing".
Add --skip-indexes; edelweiss-pilot index = write → seed → ensure_indexes.
Ship seed-edelweiss-facts.py (paired lexicon/OO/interview/QEMU facts).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:09:02 +01:00
eSliderandCursor a5a1f91d95 fix(kb): stop DROP INDEX killing HNSW via Ladybug ghost catalog
Ladybug 0.19 DROP INDEX leaves `_0_Leaf_vec_UPPER` / `0_id_docs` in catalog so
CREATE fails while SHOW_INDEXES omits the index; create_fts_and_vector used to
swallow that. Never drop FTS/VECTOR; ensure_indexes after upserts; rebuild =
delete kb.lbug. Add compose.edelweiss.yml + regression tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 16:05:14 +01:00
eSlider 80e3b7a1cf Remove curasoft references, rename to detective method
- PLAN.md: replace 'curasoft-detective' with 'detective method'
- README.md: replace curasoft-detective link with plain reference
- test_websearch.py: fix test domain from ticket.curasoft.de to example.com
- Rewrote git history with git-filter-repo to remove all traces
2026-08-12 13:46:54 +01:00
eSlider ef4189c72d kbsearch: Go implementation with daemon model serving
- New nested module bin/kbsearch with Go implementation of bin/kb/search
- Embedding model (potion-multilingual-128M) served by localhost daemon
  so repeated CLI calls reuse the loaded model
- Bash launcher bin/kb/search builds binary on first run, caches to var/bin/
- Hybrid FTS + vector search (RRF k=60) matching Python kblib behavior
- YAML output via port of yamlout.py (ordered keys, same format)
- JSON output with proper field order
- All flags: --root, --repo, -n, --json, --list-model
- Root go.mod reverted to 1.25.0 (kbsearch is isolated nested module)
- CI passes: go test ./... and go vet ./... unaffected by kbsearch
2026-08-11 23:57:39 +01:00
eSlider 60c20ed98d feat(mail): full Gmail+OnlyOffice sync, import, and brain indexing
- bin/mail/sync.go: async Go sync engine (8 workers, paginated Gmail via
  API + OnlyOffice IMAP); Gmail attachments key off body.attachmentId, not
  MIME partId; ICS sidecars Latin-1->UTF-8 normalized (TestICSToMarkdownNormalizesLatin1)
- bin/mail/import: message.json -> markdown; PDFs via pdftotext -layout
  fast path with docling subprocess fallback for the ~5% textless files
- bin/mail/index_mail: fresh-rebuild indexer (repo corpus + mail) avoiding
  ladybug WAL corruption on bulk-insert into indexed DBs; split from import
- bin/kb/index: keep FTS/VECTOR indexes across incremental runs (drop+recreate
  leaves stale backing tables killing the vector index)
- docs: README/PLAN/AGENTS cover the mail pipeline

Result: 17,835 messages -> 28,918 info leafs, FTS+HNSW healthy.
2026-08-11 21:57:38 +01:00
eSlider 9f22380e82 refactor(tools): bin/{subject}/{method} layout; Go serve+watch modules
Move serve/ (module) -> bin/server, tools/ -> bin/tools, replace bin/kb-watch
bash with bin/watch Go package; self-executing Go shebangs bin/serve.go and
bin/kb/watch.go; Docker + CI + git/import + docs repointed. Multi-stage image
builds static serve+watch binaries (no Go runtime in container).
2026-08-11 09:52:20 +01:00
eSlider e2eff3b9c7 feat(kb): CRM association proof via oo, fix ssh-tunnel self-ref + oo creds
- bin/facts/crm: prove person<->company/company<->project against ooCRM
  x corpus SoT (knowledge-mesh-seed.yaml), write 78 facts (root=facts)
- tools/crmfacts.py + test_crm_facts.py: parser under unit tests (26 pass)
- docs/crm-associations-proof.md: provable graph, mistakes, fixes
- oo merge 759->763 resolves duplicate GoldenRatio.Exchange legal entity
- bin/db/ssh-tunnel: "$0" self-check + accept-new/BatchMode ssh flags
- AGENTS.md: document bin/facts/crm
2026-08-10 23:22:34 +01:00
18 changed files with 469 additions and 60 deletions
+5 -3
View File
@@ -77,12 +77,14 @@ bin/brain/index.go --rebuild # rebuil
## Tools ## Tools
```bash ```bash
bin/facts/audit ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate bin/facts/audit.go ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
bin/facts/crm [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT) bin/facts/crm.go [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
bin/brain/search.go "query" --no-web # local graph only bin/brain/search.go "query" --no-web # local graph only
bin/brain/get.go <id> [--body] bin/brain/get.go <id> [--body] [--json] # Go read; Python bin/kb/get CI fallback
bin/brain/stats.go [--json]
bin/brain/eval.go [--json] # recall@5; questions in internal/brain/rank
bin/markdown/import.go [dir] # mistune leaves → YAML bin/markdown/import.go [dir] # mistune leaves → YAML
bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
+8 -6
View File
@@ -29,7 +29,7 @@ detective method: **a fact needs ≥2 independent sources or it is
| D3 | web search | Go client `bin/web/search.go` (`internal/websearch`). SearXNG URL is config (`BRAIN_SEARCH_URL`). Optional Compose profile `searxng` (sanitized settings). Do not run a second copy on a host that already has one. Empty/`throttled` ≠ “nothing exists”. | | D3 | web search | Go client `bin/web/search.go` (`internal/websearch`). SearXNG URL is config (`BRAIN_SEARCH_URL`). Optional Compose profile `searxng` (sanitized settings). Do not run a second copy on a host that already has one. Empty/`throttled` ≠ “nothing exists”. |
| D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. | | D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. |
| D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). | | D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). |
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`); Python remains for index/write until the Go write path is safe. | | D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`). Read path (`get.go` / `stats.go` / `eval.go`) is Go + cgo. Python `bin/kb/{get,stats,eval}` is the CI fallback (GitHub runners have no ladybug cgo). Index/write stays Python until the Go write path is safe. |
| D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). | | D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). |
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. | | D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. |
| D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. | | D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. |
@@ -52,11 +52,11 @@ detective method: **a fact needs ≥2 independent sources or it is
docs/ published docs (this conversation → docs/ as md) docs/ published docs (this conversation → docs/ as md)
skills/ in-project skills (web-search, db-yaml, brain, diataxis-docs) skills/ in-project skills (web-search, db-yaml, brain, diataxis-docs)
bin/ bin/
facts/extract auto-pair 2 sources → lexicon yaml + graph facts/extract.go audit.go crm.go # D14 shebang; Python implementation
facts/audit ["self"|"facts"|"info"|"stale"] 2-source + staleness gate
kb/index Python write path (called by bin/brain/index.go) kb/index Python write path (called by bin/brain/index.go)
brain/index.go rebuild FTS + HNSW (incl. --with-mail) brain/index.go rebuild FTS + HNSW (incl. --with-mail)
brain/get.go stats.go eval.go watch.go brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback
brain/watch.go
brain/search.go deduction: facts → info → web-search brain/search.go deduction: facts → info → web-search
brain/serve.go HTTP API in-process (internal/httpapi + internal/brain) brain/serve.go HTTP API in-process (internal/httpapi + internal/brain)
mail/import.go JSON → markdown (no brain write) mail/import.go JSON → markdown (no brain write)
@@ -132,8 +132,10 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
1. go vet + go test ./... (root module; packages without ladybug cgo) 1. go vet + go test ./... (root module; packages without ladybug cgo)
2. `go test ./internal/brain/rank` (cgo-free ranking + flag parser) 2. `go test ./internal/brain/rank` (cgo-free ranking + flag parser)
3. python -m unittest discover -s bin/tools (includes published-docs SoT) 3. python -m unittest discover -s bin/tools (includes published-docs SoT)
4. bin/facts/audit self (lexicon internal consistency) 4. `bin/facts/audit self` (lexicon internal consistency; `bin/facts/audit.go` is the D14 wrapper)
5. bin/brain/eval.go (recall@5 ≥ 0.95, gates index regressions) 5. `bin/kb/eval` (recall@5 ≥ 0.95). Local SoT is `bin/brain/eval.go`; CI uses
the Python twin until the runner has ladybug cgo. Questions live in
`internal/brain/rank`.
6. md-docs build/lint if docs tooling arrives. 6. md-docs build/lint if docs tooling arrives.
Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as
+5 -4
View File
@@ -28,8 +28,8 @@ graph TB
end end
subgraph dph["2dph tools"] subgraph dph["2dph tools"]
EX["bin/facts/extract<br/>2-source pairing"] EX["bin/facts/extract.go<br/>2-source pairing"]
AU["bin/facts/audit<br/>confidence + staleness"] AU["bin/facts/audit.go<br/>confidence + staleness"]
IDX["bin/brain/index.go<br/>chunk + embed"] IDX["bin/brain/index.go<br/>chunk + embed"]
MD["bin/markdown/import.go<br/>mistune leaves"] MD["bin/markdown/import.go<br/>mistune leaves"]
SR["bin/brain/search.go<br/>deduction"] SR["bin/brain/search.go<br/>deduction"]
@@ -126,7 +126,8 @@ bin/brain/search.go "invoice from last week" # same s
- **LadybugDB** — single `var/kb.lbug`, Cypher property graph, HNSW + BM25 - **LadybugDB** — single `var/kb.lbug`, Cypher property graph, HNSW + BM25
in one engine, embedded (no server), ACID, read-only-safe for concurrent in one engine, embedded (no server), ACID, read-only-safe for concurrent
readers. **Never `DROP INDEX` FTS/VECTOR** on Ladybug 0.19: DROP leaves readers. Read tools (`get` / `stats` / `eval`) are Go + cgo; Python
`bin/kb/{get,stats,eval}` is the CI fallback. **Never `DROP INDEX` FTS/VECTOR** on Ladybug 0.19: DROP leaves
ghost catalog tables (`_0_Leaf_vec_UPPER`) so recreate fails while ghost catalog tables (`_0_Leaf_vec_UPPER`) so recreate fails while
`SHOW_INDEXES` omits HNSW. Fresh indexes = delete `var/kb.lbug` + `SHOW_INDEXES` omits HNSW. Fresh indexes = delete `var/kb.lbug` +
`bin/brain/index.go --rebuild`. Use `ensure_indexes()` after upserts. `bin/brain/index.go --rebuild`. Use `ensure_indexes()` after upserts.
@@ -147,7 +148,7 @@ machines. Tests gate every commit. HTTP: `bin/brain/serve.go` calls
```bash ```bash
uv venv .venv # Python 3.12, uv-managed uv venv .venv # Python 3.12, uv-managed
uv pip install -r requirements.lock.txt # pinned toolchain uv pip install -r requirements.lock.txt # pinned toolchain
bin/facts/audit self # lexicon consistency gate bin/facts/audit.go self # lexicon consistency gate
go test ./... && python -m unittest discover -s bin/tools -t . go test ./... && python -m unittest discover -s bin/tools -t .
``` ```
+6 -4
View File
@@ -1,20 +1,22 @@
//usr/bin/env go run -tags=brain_eval "$0" "$@"; exit //usr/bin/env go run -tags=system_ladybug,brain_eval "$0" "$@"; exit
//go:build brain_eval //go:build cgo && system_ladybug && brain_eval
// //
// bin/brain/eval.go - recall@5 gate. // bin/brain/eval.go - recall@5 gate.
// //
// ./bin/brain/eval.go // ./bin/brain/eval.go
// ./bin/brain/eval.go --json // ./bin/brain/eval.go --json
// //
// Needs CGO + libladybug. Python bin/kb/eval is the CI fallback (no cgo).
// Control questions live in internal/brain/rank (cgo-free).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang. // NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main package main
import ( import (
"os" "os"
"github.com/eSlider/2dph/internal/cmdbin" "github.com/eSlider/2dph/internal/brain"
) )
func main() { func main() {
os.Exit(cmdbin.ExecFile("bin/kb/eval", os.Args[1:])) os.Exit(brain.MainEval(os.Args[1:]))
} }
+6 -4
View File
@@ -1,20 +1,22 @@
//usr/bin/env go run -tags=brain_get "$0" "$@"; exit //usr/bin/env go run -tags=system_ladybug,brain_get "$0" "$@"; exit
//go:build brain_get //go:build cgo && system_ladybug && brain_get
// //
// bin/brain/get.go - read one leaf by id. // bin/brain/get.go - read one leaf by id.
// //
// ./bin/brain/get.go <id> // ./bin/brain/get.go <id>
// ./bin/brain/get.go <id> --body // ./bin/brain/get.go <id> --body
// ./bin/brain/get.go <id> --json
// //
// Needs CGO + libladybug. Python bin/kb/get is the CI fallback (no cgo).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang. // NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main package main
import ( import (
"os" "os"
"github.com/eSlider/2dph/internal/cmdbin" "github.com/eSlider/2dph/internal/brain"
) )
func main() { func main() {
os.Exit(cmdbin.ExecFile("bin/kb/get", os.Args[1:])) os.Exit(brain.MainGet(os.Args[1:]))
} }
+5 -4
View File
@@ -1,20 +1,21 @@
//usr/bin/env go run -tags=brain_stats "$0" "$@"; exit //usr/bin/env go run -tags=system_ladybug,brain_stats "$0" "$@"; exit
//go:build brain_stats //go:build cgo && system_ladybug && brain_stats
// //
// bin/brain/stats.go - index health. // bin/brain/stats.go - index health.
// //
// ./bin/brain/stats.go // ./bin/brain/stats.go
// ./bin/brain/stats.go --json // ./bin/brain/stats.go --json
// //
// Needs CGO + libladybug. Python bin/kb/stats is the CI fallback (no cgo).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang. // NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main package main
import ( import (
"os" "os"
"github.com/eSlider/2dph/internal/cmdbin" "github.com/eSlider/2dph/internal/brain"
) )
func main() { func main() {
os.Exit(cmdbin.ExecFile("bin/kb/stats", os.Args[1:])) os.Exit(brain.MainStats(os.Args[1:]))
} }
+21
View File
@@ -0,0 +1,21 @@
//usr/bin/env go run -tags=facts_audit "$0" "$@"; exit
//go:build facts_audit
//
// bin/facts/audit.go - 2-source + lexicon checks.
//
// ./bin/facts/audit.go self
// ./bin/facts/audit.go db
//
// Python bin/facts/audit is the implementation (CI runs it directly).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/audit", os.Args[1:]))
}
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=facts_crm "$0" "$@"; exit
//go:build facts_crm
//
// bin/facts/crm.go - prove person↔company / company↔project (ooCRM × corpus).
//
// ./bin/facts/crm.go [--dry-run] [--mismatches]
//
// Python bin/facts/crm is the implementation. Graph write stays Python.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/crm", os.Args[1:]))
}
+20
View File
@@ -0,0 +1,20 @@
//usr/bin/env go run -tags=facts_extract "$0" "$@"; exit
//go:build facts_extract
//
// bin/facts/extract.go - acquire confirmed facts (2-source each).
//
// ./bin/facts/extract.go [--json] [--dry-run]
//
// Python bin/facts/extract is the implementation. Graph write stays Python.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/facts/extract", os.Args[1:]))
}
+38
View File
@@ -77,6 +77,44 @@ class BinLayoutTest(unittest.TestCase):
for method in ("index.go", "get.go", "stats.go", "eval.go", "watch.go"): for method in ("index.go", "get.go", "stats.go", "eval.go", "watch.go"):
self._assert_shebang(f"bin/brain/{method}") self._assert_shebang(f"bin/brain/{method}")
def test_brain_get_stats_eval_are_not_python_exec(self) -> None:
for method in ("get.go", "stats.go", "eval.go"):
text = (ROOT / "bin" / "brain" / method).read_text()
self.assertNotIn(
"ExecFile",
text,
f"bin/brain/{method} must call internal/brain, not ExecFile Python",
)
self.assertNotIn(
"cmdbin",
text,
f"bin/brain/{method} must not import internal/cmdbin",
)
self.assertIn(
"system_ladybug",
text.splitlines()[0],
f"bin/brain/{method} shebang must pass -tags=system_ladybug",
)
self.assertIn(
"github.com/eSlider/2dph/internal/brain",
text,
)
def test_eval_control_questions_live_in_rank(self) -> None:
rank = (ROOT / "internal" / "brain" / "rank" / "evalq.go").read_text()
py = (ROOT / "bin" / "kb" / "eval").read_text()
for frag in ("BM25", "DevOps", "LadybugDB"):
self.assertIn(frag, rank)
self.assertIn(frag, py)
self.assertIn("0.95", rank)
def test_facts_methods_are_shebangs(self) -> None:
for method in ("audit.go", "extract.go", "crm.go"):
self._assert_shebang(f"bin/facts/{method}")
text = (ROOT / "bin" / "facts" / method).read_text()
self.assertIn("cmdbin.ExecFile", text)
self.assertIn(f"bin/facts/{method.removesuffix('.go')}", text)
def test_mail_import_is_shebang_not_brain_write(self) -> None: def test_mail_import_is_shebang_not_brain_write(self) -> None:
self._assert_shebang("bin/mail/import.go") self._assert_shebang("bin/mail/import.go")
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text() index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
+8
View File
@@ -59,6 +59,14 @@ class PublishedDocsTest(unittest.TestCase):
self.assertNotIn("password", settings.lower()) self.assertNotIn("password", settings.lower())
self.assertIn("json", settings) self.assertIn("json", settings)
def test_readme_read_path_is_go(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("get.go", plan)
self.assertIn("CI fallback", plan)
design = (ROOT / "docs" / "design.md").read_text()
self.assertIn("internal/brain/rank", design)
self.assertIn("They do not exec Python", design)
def test_readme_search_escalates_web(self) -> None: def test_readme_search_escalates_web(self) -> None:
text = (ROOT / "README.md").read_text() text = (ROOT / "README.md").read_text()
self.assertIn("--no-web", text) self.assertIn("--no-web", text)
-33
View File
@@ -1,33 +0,0 @@
# CRM association proof (oo CLI ↔ corpus)
Proven with `oo` (eslider/go-onlyoffice) against the OnlyOffice portal
(`office.produktor.io`). Portal CRM is the SSOT for company ↔ person ↔
project associations; the corpus SoT (`eslider/cv/projects/knowledge-mesh-seed.yaml`)
is the second, independent source. Facts that can be backed by both are
written to the brain under `root=facts` by `bin/facts/crm`.
## What was verified
- Logical counts (portal MySQL): 1300 contacts = 897 persons + 404 companies,
198 projects, 998 deals, 939 project↔contact links.
- Every client company linked to a project has ≥1 person underneath.
- Every person `company_id` resolves to an existing company.
- Corpus org list (9) maps 1:1 onto CRM companies:
ProProdukt SL / produktor.io, Dyvenia, Immowelt AG, WhereGroup,
Keynote SIGOS, D2S/SYSTEMS, GRID, Pack und Cup, Markets Platform.
- 78 person↔company association facts written to the brain
(`how=crm-crosscheck`, `type=association`). Recall@5 in `bin/kb/eval` = 1.0.
## Mistakes found
| # | Mistake | Fix |
|---|---------|-----|
| 1 | Duplicate legal entity `GoldenRatio.Exchange` (contact 759) vs `Golden Ratio Exchange` (763); 3 deals (211, 287, 559) were linked to 759 | `oo contacts merge 759 763` — 763 kept, 759 removed, deal links re-pointed to 763 |
| 2 | `env/`-wide: OnlyOffice creds file used wrong UX (user `eslider`, password with `$2` suffix) making `oo` auth fail | `.env` fixed to `eslider@gmail.com` + clean password; `.env` stays gitignored |
## Gates after fix
- `uv run python -m unittest discover -s bin/tools -t .` → 26 tests OK
- `bin/facts/audit self` + `bin/facts/audit db` → ok
- `bin/kb/eval` → recall@5 = 1.0
- `go test ./...` (bin/server + bin/watch) → ok
+8
View File
@@ -62,3 +62,11 @@ corpus HEAD.
Confirmed = A×B or B×C agreement. Single source = hypothesis + `(not confirmed)`. Confirmed = A×B or B×C agreement. Single source = hypothesis + `(not confirmed)`.
Conflicting pairings (≥2 yes vs ≥2 no) = hypothesis (OQ1 → v2 resolution). Conflicting pairings (≥2 yes vs ≥2 no) = hypothesis (OQ1 → v2 resolution).
## Read path
`bin/brain/get.go`, `stats.go`, and `eval.go` call `internal/brain` with cgo
(`system_ladybug`). They do not exec Python. Control questions for recall@5
live in `internal/brain/rank` so CI can test the table without libladybug.
Python `bin/kb/{get,stats,eval}` remain for GitHub Actions until the runner
has ladybug cgo. Index/write is still `bin/kb/index`.
+3
View File
@@ -0,0 +1,3 @@
package brain
const ModelID = "minishlab/potion-multilingual-128M"
+16
View File
@@ -0,0 +1,16 @@
package rank
// Eval control questions (recall@5). Kept here so CI can test the gate
// table without ladybug cgo. The runner lives in internal/brain (cgo).
const EvalRecallThreshold = 0.95
type EvalQuestion struct {
Query string
Fragment string
}
var EvalQuestions = []EvalQuestion{
{"hybrid search fts and vector", "BM25"},
{"eslider devops engineer", "DevOps"},
{"ladybugdb graph engine storage", "LadybugDB"},
}
+17
View File
@@ -0,0 +1,17 @@
package rank
import "testing"
func TestEvalQuestionsAreThreeAndThreshold(t *testing.T) {
if EvalRecallThreshold != 0.95 {
t.Fatalf("threshold = %v", EvalRecallThreshold)
}
if len(EvalQuestions) != 3 {
t.Fatalf("questions = %d, want 3", len(EvalQuestions))
}
for _, q := range EvalQuestions {
if q.Query == "" || q.Fragment == "" {
t.Fatalf("empty control: %+v", q)
}
}
}
+281
View File
@@ -0,0 +1,281 @@
//go:build cgo && system_ladybug
package brain
import (
"encoding/json"
"fmt"
"os"
"sort"
"strings"
"unicode/utf8"
"github.com/eSlider/2dph/internal/brain/rank"
)
func MainGet(args []string) int {
id, body, jsonOut := "", false, false
for _, a := range args {
switch {
case a == "--body":
body = true
case a == "--json":
jsonOut = true
case a == "-h" || a == "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/get.go <id> [--body] [--json]`)
return 0
case strings.HasPrefix(a, "-"):
fmt.Fprintf(os.Stderr, "brain/get: unknown flag %s\n", a)
return 2
default:
id = a
}
}
if id == "" {
fmt.Fprintln(os.Stderr, "brain/get: id required")
return 2
}
if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1
}
defer closeBrain()
meta, text, err := lookupLeaf(id)
if err != nil {
fmt.Fprintf(os.Stderr, "brain/get: %v\n", err)
return 1
}
out := Dict{
{"id", meta["id"]},
{"root", meta["root"]},
{"confidence", meta["confidence"]},
{"source", meta["source"]},
{"type", meta["type"]},
}
if body {
out = append(out, KV{"text", text})
} else {
out = append(out, KV{"snippet", clip(text, 280)})
}
if jsonOut {
m := map[string]any{}
for _, kv := range out {
m[kv.K] = kv.V
}
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
return b2i(enc.Encode(m))
}
fmt.Print(toYAML(out, 0))
return 0
}
func MainStats(args []string) int {
jsonOut := false
for _, a := range args {
switch a {
case "--json":
jsonOut = true
case "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/stats.go [--json]`)
return 0
default:
if strings.HasPrefix(a, "-") {
fmt.Fprintf(os.Stderr, "brain/stats: unknown flag %s\n", a)
return 2
}
}
}
if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1
}
defer closeBrain()
s, err := leafStats()
if err != nil {
fmt.Fprintf(os.Stderr, "brain/stats: %v\n", err)
return 1
}
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
return b2i(enc.Encode(s))
}
by := s["by_root"].(map[string]int)
keys := make([]string, 0, len(by))
for k := range by {
keys = append(keys, k)
}
sort.Strings(keys)
byRoot := make(Dict, 0, len(keys))
for _, k := range keys {
byRoot = append(byRoot, KV{k, by[k]})
}
out := Dict{
{"total", s["total"]},
{"by_root", byRoot},
{"db", s["db"]},
{"model", s["model"]},
}
fmt.Print(toYAML(out, 0))
return 0
}
func MainEval(args []string) int {
jsonOut := false
for _, a := range args {
switch a {
case "--json":
jsonOut = true
case "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/brain/eval.go [--json]`)
return 0
default:
if strings.HasPrefix(a, "-") {
fmt.Fprintf(os.Stderr, "brain/eval: unknown flag %s\n", a)
return 2
}
}
}
if err := openBrain(); err != nil {
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
return 1
}
defer closeBrain()
recalled := 0
details := make([]any, 0, len(rank.EvalQuestions))
jsDetails := make([]map[string]any, 0, len(rank.EvalQuestions))
for _, q := range rank.EvalQuestions {
hits, err := queryFTS(q.Query, 5)
ok := false
if err == nil {
frag := strings.ToLower(q.Fragment)
for _, h := range hits {
if strings.Contains(strings.ToLower(h.Text), frag) {
ok = true
break
}
}
}
if ok {
recalled++
}
details = append(details, Dict{
{"q", q.Query},
{"fragment", q.Fragment},
{"in_top5", ok},
})
jsDetails = append(jsDetails, map[string]any{
"q": q.Query, "fragment": q.Fragment, "in_top5": ok,
})
}
n := len(rank.EvalQuestions)
recall := 0.0
if n > 0 {
recall = float64(recalled) / float64(n)
}
passed := recall >= rank.EvalRecallThreshold
if jsonOut {
enc := json.NewEncoder(os.Stdout)
enc.SetIndent("", " ")
enc.SetEscapeHTML(false)
_ = enc.Encode(map[string]any{
"recall@5": round3(recall),
"passed": passed,
"gate": n,
"details": jsDetails,
})
} else {
out := Dict{
{"recall@5", round3(recall)},
{"passed", passed},
{"gate", n},
{"details", details},
}
fmt.Print(toYAML(out, 0))
}
if !passed {
return 2
}
return 0
}
func lookupLeaf(id string) (map[string]string, string, error) {
if conn == nil {
return nil, "", fmt.Errorf("brain not open")
}
stmt, err := conn.Prepare(
"MATCH (l:Leaf {id:$id}) RETURN l.id, l.text, l.root, l.confidence, l.source, l.type",
)
if err != nil {
return nil, "", err
}
defer stmt.Close()
res, err := conn.Execute(stmt, map[string]any{"id": id})
if err != nil {
return nil, "", err
}
if !res.HasNext() {
return nil, "", fmt.Errorf("no leaf %s", id)
}
row, err := res.Next()
if err != nil {
return nil, "", err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 6 {
return nil, "", fmt.Errorf("leaf row")
}
meta := map[string]string{
"id": fmt.Sprint(vals[0]),
"root": fmt.Sprint(vals[2]),
"confidence": fmt.Sprint(vals[3]),
"source": fmt.Sprint(vals[4]),
"type": fmt.Sprint(vals[5]),
}
return meta, fmt.Sprint(vals[1]), nil
}
func leafStats() (map[string]any, error) {
if conn == nil {
return nil, fmt.Errorf("brain not open")
}
res, err := conn.Query("MATCH (l:Leaf) RETURN l.root, count(*)")
if err != nil {
return nil, err
}
byRoot := map[string]int{}
total := 0
for res.HasNext() {
row, err := res.Next()
if err != nil {
return nil, err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 2 {
continue
}
n := int(asInt(vals[1]))
byRoot[fmt.Sprint(vals[0])] = n
total += n
}
return map[string]any{
"total": total,
"by_root": byRoot,
"db": dbPath(),
"model": ModelID,
}, nil
}
func clip(s string, n int) string {
if utf8.RuneCountInString(s) <= n {
return s
}
return string([]rune(s)[:n])
}
func round3(f float64) float64 {
return float64(int(f*1000+0.5)) / 1000
}
+1 -1
View File
@@ -28,7 +28,7 @@ bin/brain/search.go "onlyoffice postgres" --root facts # restrict to confirmed
bin/brain/search.go "where is cs-lexicon" --json | yq '.[].ref' bin/brain/search.go "where is cs-lexicon" --json | yq '.[].ref'
bin/brain/get.go <id> --body # full chunk only when needed bin/brain/get.go <id> --body # full chunk only when needed
bin/brain/stats.go # index health bin/brain/stats.go # index health
bin/brain/eval.go # recall@5 >= 0.95 gate bin/brain/eval.go # recall@5 >= 0.95 gate (Go; Python bin/kb/eval is CI fallback)
``` ```
`bin/kb/search` is a deprecated wrapper. `--hop` errors (File/FROM_FILE edges `bin/kb/search` is a deprecated wrapper. `--hop` errors (File/FROM_FILE edges