Compare commits

...
6 Commits
Author SHA1 Message Date
eSlider 2332613c43 ci: point eval at the checkout and keep DevOps in the corpus.
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
Tests / Test (pull_request) Failing after 39s
Tests / Release (semver) (pull_request) Skipped
/tmp/brain-eval cannot walk up to var/kb.lbug; KB_ROOT is required.
Recall fragments must exist in README/docs so a repo-only rebuild gates.
2026-08-14 11:00:39 +01:00
eSlider 8feb93da78 ci: recall@5 SoT is Zig bin/brain/eval.go.
Tests / Test (push) Skipped
Tests / Release (semver) (push) Skipped
GitHub Actions rebuilds the repo corpus and runs the Go eval gate
instead of Python bin/kb/eval. Audit self is a hard fail (Gitea #19).
2026-08-14 10:58:12 +01:00
eSliderandGitHub c5be3f19be feat: index facts extract and chats markdown on rebuild. (#31)
Tests / Test (push) Failing after 5s
Tests / Release (semver) (push) Skipped
--with-facts / --facts-json write root=facts leafs; --with-chats
picks up var/chats/md as info. WhatsApp sync stays out of v1 (Gitea #18).
2026-08-14 10:57:40 +01:00
eSliderandGitHub a88dbb490c feat: search --hop walks FROM_FILE to Person. (#30)
Tests / Test (push) Failing after 4s
Tests / Release (semver) (push) Skipped
Parser no longer errors; hop 1 returns File, hop 3 reaches Person.
Rebuild writes Leaf-[:FROM_FILE]->File so the walk is not empty
on a fresh index (Gitea #17).
2026-08-14 10:51:35 +01:00
eSliderandGitHub 9d1a3f3c70 feat: write leafs incrementally without rebuilding the graph. (#29)
Tests / Test (push) Failing after 4s
Tests / Release (semver) (push) Skipped
Ladybug 0.19 stays FTS/HNSW queryable on MERGE of new ids; DROP INDEX
was the fatal path. bin/brain/add.go and POST /ingest land facts+info
in one transaction so watch/mail/git can become leafs now (Gitea #14).
2026-08-14 10:46:59 +01:00
eSliderandGitHub 8b5be9b659 docs: record v1 detective epic and remaining gaps. (#28)
Tests / Test (push) Failing after 4s
Tests / Release (semver) (push) Skipped
Read path and MCP are in; write, hops, corpus, and CI eval are still open. Board is Gitea epic #16.
2026-08-14 10:32:19 +01:00
34 changed files with 1045 additions and 115 deletions
+13 -3
View File
@@ -54,13 +54,23 @@ jobs:
run: go test ./internal/brain/rank -count=1
- name: facts/audit self (lexicon consistency, no network)
run: |
./bin/facts/audit self 2>/dev/null || echo "audit: not yet implemented; gate skipped"
run: ./bin/facts/audit self
- name: CGO via Zig (compile brain/search)
- name: CGO via Zig (compile brain/search + eval)
run: |
chmod +x bin/cgo/zig bin/cgo/zcc bin/cgo/zc++
bin/cgo/zig go build -tags system_ladybug -o /tmp/brain-search ./bin/brain/search.go
bin/cgo/zig go build -tags 'system_ladybug,brain_eval' -o /tmp/brain-eval ./bin/brain/eval.go
- uses: actions/cache@v4
with:
path: ~/.cache/huggingface
key: ${{ runner.os }}-hf-potion-multilingual-128M
- name: recall@5 SoT (Zig bin/brain/eval.go)
run: |
uv run python bin/kb/index --rebuild --json
KB_ROOT="$PWD" /tmp/brain-eval --json
release:
name: Release (semver)
+10 -6
View File
@@ -3,7 +3,8 @@
Evidence-first brain over the ops/eSlider stack. Facts need proof or they are
`(not confirmed)`.
Read first: [PLAN](PLAN.md) → [docs](docs/).
Read first: [PLAN](PLAN.md) → [docs](docs/) → [roadmap](docs/roadmap.md)
(epic [#16](https://git.produktor.io/eSlider/2dph/issues/16)).
## Method (detective, no fork)
@@ -39,7 +40,7 @@ PLAN.md decisions + execution + open questions
docs/ published docs
skills/ in-project agent skills (vendored, no external links)
bin/ self-describing tools bin/{subject}/{method}.go (shebang)
bin/brain/ search.go serve.go index.go get.go stats.go eval.go watch.go
bin/brain/ search.go serve.go index.go add.go get.go stats.go eval.go watch.go
bin/chats/ sync.go import.go facts.go apply.go; libs in internal/chats
bin/mail/ sync.go import.go (index_mail → brain/index.go)
bin/markdown/ import.go (H2 leaf split; Python bin/md/import fallback)
@@ -64,7 +65,7 @@ var/ kb.lbug, var/mail/*, caches (gitignored)
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
bin/mail/sync.go --source gmail --query 'from:example.com' --out var/mail # Gmail search (default in:inbox)
bin/mail/import.go --from-raw var/mail # message.json → message.md (convert only)
bin/brain/index.go --rebuild # rebuild brain incl. all mail (fresh DB)
bin/brain/index.go --rebuild --with-facts --with-chats
```
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
@@ -73,9 +74,9 @@ bin/brain/index.go --rebuild # rebuil
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs fall back to
docling (isolated subprocess — its native onnx can segfault the parent).
Conversion never touches the brain DB (crash safety).
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Ladybug
corrupts its WAL when brand-new leafs are bulk-inserted while FTS/vector
indexes exist; a fresh DB with indexes created last is the only safe path.
- `index_mail` is a deprecation shim for `bin/brain/index.go --rebuild`. Bulk
rebuild still deletes `var/kb.lbug` and creates FTS/HNSW last. Single-leaf
write is `bin/brain/add.go` (safe while indexes exist; do not DROP INDEX).
Keep conversion + indexing separate so a conversion crash can't leave the
DB mid-transaction.
@@ -88,6 +89,9 @@ bin/kb/search "query" [--repo X] # deprecated wrapper → bin/b
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
bin/brain/search.go "query" --no-web # local graph only
eval "$(bin/cgo/zig env)" # Zig cc + liblbug (not gcc)
bin/brain/index.go --rebuild [--with-mail] [--with-facts] [--with-chats]
bin/brain/add.go --text T --root facts --source "a.md x b.md" # incremental write
bin/brain/add.go --json # stdin leaf or {leafs:[...]}
bin/brain/get.go <id> [--body] [--json] # Go read; Python bin/kb/get CI fallback
bin/brain/stats.go [--json]
bin/brain/eval.go [--json] # recall@5; questions in internal/brain/rank
+37 -12
View File
@@ -4,7 +4,9 @@ A brain that loves facts and deduction. Evidence-first knowledge graph + hybrid
RAG over the operational Brain/ops/eSlider stack. Built like Sherlock
Holmes: nothing is asserted unless it has proof.
Status: **in progress**this file is the plan and the record of decisions.
Status: **in progress**read path + MCP work; v1 goal is [epic #16](https://git.produktor.io/eSlider/2dph/issues/16)
(milestone [v1 detective brain](https://git.produktor.io/eSlider/2dph/milestone/12)).
Gap: [docs/roadmap.md](docs/roadmap.md).
## What
@@ -29,9 +31,9 @@ detective method: **a fact needs ≥2 independent sources or it is
| D3 | web search | Go client `bin/web/search.go` (`internal/websearch`). SearXNG URL is config (`BRAIN_SEARCH_URL`). Optional Compose profile `searxng` (sanitized settings). Do not run a second copy on a host that already has one. Empty/`throttled` ≠ “nothing exists”. |
| D4 | embeddings | **model2vec** `minishlab/potion-multilingual-128M` instead of embeddinggemma. |
| D5 | parser | **mistune** for MD → leaf extraction (duckdb-md documented as future optional SQL/export layer, not v1). |
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`). Read path is Go + Zig CGO (D21). Python `bin/kb/{get,stats,eval}` is the CI fallback when Zig/libs are not fetched. Index/write stays Python (`compose --profile index`) until the Go write path is safe. |
| D6 | graph engine | **LadybugDB**. Go is the service (`bin/brain/search.go`, `bin/brain/serve.go` in-process, `internal/brain`). Read path is Go + Zig CGO (D21). Python `bin/kb/{get,stats,eval}` is the CI fallback when Zig/libs are not fetched. Incremental write is Python `bin/kb/add` (`bin/brain/add.go`). Bulk rebuild stays `compose --profile index` until the Go write path is safe. |
| D7 | db access | `db-yaml`/`psql-yq`-style, read-only, YAML out. OnlyOffice Postgres via SSH tunnel (`127.0.0.1:5433`). |
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. Auto-pair docker ps × compose × ssh-config × docs. |
| D8 | evidence | detective method: ≥2 independent sources or `(not confirmed)`. 2-source auto-pair docker ps × compose × ssh-config × docs. |
| D9 | facts/goal model | Who / What / How / Where / When + evidence + confidence on every edge. |
| D10 | versioning | everything is a leaf with `sha256 + observed_at + source_rev`; `File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`. Stale = `source_rev` < git HEAD. |
| D11 | strong/weak | `root` column: `facts` (strong) vs `info` (weak). Answer is `confirmed` only from facts root. |
@@ -41,9 +43,9 @@ detective method: **a fact needs ≥2 independent sources or it is
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. |
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
| D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is not shipped; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. |
| D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is compose profile `picoclaw`; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. Agent lever/loop: [#15](https://git.produktor.io/eSlider/2dph/issues/15). |
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
| D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`). |
| D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`/`ingest`). |
| D21 | CGO | Ladybug/tokenizers CGO is compiled with **Zig** (`bin/cgo/zcc``zig cc -target …-linux-gnu`), not gcc. `bin/cgo/zig` pins Zig 0.14.1 + liblbug 0.19.1 + libtokenizers 1.27.0. Compose `target: api` has no CPython; write/rebuild is profile `index`. |
## Architecture
@@ -55,8 +57,10 @@ detective method: **a fact needs ≥2 independent sources or it is
skills/ in-project skills (web-search, postgres, brain, picoclaw, diataxis-docs)
bin/
facts/extract.go audit.go crm.go # D14 shebang; Python implementation
kb/index Python write path (called by bin/brain/index.go)
kb/index Python bulk write (called by bin/brain/index.go)
kb/add Python incremental write (called by bin/brain/add.go)
brain/index.go rebuild FTS + HNSW (incl. --with-mail)
brain/add.go incremental leaf write (no rebuild)
brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback
brain/watch.go
brain/search.go deduction: facts → info → web-search
@@ -83,7 +87,10 @@ detective method: **a fact needs ≥2 independent sources or it is
Node tables: `Person, Service, Host, Container, Repo, File, Commit, Leaf`.
`Leaf(embedding FLOAT[N])` — FTS on `text`, HNSW vector index on `embedding`.
Edges: `RUNS / USES / HAS_VERSION / AUTHORED / ABOUT / ASSOCIATED / SIMILAR_0.85`.
Edges: `RUNS / USES / FROM_FILE / HAS_VERSION / AUTHORED / ABOUT / ASSOCIATED / SIMILAR_0.85`.
`FROM_FILE` / `HAS_VERSION` / `AUTHORED`: `bin/brain/search.go --hop N` walks
them from each hit (1=File, 2=Commit, 3=Person). Rebuild writes
`Leaf-[:FROM_FILE]->File`; git import writes the rest.
Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
`where`, `when`, `source_rev`.
@@ -108,9 +115,9 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
## Open questions (v2)
- OQ1: mutually-contradicting evidence — how to resolve (authority weighting,
temporal freshness, audit adjudication).
- OQ2: OCR pipeline for pdfs/images/docs — mostly solved: poppler pdftotext
fast-path for born-digital PDFs, docling fallback for the ~5% textless ones.
temporal freshness, audit adjudication). **v2**; does not block epic #16.
- OQ2: OCR — poppler `pdftotext` fast-path exists; scans still docling.
[#6](https://git.produktor.io/eSlider/2dph/issues/6) (v2, does not block #16).
- OQ3: optional duckdb-md layer for `SELECT … FORMAT MARKDOWN` export/write-back.
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
serialize and unambiguous; YAML only where humans edit files.
@@ -137,7 +144,8 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
2. `go test ./internal/brain/rank` (cgo-free ranking + flag parser)
3. python -m unittest discover -s bin/tools (includes published-docs SoT)
4. `bin/facts/audit self` (lexicon internal consistency; `bin/facts/audit.go` is the D14 wrapper)
5. `bin/kb/eval` (recall@5 ≥ 0.95). Local SoT is `bin/brain/eval.go` via Zig CGO.
5. `bin/brain/eval.go` via Zig (recall@5 ≥ 0.95). Python `bin/kb/eval` is an
explicit fallback, not the CI SoT.
6. `bin/cgo/zig go build -tags system_ladybug` (compile search with zig cc; fetches pinned zig+libs).
Feedback loop: every commit → PR → CI → green/gate → merge. Same discipline as
@@ -151,5 +159,22 @@ Feedback loop: every commit → PR → CI → green/gate → merge. Same discipl
4. .venv: ladybug + model2vec + mistune
5. schema + tools with TDD (kb + md + facts + brain)
6. ~/.config/brain config
7. corpus extraction (facts/info)
7. corpus extraction (facts/info)**in**: [#18](https://git.produktor.io/eSlider/2dph/issues/18)
8. verify: web-search smoke, onlyoffice pg, md-db round-trip, eval, audit
## Gap to v1 (epic #16)
Remaining: none for epic #16 (v1). Board:
[epic #16](https://git.produktor.io/eSlider/2dph/issues/16),
milestone [v1 detective brain](https://git.produktor.io/eSlider/2dph/milestone/12).
Narrative: [docs/roadmap.md](docs/roadmap.md).
| Order | Issue | Gap |
|-------|-------|-----|
| 1 | [#14](https://git.produktor.io/eSlider/2dph/issues/14) | **in**`bin/brain/add.go` / `POST /ingest` write facts+info without deleting `kb.lbug`. Bulk corpus still `--rebuild`. Leftover Python (mail/facts) is not the living-graph blocker. |
| 2 | [#17](https://git.produktor.io/eSlider/2dph/issues/17) | **in**`--hop N` walks `FROM_FILE``HAS_VERSION``AUTHORED` (max 3). |
| 3 | [#18](https://git.produktor.io/eSlider/2dph/issues/18) | **in**`--with-facts` / `--facts-json` land `root=facts`; `--with-chats` indexes `var/chats/md`. WhatsApp sync is out of v1. |
| 4 | [#15](https://git.produktor.io/eSlider/2dph/issues/15) | **in** — lever/loop documented (`search``get``audit`). |
| 5 | [#19](https://git.produktor.io/eSlider/2dph/issues/19) | **in** — CI recall SoT is `bin/brain/eval.go` via Zig. Python `bin/kb/eval` stays as an explicit fallback. |
Does **not** block epic close: [#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR, OQ1, OQ3, OQ4.
+13 -5
View File
@@ -96,7 +96,7 @@ bin/brain/stats.go # index health
bin/brain/eval.go # recall@5 gate
```
`--hop` is not implemented (needs File/FROM_FILE edges); the flag errors instead of walking. `bin/kb/search` is a deprecated wrapper around `bin/brain/search.go`.
`--hop N` walks File/Commit/Person from each hit (max 3). `bin/kb/search` is a deprecated wrapper around `bin/brain/search.go`.
Git history is read with [go-git](https://github.com/go-git/go-git) (no git binary):
@@ -120,6 +120,8 @@ Mail is a first-class corpus (retrievable through the same search):
```bash
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
bin/mail/import.go --from-raw var/mail # JSON → markdown
bin/brain/add.go --text T --root facts --source "a.md x b.md"
bin/brain/index.go --rebuild --with-facts --with-chats # facts extract + chats md
bin/brain/index.go --rebuild # rebuild brain (incl. mail)
bin/brain/search.go "invoice from last week" # same search over mail leafs
```
@@ -128,8 +130,9 @@ bin/brain/search.go "invoice from last week" # same s
- **LadybugDB** — single `var/kb.lbug`, Cypher + HNSW + BM25, embedded.
Read tools (`get` / `stats` / `eval`) are Go + Zig CGO (`bin/cgo/zcc`).
Python fallbacks stay for CI until the runner fetches Zig. Write is
Compose profile `index` (`bin/brain/index.go`).
Python fallbacks stay for CI until the runner fetches Zig. Incremental
write is `bin/brain/add.go` (Python `kblib.add_leafs`). Bulk rebuild is
Compose profile `index` (`bin/brain/index.go --rebuild`).
- **model2vec** — `potion-multilingual-128M` (256-dim), CPU, no Ollama
runtime dependency.
- facts and info split by `root` but written in the same transaction.
@@ -166,13 +169,18 @@ docker compose up brain-watch # auto re-index on change
## Related
eSlider DevOps engineer practice: ops, OnlyOffice, and mail feed the facts
root through `bin/facts/extract` (two-source pairing).
- [go-second-brain](https://github.com/eSlider/go-second-brain) — the earlier
Neo4j + Qdrant + Matrix RAG brain
- [agent-skills](https://github.com/eSlider/agent-skills) — upstream
skills (`web-search`, `postgres`, …) that 2dph integrates
- detective method — the two-source method
Work board (issues): [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
Work board (issues): [epic #16](https://git.produktor.io/eSlider/2dph/issues/16)
on [git.produktor.io/eSlider/2dph/issues](https://git.produktor.io/eSlider/2dph/issues).
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
See [PLAN.md](PLAN.md) for decisions, execution status, and v2 open questions.
See [PLAN.md](PLAN.md) for decisions, [docs/roadmap.md](docs/roadmap.md) for
the gap to v1, and v2 open questions.
+21
View File
@@ -0,0 +1,21 @@
//usr/bin/env go run -tags=brain_add "$0" "$@"; exit
//go:build brain_add
//
// bin/brain/add.go - incremental leaf write (Python kblib, no rebuild).
//
// ./bin/brain/add.go --text T --root facts --source "a.md x b.md"
// ./bin/brain/add.go --json
//
// D6: write stays Python. Does not delete var/kb.lbug.
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
import (
"os"
"github.com/eSlider/2dph/internal/cmdbin"
)
func main() {
os.Exit(cmdbin.ExecFile("bin/kb/add", os.Args[1:]))
}
+3 -3
View File
@@ -3,12 +3,12 @@
//
// bin/brain/index.go - rebuild the Ladybug graph (Python write path).
//
// ./bin/brain/index.go --rebuild
// ./bin/brain/index.go --rebuild --with-facts --with-chats
// ./bin/brain/index.go --rebuild --with-mail
// ./bin/brain/index.go --dry-run --with-mail
//
// v1 write is always a rebuild when mail is included (live FTS/HNSW + bulk
// insert corrupts Ladybug 0.19 WAL). `add` is v2.
// v1 write: bin/brain/add.go for one/few leafs (indexes may already exist).
// Bulk mail/corpus still --rebuild (fresh file, indexes last).
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
package main
+1 -1
View File
@@ -3,7 +3,7 @@
//
// bin/brain/search.go - deduction search over the 2dph brain.
//
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--json] [--no-web]
// ./bin/brain/search.go "query" [--root facts|info] [--repo P] [-n N] [--hop N] [--json] [--no-web]
// ./bin/brain/search.go serve [port]
// ./bin/brain/search.go --list-model
//
+3 -2
View File
@@ -29,10 +29,11 @@ func main() {
case "linkedin":
os.Exit(chats.RunSyncLinkedIn(args))
case "whatsapp":
fmt.Fprintln(os.Stderr, "chats: WhatsApp not implemented yet")
fmt.Fprintln(os.Stderr, "chats: WhatsApp sync is out of v1")
os.Exit(1)
case "help", "-h", "--help":
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]`)
fmt.Fprintln(os.Stderr, `usage: bin/chats/sync.go telegram|linkedin [flags]
WhatsApp sync is out of v1.`)
return
default:
fmt.Fprintf(os.Stderr, "chats: unknown platform %q\n", platform)
Executable
+114
View File
@@ -0,0 +1,114 @@
#!/usr/bin/env python3
"""kb/add - incremental leaf write (no rebuild).
bin/kb/add --text T --root facts|info --source S
bin/kb/add --json # stdin: one object or {"leafs":[...]}
bin/kb/add --db PATH --json
Writes facts+info in one Ladybug transaction. Does not delete kb.lbug.
Embedding is used when provided; otherwise model2vec encodes the text.
"""
from __future__ import annotations
import json
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import ( # noqa: E402
EMBED_DIM,
add_leafs,
connect,
ensure_indexes,
init_schema,
)
def _as_leafs(payload: object) -> list[dict]:
if isinstance(payload, list):
return [dict(x) for x in payload]
if isinstance(payload, dict):
if "leafs" in payload:
return [dict(x) for x in payload["leafs"]]
return [dict(payload)]
raise ValueError("json must be an object, a list, or {leafs:[...]}")
def _embed_missing(leafs: list[dict]) -> None:
missing = [lf for lf in leafs if not lf.get("embedding")]
if not missing:
return
from model2vec import StaticModel
model = StaticModel.from_pretrained("minishlab/potion-multilingual-128M")
for lf in missing:
text = str(lf.get("text") or "")
vec = model.encode([text])[0].astype(float).tolist()
if len(vec) != EMBED_DIM:
vec = (vec + [0.0] * EMBED_DIM)[:EMBED_DIM]
lf["embedding"] = vec
def main(argv: list[str]) -> int:
import argparse
p = argparse.ArgumentParser(description="add leafs without rebuilding the brain")
p.add_argument("--db", default="", help="path to kb.lbug (default var/kb.lbug)")
p.add_argument("--json", action="store_true", help="read leaf JSON from stdin")
p.add_argument("--text", default="", help="leaf text")
p.add_argument("--root", default="info", choices=("facts", "info"))
p.add_argument("--source", default="")
p.add_argument("--confidence", default="confirmed")
p.add_argument("--source-rev", default="working-tree")
p.add_argument("--how", default="brain/add")
p.add_argument("--loc", default="")
p.add_argument("--type", default="reference", dest="type_")
args = p.parse_args(argv)
if args.json:
raw = sys.stdin.read()
if not raw.strip():
print("kb/add: empty stdin", file=sys.stderr)
return 2
leafs = _as_leafs(json.loads(raw))
else:
if not args.text or not args.source:
print("kb/add: --text and --source are required (or --json)", file=sys.stderr)
return 2
leafs = [{
"text": args.text,
"root": args.root,
"source": args.source,
"confidence": args.confidence,
"source_rev": args.source_rev,
"how": args.how,
"loc": args.loc or args.source,
"type": args.type_,
}]
for lf in leafs:
if not lf.get("text") or not lf.get("source"):
print("kb/add: each leaf needs text and source", file=sys.stderr)
return 2
_embed_missing(leafs)
from kblib import DB_PATH, VAR
dbpath = Path(args.db) if args.db else DB_PATH
dbpath.parent.mkdir(parents=True, exist_ok=True)
VAR.mkdir(exist_ok=True)
db, conn = connect(dbpath, read_only=False)
init_schema(conn)
ids = add_leafs(conn, leafs)
ensure_indexes(conn)
conn.close()
db.close()
print(json.dumps({"mode": "add", "ids": ids, "db": str(dbpath)}))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+98 -14
View File
@@ -2,12 +2,15 @@
"""kb/index - build the 2dph brain from markdown + factual leafs.
bin/kb/index [--corpus DIR] [--rebuild] [--limit N]
bin/kb/index --rebuild --with-facts --with-chats
bin/kb/index --json # emit stats as JSON
Reads every .md under the corpus (default: repo root docs, skills, READMEs)
as `info` leafs, embeds them with model2vec (potion-multilingual-128M), and
writes them into var/kb.lbug with FTS + HNSW indexes. `facts` leafs come
from bin/facts/extract (docker x compose x ssh-config pairing).
from bin/facts/extract (docker × compose × ssh-config pairing) when
`--with-facts` is set. `--with-chats` indexes markdown under var/chats/md
(or a given dir) as info. WhatsApp sync stays out of v1.
--rebuild drops the database file and indexes from scratch. Without it a run
is idempotent (MERGE by (source,text) id).
@@ -22,7 +25,7 @@ ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT / "bin" / "tools"))
from kblib import ( # noqa: E402
connect, ensure_indexes, init_schema, upsert_leaf,
add_leafs, connect, ensure_indexes, init_schema, upsert_leaf, link_from_file,
open_readonly, stats,
)
from mdleaves import read_markdown, to_all, walk_markdown # noqa: E402
@@ -81,10 +84,11 @@ def index_leafs(conn, leafs: list[dict], embed_fn, limit: int) -> tuple[int, int
for lf in leafs[:limit] if limit else leafs:
query = f"{lf['heading']}\n\n{lf['text']}"
emb = embed_fn(lf["text"]) if lf["text"] else None
upsert_leaf(conn, text=query, root="info", confidence="confirmed",
lid = upsert_leaf(conn, text=query, root="info", confidence="confirmed",
source=lf["source"], source_rev="working-tree",
how="kb/index", loc=lf["source"], type_=lf.get("type", "reference"),
embedding=emb)
link_from_file(conn, lid, lf["source"], repo=str(lf.get("repo") or ""))
count += 1
return count, len(leafs)
@@ -95,12 +99,65 @@ def embedder():
return lambda text: model.encode([text])[0].astype(float).tolist()
def index_fact_dicts(conn, facts: list[dict], embed_fn) -> int:
"""Write extract-shaped dicts as root=facts leafs (2-source source field)."""
leafs = []
for f in facts:
text = str(f.get("text") or "")
source = str(f.get("source") or "")
if not text or not source:
continue
leafs.append({
"text": text,
"root": "facts",
"confidence": "confirmed",
"source": source,
"source_rev": f.get("source_rev") or "working-tree",
"how": f.get("how") or "facts/extract",
"loc": f.get("loc") or source,
"type": "fact",
"embedding": embed_fn(text) if text else None,
})
return len(add_leafs(conn, leafs))
def facts_from_extract() -> list[dict]:
import subprocess
proc = subprocess.run(
[sys.executable, str(ROOT / "bin" / "facts" / "extract"), "--json", "--dry-run"],
cwd=ROOT,
capture_output=True,
text=True,
check=False,
)
if proc.returncode != 0:
print(f"kb/index: facts/extract failed: {proc.stderr}", file=sys.stderr)
return []
try:
payload = json.loads(proc.stdout)
except json.JSONDecodeError:
print("kb/index: facts/extract produced non-JSON", file=sys.stderr)
return []
return list(payload.get("facts") or [])
def main(argv: list[str]) -> int:
import argparse
p = argparse.ArgumentParser(description="build the 2dph brain index")
p.add_argument("--corpus", action="append", help="extra markdown dir/file to index (may repeat)")
p.add_argument("--rebuild", action="store_true", help="fresh db + indexes")
p.add_argument("--db", default="", help="path to kb.lbug (default var/kb.lbug)")
p.add_argument("--no-defaults", action="store_true", help="do not index repo README/docs/skills")
p.add_argument("--with-mail", action="store_true", help="include var/mail message.md leafs")
p.add_argument("--with-facts", action="store_true", help="run facts/extract into root=facts")
p.add_argument("--facts-json", default="", help="JSON list (or {facts:[...]}) of fact dicts")
p.add_argument(
"--with-chats",
nargs="?",
const=str(ROOT / "var" / "chats" / "md"),
default="",
help="index chat markdown as info (default var/chats/md)",
)
p.add_argument("--since", default="", help="with --with-mail, only messages dated >= YYYY-MM-DD")
p.add_argument("--dry-run", action="store_true", help="count leafs, write nothing")
p.add_argument(
@@ -114,44 +171,71 @@ def main(argv: list[str]) -> int:
from kblib import DB_PATH, VAR
leafs = load_corpus(ROOT)
dbpath = Path(a.db) if a.db else DB_PATH
leafs: list[dict] = [] if a.no_defaults else load_corpus(ROOT)
if a.corpus:
for source in a.corpus:
leafs.extend(load_corpus_glob(source))
chat_n = 0
if a.with_chats:
chats = load_corpus_glob(a.with_chats)
chat_n = len(chats)
leafs.extend(chats)
mail_n = 0
if a.with_mail:
mail = from_mail_root(ROOT / "var" / "mail", since=a.since)
mail_n = len(mail)
leafs.extend(mail)
facts: list[dict] = []
if a.facts_json:
raw = Path(a.facts_json).read_text(encoding="utf-8")
payload = json.loads(raw)
facts = list(payload.get("facts") if isinstance(payload, dict) else payload)
if a.with_facts:
facts.extend(facts_from_extract())
if a.dry_run:
msg = {"indexed": 0, "corpus_total": len(leafs), "mail_leafs": mail_n, "dry_run": True}
msg = {
"indexed": 0,
"corpus_total": len(leafs),
"mail_leafs": mail_n,
"chat_leafs": chat_n,
"facts_leafs": len(facts),
"dry_run": True,
}
print(json.dumps(msg, indent=2) if a.json else
f"brain/index: {len(leafs)} leafs would be indexed (mail={mail_n})")
f"brain/index: {len(leafs)} info + {len(facts)} facts would be indexed")
return 0
VAR.mkdir(exist_ok=True)
if a.rebuild and DB_PATH.exists():
DB_PATH.unlink()
dbpath.parent.mkdir(parents=True, exist_ok=True)
if a.rebuild and dbpath.exists():
dbpath.unlink()
db, conn = connect(DB_PATH, read_only=False)
db, conn = connect(dbpath, read_only=False)
init_schema(conn)
# Never DROP FTS/VECTOR (ghost catalog). Write leafs, then ensure indexes
# unless --skip-indexes (seed facts first — MERGE under live FTS corrupts it).
# --rebuild already deleted kb.lbug above, so CREATE runs on a clean DB.
embed = embedder()
done, total = index_leafs(conn, leafs, embed, a.limit)
fact_n = index_fact_dicts(conn, facts, embed) if facts else 0
if not a.skip_indexes:
ensure_indexes(conn)
s = stats(conn)
conn.close()
db.close()
result = {"indexed": done, "corpus_total": total, **{k: v for k, v in s.items() if k in ("total", "by_root")}}
result = {
"indexed": done,
"corpus_total": total,
"facts_leafs": fact_n,
"chat_leafs": chat_n,
**{k: v for k, v in s.items() if k in ("total", "by_root")},
}
if a.skip_indexes:
result["indexes"] = "skipped"
print(json.dumps(result, indent=2) if a.json else f"indexed {done}/{total} leafs; db total {s['total']}")
print(json.dumps(result, indent=2) if a.json else
f"indexed {done}/{total} info + {fact_n} facts; db total {s['total']}")
return 0
+92 -1
View File
@@ -4,7 +4,7 @@ Single embedded graph `var/kb.lbug`. Two roots: facts (assertions backed by
>=2 independent sources) and info (narrative leafs). Hybrid retrieval: BM25
(FTS extension) + HNSW cosine (VECTOR extension) + Cypher graph hops.
All access is read-only unless `--rebuild` is passed to kb/index.
All access is read-only unless `--rebuild` (kb/index) or `kb/add`.
"""
from __future__ import annotations
@@ -118,6 +118,97 @@ def upsert_leaf(conn: ladybug.Connection, *, text: str, root: str, confidence: s
return lid
def add_leafs(conn: ladybug.Connection, leafs: list[dict]) -> list[str]:
"""Write facts+info leafs in one transaction. Safe while FTS/HNSW exist.
Each leaf dict: text, source, optional root/confidence/source_rev/how/loc/type/embedding.
Does not delete the database file. Measured on Ladybug 0.19: MERGE of new
ids (and updates) stays FTS+HNSW queryable; DROP INDEX is the fatal path.
"""
if not leafs:
return []
started = False
try:
conn.execute("BEGIN TRANSACTION")
started = True
except Exception:
started = False
ids: list[str] = []
try:
for lf in leafs:
ids.append(
upsert_leaf(
conn,
text=str(lf["text"]),
root=str(lf.get("root") or ROOT_INFO),
confidence=str(lf.get("confidence") or CONF_CONFIRMED),
source=str(lf["source"]),
source_rev=str(lf.get("source_rev") or "working-tree"),
how=str(lf.get("how") or "brain/add"),
loc=str(lf.get("loc") or lf.get("source") or ""),
type_=str(lf.get("type") or lf.get("type_") or "reference"),
embedding=lf.get("embedding"),
)
)
if started:
conn.execute("COMMIT")
except Exception:
if started:
try:
conn.execute("ROLLBACK")
except Exception:
pass
raise
return ids
def file_id(repo: str, path: str) -> str:
"""Stable File.id matching gitimport (`repo:path`)."""
return f"{repo}:{path}" if repo else path
def link_from_file(conn: ladybug.Connection, leaf_id: str, path: str,
repo: str = "", mtime: str = "") -> str:
"""MERGE File and Leaf-[:FROM_FILE]->File so --hop 1 can walk."""
fid = file_id(repo, path)
conn.execute(
"MERGE (f:File {id:$id}) SET f.path=$path, f.repo=$repo, f.mtime=$mtime",
parameters={"id": fid, "path": path, "repo": repo, "mtime": mtime},
)
conn.execute(
"MATCH (l:Leaf {id:$lid}), (f:File {id:$fid}) "
"MERGE (l)-[:FROM_FILE]->(f)",
parameters={"lid": leaf_id, "fid": fid},
)
return fid
HOP_STMTS = {
1: "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File) RETURN f.id, f.path, 1",
2: ("MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit) "
"RETURN c.id, c.subject, 2"),
3: ("MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit)"
"-[:AUTHORED]->(p:Person) RETURN p.id, p.name, 3"),
}
HOP_LABELS = {1: "File", 2: "Commit", 3: "Person"}
def hop_walk(conn: ladybug.Connection, leaf_id: str, n: int) -> list[dict]:
"""Walk Leaf → File → Commit → Person up to n hops (max 3)."""
depth = min(max(int(n), 0), 3)
out: list[dict] = []
for d in range(1, depth + 1):
rows = conn.execute(HOP_STMTS[d], parameters={"id": leaf_id}).get_all()
for row in rows:
out.append({
"id": row[0],
"label": HOP_LABELS[d],
"name": row[1],
"depth": int(row[2]),
})
return out
def leaf_index_names(conn: ladybug.Connection) -> set[str]:
"""Return index names on the Leaf table (e.g. {'id', 'Leaf_vec', '_PK'})."""
rows = conn.execute("CALL SHOW_INDEXES() RETURN *").get_all()
+34 -1
View File
@@ -75,9 +75,20 @@ class BinLayoutTest(unittest.TestCase):
)
def test_brain_methods_are_shebangs(self) -> None:
for method in ("index.go", "get.go", "stats.go", "eval.go", "watch.go"):
for method in ("index.go", "add.go", "get.go", "stats.go", "eval.go", "watch.go"):
self._assert_shebang(f"bin/brain/{method}")
def test_brain_add_is_python_write_not_rebuild(self) -> None:
self._assert_shebang("bin/brain/add.go")
text = (ROOT / "bin" / "brain" / "add.go").read_text()
self.assertIn("cmdbin.ExecFile", text)
self.assertIn("bin/kb/add", text)
self.assertNotIn("--rebuild", text)
py = (ROOT / "bin" / "kb" / "add").read_text()
self.assertIn("add_leafs", py)
self.assertIn("--json", py)
self.assertNotIn("unlink", py.lower())
def test_brain_get_stats_eval_are_not_python_exec(self) -> None:
for method in ("get.go", "stats.go", "eval.go"):
text = (ROOT / "bin" / "brain" / method).read_text()
@@ -196,3 +207,25 @@ class BinLayoutTest(unittest.TestCase):
search = (ROOT / "bin" / "kb" / "search").read_text()
self.assertIn("bin/cgo/zig", search)
self.assertNotIn("command -v gcc", search)
def test_ci_recall_sot_is_zig_brain_eval(self) -> None:
ci = (ROOT / ".github" / "workflows" / "ci.yml").read_text()
self.assertIn("bin/brain/eval.go", ci)
self.assertIn("system_ladybug,brain_eval", ci)
self.assertIn("/tmp/brain-eval", ci)
self.assertIn("KB_ROOT", ci)
self.assertNotIn("bin/kb/eval", ci)
self.assertNotIn("gate skipped", ci)
self.assertIn("./bin/facts/audit self", ci)
def test_eval_fragments_live_in_default_corpus(self) -> None:
"""CI --rebuild indexes README/PLAN/docs/skills; fragments must be there."""
corpus = []
for rel in ("README.md", "PLAN.md", "AGENTS.md"):
corpus.append((ROOT / rel).read_text())
for d in ("docs", "skills"):
for p in (ROOT / d).rglob("*.md"):
corpus.append(p.read_text())
blob = "\n".join(corpus)
for frag in ("BM25", "DevOps", "LadybugDB"):
self.assertIn(frag, blob, f"{frag} must appear in default index corpus")
+56
View File
@@ -39,3 +39,59 @@ class IndexAdapterTest(unittest.TestCase):
self.assertTrue(msg.get("dry_run"))
self.assertGreaterEqual(msg.get("corpus_total", 0), 1)
self.assertFalse(lbug.exists(), "dry-run must not create a Ladybug file")
def test_facts_json_and_chats_land_on_rebuild(self) -> None:
"""Gitea #18: facts (2-source) + chats markdown become leafs on rebuild."""
tmp = Path(tempfile.mkdtemp())
dbpath = tmp / "kb.lbug"
chats = tmp / "chats"
chats.mkdir()
(chats / "alice.md").write_text(
"# Chat\n\n## Alice and Bob\n\nhello from chats fixture unique-chat-token\n",
encoding="utf-8",
)
facts_path = tmp / "facts.json"
facts_path.write_text(json.dumps([{
"text": "container 'brain' unique-fact-token is running and declared in compose.yaml",
"source": "docker ps x compose.yaml",
"loc": "compose.yaml:brain",
"how": "facts/extract",
}]), encoding="utf-8")
venv_py = ROOT / ".venv" / "bin" / "python"
py = str(venv_py) if venv_py.is_file() else sys.executable
proc = subprocess.run(
[
py, str(ROOT / "bin" / "kb" / "index"),
"--rebuild", "--db", str(dbpath), "--no-defaults",
"--with-chats", str(chats),
"--facts-json", str(facts_path),
"--json",
],
cwd=ROOT,
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
msg = json.loads(proc.stdout)
self.assertGreaterEqual(msg.get("facts_leafs", 0), 1)
self.assertGreaterEqual(msg.get("chat_leafs", 0), 1)
self.assertTrue(dbpath.exists())
sys.path.insert(0, str(ROOT / "bin" / "tools"))
import kblib
db, conn = kblib.connect(dbpath, read_only=True)
try:
stats = kblib.stats(conn)
self.assertGreaterEqual(stats["by_root"].get("facts", 0), 1)
fts = kblib.query_fts(conn, "unique-chat-token", 5)
self.assertTrue(fts, "chats markdown must be FTS-searchable")
fact_hits = kblib.query_fts(conn, "unique-fact-token", 5)
self.assertTrue(any(h.get("root") == "facts" for h in fact_hits))
src = conn.execute(
"MATCH (l:Leaf {root:'facts'}) RETURN l.source"
).get_all()
self.assertTrue(any(" x " in str(r[0]) for r in src))
finally:
conn.close()
db.close()
+69
View File
@@ -0,0 +1,69 @@
"""Incremental add writes leafs without deleting kb.lbug."""
from __future__ import annotations
import json
import os
import subprocess
import sys
import tempfile
import unittest
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
class KbAddCLITest(unittest.TestCase):
def test_json_add_does_not_delete_db(self) -> None:
tmp = Path(tempfile.mkdtemp())
dbpath = tmp / "kb.lbug"
py = sys.executable
venv_py = ROOT / ".venv" / "bin" / "python"
if venv_py.is_file():
py = str(venv_py)
payload = {
"text": "cli zebra leaf",
"root": "info",
"source": "cli-test",
"confidence": "confirmed",
"how": "test",
"loc": str(tmp),
"type": "reference",
"embedding": [0.0] * 256,
}
payload["embedding"][0] = 0.3
proc = subprocess.run(
[py, str(ROOT / "bin" / "kb" / "add"), "--db", str(dbpath), "--json"],
cwd=ROOT,
input=json.dumps(payload),
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(proc.returncode, 0, proc.stderr)
self.assertTrue(dbpath.exists(), "add must create the db, not skip write")
out = json.loads(proc.stdout)
self.assertEqual(out.get("mode"), "add")
self.assertEqual(len(out.get("ids") or []), 1)
again = subprocess.run(
[py, str(ROOT / "bin" / "kb" / "add"), "--db", str(dbpath), "--json"],
cwd=ROOT,
input=json.dumps({
**payload,
"text": "second moose leaf",
"source": "cli-test-2",
}),
capture_output=True,
text=True,
env=os.environ.copy(),
check=False,
)
self.assertEqual(again.returncode, 0, again.stderr)
self.assertTrue(dbpath.exists())
second = json.loads(again.stdout)
self.assertEqual(len(second.get("ids") or []), 1)
self.assertNotEqual(out["ids"][0], second["ids"][0])
if __name__ == "__main__":
unittest.main()
+96
View File
@@ -71,6 +71,69 @@ class KblibTest(unittest.TestCase):
self.assertTrue(hits)
self.assertIn("Leaf_vec", kblib.leaf_index_names(self.conn))
def test_add_after_indexes_keeps_fts_queryable(self):
"""Incremental add after FTS+HNSW must find the new leaf on both indexes."""
kblib.upsert_leaf(self.conn, text="seed fox leaf", root="info",
confidence="confirmed", source="s", source_rev="r1",
how="test", loc="/tmp", type_="reference",
embedding=make_emb(0.1))
kblib.ensure_indexes(self.conn)
ids = kblib.add_leafs(self.conn, [{
"text": "added zebra after index",
"root": "facts",
"confidence": "confirmed",
"source": "a.md x b.md",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "fact",
"embedding": make_emb(0.9),
}])
self.assertEqual(len(ids), 1)
fts = kblib.query_fts(self.conn, "zebra", 5)
self.assertTrue(fts)
self.assertIn("zebra", fts[0]["text"])
self.assertEqual(fts[0]["root"], "facts")
vec = kblib.query_vector(self.conn, make_emb(0.9), 5)
self.assertTrue(any("zebra" in h["text"] for h in vec))
fox = kblib.query_fts(self.conn, "fox", 5)
self.assertTrue(fox)
self.assertIn("fox", fox[0]["text"])
def test_add_facts_and_info_one_transaction(self):
"""D12: facts and info land in the same transaction."""
kblib.ensure_indexes(self.conn)
ids = kblib.add_leafs(self.conn, [
{
"text": "tx fact leaf two-source",
"root": "facts",
"confidence": "confirmed",
"source": "compose.yml x docker ps",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "fact",
"embedding": make_emb(0.4),
},
{
"text": "tx info narrative",
"root": "info",
"confidence": "confirmed",
"source": "note.md",
"source_rev": "r1",
"how": "test",
"loc": "/tmp",
"type": "reference",
"embedding": make_emb(0.5),
},
])
self.assertEqual(len(ids), 2)
stats = kblib.stats(self.conn)
self.assertEqual(stats["by_root"].get("facts"), 1)
self.assertEqual(stats["by_root"].get("info"), 1)
self.assertTrue(kblib.query_fts(self.conn, "two-source", 5))
self.assertTrue(kblib.query_fts(self.conn, "narrative", 5))
def test_drop_vector_then_create_raises_clear_error(self):
"""DROP INDEX leaves ghost catalog; create_fts_and_vector must raise."""
kblib.upsert_leaf(self.conn, text="seed", root="info",
@@ -98,6 +161,39 @@ class KblibTest(unittest.TestCase):
self.assertEqual(stats["total"], 2)
self.assertEqual(stats["by_root"], {"facts": 1, "info": 1})
def test_hop_1_returns_file_hop_3_reaches_person(self):
"""--hop walks FROM_FILE / HAS_VERSION / AUTHORED (Gitea #17)."""
import gitimport
lid = kblib.upsert_leaf(
self.conn, text="readme hop fixture", root="info",
confidence="confirmed", source="README.md", source_rev="r1",
how="test", loc="README.md", type_="reference",
embedding=make_emb(0.3),
)
kblib.link_from_file(self.conn, lid, "README.md", repo="sample-repo")
gitimport.index_commits(self.conn, [gitimport.Commit(
sha="a1b2c3d",
author="Ada Lovelace",
email="ada@example.com",
date="2026-08-10T12:00:00Z",
subject="feat: first commit",
files=["README.md"],
)], "sample-repo")
hop1 = kblib.hop_walk(self.conn, lid, 1)
self.assertEqual(len(hop1), 1)
self.assertEqual(hop1[0]["label"], "File")
self.assertEqual(hop1[0]["name"], "README.md")
self.assertEqual(hop1[0]["depth"], 1)
hop3 = kblib.hop_walk(self.conn, lid, 3)
labels = {n["label"] for n in hop3}
self.assertIn("File", labels)
self.assertIn("Commit", labels)
self.assertIn("Person", labels)
person = [n for n in hop3 if n["label"] == "Person"][0]
self.assertEqual(person["name"], "Ada Lovelace")
self.assertEqual(person["depth"], 3)
if __name__ == "__main__":
unittest.main()
+23 -10
View File
@@ -1,7 +1,6 @@
"""Published docs must match live commands (Gitea SoT, brain/search, no fake --hop)."""
"""Published docs must match live commands (Gitea SoT, brain/search)."""
from __future__ import annotations
import re
import unittest
from pathlib import Path
@@ -139,23 +138,21 @@ class PublishedDocsTest(unittest.TestCase):
skill = (ROOT / "skills" / "brain" / "SKILL.md").read_text()
self.assertIn("`web` block", skill)
def test_docs_do_not_claim_hop_walks(self) -> None:
def test_docs_say_hop_walks_from_file(self) -> None:
paths = [
ROOT / "README.md",
ROOT / "docs" / "design.md",
ROOT / "skills" / "brain" / "SKILL.md",
ROOT / "skills" / "diataxis-docs" / "SKILL.md",
ROOT / "docs" / "runbook.md",
ROOT / "docs" / "README.md",
]
# Command-style `--hop 1` / `--hop N` plus follow/walk = the old lie.
# Honest "not implemented" notes must not match.
lie = re.compile(r"--hop (?:N|1).*(?:follow|walk)", re.I | re.S)
for path in paths:
text = path.read_text()
self.assertIsNone(
lie.search(text),
f"{path.relative_to(ROOT)} still claims --hop walks the graph",
self.assertIn("--hop", text, f"{path.relative_to(ROOT)} must document --hop")
self.assertNotIn(
"not implemented",
text.lower(),
f"{path.relative_to(ROOT)} still says hop is not implemented",
)
def test_docs_are_portable_diataxis(self) -> None:
@@ -173,3 +170,19 @@ class PublishedDocsTest(unittest.TestCase):
readme = (ROOT / "README.md").read_text()
self.assertIn("docs/runbook.md", readme)
self.assertNotIn("search.ops.io", readme)
def test_v1_epic_is_named_in_docs(self) -> None:
plan = (ROOT / "PLAN.md").read_text()
self.assertIn("Gap to v1", plan)
self.assertIn("eSlider/2dph/issues/16", plan)
self.assertIn("eSlider/2dph/issues/17", plan)
self.assertIn("eSlider/2dph/milestone/12", plan)
road = (ROOT / "docs" / "roadmap.md").read_text()
self.assertIn("type: explanation", road)
self.assertIn("issues/16", road)
self.assertIn("issues/14", road)
index = (ROOT / "docs" / "README.md").read_text()
self.assertIn("roadmap.md", index)
self.assertIn("epic #16", index)
agents = (ROOT / "AGENTS.md").read_text()
self.assertIn("roadmap.md", agents)
+7 -3
View File
@@ -5,6 +5,7 @@ related:
- docs/runbook.md
- docs/design.md
- PLAN.md
- docs/roadmap.md
---
# 2dph docs (Diataxis)
@@ -16,6 +17,7 @@ Evidence-first knowledge graph. Facts need proof or they are
|------|-----|
| tutorial / howto | [runbook](runbook.md) — run anywhere (uv, Go, Docker) |
| explanation | [design](design.md) — two roots, deduction, D17/D20/D18 |
| explanation | [roadmap](roadmap.md) — gap to v1 (epic #16) |
| howto | [picoclaw](picoclaw.md) — MCP agent profile |
| howto | [reasoner](reasoner.md) — CPU bake-off (D18) |
| reference | [PLAN.md](../PLAN.md) — decisions D1D21 |
@@ -25,10 +27,12 @@ Python write sidecar, **D14** `bin/{subject}/{method}.go`, **D15** Gitea origin,
**D17** assertion gate (facts → info → web), **D18** pluggable reasoner.
Search: `bin/brain/search.go "query"` (HTTP: `bin/brain/serve.go`
`/health` `/search` `/get` `/stats` `/audit` `/ingest`). `--hop` is
not a walk; the flag errors until File/FROM_FILE edges exist.
`/health` `/search` `/get` `/stats` `/audit` `/ingest`). `--hop N` walks
`FROM_FILE` → Commit → Person from each hit (max 3). Rebuild writes
File edges ([#17](https://git.produktor.io/eSlider/2dph/issues/17)).
Work board: [Gitea issues](https://git.produktor.io/eSlider/2dph/issues).
Work board: [Gitea issues](https://git.produktor.io/eSlider/2dph/issues)
([epic #16](https://git.produktor.io/eSlider/2dph/issues/16)).
PRs and CI: GitHub [`eSlider/2dph`](https://github.com/eSlider/2dph).
Published docs live here and match live commands.
+2 -1
View File
@@ -21,4 +21,5 @@ OO_CLI (default: $HOME/go/bin/oo)
./bin/chats/apply.go --dry-run
```
JSONL → markdown only. Brain ingest is `bin/brain/index.go` (not a `chats index`).
JSONL → markdown only. Brain ingest is `bin/brain/index.go --with-chats`
(default `var/chats/md`). WhatsApp sync is out of v1.
+11 -4
View File
@@ -4,6 +4,7 @@ status: current
related:
- docs/README.md
- docs/runbook.md
- docs/roadmap.md
---
# Design — facts, info, deduction
@@ -33,8 +34,9 @@ bin/brain/search.go "question"
is not evidence of absence; `--no-web` / `--root` skip it)
```
`--hop` is not implemented yet (needs File/FROM_FILE edges). The flag is an
error; it is not a graph walk.
`--hop N` walks `Leaf-[:FROM_FILE]->File-[:HAS_VERSION]->Commit-[:AUTHORED]->Person`
from each hit (1=File, 2=Commit, 3=Person). Rebuild writes FROM_FILE;
git import writes HAS_VERSION/AUTHORED ([#17](https://git.produktor.io/eSlider/2dph/issues/17)).
## Who / What / How / Where / When + evidence
@@ -78,14 +80,16 @@ Conflicting pairings (≥2 yes vs ≥2 no) = hypothesis (OQ1 → v2 resolution).
They do not exec Python. Control questions for recall@5 live in
`internal/brain/rank` so CI can test the table without libladybug.
Python `bin/kb/{get,stats,eval}` remain for GitHub Actions until the runner
fetches Zig + libs (`bin/cgo/zig`). Index/write is still `bin/kb/index`
fetches Zig + libs (`bin/cgo/zig`). Incremental write is `bin/kb/add`
(`bin/brain/add.go`). Bulk index/write is still `bin/kb/index`
(`docker compose --profile index`).
## Agent API (D20)
`bin/brain/serve.go` exposes the same `internal/httpapi.Ops` table as OpenAPI
(`GET /openapi.json`) and MCP (`POST /mcp` JSON-RPC `tools/list` +
`tools/call`). Tool names match paths: `search`, `get`, `stats`, `audit`.
`tools/call`). Tool names match paths: `search`, `get`, `stats`, `audit`,
`ingest` (add a leaf; omit body for the CLI hint).
Agents should use these endpoints instead of shebang CLIs.
## Reasoner (D18)
@@ -95,3 +99,6 @@ Pluggable OpenAI-compatible URL. RAM: `Qwen/Qwen3.5-9B`. Quality:
CPU sidecar: compose profile `reasoner` (`OLLAMA_NUM_GPU=0`,
`127.0.0.1:11435`). Bake-off: `bin/reasoner/bakeoff.go`. Weights stay out
of the 2dph image. See [docs/reasoner.md](reasoner.md).
Gap to v1 (hops, corpus, CI eval): [roadmap](roadmap.md),
[epic #16](https://git.produktor.io/eSlider/2dph/issues/16).
+59
View File
@@ -0,0 +1,59 @@
---
type: explanation
status: current
related:
- PLAN.md
- docs/design.md
- docs/runbook.md
---
# Gap to v1 — detective brain
Goal: a brain that does not assert without proof. Search is deduction
(`facts` ≥2 sources → `info``web`). `confirmed` only from the facts root.
**v1 is a living graph the agent can write and walk**, not “more RAG”.
Epic: [Gitea #16](https://git.produktor.io/eSlider/2dph/issues/16).
Milestone: [v1 detective brain](https://git.produktor.io/eSlider/2dph/milestone/12).
Decisions: [PLAN.md](../PLAN.md).
## In (do not reopen)
Read path Go + Zig CGO (D21). HTTP + OpenAPI + MCP (D20). PicoClaw compose
profile + CPU reasoner (D18). Mail sync → import → rebuild. D14 shebangs.
Compose `api` (no CPython) / `index` (Python write). Issues #1#5, #7#13.
[#15](https://git.produktor.io/eSlider/2dph/issues/15) lever/loop.
[#14](https://git.produktor.io/eSlider/2dph/issues/14) `bin/brain/add.go` /
`POST /ingest` (Python `kblib.add_leafs`; no Go upsert port).
[#17](https://git.produktor.io/eSlider/2dph/issues/17) `--hop N` walks
FROM_FILE / HAS_VERSION / AUTHORED.
[#18](https://git.produktor.io/eSlider/2dph/issues/18) `--with-facts` /
`--with-chats` on rebuild (WhatsApp out of v1).
[#19](https://git.produktor.io/eSlider/2dph/issues/19) CI recall SoT =
`bin/brain/eval.go` via Zig.
## Blockers
None for epic #16. v2: [#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR,
OQ1 contradiction resolution, OQ3 duckdb-md, OQ4 YAML-first leafs.
```
question
├─ FTS + HNSW ← in
├─ facts / info roots ← in
├─ web (D17) ← in
├─ brain/add ACID ← in
├─ Cypher hop ← in
└─ facts+chats corpus ← in
```
## Not v1
[#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR (OQ2), OQ1
contradiction resolution, OQ3 duckdb-md export, OQ4 YAML-first leafs.
## Close epic #16 when
Children #14, #15, #17, #18, #19 are closed. MCP tool order stays gated by tests.
+7 -4
View File
@@ -43,18 +43,21 @@ That binds `127.0.0.1:8888`. JSON format must stay enabled.
## Index then search
Write path is Compose profile `index` (Python Ladybug rebuild) until
`brain/add` is v2. The operator command is `bin/brain/index.go`.
Write path is `bin/brain/add.go` for a leaf (or `POST /ingest`). Bulk
corpus rebuild remains `bin/brain/index.go --rebuild` (Compose profile
`index`). Do not DROP INDEX on Ladybug 0.19.
```bash
bin/brain/index.go --rebuild
bin/brain/add.go --text "arc-1 runs Matrix" --root facts --source "compose.yml x docker ps"
bin/brain/index.go --rebuild --with-facts --with-chats
bin/brain/search.go "LadybugDB vector index" # facts → info → web (D17)
bin/brain/search.go "upstream flag" --no-web
bin/brain/get.go <id> --body
bin/brain/stats.go
```
`--hop` is not implemented. Empty web results are `throttled`, not absence.
`--hop N` walks File → Commit → Person from each hit. Empty web results are `throttled`, not absence.
Gap to v1: [roadmap](roadmap.md) / [epic #16](https://git.produktor.io/eSlider/2dph/issues/16).
Ladybug 0.19: never `DROP INDEX` FTS/VECTOR (ghost catalog). Fresh indexes =
delete `var/kb.lbug` then `--rebuild`.
+16 -4
View File
@@ -7,6 +7,8 @@ import (
"context"
"encoding/json"
"fmt"
"os/exec"
"path/filepath"
"github.com/eSlider/2dph/internal/brain/rank"
)
@@ -137,12 +139,22 @@ func (HTTP) Audit(context.Context) ([]byte, error) {
return json.Marshal(map[string]any{"status": "ok", "by_confidence": rows})
}
func (HTTP) Ingest(context.Context) ([]byte, error) {
func (HTTP) Ingest(ctx context.Context, body []byte) ([]byte, error) {
if len(bytes.TrimSpace(body)) == 0 {
return json.Marshal(map[string]any{
"mode": "rebuild",
"command": "bin/brain/index.go --rebuild",
"add": "v2",
"mode": "add",
"command": "bin/brain/add.go",
"rebuild": "bin/brain/index.go --rebuild",
})
}
cmd := exec.CommandContext(ctx, filepath.Join(repoRoot(), "bin", "kb", "add"), "--json")
cmd.Stdin = bytes.NewReader(body)
cmd.Dir = repoRoot()
out, err := cmd.Output()
if err != nil {
return nil, fmt.Errorf("add: %w", err)
}
return out, nil
}
func asInt(v any) int64 {
+11 -4
View File
@@ -6,7 +6,7 @@ import (
"strings"
)
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--json] [--no-web]
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--hop N] [--json] [--no-web]
bin/brain/search.go serve [port]
bin/brain/search.go --list-model`
@@ -15,6 +15,7 @@ type Options struct {
Root string
Repo string
Limit int
Hop int
JSONOut bool
ListModel bool
NoWeb bool
@@ -22,8 +23,6 @@ type Options struct {
// ParseArgs reads flags. Unknown flags are an error: silently dropping them
// meant `--hop 1` vanished and its argument `1` was appended to the query.
// --hop is recognised so it cannot be swallowed; it is not implemented until
// File/FROM_FILE edges exist.
func ParseArgs(args []string) (Options, error) {
opt := Options{Limit: 20}
var queryArgs []string
@@ -52,7 +51,15 @@ func ParseArgs(args []string) (Options, error) {
}
opt.Limit = n
case "--hop":
return opt, fmt.Errorf("--hop is not implemented yet (needs File/FROM_FILE edges)")
i++
n, err := strconv.Atoi(args[i])
if err != nil || n < 1 {
return opt, fmt.Errorf("--hop must be a positive integer, got %q", args[i])
}
if n > 3 {
return opt, fmt.Errorf("--hop max is 3 (File → Commit → Person)")
}
opt.Hop = n
case "--json":
opt.JSONOut = true
case "--no-web":
+27
View File
@@ -7,3 +7,30 @@ const FTSStmt = "CALL QUERY_FTS_INDEX('Leaf', 'id', $q) " +
const VecStmt = "CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) " +
"RETURN node.id, node.text, node.root, node.source, distance ORDER BY distance LIMIT $n"
// HopStmt is the Cypher walk from a search hit. Depth 1 = File, 2 = Commit, 3 = Person.
func HopStmt(depth int) string {
switch depth {
case 1:
return "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File) RETURN f.id, f.path, 1"
case 2:
return "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit) RETURN c.id, c.subject, 2"
case 3:
return "MATCH (l:Leaf {id:$id})-[:FROM_FILE]->(f:File)-[:HAS_VERSION]->(c:Commit)-[:AUTHORED]->(p:Person) RETURN p.id, p.name, 3"
default:
return ""
}
}
func HopLabel(depth int) string {
switch depth {
case 1:
return "File"
case 2:
return "Commit"
case 3:
return "Person"
default:
return ""
}
}
+8
View File
@@ -7,6 +7,13 @@ import (
"strings"
)
type HopNode struct {
ID string `json:"id"`
Label string `json:"label"`
Name string `json:"name"`
Depth int `json:"depth"`
}
// Hit is one search result, mirroring the python script's dict shape.
type Hit struct {
ID string `json:"id"`
@@ -15,6 +22,7 @@ type Hit struct {
Source string `json:"-"`
Score float64 `json:"score"`
Snippet string `json:"snippet,omitempty"`
Hops []HopNode `json:"hops,omitempty"`
}
// rrfK dampens the contribution of low ranks; same constant as kblib.py.
+33 -7
View File
@@ -91,15 +91,41 @@ func TestHybridKeepsVectorScoreForSharedHit(t *testing.T) {
}
// The old parser dropped unknown flags and appended their arguments to the
// query, so `search "q" --hop 1` searched for "q 1". --hop is not implemented
// here (needs File edges); it must still fail closed instead of changing q.
// query, so `search "q" --hop 1` searched for "q 1". --hop must stay a flag.
func TestParseHopIsNotSwallowedIntoTheQuery(t *testing.T) {
_, err := ParseArgs([]string{"what runs on arc-2", "--hop", "1"})
if err == nil {
t.Fatal("expected --hop to error (not implemented), not be swallowed")
opt, err := ParseArgs([]string{"what runs on arc-2", "--hop", "1"})
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
if !strings.Contains(err.Error(), "--hop") {
t.Fatalf("error should name --hop, got %v", err)
if opt.Query != "what runs on arc-2" {
t.Fatalf("query swallowed hop arg: %q", opt.Query)
}
if opt.Hop != 1 {
t.Fatalf("hop = %d, want 1", opt.Hop)
}
}
func TestParseHopMaxIsThree(t *testing.T) {
if _, err := ParseArgs([]string{"q", "--hop", "4"}); err == nil {
t.Fatal("expected --hop 4 to error")
}
opt, err := ParseArgs([]string{"q", "--hop", "3"})
if err != nil || opt.Hop != 3 {
t.Fatalf("hop 3: %+v err=%v", opt, err)
}
}
func TestHopStmtWalksFromFile(t *testing.T) {
s := HopStmt(1)
if !strings.Contains(s, "FROM_FILE") || !strings.Contains(s, "File") {
t.Fatalf("hop 1 must walk FROM_FILE, got %q", s)
}
s3 := HopStmt(3)
if !strings.Contains(s3, "HAS_VERSION") || !strings.Contains(s3, "AUTHORED") || !strings.Contains(s3, "Person") {
t.Fatalf("hop 3 must reach Person, got %q", s3)
}
if HopLabel(1) != "File" || HopLabel(3) != "Person" {
t.Fatal("hop labels")
}
}
+58
View File
@@ -56,6 +56,12 @@ func runSearch(args []string) int {
fmt.Fprintf(os.Stderr, "search: %v\n", err)
return 1
}
if opt.Hop > 0 {
if err := attachHops(hits, opt.Hop); err != nil {
fmt.Fprintf(os.Stderr, "hop: %v\n", err)
return 1
}
}
results := hits
for i := range results {
@@ -108,6 +114,44 @@ func searchHits(query, root, repo string, limit int) ([]Hit, error) {
return rank.RankAndFilter(fts, vec, root, repo, limit), nil
}
func attachHops(hits []Hit, n int) error {
if conn == nil {
return fmt.Errorf("brain not open")
}
for i := range hits {
var hops []rank.HopNode
for d := 1; d <= n; d++ {
stmt, err := conn.Prepare(rank.HopStmt(d))
if err != nil {
return err
}
res, err := conn.Execute(stmt, map[string]any{"id": hits[i].ID})
stmt.Close()
if err != nil {
return err
}
for res.HasNext() {
row, err := res.Next()
if err != nil {
return err
}
vals, err := row.GetAsSlice()
if err != nil || len(vals) < 3 {
continue
}
hops = append(hops, rank.HopNode{
ID: fmt.Sprint(vals[0]),
Label: rank.HopLabel(d),
Name: fmt.Sprint(vals[1]),
Depth: int(asInt(vals[2])),
})
}
}
hits[i].Hops = hops
}
return nil
}
func b2i(err error) int {
if err != nil {
return 1
@@ -188,6 +232,7 @@ type jsonHit struct {
Root string `json:"root"`
Score float64 `json:"score"`
Snippet string `json:"snippet,omitempty"`
Hops []rank.HopNode `json:"hops,omitempty"`
}
func toJSONOut(hits []Hit, query, rootFilter string, web *rank.SecondSource) *jsonOut {
@@ -199,6 +244,7 @@ func toJSONOut(hits []Hit, query, rootFilter string, web *rank.SecondSource) *js
Root: h.Root,
Score: h.Score,
Snippet: h.Snippet,
Hops: h.Hops,
}
}
return &jsonOut{
@@ -222,6 +268,18 @@ func resultsToDicts(hits []Hit) []any {
if h.Snippet != "" {
d = append(d, KV{"snippet", h.Snippet})
}
if len(h.Hops) > 0 {
nodes := make([]any, len(h.Hops))
for j, n := range h.Hops {
nodes[j] = Dict{
{"id", n.ID},
{"label", n.Label},
{"name", n.Name},
{"depth", n.Depth},
}
}
d = append(d, KV{"hops", nodes})
}
out[i] = d
}
return out
+9 -1
View File
@@ -144,7 +144,15 @@ func (s *Server) mcpCall(r *http.Request, params json.RawMessage) (any, error) {
return nil, fmt.Errorf("cancelled")
}
defer s.release()
body, err = s.api.Ingest(r.Context())
var payload []byte
text := strings.TrimSpace(fmt.Sprint(p.Arguments["text"]))
if text != "" && text != "<nil>" {
payload, err = json.Marshal(p.Arguments)
if err != nil {
return nil, err
}
}
body, err = s.api.Ingest(r.Context(), payload)
default:
return nil, fmt.Errorf("unknown tool %s", p.Name)
}
+43 -5
View File
@@ -11,6 +11,7 @@ import (
"context"
"encoding/json"
"errors"
"io"
"log"
"net/http"
"os"
@@ -27,7 +28,7 @@ type API interface {
Get(ctx context.Context, id string, body bool) ([]byte, error)
Stats(ctx context.Context) ([]byte, error)
Audit(ctx context.Context) ([]byte, error)
Ingest(ctx context.Context) ([]byte, error)
Ingest(ctx context.Context, body []byte) ([]byte, error)
}
type Server struct {
@@ -59,7 +60,7 @@ func (s *Server) ServeHTTP(w http.ResponseWriter, r *http.Request) {
case PathAudit:
s.handleJSON(w, r, s.api.Audit)
case PathIngest:
s.handleJSON(w, r, s.api.Ingest)
s.handleIngest(w, r)
case PathOpenAPI:
s.handleOpenAPI(w, r)
case PathMCP:
@@ -116,6 +117,24 @@ func (s *Server) handleJSON(w http.ResponseWriter, r *http.Request, fn func(cont
writeAPI(w, body, err)
}
func (s *Server) handleIngest(w http.ResponseWriter, r *http.Request) {
var raw []byte
if r.Method == http.MethodPost {
b, err := io.ReadAll(io.LimitReader(r.Body, 1<<20))
if err != nil {
writeJSON(w, http.StatusBadRequest, map[string]any{"error": "read body"})
return
}
raw = b
}
if !s.acquire(w, r) {
return
}
defer s.release()
body, err := s.api.Ingest(r.Context(), raw)
writeAPI(w, body, err)
}
func (s *Server) tryAcquire(r *http.Request) bool {
return s.acquire(nopWriter{}, r)
}
@@ -191,11 +210,30 @@ func (ExecSearcher) Get(context.Context, string, bool) ([]byte, error) {
}
func (ExecSearcher) Stats(context.Context) ([]byte, error) { return nil, errUnimplemented }
func (ExecSearcher) Audit(context.Context) ([]byte, error) { return nil, errUnimplemented }
func (ExecSearcher) Ingest(context.Context) ([]byte, error) {
func (b ExecSearcher) Ingest(ctx context.Context, body []byte) ([]byte, error) {
if len(strings.TrimSpace(string(body))) == 0 {
return json.Marshal(map[string]any{
"mode": "rebuild",
"command": "bin/brain/index.go --rebuild",
"mode": "add",
"command": "bin/brain/add.go",
"rebuild": "bin/brain/index.go --rebuild",
})
}
root := os.Getenv("KB_ROOT")
if root == "" {
root = "."
}
cmd := exec.CommandContext(ctx, filepath.Join(root, "bin/kb/add"), "--json")
cmd.Stdin = strings.NewReader(string(body))
cmd.Dir = root
out, err := cmd.Output()
if err != nil {
var exitErr *exec.ExitError
if errors.As(err, &exitErr) {
return nil, errors.New("add failed: " + strings.TrimSpace(string(exitErr.Stderr)))
}
return nil, err
}
return out, nil
}
func defaultSearchCmd(root string) string {
+27 -2
View File
@@ -1,6 +1,7 @@
package httpapi
import (
"bytes"
"context"
"encoding/json"
"net/http"
@@ -64,8 +65,11 @@ func (f *fakeSearcher) Audit(context.Context) ([]byte, error) {
return []byte(`{"status":"ok"}`), nil
}
func (f *fakeSearcher) Ingest(context.Context) ([]byte, error) {
return []byte(`{"mode":"rebuild","command":"bin/brain/index.go --rebuild"}`), nil
func (f *fakeSearcher) Ingest(_ context.Context, body []byte) ([]byte, error) {
if len(bytes.TrimSpace(body)) == 0 {
return []byte(`{"mode":"add","command":"bin/brain/add.go"}`), nil
}
return []byte(`{"mode":"add","ids":["fake-leaf"]}`), nil
}
func (f *fakeSearcher) count() int {
@@ -193,6 +197,27 @@ func TestStatsAuditIngest(t *testing.T) {
}
}
func TestIngestIsAddNotRebuildHint(t *testing.T) {
h := NewServer(&fakeSearcher{}, 1)
code, body := get(t, h, "/ingest")
if code != http.StatusOK {
t.Fatalf("GET /ingest code = %d body=%s", code, body)
}
if strings.Contains(string(body), `"add":"v2"`) || strings.Contains(string(body), "write is v2") {
t.Fatalf("GET /ingest still a v2 hint: %s", body)
}
if !strings.Contains(string(body), "bin/brain/add.go") {
t.Fatalf("GET /ingest should name add.go: %s", body)
}
code, body = postJSON(t, h, "/ingest", `{"text":"hello","root":"info","source":"t"}`)
if code != http.StatusOK {
t.Fatalf("POST /ingest code = %d body=%s", code, body)
}
if !strings.Contains(string(body), "fake-leaf") {
t.Fatalf("POST /ingest should add: %s", body)
}
}
func TestHTTPPackageDoesNotExecPython(t *testing.T) {
raw, err := os.ReadFile("server.go")
if err != nil {
+9 -1
View File
@@ -47,7 +47,15 @@ var Ops = []Op{
},
{Path: PathStats, Method: "get", ID: "stats", Summary: "index health", MCP: true},
{Path: PathAudit, Method: "get", ID: "audit", Summary: "facts confidence histogram", MCP: true},
{Path: PathIngest, Method: "get", ID: "ingest", Summary: "rebuild hint (write is v2)", MCP: true},
{
Path: PathIngest, Method: "post", ID: "ingest", Summary: "add a leaf without rebuild",
MCP: true,
Params: []Param{
{Name: "text", In: "query", Type: "string", Description: "leaf text (omit for CLI hint)"},
{Name: "root", In: "query", Type: "string", Description: "facts or info (default info)"},
{Name: "source", In: "query", Type: "string", Description: "evidence pointer; facts need two sources"},
},
},
{Path: PathOpenAPI, Method: "get", ID: "openapi", Summary: "OpenAPI 3 document for this server"},
}
+14 -1
View File
@@ -38,11 +38,24 @@ func TestMCPToolsMatchOpenAPIPaths(t *testing.T) {
t.Fatalf("MCP tool %s has no OpenAPI path %s", tool.Name, path)
}
}
for _, need := range []string{"search", "get", "stats", "audit"} {
for _, need := range []string{"search", "get", "stats", "audit", "ingest"} {
if !names[need] {
t.Fatalf("MCP tools missing %s: %v", need, names)
}
}
var ingest MCPTool
for _, tool := range tools {
if tool.Name == "ingest" {
ingest = tool
break
}
}
if strings.Contains(ingest.Description, "v2") {
t.Fatalf("ingest still a v2 hint: %s", ingest.Description)
}
if !strings.Contains(ingest.Description, "add") {
t.Fatalf("ingest should describe add: %s", ingest.Description)
}
}
func TestOpenAPIHTTP(t *testing.T) {
+3 -2
View File
@@ -26,13 +26,14 @@ second independent source when local roots cannot confirm. An answer is
bin/brain/search.go "Matrix federation" # pointers + snippets, YAML
bin/brain/search.go "onlyoffice postgres" --root facts # restrict to confirmed
bin/brain/search.go "where is cs-lexicon" --json | yq '.[].ref'
bin/brain/add.go --text T --root facts --source "a.md x b.md"
bin/brain/get.go <id> --body # full chunk only when needed
bin/brain/stats.go # index health
bin/brain/eval.go # recall@5 >= 0.95 gate (Go; Python bin/kb/eval is CI fallback)
```
`bin/kb/search` is a deprecated wrapper. `--hop` errors (File/FROM_FILE edges
are not wired yet); do not treat it as a graph walk.
`bin/kb/search` is a deprecated wrapper. `--hop N` walks
`FROM_FILE` / `HAS_VERSION` / `AUTHORED` from each hit (1=File, 3=Person).
## Rules
+1 -1
View File
@@ -8,4 +8,4 @@ Serve: `bin/brain/serve.go` (`GET /openapi.json`, `POST /mcp`).
- `get` — read one leaf by id
- `stats` — index health
- `audit` — facts confidence histogram
- `ingest`rebuild hint (write is v2)
- `ingest`add a leaf without rebuild