Compare commits
6
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b612d3cc97 | ||
|
|
5ccca2fac9 | ||
|
|
fc2723c39f | ||
|
|
0c4bb87001 | ||
|
|
0065655e03 | ||
|
|
a331042488 |
@@ -14,3 +14,9 @@ go.work.local
|
|||||||
models/
|
models/
|
||||||
# Purged from git history. Do not re-add.
|
# Purged from git history. Do not re-add.
|
||||||
docs/crm-associations-proof.md
|
docs/crm-associations-proof.md
|
||||||
|
|
||||||
|
# mount scaffold for the 8TB volume, never part of the repo
|
||||||
|
mnt/
|
||||||
|
|
||||||
|
# go build ./bin/mail/sync.go drops a binary named `sync` in cwd
|
||||||
|
/sync
|
||||||
|
|||||||
@@ -48,11 +48,12 @@ bin/postgres/ query.go (read-only YAML)
|
|||||||
bin/git/ import.go (go-git history; Python shim execs it)
|
bin/git/ import.go (go-git history; Python shim execs it)
|
||||||
bin/web/ search.go (SearXNG; Python shim execs it)
|
bin/web/ search.go (SearXNG; Python shim execs it)
|
||||||
bin/reasoner/ bakeoff.go (D18 CPU OpenAI tool-call bake-off)
|
bin/reasoner/ bakeoff.go (D18 CPU OpenAI tool-call bake-off)
|
||||||
internal/ shared Go (brain/rank is cgo-free; chats parsers; gitlog; websearch; reasoner; duckstats)
|
internal/ shared Go (brain/rank is cgo-free; facts D16; cli flaggy D23; chats; gitlog; websearch; reasoner; duckstats)
|
||||||
bin/qa/ stats.go (DuckDB quantiles / JSONL count; gcc CGO, not Zig)
|
bin/qa/ stats.go (DuckDB quantiles / JSONL count; gcc CGO, not Zig)
|
||||||
bin/watch/ corpus watcher (used by bin/brain/watch.go)
|
bin/watch/ corpus watcher (used by bin/brain/watch.go)
|
||||||
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
|
bin/tools/ vendored python libs behind bin/* (kblib, yamlout, websearch)
|
||||||
bin/cgo/ zig zcc zc++ (CGO via zig cc, not gcc)
|
bin/cgo/ zig zcc zc++ (CGO via zig cc, not gcc)
|
||||||
|
bin/stack/ start start-assistant stop status (compose + PicoClaw agent)
|
||||||
bin/docker-entrypoint container entrypoint (api: serve|search|watch; index: python)
|
bin/docker-entrypoint container entrypoint (api: serve|search|watch; index: python)
|
||||||
compose.yaml docker composition (root level, not docker/)
|
compose.yaml docker composition (root level, not docker/)
|
||||||
Dockerfile api (Zig CGO, no Python) + index (Python write)
|
Dockerfile api (Zig CGO, no Python) + index (Python write)
|
||||||
@@ -65,12 +66,20 @@ var/ kb.lbug, var/mail/*, caches (gitignored)
|
|||||||
```bash
|
```bash
|
||||||
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
|
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw message.json + attachments
|
||||||
bin/mail/sync.go --source gmail --query 'from:example.com' --out var/mail # Gmail search (default in:inbox)
|
bin/mail/sync.go --source gmail --query 'from:example.com' --out var/mail # Gmail search (default in:inbox)
|
||||||
|
bin/mail/sync.go --source m365 --env ~/.config/brain/mail.env --out var/mail # Microsoft Graph (delta)
|
||||||
bin/mail/import.go --from-raw var/mail # message.json → message.md (convert only)
|
bin/mail/import.go --from-raw var/mail # message.json → message.md (convert only)
|
||||||
bin/brain/index.go --rebuild --with-facts --with-chats
|
bin/brain/index.go --rebuild --with-facts --with-chats
|
||||||
|
bin/stack/start-mail-sync # compose ETL: sync→import every 300s
|
||||||
```
|
```
|
||||||
|
|
||||||
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
|
- `sync` (Go) downloads messages + attachments; Gmail uses paginated list +
|
||||||
`body.attachmentId` (not partId) for attachments.
|
`body.attachmentId` (not partId) for attachments. Sources: `onlyoffice`,
|
||||||
|
`gmail`, `m365` (client-credentials + delta link; commit after success).
|
||||||
|
- Compose `mail-sync` / `bin/stack/start-mail-sync`: ETL loop (default
|
||||||
|
`onlyoffice,gmail`, 300s). On `new>0` runs import; full `--rebuild` only if
|
||||||
|
`MAIL_SYNC_INDEX=1`. Secrets: `~/.config/brain/mail.env` + `~/.gmail-mcp`.
|
||||||
|
Case wrappers (e.g. family `gmail-sync-la-quinta.sh`) and ai-bot
|
||||||
|
`gmail-reauth.sh` reuse this sync/OAuth — do not fork corpus download.
|
||||||
- `import` converts body + attachments to markdown. PDFs use poppler
|
- `import` converts body + attachments to markdown. PDFs use poppler
|
||||||
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs use
|
`pdftotext -layout` fast path (~15ms); textless/scanned PDFs use
|
||||||
`pdftoppm` + tesseract `eng+deu` (`bin/mail/ocr.go`). Optional
|
`pdftoppm` + tesseract `eng+deu` (`bin/mail/ocr.go`). Optional
|
||||||
@@ -84,11 +93,13 @@ bin/brain/index.go --rebuild --with-facts --with-chats
|
|||||||
## Tools
|
## Tools
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
bin/facts/audit.go ["self"|"facts"|"info"|"stale"] # 2-source + staleness gate
|
bin/facts/audit.go ["self"|"db"|"contradict"] # 2-source + D16 adjudication
|
||||||
bin/facts/crm.go [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
|
bin/facts/crm.go [--dry-run] # proof person↔company/company↔project (ooCRM × corpus SoT)
|
||||||
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
|
bin/kb/search "query" [--repo X] # deprecated wrapper → bin/brain/search.go
|
||||||
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
|
bin/brain/search.go "query" [--root facts|info] # deduction search → YAML
|
||||||
|
bin/brain/search.go "query" --as-of 2025-01-01 # D24 fact intervals
|
||||||
bin/brain/search.go "query" --no-web # local graph only
|
bin/brain/search.go "query" --no-web # local graph only
|
||||||
|
source <(./bin/cli/complete.go bash) # flaggy completions (D23)
|
||||||
eval "$(bin/cgo/zig env)" # Zig cc + liblbug (not gcc)
|
eval "$(bin/cgo/zig env)" # Zig cc + liblbug (not gcc)
|
||||||
bin/brain/index.go --rebuild [--with-mail] [--with-facts] [--with-chats]
|
bin/brain/index.go --rebuild [--with-mail] [--with-facts] [--with-chats]
|
||||||
bin/brain/add.go --text T --root facts --source "a.md x b.md" # incremental write
|
bin/brain/add.go --text T --root facts --source "a.md x b.md" # incremental write
|
||||||
@@ -97,6 +108,10 @@ bin/brain/get.go <id> [--body] [--json] # Go read; Python bin/kb/get CI
|
|||||||
bin/brain/stats.go [--json]
|
bin/brain/stats.go [--json]
|
||||||
bin/brain/eval.go [--json] # recall@5; questions in internal/brain/rank
|
bin/brain/eval.go [--json] # recall@5; questions in internal/brain/rank
|
||||||
bin/brain/serve.go # HTTP :8630; GET /openapi.json POST /mcp
|
bin/brain/serve.go # HTTP :8630; GET /openapi.json POST /mcp
|
||||||
|
bin/stack/start # brain HTTP/MCP (reuse healthy :8630)
|
||||||
|
bin/stack/start-assistant # + reasoner + PicoClaw agent
|
||||||
|
bin/stack/status # YAML health
|
||||||
|
bin/stack/stop # compose stop; volumes kept
|
||||||
bin/markdown/import.go [dir] # H2 leafs → YAML; Python bin/md/import fallback
|
bin/markdown/import.go [dir] # H2 leafs → YAML; Python bin/md/import fallback
|
||||||
bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs
|
bin/git/import.go [REPO] [--json] [--limit N] # go-git history → commit leafs
|
||||||
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
|
bin/web/search.go "query" [--json] # SearXNG; throttled ≠ absence
|
||||||
|
|||||||
+10
@@ -6,6 +6,15 @@
|
|||||||
# API: Go + ladybug via Zig CGO (no CPython).
|
# API: Go + ladybug via Zig CGO (no CPython).
|
||||||
# Index: Python write path (profile `index` until brain/add is v2).
|
# Index: Python write path (profile `index` until brain/add is v2).
|
||||||
|
|
||||||
|
# --- mail-sync: standalone M365/OnlyOffice/Gmail puller (pure Go, no CGO) ---
|
||||||
|
FROM golang:1.26-bookworm AS mail-build
|
||||||
|
WORKDIR /src
|
||||||
|
COPY go.mod go.sum ./
|
||||||
|
RUN go mod download
|
||||||
|
COPY bin/mail ./bin/mail
|
||||||
|
COPY internal ./internal
|
||||||
|
RUN CGO_ENABLED=0 go build -o /mail-sync ./bin/mail/sync.go
|
||||||
|
|
||||||
# --- Python sidecar (Ladybug write / rebuild) ---
|
# --- Python sidecar (Ladybug write / rebuild) ---
|
||||||
FROM python:3.12-slim AS index
|
FROM python:3.12-slim AS index
|
||||||
|
|
||||||
@@ -28,6 +37,7 @@ RUN python -m pip install --no-cache-dir -r /tmp/requirements.lock.txt \
|
|||||||
COPY . .
|
COPY . .
|
||||||
RUN chmod +x /app/bin/docker-entrypoint \
|
RUN chmod +x /app/bin/docker-entrypoint \
|
||||||
&& chown -R 2dph:2dph /app
|
&& chown -R 2dph:2dph /app
|
||||||
|
COPY --from=mail-build /mail-sync /app/bin/mail-sync
|
||||||
USER 2dph
|
USER 2dph
|
||||||
|
|
||||||
ENV PATH="/app/bin:${PATH}" \
|
ENV PATH="/app/bin:${PATH}" \
|
||||||
|
|||||||
@@ -5,8 +5,12 @@ RAG over the operational Brain/ops/eSlider stack. Built like Sherlock
|
|||||||
Holmes: nothing is asserted unless it has proof.
|
Holmes: nothing is asserted unless it has proof.
|
||||||
|
|
||||||
Status: **v1 in** (epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed).
|
Status: **v1 in** (epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed).
|
||||||
v2 board: milestone [v2](https://git.produktor.io/eSlider/2dph/milestone/13) — OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6),
|
v2 board: milestone [v2](https://git.produktor.io/eSlider/2dph/milestone/13) —
|
||||||
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1, [#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3.
|
OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6) in,
|
||||||
|
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1 in,
|
||||||
|
[#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3 in,
|
||||||
|
[#34](https://git.produktor.io/eSlider/2dph/issues/34) D23 in,
|
||||||
|
[#36](https://git.produktor.io/eSlider/2dph/issues/36) OQ5/D24 in.
|
||||||
Gap: [docs/roadmap.md](docs/roadmap.md).
|
Gap: [docs/roadmap.md](docs/roadmap.md).
|
||||||
|
|
||||||
## What
|
## What
|
||||||
@@ -42,13 +46,15 @@ detective method: **a fact needs ≥2 independent sources or it is
|
|||||||
| D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. |
|
| D13 | portfolio | start graph `(Person:eslider)-[:HAS]->(Portfolio)`, associate other natural/juristic persons later. |
|
||||||
| D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. |
|
| D14 | tooling style | `bin/{subject}/{method}.go` shebang (e.g. `bin/brain/search.go`). Shared code in `internal/`. One root `go.mod` + `go.work`. No `bin/*/main.go`, no nested modules. |
|
||||||
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
|
| D15 | repo | Gitea [`eSlider/2dph`](https://git.produktor.io/eSlider/2dph) is origin + [issues](https://git.produktor.io/eSlider/2dph/issues). GitHub `eSlider/2dph` is the public clone (PRs + Actions CI). No direct `main` pushes. TDD → PR → CI green → merge. |
|
||||||
| D16 | contradictions | ≥2 yes vs ≥2 no → unrelated sources conflict → hypothesis → `(not confirmed)`. Resolution (authority, staleness adjudication) = **v2**, tracked as open question. |
|
| D16 | contradictions | ≥2 yes vs ≥2 no → hypothesis → `(not confirmed)` until a rule fires. Order: **temporal_freshness** (fresh ≥2 vs stale minority), then **authority_pairing** (runtime/config A×B beats narrative C). Store as `a x b vs c x d` on hypothesis leafs. `bin/facts/audit contradict`. [#29](https://git.produktor.io/eSlider/2dph/issues/29). |
|
||||||
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
|
| D17 | assertion gate | Fact-check every *claim* (facts → info → live → web), not every edit. `bin/brain/search.go` adds a `web` block when there is no facts hit (`throttled`/`skipped`/`refused` ≠ absence). `--root` and `--no-web` stay local. Missing graph ≠ “does not exist”. |
|
||||||
| D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is compose profile `picoclaw`; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. Agent lever/loop: [#15](https://git.produktor.io/eSlider/2dph/issues/15). |
|
| D18 | reasoner | Pluggable OpenAI-compatible URL (`REASONER_BASE_URL`). RAM: `Qwen/Qwen3.5-9B`. Quality: `prism-ml/Bonsai-27B-gguf` or `Qwen/Qwen3.6-27B`. No official Qwen3.6-9B. CPU bake-off: `bin/reasoner/bakeoff.go` + compose profile `reasoner` (`OLLAMA_NUM_GPU=0`, `:11435`). PicoClaw is compose profile `picoclaw`; tools are `search`/`get`/`audit`. Weights are not copied into the 2dph image. Agent lever/loop: [#15](https://git.produktor.io/eSlider/2dph/issues/15). |
|
||||||
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
|
| D19 | git history | [go-git](https://github.com/go-git/go-git) via `bin/git/import.go`. No subprocess of the git binary. Conversion prints commit leafs; brain write is `bin/brain/index.go`. |
|
||||||
| D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`/`ingest`). |
|
| D20 | agent API | OpenAPI + MCP are generated from the same `internal/httpapi.Ops` table as `bin/brain/serve.go` handlers. `GET /openapi.json`, `POST /mcp` (JSON-RPC tools/list + tools/call). Tool names match OpenAPI paths (`search`/`get`/`stats`/`audit`/`ingest`). |
|
||||||
| D21 | CGO | Ladybug/tokenizers CGO is compiled with **Zig** (`bin/cgo/zcc` → `zig cc -target …-linux-gnu`), not gcc. `bin/cgo/zig` pins Zig 0.14.1 + liblbug 0.19.1 + libtokenizers 1.27.0. Compose `target: api` has no CPython; write/rebuild is profile `index`. |
|
| D21 | CGO | Ladybug/tokenizers CGO is compiled with **Zig** (`bin/cgo/zcc` → `zig cc -target …-linux-gnu`), not gcc. `bin/cgo/zig` pins Zig 0.14.1 + liblbug 0.19.1 + libtokenizers 1.27.0. Compose `target: api` has no CPython; write/rebuild is profile `index`. |
|
||||||
| D22 | analytics | **duckdb-go** in-process (`internal/duckstats`, `bin/qa/stats.go`) for quantiles/JSONL. Links with **gcc/g++**, not Zig. Ladybug stays the graph; web-search cache stays modernc sqlite. Slice small structured docs with **mikefarah/yq**, not kislyuk/jq. [#30](https://git.produktor.io/eSlider/2dph/issues/30). |
|
| D22 | analytics | **duckdb-go** in-process (`internal/duckstats`, `bin/qa/stats.go`) for quantiles/JSONL. Links with **gcc/g++**, not Zig. Ladybug stays the graph; web-search cache stays modernc sqlite. Slice small structured docs with **mikefarah/yq**, not kislyuk/jq. [#30](https://git.produktor.io/eSlider/2dph/issues/30). |
|
||||||
|
| D23 | CLI | **flaggy** (`github.com/integrii/flaggy`, 0 deps). Flags at any position. Wrapper `internal/cli`. Bash complete: `source <(./bin/cli/complete.go bash)`. No cobra, no stdlib `flag` in Go tools. Search does not intercept the word `completion`. [#34](https://git.produktor.io/eSlider/2dph/issues/34). |
|
||||||
|
| D24 | fact intervals | Leaf `valid_from` / `valid_to` (YYYY-MM-DD, inclusive; empty = open/legacy). Search `--as-of` / MCP `as_of` keeps facts active that day. Not D16 `temporal_freshness` (source stale vs HEAD). Empty interval = always visible. [#36](https://git.produktor.io/eSlider/2dph/issues/36). |
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
|
|
||||||
@@ -65,7 +71,8 @@ detective method: **a fact needs ≥2 independent sources or it is
|
|||||||
brain/add.go incremental leaf write (no rebuild)
|
brain/add.go incremental leaf write (no rebuild)
|
||||||
brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback
|
brain/get.go stats.go eval.go # Go read (cgo); Python bin/kb/* CI fallback
|
||||||
brain/watch.go
|
brain/watch.go
|
||||||
brain/search.go deduction: facts → info → web-search
|
brain/search.go deduction: facts → info → web
|
||||||
|
cli/complete.go flaggy bash/zsh/fish complete (D23)
|
||||||
brain/serve.go HTTP API in-process + OpenAPI/MCP (D20); Zig CGO (D21)
|
brain/serve.go HTTP API in-process + OpenAPI/MCP (D20); Zig CGO (D21)
|
||||||
cgo/zig zcc zc++ CGO toolchain (zig cc, not gcc)
|
cgo/zig zcc zc++ CGO toolchain (zig cc, not gcc)
|
||||||
mail/import.go JSON → markdown (no brain write)
|
mail/import.go JSON → markdown (no brain write)
|
||||||
@@ -78,7 +85,8 @@ detective method: **a fact needs ≥2 independent sources or it is
|
|||||||
(libs in internal/chats; no chats index)
|
(libs in internal/chats; no chats index)
|
||||||
mail/ocr.go tesseract eng+deu (pdftoppm scans)
|
mail/ocr.go tesseract eng+deu (pdftoppm scans)
|
||||||
md/import (deprecated; bin/markdown/import.go)
|
md/import (deprecated; bin/markdown/import.go)
|
||||||
brain/extract brain/audit brain/deduce (thinking wrapper)
|
brain/extract brain/audit brain/deduce (thinking wrapper)
|
||||||
|
stack/start start-assistant start-mail-sync stop status
|
||||||
web/search (deprecated shim → web/search.go)
|
web/search (deprecated shim → web/search.go)
|
||||||
db/psql-yq (vendored)
|
db/psql-yq (vendored)
|
||||||
ssh-tunnel onlyoffice pg tunnel 5433
|
ssh-tunnel onlyoffice pg tunnel 5433
|
||||||
@@ -96,7 +104,8 @@ them from each hit (1=File, 2=Commit, 3=Person). Rebuild writes
|
|||||||
`Leaf-[:FROM_FILE]->File`; git import writes the rest.
|
`Leaf-[:FROM_FILE]->File`; git import writes the rest.
|
||||||
|
|
||||||
Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
|
Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
|
||||||
`where`, `when`, `source_rev`.
|
`where`, `when`, `source_rev`. Leaf interval of truth (D24): `valid_from`,
|
||||||
|
`valid_to`.
|
||||||
|
|
||||||
## Config
|
## Config
|
||||||
|
|
||||||
@@ -117,8 +126,8 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
|
|||||||
|
|
||||||
## Open questions (v2)
|
## Open questions (v2)
|
||||||
|
|
||||||
- OQ1: mutually-contradicting evidence — how to resolve (authority weighting,
|
- OQ1: **in** — D16 adjudication: `temporal_freshness` then `authority_pairing`.
|
||||||
temporal freshness, audit adjudication). **v2**; [#29](https://git.produktor.io/eSlider/2dph/issues/29).
|
Unresolved 2v2 stays hypothesis. [#29](https://git.produktor.io/eSlider/2dph/issues/29).
|
||||||
- OQ2: OCR — **in**. `pdftotext -layout` first; scans `pdftoppm` + tesseract
|
- OQ2: OCR — **in**. `pdftotext -layout` first; scans `pdftoppm` + tesseract
|
||||||
`eng+deu` (`bin/mail/ocr.go`, `internal/ocr`). No gocv, no gosseract CGO
|
`eng+deu` (`bin/mail/ocr.go`, `internal/ocr`). No gocv, no gosseract CGO
|
||||||
(D21 Zig owns Ladybug CGO). Optional `OCR_ENGINE=paddle` / compose profile
|
(D21 Zig owns Ladybug CGO). Optional `OCR_ENGINE=paddle` / compose profile
|
||||||
@@ -127,11 +136,14 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
|
|||||||
quantiles / JSONL count. Not a second graph. [#30](https://git.produktor.io/eSlider/2dph/issues/30).
|
quantiles / JSONL count. Not a second graph. [#30](https://git.produktor.io/eSlider/2dph/issues/30).
|
||||||
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
|
- OQ4: YAML-first storage for leafs — deferred: JSON is ~10x faster to
|
||||||
serialize and unambiguous; YAML only where humans edit files.
|
serialize and unambiguous; YAML only where humans edit files.
|
||||||
|
- OQ5: **in** — fact `valid_from` / `valid_to` + `--as-of` / MCP `as_of` (D24).
|
||||||
|
Not D16 `temporal_freshness`. [#36](https://git.produktor.io/eSlider/2dph/issues/36).
|
||||||
|
|
||||||
## Mail pipeline (done)
|
## Mail pipeline (done)
|
||||||
|
|
||||||
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail/OnlyOffice download.
|
1. `bin/mail/sync.go` (Go, 8 workers) — paginated Gmail / OnlyOffice / M365
|
||||||
Gmail attachments key off `body.attachmentId`, not MIME `partId`.
|
Graph download. Gmail attachments key off `body.attachmentId`, not MIME
|
||||||
|
`partId`. M365 uses client-credentials + delta link (commit after success).
|
||||||
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
|
2. `bin/mail/import.go --from-raw` — message.json → message.md; PDFs via
|
||||||
`pdftotext -layout` (~15ms); textless/scanned PDFs `pdftoppm` + tesseract
|
`pdftotext -layout` (~15ms); textless/scanned PDFs `pdftoppm` + tesseract
|
||||||
`eng+deu`. ICS sidecars
|
`eng+deu`. ICS sidecars
|
||||||
@@ -140,7 +152,11 @@ Common props on every node/edge: `root`, `confidence`, `evidence[]`, `how`,
|
|||||||
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
|
corrupts its WAL on bulk-insert into an already-indexed DB. Conversion and
|
||||||
indexing stay separate for crash safety. `bin/mail/index_mail` is a
|
indexing stay separate for crash safety. `bin/mail/index_mail` is a
|
||||||
deprecation shim.
|
deprecation shim.
|
||||||
4. Result: 17,835 messages → 28,918 info leafs, FTS + HNSW healthy, searchable
|
4. Compose `mail-sync` / `bin/stack/start-mail-sync` — ETL loop (default
|
||||||
|
`onlyoffice,gmail`, 300s): sync → import on `new>0`; full rebuild only if
|
||||||
|
`MAIL_SYNC_INDEX=1`. Bot digests (ai-bot) and case wrappers reuse sync/OAuth;
|
||||||
|
they do not replace the corpus path.
|
||||||
|
5. Result: 17,835 messages → 28,918 info leafs, FTS + HNSW healthy, searchable
|
||||||
via `bin/brain/search.go`.
|
via `bin/brain/search.go`.
|
||||||
|
|
||||||
## CI/CD pipeline (D15)
|
## CI/CD pipeline (D15)
|
||||||
@@ -184,4 +200,4 @@ Narrative: [docs/roadmap.md](docs/roadmap.md).
|
|||||||
| 4 | [#15](https://git.produktor.io/eSlider/2dph/issues/15) | **in** — lever/loop documented (`search` → `get` → `audit`). |
|
| 4 | [#15](https://git.produktor.io/eSlider/2dph/issues/15) | **in** — lever/loop documented (`search` → `get` → `audit`). |
|
||||||
| 5 | [#19](https://git.produktor.io/eSlider/2dph/issues/19) | **in** — CI recall SoT is `bin/brain/eval.go` via Zig. Python `bin/kb/eval` stays as an explicit fallback. |
|
| 5 | [#19](https://git.produktor.io/eSlider/2dph/issues/19) | **in** — CI recall SoT is `bin/brain/eval.go` via Zig. Python `bin/kb/eval` stays as an explicit fallback. |
|
||||||
|
|
||||||
Does **not** block epic close: OQ1 [#29](https://git.produktor.io/eSlider/2dph/issues/29), OQ3 [#30](https://git.produktor.io/eSlider/2dph/issues/30), OQ4. OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6) is **in**.
|
Does **not** block epic close: OQ4. OCR [#6](https://git.produktor.io/eSlider/2dph/issues/6), OQ1 [#29](https://git.produktor.io/eSlider/2dph/issues/29), OQ3 [#30](https://git.produktor.io/eSlider/2dph/issues/30) are **in**.
|
||||||
@@ -119,6 +119,8 @@ Mail is a first-class corpus (retrievable through the same search):
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
|
bin/mail/sync.go --source onlyoffice,gmail --workers 8 --out var/mail # raw sync (Go)
|
||||||
|
bin/mail/sync.go --source m365 --env ~/.config/brain/mail.env # Microsoft 365 Graph
|
||||||
|
bin/stack/start-mail-sync # compose ETL (300s; no auto-rebuild)
|
||||||
bin/mail/import.go --from-raw var/mail # JSON → markdown
|
bin/mail/import.go --from-raw var/mail # JSON → markdown
|
||||||
bin/brain/add.go --text T --root facts --source "a.md x b.md"
|
bin/brain/add.go --text T --root facts --source "a.md x b.md"
|
||||||
bin/brain/index.go --rebuild --with-facts --with-chats # facts extract + chats md
|
bin/brain/index.go --rebuild --with-facts --with-chats # facts extract + chats md
|
||||||
@@ -160,6 +162,10 @@ go test ./... && uv run python -m unittest discover -s bin/tools -t .
|
|||||||
Docker (optional, cached model + var volumes):
|
Docker (optional, cached model + var volumes):
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
|
bin/stack/start # brain HTTP/MCP :8630
|
||||||
|
bin/stack/start-assistant # + qwen3.5:9b + PicoClaw agent
|
||||||
|
bin/stack/status
|
||||||
|
bin/stack/stop
|
||||||
docker compose up -d brain # API (Zig CGO serve :8630)
|
docker compose up -d brain # API (Zig CGO serve :8630)
|
||||||
docker compose --profile index run --rm index # Python Ladybug rebuild
|
docker compose --profile index run --rm index # Python Ladybug rebuild
|
||||||
docker compose --profile picoclaw up brain-mcp # MCP on 127.0.0.1:8630
|
docker compose --profile picoclaw up brain-mcp # MCP on 127.0.0.1:8630
|
||||||
|
|||||||
Executable
+91
@@ -0,0 +1,91 @@
|
|||||||
|
//usr/bin/env go run "$0" "$@"; exit
|
||||||
|
//
|
||||||
|
// bin/cli/complete.go - dump flaggy shell completions for all Go shebang tools (D23).
|
||||||
|
//
|
||||||
|
// source <(./bin/cli/complete.go bash)
|
||||||
|
// ./bin/cli/complete.go zsh|fish|powershell|nushell
|
||||||
|
//
|
||||||
|
// Search does not steal the word "completion"; this binary dumps scripts.
|
||||||
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
package main
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
"os"
|
||||||
|
"strings"
|
||||||
|
|
||||||
|
mailsync "github.com/eSlider/2dph/bin/mail/sync"
|
||||||
|
"github.com/eSlider/2dph/internal/brain/rank"
|
||||||
|
"github.com/eSlider/2dph/internal/chats"
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/eSlider/2dph/internal/gitlog"
|
||||||
|
"github.com/eSlider/2dph/internal/mdleaves"
|
||||||
|
"github.com/eSlider/2dph/internal/ocr"
|
||||||
|
"github.com/eSlider/2dph/internal/reasoner"
|
||||||
|
"github.com/eSlider/2dph/internal/websearch"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
func tools() []cli.Tool {
|
||||||
|
return []cli.Tool{
|
||||||
|
{Path: "bin/brain/search.go", Name: "brain-search", New: rank.Parser},
|
||||||
|
{Path: "bin/brain/get.go", Name: "brain-get", New: func() *flaggy.Parser {
|
||||||
|
o := rank.GetOptions{}
|
||||||
|
return rank.GetParser(&o)
|
||||||
|
}},
|
||||||
|
{Path: "bin/brain/stats.go", Name: "brain-stats", New: rank.StatsParser},
|
||||||
|
{Path: "bin/brain/eval.go", Name: "brain-eval", New: rank.EvalParser},
|
||||||
|
{Path: "bin/web/search.go", Name: "web-search", New: websearch.Parser},
|
||||||
|
{Path: "bin/git/import.go", Name: "git-import", New: gitlog.Parser},
|
||||||
|
{Path: "bin/markdown/import.go", Name: "markdown-import", New: mdleaves.Parser},
|
||||||
|
{Path: "bin/qa/stats.go", Name: "qa-stats", New: cli.QAParser},
|
||||||
|
{Path: "bin/reasoner/bakeoff.go", Name: "reasoner-bakeoff", New: reasoner.Parser},
|
||||||
|
{Path: "bin/mail/ocr.go", Name: "mail-ocr", New: ocr.Parser},
|
||||||
|
{Path: "bin/mail/sync.go", Name: "mail-sync", New: mailsync.Parser},
|
||||||
|
{Path: "bin/chats/sync.go", Name: "chats-sync", New: chats.SyncParser},
|
||||||
|
{Path: "bin/chats/import.go", Name: "chats-import", New: chats.ImportParser},
|
||||||
|
{Path: "bin/chats/facts.go", Name: "chats-facts", New: chats.FactsParser},
|
||||||
|
{Path: "bin/chats/apply.go", Name: "chats-apply", New: chats.ApplyParser},
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func main() {
|
||||||
|
os.Exit(run(os.Args[1:]))
|
||||||
|
}
|
||||||
|
|
||||||
|
func run(args []string) int {
|
||||||
|
shell := "bash"
|
||||||
|
if len(args) > 0 {
|
||||||
|
switch args[0] {
|
||||||
|
case "bash", "zsh", "fish", "powershell", "nushell":
|
||||||
|
shell = args[0]
|
||||||
|
case "-h", "--help", "help":
|
||||||
|
fmt.Fprintln(os.Stderr, "usage: bin/cli/complete.go [bash|zsh|fish|powershell|nushell]")
|
||||||
|
return 0
|
||||||
|
default:
|
||||||
|
fmt.Fprintf(os.Stderr, "cli/complete: unknown shell %q\n", args[0])
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if shell == "bash" {
|
||||||
|
fmt.Print(cli.BashScript(tools()))
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
var b strings.Builder
|
||||||
|
for _, t := range tools() {
|
||||||
|
p := t.New()
|
||||||
|
p.Name = t.Name
|
||||||
|
switch shell {
|
||||||
|
case "zsh":
|
||||||
|
b.WriteString(flaggy.GenerateZshCompletion(p))
|
||||||
|
case "fish":
|
||||||
|
b.WriteString(flaggy.GenerateFishCompletion(p))
|
||||||
|
case "powershell":
|
||||||
|
b.WriteString(flaggy.GeneratePowerShellCompletion(p))
|
||||||
|
case "nushell":
|
||||||
|
b.WriteString(flaggy.GenerateNushellCompletion(p))
|
||||||
|
}
|
||||||
|
}
|
||||||
|
fmt.Print(b.String())
|
||||||
|
return 0
|
||||||
|
}
|
||||||
@@ -5,6 +5,7 @@
|
|||||||
# serve | search | watch
|
# serve | search | watch
|
||||||
# Index image (Python write path, compose profile `index`):
|
# Index image (Python write path, compose profile `index`):
|
||||||
# index | extract | audit | search (deprecated python wrapper)
|
# index | extract | audit | search (deprecated python wrapper)
|
||||||
|
# mail-sync [N] ETL loop: sync -> import; optional index (default 300s)
|
||||||
#
|
#
|
||||||
# Usage comment starts at line 2 (self-describing convention).
|
# Usage comment starts at line 2 (self-describing convention).
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
@@ -34,5 +35,27 @@ case "$CMD" in
|
|||||||
serve) exec /app/bin/serve "$@" ;;
|
serve) exec /app/bin/serve "$@" ;;
|
||||||
extract) exec "$KB_PY" /app/bin/facts/extract "$@" ;;
|
extract) exec "$KB_PY" /app/bin/facts/extract "$@" ;;
|
||||||
audit) exec "$KB_PY" /app/bin/facts/audit "$@" ;;
|
audit) exec "$KB_PY" /app/bin/facts/audit "$@" ;;
|
||||||
|
mail-sync)
|
||||||
|
# ETL: pull mail, convert to md when new>0. Full --rebuild is opt-in
|
||||||
|
# (MAIL_SYNC_INDEX=1) — ~29k leaf rebuild is minutes, not a 10s loop.
|
||||||
|
interval="${1:-300}"
|
||||||
|
[ "$interval" -gt 0 ] 2>/dev/null || interval=300
|
||||||
|
: "${MAIL_SYNC_ENV:=/secret/mail.env}"
|
||||||
|
: "${MAIL_SYNC_SRC:=onlyoffice,gmail}"
|
||||||
|
: "${MAIL_SYNC_OUT:=/app/var/mail}"
|
||||||
|
: "${MAIL_SYNC_INDEX:=0}"
|
||||||
|
while true; do
|
||||||
|
out="$("/app/bin/mail-sync" --source "$MAIL_SYNC_SRC" --env "$MAIL_SYNC_ENV" --out "$MAIL_SYNC_OUT" 2>&1)" || true
|
||||||
|
echo "$out"
|
||||||
|
new="$(printf '%s\n' "$out" | sed -n 's/.*new=\([0-9]*\).*/\1/p' | tail -1)"
|
||||||
|
if [ -n "$new" ] && [ "$new" -gt 0 ] 2>/dev/null; then
|
||||||
|
"$KB_PY" /app/bin/mail/import --from-raw "$MAIL_SYNC_OUT" 2>&1 | tail -1
|
||||||
|
if [ "$MAIL_SYNC_INDEX" = "1" ]; then
|
||||||
|
"$KB_PY" /app/bin/kb/index --rebuild --with-mail 2>&1 | tail -1
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
sleep "$interval"
|
||||||
|
done
|
||||||
|
;;
|
||||||
*) echo "unknown command: $CMD" >&2; exit 2 ;;
|
*) echo "unknown command: $CMD" >&2; exit 2 ;;
|
||||||
esac
|
esac
|
||||||
|
|||||||
+45
-18
@@ -1,14 +1,15 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""facts/audit - evidence & lexicon checks for the 2dph brain.
|
"""facts/audit - evidence & lexicon checks for the 2dph brain.
|
||||||
|
|
||||||
bin/facts/audit self # lexicon: every fact in db has >=2 sources
|
bin/facts/audit self # lexicon: docs + two-source rule
|
||||||
bin/facts/audit db # evidence gate: run against var/kb.lbug
|
bin/facts/audit db # evidence gate against var/kb.lbug
|
||||||
|
bin/facts/audit contradict # D16 adjudication (JSON claim(s) on stdin)
|
||||||
|
|
||||||
`self` mode checks the repo itself (no network, no runtime deps). It greps
|
`self` mode checks the repo itself (no network, no runtime deps).
|
||||||
for known-good two-source pairings and confirms the docs are consistent.
|
`db` mode loads every Leaf with root=facts. Confirmed facts need ` x `;
|
||||||
`db` mode loads every Leaf with root=facts and asserts each has source_rev
|
hypothesis contradictions need `a x b vs c x d` (both sides ≥2).
|
||||||
and a non-empty `loc` (the "where did you see it" evidence pointer) and that
|
`contradict` applies temporal_freshness then authority_pairing; ≥2 vs ≥2
|
||||||
'confirmed' facts carry a two-source `source` field.
|
with no rule stays hypothesis / `(not confirmed)`.
|
||||||
|
|
||||||
Exit 0 = all checks pass, 1 = audit failures, 2 = could not evaluate.
|
Exit 0 = all checks pass, 1 = audit failures, 2 = could not evaluate.
|
||||||
"""
|
"""
|
||||||
@@ -22,6 +23,8 @@ from pathlib import Path
|
|||||||
ROOT = Path(__file__).resolve().parents[2]
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
sys.path.insert(0, str(ROOT / "bin" / "tools"))
|
||||||
|
|
||||||
|
from contradict import adjudicate, check_fact_row # noqa: E402
|
||||||
|
|
||||||
|
|
||||||
def audit_db() -> list[str]:
|
def audit_db() -> list[str]:
|
||||||
from kblib import connect
|
from kblib import connect
|
||||||
@@ -33,14 +36,8 @@ def audit_db() -> list[str]:
|
|||||||
r = conn.execute("MATCH (l:Leaf {root:'facts'}) RETURN l.id, l.source, l.loc, l.how, l.confidence")
|
r = conn.execute("MATCH (l:Leaf {root:'facts'}) RETURN l.id, l.source, l.loc, l.how, l.confidence")
|
||||||
problems: list[str] = []
|
problems: list[str] = []
|
||||||
for lid, source, loc, how, conf in r.get_all():
|
for lid, source, loc, how, conf in r.get_all():
|
||||||
if conf != "confirmed":
|
problems.extend(check_fact_row(str(lid), str(source or ""), str(loc or ""),
|
||||||
problems.append(f"{lid}: facts require confidence='confirmed', got '{conf}'")
|
str(how or ""), str(conf or "")))
|
||||||
if not source or " x " not in source:
|
|
||||||
problems.append(f"{lid}: needs 2-source evidence in source, got '{source}'")
|
|
||||||
if not loc:
|
|
||||||
problems.append(f"{lid}: missing loc (evidence pointer)")
|
|
||||||
if not how:
|
|
||||||
problems.append(f"{lid}: missing how")
|
|
||||||
conn.close()
|
conn.close()
|
||||||
db.close()
|
db.close()
|
||||||
return problems
|
return problems
|
||||||
@@ -54,20 +51,50 @@ def audit_self() -> list[str]:
|
|||||||
problems.append("PLAN.md missing recall@5 gate")
|
problems.append("PLAN.md missing recall@5 gate")
|
||||||
if re.search(r"(?i)facts must have.*2 sources|2.source", plan) is None:
|
if re.search(r"(?i)facts must have.*2 sources|2.source", plan) is None:
|
||||||
problems.append("PLAN.md missing the two-source evidence rule for facts")
|
problems.append("PLAN.md missing the two-source evidence rule for facts")
|
||||||
|
if "temporal_freshness" not in plan or "authority_pairing" not in plan:
|
||||||
|
problems.append("PLAN.md missing D16 adjudication rules")
|
||||||
if re.search(r"(?i)HNSW|BM25|deduction", (ROOT / "README.md").read_text()) is None:
|
if re.search(r"(?i)HNSW|BM25|deduction", (ROOT / "README.md").read_text()) is None:
|
||||||
problems.append("README.md missing search/retrieval description")
|
problems.append("README.md missing search/retrieval description")
|
||||||
return problems
|
return problems
|
||||||
|
|
||||||
|
|
||||||
|
def audit_contradict(raw: str) -> tuple[list[str], list[dict]]:
|
||||||
|
raw = raw.strip()
|
||||||
|
if not raw:
|
||||||
|
return ["contradict: empty stdin (JSON claim or {claims:[...]})"], []
|
||||||
|
try:
|
||||||
|
payload = json.loads(raw)
|
||||||
|
except json.JSONDecodeError as e:
|
||||||
|
return [f"contradict: invalid JSON: {e}"], []
|
||||||
|
if isinstance(payload, dict) and "claims" in payload:
|
||||||
|
claims = list(payload.get("claims") or [])
|
||||||
|
elif isinstance(payload, dict):
|
||||||
|
claims = [payload]
|
||||||
|
elif isinstance(payload, list):
|
||||||
|
claims = payload
|
||||||
|
else:
|
||||||
|
return ["contradict: expected object or list"], []
|
||||||
|
details = [adjudicate(c) for c in claims]
|
||||||
|
return [], details
|
||||||
|
|
||||||
|
|
||||||
def main(argv: list[str]) -> int:
|
def main(argv: list[str]) -> int:
|
||||||
import argparse
|
import argparse
|
||||||
p = argparse.ArgumentParser(description="evidence & lexicon audit")
|
p = argparse.ArgumentParser(description="evidence & lexicon audit")
|
||||||
p.add_argument("mode", choices=("self", "db"))
|
p.add_argument("mode", choices=("self", "db", "contradict"))
|
||||||
p.add_argument("--json", action="store_true")
|
p.add_argument("--json", action="store_true")
|
||||||
a = p.parse_args(argv)
|
a = p.parse_args(argv)
|
||||||
|
|
||||||
problems = audit_self() if a.mode == "self" else audit_db()
|
details: list[dict] = []
|
||||||
out = {"mode": a.mode, "ok": not problems, "problems": problems}
|
if a.mode == "self":
|
||||||
|
problems = audit_self()
|
||||||
|
elif a.mode == "db":
|
||||||
|
problems = audit_db()
|
||||||
|
else:
|
||||||
|
problems, details = audit_contradict(sys.stdin.read())
|
||||||
|
out: dict = {"mode": a.mode, "ok": not problems, "problems": problems}
|
||||||
|
if details:
|
||||||
|
out["contradictions"] = details
|
||||||
if a.json:
|
if a.json:
|
||||||
print(json.dumps(out, indent=2))
|
print(json.dumps(out, indent=2))
|
||||||
else:
|
else:
|
||||||
|
|||||||
@@ -5,6 +5,7 @@
|
|||||||
//
|
//
|
||||||
// ./bin/facts/audit.go self
|
// ./bin/facts/audit.go self
|
||||||
// ./bin/facts/audit.go db
|
// ./bin/facts/audit.go db
|
||||||
|
// ./bin/facts/audit.go contradict --json < claim.json
|
||||||
//
|
//
|
||||||
// Python bin/facts/audit is the implementation (CI runs it directly).
|
// Python bin/facts/audit is the implementation (CI runs it directly).
|
||||||
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
// NOTE: never run `gofmt -w` on this file — it breaks the shebang.
|
||||||
|
|||||||
+9
-52
@@ -16,9 +16,8 @@ import (
|
|||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"strconv"
|
|
||||||
"time"
|
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
"github.com/eSlider/2dph/internal/cmdbin"
|
"github.com/eSlider/2dph/internal/cmdbin"
|
||||||
"github.com/eSlider/2dph/internal/gitlog"
|
"github.com/eSlider/2dph/internal/gitlog"
|
||||||
)
|
)
|
||||||
@@ -28,49 +27,16 @@ func main() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func run(args []string) int {
|
func run(args []string) int {
|
||||||
var repo, root, since string
|
c, err := gitlog.ParseArgs(args)
|
||||||
limit := 0
|
if err != nil {
|
||||||
jsonOut := false
|
return cliparse.Fail(err)
|
||||||
i := 0
|
|
||||||
for i < len(args) {
|
|
||||||
a := args[i]
|
|
||||||
switch {
|
|
||||||
case a == "--json":
|
|
||||||
jsonOut = true
|
|
||||||
case a == "--limit" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
n, err := strconv.Atoi(args[i])
|
|
||||||
if err != nil || n < 0 {
|
|
||||||
fmt.Fprintf(os.Stderr, "git/import: --limit must be a non-negative integer\n")
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
limit = n
|
|
||||||
case a == "--since" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
since = args[i]
|
|
||||||
case a == "--root" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
root = args[i]
|
|
||||||
case a == "-h" || a == "--help":
|
|
||||||
fmt.Fprintln(os.Stderr, `usage: bin/git/import.go [REPO] [--json] [--limit N] [--since DATE] [--root DIR]`)
|
|
||||||
return 0
|
|
||||||
case len(a) > 0 && a[0] != '-':
|
|
||||||
repo = a
|
|
||||||
default:
|
|
||||||
fmt.Fprintf(os.Stderr, "git/import: unknown flag %s\n", a)
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
i++
|
|
||||||
}
|
}
|
||||||
|
repo, root, since, limit, jsonOut := c.Repo, c.Root, c.Since, c.Limit, c.JSONOut
|
||||||
|
|
||||||
var sinceT time.Time
|
sinceT, err := gitlog.ParseSince(since)
|
||||||
if since != "" {
|
if err != nil {
|
||||||
var err error
|
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
|
||||||
sinceT, err = parseSince(since)
|
return 2
|
||||||
if err != nil {
|
|
||||||
fmt.Fprintf(os.Stderr, "git/import: %v\n", err)
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
repos := []string{}
|
repos := []string{}
|
||||||
@@ -132,12 +98,3 @@ func run(args []string) int {
|
|||||||
}
|
}
|
||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
|
|
||||||
func parseSince(s string) (time.Time, error) {
|
|
||||||
for _, layout := range []string{time.RFC3339, "2006-01-02"} {
|
|
||||||
if t, err := time.Parse(layout, s); err == nil {
|
|
||||||
return t, nil
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return time.Time{}, fmt.Errorf("cannot parse --since %q", s)
|
|
||||||
}
|
|
||||||
|
|||||||
@@ -65,6 +65,10 @@ def main(argv: list[str]) -> int:
|
|||||||
p.add_argument("--how", default="brain/add")
|
p.add_argument("--how", default="brain/add")
|
||||||
p.add_argument("--loc", default="")
|
p.add_argument("--loc", default="")
|
||||||
p.add_argument("--type", default="reference", dest="type_")
|
p.add_argument("--type", default="reference", dest="type_")
|
||||||
|
p.add_argument("--valid-from", default="", dest="valid_from",
|
||||||
|
help="fact interval start YYYY-MM-DD (D24)")
|
||||||
|
p.add_argument("--valid-to", default="", dest="valid_to",
|
||||||
|
help="fact interval end YYYY-MM-DD inclusive; empty=open (D24)")
|
||||||
args = p.parse_args(argv)
|
args = p.parse_args(argv)
|
||||||
|
|
||||||
if args.json:
|
if args.json:
|
||||||
@@ -86,6 +90,8 @@ def main(argv: list[str]) -> int:
|
|||||||
"how": args.how,
|
"how": args.how,
|
||||||
"loc": args.loc or args.source,
|
"loc": args.loc or args.source,
|
||||||
"type": args.type_,
|
"type": args.type_,
|
||||||
|
"valid_from": args.valid_from,
|
||||||
|
"valid_to": args.valid_to,
|
||||||
}]
|
}]
|
||||||
|
|
||||||
for lf in leafs:
|
for lf in leafs:
|
||||||
|
|||||||
+6
-8
@@ -17,6 +17,7 @@ import (
|
|||||||
"os"
|
"os"
|
||||||
"strings"
|
"strings"
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
"github.com/eSlider/2dph/internal/ocr"
|
"github.com/eSlider/2dph/internal/ocr"
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -25,15 +26,12 @@ func main() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func run(args []string) int {
|
func run(args []string) int {
|
||||||
if len(args) != 1 || strings.HasPrefix(args[0], "-") {
|
c, err := ocr.ParseArgs(args)
|
||||||
fmt.Fprintln(os.Stderr, `usage: bin/mail/ocr.go <image|pdf>`)
|
if err != nil {
|
||||||
return 2
|
return cliparse.Fail(err)
|
||||||
}
|
}
|
||||||
path := args[0]
|
path := c.Path
|
||||||
var (
|
var text string
|
||||||
text string
|
|
||||||
err error
|
|
||||||
)
|
|
||||||
if strings.HasSuffix(strings.ToLower(path), ".pdf") {
|
if strings.HasSuffix(strings.ToLower(path), ".pdf") {
|
||||||
text, err = ocr.PDFFile(path)
|
text, err = ocr.PDFFile(path)
|
||||||
} else {
|
} else {
|
||||||
|
|||||||
+3
-2
@@ -1,8 +1,9 @@
|
|||||||
//usr/bin/env go run "$0" "$@"; exit
|
//usr/bin/env go run "$0" "$@"; exit
|
||||||
// bin/mail/sync.go - async download of OnlyOffice and Gmail mail to var/mail/.
|
// bin/mail/sync.go - async download of OnlyOffice, Gmail and M365 mail to var/mail/.
|
||||||
//
|
//
|
||||||
// ./bin/mail/sync.go --source onlyoffice,gmail --limit 50 --workers 8
|
// ./bin/mail/sync.go --source onlyoffice,gmail,m365 --limit 50 --workers 8
|
||||||
// ./bin/mail/sync.go --source gmail --force
|
// ./bin/mail/sync.go --source gmail --force
|
||||||
|
// ./bin/mail/sync.go --source m365 --env .secrets/m365.env
|
||||||
// ./bin/mail/sync.go --dry-run
|
// ./bin/mail/sync.go --dry-run
|
||||||
//
|
//
|
||||||
// Writes raw message.json + attachments under var/mail/<folder>/<id>/; run
|
// Writes raw message.json + attachments under var/mail/<folder>/<id>/; run
|
||||||
|
|||||||
+92
-45
@@ -5,74 +5,103 @@ package sync
|
|||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
"flag"
|
"errors"
|
||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"strings"
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
)
|
)
|
||||||
|
|
||||||
// CLIConfig is a superset of SyncConfig plus flag parsing results.
|
// CLIConfig is a superset of SyncConfig plus flag parsing results.
|
||||||
type CLIConfig struct {
|
type CLIConfig struct {
|
||||||
Sync SyncConfig
|
Sync SyncConfig
|
||||||
Env string // .env path; default <cwd>/.env
|
Env string // .env path; default <cwd>/.env
|
||||||
Sources string
|
Sources string
|
||||||
Help bool
|
Help bool
|
||||||
}
|
}
|
||||||
|
|
||||||
// ParseCLI reads os.Args into a CLIConfig. Exit codes: 0 ok, 2 usage.
|
type flagVals struct {
|
||||||
|
env, out, srcs, query string
|
||||||
|
workers, limit, offset int
|
||||||
|
force, dryRun bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func Parser() *flaggy.Parser {
|
||||||
|
v := flagVals{workers: 4, query: "in:inbox", srcs: "onlyoffice"}
|
||||||
|
return bind(&v)
|
||||||
|
}
|
||||||
|
|
||||||
|
func bind(v *flagVals) *flaggy.Parser {
|
||||||
|
if v.workers == 0 {
|
||||||
|
v.workers = 4
|
||||||
|
}
|
||||||
|
if v.query == "" {
|
||||||
|
v.query = "in:inbox"
|
||||||
|
}
|
||||||
|
if v.srcs == "" {
|
||||||
|
v.srcs = "onlyoffice"
|
||||||
|
}
|
||||||
|
p := cliparse.New("mail-sync")
|
||||||
|
p.Description = "download mail to var/mail"
|
||||||
|
p.String(&v.env, "", "env", ".env file")
|
||||||
|
p.String(&v.out, "", "out", "var/mail root")
|
||||||
|
p.Int(&v.workers, "", "workers", "concurrent downloads")
|
||||||
|
p.Int(&v.limit, "", "limit", "max messages per source (0 = all)")
|
||||||
|
p.Int(&v.offset, "", "offset", "skip first N messages per source")
|
||||||
|
p.Bool(&v.force, "", "force", "overwrite existing message.json")
|
||||||
|
p.Bool(&v.dryRun, "", "dry-run", "list counts without writing")
|
||||||
|
p.String(&v.query, "", "query", "Gmail search query")
|
||||||
|
p.String(&v.srcs, "", "source", "comma list: onlyoffice,gmail,m365")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
// ParseCLI reads args into a CLIConfig. Exit codes: 0 ok, 2 usage.
|
||||||
func ParseCLI(args []string) (CLIConfig, int, error) {
|
func ParseCLI(args []string) (CLIConfig, int, error) {
|
||||||
fs := flag.NewFlagSet("mail/sync", flag.ContinueOnError)
|
v := flagVals{workers: 4, query: "in:inbox", srcs: "onlyoffice"}
|
||||||
var (
|
p := bind(&v)
|
||||||
env = fs.String("env", "", ".env file (default: <cwd>/.env)")
|
if err := cliparse.Parse(p, args); err != nil {
|
||||||
out = fs.String("out", "", "var/mail root (default: <cwd>/var/mail)")
|
if errors.Is(err, cliparse.ErrHelp) {
|
||||||
workers = fs.Int("workers", 4, "concurrent downloads")
|
return CLIConfig{Help: true}, 0, nil
|
||||||
limit = fs.Int("limit", 0, "max messages per source (0 = all)")
|
}
|
||||||
offset = fs.Int("offset", 0, "skip first N messages per source")
|
|
||||||
force = fs.Bool("force", false, "overwrite existing message.json + attachments")
|
|
||||||
dryRun = fs.Bool("dry-run", false, "list message counts without writing")
|
|
||||||
query = fs.String("query", "in:inbox", "Gmail search query (gmail source only)")
|
|
||||||
srcs = fs.String("source", "onlyoffice", "comma list: onlyoffice,gmail (default onlyoffice)")
|
|
||||||
help = fs.Bool("help", false, "usage")
|
|
||||||
)
|
|
||||||
fs.SetOutput(os.Stderr)
|
|
||||||
if err := fs.Parse(args); err != nil {
|
|
||||||
return CLIConfig{}, 2, err
|
return CLIConfig{}, 2, err
|
||||||
}
|
}
|
||||||
if *help || fs.NArg() > 0 {
|
if len(p.TrailingArguments) > 0 {
|
||||||
return CLIConfig{Help: true}, 0, nil
|
return CLIConfig{Help: true}, 0, nil
|
||||||
}
|
}
|
||||||
wd, err := os.Getwd()
|
wd, err := os.Getwd()
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return CLIConfig{}, 2, err
|
return CLIConfig{}, 2, err
|
||||||
}
|
}
|
||||||
if *env == "" {
|
if v.env == "" {
|
||||||
*env = filepath.Join(wd, ".env")
|
v.env = filepath.Join(wd, ".env")
|
||||||
}
|
}
|
||||||
if *out == "" {
|
if v.out == "" {
|
||||||
*out = filepath.Join(wd, "var", "mail")
|
v.out = filepath.Join(wd, "var", "mail")
|
||||||
}
|
}
|
||||||
envVars := readEnv(*env)
|
envVars := readEnv(v.env)
|
||||||
cfg := SyncConfig{
|
cfg := SyncConfig{
|
||||||
Out: *out,
|
Out: v.out,
|
||||||
Workers: *workers,
|
Workers: v.workers,
|
||||||
Limit: *limit,
|
Limit: v.limit,
|
||||||
Offset: *offset,
|
Offset: v.offset,
|
||||||
Force: *force,
|
Force: v.force,
|
||||||
DryRun: *dryRun,
|
DryRun: v.dryRun,
|
||||||
Query: *query,
|
Query: v.query,
|
||||||
Policy: RetryPolicy{},
|
Policy: RetryPolicy{},
|
||||||
}
|
}
|
||||||
cli := CLIConfig{Sync: cfg, Env: *env, Sources: *srcs}
|
out := CLIConfig{Sync: cfg, Env: v.env, Sources: v.srcs}
|
||||||
for _, s := range strings.Split(*srcs, ",") {
|
for _, s := range strings.Split(v.srcs, ",") {
|
||||||
switch strings.TrimSpace(s) {
|
switch strings.TrimSpace(s) {
|
||||||
case "onlyoffice":
|
case "onlyoffice":
|
||||||
u := pick(envVars["ONLYOFFICE_URL"], envVars["OO_URL"])
|
u := pick(envVars["ONLYOFFICE_URL"], envVars["OO_URL"])
|
||||||
user := pick(envVars["ONLYOFFICE_USER"], envVars["OO_USER"])
|
user := pick(envVars["ONLYOFFICE_USER"], envVars["OO_USER"])
|
||||||
pass := pick(envVars["ONLYOFFICE_PASS"], envVars["OO_PASSWORD"])
|
pass := pick(envVars["ONLYOFFICE_PASS"], envVars["OO_PASSWORD"])
|
||||||
if u == "" || user == "" || pass == "" {
|
if u == "" || user == "" || pass == "" {
|
||||||
return CLIConfig{}, 2, fmt.Errorf("onlyoffice source needs ONLYOFFICE_URL/USER/PASS in %s", *env)
|
return CLIConfig{}, 2, fmt.Errorf("onlyoffice source needs ONLYOFFICE_URL/USER/PASS in %s", v.env)
|
||||||
}
|
}
|
||||||
cfg.OO = &OOConfig{URL: u, User: user, Password: pass}
|
cfg.OO = &OOConfig{URL: u, User: user, Password: pass}
|
||||||
case "gmail":
|
case "gmail":
|
||||||
@@ -81,34 +110,52 @@ func ParseCLI(args []string) (CLIConfig, int, error) {
|
|||||||
CredentialsPath: filepath.Join(home, ".gmail-mcp", "credentials.json"),
|
CredentialsPath: filepath.Join(home, ".gmail-mcp", "credentials.json"),
|
||||||
KeysPath: filepath.Join(home, ".gmail-mcp", "gcp-oauth.keys.json"),
|
KeysPath: filepath.Join(home, ".gmail-mcp", "gcp-oauth.keys.json"),
|
||||||
}
|
}
|
||||||
|
case "m365":
|
||||||
|
tenant := pick(envVars["M365_TENANT"], envVars["MS_TENANT"])
|
||||||
|
cid := pick(envVars["M365_CLIENT_ID"], envVars["MS_CLIENT_ID"])
|
||||||
|
sec := pick(envVars["M365_CLIENT_SECRET"], envVars["MS_CLIENT_SECRET"])
|
||||||
|
users := pick(envVars["M365_USERS"], envVars["MS_USERS"])
|
||||||
|
if tenant == "" || cid == "" || sec == "" || users == "" {
|
||||||
|
return CLIConfig{}, 2, fmt.Errorf("m365 source needs M365_TENANT/CLIENT_ID/CLIENT_SECRET/USERS in %s", v.env)
|
||||||
|
}
|
||||||
|
var userList []string
|
||||||
|
for _, u := range strings.Split(users, ",") {
|
||||||
|
if u = strings.TrimSpace(u); u != "" {
|
||||||
|
userList = append(userList, u)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if len(userList) == 0 {
|
||||||
|
return CLIConfig{}, 2, fmt.Errorf("m365 source: M365_USERS empty")
|
||||||
|
}
|
||||||
|
cfg.M365 = &M365Credentials{Tenant: tenant, ClientID: cid, ClientSecret: sec, Users: userList}
|
||||||
default:
|
default:
|
||||||
return CLIConfig{}, 2, fmt.Errorf("unknown source %q", s)
|
return CLIConfig{}, 2, fmt.Errorf("unknown source %q", s)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
cli.Sync = cfg
|
out.Sync = cfg
|
||||||
return cli, 0, nil
|
return out, 0, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// Main is the CLI entry: returns process exit code.
|
// Main is the CLI entry: returns process exit code.
|
||||||
func Main(args []string) int {
|
func Main(args []string) int {
|
||||||
cli, code, err := ParseCLI(args)
|
cfg, code, err := ParseCLI(args)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
fmt.Fprintln(os.Stderr, "mail/sync:", err)
|
fmt.Fprintln(os.Stderr, "mail/sync:", err)
|
||||||
return code
|
return code
|
||||||
}
|
}
|
||||||
if cli.Help {
|
if cfg.Help {
|
||||||
fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail] [--query GMAIL_Q] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]")
|
fmt.Fprintln(os.Stderr, "usage: bin/mail/sync.go [--source onlyoffice,gmail,m365] [--query GMAIL_Q] [--limit N] [--offset N] [--workers N] [--force] [--dry-run]")
|
||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
|
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
|
||||||
defer cancel()
|
defer cancel()
|
||||||
start := time.Now()
|
start := time.Now()
|
||||||
stats, err := Run(ctx, cli.Sync)
|
stats, err := Run(ctx, cfg.Sync)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
fmt.Fprintln(os.Stderr, "mail/sync:", err)
|
fmt.Fprintln(os.Stderr, "mail/sync:", err)
|
||||||
return 1
|
return 1
|
||||||
}
|
}
|
||||||
if cli.Sync.DryRun {
|
if cfg.Sync.DryRun {
|
||||||
fmt.Printf("mail/sync: dry-run checked=%d (no writes)\n", stats.Checked)
|
fmt.Printf("mail/sync: dry-run checked=%d (no writes)\n", stats.Checked)
|
||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
@@ -135,13 +182,13 @@ func readEnv(path string) map[string]string {
|
|||||||
k, v, _ := strings.Cut(line, "=")
|
k, v, _ := strings.Cut(line, "=")
|
||||||
out[strings.TrimSpace(k)] = strings.Trim(strings.TrimSpace(v), "\"'")
|
out[strings.TrimSpace(k)] = strings.Trim(strings.TrimSpace(v), "\"'")
|
||||||
}
|
}
|
||||||
// env overrides file
|
|
||||||
for _, kv := range os.Environ() {
|
for _, kv := range os.Environ() {
|
||||||
k, v, ok := strings.Cut(kv, "=")
|
k, v, ok := strings.Cut(kv, "=")
|
||||||
if !ok {
|
if !ok {
|
||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
if strings.HasPrefix(k, "ONLYOFFICE_") || strings.HasPrefix(k, "OO_") {
|
if strings.HasPrefix(k, "ONLYOFFICE_") || strings.HasPrefix(k, "OO_") ||
|
||||||
|
strings.HasPrefix(k, "M365_") || strings.HasPrefix(k, "MS_") {
|
||||||
out[k] = v
|
out[k] = v
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,382 @@
|
|||||||
|
package sync
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"encoding/json"
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
"io"
|
||||||
|
"net/http"
|
||||||
|
"net/url"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"strings"
|
||||||
|
"time"
|
||||||
|
)
|
||||||
|
|
||||||
|
// M365Credentials holds a Microsoft Graph app registration with the Mail.Read
|
||||||
|
// application permission. Client credentials are read from env/.env, never
|
||||||
|
// committed.
|
||||||
|
type M365Credentials struct {
|
||||||
|
Tenant string
|
||||||
|
ClientID string
|
||||||
|
ClientSecret string
|
||||||
|
Users []string // mailbox addresses to sync, e.g. info@example.com
|
||||||
|
}
|
||||||
|
|
||||||
|
// m365Token is the cached access token with expiry.
|
||||||
|
type m365Token struct {
|
||||||
|
AccessToken string
|
||||||
|
Expiry time.Time
|
||||||
|
}
|
||||||
|
|
||||||
|
// M365Client talks to the Microsoft Graph API using the client-credentials
|
||||||
|
// flow (app registration with Mail.Read application permission). GET-only:
|
||||||
|
// messages are never marked as read or deleted.
|
||||||
|
type M365Client struct {
|
||||||
|
creds M365Credentials
|
||||||
|
base string // graph base URL; default https://graph.microsoft.com
|
||||||
|
tokenEndpoint string // login endpoint; default https://login.microsoftonline.com
|
||||||
|
client *http.Client
|
||||||
|
mu chan struct{}
|
||||||
|
token *m365Token
|
||||||
|
}
|
||||||
|
|
||||||
|
func NewM365Client(creds M365Credentials) (*M365Client, error) {
|
||||||
|
if creds.Tenant == "" || creds.ClientID == "" || creds.ClientSecret == "" {
|
||||||
|
return nil, errors.New("m365 needs tenant, client id and client secret")
|
||||||
|
}
|
||||||
|
c := &M365Client{
|
||||||
|
creds: creds,
|
||||||
|
base: "https://graph.microsoft.com",
|
||||||
|
client: &http.Client{Timeout: 90 * time.Second},
|
||||||
|
mu: make(chan struct{}, 1),
|
||||||
|
}
|
||||||
|
c.mu <- struct{}{}
|
||||||
|
return c, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// accessToken returns a fresh bearer token, refreshing via the Azure AD token
|
||||||
|
// endpoint when the cached one is missing or about to expire (within 2 min).
|
||||||
|
func (c *M365Client) accessToken(ctx context.Context) (string, error) {
|
||||||
|
select {
|
||||||
|
case <-c.mu:
|
||||||
|
case <-ctx.Done():
|
||||||
|
return "", ctx.Err()
|
||||||
|
}
|
||||||
|
defer func() { c.mu <- struct{}{} }()
|
||||||
|
if c.token != nil && c.token.AccessToken != "" && time.Now().Before(c.token.Expiry.Add(-2*time.Minute)) {
|
||||||
|
return c.token.AccessToken, nil
|
||||||
|
}
|
||||||
|
return c.refreshLocked(ctx)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *M365Client) refreshLocked(ctx context.Context) (string, error) {
|
||||||
|
form := url.Values{}
|
||||||
|
form.Set("grant_type", "client_credentials")
|
||||||
|
form.Set("client_id", c.creds.ClientID)
|
||||||
|
form.Set("client_secret", c.creds.ClientSecret)
|
||||||
|
form.Set("scope", "https://graph.microsoft.com/.default")
|
||||||
|
|
||||||
|
endpoint := c.tokenEndpoint
|
||||||
|
if endpoint == "" {
|
||||||
|
endpoint = fmt.Sprintf("https://login.microsoftonline.com/%s/oauth2/v2.0/token", c.creds.Tenant)
|
||||||
|
}
|
||||||
|
req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, strings.NewReader(form.Encode()))
|
||||||
|
if err != nil {
|
||||||
|
return "", err
|
||||||
|
}
|
||||||
|
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
|
||||||
|
resp, err := c.client.Do(req)
|
||||||
|
if err != nil {
|
||||||
|
return "", fmt.Errorf("m365 token: %w", err)
|
||||||
|
}
|
||||||
|
defer resp.Body.Close()
|
||||||
|
body, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
|
||||||
|
if resp.StatusCode != http.StatusOK {
|
||||||
|
var e struct {
|
||||||
|
Error string `json:"error"`
|
||||||
|
Desc string `json:"error_description"`
|
||||||
|
}
|
||||||
|
_ = json.Unmarshal(body, &e)
|
||||||
|
return "", fmt.Errorf("m365 token status %d: %s (%s)", resp.StatusCode, e.Error, truncate(e.Desc, 200))
|
||||||
|
}
|
||||||
|
var out struct {
|
||||||
|
AccessToken string `json:"access_token"`
|
||||||
|
ExpiresIn int64 `json:"expires_in"`
|
||||||
|
}
|
||||||
|
if err := json.Unmarshal(body, &out); err != nil {
|
||||||
|
return "", fmt.Errorf("m365 token parse: %w", err)
|
||||||
|
}
|
||||||
|
c.token = &m365Token{
|
||||||
|
AccessToken: out.AccessToken,
|
||||||
|
Expiry: time.Now().Add(time.Duration(out.ExpiresIn) * time.Second),
|
||||||
|
}
|
||||||
|
return out.AccessToken, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// deltaPage is one response page of the Graph delta query.
|
||||||
|
type deltaPage struct {
|
||||||
|
Value []struct {
|
||||||
|
ID string `json:"id"`
|
||||||
|
RemovedReason string `json:"@odata.removedReason"`
|
||||||
|
} `json:"value"`
|
||||||
|
NextLink string `json:"@odata.nextLink"`
|
||||||
|
DeltaLink string `json:"@odata.deltaLink"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// ListDeltaIDs walks the inbox delta query and returns live message ids since
|
||||||
|
// the previous deltaLink (or the full inbox when deltaLink is empty). Returns
|
||||||
|
// the new deltaLink for the next run. GET-only; nothing is mutated server-side.
|
||||||
|
func (c *M365Client) ListDeltaIDs(ctx context.Context, mailbox, deltaLink string, limit int) ([]string, string, error) {
|
||||||
|
var (
|
||||||
|
ids []string
|
||||||
|
url string
|
||||||
|
link = deltaLink
|
||||||
|
)
|
||||||
|
if link == "" {
|
||||||
|
url = fmt.Sprintf("/v1.0/users/%s/mailFolders/inbox/messages/delta", pathEscape(mailbox))
|
||||||
|
} else {
|
||||||
|
url = link
|
||||||
|
}
|
||||||
|
for url != "" {
|
||||||
|
var page deltaPage
|
||||||
|
if err := c.getJSON(ctx, url, &page); err != nil {
|
||||||
|
return ids, link, err
|
||||||
|
}
|
||||||
|
for _, m := range page.Value {
|
||||||
|
if m.RemovedReason != "" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if m.ID == "" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
ids = append(ids, m.ID)
|
||||||
|
if limit > 0 && len(ids) >= limit {
|
||||||
|
if page.DeltaLink != "" {
|
||||||
|
link = page.DeltaLink
|
||||||
|
}
|
||||||
|
return ids, link, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if page.DeltaLink != "" {
|
||||||
|
link = page.DeltaLink
|
||||||
|
url = ""
|
||||||
|
break
|
||||||
|
}
|
||||||
|
url = page.NextLink
|
||||||
|
}
|
||||||
|
return ids, link, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// GetMessage fetches a single message by id and normalizes it to the Message
|
||||||
|
// contract. GET-only.
|
||||||
|
func (c *M365Client) GetMessage(ctx context.Context, mailbox, id string) (*Message, error) {
|
||||||
|
path := fmt.Sprintf("/v1.0/users/%s/messages/%s?$expand=attachments($select=id,name,contentType,size,isInline)",
|
||||||
|
pathEscape(mailbox), pathEscape(id))
|
||||||
|
var raw struct {
|
||||||
|
ID string `json:"id"`
|
||||||
|
Subject string `json:"subject"`
|
||||||
|
From m365Recipient `json:"from"`
|
||||||
|
ToRecipients []m365Recipient `json:"toRecipients"`
|
||||||
|
CCRecipients []m365Recipient `json:"ccRecipients"`
|
||||||
|
BCCRecipients []m365Recipient `json:"bccRecipients"`
|
||||||
|
ReceivedDateTime string `json:"receivedDateTime"`
|
||||||
|
Body m365Body `json:"body"`
|
||||||
|
BodyPreview string `json:"bodyPreview"`
|
||||||
|
InternetMessageID string `json:"internetMessageId"`
|
||||||
|
Attachments []m365Attachment `json:"attachments"`
|
||||||
|
}
|
||||||
|
if err := c.getJSON(ctx, path, &raw); err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
m := &Message{
|
||||||
|
Source: "m365",
|
||||||
|
ID: raw.ID,
|
||||||
|
Folder: "m365",
|
||||||
|
Subject: raw.Subject,
|
||||||
|
From: formatRecipient(raw.From),
|
||||||
|
To: formatRecipients(raw.ToRecipients),
|
||||||
|
CC: formatRecipients(raw.CCRecipients),
|
||||||
|
BCC: formatRecipients(raw.BCCRecipients),
|
||||||
|
MimeMessageID: raw.InternetMessageID,
|
||||||
|
}
|
||||||
|
if t, err := time.Parse(time.RFC3339, raw.ReceivedDateTime); err == nil {
|
||||||
|
m.ReceivedAt = t
|
||||||
|
}
|
||||||
|
switch strings.ToLower(raw.Body.ContentType) {
|
||||||
|
case "html":
|
||||||
|
m.HTMLBody = raw.Body.Content
|
||||||
|
if raw.BodyPreview != "" {
|
||||||
|
m.TextBody = raw.BodyPreview
|
||||||
|
}
|
||||||
|
default:
|
||||||
|
m.TextBody = raw.Body.Content
|
||||||
|
if raw.BodyPreview != "" {
|
||||||
|
m.HTMLBody = raw.BodyPreview
|
||||||
|
}
|
||||||
|
}
|
||||||
|
for _, a := range raw.Attachments {
|
||||||
|
if a.IsInline || a.ID == "" || a.Name == "" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
m.Attachments = append(m.Attachments, Attachment{
|
||||||
|
FileID: a.ID,
|
||||||
|
FileName: a.Name,
|
||||||
|
StoredName: a.Name,
|
||||||
|
Size: a.Size,
|
||||||
|
ContentType: a.ContentType,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
m.HasAttachments = len(m.Attachments) > 0
|
||||||
|
return m, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
type m365Recipient struct {
|
||||||
|
EmailAddress struct {
|
||||||
|
Name string `json:"name"`
|
||||||
|
Address string `json:"address"`
|
||||||
|
} `json:"emailAddress"`
|
||||||
|
}
|
||||||
|
|
||||||
|
type m365Body struct {
|
||||||
|
ContentType string `json:"contentType"`
|
||||||
|
Content string `json:"content"`
|
||||||
|
}
|
||||||
|
|
||||||
|
type m365Attachment struct {
|
||||||
|
ID string `json:"id"`
|
||||||
|
Name string `json:"name"`
|
||||||
|
ContentType string `json:"contentType"`
|
||||||
|
Size int64 `json:"size"`
|
||||||
|
IsInline bool `json:"isInline"`
|
||||||
|
ContentID string `json:"contentId"`
|
||||||
|
}
|
||||||
|
|
||||||
|
func formatRecipients(rs []m365Recipient) string {
|
||||||
|
var parts []string
|
||||||
|
for _, r := range rs {
|
||||||
|
if s := formatRecipient(r); s != "" {
|
||||||
|
parts = append(parts, s)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return strings.Join(parts, ", ")
|
||||||
|
}
|
||||||
|
|
||||||
|
func formatRecipient(r m365Recipient) string { a := r.EmailAddress.Address
|
||||||
|
n := r.EmailAddress.Name
|
||||||
|
switch {
|
||||||
|
case n == "" || n == a:
|
||||||
|
return a
|
||||||
|
case a == "":
|
||||||
|
return n
|
||||||
|
default:
|
||||||
|
return fmt.Sprintf("%s <%s>", n, a)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// DownloadAttachment fetches an attachment's raw bytes via the /$value stream.
|
||||||
|
func (c *M365Client) DownloadAttachment(ctx context.Context, mailbox, msgID, attID string) ([]byte, error) {
|
||||||
|
path := fmt.Sprintf("/v1.0/users/%s/messages/%s/attachments/%s/$value",
|
||||||
|
pathEscape(mailbox), pathEscape(msgID), pathEscape(attID))
|
||||||
|
tok, err := c.accessToken(ctx)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.base+path, nil)
|
||||||
|
if err != nil {
|
||||||
|
return nil, err
|
||||||
|
}
|
||||||
|
req.Header.Set("Authorization", "Bearer "+tok)
|
||||||
|
resp, err := c.client.Do(req)
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("m365 attachment: %w", err)
|
||||||
|
}
|
||||||
|
defer resp.Body.Close()
|
||||||
|
body, _ := io.ReadAll(io.LimitReader(resp.Body, 256<<20))
|
||||||
|
if resp.StatusCode != http.StatusOK {
|
||||||
|
return nil, fmt.Errorf("m365 attachment %s: status %d: %s", attID, resp.StatusCode, truncate(string(body), 300))
|
||||||
|
}
|
||||||
|
return body, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (c *M365Client) getJSON(ctx context.Context, path string, out any) error {
|
||||||
|
tok, err := c.accessToken(ctx)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
u := path
|
||||||
|
if !strings.HasPrefix(u, "http") {
|
||||||
|
u = c.base + u
|
||||||
|
}
|
||||||
|
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
req.Header.Set("Authorization", "Bearer "+tok)
|
||||||
|
resp, err := c.client.Do(req)
|
||||||
|
if err != nil {
|
||||||
|
return fmt.Errorf("m365 %s: %w", path, err)
|
||||||
|
}
|
||||||
|
defer resp.Body.Close()
|
||||||
|
body, _ := io.ReadAll(io.LimitReader(resp.Body, 16<<20))
|
||||||
|
if resp.StatusCode != http.StatusOK {
|
||||||
|
return fmt.Errorf("m365 %s: status %d: %s", path, resp.StatusCode, truncate(string(body), 300))
|
||||||
|
}
|
||||||
|
if out != nil {
|
||||||
|
return json.Unmarshal(body, out)
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// m365Source adapts a mailbox to the Source worker-pool contract. Each mailbox
|
||||||
|
// gets its own folder under var/mail/m365/<localpart>/ and a delta state file.
|
||||||
|
type m365Source struct {
|
||||||
|
c *M365Client
|
||||||
|
mailbox string
|
||||||
|
localpart string
|
||||||
|
stateDir string
|
||||||
|
pending string // delta link to persist on Commit()
|
||||||
|
hasPending bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *m365Source) Folder() string { return filepath.Join("m365", s.localpart) }
|
||||||
|
|
||||||
|
func (s *m365Source) ListIDs(ctx context.Context, limit int, cursor string) ([]string, string, error) {
|
||||||
|
link, _ := os.ReadFile(filepath.Join(s.stateDir, s.localpart+".deltalink"))
|
||||||
|
ids, newLink, err := s.c.ListDeltaIDs(ctx, s.mailbox, strings.TrimSpace(string(link)), limit)
|
||||||
|
if err != nil {
|
||||||
|
return nil, "", err
|
||||||
|
}
|
||||||
|
// Buffer the new delta link; persist it only in Commit() after the full
|
||||||
|
// batch downloaded, so a failed run stays retryable without gaps.
|
||||||
|
if newLink != "" {
|
||||||
|
s.pending = newLink
|
||||||
|
s.hasPending = true
|
||||||
|
}
|
||||||
|
return ids, "", nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// Commit persists the buffered delta link. Called by the sync runner only when
|
||||||
|
// every listed message downloaded successfully.
|
||||||
|
func (s *m365Source) Commit() error {
|
||||||
|
if !s.hasPending || s.pending == "" {
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
if err := os.MkdirAll(s.stateDir, 0o755); err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
return os.WriteFile(filepath.Join(s.stateDir, s.localpart+".deltalink"), []byte(s.pending), 0o644)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *m365Source) Get(ctx context.Context, id string) (*Message, error) {
|
||||||
|
return s.c.GetMessage(ctx, s.mailbox, id)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (s *m365Source) DownloadAttachment(ctx context.Context, msg *Message, att Attachment) ([]byte, error) {
|
||||||
|
return s.c.DownloadAttachment(ctx, s.mailbox, msg.ID, att.FileID)
|
||||||
|
}
|
||||||
|
|
||||||
|
func pathEscape(s string) string {
|
||||||
|
return url.PathEscape(s)
|
||||||
|
}
|
||||||
@@ -0,0 +1,188 @@
|
|||||||
|
package sync
|
||||||
|
|
||||||
|
import (
|
||||||
|
"context"
|
||||||
|
"net/http"
|
||||||
|
"net/http/httptest"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
|
"testing"
|
||||||
|
)
|
||||||
|
|
||||||
|
// newM365TestClient serves the Graph delta + message endpoints against a fake
|
||||||
|
// token endpoint, so unit tests never touch the network.
|
||||||
|
func newM365TestClient(t *testing.T, graph http.Handler) *M365Client {
|
||||||
|
t.Helper()
|
||||||
|
tok := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
w.Header().Set("Content-Type", "application/json")
|
||||||
|
w.Write([]byte(`{"access_token":"test-token","expires_in":3600}`))
|
||||||
|
}))
|
||||||
|
gr := httptest.NewServer(graph)
|
||||||
|
t.Cleanup(func() {
|
||||||
|
tok.Close()
|
||||||
|
gr.Close()
|
||||||
|
})
|
||||||
|
client, err := NewM365Client(M365Credentials{Tenant: "t.onmicrosoft.com", ClientID: "c", ClientSecret: "s"})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
client.base = gr.URL
|
||||||
|
client.tokenEndpoint = tok.URL
|
||||||
|
return client
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestM365AccessToken(t *testing.T) {
|
||||||
|
c := newM365TestClient(t, http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {}))
|
||||||
|
got, err := c.accessToken(context.Background())
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("accessToken: %v", err)
|
||||||
|
}
|
||||||
|
if got != "test-token" {
|
||||||
|
t.Errorf("token = %q", got)
|
||||||
|
}
|
||||||
|
// Second call must reuse the cached token (no token request).
|
||||||
|
again, err := c.accessToken(context.Background())
|
||||||
|
if err != nil || again != "test-token" {
|
||||||
|
t.Fatalf("cached token: %q, %v", again, err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestM365DeltaSkipsTombstones(t *testing.T) {
|
||||||
|
first := true
|
||||||
|
graph := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
w.Header().Set("Content-Type", "application/json")
|
||||||
|
if first {
|
||||||
|
first = false
|
||||||
|
w.Write([]byte(`{"value":[
|
||||||
|
{"id":"m1"},
|
||||||
|
{"id":"m2","@odata.removedReason":"deleted"},
|
||||||
|
{"id":"m3"}
|
||||||
|
],"@odata.deltaLink":"` + deltaNext + `"}`))
|
||||||
|
return
|
||||||
|
}
|
||||||
|
// Second call must use the stored deltaLink (points at this server).
|
||||||
|
if r.URL.Path != "/v1.0/delta-next" {
|
||||||
|
w.WriteHeader(500)
|
||||||
|
w.Write([]byte(`{"error":{"message":"unexpected path"}}`))
|
||||||
|
return
|
||||||
|
}
|
||||||
|
w.Write([]byte(`{"value":[{"id":"m4"}],"@odata.deltaLink":"` + deltaFinal + `"}`))
|
||||||
|
})
|
||||||
|
c := newM365TestClient(t, graph)
|
||||||
|
deltaNext = c.base + "/v1.0/delta-next"
|
||||||
|
deltaFinal = c.base + "/v1.0/delta-final"
|
||||||
|
|
||||||
|
ids, link, err := c.ListDeltaIDs(context.Background(), "a@x.de", "", 0)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("delta: %v", err)
|
||||||
|
}
|
||||||
|
if len(ids) != 2 || ids[0] != "m1" || ids[1] != "m3" {
|
||||||
|
t.Errorf("ids = %v", ids)
|
||||||
|
}
|
||||||
|
if link == "" {
|
||||||
|
t.Error("expected new deltaLink")
|
||||||
|
}
|
||||||
|
// Incremental: pass the deltaLink, get only the new id.
|
||||||
|
ids2, link2, err := c.ListDeltaIDs(context.Background(), "a@x.de", link, 0)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("delta incremental: %v", err)
|
||||||
|
}
|
||||||
|
if len(ids2) != 1 || ids2[0] != "m4" {
|
||||||
|
t.Errorf("ids2 = %v", ids2)
|
||||||
|
}
|
||||||
|
if link2 == "" {
|
||||||
|
t.Error("expected updated deltaLink")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestM365GetMessageNormalizes(t *testing.T) {
|
||||||
|
graph := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
w.Header().Set("Content-Type", "application/json")
|
||||||
|
if pathLast(r.URL.Path) == "messages" {
|
||||||
|
w.Write([]byte(`{"value":[{"id":"m1"}]}`))
|
||||||
|
return
|
||||||
|
}
|
||||||
|
w.Write([]byte(`{
|
||||||
|
"id":"m1",
|
||||||
|
"subject":"Hallo",
|
||||||
|
"from":{"emailAddress":{"name":"Max","address":"max@x.de"}},
|
||||||
|
"toRecipients":[{"emailAddress":{"address":"a@x.de"}}],
|
||||||
|
"receivedDateTime":"2026-08-14T08:15:00Z",
|
||||||
|
"body":{"contentType":"html","content":"<p>body</p>"},
|
||||||
|
"bodyPreview":"body",
|
||||||
|
"internetMessageId":"<mid@x.de>",
|
||||||
|
"attachments":[
|
||||||
|
{"id":"att1","name":"doc.pdf","contentType":"application/pdf","size":10},
|
||||||
|
{"id":"img1","name":"logo.png","contentType":"image/png","isInline":true}
|
||||||
|
]
|
||||||
|
}`))
|
||||||
|
})
|
||||||
|
c := newM365TestClient(t, graph)
|
||||||
|
m, err := c.GetMessage(context.Background(), "a@x.de", "m1")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("GetMessage: %v", err)
|
||||||
|
}
|
||||||
|
if m.ID != "m1" || m.Subject != "Hallo" || m.From != "Max <max@x.de>" || m.To != "a@x.de" {
|
||||||
|
t.Errorf("headers mismatch: %+v", m)
|
||||||
|
}
|
||||||
|
if m.HTMLBody != "<p>body</p>" {
|
||||||
|
t.Errorf("html = %q", m.HTMLBody)
|
||||||
|
}
|
||||||
|
if m.ReceivedAt.IsZero() {
|
||||||
|
t.Error("receivedAt zero")
|
||||||
|
}
|
||||||
|
if m.MimeMessageID != "<mid@x.de>" {
|
||||||
|
t.Errorf("mime id = %q", m.MimeMessageID)
|
||||||
|
}
|
||||||
|
if len(m.Attachments) != 1 || m.Attachments[0].FileName != "doc.pdf" || m.Attachments[0].FileID != "att1" {
|
||||||
|
t.Errorf("atts = %+v", m.Attachments)
|
||||||
|
}
|
||||||
|
if !m.HasAttachments {
|
||||||
|
t.Error("expected hasAttachments")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestM365SourceDeltaState(t *testing.T) {
|
||||||
|
graph := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||||
|
w.Header().Set("Content-Type", "application/json")
|
||||||
|
w.Write([]byte(`{"value":[{"id":"m1"}],"@odata.deltaLink":"` + deltaNext + `"}`))
|
||||||
|
})
|
||||||
|
c := newM365TestClient(t, graph)
|
||||||
|
deltaNext = c.base + "/v1.0/delta-next"
|
||||||
|
stateDir := filepath.Join(t.TempDir(), ".m365")
|
||||||
|
s := &m365Source{c: c, mailbox: "info@x.de", localpart: "info", stateDir: stateDir}
|
||||||
|
|
||||||
|
if s.Folder() != "m365/info" {
|
||||||
|
t.Errorf("folder = %q", s.Folder())
|
||||||
|
}
|
||||||
|
ids, _, err := s.ListIDs(context.Background(), 0, "")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("ListIDs: %v", err)
|
||||||
|
}
|
||||||
|
if len(ids) != 1 || ids[0] != "m1" {
|
||||||
|
t.Errorf("ids = %v", ids)
|
||||||
|
}
|
||||||
|
if err := s.Commit(); err != nil {
|
||||||
|
t.Fatalf("Commit: %v", err)
|
||||||
|
}
|
||||||
|
data, err := os.ReadFile(filepath.Join(stateDir, "info.deltalink"))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("read delta state: %v", err)
|
||||||
|
}
|
||||||
|
if string(data) != deltaNext {
|
||||||
|
t.Errorf("delta state = %q, want %q", string(data), deltaNext)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func pathLast(p string) string {
|
||||||
|
for i := len(p) - 1; i >= 0; i-- {
|
||||||
|
if p[i] == '/' {
|
||||||
|
return p[i+1:]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
// deltaNext/deltaFinal are set per-test from the fake graph server URL so
|
||||||
|
// deltaLink values always point back at the fake (never the real Graph).
|
||||||
|
var deltaNext, deltaFinal string
|
||||||
+43
-1
@@ -106,6 +106,7 @@ func Retry(ctx context.Context, policy RetryPolicy, fn func() error) error {
|
|||||||
type SyncConfig struct {
|
type SyncConfig struct {
|
||||||
OO *OOConfig // OnlyOffice source (optional)
|
OO *OOConfig // OnlyOffice source (optional)
|
||||||
Gmail *GmailCredentials // Gmail source (optional)
|
Gmail *GmailCredentials // Gmail source (optional)
|
||||||
|
M365 *M365Credentials // Microsoft 365 Graph source (optional)
|
||||||
Out string // var/mail root; default <repo>/var/mail
|
Out string // var/mail root; default <repo>/var/mail
|
||||||
Workers int // concurrency; default 4
|
Workers int // concurrency; default 4
|
||||||
Limit int // max messages per source (0 = all)
|
Limit int // max messages per source (0 = all)
|
||||||
@@ -133,6 +134,14 @@ type Source interface {
|
|||||||
Folder() string
|
Folder() string
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Committer is an optional Source capability: Commit is called after all listed
|
||||||
|
// ids have been downloaded successfully. Sources that only advance durable state
|
||||||
|
// on success (e.g. a Graph delta link) implement this so a killed or failed run
|
||||||
|
// stays retryable without gaps.
|
||||||
|
type Committer interface {
|
||||||
|
Commit() error
|
||||||
|
}
|
||||||
|
|
||||||
type ooSource struct {
|
type ooSource struct {
|
||||||
c *OOClient
|
c *OOClient
|
||||||
page int
|
page int
|
||||||
@@ -222,8 +231,22 @@ func Run(ctx context.Context, cfg SyncConfig) (*SyncStats, error) {
|
|||||||
}
|
}
|
||||||
sources = append(sources, &gmailSource{c: gm, query: cfg.Query})
|
sources = append(sources, &gmailSource{c: gm, query: cfg.Query})
|
||||||
}
|
}
|
||||||
|
if cfg.M365 != nil {
|
||||||
|
stateDir := filepath.Join(cfg.Out, ".m365")
|
||||||
|
for _, mb := range cfg.M365.Users {
|
||||||
|
if !strings.Contains(mb, "@") {
|
||||||
|
return nil, fmt.Errorf("m365 user %q is not an email address", mb)
|
||||||
|
}
|
||||||
|
c, err := NewM365Client(*cfg.M365)
|
||||||
|
if err != nil {
|
||||||
|
return nil, fmt.Errorf("m365 init for %s: %w", mb, err)
|
||||||
|
}
|
||||||
|
local := strings.SplitN(mb, "@", 2)[0]
|
||||||
|
sources = append(sources, &m365Source{c: c, mailbox: mb, localpart: strings.ToLower(local), stateDir: stateDir})
|
||||||
|
}
|
||||||
|
}
|
||||||
if len(sources) == 0 {
|
if len(sources) == 0 {
|
||||||
return nil, errors.New("sync: no source configured (need OO, Gmail, or both)")
|
return nil, errors.New("sync: no source configured (need OO, Gmail, M365, or a combination)")
|
||||||
}
|
}
|
||||||
|
|
||||||
stats := &SyncStats{}
|
stats := &SyncStats{}
|
||||||
@@ -296,6 +319,25 @@ func Run(ctx context.Context, cfg SyncConfig) (*SyncStats, error) {
|
|||||||
close(jobsCh)
|
close(jobsCh)
|
||||||
wg.Wait()
|
wg.Wait()
|
||||||
|
|
||||||
|
// Only advance durable source state (e.g. delta links) when everything
|
||||||
|
// downloaded. A killed or failed run must be retryable without gaps.
|
||||||
|
if len(failures) == 0 {
|
||||||
|
seen := map[Source]bool{}
|
||||||
|
for _, j := range jobs {
|
||||||
|
if seen[j.src] {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
seen[j.src] = true
|
||||||
|
if c, ok := j.src.(Committer); ok {
|
||||||
|
if err := c.Commit(); err != nil {
|
||||||
|
mu.Lock()
|
||||||
|
failures = append(failures, j.src.Folder()+"/commit: "+err.Error())
|
||||||
|
mu.Unlock()
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
if len(failures) > 0 {
|
if len(failures) > 0 {
|
||||||
fmt.Fprintf(os.Stderr, "sync: %d failures:\n %s\n", len(failures), strings.Join(failures, "\n "))
|
fmt.Fprintf(os.Stderr, "sync: %d failures:\n %s\n", len(failures), strings.Join(failures, "\n "))
|
||||||
}
|
}
|
||||||
|
|||||||
+7
-22
@@ -15,6 +15,7 @@ import (
|
|||||||
"os"
|
"os"
|
||||||
"strings"
|
"strings"
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
"github.com/eSlider/2dph/internal/mdleaves"
|
"github.com/eSlider/2dph/internal/mdleaves"
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -23,29 +24,13 @@ func main() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func run(args []string) int {
|
func run(args []string) int {
|
||||||
jsonOut := false
|
c, err := mdleaves.ParseArgs(args)
|
||||||
files := ""
|
if err != nil {
|
||||||
root := "."
|
return cliparse.Fail(err)
|
||||||
for i := 0; i < len(args); i++ {
|
|
||||||
a := args[i]
|
|
||||||
switch {
|
|
||||||
case a == "--json":
|
|
||||||
jsonOut = true
|
|
||||||
case a == "--files" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
files = args[i]
|
|
||||||
case strings.HasPrefix(a, "--files="):
|
|
||||||
files = strings.TrimPrefix(a, "--files=")
|
|
||||||
case a == "-h" || a == "--help":
|
|
||||||
fmt.Fprintln(os.Stderr, "bin/markdown/import.go [dir] [--files a.md,b.md] [--json]")
|
|
||||||
return 0
|
|
||||||
case strings.HasPrefix(a, "-"):
|
|
||||||
fmt.Fprintln(os.Stderr, "unknown arg:", a)
|
|
||||||
return 2
|
|
||||||
default:
|
|
||||||
root = a
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
jsonOut := c.JSONOut
|
||||||
|
files := c.Files
|
||||||
|
root := c.Root
|
||||||
|
|
||||||
var paths []string
|
var paths []string
|
||||||
if files != "" {
|
if files != "" {
|
||||||
|
|||||||
+5
-17
@@ -16,8 +16,8 @@ import (
|
|||||||
"fmt"
|
"fmt"
|
||||||
"io"
|
"io"
|
||||||
"os"
|
"os"
|
||||||
"strings"
|
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
"github.com/eSlider/2dph/internal/duckstats"
|
"github.com/eSlider/2dph/internal/duckstats"
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -26,23 +26,11 @@ func main() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func run(args []string) int {
|
func run(args []string) int {
|
||||||
jsonl := ""
|
c, err := cliparse.ParseQAStats(args)
|
||||||
for i := 0; i < len(args); i++ {
|
if err != nil {
|
||||||
a := args[i]
|
return cliparse.Fail(err)
|
||||||
switch {
|
|
||||||
case a == "--jsonl" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
jsonl = args[i]
|
|
||||||
case strings.HasPrefix(a, "--jsonl="):
|
|
||||||
jsonl = strings.TrimPrefix(a, "--jsonl=")
|
|
||||||
case a == "-h" || a == "--help":
|
|
||||||
fmt.Fprintln(os.Stderr, "bin/qa/stats.go [--jsonl FILE] # stdin = JSON [float,…]")
|
|
||||||
return 0
|
|
||||||
default:
|
|
||||||
fmt.Fprintln(os.Stderr, "unknown arg:", a)
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
jsonl := c.JSONL
|
||||||
if jsonl != "" {
|
if jsonl != "" {
|
||||||
n, err := duckstats.CountJSONL(jsonl)
|
n, err := duckstats.CountJSONL(jsonl)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
|
|||||||
+7
-36
@@ -15,8 +15,8 @@ import (
|
|||||||
"encoding/json"
|
"encoding/json"
|
||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"strings"
|
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
"github.com/eSlider/2dph/internal/duckstats"
|
"github.com/eSlider/2dph/internal/duckstats"
|
||||||
"github.com/eSlider/2dph/internal/reasoner"
|
"github.com/eSlider/2dph/internal/reasoner"
|
||||||
)
|
)
|
||||||
@@ -26,42 +26,13 @@ func main() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func run(args []string) int {
|
func run(args []string) int {
|
||||||
base := os.Getenv("REASONER_BASE_URL")
|
c, err := reasoner.ParseArgs(args)
|
||||||
if base == "" {
|
if err != nil {
|
||||||
base = "http://127.0.0.1:11435/v1"
|
return cliparse.Fail(err)
|
||||||
}
|
}
|
||||||
model := os.Getenv("REASONER_MODEL")
|
base, model, jsonOut, device := c.Base, c.Model, c.JSONOut, c.Device
|
||||||
if model == "" {
|
client := reasoner.Client{BaseURL: base, Model: model, Device: device}
|
||||||
model = reasoner.OllamaRAM
|
rep := reasoner.Run(client)
|
||||||
}
|
|
||||||
jsonOut := false
|
|
||||||
device := "cpu"
|
|
||||||
for i := 0; i < len(args); i++ {
|
|
||||||
a := args[i]
|
|
||||||
switch {
|
|
||||||
case a == "--json":
|
|
||||||
jsonOut = true
|
|
||||||
case a == "--model" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
model = args[i]
|
|
||||||
case strings.HasPrefix(a, "--model="):
|
|
||||||
model = strings.TrimPrefix(a, "--model=")
|
|
||||||
case a == "--base-url" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
base = args[i]
|
|
||||||
case a == "--device" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
device = args[i]
|
|
||||||
case a == "-h" || a == "--help":
|
|
||||||
fmt.Fprintln(os.Stderr, "bin/reasoner/bakeoff.go [--model ID] [--base-url URL] [--device cpu] [--json]")
|
|
||||||
return 0
|
|
||||||
default:
|
|
||||||
fmt.Fprintln(os.Stderr, "unknown arg:", a)
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
}
|
|
||||||
c := reasoner.Client{BaseURL: base, Model: model, Device: device}
|
|
||||||
rep := reasoner.Run(c)
|
|
||||||
lat := make([]float64, 0, len(rep.Prompts))
|
lat := make([]float64, 0, len(rep.Prompts))
|
||||||
for _, p := range rep.Prompts {
|
for _, p := range rep.Prompts {
|
||||||
lat = append(lat, float64(p.LatencyMS))
|
lat = append(lat, float64(p.LatencyMS))
|
||||||
|
|||||||
@@ -0,0 +1,248 @@
|
|||||||
|
# bin/stack/lib.sh — compose helpers for start / start-assistant / stop / status.
|
||||||
|
# Sourced, not executed. No secrets. No host-absolute paths.
|
||||||
|
|
||||||
|
BRAIN_URL="${BRAIN_URL:-http://127.0.0.1:8630}"
|
||||||
|
REASONER_URL="${REASONER_URL:-http://127.0.0.1:11435}"
|
||||||
|
PICOCLAW_URL="${PICOCLAW_URL:-http://127.0.0.1:18790}"
|
||||||
|
REASONER_MODEL="${REASONER_MODEL:-qwen3.5:9b}"
|
||||||
|
STACK_WAIT_SECS="${STACK_WAIT_SECS:-90}"
|
||||||
|
STACK_WAIT_INTERVAL="${STACK_WAIT_INTERVAL:-2}"
|
||||||
|
STACK_PULL_SECS="${STACK_PULL_SECS:-600}"
|
||||||
|
|
||||||
|
if [[ -z "${ROOT:-}" ]]; then
|
||||||
|
STACK_DIR="$(CDPATH= cd -- "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||||
|
ROOT="$(CDPATH= cd -- "$STACK_DIR/../.." && pwd)"
|
||||||
|
fi
|
||||||
|
|
||||||
|
stack_usage() {
|
||||||
|
awk 'NR == 1 { next } /^#/ { sub(/^# ?/, ""); print; next } { exit }' "$1"
|
||||||
|
}
|
||||||
|
|
||||||
|
stack_die() {
|
||||||
|
echo "bin/stack: $*" >&2
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
compose() {
|
||||||
|
docker compose -f "$ROOT/compose.yaml" --project-directory "$ROOT" "$@"
|
||||||
|
}
|
||||||
|
|
||||||
|
http_get() {
|
||||||
|
local url=$1
|
||||||
|
local timeout=${2:-5}
|
||||||
|
curl -sS --max-time "$timeout" "$url" 2>/dev/null || return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
health_ok() {
|
||||||
|
local url=$1
|
||||||
|
local timeout=${2:-5}
|
||||||
|
local body
|
||||||
|
body=$(http_get "$url" "$timeout") || return 1
|
||||||
|
printf '%s' "$body" | grep -q '"status":"ok"'
|
||||||
|
}
|
||||||
|
|
||||||
|
wait_health() {
|
||||||
|
local url=$1
|
||||||
|
local n=0
|
||||||
|
while ((n <= STACK_WAIT_SECS)); do
|
||||||
|
if health_ok "$url"; then
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
n=$((n + 1))
|
||||||
|
if ((n <= STACK_WAIT_SECS)); then
|
||||||
|
sleep "$STACK_WAIT_INTERVAL"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
wait_http() {
|
||||||
|
local url=$1
|
||||||
|
local n=0
|
||||||
|
while ((n <= STACK_WAIT_SECS)); do
|
||||||
|
if http_get "$url" 5 >/dev/null; then
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
n=$((n + 1))
|
||||||
|
if ((n <= STACK_WAIT_SECS)); then
|
||||||
|
sleep "$STACK_WAIT_INTERVAL"
|
||||||
|
fi
|
||||||
|
done
|
||||||
|
return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
mcp_body() {
|
||||||
|
curl -sS --max-time 10 \
|
||||||
|
-H 'Content-Type: application/json' \
|
||||||
|
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' \
|
||||||
|
"$BRAIN_URL/mcp" 2>/dev/null || return 1
|
||||||
|
}
|
||||||
|
|
||||||
|
mcp_ok() {
|
||||||
|
local body
|
||||||
|
body=$(mcp_body) || return 1
|
||||||
|
printf '%s' "$body" | grep -Eq '"name": ?"search"' || return 1
|
||||||
|
printf '%s' "$body" | grep -Eq '"name": ?"get"' || return 1
|
||||||
|
printf '%s' "$body" | grep -Eq '"name": ?"audit"' || return 1
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
|
||||||
|
reasoner_tags() {
|
||||||
|
http_get "$REASONER_URL/api/tags" 5
|
||||||
|
}
|
||||||
|
|
||||||
|
reasoner_has_model() {
|
||||||
|
local body
|
||||||
|
body=$(reasoner_tags) || return 1
|
||||||
|
printf '%s' "$body" | grep -Fq "$REASONER_MODEL"
|
||||||
|
}
|
||||||
|
|
||||||
|
ensure_mcp() {
|
||||||
|
mcp_ok || stack_die "MCP tools/list missing search/get/audit at $BRAIN_URL/mcp"
|
||||||
|
}
|
||||||
|
|
||||||
|
ensure_brain() {
|
||||||
|
if health_ok "$BRAIN_URL/health"; then
|
||||||
|
echo "brain: reuse $BRAIN_URL" >&2
|
||||||
|
else
|
||||||
|
echo "brain: compose up" >&2
|
||||||
|
compose up -d brain
|
||||||
|
wait_health "$BRAIN_URL/health" || stack_die "brain health failed at $BRAIN_URL/health"
|
||||||
|
fi
|
||||||
|
ensure_mcp
|
||||||
|
}
|
||||||
|
|
||||||
|
pull_reasoner_model() {
|
||||||
|
echo "reasoner: pulling $REASONER_MODEL (CPU, may take minutes)" >&2
|
||||||
|
curl -sS --max-time "$STACK_PULL_SECS" \
|
||||||
|
-H 'Content-Type: application/json' \
|
||||||
|
-d "{\"name\":\"$REASONER_MODEL\"}" \
|
||||||
|
"$REASONER_URL/api/pull" >/dev/null
|
||||||
|
}
|
||||||
|
|
||||||
|
ensure_reasoner() {
|
||||||
|
if reasoner_has_model; then
|
||||||
|
echo "reasoner: reuse $REASONER_URL model $REASONER_MODEL" >&2
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
if ! reasoner_tags >/dev/null; then
|
||||||
|
echo "reasoner: compose up" >&2
|
||||||
|
compose --profile reasoner up -d reasoner
|
||||||
|
wait_http "$REASONER_URL/api/tags" || stack_die "reasoner not listening at $REASONER_URL"
|
||||||
|
fi
|
||||||
|
if reasoner_has_model; then
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
pull_reasoner_model
|
||||||
|
reasoner_has_model || stack_die "reasoner missing model $REASONER_MODEL"
|
||||||
|
}
|
||||||
|
|
||||||
|
ensure_picoclaw() {
|
||||||
|
echo "picoclaw: compose up --no-deps (reuse healthy :8630/:11435)" >&2
|
||||||
|
compose --profile picoclaw up -d --no-deps picoclaw
|
||||||
|
wait_health "$PICOCLAW_URL/health" || stack_die "picoclaw health failed at $PICOCLAW_URL/health"
|
||||||
|
}
|
||||||
|
|
||||||
|
mail_sync_running() {
|
||||||
|
compose ps --status running --services 2>/dev/null | grep -qx mail-sync
|
||||||
|
}
|
||||||
|
|
||||||
|
stack_status() {
|
||||||
|
local bh=down mcp=down ph=down present=false ms=down
|
||||||
|
health_ok "$BRAIN_URL/health" && bh=ok
|
||||||
|
mcp_ok && mcp=ok
|
||||||
|
reasoner_has_model && present=true
|
||||||
|
health_ok "$PICOCLAW_URL/health" && ph=ok
|
||||||
|
mail_sync_running && ms=ok
|
||||||
|
cat <<EOF
|
||||||
|
brain:
|
||||||
|
url: $BRAIN_URL
|
||||||
|
health: $bh
|
||||||
|
mcp: $mcp
|
||||||
|
reasoner:
|
||||||
|
url: $REASONER_URL
|
||||||
|
model: $REASONER_MODEL
|
||||||
|
present: $present
|
||||||
|
picoclaw:
|
||||||
|
url: $PICOCLAW_URL
|
||||||
|
health: $ph
|
||||||
|
mail_sync:
|
||||||
|
service: mail-sync
|
||||||
|
running: $ms
|
||||||
|
EOF
|
||||||
|
}
|
||||||
|
|
||||||
|
stack_start() {
|
||||||
|
ensure_brain
|
||||||
|
}
|
||||||
|
|
||||||
|
stack_start_mail_sync() {
|
||||||
|
echo "mail-sync: compose up (ETL sync→import; index only if MAIL_SYNC_INDEX=1)" >&2
|
||||||
|
compose up -d mail-sync
|
||||||
|
}
|
||||||
|
|
||||||
|
stack_attach_agent() {
|
||||||
|
local opts=()
|
||||||
|
if [[ -t 0 && -t 1 ]]; then
|
||||||
|
opts+=(-it)
|
||||||
|
else
|
||||||
|
opts+=(-T)
|
||||||
|
fi
|
||||||
|
if [[ (! -t 0 || ! -t 1) && $# -eq 0 ]]; then
|
||||||
|
echo "picoclaw: no TTY. Attach with:" >&2
|
||||||
|
echo " $ROOT/bin/stack/start-assistant" >&2
|
||||||
|
echo " docker compose --profile picoclaw exec -it picoclaw picoclaw agent" >&2
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
echo "picoclaw: agent (search → get → audit before a factual reply)" >&2
|
||||||
|
exec docker compose -f "$ROOT/compose.yaml" --project-directory "$ROOT" \
|
||||||
|
--profile picoclaw exec "${opts[@]}" picoclaw picoclaw agent "$@"
|
||||||
|
}
|
||||||
|
|
||||||
|
stack_start_assistant() {
|
||||||
|
local attach=1
|
||||||
|
local agent_args=()
|
||||||
|
while (($#)); do
|
||||||
|
case "$1" in
|
||||||
|
-h | --help)
|
||||||
|
stack_usage "$ROOT/bin/stack/start-assistant"
|
||||||
|
return 0
|
||||||
|
;;
|
||||||
|
--no-attach)
|
||||||
|
attach=0
|
||||||
|
shift
|
||||||
|
;;
|
||||||
|
--)
|
||||||
|
shift
|
||||||
|
agent_args+=("$@")
|
||||||
|
break
|
||||||
|
;;
|
||||||
|
*)
|
||||||
|
agent_args+=("$1")
|
||||||
|
shift
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
stack_start
|
||||||
|
ensure_reasoner
|
||||||
|
ensure_picoclaw
|
||||||
|
stack_status
|
||||||
|
if ((attach == 0)); then
|
||||||
|
echo "picoclaw: gateway $PICOCLAW_URL (agent not attached)" >&2
|
||||||
|
echo "ask the brain: $ROOT/bin/stack/start-assistant" >&2
|
||||||
|
echo "one-shot: $ROOT/bin/stack/start-assistant -- -m \"search the 2dph brain for LadybugDB\"" >&2
|
||||||
|
return 0
|
||||||
|
fi
|
||||||
|
stack_attach_agent "${agent_args[@]}"
|
||||||
|
}
|
||||||
|
|
||||||
|
stack_stop() {
|
||||||
|
case "${1:-}" in
|
||||||
|
-h | --help)
|
||||||
|
stack_usage "$ROOT/bin/stack/stop"
|
||||||
|
return 0
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
echo "stack: stop brain brain-mcp reasoner picoclaw mail-sync (volumes kept)" >&2
|
||||||
|
compose --profile picoclaw --profile reasoner stop picoclaw brain-mcp reasoner brain mail-sync
|
||||||
|
}
|
||||||
Executable
+21
@@ -0,0 +1,21 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# bin/stack/start - bring up brain HTTP/MCP and wait until search/get/audit respond.
|
||||||
|
#
|
||||||
|
# bin/stack/start
|
||||||
|
# bin/stack/status
|
||||||
|
#
|
||||||
|
# Reuses a healthy process on :8630 (host serve or compose). Does not start
|
||||||
|
# PicoClaw. Does not rebuild Ladybug.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
|
||||||
|
# shellcheck source=lib.sh
|
||||||
|
source "$STACK_DIR/lib.sh"
|
||||||
|
case "${1:-}" in
|
||||||
|
-h | --help)
|
||||||
|
stack_usage "$0"
|
||||||
|
exit 0
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
stack_start "$@"
|
||||||
|
stack_status
|
||||||
Executable
+15
@@ -0,0 +1,15 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# bin/stack/start-assistant - start + CPU reasoner + PicoClaw, then attach agent.
|
||||||
|
#
|
||||||
|
# bin/stack/start-assistant
|
||||||
|
# bin/stack/start-assistant --no-attach
|
||||||
|
# bin/stack/start-assistant -- -m "search the 2dph brain for LadybugDB"
|
||||||
|
#
|
||||||
|
# Pulls qwen3.5:9b if missing. Gateway :18790. Agent uses MCP search → get → audit.
|
||||||
|
# --no-attach leaves the gateway up without exec.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
|
||||||
|
# shellcheck source=lib.sh
|
||||||
|
source "$STACK_DIR/lib.sh"
|
||||||
|
stack_start_assistant "$@"
|
||||||
Executable
+21
@@ -0,0 +1,21 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# bin/stack/start-mail-sync - compose up mail-sync ETL (sync → import; optional index).
|
||||||
|
#
|
||||||
|
# bin/stack/start-mail-sync
|
||||||
|
#
|
||||||
|
# Default: onlyoffice,gmail every 300s into kb-var. Full --rebuild only if
|
||||||
|
# MAIL_SYNC_INDEX=1 in compose/env. Secrets: ~/.config/brain/mail.env +
|
||||||
|
# ~/.gmail-mcp (mounted). Does not start brain/picoclaw.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
|
||||||
|
# shellcheck source=lib.sh
|
||||||
|
source "$STACK_DIR/lib.sh"
|
||||||
|
case "${1:-}" in
|
||||||
|
-h | --help)
|
||||||
|
stack_usage "$0"
|
||||||
|
exit 0
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
stack_start_mail_sync "$@"
|
||||||
|
stack_status
|
||||||
Executable
+17
@@ -0,0 +1,17 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# bin/stack/status - YAML health for brain MCP, reasoner model, PicoClaw gateway.
|
||||||
|
#
|
||||||
|
# bin/stack/status
|
||||||
|
# bin/stack/status | yq '.picoclaw'
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
|
||||||
|
# shellcheck source=lib.sh
|
||||||
|
source "$STACK_DIR/lib.sh"
|
||||||
|
case "${1:-}" in
|
||||||
|
-h | --help)
|
||||||
|
stack_usage "$0"
|
||||||
|
exit 0
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
stack_status
|
||||||
Executable
+13
@@ -0,0 +1,13 @@
|
|||||||
|
#!/usr/bin/env bash
|
||||||
|
# bin/stack/stop - stop compose brain / brain-mcp / reasoner / picoclaw / mail-sync.
|
||||||
|
#
|
||||||
|
# bin/stack/stop
|
||||||
|
#
|
||||||
|
# Volumes kept (kb, reasoner weights, picoclaw-home). Does not kill a host
|
||||||
|
# bin/brain/serve.go that is not a compose service.
|
||||||
|
set -euo pipefail
|
||||||
|
|
||||||
|
STACK_DIR="$(CDPATH= cd -- "$(dirname "$0")" && pwd)"
|
||||||
|
# shellcheck source=lib.sh
|
||||||
|
source "$STACK_DIR/lib.sh"
|
||||||
|
stack_stop "$@"
|
||||||
@@ -0,0 +1,103 @@
|
|||||||
|
"""D16 contradiction adjudication (same rules as internal/facts)."""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
CONF_CONFIRMED = "confirmed"
|
||||||
|
CONF_HYPOTHESIS = "hypothesis"
|
||||||
|
|
||||||
|
RULE_UNRESOLVED = "unresolved"
|
||||||
|
RULE_TEMPORAL = "temporal_freshness"
|
||||||
|
RULE_AUTHORITY = "authority_pairing"
|
||||||
|
RULE_TWO_SOURCE = "two_source"
|
||||||
|
RULE_SINGLE = "single_source"
|
||||||
|
|
||||||
|
KIND_RUNTIME = "runtime"
|
||||||
|
KIND_CONFIG = "config"
|
||||||
|
KIND_NARRATIVE = "narrative"
|
||||||
|
|
||||||
|
|
||||||
|
def _independent(sources: list[dict]) -> int:
|
||||||
|
seen: set[str] = set()
|
||||||
|
for i, s in enumerate(sources):
|
||||||
|
sid = str(s.get("id") or "") or f"{s.get('kind', '')}#{i}"
|
||||||
|
seen.add(sid)
|
||||||
|
return len(seen)
|
||||||
|
|
||||||
|
|
||||||
|
def _fresh_n(sources: list[dict]) -> int:
|
||||||
|
return sum(1 for s in sources if not s.get("stale"))
|
||||||
|
|
||||||
|
|
||||||
|
def _strong_n(sources: list[dict]) -> int:
|
||||||
|
return sum(1 for s in sources if s.get("kind") in (KIND_RUNTIME, KIND_CONFIG))
|
||||||
|
|
||||||
|
|
||||||
|
def adjudicate(claim: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
yes = list(claim.get("yes") or [])
|
||||||
|
no = list(claim.get("no") or [])
|
||||||
|
yes_n, no_n = _independent(yes), _independent(no)
|
||||||
|
text = str(claim.get("text") or "")
|
||||||
|
|
||||||
|
def out(conf: str, rule: str, winner: str = "") -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"text": text,
|
||||||
|
"confidence": conf,
|
||||||
|
"confirmed": conf == CONF_CONFIRMED,
|
||||||
|
"rule": rule,
|
||||||
|
"winner": winner,
|
||||||
|
"yes": yes_n,
|
||||||
|
"no": no_n,
|
||||||
|
}
|
||||||
|
|
||||||
|
if yes_n < 2 or no_n < 2:
|
||||||
|
if yes_n >= 2:
|
||||||
|
return out(CONF_CONFIRMED, RULE_TWO_SOURCE, "yes")
|
||||||
|
if no_n >= 2:
|
||||||
|
return out(CONF_CONFIRMED, RULE_TWO_SOURCE, "no")
|
||||||
|
return out(CONF_HYPOTHESIS, RULE_SINGLE)
|
||||||
|
yf, nf = _fresh_n(yes), _fresh_n(no)
|
||||||
|
if yf >= 2 and nf < 2:
|
||||||
|
return out(CONF_CONFIRMED, RULE_TEMPORAL, "yes")
|
||||||
|
if nf >= 2 and yf < 2:
|
||||||
|
return out(CONF_CONFIRMED, RULE_TEMPORAL, "no")
|
||||||
|
ys, ns = _strong_n(yes), _strong_n(no)
|
||||||
|
if ys >= 2 and ns < 2:
|
||||||
|
return out(CONF_CONFIRMED, RULE_AUTHORITY, "yes")
|
||||||
|
if ns >= 2 and ys < 2:
|
||||||
|
return out(CONF_CONFIRMED, RULE_AUTHORITY, "no")
|
||||||
|
return out(CONF_HYPOTHESIS, RULE_UNRESOLVED)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_source_field(source: str) -> tuple[str, str]:
|
||||||
|
"""Split `a x b vs c x d` into (yes, no). Empty no if no ` vs `."""
|
||||||
|
if " vs " not in source:
|
||||||
|
return source, ""
|
||||||
|
yes, _, no = source.partition(" vs ")
|
||||||
|
return yes.strip(), no.strip()
|
||||||
|
|
||||||
|
|
||||||
|
def check_fact_row(lid: str, source: str, loc: str, how: str, conf: str) -> list[str]:
|
||||||
|
"""Lexicon checks for one facts leaf (no Ladybug)."""
|
||||||
|
problems: list[str] = []
|
||||||
|
src = source or ""
|
||||||
|
if conf == CONF_CONFIRMED:
|
||||||
|
if " vs " in src:
|
||||||
|
problems.append(f"{lid}: confirmed fact cannot keep a vs-contradiction")
|
||||||
|
if " x " not in src:
|
||||||
|
problems.append(f"{lid}: needs 2-source evidence in source, got '{source}'")
|
||||||
|
elif conf == CONF_HYPOTHESIS:
|
||||||
|
yes, no = parse_source_field(src)
|
||||||
|
if not no or " x " not in yes or " x " not in no:
|
||||||
|
problems.append(
|
||||||
|
f"{lid}: hypothesis contradiction needs 'a x b vs c x d', got '{source}'"
|
||||||
|
)
|
||||||
|
elif conf == "partial":
|
||||||
|
pass
|
||||||
|
else:
|
||||||
|
problems.append(f"{lid}: unknown confidence '{conf}'")
|
||||||
|
if not loc:
|
||||||
|
problems.append(f"{lid}: missing loc (evidence pointer)")
|
||||||
|
if not how:
|
||||||
|
problems.append(f"{lid}: missing how")
|
||||||
|
return problems
|
||||||
+76
-14
@@ -62,7 +62,9 @@ def init_schema(conn: ladybug.Connection) -> None:
|
|||||||
"CREATE NODE TABLE IF NOT EXISTS Leaf ("
|
"CREATE NODE TABLE IF NOT EXISTS Leaf ("
|
||||||
" id STRING, text STRING, root STRING, confidence STRING, "
|
" id STRING, text STRING, root STRING, confidence STRING, "
|
||||||
" sha256 STRING, source STRING, source_rev STRING, observed_at STRING, "
|
" sha256 STRING, source STRING, source_rev STRING, observed_at STRING, "
|
||||||
" how STRING, loc STRING, type STRING, embedding FLOAT[256], "
|
" how STRING, loc STRING, type STRING, "
|
||||||
|
" valid_from STRING, valid_to STRING, "
|
||||||
|
" embedding FLOAT[256], "
|
||||||
" PRIMARY KEY(id))"
|
" PRIMARY KEY(id))"
|
||||||
)
|
)
|
||||||
conn.execute(
|
conn.execute(
|
||||||
@@ -91,6 +93,46 @@ def init_schema(conn: ladybug.Connection) -> None:
|
|||||||
conn.execute(
|
conn.execute(
|
||||||
"CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
|
"CREATE REL TABLE IF NOT EXISTS AUTHORED (FROM Commit TO Person)"
|
||||||
)
|
)
|
||||||
|
ensure_interval_columns(conn)
|
||||||
|
|
||||||
|
|
||||||
|
def ensure_interval_columns(conn: ladybug.Connection) -> None:
|
||||||
|
"""D24: add valid_from/valid_to on older Leaf tables (idempotent ALTER)."""
|
||||||
|
for col in ("valid_from", "valid_to"):
|
||||||
|
try:
|
||||||
|
conn.execute(f"ALTER TABLE Leaf ADD {col} STRING")
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_day(s: str) -> str:
|
||||||
|
s = (s or "").strip()
|
||||||
|
if len(s) >= 10 and s[4] == "-" and s[7] == "-":
|
||||||
|
return s[:10]
|
||||||
|
return s
|
||||||
|
|
||||||
|
|
||||||
|
def active_at(valid_from: str, valid_to: str, as_of: str) -> bool:
|
||||||
|
"""D24: fact interval of truth. Empty ends = always; empty as_of = no filter."""
|
||||||
|
as_of = normalize_day(as_of)
|
||||||
|
if not as_of:
|
||||||
|
return True
|
||||||
|
fro = normalize_day(valid_from)
|
||||||
|
to = normalize_day(valid_to)
|
||||||
|
if fro and as_of < fro:
|
||||||
|
return False
|
||||||
|
if to and as_of > to:
|
||||||
|
return False
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def filter_as_of(hits: list[dict], as_of: str) -> list[dict]:
|
||||||
|
if not as_of:
|
||||||
|
return hits
|
||||||
|
return [
|
||||||
|
h for h in hits
|
||||||
|
if active_at(str(h.get("valid_from") or ""), str(h.get("valid_to") or ""), as_of)
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
def leaf_id(text: str, source: str) -> str:
|
def leaf_id(text: str, source: str) -> str:
|
||||||
@@ -99,19 +141,24 @@ def leaf_id(text: str, source: str) -> str:
|
|||||||
|
|
||||||
def upsert_leaf(conn: ladybug.Connection, *, text: str, root: str, confidence: str,
|
def upsert_leaf(conn: ladybug.Connection, *, text: str, root: str, confidence: str,
|
||||||
source: str, source_rev: str, how: str, loc: str, type_: str,
|
source: str, source_rev: str, how: str, loc: str, type_: str,
|
||||||
embedding: list[float] | None) -> str:
|
embedding: list[float] | None,
|
||||||
|
valid_from: str = "", valid_to: str = "") -> str:
|
||||||
lid = leaf_id(text, source)
|
lid = leaf_id(text, source)
|
||||||
obs = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
obs = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
|
||||||
|
vf = normalize_day(valid_from)
|
||||||
|
vt = normalize_day(valid_to)
|
||||||
conn.execute(
|
conn.execute(
|
||||||
"MERGE (l:Leaf {id:$id}) "
|
"MERGE (l:Leaf {id:$id}) "
|
||||||
"SET l.text=$text, l.root=$root, l.confidence=$confidence, "
|
"SET l.text=$text, l.root=$root, l.confidence=$confidence, "
|
||||||
" l.sha256=$sha, l.source=$source, l.source_rev=$rev, l.observed_at=$obs, "
|
" l.sha256=$sha, l.source=$source, l.source_rev=$rev, l.observed_at=$obs, "
|
||||||
" l.how=$how, l.loc=$location, l.type=$type"
|
" l.how=$how, l.loc=$location, l.type=$type, "
|
||||||
|
" l.valid_from=$vf, l.valid_to=$vt"
|
||||||
+ (", l.embedding=$emb" if embedding else ""),
|
+ (", l.embedding=$emb" if embedding else ""),
|
||||||
parameters={
|
parameters={
|
||||||
"id": lid, "text": text, "root": root, "confidence": confidence,
|
"id": lid, "text": text, "root": root, "confidence": confidence,
|
||||||
"sha": sha256_b64(text), "source": source, "rev": source_rev,
|
"sha": sha256_b64(text), "source": source, "rev": source_rev,
|
||||||
"obs": obs, "how": how, "location": loc, "type": type_,
|
"obs": obs, "how": how, "location": loc, "type": type_,
|
||||||
|
"vf": vf, "vt": vt,
|
||||||
"emb": (embedding if embedding else None),
|
"emb": (embedding if embedding else None),
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
@@ -121,9 +168,10 @@ def upsert_leaf(conn: ladybug.Connection, *, text: str, root: str, confidence: s
|
|||||||
def add_leafs(conn: ladybug.Connection, leafs: list[dict]) -> list[str]:
|
def add_leafs(conn: ladybug.Connection, leafs: list[dict]) -> list[str]:
|
||||||
"""Write facts+info leafs in one transaction. Safe while FTS/HNSW exist.
|
"""Write facts+info leafs in one transaction. Safe while FTS/HNSW exist.
|
||||||
|
|
||||||
Each leaf dict: text, source, optional root/confidence/source_rev/how/loc/type/embedding.
|
Each leaf dict: text, source, optional root/confidence/source_rev/how/loc/type/
|
||||||
Does not delete the database file. Measured on Ladybug 0.19: MERGE of new
|
embedding/valid_from/valid_to. Does not delete the database file. Measured on
|
||||||
ids (and updates) stays FTS+HNSW queryable; DROP INDEX is the fatal path.
|
Ladybug 0.19: MERGE of new ids (and updates) stays FTS+HNSW queryable; DROP
|
||||||
|
INDEX is the fatal path.
|
||||||
"""
|
"""
|
||||||
if not leafs:
|
if not leafs:
|
||||||
return []
|
return []
|
||||||
@@ -148,6 +196,8 @@ def add_leafs(conn: ladybug.Connection, leafs: list[dict]) -> list[str]:
|
|||||||
loc=str(lf.get("loc") or lf.get("source") or ""),
|
loc=str(lf.get("loc") or lf.get("source") or ""),
|
||||||
type_=str(lf.get("type") or lf.get("type_") or "reference"),
|
type_=str(lf.get("type") or lf.get("type_") or "reference"),
|
||||||
embedding=lf.get("embedding"),
|
embedding=lf.get("embedding"),
|
||||||
|
valid_from=str(lf.get("valid_from") or ""),
|
||||||
|
valid_to=str(lf.get("valid_to") or ""),
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
if started:
|
if started:
|
||||||
@@ -285,29 +335,39 @@ def drop_indexes(conn: ladybug.Connection) -> None:
|
|||||||
def query_fts(conn: ladybug.Connection, text: str, limit: int = 10) -> list[dict]:
|
def query_fts(conn: ladybug.Connection, text: str, limit: int = 10) -> list[dict]:
|
||||||
r = conn.execute(
|
r = conn.execute(
|
||||||
"CALL QUERY_FTS_INDEX('Leaf', 'id', $q) "
|
"CALL QUERY_FTS_INDEX('Leaf', 'id', $q) "
|
||||||
"RETURN node.id, node.text, node.root, score ORDER BY score DESC LIMIT $n",
|
"RETURN node.id, node.text, node.root, score, node.valid_from, node.valid_to "
|
||||||
|
"ORDER BY score DESC LIMIT $n",
|
||||||
parameters={"q": text, "n": limit},
|
parameters={"q": text, "n": limit},
|
||||||
)
|
)
|
||||||
return [{"id": row[0], "text": row[1], "root": row[2], "score": row[3]} for row in r.get_all()]
|
return [
|
||||||
|
{
|
||||||
|
"id": row[0], "text": row[1], "root": row[2], "score": row[3],
|
||||||
|
"valid_from": row[4] or "", "valid_to": row[5] or "",
|
||||||
|
}
|
||||||
|
for row in r.get_all()
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
def query_vector(conn: ladybug.Connection, embedding: list[float], limit: int = 10) -> list[dict]:
|
def query_vector(conn: ladybug.Connection, embedding: list[float], limit: int = 10) -> list[dict]:
|
||||||
r = conn.execute(
|
r = conn.execute(
|
||||||
"CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) "
|
"CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) "
|
||||||
"RETURN node.id, node.text, node.root, distance ORDER BY distance LIMIT $n",
|
"RETURN node.id, node.text, node.root, distance, node.valid_from, node.valid_to "
|
||||||
|
"ORDER BY distance LIMIT $n",
|
||||||
parameters={"q": embedding, "n": limit},
|
parameters={"q": embedding, "n": limit},
|
||||||
)
|
)
|
||||||
out = []
|
out = []
|
||||||
for row in r.get_all():
|
for row in r.get_all():
|
||||||
# distance -> similarity reasonable for cosine
|
|
||||||
score = 1.0 - row[3] if row[3] is not None else 0.0
|
score = 1.0 - row[3] if row[3] is not None else 0.0
|
||||||
out.append({"id": row[0], "text": row[1], "root": row[2], "score": score})
|
out.append({
|
||||||
|
"id": row[0], "text": row[1], "root": row[2], "score": score,
|
||||||
|
"valid_from": row[4] or "", "valid_to": row[5] or "",
|
||||||
|
})
|
||||||
return out
|
return out
|
||||||
|
|
||||||
|
|
||||||
def hybrid_search(conn: ladybug.Connection, embedding: list[float], fts_hits: list[dict],
|
def hybrid_search(conn: ladybug.Connection, embedding: list[float], fts_hits: list[dict],
|
||||||
limit: int = 10) -> list[dict]:
|
limit: int = 10, as_of: str = "") -> list[dict]:
|
||||||
"""Merge FTS + vector by reciprocal rank fusion."""
|
"""Merge FTS + vector by reciprocal rank fusion; optional D24 as-of filter."""
|
||||||
fused: dict[str, dict] = {}
|
fused: dict[str, dict] = {}
|
||||||
for rank, hit in enumerate(fts_hits):
|
for rank, hit in enumerate(fts_hits):
|
||||||
fused.setdefault(hit["id"], {**hit, "rrf": 0.0})["rrf"] = 1.0 / (60 + rank + 1)
|
fused.setdefault(hit["id"], {**hit, "rrf": 0.0})["rrf"] = 1.0 / (60 + rank + 1)
|
||||||
@@ -315,8 +375,10 @@ def hybrid_search(conn: ladybug.Connection, embedding: list[float], fts_hits: li
|
|||||||
entry = fused.setdefault(hit["id"], {**hit, "rrf": 0.0})
|
entry = fused.setdefault(hit["id"], {**hit, "rrf": 0.0})
|
||||||
entry["rrf"] += 1.0 / (60 + rank + 1)
|
entry["rrf"] += 1.0 / (60 + rank + 1)
|
||||||
entry.setdefault("score", hit.get("score", 0.0))
|
entry.setdefault("score", hit.get("score", 0.0))
|
||||||
|
entry.setdefault("valid_from", hit.get("valid_from") or "")
|
||||||
|
entry.setdefault("valid_to", hit.get("valid_to") or "")
|
||||||
ranked = sorted(fused.values(), key=lambda h: h.get("rrf", 0.0), reverse=True)
|
ranked = sorted(fused.values(), key=lambda h: h.get("rrf", 0.0), reverse=True)
|
||||||
return ranked[:limit]
|
return filter_as_of(ranked, as_of)[:limit]
|
||||||
|
|
||||||
|
|
||||||
def stats(conn: ladybug.Connection) -> dict:
|
def stats(conn: ladybug.Connection) -> dict:
|
||||||
|
|||||||
@@ -127,6 +127,38 @@ class BinLayoutTest(unittest.TestCase):
|
|||||||
self.assertIn("cmdbin.ExecFile", text)
|
self.assertIn("cmdbin.ExecFile", text)
|
||||||
self.assertIn(f"bin/facts/{method.removesuffix('.go')}", text)
|
self.assertIn(f"bin/facts/{method.removesuffix('.go')}", text)
|
||||||
|
|
||||||
|
def test_d16_adjudication_is_cgo_free(self) -> None:
|
||||||
|
self.assertTrue((ROOT / "internal" / "facts" / "contradict.go").is_file())
|
||||||
|
go = (ROOT / "internal" / "facts" / "contradict.go").read_text()
|
||||||
|
py = (ROOT / "bin" / "tools" / "contradict.py").read_text()
|
||||||
|
audit = (ROOT / "bin" / "facts" / "audit").read_text()
|
||||||
|
for token in ("temporal_freshness", "authority_pairing", "unresolved"):
|
||||||
|
self.assertIn(token, go)
|
||||||
|
self.assertIn(token, py)
|
||||||
|
self.assertIn("contradict", audit)
|
||||||
|
self.assertIn(" vs ", py)
|
||||||
|
plan = (ROOT / "PLAN.md").read_text()
|
||||||
|
self.assertIn("temporal_freshness", plan)
|
||||||
|
self.assertIn("authority_pairing", plan)
|
||||||
|
shebang = (ROOT / "bin" / "facts" / "audit.go").read_text()
|
||||||
|
self.assertIn("contradict", shebang)
|
||||||
|
|
||||||
|
def test_d23_flaggy_cli(self) -> None:
|
||||||
|
self.assertTrue((ROOT / "internal" / "cli" / "cli.go").is_file())
|
||||||
|
self.assertIn("github.com/integrii/flaggy", (ROOT / "go.mod").read_text())
|
||||||
|
plan = (ROOT / "PLAN.md").read_text()
|
||||||
|
self.assertIn("D23", plan)
|
||||||
|
self.assertIn("flaggy", plan)
|
||||||
|
complete = (ROOT / "bin" / "cli" / "complete.go").read_text()
|
||||||
|
first = complete.splitlines()[0]
|
||||||
|
self.assertTrue(first.startswith("//usr/bin/env go run"), first)
|
||||||
|
self.assertIn("complete.go bash", complete)
|
||||||
|
self.assertIn("brain-search", complete)
|
||||||
|
chats_import = (ROOT / "internal" / "chats" / "import.go").read_text()
|
||||||
|
self.assertNotIn("flag.NewFlagSet", chats_import)
|
||||||
|
args = (ROOT / "internal" / "brain" / "rank" / "args.go").read_text()
|
||||||
|
self.assertIn("internal/cli", args)
|
||||||
|
|
||||||
def test_mail_import_is_shebang_not_brain_write(self) -> None:
|
def test_mail_import_is_shebang_not_brain_write(self) -> None:
|
||||||
self._assert_shebang("bin/mail/import.go")
|
self._assert_shebang("bin/mail/import.go")
|
||||||
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
|
index_mail = (ROOT / "bin" / "mail" / "index_mail").read_text()
|
||||||
|
|||||||
@@ -0,0 +1,104 @@
|
|||||||
|
import os
|
||||||
|
import sys
|
||||||
|
import unittest
|
||||||
|
|
||||||
|
sys.path.insert(0, os.path.dirname(__file__))
|
||||||
|
|
||||||
|
from contradict import ( # noqa: E402
|
||||||
|
RULE_AUTHORITY,
|
||||||
|
RULE_SINGLE,
|
||||||
|
RULE_TEMPORAL,
|
||||||
|
RULE_TWO_SOURCE,
|
||||||
|
RULE_UNRESOLVED,
|
||||||
|
adjudicate,
|
||||||
|
check_fact_row,
|
||||||
|
parse_source_field,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def src(i, kind, stale=False):
|
||||||
|
return {"id": i, "kind": kind, "stale": stale}
|
||||||
|
|
||||||
|
|
||||||
|
class TestContradict(unittest.TestCase):
|
||||||
|
def test_two_vs_two_stays_hypothesis(self):
|
||||||
|
r = adjudicate({
|
||||||
|
"text": "svc listens on 443",
|
||||||
|
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
|
||||||
|
"no": [src("docker-old", "runtime"), src("compose-old", "config")],
|
||||||
|
})
|
||||||
|
self.assertFalse(r["confirmed"])
|
||||||
|
self.assertEqual(r["rule"], RULE_UNRESOLVED)
|
||||||
|
self.assertEqual(r["winner"], "")
|
||||||
|
|
||||||
|
def test_temporal_freshness(self):
|
||||||
|
r = adjudicate({
|
||||||
|
"text": "svc listens on 443",
|
||||||
|
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
|
||||||
|
"no": [src("old-readme", "narrative", True), src("old-wiki", "narrative", True)],
|
||||||
|
})
|
||||||
|
self.assertTrue(r["confirmed"])
|
||||||
|
self.assertEqual(r["rule"], RULE_TEMPORAL)
|
||||||
|
self.assertEqual(r["winner"], "yes")
|
||||||
|
|
||||||
|
def test_authority_pairing(self):
|
||||||
|
r = adjudicate({
|
||||||
|
"text": "svc listens on 443",
|
||||||
|
"yes": [src("docker-ps", "runtime"), src("compose", "config")],
|
||||||
|
"no": [src("readme", "narrative"), src("wiki", "narrative")],
|
||||||
|
})
|
||||||
|
self.assertTrue(r["confirmed"])
|
||||||
|
self.assertEqual(r["rule"], RULE_AUTHORITY)
|
||||||
|
self.assertEqual(r["winner"], "yes")
|
||||||
|
|
||||||
|
def test_two_source_and_single(self):
|
||||||
|
two = adjudicate({
|
||||||
|
"text": "arc-1 runs Matrix",
|
||||||
|
"yes": [src("compose", "config"), src("docker-ps", "runtime")],
|
||||||
|
})
|
||||||
|
self.assertTrue(two["confirmed"])
|
||||||
|
self.assertEqual(two["rule"], RULE_TWO_SOURCE)
|
||||||
|
one = adjudicate({"text": "maybe", "yes": [src("readme", "narrative")]})
|
||||||
|
self.assertFalse(one["confirmed"])
|
||||||
|
self.assertEqual(one["rule"], RULE_SINGLE)
|
||||||
|
|
||||||
|
def test_parse_source_field(self):
|
||||||
|
yes, no = parse_source_field("docker ps x compose.yml vs old.md x wiki.md")
|
||||||
|
self.assertIn(" x ", yes)
|
||||||
|
self.assertIn(" x ", no)
|
||||||
|
|
||||||
|
def test_check_fact_row_allows_hypothesis_vs(self):
|
||||||
|
p = check_fact_row(
|
||||||
|
"L1", "a.md x b.md vs c.md x d.md", "var/", "audit", "hypothesis",
|
||||||
|
)
|
||||||
|
self.assertEqual(p, [])
|
||||||
|
p = check_fact_row("L2", "a.md x b.md", "var/", "audit", "confirmed")
|
||||||
|
self.assertEqual(p, [])
|
||||||
|
p = check_fact_row("L3", "a.md x b.md vs c.md x d.md", "var/", "audit", "confirmed")
|
||||||
|
self.assertTrue(any("vs-contradiction" in x for x in p))
|
||||||
|
p = check_fact_row("L4", "only-one.md", "var/", "audit", "hypothesis")
|
||||||
|
self.assertTrue(any("a x b vs" in x for x in p))
|
||||||
|
|
||||||
|
def test_audit_contradict_cli_unresolved(self):
|
||||||
|
import json
|
||||||
|
import subprocess
|
||||||
|
from pathlib import Path
|
||||||
|
root = Path(__file__).resolve().parents[2]
|
||||||
|
payload = json.dumps({
|
||||||
|
"text": "svc 443",
|
||||||
|
"yes": [src("a", "runtime"), src("b", "config")],
|
||||||
|
"no": [src("c", "runtime"), src("d", "config")],
|
||||||
|
})
|
||||||
|
proc = subprocess.run(
|
||||||
|
[sys.executable, str(root / "bin" / "facts" / "audit"), "contradict", "--json"],
|
||||||
|
input=payload, capture_output=True, text=True, check=False,
|
||||||
|
)
|
||||||
|
self.assertEqual(proc.returncode, 0, proc.stderr)
|
||||||
|
out = json.loads(proc.stdout)
|
||||||
|
self.assertTrue(out["ok"])
|
||||||
|
self.assertEqual(out["contradictions"][0]["rule"], RULE_UNRESOLVED)
|
||||||
|
self.assertFalse(out["contradictions"][0]["confirmed"])
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
@@ -194,6 +194,38 @@ class KblibTest(unittest.TestCase):
|
|||||||
self.assertEqual(person["name"], "Ada Lovelace")
|
self.assertEqual(person["name"], "Ada Lovelace")
|
||||||
self.assertEqual(person["depth"], 3)
|
self.assertEqual(person["depth"], 3)
|
||||||
|
|
||||||
|
def test_as_of_keeps_x_drops_y(self) -> None:
|
||||||
|
"""OQ5/#36: as of 2025-01-01 → works-at-X, not works-at-Y."""
|
||||||
|
kblib.upsert_leaf(
|
||||||
|
self.conn, text="Andrey works at X", root="facts",
|
||||||
|
confidence="confirmed", source="crm.md x contract.md",
|
||||||
|
source_rev="r1", how="test", loc="/tmp", type_="fact",
|
||||||
|
embedding=make_emb(0.5),
|
||||||
|
valid_from="2024-03-01", valid_to="2025-07-15",
|
||||||
|
)
|
||||||
|
kblib.upsert_leaf(
|
||||||
|
self.conn, text="Andrey works at Y", root="facts",
|
||||||
|
confidence="confirmed", source="offer.md x payroll.md",
|
||||||
|
source_rev="r1", how="test", loc="/tmp", type_="fact",
|
||||||
|
embedding=make_emb(0.6),
|
||||||
|
valid_from="2025-07-16", valid_to="",
|
||||||
|
)
|
||||||
|
kblib.ensure_indexes(self.conn)
|
||||||
|
hits = kblib.query_fts(self.conn, "Andrey works", 10)
|
||||||
|
kept = kblib.filter_as_of(hits, "2025-01-01")
|
||||||
|
texts = [h["text"] for h in kept]
|
||||||
|
self.assertTrue(any("works at X" in t for t in texts), texts)
|
||||||
|
self.assertFalse(any("works at Y" in t for t in texts), texts)
|
||||||
|
later = kblib.filter_as_of(hits, "2025-08-01")
|
||||||
|
later_texts = [h["text"] for h in later]
|
||||||
|
self.assertTrue(any("works at Y" in t for t in later_texts), later_texts)
|
||||||
|
self.assertFalse(any("works at X" in t for t in later_texts), later_texts)
|
||||||
|
|
||||||
|
def test_active_at_pure(self) -> None:
|
||||||
|
self.assertTrue(kblib.active_at("2024-03-01", "2025-07-15", "2025-01-01"))
|
||||||
|
self.assertFalse(kblib.active_at("2025-07-16", "", "2025-01-01"))
|
||||||
|
self.assertTrue(kblib.active_at("", "", "2025-01-01"))
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
unittest.main()
|
unittest.main()
|
||||||
|
|||||||
@@ -186,3 +186,21 @@ class PublishedDocsTest(unittest.TestCase):
|
|||||||
self.assertIn("epic #16", index)
|
self.assertIn("epic #16", index)
|
||||||
agents = (ROOT / "AGENTS.md").read_text()
|
agents = (ROOT / "AGENTS.md").read_text()
|
||||||
self.assertIn("roadmap.md", agents)
|
self.assertIn("roadmap.md", agents)
|
||||||
|
|
||||||
|
def test_oq5_intervals_are_not_d16_stale(self) -> None:
|
||||||
|
plan = (ROOT / "PLAN.md").read_text()
|
||||||
|
design = (ROOT / "docs" / "design.md").read_text()
|
||||||
|
road = (ROOT / "docs" / "roadmap.md").read_text()
|
||||||
|
contradict = (ROOT / "internal" / "facts" / "contradict.go").read_text()
|
||||||
|
interval = (ROOT / "internal" / "facts" / "interval.go").read_text()
|
||||||
|
self.assertIn("OQ5", plan)
|
||||||
|
self.assertIn("D24", plan)
|
||||||
|
self.assertIn("issues/36", plan)
|
||||||
|
self.assertIn("valid_from", plan)
|
||||||
|
self.assertIn("--as-of", design)
|
||||||
|
self.assertIn("issues/36", road)
|
||||||
|
self.assertIn("**in**", road[road.index("OQ5"):road.index("OQ5") + 80])
|
||||||
|
self.assertIn("temporal_freshness", contradict)
|
||||||
|
self.assertNotIn("valid_from", contradict)
|
||||||
|
self.assertIn("ActiveAt", interval)
|
||||||
|
self.assertIn("NormalizeDay", interval)
|
||||||
|
|||||||
@@ -0,0 +1,228 @@
|
|||||||
|
"""bin/stack/{start,start-assistant,stop,status} — offline contract + fake PATH."""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
import stat
|
||||||
|
import subprocess
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
|
METHODS = ("start", "start-assistant", "start-mail-sync", "stop", "status")
|
||||||
|
|
||||||
|
|
||||||
|
class StackLayoutTest(unittest.TestCase):
|
||||||
|
def test_methods_are_bash_with_usage_comment(self) -> None:
|
||||||
|
lib = ROOT / "bin" / "stack" / "lib.sh"
|
||||||
|
self.assertTrue(lib.is_file(), "missing bin/stack/lib.sh")
|
||||||
|
for name in METHODS:
|
||||||
|
p = ROOT / "bin" / "stack" / name
|
||||||
|
self.assertTrue(p.is_file(), f"missing bin/stack/{name}")
|
||||||
|
self.assertTrue(os.access(p, os.X_OK), f"bin/stack/{name} must be executable")
|
||||||
|
lines = p.read_text().splitlines()
|
||||||
|
self.assertEqual(lines[0], "#!/usr/bin/env bash", name)
|
||||||
|
self.assertTrue(lines[1].startswith("# bin/stack/"), name)
|
||||||
|
text = "\n".join(lines)
|
||||||
|
self.assertIn("lib.sh", text, name)
|
||||||
|
self.assertNotIn("/mnt/", text, name)
|
||||||
|
self.assertNotIn("/home/", text, name)
|
||||||
|
|
||||||
|
def test_lib_has_no_host_paths_or_secrets(self) -> None:
|
||||||
|
lib = (ROOT / "bin" / "stack" / "lib.sh").read_text()
|
||||||
|
self.assertIn("stack_start", lib)
|
||||||
|
self.assertIn("stack_start_assistant", lib)
|
||||||
|
self.assertIn("stack_start_mail_sync", lib)
|
||||||
|
self.assertIn("stack_stop", lib)
|
||||||
|
self.assertIn("stack_status", lib)
|
||||||
|
self.assertIn("mail-sync", lib)
|
||||||
|
self.assertIn("qwen3.5:9b", lib)
|
||||||
|
self.assertIn("picoclaw agent", lib)
|
||||||
|
self.assertIn("--no-deps", lib)
|
||||||
|
self.assertIn("tools/list", lib)
|
||||||
|
self.assertNotIn("/mnt/", lib)
|
||||||
|
self.assertNotIn("/home/", lib)
|
||||||
|
self.assertNotIn("password", lib.lower())
|
||||||
|
self.assertNotIn("GITEA_TOKEN", lib)
|
||||||
|
|
||||||
|
def test_start_does_not_launch_picoclaw(self) -> None:
|
||||||
|
start = (ROOT / "bin" / "stack" / "start").read_text()
|
||||||
|
self.assertIn("stack_start", start)
|
||||||
|
self.assertNotIn("stack_start_assistant", start)
|
||||||
|
self.assertNotIn("picoclaw agent", start)
|
||||||
|
|
||||||
|
def test_start_mail_sync_is_etl_only(self) -> None:
|
||||||
|
src = (ROOT / "bin" / "stack" / "start-mail-sync").read_text()
|
||||||
|
self.assertIn("stack_start_mail_sync", src)
|
||||||
|
self.assertNotIn("stack_start_assistant", src)
|
||||||
|
ep = (ROOT / "bin" / "docker-entrypoint").read_text()
|
||||||
|
self.assertIn('MAIL_SYNC_SRC:=onlyoffice,gmail', ep)
|
||||||
|
self.assertIn('MAIL_SYNC_INDEX:=0', ep)
|
||||||
|
self.assertIn('interval="${1:-300}"', ep)
|
||||||
|
self.assertIn('MAIL_SYNC_INDEX" = "1"', ep)
|
||||||
|
|
||||||
|
def test_start_assistant_attaches_agent(self) -> None:
|
||||||
|
src = (ROOT / "bin" / "stack" / "start-assistant").read_text()
|
||||||
|
self.assertIn("stack_start_assistant", src)
|
||||||
|
self.assertIn("--no-attach", src)
|
||||||
|
|
||||||
|
def test_stop_does_not_down_volumes(self) -> None:
|
||||||
|
lib = (ROOT / "bin" / "stack" / "lib.sh").read_text()
|
||||||
|
self.assertIn(" compose ", lib)
|
||||||
|
self.assertRegex(lib, r"\bstop\b")
|
||||||
|
self.assertNotIn(" compose down", lib)
|
||||||
|
self.assertNotIn("compose down", lib)
|
||||||
|
|
||||||
|
def test_docs_name_stack_commands(self) -> None:
|
||||||
|
runbook = (ROOT / "docs" / "runbook.md").read_text()
|
||||||
|
pico = (ROOT / "docs" / "picoclaw.md").read_text()
|
||||||
|
agents = (ROOT / "AGENTS.md").read_text()
|
||||||
|
for text in (runbook, pico, agents):
|
||||||
|
self.assertIn("bin/stack/start", text)
|
||||||
|
self.assertIn("bin/stack/start-assistant", text)
|
||||||
|
self.assertIn("bin/stack/status", text)
|
||||||
|
self.assertIn("bin/stack/stop", text)
|
||||||
|
|
||||||
|
|
||||||
|
def test_help_prints_comments_not_source(self) -> None:
|
||||||
|
r = subprocess.run(
|
||||||
|
[str(ROOT / "bin" / "stack" / "start"), "--help"],
|
||||||
|
cwd=str(ROOT),
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
check=False,
|
||||||
|
)
|
||||||
|
self.assertEqual(r.returncode, 0, r.stderr)
|
||||||
|
self.assertIn("bin/stack/start", r.stdout)
|
||||||
|
self.assertNotIn("set -euo pipefail", r.stdout)
|
||||||
|
self.assertNotIn("source ", r.stdout)
|
||||||
|
|
||||||
|
|
||||||
|
class StackFakePathTest(unittest.TestCase):
|
||||||
|
def _fake_bin(self, tmp: Path, *, health_ok: bool) -> Path:
|
||||||
|
bindir = tmp / "bin"
|
||||||
|
bindir.mkdir()
|
||||||
|
curl = bindir / "curl"
|
||||||
|
docker = bindir / "docker"
|
||||||
|
log = tmp / "docker.log"
|
||||||
|
curl.write_text(
|
||||||
|
f"""#!/usr/bin/env bash
|
||||||
|
url=""
|
||||||
|
for a in "$@"; do
|
||||||
|
case "$a" in http*) url=$a ;;
|
||||||
|
esac
|
||||||
|
done
|
||||||
|
if [ "{int(health_ok)}" = "0" ] && [[ "$url" == */health ]]; then
|
||||||
|
echo '{{"status":"down"}}'
|
||||||
|
exit 7
|
||||||
|
fi
|
||||||
|
case "$url" in
|
||||||
|
*/mcp)
|
||||||
|
echo '{{"jsonrpc":"2.0","id":1,"result":{{"tools":[{{"name":"search"}},{{"name":"get"}},{{"name":"audit"}}]}}}}'
|
||||||
|
;;
|
||||||
|
*/api/tags)
|
||||||
|
echo '{{"models":[{{"name":"qwen3.5:9b"}}]}}'
|
||||||
|
;;
|
||||||
|
*/api/pull)
|
||||||
|
echo '{{"status":"success"}}'
|
||||||
|
;;
|
||||||
|
*)
|
||||||
|
echo '{{"status":"ok"}}'
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
docker.write_text(
|
||||||
|
f"""#!/usr/bin/env bash
|
||||||
|
echo "$*" >> "{log}"
|
||||||
|
exit 0
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
curl.chmod(curl.stat().st_mode | stat.S_IEXEC)
|
||||||
|
docker.chmod(docker.stat().st_mode | stat.S_IEXEC)
|
||||||
|
return bindir
|
||||||
|
|
||||||
|
def _env(self, bindir: Path) -> dict[str, str]:
|
||||||
|
env = os.environ.copy()
|
||||||
|
env["PATH"] = f"{bindir}:{env.get('PATH', '')}"
|
||||||
|
env["STACK_WAIT_SECS"] = "1"
|
||||||
|
env["STACK_WAIT_INTERVAL"] = "0"
|
||||||
|
return env
|
||||||
|
|
||||||
|
def test_start_skips_compose_when_brain_healthy(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as raw:
|
||||||
|
tmp = Path(raw)
|
||||||
|
bindir = self._fake_bin(tmp, health_ok=True)
|
||||||
|
log = tmp / "docker.log"
|
||||||
|
r = subprocess.run(
|
||||||
|
[str(ROOT / "bin" / "stack" / "start")],
|
||||||
|
cwd=str(ROOT),
|
||||||
|
env=self._env(bindir),
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
check=False,
|
||||||
|
)
|
||||||
|
self.assertEqual(r.returncode, 0, r.stderr)
|
||||||
|
if log.exists():
|
||||||
|
logged = log.read_text()
|
||||||
|
self.assertNotIn("up -d", logged, "healthy brain must not docker compose up")
|
||||||
|
self.assertIn("ps", logged) # status probes mail-sync via compose ps
|
||||||
|
|
||||||
|
def test_start_ups_brain_when_unhealthy(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as raw:
|
||||||
|
tmp = Path(raw)
|
||||||
|
bindir = self._fake_bin(tmp, health_ok=False)
|
||||||
|
log = tmp / "docker.log"
|
||||||
|
r = subprocess.run(
|
||||||
|
[str(ROOT / "bin" / "stack" / "start")],
|
||||||
|
cwd=str(ROOT),
|
||||||
|
env=self._env(bindir),
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
check=False,
|
||||||
|
)
|
||||||
|
self.assertNotEqual(r.returncode, 0, "unhealthy brain without recovering compose must fail")
|
||||||
|
self.assertTrue(log.exists(), r.stderr)
|
||||||
|
logged = log.read_text()
|
||||||
|
self.assertIn("up -d", logged)
|
||||||
|
self.assertIn("brain", logged)
|
||||||
|
self.assertNotIn("picoclaw", logged)
|
||||||
|
|
||||||
|
def test_stop_stops_named_services(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as raw:
|
||||||
|
tmp = Path(raw)
|
||||||
|
bindir = self._fake_bin(tmp, health_ok=True)
|
||||||
|
log = tmp / "docker.log"
|
||||||
|
r = subprocess.run(
|
||||||
|
[str(ROOT / "bin" / "stack" / "stop")],
|
||||||
|
cwd=str(ROOT),
|
||||||
|
env=self._env(bindir),
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
check=False,
|
||||||
|
)
|
||||||
|
self.assertEqual(r.returncode, 0, r.stderr)
|
||||||
|
logged = log.read_text()
|
||||||
|
self.assertIn("stop", logged)
|
||||||
|
self.assertNotIn(" down", logged)
|
||||||
|
for svc in ("brain", "brain-mcp", "reasoner", "picoclaw", "mail-sync"):
|
||||||
|
self.assertIn(svc, logged)
|
||||||
|
|
||||||
|
def test_start_assistant_no_attach_starts_picoclaw(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as raw:
|
||||||
|
tmp = Path(raw)
|
||||||
|
bindir = self._fake_bin(tmp, health_ok=True)
|
||||||
|
log = tmp / "docker.log"
|
||||||
|
r = subprocess.run(
|
||||||
|
[str(ROOT / "bin" / "stack" / "start-assistant"), "--no-attach"],
|
||||||
|
cwd=str(ROOT),
|
||||||
|
env=self._env(bindir),
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
check=False,
|
||||||
|
)
|
||||||
|
self.assertEqual(r.returncode, 0, r.stderr + r.stdout)
|
||||||
|
logged = log.read_text() if log.exists() else ""
|
||||||
|
self.assertIn("picoclaw", logged)
|
||||||
|
self.assertIn("--no-deps", logged)
|
||||||
|
self.assertNotIn("picoclaw agent", logged)
|
||||||
+9
-71
@@ -16,9 +16,9 @@ import (
|
|||||||
"fmt"
|
"fmt"
|
||||||
"net/http"
|
"net/http"
|
||||||
"os"
|
"os"
|
||||||
"strconv"
|
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
"github.com/eSlider/2dph/internal/websearch"
|
"github.com/eSlider/2dph/internal/websearch"
|
||||||
"golang.org/x/sys/unix"
|
"golang.org/x/sys/unix"
|
||||||
)
|
)
|
||||||
@@ -28,77 +28,15 @@ func main() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func run(args []string) int {
|
func run(args []string) int {
|
||||||
var (
|
c, err := websearch.ParseArgs(args)
|
||||||
query, site, lang, fresh, category, engines string
|
if err != nil {
|
||||||
limit = websearch.DefaultLimit
|
return cli.Fail(err)
|
||||||
jsonOut, refresh, force bool
|
|
||||||
ttl = float64(websearch.CacheTTL)
|
|
||||||
timeout = 25
|
|
||||||
)
|
|
||||||
i := 0
|
|
||||||
for i < len(args) {
|
|
||||||
a := args[i]
|
|
||||||
switch {
|
|
||||||
case a == "--json":
|
|
||||||
jsonOut = true
|
|
||||||
case a == "--refresh":
|
|
||||||
refresh = true
|
|
||||||
case a == "--force":
|
|
||||||
force = true
|
|
||||||
case (a == "-n" || a == "--limit") && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
n, err := strconv.Atoi(args[i])
|
|
||||||
if err != nil || n < 0 {
|
|
||||||
fmt.Fprintln(os.Stderr, "web/search: --limit must be a non-negative integer")
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
limit = n
|
|
||||||
case a == "--site" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
site = args[i]
|
|
||||||
case a == "--lang" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
lang = args[i]
|
|
||||||
case a == "--fresh" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
fresh = args[i]
|
|
||||||
case a == "--category" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
category = args[i]
|
|
||||||
case a == "--engines" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
engines = args[i]
|
|
||||||
case a == "--ttl" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
v, err := strconv.ParseFloat(args[i], 64)
|
|
||||||
if err != nil {
|
|
||||||
fmt.Fprintln(os.Stderr, "web/search: --ttl must be a number")
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
ttl = v
|
|
||||||
case a == "--timeout" && i+1 < len(args):
|
|
||||||
i++
|
|
||||||
n, err := strconv.Atoi(args[i])
|
|
||||||
if err != nil || n <= 0 {
|
|
||||||
fmt.Fprintln(os.Stderr, "web/search: --timeout must be a positive integer")
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
timeout = n
|
|
||||||
case a == "-h" || a == "--help":
|
|
||||||
fmt.Fprintln(os.Stderr, `usage: bin/web/search.go QUERY [--json] [-n N] [--site HOST] [--lang LANG] [--fresh day|week|month|year] [--category CAT] [--engines LIST] [--refresh] [--force]`)
|
|
||||||
return 0
|
|
||||||
case len(a) > 0 && a[0] != '-' && query == "":
|
|
||||||
query = a
|
|
||||||
default:
|
|
||||||
fmt.Fprintf(os.Stderr, "web/search: unknown flag %s\n", a)
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
i++
|
|
||||||
}
|
|
||||||
if query == "" {
|
|
||||||
fmt.Fprintln(os.Stderr, "web/search: query required")
|
|
||||||
return 2
|
|
||||||
}
|
}
|
||||||
|
query, site, lang, fresh, category, engines := c.Query, c.Site, c.Lang, c.Fresh, c.Category, c.Engines
|
||||||
|
limit := c.Limit
|
||||||
|
jsonOut, refresh, force := c.JSONOut, c.Refresh, c.Force
|
||||||
|
ttl := c.TTL
|
||||||
|
timeout := c.Timeout
|
||||||
if site != "" {
|
if site != "" {
|
||||||
query = "site:" + site + " " + query
|
query = "site:" + site + " " + query
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,5 +1,6 @@
|
|||||||
# 2dph — docker composition
|
# 2dph — docker composition
|
||||||
#
|
#
|
||||||
|
# bin/stack/start / start-assistant / status / stop
|
||||||
# docker compose up -d brain # API (Zig CGO serve)
|
# docker compose up -d brain # API (Zig CGO serve)
|
||||||
# docker compose --profile index run --rm index # Python rebuild
|
# docker compose --profile index run --rm index # Python rebuild
|
||||||
# docker compose --profile picoclaw up -d # brain-mcp + CPU reasoner + PicoClaw gateway
|
# docker compose --profile picoclaw up -d # brain-mcp + CPU reasoner + PicoClaw gateway
|
||||||
@@ -91,6 +92,38 @@ services:
|
|||||||
tmpfs:
|
tmpfs:
|
||||||
- /tmp
|
- /tmp
|
||||||
|
|
||||||
|
# mail-sync ETL: sync → import every 300s into shared var. Full brain rebuild
|
||||||
|
# only when MAIL_SYNC_INDEX=1 (expensive on large mail corpora). Default
|
||||||
|
# sources: onlyoffice,gmail. For M365 set MAIL_SYNC_SRC=m365 and put
|
||||||
|
# M365_TENANT/CLIENT_ID/CLIENT_SECRET/USERS in ~/.config/brain/mail.env
|
||||||
|
# (or m365.env + MAIL_SYNC_ENV). Gmail OAuth: mount ~/.gmail-mcp.
|
||||||
|
# docker compose up -d mail-sync
|
||||||
|
# bin/stack/start-mail-sync
|
||||||
|
mail-sync:
|
||||||
|
image: ghcr.io/eslider/2dph:index
|
||||||
|
build:
|
||||||
|
context: .
|
||||||
|
dockerfile: Dockerfile
|
||||||
|
target: index
|
||||||
|
environment:
|
||||||
|
HF_HOME: /data/hf
|
||||||
|
KB_PY: python3
|
||||||
|
MAIL_SYNC_SRC: onlyoffice,gmail
|
||||||
|
MAIL_SYNC_ENV: /secret/mail.env
|
||||||
|
MAIL_SYNC_INDEX: "0"
|
||||||
|
HOME: /home/2dph
|
||||||
|
volumes:
|
||||||
|
- kb-model:/data/hf
|
||||||
|
- kb-var:/app/var
|
||||||
|
- ~/.config/brain:/secret:ro
|
||||||
|
- ~/.gmail-mcp:/home/2dph/.gmail-mcp:ro
|
||||||
|
command: ["mail-sync", "300"]
|
||||||
|
read_only: true
|
||||||
|
tmpfs:
|
||||||
|
- /tmp
|
||||||
|
restart: unless-stopped
|
||||||
|
stop_grace_period: 20s
|
||||||
|
|
||||||
# Optional local SearXNG (D3). Skip if BRAIN_SEARCH_URL already points at a
|
# Optional local SearXNG (D3). Skip if BRAIN_SEARCH_URL already points at a
|
||||||
# live instance — do not run a second copy on that host.
|
# live instance — do not run a second copy on that host.
|
||||||
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
|
# SEARXNG_SECRET=$(openssl rand -hex 32) docker compose --profile searxng up -d
|
||||||
|
|||||||
+2
-2
@@ -18,9 +18,9 @@ Evidence-first knowledge graph. Facts need proof or they are
|
|||||||
| tutorial / howto | [runbook](runbook.md) — run anywhere (uv, Go, Docker) |
|
| tutorial / howto | [runbook](runbook.md) — run anywhere (uv, Go, Docker) |
|
||||||
| explanation | [design](design.md) — two roots, deduction, D17/D20/D18 |
|
| explanation | [design](design.md) — two roots, deduction, D17/D20/D18 |
|
||||||
| explanation | [roadmap](roadmap.md) — gap to v1 (epic #16) |
|
| explanation | [roadmap](roadmap.md) — gap to v1 (epic #16) |
|
||||||
| howto | [picoclaw](picoclaw.md) — MCP agent profile |
|
| howto | [picoclaw](picoclaw.md) — MCP agent (`bin/stack/start-assistant`) |
|
||||||
| howto | [reasoner](reasoner.md) — CPU bake-off (D18) |
|
| howto | [reasoner](reasoner.md) — CPU bake-off (D18) |
|
||||||
| reference | [PLAN.md](../PLAN.md) — decisions D1–D22 |
|
| reference | [PLAN.md](../PLAN.md) — decisions D1–D24 |
|
||||||
|
|
||||||
Decisions the public face must name: **D3** SearXNG compose, **D6** Go service /
|
Decisions the public face must name: **D3** SearXNG compose, **D6** Go service /
|
||||||
Python write sidecar, **D14** `bin/{subject}/{method}.go`, **D15** Gitea origin,
|
Python write sidecar, **D14** `bin/{subject}/{method}.go`, **D15** Gitea origin,
|
||||||
|
|||||||
+16
-1
@@ -38,6 +38,10 @@ bin/brain/search.go "question"
|
|||||||
from each hit (1=File, 2=Commit, 3=Person). Rebuild writes FROM_FILE;
|
from each hit (1=File, 2=Commit, 3=Person). Rebuild writes FROM_FILE;
|
||||||
git import writes HAS_VERSION/AUTHORED ([#17](https://git.produktor.io/eSlider/2dph/issues/17)).
|
git import writes HAS_VERSION/AUTHORED ([#17](https://git.produktor.io/eSlider/2dph/issues/17)).
|
||||||
|
|
||||||
|
Go CLIs parse with **flaggy** via `internal/cli` (D23). Flags may appear
|
||||||
|
after positionals (`search q --hop 1`). Completions:
|
||||||
|
`source <(./bin/cli/complete.go bash)`.
|
||||||
|
|
||||||
## Who / What / How / Where / When + evidence
|
## Who / What / How / Where / When + evidence
|
||||||
|
|
||||||
Every assertion edge carries:
|
Every assertion edge carries:
|
||||||
@@ -64,6 +68,15 @@ binary); conversion prints leafs, brain write is `bin/brain/index.go`.
|
|||||||
`bin/facts/audit stale` flags leafs whose observed revision is behind the
|
`bin/facts/audit stale` flags leafs whose observed revision is behind the
|
||||||
corpus HEAD.
|
corpus HEAD.
|
||||||
|
|
||||||
|
Fact **interval of truth** (D24 / OQ5): leaf props `valid_from` /
|
||||||
|
`valid_to` (YYYY-MM-DD, inclusive; empty end = open; both empty = legacy
|
||||||
|
always-active). `bin/brain/search.go --as-of YYYY-MM-DD` and MCP/HTTP
|
||||||
|
`as_of` keep hits whose interval covers that day. This is not D16
|
||||||
|
`temporal_freshness` (source freshness vs HEAD). [#36](https://git.produktor.io/eSlider/2dph/issues/36).
|
||||||
|
Existing `kb.lbug` without the columns: open/search/add runs an idempotent
|
||||||
|
`ALTER TABLE Leaf ADD …` (no full rebuild required). Fresh `--rebuild` still
|
||||||
|
creates them in `CREATE NODE TABLE`.
|
||||||
|
|
||||||
## Sources (auto-pairing)
|
## Sources (auto-pairing)
|
||||||
|
|
||||||
- A: runtime state — `docker ps` (container running), ports actually bound
|
- A: runtime state — `docker ps` (container running), ports actually bound
|
||||||
@@ -71,7 +84,9 @@ corpus HEAD.
|
|||||||
- C: narrative — READMEs, AGENTS.md, docs
|
- C: narrative — READMEs, AGENTS.md, docs
|
||||||
|
|
||||||
Confirmed = A×B or B×C agreement. Single source = hypothesis + `(not confirmed)`.
|
Confirmed = A×B or B×C agreement. Single source = hypothesis + `(not confirmed)`.
|
||||||
Conflicting pairings (≥2 yes vs ≥2 no) = hypothesis (OQ1 → v2 resolution).
|
Conflicting pairings (≥2 yes vs ≥2 no) stay hypothesis until
|
||||||
|
`temporal_freshness` or `authority_pairing` fires (`bin/facts/audit contradict`,
|
||||||
|
[#29](https://git.produktor.io/eSlider/2dph/issues/29)).
|
||||||
|
|
||||||
## Read path
|
## Read path
|
||||||
|
|
||||||
|
|||||||
@@ -6,6 +6,18 @@ CPU reasoner. Default agent model is `qwen3.5:9b` (RAM path, D18). Weights stay
|
|||||||
in the reasoner volume, not in the 2dph image.
|
in the reasoner volume, not in the 2dph image.
|
||||||
No secrets in git: Ollama needs no key; MCP is local HTTP.
|
No secrets in git: Ollama needs no key; MCP is local HTTP.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/stack/start-assistant
|
||||||
|
bin/stack/start-assistant --no-attach
|
||||||
|
bin/stack/start-assistant -- -m "search the 2dph brain for LadybugDB"
|
||||||
|
bin/stack/status
|
||||||
|
bin/stack/stop
|
||||||
|
```
|
||||||
|
|
||||||
|
`start-assistant` reuses a healthy brain on `:8630`, starts the CPU reasoner,
|
||||||
|
pulls `qwen3.5:9b` if missing, brings up the gateway with `--no-deps picoclaw`,
|
||||||
|
then `picoclaw agent` (MCP `search` → `get` → `audit`). Gateway-only Compose:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
docker compose --profile picoclaw up -d
|
docker compose --profile picoclaw up -d
|
||||||
# already serving :8630 / :11435:
|
# already serving :8630 / :11435:
|
||||||
|
|||||||
+9
-4
@@ -39,11 +39,14 @@ Epic [#16](https://git.produktor.io/eSlider/2dph/issues/16) closed.
|
|||||||
[#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR — **in**.
|
[#6](https://git.produktor.io/eSlider/2dph/issues/6) OCR — **in**.
|
||||||
[#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3 duckdb-go — **in**.
|
[#30](https://git.produktor.io/eSlider/2dph/issues/30) OQ3 duckdb-go — **in**.
|
||||||
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1 contradiction
|
[#29](https://git.produktor.io/eSlider/2dph/issues/29) OQ1 contradiction
|
||||||
resolution.
|
resolution — **in** (`temporal_freshness`, `authority_pairing`).
|
||||||
|
[#34](https://git.produktor.io/eSlider/2dph/issues/34) D23 flaggy CLI — **in**.
|
||||||
|
[#36](https://git.produktor.io/eSlider/2dph/issues/36) OQ5/D24 fact intervals /
|
||||||
|
as-of — **in** (not D16 `temporal_freshness`).
|
||||||
|
|
||||||
## Blockers
|
## Blockers
|
||||||
|
|
||||||
None for epic #16 (closed). Remaining v2: OQ1, OQ4.
|
None for epic #16 (closed). Remaining v2: OQ4 (deferred).
|
||||||
|
|
||||||
```
|
```
|
||||||
question
|
question
|
||||||
@@ -58,8 +61,10 @@ question
|
|||||||
|
|
||||||
## Not v1
|
## Not v1
|
||||||
|
|
||||||
OQ1 contradiction resolution, OQ4 YAML-first leafs.
|
OQ4 YAML-first leafs.
|
||||||
OCR (OQ2) and duckdb-go (OQ3/D22) are in.
|
OQ5/D24 fact intervals + as-of — **in**
|
||||||
|
([#36](https://git.produktor.io/eSlider/2dph/issues/36)).
|
||||||
|
OCR (OQ2), duckdb-go (OQ3/D22), and D16 adjudication (OQ1) are in.
|
||||||
|
|
||||||
## Close epic #16 when
|
## Close epic #16 when
|
||||||
|
|
||||||
|
|||||||
+18
-1
@@ -53,11 +53,15 @@ bin/brain/add.go --text "arc-1 runs Matrix" --root facts --source "compose.yml x
|
|||||||
bin/brain/index.go --rebuild --with-facts --with-chats
|
bin/brain/index.go --rebuild --with-facts --with-chats
|
||||||
bin/brain/search.go "LadybugDB vector index" # facts → info → web (D17)
|
bin/brain/search.go "LadybugDB vector index" # facts → info → web (D17)
|
||||||
bin/brain/search.go "upstream flag" --no-web
|
bin/brain/search.go "upstream flag" --no-web
|
||||||
|
bin/brain/search.go "who works where" --as-of 2025-01-01 # D24 intervals
|
||||||
|
source <(./bin/cli/complete.go bash) # D23 flaggy complete
|
||||||
bin/brain/get.go <id> --body
|
bin/brain/get.go <id> --body
|
||||||
bin/brain/stats.go
|
bin/brain/stats.go
|
||||||
```
|
```
|
||||||
|
|
||||||
`--hop N` walks File → Commit → Person from each hit. Empty web results are `throttled`, not absence.
|
`--hop N` walks File → Commit → Person from each hit. `--as-of YYYY-MM-DD`
|
||||||
|
keeps leafs whose `valid_from`/`valid_to` cover that day (empty interval =
|
||||||
|
legacy always-on). Empty web results are `throttled`, not absence.
|
||||||
Gap to v1: [roadmap](roadmap.md) / [epic #16](https://git.produktor.io/eSlider/2dph/issues/16).
|
Gap to v1: [roadmap](roadmap.md) / [epic #16](https://git.produktor.io/eSlider/2dph/issues/16).
|
||||||
|
|
||||||
Ladybug 0.19: never `DROP INDEX` FTS/VECTOR (ghost catalog). Fresh indexes =
|
Ladybug 0.19: never `DROP INDEX` FTS/VECTOR (ghost catalog). Fresh indexes =
|
||||||
@@ -65,13 +69,26 @@ delete `var/kb.lbug` then `--rebuild`.
|
|||||||
|
|
||||||
## HTTP / MCP
|
## HTTP / MCP
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/stack/start # brain :8630, wait until MCP search/get/audit
|
||||||
|
bin/stack/status # YAML: brain / reasoner / picoclaw / mail_sync
|
||||||
|
bin/stack/start-mail-sync # compose ETL: OO+Gmail sync→import (300s; no auto-rebuild)
|
||||||
|
bin/stack/start-assistant # + qwen3.5:9b + PicoClaw agent (ask the brain)
|
||||||
|
bin/stack/start-assistant --no-attach
|
||||||
|
bin/stack/stop # compose stop; volumes kept
|
||||||
|
```
|
||||||
|
|
||||||
|
Same Compose services by hand:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
docker compose up -d brain # :8630 Zig CGO serve
|
docker compose up -d brain # :8630 Zig CGO serve
|
||||||
|
docker compose up -d mail-sync # ETL loop into kb-var
|
||||||
docker compose --profile index run --rm index # rebuild
|
docker compose --profile index run --rm index # rebuild
|
||||||
docker compose --profile picoclaw up brain-mcp # MCP 127.0.0.1:8630
|
docker compose --profile picoclaw up brain-mcp # MCP 127.0.0.1:8630
|
||||||
```
|
```
|
||||||
|
|
||||||
`GET /openapi.json`, `POST /mcp`. Agent tool order: `search` → `get` → `audit`.
|
`GET /openapi.json`, `POST /mcp`. Agent tool order: `search` → `get` → `audit`.
|
||||||
|
See [picoclaw.md](picoclaw.md).
|
||||||
|
|
||||||
## Reasoner (optional, D18)
|
## Reasoner (optional, D18)
|
||||||
|
|
||||||
|
|||||||
@@ -9,6 +9,7 @@ require (
|
|||||||
github.com/daulet/tokenizers v1.27.0
|
github.com/daulet/tokenizers v1.27.0
|
||||||
github.com/duckdb/duckdb-go/v2 v2.10505.0
|
github.com/duckdb/duckdb-go/v2 v2.10505.0
|
||||||
github.com/go-git/go-git/v5 v5.19.2
|
github.com/go-git/go-git/v5 v5.19.2
|
||||||
|
github.com/integrii/flaggy v1.8.0
|
||||||
golang.org/x/sys v0.47.0
|
golang.org/x/sys v0.47.0
|
||||||
golang.org/x/text v0.40.0
|
golang.org/x/text v0.40.0
|
||||||
modernc.org/sqlite v1.56.0
|
modernc.org/sqlite v1.56.0
|
||||||
|
|||||||
@@ -77,6 +77,8 @@ github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
|
|||||||
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
|
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
|
||||||
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
|
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
|
||||||
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
|
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
|
||||||
|
github.com/integrii/flaggy v1.8.0 h1:tC1qWwg4fhF2Qdaj+MpPK04cxlOSq0+HoMZqAW6Arao=
|
||||||
|
github.com/integrii/flaggy v1.8.0/go.mod h1:QS4c80m87SXG0pmVUT/Lx2RY5EbkLvLp7IKBD2jwcFA=
|
||||||
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99 h1:BQSFePA1RWJOlocH6Fxy8MmwDt+yVQYULKfN0RoTN8A=
|
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99 h1:BQSFePA1RWJOlocH6Fxy8MmwDt+yVQYULKfN0RoTN8A=
|
||||||
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99/go.mod h1:1lJo3i6rXxKeerYnT8Nvf0QmHCRC1n8sfWVwXF2Frvo=
|
github.com/jbenet/go-context v0.0.0-20150711004518-d14ea06fba99/go.mod h1:1lJo3i6rXxKeerYnT8Nvf0QmHCRC1n8sfWVwXF2Frvo=
|
||||||
github.com/kevinburke/ssh_config v1.2.0 h1:x584FjTGwHzMwvHx18PXxbBVzfnxogHaAReU4gf13a4=
|
github.com/kevinburke/ssh_config v1.2.0 h1:x584FjTGwHzMwvHx18PXxbBVzfnxogHaAReU4gf13a4=
|
||||||
|
|||||||
@@ -85,9 +85,18 @@ func openWithSandbox(epsv string) error {
|
|||||||
closeBrain()
|
closeBrain()
|
||||||
return fmt.Errorf("LOAD EXTENSION VECTOR: %w", err)
|
return fmt.Errorf("LOAD EXTENSION VECTOR: %w", err)
|
||||||
}
|
}
|
||||||
|
migrateIntervalColumns()
|
||||||
return nil
|
return nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// migrateIntervalColumns adds D24 valid_from/valid_to on existing Leaf tables.
|
||||||
|
// Fresh CREATE already has them; ALTER is a no-op when the column exists.
|
||||||
|
func migrateIntervalColumns() {
|
||||||
|
for _, col := range []string{"valid_from", "valid_to"} {
|
||||||
|
_, _ = conn.Query("ALTER TABLE Leaf ADD " + col + " STRING")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func closeBrain() {
|
func closeBrain() {
|
||||||
if conn != nil {
|
if conn != nil {
|
||||||
conn.Close()
|
conn.Close()
|
||||||
|
|||||||
@@ -21,8 +21,8 @@ func Ready() error {
|
|||||||
// HTTP is the in-process API used by bin/brain/serve.go.
|
// HTTP is the in-process API used by bin/brain/serve.go.
|
||||||
type HTTP struct{}
|
type HTTP struct{}
|
||||||
|
|
||||||
func (HTTP) Search(ctx context.Context, query string, limit int) ([]byte, error) {
|
func (HTTP) Search(ctx context.Context, query string, limit int, asOf string) ([]byte, error) {
|
||||||
hits, err := searchHits(query, "", "", limit)
|
hits, err := searchHits(query, "", "", limit, asOf)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
@@ -41,7 +41,7 @@ func (HTTP) Search(ctx context.Context, query string, limit int) ([]byte, error)
|
|||||||
var buf bytes.Buffer
|
var buf bytes.Buffer
|
||||||
enc := json.NewEncoder(&buf)
|
enc := json.NewEncoder(&buf)
|
||||||
enc.SetEscapeHTML(false)
|
enc.SetEscapeHTML(false)
|
||||||
if err := enc.Encode(toJSONOut(hits, query, "", webOut)); err != nil {
|
if err := enc.Encode(toJSONOut(hits, query, "", asOf, webOut)); err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
return buf.Bytes(), nil
|
return buf.Bytes(), nil
|
||||||
|
|||||||
+59
-51
@@ -3,12 +3,16 @@ package rank
|
|||||||
import (
|
import (
|
||||||
"fmt"
|
"fmt"
|
||||||
"strconv"
|
"strconv"
|
||||||
"strings"
|
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/eSlider/2dph/internal/facts"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
)
|
)
|
||||||
|
|
||||||
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--hop N] [--json] [--no-web]
|
const Usage = `usage: bin/brain/search.go "query" [--root facts|info] [--repo REPO] [-n N] [--hop N] [--as-of YYYY-MM-DD] [--json] [--no-web]
|
||||||
bin/brain/search.go serve [port]
|
bin/brain/search.go serve [port]
|
||||||
bin/brain/search.go --list-model`
|
bin/brain/search.go --list-model
|
||||||
|
source <(./bin/cli/complete.go bash)`
|
||||||
|
|
||||||
type Options struct {
|
type Options struct {
|
||||||
Query string
|
Query string
|
||||||
@@ -16,67 +20,71 @@ type Options struct {
|
|||||||
Repo string
|
Repo string
|
||||||
Limit int
|
Limit int
|
||||||
Hop int
|
Hop int
|
||||||
|
AsOf string
|
||||||
JSONOut bool
|
JSONOut bool
|
||||||
ListModel bool
|
ListModel bool
|
||||||
NoWeb bool
|
NoWeb bool
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// NewParser is the flaggy schema for search (also used by bin/cli/complete.go).
|
||||||
|
func NewParser(opt *Options) *flaggy.Parser {
|
||||||
|
if opt.Limit == 0 {
|
||||||
|
opt.Limit = 20
|
||||||
|
}
|
||||||
|
p := cli.New("brain-search")
|
||||||
|
p.Description = "deduction search: facts → info → web"
|
||||||
|
p.String(&opt.Root, "", "root", "facts or info")
|
||||||
|
p.String(&opt.Repo, "", "repo", "filter by repo")
|
||||||
|
p.Int(&opt.Limit, "n", "n", "max hits")
|
||||||
|
p.Int(&opt.Hop, "", "hop", "walk FROM_FILE depth 1-3")
|
||||||
|
p.String(&opt.AsOf, "", "as-of", "keep facts active on YYYY-MM-DD (D24)")
|
||||||
|
p.Bool(&opt.JSONOut, "", "json", "JSON output")
|
||||||
|
p.Bool(&opt.NoWeb, "", "no-web", "stay local")
|
||||||
|
p.Bool(&opt.ListModel, "", "list-model", "print embedding model")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
// ParseArgs reads flags. Unknown flags are an error: silently dropping them
|
// ParseArgs reads flags. Unknown flags are an error: silently dropping them
|
||||||
// meant `--hop 1` vanished and its argument `1` was appended to the query.
|
// meant `--hop 1` vanished and its argument `1` was appended to the query.
|
||||||
func ParseArgs(args []string) (Options, error) {
|
func ParseArgs(args []string) (Options, error) {
|
||||||
opt := Options{Limit: 20}
|
opt := Options{Limit: 20}
|
||||||
var queryArgs []string
|
p := NewParser(&opt)
|
||||||
|
var q string
|
||||||
for i := 0; i < len(args); i++ {
|
p.AddPositionalValue(&q, "query", 1, false, "search query")
|
||||||
arg := args[i]
|
if err := cli.Parse(p, args); err != nil {
|
||||||
wantsValue := arg == "--root" || arg == "--repo" || arg == "-n" || arg == "--hop"
|
return opt, err
|
||||||
if wantsValue && i+1 >= len(args) {
|
}
|
||||||
return opt, fmt.Errorf("%s needs a value", arg)
|
opt.Query = cli.Query(q, p.TrailingArguments)
|
||||||
|
if opt.Root != "" && opt.Root != "facts" && opt.Root != "info" {
|
||||||
|
return opt, fmt.Errorf("--root must be facts or info, got %q", opt.Root)
|
||||||
|
}
|
||||||
|
if opt.Limit < 1 {
|
||||||
|
return opt, fmt.Errorf("-n must be a positive integer, got %q", strconv.Itoa(opt.Limit))
|
||||||
|
}
|
||||||
|
if opt.Hop < 0 {
|
||||||
|
return opt, fmt.Errorf("--hop must be a positive integer, got %q", strconv.Itoa(opt.Hop))
|
||||||
|
}
|
||||||
|
if opt.Hop > 3 {
|
||||||
|
return opt, fmt.Errorf("--hop max is 3 (File → Commit → Person)")
|
||||||
|
}
|
||||||
|
if opt.AsOf != "" {
|
||||||
|
day := facts.NormalizeDay(opt.AsOf)
|
||||||
|
if len(day) != 10 || day[4] != '-' || day[7] != '-' {
|
||||||
|
return opt, fmt.Errorf("--as-of must be YYYY-MM-DD, got %q", opt.AsOf)
|
||||||
}
|
}
|
||||||
switch arg {
|
opt.AsOf = day
|
||||||
case "--root":
|
|
||||||
i++
|
|
||||||
opt.Root = args[i]
|
|
||||||
if opt.Root != "facts" && opt.Root != "info" {
|
|
||||||
return opt, fmt.Errorf("--root must be facts or info, got %q", opt.Root)
|
|
||||||
}
|
|
||||||
case "--repo":
|
|
||||||
i++
|
|
||||||
opt.Repo = args[i]
|
|
||||||
case "-n":
|
|
||||||
i++
|
|
||||||
n, err := strconv.Atoi(args[i])
|
|
||||||
if err != nil || n < 1 {
|
|
||||||
return opt, fmt.Errorf("-n must be a positive integer, got %q", args[i])
|
|
||||||
}
|
|
||||||
opt.Limit = n
|
|
||||||
case "--hop":
|
|
||||||
i++
|
|
||||||
n, err := strconv.Atoi(args[i])
|
|
||||||
if err != nil || n < 1 {
|
|
||||||
return opt, fmt.Errorf("--hop must be a positive integer, got %q", args[i])
|
|
||||||
}
|
|
||||||
if n > 3 {
|
|
||||||
return opt, fmt.Errorf("--hop max is 3 (File → Commit → Person)")
|
|
||||||
}
|
|
||||||
opt.Hop = n
|
|
||||||
case "--json":
|
|
||||||
opt.JSONOut = true
|
|
||||||
case "--no-web":
|
|
||||||
opt.NoWeb = true
|
|
||||||
case "--list-model":
|
|
||||||
opt.ListModel = true
|
|
||||||
default:
|
|
||||||
if strings.HasPrefix(arg, "-") {
|
|
||||||
return opt, fmt.Errorf("unknown flag %q", arg)
|
|
||||||
}
|
|
||||||
queryArgs = append(queryArgs, arg)
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
opt.Query = strings.TrimSpace(strings.Join(queryArgs, " "))
|
|
||||||
if opt.Query == "" && !opt.ListModel {
|
if opt.Query == "" && !opt.ListModel {
|
||||||
return opt, fmt.Errorf("no query given")
|
return opt, fmt.Errorf("no query given")
|
||||||
}
|
}
|
||||||
return opt, nil
|
return opt, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Parser is the search schema for bin/cli/complete.go.
|
||||||
|
func Parser() *flaggy.Parser {
|
||||||
|
opt := Options{Limit: 20}
|
||||||
|
p := NewParser(&opt)
|
||||||
|
var q string
|
||||||
|
p.AddPositionalValue(&q, "query", 1, false, "search query")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,34 @@
|
|||||||
|
package rank
|
||||||
|
|
||||||
|
import "testing"
|
||||||
|
|
||||||
|
func TestFilterAsOfKeepsXDropsY(t *testing.T) {
|
||||||
|
hits := []Hit{
|
||||||
|
{ID: "x", Text: "Andrey works at X", ValidFrom: "2024-03-01", ValidTo: "2025-07-15"},
|
||||||
|
{ID: "y", Text: "Andrey works at Y", ValidFrom: "2025-07-16", ValidTo: ""},
|
||||||
|
{ID: "legacy", Text: "always true claim", ValidFrom: "", ValidTo: ""},
|
||||||
|
}
|
||||||
|
out := FilterAsOf(hits, "2025-01-01")
|
||||||
|
if len(out) != 2 {
|
||||||
|
t.Fatalf("len=%d want 2: %+v", len(out), out)
|
||||||
|
}
|
||||||
|
if out[0].ID != "x" || out[1].ID != "legacy" {
|
||||||
|
t.Fatalf("got %+v", out)
|
||||||
|
}
|
||||||
|
if FilterAsOf(hits, "") == nil || len(FilterAsOf(hits, "")) != 3 {
|
||||||
|
t.Fatal("empty as-of must keep all")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseArgsAsOf(t *testing.T) {
|
||||||
|
opt, err := ParseArgs([]string{"who works where", "--as-of", "2025-01-01", "--json"})
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
if opt.AsOf != "2025-01-01" {
|
||||||
|
t.Fatalf("AsOf=%q", opt.AsOf)
|
||||||
|
}
|
||||||
|
if _, err := ParseArgs([]string{"q", "--as-of", "not-a-date"}); err == nil {
|
||||||
|
t.Fatal("expected bad as-of error")
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -19,20 +19,35 @@ type SecondSourceHit struct {
|
|||||||
|
|
||||||
type WebFn func(query string) SecondSource
|
type WebFn func(query string) SecondSource
|
||||||
|
|
||||||
// ShouldEscalate is true when the default deduction path has no facts hit.
|
// ShouldEscalate is true when the default deduction path has no confirmed
|
||||||
|
// facts hit. Hypothesis/partial facts are `(not confirmed)` (D16).
|
||||||
// `--root facts|info` is a single-root ask: do not mix in the web.
|
// `--root facts|info` is a single-root ask: do not mix in the web.
|
||||||
func ShouldEscalate(hits []Hit, rootFilter string) bool {
|
func ShouldEscalate(hits []Hit, rootFilter string) bool {
|
||||||
if rootFilter != "" {
|
if rootFilter != "" {
|
||||||
return false
|
return false
|
||||||
}
|
}
|
||||||
for _, h := range hits {
|
for _, h := range hits {
|
||||||
if h.Root == "facts" {
|
if ConfirmedFact(h) {
|
||||||
return false
|
return false
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return true
|
return true
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// ConfirmedFact is a facts-root hit that is not hypothesis/partial.
|
||||||
|
// Empty confidence is treated as confirmed (legacy leafs).
|
||||||
|
func ConfirmedFact(h Hit) bool {
|
||||||
|
if h.Root != "facts" {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
switch h.Confidence {
|
||||||
|
case "hypothesis", "partial":
|
||||||
|
return false
|
||||||
|
default:
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// Deduce returns the second-source block, or nil when web must not run.
|
// Deduce returns the second-source block, or nil when web must not run.
|
||||||
func Deduce(hits []Hit, query, rootFilter string, noWeb bool, web WebFn) *SecondSource {
|
func Deduce(hits []Hit, query, rootFilter string, noWeb bool, web WebFn) *SecondSource {
|
||||||
if noWeb || web == nil || !ShouldEscalate(hits, rootFilter) {
|
if noWeb || web == nil || !ShouldEscalate(hits, rootFilter) {
|
||||||
|
|||||||
@@ -14,6 +14,16 @@ func TestShouldEscalateWhenNoFacts(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestShouldEscalateWhenHypothesisFacts(t *testing.T) {
|
||||||
|
hyp := Hit{ID: "c", Root: "facts", Confidence: "hypothesis", Source: "a x b vs c x d"}
|
||||||
|
if !ShouldEscalate([]Hit{hyp}, "") {
|
||||||
|
t.Fatal("hypothesis facts are (not confirmed); escalate")
|
||||||
|
}
|
||||||
|
if ConfirmedFact(hyp) {
|
||||||
|
t.Fatal("hypothesis is not confirmed")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestShouldNotEscalateWhenFactsConfirm(t *testing.T) {
|
func TestShouldNotEscalateWhenFactsConfirm(t *testing.T) {
|
||||||
hits := []Hit{h("f", "facts", "docker ps x compose"), h("i", "info", "docs/a.md")}
|
hits := []Hit{h("f", "facts", "docker ps x compose"), h("i", "info", "docs/a.md")}
|
||||||
if ShouldEscalate(hits, "") {
|
if ShouldEscalate(hits, "") {
|
||||||
|
|||||||
@@ -3,10 +3,12 @@ package rank
|
|||||||
// BM25 ranks best-first, so the top hits are the *highest* scores; cosine
|
// BM25 ranks best-first, so the top hits are the *highest* scores; cosine
|
||||||
// distance ranks best-first ascending. Both mirror kblib.py.
|
// distance ranks best-first ascending. Both mirror kblib.py.
|
||||||
const FTSStmt = "CALL QUERY_FTS_INDEX('Leaf', 'id', $q) " +
|
const FTSStmt = "CALL QUERY_FTS_INDEX('Leaf', 'id', $q) " +
|
||||||
"RETURN node.id, node.text, node.root, node.source, score ORDER BY score DESC LIMIT $n"
|
"RETURN node.id, node.text, node.root, node.source, score, node.confidence, " +
|
||||||
|
"node.valid_from, node.valid_to ORDER BY score DESC LIMIT $n"
|
||||||
|
|
||||||
const VecStmt = "CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) " +
|
const VecStmt = "CALL QUERY_VECTOR_INDEX('Leaf', 'Leaf_vec', $q, $n) " +
|
||||||
"RETURN node.id, node.text, node.root, node.source, distance ORDER BY distance LIMIT $n"
|
"RETURN node.id, node.text, node.root, node.source, distance, node.confidence, " +
|
||||||
|
"node.valid_from, node.valid_to ORDER BY distance LIMIT $n"
|
||||||
|
|
||||||
// HopStmt is the Cypher walk from a search hit. Depth 1 = File, 2 = Commit, 3 = Person.
|
// HopStmt is the Cypher walk from a search hit. Depth 1 = File, 2 = Commit, 3 = Person.
|
||||||
func HopStmt(depth int) string {
|
func HopStmt(depth int) string {
|
||||||
|
|||||||
+39
-11
@@ -5,6 +5,8 @@ package rank
|
|||||||
import (
|
import (
|
||||||
"sort"
|
"sort"
|
||||||
"strings"
|
"strings"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/facts"
|
||||||
)
|
)
|
||||||
|
|
||||||
type HopNode struct {
|
type HopNode struct {
|
||||||
@@ -16,23 +18,31 @@ type HopNode struct {
|
|||||||
|
|
||||||
// Hit is one search result, mirroring the python script's dict shape.
|
// Hit is one search result, mirroring the python script's dict shape.
|
||||||
type Hit struct {
|
type Hit struct {
|
||||||
ID string `json:"id"`
|
ID string `json:"id"`
|
||||||
Text string `json:"text"`
|
Text string `json:"text"`
|
||||||
Root string `json:"root"`
|
Root string `json:"root"`
|
||||||
Source string `json:"-"`
|
Confidence string `json:"confidence,omitempty"`
|
||||||
Score float64 `json:"score"`
|
Source string `json:"-"`
|
||||||
Snippet string `json:"snippet,omitempty"`
|
Score float64 `json:"score"`
|
||||||
Hops []HopNode `json:"hops,omitempty"`
|
Snippet string `json:"snippet,omitempty"`
|
||||||
|
ValidFrom string `json:"valid_from,omitempty"`
|
||||||
|
ValidTo string `json:"valid_to,omitempty"`
|
||||||
|
Hops []HopNode `json:"hops,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
// rrfK dampens the contribution of low ranks; same constant as kblib.py.
|
// rrfK dampens the contribution of low ranks; same constant as kblib.py.
|
||||||
const rrfK = 60
|
const rrfK = 60
|
||||||
|
|
||||||
// RankAndFilter fuses the two hit lists, applies --root/--repo, then cuts to
|
// RankAndFilter fuses the two hit lists, applies --root/--repo/--as-of, then
|
||||||
// limit. Cutting first dropped every matching leaf ranked below the cut, so
|
// cuts to limit. Cutting first dropped every matching leaf ranked below the
|
||||||
// `--root facts` came back empty whenever info leafs filled the top N.
|
// cut, so `--root facts` came back empty whenever info leafs filled the top N.
|
||||||
// limit <= 0 keeps everything.
|
// limit <= 0 keeps everything. asOf empty skips interval filter (D24).
|
||||||
func RankAndFilter(fts, vec []Hit, root, repo string, limit int) []Hit {
|
func RankAndFilter(fts, vec []Hit, root, repo string, limit int) []Hit {
|
||||||
|
return RankAndFilterAsOf(fts, vec, root, repo, "", limit)
|
||||||
|
}
|
||||||
|
|
||||||
|
// RankAndFilterAsOf is RankAndFilter with D24 fact-interval filter.
|
||||||
|
func RankAndFilterAsOf(fts, vec []Hit, root, repo, asOf string, limit int) []Hit {
|
||||||
out := Hybrid(fts, vec, 0)
|
out := Hybrid(fts, vec, 0)
|
||||||
if root != "" {
|
if root != "" {
|
||||||
out = FilterRoot(out, root)
|
out = FilterRoot(out, root)
|
||||||
@@ -40,12 +50,30 @@ func RankAndFilter(fts, vec []Hit, root, repo string, limit int) []Hit {
|
|||||||
if repo != "" {
|
if repo != "" {
|
||||||
out = FilterRepo(out, repo)
|
out = FilterRepo(out, repo)
|
||||||
}
|
}
|
||||||
|
if asOf != "" {
|
||||||
|
out = FilterAsOf(out, asOf)
|
||||||
|
}
|
||||||
if limit > 0 && len(out) > limit {
|
if limit > 0 && len(out) > limit {
|
||||||
out = out[:limit]
|
out = out[:limit]
|
||||||
}
|
}
|
||||||
return out
|
return out
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// FilterAsOf keeps hits whose [valid_from, valid_to] covers asOf (D24).
|
||||||
|
// Empty intervals stay (legacy leafs). Empty asOf keeps all.
|
||||||
|
func FilterAsOf(hits []Hit, asOf string) []Hit {
|
||||||
|
if asOf == "" {
|
||||||
|
return hits
|
||||||
|
}
|
||||||
|
var out []Hit
|
||||||
|
for _, h := range hits {
|
||||||
|
if facts.ActiveAt(h.ValidFrom, h.ValidTo, asOf) {
|
||||||
|
out = append(out, h)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
// Hybrid merges FTS and vector hits by reciprocal rank fusion.
|
// Hybrid merges FTS and vector hits by reciprocal rank fusion.
|
||||||
// limit <= 0 returns the full fused list.
|
// limit <= 0 returns the full fused list.
|
||||||
func Hybrid(fts, vec []Hit, limit int) []Hit {
|
func Hybrid(fts, vec []Hit, limit int) []Hit {
|
||||||
|
|||||||
@@ -172,4 +172,7 @@ func TestFTSQueryOrdersByScoreDescending(t *testing.T) {
|
|||||||
if !strings.Contains(FTSStmt, "ORDER BY score DESC") {
|
if !strings.Contains(FTSStmt, "ORDER BY score DESC") {
|
||||||
t.Fatalf("FTS query must order by score DESC, got:\n%s", FTSStmt)
|
t.Fatalf("FTS query must order by score DESC, got:\n%s", FTSStmt)
|
||||||
}
|
}
|
||||||
|
if !strings.Contains(FTSStmt, "node.confidence") {
|
||||||
|
t.Fatal("FTS must return confidence for D16")
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,59 @@
|
|||||||
|
package rank
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
type GetOptions struct {
|
||||||
|
ID string
|
||||||
|
Body bool
|
||||||
|
JSONOut bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func GetParser(opt *GetOptions) *flaggy.Parser {
|
||||||
|
p := cli.New("brain-get")
|
||||||
|
p.Description = "read one leaf"
|
||||||
|
p.Bool(&opt.Body, "", "body", "full text instead of snippet")
|
||||||
|
p.Bool(&opt.JSONOut, "", "json", "JSON output")
|
||||||
|
p.AddPositionalValue(&opt.ID, "id", 1, false, "leaf id")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseGet(args []string) (GetOptions, error) {
|
||||||
|
var opt GetOptions
|
||||||
|
if err := cli.Parse(GetParser(&opt), args); err != nil {
|
||||||
|
return opt, err
|
||||||
|
}
|
||||||
|
if opt.ID == "" {
|
||||||
|
return opt, fmt.Errorf("id required")
|
||||||
|
}
|
||||||
|
return opt, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
type JSONFlag struct {
|
||||||
|
JSONOut bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func bindJSON(name string, opt *JSONFlag) *flaggy.Parser {
|
||||||
|
p := cli.New(name)
|
||||||
|
p.Bool(&opt.JSONOut, "", "json", "JSON output")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func StatsParser() *flaggy.Parser {
|
||||||
|
opt := JSONFlag{}
|
||||||
|
return bindJSON("brain-stats", &opt)
|
||||||
|
}
|
||||||
|
|
||||||
|
func EvalParser() *flaggy.Parser {
|
||||||
|
opt := JSONFlag{}
|
||||||
|
return bindJSON("brain-eval", &opt)
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseJSONFlag(name string, args []string) (JSONFlag, error) {
|
||||||
|
var opt JSONFlag
|
||||||
|
return opt, cli.Parse(bindJSON(name, &opt), args)
|
||||||
|
}
|
||||||
+13
-48
@@ -11,30 +11,15 @@ import (
|
|||||||
"unicode/utf8"
|
"unicode/utf8"
|
||||||
|
|
||||||
"github.com/eSlider/2dph/internal/brain/rank"
|
"github.com/eSlider/2dph/internal/brain/rank"
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
)
|
)
|
||||||
|
|
||||||
func MainGet(args []string) int {
|
func MainGet(args []string) int {
|
||||||
id, body, jsonOut := "", false, false
|
opt, err := rank.ParseGet(args)
|
||||||
for _, a := range args {
|
if err != nil {
|
||||||
switch {
|
return cli.Fail(err)
|
||||||
case a == "--body":
|
|
||||||
body = true
|
|
||||||
case a == "--json":
|
|
||||||
jsonOut = true
|
|
||||||
case a == "-h" || a == "--help":
|
|
||||||
fmt.Fprintln(os.Stderr, `usage: bin/brain/get.go <id> [--body] [--json]`)
|
|
||||||
return 0
|
|
||||||
case strings.HasPrefix(a, "-"):
|
|
||||||
fmt.Fprintf(os.Stderr, "brain/get: unknown flag %s\n", a)
|
|
||||||
return 2
|
|
||||||
default:
|
|
||||||
id = a
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if id == "" {
|
|
||||||
fmt.Fprintln(os.Stderr, "brain/get: id required")
|
|
||||||
return 2
|
|
||||||
}
|
}
|
||||||
|
id, body, jsonOut := opt.ID, opt.Body, opt.JSONOut
|
||||||
if err := openBrain(); err != nil {
|
if err := openBrain(); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
|
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
@@ -72,21 +57,11 @@ func MainGet(args []string) int {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func MainStats(args []string) int {
|
func MainStats(args []string) int {
|
||||||
jsonOut := false
|
opt, err := rank.ParseJSONFlag("brain-stats", args)
|
||||||
for _, a := range args {
|
if err != nil {
|
||||||
switch a {
|
return cli.Fail(err)
|
||||||
case "--json":
|
|
||||||
jsonOut = true
|
|
||||||
case "-h", "--help":
|
|
||||||
fmt.Fprintln(os.Stderr, `usage: bin/brain/stats.go [--json]`)
|
|
||||||
return 0
|
|
||||||
default:
|
|
||||||
if strings.HasPrefix(a, "-") {
|
|
||||||
fmt.Fprintf(os.Stderr, "brain/stats: unknown flag %s\n", a)
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
jsonOut := opt.JSONOut
|
||||||
if err := openBrain(); err != nil {
|
if err := openBrain(); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
|
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
@@ -124,21 +99,11 @@ func MainStats(args []string) int {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func MainEval(args []string) int {
|
func MainEval(args []string) int {
|
||||||
jsonOut := false
|
opt, err := rank.ParseJSONFlag("brain-eval", args)
|
||||||
for _, a := range args {
|
if err != nil {
|
||||||
switch a {
|
return cli.Fail(err)
|
||||||
case "--json":
|
|
||||||
jsonOut = true
|
|
||||||
case "-h", "--help":
|
|
||||||
fmt.Fprintln(os.Stderr, `usage: bin/brain/eval.go [--json]`)
|
|
||||||
return 0
|
|
||||||
default:
|
|
||||||
if strings.HasPrefix(a, "-") {
|
|
||||||
fmt.Fprintf(os.Stderr, "brain/eval: unknown flag %s\n", a)
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
jsonOut := opt.JSONOut
|
||||||
if err := openBrain(); err != nil {
|
if err := openBrain(); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
|
fmt.Fprintf(os.Stderr, "open brain: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
|
|||||||
+63
-18
@@ -20,6 +20,7 @@ import (
|
|||||||
|
|
||||||
lbug "github.com/LadybugDB/go-ladybug"
|
lbug "github.com/LadybugDB/go-ladybug"
|
||||||
"github.com/eSlider/2dph/internal/brain/rank"
|
"github.com/eSlider/2dph/internal/brain/rank"
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
)
|
)
|
||||||
|
|
||||||
const defaultPort = 17830
|
const defaultPort = 17830
|
||||||
@@ -29,6 +30,9 @@ const healthPath = "/health"
|
|||||||
func runSearch(args []string) int {
|
func runSearch(args []string) int {
|
||||||
opt, err := rank.ParseArgs(args)
|
opt, err := rank.ParseArgs(args)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
|
if errors.Is(err, cli.ErrHelp) {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
fmt.Fprintf(os.Stderr, "brain/search: %v\n%s\n", err, rank.Usage)
|
fmt.Fprintf(os.Stderr, "brain/search: %v\n%s\n", err, rank.Usage)
|
||||||
return 2
|
return 2
|
||||||
}
|
}
|
||||||
@@ -51,7 +55,7 @@ func runSearch(args []string) int {
|
|||||||
}
|
}
|
||||||
defer closeBrain()
|
defer closeBrain()
|
||||||
|
|
||||||
hits, err := searchHits(query, root, repo, limit)
|
hits, err := searchHits(query, root, repo, limit, opt.AsOf)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "search: %v\n", err)
|
fmt.Fprintf(os.Stderr, "search: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
@@ -81,6 +85,7 @@ func runSearch(args []string) int {
|
|||||||
out := Dict{
|
out := Dict{
|
||||||
{"query", query},
|
{"query", query},
|
||||||
{"root_filter", root},
|
{"root_filter", root},
|
||||||
|
{"as_of", opt.AsOf},
|
||||||
{"count", len(results)},
|
{"count", len(results)},
|
||||||
{"results", resultsToDicts(results)},
|
{"results", resultsToDicts(results)},
|
||||||
}
|
}
|
||||||
@@ -92,13 +97,13 @@ func runSearch(args []string) int {
|
|||||||
enc := json.NewEncoder(os.Stdout)
|
enc := json.NewEncoder(os.Stdout)
|
||||||
enc.SetIndent("", " ")
|
enc.SetIndent("", " ")
|
||||||
enc.SetEscapeHTML(false)
|
enc.SetEscapeHTML(false)
|
||||||
return b2i(enc.Encode(toJSONOut(results, query, root, webOut)))
|
return b2i(enc.Encode(toJSONOut(results, query, root, opt.AsOf, webOut)))
|
||||||
}
|
}
|
||||||
fmt.Print(toYAML(out, 0))
|
fmt.Print(toYAML(out, 0))
|
||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
|
|
||||||
func searchHits(query, root, repo string, limit int) ([]Hit, error) {
|
func searchHits(query, root, repo string, limit int, asOf string) ([]Hit, error) {
|
||||||
emb, err := embedQuery(query)
|
emb, err := embedQuery(query)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, fmt.Errorf("embed: %w", err)
|
return nil, fmt.Errorf("embed: %w", err)
|
||||||
@@ -111,7 +116,7 @@ func searchHits(query, root, repo string, limit int) ([]Hit, error) {
|
|||||||
if vec, err = queryVector(emb, limit*3); err != nil {
|
if vec, err = queryVector(emb, limit*3); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "vec: %v\n", err)
|
fmt.Fprintf(os.Stderr, "vec: %v\n", err)
|
||||||
}
|
}
|
||||||
return rank.RankAndFilter(fts, vec, root, repo, limit), nil
|
return rank.RankAndFilterAsOf(fts, vec, root, repo, asOf, limit), nil
|
||||||
}
|
}
|
||||||
|
|
||||||
func attachHops(hits []Hit, n int) error {
|
func attachHops(hits []Hit, n int) error {
|
||||||
@@ -212,44 +217,75 @@ func rowsToHits(res *lbug.QueryResult) ([]Hit, error) {
|
|||||||
root := fmt.Sprint(vals[2])
|
root := fmt.Sprint(vals[2])
|
||||||
source := fmt.Sprint(vals[3])
|
source := fmt.Sprint(vals[3])
|
||||||
score := float64(vals[4].(float64))
|
score := float64(vals[4].(float64))
|
||||||
hits = append(hits, Hit{ID: id, Text: text, Root: root, Source: source, Score: score})
|
conf := ""
|
||||||
|
if len(vals) >= 6 {
|
||||||
|
conf = fmt.Sprint(vals[5])
|
||||||
|
}
|
||||||
|
vf, vt := "", ""
|
||||||
|
if len(vals) >= 8 {
|
||||||
|
vf = nullStr(vals[6])
|
||||||
|
vt = nullStr(vals[7])
|
||||||
|
}
|
||||||
|
hits = append(hits, Hit{
|
||||||
|
ID: id, Text: text, Root: root, Source: source, Score: score,
|
||||||
|
Confidence: conf, ValidFrom: vf, ValidTo: vt,
|
||||||
|
})
|
||||||
}
|
}
|
||||||
return hits, nil
|
return hits, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func nullStr(v any) string {
|
||||||
|
if v == nil {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
s := fmt.Sprint(v)
|
||||||
|
if s == "<nil>" {
|
||||||
|
return ""
|
||||||
|
}
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
// JSON output types
|
// JSON output types
|
||||||
type jsonOut struct {
|
type jsonOut struct {
|
||||||
Query string `json:"query"`
|
Query string `json:"query"`
|
||||||
RootFilter string `json:"root_filter"`
|
RootFilter string `json:"root_filter"`
|
||||||
|
AsOf string `json:"as_of,omitempty"`
|
||||||
Count int `json:"count"`
|
Count int `json:"count"`
|
||||||
Results []jsonHit `json:"results"`
|
Results []jsonHit `json:"results"`
|
||||||
Web *rank.SecondSource `json:"web,omitempty"`
|
Web *rank.SecondSource `json:"web,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
type jsonHit struct {
|
type jsonHit struct {
|
||||||
ID string `json:"id"`
|
ID string `json:"id"`
|
||||||
Text string `json:"text"`
|
Text string `json:"text"`
|
||||||
Root string `json:"root"`
|
Root string `json:"root"`
|
||||||
Score float64 `json:"score"`
|
Confidence string `json:"confidence,omitempty"`
|
||||||
Snippet string `json:"snippet,omitempty"`
|
Score float64 `json:"score"`
|
||||||
Hops []rank.HopNode `json:"hops,omitempty"`
|
Snippet string `json:"snippet,omitempty"`
|
||||||
|
ValidFrom string `json:"valid_from,omitempty"`
|
||||||
|
ValidTo string `json:"valid_to,omitempty"`
|
||||||
|
Hops []rank.HopNode `json:"hops,omitempty"`
|
||||||
}
|
}
|
||||||
|
|
||||||
func toJSONOut(hits []Hit, query, rootFilter string, web *rank.SecondSource) *jsonOut {
|
func toJSONOut(hits []Hit, query, rootFilter, asOf string, web *rank.SecondSource) *jsonOut {
|
||||||
out := make([]jsonHit, len(hits))
|
out := make([]jsonHit, len(hits))
|
||||||
for i, h := range hits {
|
for i, h := range hits {
|
||||||
out[i] = jsonHit{
|
out[i] = jsonHit{
|
||||||
ID: h.ID,
|
ID: h.ID,
|
||||||
Text: h.Text,
|
Text: h.Text,
|
||||||
Root: h.Root,
|
Root: h.Root,
|
||||||
Score: h.Score,
|
Confidence: h.Confidence,
|
||||||
Snippet: h.Snippet,
|
Score: h.Score,
|
||||||
Hops: h.Hops,
|
Snippet: h.Snippet,
|
||||||
|
ValidFrom: h.ValidFrom,
|
||||||
|
ValidTo: h.ValidTo,
|
||||||
|
Hops: h.Hops,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return &jsonOut{
|
return &jsonOut{
|
||||||
Query: query,
|
Query: query,
|
||||||
RootFilter: rootFilter,
|
RootFilter: rootFilter,
|
||||||
|
AsOf: asOf,
|
||||||
Count: len(hits),
|
Count: len(hits),
|
||||||
Results: out,
|
Results: out,
|
||||||
Web: web,
|
Web: web,
|
||||||
@@ -265,6 +301,15 @@ func resultsToDicts(hits []Hit) []any {
|
|||||||
{"root", h.Root},
|
{"root", h.Root},
|
||||||
{"score", h.Score},
|
{"score", h.Score},
|
||||||
}
|
}
|
||||||
|
if h.Confidence != "" {
|
||||||
|
d = append(d, KV{"confidence", h.Confidence})
|
||||||
|
}
|
||||||
|
if h.ValidFrom != "" {
|
||||||
|
d = append(d, KV{"valid_from", h.ValidFrom})
|
||||||
|
}
|
||||||
|
if h.ValidTo != "" {
|
||||||
|
d = append(d, KV{"valid_to", h.ValidTo})
|
||||||
|
}
|
||||||
if h.Snippet != "" {
|
if h.Snippet != "" {
|
||||||
d = append(d, KV{"snippet", h.Snippet})
|
d = append(d, KV{"snippet", h.Snippet})
|
||||||
}
|
}
|
||||||
|
|||||||
+15
-21
@@ -3,21 +3,22 @@ package chats
|
|||||||
import (
|
import (
|
||||||
"bytes"
|
"bytes"
|
||||||
"encoding/json"
|
"encoding/json"
|
||||||
"flag"
|
|
||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"os/exec"
|
"os/exec"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"strings"
|
"strings"
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
)
|
)
|
||||||
|
|
||||||
type ooContact struct {
|
type ooContact struct {
|
||||||
ID int `json:"id"`
|
ID int `json:"id"`
|
||||||
DisplayName string `json:"displayName"`
|
DisplayName string `json:"displayName"`
|
||||||
FirstName string `json:"firstName"`
|
FirstName string `json:"firstName"`
|
||||||
LastName string `json:"lastName"`
|
LastName string `json:"lastName"`
|
||||||
About string `json:"about"`
|
About string `json:"about"`
|
||||||
CommonData []struct {
|
CommonData []struct {
|
||||||
InfoType int `json:"infoType"`
|
InfoType int `json:"infoType"`
|
||||||
Data string `json:"data"`
|
Data string `json:"data"`
|
||||||
Category string `json:"categoryName"`
|
Category string `json:"categoryName"`
|
||||||
@@ -25,16 +26,9 @@ type ooContact struct {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func RunApply(args []string) int {
|
func RunApply(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats apply", flag.ContinueOnError)
|
dryRun, err := parseApplyFlags(args)
|
||||||
dryRun := fs.Bool("dry-run", false, "show what would be done without writing")
|
if err != nil {
|
||||||
help := fs.Bool("help", false, "")
|
return cliparse.Fail(err)
|
||||||
fs.SetOutput(os.Stderr)
|
|
||||||
if err := fs.Parse(args); err != nil {
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
if *help {
|
|
||||||
fmt.Fprintln(os.Stderr, "usage: chats apply [--dry-run]")
|
|
||||||
return 0
|
|
||||||
}
|
}
|
||||||
|
|
||||||
ooCLI := findOO()
|
ooCLI := findOO()
|
||||||
@@ -60,10 +54,10 @@ func RunApply(args []string) int {
|
|||||||
emailFacts = dedupeFacts(emailFacts)
|
emailFacts = dedupeFacts(emailFacts)
|
||||||
|
|
||||||
type resolvedFact struct {
|
type resolvedFact struct {
|
||||||
Fact ExtractedFact
|
Fact ExtractedFact
|
||||||
OoID int
|
OoID int
|
||||||
OoName string
|
OoName string
|
||||||
Action string // "info-add" or "persons-create"
|
Action string // "info-add" or "persons-create"
|
||||||
}
|
}
|
||||||
|
|
||||||
var resolved []resolvedFact
|
var resolved []resolvedFact
|
||||||
@@ -128,7 +122,7 @@ func RunApply(args []string) int {
|
|||||||
|
|
||||||
fmt.Printf("\nchats apply: %d actions to apply\n", len(resolved))
|
fmt.Printf("\nchats apply: %d actions to apply\n", len(resolved))
|
||||||
|
|
||||||
if *dryRun {
|
if dryRun {
|
||||||
for _, r := range resolved {
|
for _, r := range resolved {
|
||||||
switch r.Action {
|
switch r.Action {
|
||||||
case "info-add":
|
case "info-add":
|
||||||
|
|||||||
@@ -0,0 +1,74 @@
|
|||||||
|
package chats
|
||||||
|
|
||||||
|
import (
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
type syncTelegramFlags struct {
|
||||||
|
Limit int
|
||||||
|
Phone string
|
||||||
|
}
|
||||||
|
|
||||||
|
type syncLinkedInFlags struct {
|
||||||
|
Limit int
|
||||||
|
Refresh bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func SyncParser() *flaggy.Parser {
|
||||||
|
p := cliparse.New("chats-sync")
|
||||||
|
p.Description = "download chats to var/chats"
|
||||||
|
tg := flaggy.NewSubcommand("telegram")
|
||||||
|
li := flaggy.NewSubcommand("linkedin")
|
||||||
|
var limit int
|
||||||
|
var phone string
|
||||||
|
var refresh bool
|
||||||
|
tg.Int(&limit, "", "limit", "max messages per chat")
|
||||||
|
tg.String(&phone, "", "phone", "phone (default TELEGRAM_PHONE)")
|
||||||
|
li.Int(&limit, "", "limit", "max messages per conversation")
|
||||||
|
li.Bool(&refresh, "", "refresh", "refresh webtop session")
|
||||||
|
p.AttachSubcommand(tg, 1)
|
||||||
|
p.AttachSubcommand(li, 1)
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func ImportParser() *flaggy.Parser {
|
||||||
|
return cliparse.New("chats-import")
|
||||||
|
}
|
||||||
|
|
||||||
|
func FactsParser() *flaggy.Parser {
|
||||||
|
return cliparse.New("chats-facts")
|
||||||
|
}
|
||||||
|
|
||||||
|
func ApplyParser() *flaggy.Parser {
|
||||||
|
p := cliparse.New("chats-apply")
|
||||||
|
dry := false
|
||||||
|
p.Bool(&dry, "", "dry-run", "show without writing")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func parseTelegramFlags(args []string) (syncTelegramFlags, error) {
|
||||||
|
var f syncTelegramFlags
|
||||||
|
p := cliparse.New("chats-sync-telegram")
|
||||||
|
p.Int(&f.Limit, "", "limit", "max messages per chat")
|
||||||
|
p.String(&f.Phone, "", "phone", "phone (default TELEGRAM_PHONE)")
|
||||||
|
return f, cliparse.Parse(p, args)
|
||||||
|
}
|
||||||
|
|
||||||
|
func parseLinkedInFlags(args []string) (syncLinkedInFlags, error) {
|
||||||
|
var f syncLinkedInFlags
|
||||||
|
p := cliparse.New("chats-sync-linkedin")
|
||||||
|
p.Int(&f.Limit, "", "limit", "max messages per conversation")
|
||||||
|
p.Bool(&f.Refresh, "", "refresh", "refresh webtop session")
|
||||||
|
return f, cliparse.Parse(p, args)
|
||||||
|
}
|
||||||
|
|
||||||
|
func parseApplyFlags(args []string) (dryRun bool, err error) {
|
||||||
|
p := cliparse.New("chats-apply")
|
||||||
|
p.Bool(&dryRun, "", "dry-run", "show without writing")
|
||||||
|
return dryRun, cliparse.Parse(p, args)
|
||||||
|
}
|
||||||
|
|
||||||
|
func parseNoFlags(name string, args []string) error {
|
||||||
|
return cliparse.Parse(cliparse.New(name), args)
|
||||||
|
}
|
||||||
+4
-10
@@ -3,12 +3,13 @@ package chats
|
|||||||
import (
|
import (
|
||||||
"bufio"
|
"bufio"
|
||||||
"encoding/json"
|
"encoding/json"
|
||||||
"flag"
|
|
||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"regexp"
|
"regexp"
|
||||||
"strings"
|
"strings"
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
)
|
)
|
||||||
|
|
||||||
var (
|
var (
|
||||||
@@ -76,15 +77,8 @@ type ExtractedFact struct {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func RunFacts(args []string) int {
|
func RunFacts(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats facts", flag.ContinueOnError)
|
if err := parseNoFlags("chats-facts", args); err != nil {
|
||||||
help := fs.Bool("help", false, "")
|
return cliparse.Fail(err)
|
||||||
fs.SetOutput(os.Stderr)
|
|
||||||
if err := fs.Parse(args); err != nil {
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
if *help {
|
|
||||||
fmt.Fprintln(os.Stderr, "usage: chats facts")
|
|
||||||
return 0
|
|
||||||
}
|
}
|
||||||
|
|
||||||
root := Dir()
|
root := Dir()
|
||||||
|
|||||||
@@ -4,25 +4,19 @@ import (
|
|||||||
"bufio"
|
"bufio"
|
||||||
"bytes"
|
"bytes"
|
||||||
"encoding/json"
|
"encoding/json"
|
||||||
"flag"
|
|
||||||
"fmt"
|
"fmt"
|
||||||
"html"
|
"html"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"sort"
|
"sort"
|
||||||
"strings"
|
"strings"
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
)
|
)
|
||||||
|
|
||||||
func RunImport(args []string) int {
|
func RunImport(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats import", flag.ContinueOnError)
|
if err := parseNoFlags("chats-import", args); err != nil {
|
||||||
help := fs.Bool("help", false, "")
|
return cliparse.Fail(err)
|
||||||
fs.SetOutput(os.Stderr)
|
|
||||||
if err := fs.Parse(args); err != nil {
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
if *help {
|
|
||||||
fmt.Fprintln(os.Stderr, "usage: chats import")
|
|
||||||
return 0
|
|
||||||
}
|
}
|
||||||
|
|
||||||
root := Dir()
|
root := Dir()
|
||||||
|
|||||||
@@ -2,12 +2,13 @@ package chats
|
|||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
"flag"
|
|
||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"os/exec"
|
"os/exec"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
)
|
)
|
||||||
|
|
||||||
func checkLinkedInSession(userDataDir string) (bool, error) {
|
func checkLinkedInSession(userDataDir string) (bool, error) {
|
||||||
@@ -29,18 +30,12 @@ func checkLinkedInSession(userDataDir string) (bool, error) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func RunSyncLinkedIn(args []string) int {
|
func RunSyncLinkedIn(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats sync linkedin", flag.ContinueOnError)
|
f, err := parseLinkedInFlags(args)
|
||||||
limit := fs.Int("limit", 0, "max messages per conversation (0 = all)")
|
if err != nil {
|
||||||
refresh := fs.Bool("refresh", false, "refresh session from live webtop browser before sync")
|
return cliparse.Fail(err)
|
||||||
help := fs.Bool("help", false, "")
|
|
||||||
fs.SetOutput(os.Stderr)
|
|
||||||
if err := fs.Parse(args); err != nil {
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
if *help {
|
|
||||||
fmt.Fprintln(os.Stderr, "usage: chats sync linkedin [--limit N] [--refresh]")
|
|
||||||
return 0
|
|
||||||
}
|
}
|
||||||
|
limit := f.Limit
|
||||||
|
refresh := f.Refresh
|
||||||
|
|
||||||
userDataDir := envVar("LINKEDIN_USER_DATA_DIR", "")
|
userDataDir := envVar("LINKEDIN_USER_DATA_DIR", "")
|
||||||
if userDataDir == "" {
|
if userDataDir == "" {
|
||||||
@@ -48,7 +43,7 @@ func RunSyncLinkedIn(args []string) int {
|
|||||||
userDataDir = home + "/.linkedin-mcp/profile"
|
userDataDir = home + "/.linkedin-mcp/profile"
|
||||||
}
|
}
|
||||||
|
|
||||||
if *refresh {
|
if refresh {
|
||||||
if code := refreshLinkedInSession(userDataDir); code != 0 {
|
if code := refreshLinkedInSession(userDataDir); code != 0 {
|
||||||
return code
|
return code
|
||||||
}
|
}
|
||||||
@@ -72,7 +67,7 @@ func RunSyncLinkedIn(args []string) int {
|
|||||||
defer cancel()
|
defer cancel()
|
||||||
|
|
||||||
start := time.Now()
|
start := time.Now()
|
||||||
if err := src.Sync(ctx, Dir(), *limit); err != nil {
|
if err := src.Sync(ctx, Dir(), limit); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "chats sync linkedin: %v\n", err)
|
fmt.Fprintf(os.Stderr, "chats sync linkedin: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -2,33 +2,28 @@ package chats
|
|||||||
|
|
||||||
import (
|
import (
|
||||||
"context"
|
"context"
|
||||||
"flag"
|
|
||||||
"fmt"
|
"fmt"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"strconv"
|
"strconv"
|
||||||
"strings"
|
"strings"
|
||||||
"time"
|
"time"
|
||||||
|
|
||||||
|
cliparse "github.com/eSlider/2dph/internal/cli"
|
||||||
)
|
)
|
||||||
|
|
||||||
func RunSyncTelegram(args []string) int {
|
func RunSyncTelegram(args []string) int {
|
||||||
fs := flag.NewFlagSet("chats sync telegram", flag.ContinueOnError)
|
f, err := parseTelegramFlags(args)
|
||||||
limit := fs.Int("limit", 0, "max messages per chat (0 = all)")
|
if err != nil {
|
||||||
phone := fs.String("phone", "", "phone number (default env TELEGRAM_PHONE)")
|
return cliparse.Fail(err)
|
||||||
help := fs.Bool("help", false, "")
|
|
||||||
fs.SetOutput(os.Stderr)
|
|
||||||
if err := fs.Parse(args); err != nil {
|
|
||||||
return 2
|
|
||||||
}
|
|
||||||
if *help {
|
|
||||||
fmt.Fprintln(os.Stderr, "usage: chats sync telegram [--limit N] [--phone PHONE]")
|
|
||||||
return 0
|
|
||||||
}
|
}
|
||||||
|
limit := f.Limit
|
||||||
|
phone := f.Phone
|
||||||
|
|
||||||
apiIDStr := envVar("TELEGRAM_API_ID", "")
|
apiIDStr := envVar("TELEGRAM_API_ID", "")
|
||||||
apiHash := envVar("TELEGRAM_API_HASH", "")
|
apiHash := envVar("TELEGRAM_API_HASH", "")
|
||||||
sessionStr := envVar("TELEGRAM_SESSION_STRING", "")
|
sessionStr := envVar("TELEGRAM_SESSION_STRING", "")
|
||||||
phoneNum := *phone
|
phoneNum := phone
|
||||||
if phoneNum == "" {
|
if phoneNum == "" {
|
||||||
phoneNum = envVar("TELEGRAM_PHONE", "")
|
phoneNum = envVar("TELEGRAM_PHONE", "")
|
||||||
}
|
}
|
||||||
@@ -78,7 +73,7 @@ func RunSyncTelegram(args []string) int {
|
|||||||
defer cancel()
|
defer cancel()
|
||||||
|
|
||||||
start := time.Now()
|
start := time.Now()
|
||||||
if err := src.Sync(ctx, Dir(), *limit); err != nil {
|
if err := src.Sync(ctx, Dir(), limit); err != nil {
|
||||||
fmt.Fprintf(os.Stderr, "chats sync telegram: %v\n", err)
|
fmt.Fprintf(os.Stderr, "chats sync telegram: %v\n", err)
|
||||||
return 1
|
return 1
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,190 @@
|
|||||||
|
// Package cli is the shared flaggy wrapper (D23).
|
||||||
|
//
|
||||||
|
// flaggy: zero deps, flags at any position, shell completion scripts.
|
||||||
|
// Individual tools keep ShowCompletion off so a query like "completion" is
|
||||||
|
// not stolen; dump scripts with bin/cli/complete.go.
|
||||||
|
package cli
|
||||||
|
|
||||||
|
import (
|
||||||
|
"errors"
|
||||||
|
"fmt"
|
||||||
|
"os"
|
||||||
|
"strings"
|
||||||
|
"sync"
|
||||||
|
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
// ErrHelp means -h/--help was requested (exit 0).
|
||||||
|
var ErrHelp = errors.New("help")
|
||||||
|
|
||||||
|
var parseMu sync.Mutex
|
||||||
|
|
||||||
|
// New returns a per-call parser. Never reuse: flaggy parses once.
|
||||||
|
func New(name string) *flaggy.Parser {
|
||||||
|
p := flaggy.NewParser(name)
|
||||||
|
p.ShowVersionWithVersionFlag = false
|
||||||
|
p.ShowCompletion = false
|
||||||
|
// Extra positionals become TrailingArguments (search "two words --json").
|
||||||
|
// Unknown dash tokens are rejected in Parse after flaggy returns.
|
||||||
|
p.ShowHelpOnUnexpected = false
|
||||||
|
p.ShowHelpWithHFlag = true
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
// Parse runs p.ParseArgs and turns flaggy's os.Exit into an error.
|
||||||
|
// Not safe to call in parallel (flaggy.PanicInsteadOfExit is process-global).
|
||||||
|
func Parse(p *flaggy.Parser, args []string) error {
|
||||||
|
parseMu.Lock()
|
||||||
|
defer parseMu.Unlock()
|
||||||
|
prev := flaggy.PanicInsteadOfExit
|
||||||
|
flaggy.PanicInsteadOfExit = true
|
||||||
|
defer func() { flaggy.PanicInsteadOfExit = prev }()
|
||||||
|
|
||||||
|
var exitMsg string
|
||||||
|
err := func() error {
|
||||||
|
defer func() {
|
||||||
|
if r := recover(); r != nil {
|
||||||
|
exitMsg = fmt.Sprint(r)
|
||||||
|
}
|
||||||
|
}()
|
||||||
|
return p.ParseArgs(args)
|
||||||
|
}()
|
||||||
|
if err != nil {
|
||||||
|
return err
|
||||||
|
}
|
||||||
|
if exitMsg != "" {
|
||||||
|
if strings.Contains(exitMsg, "code: 0") {
|
||||||
|
return ErrHelp
|
||||||
|
}
|
||||||
|
return errors.New(exitMsg)
|
||||||
|
}
|
||||||
|
if u := unknownFlags(p, args); len(u) > 0 {
|
||||||
|
return fmt.Errorf("unknown flag %q", u[0])
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func unknownFlags(p *flaggy.Parser, args []string) []string {
|
||||||
|
flags := collectFlags(&p.Subcommand)
|
||||||
|
var out []string
|
||||||
|
skipNext := false
|
||||||
|
for _, a := range args {
|
||||||
|
if skipNext {
|
||||||
|
skipNext = false
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if a == "--" {
|
||||||
|
break
|
||||||
|
}
|
||||||
|
name, inline := flagName(a)
|
||||||
|
if name == "" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if name == "h" || name == "help" {
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
f := findFlag(flags, name)
|
||||||
|
if f == nil {
|
||||||
|
out = append(out, a)
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
if !inline && !isBoolFlag(f) {
|
||||||
|
skipNext = true
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
func flagName(a string) (name string, inline bool) {
|
||||||
|
if a == "-" || !strings.HasPrefix(a, "-") {
|
||||||
|
return "", false
|
||||||
|
}
|
||||||
|
rest := strings.TrimLeft(a, "-")
|
||||||
|
name, _, inline = strings.Cut(rest, "=")
|
||||||
|
return name, inline
|
||||||
|
}
|
||||||
|
|
||||||
|
func collectFlags(sc *flaggy.Subcommand) []*flaggy.Flag {
|
||||||
|
out := append([]*flaggy.Flag{}, sc.Flags...)
|
||||||
|
for _, sub := range sc.Subcommands {
|
||||||
|
out = append(out, collectFlags(sub)...)
|
||||||
|
}
|
||||||
|
return out
|
||||||
|
}
|
||||||
|
|
||||||
|
func findFlag(flags []*flaggy.Flag, name string) *flaggy.Flag {
|
||||||
|
for _, f := range flags {
|
||||||
|
if f.HasName(name) {
|
||||||
|
return f
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func isBoolFlag(f *flaggy.Flag) bool {
|
||||||
|
_, ok := f.AssignmentVar.(*bool)
|
||||||
|
return ok
|
||||||
|
}
|
||||||
|
|
||||||
|
// Query joins the first positional with leftover trailing words.
|
||||||
|
func Query(first string, trailing []string) string {
|
||||||
|
parts := make([]string, 0, 1+len(trailing))
|
||||||
|
if s := strings.TrimSpace(first); s != "" {
|
||||||
|
parts = append(parts, s)
|
||||||
|
}
|
||||||
|
for _, t := range trailing {
|
||||||
|
if s := strings.TrimSpace(t); s != "" {
|
||||||
|
parts = append(parts, s)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return strings.Join(parts, " ")
|
||||||
|
}
|
||||||
|
|
||||||
|
// Code maps parse errors to process exit codes (0 help, 2 usage).
|
||||||
|
func Code(err error) int {
|
||||||
|
if err == nil || errors.Is(err, ErrHelp) {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
|
||||||
|
// Fail prints err unless it is help or a flaggy exit that already wrote stderr.
|
||||||
|
func Fail(err error) int {
|
||||||
|
if err == nil || errors.Is(err, ErrHelp) {
|
||||||
|
return 0
|
||||||
|
}
|
||||||
|
if strings.HasPrefix(err.Error(), "Panic instead of exit") {
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
fmt.Fprintln(os.Stderr, err)
|
||||||
|
return 2
|
||||||
|
}
|
||||||
|
|
||||||
|
// Tool is one shebang CLI for completion dump.
|
||||||
|
type Tool struct {
|
||||||
|
Path string
|
||||||
|
Name string
|
||||||
|
New func() *flaggy.Parser
|
||||||
|
}
|
||||||
|
|
||||||
|
// BashScript concatenates flaggy bash complete scripts and binds each
|
||||||
|
// function to the shebang path (./bin/subject/method.go).
|
||||||
|
func BashScript(tools []Tool) string {
|
||||||
|
var b strings.Builder
|
||||||
|
b.WriteString("# 2dph flaggy completions (D23). source <(./bin/cli/complete.go bash)\n")
|
||||||
|
for _, t := range tools {
|
||||||
|
p := t.New()
|
||||||
|
p.Name = t.Name
|
||||||
|
script := flaggy.GenerateBashCompletion(p)
|
||||||
|
b.WriteString(script)
|
||||||
|
fn := "_" + strings.ReplaceAll(t.Name, "-", "_") + "_complete"
|
||||||
|
if t.Path != "" && t.Path != t.Name {
|
||||||
|
fmt.Fprintf(&b, "complete -F %s %s\n", fn, t.Path)
|
||||||
|
if !strings.HasPrefix(t.Path, "./") {
|
||||||
|
fmt.Fprintf(&b, "complete -F %s ./%s\n", fn, t.Path)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return b.String()
|
||||||
|
}
|
||||||
@@ -0,0 +1,77 @@
|
|||||||
|
package cli
|
||||||
|
|
||||||
|
import (
|
||||||
|
"errors"
|
||||||
|
"strings"
|
||||||
|
"testing"
|
||||||
|
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
func TestParseBoolAndIntAnyPosition(t *testing.T) {
|
||||||
|
p := New("t")
|
||||||
|
jsonOut := false
|
||||||
|
n := 20
|
||||||
|
q := ""
|
||||||
|
p.Bool(&jsonOut, "", "json", "JSON")
|
||||||
|
p.Int(&n, "n", "n", "limit")
|
||||||
|
p.AddPositionalValue(&q, "query", 1, false, "q")
|
||||||
|
if err := Parse(p, []string{"two", "words", "--json", "-n", "5"}); err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
got := Query(q, p.TrailingArguments)
|
||||||
|
if got != "two words" || !jsonOut || n != 5 {
|
||||||
|
t.Fatalf("q=%q json=%v n=%d", got, jsonOut, n)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseUnknownFlagIsError(t *testing.T) {
|
||||||
|
p := New("t")
|
||||||
|
jsonOut := false
|
||||||
|
p.Bool(&jsonOut, "", "json", "JSON")
|
||||||
|
if err := Parse(p, []string{"--nope"}); err == nil {
|
||||||
|
t.Fatal("unknown flag accepted")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseHelpIsErrHelp(t *testing.T) {
|
||||||
|
p := New("t")
|
||||||
|
jsonOut := false
|
||||||
|
p.Bool(&jsonOut, "", "json", "JSON")
|
||||||
|
err := Parse(p, []string{"--help"})
|
||||||
|
if !errors.Is(err, ErrHelp) {
|
||||||
|
t.Fatalf("got %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestParseMissingFlagValueIsError(t *testing.T) {
|
||||||
|
p := New("t")
|
||||||
|
n := 0
|
||||||
|
p.Int(&n, "", "hop", "hop")
|
||||||
|
if err := Parse(p, []string{"--hop"}); err == nil {
|
||||||
|
t.Fatal("expected missing value error")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestBashScriptNamesShebangPath(t *testing.T) {
|
||||||
|
out := BashScript([]Tool{{
|
||||||
|
Path: "bin/brain/search.go",
|
||||||
|
Name: "brain-search",
|
||||||
|
New: newSearchLike,
|
||||||
|
}})
|
||||||
|
if !strings.Contains(out, "--json") || !strings.Contains(out, "--hop") {
|
||||||
|
t.Fatalf("flags missing:\n%s", out)
|
||||||
|
}
|
||||||
|
if !strings.Contains(out, "complete -F") || !strings.Contains(out, "bin/brain/search.go") {
|
||||||
|
t.Fatalf("shebang complete missing:\n%s", out)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func newSearchLike() *flaggy.Parser {
|
||||||
|
p := New("brain-search")
|
||||||
|
jsonOut := false
|
||||||
|
hop := 0
|
||||||
|
p.Bool(&jsonOut, "", "json", "JSON")
|
||||||
|
p.Int(&hop, "", "hop", "graph hop")
|
||||||
|
return p
|
||||||
|
}
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
package cli
|
||||||
|
|
||||||
|
import "github.com/integrii/flaggy"
|
||||||
|
|
||||||
|
type QAStats struct {
|
||||||
|
JSONL string
|
||||||
|
}
|
||||||
|
|
||||||
|
func QAParser() *flaggy.Parser {
|
||||||
|
c := QAStats{}
|
||||||
|
return BindQA(&c)
|
||||||
|
}
|
||||||
|
|
||||||
|
func BindQA(c *QAStats) *flaggy.Parser {
|
||||||
|
p := New("qa-stats")
|
||||||
|
p.Description = "DuckDB quantiles / JSONL count"
|
||||||
|
p.String(&c.JSONL, "", "jsonl", "JSONL file (else stdin JSON [float,…])")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseQAStats(args []string) (QAStats, error) {
|
||||||
|
var c QAStats
|
||||||
|
return c, Parse(BindQA(&c), args)
|
||||||
|
}
|
||||||
@@ -0,0 +1,120 @@
|
|||||||
|
// Package facts is cgo-free evidence rules (D16 contradictions).
|
||||||
|
package facts
|
||||||
|
|
||||||
|
import "strconv"
|
||||||
|
|
||||||
|
const (
|
||||||
|
ConfConfirmed = "confirmed"
|
||||||
|
ConfHypothesis = "hypothesis"
|
||||||
|
|
||||||
|
RuleUnresolved = "unresolved"
|
||||||
|
RuleTemporalFreshness = "temporal_freshness"
|
||||||
|
RuleAuthorityPairing = "authority_pairing"
|
||||||
|
RuleTwoSource = "two_source"
|
||||||
|
RuleSingleSource = "single_source"
|
||||||
|
|
||||||
|
KindRuntime = "runtime"
|
||||||
|
KindConfig = "config"
|
||||||
|
KindNarrative = "narrative"
|
||||||
|
)
|
||||||
|
|
||||||
|
// Source is one independent pointer on a yes or no side.
|
||||||
|
type Source struct {
|
||||||
|
ID string `json:"id"`
|
||||||
|
Kind string `json:"kind"`
|
||||||
|
When string `json:"when,omitempty"`
|
||||||
|
Stale bool `json:"stale,omitempty"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// Claim is one assertion with yes/no evidence lists.
|
||||||
|
type Claim struct {
|
||||||
|
Text string `json:"text"`
|
||||||
|
Yes []Source `json:"yes"`
|
||||||
|
No []Source `json:"no"`
|
||||||
|
}
|
||||||
|
|
||||||
|
// Result is audit output. Confirmed=false means `(not confirmed)`.
|
||||||
|
type Result struct {
|
||||||
|
Text string `json:"text"`
|
||||||
|
Confidence string `json:"confidence"`
|
||||||
|
Confirmed bool `json:"confirmed"`
|
||||||
|
Rule string `json:"rule"`
|
||||||
|
Winner string `json:"winner,omitempty"`
|
||||||
|
YesN int `json:"yes"`
|
||||||
|
NoN int `json:"no"`
|
||||||
|
}
|
||||||
|
|
||||||
|
func independent(ss []Source) int {
|
||||||
|
seen := map[string]struct{}{}
|
||||||
|
for i, s := range ss {
|
||||||
|
id := s.ID
|
||||||
|
if id == "" {
|
||||||
|
id = s.Kind + "#" + strconv.Itoa(i)
|
||||||
|
}
|
||||||
|
seen[id] = struct{}{}
|
||||||
|
}
|
||||||
|
return len(seen)
|
||||||
|
}
|
||||||
|
|
||||||
|
func freshN(ss []Source) int {
|
||||||
|
n := 0
|
||||||
|
for _, s := range ss {
|
||||||
|
if !s.Stale {
|
||||||
|
n++
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return n
|
||||||
|
}
|
||||||
|
|
||||||
|
func strongN(ss []Source) int {
|
||||||
|
n := 0
|
||||||
|
for _, s := range ss {
|
||||||
|
if s.Kind == KindRuntime || s.Kind == KindConfig {
|
||||||
|
n++
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return n
|
||||||
|
}
|
||||||
|
|
||||||
|
func out(c Claim, conf, rule, winner string) Result {
|
||||||
|
return Result{
|
||||||
|
Text: c.Text,
|
||||||
|
Confidence: conf,
|
||||||
|
Confirmed: conf == ConfConfirmed,
|
||||||
|
Rule: rule,
|
||||||
|
Winner: winner,
|
||||||
|
YesN: independent(c.Yes),
|
||||||
|
NoN: independent(c.No),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Adjudicate applies D16: ≥2 yes vs ≥2 no stays hypothesis until a rule fires.
|
||||||
|
// Order: temporal_freshness, then authority_pairing (A/B beats narrative C).
|
||||||
|
func Adjudicate(c Claim) Result {
|
||||||
|
yesN := independent(c.Yes)
|
||||||
|
noN := independent(c.No)
|
||||||
|
if yesN < 2 || noN < 2 {
|
||||||
|
if yesN >= 2 {
|
||||||
|
return out(c, ConfConfirmed, RuleTwoSource, "yes")
|
||||||
|
}
|
||||||
|
if noN >= 2 {
|
||||||
|
return out(c, ConfConfirmed, RuleTwoSource, "no")
|
||||||
|
}
|
||||||
|
return out(c, ConfHypothesis, RuleSingleSource, "")
|
||||||
|
}
|
||||||
|
yf, nf := freshN(c.Yes), freshN(c.No)
|
||||||
|
if yf >= 2 && nf < 2 {
|
||||||
|
return out(c, ConfConfirmed, RuleTemporalFreshness, "yes")
|
||||||
|
}
|
||||||
|
if nf >= 2 && yf < 2 {
|
||||||
|
return out(c, ConfConfirmed, RuleTemporalFreshness, "no")
|
||||||
|
}
|
||||||
|
ys, ns := strongN(c.Yes), strongN(c.No)
|
||||||
|
if ys >= 2 && ns < 2 {
|
||||||
|
return out(c, ConfConfirmed, RuleAuthorityPairing, "yes")
|
||||||
|
}
|
||||||
|
if ns >= 2 && ys < 2 {
|
||||||
|
return out(c, ConfConfirmed, RuleAuthorityPairing, "no")
|
||||||
|
}
|
||||||
|
return out(c, ConfHypothesis, RuleUnresolved, "")
|
||||||
|
}
|
||||||
@@ -0,0 +1,86 @@
|
|||||||
|
package facts
|
||||||
|
|
||||||
|
import "testing"
|
||||||
|
|
||||||
|
func src(id, kind string, stale bool) Source {
|
||||||
|
return Source{ID: id, Kind: kind, Stale: stale}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestTwoVsTwoStaysHypothesis(t *testing.T) {
|
||||||
|
c := Claim{
|
||||||
|
Text: "svc listens on 443",
|
||||||
|
Yes: []Source{
|
||||||
|
src("docker-ps", KindRuntime, false),
|
||||||
|
src("compose", KindConfig, false),
|
||||||
|
},
|
||||||
|
No: []Source{
|
||||||
|
src("docker-ps-old", KindRuntime, false),
|
||||||
|
src("compose-old", KindConfig, false),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
r := Adjudicate(c)
|
||||||
|
if r.Confirmed || r.Confidence != ConfHypothesis || r.Rule != RuleUnresolved {
|
||||||
|
t.Fatalf("2v2 must stay (not confirmed): %+v", r)
|
||||||
|
}
|
||||||
|
if r.Winner != "" {
|
||||||
|
t.Fatalf("unresolved must not name a winner: %+v", r)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestTemporalFreshnessResolvesStaleSide(t *testing.T) {
|
||||||
|
c := Claim{
|
||||||
|
Text: "svc listens on 443",
|
||||||
|
Yes: []Source{
|
||||||
|
src("docker-ps", KindRuntime, false),
|
||||||
|
src("compose", KindConfig, false),
|
||||||
|
},
|
||||||
|
No: []Source{
|
||||||
|
src("old-readme", KindNarrative, true),
|
||||||
|
src("old-wiki", KindNarrative, true),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
r := Adjudicate(c)
|
||||||
|
if !r.Confirmed || r.Rule != RuleTemporalFreshness || r.Winner != "yes" {
|
||||||
|
t.Fatalf("fresh yes vs stale no: %+v", r)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestAuthorityPairingBeatsNarrative(t *testing.T) {
|
||||||
|
c := Claim{
|
||||||
|
Text: "svc listens on 443",
|
||||||
|
Yes: []Source{
|
||||||
|
src("docker-ps", KindRuntime, false),
|
||||||
|
src("compose", KindConfig, false),
|
||||||
|
},
|
||||||
|
No: []Source{
|
||||||
|
src("readme", KindNarrative, false),
|
||||||
|
src("wiki", KindNarrative, false),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
r := Adjudicate(c)
|
||||||
|
if !r.Confirmed || r.Rule != RuleAuthorityPairing || r.Winner != "yes" {
|
||||||
|
t.Fatalf("A×B vs C×C: %+v", r)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestTwoSourceYesIsConfirmed(t *testing.T) {
|
||||||
|
c := Claim{
|
||||||
|
Text: "arc-1 runs Matrix",
|
||||||
|
Yes: []Source{
|
||||||
|
src("compose", KindConfig, false),
|
||||||
|
src("docker-ps", KindRuntime, false),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
r := Adjudicate(c)
|
||||||
|
if !r.Confirmed || r.Rule != RuleTwoSource || r.Winner != "yes" {
|
||||||
|
t.Fatalf("%+v", r)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestSingleSourceIsHypothesis(t *testing.T) {
|
||||||
|
c := Claim{Text: "maybe", Yes: []Source{src("readme", KindNarrative, false)}}
|
||||||
|
r := Adjudicate(c)
|
||||||
|
if r.Confirmed || r.Rule != RuleSingleSource {
|
||||||
|
t.Fatalf("%+v", r)
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
package facts
|
||||||
|
|
||||||
|
// Interval of truth for a fact leaf (D24 / OQ5). Not D16 source staleness.
|
||||||
|
//
|
||||||
|
// Empty valid_from and valid_to means "always" (legacy leafs). Empty asOf
|
||||||
|
// means "do not filter". Dates compare as YYYY-MM-DD (lexicographic).
|
||||||
|
|
||||||
|
// NormalizeDay keeps the calendar day from ISO-8601 or bare dates.
|
||||||
|
func NormalizeDay(s string) string {
|
||||||
|
s = trimSpace(s)
|
||||||
|
if len(s) >= 10 && s[4] == '-' && s[7] == '-' {
|
||||||
|
return s[:10]
|
||||||
|
}
|
||||||
|
return s
|
||||||
|
}
|
||||||
|
|
||||||
|
func trimSpace(s string) string {
|
||||||
|
i, j := 0, len(s)
|
||||||
|
for i < j && (s[i] == ' ' || s[i] == '\t' || s[i] == '\n' || s[i] == '\r') {
|
||||||
|
i++
|
||||||
|
}
|
||||||
|
for j > i && (s[j-1] == ' ' || s[j-1] == '\t' || s[j-1] == '\n' || s[j-1] == '\r') {
|
||||||
|
j--
|
||||||
|
}
|
||||||
|
return s[i:j]
|
||||||
|
}
|
||||||
|
|
||||||
|
// ActiveAt reports whether a fact with [validFrom, validTo] holds at asOf.
|
||||||
|
// validTo empty = open-ended. Both ends inclusive.
|
||||||
|
func ActiveAt(validFrom, validTo, asOf string) bool {
|
||||||
|
asOf = NormalizeDay(asOf)
|
||||||
|
if asOf == "" {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
from := NormalizeDay(validFrom)
|
||||||
|
to := NormalizeDay(validTo)
|
||||||
|
if from != "" && asOf < from {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
if to != "" && asOf > to {
|
||||||
|
return false
|
||||||
|
}
|
||||||
|
return true
|
||||||
|
}
|
||||||
@@ -0,0 +1,73 @@
|
|||||||
|
package facts
|
||||||
|
|
||||||
|
import "testing"
|
||||||
|
|
||||||
|
func TestActiveAtOpenEnded(t *testing.T) {
|
||||||
|
// works at Y from 2025-07-16, no end
|
||||||
|
if !ActiveAt("2025-07-16", "", "2025-07-16") {
|
||||||
|
t.Fatal("inclusive valid_from")
|
||||||
|
}
|
||||||
|
if !ActiveAt("2025-07-16", "", "2026-01-01") {
|
||||||
|
t.Fatal("open-ended valid_to")
|
||||||
|
}
|
||||||
|
if ActiveAt("2025-07-16", "", "2025-07-15") {
|
||||||
|
t.Fatal("before valid_from must be inactive")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestActiveAtClosedInterval(t *testing.T) {
|
||||||
|
// works at X 2024-03-01 .. 2025-07-15
|
||||||
|
if !ActiveAt("2024-03-01", "2025-07-15", "2025-01-01") {
|
||||||
|
t.Fatal("mid interval")
|
||||||
|
}
|
||||||
|
if !ActiveAt("2024-03-01", "2025-07-15", "2024-03-01") {
|
||||||
|
t.Fatal("inclusive start")
|
||||||
|
}
|
||||||
|
if !ActiveAt("2024-03-01", "2025-07-15", "2025-07-15") {
|
||||||
|
t.Fatal("inclusive end")
|
||||||
|
}
|
||||||
|
if ActiveAt("2024-03-01", "2025-07-15", "2025-07-16") {
|
||||||
|
t.Fatal("day after end")
|
||||||
|
}
|
||||||
|
if ActiveAt("2024-03-01", "2025-07-15", "2024-02-28") {
|
||||||
|
t.Fatal("day before start")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestActiveAtEmptyIntervalAlwaysTrue(t *testing.T) {
|
||||||
|
// legacy leafs without intervals stay visible for any as-of
|
||||||
|
if !ActiveAt("", "", "2025-01-01") {
|
||||||
|
t.Fatal("empty interval must remain active")
|
||||||
|
}
|
||||||
|
if !ActiveAt("", "", "") {
|
||||||
|
t.Fatal("no as-of means all active")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestActiveAtEmptyAsOfKeepsAll(t *testing.T) {
|
||||||
|
if !ActiveAt("2099-01-01", "2099-12-31", "") {
|
||||||
|
t.Fatal("empty as-of must not filter")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestAsOfPickXNotY(t *testing.T) {
|
||||||
|
// Acceptance from #36: as of 2025-01-01 → X, not Y
|
||||||
|
xFrom, xTo := "2024-03-01", "2025-07-15"
|
||||||
|
yFrom, yTo := "2025-07-16", ""
|
||||||
|
asOf := "2025-01-01"
|
||||||
|
if !ActiveAt(xFrom, xTo, asOf) {
|
||||||
|
t.Fatal("X must be active as of 2025-01-01")
|
||||||
|
}
|
||||||
|
if ActiveAt(yFrom, yTo, asOf) {
|
||||||
|
t.Fatal("Y must be inactive as of 2025-01-01")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestNormalizeDayTrimsTime(t *testing.T) {
|
||||||
|
if NormalizeDay("2025-01-01T12:00:00Z") != "2025-01-01" {
|
||||||
|
t.Fatalf("got %q", NormalizeDay("2025-01-01T12:00:00Z"))
|
||||||
|
}
|
||||||
|
if NormalizeDay("2025-01-01") != "2025-01-01" {
|
||||||
|
t.Fatalf("got %q", NormalizeDay("2025-01-01"))
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
package gitlog
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
"time"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
type CLI struct {
|
||||||
|
Repo, Root, Since string
|
||||||
|
Limit int
|
||||||
|
JSONOut bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func Parser() *flaggy.Parser {
|
||||||
|
c := CLI{}
|
||||||
|
return Bind(&c)
|
||||||
|
}
|
||||||
|
|
||||||
|
func Bind(c *CLI) *flaggy.Parser {
|
||||||
|
p := cli.New("git-import")
|
||||||
|
p.Description = "go-git history → commit leafs"
|
||||||
|
p.Bool(&c.JSONOut, "", "json", "JSON output")
|
||||||
|
p.Int(&c.Limit, "", "limit", "max commits (0 = all)")
|
||||||
|
p.String(&c.Since, "", "since", "RFC3339 or YYYY-MM-DD")
|
||||||
|
p.String(&c.Root, "", "root", "scan dir for git repos")
|
||||||
|
p.AddPositionalValue(&c.Repo, "repo", 1, false, "git repo path")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseArgs(args []string) (CLI, error) {
|
||||||
|
var c CLI
|
||||||
|
if err := cli.Parse(Bind(&c), args); err != nil {
|
||||||
|
return c, err
|
||||||
|
}
|
||||||
|
return c, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseSince(s string) (time.Time, error) {
|
||||||
|
if s == "" {
|
||||||
|
return time.Time{}, nil
|
||||||
|
}
|
||||||
|
for _, layout := range []string{time.RFC3339, "2006-01-02"} {
|
||||||
|
if t, err := time.Parse(layout, s); err == nil {
|
||||||
|
return t, nil
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return time.Time{}, fmt.Errorf("cannot parse --since %q", s)
|
||||||
|
}
|
||||||
@@ -105,11 +105,18 @@ func (s *Server) mcpCall(r *http.Request, params json.RawMessage) (any, error) {
|
|||||||
if limit < 1 || limit > 100 {
|
if limit < 1 || limit > 100 {
|
||||||
return mcpText(`{"error":"n must be int 1..100"}`, true), nil
|
return mcpText(`{"error":"n must be int 1..100"}`, true), nil
|
||||||
}
|
}
|
||||||
|
asOf := ""
|
||||||
|
if raw, ok := p.Arguments["as_of"]; ok {
|
||||||
|
asOf = strings.TrimSpace(fmt.Sprint(raw))
|
||||||
|
if asOf == "<nil>" {
|
||||||
|
asOf = ""
|
||||||
|
}
|
||||||
|
}
|
||||||
if !s.tryAcquire(r) {
|
if !s.tryAcquire(r) {
|
||||||
return nil, fmt.Errorf("cancelled")
|
return nil, fmt.Errorf("cancelled")
|
||||||
}
|
}
|
||||||
defer s.release()
|
defer s.release()
|
||||||
body, err = s.api.Search(r.Context(), q, limit)
|
body, err = s.api.Search(r.Context(), q, limit, asOf)
|
||||||
case "get":
|
case "get":
|
||||||
id := strings.TrimSpace(fmt.Sprint(p.Arguments["id"]))
|
id := strings.TrimSpace(fmt.Sprint(p.Arguments["id"]))
|
||||||
if id == "" || id == "<nil>" {
|
if id == "" || id == "<nil>" {
|
||||||
|
|||||||
@@ -24,7 +24,7 @@ import (
|
|||||||
|
|
||||||
// API is the in-process brain surface. Production serve.go wires internal/brain.
|
// API is the in-process brain surface. Production serve.go wires internal/brain.
|
||||||
type API interface {
|
type API interface {
|
||||||
Search(ctx context.Context, query string, limit int) ([]byte, error)
|
Search(ctx context.Context, query string, limit int, asOf string) ([]byte, error)
|
||||||
Get(ctx context.Context, id string, body bool) ([]byte, error)
|
Get(ctx context.Context, id string, body bool) ([]byte, error)
|
||||||
Stats(ctx context.Context) ([]byte, error)
|
Stats(ctx context.Context) ([]byte, error)
|
||||||
Audit(ctx context.Context) ([]byte, error)
|
Audit(ctx context.Context) ([]byte, error)
|
||||||
@@ -85,11 +85,12 @@ func (s *Server) handleSearch(w http.ResponseWriter, r *http.Request) {
|
|||||||
}
|
}
|
||||||
limit = n
|
limit = n
|
||||||
}
|
}
|
||||||
|
asOf := strings.TrimSpace(r.URL.Query().Get("as_of"))
|
||||||
if !s.acquire(w, r) {
|
if !s.acquire(w, r) {
|
||||||
return
|
return
|
||||||
}
|
}
|
||||||
defer s.release()
|
defer s.release()
|
||||||
body, err := s.api.Search(r.Context(), q, limit)
|
body, err := s.api.Search(r.Context(), q, limit, asOf)
|
||||||
writeAPI(w, body, err)
|
writeAPI(w, body, err)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -187,13 +188,18 @@ type ExecSearcher struct {
|
|||||||
Timeout time.Duration
|
Timeout time.Duration
|
||||||
}
|
}
|
||||||
|
|
||||||
func (b ExecSearcher) Search(ctx context.Context, query string, limit int) ([]byte, error) {
|
func (b ExecSearcher) Search(ctx context.Context, query string, limit int, asOf string) ([]byte, error) {
|
||||||
if b.Timeout == 0 {
|
if b.Timeout == 0 {
|
||||||
b.Timeout = 60 * time.Second
|
b.Timeout = 60 * time.Second
|
||||||
}
|
}
|
||||||
ctx, cancel := context.WithTimeout(ctx, b.Timeout)
|
ctx, cancel := context.WithTimeout(ctx, b.Timeout)
|
||||||
defer cancel()
|
defer cancel()
|
||||||
cmd := exec.CommandContext(ctx, b.CmdPath, "--json", "-n", strconv.Itoa(limit), query)
|
args := []string{"--json", "-n", strconv.Itoa(limit)}
|
||||||
|
if asOf != "" {
|
||||||
|
args = append(args, "--as-of", asOf)
|
||||||
|
}
|
||||||
|
args = append(args, query)
|
||||||
|
cmd := exec.CommandContext(ctx, b.CmdPath, args...)
|
||||||
out, err := cmd.Output()
|
out, err := cmd.Output()
|
||||||
if err != nil {
|
if err != nil {
|
||||||
var exitErr *exec.ExitError
|
var exitErr *exec.ExitError
|
||||||
|
|||||||
@@ -21,10 +21,10 @@ type fakeSearcher struct {
|
|||||||
calls int
|
calls int
|
||||||
active atomic.Int32
|
active atomic.Int32
|
||||||
maxSeen atomic.Int32
|
maxSeen atomic.Int32
|
||||||
callback func(q string, limit int) ([]byte, error)
|
callback func(q string, limit int, asOf string) ([]byte, error)
|
||||||
}
|
}
|
||||||
|
|
||||||
func (f *fakeSearcher) Search(ctx context.Context, query string, limit int) ([]byte, error) {
|
func (f *fakeSearcher) Search(ctx context.Context, query string, limit int, asOf string) ([]byte, error) {
|
||||||
f.mu.Lock()
|
f.mu.Lock()
|
||||||
f.calls++
|
f.calls++
|
||||||
f.mu.Unlock()
|
f.mu.Unlock()
|
||||||
@@ -44,7 +44,7 @@ func (f *fakeSearcher) Search(ctx context.Context, query string, limit int) ([]b
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
if f.callback != nil {
|
if f.callback != nil {
|
||||||
return f.callback(query, limit)
|
return f.callback(query, limit, asOf)
|
||||||
}
|
}
|
||||||
return []byte(`{"query":"` + query + `","count":0,"results":[]}`), nil
|
return []byte(`{"query":"` + query + `","count":0,"results":[]}`), nil
|
||||||
}
|
}
|
||||||
@@ -109,7 +109,7 @@ func TestSearchMissingQuery(t *testing.T) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func TestSearchReturnsSearcherResult(t *testing.T) {
|
func TestSearchReturnsSearcherResult(t *testing.T) {
|
||||||
fs := &fakeSearcher{callback: func(q string, limit int) ([]byte, error) {
|
fs := &fakeSearcher{callback: func(q string, limit int, asOf string) ([]byte, error) {
|
||||||
return []byte(`{"query":"` + q + `","count":1,"results":[{"id":"x"}]}`), nil
|
return []byte(`{"query":"` + q + `","count":1,"results":[{"id":"x"}]}`), nil
|
||||||
}}
|
}}
|
||||||
h := NewServer(fs, 1)
|
h := NewServer(fs, 1)
|
||||||
@@ -168,7 +168,7 @@ func TestSearchRejectsBadLimit(t *testing.T) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
func TestGetLeaf(t *testing.T) {
|
func TestGetLeaf(t *testing.T) {
|
||||||
fs := &fakeSearcher{callback: func(q string, limit int) ([]byte, error) {
|
fs := &fakeSearcher{callback: func(q string, limit int, asOf string) ([]byte, error) {
|
||||||
return []byte(`{}`), nil
|
return []byte(`{}`), nil
|
||||||
}}
|
}}
|
||||||
h := NewServer(fs, 1)
|
h := NewServer(fs, 1)
|
||||||
|
|||||||
@@ -35,6 +35,7 @@ var Ops = []Op{
|
|||||||
Params: []Param{
|
Params: []Param{
|
||||||
{Name: "q", In: "query", Type: "string", Description: "search query", Required: true},
|
{Name: "q", In: "query", Type: "string", Description: "search query", Required: true},
|
||||||
{Name: "n", In: "query", Type: "integer", Description: "hit limit 1..100 (default 10)"},
|
{Name: "n", In: "query", Type: "integer", Description: "hit limit 1..100 (default 10)"},
|
||||||
|
{Name: "as_of", In: "query", Type: "string", Description: "YYYY-MM-DD; keep facts active on that day (D24)"},
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -54,6 +55,8 @@ var Ops = []Op{
|
|||||||
{Name: "text", In: "query", Type: "string", Description: "leaf text (omit for CLI hint)"},
|
{Name: "text", In: "query", Type: "string", Description: "leaf text (omit for CLI hint)"},
|
||||||
{Name: "root", In: "query", Type: "string", Description: "facts or info (default info)"},
|
{Name: "root", In: "query", Type: "string", Description: "facts or info (default info)"},
|
||||||
{Name: "source", In: "query", Type: "string", Description: "evidence pointer; facts need two sources"},
|
{Name: "source", In: "query", Type: "string", Description: "evidence pointer; facts need two sources"},
|
||||||
|
{Name: "valid_from", In: "query", Type: "string", Description: "fact interval start YYYY-MM-DD (D24)"},
|
||||||
|
{Name: "valid_to", In: "query", Type: "string", Description: "fact interval end YYYY-MM-DD inclusive (D24)"},
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
{Path: PathOpenAPI, Method: "get", ID: "openapi", Summary: "OpenAPI 3 document for this server"},
|
{Path: PathOpenAPI, Method: "get", ID: "openapi", Summary: "OpenAPI 3 document for this server"},
|
||||||
|
|||||||
@@ -0,0 +1,41 @@
|
|||||||
|
package mdleaves
|
||||||
|
|
||||||
|
import (
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
type CLI struct {
|
||||||
|
Root string
|
||||||
|
Files string
|
||||||
|
JSONOut bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func Parser() *flaggy.Parser {
|
||||||
|
c := CLI{Root: "."}
|
||||||
|
return Bind(&c)
|
||||||
|
}
|
||||||
|
|
||||||
|
func Bind(c *CLI) *flaggy.Parser {
|
||||||
|
if c.Root == "" {
|
||||||
|
c.Root = "."
|
||||||
|
}
|
||||||
|
p := cli.New("markdown-import")
|
||||||
|
p.Description = "split markdown H2 leafs"
|
||||||
|
p.Bool(&c.JSONOut, "", "json", "JSON output")
|
||||||
|
p.String(&c.Files, "", "files", "comma-separated paths")
|
||||||
|
p.AddPositionalValue(&c.Root, "dir", 1, false, "markdown root")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseArgs(args []string) (CLI, error) {
|
||||||
|
c := CLI{Root: "."}
|
||||||
|
p := Bind(&c)
|
||||||
|
if err := cli.Parse(p, args); err != nil {
|
||||||
|
return c, err
|
||||||
|
}
|
||||||
|
if extra := cli.Query("", p.TrailingArguments); extra != "" && c.Root == "." {
|
||||||
|
c.Root = extra
|
||||||
|
}
|
||||||
|
return c, nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
package ocr
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
type CLI struct {
|
||||||
|
Path string
|
||||||
|
}
|
||||||
|
|
||||||
|
func Parser() *flaggy.Parser {
|
||||||
|
c := CLI{}
|
||||||
|
return Bind(&c)
|
||||||
|
}
|
||||||
|
|
||||||
|
func Bind(c *CLI) *flaggy.Parser {
|
||||||
|
p := cli.New("mail-ocr")
|
||||||
|
p.Description = "tesseract eng+deu on image or scanned PDF"
|
||||||
|
p.AddPositionalValue(&c.Path, "file", 1, false, "image or pdf")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseArgs(args []string) (CLI, error) {
|
||||||
|
var c CLI
|
||||||
|
if err := cli.Parse(Bind(&c), args); err != nil {
|
||||||
|
return c, err
|
||||||
|
}
|
||||||
|
if c.Path == "" {
|
||||||
|
return c, fmt.Errorf("usage: bin/mail/ocr.go <image|pdf>")
|
||||||
|
}
|
||||||
|
return c, nil
|
||||||
|
}
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
package reasoner
|
||||||
|
|
||||||
|
import (
|
||||||
|
"os"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
type CLI struct {
|
||||||
|
Base string
|
||||||
|
Model string
|
||||||
|
Device string
|
||||||
|
JSONOut bool
|
||||||
|
}
|
||||||
|
|
||||||
|
func Parser() *flaggy.Parser {
|
||||||
|
c := NewCLI()
|
||||||
|
return Bind(&c)
|
||||||
|
}
|
||||||
|
|
||||||
|
func NewCLI() CLI {
|
||||||
|
base := os.Getenv("REASONER_BASE_URL")
|
||||||
|
if base == "" {
|
||||||
|
base = "http://127.0.0.1:11435/v1"
|
||||||
|
}
|
||||||
|
model := os.Getenv("REASONER_MODEL")
|
||||||
|
if model == "" {
|
||||||
|
model = OllamaRAM
|
||||||
|
}
|
||||||
|
return CLI{Base: base, Model: model, Device: "cpu"}
|
||||||
|
}
|
||||||
|
|
||||||
|
func Bind(c *CLI) *flaggy.Parser {
|
||||||
|
p := cli.New("reasoner-bakeoff")
|
||||||
|
p.Description = "CPU tool-call bake-off"
|
||||||
|
p.Bool(&c.JSONOut, "", "json", "JSON output")
|
||||||
|
p.String(&c.Model, "", "model", "Ollama/HF model id")
|
||||||
|
p.String(&c.Base, "", "base-url", "OpenAI-compatible URL")
|
||||||
|
p.String(&c.Device, "", "device", "cpu")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseArgs(args []string) (CLI, error) {
|
||||||
|
c := NewCLI()
|
||||||
|
return c, cli.Parse(Bind(&c), args)
|
||||||
|
}
|
||||||
@@ -0,0 +1,63 @@
|
|||||||
|
package websearch
|
||||||
|
|
||||||
|
import (
|
||||||
|
"fmt"
|
||||||
|
|
||||||
|
"github.com/eSlider/2dph/internal/cli"
|
||||||
|
"github.com/integrii/flaggy"
|
||||||
|
)
|
||||||
|
|
||||||
|
type CLI struct {
|
||||||
|
Query, Site, Lang, Fresh, Category, Engines string
|
||||||
|
Limit int
|
||||||
|
JSONOut, Refresh, Force bool
|
||||||
|
TTL float64
|
||||||
|
Timeout int
|
||||||
|
}
|
||||||
|
|
||||||
|
func NewCLI() CLI {
|
||||||
|
return CLI{Limit: DefaultLimit, TTL: float64(CacheTTL), Timeout: 25}
|
||||||
|
}
|
||||||
|
|
||||||
|
func Parser() *flaggy.Parser {
|
||||||
|
c := NewCLI()
|
||||||
|
return Bind(&c)
|
||||||
|
}
|
||||||
|
|
||||||
|
func Bind(c *CLI) *flaggy.Parser {
|
||||||
|
p := cli.New("web-search")
|
||||||
|
p.Description = "SearXNG second source (throttled ≠ absence)"
|
||||||
|
p.Bool(&c.JSONOut, "", "json", "JSON output")
|
||||||
|
p.Bool(&c.Refresh, "", "refresh", "bypass cache")
|
||||||
|
p.Bool(&c.Force, "", "force", "allow PII in query")
|
||||||
|
p.Int(&c.Limit, "n", "limit", "max hits")
|
||||||
|
p.String(&c.Site, "", "site", "restrict to host")
|
||||||
|
p.String(&c.Lang, "", "lang", "language")
|
||||||
|
p.String(&c.Fresh, "", "fresh", "day|week|month|year")
|
||||||
|
p.String(&c.Category, "", "category", "searx category")
|
||||||
|
p.String(&c.Engines, "", "engines", "engine list")
|
||||||
|
p.Float64(&c.TTL, "", "ttl", "cache ttl seconds")
|
||||||
|
p.Int(&c.Timeout, "", "timeout", "http timeout seconds")
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
func ParseArgs(args []string) (CLI, error) {
|
||||||
|
c := NewCLI()
|
||||||
|
p := Bind(&c)
|
||||||
|
var q string
|
||||||
|
p.AddPositionalValue(&q, "query", 1, false, "search query")
|
||||||
|
if err := cli.Parse(p, args); err != nil {
|
||||||
|
return c, err
|
||||||
|
}
|
||||||
|
c.Query = cli.Query(q, p.TrailingArguments)
|
||||||
|
if c.Query == "" {
|
||||||
|
return c, fmt.Errorf("query required")
|
||||||
|
}
|
||||||
|
if c.Limit < 0 {
|
||||||
|
return c, fmt.Errorf("--limit must be a non-negative integer")
|
||||||
|
}
|
||||||
|
if c.Timeout <= 0 {
|
||||||
|
return c, fmt.Errorf("--timeout must be a positive integer")
|
||||||
|
}
|
||||||
|
return c, nil
|
||||||
|
}
|
||||||
@@ -26,6 +26,7 @@ second independent source when local roots cannot confirm. An answer is
|
|||||||
bin/brain/search.go "Matrix federation" # pointers + snippets, YAML
|
bin/brain/search.go "Matrix federation" # pointers + snippets, YAML
|
||||||
bin/brain/search.go "onlyoffice postgres" --root facts # restrict to confirmed
|
bin/brain/search.go "onlyoffice postgres" --root facts # restrict to confirmed
|
||||||
bin/brain/search.go "where is cs-lexicon" --json | yq '.[].ref'
|
bin/brain/search.go "where is cs-lexicon" --json | yq '.[].ref'
|
||||||
|
bin/brain/search.go "who works where" --as-of 2025-01-01 # D24 intervals
|
||||||
bin/brain/add.go --text T --root facts --source "a.md x b.md"
|
bin/brain/add.go --text T --root facts --source "a.md x b.md"
|
||||||
bin/brain/get.go <id> --body # full chunk only when needed
|
bin/brain/get.go <id> --body # full chunk only when needed
|
||||||
bin/brain/stats.go # index health
|
bin/brain/stats.go # index health
|
||||||
@@ -34,6 +35,8 @@ bin/brain/eval.go # recall@5 >= 0.95 gate (
|
|||||||
|
|
||||||
`bin/kb/search` is a deprecated wrapper. `--hop N` walks
|
`bin/kb/search` is a deprecated wrapper. `--hop N` walks
|
||||||
`FROM_FILE` / `HAS_VERSION` / `AUTHORED` from each hit (1=File, 3=Person).
|
`FROM_FILE` / `HAS_VERSION` / `AUTHORED` from each hit (1=File, 3=Person).
|
||||||
|
`--as-of YYYY-MM-DD` keeps leafs whose `valid_from`/`valid_to` cover that day
|
||||||
|
(empty interval = always; not D16 source staleness).
|
||||||
|
|
||||||
## Rules
|
## Rules
|
||||||
|
|
||||||
@@ -45,6 +48,8 @@ bin/brain/eval.go # recall@5 >= 0.95 gate (
|
|||||||
are not evidence of absence. `--root facts|info` and `--no-web` skip the web.
|
are not evidence of absence. `--root facts|info` and `--no-web` skip the web.
|
||||||
- If recall looks wrong, run `bin/brain/eval.go`; it gates control questions and
|
- If recall looks wrong, run `bin/brain/eval.go`; it gates control questions and
|
||||||
should stay at or above 95% recall@5.
|
should stay at or above 95% recall@5.
|
||||||
|
- Contradictions (≥2 yes vs ≥2 no) stay `(not confirmed)` until
|
||||||
|
`bin/facts/audit contradict` fires `temporal_freshness` or `authority_pairing`.
|
||||||
- Agents: `GET /openapi.json` and `POST /mcp` on `bin/brain/serve.go` (same
|
- Agents: `GET /openapi.json` and `POST /mcp` on `bin/brain/serve.go` (same
|
||||||
handlers; tool names match paths `search`/`get`/`stats`/`audit`). Generated
|
handlers; tool names match paths `search`/`get`/`stats`/`audit`). Generated
|
||||||
list: [tools.md](tools.md).
|
list: [tools.md](tools.md).
|
||||||
|
|||||||
@@ -12,6 +12,12 @@ PicoClaw speaks MCP at `POST /mcp` on `bin/brain/serve.go`. Compose profile
|
|||||||
`picoclaw` runs the official `sipeed/picoclaw` gateway plus `brain-mcp`
|
`picoclaw` runs the official `sipeed/picoclaw` gateway plus `brain-mcp`
|
||||||
(see [docs/picoclaw.md](../../docs/picoclaw.md)).
|
(see [docs/picoclaw.md](../../docs/picoclaw.md)).
|
||||||
|
|
||||||
|
```bash
|
||||||
|
bin/stack/start-assistant # brain + qwen3.5:9b + gateway + picoclaw agent
|
||||||
|
bin/stack/status
|
||||||
|
bin/stack/stop
|
||||||
|
```
|
||||||
|
|
||||||
## Tool order (before a factual reply)
|
## Tool order (before a factual reply)
|
||||||
|
|
||||||
1. **`search`** — facts root first, then info. The `web` block is a second
|
1. **`search`** — facts root first, then info. The `web` block is a second
|
||||||
|
|||||||
Reference in New Issue
Block a user