Compare commits
10
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
33b5c1ea97 | ||
|
|
9a6da1a4f7 | ||
|
|
504d13ed08 | ||
|
|
efc0864d42 | ||
|
|
59caf5e560 | ||
|
|
e7803c5269 | ||
|
|
a08e7c49ad | ||
|
|
4310aa7002 | ||
|
|
1a7a962183 | ||
|
|
af8e3b053a |
@@ -34,3 +34,10 @@ ONLYOFFICE_PROJECT_ID=33
|
||||
# MINIO_BUCKET=office
|
||||
# MINIO_ACCESS_KEY=
|
||||
# MINIO_SECRET_KEY=
|
||||
|
||||
# oo search — direct Elasticsearch access for name + content search. ES lives
|
||||
# inside the OnlyOffice VM on localhost:9200; expose it with an SSH tunnel
|
||||
# (see docs/elasticsearch.md). ONLYOFFICE_ES_INDEX defaults to files_file.
|
||||
# ONLYOFFICE_ES_URL=http://127.0.0.1:9200
|
||||
# ONLYOFFICE_ES_INDEX=files_file
|
||||
# ONLYOFFICE_TENANT=
|
||||
|
||||
@@ -9,16 +9,17 @@ Canonical Go client for OnlyOffice Workspace (Projects + Calendar + CRM) and the
|
||||
- `request.go` — `Request`, `Query`, `Time`, `Token`, `MetaResponse`, `Permissions`.
|
||||
- `auth.go` — `Authenticate`, `AuthenticateContext`, `InvalidateToken`, `Auth`, token lifecycle.
|
||||
- `http.go` — transport + DRY response decoders (`ResponseArray`/`ResponseObject`/`postFormObject`/`putFormObject`/`deleteObject`).
|
||||
- `projects.go`, `tasks.go`, `users.go`, `calendar.go`, `crm.go`, `files.go`, `mails.go`, `invoices.go` — typed / untyped domain methods. **`files.go`** — CRM opportunity upload plus **project/task Documents**. **`mails.go`** — OnlyOffice Workspace Mail. **`invoices.go`** — CRM invoices, PDF regen/cleanup, status. Association rules: [`docs/crm-associations.md`](docs/crm-associations.md).
|
||||
- `projects.go`, `tasks.go`, `users.go`, `calendar.go`, `crm.go`, `files.go`, `files_webdav.go`, `files_stem.go`, `retry.go`, `mails.go`, `invoices.go` — typed / untyped domain methods. **`files.go`** — CRM opportunity upload plus **project/task Documents** (`UpdateFile`, `UploadToFolderReplacing`). **`files_webdav.go`** — Documents module by id (`ListDavFolder`, `MoveDavItems`/`CopyDavItems` with per-operation error surfacing, `ListFileOps`). **`retry.go`** — `DoRetry`: deterministic linear backoff (no jitter) on 429/502/503/504; every bulk tool routes API calls through it. **`mails.go`** — OnlyOffice Workspace Mail. **`invoices.go`** — CRM invoices, PDF regen/cleanup, status. Association rules: [`docs/crm-associations.md`](docs/crm-associations.md).
|
||||
- Pure stdlib + `google/go-querystring`; no UI, no dotenv.
|
||||
- **CLI — `cmd/oo/` as `package main`.** Cobra wrapper that loads `.env` via `godotenv` at startup. **Subject-based command tree** mirroring [`tea`](https://gitea.com/gitea/tea):
|
||||
- `main.go` — entry point (docstring lists the command tree).
|
||||
- `common.go` — `rootCmd`, `newOO`, `printTable`/`printObject`, `--output table|json` flag.
|
||||
- `calendar.go`, `projects.go`, `projects_files.go`, `tasks.go`, `tasks_files.go`, `users.go`, `contacts.go`, `opportunities.go`, `cases.go`, `crm_tasks.go`, `mails.go`, `invoices.go` — one file per subject (or per subject facet), each registers in `init()`.
|
||||
- `calendar.go`, `projects.go`, `projects_files.go`, `tasks.go`, `tasks_files.go`, `users.go`, `contacts.go`, `opportunities.go`, `cases.go`, `crm.go`, `crm_tasks.go`, `catalog.go`, `docs.go`, `dav.go`, `mails.go`, `invoices.go` — one file per subject (or per subject facet), each registers in `init()`. `dav.go` exposes the Documents module by id (`oo dav ls|move|copy|mkdir|rename-file|rename-folder|download|fileops`).
|
||||
- CLI-only deps (`spf13/cobra`, `joho/godotenv`) stay out of the library.
|
||||
- **TUI — `cmd/office/` as `package main`.** Bubble Tea three-pane browser (module tree, selectable list, markdown preview). Reuses `cmd/internal/bootstrap` for env/auth and the root `onlyoffice` library for all API calls. UI logic in `cmd/office/ui/`; preview/formatting in `cmd/office/preview/`; list loaders in `cmd/office/fetch/`.
|
||||
- **List table (`DataTable`)** — `cmd/office/ui/table*.go`. Column layout policies live in `cmd/office/model/table_layout.go` (`TableFlexLayoutFor`); cell rendering uses the bubbles/table inline pattern in `table_render.go` (`renderTableCell`, `padANSIWidth`). See `.cursor/skills/office-tui-table/SKILL.md` before changing center-pane tables.
|
||||
- **Shared bootstrap — `cmd/internal/bootstrap/`.** `LoadEnv()` + `NewClient(ctx)` extracted from `oo`; both binaries import it.
|
||||
- **Bulk Documents tools — `cmd/ooscan/`, `cmd/pdfamount/`, `cmd/kontoblatt/`, `cmd/kontolink/`.** Single-purpose binaries (folder index, PDF amounts, Kontoblatt summary/linking). Pace requests, route API calls through `DoRetry`; usage in README.
|
||||
- **Personal ops tooling** (disk inventory, dossier→CRM sync, SearXNG) lives in private [`eSlider/oo-workspace`](https://git.produktor.io/eSlider/oo-workspace) (`oow`), not in this public tree.
|
||||
|
||||
## Rules
|
||||
@@ -27,7 +28,7 @@ Canonical Go client for OnlyOffice Workspace (Projects + Calendar + CRM) and the
|
||||
- New endpoints go into the library first; CLI commands are thin wrappers.
|
||||
- Prefer `ResponseObject` / `postFormObject` / `putFormObject` / `deleteObject` over hand-rolled `json.Unmarshal(responseField(...))` blocks — they exist for DRY, use them.
|
||||
- Domain split is by file, **not** by subpackage. Don't introduce `internal/` or `pkg/*` subpackages inside the library — it flattens the `*Client` call surface for a reason.
|
||||
- CLI commands follow **subject → verb** structure (`oo <subject> <verb>`), never `oo <verb>-<subject>`. Add new commands to the existing subject file if one fits; create a new `cmd/oo/<subject>.go` for a genuinely new domain.
|
||||
- CLI commands follow **subject → verb** structure (`oo <subject> <verb>`), never `oo <verb>-<subject>`. Add new commands to the existing subject file if one fits; create a new `cmd/oo/<subject>.go` for a genuinely new domain. The subject→verb tree in `cmd/oo/main.go` and the README table are documentation — update them with the code.
|
||||
- **Documents for agents:** prefer Markdown in git; OnlyOffice UI is weak for `.md`/`.txt`. Use `oo docs put-md` (md→docx) and `oo docs put-txt` (txt→docx, preserves line breaks). All upload paths default to **upsert** by `stem|ext` (`--replace`, default true); `--no-replace` fails on conflict; `--allow-duplicate` opts into raw OO append. `oo projects files dedupe PROJECT_ID` reports/removes duplicate stem|ext copies (`--apply`, `--cross`; includes project root folder).
|
||||
- Every table output goes through `printTable(headers, rows)`; every single-object through `printObject(v)`. Do not `fmt.Println` rows ad-hoc or the `--output json` flag breaks for that command.
|
||||
- No secrets in the repo; use `.env` (gitignored). Commit `.env.example` only.
|
||||
|
||||
@@ -503,6 +503,29 @@ type Task struct {
|
||||
|---|---|
|
||||
| `GetUsers()` | List all users with profiles |
|
||||
|
||||
### Documents Files
|
||||
|
||||
| Method | Description |
|
||||
|---|---|
|
||||
| `ListDavFolder(ctx, id)` | List a Documents folder (`@root` for virtual sections) |
|
||||
| `ListDavSections(ctx)` | Virtual sections (Documents, Projects, …) |
|
||||
| `CreateDavFolder(ctx, parentID, title)` | Create a subfolder |
|
||||
| `RenameDavFolder(ctx, id, title)` / `RenameDavFile(ctx, id, title)` | Rename folder / file |
|
||||
| `DownloadFile(ctx, id, dst)` / `DownloadDavFile(ctx, id, w)` | Download file bytes |
|
||||
| `UploadDavFile(ctx, folderID, fileName, src)` | Upload from a reader |
|
||||
| `UploadToFolder(ctx, folderID, localPath)` | Upload a local file into a folder |
|
||||
| `UploadToFolderReplacing(ctx, folderID, localPath)` | Upsert by `stem\|ext`; returns replaced ids |
|
||||
| `UpdateFile(ctx, fileID, localPath)` | New version of an existing file (same id, no copy) |
|
||||
| `MoveDavItems(ctx, folderIDs, fileIDs, dest)` | Move (`resolveType=Skip`); per-operation errors surfaced, not silent nil |
|
||||
| `CopyDavItems(ctx, folderIDs, fileIDs, dest)` | Copy (`conflictResolveType=Skip`); errors surfaced |
|
||||
| `MoveFiles(ctx, destFolderID, fileIDs)` | Move with `resolveType=Skip` + `holdResult`; errors surfaced |
|
||||
| `ListFileOps(ctx)` | Active file operations (move/copy status polling) |
|
||||
| `FolderFiles(ctx, folderID)` | Flat file list of a folder (stem helpers) |
|
||||
| `DeleteFilesByStem(ctx, folderID, stem)` | Remove `stem\|ext` copies |
|
||||
| `DoRetry(ctx, policy, fn)` | Deterministic linear backoff (N·Base, no jitter) on 429/502/503/504 |
|
||||
| `DefaultRetryPolicy()` | 5 attempts, 1s·2s·3s·4s waits, 30s cap |
|
||||
| `Transient(err)` | True for retriable OnlyOffice answers |
|
||||
|
||||
### Helper Types
|
||||
|
||||
| Type | Description |
|
||||
@@ -622,6 +645,8 @@ oo docs convert ./note.docx # → note.md
|
||||
oo docs ocr ./scan.jpg --md ./scan.md # searchable PDF + markdown
|
||||
oo docs hocr ./scan.jpg --lang spa --md ./scan.hocr.md --yaml ./scan.yml
|
||||
oo docs put-md 7 ./OO-HONDA-7-INDEX.md --folder 490
|
||||
oo docs put-txt 7 ./notes.txt --folder 490
|
||||
oo docs put-xlsx 7 ./table.xlsx --folder 490
|
||||
oo docs as-md 2815 --to ./parte.md # download OO file as MD (OCR if needed)
|
||||
oo docs as-md 307 --hocr --lang spa # OO download via go-hocr structure
|
||||
oo projects files put-md 7 ./note.md # alias
|
||||
@@ -631,21 +656,84 @@ oo tasks files upload 208 ./notes.pdf
|
||||
oo tasks files detach 208 12345
|
||||
```
|
||||
|
||||
### Documents module (`oo dav`)
|
||||
|
||||
Direct access to the Documents module by folder/file id — the same calls that
|
||||
back `oo-webdav` and the project/task file commands. `move` sends
|
||||
`resolveType=Skip` + `holdResult=true`: without those params the legacy
|
||||
`fileops/move` endpoint answers 200 without moving anything, and the library
|
||||
surfaces such per-operation errors instead of a silent nil
|
||||
(`MoveDavItems` / `CopyDavItems` / `MoveFiles`).
|
||||
|
||||
```bash
|
||||
oo dav ls 659
|
||||
oo dav ls @root # virtual sections (Documents, Projects, …)
|
||||
oo dav mkdir 659 "2026 inbox"
|
||||
oo dav move 659 22881 22882 # DEST_FOLDER_ID FILE_ID…
|
||||
oo dav move 659 22881 --folders 670 # move folders along with files
|
||||
oo dav copy 659 22881
|
||||
oo dav rename-file 22881 invoice-v2.pdf
|
||||
oo dav rename-folder 671 o2-archive
|
||||
oo dav download 22881 --to ./copy.pdf # default path: ./<server title>
|
||||
oo dav fileops # active move/copy operations (status polling)
|
||||
```
|
||||
|
||||
### Search (`oo search`)
|
||||
|
||||
Full-text search over the Documents index. The REST endpoint
|
||||
`/api/2.0/files/@search/{query}` only searches file names in the database, so
|
||||
`oo search` talks to the OnlyOffice **Elasticsearch** directly (index
|
||||
`files_file`). Name search is default; `--content` also matches extracted
|
||||
document text (`document.attachment.content`, Office formats only).
|
||||
See [`docs/elasticsearch.md`](docs/elasticsearch.md) for the tunnel setup.
|
||||
|
||||
```bash
|
||||
oo search "Rechnung" # names only
|
||||
oo search "Mahngebühr" --content # names + document text
|
||||
oo search "Rechnung" --folder 649 --limit 50
|
||||
oo search "Rechnung" --json # shorthand for -o json
|
||||
```
|
||||
|
||||
Requires `ONLYOFFICE_ES_URL` (plus optional `ONLYOFFICE_ES_INDEX`,
|
||||
`ONLYOFFICE_TENANT`).
|
||||
|
||||
### Bulk tools (`cmd/`)
|
||||
|
||||
Small single-purpose binaries for bulk Documents work. All of them pace
|
||||
requests and retry transient OnlyOffice answers (429/502/503/504) with a
|
||||
deterministic linear backoff — no jitter, same waits on every run
|
||||
(see `DoRetry` below). Build with `go build ./cmd/<tool>`.
|
||||
|
||||
```bash
|
||||
ooscan 659 # recursive index → TSV: file_id, folder_id, path, title
|
||||
ooscan 659 666 > oo-index.tsv # several roots into one index
|
||||
pdfamount 671 # "Zu zahlender Betrag" per PDF → TSV: file_id, title, amount
|
||||
kontoblatt 3906 ./kontoblatt.xlsx # summary (Gegenkonto/Monat) uploaded next to source
|
||||
kontolink IN.xlsx oo-index.tsv OUT.xlsx [FILE_ID] [AMOUNTS_TSV]
|
||||
# kontolink writes DocEditor links into the Link column: Beleg → supplier+month
|
||||
# → amount+date (5th arg = pdfamount output); with FILE_ID it updates the
|
||||
# source file in place, else uploads an "(links)" copy next to it.
|
||||
```
|
||||
|
||||
| Subject | Verbs |
|
||||
|---|---|
|
||||
| `calendar` | `list`, `events`, `add`, `delete` |
|
||||
| `projects` | `list`, `get`, `milestones`, `create`, `update`, `delete`, **`files`** (`list`, `upload`, `download`, `rename`, `delete`) |
|
||||
| `projects` | `list`, `get`, `milestones`, `milestone-create`, `create`, `update`, `delete`, `contacts` (`add`, `remove`), `link-authors`, `link-git`, **`files`** (`list`, `upload`, `download`, `rename`, `delete`, `dedupe`, `as-md`, `put-md`, `put-txt`, `put-xlsx`) |
|
||||
| `tasks` | `list`, `get`, `create`, `update`, `delete`, `subtask add`, **`files`** (`list`, `upload`, `detach`) |
|
||||
| `users` | `list`, `self` (alias: `oo whoami`) |
|
||||
| `contacts` | `list`, `get`, `delete`, `info-add`, `merge`, `dedupe-info` |
|
||||
| `contacts` | `list`, `get`, `delete`, `info-add`, `merge`, `dedupe-info`, `tags`, `tag-add`, `tag-create`, `tag-remove` |
|
||||
| `persons` | `list`, `create`, `delete`, `dedupe` |
|
||||
| `companies` | `list`, `create`, `delete`, `dedupe`, `dedupe-persons` |
|
||||
| `opportunities` | `list`, `get`, `create`, `delete`, `stages`, `member-add`, `dedupe`, `dedupe-members`, `fix-titles` |
|
||||
| `opportunities` | `list`, `get`, `create`, `update`, `delete`, `stages`, `member-add`, `dedupe`, `dedupe-members`, `fix-titles` |
|
||||
| `invoices` | `list`, `get`, `create`, `update`, `pdf`, `pdf-cleanup`, `status`, `delete`, `items …` |
|
||||
| `crm` | `cleanup` |
|
||||
| `mails` | `accounts`, `folders`, `list`, `get`, `draft`, `attach`, `draft-invoice`, `delete` |
|
||||
| `mails` | `accounts`, `folders`, `list`, `get`, `download-attachment`, `draft`, `attach`, `draft-invoice`, `send`, `delete` |
|
||||
| `cases` | `list`, `create`, `delete`, `member-add` |
|
||||
| `crm-tasks` | `list`, `create`, `delete`, `categories` |
|
||||
| `crm-tasks` | `list`, `create`, `delete`, `categories`, `reassign-self` |
|
||||
| `docs` | `tools`, `convert`, `optimize`, `ocr`, `hocr`, `as-md`, `put-md`, `put-txt`, `put-xlsx` |
|
||||
| `catalog` | `match`, `merge`, `apply`, `scan-contacts`, `scan-projects`, `scan-thunderbird` |
|
||||
| `dav` | `ls`, `move`, `copy`, `mkdir`, `rename-file`, `rename-folder`, `download`, `fileops` |
|
||||
| `search` | `QUERY` (`--content`, `--folder ID`, `--limit N`, `--json`) |
|
||||
|
||||
The CLI reads only `.env` from the current working directory (godotenv is a
|
||||
CLI-only concern — the library itself never loads dotfiles).
|
||||
|
||||
@@ -0,0 +1,220 @@
|
||||
// Command kontoblatt builds a summary ("сводная таблица") of a Kontoblatt XLSX
|
||||
// (Datum, Gegenkonto, Buchungstext, Beleg, Soll, Haben, Bemerkung) and uploads
|
||||
// it back to the same OnlyOffice folder as the source file.
|
||||
//
|
||||
// Usage: kontoblatt <FILE_ID> <LOCAL_XLSX>
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
onlyoffice "github.com/eslider/go-onlyoffice"
|
||||
"github.com/xuri/excelize/v2"
|
||||
)
|
||||
|
||||
type agg struct {
|
||||
count int
|
||||
soll float64
|
||||
haben float64
|
||||
reFehlt int
|
||||
}
|
||||
|
||||
type rec struct {
|
||||
date, month, konto, text string
|
||||
soll, haben float64
|
||||
reFehlt bool
|
||||
}
|
||||
|
||||
var dateRe = regexp.MustCompile(`^\d{2}\.\d{2}\.\d{4}$`)
|
||||
|
||||
func parseAmount(s string) float64 {
|
||||
s = strings.TrimSpace(s)
|
||||
s = strings.ReplaceAll(s, "€", "")
|
||||
s = strings.ReplaceAll(s, " ", "")
|
||||
s = strings.ReplaceAll(s, ",", "") // German thousands separator
|
||||
s = strings.TrimSpace(s)
|
||||
if s == "" {
|
||||
return 0
|
||||
}
|
||||
v, err := strconv.ParseFloat(s, 64)
|
||||
if err != nil {
|
||||
return 0
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
func cell(row []string, i int) string {
|
||||
if i < len(row) {
|
||||
return strings.TrimSpace(row[i])
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
func main() {
|
||||
if len(os.Args) < 3 {
|
||||
fmt.Fprintln(os.Stderr, "usage: kontoblatt <FILE_ID> <LOCAL_XLSX>")
|
||||
os.Exit(2)
|
||||
}
|
||||
fileID, path := os.Args[1], os.Args[2]
|
||||
ctx := context.Background()
|
||||
|
||||
f, err := excelize.OpenFile(path)
|
||||
if err != nil {
|
||||
panic(err)
|
||||
}
|
||||
defer f.Close()
|
||||
|
||||
var recs []rec
|
||||
for _, sh := range f.GetSheetList() {
|
||||
rows, err := f.GetRows(sh)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
for _, r := range rows {
|
||||
d := cell(r, 0)
|
||||
if !dateRe.MatchString(d) {
|
||||
continue
|
||||
}
|
||||
text := cell(r, 2)
|
||||
recs = append(recs, rec{
|
||||
date: d,
|
||||
month: d[3:10],
|
||||
konto: cell(r, 1),
|
||||
text: text,
|
||||
soll: parseAmount(cell(r, 4)),
|
||||
haben: parseAmount(cell(r, 5)),
|
||||
reFehlt: strings.Contains(strings.ToUpper(text), "FEHLT"),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
byKonto := map[string]*agg{}
|
||||
byMonth := map[string]*agg{}
|
||||
getK := func(k string) *agg {
|
||||
if byKonto[k] == nil {
|
||||
byKonto[k] = &agg{}
|
||||
}
|
||||
return byKonto[k]
|
||||
}
|
||||
getM := func(k string) *agg {
|
||||
if byMonth[k] == nil {
|
||||
byMonth[k] = &agg{}
|
||||
}
|
||||
return byMonth[k]
|
||||
}
|
||||
var tot agg
|
||||
for _, r := range recs {
|
||||
k := getK(r.konto)
|
||||
k.count++
|
||||
k.soll += r.soll
|
||||
k.haben += r.haben
|
||||
if r.reFehlt {
|
||||
k.reFehlt++
|
||||
}
|
||||
m := getM(r.month)
|
||||
m.count++
|
||||
m.soll += r.soll
|
||||
m.haben += r.haben
|
||||
if r.reFehlt {
|
||||
m.reFehlt++
|
||||
}
|
||||
tot.count++
|
||||
tot.soll += r.soll
|
||||
tot.haben += r.haben
|
||||
if r.reFehlt {
|
||||
tot.reFehlt++
|
||||
}
|
||||
}
|
||||
|
||||
out := excelize.NewFile()
|
||||
defer out.Close()
|
||||
writeSheet(out, "Nach Gegenkonto", "Gegenkonto", byKonto, tot)
|
||||
writeSheet(out, "Nach Monat", "Monat", byMonth, tot)
|
||||
outPath := "/tmp/opencode/kontoblatt-zusammenfassung.xlsx"
|
||||
if err := out.SaveAs(outPath); err != nil {
|
||||
panic(err)
|
||||
}
|
||||
|
||||
// upload next to the source file
|
||||
creds := onlyoffice.GetEnvironmentCredentials()
|
||||
c := onlyoffice.NewClient(creds)
|
||||
var src *onlyoffice.FileEntry
|
||||
if derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
|
||||
var err error
|
||||
src, err = c.GetFile(ctx, fileID)
|
||||
return err
|
||||
}); derr != nil {
|
||||
panic(derr)
|
||||
}
|
||||
folder := ""
|
||||
if src.FolderID != nil {
|
||||
folder = src.FolderID.String()
|
||||
}
|
||||
title := ""
|
||||
if src.Title != nil {
|
||||
title = *src.Title
|
||||
}
|
||||
fmt.Printf("source: id=%s title=%q folder=%s\n", fileID, title, folder)
|
||||
|
||||
name := "Kontoblatt-1591-2025-Zusammenfassung.xlsx"
|
||||
tmp := "/tmp/opencode/" + name
|
||||
data, _ := os.ReadFile(outPath)
|
||||
if err := os.WriteFile(tmp, data, 0o600); err != nil {
|
||||
panic(err)
|
||||
}
|
||||
var entry *onlyoffice.FileEntry
|
||||
if derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
|
||||
var err error
|
||||
entry, _, err = c.UploadToFolderReplacing(ctx, folder, tmp)
|
||||
return err
|
||||
}); derr != nil {
|
||||
panic(derr)
|
||||
}
|
||||
fmt.Printf("uploaded: %s -> folder %s (id %v)\n", name, folder, entry.ID)
|
||||
|
||||
// print the summary
|
||||
printAgg("Nach Gegenkonto", byKonto, tot)
|
||||
printAgg("Nach Monat", byMonth, tot)
|
||||
}
|
||||
|
||||
func writeSheet(f *excelize.File, sheet, key string, m map[string]*agg, tot agg) {
|
||||
f.NewSheet(sheet)
|
||||
rows := [][]any{{key, "Anzahl", "Soll", "Haben", "Saldo", `davon "fehlt"`}}
|
||||
keys := make([]string, 0, len(m))
|
||||
for k := range m {
|
||||
keys = append(keys, k)
|
||||
}
|
||||
sort.Strings(keys)
|
||||
for _, k := range keys {
|
||||
a := m[k]
|
||||
rows = append(rows, []any{k, a.count, a.soll, a.haben, a.soll - a.haben, a.reFehlt})
|
||||
}
|
||||
rows = append(rows, []any{"GESAMT", tot.count, tot.soll, tot.haben, tot.soll - tot.haben, tot.reFehlt})
|
||||
for i, row := range rows {
|
||||
for j, v := range row {
|
||||
cellRef, _ := excelize.CoordinatesToCellName(j+1, i+1)
|
||||
_ = f.SetCellValue(sheet, cellRef, v)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func printAgg(title string, m map[string]*agg, tot agg) {
|
||||
fmt.Printf("\n== %s ==\n", title)
|
||||
keys := make([]string, 0, len(m))
|
||||
for k := range m {
|
||||
keys = append(keys, k)
|
||||
}
|
||||
sort.Strings(keys)
|
||||
fmt.Printf("%-12s %6s %12s %12s %12s %7s\n", "key", "count", "soll", "haben", "saldo", "fehlt")
|
||||
for _, k := range keys {
|
||||
a := m[k]
|
||||
fmt.Printf("%-12s %6d %12.2f %12.2f %12.2f %7d\n", k, a.count, a.soll, a.haben, a.soll-a.haben, a.reFehlt)
|
||||
}
|
||||
fmt.Printf("%-12s %6d %12.2f %12.2f %12.2f %7d\n", "GESAMT", tot.count, tot.soll, tot.haben, tot.soll-tot.haben, tot.reFehlt)
|
||||
}
|
||||
@@ -0,0 +1,478 @@
|
||||
// Command kontolink fills the "Link" column of a Kontoblatt ("ungeklärte
|
||||
// Posten") XLSX by matching each row to an OnlyOffice document.
|
||||
//
|
||||
// Strategy (deterministic, conservative — no LLM):
|
||||
// 1. Beleg token (letters/digits from the "Beleg" column) appears in the file
|
||||
// title; among candidates prefer (a) the row's month, (b) real invoices over
|
||||
// copies/dupes, and require the result to be unique;
|
||||
// 2. else supplier + row month + "rechnung", again unique.
|
||||
//
|
||||
// A file is linked at most once (rows already carrying a link are kept and their
|
||||
// file counts as used). Ambiguous rows are left UNLINKED for manual review.
|
||||
//
|
||||
// Usage: kontolink <IN_XLSX> <INDEX_TSV> <OUT_XLSX>
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
onlyoffice "github.com/eslider/go-onlyoffice"
|
||||
"github.com/xuri/excelize/v2"
|
||||
)
|
||||
|
||||
var (
|
||||
dateRe = regexp.MustCompile(`^\d{2}\.\d{2}\.\d{4}$`)
|
||||
nonAln = regexp.MustCompile(`[^0-9a-z]+`)
|
||||
fileID = regexp.MustCompile(`fileid=(\d+)`)
|
||||
)
|
||||
|
||||
func parseDay(s string) (time.Time, bool) {
|
||||
t, err := time.Parse("02.01.2006", strings.TrimSpace(s))
|
||||
return t, err == nil
|
||||
}
|
||||
|
||||
func titleDay(title string) (time.Time, bool) {
|
||||
if len(title) >= 10 {
|
||||
if t, err := time.Parse("2006-01-02", title[:10]); err == nil {
|
||||
return t, true
|
||||
}
|
||||
}
|
||||
return time.Time{}, false
|
||||
}
|
||||
|
||||
// nearest picks the candidate whose title date is closest to rd. Ties and
|
||||
// undated candidates (when >1) are rejected.
|
||||
func nearest(cands []entry, rd time.Time) (entry, bool) {
|
||||
if len(cands) == 1 {
|
||||
return cands[0], true
|
||||
}
|
||||
best, bestD, tie := -1, 0.0, false
|
||||
for i, e := range cands {
|
||||
td, ok := titleDay(e.title)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
d := td.Sub(rd).Hours() / 24
|
||||
if d < 0 {
|
||||
d = -d
|
||||
}
|
||||
if best < 0 || d < bestD {
|
||||
best, bestD, tie = i, d, false
|
||||
} else if d == bestD {
|
||||
tie = true
|
||||
}
|
||||
}
|
||||
if best < 0 || tie {
|
||||
return entry{}, false
|
||||
}
|
||||
return cands[best], true
|
||||
}
|
||||
|
||||
const linkPrefix = "https://office.pro-dukt.de/Products/Files/DocEditor.aspx?fileid="
|
||||
|
||||
type entry struct {
|
||||
id, path, title, norm string
|
||||
}
|
||||
|
||||
func norm(s string) string { return nonAln.ReplaceAllString(strings.ToLower(s), "") }
|
||||
|
||||
func main() {
|
||||
if len(os.Args) < 4 {
|
||||
fmt.Fprintln(os.Stderr, "usage: kontolink <IN_XLSX> <INDEX_TSV> <OUT_XLSX>")
|
||||
os.Exit(2)
|
||||
}
|
||||
in, idxPath, out := os.Args[1], os.Args[2], os.Args[3]
|
||||
|
||||
idxRaw, err := os.ReadFile(idxPath)
|
||||
if err != nil {
|
||||
panic(err)
|
||||
}
|
||||
var entries []entry
|
||||
for _, line := range strings.Split(string(idxRaw), "\n") {
|
||||
parts := strings.Split(line, "\t")
|
||||
if len(parts) < 4 || parts[0] == "" {
|
||||
continue
|
||||
}
|
||||
entries = append(entries, entry{id: parts[0], path: parts[2], title: parts[3], norm: norm(parts[3])})
|
||||
}
|
||||
|
||||
f, err := excelize.OpenFile(in)
|
||||
if err != nil {
|
||||
panic(err)
|
||||
}
|
||||
defer f.Close()
|
||||
sheet := f.GetSheetList()[0]
|
||||
rows, err := f.GetRows(sheet)
|
||||
if err != nil {
|
||||
panic(err)
|
||||
}
|
||||
|
||||
// optional 5th arg: amounts TSV "file_id\ttitle\tamount" (see cmd/pdfamount)
|
||||
var amts []amtEntry
|
||||
if len(os.Args) >= 6 && os.Args[5] != "" {
|
||||
amts = loadAmounts(os.Args[5])
|
||||
}
|
||||
|
||||
used := map[string]bool{}
|
||||
for _, r := range rows {
|
||||
if m := fileID.FindStringSubmatch(cell(r, 7)); m != nil {
|
||||
used[m[1]] = true
|
||||
}
|
||||
}
|
||||
|
||||
var linked, byBeleg, bySupplier, byAmount, unmatched, ambiguous int
|
||||
for i, r := range rows {
|
||||
if i == 0 || !dateRe.MatchString(cell(r, 0)) || strings.TrimSpace(cell(r, 7)) != "" {
|
||||
continue
|
||||
}
|
||||
beleg := norm(cell(r, 3))
|
||||
supplier := supplierNorm(cell(r, 2))
|
||||
month := monthYear(cell(r, 0))
|
||||
rd, _ := parseDay(cell(r, 0))
|
||||
|
||||
e, kind, ok := pick(entries, used, beleg, supplier, month, rd)
|
||||
if !ok {
|
||||
if ae, aok := amountPick(amts, used, supplier, rowAmount(r), rd); aok {
|
||||
e, kind, ok = entry{id: ae.id, title: ae.title}, "amount", true
|
||||
}
|
||||
}
|
||||
if !ok {
|
||||
if beleg != "" {
|
||||
ambiguous++
|
||||
} else {
|
||||
unmatched++
|
||||
}
|
||||
continue
|
||||
}
|
||||
ref, _ := excelize.CoordinatesToCellName(8, i+1)
|
||||
if err := f.SetCellValue(sheet, ref, linkPrefix+e.id); err != nil {
|
||||
panic(err)
|
||||
}
|
||||
used[e.id] = true
|
||||
linked++
|
||||
switch kind {
|
||||
case "beleg":
|
||||
byBeleg++
|
||||
case "supplier":
|
||||
bySupplier++
|
||||
case "amount":
|
||||
byAmount++
|
||||
}
|
||||
fmt.Printf("row %3d %-30s -> %s [%s]\n", i+1, cell(r, 2), e.title, kind)
|
||||
}
|
||||
|
||||
if err := f.SaveAs(out); err != nil {
|
||||
panic(err)
|
||||
}
|
||||
fmt.Printf("\nlinked=%d (beleg=%d, supplier=%d, amount=%d), ambiguous=%d, no-candidate=%d\n",
|
||||
linked, byBeleg, bySupplier, byAmount, ambiguous, unmatched)
|
||||
|
||||
// Optional 4th arg: source OnlyOffice file id. Try to update it in place;
|
||||
// if it is locked (OnlyOffice 500), upload a "(links)" copy next to it.
|
||||
if len(os.Args) >= 5 && os.Args[4] != "" {
|
||||
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 120*time.Second)
|
||||
defer cancel()
|
||||
var src *onlyoffice.FileEntry
|
||||
if derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
|
||||
var err error
|
||||
src, err = c.GetFile(ctx, os.Args[4])
|
||||
return err
|
||||
}); derr != nil {
|
||||
panic(derr)
|
||||
}
|
||||
folder, title := "", ""
|
||||
if src.FolderID != nil {
|
||||
folder = src.FolderID.String()
|
||||
}
|
||||
if src.Title != nil {
|
||||
title = *src.Title
|
||||
}
|
||||
uderr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
|
||||
_, err := c.UpdateFile(ctx, os.Args[4], out)
|
||||
return err
|
||||
})
|
||||
if uderr == nil {
|
||||
fmt.Printf("updated file %s in place\n", os.Args[4])
|
||||
return
|
||||
}
|
||||
fmt.Printf("in-place update failed (locked?); uploading a copy to folder %s\n", folder)
|
||||
ext := filepath.Ext(title)
|
||||
name := strings.TrimSuffix(title, ext) + " (links)" + ext
|
||||
tmp := filepath.Join(os.TempDir(), name)
|
||||
data, _ := os.ReadFile(out)
|
||||
if err := os.WriteFile(tmp, data, 0o600); err != nil {
|
||||
panic(err)
|
||||
}
|
||||
if derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
|
||||
_, _, err := c.UploadToFolderReplacing(ctx, folder, tmp)
|
||||
return err
|
||||
}); derr != nil {
|
||||
panic(derr)
|
||||
}
|
||||
fmt.Printf("uploaded copy: %s -> folder %s\n", name, folder)
|
||||
}
|
||||
}
|
||||
|
||||
func cell(r []string, i int) string {
|
||||
if i < len(r) {
|
||||
return strings.TrimSpace(r[i])
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
func supplierNorm(s string) string {
|
||||
s = strings.ToUpper(s)
|
||||
if i := strings.Index(s, ","); i >= 0 {
|
||||
s = s[:i]
|
||||
}
|
||||
for _, w := range []string{"RE FEHLT", "GS FEHLT", "WOFR", "WOFÜR"} {
|
||||
s = strings.ReplaceAll(s, w, "")
|
||||
}
|
||||
return norm(s)
|
||||
}
|
||||
|
||||
func monthYear(date string) string {
|
||||
if len(date) == 10 {
|
||||
return date[6:10] + "-" + date[3:5]
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// pick returns an unused candidate. Beleg match wins; supplier+month is a
|
||||
// fallback. When several candidates qualify, the one closest in time to the row
|
||||
// date wins; a tie is rejected (ambiguous) rather than guessed.
|
||||
func pick(entries []entry, used map[string]bool, beleg, supplier, month string, rd time.Time) (entry, string, bool) {
|
||||
free := func(e entry) bool { return !used[e.id] }
|
||||
|
||||
if len(beleg) >= 5 {
|
||||
var inMonth []entry
|
||||
for _, e := range entries {
|
||||
if free(e) && belegMatches(e.norm, beleg) &&
|
||||
(month == "" || strings.Contains(e.title, month)) {
|
||||
inMonth = append(inMonth, e)
|
||||
}
|
||||
}
|
||||
if supplier != "" {
|
||||
var s []entry
|
||||
for _, e := range inMonth {
|
||||
if strings.Contains(e.norm, supplier) {
|
||||
s = append(s, e)
|
||||
}
|
||||
}
|
||||
if len(s) > 0 {
|
||||
inMonth = s
|
||||
}
|
||||
}
|
||||
inMonth = topRank(inMonth)
|
||||
if e, ok := nearest(inMonth, rd); ok {
|
||||
return e, "beleg", true
|
||||
}
|
||||
// A Beleg is present but no file carries it: do NOT fall back to a
|
||||
// supplier guess (that links the wrong invoice).
|
||||
return entry{}, "", false
|
||||
}
|
||||
|
||||
if supplier != "" && month != "" {
|
||||
var c []entry
|
||||
for _, e := range entries {
|
||||
if free(e) && strings.Contains(e.norm, supplier) &&
|
||||
strings.Contains(e.title, month) && strings.Contains(e.norm, "rechnung") {
|
||||
c = append(c, e)
|
||||
}
|
||||
}
|
||||
c = topRank(c)
|
||||
if e, ok := nearest(c, rd); ok {
|
||||
return e, "supplier", true
|
||||
}
|
||||
}
|
||||
return entry{}, "", false
|
||||
}
|
||||
|
||||
// topRank keeps only the highest-ranked candidates (real invoice over copy /
|
||||
// dupe / op), so a tie with a duplicate does not mask the real file.
|
||||
func topRank(cands []entry) []entry {
|
||||
if len(cands) < 2 {
|
||||
return cands
|
||||
}
|
||||
best := 0
|
||||
for _, e := range cands {
|
||||
if rank(e) > best {
|
||||
best = rank(e)
|
||||
}
|
||||
}
|
||||
out := cands[:0]
|
||||
for _, e := range cands {
|
||||
if rank(e) == best {
|
||||
out = append(out, e)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func rank(e entry) int {
|
||||
s := 0
|
||||
if strings.Contains(e.path, "/2025") || strings.Contains(e.path, "/2024") {
|
||||
s += 4
|
||||
}
|
||||
if strings.Contains(e.norm, "rechnung") {
|
||||
s += 2
|
||||
}
|
||||
if strings.Contains(e.norm, "dupe") || strings.Contains(e.norm, "copy") ||
|
||||
strings.Contains(e.norm, "op") {
|
||||
s--
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// belegMatches reports whether a Beleg identifies the file: the whole normalized
|
||||
// Beleg appears, or (for long numeric Belege, e.g. "24/641393110") an 8-digit
|
||||
// window of its longest digit run appears.
|
||||
func belegMatches(titleNorm, beleg string) bool {
|
||||
if strings.Contains(titleNorm, beleg) {
|
||||
return true
|
||||
}
|
||||
run := longestDigitRun(beleg)
|
||||
for i := 0; i+8 <= len(run); i++ {
|
||||
if strings.Contains(titleNorm, run[i:i+8]) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
func longestDigitRun(s string) string {
|
||||
var best, cur strings.Builder
|
||||
for _, r := range s {
|
||||
if r >= '0' && r <= '9' {
|
||||
cur.WriteRune(r)
|
||||
if cur.Len() > best.Len() {
|
||||
best.Reset()
|
||||
best.WriteString(cur.String())
|
||||
}
|
||||
} else {
|
||||
cur.Reset()
|
||||
}
|
||||
}
|
||||
return best.String()
|
||||
}
|
||||
|
||||
type amtEntry struct {
|
||||
id string
|
||||
title string
|
||||
norm string
|
||||
amount float64
|
||||
date time.Time
|
||||
hasDate bool
|
||||
}
|
||||
|
||||
func loadAmounts(path string) []amtEntry {
|
||||
raw, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil
|
||||
}
|
||||
var out []amtEntry
|
||||
for _, line := range strings.Split(string(raw), "\n") {
|
||||
p := strings.Split(line, "\t")
|
||||
if len(p) < 3 {
|
||||
continue
|
||||
}
|
||||
v, err := strconv.ParseFloat(strings.TrimSpace(p[2]), 64)
|
||||
if err != nil {
|
||||
continue
|
||||
}
|
||||
e := amtEntry{id: p[0], title: p[1], norm: norm(p[1]), amount: v}
|
||||
if len(p[1]) >= 10 {
|
||||
if t, err := time.Parse("2006-01-02", p[1][:10]); err == nil {
|
||||
e.date, e.hasDate = t, true
|
||||
}
|
||||
}
|
||||
out = append(out, e)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func rowAmount(r []string) float64 {
|
||||
if v := parseAmount(cell(r, 4)); v != 0 {
|
||||
return v
|
||||
}
|
||||
return parseAmount(cell(r, 5))
|
||||
}
|
||||
|
||||
func parseAmount(s string) float64 {
|
||||
s = strings.ReplaceAll(s, "€", "")
|
||||
s = strings.ReplaceAll(s, " ", "")
|
||||
s = strings.ReplaceAll(s, ",", ".")
|
||||
if s == "" {
|
||||
return 0
|
||||
}
|
||||
v, err := strconv.ParseFloat(s, 64)
|
||||
if err != nil {
|
||||
return 0
|
||||
}
|
||||
return v
|
||||
}
|
||||
|
||||
// amountPick matches a row to an O2 invoice by amount + nearest date. Scoped to
|
||||
// Telefonica/O2 rows and O2 files, so it cannot cross-link other suppliers.
|
||||
func amountPick(amts []amtEntry, used map[string]bool, supplier string, amt float64, rd time.Time) (amtEntry, bool) {
|
||||
if amt <= 0 || len(amts) == 0 {
|
||||
return amtEntry{}, false
|
||||
}
|
||||
if !strings.Contains(supplier, "telefonica") && !strings.Contains(supplier, "o2") {
|
||||
return amtEntry{}, false
|
||||
}
|
||||
var cands []amtEntry
|
||||
for _, a := range amts {
|
||||
if used[a.id] || !strings.Contains(a.norm, "o2") {
|
||||
continue
|
||||
}
|
||||
d := a.amount - amt
|
||||
if d < 0 {
|
||||
d = -d
|
||||
}
|
||||
if d > 0.005 {
|
||||
continue
|
||||
}
|
||||
if a.hasDate && !rd.IsZero() {
|
||||
days := a.date.Sub(rd).Hours() / 24
|
||||
if days < 0 {
|
||||
days = -days
|
||||
}
|
||||
if days > 75 {
|
||||
continue
|
||||
}
|
||||
}
|
||||
cands = append(cands, a)
|
||||
}
|
||||
if len(cands) == 1 {
|
||||
return cands[0], true
|
||||
}
|
||||
best, bestD, tie := -1, 0.0, false
|
||||
for i, a := range cands {
|
||||
if !a.hasDate {
|
||||
continue
|
||||
}
|
||||
d := a.date.Sub(rd).Hours() / 24
|
||||
if d < 0 {
|
||||
d = -d
|
||||
}
|
||||
if best < 0 || d < bestD {
|
||||
best, bestD, tie = i, d, false
|
||||
} else if d == bestD {
|
||||
tie = true
|
||||
}
|
||||
}
|
||||
if best < 0 || tie {
|
||||
return amtEntry{}, false
|
||||
}
|
||||
return cands[best], true
|
||||
}
|
||||
+8
-6
@@ -3,20 +3,22 @@
|
||||
// Command tree is subject-based (mirrors the library split and the `tea` CLI):
|
||||
//
|
||||
// oo calendar list | events | add | delete
|
||||
// oo projects list | get | milestones | create | update | delete | files (list|upload|download|rename|delete|as-md|put-md)
|
||||
// oo projects list | get | milestones | milestone-create | create | update | delete | contacts (add|remove) | link-authors | link-git | files (list|upload|download|rename|delete|dedupe|as-md|put-md|put-txt|put-xlsx)
|
||||
// oo tasks list | get | create | update | delete | subtask add | files (list|upload|detach)
|
||||
// oo users list | self (alias: oo whoami)
|
||||
// oo contacts list | get | delete | info-add | merge | dedupe-info
|
||||
// oo contacts list | get | delete | info-add | merge | dedupe-info | tags | tag-add | tag-create | tag-remove
|
||||
// oo persons list | create | delete | dedupe
|
||||
// oo companies list | create | delete | dedupe | dedupe-persons
|
||||
// oo opportunities list | get | create | delete | stages | member-add | dedupe | dedupe-members | fix-titles
|
||||
// oo opportunities list | get | create | update | delete | stages | member-add | dedupe | dedupe-members | fix-titles
|
||||
// oo cases list | create | delete | member-add
|
||||
// oo crm-tasks list | create | delete | categories
|
||||
// oo crm-tasks list | create | delete | categories | reassign-self
|
||||
// oo crm cleanup
|
||||
// oo mails accounts | folders | list | get | download-attachment | draft | attach | draft-invoice | delete
|
||||
// oo mails accounts | folders | list | get | download-attachment | draft | attach | draft-invoice | send | delete
|
||||
// oo invoices list | get | create | update | pdf | pdf-cleanup | status | delete | items …
|
||||
// oo docs tools | convert | ocr | as-md | put-md
|
||||
// oo docs tools | convert | optimize | ocr | hocr | as-md | put-md | put-txt | put-xlsx
|
||||
// oo catalog match | merge | apply | scan-contacts | scan-projects | scan-thunderbird
|
||||
// oo dav ls | move | copy | mkdir | rename-file | rename-folder | download | fileops
|
||||
// oo search QUERY [--content] [--folder ID] [--limit N] [--json]
|
||||
//
|
||||
// CRM association rules: docs/crm-associations.md
|
||||
//
|
||||
|
||||
@@ -0,0 +1,72 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
onlyoffice "github.com/eslider/go-onlyoffice"
|
||||
"github.com/spf13/cobra"
|
||||
)
|
||||
|
||||
func init() {
|
||||
rootCmd.AddCommand(searchCmd())
|
||||
}
|
||||
|
||||
// searchCmd queries the OnlyOffice Elasticsearch index directly. The REST
|
||||
// /api/2.0/files/@search endpoint only searches file names in the database;
|
||||
// content search needs ES (see docs/elasticsearch.md).
|
||||
func searchCmd() *cobra.Command {
|
||||
var (
|
||||
content bool
|
||||
folder string
|
||||
limit int
|
||||
asJSON bool
|
||||
)
|
||||
cmd := &cobra.Command{
|
||||
Use: "search QUERY",
|
||||
Short: "Full-text search over documents by name, optionally by content (Elasticsearch)",
|
||||
Long: "Search the OnlyOffice Documents index.\n\n" +
|
||||
"By default only file names are matched. With --content the query also\n" +
|
||||
"matches extracted document text (document.attachment.content); this covers\n" +
|
||||
"Office formats (docx/xlsx/pptx) and is slower.\n\n" +
|
||||
"Requires ONLYOFFICE_ES_URL (and optionally ONLYOFFICE_ES_INDEX,\n" +
|
||||
"ONLYOFFICE_TENANT). See docs/elasticsearch.md for the tunnel setup.",
|
||||
Args: cobra.ExactArgs(1),
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
if asJSON {
|
||||
outputFormat = "json"
|
||||
}
|
||||
es, err := onlyoffice.NewESSearcher(onlyoffice.ESConfigFromEnv())
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
hits, err := es.Search(cmd.Context(), onlyoffice.SearchQuery{
|
||||
Text: args[0],
|
||||
InContent: content,
|
||||
FolderID: folder,
|
||||
Limit: limit,
|
||||
})
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
rows := make([]map[string]any, 0, len(hits))
|
||||
for _, h := range hits {
|
||||
rows = append(rows, map[string]any{
|
||||
"id": h.ID,
|
||||
"title": h.Title,
|
||||
"folder": h.ParentID,
|
||||
"score": h.Score,
|
||||
"highlight": h.Highlight,
|
||||
})
|
||||
}
|
||||
if outputFormat == "json" {
|
||||
printJSON(rows)
|
||||
return nil
|
||||
}
|
||||
printTable([]string{"id", "title", "folder", "score", "highlight"}, rows)
|
||||
return nil
|
||||
},
|
||||
}
|
||||
cmd.Flags().BoolVar(&content, "content", false, "also match extracted document content")
|
||||
cmd.Flags().StringVar(&folder, "folder", "", "limit to a Documents folder id")
|
||||
cmd.Flags().IntVar(&limit, "limit", 20, "maximum number of results")
|
||||
cmd.Flags().BoolVar(&asJSON, "json", false, "shorthand for --output json")
|
||||
return cmd
|
||||
}
|
||||
@@ -0,0 +1,46 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestSearchCommandRegisteredWithFlags(t *testing.T) {
|
||||
cmd, _, err := rootCmd.Find([]string{"search"})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if cmd.Name() != "search" {
|
||||
t.Fatalf("search resolved to %q", cmd.Name())
|
||||
}
|
||||
for _, name := range []string{"content", "folder", "limit", "json"} {
|
||||
if cmd.Flags().Lookup(name) == nil {
|
||||
t.Errorf("search: missing --%s flag", name)
|
||||
}
|
||||
}
|
||||
if got := cmd.Flags().Lookup("limit").DefValue; got != "20" {
|
||||
t.Errorf("--limit default = %q, want 20", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSearchWithoutESURLIsClearError(t *testing.T) {
|
||||
clearEnv(t, "ONLYOFFICE_ES_URL", "ONLYOFFICE_ES_INDEX", "ONLYOFFICE_TENANT")
|
||||
errBuf := &bytes.Buffer{}
|
||||
rootCmd.SetErr(errBuf)
|
||||
rootCmd.SetOut(&bytes.Buffer{})
|
||||
rootCmd.SetArgs([]string{"search", "Rechnung"})
|
||||
t.Cleanup(func() {
|
||||
rootCmd.SetArgs(nil)
|
||||
rootCmd.SetOut(nil)
|
||||
rootCmd.SetErr(nil)
|
||||
})
|
||||
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error without ONLYOFFICE_ES_URL")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "ONLYOFFICE_ES_URL") {
|
||||
t.Fatalf("error %q missing ONLYOFFICE_ES_URL", err.Error())
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,50 @@
|
||||
// Command ooscan recursively lists OnlyOffice Documents folders into a TSV
|
||||
// index: file_id, folder_id, path, title.
|
||||
//
|
||||
// Usage: ooscan <FOLDER_ID> [<FOLDER_ID>...]
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
"time"
|
||||
|
||||
onlyoffice "github.com/eslider/go-onlyoffice"
|
||||
)
|
||||
|
||||
func main() {
|
||||
ctx := context.Background()
|
||||
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
|
||||
seen := map[string]bool{}
|
||||
for _, root := range os.Args[1:] {
|
||||
walk(ctx, c, root, "", 0, seen)
|
||||
}
|
||||
}
|
||||
|
||||
func walk(ctx context.Context, c *onlyoffice.Client, folderID, path string, depth int, seen map[string]bool) {
|
||||
if depth > 8 || seen[folderID] {
|
||||
return
|
||||
}
|
||||
seen[folderID] = true
|
||||
// Throttle: OnlyOffice rate-limits (429) and the host must not be flooded.
|
||||
time.Sleep(350 * time.Millisecond)
|
||||
ctx, cancel := context.WithTimeout(ctx, 60*time.Second)
|
||||
defer cancel()
|
||||
var l *onlyoffice.DavListing
|
||||
derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
|
||||
var err error
|
||||
l, err = c.ListDavFolder(ctx, folderID)
|
||||
return err
|
||||
})
|
||||
if derr != nil {
|
||||
fmt.Fprintf(os.Stderr, "list %s (%s): %v\n", path, folderID, derr)
|
||||
return
|
||||
}
|
||||
for _, f := range l.Files {
|
||||
fmt.Printf("%s\t%s\t%s\t%s\n", f.ID, folderID, path, f.Title)
|
||||
}
|
||||
for _, sub := range l.Folders {
|
||||
walk(ctx, c, sub.ID, path+"/"+sub.Title, depth+1, seen)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,124 @@
|
||||
// Command pdfamount walks a Documents folder, downloads matching PDFs and
|
||||
// extracts the payable amount, printing "file_id\ttitle\tamount".
|
||||
//
|
||||
// Usage: pdfamount <FOLDER_ID> [TITLE_FILTER_REGEX]
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
"os/exec"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
onlyoffice "github.com/eslider/go-onlyoffice"
|
||||
)
|
||||
|
||||
var amountRe = regexp.MustCompile(`(?i)(zu zahlender betrag|rechnungsbetrag)\s*[:\s]*([0-9][0-9.]*,[0-9]{2})`)
|
||||
|
||||
func main() {
|
||||
if len(os.Args) < 2 {
|
||||
fmt.Fprintln(os.Stderr, "usage: pdfamount <FOLDER_ID> [TITLE_FILTER_REGEX]")
|
||||
os.Exit(2)
|
||||
}
|
||||
folder := os.Args[1]
|
||||
filter := regexp.MustCompile(`(?i)rechnung`)
|
||||
if len(os.Args) >= 3 {
|
||||
filter = regexp.MustCompile(os.Args[2])
|
||||
}
|
||||
ctx := context.Background()
|
||||
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
|
||||
|
||||
files := listAll(ctx, c, folder)
|
||||
for _, f := range files {
|
||||
if !filter.MatchString(f.title) {
|
||||
continue
|
||||
}
|
||||
if !strings.HasSuffix(strings.ToLower(f.title), ".pdf") {
|
||||
continue
|
||||
}
|
||||
amount, err := pdfAmount(ctx, c, f.id)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "%s: %v\n", f.title, err)
|
||||
continue
|
||||
}
|
||||
if amount == "" {
|
||||
continue
|
||||
}
|
||||
fmt.Printf("%s\t%s\t%s\n", f.id, f.title, amount)
|
||||
}
|
||||
}
|
||||
|
||||
type file struct{ id, title string }
|
||||
|
||||
func listAll(ctx context.Context, c *onlyoffice.Client, folder string) []file {
|
||||
seen := map[string]bool{}
|
||||
var out []file
|
||||
var walk func(string)
|
||||
walk = func(id string) {
|
||||
if seen[id] {
|
||||
return
|
||||
}
|
||||
seen[id] = true
|
||||
time.Sleep(300 * time.Millisecond)
|
||||
l, err := c.ListDavFolder(ctx, id)
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "list %s: %v\n", id, err)
|
||||
return
|
||||
}
|
||||
for _, f := range l.Files {
|
||||
out = append(out, file{f.ID, f.Title})
|
||||
}
|
||||
for _, sub := range l.Folders {
|
||||
walk(sub.ID)
|
||||
}
|
||||
}
|
||||
walk(folder)
|
||||
return out
|
||||
}
|
||||
|
||||
func pdfAmount(ctx context.Context, c *onlyoffice.Client, id string) (string, error) {
|
||||
time.Sleep(time.Second)
|
||||
tmp, err := os.CreateTemp("", "pdf-*.pdf")
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
defer os.Remove(tmp.Name())
|
||||
derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
|
||||
_ = tmp.Truncate(0)
|
||||
_, _ = tmp.Seek(0, 0)
|
||||
_, err := c.DownloadFile(ctx, id, tmp)
|
||||
return err
|
||||
})
|
||||
if derr != nil {
|
||||
tmp.Close()
|
||||
return "", derr
|
||||
}
|
||||
tmp.Close()
|
||||
var buf bytes.Buffer
|
||||
cmd := exec.CommandContext(ctx, "pdftotext", "-layout", tmp.Name(), "-")
|
||||
cmd.Stdout = &buf
|
||||
if err := cmd.Run(); err != nil {
|
||||
return "", err
|
||||
}
|
||||
m := amountRe.FindStringSubmatch(buf.String())
|
||||
if m == nil {
|
||||
return "", nil
|
||||
}
|
||||
return parseDe(m[2]), nil
|
||||
}
|
||||
|
||||
// parseDe turns "1.234,56" into 1234.56.
|
||||
func parseDe(s string) string {
|
||||
s = strings.ReplaceAll(s, ".", "")
|
||||
s = strings.ReplaceAll(s, ",", ".")
|
||||
v, err := strconv.ParseFloat(s, 64)
|
||||
if err != nil {
|
||||
return s
|
||||
}
|
||||
return strconv.FormatFloat(v, 'f', 2, 64)
|
||||
}
|
||||
@@ -0,0 +1,125 @@
|
||||
---
|
||||
type: reference
|
||||
status: current
|
||||
related:
|
||||
- README.md
|
||||
- file_es.go
|
||||
---
|
||||
|
||||
# Elasticsearch — полнотекстовый поиск OnlyOffice
|
||||
|
||||
## Что это
|
||||
|
||||
Полнотекстовый поиск OnlyOffice Workspace работает на **Elasticsearch**.
|
||||
Клиент на сервере — NEST. Индекс — имя таблицы.
|
||||
|
||||
Для файлов индекс `files_file`:
|
||||
|
||||
| поле | тип | смысл |
|
||||
|------|-----|-------|
|
||||
| `id` | integer | id файла (тот же, что в REST/Documents) |
|
||||
| `title` | text (`whitespacecustom`) | имя файла |
|
||||
| `tenantId` | integer | тенант (портал) |
|
||||
| `folders` | nested | список папок: `folderId` (строка), `id`, `tenantId` |
|
||||
| `document.attachment.content` | text (`document`) | извлеченный текст (ingest-attachment) |
|
||||
| `document.attachment.content_type` | text | MIME |
|
||||
|
||||
Важно:
|
||||
- Живой сервер — **Elasticsearch 7.16.3**, кластер `elasticsearch`.
|
||||
- REST `GET /api/2.0/files/@search/{query}` ищет **только по имени в БД**
|
||||
(`fileDao.Search`), ES не задействует. Для поиска по содержимому нужен
|
||||
прямой ES — это и делает `oo search`.
|
||||
- `title` analyzer `whitespacecustom` режет по пробелам и lower-case. Полное
|
||||
имя файла — один токен (`Rechnung-4711.pdf`), поэтому поиск по имени ищет
|
||||
слово целиком, а не подстроку.
|
||||
- `document.attachment.content` заполняется **только для Office-форматов**
|
||||
(docx / xlsx / pptx). У PDF/txt, залитых через API, контент не извлекается.
|
||||
- Индексация асинхронная (TeamLabSvc) — файл появляется в ES не мгновенно.
|
||||
|
||||
## Доступ
|
||||
|
||||
ES слушает `127.0.0.1:9200` **внутри** VM OnlyOffice. Снаружи порт закрыт,
|
||||
SSH в VM открыт на хосте как `127.0.0.1:32` (контейнер `onlyoffice-v2`,
|
||||
QEMU). Схема — SSH-туннель.
|
||||
|
||||
```bash
|
||||
# из корня go-onlyoffice (ключ и хост — как в infra-доках)
|
||||
ssh -f -N -o ControlMaster=no -o ControlPath=none \
|
||||
-p 32 -i ~/.ssh/id_ed25519 \
|
||||
-L 9200:127.0.0.1:9200 root@127.0.0.1
|
||||
|
||||
curl -s http://127.0.0.1:9200/ | head # tagline + version
|
||||
curl -s 'http://127.0.0.1:9200/_cat/indices?h=index,docs.count'
|
||||
```
|
||||
|
||||
`-o ControlMaster=no -o ControlPath=none` обязательны: иначе forward уходит
|
||||
в persistent master-соединение из `~/.ssh/config` и порт остаётся занят.
|
||||
|
||||
Проверить, что туннель жив:
|
||||
|
||||
```bash
|
||||
curl -s http://127.0.0.1:9200/files_file/_count
|
||||
```
|
||||
|
||||
## Переменные
|
||||
|
||||
| env | default | смысл |
|
||||
|-----|---------|-------|
|
||||
| `ONLYOFFICE_ES_URL` | — (обязателен) | `scheme://host:port` ES |
|
||||
| `ONLYOFFICE_ES_INDEX` | `files_file` | индекс |
|
||||
| `ONLYOFFICE_TENANT` | пусто (все) | фильтр `tenantId` |
|
||||
|
||||
Имена — в [`.env.example`](../.env.example). Секретов нет: ES без пароля.
|
||||
|
||||
## CLI
|
||||
|
||||
```bash
|
||||
ONLYOFFICE_ES_URL=http://127.0.0.1:9200 oo search "Rechnung"
|
||||
ONLYOFFICE_ES_URL=http://127.0.0.1:9200 oo search "Mahngebühr" --content
|
||||
oo search "Rechnung" --folder 649 --limit 50 --json
|
||||
```
|
||||
|
||||
Флаги: `--content` (искать и по тексту), `--folder ID` (папка
|
||||
`folders.folderId`), `--limit N` (по умолчанию 20, максимум 200),
|
||||
`--json` = `-o json`.
|
||||
|
||||
## Библиотека
|
||||
|
||||
`file_es.go` — `ESSearcher` (`Name() = "elasticsearch"`), прямой ES REST на
|
||||
stdlib `net/http`:
|
||||
|
||||
```go
|
||||
es, _ := onlyoffice.NewESSearcher(onlyoffice.ESConfigFromEnv())
|
||||
hits, _ := es.Search(ctx, onlyoffice.SearchQuery{
|
||||
Text: "Rechnung", InContent: true, Limit: 20,
|
||||
})
|
||||
```
|
||||
|
||||
Запрос: `multi_match` по `title^2` (+ `document.attachment.content` при
|
||||
`InContent`), фильтры `tenantId` и `folders.folderId`, `_source`
|
||||
id/title/folders, `highlight` для фрагмента. Ответ → `[]SearchHit` (модель из
|
||||
эпика #34; пока объявлена в `file_es.go`, переедет в `file_core.go` с F1 #35).
|
||||
|
||||
## Тесты
|
||||
|
||||
```bash
|
||||
# unit — чистые builders/парсеры, без сети
|
||||
go test ./ -run ES
|
||||
|
||||
# integration — нужен ONLYOFFICE_ES_URL (+ креды REST для залива)
|
||||
set -a; . .env; set +a
|
||||
ONLYOFFICE_ES_URL=http://127.0.0.1:9200 ONLYOFFICE_TENANT=1 \
|
||||
go test -tags=integration -run TestIntegrationESSearch -v .
|
||||
```
|
||||
|
||||
Интеграционный тест заливает временный xlsx (в имени и в ячейке — уникальные
|
||||
токены), ждёт индексации, проверяет поиск по имени и по содержимому, затем
|
||||
удаляет проект.
|
||||
|
||||
## Грабли
|
||||
|
||||
- `locale`/версия ES: 7.16.3, `_search` совместим с REST 7.x.
|
||||
- ES без auth и слушает только localhost — туннель обязателен.
|
||||
- Фильтр `tenantId` сузит выдачу; без него видны документы всех тенантов.
|
||||
- Поиск по содержимому PDF, залитых через API, не работает (нет
|
||||
`attachment.content`) — только Office-форматы.
|
||||
+329
@@ -0,0 +1,329 @@
|
||||
package onlyoffice
|
||||
|
||||
// Elasticsearch backend of the unified file client (epic #34, F3 #37).
|
||||
//
|
||||
// OnlyOffice full-text search runs on Elasticsearch (index `files_file`, NEST
|
||||
// client on the server). The REST endpoint GET /api/2.0/files/@search/{query}
|
||||
// only searches file names in the database, so content search needs a direct
|
||||
// ES query. The live server is Elasticsearch 7.16.3; the request shape below
|
||||
// is plain REST and stays stdlib-only, matching the repo's no-extra-deps rule.
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"net/http"
|
||||
"os"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Canonical file/search model (epic #34, F1 #35). Declared here because F1 is
|
||||
// not merged yet; move to file_core.go and drop these when it lands. Keep the
|
||||
// shape identical to the contract in #35.
|
||||
type (
|
||||
// Kind distinguishes a file from a folder.
|
||||
Kind int
|
||||
// Entry is a canonical file/folder record.
|
||||
Entry struct {
|
||||
ID string
|
||||
ParentID string
|
||||
Title string
|
||||
Kind Kind
|
||||
Size int64
|
||||
MIME string
|
||||
Created time.Time
|
||||
Modified time.Time
|
||||
Version int
|
||||
Provider string
|
||||
}
|
||||
// SearchQuery is a backend-agnostic search request.
|
||||
SearchQuery struct {
|
||||
Text string
|
||||
InContent bool
|
||||
FolderID string
|
||||
Extensions []string
|
||||
Limit int
|
||||
}
|
||||
// SearchHit is a search result entry plus its relevance data.
|
||||
SearchHit struct {
|
||||
Entry
|
||||
Score float64
|
||||
Highlight string
|
||||
Path []string
|
||||
}
|
||||
// Searcher searches a document store by name and optionally content.
|
||||
Searcher interface {
|
||||
Search(ctx context.Context, q SearchQuery) ([]SearchHit, error)
|
||||
Name() string
|
||||
}
|
||||
)
|
||||
|
||||
// Kind values (epic #34).
|
||||
const (
|
||||
File Kind = iota
|
||||
Folder
|
||||
)
|
||||
|
||||
const (
|
||||
defaultESIndex = "files_file"
|
||||
defaultESLimit = 20
|
||||
maxESLimit = 200
|
||||
maxESResponseSize = 8 << 20
|
||||
)
|
||||
|
||||
// ESConfig configures the direct Elasticsearch searcher.
|
||||
type ESConfig struct {
|
||||
URL string // scheme://host:port of the ES HTTP endpoint
|
||||
Index string // index name, default files_file
|
||||
Tenant string // tenantId filter, empty means all tenants
|
||||
}
|
||||
|
||||
// ESConfigFromEnv reads ONLYOFFICE_ES_URL, ONLYOFFICE_ES_INDEX (default
|
||||
// files_file) and ONLYOFFICE_TENANT. The library never loads dotfiles — the
|
||||
// CLI does that.
|
||||
func ESConfigFromEnv() ESConfig {
|
||||
return ESConfig{
|
||||
URL: strings.TrimRight(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL")), "/"),
|
||||
Index: firstNonEmpty(os.Getenv("ONLYOFFICE_ES_INDEX"), defaultESIndex),
|
||||
Tenant: strings.TrimSpace(os.Getenv("ONLYOFFICE_TENANT")),
|
||||
}
|
||||
}
|
||||
|
||||
// ESSearcher queries OnlyOffice's Elasticsearch index directly for file name
|
||||
// and document content.
|
||||
type ESSearcher struct {
|
||||
cfg ESConfig
|
||||
http *http.Client
|
||||
}
|
||||
|
||||
// NewESSearcher returns a searcher for the OnlyOffice Elasticsearch index.
|
||||
// The URL is required; an empty index falls back to files_file.
|
||||
func NewESSearcher(cfg ESConfig) (*ESSearcher, error) {
|
||||
if strings.TrimSpace(cfg.URL) == "" {
|
||||
return nil, fmt.Errorf("onlyoffice: elasticsearch URL is empty (set ONLYOFFICE_ES_URL)")
|
||||
}
|
||||
cfg.URL = strings.TrimRight(cfg.URL, "/")
|
||||
if cfg.Index == "" {
|
||||
cfg.Index = defaultESIndex
|
||||
}
|
||||
return &ESSearcher{cfg: cfg, http: &http.Client{Timeout: 30 * time.Second}}, nil
|
||||
}
|
||||
|
||||
// Name implements Searcher.
|
||||
func (s *ESSearcher) Name() string { return "elasticsearch" }
|
||||
|
||||
// Search runs a multi_match over title (and, when q.InContent is set,
|
||||
// document.attachment.content), filtered by tenant and optional folder.
|
||||
func (s *ESSearcher) Search(ctx context.Context, q SearchQuery) ([]SearchHit, error) {
|
||||
q.Text = strings.TrimSpace(q.Text)
|
||||
if q.Text == "" {
|
||||
return nil, fmt.Errorf("onlyoffice: empty search query")
|
||||
}
|
||||
body, err := json.Marshal(esSearchRequest(q, s.cfg.Tenant))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("onlyoffice: build elasticsearch query: %w", err)
|
||||
}
|
||||
endpoint := s.cfg.URL + "/" + s.cfg.Index + "/_search"
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(body))
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
req.Header.Set("Content-Type", "application/json")
|
||||
req.Header.Set("Accept", "application/json")
|
||||
resp, err := s.http.Do(req)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("onlyoffice: elasticsearch search: %w", err)
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
raw, err := io.ReadAll(io.LimitReader(resp.Body, maxESResponseSize))
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if resp.StatusCode >= 400 {
|
||||
return nil, fmt.Errorf("onlyoffice: elasticsearch search: %d %s", resp.StatusCode, truncate(string(raw), 400))
|
||||
}
|
||||
return parseESSearchResponse(raw)
|
||||
}
|
||||
|
||||
// esSearchRequest builds the ES query body. Pure, so it is unit-tested.
|
||||
func esSearchRequest(q SearchQuery, tenant string) esRequest {
|
||||
limit := q.Limit
|
||||
if limit <= 0 {
|
||||
limit = defaultESLimit
|
||||
}
|
||||
if limit > maxESLimit {
|
||||
limit = maxESLimit
|
||||
}
|
||||
fields := []string{"title^2"}
|
||||
if q.InContent {
|
||||
fields = append(fields, "document.attachment.content")
|
||||
}
|
||||
must := []esClause{{MultiMatch: &esMultiMatch{Query: q.Text, Fields: fields}}}
|
||||
|
||||
var filter []esClause
|
||||
if t := strings.TrimSpace(tenant); t != "" {
|
||||
filter = append(filter, esClause{Term: map[string]any{"tenantId": numericOrString(t)}})
|
||||
}
|
||||
if f := strings.TrimSpace(q.FolderID); f != "" {
|
||||
filter = append(filter, esClause{Term: map[string]any{"folders.folderId": f}})
|
||||
}
|
||||
for _, ext := range normalizeExtensions(q.Extensions) {
|
||||
filter = append(filter, esClause{Wildcard: map[string]any{"title": "*." + ext}})
|
||||
}
|
||||
|
||||
highlightFields := map[string]struct{}{"title": {}}
|
||||
if q.InContent {
|
||||
highlightFields["document.attachment.content"] = struct{}{}
|
||||
}
|
||||
return esRequest{
|
||||
Size: limit,
|
||||
Source: []string{"id", "title", "folders"},
|
||||
Query: esQuery{Bool: esBool{Must: must, Filter: filter}},
|
||||
Highlight: esHighlight{PreTags: []string{"<em>"}, PostTags: []string{"</em>"}, Fields: highlightFields},
|
||||
}
|
||||
}
|
||||
|
||||
// normalizeExtensions lowercases, trims leading dots and drops empties.
|
||||
func normalizeExtensions(exts []string) []string {
|
||||
out := make([]string, 0, len(exts))
|
||||
seen := map[string]bool{}
|
||||
for _, e := range exts {
|
||||
e = strings.ToLower(strings.TrimSpace(strings.TrimPrefix(strings.TrimSpace(e), ".")))
|
||||
if e == "" || seen[e] {
|
||||
continue
|
||||
}
|
||||
seen[e] = true
|
||||
out = append(out, e)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// numericOrString keeps an integer-looking filter value numeric (tenantId is
|
||||
// a long) and leaves anything else as a string (folderId is a text token).
|
||||
func numericOrString(s string) any {
|
||||
if n, err := strconv.ParseInt(s, 10, 64); err == nil {
|
||||
return n
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// esRequest is the subset of the ES query DSL this client emits.
|
||||
type esRequest struct {
|
||||
Size int `json:"size"`
|
||||
Source []string `json:"_source"`
|
||||
Query esQuery `json:"query"`
|
||||
Highlight esHighlight `json:"highlight"`
|
||||
}
|
||||
|
||||
type esQuery struct {
|
||||
Bool esBool `json:"bool"`
|
||||
}
|
||||
|
||||
type esBool struct {
|
||||
Must []esClause `json:"must,omitempty"`
|
||||
Filter []esClause `json:"filter,omitempty"`
|
||||
}
|
||||
|
||||
type esClause struct {
|
||||
MultiMatch *esMultiMatch `json:"multi_match,omitempty"`
|
||||
Term map[string]any `json:"term,omitempty"`
|
||||
Wildcard map[string]any `json:"wildcard,omitempty"`
|
||||
}
|
||||
|
||||
type esMultiMatch struct {
|
||||
Query string `json:"query"`
|
||||
Fields []string `json:"fields"`
|
||||
}
|
||||
|
||||
type esHighlight struct {
|
||||
PreTags []string `json:"pre_tags,omitempty"`
|
||||
PostTags []string `json:"post_tags,omitempty"`
|
||||
Fields map[string]struct{} `json:"fields"`
|
||||
}
|
||||
|
||||
// esResponse is the subset of an ES search response we consume.
|
||||
type esResponse struct {
|
||||
Took int `json:"took"`
|
||||
Hits struct {
|
||||
Total struct {
|
||||
Value int `json:"value"`
|
||||
Relation string `json:"relation"`
|
||||
} `json:"total"`
|
||||
Hits []esResponseHit `json:"hits"`
|
||||
} `json:"hits"`
|
||||
}
|
||||
|
||||
type esResponseHit struct {
|
||||
ID string `json:"_id"`
|
||||
Score float64 `json:"_score"`
|
||||
Source struct {
|
||||
ID int `json:"id"`
|
||||
Title string `json:"title"`
|
||||
Folders []struct {
|
||||
FolderID string `json:"folderId"`
|
||||
ID int `json:"id"`
|
||||
} `json:"folders"`
|
||||
} `json:"_source"`
|
||||
Highlight map[string][]string `json:"highlight"`
|
||||
}
|
||||
|
||||
// parseESSearchResponse converts an ES search response into SearchHit values.
|
||||
// Pure, so it is unit-tested.
|
||||
func parseESSearchResponse(raw []byte) ([]SearchHit, error) {
|
||||
var r esResponse
|
||||
if err := json.Unmarshal(raw, &r); err != nil {
|
||||
return nil, fmt.Errorf("onlyoffice: decode elasticsearch response: %w", err)
|
||||
}
|
||||
hits := make([]SearchHit, 0, len(r.Hits.Hits))
|
||||
for _, h := range r.Hits.Hits {
|
||||
id := strconv.Itoa(h.Source.ID)
|
||||
if h.Source.ID == 0 {
|
||||
id = h.ID
|
||||
}
|
||||
var parent string
|
||||
path := make([]string, 0, len(h.Source.Folders))
|
||||
for i, f := range h.Source.Folders {
|
||||
path = append(path, f.FolderID)
|
||||
if i == 0 {
|
||||
parent = f.FolderID
|
||||
}
|
||||
}
|
||||
hits = append(hits, SearchHit{
|
||||
Entry: Entry{
|
||||
ID: id,
|
||||
ParentID: parent,
|
||||
Title: h.Source.Title,
|
||||
Kind: File,
|
||||
Provider: "elasticsearch",
|
||||
},
|
||||
Score: h.Score,
|
||||
Highlight: esHighlightText(h.Highlight),
|
||||
Path: path,
|
||||
})
|
||||
}
|
||||
return hits, nil
|
||||
}
|
||||
|
||||
var esHighlightTag = regexp.MustCompile(`</?em[^>]*>`)
|
||||
|
||||
// esHighlightText flattens a highlight map into one plain-text snippet,
|
||||
// preferring the content fragment over the title.
|
||||
func esHighlightText(hl map[string][]string) string {
|
||||
for _, key := range []string{"document.attachment.content", "title"} {
|
||||
frags := hl[key]
|
||||
if len(frags) == 0 {
|
||||
continue
|
||||
}
|
||||
clean := make([]string, 0, len(frags))
|
||||
for _, f := range frags {
|
||||
clean = append(clean, esHighlightTag.ReplaceAllString(f, ""))
|
||||
}
|
||||
return strings.Join(clean, " … ")
|
||||
}
|
||||
return ""
|
||||
}
|
||||
@@ -0,0 +1,135 @@
|
||||
//go:build integration
|
||||
|
||||
package onlyoffice
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/xuri/excelize/v2"
|
||||
)
|
||||
|
||||
// TestIntegrationESSearch uploads a throwaway workbook and verifies that the
|
||||
// direct Elasticsearch search finds it by file name and by content.
|
||||
//
|
||||
// The content index (document.attachment.content) is only populated for Office
|
||||
// formats (docx/xlsx/pptx), so the fixture is an xlsx whose cell carries a
|
||||
// unique token. Requires ONLYOFFICE_ES_URL (a reachable ES endpoint — in the
|
||||
// current setup a tunnel to the ES inside the OnlyOffice VM, see
|
||||
// docs/elasticsearch.md) plus the regular REST credentials for the upload.
|
||||
// Skips when either is missing.
|
||||
func TestIntegrationESSearch(t *testing.T) {
|
||||
esURL := strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL"))
|
||||
if esURL == "" {
|
||||
t.Skip("ONLYOFFICE_ES_URL not set — skipping Elasticsearch integration test")
|
||||
}
|
||||
c := liveClient(t)
|
||||
t.Cleanup(func() { cleanupTestProjects(t, c) })
|
||||
|
||||
stamp := time.Now().UTC().Format("20060102-150405")
|
||||
nameToken := "goesname" + stamp
|
||||
contentToken := "goescontent" + stamp
|
||||
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
|
||||
defer cancel()
|
||||
|
||||
project, err := c.CreateProject(NewProjectRequest{
|
||||
Title: testProjectPrefix + "es-" + stamp,
|
||||
Description: "go-onlyoffice elasticsearch integration",
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("CreateProject: %v", err)
|
||||
}
|
||||
if project.ID == nil {
|
||||
t.Fatal("created project without id")
|
||||
}
|
||||
pid := strconv.Itoa(*project.ID)
|
||||
|
||||
title := nameToken + ".xlsx"
|
||||
localPath := filepath.Join(t.TempDir(), title)
|
||||
book := excelize.NewFile()
|
||||
if err := book.SetCellValue("Sheet1", "A1", "OnlyOffice Elasticsearch content fixture "+contentToken); err != nil {
|
||||
t.Fatalf("SetCellValue: %v", err)
|
||||
}
|
||||
if err := book.SaveAs(localPath); err != nil {
|
||||
t.Fatalf("SaveAs: %v", err)
|
||||
}
|
||||
|
||||
entry, err := c.UploadProjectFile(ctx, pid, localPath)
|
||||
if err != nil {
|
||||
t.Fatalf("UploadProjectFile: %v", err)
|
||||
}
|
||||
fileID := strconv.Itoa(int(FileEntryNumericID(entry)))
|
||||
if fileID == "0" {
|
||||
t.Fatalf("upload returned no file id: %+v", entry)
|
||||
}
|
||||
|
||||
es, err := NewESSearcher(ESConfig{
|
||||
URL: esURL,
|
||||
Index: os.Getenv("ONLYOFFICE_ES_INDEX"),
|
||||
Tenant: os.Getenv("ONLYOFFICE_TENANT"),
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("NewESSearcher: %v", err)
|
||||
}
|
||||
|
||||
// Indexing is asynchronous on the server; poll until the file shows up.
|
||||
// The server's title analyzer splits on whitespace, so the name query is
|
||||
// the full file name token (including extension), as a user would type it.
|
||||
nameHit := waitForHit(t, ctx, es, SearchQuery{Text: title}, fileID)
|
||||
if nameHit.Title != title {
|
||||
t.Errorf("name hit title = %q, want %q", nameHit.Title, title)
|
||||
}
|
||||
contentHit := waitForHit(t, ctx, es, SearchQuery{Text: contentToken, InContent: true}, fileID)
|
||||
if contentHit.Highlight == "" {
|
||||
t.Error("content hit has no highlight fragment")
|
||||
}
|
||||
if !strings.Contains(contentHit.Title, nameToken) {
|
||||
t.Errorf("content hit title = %q, want the uploaded workbook", contentHit.Title)
|
||||
}
|
||||
|
||||
// The content token is absent from the title, so a name-only search must
|
||||
// not return the file — this proves the content field is really queried.
|
||||
if hits := searchQuiet(t, es, SearchQuery{Text: contentToken}); len(hits) != 0 {
|
||||
t.Errorf("name-only search for content token returned %d hits, want 0", len(hits))
|
||||
}
|
||||
}
|
||||
|
||||
// waitForHit polls ES until the file with fileID appears and returns that hit.
|
||||
func waitForHit(t *testing.T, ctx context.Context, s *ESSearcher, q SearchQuery, fileID string) SearchHit {
|
||||
t.Helper()
|
||||
var lastErr error
|
||||
for {
|
||||
hits, err := s.Search(ctx, q)
|
||||
if err != nil {
|
||||
lastErr = err
|
||||
} else {
|
||||
for _, h := range hits {
|
||||
if h.ID == fileID {
|
||||
return h
|
||||
}
|
||||
}
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
t.Fatalf("search %q: file %s not indexed in time (last err: %v)", q.Text, fileID, lastErr)
|
||||
case <-time.After(3 * time.Second):
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func searchQuiet(t *testing.T, s *ESSearcher, q SearchQuery) []SearchHit {
|
||||
t.Helper()
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
|
||||
defer cancel()
|
||||
hits, err := s.Search(ctx, q)
|
||||
if err != nil {
|
||||
t.Fatalf("Search(%q): %v", q.Text, err)
|
||||
}
|
||||
return hits
|
||||
}
|
||||
+188
@@ -0,0 +1,188 @@
|
||||
package onlyoffice
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"reflect"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestESSearchRequestNameOnly(t *testing.T) {
|
||||
got := esSearchRequest(SearchQuery{Text: "Rechnung"}, "1")
|
||||
if got.Size != defaultESLimit {
|
||||
t.Errorf("size = %d, want %d", got.Size, defaultESLimit)
|
||||
}
|
||||
if !reflect.DeepEqual(got.Source, []string{"id", "title", "folders"}) {
|
||||
t.Errorf("_source = %v", got.Source)
|
||||
}
|
||||
if len(got.Query.Bool.Must) != 1 || got.Query.Bool.Must[0].MultiMatch == nil {
|
||||
t.Fatalf("must = %+v, want one multi_match", got.Query.Bool.Must)
|
||||
}
|
||||
mm := got.Query.Bool.Must[0].MultiMatch
|
||||
if mm.Query != "Rechnung" {
|
||||
t.Errorf("query = %q", mm.Query)
|
||||
}
|
||||
if !reflect.DeepEqual(mm.Fields, []string{"title^2"}) {
|
||||
t.Errorf("fields = %v, want title only", mm.Fields)
|
||||
}
|
||||
if _, ok := got.Highlight.Fields["document.attachment.content"]; ok {
|
||||
t.Error("content highlight present without InContent")
|
||||
}
|
||||
if _, ok := got.Highlight.Fields["title"]; !ok {
|
||||
t.Error("title highlight missing")
|
||||
}
|
||||
if len(got.Query.Bool.Filter) != 1 || got.Query.Bool.Filter[0].Term["tenantId"] != int64(1) {
|
||||
t.Errorf("tenant filter = %+v, want numeric tenantId=1", got.Query.Bool.Filter)
|
||||
}
|
||||
}
|
||||
|
||||
func TestESSearchRequestContentFields(t *testing.T) {
|
||||
got := esSearchRequest(SearchQuery{Text: "Mahnung", InContent: true}, "")
|
||||
mm := got.Query.Bool.Must[0].MultiMatch
|
||||
want := []string{"title^2", "document.attachment.content"}
|
||||
if !reflect.DeepEqual(mm.Fields, want) {
|
||||
t.Errorf("fields = %v, want %v", mm.Fields, want)
|
||||
}
|
||||
if _, ok := got.Highlight.Fields["document.attachment.content"]; !ok {
|
||||
t.Error("content highlight missing with InContent")
|
||||
}
|
||||
if len(got.Query.Bool.Filter) != 0 {
|
||||
t.Errorf("filter = %+v, want none without tenant/folder", got.Query.Bool.Filter)
|
||||
}
|
||||
}
|
||||
|
||||
func TestESSearchRequestFiltersAndLimit(t *testing.T) {
|
||||
got := esSearchRequest(SearchQuery{
|
||||
Text: "Storchen",
|
||||
FolderID: "649",
|
||||
Extensions: []string{".PDF", "pdf", "docx"},
|
||||
Limit: 999,
|
||||
}, "42")
|
||||
if got.Size != maxESLimit {
|
||||
t.Errorf("size = %d, want cap %d", got.Size, maxESLimit)
|
||||
}
|
||||
var tenant, folder, wildcards int
|
||||
for _, f := range got.Query.Bool.Filter {
|
||||
switch {
|
||||
case f.Term != nil && f.Term["tenantId"] != nil:
|
||||
tenant++
|
||||
case f.Term != nil && f.Term["folders.folderId"] != nil:
|
||||
folder++
|
||||
if f.Term["folders.folderId"] != "649" {
|
||||
t.Errorf("folder filter = %+v", f.Term)
|
||||
}
|
||||
case f.Wildcard != nil:
|
||||
wildcards++
|
||||
}
|
||||
}
|
||||
if tenant != 1 || folder != 1 {
|
||||
t.Errorf("term filters tenant=%d folder=%d, want 1 each", tenant, folder)
|
||||
}
|
||||
if wildcards != 2 {
|
||||
t.Errorf("wildcard filters = %d, want deduped PDF+docx", wildcards)
|
||||
}
|
||||
}
|
||||
|
||||
func TestESSearchRequestRejectsEmptyTextAtSearch(t *testing.T) {
|
||||
s, err := NewESSearcher(ESConfig{URL: "http://localhost:9200"})
|
||||
if err != nil {
|
||||
t.Fatalf("NewESSearcher: %v", err)
|
||||
}
|
||||
if _, err := s.Search(t.Context(), SearchQuery{Text: " "}); err == nil {
|
||||
t.Error("empty query: want error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestNewESSearcherRequiresURL(t *testing.T) {
|
||||
if _, err := NewESSearcher(ESConfig{}); err == nil {
|
||||
t.Error("empty URL: want error")
|
||||
}
|
||||
s, err := NewESSearcher(ESConfig{URL: "http://es:9200/"})
|
||||
if err != nil {
|
||||
t.Fatalf("NewESSearcher: %v", err)
|
||||
}
|
||||
if s.cfg.Index != defaultESIndex {
|
||||
t.Errorf("index = %q, want %q", s.cfg.Index, defaultESIndex)
|
||||
}
|
||||
if s.cfg.URL != "http://es:9200" {
|
||||
t.Errorf("url = %q, want trimmed", s.cfg.URL)
|
||||
}
|
||||
if s.Name() != "elasticsearch" {
|
||||
t.Errorf("Name() = %q", s.Name())
|
||||
}
|
||||
}
|
||||
|
||||
func TestNormalizeExtensions(t *testing.T) {
|
||||
got := normalizeExtensions([]string{" .PDF ", "pdf", "", "xlsx"})
|
||||
want := []string{"pdf", "xlsx"}
|
||||
if !reflect.DeepEqual(got, want) {
|
||||
t.Errorf("normalizeExtensions = %v, want %v", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseESSearchResponse(t *testing.T) {
|
||||
raw := []byte(`{
|
||||
"took": 12,
|
||||
"hits": {
|
||||
"total": {"value": 2, "relation": "eq"},
|
||||
"hits": [
|
||||
{
|
||||
"_id": "2395",
|
||||
"_score": 7.31,
|
||||
"_source": {"id": 2395, "title": "Rechnung-4711.pdf",
|
||||
"folders": [{"folderId": "438", "id": 0}, {"folderId": "11", "id": 0}]},
|
||||
"highlight": {
|
||||
"title": ["<em>Rechnung</em>-4711.pdf"],
|
||||
"document.attachment.content": ["… Zahlung der <em>Rechnung</em> …"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"_id": "2318",
|
||||
"_score": 6.02,
|
||||
"_source": {"id": 2318, "title": "Mahnung.pdf", "folders": []},
|
||||
"highlight": {"title": ["<em>Mahnung</em>.pdf"]}
|
||||
}
|
||||
]
|
||||
}
|
||||
}`)
|
||||
hits, err := parseESSearchResponse(raw)
|
||||
if err != nil {
|
||||
t.Fatalf("parseESSearchResponse: %v", err)
|
||||
}
|
||||
if len(hits) != 2 {
|
||||
t.Fatalf("hits = %d, want 2", len(hits))
|
||||
}
|
||||
h0 := hits[0]
|
||||
if h0.ID != "2395" || h0.Title != "Rechnung-4711.pdf" || h0.Kind != File {
|
||||
t.Errorf("hit0 entry = %+v", h0.Entry)
|
||||
}
|
||||
if h0.ParentID != "438" || !reflect.DeepEqual(h0.Path, []string{"438", "11"}) {
|
||||
t.Errorf("hit0 path = %v parent = %q", h0.Path, h0.ParentID)
|
||||
}
|
||||
if h0.Score != 7.31 {
|
||||
t.Errorf("hit0 score = %v", h0.Score)
|
||||
}
|
||||
if h0.Highlight != "… Zahlung der Rechnung …" {
|
||||
t.Errorf("hit0 highlight = %q, want content fragment", h0.Highlight)
|
||||
}
|
||||
if hits[1].Highlight != "Mahnung.pdf" {
|
||||
t.Errorf("hit1 highlight = %q, want title without tags", hits[1].Highlight)
|
||||
}
|
||||
if hits[1].ParentID != "" || len(hits[1].Path) != 0 {
|
||||
t.Errorf("hit1 path = %v", hits[1].Path)
|
||||
}
|
||||
}
|
||||
|
||||
func TestESSearchRequestJSONShape(t *testing.T) {
|
||||
got := esSearchRequest(SearchQuery{Text: "Rechnung", InContent: true}, "1")
|
||||
b, err := json.Marshal(got)
|
||||
if err != nil {
|
||||
t.Fatalf("marshal: %v", err)
|
||||
}
|
||||
var back map[string]any
|
||||
if err := json.Unmarshal(b, &back); err != nil {
|
||||
t.Fatalf("unmarshal: %v", err)
|
||||
}
|
||||
if _, ok := back["query"].(map[string]any)["bool"]; !ok {
|
||||
t.Errorf("query.bool missing: %s", b)
|
||||
}
|
||||
}
|
||||
@@ -5,6 +5,7 @@ import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"path"
|
||||
"strconv"
|
||||
@@ -357,20 +358,31 @@ func (c *Client) UploadToFolder(ctx context.Context, folderID, localPath string)
|
||||
|
||||
// UpdateFile uploads a new version of an existing file (same id, name and
|
||||
// folder). It does not delete and does not create a second file.
|
||||
//
|
||||
// The Documents API method is PUT /api/2.0/files/{id}/update; POST is kept as
|
||||
// a fallback for older servers. The path is tried with and without .json.
|
||||
func (c *Client) UpdateFile(ctx context.Context, fileID, localPath string) (*FileEntry, error) {
|
||||
if fileID == "" || localPath == "" {
|
||||
return nil, fmt.Errorf("file id and local path are required")
|
||||
}
|
||||
uploadPath := fmt.Sprintf("/api/2.0/files/%s/update", url.PathEscape(fileID))
|
||||
raw, err := c.uploadMultipart(ctx, uploadPath, "file", localPath)
|
||||
if err != nil {
|
||||
uploadPath = fmt.Sprintf("/api/2.0/files/%s/update.json", url.PathEscape(fileID))
|
||||
raw, err = c.uploadMultipart(ctx, uploadPath, "file", localPath)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
base := fmt.Sprintf("/api/2.0/files/%s/update", url.PathEscape(fileID))
|
||||
attempts := []struct {
|
||||
method, path string
|
||||
}{
|
||||
{http.MethodPut, base},
|
||||
{http.MethodPut, base + ".json"},
|
||||
{http.MethodPost, base},
|
||||
{http.MethodPost, base + ".json"},
|
||||
}
|
||||
return decodeResponseFileEntry(raw)
|
||||
var lastErr error
|
||||
for _, a := range attempts {
|
||||
raw, err := c.uploadMultipartMethod(ctx, a.method, a.path, "file", localPath)
|
||||
if err == nil {
|
||||
return decodeResponseFileEntry(raw)
|
||||
}
|
||||
lastErr = err
|
||||
}
|
||||
return nil, lastErr
|
||||
}
|
||||
|
||||
// FileFolderID returns the parent folder id string for a file entry, if known.
|
||||
|
||||
+6
-6
@@ -410,12 +410,12 @@ func (f *DavFolder) UnmarshalJSON(b []byte) error {
|
||||
// UnmarshalJSON decodes a file row, capturing size and timestamps.
|
||||
func (f *DavFile) UnmarshalJSON(b []byte) error {
|
||||
var raw struct {
|
||||
ID *json.Number `json:"id"`
|
||||
Title *string `json:"title"`
|
||||
PureSize *int64 `json:"pureContentLength"`
|
||||
SizeStr *string `json:"contentLength"`
|
||||
Updated *string `json:"updated"`
|
||||
ViewURL *string `json:"viewUrl"`
|
||||
ID *json.Number `json:"id"`
|
||||
Title *string `json:"title"`
|
||||
PureSize *int64 `json:"pureContentLength"`
|
||||
SizeStr *string `json:"contentLength"`
|
||||
Updated *string `json:"updated"`
|
||||
ViewURL *string `json:"viewUrl"`
|
||||
}
|
||||
if err := json.Unmarshal(b, &raw); err != nil {
|
||||
return err
|
||||
|
||||
@@ -346,6 +346,14 @@ func (c *Client) putJSON(ctx context.Context, path string, body any) (json.RawMe
|
||||
|
||||
// uploadMultipart posts a single file to path under the given form field name.
|
||||
func (c *Client) uploadMultipart(ctx context.Context, path, fieldName, filePath string) (json.RawMessage, error) {
|
||||
return c.uploadMultipartMethod(ctx, http.MethodPost, path, fieldName, filePath)
|
||||
}
|
||||
|
||||
// uploadMultipartMethod sends a single-file multipart request with the given
|
||||
// HTTP method. The OnlyOffice Documents API needs PUT for /update (a new
|
||||
// version) and POST for /upload (a new file); sending POST to /update answers
|
||||
// 500 on current servers.
|
||||
func (c *Client) uploadMultipartMethod(ctx context.Context, method, path, fieldName, filePath string) (json.RawMessage, error) {
|
||||
auth, err := c.authHeader()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -368,7 +376,7 @@ func (c *Client) uploadMultipart(ctx context.Context, path, fieldName, filePath
|
||||
if err := mw.Close(); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL()+path, &buf)
|
||||
req, err := http.NewRequestWithContext(ctx, method, c.baseURL()+path, &buf)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
@@ -0,0 +1,60 @@
|
||||
package onlyoffice
|
||||
|
||||
import (
|
||||
"context"
|
||||
"regexp"
|
||||
"time"
|
||||
)
|
||||
|
||||
// RetryPolicy controls deterministic retries against OnlyOffice: fixed linear
|
||||
// backoff without jitter, so repeated runs wait exactly the same schedule.
|
||||
// OnlyOffice throttles bulk reads/writes with 429 (and occasional 502/503/504
|
||||
// from openresty), so every bulk tool routes API calls through DoRetry.
|
||||
type RetryPolicy struct {
|
||||
Attempts int // total attempts, including the first try
|
||||
Base time.Duration // wait before retry N is N*Base
|
||||
Max time.Duration // per-wait cap
|
||||
}
|
||||
|
||||
// DefaultRetryPolicy retries up to 5 times with 1s, 2s, 3s, 4s waits.
|
||||
func DefaultRetryPolicy() RetryPolicy {
|
||||
return RetryPolicy{Attempts: 5, Base: time.Second, Max: 30 * time.Second}
|
||||
}
|
||||
|
||||
var transientRe = regexp.MustCompile(`:\s*(429|502|503|504)\b`)
|
||||
|
||||
// Transient reports whether err looks like a transient OnlyOffice answer
|
||||
// (an HTTP 429/502/503/504 surfaced as "...: <code> ...").
|
||||
func Transient(err error) bool {
|
||||
if err == nil {
|
||||
return false
|
||||
}
|
||||
return transientRe.MatchString(err.Error())
|
||||
}
|
||||
|
||||
// DoRetry runs fn until it succeeds, fails non-transiently, or attempts run
|
||||
// out. Waits are deterministic: N*Base capped at Max, no jitter.
|
||||
func DoRetry(ctx context.Context, p RetryPolicy, fn func() error) error {
|
||||
if p.Attempts < 1 {
|
||||
p.Attempts = 1
|
||||
}
|
||||
var err error
|
||||
for attempt := 1; attempt <= p.Attempts; attempt++ {
|
||||
if ctx.Err() != nil {
|
||||
return ctx.Err()
|
||||
}
|
||||
if err = fn(); err == nil || !Transient(err) || attempt == p.Attempts {
|
||||
return err
|
||||
}
|
||||
wait := time.Duration(attempt) * p.Base
|
||||
if wait > p.Max {
|
||||
wait = p.Max
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return ctx.Err()
|
||||
case <-time.After(wait):
|
||||
}
|
||||
}
|
||||
return err
|
||||
}
|
||||
Reference in New Issue
Block a user