Compare commits

...
62 Commits
Author SHA1 Message Date
eSlider aa069bea56 Merge pull request 'chore(release): v0.21.0 (#78)' (#79) from chore/release-v0.21.0#78 into main
Release / GoReleaser (push) Skipped
Tests / Test (Go 1.25) (push) Successful in 1m23s
Tests / Secret scan (gitleaks) (push) Successful in 3s
Tests / Test (Go stable) (push) Successful in 1m28s
2026-09-23 18:25:08 +01:00
eSlider a8fa202496 chore(release): v0.21.0 (#78)
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 1m29s
Tests / Test (Go stable) (pull_request) Successful in 1m33s
2026-09-23 18:25:00 +01:00
eSlider 87ff317166 Merge pull request 'chore: move business tooling out of the public library; tidy filestore naming' (#77) from chore/split-business-tools into main
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go 1.25) (push) Successful in 1m25s
Tests / Test (Go stable) (push) Successful in 1m28s
2026-09-23 09:05:48 +01:00
eSlider dc62570a5c chore: move business tooling out of the public library; tidy filestore naming
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m19s
Tests / Test (Go 1.25) (pull_request) Successful in 1m24s
Keep the public tree project-generic. Business/one-off tools, deployment
and business docs move to the private oo-workspace repo.

Moved to oo-workspace:
- cmd/ooscan, cmd/pdfamount, cmd/kontoblatt, cmd/kontolink
- internal/xlspipe (cutover-portugal workbook) -> oow workbook build
  (drops the --template/--title flags from oo docs put-xlsx)
- deploy/docker-compose.rclone-webdav.yml + docs/rclone-webdav.md
- docs/crm-associations.md

Removed GitHub-era leftovers:
- .github/workflows/release-please.yml, release-please-config.json,
  .release-please-manifest.json (tags are created on Gitea per SemVer)

Naming: the FileStore subsystem is now filestore_*.go (was file_*.go) to
match files_*.go (project Documents). Docs/AGENTS/README updated.
2026-09-23 09:05:40 +01:00
eSlider a9a77e110d Merge pull request 'feat: upstream generic workspace tooling (board-sync, crm audit, catalog names)' (#76) from feat/upstream-oow-workspace into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 1m23s
Tests / Test (Go 1.25) (push) Successful in 1m23s
2026-09-22 23:13:49 +01:00
eSlider 82da45dc4d feat: upstream generic workspace tooling (board-sync, crm audit, catalog names)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 1m16s
Tests / Test (Go stable) (pull_request) Successful in 1m36s
Promote the generic, reusable parts of the private oo-workspace into the
public library/CLI, so oo-workspace can shrink to business glue.

- board.go: Board types + (c *Client) SyncBoard — upsert project
  milestones/tasks by exact title from a YAML board (dry-run when
  apply=false). CLI: oo projects board-sync.
- crm_audit.go: OpportunityAudit + (c *Client) AuditOpportunities —
  file/task/member counts with a generic class (ok|dup|empty|junk-title).
  CLI: oo crm audit [--out].
- catalog/names.go: CleanPersonNames / GuessNameFromEmail /
  FormatProjectTitle (ported from oo-workspace).
- catalog: adopt the newer apply/match logic — name repair (never encode
  company in lastName), company grouping key, contact-info type
  normalization, preserve an already-applied oo_id. Keeps the local
  Address/OOProjects fields and the config-driven classifier.

Business-only oo-workspace commands (crm clean/migrate, folders, search)
deliberately stay private.
2026-09-22 23:13:40 +01:00
eSlider 8848ee00f1 Merge pull request 'feat(oo): native conversion, deep links, sheet export (re-land)' (#75) from feat/oo-conversion-links into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go 1.25) (push) Successful in 1m29s
Tests / Test (Go stable) (push) Successful in 1m30s
2026-09-22 22:42:07 +01:00
eSlider ff46aad615 chore: keep the re-landed feature files showroom-safe
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m21s
Tests / Test (Go 1.25) (pull_request) Successful in 1m27s
The conversion / link / sheet commits predate the tree cleanup, so their
comments and README/test examples still carried internal refs. Re-apply
the genericization (no internal hosts, personal names or client file ids).
2026-09-22 22:41:58 +01:00
eSlider 5e73f6c335 feat(oo): sheet-aware spreadsheet export (docs csv/json)
- lib: WorkbookSheetNames/CSV/JSON (excelize) — sheet selection + JSON rows
- oo docs csv  SRC... [--sheet N] [--delimiter ,|;|||tab] [--out]
- oo docs json SRC... [--sheet N] [--out]
- SRC = OO file id or local path; --sheet is 1-based (default first)
- DS csv output is first-sheet-only (verified) -> local reader for selection/JSON
- test: TestWorkbookSheetExport
2026-09-22 22:41:14 +01:00
eSlider 7c089a6439 feat(oo): docs pdf --stream/--pipe (converted bytes to stdout) 2026-09-22 22:41:14 +01:00
eSlider 4515a88fe8 docs(readme): examples for links, conversion, team/users CRUD 2026-09-22 22:41:14 +01:00
eSlider 58b905cc49 feat(oo): native document conversion (docs pdf/presigned)
- lib: PresignedURI, SignJWT (HS256, stdlib), ConvertDocument (DocumentServer
  /converter; legacy /ConvertService.ashx), DownloadURLTo
- oo docs pdf FILE_ID|PATH...: OO file ids or local files (temp upload to
  --folder, convert, download, cleanup) -> PDF/other
- oo docs presigned FILE_ID
- docs base from $ONLYOFFICE_DOCS_URL else $ONLYOFFICE_URL/ds-vpath; JWT secret
  from $ONLYOFFICE_DS_SECRET (DocumentServer CoAuthoring secret)
- test: SignJWT
2026-09-22 22:41:14 +01:00
eSlider 520d6d05c8 docs(oo): comment where the new helpers are used
- link.go/links.go: third-party doc packs (fileid links), id-stability caveat
- auth.go AuthenticateAs / users check: email-vs-userName login finding
- projects files replace-in: version-history rationale
2026-09-22 22:41:14 +01:00
eSlider df38b75011 feat(oo): file deep links, login check, replace-in (fresh upload)
- link: print Products/Files/DocEditor.aspx?fileid=… deep links (title+url)
- lib: FileEditorURL/FolderURL helpers (+tests)
- users check: verify credentials (userName vs email) via authentication.json
- projects files replace-in FOLDER_ID FILE...: hard delete same stem|ext in a
  folder + fresh upload → single clean version (no version history)
2026-09-22 22:41:14 +01:00
eSlider ee8e88fb9b feat(oo): projects files update (overwrite existing file content) 2026-09-22 22:41:14 +01:00
eSlider e323a43317 feat(oo): projects milestone-delete 2026-09-22 22:41:14 +01:00
eSlider 71d421fb41 Merge pull request 'chore: keep host/client specifics out of the tree (env & config)' (#74) from chore/showroom-generic into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 55s
Tests / Test (Go stable) (push) Successful in 1m6s
2026-09-22 22:40:49 +01:00
eSlider afb93feae5 chore: keep host/client specifics out of the tree (env & config)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 59s
Tests / Test (Go stable) (pull_request) Successful in 1m7s
Showroom-safe: the tree no longer carries internal hosts, IPs, ports,
personal names, client domains or client file names. Behaviour is
unchanged and now supplied per deployment.

- catalog: hardcoded mail-org / project classifiers become a YAML-driven
  Classifier (OO_CATALOG_CONFIG or --config); neutral default classifies
  nothing as work. New catalog/classify.go + example + tests.
- storage_fallback: drop the baked-in MinIO endpoint IP; require
  MINIO_ENDPOINT (+ keys) from the env.
- kontolink: build DocEditor links from $ONLYOFFICE_URL instead of a
  hardcoded portal host; kontoblatt: no client file id in the output name.
- oo: load .env CLI-wide (bootstrap.LoadEnv in execute) so
  non-authenticating commands (catalog scans) also see config.
- genericize comments/docs/fixtures (AGENTS, README, .env.example,
  crm-associations, catalog tests, ES/pdfattach tests, mails).
2026-09-22 22:03:36 +01:00
eSlider b493472f2e Merge pull request 'feat(oo): project team CRUD and user lifecycle' (#73) from feat/oo-team-users-crud into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 1m1s
Tests / Test (Go 1.25) (push) Successful in 1m2s
Reviewed-on: #73
2026-09-22 14:29:47 +01:00
eSlider 1dc99d6997 feat(oo): project team CRUD and user lifecycle (block/unblock/password/delete)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 1m2s
Tests / Test (Go stable) (pull_request) Successful in 1m40s
- project team: oo projects team list|add|remove|set (portal users)
- users: get|create|update|delete|block|unblock|password
- lib: CreateUser, DeleteUser (auto-suspend before delete), Add/Remove/
  SetProjectTeam, ListProjectTeam; deleteJSON reused; jsonBodyReader helper
- tests: unmarshalResponseArray
2026-09-22 14:26:23 +01:00
eSlider 9b21d7e52a Merge pull request 'fix(crm): адреса контактов через ContactInfo типа Address (#288)' (#72) from feat/catalog-addresses#288 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 58s
Tests / Test (Go 1.25) (push) Successful in 59s
2026-09-22 08:32:25 +01:00
eSlider c4beba5173 fix(crm): store contact addresses via ContactInfo Address type (#288)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m1s
Tests / Test (Go 1.25) (pull_request) Successful in 1m1s
2026-09-22 07:08:01 +00:00
eSlider 2a187ae132 Merge pull request 'fix(retry): глобальный rate-limit + Retry-After + cooldown (#70)' (#71) from fix/oo-backoff#70 into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 21s
Tests / Test (Go 1.25) (push) Successful in 22s
2026-09-17 13:39:56 +01:00
eSlider c9e16c7169 fix(retry): глобальный rate-limit + Retry-After + cooldown (#70)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 25s
Tests / Test (Go stable) (pull_request) Successful in 27s
2026-09-17 12:39:25 +00:00
eSlider 194ae62720 Merge pull request 'docs: убрать потребительские детали match из index-and-search (#34)' (#69) from docs/index-and-search-scope into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 26s
Tests / Test (Go 1.25) (push) Successful in 26s
2026-09-17 08:51:45 +01:00
eSlider 0d73146ebc docs: убрать потребительские детали match из index-and-search (#34)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 3s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
Tests / Test (Go stable) (pull_request) Successful in 29s
2026-09-17 07:51:02 +00:00
eSlider 46661f51f8 Merge pull request 'docs: карта поиска и обновления индексов (#34)' (#68) from docs/index-and-search into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 25s
Tests / Test (Go 1.25) (push) Successful in 29s
2026-09-17 08:49:31 +01:00
eSlider f22d714be7 docs: карта поиска и обновления индексов (#34)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
Tests / Test (Go stable) (pull_request) Successful in 21s
2026-09-17 07:48:40 +00:00
eSlider 6f85c206cf Merge pull request 'test(files): live-тест файлового dedup (#63)' (#65) from test/w6-file-dedup#63 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Test (Go stable) (push) Failing after 3s
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 28s
2026-09-16 23:03:23 +01:00
eSlider 1c00edbd0b Merge pull request 'test(files): live CRUD папок + фасадный CRUD (#62)' (#66) from test/w5-folder-crud#62 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Test (Go stable) (push) Failing after 4s
Tests / Secret scan (gitleaks) (push) Successful in 6s
Tests / Test (Go 1.25) (push) Successful in 23s
2026-09-16 23:03:17 +01:00
eSlider 9cd157c59f test(files): live CRUD папок + фасадный CRUD (#62)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Test (Go stable) (pull_request) Failing after 4s
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
2026-09-16 21:49:10 +00:00
eSlider 932cd2895a fix(files): REST FileStore resolves folders for stat/rename/move/delete (#62) 2026-09-16 21:49:08 +00:00
eSlider 5588c2d3f7 fix(files): ретраить transient-ответы при удалении файлов (#63)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 20s
Tests / Test (Go 1.25) (pull_request) Successful in 20s
DeleteDavItems (через ApplyDedupGroups/DeleteFilesByDedupKey) ходил в
deleteJSON напрямую и падал на 429/502/503/504. Теперь каждый delete идёт
через DoRetry, как остальные bulk-пути.
2026-09-16 21:47:37 +00:00
eSlider f8783b86a8 fix(files): пропускать пустой dedup-ключ (dotfiles) (#63)
Файлы без stem (.env, .gitignore, .npmrc) давали FileDedupKey="", и
findWithinFolderDuplicates/findCrossFolderDuplicates считали их одной
группой дублей и удаляли лишние. Пустой ключ больше не образует группу.
2026-09-16 21:47:36 +00:00
eSlider d9db1c5e56 test(files): live-тест файлового dedup (#63)
TestIntegrationFileDedup: throwaway-проект, две папки + root, дубли по
stem|ext, FindProjectDuplicates/mergeProjectRootForDedupe,
ApplyDedupGroups/DeleteFilesByDedupKey, survivor и удаления подтверждены
polling-ом. Папки/аплоады ретраят временный post-create 500 портала.
2026-09-16 21:47:33 +00:00
eSlider c551634e55 Merge pull request 'docs: Testing + rclone/SQL + docs index (#53)' (#64) from docs/w1-w3#53 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 19s
Tests / Test (Go 1.25) (push) Successful in 24s
2026-09-16 22:41:12 +01:00
eSlider a88fd5b40c docs: Testing section, rclone/SQL refs, docs index (#53)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
Tests / Test (Go stable) (pull_request) Successful in 22s
2026-09-16 21:39:41 +00:00
eSlider 9b149281a8 Merge pull request 'fix(retry): починка красного, полный прогон тестов (#57)' (#61) from chore/test-green#57 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 23s
Tests / Test (Go stable) (push) Successful in 24s
2026-09-16 22:37:05 +01:00
eSlider d6b8eb7777 fix(retry): retry transient edge answers centrally, fix stale task test (#57)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 21s
Tests / Test (Go stable) (pull_request) Successful in 21s
- Route HTTP helpers (getJSON/formRequest/deleteReq/postJSON/putJSON/
  multipart upload), Query and AuthenticateContext/ensureToken through the
  deterministic 429/502/503/504 retry (retryRaw), so the integration suite
  no longer fails on the shared openresty rate limit under parallel runs.
- Fix stale cmd/office/fetch integration test: loader.TaskFields was
  replaced by loader.DetailForm (broke go vet -tags=integration).
- Add unit tests for Transient/DoRetry.
2026-09-16 21:32:14 +00:00
eSlider 8a17a32269 Merge pull request 'feat(deploy): rclone WebDAV mount (docker compose) + smoke + docs (#54)' (#59) from feat/rclone-webdav#54 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 8s
Tests / Test (Go 1.25) (push) Successful in 25s
Tests / Test (Go stable) (push) Successful in 27s
2026-09-16 22:30:16 +01:00
eSlider 43a1a47faa Merge pull request 'test(es): доказать, что oo search идёт в Elasticsearch (#56)' (#58) from feat/es-integration#56 into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 26s
Tests / Test (Go stable) (push) Successful in 28s
2026-09-16 22:30:14 +01:00
eSlider 620505cb24 Merge pull request 'feat(files): SQL backend via Client.FileStore + live MySQL facade test (#55)' (#60) from feat/db-integration#55 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 30s
Tests / Test (Go 1.25) (push) Successful in 37s
2026-09-16 22:30:05 +01:00
eSlider a6bba30438 feat(files): SQL file backend via Client.FileStore/SQLFileStore + live MySQL integration (#55)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 24s
Tests / Test (Go 1.25) (pull_request) Successful in 26s
2026-09-16 21:26:15 +00:00
SE 6549fbf7e2 feat(deploy): rclone WebDAV mount compose + smoke + docs (#54)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 20s
Tests / Test (Go stable) (pull_request) Successful in 22s
2026-09-16 21:25:38 +00:00
eSlider 1a8bf0f83c test(es): prove oo search uses the Elasticsearch backend (#56)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 3s
Tests / Test (Go stable) (pull_request) Successful in 24s
Tests / Test (Go 1.25) (pull_request) Successful in 25s
2026-09-16 21:22:43 +00:00
eSlider 046882b8ef Merge pull request 'feat(search): полный уникальный path первым + ближайшая папка в результатах (#51)' (#52) from feat/search-path#51 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 22s
Tests / Test (Go 1.25) (push) Successful in 22s
2026-09-16 22:19:46 +01:00
Yftyr a3ef83a961 feat(search): full unique path first + immediate folder in results (#51)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 21s
Tests / Test (Go stable) (pull_request) Successful in 26s
2026-09-16 21:19:32 +00:00
eSlider 54f825f8e8 Merge pull request 'feat(search): подстрока/AND, scope по поддереву, лимит 1000 (#49)' (#50) from fix/es-substring#49 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 24s
Tests / Test (Go 1.25) (push) Successful in 26s
2026-09-16 22:16:01 +01:00
Yftyr 6a4940ba82 feat(search): substring/AND terms, nested folder scope, limit 1000 (#49)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 26s
Tests / Test (Go 1.25) (pull_request) Successful in 39s
2026-09-16 21:15:44 +00:00
eSlider 1da368fd5d Merge pull request 'fix(docpipe): индекс вложений .yaml и без расширения (#47)' (#48) from fix/attachments-yaml#47 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 26s
Tests / Test (Go 1.25) (push) Successful in 27s
2026-09-16 19:54:20 +01:00
Yftyr 66a6b57401 fix(docpipe): index .yaml and extensionless PDF attachments (#47)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 21s
Tests / Test (Go 1.25) (pull_request) Successful in 22s
2026-09-16 18:54:09 +00:00
eSlider 075c66d4d7 Merge pull request 'docs: единый контракт файлового клиента (#39)' (#46) from docs/file-client#39 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 6s
Tests / Test (Go stable) (push) Successful in 21s
Tests / Test (Go 1.25) (push) Successful in 1m20s
2026-09-16 18:37:36 +01:00
eSlider c59ad645ea docs(files): add unified file client contract (#39)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 17s
Tests / Test (Go stable) (pull_request) Successful in 1m20s
2026-09-16 17:37:15 +00:00
eSlider c76ab7267d Merge pull request 'feat(search): PDF-контент в поиске — свой индекс oo_docs_text (#42)' (#44) from feat/pdf-content#42 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 1m21s
Tests / Test (Go 1.25) (push) Successful in 1m23s
2026-09-16 18:34:54 +01:00
eSlider 3cc288d281 feat(search): index embedded PDF attachment text (#42)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 25s
Tests / Test (Go 1.25) (pull_request) Successful in 28s
2026-09-16 17:34:04 +00:00
eSlider ce4778bdf1 fix(files): dedupe ProviderPG after facade/SQL merge, rename test fake 2026-09-16 17:29:01 +00:00
eSlider 3fe43ee82c feat(search): PDF content via own ES index and oo index (#42) 2026-09-16 17:28:05 +00:00
eSlider 6d93ab5b1f Merge pull request 'feat(files): read-only SQL file store (Community Server DB) (#36)' (#45) from feat/pg-store#36 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 3s
Tests / Test (Go stable) (push) Failing after 26s
Tests / Test (Go 1.25) (push) Failing after 28s
2026-09-16 18:27:21 +01:00
eSlider 95d79925ea Merge pull request 'refactor(files): единый файловый фасад + миграция CLI/TUI (#38)' (#43) from feat/file-facade#38 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Failing after 37s
Tests / Test (Go stable) (push) Failing after 38s
2026-09-16 18:27:21 +01:00
eSlider bf3ef025aa feat(files): read-only SQL file store over Community Server DB (#36)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 1m56s
Tests / Test (Go stable) (pull_request) Successful in 1m57s
2026-09-16 17:24:22 +00:00
eSlider ba29738482 refactor(files): single file client facade + CLI/TUI migration (#38)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 19s
Tests / Test (Go 1.25) (pull_request) Successful in 22s
- FileClient composes FileStore/Searcher backends; Read()/Write()/Search()
  select REST/DAV/PG/ES with transient read fallback. Client.Files() returns
  it and *FileClient implements FileStore, so existing callers keep working.
- Entry gains backend-native Updated + folder FilesCount/FoldersCount so
  dav ls output round-trips.
- cmd/oo dav/projects files and cmd/office/fetch download/preview/delete go
  through FileStore; oo search goes through the facade.
- Mark transport methods that FileStore now abstracts as deprecated.
2026-09-16 17:20:41 +00:00
eSlider 72c7cd5173 Merge pull request 'feat(search): Elasticsearch-клиент (имя + контент) + oo search (#37)' (#41) from feat/es-search#37 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 3s
Tests / Test (Go stable) (push) Successful in 17s
Tests / Test (Go 1.25) (push) Successful in 24s
2026-09-16 17:43:21 +01:00
116 changed files with 9368 additions and 2531 deletions
+33 -3
View File
@@ -22,22 +22,52 @@ ONLYOFFICE_PROJECT_ID=33
# oo mails uses ONLYOFFICE_URL/USER/PASS above (Workspace Mail addon). # oo mails uses ONLYOFFICE_URL/USER/PASS above (Workspace Mail addon).
# deploy/docker-compose.rclone-webdav.yml — rclone FUSE mount of the oo-webdav
# sidecar. Reuses ONLYOFFICE_USER and accepts ONLYOFFICE_PASSWORD (alias
# ONLYOFFICE_PASS). Set ONLYOFFICE_WEBDAV_URL only if the sidecar is not on
# the default docker bridge address. See docs/rclone-webdav.md.
# ONLYOFFICE_WEBDAV_URL=http://172.17.0.1:8098/webdav
# cmd/office TUI — optional Document Server for DOCX→HTML preview: # cmd/office TUI — optional Document Server for DOCX→HTML preview:
# ONLYOFFICE_DOCS_URL=https://docs.example.com # ONLYOFFICE_DOCS_URL=https://docs.example.com
# ONLYOFFICE_DOCS_SECRET= # ONLYOFFICE_DOCS_SECRET=
# MinIO download fallback for the portal's stale AWS S3 consumer (older # MinIO download fallback for the portal's stale AWS S3 consumer (older
# Documents folders). When the portal redirects to amazonaws.com with access # Documents folders). When the portal redirects to amazonaws.com with access
# key "minio" (403 InvalidAccessKeyId), files are fetched from the local MinIO # key "minio" (403 InvalidAccessKeyId), files are fetched from the configured
# store instead. Without a key/secret the fallback is disabled. # MinIO store instead. Without endpoint + key/secret the fallback is disabled.
# MINIO_ENDPOINT=http://192.168.188.10:9000 # MINIO_ENDPOINT=http://minio.example.com:9000
# MINIO_BUCKET=office # MINIO_BUCKET=office
# MINIO_ACCESS_KEY= # MINIO_ACCESS_KEY=
# MINIO_SECRET_KEY= # MINIO_SECRET_KEY=
# Catalog scan classification rules (client mail domains, project remotes/names).
# Copy catalog/classify.example.yaml and point this at it; no rules means
# nothing is classified as work.
# OO_CATALOG_CONFIG=~/.config/oo/catalog-classify.yaml
# oo search — direct Elasticsearch access for name + content search. ES lives # oo search — direct Elasticsearch access for name + content search. ES lives
# inside the OnlyOffice VM on localhost:9200; expose it with an SSH tunnel # inside the OnlyOffice VM on localhost:9200; expose it with an SSH tunnel
# (see docs/elasticsearch.md). ONLYOFFICE_ES_INDEX defaults to files_file. # (see docs/elasticsearch.md). ONLYOFFICE_ES_INDEX defaults to files_file.
# ONLYOFFICE_ES_URL=http://127.0.0.1:9200 # ONLYOFFICE_ES_URL=http://127.0.0.1:9200
# ONLYOFFICE_ES_INDEX=files_file # ONLYOFFICE_ES_INDEX=files_file
# ONLYOFFICE_TENANT= # ONLYOFFICE_TENANT=
# Read-only SQL file store over the Community Server database (see
# docs/community-server-db.md). The live portal runs MySQL; ONLYOFFICE_DSN is
# `user:pass@tcp(host:port)/onlyoffice?parseTime=true`, or a `postgres://` URL.
# ONLYOFFICE_DSN=
# ONLYOFFICE_PG_DRIVER= # postgres | mysql (auto-detected from DSN)
# ONLYOFFICE_PG_TENANT=1 # falls back to ONLYOFFICE_TENANT
# Alternatively build a PostgreSQL DSN from parts:
# ONLYOFFICE_PG_HOST=
# ONLYOFFICE_PG_PORT=5432
# ONLYOFFICE_PG_USER=
# ONLYOFFICE_PG_PASSWORD=
# ONLYOFFICE_PG_DBNAME=onlyoffice
# ONLYOFFICE_PG_SSLMODE=disable
# oo index / oo search --backend own — own full-text index for PDF/scans,
# filled by `oo index` from internal/docpipe (pdftotext + OCR). Defaults to
# oo_docs_text. Uses the same ONLYOFFICE_ES_URL.
# ONLYOFFICE_ES_TEXT_INDEX=oo_docs_text
-42
View File
@@ -1,42 +0,0 @@
name: Release Please
on:
push:
branches: [main]
workflow_dispatch:
permissions:
contents: write
pull-requests: write
issues: write
actions: write
jobs:
release-please:
name: Release Please
if: github.server_url == 'https://github.com'
runs-on: ubuntu-latest
steps:
- name: Run release-please
uses: googleapis/release-please-action@v4
id: release
with:
config-file: release-please-config.json
manifest-file: .release-please-manifest.json
token: ${{ secrets.RELEASE_PLEASE_TOKEN != '' && secrets.RELEASE_PLEASE_TOKEN || secrets.GITHUB_TOKEN }}
# Tags/releases created with GITHUB_TOKEN do not trigger other workflows
# (push:tags on Release.yml never fires). Dispatch GoReleaser explicitly.
- name: Trigger GoReleaser
if: ${{ steps.release.outputs.release_created == 'true' }}
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
TAG: ${{ steps.release.outputs.tag_name }}
run: |
set -euo pipefail
echo "Dispatching Release workflow for $TAG"
gh workflow run release.yml --repo "${{ github.repository }}" -f "tag=${TAG}"
outputs:
release_created: ${{ steps.release.outputs.release_created }}
tag_name: ${{ steps.release.outputs.tag_name }}
+2 -2
View File
@@ -20,8 +20,8 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
# Always clone default branch first. workflow_dispatch often races with # Always clone default branch first. workflow_dispatch often races with
# release-please creating the tag; fetching refs/tags/X before it exists # tag creation; fetching refs/tags/X before it exists fails checkout.
# fails checkout. Retry fetch, then check out the tag. # Retry fetch, then check out the tag.
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with: with:
fetch-depth: 0 fetch-depth: 0
-3
View File
@@ -1,3 +0,0 @@
{
".": "0.18.0"
}
+14 -7
View File
@@ -9,18 +9,19 @@ Canonical Go client for OnlyOffice Workspace (Projects + Calendar + CRM) and the
- `request.go` — `Request`, `Query`, `Time`, `Token`, `MetaResponse`, `Permissions`. - `request.go` — `Request`, `Query`, `Time`, `Token`, `MetaResponse`, `Permissions`.
- `auth.go` — `Authenticate`, `AuthenticateContext`, `InvalidateToken`, `Auth`, token lifecycle. - `auth.go` — `Authenticate`, `AuthenticateContext`, `InvalidateToken`, `Auth`, token lifecycle.
- `http.go` — transport + DRY response decoders (`ResponseArray`/`ResponseObject`/`postFormObject`/`putFormObject`/`deleteObject`). - `http.go` — transport + DRY response decoders (`ResponseArray`/`ResponseObject`/`postFormObject`/`putFormObject`/`deleteObject`).
- `projects.go`, `tasks.go`, `users.go`, `calendar.go`, `crm.go`, `files.go`, `files_webdav.go`, `files_stem.go`, `retry.go`, `mails.go`, `invoices.go` — typed / untyped domain methods. **`files.go`** — CRM opportunity upload plus **project/task Documents** (`UpdateFile`, `UploadToFolderReplacing`). **`files_webdav.go`** — Documents module by id (`ListDavFolder`, `MoveDavItems`/`CopyDavItems` with per-operation error surfacing, `ListFileOps`). **`retry.go`** — `DoRetry`: deterministic linear backoff (no jitter) on 429/502/503/504; every bulk tool routes API calls through it. **`mails.go`** — OnlyOffice Workspace Mail. **`invoices.go`** — CRM invoices, PDF regen/cleanup, status. Association rules: [`docs/crm-associations.md`](docs/crm-associations.md). - `projects.go`, `tasks.go`, `users.go`, `calendar.go`, `crm.go`, `files.go`, `files_webdav.go`, `files_stem.go`, `retry.go`, `mails.go`, `invoices.go` — typed / untyped domain methods. **`files.go`** — CRM opportunity upload plus **project/task Documents** (`UpdateFile`, `UploadToFolderReplacing`). **`files_webdav.go`** — Documents module by id (`ListDavFolder`, `MoveDavItems`/`CopyDavItems` with per-operation error surfacing, `ListFileOps`). **`retry.go`** — `DoRetry`: deterministic exponential backoff (no jitter) on 429/502/503/504; every bulk tool routes API calls through it, and the HTTP transport + auth (`retryRaw`, `AuthenticateContext`) retry transient answers centrally. `ratelimit.go` adds a process-wide token bucket (`OO_RATE_LIMIT`/`OO_BURST`) and a shared 429 cooldown gate, installed via `pacedTransport` in `NewClient`; `Retry-After` is parsed into `*TransientError` and honoured. See [`docs/rate-limiting.md`](docs/rate-limiting.md). **`mails.go`** — OnlyOffice Workspace Mail. **`invoices.go`** — CRM invoices, PDF regen/cleanup, status. Association rules live with the private `oo-workspace` tooling.
- **Unified file client (epic #34) — `filestore_core.go`, `filestore_rest.go`, `filestore_dav.go`, `filestore_pg.go`, `filestore_es.go`, `filestore_es_text.go`, `filestore_text_index.go`, `filestore_facade.go`.** `filestore_core.go` — model (`Entry`, `Kind`) + `FileStore`/`Searcher`; `filestore_rest.go`/`filestore_dav.go` — REST/WebDAV adapters; `filestore_pg.go` — **read-only** SQL store (PostgreSQL/MySQL, `ErrReadOnly` on writes); `filestore_es.go` — OnlyOffice Elasticsearch searcher; `filestore_es_text.go`/`filestore_text_index.go` — own PDF/scan index (`oo_docs_text`, PDF attachments via pdfdetach); `filestore_facade.go` — `FileClient` with read/write/search order and transient fallback. Use `c.Files()` (facade), `c.FileStore("rest"|"dav"|"pg"|"sql")` or `c.SQLFileStore()`; contract and how to add a backend: [`docs/unified-file-client.md`](docs/unified-file-client.md).
- Pure stdlib + `google/go-querystring`; no UI, no dotenv. - Pure stdlib + `google/go-querystring`; no UI, no dotenv.
- **CLI — `cmd/oo/` as `package main`.** Cobra wrapper that loads `.env` via `godotenv` at startup. **Subject-based command tree** mirroring [`tea`](https://gitea.com/gitea/tea): - **CLI — `cmd/oo/` as `package main`.** Cobra wrapper that loads `.env` via `godotenv` at startup. **Subject-based command tree** mirroring [`tea`](https://gitea.com/gitea/tea):
- `main.go` — entry point (docstring lists the command tree). - `main.go` — entry point (docstring lists the command tree).
- `common.go` — `rootCmd`, `newOO`, `printTable`/`printObject`, `--output table|json` flag. - `common.go` — `rootCmd`, `newOO`, `printTable`/`printObject`, `--output table|json` flag.
- `calendar.go`, `projects.go`, `projects_files.go`, `tasks.go`, `tasks_files.go`, `users.go`, `contacts.go`, `opportunities.go`, `cases.go`, `crm.go`, `crm_tasks.go`, `catalog.go`, `docs.go`, `dav.go`, `mails.go`, `invoices.go` — one file per subject (or per subject facet), each registers in `init()`. `dav.go` exposes the Documents module by id (`oo dav ls|move|copy|mkdir|rename-file|rename-folder|download|fileops`). - `calendar.go`, `projects.go`, `projects_files.go`, `tasks.go`, `tasks_files.go`, `users.go`, `contacts.go`, `opportunities.go`, `cases.go`, `crm.go`, `crm_tasks.go`, `catalog.go`, `docs.go`, `dav.go`, `search.go`, `index.go`, `mails.go`, `invoices.go` — one file per subject (or per subject facet), each registers in `init()`. `dav.go` exposes the Documents module by id (`oo dav ls|move|copy|mkdir|rename-file|rename-folder|download|fileops`); `search.go` runs `oo search QUERY` (name/content, `--backend oo|own`); `index.go` fills the own full-text index (`oo index folder|files`, see [`docs/unified-file-client.md`](docs/unified-file-client.md)).
- CLI-only deps (`spf13/cobra`, `joho/godotenv`) stay out of the library. - CLI-only deps (`spf13/cobra`, `joho/godotenv`) stay out of the library.
- **TUI — `cmd/office/` as `package main`.** Bubble Tea three-pane browser (module tree, selectable list, markdown preview). Reuses `cmd/internal/bootstrap` for env/auth and the root `onlyoffice` library for all API calls. UI logic in `cmd/office/ui/`; preview/formatting in `cmd/office/preview/`; list loaders in `cmd/office/fetch/`. - **TUI — `cmd/office/` as `package main`.** Bubble Tea three-pane browser (module tree, selectable list, markdown preview). Reuses `cmd/internal/bootstrap` for env/auth and the root `onlyoffice` library for all API calls. UI logic in `cmd/office/ui/`; preview/formatting in `cmd/office/preview/`; list loaders in `cmd/office/fetch/`.
- **List table (`DataTable`)** — `cmd/office/ui/table*.go`. Column layout policies live in `cmd/office/model/table_layout.go` (`TableFlexLayoutFor`); cell rendering uses the bubbles/table inline pattern in `table_render.go` (`renderTableCell`, `padANSIWidth`). See `.cursor/skills/office-tui-table/SKILL.md` before changing center-pane tables. - **List table (`DataTable`)** — `cmd/office/ui/table*.go`. Column layout policies live in `cmd/office/model/table_layout.go` (`TableFlexLayoutFor`); cell rendering uses the bubbles/table inline pattern in `table_render.go` (`renderTableCell`, `padANSIWidth`). See `.cursor/skills/office-tui-table/SKILL.md` before changing center-pane tables.
- **Shared bootstrap — `cmd/internal/bootstrap/`.** `LoadEnv()` + `NewClient(ctx)` extracted from `oo`; both binaries import it. - **Shared bootstrap — `cmd/internal/bootstrap/`.** `LoadEnv()` + `NewClient(ctx)` extracted from `oo`; both binaries import it.
- **Bulk Documents tools — `cmd/ooscan/`, `cmd/pdfamount/`, `cmd/kontoblatt/`, `cmd/kontolink/`.** Single-purpose binaries (folder index, PDF amounts, Kontoblatt summary/linking). Pace requests, route API calls through `DoRetry`; usage in README. - **Bulk Documents tools live in the private `oo-workspace` repo** (`ooscan`, `pdfamount`, `kontoblatt`, `kontolink`), not in this public tree. They use the public client and its `DoRetry` pacing.
- **Personal ops tooling** (disk inventory, dossier→CRM sync, SearXNG) lives in private [`eSlider/oo-workspace`](https://git.produktor.io/eSlider/oo-workspace) (`oow`), not in this public tree. - **Personal ops tooling** (disk inventory, dossier→CRM sync, SearXNG) lives in a private companion repo `eSlider/oo-workspace` (the `oow` CLI), not in this public tree.
## Rules ## Rules
@@ -32,6 +33,11 @@ Canonical Go client for OnlyOffice Workspace (Projects + Calendar + CRM) and the
- **Documents for agents:** prefer Markdown in git; OnlyOffice UI is weak for `.md`/`.txt`. Use `oo docs put-md` (md→docx) and `oo docs put-txt` (txt→docx, preserves line breaks). All upload paths default to **upsert** by `stem|ext` (`--replace`, default true); `--no-replace` fails on conflict; `--allow-duplicate` opts into raw OO append. `oo projects files dedupe PROJECT_ID` reports/removes duplicate stem|ext copies (`--apply`, `--cross`; includes project root folder). - **Documents for agents:** prefer Markdown in git; OnlyOffice UI is weak for `.md`/`.txt`. Use `oo docs put-md` (md→docx) and `oo docs put-txt` (txt→docx, preserves line breaks). All upload paths default to **upsert** by `stem|ext` (`--replace`, default true); `--no-replace` fails on conflict; `--allow-duplicate` opts into raw OO append. `oo projects files dedupe PROJECT_ID` reports/removes duplicate stem|ext copies (`--apply`, `--cross`; includes project root folder).
- Every table output goes through `printTable(headers, rows)`; every single-object through `printObject(v)`. Do not `fmt.Println` rows ad-hoc or the `--output json` flag breaks for that command. - Every table output goes through `printTable(headers, rows)`; every single-object through `printObject(v)`. Do not `fmt.Println` rows ad-hoc or the `--output json` flag breaks for that command.
- No secrets in the repo; use `.env` (gitignored). Commit `.env.example` only. - No secrets in the repo; use `.env` (gitignored). Commit `.env.example` only.
- **No host/client specifics in the tree.** Endpoints, IPs/ports, client mail
domains, project remotes/names and personal names stay out of source and
fixtures — they come from env/config (`MINIO_*`, `OO_CATALOG_CONFIG`, see
[`catalog/classify.example.yaml`](catalog/classify.example.yaml)). This repo is
mirrored to GitHub as a public showroom, so the tree must stay project-generic.
- Follow SemVer on tags; this repo is tagged at GitHub under `git@github.com:eSlider/go-onlyoffice.git`. - Follow SemVer on tags; this repo is tagged at GitHub under `git@github.com:eSlider/go-onlyoffice.git`.
### Testing policy (2026-04-24) ### Testing policy (2026-04-24)
@@ -55,6 +61,7 @@ write `mux.HandleFunc("/api/2.0/...")` to emulate OnlyOffice, we write an
## Related ## Related
- [`eSlider/inventar`](https://git.produktor.io/eSlider/inventar) — ASR/ADR (see ASR-0008 Go library module conventions). - [`docs/README.md`](docs/README.md) — reference index (file client, ES, SQL, rclone).
- [`eSlider/inventar-sync`](https://git.produktor.io/eSlider/inventar-sync) — OnlyOffice → Gitea issue sync, consumes this library. - `eSlider/inventar` — ASR/ADR (see ASR-0008 Go library module conventions).
- [`produktor.io/vidarr`](https://git.produktor.io/produktor.io/vidarr) — legacy consumer being migrated from `pkg/onlyoffice` to this module. - `eSlider/inventar-sync` — OnlyOffice → Gitea issue sync, consumes this library.
- `vidarr` — legacy consumer being migrated from `pkg/onlyoffice` to this module.
+22 -1
View File
@@ -4,7 +4,28 @@ All notable changes to this project are documented here. The format is based on
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project
adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## Unreleased ## [0.21.0](https://git.produktor.io/eSlider/go-onlyoffice/compare/v0.20.0...v0.21.0) (2026-09-23)
### Features
* **retry:** global token-bucket rate limit (`OO_RATE_LIMIT`/`OO_BURST`), typed
`TransientError` with `Retry-After`, exponential backoff
(`OO_RETRY_ATTEMPTS`/`_BASE`/`_MAX`) and a process-wide 429 cooldown gate.
All HTTP paths are paced via `pacedTransport` (#70).
* **oo:** native document conversion (`oo docs pdf`), deep links, sheet-aware
export (`oo docs csv`/`json`), `oo docs presigned`, plus project team and
user lifecycle, milestone-delete and file update/replace (#73, #75).
* **catalog:** upstream generic workspace tooling (board-sync, CRM audit,
catalog names) (#76).
### Chores
* move business tooling (`ooscan`, `pdfamount`, `kontoblatt`, `kontolink`) out
of the public library into the private `oo-workspace` repo; tidy filestore
naming (#77).
* keep host/client specifics out of the tree (#74).
## [0.18.0](https://github.com/eSlider/go-onlyoffice/compare/v0.17.0...v0.18.0) (2026-09-04) ## [0.18.0](https://github.com/eSlider/go-onlyoffice/compare/v0.17.0...v0.18.0) (2026-09-04)
+171 -29
View File
@@ -476,6 +476,11 @@ type Task struct {
| `UpdateProject(req)` | Update project details | | `UpdateProject(req)` | Update project details |
| `DeleteProject(id)` | Delete a project | | `DeleteProject(id)` | Delete a project |
| `GetProjectMilestones(project)` | Get milestones with task counts | | `GetProjectMilestones(project)` | Get milestones with task counts |
| `DeleteMilestone(id)` | Remove a milestone |
| `ListProjectTeam(ctx, id)` | Portal users on the project team |
| `AddProjectTeamUser(ctx, id, userID)` | Add a portal user to the team |
| `RemoveProjectTeamUser(ctx, id, userID)` | Remove a portal user from the team |
| `SetProjectTeam(ctx, id, participants, notify)` | Replace the team with the given user ids |
### Tasks ### Tasks
@@ -502,6 +507,14 @@ type Task struct {
| Method | Description | | Method | Description |
|---|---| |---|---|
| `GetUsers()` | List all users with profiles | | `GetUsers()` | List all users with profiles |
| `GetUser(ctx, id)` | One user profile by id |
| `CreateUser(ctx, NewUserRequest)` | Add a portal user |
| `UpdateUser(ctx, id, body)` | Update profile fields (JSON PUT) |
| `DeleteUser(ctx, id)` | Delete permanently (auto-terminates first — OO refuses active users) |
| `BlockUser(ctx, id)` / `UnblockUser(ctx, id)` | Terminate / reactivate (login kept/denied) |
| `ChangeUserPassword(ctx, id, pw)` | Set a new password |
| `ChangeUserStatus(ctx, id, active)` | Activate / Terminate via `people/status` |
| `AuthenticateAs(ctx, login, pw)` | Verify a login without mutating the cached token |
### Documents Files ### Documents Files
@@ -522,9 +535,33 @@ type Task struct {
| `ListFileOps(ctx)` | Active file operations (move/copy status polling) | | `ListFileOps(ctx)` | Active file operations (move/copy status polling) |
| `FolderFiles(ctx, folderID)` | Flat file list of a folder (stem helpers) | | `FolderFiles(ctx, folderID)` | Flat file list of a folder (stem helpers) |
| `DeleteFilesByStem(ctx, folderID, stem)` | Remove `stem\|ext` copies | | `DeleteFilesByStem(ctx, folderID, stem)` | Remove `stem\|ext` copies |
| `DoRetry(ctx, policy, fn)` | Deterministic linear backoff (N·Base, no jitter) on 429/502/503/504 | | `DoRetry(ctx, policy, fn)` | Deterministic exponential backoff (`Base·2^(N-1)`, no jitter) on 429/502/503/504; honours `Retry-After` and the process-wide cooldown gate |
| `DefaultRetryPolicy()` | 5 attempts, 1s·2s·3s·4s waits, 30s cap | | `DefaultRetryPolicy()` | From env: 7 attempts, 2s base, 2m cap (`OO_RETRY_ATTEMPTS/_BASE/_MAX`) |
| `Transient(err)` | True for retriable OnlyOffice answers | | `Transient(err)` | True for retriable OnlyOffice answers (`*TransientError` or HTTP 429/502/503/504 text) |
### Deep links & conversion
| Method | Description |
|---|---|
| `FileEditorURL(portalBase, id)` / `(c *Client).FileEditorURL(id)` | DocEditor deep link `/Products/Files/DocEditor.aspx?fileid=` |
| `FolderURL(portalBase, id)` / `(c *Client).FolderURL(id)` | Documents folder link |
| `PresignedURI(ctx, fileID)` | Short-lived fetchable URL of a portal file (`presigneduri`) |
| `SignJWT(secret, payload)` | HS256 JWT, stdlib only |
| `ConvertDocument(ctx, docsBase, secret, req)` | OnlyOffice DocumentServer conversion (`/converter`; legacy `/ConvertService.ashx`) |
| `DownloadURLTo(ctx, url, w)` | Stream an absolute URL into a writer |
| `SyncBoard(ctx, board, apply)` | Upsert project milestones/tasks from a YAML board (dry-run when apply=false) |
| `AuditOpportunities(ctx)` | Opportunities with file/task/member counts and a coarse class |
| `WorkbookSheetNames(data)` | Worksheet names of an XLS/XLSX/ODS workbook |
| `WorkbookSheetCSV(data, sheet, delim)` | One worksheet → CSV (sheet-aware, excelize) |
| `WorkbookSheetJSON(data, sheet)` | One worksheet → rows as objects (first row = header) |
Example — convert a portal file (or a local file) to PDF with the native engine:
```bash
export ONLYOFFICE_DS_SECRET=<DocumentServer CoAuthoring secret>
oo docs pdf 1234 --out out.pdf # OO file id → PDF
oo docs pdf ./report.docx --to pdf # local file → PDF (temp upload, auto-cleanup)
```
### Helper Types ### Helper Types
@@ -582,6 +619,41 @@ oo opportunities list
oo opportunities stages oo opportunities stages
oo cases list oo cases list
oo crm-tasks categories oo crm-tasks categories
# Deep links & native document conversion
oo link 1234 2345 # DocEditor URL for file ids (title + url)
oo docs presigned 1234 # short-lived fetchable URL of an OO file
oo docs pdf 1234 --out out.pdf # OO file → PDF (via DocumentServer)
oo docs pdf ./report.docx --to pdf # local file → PDF on the fly (temp upload+cleanup)
oo docs pdf 1234 --stream > out.pdf # pipe: bytes to stdout (alias --pipe)
# Spreadsheet export (sheet-aware; local reader — the DS csv output is first-sheet-only)
oo docs csv 1234 --sheet 2 # XLS/XLSX/ODS worksheet → CSV (--sheet N, 1-based)
oo docs csv 1234 --delimiter ';' # ; | | \t via --delimiter
oo docs json 1234 --sheet 2 # worksheet → JSON rows (first row = header)
oo docs csv ./book.xlsx --sheet 1 --out sheet1.csv
# export ONLYOFFICE_DS_SECRET=<DocumentServer CoAuthoring secret>
# docs base: $ONLYOFFICE_DOCS_URL, else $ONLYOFFICE_URL + /ds-vpath
# Project team (portal users) CRUD
oo projects team list 42
oo projects team add 42 <user_id> [<user_id>...]
oo projects team remove 42 <user_id>
oo projects team set 42 <user_id> [...] # replace team (may lag; verify with list)
oo projects milestone-delete 7
# Project documents: fresh re-upload (single clean version) / new version
oo projects files replace-in 1 ./contract.pdf # hard delete same stem|ext in folder + upload
oo projects files update 1 ./contract.docx # overwrite content, same file id
# Users lifecycle
oo users list ; oo users get <user_id>
oo users create --first Jane --last Doe --email jane.doe@example.com --password '…'
oo users check --login jane.doe@example.com # verify login (email works when userName 500s)
oo users update <user_id> --title "…" --location "…"
oo users block <user_id> ; oo users unblock <user_id>
oo users password <user_id> # reads the new password from stdin
oo users delete <user_id> [<user_id>...]
``` ```
### office (TUI) ### office (TUI)
@@ -678,7 +750,7 @@ oo dav download 22881 --to ./copy.pdf # default path: ./<server title>
oo dav fileops # active move/copy operations (status polling) oo dav fileops # active move/copy operations (status polling)
``` ```
### Search (`oo search`) ### Search and index (`oo search`, `oo index`)
Full-text search over the Documents index. The REST endpoint Full-text search over the Documents index. The REST endpoint
`/api/2.0/files/@search/{query}` only searches file names in the database, so `/api/2.0/files/@search/{query}` only searches file names in the database, so
@@ -697,28 +769,48 @@ oo search "Rechnung" --json # shorthand for -o json
Requires `ONLYOFFICE_ES_URL` (plus optional `ONLYOFFICE_ES_INDEX`, Requires `ONLYOFFICE_ES_URL` (plus optional `ONLYOFFICE_ES_INDEX`,
`ONLYOFFICE_TENANT`). `ONLYOFFICE_TENANT`).
### Bulk tools (`cmd/`) #### PDF/scans: own index (`oo index` + `--backend own`)
Small single-purpose binaries for bulk Documents work. All of them pace The OnlyOffice index covers Office formats only, so PDFs (`S1019`-style invoice
requests and retry transient OnlyOffice answers (429/502/503/504) with a numbers) are not searchable by content. `oo index` extracts PDF text with
deterministic linear backoff — no jitter, same waits on every run `internal/docpipe` (pdftotext, OCR for scans) — including the text of embedded
(see `DoRetry` below). Build with `go build ./cmd/<tool>`. PDF attachments (`pdfdetach`: `<doc>.md`, `.xml`, covers the original/scan and
ZUGFeRD e-invoice XML) — into a separate index (`ONLYOFFICE_ES_TEXT_INDEX`,
default `oo_docs_text`); the OnlyOffice server and its index are **not**
modified. Then search it with `--backend own`.
```bash ```bash
ooscan 659 # recursive index → TSV: file_id, folder_id, path, title oo index folder 634 --recursive --exts pdf # populate (idempotent upsert)
ooscan 659 666 > oo-index.tsv # several roots into one index oo index files 3576 3578 # specific files
pdfamount 671 # "Zu zahlender Betrag" per PDF → TSV: file_id, title, amount oo index folder 634 --dry-run # plan only
kontoblatt 3906 ./kontoblatt.xlsx # summary (Gegenkonto/Monat) uploaded next to source oo search "S1021" --content --backend own # finds the PDF
kontolink IN.xlsx oo-index.tsv OUT.xlsx [FILE_ID] [AMOUNTS_TSV] oo search "Rechnung" --backend own --folder 634 --json
# kontolink writes DocEditor links into the Link column: Beleg → supplier+month
# → amount+date (5th arg = pdfamount output); with FILE_ID it updates the
# source file in place, else uploads an "(links)" copy next to it.
``` ```
See [`docs/elasticsearch.md`](docs/elasticsearch.md) for the decision and
trade-offs.
### Unified file client
All file backends (REST, WebDAV, read-only SQL, Elasticsearch) sit behind one
facade: `c.Files()` returns a `*FileClient` that also implements `FileStore`,
so old call sites keep working. Pick a transport per call with
`c.FileStore("rest"|"dav"|"pg"|"sql")` (SQL is read-only), open the SQL store
with `c.SQLFileStore()`, or register a backend on the facade
(`RegisterStore`/`RegisterSearcher`). Contract, model (`Entry`/`Kind`),
fallback rules, env names and how to add a backend:
[`docs/unified-file-client.md`](docs/unified-file-client.md).
### Bulk tools
Business / one-off bulk tools (`ooscan`, `pdfamount`, `kontoblatt`,
`kontolink`) live in the private `oo-workspace` repo, not in this public
library. They build on the public client and the same `DoRetry` pacing.
| Subject | Verbs | | Subject | Verbs |
|---|---| |---|---|
| `calendar` | `list`, `events`, `add`, `delete` | | `calendar` | `list`, `events`, `add`, `delete` |
| `projects` | `list`, `get`, `milestones`, `milestone-create`, `create`, `update`, `delete`, `contacts` (`add`, `remove`), `link-authors`, `link-git`, **`files`** (`list`, `upload`, `download`, `rename`, `delete`, `dedupe`, `as-md`, `put-md`, `put-txt`, `put-xlsx`) | | `projects` | `list`, `get`, `milestones`, `milestone-create`, `board-sync`, `create`, `update`, `delete`, `contacts` (`add`, `remove`), `link-authors`, `link-git`, **`files`** (`list`, `upload`, `download`, `rename`, `delete`, `dedupe`, `as-md`, `put-md`, `put-txt`, `put-xlsx`) |
| `tasks` | `list`, `get`, `create`, `update`, `delete`, `subtask add`, **`files`** (`list`, `upload`, `detach`) | | `tasks` | `list`, `get`, `create`, `update`, `delete`, `subtask add`, **`files`** (`list`, `upload`, `detach`) |
| `users` | `list`, `self` (alias: `oo whoami`) | | `users` | `list`, `self` (alias: `oo whoami`) |
| `contacts` | `list`, `get`, `delete`, `info-add`, `merge`, `dedupe-info`, `tags`, `tag-add`, `tag-create`, `tag-remove` | | `contacts` | `list`, `get`, `delete`, `info-add`, `merge`, `dedupe-info`, `tags`, `tag-add`, `tag-create`, `tag-remove` |
@@ -726,14 +818,15 @@ kontolink IN.xlsx oo-index.tsv OUT.xlsx [FILE_ID] [AMOUNTS_TSV]
| `companies` | `list`, `create`, `delete`, `dedupe`, `dedupe-persons` | | `companies` | `list`, `create`, `delete`, `dedupe`, `dedupe-persons` |
| `opportunities` | `list`, `get`, `create`, `update`, `delete`, `stages`, `member-add`, `dedupe`, `dedupe-members`, `fix-titles` | | `opportunities` | `list`, `get`, `create`, `update`, `delete`, `stages`, `member-add`, `dedupe`, `dedupe-members`, `fix-titles` |
| `invoices` | `list`, `get`, `create`, `update`, `pdf`, `pdf-cleanup`, `status`, `delete`, `items …` | | `invoices` | `list`, `get`, `create`, `update`, `pdf`, `pdf-cleanup`, `status`, `delete`, `items …` |
| `crm` | `cleanup` | | `crm` | `audit`, `cleanup` |
| `mails` | `accounts`, `folders`, `list`, `get`, `download-attachment`, `draft`, `attach`, `draft-invoice`, `send`, `delete` | | `mails` | `accounts`, `folders`, `list`, `get`, `download-attachment`, `draft`, `attach`, `draft-invoice`, `send`, `delete` |
| `cases` | `list`, `create`, `delete`, `member-add` | | `cases` | `list`, `create`, `delete`, `member-add` |
| `crm-tasks` | `list`, `create`, `delete`, `categories`, `reassign-self` | | `crm-tasks` | `list`, `create`, `delete`, `categories`, `reassign-self` |
| `docs` | `tools`, `convert`, `optimize`, `ocr`, `hocr`, `as-md`, `put-md`, `put-txt`, `put-xlsx` | | `docs` | `tools`, `convert`, `pdf`, `presigned`, `csv`, `json`, `optimize`, `ocr`, `hocr`, `as-md`, `put-md`, `put-txt`, `put-xlsx` |
| `catalog` | `match`, `merge`, `apply`, `scan-contacts`, `scan-projects`, `scan-thunderbird` | | `catalog` | `match`, `merge`, `apply`, `scan-contacts`, `scan-projects`, `scan-thunderbird` |
| `dav` | `ls`, `move`, `copy`, `mkdir`, `rename-file`, `rename-folder`, `download`, `fileops` | | `dav` | `ls`, `move`, `copy`, `mkdir`, `rename-file`, `rename-folder`, `download`, `fileops` |
| `search` | `QUERY` (`--content`, `--folder ID`, `--limit N`, `--json`) | | `search` | `QUERY` (`--content`, `--folder ID`, `--limit N`, `--backend oo\|own`, `--json`) |
| `index` | `folder FOLDER_ID`, `files FILE_ID...` (`--recursive`, `--exts pdf`, `--limit N`, `--dry-run`) |
The CLI reads only `.env` from the current working directory (godotenv is a The CLI reads only `.env` from the current working directory (godotenv is a
CLI-only concern — the library itself never loads dotfiles). CLI-only concern — the library itself never loads dotfiles).
@@ -741,6 +834,12 @@ CLI-only concern — the library itself never loads dotfiles).
Canonical `ONLYOFFICE_*` variables win over aliases. Optional CLI-only aliases: Canonical `ONLYOFFICE_*` variables win over aliases. Optional CLI-only aliases:
`OO_URL` / `OO_USER` / `OO_PASS` → `ONLYOFFICE_URL` / `ONLYOFFICE_USER` / `ONLYOFFICE_PASS`. `OO_URL` / `OO_USER` / `OO_PASS` → `ONLYOFFICE_URL` / `ONLYOFFICE_USER` / `ONLYOFFICE_PASS`.
Catalog scanning (`oo catalog scan-projects` / `scan-thunderbird`) classifies
clients from rules in `$OO_CATALOG_CONFIG` (or `--config`); see
[`catalog/classify.example.yaml`](catalog/classify.example.yaml). Without rules
nothing is classified as work. The MinIO download fallback is off unless
`MINIO_ENDPOINT` + `MINIO_ACCESS_KEY` + `MINIO_SECRET_KEY` are set.
Run `oo --help` or `oo <subject> --help` for the full command reference. Run `oo --help` or `oo <subject> --help` for the full command reference.
> **0.5.0 migration note:** the command tree was flattened per-subject. Old > **0.5.0 migration note:** the command tree was flattened per-subject. Old
@@ -820,8 +919,8 @@ Merge two known company ids (keeps `INTO`):
oo contacts merge FROM_ID INTO_ID oo contacts merge FROM_ID INTO_ID
``` ```
Company ↔ person ↔ deal ↔ project ↔ invoice ↔ mail rules and OO quirks: Company ↔ person ↔ deal ↔ project ↔ invoice ↔ mail rules and OO quirks live
[docs/crm-associations.md](docs/crm-associations.md). with the private `oo-workspace` tooling.
### Invoices (`oo invoices`) ### Invoices (`oo invoices`)
@@ -966,22 +1065,65 @@ oo projects files list 33
| `ONLYOFFICE_CALENDAR_ID` | Default calendar id used when omitted (default `1`) | | `ONLYOFFICE_CALENDAR_ID` | Default calendar id used when omitted (default `1`) |
| `ONLYOFFICE_PROJECT_ID` | Default project id used when omitted (default `33`) | | `ONLYOFFICE_PROJECT_ID` | Default project id used when omitted (default `33`) |
| `OO_URL`, `OO_USER`, `OO_PASS` | Optional CLI-only aliases for `ONLYOFFICE_*` | | `OO_URL`, `OO_USER`, `OO_PASS` | Optional CLI-only aliases for `ONLYOFFICE_*` |
| `OO_RATE_LIMIT` | Process-wide request pacing, req/s (default `4`; `0` disables) |
| `OO_BURST` | Token-bucket burst (default `1`) |
| `OO_RETRY_ATTEMPTS` | Transient retries, total attempts (default `7`) |
| `OO_RETRY_BASE` | Exponential backoff base (default `2s`) |
| `OO_RETRY_MAX` | Backoff cap (default `2m`) |
| `ONLYOFFICE_ES_URL` | Elasticsearch base URL (`oo search`, own index); see [`docs/elasticsearch.md`](docs/elasticsearch.md) |
| `ONLYOFFICE_ES_INDEX` | OnlyOffice index (default `files_file`) |
| `ONLYOFFICE_ES_TEXT_INDEX` | Own PDF/scan index (default `oo_docs_text`) |
| `ONLYOFFICE_TENANT` | `tenantId` filter for ES/SQL (empty = all) |
| `ONLYOFFICE_DSN` | Read-only SQL DSN (MySQL or `postgres://`); see [`docs/community-server-db.md`](docs/community-server-db.md) |
| `ONLYOFFICE_PG_DRIVER`, `ONLYOFFICE_PG_TENANT`, `ONLYOFFICE_PG_HOST/_PORT/_USER/_PASSWORD/_DBNAME/_SSLMODE` | SQL store override / DSN by parts (PostgreSQL) |
| `MINIO_ENDPOINT`, `MINIO_BUCKET`, `MINIO_ACCESS_KEY`, `MINIO_SECRET_KEY` | Object-store layout for SQL `Download` |
| `ONLYOFFICE_WEBDAV_URL` | rclone WebDAV sidecar URL (default `http://172.17.0.1:8098/webdav`) |
Mail and CRM cleanup are documented in [oo CLI use cases](#oo-cli-use-cases) above. Personal disk inventory / dossier sync lives in the private `oo-workspace` (`oow`) tooling. Mail and CRM cleanup are documented in [oo CLI use cases](#oo-cli-use-cases) above. Personal disk inventory / dossier sync lives in the private `oo-workspace` (`oow`) tooling.
### CI / releases ## Testing
GitHub Actions (pattern from [`eSlider/go-config`](https://github.com/eSlider/go-config)): ```bash
go build ./... && go vet ./...
go test ./... # unit — no network, no vendor mocks
go test -race ./...
go test -tags=integration ./... # live OnlyOffice (skip without creds)
```
Unit tests are pure Go (parsers, encoders, conversions). Integration tests
(`//go:build integration`) hit a live instance and **skip** cleanly when the
env is missing, so `go test ./...` stays green offline. New endpoints ship with
an integration test before merge (policy in [`AGENTS.md`](AGENTS.md)).
Live runs need credentials (`ONLYOFFICE_URL`, `ONLYOFFICE_USER`,
`ONLYOFFICE_PASS`) and, per backend:
- **Elasticsearch** (`oo search`, own index) — ES lives on `127.0.0.1:9200`
inside the OnlyOffice VM; expose it over SSH
(`-L 9200:127.0.0.1:9200`) and set `ONLYOFFICE_ES_URL`
(see [`docs/elasticsearch.md`](docs/elasticsearch.md)).
- **SQL backend** (`FileClient`, `SQLFileStore`) — MySQL on `127.0.0.1:3306`
in the same VM; tunnel `-L 3306:127.0.0.1:3306`, then set `ONLYOFFICE_DSN`
(see [`docs/community-server-db.md`](docs/community-server-db.md)).
`ONLYOFFICE_PG_TEST_FILE_ID` / `ONLYOFFICE_PG_TEST_FOLDER_ID` select a real
file for the REST cross-check; `MINIO_*` enable the download check.
### rclone WebDAV mount
The rclone WebDAV mount (compose + `oo-webdav` sidecar) is deployment
tooling and lives in the private `oo-workspace` repo. Set
`ONLYOFFICE_WEBDAV_URL` to point the client at it.
### CI / releases
| Workflow | Trigger | Purpose | | Workflow | Trigger | Purpose |
|---|---|---| |---|---|---|
| `test.yml` | push / PR | `go vet`, unit tests, build `oo` + `office` | | `test.yml` | push / PR | `go vet`, unit tests, build `oo` + `office` |
| `release-please.yml` | push to `main` | semver PR from conventional commits |
| `release.yml` | tag `v*` | GoReleaser cross-platform `oo` + `office` binaries | | `release.yml` | tag `v*` | GoReleaser cross-platform `oo` + `office` binaries |
Repo setting required once: **Settings → Actions → General → Allow GitHub Actions to create and approve pull requests**. Gitea is canonical; tags are created there per SemVer (`fix:` → patch,
`feat:` → minor, `!` → major). GoReleaser publishes assets to
Merge the release-please PR to tag a version; GoReleaser publishes assets to [GitHub Releases](https://github.com/eSlider/go-onlyoffice/releases). [GitHub Releases](https://github.com/eSlider/go-onlyoffice/releases).
## Examples ## Examples
+61 -8
View File
@@ -43,6 +43,11 @@ func (c *Client) Authenticate() error { return c.ensureToken() }
// cached token is still valid it returns immediately; otherwise it performs // cached token is still valid it returns immediately; otherwise it performs
// a POST to /api/2.0/authentication.json that is cancellable via ctx. // a POST to /api/2.0/authentication.json that is cancellable via ctx.
// //
// Transient answers from the edge (openresty 429/502/503/504) are retried with
// the same deterministic policy as every other request (see retry.go), because
// the server rate-limits authentication and the integration suite otherwise
// fails with a raw HTML 429 page.
//
// This is the recommended entry point for long-running syncs (cron, // This is the recommended entry point for long-running syncs (cron,
// watchers) because it guarantees that a stalled auth call will not block // watchers) because it guarantees that a stalled auth call will not block
// the caller past its deadline. // the caller past its deadline.
@@ -50,6 +55,14 @@ func (c *Client) AuthenticateContext(ctx context.Context) error {
if c.tokenValid() { if c.tokenValid() {
return nil return nil
} }
return DoRetry(ctx, DefaultRetryPolicy(), func() error {
return c.authenticateOnce(ctx)
})
}
// authenticateOnce performs a single authentication POST. Callers must handle
// retries; use AuthenticateContext.
func (c *Client) authenticateOnce(ctx context.Context) error {
body, err := json.Marshal(c.credentials) body, err := json.Marshal(c.credentials)
if err != nil { if err != nil {
return fmt.Errorf("marshal credentials: %w", err) return fmt.Errorf("marshal credentials: %w", err)
@@ -69,6 +82,51 @@ func (c *Client) AuthenticateContext(ctx context.Context) error {
if err != nil { if err != nil {
return err return err
} }
if resp.StatusCode >= 400 {
return statusError(resp.StatusCode, retryAfterOf(resp), "auth: %d %s", resp.StatusCode, truncate(string(raw), 400))
}
var env struct {
Response *Token `json:"response"`
}
if err := json.Unmarshal(raw, &env); err != nil {
return fmt.Errorf("auth decode: %w", err)
}
if env.Response == nil || env.Response.Value == "" {
return fmt.Errorf("auth: empty token in response")
}
c.token = env.Response
return nil
}
// AuthenticateAs verifies a login/password pair against the portal WITHOUT
// mutating the client's cached token. It returns nil when the portal issues a
// token, and the portal error otherwise.
//
// OnlyOffice accepts either the userName or the account email as the login. On
// some portals the account userName login returns HTTP 500 "User authentication
// failed" while the account email succeeds — confirmed for a freshly created
// guest user. Use this probe before sharing credentials (see `oo users check`),
// and prefer the email as the login.
func (c *Client) AuthenticateAs(ctx context.Context, login, password string) error {
body, err := json.Marshal(Credentials{User: login, Password: password})
if err != nil {
return fmt.Errorf("marshal credentials: %w", err)
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL()+"/api/2.0/authentication.json", bytes.NewReader(body))
if err != nil {
return err
}
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Accept", "application/json")
resp, err := c.client.Do(req)
if err != nil {
return fmt.Errorf("auth request: %w", err)
}
defer resp.Body.Close()
raw, err := io.ReadAll(resp.Body)
if err != nil {
return err
}
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return fmt.Errorf("auth: %d %s", resp.StatusCode, truncate(string(raw), 400)) return fmt.Errorf("auth: %d %s", resp.StatusCode, truncate(string(raw), 400))
} }
@@ -81,7 +139,6 @@ func (c *Client) AuthenticateContext(ctx context.Context) error {
if env.Response == nil || env.Response.Value == "" { if env.Response == nil || env.Response.Value == "" {
return fmt.Errorf("auth: empty token in response") return fmt.Errorf("auth: empty token in response")
} }
c.token = env.Response
return nil return nil
} }
@@ -99,17 +156,13 @@ func (c *Client) tokenValid() bool {
// ensureToken refreshes the authentication token when missing or expired. // ensureToken refreshes the authentication token when missing or expired.
// Mirrors the logic inline in Query() but is safe to call from helpers that // Mirrors the logic inline in Query() but is safe to call from helpers that
// bypass the typed Request abstraction. // bypass the typed Request abstraction. It shares AuthenticateContext so the
// transient-retry policy applies to every code path.
func (c *Client) ensureToken() error { func (c *Client) ensureToken() error {
if c.tokenValid() { if c.tokenValid() {
return nil return nil
} }
tok, err := c.Auth(c.credentials) return c.AuthenticateContext(context.Background())
if err != nil {
return err
}
c.token = tok
return nil
} }
// authHeader returns the value for the Authorization header, ensuring a token. // authHeader returns the value for the Authorization header, ensuring a token.
+205
View File
@@ -0,0 +1,205 @@
package onlyoffice
// Project board (Gantt) upsert from a YAML board file.
//
// A board describes projects, their milestones and tasks by exact title. Sync
// creates only what is missing: existing milestones/tasks (matched by title)
// are left untouched, so the file can be the source of truth for a project
// plan and re-applied safely. Dry-run (apply=false) reports counts without
// writing.
//
// The file format is deliberately small and presentation-free:
//
// projects:
// - id: 42
// name: "Example"
// milestones:
// - title: "Kickoff"
// deadline: "2026-01-15"
// key: true
// tasks:
// - title: "Draft"
// start: "2026-01-02"
// deadline: "2026-01-10"
// description: "…"
import (
"context"
"fmt"
"os"
"strconv"
"time"
"gopkg.in/yaml.v3"
)
// Board is a YAML mapping from a project plan onto OnlyOffice milestones/tasks.
type Board struct {
Projects []BoardProject `yaml:"projects"`
}
// BoardProject is one project with its milestones.
type BoardProject struct {
ID int `yaml:"id"`
Name string `yaml:"name,omitempty"`
Milestones []BoardMilestone `yaml:"milestones"`
}
// BoardMilestone is a milestone ("key" marks it as a key milestone).
type BoardMilestone struct {
Title string `yaml:"title"`
Deadline string `yaml:"deadline"`
Key bool `yaml:"key,omitempty"`
Tasks []BoardTask `yaml:"tasks"`
}
// BoardTask is a task inside a milestone. Start falls back to Deadline.
type BoardTask struct {
Title string `yaml:"title"`
Start string `yaml:"start,omitempty"`
Deadline string `yaml:"deadline"`
Description string `yaml:"description,omitempty"`
}
// BoardSyncResult counts what SyncBoard created or skipped.
type BoardSyncResult struct {
CreatedMilestones int
SkippedMilestones int
CreatedTasks int
SkippedTasks int
DryRun bool
}
// LoadBoard reads a board YAML file.
func LoadBoard(path string) (*Board, error) {
b, err := os.ReadFile(path)
if err != nil {
return nil, err
}
return ParseBoard(b)
}
// ParseBoard decodes a board from YAML bytes.
func ParseBoard(data []byte) (*Board, error) {
var board Board
if err := yaml.Unmarshal(data, &board); err != nil {
return nil, fmt.Errorf("board: %w", err)
}
if len(board.Projects) == 0 {
return nil, fmt.Errorf("board: no projects")
}
return &board, nil
}
func boardDay(s string) (Time, error) {
t, err := time.Parse("2006-01-02", s)
if err != nil {
return Time{}, err
}
return Time(t), nil
}
func boardMilestoneIDs(ms []*Milestone) map[string]int64 {
out := map[string]int64{}
for _, m := range ms {
if m == nil || m.Title == nil || m.ID == nil {
continue
}
out[*m.Title] = *m.ID
}
return out
}
func boardTaskTitles(rows []map[string]any) map[string]struct{} {
out := map[string]struct{}{}
for _, r := range rows {
if t, _ := r["title"].(string); t != "" {
out[t] = struct{}{}
}
}
return out
}
// SyncBoard upserts milestones and tasks by exact title. With apply=false it
// only counts what would be created.
func (c *Client) SyncBoard(ctx context.Context, board *Board, apply bool) (*BoardSyncResult, error) {
if board == nil || len(board.Projects) == 0 {
return nil, fmt.Errorf("board: no projects")
}
res := &BoardSyncResult{DryRun: !apply}
for _, p := range board.Projects {
pid := p.ID
existing, err := c.GetProjectMilestones(&Project{ID: &pid})
if err != nil {
return res, fmt.Errorf("project %d milestones: %w", pid, err)
}
haveMS := boardMilestoneIDs(existing)
tasks, err := c.ListTasks(ctx, strconv.Itoa(pid), "")
if err != nil {
return res, fmt.Errorf("project %d tasks: %w", pid, err)
}
haveTask := boardTaskTitles(tasks)
for _, m := range p.Milestones {
msID, ok := haveMS[m.Title]
if !ok {
res.CreatedMilestones++
if apply {
dl, err := boardDay(m.Deadline)
if err != nil {
return res, fmt.Errorf("milestone %q deadline: %w", m.Title, err)
}
created, err := c.CreateMilestone(NewMilestoneRequest{
ProjectID: pid,
Title: m.Title,
Deadline: dl,
IsKey: m.Key,
})
if err != nil {
return res, fmt.Errorf("create milestone %q: %w", m.Title, err)
}
if created.ID != nil {
msID = *created.ID
}
haveMS[m.Title] = msID
}
} else {
res.SkippedMilestones++
}
for _, t := range m.Tasks {
if _, exists := haveTask[t.Title]; exists {
res.SkippedTasks++
continue
}
res.CreatedTasks++
if !apply {
continue
}
start := t.Start
if start == "" {
start = t.Deadline
}
st, err := boardDay(start)
if err != nil {
return res, fmt.Errorf("task %q start: %w", t.Title, err)
}
dl, err := boardDay(t.Deadline)
if err != nil {
return res, fmt.Errorf("task %q deadline: %w", t.Title, err)
}
if _, err := c.CreateProjectTask(NewProjectTaskRequest{
ProjectId: pid,
Title: t.Title,
Description: t.Description,
StartDate: st,
Deadline: dl,
MilestoneId: int(msID),
}); err != nil {
return res, fmt.Errorf("create task %q: %w", t.Title, err)
}
haveTask[t.Title] = struct{}{}
}
}
}
return res, nil
}
+70
View File
@@ -0,0 +1,70 @@
package onlyoffice
import "testing"
func TestParseBoard(t *testing.T) {
b, err := ParseBoard([]byte(`
projects:
- id: 13
name: Example
milestones:
- title: "[lq] Test"
deadline: "2026-01-15"
key: true
tasks:
- title: Draft
start: "2026-01-02"
deadline: "2026-01-10"
description: "…"
`))
if err != nil {
t.Fatal(err)
}
if len(b.Projects) != 1 || b.Projects[0].ID != 13 {
t.Fatalf("projects: %+v", b.Projects)
}
ms := b.Projects[0].Milestones[0]
if ms.Title != "[lq] Test" || !ms.Key || ms.Deadline != "2026-01-15" {
t.Fatalf("milestone: %+v", ms)
}
if len(ms.Tasks) != 1 || ms.Tasks[0].Title != "Draft" || ms.Tasks[0].Start != "2026-01-02" {
t.Fatalf("task: %+v", ms.Tasks)
}
}
func TestParseBoardEmpty(t *testing.T) {
if _, err := ParseBoard([]byte("projects: []")); err == nil {
t.Fatal("expected error for empty board")
}
}
func TestBoardMilestoneIDs(t *testing.T) {
title := "[lq] Test"
id := int64(9)
got := boardMilestoneIDs([]*Milestone{{Title: &title, ID: &id}, nil})
if got[title] != 9 {
t.Fatalf("%v", got)
}
}
func TestBoardTaskTitles(t *testing.T) {
got := boardTaskTitles([]map[string]any{{"title": "a"}, {"title": "b"}, {"nope": 1}})
if _, ok := got["a"]; !ok {
t.Fatal("missing a")
}
if _, ok := got["b"]; !ok {
t.Fatal("missing b")
}
if len(got) != 2 {
t.Fatalf("%v", got)
}
}
func TestBoardDay(t *testing.T) {
if _, err := boardDay("2026-01-15"); err != nil {
t.Fatal(err)
}
if _, err := boardDay("15.01.2026"); err == nil {
t.Fatal("expected error for non-ISO date")
}
}
+28
View File
@@ -0,0 +1,28 @@
package catalog
import "testing"
func TestMergeAddressesDedup(t *testing.T) {
dst := []Address{{Street: "Weg 1", City: "Stadt", Zip: "1"}}
got := mergeAddresses(dst, []Address{
{Street: "weg 1", City: "stadt", Zip: "1"}, // duplicate (case-insensitive)
{Street: "Weg 2", City: "Stadt", Zip: "2"}, // new
})
if len(got) != 2 {
t.Fatalf("got %d addresses, want 2: %+v", len(got), got)
}
if got[1].Street != "Weg 2" {
t.Errorf("second = %+v", got[1])
}
}
func TestEntryAddressesYAML(t *testing.T) {
doc := &Document{Entries: []Entry{{
ID: "person:max maier", Kind: "person", Name: "Max Maier",
Addresses: []Address{{Street: "Weg 1", City: "Stadt", Zip: "12345", Category: "Billing", Primary: true}},
}}}
merged := MergeDocs(doc)
if len(merged.Entries) != 1 || len(merged.Entries[0].Addresses) != 1 {
t.Fatalf("merged = %+v", merged.Entries)
}
}
+25 -15
View File
@@ -115,20 +115,13 @@ func applyCompany(ctx context.Context, client *onlyoffice.Client, e *Entry) (boo
} }
func applyPerson(ctx context.Context, client *onlyoffice.Client, e *Entry) (bool, error) { func applyPerson(ctx context.Context, client *onlyoffice.Client, e *Entry) (bool, error) {
first := strings.TrimSpace(e.First) org := strings.TrimSpace(e.Org)
last := strings.TrimSpace(e.Last) first, last := CleanPersonNames(e.First, e.Last, e.Name, org, e.Emails)
if first == "" && last == "" { e.First, e.Last = first, last
first, last = SplitDisplayName(e.Name)
}
if first == "" {
first = strings.TrimSpace(e.Name)
}
if first == "" { if first == "" {
return false, fmt.Errorf("person missing name") return false, fmt.Errorf("person missing name")
} }
if last == "" { e.Name = strings.TrimSpace(first + " " + strings.Trim(last, "-"))
last = "-"
}
var p map[string]any var p map[string]any
var err error var err error
@@ -144,21 +137,27 @@ func applyPerson(ctx context.Context, client *onlyoffice.Client, e *Entry) (bool
} }
created := false created := false
companyID := 0 companyID := 0
if e.Org != "" { if org != "" {
if co, ferr := client.FindCompany(ctx, e.Org); ferr == nil && co != nil { if co, ferr := client.FindCompany(ctx, org); ferr == nil && co != nil {
companyID, _ = strconv.Atoi(contactIDString(co)) companyID, _ = strconv.Atoi(contactIDString(co))
} }
} }
if p == nil { if p == nil {
about := "" about := ""
if e.Org != "" { if org != "" {
about = "org: " + e.Org about = "org: " + org
} }
p, err = client.CreatePerson(ctx, first, last, companyID, "", about) p, err = client.CreatePerson(ctx, first, last, companyID, "", about)
if err != nil { if err != nil {
return false, err return false, err
} }
created = true created = true
} else {
// Repair names + ensure company link (never encode company in lastName).
id := contactIDString(p)
if _, err := client.UpdatePerson(ctx, id, first, last, companyID, "", ""); err != nil {
return false, fmt.Errorf("update person %s: %w", id, err)
}
} }
id := contactIDString(p) id := contactIDString(p)
e.OOID = id e.OOID = id
@@ -195,5 +194,16 @@ func ensureContactInfos(ctx context.Context, client *onlyoffice.Client, contactI
return fmt.Errorf("add phone %s: %w", ph, err) return fmt.Errorf("add phone %s: %w", ph, err)
} }
} }
for _, a := range e.Addresses {
if strings.TrimSpace(a.Street) == "" && strings.TrimSpace(a.City) == "" && strings.TrimSpace(a.Zip) == "" {
continue
}
if onlyoffice.HasContactAddress(existing, a.Street, a.City, a.Zip, a.Category) {
continue
}
if _, err := client.AddContactAddress(ctx, contactID, a.Street, a.City, a.State, a.Zip, a.Country, a.Category, a.Primary); err != nil {
return fmt.Errorf("add address %s: %w", strings.TrimSpace(a.Street), err)
}
}
return nil return nil
} }
+17 -5
View File
@@ -100,26 +100,38 @@ END:VCARD
func TestScanProjectsRoot(t *testing.T) { func TestScanProjectsRoot(t *testing.T) {
root := t.TempDir() root := t.TempDir()
repo := filepath.Join(root, "produktor-demo") repo := filepath.Join(root, "acme-demo")
if err := os.MkdirAll(filepath.Join(repo, ".git"), 0o755); err != nil { if err := os.MkdirAll(filepath.Join(repo, ".git"), 0o755); err != nil {
t.Fatal(err) t.Fatal(err)
} }
doc, err := ScanProjectsRoot(root, 3) cl := &Classifier{WorkNames: []string{"acme"}}
doc, err := ScanProjectsRootOpts(root, 3, ScanOptions{Classifier: cl})
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
found := false found := false
for _, e := range doc.Entries { for _, e := range doc.Entries {
if e.Kind == "company" && e.Name == "produktor-demo" { if e.Kind == "company" && e.Name == "acme-demo" {
found = true found = true
if e.Role != "work" { if e.Role != "work" || e.Zone != "warm" {
t.Fatalf("role=%q", e.Role) t.Fatalf("role=%q zone=%q", e.Role, e.Zone)
} }
} }
} }
if !found { if !found {
t.Fatalf("missing company: %+v", doc.Entries) t.Fatalf("missing company: %+v", doc.Entries)
} }
// Neutral default leaves it unclassified.
doc, err = ScanProjectsRoot(root, 3)
if err != nil {
t.Fatal(err)
}
for _, e := range doc.Entries {
if e.Kind == "company" && e.Name == "acme-demo" && e.Role != "unknown" {
t.Fatalf("neutral role=%q", e.Role)
}
}
} }
func TestEntryID(t *testing.T) { func TestEntryID(t *testing.T) {
+26
View File
@@ -0,0 +1,26 @@
# Deployment classification rules for `oo catalog scan-projects` /
# `oo catalog scan-thunderbird`. Point OO_CATALOG_CONFIG (or --config) at a copy
# of this file. Keep your real rules out of the repository — they name your
# clients and hosts. The library defaults to no rules (nothing is "work").
#
# work_remotes: a git remote containing any of these substrings → work/hot.
work_remotes:
- git.internal.example
- github.com/acme
#
# work_names: a project directory name containing any of these substrings → work/warm.
work_names:
- acme
#
# mail_orgs: ordered rules for Thunderbird/mbox identities; the first match wins.
# Match by exact `domain`, `suffix` (e.g. ".example.com") and/or display `name`.
# `zone` defaults to hot, `role` to work.
mail_orgs:
- domain: acme.example
org: Acme GmbH
zone: hot
role: work
- suffix: .gov.example
org: Public Sector
zone: warm
role: work
+127
View File
@@ -0,0 +1,127 @@
package catalog
import (
"fmt"
"os"
"strings"
"gopkg.in/yaml.v3"
)
// Classifier maps project trees and mail identities to catalog org/zone/role.
//
// The library ships with neutral defaults: nothing is classified as work unless
// the deployment supplies rules. Those rules are deployment-specific, so they
// live in a YAML config file (path from OO_CATALOG_CONFIG or the --config flag),
// not in the code. See catalog/classify.example.yaml.
type Classifier struct {
// WorkRemotes: a git remote containing any of these substrings → work/hot.
WorkRemotes []string `yaml:"work_remotes,omitempty"`
// WorkNames: a project name containing any of these substrings → work/warm.
WorkNames []string `yaml:"work_names,omitempty"`
// MailOrgs: ordered mail-identity rules; the first match wins.
MailOrgs []MailRule `yaml:"mail_orgs,omitempty"`
}
// MailRule maps an email domain and/or a display-name substring to an org with
// a zone/role. At least one of Domain, Suffix or Name must be set.
type MailRule struct {
Domain string `yaml:"domain,omitempty"` // exact domain, case-insensitive
Suffix string `yaml:"suffix,omitempty"` // domain suffix, e.g. ".example.com"
Name string `yaml:"name,omitempty"` // substring of the display name
Org string `yaml:"org"`
Zone string `yaml:"zone,omitempty"` // default "hot"
Role string `yaml:"role,omitempty"` // default "work"
}
// DefaultClassifier returns the neutral classifier (no deployment rules).
func DefaultClassifier() *Classifier { return &Classifier{} }
// LoadClassifier reads a classifier config from a YAML file.
func LoadClassifier(path string) (*Classifier, error) {
b, err := os.ReadFile(path)
if err != nil {
return nil, err
}
var c Classifier
if err := yaml.Unmarshal(b, &c); err != nil {
return nil, fmt.Errorf("parse classifier config %s: %w", path, err)
}
return &c, nil
}
// LoadClassifierFromEnv loads the classifier named by OO_CATALOG_CONFIG. An
// empty variable yields the neutral classifier.
func LoadClassifierFromEnv() (*Classifier, error) {
path := strings.TrimSpace(os.Getenv("OO_CATALOG_CONFIG"))
if path == "" {
return DefaultClassifier(), nil
}
return LoadClassifier(path)
}
// ClassifyProject returns (role, zone) for a project name and git remote.
// Generic name heuristics come first; deployment rules supply the work cases.
func (c *Classifier) ClassifyProject(name, remote string) (role, zone string) {
lower := strings.ToLower(name)
remoteL := strings.ToLower(remote)
switch {
case strings.Contains(lower, "experiment") || strings.HasPrefix(lower, "test"):
return "experiment", "cold"
case lower == "mama" || lower == "personal" || strings.Contains(lower, "private"):
return "personal", "private"
}
if c != nil {
for _, r := range c.WorkRemotes {
if r != "" && strings.Contains(remoteL, strings.ToLower(r)) {
return "work", "hot"
}
}
for _, n := range c.WorkNames {
if n != "" && strings.Contains(lower, strings.ToLower(n)) {
return "work", "warm"
}
}
}
return "unknown", "warm"
}
// ClassifyMail returns (org, zone, role) for a mail identity. name is the
// display name (may be empty); email is the address.
func (c *Classifier) ClassifyMail(name, email string) (org, zone, role string) {
em := NormalizeEmail(email)
_, domain, _ := strings.Cut(em, "@")
nameL := strings.ToLower(strings.TrimSpace(name))
if c != nil {
for _, r := range c.MailOrgs {
if !mailRuleMatches(r, domain, nameL) {
continue
}
z, ro := r.Zone, r.Role
if z == "" {
z = "hot"
}
if ro == "" {
ro = "work"
}
return r.Org, z, ro
}
}
if strings.HasSuffix(domain, ".de") && looksPublicSector(domain) {
return domain, "warm", "work"
}
return "", "private", "unknown"
}
func mailRuleMatches(r MailRule, domain, nameL string) bool {
if r.Domain != "" && domain == strings.ToLower(strings.TrimSpace(r.Domain)) {
return true
}
if r.Suffix != "" && strings.HasSuffix(domain, strings.ToLower(strings.TrimSpace(r.Suffix))) {
return true
}
if r.Name != "" && nameL != "" && strings.Contains(nameL, strings.ToLower(strings.TrimSpace(r.Name))) {
return true
}
return false
}
+64
View File
@@ -0,0 +1,64 @@
package catalog
import (
"os"
"path/filepath"
"testing"
)
func TestClassifierRules(t *testing.T) {
cl := &Classifier{
WorkRemotes: []string{"git.internal.example"},
WorkNames: []string{"acme"},
MailOrgs: []MailRule{
{Domain: "acme.example", Org: "Acme", Zone: "hot", Role: "work"},
{Suffix: ".gov.example", Org: "Public", Zone: "warm", Role: "work"},
},
}
if role, zone := cl.ClassifyProject("acme-app", "git@git.internal.example:team/acme-app.git"); role != "work" || zone != "hot" {
t.Fatalf("remote: %s/%s", role, zone)
}
if role, zone := cl.ClassifyProject("acme-demo", ""); role != "work" || zone != "warm" {
t.Fatalf("name: %s/%s", role, zone)
}
if role, zone := cl.ClassifyProject("experiment-x", ""); role != "experiment" || zone != "cold" {
t.Fatalf("experiment: %s/%s", role, zone)
}
if org, zone, role := cl.ClassifyMail("", "bob@acme.example"); org != "Acme" || zone != "hot" || role != "work" {
t.Fatalf("mail: %s/%s/%s", org, zone, role)
}
if org, _, _ := cl.ClassifyMail("X", "x@team.gov.example"); org != "Public" {
t.Fatalf("suffix: %s", org)
}
if org, zone, role := cl.ClassifyMail("", "someone@unknown.example"); org != "" || zone != "private" || role != "unknown" {
t.Fatalf("neutral: %s/%s/%s", org, zone, role)
}
}
func TestLoadClassifier(t *testing.T) {
dir := t.TempDir()
path := filepath.Join(dir, "classify.yaml")
if err := os.WriteFile(path, []byte("work_names:\n - acme\nmail_orgs:\n - domain: acme.example\n org: Acme\n"), 0o644); err != nil {
t.Fatal(err)
}
cl, err := LoadClassifier(path)
if err != nil {
t.Fatal(err)
}
if role, _ := cl.ClassifyProject("acme-app", ""); role != "work" {
t.Fatalf("role=%s", role)
}
if org, zone, _ := cl.ClassifyMail("", "a@acme.example"); org != "Acme" || zone != "hot" {
t.Fatalf("org=%s zone=%s", org, zone)
}
t.Setenv("OO_CATALOG_CONFIG", path)
if envCl, err := LoadClassifierFromEnv(); err != nil || envCl == nil || len(envCl.WorkNames) == 0 {
t.Fatalf("env: %+v %v", envCl, err)
}
t.Setenv("OO_CATALOG_CONFIG", "")
neutral, err := LoadClassifierFromEnv()
if err != nil || len(neutral.WorkNames) != 0 || len(neutral.MailOrgs) != 0 {
t.Fatalf("neutral: %+v %v", neutral, err)
}
}
+9 -4
View File
@@ -27,7 +27,7 @@ func MatchAgainstOO(ctx context.Context, client *onlyoffice.Client, doc *Documen
byEmail[NormalizeEmail(em)] = c byEmail[NormalizeEmail(em)] = c
} }
if isCo { if isCo {
key := NormalizeName(fmt.Sprint(c["displayName"])) key := onlyoffice.CompanyGroupingKey(fmt.Sprint(c["displayName"]))
if key != "" { if key != "" {
byCompanyName[key] = c byCompanyName[key] = c
} }
@@ -63,7 +63,7 @@ func MatchAgainstOO(ctx context.Context, client *onlyoffice.Client, doc *Documen
} }
if !matched { if !matched {
if e.Kind == "company" { if e.Kind == "company" {
if c, ok := byCompanyName[NormalizeName(e.Name)]; ok { if c, ok := byCompanyName[onlyoffice.CompanyGroupingKey(e.Name)]; ok {
oo = c oo = c
matched = true matched = true
} }
@@ -92,6 +92,12 @@ func MatchAgainstOO(ctx context.Context, client *onlyoffice.Client, doc *Documen
e.OOID = contactIDString(oo) e.OOID = contactIDString(oo)
continue continue
} }
// Keep a previously applied oo_id (list payloads often omit emails, so
// email match can miss persons that already exist in CRM).
if strings.TrimSpace(e.OOID) != "" {
e.Status = "exists"
continue
}
e.Status = "new" e.Status = "new"
e.OOID = "" e.OOID = ""
} }
@@ -120,8 +126,7 @@ func contactEmails(c map[string]any) []string {
out = append(out, em) out = append(out, em)
} }
for _, row := range onlyoffice.ContactInfoRows(c) { for _, row := range onlyoffice.ContactInfoRows(c) {
t := strings.ToLower(fmt.Sprint(row["infoType"])) if onlyoffice.NormalizeContactInfoType(fmt.Sprint(row["infoType"])) != "email" {
if t != "email" {
continue continue
} }
data := strings.TrimSpace(fmt.Sprint(row["data"])) data := strings.TrimSpace(fmt.Sprint(row["data"]))
+11 -5
View File
@@ -16,6 +16,8 @@ type ScanOptions struct {
MboxHeaders bool MboxHeaders bool
// MboxMaxBytes skips individual mbox files larger than this (0 = 256 MiB default). // MboxMaxBytes skips individual mbox files larger than this (0 = 256 MiB default).
MboxMaxBytes int64 MboxMaxBytes int64
// Classifier supplies deployment classification rules; nil → neutral default.
Classifier *Classifier
} }
// ScanThunderbirdRoot finds Thunderbird profiles under root and emits person rows // ScanThunderbirdRoot finds Thunderbird profiles under root and emits person rows
@@ -37,6 +39,10 @@ func ScanThunderbirdRootOpts(root string, opts ScanOptions) (*Document, error) {
if opts.MboxMaxBytes <= 0 { if opts.MboxMaxBytes <= 0 {
opts.MboxMaxBytes = 256 << 20 opts.MboxMaxBytes = 256 << 20
} }
cl := opts.Classifier
if cl == nil {
cl = DefaultClassifier()
}
var entries []Entry var entries []Entry
seenDB := map[string]struct{}{} seenDB := map[string]struct{}{}
@@ -64,7 +70,7 @@ func ScanThunderbirdRootOpts(root string, opts ScanOptions) (*Document, error) {
return nil return nil
} }
seenDB[path] = struct{}{} seenDB[path] = struct{}{}
parsed, perr := parseGlodaContacts(path) parsed, perr := parseGlodaContacts(path, cl)
if perr != nil { if perr != nil {
entries = append(entries, Entry{ entries = append(entries, Entry{
ID: EntryID("person", "", filepath.Base(path)), ID: EntryID("person", "", filepath.Base(path)),
@@ -84,7 +90,7 @@ func ScanThunderbirdRootOpts(root string, opts ScanOptions) (*Document, error) {
return nil return nil
} }
seenMAB[path] = struct{}{} seenMAB[path] = struct{}{}
parsed, perr := parseMABEmails(path) parsed, perr := parseMABEmails(path, cl)
if perr != nil { if perr != nil {
return nil return nil
} }
@@ -101,7 +107,7 @@ func ScanThunderbirdRootOpts(root string, opts ScanOptions) (*Document, error) {
if info.Size() > opts.MboxMaxBytes { if info.Size() > opts.MboxMaxBytes {
return nil return nil
} }
parsed, perr := parseMboxHeaderEmails(path) parsed, perr := parseMboxHeaderEmails(path, cl)
if perr != nil { if perr != nil {
return nil return nil
} }
@@ -150,7 +156,7 @@ func isLikelyMboxFile(name, path string) bool {
} }
// parseMboxHeaderEmails extracts addresses from From/To/Cc/Reply-To headers only. // parseMboxHeaderEmails extracts addresses from From/To/Cc/Reply-To headers only.
func parseMboxHeaderEmails(path string) ([]Entry, error) { func parseMboxHeaderEmails(path string, cl *Classifier) ([]Entry, error) {
f, err := os.Open(path) f, err := os.Open(path)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -220,7 +226,7 @@ func parseMboxHeaderEmails(path string) ([]Entry, error) {
var out []Entry var out []Entry
for em := range emails { for em := range emails {
org, zone, role := classifyMailIdentity("", em) org, zone, role := cl.ClassifyMail("", em)
out = append(out, Entry{ out = append(out, Entry{
ID: EntryID("person", em, ""), ID: EntryID("person", em, ""),
Kind: "person", Kind: "person",
+7 -7
View File
@@ -10,22 +10,22 @@ func TestParseMboxHeaderEmails(t *testing.T) {
dir := t.TempDir() dir := t.TempDir()
path := filepath.Join(dir, "INBOX") path := filepath.Join(dir, "INBOX")
body := `From - Mon Jul 1 00:00:00 2016 body := `From - Mon Jul 1 00:00:00 2016
From: Axel Schaefer <axel.schaefer@wheregroup.com> From: Alice Smith <alice.smith@acme.example>
To: Andriy Oblivantsev <andriy.oblivantsev@wheregroup.com> To: Bob Jones <bob.jones@acme.example>
Cc: noreply@example.com, client@stadt-example.de Cc: noreply@example.com, client@stadt-example.de
Subject: test Subject: test
Body line ignored Body line ignored
From - Mon Jul 2 00:00:00 2016 From - Mon Jul 2 00:00:00 2016
From: Someone <paul.schmidt@wheregroup.com> From: Someone <paul.schmidt@acme.example>
To: list@wheregroup.com To: list@acme.example
more body more body
` `
if err := os.WriteFile(path, []byte(body), 0o644); err != nil { if err := os.WriteFile(path, []byte(body), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
} }
ents, err := parseMboxHeaderEmails(path) ents, err := parseMboxHeaderEmails(path, DefaultClassifier())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -35,7 +35,7 @@ more body
got[e.Emails[0]] = true got[e.Emails[0]] = true
} }
} }
if !got["axel.schaefer@wheregroup.com"] || !got["andriy.oblivantsev@wheregroup.com"] { if !got["alice.smith@acme.example"] || !got["bob.jones@acme.example"] {
t.Fatalf("%v", got) t.Fatalf("%v", got)
} }
if got["noreply@example.com"] { if got["noreply@example.com"] {
@@ -53,7 +53,7 @@ func TestScanThunderbirdRootOptsMbox(t *testing.T) {
t.Fatal(err) t.Fatal(err)
} }
if err := os.WriteFile(filepath.Join(imap, "INBOX"), []byte( if err := os.WriteFile(filepath.Join(imap, "INBOX"), []byte(
"From - x\nFrom: a@wheregroup.com\nTo: b@wheregroup.com\n\nbody\n", "From - x\nFrom: a@acme.example\nTo: b@acme.example\n\nbody\n",
), 0o644); err != nil { ), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
} }
+152
View File
@@ -0,0 +1,152 @@
package catalog
import (
"regexp"
"strings"
"unicode"
)
var (
parenSuffixRE = regexp.MustCompile(`(?i)\s*[\(\[\{][^)\]\}]*[\)\]\}]\s*$`)
dashCompanyRE = regexp.MustCompile(`(?i)\s+[-–—]\s+[A-Za-z0-9].*$`)
emailLocalRE = regexp.MustCompile(`(?i)^[a-z0-9._%+\-]+@[a-z0-9.\-]+\.[a-z]{2,}$`)
nonNameTokenRE = regexp.MustCompile(`[^a-zA-ZÀ-öø-ÿĀ-ž0-9'’.\-]+`)
)
// CleanPersonNames strips company annotations from display names and fills
// first/last from the email local-part when the source used an address as the
// name. Company affiliation belongs on Org / the CRM companyId — never in LastName.
func CleanPersonNames(first, last, display, org string, emails []string) (cleanFirst, cleanLast string) {
first = strings.TrimSpace(first)
last = strings.TrimSpace(last)
display = strings.TrimSpace(display)
org = strings.TrimSpace(org)
if looksLikeEmail(first) {
ef, el := GuessNameFromEmail(first)
first, last = ef, el
}
if looksLikeEmail(display) && first == "" && last == "" {
display = ""
}
if first == "" && last == "" && display != "" {
first, last = SplitDisplayName(display)
}
first = stripCompanyAnnotation(first, org)
last = stripCompanyAnnotation(last, org)
// "Smith - Acme" / "Jones (Acme)" landed in last.
last = stripCompanyAnnotation(last, org)
if i := strings.IndexAny(first, "(["); i > 0 {
first = strings.TrimSpace(first[:i])
}
// Entire last name is just the company (e.g. last="Acme").
if org != "" && personLastIsOrg(last, org) {
last = ""
}
if (first == "" || looksLikeEmail(first)) && len(emails) > 0 {
ef, el := GuessNameFromEmail(emails[0])
if first == "" || looksLikeEmail(first) {
first = ef
}
if last == "" || last == "-" {
last = el
}
}
first = strings.TrimSpace(first)
last = strings.TrimSpace(last)
if last == "" {
last = "-"
}
return first, last
}
func stripCompanyAnnotation(s, org string) string {
s = strings.TrimSpace(s)
if s == "" {
return ""
}
s = parenSuffixRE.ReplaceAllString(s, "")
s = strings.TrimSpace(s)
s = dashCompanyRE.ReplaceAllString(s, "")
s = strings.TrimSpace(s)
if org != "" {
for _, sep := range []string{" - ", " – ", " — ", " / "} {
if i := strings.LastIndex(strings.ToLower(s), strings.ToLower(sep+org)); i >= 0 {
s = strings.TrimSpace(s[:i])
}
}
suf := " (" + org + ")"
if strings.HasSuffix(strings.ToLower(s), strings.ToLower(suf)) {
s = strings.TrimSpace(s[:len(s)-len(suf)])
}
}
return strings.TrimSpace(s)
}
func personLastIsOrg(last, org string) bool {
last = NormalizeName(last)
org = NormalizeName(org)
if last == "" || org == "" {
return false
}
if last == org {
return true
}
// "Acme" vs "Acme GmbH & Co. KG"
return strings.HasPrefix(org, last+" ") || strings.HasPrefix(org, last+",")
}
func looksLikeEmail(s string) bool {
return emailLocalRE.MatchString(strings.TrimSpace(s))
}
// GuessNameFromEmail turns local@domain into Title-Case first/last when the
// local part looks like first.last / first_last / first-last.
func GuessNameFromEmail(email string) (first, last string) {
email = NormalizeEmail(email)
local, _, ok := strings.Cut(email, "@")
if !ok || local == "" {
return "", ""
}
local = strings.Split(local, "+")[0]
parts := strings.FieldsFunc(local, func(r rune) bool {
return r == '.' || r == '_' || r == '-'
})
if len(parts) == 0 {
return titleToken(local), ""
}
if len(parts) == 1 {
return titleToken(parts[0]), ""
}
return titleToken(parts[0]), titleToken(strings.Join(parts[1:], " "))
}
func titleToken(s string) string {
s = nonNameTokenRE.ReplaceAllString(s, " ")
s = strings.TrimSpace(s)
if s == "" {
return ""
}
runes := []rune(strings.ToLower(s))
runes[0] = unicode.ToTitle(runes[0])
return string(runes)
}
// FormatProjectTitle builds "CC | Company | Title" (spaces around |).
// Country should be a short code (DE, TF, UA, …). Empty segments are dropped.
func FormatProjectTitle(country, company, title string) string {
parts := make([]string, 0, 3)
for _, p := range []string{country, company, title} {
p = strings.TrimSpace(p)
p = strings.ReplaceAll(p, "|", "/")
if p != "" {
parts = append(parts, p)
}
}
return strings.Join(parts, " | ")
}
+41
View File
@@ -0,0 +1,41 @@
package catalog
import "testing"
func TestCleanPersonNamesStripsCompanyParen(t *testing.T) {
f, l := CleanPersonNames("John", "Smith (Acme)", "John Smith (Acme)", "Acme", nil)
if f != "John" || l != "Smith" {
t.Fatalf("got %q %q", f, l)
}
}
func TestCleanPersonNamesStripsDashCompany(t *testing.T) {
f, l := CleanPersonNames("Jens", "Meyer - Acme", "", "Acme", nil)
if f != "Jens" || l != "Meyer" {
t.Fatalf("got %q %q", f, l)
}
}
func TestCleanPersonNamesFromEmail(t *testing.T) {
f, l := CleanPersonNames("david.patzke@acme.example", "-", "", "Acme",
[]string{"david.patzke@acme.example"})
if f != "David" || l != "Patzke" {
t.Fatalf("got %q %q", f, l)
}
}
func TestCleanPersonNamesLastIsCompany(t *testing.T) {
f, l := CleanPersonNames("Thorsten", "Acme", "", "Acme GmbH & Co. KG", nil)
if f != "Thorsten" || l != "-" {
t.Fatalf("got %q %q", f, l)
}
}
func TestFormatProjectTitle(t *testing.T) {
if got := FormatProjectTitle("DE", "Acme", "Golang"); got != "DE | Acme | Golang" {
t.Fatalf("got %q", got)
}
if got := FormatProjectTitle("", "Acme", ""); got != "Acme" {
t.Fatalf("got %q", got)
}
}
+14 -24
View File
@@ -10,10 +10,19 @@ import (
// ScanProjectsRoot finds git roots under root (max depth) and emits company rows. // ScanProjectsRoot finds git roots under root (max depth) and emits company rows.
func ScanProjectsRoot(root string, maxDepth int) (*Document, error) { func ScanProjectsRoot(root string, maxDepth int) (*Document, error) {
return ScanProjectsRootOpts(root, maxDepth, ScanOptions{})
}
// ScanProjectsRootOpts is ScanProjectsRoot with deployment classification rules.
func ScanProjectsRootOpts(root string, maxDepth int, opts ScanOptions) (*Document, error) {
root = filepath.Clean(root) root = filepath.Clean(root)
if maxDepth <= 0 { if maxDepth <= 0 {
maxDepth = 4 maxDepth = 4
} }
cl := opts.Classifier
if cl == nil {
cl = DefaultClassifier()
}
st, err := os.Stat(root) st, err := os.Stat(root)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -23,7 +32,7 @@ func ScanProjectsRoot(root string, maxDepth int) (*Document, error) {
} }
var entries []Entry var entries []Entry
err = walkGitRoots(root, root, 0, maxDepth, &entries) err = walkGitRoots(root, root, 0, maxDepth, cl, &entries)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -43,7 +52,7 @@ func ScanProjectsRoot(root string, maxDepth int) (*Document, error) {
} }
path := filepath.Join(root, name) path := filepath.Join(root, name)
id := EntryID("company", "", name) id := EntryID("company", "", name)
role, zone := classifyProjectName(name, "") role, zone := cl.ClassifyProject(name, "")
entries = append(entries, Entry{ entries = append(entries, Entry{
ID: id, ID: id,
Kind: "company", Kind: "company",
@@ -59,7 +68,7 @@ func ScanProjectsRoot(root string, maxDepth int) (*Document, error) {
return MergeDocs(&Document{Entries: entries}), nil return MergeDocs(&Document{Entries: entries}), nil
} }
func walkGitRoots(root, dir string, depth, maxDepth int, out *[]Entry) error { func walkGitRoots(root, dir string, depth, maxDepth int, cl *Classifier, out *[]Entry) error {
if depth > maxDepth { if depth > maxDepth {
return nil return nil
} }
@@ -67,7 +76,7 @@ func walkGitRoots(root, dir string, depth, maxDepth int, out *[]Entry) error {
if st, err := os.Stat(filepath.Join(dir, ".git")); err == nil && (st.IsDir() || st.Mode().IsRegular()) { if st, err := os.Stat(filepath.Join(dir, ".git")); err == nil && (st.IsDir() || st.Mode().IsRegular()) {
name := filepath.Base(dir) name := filepath.Base(dir)
remote := gitRemoteOrigin(dir) remote := gitRemoteOrigin(dir)
role, zone := classifyProjectName(name, remote) role, zone := cl.ClassifyProject(name, remote)
*out = append(*out, Entry{ *out = append(*out, Entry{
ID: EntryID("company", "", name), ID: EntryID("company", "", name),
Kind: "company", Kind: "company",
@@ -93,7 +102,7 @@ func walkGitRoots(root, dir string, depth, maxDepth int, out *[]Entry) error {
if name == ".git" || name == "node_modules" || name == "vendor" || name == ".venv" || name == "dist" { if name == ".git" || name == "node_modules" || name == "vendor" || name == ".venv" || name == "dist" {
continue continue
} }
_ = walkGitRoots(root, filepath.Join(dir, name), depth+1, maxDepth, out) _ = walkGitRoots(root, filepath.Join(dir, name), depth+1, maxDepth, cl, out)
} }
return nil return nil
} }
@@ -106,22 +115,3 @@ func gitRemoteOrigin(dir string) string {
} }
return strings.TrimSpace(string(b)) return strings.TrimSpace(string(b))
} }
func classifyProjectName(name, remote string) (role, zone string) {
lower := strings.ToLower(name)
remoteL := strings.ToLower(remote)
switch {
case strings.Contains(lower, "experiment") || strings.HasPrefix(lower, "test"):
return "experiment", "cold"
case lower == "mama" || lower == "personal" || strings.Contains(lower, "private"):
return "personal", "private"
case strings.Contains(remoteL, "git.produktor.io") || strings.Contains(remoteL, "github.com/eslider"):
return "work", "hot"
case strings.Contains(lower, "produktor") || strings.Contains(lower, "eslider") ||
strings.Contains(lower, "asesoria") || strings.Contains(lower, "dyvenia") ||
strings.Contains(lower, "onlyoffice"):
return "work", "warm"
default:
return "unknown", "warm"
}
}
+4 -23
View File
@@ -64,7 +64,7 @@ func noisyEmail(email string) bool {
return false return false
} }
func parseMABEmails(path string) ([]Entry, error) { func parseMABEmails(path string, cl *Classifier) ([]Entry, error) {
b, err := os.ReadFile(path) b, err := os.ReadFile(path)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -76,7 +76,7 @@ func parseMABEmails(path string) ([]Entry, error) {
if noisyEmail(em) { if noisyEmail(em) {
continue continue
} }
org, zone, role := classifyMailIdentity("", em) org, zone, role := cl.ClassifyMail("", em)
out = append(out, Entry{ out = append(out, Entry{
ID: EntryID("person", em, ""), ID: EntryID("person", em, ""),
Kind: "person", Kind: "person",
@@ -93,7 +93,7 @@ func parseMABEmails(path string) ([]Entry, error) {
return out, nil return out, nil
} }
func parseGlodaContacts(dbPath string) ([]Entry, error) { func parseGlodaContacts(dbPath string, cl *Classifier) ([]Entry, error) {
// read-only URI; immutable=1 helps when WAL/shm are missing // read-only URI; immutable=1 helps when WAL/shm are missing
dsn := "file:" + dbPath + "?mode=ro&_pragma=query_only(1)" dsn := "file:" + dbPath + "?mode=ro&_pragma=query_only(1)"
db, err := sql.Open("sqlite", dsn) db, err := sql.Open("sqlite", dsn)
@@ -137,7 +137,7 @@ func parseGlodaContacts(dbPath string) ([]Entry, error) {
if display != "" { if display != "" {
first, last = SplitDisplayName(display) first, last = SplitDisplayName(display)
} }
org, zone, role := classifyMailIdentity(display, em) org, zone, role := cl.ClassifyMail(display, em)
out = append(out, Entry{ out = append(out, Entry{
ID: EntryID("person", em, display), ID: EntryID("person", em, display),
Kind: "person", Kind: "person",
@@ -157,25 +157,6 @@ func parseGlodaContacts(dbPath string) ([]Entry, error) {
return out, rows.Err() return out, rows.Err()
} }
func classifyMailIdentity(name, email string) (org, zone, role string) {
em := NormalizeEmail(email)
_, domain, _ := strings.Cut(em, "@")
switch {
case domain == "wheregroup.com" || strings.Contains(strings.ToLower(name), "wheregroup"):
return "WhereGroup", "warm", "work"
case domain == "produktor.io" || domain == "eslider.de" || strings.HasSuffix(domain, ".produktor.io"):
return "produktor.io", "hot", "work"
case domain == "dyvenia.com":
return "Dyvenia", "warm", "work"
case domain == "immowelt.de" || domain == "immowelt.com":
return "Immowelt", "warm", "work"
case strings.HasSuffix(domain, ".de") && looksPublicSector(domain):
return domain, "warm", "work"
default:
return "", "private", "unknown"
}
}
func looksPublicSector(domain string) bool { func looksPublicSector(domain string) bool {
d := strings.ToLower(domain) d := strings.ToLower(domain)
hints := []string{ hints := []string{
+23 -11
View File
@@ -16,8 +16,8 @@ func TestNoisyEmail(t *testing.T) {
if !noisyEmail("x@marketplace.amazon.de") { if !noisyEmail("x@marketplace.amazon.de") {
t.Fatal("amazon marketplace") t.Fatal("amazon marketplace")
} }
if noisyEmail("andriy.oblivantsev@wheregroup.com") { if noisyEmail("alice.smith@acme.example") {
t.Fatal("should keep wheregroup") t.Fatal("should keep a human work address")
} }
} }
@@ -25,14 +25,15 @@ func TestParseMABEmails(t *testing.T) {
dir := t.TempDir() dir := t.TempDir()
path := filepath.Join(dir, "abook.mab") path := filepath.Join(dir, "abook.mab")
body := `// mork junk body := `// mork junk
PrimaryEmail=andriy.oblivantsev@wheregroup.com PrimaryEmail=alice.smith@acme.example
noreply@github.com noreply@github.com
axel.schaefer@wheregroup.com bob.jones@acme.example
` `
if err := os.WriteFile(path, []byte(body), 0o644); err != nil { if err := os.WriteFile(path, []byte(body), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
} }
ents, err := parseMABEmails(path) // Neutral default: no deployment rules → unclassified.
ents, err := parseMABEmails(path, DefaultClassifier())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -40,7 +41,18 @@ func TestParseMABEmails(t *testing.T) {
t.Fatalf("got %d: %+v", len(ents), ents) t.Fatalf("got %d: %+v", len(ents), ents)
} }
for _, e := range ents { for _, e := range ents {
if e.Org != "WhereGroup" || e.Role != "work" { if e.Org != "" || e.Role != "unknown" {
t.Fatalf("%+v", e)
}
}
// A deployment rule classifies the domain as work.
cl := &Classifier{MailOrgs: []MailRule{{Domain: "acme.example", Org: "Acme", Zone: "warm", Role: "work"}}}
ents, err = parseMABEmails(path, cl)
if err != nil {
t.Fatal(err)
}
for _, e := range ents {
if e.Org != "Acme" || e.Role != "work" || e.Zone != "warm" {
t.Fatalf("%+v", e) t.Fatalf("%+v", e)
} }
} }
@@ -56,8 +68,8 @@ func TestParseGlodaContacts(t *testing.T) {
_, err = db.Exec(` _, err = db.Exec(`
CREATE TABLE contacts (id INTEGER PRIMARY KEY, name TEXT); CREATE TABLE contacts (id INTEGER PRIMARY KEY, name TEXT);
CREATE TABLE identities (id INTEGER PRIMARY KEY, contactID INTEGER, kind TEXT, value TEXT); CREATE TABLE identities (id INTEGER PRIMARY KEY, contactID INTEGER, kind TEXT, value TEXT);
INSERT INTO contacts VALUES (1, 'Axel Schaefer'); INSERT INTO contacts VALUES (1, 'Alice Smith');
INSERT INTO identities VALUES (1, 1, 'email', 'axel.schaefer@wheregroup.com'); INSERT INTO identities VALUES (1, 1, 'email', 'alice.smith@acme.example');
INSERT INTO contacts VALUES (2, 'Noise Bot'); INSERT INTO contacts VALUES (2, 'Noise Bot');
INSERT INTO identities VALUES (2, 2, 'email', 'noreply@example.com'); INSERT INTO identities VALUES (2, 2, 'email', 'noreply@example.com');
`) `)
@@ -66,14 +78,14 @@ func TestParseGlodaContacts(t *testing.T) {
} }
_ = db.Close() _ = db.Close()
ents, err := parseGlodaContacts(dbPath) ents, err := parseGlodaContacts(dbPath, DefaultClassifier())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
if len(ents) != 1 { if len(ents) != 1 {
t.Fatalf("got %d %+v", len(ents), ents) t.Fatalf("got %d %+v", len(ents), ents)
} }
if ents[0].First != "Axel" || ents[0].Emails[0] != "axel.schaefer@wheregroup.com" { if ents[0].First != "Alice" || ents[0].Emails[0] != "alice.smith@acme.example" {
t.Fatalf("%+v", ents[0]) t.Fatalf("%+v", ents[0])
} }
} }
@@ -84,7 +96,7 @@ func TestScanThunderbirdRoot(t *testing.T) {
if err := os.MkdirAll(prof, 0o755); err != nil { if err := os.MkdirAll(prof, 0o755); err != nil {
t.Fatal(err) t.Fatal(err)
} }
if err := os.WriteFile(filepath.Join(prof, "abook.mab"), []byte("mail=paul.schmidt@wheregroup.com\n"), 0o644); err != nil { if err := os.WriteFile(filepath.Join(prof, "abook.mab"), []byte("mail=paul.schmidt@acme.example\n"), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
} }
doc, err := ScanThunderbirdRoot(root) doc, err := ScanThunderbirdRoot(root)
+33
View File
@@ -12,6 +12,18 @@ import (
"gopkg.in/yaml.v3" "gopkg.in/yaml.v3"
) )
// Address is one postal address of a catalog row. Category is the OO
// AddressCategory label (Home|Postal|Office|Billing|Other|Work).
type Address struct {
Street string `yaml:"street,omitempty"`
City string `yaml:"city,omitempty"`
State string `yaml:"state,omitempty"`
Zip string `yaml:"zip,omitempty"`
Country string `yaml:"country,omitempty"`
Category string `yaml:"category,omitempty"`
Primary bool `yaml:"primary,omitempty"`
}
// Entry is one catalog row (person or company). // Entry is one catalog row (person or company).
type Entry struct { type Entry struct {
ID string `yaml:"id"` ID string `yaml:"id"`
@@ -21,6 +33,7 @@ type Entry struct {
Last string `yaml:"last,omitempty"` Last string `yaml:"last,omitempty"`
Emails []string `yaml:"emails,omitempty"` Emails []string `yaml:"emails,omitempty"`
Phones []string `yaml:"phones,omitempty"` Phones []string `yaml:"phones,omitempty"`
Addresses []Address `yaml:"addresses,omitempty"`
Org string `yaml:"org,omitempty"` Org string `yaml:"org,omitempty"`
Sources []string `yaml:"sources,omitempty"` Sources []string `yaml:"sources,omitempty"`
Zone string `yaml:"zone"` Zone string `yaml:"zone"`
@@ -136,6 +149,7 @@ func mergeEntry(dst, src *Entry) {
dst.Sources = uniqueStrings(append(dst.Sources, src.Sources...)) dst.Sources = uniqueStrings(append(dst.Sources, src.Sources...))
dst.Emails = uniqueEmails(append(dst.Emails, src.Emails...)) dst.Emails = uniqueEmails(append(dst.Emails, src.Emails...))
dst.Phones = uniqueStrings(append(dst.Phones, src.Phones...)) dst.Phones = uniqueStrings(append(dst.Phones, src.Phones...))
dst.Addresses = mergeAddresses(dst.Addresses, src.Addresses)
if dst.First == "" { if dst.First == "" {
dst.First = src.First dst.First = src.First
} }
@@ -176,6 +190,25 @@ func mergeEntry(dst, src *Entry) {
} }
} }
// mergeAddresses appends src addresses not already present (by street/city/zip).
func mergeAddresses(dst, src []Address) []Address {
for _, a := range src {
exists := false
for _, b := range dst {
if strings.EqualFold(strings.TrimSpace(a.Street), strings.TrimSpace(b.Street)) &&
strings.EqualFold(strings.TrimSpace(a.City), strings.TrimSpace(b.City)) &&
strings.EqualFold(strings.TrimSpace(a.Zip), strings.TrimSpace(b.Zip)) {
exists = true
break
}
}
if !exists {
dst = append(dst, a)
}
}
return dst
}
func uniqueEmails(in []string) []string { func uniqueEmails(in []string) []string {
seen := map[string]struct{}{} seen := map[string]struct{}{}
var out []string var out []string
+10 -2
View File
@@ -14,6 +14,7 @@ import (
"net/http/cookiejar" "net/http/cookiejar"
"os" "os"
"strings" "strings"
"sync"
) )
// Client of OnlyOffice API uses credentials to get a token and query the API // Client of OnlyOffice API uses credentials to get a token and query the API
@@ -32,13 +33,20 @@ type Client struct {
defaults Defaults // optional fallbacks for calendar/project IDs defaults Defaults // optional fallbacks for calendar/project IDs
selfID string // cached /api/2.0/people/@self id selfID string // cached /api/2.0/people/@self id
noteCatID int // cached CRM history category id for "note" noteCatID int // cached CRM history category id for "note"
folderTitles map[string]string // cached Documents folder id -> title (F9)
folderTitlesMu sync.Mutex
} }
// NewClient returns a new Client backed by http.DefaultClient. // NewClient returns a new Client whose transport is paced by the process-wide
// rate limiter and 429 cooldown gate (OO_RATE_LIMIT/OO_BURST; see ratelimit.go).
func NewClient(c Credentials) *Client { func NewClient(c Credentials) *Client {
jar, _ := cookiejar.New(nil) jar, _ := cookiejar.New(nil)
return &Client{ return &Client{
client: &http.Client{Jar: jar}, client: &http.Client{
Jar: jar,
Transport: &pacedTransport{base: http.DefaultTransport},
},
credentials: &c, credentials: &c,
} }
} }
-220
View File
@@ -1,220 +0,0 @@
// Command kontoblatt builds a summary ("сводная таблица") of a Kontoblatt XLSX
// (Datum, Gegenkonto, Buchungstext, Beleg, Soll, Haben, Bemerkung) and uploads
// it back to the same OnlyOffice folder as the source file.
//
// Usage: kontoblatt <FILE_ID> <LOCAL_XLSX>
package main
import (
"context"
"fmt"
"os"
"regexp"
"sort"
"strconv"
"strings"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/xuri/excelize/v2"
)
type agg struct {
count int
soll float64
haben float64
reFehlt int
}
type rec struct {
date, month, konto, text string
soll, haben float64
reFehlt bool
}
var dateRe = regexp.MustCompile(`^\d{2}\.\d{2}\.\d{4}$`)
func parseAmount(s string) float64 {
s = strings.TrimSpace(s)
s = strings.ReplaceAll(s, "€", "")
s = strings.ReplaceAll(s, " ", "")
s = strings.ReplaceAll(s, ",", "") // German thousands separator
s = strings.TrimSpace(s)
if s == "" {
return 0
}
v, err := strconv.ParseFloat(s, 64)
if err != nil {
return 0
}
return v
}
func cell(row []string, i int) string {
if i < len(row) {
return strings.TrimSpace(row[i])
}
return ""
}
func main() {
if len(os.Args) < 3 {
fmt.Fprintln(os.Stderr, "usage: kontoblatt <FILE_ID> <LOCAL_XLSX>")
os.Exit(2)
}
fileID, path := os.Args[1], os.Args[2]
ctx := context.Background()
f, err := excelize.OpenFile(path)
if err != nil {
panic(err)
}
defer f.Close()
var recs []rec
for _, sh := range f.GetSheetList() {
rows, err := f.GetRows(sh)
if err != nil {
continue
}
for _, r := range rows {
d := cell(r, 0)
if !dateRe.MatchString(d) {
continue
}
text := cell(r, 2)
recs = append(recs, rec{
date: d,
month: d[3:10],
konto: cell(r, 1),
text: text,
soll: parseAmount(cell(r, 4)),
haben: parseAmount(cell(r, 5)),
reFehlt: strings.Contains(strings.ToUpper(text), "FEHLT"),
})
}
}
byKonto := map[string]*agg{}
byMonth := map[string]*agg{}
getK := func(k string) *agg {
if byKonto[k] == nil {
byKonto[k] = &agg{}
}
return byKonto[k]
}
getM := func(k string) *agg {
if byMonth[k] == nil {
byMonth[k] = &agg{}
}
return byMonth[k]
}
var tot agg
for _, r := range recs {
k := getK(r.konto)
k.count++
k.soll += r.soll
k.haben += r.haben
if r.reFehlt {
k.reFehlt++
}
m := getM(r.month)
m.count++
m.soll += r.soll
m.haben += r.haben
if r.reFehlt {
m.reFehlt++
}
tot.count++
tot.soll += r.soll
tot.haben += r.haben
if r.reFehlt {
tot.reFehlt++
}
}
out := excelize.NewFile()
defer out.Close()
writeSheet(out, "Nach Gegenkonto", "Gegenkonto", byKonto, tot)
writeSheet(out, "Nach Monat", "Monat", byMonth, tot)
outPath := "/tmp/opencode/kontoblatt-zusammenfassung.xlsx"
if err := out.SaveAs(outPath); err != nil {
panic(err)
}
// upload next to the source file
creds := onlyoffice.GetEnvironmentCredentials()
c := onlyoffice.NewClient(creds)
var src *onlyoffice.FileEntry
if derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
var err error
src, err = c.GetFile(ctx, fileID)
return err
}); derr != nil {
panic(derr)
}
folder := ""
if src.FolderID != nil {
folder = src.FolderID.String()
}
title := ""
if src.Title != nil {
title = *src.Title
}
fmt.Printf("source: id=%s title=%q folder=%s\n", fileID, title, folder)
name := "Kontoblatt-1591-2025-Zusammenfassung.xlsx"
tmp := "/tmp/opencode/" + name
data, _ := os.ReadFile(outPath)
if err := os.WriteFile(tmp, data, 0o600); err != nil {
panic(err)
}
var entry *onlyoffice.FileEntry
if derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
var err error
entry, _, err = c.UploadToFolderReplacing(ctx, folder, tmp)
return err
}); derr != nil {
panic(derr)
}
fmt.Printf("uploaded: %s -> folder %s (id %v)\n", name, folder, entry.ID)
// print the summary
printAgg("Nach Gegenkonto", byKonto, tot)
printAgg("Nach Monat", byMonth, tot)
}
func writeSheet(f *excelize.File, sheet, key string, m map[string]*agg, tot agg) {
f.NewSheet(sheet)
rows := [][]any{{key, "Anzahl", "Soll", "Haben", "Saldo", `davon "fehlt"`}}
keys := make([]string, 0, len(m))
for k := range m {
keys = append(keys, k)
}
sort.Strings(keys)
for _, k := range keys {
a := m[k]
rows = append(rows, []any{k, a.count, a.soll, a.haben, a.soll - a.haben, a.reFehlt})
}
rows = append(rows, []any{"GESAMT", tot.count, tot.soll, tot.haben, tot.soll - tot.haben, tot.reFehlt})
for i, row := range rows {
for j, v := range row {
cellRef, _ := excelize.CoordinatesToCellName(j+1, i+1)
_ = f.SetCellValue(sheet, cellRef, v)
}
}
}
func printAgg(title string, m map[string]*agg, tot agg) {
fmt.Printf("\n== %s ==\n", title)
keys := make([]string, 0, len(m))
for k := range m {
keys = append(keys, k)
}
sort.Strings(keys)
fmt.Printf("%-12s %6s %12s %12s %12s %7s\n", "key", "count", "soll", "haben", "saldo", "fehlt")
for _, k := range keys {
a := m[k]
fmt.Printf("%-12s %6d %12.2f %12.2f %12.2f %7d\n", k, a.count, a.soll, a.haben, a.soll-a.haben, a.reFehlt)
}
fmt.Printf("%-12s %6d %12.2f %12.2f %12.2f %7d\n", "GESAMT", tot.count, tot.soll, tot.haben, tot.soll-tot.haben, tot.reFehlt)
}
-478
View File
@@ -1,478 +0,0 @@
// Command kontolink fills the "Link" column of a Kontoblatt ("ungeklärte
// Posten") XLSX by matching each row to an OnlyOffice document.
//
// Strategy (deterministic, conservative — no LLM):
// 1. Beleg token (letters/digits from the "Beleg" column) appears in the file
// title; among candidates prefer (a) the row's month, (b) real invoices over
// copies/dupes, and require the result to be unique;
// 2. else supplier + row month + "rechnung", again unique.
//
// A file is linked at most once (rows already carrying a link are kept and their
// file counts as used). Ambiguous rows are left UNLINKED for manual review.
//
// Usage: kontolink <IN_XLSX> <INDEX_TSV> <OUT_XLSX>
package main
import (
"context"
"fmt"
"os"
"path/filepath"
"regexp"
"strconv"
"strings"
"time"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/xuri/excelize/v2"
)
var (
dateRe = regexp.MustCompile(`^\d{2}\.\d{2}\.\d{4}$`)
nonAln = regexp.MustCompile(`[^0-9a-z]+`)
fileID = regexp.MustCompile(`fileid=(\d+)`)
)
func parseDay(s string) (time.Time, bool) {
t, err := time.Parse("02.01.2006", strings.TrimSpace(s))
return t, err == nil
}
func titleDay(title string) (time.Time, bool) {
if len(title) >= 10 {
if t, err := time.Parse("2006-01-02", title[:10]); err == nil {
return t, true
}
}
return time.Time{}, false
}
// nearest picks the candidate whose title date is closest to rd. Ties and
// undated candidates (when >1) are rejected.
func nearest(cands []entry, rd time.Time) (entry, bool) {
if len(cands) == 1 {
return cands[0], true
}
best, bestD, tie := -1, 0.0, false
for i, e := range cands {
td, ok := titleDay(e.title)
if !ok {
continue
}
d := td.Sub(rd).Hours() / 24
if d < 0 {
d = -d
}
if best < 0 || d < bestD {
best, bestD, tie = i, d, false
} else if d == bestD {
tie = true
}
}
if best < 0 || tie {
return entry{}, false
}
return cands[best], true
}
const linkPrefix = "https://office.pro-dukt.de/Products/Files/DocEditor.aspx?fileid="
type entry struct {
id, path, title, norm string
}
func norm(s string) string { return nonAln.ReplaceAllString(strings.ToLower(s), "") }
func main() {
if len(os.Args) < 4 {
fmt.Fprintln(os.Stderr, "usage: kontolink <IN_XLSX> <INDEX_TSV> <OUT_XLSX>")
os.Exit(2)
}
in, idxPath, out := os.Args[1], os.Args[2], os.Args[3]
idxRaw, err := os.ReadFile(idxPath)
if err != nil {
panic(err)
}
var entries []entry
for _, line := range strings.Split(string(idxRaw), "\n") {
parts := strings.Split(line, "\t")
if len(parts) < 4 || parts[0] == "" {
continue
}
entries = append(entries, entry{id: parts[0], path: parts[2], title: parts[3], norm: norm(parts[3])})
}
f, err := excelize.OpenFile(in)
if err != nil {
panic(err)
}
defer f.Close()
sheet := f.GetSheetList()[0]
rows, err := f.GetRows(sheet)
if err != nil {
panic(err)
}
// optional 5th arg: amounts TSV "file_id\ttitle\tamount" (see cmd/pdfamount)
var amts []amtEntry
if len(os.Args) >= 6 && os.Args[5] != "" {
amts = loadAmounts(os.Args[5])
}
used := map[string]bool{}
for _, r := range rows {
if m := fileID.FindStringSubmatch(cell(r, 7)); m != nil {
used[m[1]] = true
}
}
var linked, byBeleg, bySupplier, byAmount, unmatched, ambiguous int
for i, r := range rows {
if i == 0 || !dateRe.MatchString(cell(r, 0)) || strings.TrimSpace(cell(r, 7)) != "" {
continue
}
beleg := norm(cell(r, 3))
supplier := supplierNorm(cell(r, 2))
month := monthYear(cell(r, 0))
rd, _ := parseDay(cell(r, 0))
e, kind, ok := pick(entries, used, beleg, supplier, month, rd)
if !ok {
if ae, aok := amountPick(amts, used, supplier, rowAmount(r), rd); aok {
e, kind, ok = entry{id: ae.id, title: ae.title}, "amount", true
}
}
if !ok {
if beleg != "" {
ambiguous++
} else {
unmatched++
}
continue
}
ref, _ := excelize.CoordinatesToCellName(8, i+1)
if err := f.SetCellValue(sheet, ref, linkPrefix+e.id); err != nil {
panic(err)
}
used[e.id] = true
linked++
switch kind {
case "beleg":
byBeleg++
case "supplier":
bySupplier++
case "amount":
byAmount++
}
fmt.Printf("row %3d %-30s -> %s [%s]\n", i+1, cell(r, 2), e.title, kind)
}
if err := f.SaveAs(out); err != nil {
panic(err)
}
fmt.Printf("\nlinked=%d (beleg=%d, supplier=%d, amount=%d), ambiguous=%d, no-candidate=%d\n",
linked, byBeleg, bySupplier, byAmount, ambiguous, unmatched)
// Optional 4th arg: source OnlyOffice file id. Try to update it in place;
// if it is locked (OnlyOffice 500), upload a "(links)" copy next to it.
if len(os.Args) >= 5 && os.Args[4] != "" {
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
ctx, cancel := context.WithTimeout(context.Background(), 120*time.Second)
defer cancel()
var src *onlyoffice.FileEntry
if derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
var err error
src, err = c.GetFile(ctx, os.Args[4])
return err
}); derr != nil {
panic(derr)
}
folder, title := "", ""
if src.FolderID != nil {
folder = src.FolderID.String()
}
if src.Title != nil {
title = *src.Title
}
uderr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
_, err := c.UpdateFile(ctx, os.Args[4], out)
return err
})
if uderr == nil {
fmt.Printf("updated file %s in place\n", os.Args[4])
return
}
fmt.Printf("in-place update failed (locked?); uploading a copy to folder %s\n", folder)
ext := filepath.Ext(title)
name := strings.TrimSuffix(title, ext) + " (links)" + ext
tmp := filepath.Join(os.TempDir(), name)
data, _ := os.ReadFile(out)
if err := os.WriteFile(tmp, data, 0o600); err != nil {
panic(err)
}
if derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
_, _, err := c.UploadToFolderReplacing(ctx, folder, tmp)
return err
}); derr != nil {
panic(derr)
}
fmt.Printf("uploaded copy: %s -> folder %s\n", name, folder)
}
}
func cell(r []string, i int) string {
if i < len(r) {
return strings.TrimSpace(r[i])
}
return ""
}
func supplierNorm(s string) string {
s = strings.ToUpper(s)
if i := strings.Index(s, ","); i >= 0 {
s = s[:i]
}
for _, w := range []string{"RE FEHLT", "GS FEHLT", "WOFR", "WOFÜR"} {
s = strings.ReplaceAll(s, w, "")
}
return norm(s)
}
func monthYear(date string) string {
if len(date) == 10 {
return date[6:10] + "-" + date[3:5]
}
return ""
}
// pick returns an unused candidate. Beleg match wins; supplier+month is a
// fallback. When several candidates qualify, the one closest in time to the row
// date wins; a tie is rejected (ambiguous) rather than guessed.
func pick(entries []entry, used map[string]bool, beleg, supplier, month string, rd time.Time) (entry, string, bool) {
free := func(e entry) bool { return !used[e.id] }
if len(beleg) >= 5 {
var inMonth []entry
for _, e := range entries {
if free(e) && belegMatches(e.norm, beleg) &&
(month == "" || strings.Contains(e.title, month)) {
inMonth = append(inMonth, e)
}
}
if supplier != "" {
var s []entry
for _, e := range inMonth {
if strings.Contains(e.norm, supplier) {
s = append(s, e)
}
}
if len(s) > 0 {
inMonth = s
}
}
inMonth = topRank(inMonth)
if e, ok := nearest(inMonth, rd); ok {
return e, "beleg", true
}
// A Beleg is present but no file carries it: do NOT fall back to a
// supplier guess (that links the wrong invoice).
return entry{}, "", false
}
if supplier != "" && month != "" {
var c []entry
for _, e := range entries {
if free(e) && strings.Contains(e.norm, supplier) &&
strings.Contains(e.title, month) && strings.Contains(e.norm, "rechnung") {
c = append(c, e)
}
}
c = topRank(c)
if e, ok := nearest(c, rd); ok {
return e, "supplier", true
}
}
return entry{}, "", false
}
// topRank keeps only the highest-ranked candidates (real invoice over copy /
// dupe / op), so a tie with a duplicate does not mask the real file.
func topRank(cands []entry) []entry {
if len(cands) < 2 {
return cands
}
best := 0
for _, e := range cands {
if rank(e) > best {
best = rank(e)
}
}
out := cands[:0]
for _, e := range cands {
if rank(e) == best {
out = append(out, e)
}
}
return out
}
func rank(e entry) int {
s := 0
if strings.Contains(e.path, "/2025") || strings.Contains(e.path, "/2024") {
s += 4
}
if strings.Contains(e.norm, "rechnung") {
s += 2
}
if strings.Contains(e.norm, "dupe") || strings.Contains(e.norm, "copy") ||
strings.Contains(e.norm, "op") {
s--
}
return s
}
// belegMatches reports whether a Beleg identifies the file: the whole normalized
// Beleg appears, or (for long numeric Belege, e.g. "24/641393110") an 8-digit
// window of its longest digit run appears.
func belegMatches(titleNorm, beleg string) bool {
if strings.Contains(titleNorm, beleg) {
return true
}
run := longestDigitRun(beleg)
for i := 0; i+8 <= len(run); i++ {
if strings.Contains(titleNorm, run[i:i+8]) {
return true
}
}
return false
}
func longestDigitRun(s string) string {
var best, cur strings.Builder
for _, r := range s {
if r >= '0' && r <= '9' {
cur.WriteRune(r)
if cur.Len() > best.Len() {
best.Reset()
best.WriteString(cur.String())
}
} else {
cur.Reset()
}
}
return best.String()
}
type amtEntry struct {
id string
title string
norm string
amount float64
date time.Time
hasDate bool
}
func loadAmounts(path string) []amtEntry {
raw, err := os.ReadFile(path)
if err != nil {
return nil
}
var out []amtEntry
for _, line := range strings.Split(string(raw), "\n") {
p := strings.Split(line, "\t")
if len(p) < 3 {
continue
}
v, err := strconv.ParseFloat(strings.TrimSpace(p[2]), 64)
if err != nil {
continue
}
e := amtEntry{id: p[0], title: p[1], norm: norm(p[1]), amount: v}
if len(p[1]) >= 10 {
if t, err := time.Parse("2006-01-02", p[1][:10]); err == nil {
e.date, e.hasDate = t, true
}
}
out = append(out, e)
}
return out
}
func rowAmount(r []string) float64 {
if v := parseAmount(cell(r, 4)); v != 0 {
return v
}
return parseAmount(cell(r, 5))
}
func parseAmount(s string) float64 {
s = strings.ReplaceAll(s, "€", "")
s = strings.ReplaceAll(s, " ", "")
s = strings.ReplaceAll(s, ",", ".")
if s == "" {
return 0
}
v, err := strconv.ParseFloat(s, 64)
if err != nil {
return 0
}
return v
}
// amountPick matches a row to an O2 invoice by amount + nearest date. Scoped to
// Telefonica/O2 rows and O2 files, so it cannot cross-link other suppliers.
func amountPick(amts []amtEntry, used map[string]bool, supplier string, amt float64, rd time.Time) (amtEntry, bool) {
if amt <= 0 || len(amts) == 0 {
return amtEntry{}, false
}
if !strings.Contains(supplier, "telefonica") && !strings.Contains(supplier, "o2") {
return amtEntry{}, false
}
var cands []amtEntry
for _, a := range amts {
if used[a.id] || !strings.Contains(a.norm, "o2") {
continue
}
d := a.amount - amt
if d < 0 {
d = -d
}
if d > 0.005 {
continue
}
if a.hasDate && !rd.IsZero() {
days := a.date.Sub(rd).Hours() / 24
if days < 0 {
days = -days
}
if days > 75 {
continue
}
}
cands = append(cands, a)
}
if len(cands) == 1 {
return cands[0], true
}
best, bestD, tie := -1, 0.0, false
for i, a := range cands {
if !a.hasDate {
continue
}
d := a.date.Sub(rd).Hours() / 24
if d < 0 {
d = -d
}
if best < 0 || d < bestD {
best, bestD, tie = i, d, false
} else if d == bestD {
tie = true
}
}
if best < 0 || tie {
return amtEntry{}, false
}
return cands[best], true
}
+14 -6
View File
@@ -17,6 +17,18 @@ const MailListPageSize = 25
// Loader fetches list items for a menu subject using the OnlyOffice client. // Loader fetches list items for a menu subject using the OnlyOffice client.
type Loader struct { type Loader struct {
Client *onlyoffice.Client Client *onlyoffice.Client
// Files is the backend-agnostic file store used for file download, preview
// and delete. When nil it falls back to Client.FileStore(ProviderREST).
Files onlyoffice.FileStore
}
// fileStore returns the configured file store, defaulting to REST.
func (l *Loader) fileStore() onlyoffice.FileStore {
if l.Files != nil {
return l.Files
}
return l.Client.FileStore(onlyoffice.ProviderREST)
} }
// List returns items for the given list spec (nav leaf). // List returns items for the given list spec (nav leaf).
@@ -171,11 +183,7 @@ func (l *Loader) executeDelete(ctx context.Context, item model.Item) (string, er
} }
return fmt.Sprintf("Deleted message %s", item.Title), nil return fmt.Sprintf("Deleted message %s", item.Title), nil
case model.KindFile: case model.KindFile:
id, err := strconv.Atoi(item.ID) if err := l.fileStore().Delete(ctx, []string{item.ID}); err != nil {
if err != nil {
return "", err
}
if err := l.Client.DeleteFiles(ctx, []int{id}); err != nil {
return "", err return "", err
} }
return fmt.Sprintf("Deleted file %s", item.Title), nil return fmt.Sprintf("Deleted file %s", item.Title), nil
@@ -199,7 +207,7 @@ func (l *Loader) executeDownload(ctx context.Context, item model.Item, destPath
return "", err return "", err
} }
defer f.Close() defer f.Close()
if _, err := l.Client.DownloadFile(ctx, item.ID, f); err != nil { if _, err := l.fileStore().Download(ctx, item.ID, f); err != nil {
return "", err return "", err
} }
return fmt.Sprintf("Downloaded to %s", destPath), nil return fmt.Sprintf("Downloaded to %s", destPath), nil
+3 -6
View File
@@ -5,7 +5,6 @@ import (
"context" "context"
"fmt" "fmt"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/eslider/go-onlyoffice/cmd/office/model" "github.com/eslider/go-onlyoffice/cmd/office/model"
"github.com/eslider/go-onlyoffice/cmd/office/preview" "github.com/eslider/go-onlyoffice/cmd/office/preview"
) )
@@ -30,13 +29,11 @@ func (l *Loader) filePreviewMarkdown(ctx context.Context, item model.Item) (stri
return "", fmt.Errorf("file id missing") return "", fmt.Errorf("file id missing")
} }
name := item.Title name := item.Title
if meta, err := l.Client.GetFile(ctx, item.ID); err == nil && meta != nil { if e, err := l.fileStore().Stat(ctx, item.ID); err == nil && e.Title != "" {
if t := onlyoffice.FileEntryTitle(meta); t != "" { name = e.Title
name = t
}
} }
var buf bytes.Buffer var buf bytes.Buffer
if _, err := l.Client.DownloadFile(ctx, item.ID, &buf); err != nil { if _, err := l.fileStore().Download(ctx, item.ID, &buf); err != nil {
return "", err return "", err
} }
return preview.FileBytesToMarkdown(name, buf.Bytes()) return preview.FileBytesToMarkdown(name, buf.Bytes())
+3 -7
View File
@@ -4,7 +4,6 @@ package fetch_test
import ( import (
"context" "context"
"os"
"testing" "testing"
"github.com/eslider/go-onlyoffice/cmd/office/model" "github.com/eslider/go-onlyoffice/cmd/office/model"
@@ -46,9 +45,6 @@ func TestIntegrationUpdateTaskTitleDescription(t *testing.T) {
} }
func TestIntegrationTaskFieldsFromLiveAPI(t *testing.T) { func TestIntegrationTaskFieldsFromLiveAPI(t *testing.T) {
if os.Getenv("ONLYOFFICE_URL") == "" && os.Getenv("ONLYOFFICE_HOST") == "" {
t.Skip("ONLYOFFICE_URL not set")
}
loader, ctx := liveLoader(t) loader, ctx := liveLoader(t)
items, err := loader.List(ctx, model.ListSpec{Subject: model.SubjectTasks}) items, err := loader.List(ctx, model.ListSpec{Subject: model.SubjectTasks})
if err != nil { if err != nil {
@@ -57,12 +53,12 @@ func TestIntegrationTaskFieldsFromLiveAPI(t *testing.T) {
if len(items) == 0 { if len(items) == 0 {
t.Skip("no tasks") t.Skip("no tasks")
} }
title, desc, err := loader.TaskFields(ctx, items[0]) fields, err := loader.DetailForm(ctx, items[0])
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
if title == "" { if fields.Primary == "" {
t.Fatal("empty title") t.Fatal("empty title")
} }
_ = desc _ = fields.Secondary
} }
+25 -3
View File
@@ -64,6 +64,7 @@ func catalogScanContactsCmd() *cobra.Command {
func catalogScanProjectsCmd() *cobra.Command { func catalogScanProjectsCmd() *cobra.Command {
var outPath string var outPath string
var maxDepth int var maxDepth int
var configPath string
cmd := &cobra.Command{ cmd := &cobra.Command{
Use: "scan-projects", Use: "scan-projects",
Short: "Git roots / remotes / top-level dirs → company rows", Short: "Git roots / remotes / top-level dirs → company rows",
@@ -72,7 +73,11 @@ func catalogScanProjectsCmd() *cobra.Command {
if root == "" { if root == "" {
return fmt.Errorf("--root is required") return fmt.Errorf("--root is required")
} }
doc, err := catalog.ScanProjectsRoot(root, maxDepth) cl, err := catalogClassifier(configPath)
if err != nil {
return err
}
doc, err := catalog.ScanProjectsRootOpts(root, maxDepth, catalog.ScanOptions{Classifier: cl})
if err != nil { if err != nil {
return err return err
} }
@@ -81,6 +86,7 @@ func catalogScanProjectsCmd() *cobra.Command {
} }
cmd.Flags().String("root", "", "projects directory (local path)") cmd.Flags().String("root", "", "projects directory (local path)")
cmd.Flags().IntVar(&maxDepth, "max-depth", 4, "max directory depth for git roots") cmd.Flags().IntVar(&maxDepth, "max-depth", 4, "max directory depth for git roots")
cmd.Flags().StringVar(&configPath, "config", "", "classification rules YAML (default $OO_CATALOG_CONFIG)")
cmd.Flags().StringVarP(&outPath, "out", "O", "", "write YAML to this path") cmd.Flags().StringVarP(&outPath, "out", "O", "", "write YAML to this path")
_ = cmd.MarkFlagRequired("root") _ = cmd.MarkFlagRequired("root")
return cmd return cmd
@@ -89,6 +95,7 @@ func catalogScanProjectsCmd() *cobra.Command {
func catalogScanThunderbirdCmd() *cobra.Command { func catalogScanThunderbirdCmd() *cobra.Command {
var outPath string var outPath string
var mboxHeaders bool var mboxHeaders bool
var configPath string
cmd := &cobra.Command{ cmd := &cobra.Command{
Use: "scan-thunderbird", Use: "scan-thunderbird",
Short: "Thunderbird profiles: abook/history.mab + Gloda SQLite contacts", Short: "Thunderbird profiles: abook/history.mab + Gloda SQLite contacts",
@@ -98,13 +105,18 @@ func catalogScanThunderbirdCmd() *cobra.Command {
- optional --mbox-headers: From/To/Cc/Reply-To from mbox folder files (no bodies) - optional --mbox-headers: From/To/Cc/Reply-To from mbox folder files (no bodies)
Noisy senders (noreply, Amazon marketplace, GitHub reply, …) are skipped. Noisy senders (noreply, Amazon marketplace, GitHub reply, …) are skipped.
Default zone is private; known work domains (e.g. wheregroup.com) get zone=warm role=work.`, Default zone is private; work domains/names are classified from the rules in
$OO_CATALOG_CONFIG (or --config) — none are hardcoded.`,
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
root, _ := cmd.Flags().GetString("root") root, _ := cmd.Flags().GetString("root")
if root == "" { if root == "" {
return fmt.Errorf("--root is required") return fmt.Errorf("--root is required")
} }
doc, err := catalog.ScanThunderbirdRootOpts(root, catalog.ScanOptions{MboxHeaders: mboxHeaders}) cl, err := catalogClassifier(configPath)
if err != nil {
return err
}
doc, err := catalog.ScanThunderbirdRootOpts(root, catalog.ScanOptions{MboxHeaders: mboxHeaders, Classifier: cl})
if err != nil { if err != nil {
return err return err
} }
@@ -113,11 +125,21 @@ Default zone is private; known work domains (e.g. wheregroup.com) get zone=warm
} }
cmd.Flags().String("root", "", "Thunderbird profile or parent directory (local path)") cmd.Flags().String("root", "", "Thunderbird profile or parent directory (local path)")
cmd.Flags().BoolVar(&mboxHeaders, "mbox-headers", false, "also extract emails from mbox From/To/Cc headers") cmd.Flags().BoolVar(&mboxHeaders, "mbox-headers", false, "also extract emails from mbox From/To/Cc headers")
cmd.Flags().StringVar(&configPath, "config", "", "classification rules YAML (default $OO_CATALOG_CONFIG)")
cmd.Flags().StringVarP(&outPath, "out", "O", "", "write YAML to this path") cmd.Flags().StringVarP(&outPath, "out", "O", "", "write YAML to this path")
_ = cmd.MarkFlagRequired("root") _ = cmd.MarkFlagRequired("root")
return cmd return cmd
} }
// catalogClassifier loads classification rules from --config, else
// $OO_CATALOG_CONFIG, else the neutral default (no rules).
func catalogClassifier(configPath string) (*catalog.Classifier, error) {
if configPath == "" {
return catalog.LoadClassifierFromEnv()
}
return catalog.LoadClassifier(configPath)
}
func catalogMergeCmd() *cobra.Command { func catalogMergeCmd() *cobra.Command {
var inputs []string var inputs []string
var outPath string var outPath string
+3 -1
View File
@@ -31,7 +31,9 @@ func init() {
} }
// execute runs the root command. Exported only to main.go in the same package. // execute runs the root command. Exported only to main.go in the same package.
func execute() error { return rootCmd.Execute() } // .env is loaded CLI-wide so non-authenticating commands (e.g. catalog scans)
// also see configuration such as OO_CATALOG_CONFIG.
func execute() error { bootstrap.LoadEnv(); return rootCmd.Execute() }
// newOO loads env (only .env in CWD) and returns an authenticated client. // newOO loads env (only .env in CWD) and returns an authenticated client.
// godotenv is a CLI-only concern; the library itself never loads dotfiles. // godotenv is a CLI-only concern; the library itself never loads dotfiles.
+1 -1
View File
@@ -7,7 +7,7 @@ import (
var crmCmd = &cobra.Command{ var crmCmd = &cobra.Command{
Use: "crm", Use: "crm",
Short: "CRM maintenance (dedupe, cleanup)", Short: "CRM maintenance (audit, dedupe, cleanup)",
} }
func init() { func init() {
+63
View File
@@ -0,0 +1,63 @@
package main
import (
"encoding/json"
"fmt"
"os"
"sort"
"github.com/spf13/cobra"
)
func init() {
crmCmd.AddCommand(crmAuditCmd())
}
func crmAuditCmd() *cobra.Command {
var outPath string
cmd := &cobra.Command{
Use: "audit",
Short: "Audit opportunities (files/tasks/members per deal) and classify",
Long: `Lists every opportunity with its file/task/member counts and a coarse class
(ok | dup | empty | junk-title). --out writes the full JSON audit for later use.`,
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
audits, err := c.AuditOpportunities(cmd.Context())
if err != nil {
return err
}
if outPath != "" {
b, merr := json.MarshalIndent(audits, "", " ")
if merr != nil {
return merr
}
if werr := os.WriteFile(outPath, b, 0o644); werr != nil {
return werr
}
}
byClass := map[string]int{}
for _, a := range audits {
byClass[a.Class]++
}
classes := make([]string, 0, len(byClass))
for k := range byClass {
classes = append(classes, k)
}
sort.Strings(classes)
rows := make([]map[string]any, 0, len(classes))
for _, k := range classes {
rows = append(rows, map[string]any{"class": k, "count": byClass[k]})
}
printTable([]string{"class", "count"}, rows)
if outPath != "" {
fmt.Fprintf(cmd.OutOrStdout(), "wrote %d audits → %s\n", len(audits), outPath)
}
return nil
},
}
cmd.Flags().StringVar(&outPath, "out", "", "write audit JSON to this file")
return cmd
}
+49 -45
View File
@@ -3,6 +3,7 @@ package main
import ( import (
"fmt" "fmt"
"os" "os"
"time"
onlyoffice "github.com/eslider/go-onlyoffice" onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra" "github.com/spf13/cobra"
@@ -12,8 +13,8 @@ func init() {
rootCmd.AddCommand(davCmd()) rootCmd.AddCommand(davCmd())
} }
// davCmd exposes the Documents module through the same Dav calls that back // davCmd exposes the Documents module through the backend-agnostic FileStore
// oo-webdav (ListDavFolder / MoveDavItems / CopyDavItems / DownloadDavFile). // (DAV backend). The underlying Dav calls are the oo-webdav proven path:
// MoveDavItems sends resolveType=Skip + holdResult=true, which the legacy // MoveDavItems sends resolveType=Skip + holdResult=true, which the legacy
// fileops/move call without those params silently ignores (200 without move). // fileops/move call without those params silently ignores (200 without move).
func davCmd() *cobra.Command { func davCmd() *cobra.Command {
@@ -61,35 +62,21 @@ func davLsCmd() *cobra.Command {
printTable([]string{"id", "title", "filesCount", "foldersCount"}, rows) printTable([]string{"id", "title", "filesCount", "foldersCount"}, rows)
return nil return nil
} }
l, err := c.ListDavFolder(ctx, args[0]) entries, err := c.FileStore(onlyoffice.ProviderDAV).List(ctx, args[0])
if err != nil { if err != nil {
return err return err
} }
if outputFormat == "json" { folders := make([]onlyoffice.Entry, 0, len(entries))
folders := make([]map[string]any, 0, len(l.Folders)) files := make([]onlyoffice.Entry, 0, len(entries))
for _, f := range l.Folders { for _, e := range entries {
folders = append(folders, map[string]any{ if e.Kind == onlyoffice.Folder {
"id": f.ID, folders = append(folders, e)
"title": f.Title, } else {
"filesCount": f.FilesCount, files = append(files, e)
"foldersCount": f.FoldersCount,
})
} }
files := make([]map[string]any, 0, len(l.Files))
for _, f := range l.Files {
files = append(files, map[string]any{
"id": f.ID,
"title": f.Title,
"size": f.Size,
"updated": f.Updated,
})
} }
printObject(map[string]any{"folders": folders, "files": files}) frows := make([]map[string]any, 0, len(folders))
return nil for _, f := range folders {
}
if len(l.Folders) > 0 {
frows := make([]map[string]any, 0, len(l.Folders))
for _, f := range l.Folders {
frows = append(frows, map[string]any{ frows = append(frows, map[string]any{
"id": f.ID, "id": f.ID,
"title": f.Title, "title": f.Title,
@@ -97,23 +84,24 @@ func davLsCmd() *cobra.Command {
"foldersCount": f.FoldersCount, "foldersCount": f.FoldersCount,
}) })
} }
if outputFormat == "table" { rows := make([]map[string]any, 0, len(files))
fmt.Println("folders:") for _, f := range files {
}
printTable([]string{"id", "title", "filesCount", "foldersCount"}, frows)
}
rows := make([]map[string]any, 0, len(l.Files))
for _, f := range l.Files {
rows = append(rows, map[string]any{ rows = append(rows, map[string]any{
"id": f.ID, "id": f.ID,
"title": f.Title, "title": f.Title,
"size": f.Size, "size": f.Size,
"updated": f.Updated, "updated": entryUpdated(f),
}) })
} }
if outputFormat == "table" { if outputFormat == "json" {
fmt.Println("files:") printObject(map[string]any{"folders": frows, "files": rows})
return nil
} }
if len(frows) > 0 {
fmt.Println("folders:")
printTable([]string{"id", "title", "filesCount", "foldersCount"}, frows)
}
fmt.Println("files:")
printTable([]string{"id", "title", "size", "updated"}, rows) printTable([]string{"id", "title", "size", "updated"}, rows)
return nil return nil
}, },
@@ -121,6 +109,18 @@ func davLsCmd() *cobra.Command {
return cmd return cmd
} }
// entryUpdated prefers the backend-native timestamp string so table/JSON output
// round-trips what the API returned.
func entryUpdated(e onlyoffice.Entry) string {
if e.Updated != "" {
return e.Updated
}
if e.Modified.IsZero() {
return ""
}
return e.Modified.Format(time.RFC3339)
}
func davMoveCmd() *cobra.Command { func davMoveCmd() *cobra.Command {
var folderIDs []string var folderIDs []string
cmd := &cobra.Command{ cmd := &cobra.Command{
@@ -132,7 +132,8 @@ func davMoveCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
if err := c.MoveDavItems(cmd.Context(), folderIDs, args[1:], args[0]); err != nil { ids := append(append([]string{}, folderIDs...), args[1:]...)
if err := c.FileStore(onlyoffice.ProviderDAV).Move(cmd.Context(), ids, args[0]); err != nil {
return err return err
} }
printObject(map[string]any{"moved_files": args[1:], "moved_folders": folderIDs, "dest": args[0]}) printObject(map[string]any{"moved_files": args[1:], "moved_folders": folderIDs, "dest": args[0]})
@@ -154,7 +155,8 @@ func davCopyCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
if err := c.CopyDavItems(cmd.Context(), folderIDs, args[1:], args[0]); err != nil { ids := append(append([]string{}, folderIDs...), args[1:]...)
if err := c.FileStore(onlyoffice.ProviderDAV).Copy(cmd.Context(), ids, args[0]); err != nil {
return err return err
} }
printObject(map[string]any{"copied_files": args[1:], "copied_folders": folderIDs, "dest": args[0]}) printObject(map[string]any{"copied_files": args[1:], "copied_folders": folderIDs, "dest": args[0]})
@@ -175,7 +177,7 @@ func davMkdirCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
f, err := c.CreateDavFolder(cmd.Context(), args[0], args[1]) f, err := c.FileStore(onlyoffice.ProviderDAV).CreateFolder(cmd.Context(), args[0], args[1])
if err != nil { if err != nil {
return err return err
} }
@@ -200,7 +202,8 @@ func davRemoveCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
if err := c.DeleteDavItems(cmd.Context(), folderIDs, args); err != nil { ids := append(append([]string{}, args...), folderIDs...)
if err := c.FileStore(onlyoffice.ProviderDAV).Delete(cmd.Context(), ids); err != nil {
return err return err
} }
printObject(map[string]any{"deleted_files": args, "deleted_folders": folderIDs}) printObject(map[string]any{"deleted_files": args, "deleted_folders": folderIDs})
@@ -221,7 +224,7 @@ func davRenameFileCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
if err := c.RenameDavFile(cmd.Context(), args[0], args[1]); err != nil { if err := c.FileStore(onlyoffice.ProviderDAV).Rename(cmd.Context(), args[0], args[1]); err != nil {
return err return err
} }
printObject(map[string]any{"id": args[0], "title": args[1]}) printObject(map[string]any{"id": args[0], "title": args[1]})
@@ -240,7 +243,7 @@ func davRenameFolderCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
if err := c.RenameDavFolder(cmd.Context(), args[0], args[1]); err != nil { if err := c.FileStore(onlyoffice.ProviderDAV).Rename(cmd.Context(), args[0], args[1]); err != nil {
return err return err
} }
printObject(map[string]any{"id": args[0], "title": args[1]}) printObject(map[string]any{"id": args[0], "title": args[1]})
@@ -261,20 +264,21 @@ func davDownloadCmd() *cobra.Command {
return err return err
} }
ctx := cmd.Context() ctx := cmd.Context()
f, err := c.GetFile(ctx, args[0]) store := c.FileStore(onlyoffice.ProviderDAV)
e, err := store.Stat(ctx, args[0])
if err != nil { if err != nil {
return err return err
} }
path := to path := to
if path == "" { if path == "" {
path = onlyoffice.SafeLocalFileName(onlyoffice.FileEntryTitle(f)) path = onlyoffice.SafeLocalFileName(e.Title)
} }
out, err := os.Create(path) out, err := os.Create(path)
if err != nil { if err != nil {
return err return err
} }
defer out.Close() defer out.Close()
n, err := c.DownloadDavFile(ctx, args[0], out) n, err := store.Download(ctx, args[0], out)
if err != nil { if err != nil {
_ = os.Remove(path) _ = os.Remove(path)
return err return err
+306 -50
View File
@@ -1,15 +1,19 @@
package main package main
import ( import (
"bytes"
"context" "context"
"encoding/json"
"fmt" "fmt"
"io"
"os" "os"
"path/filepath" "path/filepath"
"strconv"
"strings" "strings"
"time"
onlyoffice "github.com/eslider/go-onlyoffice" onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/eslider/go-onlyoffice/internal/docpipe" "github.com/eslider/go-onlyoffice/internal/docpipe"
"github.com/eslider/go-onlyoffice/internal/xlspipe"
"github.com/spf13/cobra" "github.com/spf13/cobra"
) )
@@ -26,12 +30,16 @@ func docsCmd() *cobra.Command {
OnlyOffice Documents UI is poor for .md/.txt — keep sources in git, store .docx in OO. OnlyOffice Documents UI is poor for .md/.txt — keep sources in git, store .docx in OO.
Upload Markdown as DOCX: oo docs put-md PROJECT_ID file.md Upload Markdown as DOCX: oo docs put-md PROJECT_ID file.md
Upload plain text: oo docs put-txt PROJECT_ID file.txt (preserves line breaks) Upload plain text: oo docs put-txt PROJECT_ID file.txt (preserves line breaks)
Upload/generate XLSX: oo docs put-xlsx PROJECT_ID [--template cutover-portugal | FILE.xlsx] Upload XLSX: oo docs put-xlsx PROJECT_ID FILE.xlsx
Read an OO file as MD: oo docs as-md FILE_ID Read an OO file as MD: oo docs as-md FILE_ID
OCR a scan locally: oo docs ocr scan.pdf --md out.md OCR a scan locally: oo docs ocr scan.pdf --md out.md
Structured OCR (hOCR→MD): oo docs hocr scan.jpg --md out.md --yaml out.yml`, Structured OCR (hOCR→MD): oo docs hocr scan.jpg --md out.md --yaml out.yml`,
} }
cmd.AddCommand(docsConvertCmd()) cmd.AddCommand(docsConvertCmd())
cmd.AddCommand(docsPDFCmd())
cmd.AddCommand(docsPresignedCmd())
cmd.AddCommand(docsCSVCmd())
cmd.AddCommand(docsJSONCmd())
cmd.AddCommand(docsOptimizeCmd()) cmd.AddCommand(docsOptimizeCmd())
cmd.AddCommand(docsOCRCmd()) cmd.AddCommand(docsOCRCmd())
cmd.AddCommand(docsHOCRCmd()) cmd.AddCommand(docsHOCRCmd())
@@ -61,6 +69,293 @@ func docsToolsCmd() *cobra.Command {
} }
} }
func docsPDFCmd() *cobra.Command {
var out, docsURL, secret, output, folder string
var stream, pipe bool
cmd := &cobra.Command{
Use: "pdf FILE_ID | PATH [ARG...]",
Short: "Convert files to PDF via the DocumentServer converter (OO file ids or local paths)",
Long: `Native OnlyOffice conversion (the engine behind the portal's "Download as PDF"):
1. GET /api/2.0/files/file/{id}/presigneduri → fetchable source URL
2. POST <docs>/converter with a JWT → converted file URL
3. download the result
Arguments may be OnlyOffice file ids OR local file paths. A local path is
uploaded to the scratch folder (--folder, default 2 = "My documents"), converted,
downloaded and then removed — so any local document yields a PDF on the fly.
Docs base defaults to $ONLYOFFICE_DOCS_URL, else $ONLYOFFICE_URL + "/ds-vpath"
(/ds-vpath is the usual reverse-proxy mount for the DocumentServer). The JWT secret is
$ONLYOFFICE_DS_SECRET (DocumentServer services.CoAuthoring.secret).`,
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
base := docsBaseURL(docsURL)
if base == "" {
return fmt.Errorf("docs base url unknown; set --docs-url or ONLYOFFICE_DOCS_URL")
}
sec := secret
if sec == "" {
sec = firstEnv("ONLYOFFICE_DS_SECRET", "OO_DS_SECRET")
}
if sec == "" {
return fmt.Errorf("JWT secret required: --secret or ONLYOFFICE_DS_SECRET")
}
ot := output
if ot == "" {
ot = "pdf"
}
toStdout := stream || pipe
if toStdout && len(args) > 1 {
return fmt.Errorf("--stream/--pipe writes one file to stdout; pass a single input")
}
for _, arg := range args {
id, local := arg, false
title := ""
if fi, statErr := os.Stat(arg); statErr == nil && !fi.IsDir() {
// Local file → temporary upload into the scratch folder.
ent, uerr := c.UploadToFolder(cmd.Context(), folder, arg)
if uerr != nil {
return fmt.Errorf("upload %s: %w", arg, uerr)
}
id = strconv.FormatInt(onlyoffice.FileEntryNumericID(ent), 10)
local = true
title = filepath.Base(arg)
} else {
if f, ferr := c.GetFile(cmd.Context(), id); ferr == nil && f != nil && f.Title != nil {
title = *f.Title
}
}
src, err := c.PresignedURI(cmd.Context(), id)
if err != nil {
return fmt.Errorf("presigneduri %s: %w", id, err)
}
res, err := c.ConvertDocument(cmd.Context(), base, sec, onlyoffice.ConvertRequest{
URL: src,
OutputType: ot,
FileType: strings.TrimPrefix(filepath.Ext(title), "."),
Title: title,
Key: fmt.Sprintf("oo-%s-%d", id, time.Now().UnixNano()),
})
if err != nil {
return fmt.Errorf("convert %s: %w", id, err)
}
var w io.Writer
dst := out
if toStdout {
w = os.Stdout
} else {
if dst == "" {
stem := strings.TrimSuffix(title, filepath.Ext(title))
if stem == "" {
stem = "file-" + id
}
dst = stem + "." + ot
}
fh, err := os.Create(dst)
if err != nil {
return err
}
w = fh
defer fh.Close()
}
n, derr := c.DownloadURLTo(cmd.Context(), res.FileURL, w)
if local {
// Best-effort cleanup of the temporary upload.
if nid, e := strconv.Atoi(id); e == nil {
_ = c.DeleteFiles(cmd.Context(), []int{nid})
}
}
if derr != nil {
return fmt.Errorf("download: %w", derr)
}
if toStdout {
// Keep stdout byte-clean for pipes; status goes to stderr.
fmt.Fprintf(os.Stderr, "converted %s -> stdout (%d bytes, %s)\n", arg, n, ot)
} else {
printObject(map[string]any{"source": arg, "fileid": id, "title": title, "output": dst, "bytes": n, "type": ot})
}
}
return nil
},
}
cmd.Flags().StringVar(&out, "out", "", "output path (default: ./<title>.<format>)")
cmd.Flags().StringVar(&output, "to", "pdf", "output format (pdf, docx, xlsx, …)")
cmd.Flags().StringVar(&docsURL, "docs-url", "", "DocumentServer base (default $ONLYOFFICE_DOCS_URL or $ONLYOFFICE_URL/ds-vpath)")
cmd.Flags().StringVar(&secret, "secret", "", "JWT secret (default $ONLYOFFICE_DS_SECRET)")
cmd.Flags().StringVar(&folder, "folder", "2", "scratch folder id for local-file uploads")
cmd.Flags().BoolVar(&stream, "stream", false, "write the converted bytes to stdout (pipe-friendly)")
cmd.Flags().BoolVar(&pipe, "pipe", false, "alias for --stream")
return cmd
}
func docsPresignedCmd() *cobra.Command {
return &cobra.Command{
Use: "presigned FILE_ID",
Short: "Print a short-lived fetchable URI for a portal file",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
u, err := c.PresignedURI(cmd.Context(), args[0])
if err != nil {
return err
}
printObject(map[string]any{"fileid": args[0], "uri": u})
return nil
},
}
}
// loadWorkbookBytes reads an argument that is either an OnlyOffice file id
// (downloaded via the client) or a local path.
func loadWorkbookBytes(cmd *cobra.Command, c *onlyoffice.Client, arg string) ([]byte, error) {
if fi, err := os.Stat(arg); err == nil && !fi.IsDir() {
return os.ReadFile(arg)
}
var buf bytes.Buffer
if _, err := c.DownloadFile(cmd.Context(), arg, &buf); err != nil {
return nil, err
}
return buf.Bytes(), nil
}
// parseDelimiter maps a flag value to a CSV delimiter rune ("," default).
func parseDelimiter(s string) rune {
switch s {
case "", ",":
return ','
case "\\t", "tab", "\t":
return '\t'
case ";":
return ';'
case "|":
return '|'
default:
r := []rune(s)
return r[0]
}
}
func docsCSVCmd() *cobra.Command {
var sheet int
var delim, out string
cmd := &cobra.Command{
Use: "csv SRC [SRC...]",
Short: "Export a worksheet (XLS/XLSX/ODS) to CSV (sheet-aware)",
Long: `SRC is an OnlyOffice file id or a local path. --sheet is 1-based (default 1 = first).
Why not the DocumentServer: its csv output covers only the first worksheet and
ignores a sheet selector (verified). Sheet selection and JSON use a local reader.`,
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
d := parseDelimiter(delim)
for _, arg := range args {
data, err := loadWorkbookBytes(cmd, c, arg)
if err != nil {
return fmt.Errorf("read %s: %w", arg, err)
}
text, err := onlyoffice.WorkbookSheetCSV(data, sheet-1, d)
if err != nil {
return fmt.Errorf("%s: %w", arg, err)
}
if out != "" {
if err := os.WriteFile(out, []byte(text), 0o644); err != nil {
return err
}
printObject(map[string]any{"source": arg, "sheet": sheet, "output": out, "bytes": len(text)})
continue
}
fmt.Print(text)
}
return nil
},
}
cmd.Flags().IntVar(&sheet, "sheet", 1, "worksheet number (1-based, default first)")
cmd.Flags().StringVar(&delim, "delimiter", ",", "CSV delimiter (',', ';', '|', 'tab')")
cmd.Flags().StringVar(&out, "out", "", "write to this file instead of stdout (single input)")
return cmd
}
func docsJSONCmd() *cobra.Command {
var sheet int
var out string
cmd := &cobra.Command{
Use: "json SRC [SRC...]",
Short: "Export a worksheet (XLS/XLSX/ODS) to JSON rows (first row = header)",
Long: `SRC is an OnlyOffice file id or a local path. --sheet is 1-based (default 1 = first).
Each data row becomes an object keyed by the header cells of that sheet.`,
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, arg := range args {
data, err := loadWorkbookBytes(cmd, c, arg)
if err != nil {
return fmt.Errorf("read %s: %w", arg, err)
}
rows, err := onlyoffice.WorkbookSheetJSON(data, sheet-1)
if err != nil {
return fmt.Errorf("%s: %w", arg, err)
}
b, err := json.MarshalIndent(rows, "", " ")
if err != nil {
return err
}
b = append(b, '\n')
if out != "" {
if err := os.WriteFile(out, b, 0o644); err != nil {
return err
}
printObject(map[string]any{"source": arg, "sheet": sheet, "rows": len(rows), "output": out})
continue
}
os.Stdout.Write(b)
}
return nil
},
}
cmd.Flags().IntVar(&sheet, "sheet", 1, "worksheet number (1-based, default first)")
cmd.Flags().StringVar(&out, "out", "", "write to this file instead of stdout (single input)")
return cmd
}
// docsBaseURL resolves the DocumentServer base: --docs-url, $ONLYOFFICE_DOCS_URL,
// else the standard /ds-vpath reverse-proxy mount.
func docsBaseURL(flag string) string {
if flag != "" {
return flag
}
if v := os.Getenv("ONLYOFFICE_DOCS_URL"); v != "" {
return v
}
if v := firstEnv("ONLYOFFICE_URL", "ONLYOFFICE_HOST", "OO_URL"); v != "" {
return strings.TrimRight(v, "/") + "/ds-vpath"
}
return ""
}
func firstEnv(keys ...string) string {
for _, k := range keys {
if v := os.Getenv(k); v != "" {
return v
}
}
return ""
}
func strOrNil(s string) any { func strOrNil(s string) any {
if s == "" { if s == "" {
return nil return nil
@@ -514,64 +809,27 @@ func docsPutTxtCmd() *cobra.Command {
} }
func docsPutXlsxCmd() *cobra.Command { func docsPutXlsxCmd() *cobra.Command {
var folderID, template, title, keepLocal string var folderID, keepLocal string
var replace bool var replace bool
cmd := &cobra.Command{ cmd := &cobra.Command{
Use: "put-xlsx PROJECT_ID [LOCAL_XLSX]", Use: "put-xlsx PROJECT_ID LOCAL_XLSX",
Short: "Upload or generate an XLSX workbook into a project (excelize templates with formulas)", Short: "Upload an XLSX workbook into a project (upsert by stem|ext)",
Long: `Spreadsheets live in OnlyOffice — not in git. Generate multi-sheet workbooks with Long: `Spreadsheets live in OnlyOffice — not in git. Uploads an existing .xlsx into the
formulas (SUM/AVG, cross-sheet refs, named inputs) via --template, or upload an existing .xlsx. project Documents (or --folder), replacing the same stem|ext by default.
oo docs put-xlsx 218 --template cutover-portugal
oo docs put-xlsx 218 --template cutover-portugal --title 2026-08-28-cutover-budget.xlsx
oo docs put-xlsx 218 ./my.xlsx`, oo docs put-xlsx 218 ./my.xlsx`,
Args: cobra.RangeArgs(1, 2), Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
pid := args[0] pid := args[0]
c, err := newOO(cmd) c, err := newOO(cmd)
if err != nil { if err != nil {
return err return err
} }
dir, err := os.MkdirTemp("", "oo-docs-put-xlsx-*") xlsxPath := args[1]
if err != nil {
return err
}
defer os.RemoveAll(dir)
var xlsxPath string
var srcLabel string
switch {
case template != "":
wb, err := xlspipe.BuildTemplate(template)
if err != nil {
return err
}
name := title
if name == "" {
name = "cutover-budget.xlsx"
if template == xlspipe.TemplateCutoverPortugal {
name = "2026-08-28-cutover-budget.xlsx"
}
}
if !strings.HasSuffix(strings.ToLower(name), ".xlsx") {
name += ".xlsx"
}
xlsxPath = filepath.Join(dir, name)
if err := xlspipe.Save(wb, xlsxPath); err != nil {
wb.Close()
return err
}
wb.Close()
srcLabel = "template:" + template
case len(args) == 2:
xlsxPath = args[1]
if docpipe.Ext(xlsxPath) != ".xlsx" { if docpipe.Ext(xlsxPath) != ".xlsx" {
return fmt.Errorf("expected .xlsx, got %s", docpipe.Ext(xlsxPath)) return fmt.Errorf("expected .xlsx, got %s", docpipe.Ext(xlsxPath))
} }
srcLabel = xlsxPath srcLabel := xlsxPath
default:
return fmt.Errorf("pass LOCAL_XLSX or --template")
}
if keepLocal != "" { if keepLocal != "" {
b, err := os.ReadFile(xlsxPath) b, err := os.ReadFile(xlsxPath)
@@ -607,9 +865,7 @@ formulas (SUM/AVG, cross-sheet refs, named inputs) via --template, or upload an
}, },
} }
cmd.Flags().StringVar(&folderID, "folder", "", "Documents folder id (default: project root)") cmd.Flags().StringVar(&folderID, "folder", "", "Documents folder id (default: project root)")
cmd.Flags().StringVar(&template, "template", "", "built-in workbook template (cutover-portugal)") cmd.Flags().StringVar(&keepLocal, "keep-xlsx", "", "also copy the uploaded bytes to this local path")
cmd.Flags().StringVar(&title, "title", "", "upload file name when using --template")
cmd.Flags().StringVar(&keepLocal, "keep-xlsx", "", "also write generated/uploaded bytes to this local path")
cmd.Flags().BoolVar(&replace, "replace", true, "replace same stem|ext before upload (default); false = fail if name taken") cmd.Flags().BoolVar(&replace, "replace", true, "replace same stem|ext before upload (default); false = fail if name taken")
return cmd return cmd
} }
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"strings"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra"
)
func init() {
rootCmd.AddCommand(indexCmd())
}
// indexFlags are shared by the `oo index folder` and `oo index files` verbs.
type indexFlags struct {
recursive bool
exts string
limit int
workers int
lang string
minChars int
workDir string
backend string
dryRun bool
asJSON bool
}
// indexCmd populates the own full-text index (oo_docs_text) that makes PDF and
// scanned content searchable. The OnlyOffice index is left untouched.
func indexCmd() *cobra.Command {
f := &indexFlags{}
cmd := &cobra.Command{
Use: "index",
Short: "Populate the own full-text index for PDF/scan content",
Long: "Index document text that the OnlyOffice Elasticsearch index does not\n" +
"cover (PDFs and scans) into a separate index (ONLYOFFICE_ES_TEXT_INDEX,\n" +
"default oo_docs_text). Text is extracted with docpipe (pdftotext, OCR)\n" +
"and the OnlyOffice server is never modified.\n\n" +
"Requires ONLYOFFICE_URL/USER/PASS (to download files) and ONLYOFFICE_ES_URL\n" +
"(to write the index). See docs/elasticsearch.md.",
}
cmd.PersistentFlags().BoolVar(&f.recursive, "recursive", false, "folder: descend into subfolders")
cmd.PersistentFlags().StringVar(&f.exts, "exts", "pdf", "comma-separated extensions to index")
cmd.PersistentFlags().IntVar(&f.limit, "limit", 0, "maximum number of files to index (0 = all)")
cmd.PersistentFlags().IntVar(&f.workers, "workers", 3, "parallel downloads/extractions")
cmd.PersistentFlags().StringVar(&f.lang, "lang", "deu+eng", "OCR language(s)")
cmd.PersistentFlags().IntVar(&f.minChars, "min-chars", 0, "text-layer threshold below which OCR runs")
cmd.PersistentFlags().StringVar(&f.workDir, "work-dir", "", "temp dir for downloads (default: system temp)")
cmd.PersistentFlags().StringVar(&f.backend, "backend", "rest", "file backend: rest|dav")
cmd.PersistentFlags().BoolVar(&f.dryRun, "dry-run", false, "list what would be indexed, without changes")
cmd.PersistentFlags().BoolVar(&f.asJSON, "json", false, "shorthand for --output json")
cmd.AddCommand(indexFolderCmd(f), indexFilesCmd(f))
return cmd
}
func indexFolderCmd(f *indexFlags) *cobra.Command {
return &cobra.Command{
Use: "folder FOLDER_ID",
Short: "Index every matching file in a Documents folder",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
return runIndex(cmd, f, args[0], nil)
},
}
}
func indexFilesCmd(f *indexFlags) *cobra.Command {
return &cobra.Command{
Use: "files FILE_ID...",
Short: "Index specific Documents files",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
return runIndex(cmd, f, "", args)
},
}
}
func runIndex(cmd *cobra.Command, f *indexFlags, folderID string, ids []string) error {
if f.asJSON {
outputFormat = "json"
}
c, err := newOO(cmd)
if err != nil {
return err
}
idx, err := onlyoffice.NewESTextIndex(onlyoffice.ESTextConfigFromEnv())
if err != nil {
return err
}
ti := onlyoffice.NewTextIndexer(c.FileStore(f.backend), idx)
ti.WorkDir = f.workDir
opts := onlyoffice.IndexOptions{
Recursive: f.recursive,
Extensions: splitList(f.exts),
Limit: f.limit,
Lang: f.lang,
MinChars: f.minChars,
Workers: f.workers,
}
ctx := cmd.Context()
if f.dryRun {
var entries []onlyoffice.Entry
if folderID != "" {
entries, err = ti.PlanFolder(ctx, folderID, opts)
} else {
entries, err = ti.PlanFiles(ctx, ids, opts)
}
if err != nil {
return err
}
rows := make([]map[string]any, 0, len(entries))
for _, e := range entries {
rows = append(rows, map[string]any{
"id": e.ID,
"title": e.Title,
"folder": e.ParentID,
})
}
printTable([]string{"id", "title", "folder"}, rows)
return nil
}
if err := ti.Ensure(ctx); err != nil {
return err
}
var res onlyoffice.IndexResult
if folderID != "" {
res, err = ti.IndexFolder(ctx, folderID, opts)
} else {
res, err = ti.IndexFiles(ctx, ids, opts)
}
if err != nil {
return err
}
printObject(map[string]any{
"index": idx.Index(),
"scanned": res.Scanned,
"indexed": res.Indexed,
"skipped": res.Skipped,
"failed": res.Failed,
"errors": res.Errors,
})
return nil
}
// splitList parses a comma-separated flag value, dropping blanks.
func splitList(s string) []string {
parts := strings.Split(s, ",")
out := make([]string, 0, len(parts))
for _, p := range parts {
if p = strings.TrimSpace(p); p != "" {
out = append(out, p)
}
}
return out
}
+55
View File
@@ -0,0 +1,55 @@
package main
import (
"reflect"
"strings"
"testing"
)
func TestIndexCommandRegistered(t *testing.T) {
cmd, _, err := rootCmd.Find([]string{"index"})
if err != nil {
t.Fatal(err)
}
if cmd.Name() != "index" {
t.Fatalf("index resolved to %q", cmd.Name())
}
for _, name := range []string{"exts", "limit", "workers", "lang", "min-chars", "work-dir", "backend", "dry-run", "recursive", "json"} {
if cmd.PersistentFlags().Lookup(name) == nil {
t.Errorf("index: missing --%s flag", name)
}
}
if cmd.PersistentFlags().Lookup("exts").DefValue != "pdf" {
t.Errorf("--exts default = %q, want pdf", cmd.PersistentFlags().Lookup("exts").DefValue)
}
for _, verb := range []string{"index folder", "index files"} {
sub, _, err := rootCmd.Find(strings.Fields(verb))
if err != nil {
t.Fatalf("%s: %v", verb, err)
}
if sub.Name() != strings.Fields(verb)[1] {
t.Errorf("%s resolved to %q", verb, sub.Name())
}
}
}
func TestIndexSearchBackendFlag(t *testing.T) {
cmd, _, err := rootCmd.Find([]string{"search"})
if err != nil {
t.Fatal(err)
}
if cmd.Flags().Lookup("backend") == nil {
t.Fatal("search: missing --backend flag")
}
if cmd.Flags().Lookup("backend").DefValue != "oo" {
t.Errorf("--backend default = %q, want oo", cmd.Flags().Lookup("backend").DefValue)
}
}
func TestSplitList(t *testing.T) {
got := splitList(" pdf , .PDF, docx ,, ")
want := []string{"pdf", ".PDF", "docx"}
if !reflect.DeepEqual(got, want) {
t.Errorf("splitList = %v, want %v", got, want)
}
}
+3 -3
View File
@@ -91,7 +91,7 @@ func invoiceCreateCmd() *cobra.Command {
Long: `Create a CRM invoice (Draft) with a single line. Long: `Create a CRM invoice (Draft) with a single line.
Always pass --opportunity when a deal exists (entity link at create). Updating Always pass --opportunity when a deal exists (entity link at create). Updating
--opportunity later often fails with HTTP 400 — see docs/crm-associations.md. --opportunity later often fails with HTTP 400 — see the CRM association rules.
Example: Example:
oo invoices create --number INV-2026-01 --contact CONTACT_ID --item ITEM_ID \ oo invoices create --number INV-2026-01 --contact CONTACT_ID --item ITEM_ID \
@@ -175,7 +175,7 @@ func invoiceUpdateCmd() *cobra.Command {
Long: `Update Draft invoice fields. Long: `Update Draft invoice fields.
--opportunity often returns HTTP 400 on existing invoices. Prefer --opportunity often returns HTTP 400 on existing invoices. Prefer
oo invoices create … --opportunity, or delete+recreate. See docs/crm-associations.md. oo invoices create … --opportunity, or delete+recreate. See the CRM association rules.
`, `,
Args: cobra.ExactArgs(1), Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
@@ -274,7 +274,7 @@ func invoiceStatusCmd() *cobra.Command {
Long: `PUT /api/2.0/crm/invoice/status/{id}. Long: `PUT /api/2.0/crm/invoice/status/{id}.
Billed invoices are not content-editable. Billed→Draft often does not work — Billed invoices are not content-editable. Billed→Draft often does not work —
recreate as Draft instead (docs/crm-associations.md). recreate as Draft instead (the CRM association rules).
`, `,
Args: cobra.MinimumNArgs(1), Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
+45
View File
@@ -0,0 +1,45 @@
package main
import (
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra"
)
func init() {
rootCmd.AddCommand(linkCmd())
}
// linkCmd prints deep links for file ids. Used to build third-party document
// packs whose cover embeds links to contracts and supporting files. File ids
// come from `oo projects files list` / `oo dav ls`; `oo projects files
// replace-in` keeps them clean when a document is re-uploaded.
func linkCmd() *cobra.Command {
return &cobra.Command{
Use: "link FILE_ID [FILE_ID...]",
Short: "Print OnlyOffice DocEditor deep links (Products/Files/DocEditor.aspx?fileid=…)",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, id := range args {
printObject(map[string]any{
"fileid": id,
"title": fileTitle(cmd, c, id),
"url": c.FileEditorURL(id),
})
}
return nil
},
}
}
// fileTitle best-effort resolves a file title; never fails the command.
func fileTitle(cmd *cobra.Command, c *onlyoffice.Client, id string) string {
f, err := c.GetFile(cmd.Context(), id)
if err != nil || f == nil || f.Title == nil {
return ""
}
return *f.Title
}
+8 -6
View File
@@ -3,24 +3,26 @@
// Command tree is subject-based (mirrors the library split and the `tea` CLI): // Command tree is subject-based (mirrors the library split and the `tea` CLI):
// //
// oo calendar list | events | add | delete // oo calendar list | events | add | delete
// oo projects list | get | milestones | milestone-create | create | update | delete | contacts (add|remove) | link-authors | link-git | files (list|upload|download|rename|delete|dedupe|as-md|put-md|put-txt|put-xlsx) // oo projects list | get | milestones | milestone-create | milestone-delete | board-sync | create | update | delete | contacts (add|remove) | team (list|add|remove|set) | link-authors | link-git | files (list|upload|replace-in|update|download|rename|delete|dedupe|as-md|put-md|put-txt|put-xlsx)
// oo tasks list | get | create | update | delete | subtask add | files (list|upload|detach) // oo tasks list | get | create | update | delete | subtask add | files (list|upload|detach)
// oo users list | self (alias: oo whoami) // oo users list | self | get | create | update | delete | block | unblock | password | check (alias: oo whoami)
// oo link FILE_ID [FILE_ID...] DocEditor deep links (Products/Files/DocEditor.aspx?fileid=…)
// oo contacts list | get | delete | info-add | merge | dedupe-info | tags | tag-add | tag-create | tag-remove // oo contacts list | get | delete | info-add | merge | dedupe-info | tags | tag-add | tag-create | tag-remove
// oo persons list | create | delete | dedupe // oo persons list | create | delete | dedupe
// oo companies list | create | delete | dedupe | dedupe-persons // oo companies list | create | delete | dedupe | dedupe-persons
// oo opportunities list | get | create | update | delete | stages | member-add | dedupe | dedupe-members | fix-titles // oo opportunities list | get | create | update | delete | stages | member-add | dedupe | dedupe-members | fix-titles
// oo cases list | create | delete | member-add // oo cases list | create | delete | member-add
// oo crm-tasks list | create | delete | categories | reassign-self // oo crm-tasks list | create | delete | categories | reassign-self
// oo crm cleanup // oo crm audit | cleanup
// oo mails accounts | folders | list | get | download-attachment | draft | attach | draft-invoice | send | delete // oo mails accounts | folders | list | get | download-attachment | draft | attach | draft-invoice | send | delete
// oo invoices list | get | create | update | pdf | pdf-cleanup | status | delete | items … // oo invoices list | get | create | update | pdf | pdf-cleanup | status | delete | items …
// oo docs tools | convert | optimize | ocr | hocr | as-md | put-md | put-txt | put-xlsx // oo docs tools | convert | pdf | presigned | csv | json | optimize | ocr | hocr | as-md | put-md | put-txt | put-xlsx
// oo catalog match | merge | apply | scan-contacts | scan-projects | scan-thunderbird // oo catalog match | merge | apply | scan-contacts | scan-projects | scan-thunderbird
// oo dav ls | move | copy | mkdir | rename-file | rename-folder | download | fileops // oo dav ls | move | copy | mkdir | rename-file | rename-folder | download | fileops
// oo search QUERY [--content] [--folder ID] [--limit N] [--json] // oo search QUERY [--content] [--folder ID] [--limit N] [--backend oo|own] [--json]
// oo index folder FOLDER_ID | files FILE_ID... [--recursive] [--exts pdf] [--dry-run]
// //
// CRM association rules: docs/crm-associations.md // CRM association rules live with the private oo-workspace tooling.
// //
// Every list supports `--output/-o json|table` (table is the default). // Every list supports `--output/-o json|table` (table is the default).
// //
+158
View File
@@ -24,10 +24,142 @@ func init() {
projectsCmd.AddCommand(prjGetCmd()) projectsCmd.AddCommand(prjGetCmd())
projectsCmd.AddCommand(prjMilestonesCmd()) projectsCmd.AddCommand(prjMilestonesCmd())
projectsCmd.AddCommand(prjMilestoneCreateCmd()) projectsCmd.AddCommand(prjMilestoneCreateCmd())
projectsCmd.AddCommand(prjMilestoneDeleteCmd())
projectsCmd.AddCommand(prjCreateCmd()) projectsCmd.AddCommand(prjCreateCmd())
projectsCmd.AddCommand(prjUpdateCmd()) projectsCmd.AddCommand(prjUpdateCmd())
projectsCmd.AddCommand(prjDeleteCmd()) projectsCmd.AddCommand(prjDeleteCmd())
projectsCmd.AddCommand(prjContactsCmd()) projectsCmd.AddCommand(prjContactsCmd())
projectsCmd.AddCommand(prjTeamCmd())
}
func prjTeamCmd() *cobra.Command {
cmd := &cobra.Command{
Use: "team",
Short: "Project team (portal users) — CRUD",
Long: `Project team members are portal users (People), not CRM contacts.
CRM companies/persons linked to a project live under 'oo projects contacts'.`,
}
cmd.AddCommand(prjTeamListCmd())
cmd.AddCommand(prjTeamAddCmd())
cmd.AddCommand(prjTeamRemoveCmd())
cmd.AddCommand(prjTeamSetCmd())
return cmd
}
func prjTeamListCmd() *cobra.Command {
return &cobra.Command{
Use: "list PROJECT_ID",
Short: "List portal users on the project team",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
pid, err := strconv.Atoi(args[0])
if err != nil {
return fmt.Errorf("project id: %w", err)
}
c, err := newOO(cmd)
if err != nil {
return err
}
list, err := c.ListProjectTeam(cmd.Context(), pid)
if err != nil {
return err
}
printTable([]string{"id", "displayName", "userName", "email", "isAdmin"}, teamRows(list))
return nil
},
}
}
func prjTeamAddCmd() *cobra.Command {
return &cobra.Command{
Use: "add PROJECT_ID USER_ID [USER_ID...]",
Short: "Add portal user(s) to the project team",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid, err := strconv.Atoi(args[0])
if err != nil {
return fmt.Errorf("project id: %w", err)
}
c, err := newOO(cmd)
if err != nil {
return err
}
for _, uid := range args[1:] {
if _, err := c.AddProjectTeamUser(cmd.Context(), pid, uid); err != nil {
return fmt.Errorf("add %s: %w", uid, err)
}
printObject(map[string]any{"project_id": pid, "user_id": uid, "added": true})
}
return nil
},
}
}
func prjTeamRemoveCmd() *cobra.Command {
return &cobra.Command{
Use: "remove PROJECT_ID USER_ID [USER_ID...]",
Short: "Remove portal user(s) from the project team",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid, err := strconv.Atoi(args[0])
if err != nil {
return fmt.Errorf("project id: %w", err)
}
c, err := newOO(cmd)
if err != nil {
return err
}
for _, uid := range args[1:] {
if _, err := c.RemoveProjectTeamUser(cmd.Context(), pid, uid); err != nil {
return fmt.Errorf("remove %s: %w", uid, err)
}
printObject(map[string]any{"project_id": pid, "user_id": uid, "removed": true})
}
return nil
},
}
}
func prjTeamSetCmd() *cobra.Command {
var notify bool
cmd := &cobra.Command{
Use: "set PROJECT_ID USER_ID [USER_ID...]",
Short: "Replace the project team with the given users (register several at once)",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid, err := strconv.Atoi(args[0])
if err != nil {
return fmt.Errorf("project id: %w", err)
}
c, err := newOO(cmd)
if err != nil {
return err
}
team, err := c.SetProjectTeam(cmd.Context(), pid, args[1:], notify)
if err != nil {
return err
}
printTable([]string{"id", "displayName", "userName", "email", "isAdmin"}, teamRows(team))
return nil
},
}
cmd.Flags().BoolVar(&notify, "notify", false, "notify added members")
return cmd
}
// teamRows maps raw team member maps into table rows.
func teamRows(list []map[string]any) []map[string]any {
rows := make([]map[string]any, 0, len(list))
for _, m := range list {
rows = append(rows, map[string]any{
"id": idString(m, "id"),
"displayName": m["displayName"],
"userName": m["userName"],
"email": m["email"],
"isAdmin": m["isAdmin"],
})
}
return rows
} }
func prjListCmd() *cobra.Command { func prjListCmd() *cobra.Command {
@@ -166,6 +298,32 @@ func prjMilestoneCreateCmd() *cobra.Command {
return cmd return cmd
} }
func prjMilestoneDeleteCmd() *cobra.Command {
return &cobra.Command{
Use: "milestone-delete MILESTONE_ID [MILESTONE_ID...]",
Aliases: []string{"milestone-rm"},
Short: "Delete project milestone(s) by id",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, raw := range args {
id, err := strconv.ParseInt(raw, 10, 64)
if err != nil {
return fmt.Errorf("milestone id %q must be integer: %w", raw, err)
}
if err := c.DeleteMilestone(id); err != nil {
return fmt.Errorf("delete milestone %d: %w", id, err)
}
printObject(map[string]any{"milestone_id": id, "deleted": true})
}
return nil
},
}
}
func prjCreateCmd() *cobra.Command { func prjCreateCmd() *cobra.Command {
var desc, resp string var desc, resp string
var country, company string var country, company string
+54
View File
@@ -0,0 +1,54 @@
package main
import (
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra"
)
func init() {
projectsCmd.AddCommand(prjBoardSyncCmd())
}
// prjBoardSyncCmd upserts project milestones/tasks from a board YAML.
func prjBoardSyncCmd() *cobra.Command {
var apply bool
cmd := &cobra.Command{
Use: "board-sync BOARD.yaml",
Short: "Upsert project milestones/tasks from a board YAML (dry-run by default)",
Long: `Reads a board YAML (projects → milestones → tasks) and creates only the
milestones/tasks that are missing, matching by exact title. Existing entries are
left untouched, so the same file can be re-applied safely.
Dry-run by default; pass --apply to write to OnlyOffice.`,
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
board, err := onlyoffice.LoadBoard(args[0])
if err != nil {
return err
}
c, err := newOO(cmd)
if err != nil {
return err
}
res, err := c.SyncBoard(cmd.Context(), board, apply)
if err != nil {
return err
}
mode := "dry-run"
if apply {
mode = "apply"
}
printObject(map[string]any{
"mode": mode,
"projects": len(board.Projects),
"created_milestones": res.CreatedMilestones,
"skipped_milestones": res.SkippedMilestones,
"created_tasks": res.CreatedTasks,
"skipped_tasks": res.SkippedTasks,
})
return nil
},
}
cmd.Flags().BoolVar(&apply, "apply", false, "write to OnlyOffice (default: dry-run)")
return cmd
}
+101 -6
View File
@@ -3,6 +3,7 @@ package main
import ( import (
"fmt" "fmt"
"os" "os"
"path/filepath"
"strconv" "strconv"
"time" "time"
@@ -21,6 +22,8 @@ func projectFilesCmd() *cobra.Command {
} }
cmd.AddCommand(prjFilesListCmd()) cmd.AddCommand(prjFilesListCmd())
cmd.AddCommand(prjFilesUploadCmd()) cmd.AddCommand(prjFilesUploadCmd())
cmd.AddCommand(prjFilesReplaceInCmd())
cmd.AddCommand(prjFilesUpdateCmd())
cmd.AddCommand(prjFilesDownloadCmd()) cmd.AddCommand(prjFilesDownloadCmd())
cmd.AddCommand(prjFilesRenameCmd()) cmd.AddCommand(prjFilesRenameCmd())
cmd.AddCommand(prjFilesDeleteCmd()) cmd.AddCommand(prjFilesDeleteCmd())
@@ -148,6 +151,68 @@ Pass --no-replace to fail when the name is taken; --allow-duplicate to always cr
return cmd return cmd
} }
func prjFilesReplaceInCmd() *cobra.Command {
return &cobra.Command{
Use: "replace-in FOLDER_ID LOCAL_PATH [LOCAL_PATH...]",
Short: "Replace same-named file(s) in a folder: hard delete + fresh upload (no version history)",
Long: `Deletes any file in FOLDER_ID with the same stem|ext (hard delete — the CLI
delete is permanent) and uploads the local file fresh. Unlike 'update' this
leaves a single clean version.
Why it exists: repeated 'update' of a shared document accumulated a visible
version history and left a stale id. replace-in yields one clean revision; then
point links at the returned id (or keep an nginx alias for the legacy fileid).
Note: file ids are server-assigned; a fresh upload gets a new id.`,
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
folderID := args[0]
for _, p := range args[1:] {
stem := onlyoffice.UploadStemFromLocal(p)
ext := onlyoffice.UploadExtFromLocal(p)
deleted, derr := c.DeleteFilesByDedupKey(cmd.Context(), folderID, stem, ext)
if derr != nil {
return derr
}
ent, uerr := c.UploadToFolder(cmd.Context(), folderID, p)
if uerr != nil {
return uerr
}
obj := fileEntryToMap(ent)
if len(deleted) > 0 {
obj["replaced_file_ids"] = deleted
}
printObject(obj)
}
return nil
},
}
}
func prjFilesUpdateCmd() *cobra.Command {
return &cobra.Command{
Use: "update FILE_ID LOCAL_PATH",
Short: "Overwrite an existing Documents file with new content (new version)",
Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
entry, err := c.UpdateFile(cmd.Context(), args[0], args[1])
if err != nil {
return err
}
printObject(fileEntryToMap(entry))
return nil
},
}
}
func prjFilesDownloadCmd() *cobra.Command { func prjFilesDownloadCmd() *cobra.Command {
var to string var to string
cmd := &cobra.Command{ cmd := &cobra.Command{
@@ -159,20 +224,22 @@ func prjFilesDownloadCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
f, err := c.GetFile(cmd.Context(), args[0]) ctx := cmd.Context()
store := c.Files()
e, err := store.Stat(ctx, args[0])
if err != nil { if err != nil {
return err return err
} }
path := to path := to
if path == "" { if path == "" {
path = onlyoffice.SafeLocalFileName(onlyoffice.FileEntryTitle(f)) path = onlyoffice.SafeLocalFileName(e.Title)
} }
out, err := os.Create(path) out, err := os.Create(path)
if err != nil { if err != nil {
return err return err
} }
defer out.Close() defer out.Close()
n, err := c.DownloadFile(cmd.Context(), args[0], out) n, err := store.Download(ctx, args[0], out)
if err != nil { if err != nil {
_ = os.Remove(path) _ = os.Remove(path)
return err return err
@@ -199,11 +266,16 @@ func prjFilesRenameCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
entry, err := c.RenameFile(cmd.Context(), args[0], args[1]) store := c.Files()
if err := store.Rename(cmd.Context(), args[0], args[1]); err != nil {
return err
}
entry, err := store.Stat(cmd.Context(), args[0])
if err != nil { if err != nil {
return err return err
} }
printObject(fileEntryToMap(entry)) entry.Title = args[1]
printObject(entryToMap(entry))
return nil return nil
}, },
} }
@@ -228,7 +300,7 @@ func prjFilesDeleteCmd() *cobra.Command {
} }
ids = append(ids, id) ids = append(ids, id)
} }
if err := c.DeleteFiles(cmd.Context(), ids); err != nil { if err := c.Files().Delete(cmd.Context(), args); err != nil {
return err return err
} }
printObject(map[string]any{"deleted": ids}) printObject(map[string]any{"deleted": ids})
@@ -317,6 +389,29 @@ func fileEntryToMap(f *onlyoffice.FileEntry) map[string]any {
return m return m
} }
// entryToMap renders a canonical Entry with the same keys as fileEntryToMap.
func entryToMap(e onlyoffice.Entry) map[string]any {
m := map[string]any{
"id": e.ID,
"title": e.Title,
"fileExst": filepath.Ext(e.Title),
"contentLength": contentLengthString(e.Size),
}
if e.Updated != "" {
m["updated"] = e.Updated
} else if !e.Modified.IsZero() {
m["updated"] = e.Modified.Format(time.RFC3339)
}
return m
}
func contentLengthString(n int64) string {
if n <= 0 {
return ""
}
return strconv.FormatInt(n, 10)
}
func fileIDStr(f *onlyoffice.FileEntry) string { func fileIDStr(f *onlyoffice.FileEntry) string {
if f == nil || f.ID == nil { if f == nil || f.ID == nil {
return "" return ""
+40 -10
View File
@@ -1,7 +1,11 @@
package main package main
import ( import (
"fmt"
"strings"
onlyoffice "github.com/eslider/go-onlyoffice" onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/eslider/go-onlyoffice/cmd/internal/bootstrap"
"github.com/spf13/cobra" "github.com/spf13/cobra"
) )
@@ -9,46 +13,70 @@ func init() {
rootCmd.AddCommand(searchCmd()) rootCmd.AddCommand(searchCmd())
} }
// searchCmd queries the OnlyOffice Elasticsearch index directly. The REST // searchCmd queries the OnlyOffice document index through the file facade. The
// /api/2.0/files/@search endpoint only searches file names in the database; // REST /api/2.0/files/@search endpoint only searches file names in the database;
// content search needs ES (see docs/elasticsearch.md). // content search needs Elasticsearch (see docs/elasticsearch.md).
func searchCmd() *cobra.Command { func searchCmd() *cobra.Command {
var ( var (
content bool content bool
folder string folder string
limit int limit int
backend string
asJSON bool asJSON bool
substring bool
) )
cmd := &cobra.Command{ cmd := &cobra.Command{
Use: "search QUERY", Use: "search QUERY...",
Short: "Full-text search over documents by name, optionally by content (Elasticsearch)", Short: "Full-text search over documents by name, optionally by content (Elasticsearch)",
Long: "Search the OnlyOffice Documents index.\n\n" + Long: "Search the OnlyOffice Documents index.\n\n" +
"By default only file names are matched. With --content the query also\n" + "By default only file names are matched. With --content the query also\n" +
"matches extracted document text (document.attachment.content); this covers\n" + "matches extracted document text (document.attachment.content); this covers\n" +
"Office formats (docx/xlsx/pptx) and is slower.\n\n" + "Office formats (docx/xlsx/pptx) and is slower.\n\n" +
"--backend own queries the separate index populated by `oo index`\n" +
"(ONLYOFFICE_ES_TEXT_INDEX, default oo_docs_text) instead, which also holds\n" +
"PDFs and scans (see docs/elasticsearch.md).\n\n" +
"Requires ONLYOFFICE_ES_URL (and optionally ONLYOFFICE_ES_INDEX,\n" + "Requires ONLYOFFICE_ES_URL (and optionally ONLYOFFICE_ES_INDEX,\n" +
"ONLYOFFICE_TENANT). See docs/elasticsearch.md for the tunnel setup.", "ONLYOFFICE_TENANT). See docs/elasticsearch.md for the tunnel setup.",
Args: cobra.ExactArgs(1), Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
if asJSON { if asJSON {
outputFormat = "json" outputFormat = "json"
} }
es, err := onlyoffice.NewESSearcher(onlyoffice.ESConfigFromEnv()) bootstrap.LoadEnv()
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
var (
searcher onlyoffice.Searcher
err error
)
switch strings.ToLower(strings.TrimSpace(backend)) {
case "", "oo", "elasticsearch":
searcher, err = c.Files().Search()
case "own", "es-text":
searcher, err = onlyoffice.NewESTextIndex(onlyoffice.ESTextConfigFromEnv())
default:
return fmt.Errorf("unknown search backend %q (want oo|own)", backend)
}
if err != nil { if err != nil {
return err return err
} }
hits, err := es.Search(cmd.Context(), onlyoffice.SearchQuery{ hits, err := searcher.Search(cmd.Context(), onlyoffice.SearchQuery{
Text: args[0], Text: strings.Join(args, " "),
InContent: content, InContent: content,
FolderID: folder, FolderID: folder,
Limit: limit, Limit: limit,
Substring: substring,
}) })
if err != nil { if err != nil {
return err return err
} }
rows := make([]map[string]any, 0, len(hits)) rows := make([]map[string]any, 0, len(hits))
for _, h := range hits { for _, h := range hits {
folderPath := h.Path
if len(folderPath) == 0 && h.ParentID != "" {
folderPath = []string{h.ParentID}
}
rows = append(rows, map[string]any{ rows = append(rows, map[string]any{
"path": c.UniquePath(cmd.Context(), folderPath, h.Title),
"id": h.ID, "id": h.ID,
"title": h.Title, "title": h.Title,
"folder": h.ParentID, "folder": h.ParentID,
@@ -60,13 +88,15 @@ func searchCmd() *cobra.Command {
printJSON(rows) printJSON(rows)
return nil return nil
} }
printTable([]string{"id", "title", "folder", "score", "highlight"}, rows) printTable([]string{"path", "id", "title", "folder", "score", "highlight"}, rows)
return nil return nil
}, },
} }
cmd.Flags().BoolVar(&content, "content", false, "also match extracted document content") cmd.Flags().BoolVar(&content, "content", false, "also match extracted document content")
cmd.Flags().StringVar(&folder, "folder", "", "limit to a Documents folder id") cmd.Flags().BoolVar(&substring, "substring", false, "case-insensitive *term* title match; multiple QUERY args are ANDed")
cmd.Flags().StringVar(&folder, "folder", "", "limit to a Documents folder id (matches the folder subtree)")
cmd.Flags().IntVar(&limit, "limit", 20, "maximum number of results") cmd.Flags().IntVar(&limit, "limit", 20, "maximum number of results")
cmd.Flags().StringVar(&backend, "backend", "oo", "index to query: oo (OnlyOffice) | own (oo index)")
cmd.Flags().BoolVar(&asJSON, "json", false, "shorthand for --output json") cmd.Flags().BoolVar(&asJSON, "json", false, "shorthand for --output json")
return cmd return cmd
} }
+308
View File
@@ -1,6 +1,12 @@
package main package main
import ( import (
"bufio"
"fmt"
"os"
"strings"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra" "github.com/spf13/cobra"
) )
@@ -13,6 +19,14 @@ func init() {
rootCmd.AddCommand(usersCmd) rootCmd.AddCommand(usersCmd)
usersCmd.AddCommand(usersListCmd()) usersCmd.AddCommand(usersListCmd())
usersCmd.AddCommand(usersSelfCmd()) usersCmd.AddCommand(usersSelfCmd())
usersCmd.AddCommand(usersGetCmd())
usersCmd.AddCommand(usersCreateCmd())
usersCmd.AddCommand(usersUpdateCmd())
usersCmd.AddCommand(usersDeleteCmd())
usersCmd.AddCommand(usersBlockCmd())
usersCmd.AddCommand(usersUnblockCmd())
usersCmd.AddCommand(usersPasswordCmd())
usersCmd.AddCommand(usersCheckCmd())
rootCmd.AddCommand(whoamiCmd()) rootCmd.AddCommand(whoamiCmd())
} }
@@ -65,6 +79,300 @@ func usersSelfCmd() *cobra.Command {
} }
} }
func usersGetCmd() *cobra.Command {
return &cobra.Command{
Use: "get USER_ID",
Short: "Show one portal user profile",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
u, err := c.GetUser(cmd.Context(), args[0])
if err != nil {
return err
}
printObject(u)
return nil
},
}
}
func usersCreateCmd() *cobra.Command {
var first, last, email, password, title, location, sex, comment string
var visitor bool
cmd := &cobra.Command{
Use: "create",
Short: "Create a portal user",
Long: `Create a portal user (POST /api/2.0/people).
Without --password the portal generates one and the account stays NotActivated
until the user follows the activation link. With --password the account is
Active immediately. Use --visitor for a guest account.`,
RunE: func(cmd *cobra.Command, args []string) error {
if email == "" || first == "" || last == "" {
return fmt.Errorf("--email, --first and --last are required")
}
c, err := newOO(cmd)
if err != nil {
return err
}
req := onlyoffice.NewUserRequest{
FirstName: first,
LastName: last,
Email: email,
Password: password,
Title: title,
Location: location,
Sex: sex,
Comment: comment,
}
if cmd.Flags().Changed("visitor") {
req.IsVisitor = &visitor
}
u, err := c.CreateUser(cmd.Context(), req)
if err != nil {
return err
}
printObject(map[string]any{
"id": idString(u, "id"),
"displayName": idString(u, "displayName"),
"email": idString(u, "email"),
"status": u["status"],
})
return nil
},
}
cmd.Flags().StringVar(&first, "first", "", "first name (required)")
cmd.Flags().StringVar(&last, "last", "", "last name (required)")
cmd.Flags().StringVar(&email, "email", "", "email (required)")
cmd.Flags().StringVar(&password, "password", "", "initial password (default: portal-generated)")
cmd.Flags().StringVar(&title, "title", "", "job title")
cmd.Flags().StringVar(&location, "location", "", "location")
cmd.Flags().StringVar(&sex, "sex", "", "sex: male|female")
cmd.Flags().StringVar(&comment, "comment", "", "comment")
cmd.Flags().BoolVar(&visitor, "visitor", false, "create as guest (isVisitor=true)")
return cmd
}
func usersUpdateCmd() *cobra.Command {
var first, last, email, title, location, sex, comment string
cmd := &cobra.Command{
Use: "update USER_ID",
Short: "Update portal user profile fields (only flags passed)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
body := map[string]any{}
if cmd.Flags().Changed("first") {
body["firstname"] = first
}
if cmd.Flags().Changed("last") {
body["lastname"] = last
}
if cmd.Flags().Changed("email") {
body["email"] = email
}
if cmd.Flags().Changed("title") {
body["title"] = title
}
if cmd.Flags().Changed("location") {
body["location"] = location
}
if cmd.Flags().Changed("sex") {
body["sex"] = sex
}
if cmd.Flags().Changed("comment") {
body["comment"] = comment
}
if len(body) == 0 {
return fmt.Errorf("nothing to update: pass at least one of --first/--last/--email/--title/--location/--sex/--comment")
}
c, err := newOO(cmd)
if err != nil {
return err
}
u, err := c.UpdateUser(cmd.Context(), args[0], body)
if err != nil {
return err
}
printObject(map[string]any{
"id": idString(u, "id"),
"displayName": idString(u, "displayName"),
})
return nil
},
}
cmd.Flags().StringVar(&first, "first", "", "first name")
cmd.Flags().StringVar(&last, "last", "", "last name")
cmd.Flags().StringVar(&email, "email", "", "email")
cmd.Flags().StringVar(&title, "title", "", "job title")
cmd.Flags().StringVar(&location, "location", "", "location")
cmd.Flags().StringVar(&sex, "sex", "", "sex: male|female")
cmd.Flags().StringVar(&comment, "comment", "", "comment")
return cmd
}
func usersDeleteCmd() *cobra.Command {
return &cobra.Command{
Use: "delete USER_ID [USER_ID...]",
Aliases: []string{"rm"},
Short: "Delete portal user(s) permanently",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, id := range args {
u, err := c.DeleteUser(cmd.Context(), id)
if err != nil {
return fmt.Errorf("delete %s: %w", id, err)
}
printObject(map[string]any{"id": id, "deleted": true, "displayName": idString(u, "displayName")})
}
return nil
},
}
}
func usersBlockCmd() *cobra.Command {
return &cobra.Command{
Use: "block USER_ID [USER_ID...]",
Aliases: []string{"disable"},
Short: "Block (terminate) user(s): login denied, profile kept",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, id := range args {
if err := c.BlockUser(cmd.Context(), id); err != nil {
return fmt.Errorf("block %s: %w", id, err)
}
printObject(map[string]any{"id": id, "blocked": true})
}
return nil
},
}
}
func usersUnblockCmd() *cobra.Command {
return &cobra.Command{
Use: "unblock USER_ID [USER_ID...]",
Aliases: []string{"enable", "activate"},
Short: "Unblock (activate) user(s)",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, id := range args {
if err := c.UnblockUser(cmd.Context(), id); err != nil {
return fmt.Errorf("unblock %s: %w", id, err)
}
printObject(map[string]any{"id": id, "unblocked": true})
}
return nil
},
}
}
func usersPasswordCmd() *cobra.Command {
var password string
cmd := &cobra.Command{
Use: "password USER_ID",
Short: "Set a user password (reads stdin when --password is empty)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
pwd := password
if pwd == "" {
b, err := readLine(os.Stdin)
if err != nil {
return fmt.Errorf("read password: %w", err)
}
pwd = b
}
if pwd == "" {
return fmt.Errorf("password is empty")
}
c, err := newOO(cmd)
if err != nil {
return err
}
if err := c.ChangeUserPassword(cmd.Context(), args[0], pwd); err != nil {
return err
}
printObject(map[string]any{"id": args[0], "password_changed": true})
return nil
},
}
cmd.Flags().StringVar(&password, "password", "", "new password (omit to read one line from stdin)")
return cmd
}
func usersCheckCmd() *cobra.Command {
var login, password string
cmd := &cobra.Command{
Use: "check",
Short: "Check that a login can authenticate (userName or email)",
Long: `Probes POST /api/2.0/authentication.json with the given credentials and
discards the token.
Where this is used: before handing portal credentials to an external party
(e.g. a guest given read access to a document pack), verify the login actually
works. On some portals the account email is the reliable login identifier — the
userName login fails with 500 for a freshly created user — so share the email,
not the userName.`,
RunE: func(cmd *cobra.Command, args []string) error {
if login == "" {
return fmt.Errorf("--login is required (userName or email)")
}
if password == "" {
if b, err := readLine(os.Stdin); err == nil {
password = b
}
}
c, err := newOO(cmd)
if err != nil {
return err
}
if err := c.AuthenticateAs(cmd.Context(), login, password); err != nil {
printObject(map[string]any{"login": login, "ok": false, "error": trimAuthErr(err)})
return fmt.Errorf("login failed for %s", login)
}
printObject(map[string]any{"login": login, "ok": true})
return nil
},
}
cmd.Flags().StringVar(&login, "login", "", "userName or email")
cmd.Flags().StringVar(&password, "password", "", "password (omit to read one line from stdin)")
return cmd
}
// trimAuthErr keeps the error short for table output.
func trimAuthErr(err error) string {
s := err.Error()
if len(s) > 160 {
s = s[:160] + "…"
}
return s
}
// readLine reads a single trimmed line from r.
func readLine(r *os.File) (string, error) {
sc := bufio.NewScanner(r)
if !sc.Scan() {
if err := sc.Err(); err != nil {
return "", err
}
return "", nil
}
return strings.TrimSpace(sc.Text()), nil
}
// whoamiCmd is a convenience shortcut at the root level. // whoamiCmd is a convenience shortcut at the root level.
func whoamiCmd() *cobra.Command { func whoamiCmd() *cobra.Command {
return &cobra.Command{ return &cobra.Command{
-50
View File
@@ -1,50 +0,0 @@
// Command ooscan recursively lists OnlyOffice Documents folders into a TSV
// index: file_id, folder_id, path, title.
//
// Usage: ooscan <FOLDER_ID> [<FOLDER_ID>...]
package main
import (
"context"
"fmt"
"os"
"time"
onlyoffice "github.com/eslider/go-onlyoffice"
)
func main() {
ctx := context.Background()
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
seen := map[string]bool{}
for _, root := range os.Args[1:] {
walk(ctx, c, root, "", 0, seen)
}
}
func walk(ctx context.Context, c *onlyoffice.Client, folderID, path string, depth int, seen map[string]bool) {
if depth > 8 || seen[folderID] {
return
}
seen[folderID] = true
// Throttle: OnlyOffice rate-limits (429) and the host must not be flooded.
time.Sleep(350 * time.Millisecond)
ctx, cancel := context.WithTimeout(ctx, 60*time.Second)
defer cancel()
var l *onlyoffice.DavListing
derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
var err error
l, err = c.ListDavFolder(ctx, folderID)
return err
})
if derr != nil {
fmt.Fprintf(os.Stderr, "list %s (%s): %v\n", path, folderID, derr)
return
}
for _, f := range l.Files {
fmt.Printf("%s\t%s\t%s\t%s\n", f.ID, folderID, path, f.Title)
}
for _, sub := range l.Folders {
walk(ctx, c, sub.ID, path+"/"+sub.Title, depth+1, seen)
}
}
-308
View File
@@ -1,308 +0,0 @@
// Command pdfamount walks a Documents folder, downloads matching PDFs and
// extracts the payable amount, printing "file_id\ttitle\tamount".
//
// Usage: pdfamount <FOLDER_ID> [TITLE_FILTER_REGEX]
package main
import (
"bytes"
"context"
"fmt"
"os"
"os/exec"
"regexp"
"strconv"
"strings"
"time"
onlyoffice "github.com/eslider/go-onlyoffice"
)
// amountPat is the amount capture shared by every amount regex.
const amountPat = `([0-9]+(?:[.,][0-9]+)*)`
// amountRE builds "<label> [optional (comment)] [: -] <number>".
func amountRE(label string) *regexp.Regexp {
return regexp.MustCompile(
`(?i)\b` + regexp.QuoteMeta(label) + `\b\s*(?:\([^)]*\))?\s*[:\-]?\s*` + amountPat)
}
// amountRes lists the payable-amount patterns in strict priority order: the
// first pattern with a usable amount wins, and a lower-priority label can never
// override a higher-priority one ("zu zahlender betrag" > "rechnungsbetrag" >
// "rechnungsendbetrag" > "gesamtbetrag" > "gesamtsumme (inkl. steuern)").
//
// "gesamtbetrag" and "gesamtsumme" are not in the original set but are the real
// labels on Diashop invoices ("Gesamtsumme (inkl. Steuern)"). The inclusive
// variant is matched before a plain "gesamtsumme". Everything after those
// primary labels is the broader fallback set, consulted only when no primary
// label yields an amount. Within one pattern the last usable amount is taken,
// because totals usually come last.
var amountRes = []*regexp.Regexp{
amountRE("zu zahlender betrag"),
amountRE("rechnungsbetrag"),
amountRE("rechnungsendbetrag"),
amountRE("gesamtbetrag"),
regexp.MustCompile(`(?i)\bgesamtsumme\b\s*\(\s*inkl\.?\s*steuern\s*\)\s*[:\-]?\s*` + amountPat),
amountRE("gesamtsumme"),
amountRE("endbetrag"),
amountRE("zahlbetrag"),
amountRE("bruttobetrag"),
amountRE("betrag"),
amountRE("total"),
amountRE("summe"),
}
// taxLineRe marks a line whose number is a tax rate/percentage: an explicit
// percent sign or a VAT/tax keyword. "Steuern" (plural, as in "inkl. Steuern")
// is handled separately so the inclusive total stays usable.
var taxLineRe = regexp.MustCompile(`(?i)%|\bMwSt\b|\bUSt\b|\bProzent\b`)
// steuerRe finds "Steuer"/"Umsatzsteuer" etc. RE2 has no lookahead, so the
// plural "Steuern" is excluded in isTaxLine.
var steuerRe = regexp.MustCompile(`(?i)steuer`)
// percentAfterRe detects a percent sign directly after a number (spaces ok).
var percentAfterRe = regexp.MustCompile(`^\s*%`)
func main() {
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, "usage: pdfamount <FOLDER_ID> [TITLE_FILTER_REGEX]")
os.Exit(2)
}
folder := os.Args[1]
filter := regexp.MustCompile(`(?i)rechnung`)
if len(os.Args) >= 3 {
filter = regexp.MustCompile(os.Args[2])
}
ctx := context.Background()
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
files := listAll(ctx, c, folder)
for _, f := range files {
if !filter.MatchString(f.title) {
continue
}
if !strings.HasSuffix(strings.ToLower(f.title), ".pdf") {
continue
}
amount, err := pdfAmount(ctx, c, f.id)
if err != nil {
fmt.Fprintf(os.Stderr, "%s: %v\n", f.title, err)
continue
}
if amount == "" {
continue
}
fmt.Printf("%s\t%s\t%s\n", f.id, f.title, amount)
}
}
type file struct{ id, title string }
func listAll(ctx context.Context, c *onlyoffice.Client, folder string) []file {
seen := map[string]bool{}
var out []file
var walk func(string)
walk = func(id string) {
if seen[id] {
return
}
seen[id] = true
time.Sleep(300 * time.Millisecond)
l, err := c.ListDavFolder(ctx, id)
if err != nil {
fmt.Fprintf(os.Stderr, "list %s: %v\n", id, err)
return
}
for _, f := range l.Files {
out = append(out, file{f.ID, f.Title})
}
for _, sub := range l.Folders {
walk(sub.ID)
}
}
walk(folder)
return out
}
func pdfAmount(ctx context.Context, c *onlyoffice.Client, id string) (string, error) {
time.Sleep(time.Second)
tmp, err := os.CreateTemp("", "pdf-*.pdf")
if err != nil {
return "", err
}
defer os.Remove(tmp.Name())
derr := onlyoffice.DoRetry(ctx, onlyoffice.DefaultRetryPolicy(), func() error {
_ = tmp.Truncate(0)
_, _ = tmp.Seek(0, 0)
_, err := c.DownloadFile(ctx, id, tmp)
return err
})
if derr != nil {
tmp.Close()
return "", derr
}
tmp.Close()
var buf bytes.Buffer
cmd := exec.CommandContext(ctx, "pdftotext", "-layout", tmp.Name(), "-")
cmd.Stdout = &buf
if err := cmd.Run(); err != nil {
return "", err
}
return extractAmount(buf.String()), nil
}
// extractAmount returns the normalised ("1234.56") payable amount found in
// text, or "" if no usable amount matches.
//
// DKV invoices are special-cased first: they repeat a per-vehicle "TOTAL:" line
// and carry the real total only in the "Gesamtsummenaufstellung" section.
func extractAmount(text string) string {
if v, ok := dkvGrandTotal(text); ok {
return v
}
for _, re := range amountRes {
if v, ok := lastUsableAmount(text, re); ok {
return v
}
}
return ""
}
// dkvGrandTotal extracts the total of a DKV "Gesamtsummenaufstellung" section.
//
// Rule: DKV invoices repeat a per-vehicle "TOTAL:" line, so the last TOTAL is
// not the invoice total. When a "Gesamtsummenaufstellung" section exists, its
// total wins over every "TOTAL:" line: the first amount after the "»" marker,
// or, if there is none, the last amount in the section. The section ends at the
// page break (form feed) or end of text.
func dkvGrandTotal(text string) (string, bool) {
idx := strings.Index(strings.ToLower(text), "gesamtsummenaufstellung")
if idx < 0 {
return "", false
}
section := text[idx:]
if ff := strings.IndexByte(section, '\f'); ff >= 0 {
section = section[:ff]
}
if m := strings.Index(section, "»"); m >= 0 {
if v, ok := firstAmount(section[m:]); ok {
return v, true
}
}
return lastAmount(section)
}
// lastUsableAmount returns the last amount matched by re that is not a tax rate
// or percentage. Within one label the last usable amount wins.
func lastUsableAmount(text string, re *regexp.Regexp) (string, bool) {
ms := re.FindAllStringSubmatchIndex(text, -1)
for i := len(ms) - 1; i >= 0; i-- {
m := ms[i]
if isTaxRate(text, m[2], m[3]) {
continue
}
if v, ok := normalizeAmount(text[m[2]:m[3]]); ok {
return v, true
}
}
return "", false
}
// isTaxRate reports whether the number at text[start:end] is a tax rate or a
// percentage instead of a payable amount. A candidate is rejected when the
// token right after the number is "%" or the number's line carries a percent
// sign or a tax keyword. Rejecting is deliberate: office matching treats a
// known-but-different amount as a hard disqualifier, so an empty result is
// safer than the VAT rate.
func isTaxRate(text string, start, end int) bool {
if percentAfterRe.MatchString(text[end:]) {
return true
}
lineStart := strings.LastIndexByte(text[:start], '\n') + 1
line := text[lineStart:]
if n := strings.IndexByte(text[end:], '\n'); n >= 0 {
line = text[lineStart : end+n]
}
return isTaxLine(line)
}
// isTaxLine reports whether a line looks like a tax rate rather than a payable
// amount. "Steuern" is treated as a qualifier ("inkl. Steuern"), not a rate.
func isTaxLine(line string) bool {
if taxLineRe.MatchString(line) {
return true
}
for _, loc := range steuerRe.FindAllStringIndex(line, -1) {
if loc[1] >= len(line) || (line[loc[1]] != 'n' && line[loc[1]] != 'N') {
return true
}
}
return false
}
// numberRe finds bare numbers (with optional thousands/decimal separators).
var numberRe = regexp.MustCompile(`[0-9]+(?:[.,][0-9]+)*`)
func firstAmount(s string) (string, bool) {
for _, m := range numberRe.FindAllString(s, -1) {
if v, ok := normalizeAmount(m); ok {
return v, true
}
}
return "", false
}
func lastAmount(s string) (string, bool) {
ms := numberRe.FindAllString(s, -1)
for i := len(ms) - 1; i >= 0; i-- {
if v, ok := normalizeAmount(ms[i]); ok {
return v, true
}
}
return "", false
}
// normalizeAmount turns "1.234,56" (DE), "1,234.56" (EN) or "1234.56" into
// "1234.56". The rightmost separator is decimal only when followed by one or
// two digits; otherwise every separator is a thousands separator.
func normalizeAmount(s string) (string, bool) {
last := -1
for i := 0; i < len(s); i++ {
if s[i] == '.' || s[i] == ',' {
last = i
}
}
var dec byte
if last >= 0 {
digits := 0
for i := last + 1; i < len(s); i++ {
if s[i] < '0' || s[i] > '9' {
return "", false
}
digits++
}
if digits == 1 || digits == 2 {
dec = s[last]
}
}
var b strings.Builder
for i := 0; i < len(s); i++ {
switch c := s[i]; {
case c >= '0' && c <= '9':
b.WriteByte(c)
case (c == '.' || c == ',') && c == dec:
b.WriteByte('.')
case c == '.' || c == ',':
// thousands separator
default:
return "", false
}
}
v, err := strconv.ParseFloat(b.String(), 64)
if err != nil {
return "", false
}
return strconv.FormatFloat(v, 'f', 2, 64), true
}
-175
View File
@@ -1,175 +0,0 @@
package main
import "testing"
func TestExtractAmount(t *testing.T) {
tests := []struct {
name, text, want string
}{
{
name: "rechnungsbetrag de format",
text: "Rechnungsbetrag: 1.234,56 €",
want: "1234.56",
},
{
name: "rechnungsbetrag en thousands and dot",
text: "Rechnungsbetrag: 1,234.56",
want: "1234.56",
},
{
name: "rechnungsbetrag plain dot",
text: "Rechnungsbetrag: 1234.56",
want: "1234.56",
},
{
name: "rechnungsbetrag de comma only",
text: "Rechnungsbetrag: 1234,56",
want: "1234.56",
},
{
name: "currency suffix eur",
text: "Rechnungsbetrag: 1.234,56 EUR",
want: "1234.56",
},
{
name: "zu zahlender betrag wins over rechnungsbetrag",
text: "Zu zahlender Betrag: 10,00\nRechnungsbetrag: 99,00",
want: "10.00",
},
{
name: "rechnungsbetrag wins over endbetrag",
text: "Endbetrag: 20,00\nRechnungsbetrag: 30,00",
want: "30.00",
},
{
name: "bruttobetrag wins over bare betrag",
text: "Bruttobetrag: 50,00\nBetrag: 10,00",
want: "50.00",
},
{
name: "gesamtbetrag wins over bare betrag",
text: "Gesamtbetrag: 80,00\nBetrag: 10,00",
want: "80.00",
},
{
name: "last occurrence of same label wins",
text: "Rechnungsbetrag: 10,00\nRechnungsbetrag: 20,00",
want: "20.00",
},
{
name: "endbetrag fallback",
text: "Endbetrag: 42,00",
want: "42.00",
},
{
name: "zahlbetrag fallback without colon",
text: "Zahlbetrag 7,50 €",
want: "7.50",
},
{
name: "rechnungsendbetrag beats endbetrag",
text: "Rechnungsendbetrag: 12,00\nEndbetrag: 13,00",
want: "12.00",
},
{
name: "dkv style total line",
text: "Kundenbezogene Daten\n» TOTAL: 123,45 100,00 23,45 123,45\n",
want: "123.45",
},
{
name: "dkv gesamtsummenaufstellung grand total after marker",
text: "» TOTAL: 111,11 100,00 11,11 111,11\n" +
"» TOTAL: 222,22 200,00 22,22 222,22\n" +
"Gesamtsummenaufstellung\n" +
"Netto 240,00\n" +
"MwSt 47,25\n" +
"» 287,25\n",
want: "287.25",
},
{
name: "dkv gesamtsummenaufstellung total on next line",
text: "» TOTAL: 111,11\nGesamtsummenaufstellung\n»\n287,25\n",
want: "287.25",
},
{
name: "tax rate with percent sign is not an amount",
text: "Betrag: 19,00 % MwSt",
want: "",
},
{
name: "mehrwertsteuer rate is not an amount",
text: "Gesamtsumme: 19,00% MwSt",
want: "",
},
{
name: "steuer word on the number line rejects it",
text: "Betrag: 2,83 Steuer",
want: "",
},
{
name: "rejected primary falls back to a usable label",
text: "Gesamtsumme: 19,00 % MwSt\nEndbetrag: 42,00",
want: "42.00",
},
{
name: "labeled zu zahlender betrag beats unlabeled larger number",
text: "unlabeled 999,99\nZu zahlender Betrag: 10,00",
want: "10.00",
},
{
name: "labeled zu zahlender betrag beats lower label larger number",
text: "Endbetrag: 999,99\nZu zahlender Betrag: 10,00",
want: "10.00",
},
{
name: "diashop style gesamtsumme with comment",
text: "Zwischensumme\n12,34 €\nZwischensumme\n12,34 €\nVersand & Bearbeitung\n4,95 €\nGesamtsumme (inkl. Steuern)\n17,29 €\n",
want: "17.29",
},
{
name: "diashop picks inclusive total last",
text: "Gesamtsumme (exkl. Steuern)\n12,34 €\nGesamtsumme (inkl. Steuern)\n17,29 €",
want: "17.29",
},
{
name: "no label",
text: "some text without any amount label 12,34",
want: "",
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
if got := extractAmount(tc.text); got != tc.want {
t.Fatalf("extractAmount()=%q want %q", got, tc.want)
}
})
}
}
func TestNormalizeAmount(t *testing.T) {
tests := []struct {
in string
want string
ok bool
}{
{"1.234,56", "1234.56", true},
{"1,234.56", "1234.56", true},
{"1234.56", "1234.56", true},
{"1234,56", "1234.56", true},
{"1.234.567,89", "1234567.89", true},
{"1,234,567.89", "1234567.89", true},
{"1.234", "1234.00", true},
{"12,5", "12.50", true},
{"12", "12.00", true},
{"", "0.00", false},
}
for _, tc := range tests {
got, ok := normalizeAmount(tc.in)
if ok != tc.ok {
t.Fatalf("normalizeAmount(%q) ok=%v want %v", tc.in, ok, tc.ok)
}
if ok && got != tc.want {
t.Fatalf("normalizeAmount(%q)=%q want %q", tc.in, got, tc.want)
}
}
}
+170
View File
@@ -0,0 +1,170 @@
package onlyoffice
// Document conversion via the OnlyOffice DocumentServer converter.
//
// The DocumentServer (the same engine behind the portal's "Download as PDF")
// converts any office format. From a portal-reachable host the converter is
// exposed at "<portal>/ds-vpath/converter" (reverse proxy) or directly at
// "http://<docs-server>:8083/converter" (legacy path: /ConvertService.ashx).
//
// Flow: PresignedURI(fileId) → Convert(docsBase, secret, req) → download
// result.FileURL. The JWT is HS256 signed with the DocumentServer's
// services.CoAuthoring.secret (NOT storage.fs.secretString).
import (
"bytes"
"context"
"crypto/hmac"
"crypto/sha256"
"encoding/base64"
"encoding/json"
"fmt"
"io"
"net/http"
"net/url"
"strings"
)
// PresignedURI returns a short-lived, fetchable download URI for a portal file
// (GET /api/2.0/files/file/{fileId}/presigneduri). The DocumentServer can fetch
// it without the caller's session, so it is the input for Convert.
func (c *Client) PresignedURI(ctx context.Context, fileID string) (string, error) {
if fileID == "" {
return "", fmt.Errorf("file id is required")
}
raw, err := c.getJSON(ctx, fmt.Sprintf("/api/2.0/files/file/%s/presigneduri", url.PathEscape(fileID)))
if err != nil {
return "", err
}
resp, err := responseField(raw, "response")
if err != nil {
return "", err
}
var s string
if err := json.Unmarshal(resp, &s); err == nil && s != "" {
return s, nil
}
// Some builds return an object instead of a bare string.
var o map[string]any
if err := json.Unmarshal(resp, &o); err == nil {
for _, k := range []string{"uri", "url", "Uri", "Url"} {
if v, ok := o[k].(string); ok && v != "" {
return v, nil
}
}
}
return "", fmt.Errorf("presigneduri: unexpected response %s", truncate(string(resp), 200))
}
// ConvertRequest is the DocumentServer converter body.
type ConvertRequest struct {
URL string `json:"url"`
OutputType string `json:"outputtype"`
FileType string `json:"filetype,omitempty"`
Key string `json:"key"`
Title string `json:"title,omitempty"`
}
// ConvertResult is the DocumentServer converter reply.
type ConvertResult struct {
FileURL string `json:"fileUrl"`
FileType string `json:"fileType"`
Percent int `json:"percent"`
EndConvert bool `json:"endConvert"`
Error *int `json:"error,omitempty"`
}
// SignJWT builds an HS256 JWT with the given payload (stdlib only).
func SignJWT(secret string, payload any) (string, error) {
if secret == "" {
return "", fmt.Errorf("jwt secret is empty")
}
hb, err := json.Marshal(map[string]string{"alg": "HS256", "typ": "JWT"})
if err != nil {
return "", err
}
pb, err := json.Marshal(payload)
if err != nil {
return "", err
}
enc := base64.RawURLEncoding.EncodeToString
signing := enc(hb) + "." + enc(pb)
mac := hmac.New(sha256.New, []byte(secret))
mac.Write([]byte(signing))
return signing + "." + enc(mac.Sum(nil)), nil
}
// ConvertDocument asks a DocumentServer to convert req.URL into req.OutputType.
// docsBase is e.g. "https://portal/ds-vpath" or "http://localhost:8083";
// secret is the DocumentServer CoAuthoring JWT secret. Passes the JWT both as
// the AuthorizationJwt header and as a body token.
func (c *Client) ConvertDocument(ctx context.Context, docsBase, secret string, req ConvertRequest) (*ConvertResult, error) {
if strings.TrimSpace(docsBase) == "" {
return nil, fmt.Errorf("docs base url is required")
}
if req.URL == "" {
return nil, fmt.Errorf("source url is required")
}
if req.OutputType == "" {
return nil, fmt.Errorf("outputtype is required")
}
if req.Key == "" {
return nil, fmt.Errorf("conversion key is required")
}
jwt, err := SignJWT(secret, req)
if err != nil {
return nil, err
}
body, err := json.Marshal(req)
if err != nil {
return nil, err
}
endpoint := strings.TrimRight(docsBase, "/") + "/converter"
httpReq, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(body))
if err != nil {
return nil, err
}
httpReq.Header.Set("Content-Type", "application/json")
httpReq.Header.Set("Accept", "application/json")
httpReq.Header.Set("AuthorizationJwt", "Bearer "+jwt)
resp, err := c.client.Do(httpReq)
if err != nil {
return nil, fmt.Errorf("converter request: %w", err)
}
defer resp.Body.Close()
raw, err := io.ReadAll(resp.Body)
if err != nil {
return nil, err
}
if resp.StatusCode >= 400 {
return nil, fmt.Errorf("converter: %d %s", resp.StatusCode, truncate(string(raw), 300))
}
var out ConvertResult
if err := json.Unmarshal(raw, &out); err != nil {
return nil, fmt.Errorf("converter decode: %w (%s)", err, truncate(string(raw), 200))
}
if out.Error != nil {
return &out, fmt.Errorf("converter error %d", *out.Error)
}
if out.FileURL == "" {
return &out, fmt.Errorf("converter returned no fileUrl")
}
return &out, nil
}
// DownloadURLTo streams an absolute URL (no portal auth) into dst.
func (c *Client) DownloadURLTo(ctx context.Context, rawurl string, dst io.Writer) (int64, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, rawurl, nil)
if err != nil {
return 0, err
}
resp, err := c.client.Do(req)
if err != nil {
return 0, err
}
defer resp.Body.Close()
if resp.StatusCode >= 400 {
return 0, fmt.Errorf("download: %d", resp.StatusCode)
}
return io.Copy(dst, resp.Body)
}
+24
View File
@@ -0,0 +1,24 @@
package onlyoffice
import (
"strings"
"testing"
)
func TestSignJWT(t *testing.T) {
payload := map[string]any{"url": "u", "outputtype": "pdf"}
tok, err := SignJWT("secret", payload)
if err != nil {
t.Fatal(err)
}
if n := len(strings.Split(tok, ".")); n != 3 {
t.Fatalf("JWT must have 3 parts, got %d", n)
}
tok2, _ := SignJWT("secret", payload)
if tok != tok2 {
t.Fatal("SignJWT must be deterministic for identical input")
}
if _, err := SignJWT("", payload); err == nil {
t.Fatal("expected error for empty secret")
}
}
+91
View File
@@ -0,0 +1,91 @@
package onlyoffice
import (
"fmt"
"strconv"
"strings"
)
// addressCategoryCodes maps ASC.CRM.Core.AddressCategory names to their numeric
// codes. The OO API expects the code, the UI/docs use the label.
var addressCategoryCodes = map[string]int{
"home": 0,
"postal": 1,
"office": 2,
"billing": 3,
"other": 4,
"work": 5,
}
// AddressCategoryCode returns the numeric code for an AddressCategory label
// (Home|Postal|Office|Billing|Other|Work) or a numeric string. Unknown/empty
// labels fall back to Billing, the category `oo companies create` used.
func AddressCategoryCode(category string) int {
s := strings.ToLower(strings.TrimSpace(category))
if s == "" {
return addressCategoryCodes["billing"]
}
if n, err := strconv.Atoi(s); err == nil {
if n >= 0 && n <= 5 {
return n
}
return addressCategoryCodes["billing"]
}
if n, ok := addressCategoryCodes[s]; ok {
return n
}
return addressCategoryCodes["billing"]
}
// ContactAddresses returns the postal address rows of a contact map.
func ContactAddresses(contact map[string]any) []map[string]any {
if rows, ok := contact["addresses"].([]any); ok {
return mapsFromAnySlice(rows)
}
if rows, ok := contact["addresses"].([]map[string]any); ok {
return rows
}
return nil
}
// HasContactAddress reports whether a contact already has the given postal
// address. street+city+zip+category identify it; comparison is normalized.
func HasContactAddress(contact map[string]any, street, city, zip, category string) bool {
wantStreet, wantCity, wantZip := normalizeAddressPart(street), normalizeAddressPart(city), normalizeAddressPart(zip)
wantCat := AddressCategoryCode(category)
for _, row := range ContactAddresses(contact) {
if normalizeAddressPart(fmt.Sprint(row["street"])) != wantStreet {
continue
}
if normalizeAddressPart(fmt.Sprint(row["city"])) != wantCity {
continue
}
if normalizeAddressPart(fmt.Sprint(row["zip"])) != wantZip {
continue
}
if int(anyToFloat(row["category"])) != wantCat {
continue
}
return true
}
return false
}
func normalizeAddressPart(s string) string {
s = strings.ToLower(strings.TrimSpace(s))
return strings.Join(strings.Fields(s), " ")
}
func anyToFloat(v any) float64 {
switch n := v.(type) {
case float64:
return n
case int:
return float64(n)
case string:
f, _ := strconv.ParseFloat(strings.TrimSpace(n), 64)
return f
default:
return 0
}
}
+42
View File
@@ -0,0 +1,42 @@
package onlyoffice
import "testing"
func TestAddressCategoryCode(t *testing.T) {
cases := map[string]int{
"Home": 0, "Postal": 1, "Office": 2, "Billing": 3, "Other": 4, "Work": 5,
"billing": 3, " work ": 5, "3": 3, "5": 5,
"": 3, "nonsense": 3, "99": 3,
}
for in, want := range cases {
if got := AddressCategoryCode(in); got != want {
t.Errorf("AddressCategoryCode(%q) = %d, want %d", in, got, want)
}
}
}
func TestHasContactAddress(t *testing.T) {
contact := map[string]any{
"addresses": []any{
map[string]any{
"street": "Lubanas st. 125a-25", "city": "Riga",
"zip": "LV-1021", "country": "Latvia", "category": float64(3),
},
},
}
if !HasContactAddress(contact, " Lubanas St. 125a-25 ", "riga", "lv-1021", "Billing") {
t.Error("want match (normalized, case-insensitive)")
}
if HasContactAddress(contact, "Lubanas st. 125a-25", "Riga", "LV-1021", "Work") {
t.Error("different category must not match")
}
if HasContactAddress(contact, "Lubanas st. 125a-25", "Riga", "00000", "Billing") {
t.Error("different zip must not match")
}
if HasContactAddress(map[string]any{}, "x", "y", "z", "Billing") {
t.Error("empty contact must not match")
}
if got := ContactAddresses(contact); len(got) != 1 {
t.Fatalf("ContactAddresses = %d rows", len(got))
}
}
+126
View File
@@ -0,0 +1,126 @@
package onlyoffice
import (
"context"
"fmt"
"strconv"
"strings"
)
// OpportunityAudit is one CRM opportunity with its resource counts and a coarse
// class, for hygiene reporting (see `oo crm audit`).
type OpportunityAudit struct {
ID int64 `json:"id"`
Title string `json:"title"`
Created string `json:"created,omitempty"`
Files int `json:"files"`
OpenTasks int `json:"open_tasks"`
ClosedTasks int `json:"closed_tasks"`
Members int `json:"members"`
GroupKey string `json:"group_key,omitempty"`
Class string `json:"class"` // ok | dup | empty | junk-title
}
// AuditOpportunities lists every opportunity with file/task/member counts and a
// coarse class. The classification is generic and rule-free:
//
// dup — another opportunity shares the same title key
// empty — no files, tasks or members
// junk-title — title is not of the "Role @ Company" shape
// ok — everything else
//
// Callers that need stricter business rules can post-process the result.
func (c *Client) AuditOpportunities(ctx context.Context) ([]OpportunityAudit, error) {
deals, err := c.ListAllOpportunities(ctx)
if err != nil {
return nil, err
}
tasks, _, err := c.ListCRMTasks(ctx, 5000, 0)
if err != nil {
return nil, err
}
open, closed := taskCountsByOpportunity(tasks)
out := make([]OpportunityAudit, 0, len(deals))
for _, row := range deals {
id := auditID(row["id"])
if id == 0 {
continue
}
title := auditStr(row["title"])
files := 0
if fl, ferr := c.ListOpportunityFiles(ctx, strconv.FormatInt(id, 10)); ferr == nil {
files = len(fl)
}
key := strconv.FormatInt(id, 10)
out = append(out, OpportunityAudit{
ID: id,
Title: title,
Created: auditStr(row["created"]),
Files: files,
OpenTasks: open[key],
ClosedTasks: closed[key],
Members: len(OpportunityMembers(row)),
GroupKey: DealTitleKey(title, false),
})
}
groupCount := map[string]int{}
for _, a := range out {
groupCount[a.GroupKey]++
}
for i := range out {
a := &out[i]
switch {
case groupCount[a.GroupKey] > 1:
a.Class = "dup"
case a.Files == 0 && a.OpenTasks == 0 && a.ClosedTasks == 0 && a.Members == 0:
a.Class = "empty"
case !strings.Contains(a.Title, "@") || strings.HasPrefix(strings.TrimSpace(a.Title), "@"):
a.Class = "junk-title"
default:
a.Class = "ok"
}
}
return out, nil
}
// taskCountsByOpportunity buckets CRM task statuses per opportunity id.
func taskCountsByOpportunity(tasks []map[string]any) (open, closed map[string]int) {
open, closed = map[string]int{}, map[string]int{}
for _, t := range tasks {
ent, ok := t["entity"].(map[string]any)
if !ok || auditStr(ent["entityType"]) != "opportunity" {
continue
}
eid := auditStr(ent["entityId"])
status := strings.ToLower(auditStr(t["status"]))
if status == "2" || status == "closed" {
closed[eid]++
} else {
open[eid]++
}
}
return open, closed
}
func auditID(v any) int64 {
switch x := v.(type) {
case float64:
return int64(x)
case int:
return int64(x)
case int64:
return x
default:
n, _ := strconv.ParseInt(strings.TrimSpace(fmt.Sprint(x)), 10, 64)
return n
}
}
func auditStr(v any) string {
if v == nil {
return ""
}
return fmt.Sprint(v)
}
+36
View File
@@ -0,0 +1,36 @@
package onlyoffice
import "testing"
func TestAuditID(t *testing.T) {
cases := []struct {
in any
want int64
}{
{float64(12), 12},
{7, 7},
{int64(9), 9},
{"42", 42},
{nil, 0},
}
for _, c := range cases {
if got := auditID(c.in); got != c.want {
t.Fatalf("auditID(%v) = %d, want %d", c.in, got, c.want)
}
}
}
func TestTaskCountsByOpportunity(t *testing.T) {
open, closed := taskCountsByOpportunity([]map[string]any{
{"id": 1, "status": 1, "entity": map[string]any{"entityType": "opportunity", "entityId": 10}},
{"id": 2, "status": "2", "entity": map[string]any{"entityType": "opportunity", "entityId": 10}},
{"id": 3, "status": 1, "entity": map[string]any{"entityType": "contact", "entityId": 10}},
{"id": 4, "entity": "not-a-map"},
})
if open["10"] != 1 || closed["10"] != 1 {
t.Fatalf("open=%v closed=%v", open, closed)
}
if len(open) != 1 || len(closed) != 1 {
t.Fatalf("unexpected buckets: open=%v closed=%v", open, closed)
}
}
+27
View File
@@ -0,0 +1,27 @@
---
type: reference
status: current
related:
- README.md
---
# go-onlyoffice — docs
Индекс справочников. Общее — [README.md](../README.md), правила — [AGENTS.md](../AGENTS.md).
## Файлы
- [unified-file-client.md](unified-file-client.md) — единый файловый клиент:
`Entry`/`FileStore`/`FileClient`, бэкенды REST/DAV/SQL/ES, env, как добавить
бэкенд.
- [community-server-db.md](community-server-db.md) — read-only SQL-бэкенд
(MySQL/PostgreSQL): схема, SSH-туннель, DSN, MinIO download.
- [elasticsearch.md](elasticsearch.md) — поиск: индекс OnlyOffice `files_file`
и свой `oo_docs_text` (PDF/сканы), туннель.
- [index-and-search.md](index-and-search.md) — карта контуров поиска и как
обновлять индексы (`oo index`, `oo search`).
- [rate-limiting.md](rate-limiting.md) — rate limit, exponential backoff,
`Retry-After`, общий cooldown против 429; env `OO_RATE_LIMIT`/`OO_BURST`/
`OO_RETRY_*`.
## Тесты
Команды и туннели — раздел Testing в [README.md](../README.md#testing).
+179
View File
@@ -0,0 +1,179 @@
---
type: reference
status: current
related:
- README.md
- filestore_pg.go
- docs/elasticsearch.md
---
# Community Server DB — прямой SQL-доступ (read-only)
## Что это
Бэкенд `pgStore` (`filestore_pg.go`) читает файлы и папки **напрямую из БД
Community Server**, без HTTP-слоя. Реализует `FileStore` (`List`/`Stat`/
`Download`) и `Searcher` по имени. Запись запрещена: все write-методы
возвращают `ErrReadOnly`.
## Что за БД (research, live)
Проверено на VM `onlyoffice-v2` (SSH `127.0.0.1:32`):
- Community Server работает на **MySQL 8.0**, не на PostgreSQL.
- Хост: `127.0.0.1:3306` внутри VM, база `onlyoffice`.
- Конфиг: `/etc/onlyoffice/communityserver/appsettings.production.json`,
`providerName: MySql.Data.MySqlClient`.
- Таблицы: `files_file`, `files_folder`, `files_folder_tree`,
`files_security`, тенанты — `tenants_tenants` (не `tenants`).
- PostgreSQL 16 в той же VM — **наш** контур (`edw_docs`, роли `edw`/`edw_ro`,
office-assistant), к OnlyOffice отношения не имеет. `files_file` в PG нет.
- Портал хранит файлы в **S3/MinIO** (DiscStorage только для мелочи).
Бакет `office`, объект — по ключу (см. ниже).
Вывод: бэкенд назван по issue «PostgreSQL», но живой источник — MySQL.
`database/sql` + драйвер по DSN: `mysql` для MySQL, `pgx` для PostgreSQL.
`Name()` возвращает фактический движок (`mysql` или `postgres`).
## Схема
`files_file` — одна строка **на версию** (PK `tenant_id, id, version`):
| поле | смысл |
|------|-------|
| `id` | id файла (тот же, что в REST/ES) |
| `version` | номер версии этой строки |
| `version_group` | номер версии |
| `current_version` | `1` = текущая версия, `0` = старая |
| `folder_id` | id родительской папки |
| `title` | имя файла с расширением |
| `content_length` | размер в байтах |
| `create_on`, `modified_on` | даты (UTC, без зоны) |
| `tenant_id` | тенант (портал) |
`files_folder`: `id`, `parent_id`, `title`, `create_on`, `modified_on`,
`tenant_id`. `files_folder_tree`: `folder_id`, `parent_id`, `level` — готовое
дерево, пока не используется.
Текущую строку файла берём по `current_version = 1`.
## Доступ (SSH-туннель)
MySQL слушает только `127.0.0.1:3306` внутри VM. Снаружи — SSH-туннель
(SSH в VM открыт как `127.0.0.1:32`):
```bash
ssh -f -N -o ControlMaster=no -o ControlPath=none \
-p 32 -i ~/.ssh/id_ed25519 \
-L 3306:127.0.0.1:3306 root@127.0.0.1
# MySQL DSN затем:
# root:<pw>@tcp(127.0.0.1:3306)/onlyoffice?parseTime=true
```
Любой свободный локальный порт подойдёт (напр. `13306`); тогда тот же порт —
в DSN. `-o ControlMaster=no -o ControlPath=none` обязательны: иначе forward
уходит в persistent master из `~/.ssh/config`.
Креды MySQL — в конфиге Community Server внутри VM:
`/etc/onlyoffice/communityserver/appsettings.production.json` →
`ConnectionStrings.connectionString` (поля `User ID`, `Password`), база
`onlyoffice`. В самом MySQL-контейнере (`onlyoffice-mysql-server`) база пустая;
рабочий сервер — host-mysqld на `127.0.0.1:3306` (207 таблиц). Не печатать
пароль.
## Переменные
| env | default | смысл |
|-----|---------|-------|
| `ONLYOFFICE_DSN` | — | DSN драйвера (MySQL `...@tcp(...)/...` или `postgres://...`) |
| `ONLYOFFICE_PG_DRIVER` | авто | `postgres` или `mysql`; иначе по форме DSN |
| `ONLYOFFICE_PG_TENANT` | `ONLYOFFICE_TENANT` | фильтр `tenant_id` (пусто = все) |
| `ONLYOFFICE_PG_HOST/PORT/USER/PASSWORD/DBNAME/SSLMODE` | — | собрать PG DSN, если `ONLYOFFICE_DSN` пуст |
Имена — в [`.env.example`](../.env.example). Секретов нет.
## Использование
Напрямую: `NewPGStore(PGConfigFromEnv())`.
Через фасад (эпик #34): SQL-стор регистрируется на `FileClient`. После этого
`Read()` и все чтения (`Stat`/`List`) идут в БД, `Write()` остаётся REST/DAV.
```go
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
sql, err := c.SQLFileStore() // открыть из env; caller закрывает
if err != nil { /* нет DSN / нет связи */ }
if closer, ok := sql.(interface{ Close() error }); ok { defer closer.Close() }
f := c.Files()
f.RegisterStore(onlyoffice.ProviderPG, sql)
e, _ := f.Stat(ctx, "19423") // e.Provider == "mysql" — ответил SQL
```
`Client.FileStore("pg"|"sql"|"postgres"|"mysql")` тоже отдаёт SQL-стор
(открывает из env). Если DSN нет/битый — возвращается не `nil`, а заглушка,
чей метод отдаёт ошибку открытия; ошибку как таковую даёт `SQLFileStore()`.
Различить бэкенд в ответе можно по `Entry.Provider` (`mysql` у SQL, `rest` у
REST).
## Download (MinIO)
`Download` не ходит в REST. Ключ объекта собирается из строки `files_file`:
```
00/00/<tenant>/files/folder_<shard>/file_<id>/v<version>/content.<ext>
shard = (id/1000 + 1) * 1000
```
`shard` — не `folder_id`, а следующая тысяча над `id` (файл 3727 →
`folder_4000`). Проверено live по бакету `office`.
Стриминг переиспользует `downloadMinioObject` из `storage_fallback.go`
(та же подпись SigV4 и `MINIO_*`), без дублирования.
Ограничение: схема валидна только для файлов, лежащих в **MinIO/S3** (старые
папки). Файлы в **Disc**-хранилище портала (`Data/Products/Files/...`, новые
папки) по этому ключу недоступны — `Download` вернёт `404`. Если
`MINIO_ACCESS_KEY`/`MINIO_SECRET_KEY` не заданы, `Download` вернёт явную
ошибку; `Stat`/`List`/`Search` работают и без них.
## Тесты
```bash
go test ./... # unit: rebind, csObjectKey, маппинг
go test -race ./...
# integration (нужен DSN; skip без него)
ONLYOFFICE_DSN='root:<pw>@tcp(127.0.0.1:3306)/onlyoffice?parseTime=true' \
ONLYOFFICE_PG_TENANT=1 \
ONLYOFFICE_PG_TEST_FILE_ID=19423 \
ONLYOFFICE_PG_TEST_FOLDER_ID=676 \
go test -tags=integration -run 'TestIntegrationPGStore|TestIntegrationSQLFacade' -v ./...
# плюс MINIO_* для сверки Download с REST (иначе этот шаг skip)
MINIO_ENDPOINT=http://127.0.0.1:9000 MINIO_BUCKET=office \
MINIO_ACCESS_KEY=... MINIO_SECRET_KEY=... \
go test -tags=integration -run TestIntegrationPGStore -v ./...
```
- `TestIntegrationPGStore` — `Stat`/`List`/`Download` SQL против REST и
`ErrReadOnly` у write-методов.
- `TestIntegrationSQLFacade` — SQL-стор, зарегистрированный на фасаде, реально
обслуживает чтения: `Read().Name()` = SQL-бэкенд, `Entry.Provider == "mysql"`
(у REST — `"rest"`), сверка `Stat`/`List` с REST, и прямой
`Client.FileStore("pg")`.
Без `ONLYOFFICE_DSN` оба теста делают чистый `skip`.
## Грабли
- MySQL хранит `datetime` без зоны; `parseTime=true` (ставится автоматически)
читает их как UTC. REST отдаёт `+02:00` — сравнивать моменты, не строки.
- `GetFile` (REST) не отдаёт `contentLength` — размер сверять с `Stat` SQL.
- Один файл = много строк `files_file` (по версиям). Без `current_version = 1`
получите дубликаты.
- `folder_id` не входит в ключ MinIO; ключ считает `shard` от `id`.
- Searcher SQL ищет только по имени (`LIKE`). Контент — Elasticsearch
([elasticsearch.md](elasticsearch.md)).
-124
View File
@@ -1,124 +0,0 @@
# CRM associations (company ↔ person ↔ deal ↔ project ↔ invoice ↔ mail)
Operational rules for the `oo` CLI and this library. Business SSOT remains
OnlyOffice Workspace CRM + Projects.
## Canonical graph
One **legal company** owns the relationship. Do not invent a second “bill-to”
company just for PDF layout.
```text
Company
├── Person (buyer contact) oo persons create --company-id
├── Opportunity / Deal oo opportunities … ; member-add company + person
├── Project (hub) oo projects … ; contacts add company + person
│ └── Epic + subtasks
└── Invoice (Draft → …) oo invoices create --contact COMPANY --opportunity DEAL
└── PDF file oo invoices pdf ID
└── Mail draft oo mails draft-invoice --invoice ID --to …
```
| Layer | CLI | Must link |
|-------|-----|-----------|
| Company | `oo companies create` | website, email, phone, **one** Billing address |
| Person | `oo persons create --company-id` / `oo persons update ID` | job title; never encode employer in `lastName`; **update uses JSON** (form PUT ignores `companyId`/`about`) |
| Deal | `oo opportunities create` + `member-add` | company **and** person as members |
| Project | `oo projects create` + `contacts add` | same company + person |
| Invoice | `oo invoices create --contact COMPANY --opportunity DEAL` | `entityId` at **create** |
| Mail | `oo mails draft-invoice` | attach current PDF; **do not send** until confirmed |
UI checks (same company card):
- `#contacts` → person
- `#deals` → opportunity
- `#projects` → hub project
- `#invoices` on the **deal** → invoice (needs `entity`)
- `#files` → preferably **one** current invoice PDF
**Project Team ≠ Project Contacts.** Team = portal users. CRM people/companies
show under the project **Contacts** tab (`oo projects contacts list`).
## Hard rules
1. **One company per legal entity.** Duplicate “bill-to” contacts empty Deals /
Projects / Contacts tabs and break merge. Prefer
`oo contacts merge FROM INTO` (keeps `INTO`) or `oo companies dedupe`.
2. **Link invoice → deal at create.**
`POST /crm/invoice` with `entityId` + `entityType: 0` (Opportunity).
`oo invoices update … --opportunity` often returns **400**
(“Value does not fall within the expected range”). If the link is missing,
delete the Draft and recreate with `--opportunity`.
3. **Bill To = company id**, not a throwaway contact. Person stays under the
company (`companyId`). Optional `consigneeId` for Empfänger when the portal
template prints it.
4. **Stay Draft until mail is ready.** Billed (`status id=2`) is **not editable**
via content PUT. Going Billed → Draft via `…/crm/invoice/status/1` usually
**does not work** — delete + recreate Draft instead.
5. **Do not regenerate PDF in a loop** without cleanup. Each
`GET …/crm/invoice/{id}/pdf` attaches a new file to the company (and often
the deal). Keep `invoice.fileID`; delete older PDFs with
`oo invoices pdf-cleanup ID` / Documents `fileops/delete`.
## Invoice PDF quirks
| Symptom | Workaround |
|---------|------------|
| Cached / stale PDF | Touch invoice (Draft PUT that clears `fileID`), then `GET …/pdf` — `oo invoices pdf ID --force` |
| Billing address missing on **new** PDFs | Temporary multiline `companyName` (`Line1\nLine2\n…`) on the **canonical** company → force PDF → restore clean name. Cached `fileID` keeps the multiline Bill To. |
| Separate bill-to company for newlines | **Forbidden** — merge back to the real company |
| Invoice **number** won’t change on PUT | Delete Draft and recreate with the desired number |
| Notizen / Bedingungen spacing | Leading `\n` and blank lines only — no HTML (tags print literally) |
| Issuer street lines | Organisation profile address (`street` with `\n`), not only terms |
Status ids commonly used: `1` Draft, `2` Billed, `3` Rejected, `4` Paid.
## Mail quirks
| Symptom | Workaround |
|---------|------------|
| Signature / body doubles chat URL | Put chat in **one** place only. UI drafts: signature. API send: body (API **does not** append signature). |
| Signature / body cuts URL at `#` | Plain text URLs — avoid `<a href="…#…">` (or encode `#` as `%23` in href) |
| German letter spacing | Blank `<p>&nbsp;</p>` between blocks (`MailHTMLWithBlankParagraphs`) |
| Send | `PUT /api/2.0/mail/messages/send.json` with `id/from/to/subject/body`; omit empty `cc`/`bcc`. Never auto-send; draft only until the human confirms |
Prefer OnlyOffice Mail (`/addons/mail/#drafts`) for invoice delivery until confirmed.
## Project / task quirks
- Hub title: `CC | Company` (e.g. `DE | Acme GmbH`).
- Streams = epics/tasks under the hub, not a third title segment (unless the
project itself is a named delivery stream).
- Closing a **subtask**:
`PUT /api/2.0/project/task/{epicId}/{subtaskId}/status` with `status=2`.
`oo tasks update SUBTASK -s closed` returns **404** for subtasks.
- After deleting a CRM contact, `GET /project/contact/{deletedId}` may still
return projects (ghost). Official project contact list should only show live
ids; unlink may 400 if the contact is gone.
## Merge / cleanup cheat sheet
```bash
# Keep the preferred company (INTO), drop the duplicate (FROM)
oo contacts merge FROM_ID INTO_ID
# Or by normalized name (careful — whole CRM)
oo companies dedupe
# Invoice ↔ deal must exist at create
oo invoices create --number P-YYYY-NN --contact COMPANY_ID --item ITEM_ID \
--price 300 --opportunity DEAL_ID --language de-DE …
# Fresh PDF + prune older PDFs on company/deal
oo invoices pdf INVOICE_ID --force
oo invoices pdf-cleanup INVOICE_ID
# Mail draft (no send)
oo mails draft-invoice --invoice INVOICE_ID --to billing@example.com
```
## Related
- README § invoices / mail / CRM cleanup
- Personal workspace tooling (disk inventory, dossier sync): private
`git.produktor.io/eSlider/oo-workspace` (`oow` CLI)
+140 -5
View File
@@ -3,7 +3,7 @@ type: reference
status: current status: current
related: related:
- README.md - README.md
- file_es.go - filestore_es.go
--- ---
# Elasticsearch — полнотекстовый поиск OnlyOffice # Elasticsearch — полнотекстовый поиск OnlyOffice
@@ -83,9 +83,14 @@ oo search "Rechnung" --folder 649 --limit 50 --json
`folders.folderId`), `--limit N` (по умолчанию 20, максимум 200), `folders.folderId`), `--limit N` (по умолчанию 20, максимум 200),
`--json` = `-o json`. `--json` = `-o json`.
## Обновление индекса и карта поиска
Обзор всех контуров поиска и как обновлять индексы (`oo index`) —
[index-and-search.md](index-and-search.md).
## Библиотека ## Библиотека
`file_es.go` — `ESSearcher` (`Name() = "elasticsearch"`), прямой ES REST на `filestore_es.go` — `ESSearcher` (`Name() = "elasticsearch"`), прямой ES REST на
stdlib `net/http`: stdlib `net/http`:
```go ```go
@@ -98,7 +103,7 @@ hits, _ := es.Search(ctx, onlyoffice.SearchQuery{
Запрос: `multi_match` по `title^2` (+ `document.attachment.content` при Запрос: `multi_match` по `title^2` (+ `document.attachment.content` при
`InContent`), фильтры `tenantId` и `folders.folderId`, `_source` `InContent`), фильтры `tenantId` и `folders.folderId`, `_source`
id/title/folders, `highlight` для фрагмента. Ответ → `[]SearchHit` (модель из id/title/folders, `highlight` для фрагмента. Ответ → `[]SearchHit` (модель из
эпика #34; пока объявлена в `file_es.go`, переедет в `file_core.go` с F1 #35). эпика #34; пока объявлена в `filestore_es.go`, переедет в `filestore_core.go` с F1 #35).
## Тесты ## Тесты
@@ -121,5 +126,135 @@ ONLYOFFICE_ES_URL=http://127.0.0.1:9200 ONLYOFFICE_TENANT=1 \
- `locale`/версия ES: 7.16.3, `_search` совместим с REST 7.x. - `locale`/версия ES: 7.16.3, `_search` совместим с REST 7.x.
- ES без auth и слушает только localhost — туннель обязателен. - ES без auth и слушает только localhost — туннель обязателен.
- Фильтр `tenantId` сузит выдачу; без него видны документы всех тенантов. - Фильтр `tenantId` сузит выдачу; без него видны документы всех тенантов.
- Поиск по содержимому PDF, залитых через API, не работает (нет - Поиск по содержимому PDF в индексе OnlyOffice не работает (для PDF нет
`attachment.content`) — только Office-форматы. `attachment.content`) — только Office-форматы. Решение для PDF — свой индекс
`oo_docs_text` (F6 #42), см. ниже.
# PDF и сканы — свой индекс (F6 #42)
## Проблема
`oo search --content "S1019"` не находил номер внутри PDF-счёта: в индексе
OnlyOffice PDF лежит только по имени.
## Почему PDF исключён (исходники CommunityServer)
Разобрано в `ONLYOFFICE/CommunityServer`:
- `web/core/ASC.Web.Core/Files/FileUtility.cs` — `CanIndex(fileName)` читает
серверную настройку `files.index.formats` (в `web/studio/ASC.Web.Studio/web.appsettings.config`
значение по умолчанию `".pptx|.xlsx|.docx"`).
- `web/studio/ASC.Web.Studio/Products/Files/Core/Search/FilesWrapper.cs` —
`GetDocumentStream*` возвращает `null`, если `!FileUtility.CanIndex(Title)`,
файл зашифрован или больше `MaxFileSize`.
- `module/ASC.ElasticSearch/Core/WrapperWithDoc.cs` + mapping в `Wrapper.cs` —
маппинг `document.attachment.content` и ingest-pipeline `attachments`
формат-агностичны: они распарсят любой поток.
Вывод: PDF исключён **только настройкой** `files.index.formats`; жёсткого
ограничения на формат в коде нет.
## Варианты и решение
| # | Вариант | Оценка |
|---|---------|--------|
| a | Включить `.pdf` в `files.index.formats` + reindex | Правка сервера OO; настройка может потеряться при обновлении; полный reindex 39k док-в; Tika **не OCR** — сканы без текстового слоя дадут пустой контент. Отклонён без решения PO. |
| b | Server-side ingest/attachment для PDF | По факту то же, что (a): сервер кормит поток только для `CanIndex`. |
| c | **Свой индекс** `oo_docs_text`, наполняемый `internal/docpipe` | **Выбран.** Сервер OO не трогаем; детерминированно; работает OCR для сканов; независимо от обновлений OO; любые форматы; фильтры папка/тип. |
| d | Локальный поиск без индекса | Отклонён как основной: качаем и извлекаем на каждый запрос, нет выдачи/ранжирования/highlight. |
Итог: **вариант c**. Индекс OnlyOffice (`files_file`) не изменяется; наш
индекс живёт рядом.
## Устройство
- `filestore_es_text.go` — `ESTextIndex` (`Name() = "es-text"`):
`Ensure` (создаёт индекс с явным маппингом), `Put` (bulk, `refresh`),
`Delete` (по `id`), `Search` (`multi_match` по `title^2` + `content`,
фильтры `folder`/`ext`, highlight).
- `filestore_text_index.go` — `TextIndexer`: листает папки (`FileStore.List`),
качает файлы (`FileStore.Download`), извлекает текст через
`internal/docpipe` (`pdftotext`, для сканов — `ocrmypdf`/`tesseract`),
пишет в `TextIndex`. Пул воркеров (по умолчанию 3).
- CLI: `oo index folder|files` наполняет индекс; `oo search --backend own`
ищет по нему.
### Встроенные вложения PDF
Оцифрованные PDF несут вложения (`<doc>.md` — текст/таблицы скана,
`<doc>.yaml`/`.json` — метаданные, `.xml` — EN 16931 CII eRechnung,
`factur-x.xml` у ZUGFeRD; см. `office-assistant/docs/reference/document-metadata.md`).
`TextIndexer` обходит их: `pdfdetach -list` перечисляет, `-save` сохраняет,
каждое вложение проходит штатный `docpipe.ToMarkdown` (PDF/картинки → OCR,
`.md`/`.txt` — как есть). Форматы, которые docpipe не конвертирует
(`.xml`/`.html` — снимаются теги; `.json`/`.csv` — как текст), извлекаются
текстом; нечитаемые — пропускаются.
Текст склеивается: тело, затем по секции на вложение с маркером
`[attachment: <имя>]` (функция `docpipe.JoinWithAttachments`). Индекс — тот же
`file_id`, upsert идемпотентен. Нет вложений или pdfdetach/формат нечитаем —
индексируется тело (без падения).
Поля `oo_docs_text`:
| поле | тип | смысл |
|------|-----|-------|
| `id` | keyword | id файла Documents |
| `title` | text (+`.keyword`) | имя файла |
| `folder` | keyword | id папки |
| `ext` | keyword | расширение |
| `content` | text | извлечённый текст (pdftotext/OCR) |
## CLI
```bash
set -a; . .env; set +a # ONLYOFFICE_URL/USER/PASS + ONLYOFFICE_ES_URL
oo index folder 634 # PDF в папке 634
oo index folder 634 --recursive --exts pdf,png --limit 100
oo index files 3576 3578 # точечно
oo index folder 634 --dry-run # показать план, ничего не менять
oo search "S1021" --content --backend own
oo search "S1021" --backend own --folder 634 --json
```
`--backend` у `oo search`: `oo` (по умолчанию, индекс OnlyOffice) или `own`
(наш `ONLYOFFICE_ES_TEXT_INDEX`).
## Переменные (дополнение)
| env | default | смысл |
|-----|---------|-------|
| `ONLYOFFICE_ES_TEXT_INDEX` | `oo_docs_text` | индекс своего конвейера |
`ONLYOFFICE_ES_URL` — общий для обоих индексов.
## Тесты
```bash
go test -run 'ESText|TextIndexer|Index' ./ ./cmd/oo/ # unit, без сети
ONLYOFFICE_ES_URL=http://127.0.0.1:9200 \
go test -tags=integration -run TestIntegrationESTextIndex -v .
```
Интеграционный тест создаёт временный индекс, наполняет, ищет по контенту,
проверяет фильтры и удаление, затем удаляет индекс;
`TestIntegrationESTextIndexPDFAttachment` индексирует
`testdata/pdf-with-attachment.pdf` реальным конвейером (pdfdetach + pdftotext)
и ищет токен, лежащий только во вложении. Unit-тесты используют
fake-store/fake-extractor и не требуют pdftotext/OCR (парсер списка, склейка
`JoinWithAttachments`, снятие тегов `xmlToText` — чистые).
## Грабли
- Наполнение — ручное (`oo index`); после изменения/добавления PDF повтори.
Повтор идемпотентен (upsert по id файла).
- В индексе ищется только то, что проиндексировано; `oo index` качает каждый
файл и (для сканов) гоняет OCR — это медленно, отсюда `--limit`/`--exts`.
- `folder` фильтруется как id папки, а не как путь.
- Дубликаты (напр. `S1055.pdf` и `2026-08-20-S1055-…`) дадут несколько строк —
это ожидаемо, дедуп — на стороне потребителя.
- Вложения: нужен `pdfdetach` (poppler); если его нет — индексируется только
тело. Вложенный PDF/картинка с плохим текстовым слоем проходит OCR, это
медленно. `.json`-метаданные (CuraSoft) индексируются как текст и могут
добавить шумовых токенов.
+96
View File
@@ -0,0 +1,96 @@
---
type: reference
status: current
related:
- docs/elasticsearch.md
- docs/unified-file-client.md
- docs/community-server-db.md
---
# Поиск и индексация
Четыре разных контура поиска. Не путать: у каждого свой индекс, свои входы и
свой способ обновления.
| Контур | Что ищет | Индекс | Обновление | Вход |
|--------|----------|--------|------------|------|
| REST `@search` | только имена в БД | нет | — (живой запрос) | `oo search` (по умолчанию `--backend oo`) |
| ES `files_file` | имя + текст Office | Elasticsearch портала | сервер, асинхронно | `oo search --content` |
| ES `oo_docs_text` | PDF/сканы (свой) | Elasticsearch портала | `oo index` | `oo search --backend own` |
## Карта кода
- `filestore_core.go` — интерфейсы `Searcher`, модели `SearchQuery`/`SearchHit`.
- `filestore_es.go` — `ESSearcher` (индекс OnlyOffice `files_file`).
- `filestore_es_text.go` — `ESTextIndex` (`oo_docs_text`): `Ensure`, `Put`, `Delete`,
`Search`.
- `filestore_text_index.go` — `TextIndexer`: обход папок (`FileStore.List`), download
(`FileStore.Download`), извлечение текста (`internal/docpipe`), запись в
`TextIndex`; пул воркеров.
- `filestore_facade.go` — связка бэкендов (`Files().Search()`, порядок и fallback).
- CLI: `cmd/oo/search.go`, `cmd/oo/index.go`.
- Разовые бинари для match (`ooscan`, `pdfamount`) живут в приватном
`oo-workspace`.
- `internal/docpipe` — текст из PDF (pdftotext), для сканов OCR
(ocrmypdf/tesseract), вложения PDF (pdfdetach).
## Поиск
```bash
# имя, индекс портала
oo search "Rechnung" --limit 50 --json
# имя + текст Office (docx/xlsx/pptx)
oo search "Mahngebühr" --content
# свой индекс: PDF и сканы
oo search "S1019" --backend own --folder 649
```
Флаги `oo search`: `--content`, `--folder ID`, `--limit N`, `--backend oo|own`,
`--substring`, `--json`. Требует `ONLYOFFICE_ES_URL` (см.
[elasticsearch.md](elasticsearch.md)); `--backend own` дополнительно ничего не
требует от сервера — читает `oo_docs_text`.
Почему не REST: `GET /api/2.0/files/@search/{query}` ищет только имя в БД
(`fileDao.Search`), ES не трогает. Почему PDF не в `files_file`: сервер индексит
контент только для форматов из `files.index.formats` (по умолчанию
`.pptx|.xlsx|.docx`) — отсюда свой `oo_docs_text`.
## Обновление своего индекса (`oo index`)
```bash
# одна папка
oo index folder 649 --recursive --exts pdf
# точечно по id
oo index files 3576 3578
# без записи: что было бы проиндексировано
oo index folder 649 --recursive --dry-run
```
Флаги: `--recursive`, `--exts pdf` (по умолчанию), `--limit N`,
`--workers 3`, `--lang deu+eng`, `--min-chars N` (порог текстового слоя, ниже
которого включается OCR), `--work-dir`, `--backend rest|dav`, `--dry-run`,
`--json`.
Свойства:
- Идемпотентно: upsert по `id` файла; повтор не двоит.
- Сервер OnlyOffice не меняется: индекс живёт рядом (`ONLYOFFICE_ES_TEXT_INDEX`,
по умолчанию `oo_docs_text`).
- Медленно на сканах (OCR на каждый файл). Ограничивай `--folder`/`--limit`,
не индексируй корень целиком.
- Индексация PDF в `files_file` не делается — только `oo_docs_text`.
## Bulk-инструменты
Плоские TSV-инструменты (`ooscan`, `pdfamount`) и сверка Excel живут в
приватном `oo-workspace`, не в публичной библиотеке.
## Грабли
- ES слушает только `127.0.0.1:9200` внутри VM — SSH-туннель обязателен
(`-o ControlMaster=no -o ControlPath=none`, см. [elasticsearch.md](elasticsearch.md)).
- Портальные листинги/скачивание упираются в 429; все bulk-пути идут через
`DoRetry` (линейный бэкофф), `ooscan` дополнительно спит 350 мс на папку.
- `files_file` обновляется сервером асинхронно — свежий файл виден не сразу.
- `title` analyzer `whitespacecustom`: имя — один токен, подстрока только через
`--substring` (или wildcard).
+55
View File
@@ -0,0 +1,55 @@
---
type: reference
status: current
related:
- README.md
- ../AGENTS.md
---
# Rate limit, backoff и cooldown
Устойчивость к 429 (openresty). Всё встроено в библиотеку — отдельный пакет не
нужен. Реализация: `ratelimit.go`, `retry.go`.
## Что происходит с каждым запросом
1. **Cooldown-гейт** — общий на процесс. Если недавно пришёл 429, все запросы
ждут конца окна.
2. **Rate limiter** — token bucket на процесс. Пейсит все HTTP-пути: листинг,
создание папок, загрузку, `get project`, auth.
3. Запрос уходит.
4. Ответ ≥400 → `*TransientError` (для 429/502/503/504) с `Retry-After`.
5. `DoRetry` — экспоненциальный backoff, без jitter.
6. `Retry-After` длиннее backoff → ждём его; окно уходит в общий cooldown.
Установлено в `NewClient` через `pacedTransport`; отдельный код трогать не надо.
## Env
| Переменная | Default | Смысл |
|---|---|---|
| `OO_RATE_LIMIT` | `4` | запросов/с на процесс; `0` — лимитер выключен |
| `OO_BURST` | `1` | запас токенов token bucket |
| `OO_RETRY_ATTEMPTS` | `7` | всего попыток, включая первую |
| `OO_RETRY_BASE` | `2s` | база экспоненты: ждать перед попыткой N = `Base*2^(N-1)` |
| `OO_RETRY_MAX` | `2m` | потолок ожидания |
Битые значения → default. `OO_RETRY_*` — формат `time.ParseDuration`
(`2s`, `30s`, `2m`).
## Правила
- Детерминированно, без jitter — повторный прогон ждёт столько же.
- `Retry-After` — секунды (`120`) или HTTP-date.
- Cooldown общий: параллельные и последовательные вызовы не бьют в стену.
- Backoff cap не ограничивает `Retry-After` — серверу верим больше.
- Только stdlib.
## Когда руками снять нагрузку
`OO_RATE_LIMIT` ниже (`2`), `OO_BURST=1`; при массовом apply — батчами.
## Тесты
`retry_test.go` — `Retry-After`, экспонента, cap; `ratelimit_test.go` — burst,
`OO_RATE_LIMIT=0`, cooldown. Фейковый сервер отдаёт 429 с заголовком.
+204
View File
@@ -0,0 +1,204 @@
---
type: reference
status: current
related:
- README.md
- filestore_core.go
- filestore_facade.go
- docs/elasticsearch.md
- docs/community-server-db.md
---
# Unified file client — контракт файловых бэкендов
## Что это
Один файловый клиент на все бэкенды (эпик #34). Модель и интерфейсы —
`filestore_core.go`. Фасад `FileClient` — `filestore_facade.go`. Бэкенды:
REST, WebDAV, SQL (PostgreSQL/MySQL), Elasticsearch. Правило одно:
код зовёт `c.Files()` и не знает про транспорт.
## Модель
- `Kind` — `File` (0) или `Folder` (1).
- `Entry` — бэкенд-независимая строка: `ID`, `ParentID`, `Title`, `Kind`,
`Size`, `MIME`, `Created`, `Modified`, `Updated` (сырая строка API),
`Version`, `Provider`, `FilesCount`/`FoldersCount` (папки).
Чего бэкенд не даёт — остаётся в нуле.
- `SearchQuery` — `Text`, `InContent`, `FolderID`, `Extensions`, `Limit`.
- `SearchHit` — `Entry` + `Score`, `Highlight`, `Path`.
## Интерфейсы
`FileStore` — операции с файлами:
```go
type FileStore interface {
Name() string
List(ctx, parentID) ([]Entry, error)
Stat(ctx, id) (Entry, error)
CreateFolder(ctx, parentID, title) (Entry, error)
Upload(ctx, parentID, title, r) (Entry, error)
Download(ctx, id, w) (int64, error)
Move(ctx, ids, parentID) error
Copy(ctx, ids, parentID) error
Rename(ctx, id, title) error
Delete(ctx, ids) error
}
```
`Searcher` — поиск (необязательный):
```go
type Searcher interface {
Search(ctx, q SearchQuery) ([]SearchHit, error)
Name() string
}
```
`TextIndex` (`filestore_es_text.go`) — свой индекс: `Put`, `Delete`, `Search`,
`Name`. `ESTextIndex` реализует и `Searcher`, и `TextIndex`.
## Бэкенды
| бэкенд | провайдер | файл | что умеет |
|--------|-----------|------|-----------|
| REST | `rest` | `filestore_rest.go` | read + write, Documents API |
| WebDAV | `dav` | `filestore_dav.go` | read + write, Documents fileops |
| SQL | `postgres` / `mysql` | `filestore_pg.go` | **read-only** |
| OnlyOffice ES | `elasticsearch` | `filestore_es.go` | поиск (имя + контент Office) |
| свой ES-индекс | `es-text` | `filestore_es_text.go` | поиск + запись (PDF/сканы) |
- REST: `Stat` знает только файлы; папки — через `List`.
- WebDAV: `Move`/`Copy`/`Delete` сперва `Stat`-ят id (папка/файл), потом зовут
fileops.
- SQL: `List`/`Stat`/`Download`/`Search` (по имени). Все write-методы →
`ErrReadOnly`. `Download` идёт в S3/MinIO по layout портала.
- OnlyOffice ES: индекс `files_file`, контент только для docx/xlsx/pptx.
- Свой ES: индекс `oo_docs_text`, контент из `internal/docpipe`, в т.ч.
встроенные PDF-вложения.
## Фасад `FileClient`
`c.Files()` → `*FileClient`. Он же реализует `FileStore`, старый код
компилируется.
- `Read()` — первый зарегистрированный из `readOrder`:
`postgres` → `mysql` → `rest` → `dav`.
- `Write()` — первый из `writeOrder`: `rest` → `dav`. SQL не пишет.
- `Search()` — первый из `searchOrder`: `elasticsearch`. Нет бэкенда →
ошибка (`ONLYOFFICE_ES_URL`).
- `RegisterStore(name, s)` / `RegisterSearcher(name, s)` — добавить бэкенд.
Fallback:
- `List`/`Stat` идут по `readOrder`; переходят к следующему только на
transient-ошибке (429/502/503/504). Иначе ошибка финальная.
- `Download` **без** fallback: часть байтов уже в `w`, второй бэкенд допишет.
- Запись (`CreateFolder`/`Upload`/`Move`/`Copy`/`Rename`/`Delete`) — только
`Write()`, без fallback.
`newFileClient` сам кладёт `rest` и `dav`; ES-поиск — если задан
`ONLYOFFICE_ES_URL`. SQL-стор регистрирует вызывающий: фасад создаётся на
каждый `c.Files()`, регистрируй на том же экземпляре.
```go
sql, err := c.SQLFileStore() // открыть из env (ONLYOFFICE_DSN)
if err != nil { /* нет DSN */ }
if closer, ok := sql.(interface{ Close() error }); ok { defer closer.Close() }
f := c.Files()
f.RegisterStore(onlyoffice.ProviderPG, sql) // или sql.Name() == "mysql"
e, _ := f.Stat(ctx, "19423") // e.Provider == "mysql"
entries, _ := f.List(ctx, "676") // пойдёт в SQL
```
`Client.FileStore("pg"|"sql"|"postgres"|"mysql")` — одноразовый доступ к
SQL-стору без фасада: открывает из env; при ошибке возвращает заглушку,
которая отдаёт ошибку открытия на каждом вызове (не `nil`). `SQLFileStore()`
— тот же открыватель, но с ошибкой. Отвечавший бэкенд видно по
`Entry.Provider` (`mysql` / `postgres` у SQL, `rest` у REST).
## CLI
```bash
# поиск: --backend oo (индекс OnlyOffice) | own (свой oo_docs_text)
oo search "Rechnung"
oo search "Mahngebühr" --content
oo search "S1021" --content --backend own --folder 634 --limit 50 --json
# наполнение своего индекса (PDF/сканы, idempotent upsert по file id)
oo index folder 634
oo index folder 634 --recursive --exts pdf,png --limit 100
oo index files 3576 3578
oo index folder 634 --dry-run
```
`oo index` флаги: `--recursive`, `--exts` (default `pdf`), `--limit`,
`--workers` (3), `--lang` (`deu+eng`), `--min-chars`, `--work-dir`,
`--backend rest|dav`, `--dry-run`, `--json`.
Библиотека:
```go
idx, _ := onlyoffice.NewESTextIndex(onlyoffice.ESTextConfigFromEnv())
ti := onlyoffice.NewTextIndexer(store, idx) // store = FileStore
res, _ := ti.IndexFolder(ctx, "634", onlyoffice.IndexOptions{Recursive: true})
```
## Env (только имена)
| env | default | кто читает |
|-----|---------|------------|
| `ONLYOFFICE_ES_URL` | — | ES (оба индекса), обязателен |
| `ONLYOFFICE_ES_INDEX` | `files_file` | индекс OnlyOffice |
| `ONLYOFFICE_ES_TEXT_INDEX` | `oo_docs_text` | свой индекс |
| `ONLYOFFICE_TENANT` | пусто | фильтр `tenantId` |
| `ONLYOFFICE_DSN` | — | SQL DSN (MySQL/PostgreSQL) |
| `ONLYOFFICE_PG_DRIVER` | auto | `postgres` / `mysql` |
| `ONLYOFFICE_PG_TENANT` | `ONLYOFFICE_TENANT` | SQL tenant |
| `ONLYOFFICE_PG_HOST` `_PORT` `_USER` `_PASSWORD` `_DBNAME` `_SSLMODE` | — | DSN по частям |
| `MINIO_ENDPOINT` `MINIO_BUCKET` `MINIO_ACCESS_KEY` `MINIO_SECRET_KEY` | — | download SQL-стора |
| `OO_URL` `OO_USER` `OO_PASS` | — | CLI-алиасы |
Имена — в [`.env.example`](../.env.example). Секретов в репо нет.
## Ограничения
- OnlyOffice ES: контент только Office-форматов. PDF — только по имени.
Встроенные вложения PDF сервер не индексирует.
- Свой индекс `oo_docs_text`: покрывает PDF/сканы и вложения (pdfdetach), но
наполняется вручную (`oo index`) и идемпотентен. Фильтр `folder` — id папки,
не путь. Дубли дают несколько строк — дедуп на потребителе.
- SQL: read-only. `InContent` игнорируется (только имя). Download — через
MinIO-схему, не HTTP.
- ES: без auth, слушает localhost внутри VM — нужен SSH-туннель
(см. [elasticsearch.md](elasticsearch.md)).
- `oo index` качает каждый файл и для сканов гоняет OCR — медленно; отсюда
`--limit` и `--exts`. Нужен `pdfdetach` (poppler); без него — только тело PDF.
## Как добавить бэкенд
1. Файл `file_<name>.go`. Реализуй `FileStore` (`Name` + 9 методов). Нужен
поиск — добавь `Searcher`; нужна запись своего индекса — `TextIndex`.
2. Добавь const провайдера рядом с `ProviderREST`/`ProviderDAV`.
3. Зарегистрируй: в `newFileClient` или снаружи через
`RegisterStore`/`RegisterSearcher`.
4. Внеси имя в `readOrder` / `writeOrder` / `searchOrder`.
5. Есть CLI-команда — добавь значение в `--backend`.
6. Тесты: unit (чистые builders/парсеры, без сети) + интеграционный
(`//go:build integration`, skip без кред).
## Тесты
```bash
go test ./... # unit, без сети
go test -tags=integration ./... # live (креды в .env)
go test ./ -run 'FileStore|Facade|ESText|PG'
```
## См. также
- [README.md](README.md) — индекс справочников.
- [elasticsearch.md](elasticsearch.md) — индекс OnlyOffice и свой `oo_docs_text`.
- [community-server-db.md](community-server-db.md) — SQL-стор и схема БД.
+15 -1
View File
@@ -247,6 +247,8 @@ func (c *Client) UploadProjectFileReplacing(ctx context.Context, projectID, loca
} }
// GetFile returns file metadata including viewUrl for download. // GetFile returns file metadata including viewUrl for download.
//
// Deprecated: use FileStore.Stat via Client.Files()/Client.FileStore.
func (c *Client) GetFile(ctx context.Context, fileID string) (*FileEntry, error) { func (c *Client) GetFile(ctx context.Context, fileID string) (*FileEntry, error) {
if fileID == "" { if fileID == "" {
return nil, fmt.Errorf("file id is required") return nil, fmt.Errorf("file id is required")
@@ -260,6 +262,8 @@ func (c *Client) GetFile(ctx context.Context, fileID string) (*FileEntry, error)
} }
// RenameFile sets a new title (including extension) for the file. // RenameFile sets a new title (including extension) for the file.
//
// Deprecated: use FileStore.Rename via Client.Files()/Client.FileStore.
func (c *Client) RenameFile(ctx context.Context, fileID, newTitle string) (*FileEntry, error) { func (c *Client) RenameFile(ctx context.Context, fileID, newTitle string) (*FileEntry, error) {
if fileID == "" || newTitle == "" { if fileID == "" || newTitle == "" {
return nil, fmt.Errorf("file id and new title are required") return nil, fmt.Errorf("file id and new title are required")
@@ -274,7 +278,9 @@ func (c *Client) RenameFile(ctx context.Context, fileID, newTitle string) (*File
// DeleteFiles permanently deletes files by numeric id (Documents module). // DeleteFiles permanently deletes files by numeric id (Documents module).
// Uses per-file DELETE (DeleteDavItems); fileops/delete returns 200 on some // Uses per-file DELETE (DeleteDavItems); fileops/delete returns 200 on some
// portals (e.g. produktor.io) without removing the file. // portals without actually removing the file.
//
// Deprecated: use FileStore.Delete via Client.Files()/Client.FileStore.
func (c *Client) DeleteFiles(ctx context.Context, fileIDs []int) error { func (c *Client) DeleteFiles(ctx context.Context, fileIDs []int) error {
if len(fileIDs) == 0 { if len(fileIDs) == 0 {
return fmt.Errorf("no file ids to delete") return fmt.Errorf("no file ids to delete")
@@ -288,6 +294,8 @@ func (c *Client) DeleteFiles(ctx context.Context, fileIDs []int) error {
// ListFolder returns the Documents module listing for a folder id // ListFolder returns the Documents module listing for a folder id
// (GET /api/2.0/files/{folderId}). // (GET /api/2.0/files/{folderId}).
//
// Deprecated: use FileStore.List via Client.Files()/Client.FileStore.
func (c *Client) ListFolder(ctx context.Context, folderID string) (map[string]any, error) { func (c *Client) ListFolder(ctx context.Context, folderID string) (map[string]any, error) {
if folderID == "" { if folderID == "" {
return nil, fmt.Errorf("folder id is required") return nil, fmt.Errorf("folder id is required")
@@ -313,6 +321,8 @@ func (c *Client) CreateFolder(ctx context.Context, parentFolderID, title string)
} }
// MoveFiles moves file ids into destFolderID (Documents fileops/move). // MoveFiles moves file ids into destFolderID (Documents fileops/move).
//
// Deprecated: use FileStore.Move via Client.Files()/Client.FileStore.
func (c *Client) MoveFiles(ctx context.Context, destFolderID int, fileIDs []int) (map[string]any, error) { func (c *Client) MoveFiles(ctx context.Context, destFolderID int, fileIDs []int) (map[string]any, error) {
if destFolderID == 0 || len(fileIDs) == 0 { if destFolderID == 0 || len(fileIDs) == 0 {
return nil, fmt.Errorf("dest folder and file ids are required") return nil, fmt.Errorf("dest folder and file ids are required")
@@ -343,6 +353,8 @@ func (c *Client) MoveFiles(ctx context.Context, destFolderID int, fileIDs []int)
} }
// UploadToFolder uploads a local file into an arbitrary Documents folder id. // UploadToFolder uploads a local file into an arbitrary Documents folder id.
//
// Deprecated: use FileStore.Upload via Client.Files()/Client.FileStore.
func (c *Client) UploadToFolder(ctx context.Context, folderID, localPath string) (*FileEntry, error) { func (c *Client) UploadToFolder(ctx context.Context, folderID, localPath string) (*FileEntry, error) {
if folderID == "" || localPath == "" { if folderID == "" || localPath == "" {
return nil, fmt.Errorf("folder id and local path are required") return nil, fmt.Errorf("folder id and local path are required")
@@ -400,6 +412,8 @@ func FileFolderID(f *FileEntry) string {
// as API calls. Writes into dst. When the portal serves the file from its stale // as API calls. Writes into dst. When the portal serves the file from its stale
// AWS S3 consumer, the bytes are fetched from the local MinIO store instead // AWS S3 consumer, the bytes are fetched from the local MinIO store instead
// (see storage_fallback.go). // (see storage_fallback.go).
//
// Deprecated: use FileStore.Download via Client.Files()/Client.FileStore.
func (c *Client) DownloadFile(ctx context.Context, fileID string, dst io.Writer) (int64, error) { func (c *Client) DownloadFile(ctx context.Context, fileID string, dst io.Writer) (int64, error) {
f, err := c.GetFile(ctx, fileID) f, err := c.GetFile(ctx, fileID)
if err != nil { if err != nil {
+6
View File
@@ -154,6 +154,9 @@ func findWithinFolderDuplicates(indexed []ProjectFolderFile) []DedupGroup {
byKey[k] = append(byKey[k], it.File) byKey[k] = append(byKey[k], it.File)
} }
for k, group := range byKey { for k, group := range byKey {
if k == "" {
continue // dotfiles etc. have no stem: never treat as duplicates
}
if len(group) < 2 { if len(group) < 2 {
continue continue
} }
@@ -178,6 +181,9 @@ func findCrossFolderDuplicates(indexed []ProjectFolderFile) []DedupGroup {
} }
var out []DedupGroup var out []DedupGroup
for k, items := range byKey { for k, items := range byKey {
if k == "" {
continue // dotfiles etc. have no stem: never treat as duplicates
}
if len(items) < 2 { if len(items) < 2 {
continue continue
} }
+416
View File
@@ -0,0 +1,416 @@
//go:build integration
package onlyoffice
import (
"context"
"encoding/json"
"os"
"path/filepath"
"strconv"
"testing"
"time"
)
// TestIntegrationFileDedup proves on a live OnlyOffice portal that the file
// dedup helpers find real duplicates and delete only the redundant copies.
// It creates a throwaway "go-onlyoffice-test-" project (removed by cleanup),
// places same stem|ext files in two subfolders and in the project root, then
// exercises FindProjectDuplicates, mergeProjectRootForDedupe,
// ApplyDedupGroups and DeleteFilesByDedupKey. Destructive — run only against
// an instance you own.
func TestIntegrationFileDedup(t *testing.T) {
c := liveClient(t)
t.Cleanup(func() { cleanupTestProjects(t, c) })
ctx := context.Background()
suffix := time.Now().UTC().Format("20060102-150405")
project, err := c.CreateProject(NewProjectRequest{
Title: testProjectPrefix + "dedup-" + suffix,
Description: "go-onlyoffice file dedup integration",
})
if err != nil {
t.Fatalf("CreateProject: %v", err)
}
if project.ID == nil {
t.Fatal("created project without id")
}
pid := strconv.Itoa(*project.ID)
root := projectFolderEventually(t, ctx, c, pid)
aID := createFolderLive(t, ctx, c, root, "A-"+suffix)
bID := createFolderLive(t, ctx, c, root, "B-"+suffix)
createFolderLive(t, ctx, c, root, "_trash-"+suffix)
// Check IsTrashFolderTitle against a real live folder title.
if !IsTrashFolderTitle("_trash-" + suffix) {
t.Fatalf("IsTrashFolderTitle(%q) = false for a live _trash folder", "_trash-"+suffix)
}
stem := "dedup-" + suffix
local := writeLocalFile(t, stem+".txt", []byte("dedup integration "+suffix+"\n"))
a1 := uploadFolderLive(t, ctx, c, aID, local)
a2 := uploadFolderLive(t, ctx, c, aID, local)
b1 := uploadFolderLive(t, ctx, c, bID, local)
r1 := uploadFolderLive(t, ctx, c, root, local)
r2 := uploadFolderLive(t, ctx, c, root, local)
t.Logf("uploaded a1=%d a2=%d b1=%d r1=%d r2=%d",
FileEntryNumericID(a1), FileEntryNumericID(a2), FileEntryNumericID(b1),
FileEntryNumericID(r1), FileEntryNumericID(r2))
if n := dedupWaitCount(t, ctx, c, aID, stem, ".txt", 2, 30*time.Second); n != 2 {
t.Fatalf("folder A has %d copies after upload, want 2", n)
}
if n := dedupWaitCount(t, ctx, c, bID, stem, ".txt", 1, 30*time.Second); n != 1 {
t.Fatalf("folder B has %d copies after upload, want 1", n)
}
if n := dedupWaitCount(t, ctx, c, root, stem, ".txt", 2, 30*time.Second); n != 2 {
t.Fatalf("project root has %d copies after upload, want 2", n)
}
// #2 mergeProjectRootForDedupe on the live tree: the project root that
// carries documents must be part of the scan exactly once.
folders, byFolder, rootFiles := liveProjectIndex(t, ctx, c, pid, root)
merged, mergedBy := mergeProjectRootForDedupe(root, folders, byFolder, rootFiles)
rootCount := 0
for _, folder := range merged {
if folder != nil && folder.ID != nil && folder.ID.String() == root {
rootCount++
}
}
if rootCount != 1 {
t.Fatalf("mergeProjectRootForDedupe: project root appears %d times in live tree, want 1", rootCount)
}
if got := len(FindFilesByDedupKey(mergedBy[root], stem, ".txt")); got != 2 {
t.Fatalf("mergeProjectRootForDedupe: root carries %d matching files, want 2", got)
}
// Within-folder scan: a duplicate pair in A and in the project root.
within := FindProjectDuplicates(merged, mergedBy, DedupOptions{})
removesByFolder := map[string]int{}
for _, g := range within {
removesByFolder[g.FolderID] = len(g.Remove)
}
if len(within) != 2 || removesByFolder[aID] != 1 || removesByFolder[root] != 1 {
t.Fatalf("within-folder groups = %d (%v), want exactly A:1 root:1", len(within), removesByFolder)
}
// DeleteFilesByDedupKey removes every stem|ext copy in one folder.
removed, err := c.DeleteFilesByDedupKey(ctx, aID, stem, ".txt")
if err != nil {
t.Fatalf("DeleteFilesByDedupKey(A): %v", err)
}
if len(removed) != 2 {
t.Fatalf("DeleteFilesByDedupKey(A) removed %v, want 2 ids", removed)
}
if n := dedupWaitCount(t, ctx, c, aID, stem, ".txt", 0, 30*time.Second); n != 0 {
t.Fatalf("folder A still has %d copies after DeleteFilesByDedupKey", n)
}
// Re-create the A duplicates so the cross-folder project scan can be
// applied and a single survivor proven by polling.
uploadFolderLive(t, ctx, c, aID, local)
uploadFolderLive(t, ctx, c, aID, local)
if n := dedupWaitCount(t, ctx, c, aID, stem, ".txt", 2, 30*time.Second); n != 2 {
t.Fatalf("folder A has %d re-uploaded copies, want 2", n)
}
// Cross-folder dry-run over the live project: one key, five copies, and
// the project-root copies must be part of the group (root merge live).
groups, deleted, err := dedupeProjectEventually(t, ctx, c, pid, DedupOptions{CrossFolder: true}, false)
if err != nil {
t.Fatalf("DedupeProject: %v", err)
}
if len(deleted) != 0 {
t.Fatalf("dry-run DedupeProject deleted %v", deleted)
}
if len(groups) != 1 {
t.Fatalf("DedupeProject cross-folder groups = %d, want 1 (%+v)", len(groups), groups)
}
if len(groups[0].Remove) != 4 {
t.Fatalf("cross-folder group removes %d files, want 4", len(groups[0].Remove))
}
if !dedupGroupHasID(groups[0], FileEntryNumericID(r1)) && !dedupGroupHasID(groups[0], FileEntryNumericID(r2)) {
t.Fatalf("cross-folder group does not include a project-root copy (merge not applied)")
}
keepID := FileEntryNumericID(groups[0].Keep)
deleted, err = c.ApplyDedupGroups(ctx, groups)
if err != nil {
t.Fatalf("ApplyDedupGroups: %v", err)
}
if len(deleted) != 4 {
t.Fatalf("ApplyDedupGroups deleted %v, want 4 ids", deleted)
}
survivorID, total := dedupWaitTotal(t, ctx, c, []string{aID, bID, root}, stem, ".txt", 1, 40*time.Second)
if total != 1 {
t.Fatalf("after ApplyDedupGroups %d copies survive, want 1", total)
}
if survivorID != keepID {
t.Fatalf("remaining copy id = %d, want kept id %d", survivorID, keepID)
}
// The removed ids must really be gone from every folder.
gone := map[int64]bool{
FileEntryNumericID(a2): true,
FileEntryNumericID(b1): true,
FileEntryNumericID(r1): true,
FileEntryNumericID(r2): true,
}
for _, fid := range []string{aID, bID, root} {
files, err := c.FolderFiles(ctx, fid)
if err != nil {
t.Fatalf("FolderFiles %s: %v", fid, err)
}
for _, f := range FindFilesByDedupKey(files, stem, ".txt") {
id := FileEntryNumericID(f)
if gone[id] {
t.Fatalf("removed id %d still present in folder %s", id, fid)
}
}
}
t.Logf("survivor id=%d keep id=%d, deleted=%v", survivorID, keepID, deleted)
}
// createFolderLive creates a subfolder (retrying transient 5xx) and returns
// its Documents folder id.
func createFolderLive(t *testing.T, ctx context.Context, c *Client, parentID, title string) string {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var (
m map[string]any
err error
)
for {
m, err = c.CreateFolder(ctx, parentID, title)
if err == nil || time.Now().After(deadline) {
break
}
time.Sleep(time.Second)
}
if err != nil {
t.Fatalf("CreateFolder %q: %v", title, err)
}
if m == nil {
t.Fatalf("CreateFolder %q: empty response", title)
}
switch v := m["id"].(type) {
case float64:
return strconv.FormatInt(int64(v), 10)
case json.Number:
return v.String()
case string:
if v != "" {
return v
}
}
t.Fatalf("CreateFolder %q: no id in response %#v", title, m)
return ""
}
// writeLocalFile writes content to a temp file and returns its path.
func writeLocalFile(t *testing.T, name string, content []byte) string {
t.Helper()
p := filepath.Join(t.TempDir(), name)
if err := os.WriteFile(p, content, 0o600); err != nil {
t.Fatal(err)
}
return p
}
// uploadFolderLive uploads localPath into folderID (retrying the portal's
// transient post-create 500) and returns the file entry.
func uploadFolderLive(t *testing.T, ctx context.Context, c *Client, folderID, localPath string) *FileEntry {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var (
e *FileEntry
err error
)
for {
e, err = c.UploadToFolder(ctx, folderID, localPath)
if err == nil || time.Now().After(deadline) {
break
}
time.Sleep(time.Second)
}
if err != nil {
t.Fatalf("UploadToFolder %s: %v", folderID, err)
}
if e == nil || e.ID == nil {
t.Fatalf("UploadToFolder %s: no file entry (%+v)", folderID, e)
}
return e
}
// liveProjectIndex rebuilds the project folder/file index the same way
// DedupeProject does, for direct mergeProjectRootForDedupe assertions.
func liveProjectIndex(t *testing.T, ctx context.Context, c *Client, projectID, rootID string) ([]*FolderEntry, map[string][]*FileEntry, []*FileEntry) {
t.Helper()
pf := getProjectFilesEventually(t, ctx, c, projectID)
rootFiles := folderFilesEventually(t, ctx, c, rootID)
folders := make([]*FolderEntry, 0, len(pf.Folders)+1)
byFolder := make(map[string][]*FileEntry, len(pf.Folders)+1)
for _, folder := range pf.Folders {
if folder == nil || folder.ID == nil {
continue
}
fid := folder.ID.String()
if fid == rootID {
byFolder[fid] = rootFiles
} else {
byFolder[fid] = folderFilesEventually(t, ctx, c, fid)
}
folders = append(folders, folder)
}
return folders, byFolder, rootFiles
}
// projectFolderEventually resolves the project Documents root id, retrying on
// a transient portal answer.
func projectFolderEventually(t *testing.T, ctx context.Context, c *Client, projectID string) string {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var (
root string
err error
)
for {
root, err = c.projectFolderID(ctx, projectID)
if err == nil || time.Now().After(deadline) {
break
}
time.Sleep(time.Second)
}
if err != nil {
t.Fatalf("projectFolderID: %v", err)
}
if root == "" {
t.Fatal("projectFolderID returned empty id")
}
return root
}
// getProjectFilesEventually lists a project's files/folders, retrying on a
// transient portal answer.
func getProjectFilesEventually(t *testing.T, ctx context.Context, c *Client, projectID string) *ProjectFilesResponse {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var last error
for {
pf, err := c.GetProjectFiles(ctx, projectID)
if err == nil {
return pf
}
last = err
if time.Now().After(deadline) {
t.Fatalf("GetProjectFiles %s: %v", projectID, last)
}
time.Sleep(500 * time.Millisecond)
}
}
// folderFilesEventually lists a folder, retrying while the portal answers
// transiently (a freshly created folder can 500 until its parent map settles).
func folderFilesEventually(t *testing.T, ctx context.Context, c *Client, folderID string) []*FileEntry {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var last error
for {
files, err := c.FolderFiles(ctx, folderID)
if err == nil {
return files
}
last = err
if time.Now().After(deadline) {
t.Fatalf("FolderFiles %s: %v", folderID, last)
}
time.Sleep(500 * time.Millisecond)
}
}
// dedupeProjectEventually runs a project dedup scan, retrying the whole scan on
// a transient portal error. It is used for dry-runs only (apply must stay a
// single deliberate call).
func dedupeProjectEventually(t *testing.T, ctx context.Context, c *Client, projectID string, opts DedupOptions, apply bool) ([]DedupGroup, []int, error) {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var (
groups []DedupGroup
deleted []int
err error
)
for {
groups, deleted, err = c.DedupeProject(ctx, projectID, opts, apply)
if err == nil || time.Now().After(deadline) {
return groups, deleted, err
}
time.Sleep(time.Second)
}
}
// dedupWaitCount polls folderID until want stem|ext copies are visible.
func dedupWaitCount(t *testing.T, ctx context.Context, c *Client, folderID, stem, ext string, want int, d time.Duration) int {
t.Helper()
deadline := time.Now().Add(d)
got := -1
for {
if files, err := c.FolderFiles(ctx, folderID); err == nil {
got = len(FindFilesByDedupKey(files, stem, ext))
if got == want {
return got
}
}
if time.Now().After(deadline) {
return got
}
time.Sleep(500 * time.Millisecond)
}
}
// dedupWaitTotal polls the given folders until the total number of stem|ext
// copies reaches want, returning the last seen file id and count.
func dedupWaitTotal(t *testing.T, ctx context.Context, c *Client, folderIDs []string, stem, ext string, want int, d time.Duration) (int64, int) {
t.Helper()
deadline := time.Now().Add(d)
var survivor int64
total := -1
for {
survivor, total = 0, 0
for _, fid := range folderIDs {
files, err := c.FolderFiles(ctx, fid)
if err != nil {
continue
}
for _, f := range FindFilesByDedupKey(files, stem, ext) {
total++
survivor = FileEntryNumericID(f)
}
}
if total == want {
return survivor, total
}
if time.Now().After(deadline) {
return survivor, total
}
time.Sleep(500 * time.Millisecond)
}
}
func dedupGroupHasID(g DedupGroup, id int64) bool {
if id == 0 {
return false
}
if FileEntryNumericID(g.Keep) == id {
return true
}
for _, f := range g.Remove {
if FileEntryNumericID(f) == id {
return true
}
}
return false
}
+26
View File
@@ -84,6 +84,32 @@ func TestMergeProjectRootForDedupe(t *testing.T) {
} }
} }
func TestFindDuplicatesSkipsEmptyKey(t *testing.T) {
// Dotfiles (".env", ".gitignore", ".npmrc") normalize to an empty stem, so
// FileDedupKey is "". They are not duplicates of each other and must never
// form a dedup group that would delete one of them.
env := &FileEntry{ID: jsonNum("1"), Title: strPtr(".env")}
gitignore := &FileEntry{ID: jsonNum("2"), Title: strPtr(".gitignore")}
npmrc := &FileEntry{ID: jsonNum("3"), Title: strPtr(".npmrc")}
if FileDedupKey(env) != "" || FileDedupKey(gitignore) != "" || FileDedupKey(npmrc) != "" {
t.Fatalf("dotfiles should have empty dedup key")
}
within := []ProjectFolderFile{
{FolderID: "500", FolderTitle: "Cfg", File: env},
{FolderID: "500", FolderTitle: "Cfg", File: gitignore},
}
if groups := findWithinFolderDuplicates(within); len(groups) != 0 {
t.Fatalf("within-folder empty-key groups = %d, want 0 (%+v)", len(groups), groups)
}
cross := []ProjectFolderFile{
{FolderID: "500", FolderTitle: "Cfg", File: env},
{FolderID: "501", FolderTitle: "Other", File: npmrc},
}
if groups := findCrossFolderDuplicates(cross); len(groups) != 0 {
t.Fatalf("cross-folder empty-key groups = %d, want 0 (%+v)", len(groups), groups)
}
}
func TestIsTrashFolderTitle(t *testing.T) { func TestIsTrashFolderTitle(t *testing.T) {
if !IsTrashFolderTitle("_trash-md") { if !IsTrashFolderTitle("_trash-md") {
t.Fatal("expected trash") t.Fatal("expected trash")
+79
View File
@@ -0,0 +1,79 @@
package onlyoffice
// Human-readable folder paths for search results (F9). The OnlyOffice ES
// index stores only ancestor folder ids; titles live in the Documents tree, so
// resolving a path costs one GET /api/2.0/files/{id} per distinct folder,
// cached on the client. Folders that cannot be listed (e.g. a section root)
// fall back to their id, so a path is always produced.
import (
"context"
"strings"
)
// FolderTitle returns the title of a Documents folder id, cached on the client.
// An empty id yields an empty title. Unknown/unlistable ids (section roots)
// return ("", nil) so callers can fall back to the id.
func (c *Client) FolderTitle(ctx context.Context, folderID string) (string, error) {
folderID = strings.TrimSpace(folderID)
if folderID == "" {
return "", nil
}
c.folderTitlesMu.Lock()
if c.folderTitles != nil {
if t, ok := c.folderTitles[folderID]; ok {
c.folderTitlesMu.Unlock()
return t, nil
}
}
c.folderTitlesMu.Unlock()
title := ""
out, err := c.ListFolder(ctx, folderID)
if err == nil {
if cur, ok := out["current"].(map[string]any); ok {
if s, ok := cur["title"].(string); ok {
title = strings.TrimSpace(s)
}
}
}
c.folderTitlesMu.Lock()
if c.folderTitles == nil {
c.folderTitles = map[string]string{}
}
c.folderTitles[folderID] = title
c.folderTitlesMu.Unlock()
return title, nil
}
// FolderPath resolves an ancestor folder id chain (root → leaf, as the ES
// backend reports it) into folder titles, falling back to the id when a title
// cannot be read. The result never fails on a single lookup: only the whole
// call honours ctx cancellation.
func (c *Client) FolderPath(ctx context.Context, ids []string) []string {
out := make([]string, 0, len(ids))
for _, id := range ids {
if err := ctx.Err(); err != nil {
break
}
title, err := c.FolderTitle(ctx, id)
if err != nil || title == "" {
title = id
}
out = append(out, title)
}
return out
}
// UniquePath builds a stable, human-readable, unique path for a result: the
// resolved folder chain plus the file title. "." separates nothing — the
// segments are joined with "/", matching the Documents breadcrumb the web UI
// shows.
func (c *Client) UniquePath(ctx context.Context, folderPath []string, title string) string {
parts := c.FolderPath(ctx, folderPath)
if t := strings.TrimSpace(title); t != "" {
parts = append(parts, t)
}
return strings.Join(parts, "/")
}
+35
View File
@@ -0,0 +1,35 @@
//go:build integration
package onlyoffice
import (
"context"
"strings"
"testing"
)
// TestIntegrationFolderPath resolves the real Fibu EDL folder chain
// (project root 522 → Eingangsrechnungen 647 → 2025 649) to titles.
func TestIntegrationFolderPath(t *testing.T) {
creds := GetEnvironmentCredentials()
if strings.TrimSpace(creds.Url) == "" || strings.TrimSpace(creds.User) == "" {
t.Skip("no ONLYOFFICE_URL/USER credentials")
}
c := NewClient(creds)
ctx := context.Background()
path := c.FolderPath(ctx, []string{"522", "647", "649"})
if len(path) != 3 {
t.Fatalf("FolderPath returned %v, want 3 segments", path)
}
for i, seg := range path {
if strings.TrimSpace(seg) == "" {
t.Errorf("segment %d empty: %v", i, path)
}
}
full := c.UniquePath(ctx, []string{"522", "647", "649"}, "Rechnung-x.pdf")
if !strings.HasSuffix(full, "Rechnung-x.pdf") || !strings.Contains(full, "/") {
t.Errorf("UniquePath = %q, want a slash-joined path ending in the file", full)
}
t.Logf("path=%v full=%q", path, full)
}
+31 -4
View File
@@ -52,6 +52,8 @@ type DavListing struct {
// ListDavFolder returns the contents of a folder by id, which may be a // ListDavFolder returns the contents of a folder by id, which may be a
// symbolic root such as "@my". For "@root" use ListDavSections. // symbolic root such as "@my". For "@root" use ListDavSections.
//
// Deprecated: use FileStore.List via Client.Files()/Client.FileStore.
func (c *Client) ListDavFolder(ctx context.Context, id string) (*DavListing, error) { func (c *Client) ListDavFolder(ctx context.Context, id string) (*DavListing, error) {
raw, err := c.getJSON(ctx, "/api/2.0/files/"+url.PathEscape(id)) raw, err := c.getJSON(ctx, "/api/2.0/files/"+url.PathEscape(id))
if err != nil { if err != nil {
@@ -111,6 +113,8 @@ func (c *Client) ListDavSections(ctx context.Context) ([]DavFolder, error) {
} }
// CreateDavFolder creates a folder titled title inside parentID. // CreateDavFolder creates a folder titled title inside parentID.
//
// Deprecated: use FileStore.CreateFolder via Client.Files()/Client.FileStore.
func (c *Client) CreateDavFolder(ctx context.Context, parentID, title string) (*DavFolder, error) { func (c *Client) CreateDavFolder(ctx context.Context, parentID, title string) (*DavFolder, error) {
raw, err := c.postJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(parentID), raw, err := c.postJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(parentID),
map[string]string{"title": title}) map[string]string{"title": title})
@@ -130,6 +134,8 @@ func (c *Client) CreateDavFolder(ctx context.Context, parentID, title string) (*
} }
// RenameDavFolder renames a folder. // RenameDavFolder renames a folder.
//
// Deprecated: use FileStore.Rename via Client.Files()/Client.FileStore.
func (c *Client) RenameDavFolder(ctx context.Context, id, title string) error { func (c *Client) RenameDavFolder(ctx context.Context, id, title string) error {
_, err := c.putJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(id), _, err := c.putJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(id),
map[string]string{"title": title}) map[string]string{"title": title})
@@ -137,6 +143,8 @@ func (c *Client) RenameDavFolder(ctx context.Context, id, title string) error {
} }
// RenameDavFile renames a file (title includes the extension). // RenameDavFile renames a file (title includes the extension).
//
// Deprecated: use FileStore.Rename via Client.Files()/Client.FileStore.
func (c *Client) RenameDavFile(ctx context.Context, id, title string) error { func (c *Client) RenameDavFile(ctx context.Context, id, title string) error {
_, err := c.putJSON(ctx, "/api/2.0/files/file/"+url.PathEscape(id), _, err := c.putJSON(ctx, "/api/2.0/files/file/"+url.PathEscape(id),
map[string]string{"title": title}) map[string]string{"title": title})
@@ -147,6 +155,8 @@ func (c *Client) RenameDavFile(ctx context.Context, id, title string) error {
// The fileops API answers 200 with per-operation error strings even when // The fileops API answers 200 with per-operation error strings even when
// nothing moves (e.g. missing permission), so the response is parsed and the // nothing moves (e.g. missing permission), so the response is parsed and the
// first operation error is returned instead of a silent nil. // first operation error is returned instead of a silent nil.
//
// Deprecated: use FileStore.Move via Client.Files()/Client.FileStore.
func (c *Client) MoveDavItems(ctx context.Context, folderIDs, fileIDs []string, destFolderID string) error { func (c *Client) MoveDavItems(ctx context.Context, folderIDs, fileIDs []string, destFolderID string) error {
raw, err := c.putJSON(ctx, "/api/2.0/files/fileops/move", map[string]any{ raw, err := c.putJSON(ctx, "/api/2.0/files/fileops/move", map[string]any{
"folderIds": nums(folderIDs), "folderIds": nums(folderIDs),
@@ -163,6 +173,8 @@ func (c *Client) MoveDavItems(ctx context.Context, folderIDs, fileIDs []string,
// CopyDavItems copies the given folders and/or files into destFolderID. // CopyDavItems copies the given folders and/or files into destFolderID.
// Per-operation errors are surfaced like in MoveDavItems. // Per-operation errors are surfaced like in MoveDavItems.
//
// Deprecated: use FileStore.Copy via Client.Files()/Client.FileStore.
func (c *Client) CopyDavItems(ctx context.Context, folderIDs, fileIDs []string, destFolderID string) error { func (c *Client) CopyDavItems(ctx context.Context, folderIDs, fileIDs []string, destFolderID string) error {
raw, err := c.putJSON(ctx, "/api/2.0/files/fileops/copy", map[string]any{ raw, err := c.putJSON(ctx, "/api/2.0/files/fileops/copy", map[string]any{
"folderIds": nums(folderIDs), "folderIds": nums(folderIDs),
@@ -226,22 +238,35 @@ func fileopsError(raw json.RawMessage) error {
} }
// DeleteDavItems deletes the given folders and/or files. // DeleteDavItems deletes the given folders and/or files.
//
// Deprecated: use FileStore.Delete via Client.Files()/Client.FileStore.
func (c *Client) DeleteDavItems(ctx context.Context, folderIDs, fileIDs []string) error { func (c *Client) DeleteDavItems(ctx context.Context, folderIDs, fileIDs []string) error {
body := map[string]any{"DeleteAfter": true, "Immediately": true} body := map[string]any{"DeleteAfter": true, "Immediately": true}
for _, id := range folderIDs { for _, id := range folderIDs {
if _, err := c.deleteJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(id), body); err != nil { if err := c.deleteDavItem(ctx, "/api/2.0/files/folder/"+url.PathEscape(id), body); err != nil {
return err return err
} }
} }
for _, id := range fileIDs { for _, id := range fileIDs {
if _, err := c.deleteJSON(ctx, "/api/2.0/files/file/"+url.PathEscape(id), body); err != nil { if err := c.deleteDavItem(ctx, "/api/2.0/files/file/"+url.PathEscape(id), body); err != nil {
return err return err
} }
} }
return nil return nil
} }
// deleteDavItem deletes one item, retrying transient 429/502/503/504 answers
// through DoRetry like every other bulk path (deletes are idempotent).
func (c *Client) deleteDavItem(ctx context.Context, path string, body any) error {
return DoRetry(ctx, DefaultRetryPolicy(), func() error {
_, err := c.deleteJSON(ctx, path, body)
return err
})
}
// UploadDavFile uploads src (fileName) into folderID, streaming from src. // UploadDavFile uploads src (fileName) into folderID, streaming from src.
//
// Deprecated: use FileStore.Upload via Client.Files()/Client.FileStore.
func (c *Client) UploadDavFile(ctx context.Context, folderID, fileName string, src io.Reader) (*DavFile, error) { func (c *Client) UploadDavFile(ctx context.Context, folderID, fileName string, src io.Reader) (*DavFile, error) {
raw, err := c.uploadReader(ctx, "/api/2.0/files/"+url.PathEscape(folderID)+"/upload", "file", fileName, src) raw, err := c.uploadReader(ctx, "/api/2.0/files/"+url.PathEscape(folderID)+"/upload", "file", fileName, src)
if err != nil { if err != nil {
@@ -261,6 +286,8 @@ func (c *Client) UploadDavFile(ctx context.Context, folderID, fileName string, s
// DownloadDavFile streams the file identified by id to w, returning bytes // DownloadDavFile streams the file identified by id to w, returning bytes
// copied. It shares the MinIO stale-S3 fallback with DownloadFile. // copied. It shares the MinIO stale-S3 fallback with DownloadFile.
//
// Deprecated: use FileStore.Download via Client.Files()/Client.FileStore.
func (c *Client) DownloadDavFile(ctx context.Context, id string, w io.Writer) (int64, error) { func (c *Client) DownloadDavFile(ctx context.Context, id string, w io.Writer) (int64, error) {
return c.DownloadFile(ctx, id, w) return c.DownloadFile(ctx, id, w)
} }
@@ -298,7 +325,7 @@ func (c *Client) deleteJSON(ctx context.Context, path string, body any) (json.Ra
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("DELETE %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "DELETE %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
@@ -338,7 +365,7 @@ func (c *Client) uploadReader(ctx context.Context, path, fieldName, fileName str
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("upload %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "upload %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
@@ -10,10 +10,11 @@ import (
"time" "time"
) )
// TestIntegrationFileStores runs the same operation set (create folder, upload, // TestIntegrationFileStores runs the same operation set through the REST and
// list, stat, download, move, copy, rename, delete) through the REST and DAV // DAV FileStore adapters against a throwaway project Documents folder: file
// FileStore adapters against a throwaway project Documents folder. Destructive // create/upload/list/stat/download/move/copy/rename/delete and folder
// — only run against instances you own. // create/stat/list/rename/move/delete. Destructive — only run against
// instances you own.
// //
// The Documents fileops API is asynchronous: a move/copy/delete is accepted // The Documents fileops API is asynchronous: a move/copy/delete is accepted
// immediately and becomes visible a moment later, so effects are polled. // immediately and becomes visible a moment later, so effects are polled.
@@ -102,10 +103,7 @@ func testFileStoreOps(t *testing.T, ctx context.Context, c *Client, store FileSt
t.Fatalf("moved file %s not in dst", up.ID) t.Fatalf("moved file %s not in dst", up.ID)
} }
if err := store.Copy(ctx, []string{up.ID}, src.ID); err != nil { copied := copyEventually(t, ctx, store, up.ID, src.ID, 20*time.Second)
t.Fatalf("Copy: %v", err)
}
copied := waitOtherFile(ctx, store, src.ID, up.ID, 20*time.Second)
if copied == nil { if copied == nil {
t.Fatalf("no copy found in src after Copy") t.Fatalf("no copy found in src after Copy")
} }
@@ -122,6 +120,67 @@ func testFileStoreOps(t *testing.T, ctx context.Context, c *Client, store FileSt
if !waitNoEntry(ctx, store, src.ID, copied.ID, 20*time.Second) { if !waitNoEntry(ctx, store, src.ID, copied.ID, 20*time.Second) {
t.Fatalf("copy %s still present in src after delete", copied.ID) t.Fatalf("copy %s still present in src after delete", copied.ID)
} }
// --- CRUD on the folders themselves, reusing the throwaway src/dst ---
// A child file lets us prove it survives the folder rename and move.
child, err := store.Upload(ctx, src.ID, "child-"+suffix+".txt", bytes.NewReader(content))
if err != nil {
t.Fatalf("Upload child: %v", err)
}
if !waitEntry(ctx, store, src.ID, child.ID, 15*time.Second) {
t.Fatalf("child %s not listed in src", child.ID)
}
fst, err := store.Stat(ctx, src.ID)
if err != nil {
t.Fatalf("Stat(folder): %v", err)
}
if fst.ID != src.ID || fst.Kind != Folder {
t.Fatalf("Stat(folder) = %+v", fst)
}
flist, err := store.List(ctx, src.ID)
if err != nil {
t.Fatalf("List(folder): %v", err)
}
if entryByID(flist, child.ID) == nil {
t.Fatalf("child %s not in List(src)", child.ID)
}
folderTitle := "renamed-folder-" + suffix
renameEventually(t, ctx, store, src.ID, folderTitle)
if e, err := store.Stat(ctx, src.ID); err != nil {
t.Fatalf("Stat(folder) after rename: %v", err)
} else if e.Kind != Folder || e.Title != folderTitle {
t.Fatalf("folder after rename = %+v, want title %q", e, folderTitle)
}
moveEventually(t, ctx, store, src.ID, dst.ID)
if !waitEntry(ctx, store, dst.ID, src.ID, 20*time.Second) {
t.Fatalf("moved folder %s not in dst %s", src.ID, dst.ID)
}
if !waitNoEntry(ctx, store, root, src.ID, 20*time.Second) {
t.Fatalf("folder %s still in root after move", src.ID)
}
if !waitEntry(ctx, store, src.ID, child.ID, 20*time.Second) {
t.Fatalf("child file %s lost after moving folder %s", child.ID, src.ID)
}
if err := store.Delete(ctx, []string{child.ID}); err != nil {
t.Fatalf("Delete(child): %v", err)
}
if err := store.Delete(ctx, []string{src.ID}); err != nil {
t.Fatalf("Delete(folder): %v", err)
}
if !waitNoEntry(ctx, store, dst.ID, src.ID, 20*time.Second) {
t.Fatalf("folder %s still present in dst after delete", src.ID)
}
if err := store.Delete(ctx, []string{dst.ID}); err != nil {
t.Fatalf("Delete(dst folder): %v", err)
}
if !waitNoEntry(ctx, store, root, dst.ID, 20*time.Second) {
t.Fatalf("dst folder %s still present in root after delete", dst.ID)
}
} }
// moveEventually issues Move and retries while the operation is not visible yet // moveEventually issues Move and retries while the operation is not visible yet
@@ -141,6 +200,23 @@ func moveEventually(t *testing.T, ctx context.Context, store FileStore, id, dstI
t.Fatalf("Move %s -> %s: %v", id, dstID, lastErr) t.Fatalf("Move %s -> %s: %v", id, dstID, lastErr)
} }
// copyEventually issues Copy and retries while the new copy is not visible yet
// (copy is accepted asynchronously, like move).
func copyEventually(t *testing.T, ctx context.Context, store FileStore, id, dstID string, d time.Duration) *Entry {
t.Helper()
var lastErr error
for attempt := 0; attempt < 5; attempt++ {
if lastErr = store.Copy(ctx, []string{id}, dstID); lastErr == nil {
if e := waitOtherFile(ctx, store, dstID, id, d); e != nil {
return e
}
}
time.Sleep(time.Second)
}
t.Fatalf("Copy %s -> %s: %v", id, dstID, lastErr)
return nil
}
func renameEventually(t *testing.T, ctx context.Context, store FileStore, id, title string) { func renameEventually(t *testing.T, ctx context.Context, store FileStore, id, title string) {
t.Helper() t.Helper()
var lastErr error var lastErr error
+69 -6
View File
@@ -52,8 +52,17 @@ type Entry struct {
MIME string MIME string
Created time.Time Created time.Time
Modified time.Time Modified time.Time
// Updated is the backend-native timestamp string, when the backend exposes
// one. It lets list output round-trip the API value; Modified is the
// parsed form for logic.
Updated string
Version int Version int
Provider string Provider string
// Folder-only counters. Zero for files and for backends that do not
// report them.
FilesCount int
FoldersCount int
} }
// FileStore is the operation surface every file backend implements. // FileStore is the operation surface every file backend implements.
@@ -71,13 +80,16 @@ type FileStore interface {
} }
// SearchQuery narrows a Searcher request. InContent asks the backend to match // SearchQuery narrows a Searcher request. InContent asks the backend to match
// document bodies, not just titles. // document bodies, not just titles. Substring switches title matching from the
// analyzer's whole-token match to a case-insensitive "*term*" wildcard and ANDs
// every whitespace-separated term (e.g. "rechnung 2025").
type SearchQuery struct { type SearchQuery struct {
Text string Text string
InContent bool InContent bool
FolderID string FolderID string
Extensions []string Extensions []string
Limit int Limit int
Substring bool
} }
// SearchHit is one Searcher result: the matching entry plus backend-specific // SearchHit is one Searcher result: the matching entry plus backend-specific
@@ -96,20 +108,66 @@ type Searcher interface {
Name() string Name() string
} }
// FileStore returns the adapter for a backend name: ProviderREST (default) or // FileStore returns the adapter for a backend name: ProviderREST (default),
// ProviderDAV. Unknown or empty names select the REST backend. The full facade // ProviderDAV (alias "webdav") or the read-only SQL store (ProviderPG,
// (backend composition) is deliberately left to a later change. // ProviderMySQL and the aliases "pg"/"sql"). The SQL store is opened from the
// environment (ONLYOFFICE_DSN / ONLYOFFICE_PG_*); when it cannot be opened the
// returned store surfaces that error on every operation instead of returning
// nil. Use SQLFileStore when the open error itself is needed. Unknown or empty
// names select the REST backend. The composed facade (backend
// selection/fallback) lives on FileClient in filestore_facade.go.
func (c *Client) FileStore(backend string) FileStore { func (c *Client) FileStore(backend string) FileStore {
switch strings.ToLower(strings.TrimSpace(backend)) { switch strings.ToLower(strings.TrimSpace(backend)) {
case ProviderDAV, "webdav": case ProviderDAV, "webdav":
return &davStore{c: c} return &davStore{c: c}
case ProviderPG, ProviderMySQL, "pg", "sql":
s, err := c.SQLFileStore()
if err != nil {
return &errStore{name: strings.ToLower(strings.TrimSpace(backend)), err: err}
}
return s
default: default:
return &restStore{c: c} return &restStore{c: c}
} }
} }
// Files returns the default (REST) file store. // errStore is the FileStore placeholder returned when a backend cannot be
func (c *Client) Files() FileStore { return c.FileStore(ProviderREST) } // opened (for example SQL without a DSN). Every operation returns the recorded
// error instead of panicking on a nil interface.
type errStore struct {
name string
err error
}
func (s *errStore) Name() string { return s.name }
func (s *errStore) List(context.Context, string) ([]Entry, error) { return nil, s.err }
func (s *errStore) Stat(context.Context, string) (Entry, error) { return Entry{}, s.err }
func (s *errStore) CreateFolder(context.Context, string, string) (Entry, error) {
return Entry{}, s.err
}
func (s *errStore) Upload(context.Context, string, string, io.Reader) (Entry, error) {
return Entry{}, s.err
}
func (s *errStore) Download(context.Context, string, io.Writer) (int64, error) {
return 0, s.err
}
func (s *errStore) Move(context.Context, []string, string) error { return s.err }
func (s *errStore) Copy(context.Context, []string, string) error { return s.err }
func (s *errStore) Rename(context.Context, string, string) error { return s.err }
func (s *errStore) Delete(context.Context, []string) error { return s.err }
// Files returns the composed file facade. The returned *FileClient implements
// FileStore, so callers that used Files() as the plain REST store keep working.
func (c *Client) Files() *FileClient { return c.newFileClient() }
// retryStoreOp runs one store operation under the shared deterministic // retryStoreOp runs one store operation under the shared deterministic
// transient-error policy (429/502/503/504). // transient-error policy (429/502/503/504).
@@ -140,6 +198,7 @@ func FileEntryToEntry(f *FileEntry, provider string) Entry {
e.MIME = mimeForTitle(e.Title, exst) e.MIME = mimeForTitle(e.Title, exst)
if f.Updated != nil { if f.Updated != nil {
e.Modified = *f.Updated e.Modified = *f.Updated
e.Updated = f.Updated.Format(time.RFC3339)
} }
return e return e
} }
@@ -153,6 +212,7 @@ func DavFileToEntry(f DavFile, provider string) Entry {
Size: f.Size, Size: f.Size,
MIME: mimeForTitle(f.Title, ""), MIME: mimeForTitle(f.Title, ""),
Modified: f.ModTime(), Modified: f.ModTime(),
Updated: f.Updated,
Provider: provider, Provider: provider,
} }
} }
@@ -165,7 +225,10 @@ func DavFolderToEntry(f DavFolder, provider string) Entry {
Title: f.Title, Title: f.Title,
Kind: Folder, Kind: Folder,
Modified: f.ModTime(), Modified: f.ModTime(),
Updated: f.Updated,
Provider: provider, Provider: provider,
FilesCount: f.FilesCount,
FoldersCount: f.FoldersCount,
} }
} }
View File
+50 -10
View File
@@ -23,12 +23,12 @@ import (
) )
// The canonical model (Kind, Entry, SearchQuery, SearchHit, Searcher) lives in // The canonical model (Kind, Entry, SearchQuery, SearchHit, Searcher) lives in
// file_core.go (F1 #35). // filestore_core.go (F1 #35).
const ( const (
defaultESIndex = "files_file" defaultESIndex = "files_file"
defaultESLimit = 20 defaultESLimit = 20
maxESLimit = 200 maxESLimit = 1000
maxESResponseSize = 8 << 20 maxESResponseSize = 8 << 20
) )
@@ -119,14 +119,30 @@ func esSearchRequest(q SearchQuery, tenant string) esRequest {
if q.InContent { if q.InContent {
fields = append(fields, "document.attachment.content") fields = append(fields, "document.attachment.content")
} }
must := []esClause{{MultiMatch: &esMultiMatch{Query: q.Text, Fields: fields}}} var must []esClause
if q.Substring {
for _, term := range strings.Fields(strings.ToLower(q.Text)) {
if term = escapeWildcard(term); term != "" {
must = append(must, esClause{Wildcard: map[string]any{"title": "*" + term + "*"}})
}
}
}
if len(must) == 0 {
must = []esClause{{MultiMatch: &esMultiMatch{Query: q.Text, Fields: fields}}}
}
var filter []esClause var filter []esClause
if t := strings.TrimSpace(tenant); t != "" { if t := strings.TrimSpace(tenant); t != "" {
filter = append(filter, esClause{Term: map[string]any{"tenantId": numericOrString(t)}}) filter = append(filter, esClause{Term: map[string]any{"tenantId": numericOrString(t)}})
} }
if f := strings.TrimSpace(q.FolderID); f != "" { if f := strings.TrimSpace(q.FolderID); f != "" {
filter = append(filter, esClause{Term: map[string]any{"folders.folderId": f}}) // folders is an ES nested field; a plain term on folders.folderId would
// not match. The stored Folders list holds every ancestor id, so
// filtering by a project root id scopes to its whole subtree.
filter = append(filter, esClause{Nested: &esNested{
Path: "folders",
Query: esNestedTerm{Term: map[string]any{"folders.folderId": f}},
}})
} }
for _, ext := range normalizeExtensions(q.Extensions) { for _, ext := range normalizeExtensions(q.Extensions) {
filter = append(filter, esClause{Wildcard: map[string]any{"title": "*." + ext}}) filter = append(filter, esClause{Wildcard: map[string]any{"title": "*." + ext}})
@@ -188,7 +204,24 @@ type esBool struct {
type esClause struct { type esClause struct {
MultiMatch *esMultiMatch `json:"multi_match,omitempty"` MultiMatch *esMultiMatch `json:"multi_match,omitempty"`
Term map[string]any `json:"term,omitempty"` Term map[string]any `json:"term,omitempty"`
Terms map[string]any `json:"terms,omitempty"`
Wildcard map[string]any `json:"wildcard,omitempty"` Wildcard map[string]any `json:"wildcard,omitempty"`
Nested *esNested `json:"nested,omitempty"`
}
type esNested struct {
Path string `json:"path"`
Query esNestedTerm `json:"query"`
}
type esNestedTerm struct {
Term map[string]any `json:"term,omitempty"`
}
// escapeWildcard strips ES wildcard metacharacters from a user term so a query
// cannot inject wildcard syntax. Pure, so it is unit-tested.
func escapeWildcard(s string) string {
return strings.NewReplacer("*", "", "?", "", `\`, "").Replace(s)
} }
type esMultiMatch struct { type esMultiMatch struct {
@@ -241,13 +274,19 @@ func parseESSearchResponse(raw []byte) ([]SearchHit, error) {
if h.Source.ID == 0 { if h.Source.ID == 0 {
id = h.ID id = h.ID
} }
// Folders is the ancestor breadcrumb in root → leaf order, so the last
// entry is the immediate parent (the previous "first" value was the
// project root, which made every result look like it lived in #522).
var parent string var parent string
path := make([]string, 0, len(h.Source.Folders)) path := make([]string, 0, len(h.Source.Folders))
for i, f := range h.Source.Folders { for _, f := range h.Source.Folders {
path = append(path, f.FolderID) if strings.TrimSpace(f.FolderID) == "" {
if i == 0 { continue
parent = f.FolderID
} }
path = append(path, f.FolderID)
}
if len(path) > 0 {
parent = path[len(path)-1]
} }
hits = append(hits, SearchHit{ hits = append(hits, SearchHit{
Entry: Entry{ Entry: Entry{
@@ -268,9 +307,10 @@ func parseESSearchResponse(raw []byte) ([]SearchHit, error) {
var esHighlightTag = regexp.MustCompile(`</?em[^>]*>`) var esHighlightTag = regexp.MustCompile(`</?em[^>]*>`)
// esHighlightText flattens a highlight map into one plain-text snippet, // esHighlightText flattens a highlight map into one plain-text snippet,
// preferring the content fragment over the title. // preferring the content fragment over the title. It covers both the
// OnlyOffice content field and the own-index "content" field.
func esHighlightText(hl map[string][]string) string { func esHighlightText(hl map[string][]string) string {
for _, key := range []string{"document.attachment.content", "title"} { for _, key := range []string{"document.attachment.content", "content", "title"} {
frags := hl[key] frags := hl[key]
if len(frags) == 0 { if len(frags) == 0 {
continue continue
@@ -14,6 +14,97 @@ import (
"github.com/xuri/excelize/v2" "github.com/xuri/excelize/v2"
) )
// Default known fixtures for TestIntegrationESFacadeUsesES on the live index.
const (
defaultESTestQuery = "Rechnung_986-2025.pdf"
defaultESTestSubstring = "rechnung 2025"
defaultESTestFolder = "522"
)
// TestIntegrationESFacadeUsesES proves that the public search path — `oo search`
// and Client.Files().Search() — really runs against the OnlyOffice
// Elasticsearch backend and not the REST @search endpoint, which only looks at
// file names in the database and is not a Searcher at all (see
// docs/elasticsearch.md). It pins the concrete backend and checks that a known
// document comes back with a non-empty id and folder path.
//
// Requires ONLYOFFICE_ES_URL only — the query never touches the REST API, so no
// OnlyOffice credentials are needed. Skips when it is missing. The fixture is
// overridable with ONLYOFFICE_ES_TEST_QUERY, ONLYOFFICE_ES_TEST_TITLE,
// ONLYOFFICE_ES_TEST_SUBSTRING and ONLYOFFICE_ES_TEST_FOLDER.
func TestIntegrationESFacadeUsesES(t *testing.T) {
esURL := strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL"))
if esURL == "" {
t.Skip("ONLYOFFICE_ES_URL not set — skipping Elasticsearch integration test")
}
query := firstNonEmpty(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_TEST_QUERY")), defaultESTestQuery)
wantTitle := firstNonEmpty(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_TEST_TITLE")), query)
substring := firstNonEmpty(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_TEST_SUBSTRING")), defaultESTestSubstring)
folder := firstNonEmpty(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_TEST_FOLDER")), defaultESTestFolder)
c := NewClient(Credentials{})
searcher, err := c.Files().Search()
if err != nil {
t.Fatalf("Files().Search(): %v", err)
}
if got := searcher.Name(); got != ProviderES {
t.Fatalf("searcher.Name() = %q, want %q (REST @search is not a Searcher)", got, ProviderES)
}
if _, ok := searcher.(*ESSearcher); !ok {
t.Fatalf("searcher = %T, want *ESSearcher (ES backend, not REST)", searcher)
}
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
// Known file name: multi_match over title, as `oo search <file>` does.
start := time.Now()
hits, err := searcher.Search(ctx, SearchQuery{Text: query, Limit: 20})
if err != nil {
t.Fatalf("Search(%q): %v", query, err)
}
t.Logf("ES facade query %q: %d hits in %s", query, len(hits), time.Since(start))
known := findHitByTitle(hits, wantTitle)
if known == nil {
t.Fatalf("query %q returned %d hits, none titled %q", query, len(hits), wantTitle)
}
if strings.TrimSpace(known.ID) == "" {
t.Errorf("hit %q has empty id", known.Title)
}
if len(known.Path) == 0 {
t.Errorf("hit %q has empty path", known.Title)
}
if known.Provider != ProviderES {
t.Errorf("hit provider = %q, want %q", known.Provider, ProviderES)
}
// Substring + folder subtree, as `oo search <terms> --substring --folder N`
// does: wildcard terms ANDed together, scoped to the folder's subtree.
start = time.Now()
subHits, err := searcher.Search(ctx, SearchQuery{Text: substring, Substring: true, FolderID: folder, Limit: 200})
if err != nil {
t.Fatalf("substring Search(%q, folder %s): %v", substring, folder, err)
}
t.Logf("ES facade substring %q folder %s: %d hits in %s", substring, folder, len(subHits), time.Since(start))
if len(subHits) == 0 {
t.Fatalf("substring query %q in folder %s returned no hits", substring, folder)
}
if findHitByTitle(subHits, wantTitle) == nil {
t.Errorf("substring query %q in folder %s did not return %q", substring, folder, wantTitle)
}
}
// findHitByTitle returns the first hit whose title matches, case-insensitively.
func findHitByTitle(hits []SearchHit, title string) *SearchHit {
for i := range hits {
if strings.EqualFold(strings.TrimSpace(hits[i].Title), title) {
return &hits[i]
}
}
return nil
}
// TestIntegrationESSearch uploads a throwaway workbook and verifies that the // TestIntegrationESSearch uploads a throwaway workbook and verifies that the
// direct Elasticsearch search finds it by file name and by content. // direct Elasticsearch search finds it by file name and by content.
// //
+35 -6
View File
@@ -55,7 +55,7 @@ func TestESSearchRequestFiltersAndLimit(t *testing.T) {
Text: "Storchen", Text: "Storchen",
FolderID: "649", FolderID: "649",
Extensions: []string{".PDF", "pdf", "docx"}, Extensions: []string{".PDF", "pdf", "docx"},
Limit: 999, Limit: 5000,
}, "42") }, "42")
if got.Size != maxESLimit { if got.Size != maxESLimit {
t.Errorf("size = %d, want cap %d", got.Size, maxESLimit) t.Errorf("size = %d, want cap %d", got.Size, maxESLimit)
@@ -65,10 +65,10 @@ func TestESSearchRequestFiltersAndLimit(t *testing.T) {
switch { switch {
case f.Term != nil && f.Term["tenantId"] != nil: case f.Term != nil && f.Term["tenantId"] != nil:
tenant++ tenant++
case f.Term != nil && f.Term["folders.folderId"] != nil: case f.Nested != nil:
folder++ folder++
if f.Term["folders.folderId"] != "649" { if f.Nested.Path != "folders" || f.Nested.Query.Term["folders.folderId"] != "649" {
t.Errorf("folder filter = %+v", f.Term) t.Errorf("folder filter = %+v, want nested folders term 649", f.Nested)
} }
case f.Wildcard != nil: case f.Wildcard != nil:
wildcards++ wildcards++
@@ -82,6 +82,34 @@ func TestESSearchRequestFiltersAndLimit(t *testing.T) {
} }
} }
func TestESSearchRequestSubstringAndsTerms(t *testing.T) {
got := esSearchRequest(SearchQuery{Text: "Rechnung 2025", Substring: true, FolderID: "522"}, "")
if got.Query.Bool.Must[0].MultiMatch != nil {
t.Fatalf("substring must not use multi_match: %+v", got.Query.Bool.Must)
}
if len(got.Query.Bool.Must) != 2 {
t.Fatalf("must = %+v, want two ANDed wildcard terms", got.Query.Bool.Must)
}
want := []string{"*rechnung*", "*2025*"}
for i, m := range got.Query.Bool.Must {
if m.Wildcard == nil || m.Wildcard["title"] != want[i] {
t.Errorf("must[%d] = %+v, want title wildcard %q", i, m, want[i])
}
}
if len(got.Query.Bool.Filter) != 1 || got.Query.Bool.Filter[0].Nested == nil {
t.Errorf("folder filter = %+v, want nested", got.Query.Bool.Filter)
}
}
func TestEscapeWildcard(t *testing.T) {
cases := map[string]string{"*rechnung*": "rechnung", "a?b\\c": "abc", "plain": "plain"}
for in, want := range cases {
if got := escapeWildcard(in); got != want {
t.Errorf("escapeWildcard(%q) = %q, want %q", in, got, want)
}
}
}
func TestESSearchRequestRejectsEmptyTextAtSearch(t *testing.T) { func TestESSearchRequestRejectsEmptyTextAtSearch(t *testing.T) {
s, err := NewESSearcher(ESConfig{URL: "http://localhost:9200"}) s, err := NewESSearcher(ESConfig{URL: "http://localhost:9200"})
if err != nil { if err != nil {
@@ -155,8 +183,9 @@ func TestParseESSearchResponse(t *testing.T) {
if h0.ID != "2395" || h0.Title != "Rechnung-4711.pdf" || h0.Kind != File { if h0.ID != "2395" || h0.Title != "Rechnung-4711.pdf" || h0.Kind != File {
t.Errorf("hit0 entry = %+v", h0.Entry) t.Errorf("hit0 entry = %+v", h0.Entry)
} }
if h0.ParentID != "438" || !reflect.DeepEqual(h0.Path, []string{"438", "11"}) { // folders is root → leaf; the immediate parent is the last entry.
t.Errorf("hit0 path = %v parent = %q", h0.Path, h0.ParentID) if h0.ParentID != "11" || !reflect.DeepEqual(h0.Path, []string{"438", "11"}) {
t.Errorf("hit0 path = %v parent = %q, want parent 11", h0.Path, h0.ParentID)
} }
if h0.Score != 7.31 { if h0.Score != 7.31 {
t.Errorf("hit0 score = %v", h0.Score) t.Errorf("hit0 score = %v", h0.Score)
+336
View File
@@ -0,0 +1,336 @@
package onlyoffice
// Own full-text index (epic #34, F6 #42).
//
// The OnlyOffice Elasticsearch index (files_file) only holds extracted content
// for Office formats. FileUtility.CanIndex gates extraction by the server
// setting files.index.formats, whose default is ".pptx|.xlsx|.docx", so PDFs
// are indexed by name only. Instead of patching the server (risky: lost on
// upgrade, forces a full reindex) this file implements a second, independent
// index (default oo_docs_text) that our own pipeline fills from
// internal/docpipe (pdftotext + OCR). The OnlyOffice index is never touched.
//
// See docs/elasticsearch.md for the decision and the trade-offs.
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
const defaultESTextIndex = "oo_docs_text"
// ESTextConfig configures the own full-text index.
type ESTextConfig struct {
URL string // scheme://host:port of the ES HTTP endpoint
Index string // index name, default oo_docs_text
Tenant string // reserved for future multi-tenant data; unused for now
}
// ESTextConfigFromEnv reads ONLYOFFICE_ES_URL and ONLYOFFICE_ES_TEXT_INDEX
// (default oo_docs_text). The library never loads dotfiles — the CLI does that.
func ESTextConfigFromEnv() ESTextConfig {
return ESTextConfig{
URL: strings.TrimRight(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL")), "/"),
Index: firstNonEmpty(os.Getenv("ONLYOFFICE_ES_TEXT_INDEX"), defaultESTextIndex),
Tenant: strings.TrimSpace(os.Getenv("ONLYOFFICE_TENANT")),
}
}
// TextDoc is one document in the own full-text index. It is keyed by the
// OnlyOffice file id so hits map straight back to Documents entries.
type TextDoc struct {
ID string `json:"id"`
Title string `json:"title"`
FolderID string `json:"folder,omitempty"`
Ext string `json:"ext,omitempty"`
Content string `json:"content"`
}
// TextIndex is the storage/search surface for locally extracted document text.
// It complements Searcher: ESSearcher reads OnlyOffice's index, ESTextIndex
// reads ours.
type TextIndex interface {
Put(ctx context.Context, docs []TextDoc) error
Delete(ctx context.Context, ids []string) error
Search(ctx context.Context, q SearchQuery) ([]SearchHit, error)
Name() string
}
// ESTextIndex is a TextIndex (and Searcher) over a dedicated Elasticsearch
// index filled by TextIndexer.
type ESTextIndex struct {
cfg ESTextConfig
http *http.Client
}
var (
_ TextIndex = (*ESTextIndex)(nil)
_ Searcher = (*ESTextIndex)(nil)
)
// NewESTextIndex returns a searcher/writer for the own full-text index. The URL
// is required; an empty index falls back to oo_docs_text.
func NewESTextIndex(cfg ESTextConfig) (*ESTextIndex, error) {
if strings.TrimSpace(cfg.URL) == "" {
return nil, fmt.Errorf("onlyoffice: elasticsearch URL is empty (set ONLYOFFICE_ES_URL)")
}
cfg.URL = strings.TrimRight(cfg.URL, "/")
if cfg.Index == "" {
cfg.Index = defaultESTextIndex
}
return &ESTextIndex{cfg: cfg, http: &http.Client{Timeout: 120 * time.Second}}, nil
}
// Name implements Searcher and TextIndex.
func (x *ESTextIndex) Name() string { return "es-text" }
// Index returns the configured index name.
func (x *ESTextIndex) Index() string { return x.cfg.Index }
// esTextMapping pins explicit types: content must stay a plain text field (no
// keyword subfield) and folder/ext stay exact keywords for filters.
const esTextMapping = `{
"mappings": {
"properties": {
"id": {"type": "keyword"},
"title": {"type": "text", "fields": {"keyword": {"type": "keyword", "ignore_above": 512}}},
"folder": {"type": "keyword"},
"ext": {"type": "keyword"},
"content": {"type": "text"}
}
}
}`
// Ensure creates the index with the explicit mapping. A missing index is
// created; an already existing one is left untouched.
func (x *ESTextIndex) Ensure(ctx context.Context) error {
status, raw, err := x.do(ctx, http.MethodPut, "/"+x.cfg.Index, []byte(esTextMapping), "application/json")
if err != nil {
return err
}
if status == http.StatusOK {
return nil
}
if status == http.StatusBadRequest && bytes.Contains(raw, []byte("resource_already_exists_exception")) {
return nil
}
return fmt.Errorf("onlyoffice: create text index %s: %d %s", x.cfg.Index, status, truncate(string(raw), 300))
}
// Put upserts documents via the bulk API and refreshes so they are immediately
// searchable.
func (x *ESTextIndex) Put(ctx context.Context, docs []TextDoc) error {
if len(docs) == 0 {
return nil
}
status, raw, err := x.do(ctx, http.MethodPost, "/"+x.cfg.Index+"/_bulk?refresh=true", esTextBulkBody(docs), "application/x-ndjson")
if err != nil {
return err
}
if status >= 400 {
return fmt.Errorf("onlyoffice: bulk index %s: %d %s", x.cfg.Index, status, truncate(string(raw), 400))
}
var res esBulkResponse
if err := json.Unmarshal(raw, &res); err != nil {
return fmt.Errorf("onlyoffice: decode bulk response: %w", err)
}
if !res.Errors {
return nil
}
return fmt.Errorf("onlyoffice: bulk index %s: %s", x.cfg.Index, res.firstError())
}
// Delete removes documents by OnlyOffice file id. A missing index means there
// is nothing to delete.
func (x *ESTextIndex) Delete(ctx context.Context, ids []string) error {
if len(ids) == 0 {
return nil
}
body, err := json.Marshal(map[string]any{"query": map[string]any{"terms": map[string]any{"id": ids}}})
if err != nil {
return err
}
status, raw, err := x.do(ctx, http.MethodPost, "/"+x.cfg.Index+"/_delete_by_query?refresh=true", body, "application/json")
if err != nil {
return err
}
if status == http.StatusNotFound {
return nil
}
if status >= 400 {
return fmt.Errorf("onlyoffice: delete from %s: %d %s", x.cfg.Index, status, truncate(string(raw), 400))
}
return nil
}
// Search runs a multi_match over title (boosted) and content, with optional
// folder and extension filters. A missing index yields no hits, not an error.
func (x *ESTextIndex) Search(ctx context.Context, q SearchQuery) ([]SearchHit, error) {
q.Text = strings.TrimSpace(q.Text)
if q.Text == "" {
return nil, fmt.Errorf("onlyoffice: empty search query")
}
body, err := json.Marshal(esTextSearchRequest(q))
if err != nil {
return nil, fmt.Errorf("onlyoffice: build elasticsearch query: %w", err)
}
status, raw, err := x.do(ctx, http.MethodPost, "/"+x.cfg.Index+"/_search", body, "application/json")
if err != nil {
return nil, err
}
if status == http.StatusNotFound {
return nil, nil
}
if status >= 400 {
return nil, fmt.Errorf("onlyoffice: search %s: %d %s", x.cfg.Index, status, truncate(string(raw), 400))
}
return parseESTextResponse(raw)
}
// do sends one request and returns the status and body (bounded). The caller
// decides which statuses are errors.
func (x *ESTextIndex) do(ctx context.Context, method, path string, body []byte, contentType string) (int, []byte, error) {
var r io.Reader
if body != nil {
r = bytes.NewReader(body)
}
req, err := http.NewRequestWithContext(ctx, method, x.cfg.URL+path, r)
if err != nil {
return 0, nil, err
}
req.Header.Set("Accept", "application/json")
if contentType != "" {
req.Header.Set("Content-Type", contentType)
}
resp, err := x.http.Do(req)
if err != nil {
return 0, nil, fmt.Errorf("onlyoffice: elasticsearch %s: %w", method, err)
}
defer resp.Body.Close()
raw, err := io.ReadAll(io.LimitReader(resp.Body, maxESResponseSize))
if err != nil {
return resp.StatusCode, nil, err
}
return resp.StatusCode, raw, nil
}
// esTextBulkBody renders the NDJSON bulk payload. Pure, so it is unit-tested.
func esTextBulkBody(docs []TextDoc) []byte {
var b bytes.Buffer
enc := json.NewEncoder(&b)
enc.SetEscapeHTML(false)
for _, d := range docs {
_ = enc.Encode(map[string]any{"index": map[string]any{"_id": d.ID}})
_ = enc.Encode(d)
}
return b.Bytes()
}
// esTextSearchRequest builds the own-index query. Pure, so it is unit-tested.
func esTextSearchRequest(q SearchQuery) esRequest {
limit := q.Limit
if limit <= 0 {
limit = defaultESLimit
}
if limit > maxESLimit {
limit = maxESLimit
}
fields := []string{"title^2", "content"}
must := []esClause{{MultiMatch: &esMultiMatch{Query: q.Text, Fields: fields}}}
var filter []esClause
if f := strings.TrimSpace(q.FolderID); f != "" {
filter = append(filter, esClause{Term: map[string]any{"folder": f}})
}
if exts := normalizeExtensions(q.Extensions); len(exts) > 0 {
filter = append(filter, esClause{Terms: map[string]any{"ext": exts}})
}
return esRequest{
Size: limit,
Source: []string{"id", "title", "folder", "ext"},
Query: esQuery{Bool: esBool{Must: must, Filter: filter}},
Highlight: esHighlight{PreTags: []string{"<em>"}, PostTags: []string{"</em>"}, Fields: map[string]struct{}{"title": {}, "content": {}}},
}
}
// esBulkResponse is the subset of an ES bulk response we consume.
type esBulkResponse struct {
Errors bool `json:"errors"`
Items []map[string]struct {
ID string `json:"_id"`
Status int `json:"status"`
Error *struct {
Type string `json:"type"`
Reason string `json:"reason"`
} `json:"error"`
} `json:"items"`
}
// firstError returns a compact description of the first failed bulk item.
func (r esBulkResponse) firstError() string {
for _, item := range r.Items {
for op, res := range item {
if res.Error != nil {
return fmt.Sprintf("%s %s: %s %s", op, res.ID, res.Error.Type, res.Error.Reason)
}
}
}
return "unknown bulk error"
}
// esTextResponse is the subset of an own-index search response we consume.
type esTextResponse struct {
Hits struct {
Total struct {
Value int `json:"value"`
} `json:"total"`
Hits []struct {
ID string `json:"_id"`
Score float64 `json:"_score"`
Source TextDoc `json:"_source"`
HL map[string][]string `json:"highlight"`
} `json:"hits"`
} `json:"hits"`
}
// parseESTextResponse converts an own-index search response into SearchHit
// values. Pure, so it is unit-tested.
func parseESTextResponse(raw []byte) ([]SearchHit, error) {
var r esTextResponse
if err := json.Unmarshal(raw, &r); err != nil {
return nil, fmt.Errorf("onlyoffice: decode elasticsearch response: %w", err)
}
hits := make([]SearchHit, 0, len(r.Hits.Hits))
for _, h := range r.Hits.Hits {
id := h.Source.ID
if id == "" {
id = h.ID
}
parent := h.Source.FolderID
var path []string
if parent != "" {
path = []string{parent}
}
hits = append(hits, SearchHit{
Entry: Entry{
ID: id,
ParentID: parent,
Title: h.Source.Title,
Kind: File,
Provider: "es-text",
},
Score: h.Score,
Highlight: esHighlightText(h.HL),
Path: path,
})
}
return hits, nil
}
+158
View File
@@ -0,0 +1,158 @@
//go:build integration
package onlyoffice
import (
"context"
"net/http"
"os"
"strings"
"testing"
"time"
"github.com/eslider/go-onlyoffice/internal/docpipe"
)
// TestIntegrationESTextIndex verifies the own full-text index end to end
// against a live Elasticsearch: create the index with its mapping, index a
// document, find it by content (and reject it via a folder filter and after
// deletion), then drop the throwaway index.
//
// Requires ONLYOFFICE_ES_URL (a reachable ES endpoint — in the current setup a
// tunnel to the ES inside the OnlyOffice VM, see docs/elasticsearch.md). It
// does not need OnlyOffice credentials because no file is downloaded: the
// TextIndexer write path is covered by unit tests with a fake extractor.
func TestIntegrationESTextIndex(t *testing.T) {
esURL := strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL"))
if esURL == "" {
t.Skip("ONLYOFFICE_ES_URL not set — skipping Elasticsearch integration test")
}
stamp := time.Now().UTC().Format("20060102150405")
idx, err := NewESTextIndex(ESTextConfig{URL: esURL, Index: "oo_docs_text_it_" + stamp})
if err != nil {
t.Fatalf("NewESTextIndex: %v", err)
}
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
t.Cleanup(func() {
cleanupCtx, done := context.WithTimeout(context.Background(), 30*time.Second)
defer done()
_, _, _ = idx.do(cleanupCtx, http.MethodDelete, "/"+idx.Index(), nil, "")
})
if err := idx.Ensure(ctx); err != nil {
t.Fatalf("Ensure: %v", err)
}
// Ensure is idempotent.
if err := idx.Ensure(ctx); err != nil {
t.Fatalf("Ensure (second): %v", err)
}
token := "gotes" + stamp
doc := TextDoc{
ID: "3578",
Title: "2026-07-28-S1021-acme-rechnung.pdf",
FolderID: "634",
Ext: "pdf",
Content: "Begleitzettel SGB XI — Rechnung " + token,
}
if err := idx.Put(ctx, []TextDoc{doc}); err != nil {
t.Fatalf("Put: %v", err)
}
hits, err := idx.Search(ctx, SearchQuery{Text: token})
if err != nil {
t.Fatalf("Search: %v", err)
}
if len(hits) != 1 || hits[0].ID != "3578" {
t.Fatalf("content search hits = %+v, want doc 3578", hits)
}
if !strings.Contains(hits[0].Highlight, token) {
t.Errorf("highlight = %q, want token", hits[0].Highlight)
}
if hits, err := idx.Search(ctx, SearchQuery{Text: token, FolderID: "999"}); err != nil {
t.Fatalf("Search with folder filter: %v", err)
} else if len(hits) != 0 {
t.Errorf("folder filter returned %d hits, want 0", len(hits))
}
if hits, err := idx.Search(ctx, SearchQuery{Text: token, Extensions: []string{"docx"}}); err != nil {
t.Fatalf("Search with ext filter: %v", err)
} else if len(hits) != 0 {
t.Errorf("ext filter returned %d hits, want 0", len(hits))
}
if err := idx.Delete(ctx, []string{"3578"}); err != nil {
t.Fatalf("Delete: %v", err)
}
if hits, err := idx.Search(ctx, SearchQuery{Text: token}); err != nil {
t.Fatalf("Search after delete: %v", err)
} else if len(hits) != 0 {
t.Errorf("after delete search returned %d hits, want 0", len(hits))
}
}
// TestIntegrationESTextIndexPDFAttachment indexes testdata/pdf-with-attachment.pdf
// through the real pipeline (TextIndexer + docpipe: pdfdetach + pdftotext) and
// verifies that text living only in the embedded attachment is searchable.
//
// Requires ONLYOFFICE_ES_URL plus poppler (pdfdetach/pdftotext). No OnlyOffice
// credentials are needed: a fixture FileStore serves the PDF bytes.
func TestIntegrationESTextIndexPDFAttachment(t *testing.T) {
esURL := strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL"))
if esURL == "" {
t.Skip("ONLYOFFICE_ES_URL not set — skipping Elasticsearch integration test")
}
if docpipe.LookPath().PDFDetach == "" {
t.Skip("pdfdetach not on PATH — skipping PDF attachment integration test")
}
pdf, err := os.ReadFile("testdata/pdf-with-attachment.pdf")
if err != nil {
t.Fatalf("read fixture: %v", err)
}
stamp := time.Now().UTC().Format("20060102150405")
idx, err := NewESTextIndex(ESTextConfig{URL: esURL, Index: "oo_docs_text_it_att_" + stamp})
if err != nil {
t.Fatalf("NewESTextIndex: %v", err)
}
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
t.Cleanup(func() {
cleanupCtx, done := context.WithTimeout(context.Background(), 30*time.Second)
defer done()
_, _, _ = idx.do(cleanupCtx, http.MethodDelete, "/"+idx.Index(), nil, "")
})
if err := idx.Ensure(ctx); err != nil {
t.Fatalf("Ensure: %v", err)
}
store := &textFakeStore{files: map[string][]byte{"9001": pdf}}
ix := NewTextIndexer(store, idx)
res, err := ix.IndexEntries(ctx, []Entry{{ID: "9001", Title: "scan.pdf", ParentID: "777", Kind: File}}, IndexOptions{MinChars: 1})
if err != nil {
t.Fatalf("IndexEntries: %v", err)
}
if res.Indexed != 1 || res.Failed != 0 {
t.Fatalf("result = %+v, want one indexed doc", res)
}
// Token appears only inside the embedded goo-note.txt attachment.
hits, err := idx.Search(ctx, SearchQuery{Text: "gooattachmenttoken"})
if err != nil {
t.Fatalf("Search attachment token: %v", err)
}
if len(hits) != 1 || hits[0].ID != "9001" {
t.Fatalf("attachment-token hits = %+v, want doc 9001", hits)
}
if !strings.Contains(hits[0].Highlight, "gooattachmenttoken") {
t.Errorf("highlight = %q, want attachment token", hits[0].Highlight)
}
// Body text is indexed as before.
if hits, err := idx.Search(ctx, SearchQuery{Text: "goobodytoken"}); err != nil {
t.Fatalf("Search body token: %v", err)
} else if len(hits) != 1 {
t.Errorf("body-token hits = %d, want 1", len(hits))
}
}
+275
View File
@@ -0,0 +1,275 @@
package onlyoffice
import (
"context"
"fmt"
"io"
"os"
"reflect"
"strings"
"testing"
)
func TestESTextSearchRequestShape(t *testing.T) {
got := esTextSearchRequest(SearchQuery{
Text: "S1021",
FolderID: "634",
Extensions: []string{".PDF", "pdf"},
Limit: 5,
})
if got.Size != 5 {
t.Errorf("size = %d, want 5", got.Size)
}
mm := got.Query.Bool.Must[0].MultiMatch
if mm == nil || !reflect.DeepEqual(mm.Fields, []string{"title^2", "content"}) {
t.Fatalf("multi_match = %+v, want title^2 + content", mm)
}
if _, ok := got.Highlight.Fields["content"]; !ok {
t.Error("content highlight missing")
}
if _, ok := got.Highlight.Fields["title"]; !ok {
t.Error("title highlight missing")
}
var folder, exts int
for _, f := range got.Query.Bool.Filter {
switch {
case f.Term != nil && f.Term["folder"] != nil:
folder++
if f.Term["folder"] != "634" {
t.Errorf("folder term = %+v", f.Term)
}
case f.Terms != nil:
exts++
if !reflect.DeepEqual(f.Terms["ext"], []string{"pdf"}) {
t.Errorf("ext terms = %+v, want deduped pdf", f.Terms)
}
}
}
if folder != 1 || exts != 1 {
t.Errorf("filters folder=%d exts=%d, want 1 each", folder, exts)
}
}
func TestESTextBulkBody(t *testing.T) {
body := esTextBulkBody([]TextDoc{
{ID: "3578", Title: "S1021.pdf", FolderID: "634", Ext: "pdf", Content: "Begleitzettel <S1021> & mehr"},
{ID: "3579", Title: "S1023.pdf", Ext: "pdf", Content: "x"},
})
lines := strings.Split(strings.TrimRight(string(body), "\n"), "\n")
if len(lines) != 4 {
t.Fatalf("bulk body has %d lines, want 4:\n%s", len(lines), body)
}
if !strings.Contains(lines[0], `"index"`) || !strings.Contains(lines[0], `"_id":"3578"`) {
t.Errorf("action line = %q", lines[0])
}
if !strings.Contains(lines[1], `"content":"Begleitzettel <S1021> & mehr"`) {
t.Errorf("source line should keep HTML unescaped, got %q", lines[1])
}
if !strings.Contains(lines[2], `"_id":"3579"`) {
t.Errorf("second action line = %q", lines[2])
}
}
func TestParseESTextResponse(t *testing.T) {
raw := []byte(`{
"hits": {
"total": {"value": 1, "relation": "eq"},
"hits": [
{
"_id": "3578",
"_score": 3.21,
"_source": {"id": "3578", "title": "2026-07-28-S1021-acme-rechnung.pdf", "folder": "634", "ext": "pdf"},
"highlight": {"content": ["Begleitzettel … <em>S1021</em> …"]}
}
]
}
}`)
hits, err := parseESTextResponse(raw)
if err != nil {
t.Fatalf("parseESTextResponse: %v", err)
}
if len(hits) != 1 {
t.Fatalf("hits = %d, want 1", len(hits))
}
h := hits[0]
if h.ID != "3578" || h.Title != "2026-07-28-S1021-acme-rechnung.pdf" || h.Kind != File {
t.Errorf("entry = %+v", h.Entry)
}
if h.ParentID != "634" || !reflect.DeepEqual(h.Path, []string{"634"}) {
t.Errorf("path = %v parent = %q", h.Path, h.ParentID)
}
if h.Provider != "es-text" {
t.Errorf("provider = %q", h.Provider)
}
if h.Highlight != "Begleitzettel … S1021 …" {
t.Errorf("highlight = %q, want tags stripped", h.Highlight)
}
}
func TestNewESTextIndexDefaults(t *testing.T) {
if _, err := NewESTextIndex(ESTextConfig{}); err == nil {
t.Error("empty URL: want error")
}
x, err := NewESTextIndex(ESTextConfig{URL: "http://es:9200/"})
if err != nil {
t.Fatalf("NewESTextIndex: %v", err)
}
if x.Index() != defaultESTextIndex {
t.Errorf("index = %q, want %q", x.Index(), defaultESTextIndex)
}
if x.cfg.URL != "http://es:9200" {
t.Errorf("url = %q, want trimmed", x.cfg.URL)
}
if x.Name() != "es-text" {
t.Errorf("Name() = %q", x.Name())
}
}
func TestESTextConfigFromEnvIndexDefault(t *testing.T) {
t.Setenv("ONLYOFFICE_ES_URL", "http://es:9200/")
t.Setenv("ONLYOFFICE_ES_TEXT_INDEX", "")
cfg := ESTextConfigFromEnv()
if cfg.Index != defaultESTextIndex {
t.Errorf("index = %q, want %q", cfg.Index, defaultESTextIndex)
}
}
func TestTextIndexerIndexEntries(t *testing.T) {
store := &textFakeStore{
files: map[string][]byte{"1": []byte("PDFBYTES")},
}
idx := &textFakeIndex{}
ix := NewTextIndexer(store, idx)
ix.Extractor = textFakeExtractor{prefix: "TEXT "}
res, err := ix.IndexEntries(context.Background(), []Entry{
{ID: "1", Title: "Rechnung.PDF", ParentID: "649", Kind: File},
{ID: "2", Title: "Tabelle.xlsx", ParentID: "649", Kind: File},
{ID: "3", Title: "Unterordner", Kind: Folder},
}, IndexOptions{})
if err != nil {
t.Fatalf("IndexEntries: %v", err)
}
if res.Scanned != 3 || res.Indexed != 1 || res.Skipped != 2 || res.Failed != 0 {
t.Errorf("result = %+v, want scanned=3 indexed=1 skipped=2 failed=0", res)
}
if len(idx.docs) != 1 {
t.Fatalf("indexed docs = %d, want 1", len(idx.docs))
}
got := idx.docs[0]
want := TextDoc{ID: "1", Title: "Rechnung.PDF", FolderID: "649", Ext: "pdf", Content: "TEXT PDFBYTES"}
if !reflect.DeepEqual(got, want) {
t.Errorf("doc = %+v, want %+v", got, want)
}
}
func TestTextIndexerRecordsExtractionFailure(t *testing.T) {
store := &textFakeStore{files: map[string][]byte{"1": []byte("x")}}
idx := &textFakeIndex{}
ix := NewTextIndexer(store, idx)
ix.Extractor = textFailingExtractor{}
res, err := ix.IndexEntries(context.Background(), []Entry{{ID: "1", Title: "a.pdf", Kind: File}}, IndexOptions{})
if err != nil {
t.Fatalf("IndexEntries: %v", err)
}
if res.Indexed != 0 || res.Failed != 1 || len(res.Errors) != 1 {
t.Errorf("result = %+v, want one failure recorded", res)
}
}
func TestTextIndexerPlanFolder(t *testing.T) {
store := &textFakeStore{dirs: map[string][]Entry{
"root": {
{ID: "10", Title: "a.pdf", Kind: File},
{ID: "11", Title: "sub", Kind: Folder},
},
"11": {
{ID: "12", Title: "b.PDF", Kind: File},
{ID: "13", Title: "c.xlsx", Kind: File},
},
}}
ix := NewTextIndexer(store, &textFakeIndex{})
flat, err := ix.PlanFolder(context.Background(), "root", IndexOptions{})
if err != nil {
t.Fatalf("PlanFolder: %v", err)
}
if len(flat) != 1 || flat[0].ID != "10" {
t.Errorf("flat plan = %+v, want only a.pdf", flat)
}
deep, err := ix.PlanFolder(context.Background(), "root", IndexOptions{Recursive: true, Limit: 10})
if err != nil {
t.Fatalf("PlanFolder recursive: %v", err)
}
if len(deep) != 2 {
t.Errorf("recursive plan = %d entries, want 2", len(deep))
}
}
// --- fakes -----------------------------------------------------------------
type textFakeStore struct {
dirs map[string][]Entry
files map[string][]byte
stat map[string]Entry
}
func (f *textFakeStore) Name() string { return "fake" }
func (f *textFakeStore) List(_ context.Context, parentID string) ([]Entry, error) {
return f.dirs[parentID], nil
}
func (f *textFakeStore) Stat(_ context.Context, id string) (Entry, error) {
if e, ok := f.stat[id]; ok {
return e, nil
}
return Entry{}, fmt.Errorf("not found: %s", id)
}
func (f *textFakeStore) Download(_ context.Context, id string, w io.Writer) (int64, error) {
b, ok := f.files[id]
if !ok {
return 0, fmt.Errorf("no bytes for %s", id)
}
n, err := w.Write(b)
return int64(n), err
}
func (f *textFakeStore) CreateFolder(context.Context, string, string) (Entry, error) {
return Entry{}, nil
}
func (f *textFakeStore) Upload(context.Context, string, string, io.Reader) (Entry, error) {
return Entry{}, nil
}
func (f *textFakeStore) Move(context.Context, []string, string) error { return nil }
func (f *textFakeStore) Copy(context.Context, []string, string) error { return nil }
func (f *textFakeStore) Rename(context.Context, string, string) error { return nil }
func (f *textFakeStore) Delete(context.Context, []string) error { return nil }
type textFakeIndex struct{ docs []TextDoc }
func (f *textFakeIndex) Put(_ context.Context, docs []TextDoc) error {
f.docs = append(f.docs, docs...)
return nil
}
func (f *textFakeIndex) Delete(context.Context, []string) error { return nil }
func (f *textFakeIndex) Search(context.Context, SearchQuery) ([]SearchHit, error) { return nil, nil }
func (f *textFakeIndex) Name() string { return "fake" }
type textFakeExtractor struct{ prefix string }
func (f textFakeExtractor) Extract(path, _, _ string, _ int) (string, error) {
b, err := os.ReadFile(path)
if err != nil {
return "", err
}
return f.prefix + string(b), nil
}
type textFailingExtractor struct{}
func (textFailingExtractor) Extract(string, string, string, int) (string, error) {
return "", fmt.Errorf("boom")
}
+249
View File
@@ -0,0 +1,249 @@
package onlyoffice
// Single file client (epic #34, F4 #38). FileClient composes the registered
// FileStore and Searcher backends and picks one per operation: REST/DAV for
// writes, PostgreSQL (when registered) for fast reads, Elasticsearch for name
// and content search. Client.Files returns the facade; it also implements
// FileStore, so existing callers keep compiling.
import (
"context"
"errors"
"io"
"strings"
)
// ProviderES is the composed Elasticsearch searcher. The SQL store owns
// ProviderPG/ProviderMySQL (filestore_pg.go); the facade references ProviderPG in
// readOrder.
const ProviderES = "elasticsearch"
var (
errNoReadBackend = errors.New("onlyoffice: no file backend registered for reads")
errNoWriteBackend = errors.New("onlyoffice: no file backend registered for writes")
errNoSearcher = errors.New("onlyoffice: no search backend registered (set ONLYOFFICE_ES_URL)")
)
// FileClient is the single entry point for file operations. It holds the
// registered backends and the order in which each operation tries them.
type FileClient struct {
stores map[string]FileStore
searchers map[string]Searcher
readOrder []string
writeOrder []string
searchOrder []string
}
// newFileClient builds the facade over the built-in REST and DAV stores. The
// Elasticsearch searcher is registered when ONLYOFFICE_ES_URL is set; the
// missing-credential case is left to Search so read-only commands still work.
func (c *Client) newFileClient() *FileClient {
f := &FileClient{
stores: map[string]FileStore{
ProviderREST: &restStore{c: c},
ProviderDAV: &davStore{c: c},
},
searchers: map[string]Searcher{},
readOrder: []string{ProviderPG, ProviderMySQL, ProviderREST, ProviderDAV},
writeOrder: []string{ProviderREST, ProviderDAV},
searchOrder: []string{ProviderES},
}
if cfg := ESConfigFromEnv(); cfg.URL != "" {
if es, err := NewESSearcher(cfg); err == nil {
f.searchers[ProviderES] = es
}
}
return f
}
// RegisterStore adds or replaces a named backend (for example the PostgreSQL
// read store). The name is matched case-insensitively.
func (f *FileClient) RegisterStore(name string, s FileStore) {
if f == nil || s == nil {
return
}
name = normalizeProvider(name)
if name == "" {
return
}
if f.stores == nil {
f.stores = map[string]FileStore{}
}
f.stores[name] = s
}
// RegisterSearcher adds or replaces a named search backend.
func (f *FileClient) RegisterSearcher(name string, s Searcher) {
if f == nil || s == nil {
return
}
name = normalizeProvider(name)
if name == "" {
return
}
if f.searchers == nil {
f.searchers = map[string]Searcher{}
}
f.searchers[name] = s
}
// Read returns the preferred backend for reads: the SQL store (PostgreSQL or
// MySQL) when registered, then REST, then WebDAV.
func (f *FileClient) Read() FileStore { return f.firstStore(f.readOrder) }
// Write returns the preferred backend for writes: REST, then WebDAV.
func (f *FileClient) Write() FileStore { return f.firstStore(f.writeOrder) }
// Search returns the preferred name/content searcher (Elasticsearch), or an
// error when no search backend is configured.
func (f *FileClient) Search() (Searcher, error) {
if f == nil {
return nil, errNoSearcher
}
for _, name := range f.searchOrder {
if s := f.searchers[normalizeProvider(name)]; s != nil {
return s, nil
}
}
return nil, errNoSearcher
}
// firstStore returns the first registered store in the order.
func (f *FileClient) firstStore(order []string) FileStore {
if f == nil {
return nil
}
for _, name := range order {
if s := f.stores[normalizeProvider(name)]; s != nil {
return s
}
}
return nil
}
// orderedStores returns the registered stores in the order.
func (f *FileClient) orderedStores(order []string) []FileStore {
if f == nil {
return nil
}
out := make([]FileStore, 0, len(order))
for _, name := range order {
if s := f.stores[normalizeProvider(name)]; s != nil {
out = append(out, s)
}
}
return out
}
func normalizeProvider(name string) string {
return strings.ToLower(strings.TrimSpace(name))
}
// Name implements FileStore and reports the preferred read backend.
func (f *FileClient) Name() string {
if s := f.Read(); s != nil {
return s.Name()
}
return ""
}
// List reads from the preferred backend, falling back to the next read backend
// only on a transient error (429/502/503/504).
func (f *FileClient) List(ctx context.Context, parentID string) ([]Entry, error) {
return fallbackRead(ctx, f.orderedStores(f.readOrder), func(s FileStore) ([]Entry, error) {
return s.List(ctx, parentID)
})
}
// Stat reads from the preferred backend, with the same transient fallback.
func (f *FileClient) Stat(ctx context.Context, id string) (Entry, error) {
return fallbackRead(ctx, f.orderedStores(f.readOrder), func(s FileStore) (Entry, error) {
return s.Stat(ctx, id)
})
}
// Download streams file bytes. It does not fall back: a failed attempt may have
// already written partial bytes into w, so a second backend would append.
func (f *FileClient) Download(ctx context.Context, id string, w io.Writer) (int64, error) {
s := f.Read()
if s == nil {
return 0, errNoReadBackend
}
return s.Download(ctx, id, w)
}
// CreateFolder writes to the preferred write backend.
func (f *FileClient) CreateFolder(ctx context.Context, parentID, title string) (Entry, error) {
s := f.Write()
if s == nil {
return Entry{}, errNoWriteBackend
}
return s.CreateFolder(ctx, parentID, title)
}
// Upload writes to the preferred write backend.
func (f *FileClient) Upload(ctx context.Context, parentID, title string, r io.Reader) (Entry, error) {
s := f.Write()
if s == nil {
return Entry{}, errNoWriteBackend
}
return s.Upload(ctx, parentID, title, r)
}
// Move writes to the preferred write backend.
func (f *FileClient) Move(ctx context.Context, ids []string, parentID string) error {
s := f.Write()
if s == nil {
return errNoWriteBackend
}
return s.Move(ctx, ids, parentID)
}
// Copy writes to the preferred write backend.
func (f *FileClient) Copy(ctx context.Context, ids []string, parentID string) error {
s := f.Write()
if s == nil {
return errNoWriteBackend
}
return s.Copy(ctx, ids, parentID)
}
// Rename writes to the preferred write backend.
func (f *FileClient) Rename(ctx context.Context, id, title string) error {
s := f.Write()
if s == nil {
return errNoWriteBackend
}
return s.Rename(ctx, id, title)
}
// Delete writes to the preferred write backend.
func (f *FileClient) Delete(ctx context.Context, ids []string) error {
s := f.Write()
if s == nil {
return errNoWriteBackend
}
return s.Delete(ctx, ids)
}
// fallbackRead runs op against each store in order, moving on only when the
// error is transient. Non-transient errors (not found, forbidden) are final.
func fallbackRead[T any](ctx context.Context, stores []FileStore, op func(FileStore) (T, error)) (T, error) {
var zero T
if len(stores) == 0 {
return zero, errNoReadBackend
}
var err error
for i, s := range stores {
var v T
v, err = op(s)
if err == nil {
return v, nil
}
if i == len(stores)-1 || !Transient(err) {
return zero, err
}
}
return zero, err
}
+131
View File
@@ -0,0 +1,131 @@
//go:build integration
package onlyoffice
import (
"bytes"
"context"
"strconv"
"testing"
"time"
)
// TestIntegrationFacadeCRUD drives the whole operation set through the composed
// facade c.Files(): folder create, upload, stat, list, rename, move, copy,
// delete. Writes must go to REST (the default writeOrder), reads follow
// readOrder (REST when no SQL backend is registered) and every returned Entry
// must report its provider. Destructive — throwaway project, cleaned up.
func TestIntegrationFacadeCRUD(t *testing.T) {
c := liveClient(t)
t.Cleanup(func() { cleanupTestProjects(t, c) })
ctx := context.Background()
suffix := time.Now().UTC().Format("20060102-150405")
project, err := c.CreateProject(NewProjectRequest{
Title: testProjectPrefix + "facade-" + suffix,
Description: "go-onlyoffice facade CRUD integration",
})
if err != nil {
t.Fatalf("CreateProject: %v", err)
}
if project.ID == nil {
t.Fatal("created project without id")
}
root, err := c.projectFolderID(ctx, strconv.Itoa(*project.ID))
if err != nil {
t.Fatalf("projectFolderID: %v", err)
}
f := c.Files()
if got := f.Write().Name(); got != ProviderREST {
t.Fatalf("Write().Name() = %q, want %q", got, ProviderREST)
}
if got := f.Read().Name(); got != ProviderREST {
t.Fatalf("Read().Name() = %q, want %q (no SQL backend registered)", got, ProviderREST)
}
src, err := f.CreateFolder(ctx, root, "facade-src-"+suffix)
if err != nil {
t.Fatalf("CreateFolder src: %v", err)
}
if src.Kind != Folder || src.ID == "" {
t.Fatalf("created src folder: %+v", src)
}
if src.Provider != ProviderREST {
t.Fatalf("CreateFolder provider = %q, want %q", src.Provider, ProviderREST)
}
dst, err := f.CreateFolder(ctx, root, "facade-dst-"+suffix)
if err != nil {
t.Fatalf("CreateFolder dst: %v", err)
}
if dst.Provider != ProviderREST {
t.Fatalf("CreateFolder dst provider = %q, want %q", dst.Provider, ProviderREST)
}
t.Cleanup(func() {
if err := c.DeleteDavItems(ctx, []string{src.ID, dst.ID}, nil); err != nil {
t.Logf("cleanup folders: %v", err)
}
})
content := []byte("facade crud " + suffix + "\n")
up, err := f.Upload(ctx, src.ID, "facade-doc-"+suffix+".txt", bytes.NewReader(content))
if err != nil {
t.Fatalf("Upload: %v", err)
}
if up.Kind != File || up.ID == "" {
t.Fatalf("uploaded entry: %+v", up)
}
if up.Provider != ProviderREST {
t.Fatalf("Upload provider = %q, want %q (write order REST first)", up.Provider, ProviderREST)
}
if !waitEntry(ctx, f, src.ID, up.ID, 15*time.Second) {
t.Fatalf("uploaded %s not listed in src", up.ID)
}
st, err := f.Stat(ctx, up.ID)
if err != nil {
t.Fatalf("Stat: %v", err)
}
if st.ID != up.ID || st.Kind != File {
t.Fatalf("Stat = %+v", st)
}
if st.Provider != ProviderREST {
t.Fatalf("Stat provider = %q, want %q (read order REST)", st.Provider, ProviderREST)
}
list, err := f.List(ctx, src.ID)
if err != nil {
t.Fatalf("List: %v", err)
}
if e := entryByID(list, up.ID); e == nil {
t.Fatalf("uploaded %s not in List(src)", up.ID)
} else if e.Provider != ProviderREST {
t.Fatalf("List provider = %q, want %q", e.Provider, ProviderREST)
}
renamed := "facade-renamed-" + suffix + ".txt"
renameEventually(t, ctx, f, up.ID, renamed)
moveEventually(t, ctx, f, up.ID, dst.ID)
if !waitEntry(ctx, f, dst.ID, up.ID, 20*time.Second) {
t.Fatalf("moved file %s not in dst", up.ID)
}
copied := copyEventually(t, ctx, f, up.ID, src.ID, 20*time.Second)
if copied == nil {
t.Fatalf("no copy found in src after Copy")
}
if copied.Provider != ProviderREST {
t.Fatalf("Copy provider = %q, want %q", copied.Provider, ProviderREST)
}
if err := f.Delete(ctx, []string{up.ID, copied.ID}); err != nil {
t.Fatalf("Delete: %v", err)
}
if !waitNoEntry(ctx, f, dst.ID, up.ID, 20*time.Second) {
t.Fatalf("file %s still present in dst after delete", up.ID)
}
if !waitNoEntry(ctx, f, src.ID, copied.ID, 20*time.Second) {
t.Fatalf("copy %s still present in src after delete", copied.ID)
}
}
+298
View File
@@ -0,0 +1,298 @@
package onlyoffice
import (
"context"
"errors"
"fmt"
"io"
"strings"
"testing"
)
// fakeStore is a FileStore test double; it records which backend served a call
// and returns a canned result or error.
type fakeStore struct {
name string
entries []Entry
err error
calls *[]string
}
func (f *fakeStore) record(op string) {
if f.calls != nil {
*f.calls = append(*f.calls, op+":"+f.name)
}
}
func (f *fakeStore) Name() string { return f.name }
func (f *fakeStore) List(_ context.Context, _ string) ([]Entry, error) {
f.record("list")
if f.err != nil {
return nil, f.err
}
return f.entries, nil
}
func (f *fakeStore) Stat(_ context.Context, id string) (Entry, error) {
f.record("stat")
if f.err != nil {
return Entry{}, f.err
}
return Entry{ID: id, Title: "t-" + f.name, Provider: f.name}, nil
}
func (f *fakeStore) CreateFolder(_ context.Context, _, title string) (Entry, error) {
f.record("mkdir")
if f.err != nil {
return Entry{}, f.err
}
return Entry{ID: "new", Title: title, Provider: f.name}, nil
}
func (f *fakeStore) Upload(_ context.Context, _, title string, _ io.Reader) (Entry, error) {
f.record("upload")
return Entry{ID: "up", Title: title, Provider: f.name}, f.err
}
func (f *fakeStore) Download(_ context.Context, _ string, _ io.Writer) (int64, error) {
f.record("download")
return 0, f.err
}
func (f *fakeStore) Move(_ context.Context, _ []string, _ string) error {
f.record("move")
return f.err
}
func (f *fakeStore) Copy(_ context.Context, _ []string, _ string) error {
f.record("copy")
return f.err
}
func (f *fakeStore) Rename(_ context.Context, _, _ string) error {
f.record("rename")
return f.err
}
func (f *fakeStore) Delete(_ context.Context, _ []string) error {
f.record("delete")
return f.err
}
type fakeSearcher struct{ name string }
func (s *fakeSearcher) Name() string { return s.name }
func (s *fakeSearcher) Search(_ context.Context, _ SearchQuery) ([]SearchHit, error) {
return []SearchHit{{Entry: Entry{Title: s.name}}}, nil
}
func newFacadeTestClient(stores map[string]FileStore, read, write []string) *FileClient {
return &FileClient{
stores: stores,
searchers: map[string]Searcher{},
readOrder: read,
writeOrder: write,
}
}
// TestFileClientIsFileStore guarantees the facade can stand in for the
// interface anywhere a plain FileStore is expected.
func TestFileClientIsFileStore(t *testing.T) {
var _ FileStore = (*FileClient)(nil)
}
func TestClientFilesPrefersRESTForReadsAndWrites(t *testing.T) {
c := NewClient(Credentials{})
f := c.Files()
if got := f.Read().Name(); got != ProviderREST {
t.Errorf("Read().Name() = %q, want %q", got, ProviderREST)
}
if got := f.Write().Name(); got != ProviderREST {
t.Errorf("Write().Name() = %q, want %q", got, ProviderREST)
}
if got := f.Name(); got != ProviderREST {
t.Errorf("Name() = %q, want %q", got, ProviderREST)
}
}
func TestFileClientPostgresTakesReadPriority(t *testing.T) {
pg := &fakeStore{name: ProviderPG}
f := newFacadeTestClient(
map[string]FileStore{ProviderREST: &fakeStore{name: ProviderREST}, ProviderPG: pg},
[]string{ProviderPG, ProviderREST},
[]string{ProviderREST},
)
if got := f.Read().Name(); got != ProviderPG {
t.Errorf("Read().Name() = %q, want %q", got, ProviderPG)
}
if got := f.Write().Name(); got != ProviderREST {
t.Errorf("Write().Name() = %q, want %q (PG is read-only)", got, ProviderREST)
}
}
func TestFileClientRegisterStoreNormalizesName(t *testing.T) {
pg := &fakeStore{name: "pg"}
f := newFacadeTestClient(map[string]FileStore{}, []string{ProviderPG}, nil)
f.RegisterStore(" POSTGRES ", pg)
if got := f.Read(); got != pg {
t.Fatalf("Read() = %v, want registered postgres store", got)
}
f.RegisterStore("", pg)
f.RegisterStore("pg", nil)
}
func TestFileClientReadFallsBackOnlyOnTransient(t *testing.T) {
var calls []string
primary := &fakeStore{name: "primary", err: fmt.Errorf("onlyoffice: list: 503 unavailable"), calls: &calls}
secondary := &fakeStore{name: "secondary", entries: []Entry{{ID: "1"}}, calls: &calls}
f := newFacadeTestClient(
map[string]FileStore{"primary": primary, "secondary": secondary},
[]string{"primary", "secondary"},
nil,
)
got, err := f.List(context.Background(), "root")
if err != nil {
t.Fatalf("List: %v", err)
}
if len(got) != 1 || got[0].ID != "1" {
t.Fatalf("List() = %+v, want secondary entry", got)
}
want := []string{"list:primary", "list:secondary"}
if fmt.Sprint(calls) != fmt.Sprint(want) {
t.Fatalf("call order = %v, want %v", calls, want)
}
}
func TestFileClientReadStopsOnPermanentError(t *testing.T) {
var calls []string
primary := &fakeStore{name: "primary", err: errors.New("onlyoffice: not found"), calls: &calls}
secondary := &fakeStore{name: "secondary", entries: []Entry{{ID: "1"}}, calls: &calls}
f := newFacadeTestClient(
map[string]FileStore{"primary": primary, "secondary": secondary},
[]string{"primary", "secondary"},
nil,
)
if _, err := f.List(context.Background(), "root"); err == nil {
t.Fatal("expected permanent error to be returned")
}
if len(calls) != 1 || calls[0] != "list:primary" {
t.Fatalf("secondary backend must not run on a permanent error: %v", calls)
}
}
func TestFileClientWriteUsesWriteBackend(t *testing.T) {
var calls []string
rest := &fakeStore{name: ProviderREST, calls: &calls}
dav := &fakeStore{name: ProviderDAV, calls: &calls}
f := newFacadeTestClient(
map[string]FileStore{ProviderREST: rest, ProviderDAV: dav},
[]string{ProviderREST},
[]string{ProviderREST, ProviderDAV},
)
if _, err := f.CreateFolder(context.Background(), "p", "t"); err != nil {
t.Fatalf("CreateFolder: %v", err)
}
if _, err := f.Upload(context.Background(), "p", "t", nil); err != nil {
t.Fatalf("Upload: %v", err)
}
if len(calls) != 2 || calls[0] != "mkdir:rest" || calls[1] != "upload:rest" {
t.Fatalf("write calls = %v, want REST", calls)
}
}
func TestFileClientWriteWithoutBackend(t *testing.T) {
f := newFacadeTestClient(map[string]FileStore{}, nil, nil)
if err := f.Delete(context.Background(), []string{"1"}); !errors.Is(err, errNoWriteBackend) {
t.Fatalf("Delete err = %v, want errNoWriteBackend", err)
}
if _, err := f.List(context.Background(), "root"); !errors.Is(err, errNoReadBackend) {
t.Fatalf("List err = %v, want errNoReadBackend", err)
}
}
func TestNewFileClientReadOrderIncludesSQL(t *testing.T) {
f := NewClient(Credentials{}).newFileClient()
want := []string{ProviderPG, ProviderMySQL, ProviderREST, ProviderDAV}
if fmt.Sprint(f.readOrder) != fmt.Sprint(want) {
t.Fatalf("readOrder = %v, want %v", f.readOrder, want)
}
}
// TestClientFileStoreSQLRoutingWithoutDSN checks that the SQL backend names are
// recognised and never yield nil: without a DSN the returned store surfaces the
// open error on use.
func TestClientFileStoreSQLRoutingWithoutDSN(t *testing.T) {
t.Setenv("ONLYOFFICE_DSN", "")
t.Setenv("ONLYOFFICE_PG_HOST", "")
c := NewClient(Credentials{})
for _, name := range []string{"pg", "sql", ProviderPG, ProviderMySQL} {
s := c.FileStore(name)
if s == nil {
t.Fatalf("FileStore(%q) = nil", name)
}
if _, err := s.Stat(context.Background(), "1"); err == nil {
t.Errorf("FileStore(%q).Stat without DSN: want error", name)
}
}
}
func TestFileClientMySQLStoreIsPreferredForReads(t *testing.T) {
mysql := &fakeStore{name: ProviderMySQL}
f := newFacadeTestClient(
map[string]FileStore{ProviderREST: &fakeStore{name: ProviderREST}, ProviderMySQL: mysql},
[]string{ProviderPG, ProviderMySQL, ProviderREST},
[]string{ProviderREST},
)
if got := f.Read().Name(); got != ProviderMySQL {
t.Errorf("Read().Name() = %q, want %q", got, ProviderMySQL)
}
}
// TestFileClientWriteToReadOnlyStore guarantees the facade surfaces ErrReadOnly
// when the configured write backend is the read-only SQL store.
func TestFileClientWriteToReadOnlyStore(t *testing.T) {
pg := &pgStore{driver: ProviderPG}
f := newFacadeTestClient(
map[string]FileStore{ProviderREST: &fakeStore{name: ProviderREST}, ProviderPG: pg},
[]string{ProviderPG, ProviderREST},
[]string{ProviderPG, ProviderREST},
)
ctx := context.Background()
if _, err := f.CreateFolder(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
t.Errorf("CreateFolder err = %v", err)
}
if _, err := f.Upload(ctx, "1", "x", strings.NewReader("x")); !errors.Is(err, ErrReadOnly) {
t.Errorf("Upload err = %v", err)
}
if err := f.Move(ctx, []string{"1"}, "2"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Move err = %v", err)
}
if err := f.Copy(ctx, []string{"1"}, "2"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Copy err = %v", err)
}
if err := f.Rename(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Rename err = %v", err)
}
if err := f.Delete(ctx, []string{"1"}); !errors.Is(err, ErrReadOnly) {
t.Errorf("Delete err = %v", err)
}
}
func TestFileClientSearchSelection(t *testing.T) {
f := &FileClient{searchers: map[string]Searcher{}, searchOrder: []string{ProviderES}}
_, err := f.Search()
if err == nil || !strings.Contains(err.Error(), "ONLYOFFICE_ES_URL") {
t.Fatalf("Search without backend = %v, want ONLYOFFICE_ES_URL hint", err)
}
es := &fakeSearcher{name: "fake-es"}
f.RegisterSearcher(ProviderES, es)
got, err := f.Search()
if err != nil {
t.Fatalf("Search: %v", err)
}
if got.Name() != "fake-es" {
t.Fatalf("searcher = %q, want fake-es", got.Name())
}
}
+525
View File
@@ -0,0 +1,525 @@
package onlyoffice
// Read-only SQL backend of the unified file client (epic #34, F2 #36).
//
// The goal is to read files and folders straight from the Community Server
// database, without the REST layer. Research on the live portal (VM
// `onlyoffice-v2`) showed the server runs on **MySQL 8.0** (`files_file`,
// `files_folder`, `files_folder_tree`, tenant `tenants_tenants`), not
// PostgreSQL — see docs/community-server-db.md. The store below therefore
// speaks `database/sql` and selects its driver from the DSN, so it works
// against the live MySQL today and against PostgreSQL if the portal is ever
// migrated. Every query is a SELECT; the write methods of FileStore return
// ErrReadOnly.
//
// Downloads follow the portal's S3/MinIO object layout through the shared
// MinIO helper in storage_fallback.go — no HTTP file endpoint is used.
import (
"context"
"database/sql"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"path/filepath"
"strconv"
"strings"
"time"
"github.com/go-sql-driver/mysql"
_ "github.com/jackc/pgx/v5/stdlib"
)
// Provider names for the SQL backend. ProviderPG is the value Name reports for
// a PostgreSQL connection and ProviderMySQL for MySQL.
const (
ProviderPG = "postgres"
ProviderMySQL = "mysql"
)
// ErrReadOnly is returned by every FileStore write method of the SQL backend.
var ErrReadOnly = errors.New("onlyoffice: sql file store is read-only")
const (
pgConnectTimeout = 10 * time.Second
pgSearchLimit = 50
pgSearchMaxLimit = 500
)
// PGConfig configures the read-only SQL store. DSN is a driver DSN:
// `user:pass@tcp(host:port)/onlyoffice?parseTime=true` for MySQL or a
// `postgres://` / libpq keyword string for PostgreSQL. Driver, when set,
// forces the engine ("postgres" or "mysql"); otherwise it is detected from the
// DSN. Tenant filters rows (empty means all tenants).
type PGConfig struct {
DSN string
Driver string
Tenant string
}
// PGConfigFromEnv reads ONLYOFFICE_DSN (or the ONLYOFFICE_PG_* parts),
// ONLYOFFICE_PG_DRIVER and the tenant from ONLYOFFICE_PG_TENANT /
// ONLYOFFICE_TENANT. The library never loads dotfiles — the CLI does that.
func PGConfigFromEnv() PGConfig {
dsn := strings.TrimSpace(os.Getenv("ONLYOFFICE_DSN"))
if dsn == "" {
dsn = pgDSNFromParts()
}
return PGConfig{
DSN: dsn,
Driver: strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_DRIVER")),
Tenant: firstNonEmpty(os.Getenv("ONLYOFFICE_PG_TENANT"), os.Getenv("ONLYOFFICE_TENANT")),
}
}
// pgDSNFromParts builds a libpq keyword DSN from ONLYOFFICE_PG_* variables.
// It returns "" unless a host is set, which keeps the MySQL path (ONLYOFFICE_DSN)
// the default.
func pgDSNFromParts() string {
host := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_HOST"))
if host == "" {
return ""
}
port := firstNonEmpty(os.Getenv("ONLYOFFICE_PG_PORT"), "5432")
dbname := firstNonEmpty(os.Getenv("ONLYOFFICE_PG_DBNAME"), "onlyoffice")
sslmode := firstNonEmpty(os.Getenv("ONLYOFFICE_PG_SSLMODE"), "disable")
return fmt.Sprintf("host=%s port=%s user=%s password=%s dbname=%s sslmode=%s",
host, port, os.Getenv("ONLYOFFICE_PG_USER"), os.Getenv("ONLYOFFICE_PG_PASSWORD"), dbname, sslmode)
}
// SQLFileStore opens the read-only SQL store from the environment
// (PGConfigFromEnv: ONLYOFFICE_DSN or the ONLYOFFICE_PG_* parts). It is the
// error-aware counterpart of Client.FileStore("pg"/"sql"), which returns an
// errStore when the open fails. The caller owns the returned store and should
// close it (the concrete type has a Close method).
func (c *Client) SQLFileStore() (FileStore, error) {
return NewPGStore(PGConfigFromEnv())
}
// pgStore is a read-only FileStore/Searcher over the Community Server database.
type pgStore struct {
db *sql.DB
driver string
tenantID int64
hasTenant bool
http *http.Client
}
var (
_ FileStore = (*pgStore)(nil)
_ Searcher = (*pgStore)(nil)
)
// NewPGStore opens the database and verifies connectivity. It never writes.
func NewPGStore(cfg PGConfig) (*pgStore, error) {
dsn := strings.TrimSpace(cfg.DSN)
if dsn == "" {
return nil, fmt.Errorf("onlyoffice: sql file store: empty DSN (set ONLYOFFICE_DSN)")
}
driver := pgDriver(dsn, cfg.Driver)
dsn, err := normalizeSQLDSN(driver, dsn)
if err != nil {
return nil, err
}
db, err := sql.Open(sqlDriverName(driver), dsn)
if err != nil {
return nil, fmt.Errorf("onlyoffice: sql file store: open %s: %w", driver, err)
}
ctx, cancel := context.WithTimeout(context.Background(), pgConnectTimeout)
defer cancel()
if err := db.PingContext(ctx); err != nil {
db.Close()
return nil, fmt.Errorf("onlyoffice: sql file store: ping %s: %w", driver, err)
}
s := &pgStore{db: db, driver: driver, http: &http.Client{}}
if t := strings.TrimSpace(cfg.Tenant); t != "" {
n, err := strconv.ParseInt(t, 10, 64)
if err != nil {
db.Close()
return nil, fmt.Errorf("onlyoffice: sql file store: non-numeric tenant %q", t)
}
s.tenantID, s.hasTenant = n, true
}
return s, nil
}
// Close releases the database handle.
func (s *pgStore) Close() error { return s.db.Close() }
// Name implements FileStore and Searcher.
func (s *pgStore) Name() string { return s.driver }
// pgDriver resolves the engine: the explicit value wins, otherwise the DSN
// shape decides. A leading postgres:// scheme or a libpq keyword DSN (which
// always carries '=') selects PostgreSQL; anything else is MySQL.
func pgDriver(dsn, explicit string) string {
switch strings.ToLower(strings.TrimSpace(explicit)) {
case ProviderPG, "pg", "postgresql", "pgx":
return ProviderPG
case ProviderMySQL, "mariadb":
return ProviderMySQL
}
l := strings.ToLower(strings.TrimSpace(dsn))
switch {
case strings.HasPrefix(l, "postgres://"), strings.HasPrefix(l, "postgresql://"):
return ProviderPG
case strings.HasPrefix(l, "mysql://"), strings.Contains(l, "@tcp("), strings.Contains(l, "@unix("):
return ProviderMySQL
case strings.Contains(l, "="):
return ProviderPG
default:
return ProviderMySQL
}
}
// sqlDriverName maps the engine to its registered database/sql driver.
func sqlDriverName(driver string) string {
if driver == ProviderPG {
return "pgx"
}
return "mysql"
}
// normalizeSQLDSN converts a mysql:// URL to the go-sql-driver form and forces
// parseTime so datetime columns scan into time.Time. PostgreSQL DSNs pass
// through untouched.
func normalizeSQLDSN(driver, dsn string) (string, error) {
if driver != ProviderMySQL {
return dsn, nil
}
if strings.HasPrefix(strings.ToLower(dsn), "mysql://") {
converted, err := mysqlDSNFromURL(dsn)
if err != nil {
return "", err
}
dsn = converted
}
cfg, err := mysql.ParseDSN(dsn)
if err != nil {
return "", fmt.Errorf("onlyoffice: sql file store: parse mysql DSN: %w", err)
}
cfg.ParseTime = true
return cfg.FormatDSN(), nil
}
// mysqlDSNFromURL turns mysql://user:pass@host:port/db into the driver DSN.
func mysqlDSNFromURL(raw string) (string, error) {
u, err := url.Parse(raw)
if err != nil || u.Host == "" {
return "", fmt.Errorf("onlyoffice: sql file store: bad mysql URL %q", raw)
}
user := ""
if u.User != nil {
user = u.User.Username()
if p, ok := u.User.Password(); ok {
user += ":" + p
}
}
q := u.Query()
q.Set("parseTime", "true")
return fmt.Sprintf("%s@tcp(%s)/%s?%s", user, u.Host, strings.TrimPrefix(u.Path, "/"), q.Encode()), nil
}
// rebind rewrites '?' placeholders to PostgreSQL's $1..$n. MySQL keeps them.
func rebind(query, driver string) string {
if driver != ProviderPG {
return query
}
var b strings.Builder
b.Grow(len(query) + 8)
n := 0
for _, r := range query {
if r == '?' {
n++
b.WriteByte('$')
b.WriteString(strconv.Itoa(n))
continue
}
b.WriteRune(r)
}
return b.String()
}
// List returns the folders and files directly below parentID, folders first.
func (s *pgStore) List(ctx context.Context, parentID string) ([]Entry, error) {
pid, err := parseEntryID(parentID)
if err != nil {
return nil, err
}
folders, err := s.queryFolders(ctx, "parent_id = ?", pid)
if err != nil {
return nil, err
}
files, err := s.queryFiles(ctx, "folder_id = ? AND current_version = 1", pid)
if err != nil {
return nil, err
}
out := make([]Entry, 0, len(folders)+len(files))
for _, f := range folders {
out = append(out, folderRowToEntry(f, s.Name()))
}
for _, f := range files {
out = append(out, fileRowToEntry(f, s.Name()))
}
return out, nil
}
// Stat resolves a folder or file id to an Entry. Folders win when both id
// spaces overlap (they never do on a real portal, but the lookup is cheap).
func (s *pgStore) Stat(ctx context.Context, id string) (Entry, error) {
n, err := parseEntryID(id)
if err != nil {
return Entry{}, err
}
folders, err := s.queryFolders(ctx, "id = ?", n)
if err != nil {
return Entry{}, err
}
if len(folders) > 0 {
return folderRowToEntry(folders[0], s.Name()), nil
}
files, err := s.queryFiles(ctx, "id = ? AND current_version = 1", n)
if err != nil {
return Entry{}, err
}
if len(files) == 0 {
return Entry{}, fmt.Errorf("onlyoffice: sql file store: id %s not found", id)
}
return fileRowToEntry(files[0], s.Name()), nil
}
// Download streams the file's current version from the portal's S3/MinIO store.
// The object key is reconstructed from the file id and version; the parent
// folder id is not part of the key.
func (s *pgStore) Download(ctx context.Context, id string, w io.Writer) (int64, error) {
n, err := parseEntryID(id)
if err != nil {
return 0, err
}
files, err := s.queryFiles(ctx, "id = ? AND current_version = 1", n)
if err != nil {
return 0, err
}
if len(files) == 0 {
return 0, fmt.Errorf("onlyoffice: sql file store: file %s not found", id)
}
f := files[0]
key := csObjectKey(s.tenantID, f.id, f.version, filepath.Ext(f.title))
return downloadMinioObject(ctx, s.http, key, w)
}
// CreateFolder is unavailable: the SQL backend is read-only.
func (s *pgStore) CreateFolder(context.Context, string, string) (Entry, error) {
return Entry{}, ErrReadOnly
}
// Upload is unavailable: the SQL backend is read-only.
func (s *pgStore) Upload(context.Context, string, string, io.Reader) (Entry, error) {
return Entry{}, ErrReadOnly
}
// Move is unavailable: the SQL backend is read-only.
func (s *pgStore) Move(context.Context, []string, string) error { return ErrReadOnly }
// Copy is unavailable: the SQL backend is read-only.
func (s *pgStore) Copy(context.Context, []string, string) error { return ErrReadOnly }
// Rename is unavailable: the SQL backend is read-only.
func (s *pgStore) Rename(context.Context, string, string) error { return ErrReadOnly }
// Delete is unavailable: the SQL backend is read-only.
func (s *pgStore) Delete(context.Context, []string) error { return ErrReadOnly }
// Search matches file titles by substring. Content search lives in the
// Elasticsearch backend; q.InContent is ignored here.
func (s *pgStore) Search(ctx context.Context, q SearchQuery) ([]SearchHit, error) {
text := strings.TrimSpace(q.Text)
if text == "" {
return nil, fmt.Errorf("onlyoffice: empty search query")
}
limit := q.Limit
if limit <= 0 {
limit = pgSearchLimit
}
if limit > pgSearchMaxLimit {
limit = pgSearchMaxLimit
}
where := "title LIKE ? AND current_version = 1"
args := []any{"%" + text + "%"}
if s.hasTenant {
where += " AND tenant_id = ?"
args = append(args, s.tenantID)
}
if fid := strings.TrimSpace(q.FolderID); fid != "" {
n, err := parseEntryID(fid)
if err != nil {
return nil, err
}
where += " AND folder_id = ?"
args = append(args, n)
}
for _, ext := range normalizeExtensions(q.Extensions) {
where += " AND LOWER(title) LIKE ?"
args = append(args, "%."+ext)
}
query := rebind(`SELECT id, folder_id, title, content_length, version, create_on, modified_on
FROM files_file WHERE `+where+` ORDER BY modified_on DESC, id DESC LIMIT ?`, s.driver)
args = append(args, limit)
rows, err := s.db.QueryContext(ctx, query, args...)
if err != nil {
return nil, fmt.Errorf("onlyoffice: sql search: %w", err)
}
defer rows.Close()
var hits []SearchHit
for rows.Next() {
f, err := scanFileRow(rows)
if err != nil {
return nil, err
}
hits = append(hits, SearchHit{Entry: fileRowToEntry(f, s.Name())})
}
return hits, rows.Err()
}
// queryFolders runs a folder SELECT with the tenant filter applied.
func (s *pgStore) queryFolders(ctx context.Context, where string, arg any) ([]pgFolderRow, error) {
args := []any{arg}
if s.hasTenant {
where += " AND tenant_id = ?"
args = append(args, s.tenantID)
}
query := rebind(`SELECT id, parent_id, title, create_on, modified_on
FROM files_folder WHERE `+where+` ORDER BY title, id`, s.driver)
rows, err := s.db.QueryContext(ctx, query, args...)
if err != nil {
return nil, fmt.Errorf("onlyoffice: sql list folders: %w", err)
}
defer rows.Close()
var out []pgFolderRow
for rows.Next() {
var r pgFolderRow
if err := rows.Scan(&r.id, &r.parentID, &r.title, &r.created, &r.modified); err != nil {
return nil, fmt.Errorf("onlyoffice: sql folder row: %w", err)
}
out = append(out, r)
}
return out, rows.Err()
}
// queryFiles runs a file SELECT for the current version with the tenant filter.
func (s *pgStore) queryFiles(ctx context.Context, where string, arg any) ([]pgFileRow, error) {
args := []any{arg}
if s.hasTenant {
where += " AND tenant_id = ?"
args = append(args, s.tenantID)
}
query := rebind(`SELECT id, folder_id, title, content_length, version, create_on, modified_on
FROM files_file WHERE `+where+` ORDER BY title, id`, s.driver)
rows, err := s.db.QueryContext(ctx, query, args...)
if err != nil {
return nil, fmt.Errorf("onlyoffice: sql list files: %w", err)
}
defer rows.Close()
var out []pgFileRow
for rows.Next() {
f, err := scanFileRow(rows)
if err != nil {
return nil, err
}
out = append(out, f)
}
return out, rows.Err()
}
// pgFileRow is one current files_file row.
type pgFileRow struct {
id int64
folderID int64
title string
size int64
version int
created time.Time
modified time.Time
}
// pgFolderRow is one files_folder row.
type pgFolderRow struct {
id int64
parentID int64
title string
created time.Time
modified time.Time
}
// scanFileRow reads the canonical file column order.
func scanFileRow(rows *sql.Rows) (pgFileRow, error) {
var f pgFileRow
if err := rows.Scan(&f.id, &f.folderID, &f.title, &f.size, &f.version, &f.created, &f.modified); err != nil {
return f, fmt.Errorf("onlyoffice: sql file row: %w", err)
}
return f, nil
}
// fileRowToEntry maps a files_file row to the canonical model.
func fileRowToEntry(f pgFileRow, provider string) Entry {
return Entry{
ID: strconv.FormatInt(f.id, 10),
ParentID: strconv.FormatInt(f.folderID, 10),
Title: f.title,
Kind: File,
Size: f.size,
MIME: mimeForTitle(f.title, ""),
Created: f.created.UTC(),
Modified: f.modified.UTC(),
Version: f.version,
Provider: provider,
}
}
// folderRowToEntry maps a files_folder row to the canonical model.
func folderRowToEntry(f pgFolderRow, provider string) Entry {
return Entry{
ID: strconv.FormatInt(f.id, 10),
ParentID: strconv.FormatInt(f.parentID, 10),
Title: f.title,
Kind: Folder,
Created: f.created.UTC(),
Modified: f.modified.UTC(),
Provider: provider,
}
}
// csObjectKey reconstructs the object key the portal's S3 consumer uses:
//
// 00/00/<tenant>/files/folder_<shard>/file_<id>/v<version>/content.<ext>
//
// The shard is the next thousand above the file id (file 3727 -> folder_4000),
// NOT the parent folder id — verified live against the MinIO bucket.
func csObjectKey(tenant int64, fileID int64, version int, ext string) string {
shard := (fileID/1000 + 1) * 1000
ext = strings.TrimPrefix(strings.ToLower(strings.TrimSpace(ext)), ".")
if ext == "" {
ext = "bin"
}
if version < 1 {
version = 1
}
if tenant <= 0 {
tenant = 1
}
return fmt.Sprintf("00/00/%02d/files/folder_%d/file_%d/v%d/content.%s", tenant, shard, fileID, version, ext)
}
// parseEntryID parses a numeric OnlyOffice id or returns a store error.
func parseEntryID(id string) (int64, error) {
n, err := strconv.ParseInt(strings.TrimSpace(id), 10, 64)
if err != nil {
return 0, fmt.Errorf("onlyoffice: sql file store: non-numeric id %q", id)
}
return n, nil
}
+224
View File
@@ -0,0 +1,224 @@
//go:build integration
package onlyoffice
import (
"bytes"
"context"
"errors"
"os"
"strings"
"testing"
"time"
)
// TestIntegrationPGStore exercises the read-only SQL backend against the live
// Community Server database and cross-checks list/stat/download with the REST
// FileStore. It needs ONLYOFFICE_DSN plus the usual ONLYOFFICE_URL/USER/PASS;
// ONLYOFFICE_PG_TEST_FILE_ID / ONLYOFFICE_PG_TEST_FOLDER_ID pick a real file
// (a file reachable over REST too). Download streams from MinIO, so it also
// needs MINIO_ACCESS_KEY/MINIO_SECRET_KEY.
//
// The live Community Server runs on MySQL; PostgreSQL is supported by the same
// code path when the DSN says so.
func TestIntegrationPGStore(t *testing.T) {
cfg := PGConfigFromEnv()
if strings.TrimSpace(cfg.DSN) == "" {
t.Skip("ONLYOFFICE_DSN not set — skipping SQL store integration test")
}
store, err := NewPGStore(cfg)
if err != nil {
t.Fatalf("NewPGStore: %v", err)
}
t.Cleanup(func() { _ = store.Close() })
t.Logf("sql store backend: %s", store.Name())
ctx := context.Background()
if err := testPGStoreReadOnly(ctx, store); err != nil {
t.Fatal(err)
}
fileID := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_TEST_FILE_ID"))
folderID := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_TEST_FOLDER_ID"))
if fileID == "" || folderID == "" {
t.Skip("ONLYOFFICE_PG_TEST_FILE_ID / ONLYOFFICE_PG_TEST_FOLDER_ID not set — skipping live comparison")
}
c := liveClient(t)
rest := c.Files()
dbFile, err := store.Stat(ctx, fileID)
if err != nil {
t.Fatalf("sql Stat(%s): %v", fileID, err)
}
restFile, err := rest.Stat(ctx, fileID)
if err != nil {
t.Fatalf("rest Stat(%s): %v", fileID, err)
}
if dbFile.Kind != File {
t.Errorf("sql kind = %v, want file", dbFile.Kind)
}
if dbFile.ID != restFile.ID || dbFile.Title != restFile.Title || dbFile.ParentID != restFile.ParentID {
t.Errorf("stat mismatch sql=%+v rest=%+v", dbFile, restFile)
}
// GetFile omits contentLength, so size is only comparable when REST has it.
if restFile.Size > 0 && dbFile.Size != restFile.Size {
t.Errorf("size sql=%d rest=%d", dbFile.Size, restFile.Size)
}
if d := dbFile.Modified.Sub(restFile.Modified); d > 2*time.Minute || d < -2*time.Minute {
t.Errorf("modified sql=%v rest=%v", dbFile.Modified, restFile.Modified)
}
list, err := store.List(ctx, folderID)
if err != nil {
t.Fatalf("sql List(%s): %v", folderID, err)
}
if entryByID(list, fileID) == nil {
t.Errorf("file %s not in sql List(%s)", fileID, folderID)
}
// Every file the REST layer can see in the folder must be in the SQL list
// (the SQL store sees more, so only assert this direction).
restList, err := rest.List(ctx, folderID)
if err != nil {
t.Fatalf("rest List(%s): %v", folderID, err)
}
dbIDs := make(map[string]bool, len(list))
for _, e := range list {
dbIDs[e.ID] = true
}
for _, e := range restList {
if e.Kind == File && !dbIDs[e.ID] {
t.Errorf("rest file %s (%q) missing from sql list", e.ID, e.Title)
}
}
// Download reads the object store, not the database, so it only runs with
// the MinIO credentials configured (the portal's S3 layout). Without them
// the DSN-only assertions above still prove the SQL reads.
if os.Getenv("MINIO_ACCESS_KEY") == "" || os.Getenv("MINIO_SECRET_KEY") == "" {
t.Log("MINIO_ACCESS_KEY/MINIO_SECRET_KEY not set — SQL download cross-check skipped")
return
}
var buf bytes.Buffer
n, err := store.Download(ctx, fileID, &buf)
if err != nil {
t.Fatalf("sql Download(%s): %v", fileID, err)
}
if n == 0 || n != dbFile.Size {
t.Errorf("sql Download = %d bytes, stat says %d", n, dbFile.Size)
}
var restBuf bytes.Buffer
rn, err := rest.Download(ctx, fileID, &restBuf)
if err != nil {
t.Fatalf("rest Download(%s): %v", fileID, err)
}
if rn != n || !bytes.Equal(restBuf.Bytes(), buf.Bytes()) {
t.Errorf("download mismatch sql=%d rest=%d bytes", n, rn)
}
}
// TestIntegrationSQLFacade proves that reads are served by the SQL store when
// it is part of the file client, not by REST. It needs ONLYOFFICE_DSN plus
// ONLYOFFICE_PG_TEST_FILE_ID / ONLYOFFICE_PG_TEST_FOLDER_ID and the usual REST
// credentials (for the cross-check). Every Entry served by SQL carries
// Provider "mysql"; REST entries carry "rest", so the provider is the proof of
// which backend answered.
func TestIntegrationSQLFacade(t *testing.T) {
cfg := PGConfigFromEnv()
if strings.TrimSpace(cfg.DSN) == "" {
t.Skip("ONLYOFFICE_DSN not set — skipping SQL facade integration test")
}
fileID := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_TEST_FILE_ID"))
folderID := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_TEST_FOLDER_ID"))
if fileID == "" || folderID == "" {
t.Skip("ONLYOFFICE_PG_TEST_FILE_ID / ONLYOFFICE_PG_TEST_FOLDER_ID not set — skipping SQL facade integration test")
}
c := liveClient(t)
ctx := context.Background()
sqlStore, err := c.SQLFileStore()
if err != nil {
t.Fatalf("SQLFileStore: %v", err)
}
if closer, ok := sqlStore.(interface{ Close() error }); ok {
t.Cleanup(func() { _ = closer.Close() })
}
if sqlStore.Name() == ProviderREST {
t.Fatalf("SQLFileStore returned REST")
}
// Direct constructor: Client.FileStore("pg") must not be REST either.
direct := c.FileStore("pg")
if direct == nil || direct.Name() == ProviderREST {
t.Fatalf("FileStore(\"pg\") = %v, want SQL backend", direct)
}
if closer, ok := direct.(interface{ Close() error }); ok {
t.Cleanup(func() { _ = closer.Close() })
}
f := c.Files()
f.RegisterStore(ProviderPG, sqlStore)
if got := f.Read().Name(); got != sqlStore.Name() {
t.Fatalf("facade read backend = %q, want %q (SQL)", got, sqlStore.Name())
}
got, err := f.Stat(ctx, fileID)
if err != nil {
t.Fatalf("facade Stat(%s): %v", fileID, err)
}
if got.Provider != ProviderMySQL {
t.Errorf("facade Stat provider = %q, want %q (SQL, not REST)", got.Provider, ProviderMySQL)
}
want, err := c.FileStore(ProviderREST).Stat(ctx, fileID)
if err != nil {
t.Fatalf("rest Stat(%s): %v", fileID, err)
}
if got.ID != want.ID || got.Title != want.Title || got.ParentID != want.ParentID {
t.Errorf("facade SQL stat %+v != REST %+v", got, want)
}
list, err := f.List(ctx, folderID)
if err != nil {
t.Fatalf("facade List(%s): %v", folderID, err)
}
entry := entryByID(list, fileID)
if entry == nil {
t.Fatalf("file %s not in facade List(%s)", fileID, folderID)
}
if entry.Provider != ProviderMySQL {
t.Errorf("facade List provider = %q, want %q", entry.Provider, ProviderMySQL)
}
d, err := direct.Stat(ctx, fileID)
if err != nil {
t.Fatalf("FileStore(\"pg\").Stat(%s): %v", fileID, err)
}
if d.Provider != ProviderMySQL {
t.Errorf("FileStore(\"pg\") provider = %q, want %q", d.Provider, ProviderMySQL)
}
}
// testPGStoreReadOnly asserts that every write method returns ErrReadOnly.
func testPGStoreReadOnly(ctx context.Context, s *pgStore) error {
if _, err := s.CreateFolder(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
return errors.New("CreateFolder did not return ErrReadOnly")
}
if _, err := s.Upload(ctx, "1", "x", strings.NewReader("x")); !errors.Is(err, ErrReadOnly) {
return errors.New("Upload did not return ErrReadOnly")
}
if err := s.Move(ctx, []string{"1"}, "2"); !errors.Is(err, ErrReadOnly) {
return errors.New("Move did not return ErrReadOnly")
}
if err := s.Copy(ctx, []string{"1"}, "2"); !errors.Is(err, ErrReadOnly) {
return errors.New("Copy did not return ErrReadOnly")
}
if err := s.Rename(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
return errors.New("Rename did not return ErrReadOnly")
}
if err := s.Delete(ctx, []string{"1"}); !errors.Is(err, ErrReadOnly) {
return errors.New("Delete did not return ErrReadOnly")
}
return nil
}
+165
View File
@@ -0,0 +1,165 @@
package onlyoffice
import (
"context"
"errors"
"testing"
"time"
)
func TestRebind(t *testing.T) {
mysqlQuery := "SELECT id FROM files_file WHERE folder_id = ? AND title = ? LIMIT ?"
if got := rebind(mysqlQuery, ProviderMySQL); got != mysqlQuery {
t.Errorf("mysql query changed: %q", got)
}
want := "SELECT id FROM files_file WHERE folder_id = $1 AND title = $2 LIMIT $3"
if got := rebind(mysqlQuery, ProviderPG); got != want {
t.Errorf("rebind = %q, want %q", got, want)
}
}
func TestCSPObjectKey(t *testing.T) {
cases := []struct {
tenant int64
fileID int64
version int
ext string
want string
}{
{1, 2, 1, ".docx", "00/00/01/files/folder_1000/file_2/v1/content.docx"},
{1, 999, 1, ".pdf", "00/00/01/files/folder_1000/file_999/v1/content.pdf"},
{1, 1000, 1, ".xlsx", "00/00/01/files/folder_2000/file_1000/v1/content.xlsx"},
{1, 3727, 1, ".pdf", "00/00/01/files/folder_4000/file_3727/v1/content.pdf"},
{1, 22484, 1, ".PDF", "00/00/01/files/folder_23000/file_22484/v1/content.pdf"},
{1, 4, 6, "xlsx", "00/00/01/files/folder_1000/file_4/v6/content.xlsx"},
{0, 7, 0, "", "00/00/01/files/folder_1000/file_7/v1/content.bin"},
{2, 11, 3, ".doc", "00/00/02/files/folder_1000/file_11/v3/content.doc"},
}
for _, tc := range cases {
if got := csObjectKey(tc.tenant, tc.fileID, tc.version, tc.ext); got != tc.want {
t.Errorf("csObjectKey(%d,%d,%d,%q) = %q, want %q", tc.tenant, tc.fileID, tc.version, tc.ext, got, tc.want)
}
}
}
func TestPGDriverDetection(t *testing.T) {
cases := []struct {
dsn, explicit, want string
}{
{"postgres://u:p@h:5432/onlyoffice", "", ProviderPG},
{"postgresql://u:p@h/db", "", ProviderPG},
{"host=h user=u password=p dbname=onlyoffice sslmode=disable", "", ProviderPG},
{"root:secret@tcp(127.0.0.1:3306)/onlyoffice?parseTime=true", "", ProviderMySQL},
{"mysql://root:secret@127.0.0.1:3306/onlyoffice", "", ProviderMySQL},
{"root:secret@tcp(h:3306)/db", "postgres", ProviderPG},
{"postgres://u:p@h/db", "mysql", ProviderMySQL},
}
for _, tc := range cases {
if got := pgDriver(tc.dsn, tc.explicit); got != tc.want {
t.Errorf("pgDriver(%q, %q) = %q, want %q", tc.dsn, tc.explicit, got, tc.want)
}
}
}
func TestNormalizeSQLDSNMySQL(t *testing.T) {
got, err := normalizeSQLDSN(ProviderMySQL, "mysql://root:secret@127.0.0.1:3306/onlyoffice")
if err != nil {
t.Fatalf("normalizeSQLDSN: %v", err)
}
want := "root:secret@tcp(127.0.0.1:3306)/onlyoffice?parseTime=true"
if got != want {
t.Errorf("normalize = %q, want %q", got, want)
}
// A driver DSN keeps parseTime and gains it when missing.
got, err = normalizeSQLDSN(ProviderMySQL, "root:secret@tcp(127.0.0.1:3306)/onlyoffice")
if err != nil {
t.Fatalf("normalizeSQLDSN: %v", err)
}
if got != want {
t.Errorf("normalize = %q, want %q", got, want)
}
}
func TestFileRowToEntry(t *testing.T) {
created := time.Date(2026, 9, 12, 18, 0, 37, 0, time.UTC)
modified := time.Date(2026, 9, 13, 13, 50, 36, 0, time.UTC)
e := fileRowToEntry(pgFileRow{
id: 22484, folderID: 649, title: "Rechnung.pdf",
size: 123433, version: 2, created: created, modified: modified,
}, ProviderMySQL)
if e.ID != "22484" || e.ParentID != "649" {
t.Errorf("ids = %q/%q", e.ID, e.ParentID)
}
if e.Title != "Rechnung.pdf" || e.Kind != File {
t.Errorf("title/kind = %q/%v", e.Title, e.Kind)
}
if e.Size != 123433 || e.Version != 2 {
t.Errorf("size/version = %d/%d", e.Size, e.Version)
}
if e.MIME != "application/pdf" {
t.Errorf("mime = %q", e.MIME)
}
if !e.Created.Equal(created) || !e.Modified.Equal(modified) {
t.Errorf("times = %v/%v", e.Created, e.Modified)
}
if e.Provider != ProviderMySQL {
t.Errorf("provider = %q", e.Provider)
}
}
func TestFolderRowToEntry(t *testing.T) {
modified := time.Date(2026, 8, 1, 10, 30, 0, 0, time.UTC)
e := folderRowToEntry(pgFolderRow{id: 649, parentID: 647, title: "2025", modified: modified}, ProviderMySQL)
if e.ID != "649" || e.ParentID != "647" || e.Title != "2025" {
t.Errorf("folder = %+v", e)
}
if e.Kind != Folder {
t.Errorf("kind = %v, want folder", e.Kind)
}
if e.Size != 0 || e.MIME != "" {
t.Errorf("folder size/mime = %d/%q", e.Size, e.MIME)
}
if !e.Modified.Equal(modified) {
t.Errorf("modified = %v", e.Modified)
}
}
func TestPGStoreWriteMethodsReadOnly(t *testing.T) {
s := &pgStore{driver: ProviderPG}
ctx := context.Background()
if _, err := s.CreateFolder(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
t.Errorf("CreateFolder err = %v", err)
}
if _, err := s.Upload(ctx, "1", "x", nil); !errors.Is(err, ErrReadOnly) {
t.Errorf("Upload err = %v", err)
}
if err := s.Move(ctx, nil, "1"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Move err = %v", err)
}
if err := s.Copy(ctx, nil, "1"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Copy err = %v", err)
}
if err := s.Rename(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Rename err = %v", err)
}
if err := s.Delete(ctx, nil); !errors.Is(err, ErrReadOnly) {
t.Errorf("Delete err = %v", err)
}
}
func TestPGStoreName(t *testing.T) {
if got := (&pgStore{driver: ProviderPG}).Name(); got != ProviderPG {
t.Errorf("Name = %q, want %q", got, ProviderPG)
}
if got := (&pgStore{driver: ProviderMySQL}).Name(); got != ProviderMySQL {
t.Errorf("Name = %q, want %q", got, ProviderMySQL)
}
}
func TestPGStoreStatRejectsNonNumeric(t *testing.T) {
s := &pgStore{driver: ProviderPG}
if _, err := s.Stat(context.Background(), "not-a-number"); err == nil {
t.Error("Stat accepted a non-numeric id")
}
}
+67 -41
View File
@@ -7,12 +7,9 @@ package onlyoffice
import ( import (
"context" "context"
"encoding/json" "encoding/json"
"fmt"
"io" "io"
"os" "os"
"path/filepath" "path/filepath"
"strconv"
"strings"
) )
// restStore is a FileStore over the REST Documents API. // restStore is a FileStore over the REST Documents API.
@@ -39,11 +36,25 @@ func (s *restStore) List(ctx context.Context, parentID string) ([]Entry, error)
return out, err return out, err
} }
// Stat returns file metadata. The REST adapter resolves files only; folders // Stat returns file or folder metadata. Folders are resolved through the
// are listed by their parent (use List). // listing endpoint (their own id appears as the listing's Current); other ids
// fall back to the file metadata API.
func (s *restStore) Stat(ctx context.Context, id string) (Entry, error) { func (s *restStore) Stat(ctx context.Context, id string) (Entry, error) {
return s.stat(ctx, id)
}
// stat resolves a single id to a folder or file Entry.
func (s *restStore) stat(ctx context.Context, id string) (Entry, error) {
var out Entry var out Entry
err := retryStoreOp(ctx, func() error { err := retryStoreOp(ctx, func() error {
if l, err := s.c.ListDavFolder(ctx, id); err == nil {
if l != nil && l.Current.ID != "" && l.Current.ID == id {
out = DavFolderToEntry(l.Current, ProviderREST)
return nil
}
} else if Transient(err) {
return err
}
f, err := s.c.GetFile(ctx, id) f, err := s.c.GetFile(ctx, id)
if err != nil { if err != nil {
return err return err
@@ -124,56 +135,84 @@ func (s *restStore) Download(ctx context.Context, id string, w io.Writer) (int64
return n, err return n, err
} }
// Move moves file ids into parentID. The REST MoveFiles endpoint handles files // Move moves folders and/or files into parentID. Ids are classified through
// only; folder moves are not exposed by this adapter. // stat so folder moves use folderIds and file moves use fileIds on the shared
// fileops/move endpoint.
func (s *restStore) Move(ctx context.Context, ids []string, parentID string) error { func (s *restStore) Move(ctx context.Context, ids []string, parentID string) error {
dest, err := strconv.Atoi(strings.TrimSpace(parentID)) folders, files, err := s.split(ctx, ids)
if err != nil {
return fmt.Errorf("onlyoffice: rest store: move: non-numeric destination folder id %q", parentID)
}
fileIDs, err := numericIDs(ids)
if err != nil { if err != nil {
return err return err
} }
return retryStoreOp(ctx, func() error { if len(folders) == 0 && len(files) == 0 {
_, err := s.c.MoveFiles(ctx, dest, fileIDs)
return err
})
}
// Copy copies file ids into parentID. files.go has no copy method, so the
// shared REST fileops copy endpoint (CopyDavItems) is used.
func (s *restStore) Copy(ctx context.Context, ids []string, parentID string) error {
if len(ids) == 0 {
return nil return nil
} }
return retryStoreOp(ctx, func() error { return retryStoreOp(ctx, func() error {
return s.c.CopyDavItems(ctx, nil, ids, parentID) return s.c.MoveDavItems(ctx, folders, files, parentID)
}) })
} }
// Rename sets a new title (including extension) for a file. // Copy copies folders and/or files into parentID. files.go has no copy method,
// so the shared REST fileops copy endpoint (CopyDavItems) is used.
func (s *restStore) Copy(ctx context.Context, ids []string, parentID string) error {
folders, files, err := s.split(ctx, ids)
if err != nil {
return err
}
if len(folders) == 0 && len(files) == 0 {
return nil
}
return retryStoreOp(ctx, func() error {
return s.c.CopyDavItems(ctx, folders, files, parentID)
})
}
// Rename sets a new title (including extension) for a file or folder.
func (s *restStore) Rename(ctx context.Context, id, title string) error { func (s *restStore) Rename(ctx context.Context, id, title string) error {
e, err := s.stat(ctx, id)
if err != nil {
return err
}
if e.Kind == Folder {
return retryStoreOp(ctx, func() error {
return s.c.RenameDavFolder(ctx, id, title)
})
}
return retryStoreOp(ctx, func() error { return retryStoreOp(ctx, func() error {
_, err := s.c.RenameFile(ctx, id, title) _, err := s.c.RenameFile(ctx, id, title)
return err return err
}) })
} }
// Delete permanently deletes file ids. // Delete permanently deletes folders and/or files.
func (s *restStore) Delete(ctx context.Context, ids []string) error { func (s *restStore) Delete(ctx context.Context, ids []string) error {
fileIDs, err := numericIDs(ids) folders, files, err := s.split(ctx, ids)
if err != nil { if err != nil {
return err return err
} }
if len(fileIDs) == 0 { if len(folders) == 0 && len(files) == 0 {
return nil return nil
} }
return retryStoreOp(ctx, func() error { return retryStoreOp(ctx, func() error {
return s.c.DeleteFiles(ctx, fileIDs) return s.c.DeleteDavItems(ctx, folders, files)
}) })
} }
// split classifies ids into folder and file id lists.
func (s *restStore) split(ctx context.Context, ids []string) (folders, files []string, err error) {
for _, id := range ids {
e, err := s.stat(ctx, id)
if err != nil {
return nil, nil, err
}
if e.Kind == Folder {
folders = append(folders, id)
} else {
files = append(files, id)
}
}
return folders, files, nil
}
// entriesFromFolderMap converts a ListFolder response map into canonical // entriesFromFolderMap converts a ListFolder response map into canonical
// entries, reusing the DavFile/DavFolder decoders for robust size handling. // entries, reusing the DavFile/DavFolder decoders for robust size handling.
func entriesFromFolderMap(m map[string]any, provider string) ([]Entry, error) { func entriesFromFolderMap(m map[string]any, provider string) ([]Entry, error) {
@@ -215,16 +254,3 @@ func folderEntryFromMap(m map[string]any, parentID, provider string) (Entry, err
e = DavFolderToEntry(f, provider) e = DavFolderToEntry(f, provider)
return e, nil return e, nil
} }
// numericIDs parses Documents numeric ids from strings.
func numericIDs(ids []string) ([]int, error) {
out := make([]int, 0, len(ids))
for _, id := range ids {
n, err := strconv.Atoi(strings.TrimSpace(id))
if err != nil {
return nil, fmt.Errorf("onlyoffice: rest store: non-numeric id %q", id)
}
out = append(out, n)
}
return out, nil
}
+337
View File
@@ -0,0 +1,337 @@
package onlyoffice
// Text extraction pipeline for the own full-text index (epic #34, F6 #42).
//
// TextIndexer downloads stored documents, extracts text through docpipe
// (pdftotext; OCR for scans) and writes the result to a TextIndex. For PDFs it
// also indexes the text of embedded attachments (pdfdetach), so a scan filed
// as an attachment is searchable too. It is the write side of ESTextIndex and
// never touches the OnlyOffice server's own ES index.
import (
"context"
"fmt"
"os"
"path/filepath"
"strings"
"sync"
"github.com/eslider/go-onlyoffice/internal/docpipe"
)
// defaultTextIndexExts are the formats extracted by default. The OnlyOffice
// index already covers docx/xlsx/pptx; F6 adds PDF.
var defaultTextIndexExts = []string{"pdf"}
const (
defaultTextIndexWorkers = 3
defaultTextIndexLang = "deu+eng"
maxTextIndexErrors = 20
)
// TextExtractor turns a local file into indexable plain text. The default uses
// docpipe (pdftotext + OCR); tests inject a fake to stay offline.
type TextExtractor interface {
Extract(path, workDir, lang string, minChars int) (string, error)
}
// docpipeExtractor is the production TextExtractor.
type docpipeExtractor struct{ tools docpipe.Tools }
// Extract renders the file as Markdown, OCRing PDFs/images with a weak text
// layer first and appending the text of embedded PDF attachments
// (docpipe.ToMarkdownWithAttachments).
func (d docpipeExtractor) Extract(path, workDir, lang string, minChars int) (string, error) {
text, err := d.tools.ToMarkdownWithAttachments(path, workDir, lang, minChars)
if err != nil {
return "", err
}
return strings.TrimSpace(text), nil
}
// IndexOptions controls a TextIndexer run.
type IndexOptions struct {
Recursive bool // IndexFolder: descend into subfolders
Extensions []string // empty = defaultTextIndexExts (pdf)
Limit int // max files to index, 0 = all
Lang string // OCR language(s), default deu+eng
MinChars int // OCR threshold, default docpipe.DefaultMinTextChars
Workers int // parallel downloads/extractions, default 3
}
// IndexResult summarises a run.
type IndexResult struct {
Scanned int
Indexed int
Skipped int
Failed int
Errors []string
}
// TextIndexer wires a FileStore, a TextIndex and an extractor together.
type TextIndexer struct {
Store FileStore
Index TextIndex
Extractor TextExtractor // nil = local docpipe tools
WorkDir string // temp dir for downloads/extraction
}
// NewTextIndexer returns a TextIndexer over the given store and index.
func NewTextIndexer(store FileStore, index TextIndex) *TextIndexer {
return &TextIndexer{Store: store, Index: index}
}
// textIndexEnsurer is implemented by indexes that can be created up front.
type textIndexEnsurer interface {
Ensure(ctx context.Context) error
}
// Ensure creates the backing index when the TextIndex supports it.
func (ix *TextIndexer) Ensure(ctx context.Context) error {
if e, ok := ix.Index.(textIndexEnsurer); ok {
return e.Ensure(ctx)
}
return nil
}
// IndexFiles stats the given file ids and indexes them.
func (ix *TextIndexer) IndexFiles(ctx context.Context, ids []string, opts IndexOptions) (IndexResult, error) {
entries := make([]Entry, 0, len(ids))
for _, id := range ids {
e, err := ix.Store.Stat(ctx, id)
if err != nil {
return IndexResult{}, fmt.Errorf("onlyoffice: stat %s: %w", id, err)
}
entries = append(entries, e)
}
return ix.IndexEntries(ctx, entries, opts)
}
// IndexFolder lists a folder and indexes every matching file.
func (ix *TextIndexer) IndexFolder(ctx context.Context, folderID string, opts IndexOptions) (IndexResult, error) {
entries, err := ix.collect(ctx, folderID, opts.Recursive)
if err != nil {
return IndexResult{}, err
}
return ix.IndexEntries(ctx, entries, opts)
}
// PlanFolder lists the files IndexFolder would process, without downloading or
// extracting anything.
func (ix *TextIndexer) PlanFolder(ctx context.Context, folderID string, opts IndexOptions) ([]Entry, error) {
entries, err := ix.collect(ctx, folderID, opts.Recursive)
if err != nil {
return nil, err
}
return selectEntries(entries, opts), nil
}
// PlanFiles stats the ids and returns those that would be indexed.
func (ix *TextIndexer) PlanFiles(ctx context.Context, ids []string, opts IndexOptions) ([]Entry, error) {
entries := make([]Entry, 0, len(ids))
for _, id := range ids {
e, err := ix.Store.Stat(ctx, id)
if err != nil {
return nil, fmt.Errorf("onlyoffice: stat %s: %w", id, err)
}
entries = append(entries, e)
}
return selectEntries(entries, opts), nil
}
// IndexEntries extracts and indexes the given files (folders are ignored).
func (ix *TextIndexer) IndexEntries(ctx context.Context, entries []Entry, opts IndexOptions) (IndexResult, error) {
opts = opts.withDefaults()
var res IndexResult
work := selectEntries(entries, opts)
res.Scanned = len(entries)
res.Skipped = len(entries) - len(work)
if len(work) == 0 {
return res, nil
}
workers := opts.Workers
if workers > len(work) {
workers = len(work)
}
if workers < 1 {
workers = 1
}
type outcome struct {
doc TextDoc
err error
}
jobs := make(chan Entry)
results := make(chan outcome, workers)
var wg sync.WaitGroup
for i := 0; i < workers; i++ {
wg.Add(1)
go func() {
defer wg.Done()
for e := range jobs {
if err := ctx.Err(); err != nil {
results <- outcome{err: err}
continue
}
doc, err := ix.indexOne(ctx, e, opts)
results <- outcome{doc: doc, err: err}
}
}()
}
go func() {
defer close(jobs)
for _, e := range work {
select {
case jobs <- e:
case <-ctx.Done():
return
}
}
}()
go func() {
wg.Wait()
close(results)
}()
var docs []TextDoc
for r := range results {
if r.err != nil {
res.Failed++
if len(res.Errors) < maxTextIndexErrors {
res.Errors = append(res.Errors, r.err.Error())
}
continue
}
docs = append(docs, r.doc)
}
if err := ctx.Err(); err != nil {
return res, err
}
if len(docs) > 0 {
if err := ix.Index.Put(ctx, docs); err != nil {
return res, fmt.Errorf("onlyoffice: index %d docs: %w", len(docs), err)
}
res.Indexed = len(docs)
}
return res, nil
}
// indexOne downloads and extracts a single file.
func (ix *TextIndexer) indexOne(ctx context.Context, e Entry, opts IndexOptions) (TextDoc, error) {
ext := fileExt(e.Title)
dir := ix.WorkDir
if dir == "" {
dir = os.TempDir()
}
if err := os.MkdirAll(dir, 0o755); err != nil {
return TextDoc{}, err
}
tmp, err := os.CreateTemp(dir, "ooidx-*."+ext)
if err != nil {
return TextDoc{}, err
}
tmpPath := tmp.Name()
defer os.Remove(tmpPath)
if _, err := ix.Store.Download(ctx, e.ID, tmp); err != nil {
tmp.Close()
return TextDoc{}, fmt.Errorf("download %s (%s): %w", e.ID, e.Title, err)
}
if err := tmp.Close(); err != nil {
return TextDoc{}, err
}
text, err := ix.extractor().Extract(tmpPath, dir, opts.Lang, opts.MinChars)
if err != nil {
return TextDoc{}, fmt.Errorf("extract %s: %w", e.Title, err)
}
return TextDoc{ID: e.ID, Title: e.Title, FolderID: e.ParentID, Ext: ext, Content: text}, nil
}
// collect lists files under folderID, breadth-first when recursive.
func (ix *TextIndexer) collect(ctx context.Context, folderID string, recursive bool) ([]Entry, error) {
var files []Entry
queue := []string{folderID}
for len(queue) > 0 {
if err := ctx.Err(); err != nil {
return nil, err
}
id := queue[0]
queue = queue[1:]
entries, err := ix.Store.List(ctx, id)
if err != nil {
return nil, fmt.Errorf("onlyoffice: list folder %s: %w", id, err)
}
for _, e := range entries {
if e.Kind == Folder {
if recursive {
queue = append(queue, e.ID)
}
continue
}
if e.ParentID == "" {
e.ParentID = id
}
files = append(files, e)
}
}
return files, nil
}
// selectEntries filters files by extension and applies the limit.
func selectEntries(entries []Entry, opts IndexOptions) []Entry {
allowed := extensionSet(opts.Extensions)
work := make([]Entry, 0, len(entries))
for _, e := range entries {
if e.Kind != File {
continue
}
if opts.Limit > 0 && len(work) >= opts.Limit {
break
}
if !allowed[fileExt(e.Title)] {
continue
}
work = append(work, e)
}
return work
}
// extensionSet normalises the extension allow-list (default: pdf).
func extensionSet(exts []string) map[string]bool {
if len(exts) == 0 {
exts = defaultTextIndexExts
}
set := make(map[string]bool, len(exts))
for _, e := range normalizeExtensions(exts) {
set[e] = true
}
return set
}
// fileExt returns the lower-case extension without the dot.
func fileExt(title string) string {
return strings.ToLower(strings.TrimPrefix(filepath.Ext(strings.TrimSpace(title)), "."))
}
// withDefaults fills zero-valued options.
func (o IndexOptions) withDefaults() IndexOptions {
if o.Workers <= 0 {
o.Workers = defaultTextIndexWorkers
}
if o.MinChars <= 0 {
o.MinChars = docpipe.DefaultMinTextChars
}
if strings.TrimSpace(o.Lang) == "" {
o.Lang = defaultTextIndexLang
}
return o
}
// extractor returns the configured extractor or the local docpipe default.
func (ix *TextIndexer) extractor() TextExtractor {
if ix.Extractor != nil {
return ix.Extractor
}
return docpipeExtractor{tools: docpipe.LookPath()}
}
+8
View File
@@ -13,7 +13,9 @@ require (
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3
github.com/eslider/go-hocr v0.2.2-0.20260827163626-8ff01582b002 github.com/eslider/go-hocr v0.2.2-0.20260827163626-8ff01582b002
github.com/eslider/go-xls/v2 v2.1.0 github.com/eslider/go-xls/v2 v2.1.0
github.com/go-sql-driver/mysql v1.10.1
github.com/google/go-querystring v1.2.0 github.com/google/go-querystring v1.2.0
github.com/jackc/pgx/v5 v5.11.0
github.com/joho/godotenv v1.5.1 github.com/joho/godotenv v1.5.1
github.com/mattn/go-runewidth v0.0.15 github.com/mattn/go-runewidth v0.0.15
github.com/muesli/termenv v0.16.0 github.com/muesli/termenv v0.16.0
@@ -24,6 +26,7 @@ require (
) )
require ( require (
filippo.io/edwards25519 v1.2.0 // indirect
github.com/JohannesKaufmann/dom v0.3.1 // indirect github.com/JohannesKaufmann/dom v0.3.1 // indirect
github.com/alecthomas/chroma/v2 v2.14.0 // indirect github.com/alecthomas/chroma/v2 v2.14.0 // indirect
github.com/atotto/clipboard v0.1.4 // indirect github.com/atotto/clipboard v0.1.4 // indirect
@@ -36,6 +39,10 @@ require (
github.com/google/uuid v1.6.0 // indirect github.com/google/uuid v1.6.0 // indirect
github.com/gorilla/css v1.0.1 // indirect github.com/gorilla/css v1.0.1 // indirect
github.com/inconshreveable/mousetrap v1.1.0 // indirect github.com/inconshreveable/mousetrap v1.1.0 // indirect
github.com/jackc/pgpassfile v1.0.0 // indirect
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
github.com/jackc/puddle/v2 v2.2.2 // indirect
github.com/kr/text v0.2.0 // indirect
github.com/lucasb-eyer/go-colorful v1.4.0 // indirect github.com/lucasb-eyer/go-colorful v1.4.0 // indirect
github.com/mattn/go-isatty v0.0.24 // indirect github.com/mattn/go-isatty v0.0.24 // indirect
github.com/mattn/go-localereader v0.0.1 // indirect github.com/mattn/go-localereader v0.0.1 // indirect
@@ -48,6 +55,7 @@ require (
github.com/richardlehane/mscfb v1.0.7 // indirect github.com/richardlehane/mscfb v1.0.7 // indirect
github.com/richardlehane/msoleps v1.0.6 // indirect github.com/richardlehane/msoleps v1.0.6 // indirect
github.com/rivo/uniseg v0.4.7 // indirect github.com/rivo/uniseg v0.4.7 // indirect
github.com/rogpeppe/go-internal v1.16.0 // indirect
github.com/spf13/pflag v1.0.9 // indirect github.com/spf13/pflag v1.0.9 // indirect
github.com/tiendc/go-deepcopy v1.7.2 // indirect github.com/tiendc/go-deepcopy v1.7.2 // indirect
github.com/xuri/efp v0.0.1 // indirect github.com/xuri/efp v0.0.1 // indirect
+26 -1
View File
@@ -1,3 +1,5 @@
filippo.io/edwards25519 v1.2.0 h1:crnVqOiS4jqYleHd9vaKZ+HKtHfllngJIiOpNpoJsjo=
filippo.io/edwards25519 v1.2.0/go.mod h1:xzAOLCNug/yB62zG1bQ8uziwrIqIuxhctzJT18Q77mc=
github.com/JohannesKaufmann/dom v0.3.1 h1:J16l9JAHWgkFPR3VIPbQ1gvS0cWab6laK1q7PFL3qh0= github.com/JohannesKaufmann/dom v0.3.1 h1:J16l9JAHWgkFPR3VIPbQ1gvS0cWab6laK1q7PFL3qh0=
github.com/JohannesKaufmann/dom v0.3.1/go.mod h1:BZPkf8ZeYrBgABjwJn9iiKt8aiCtkxpHkevms+Yp2DE= github.com/JohannesKaufmann/dom v0.3.1/go.mod h1:BZPkf8ZeYrBgABjwJn9iiKt8aiCtkxpHkevms+Yp2DE=
github.com/JohannesKaufmann/html-to-markdown/v2 v2.5.2 h1:XFJZFWESIWlUEHHjzBuv8RvrtCWnSGlimEX17ysSDb8= github.com/JohannesKaufmann/html-to-markdown/v2 v2.5.2 h1:XFJZFWESIWlUEHHjzBuv8RvrtCWnSGlimEX17ysSDb8=
@@ -35,6 +37,8 @@ github.com/charmbracelet/x/exp/golden v0.0.0-20240715153702-9ba8adf781c4/go.mod
github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81 h1:q2hJAaP1k2wIvVRd/hEHD7lacgqrCPS+k8g1MndzfWY= github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81 h1:q2hJAaP1k2wIvVRd/hEHD7lacgqrCPS+k8g1MndzfWY=
github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81/go.mod h1:YynlIjWYF8myEu6sdkwKIvGQq+cOckRm6So2avqoYAk= github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81/go.mod h1:YynlIjWYF8myEu6sdkwKIvGQq+cOckRm6So2avqoYAk=
github.com/cpuguy83/go-md2man/v2 v2.0.6/go.mod h1:oOW0eioCTA6cOiMLiUPZOpcVxMig6NIQQ7OS05n1F4g= github.com/cpuguy83/go-md2man/v2 v2.0.6/go.mod h1:oOW0eioCTA6cOiMLiUPZOpcVxMig6NIQQ7OS05n1F4g=
github.com/creack/pty v1.1.9/go.mod h1:oKZEueFk5CKHvIhNR5MUki03XCEU+Q6VDXinZuGJ33E=
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c= github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dlclark/regexp2 v1.11.0 h1:G/nrcoOa7ZXlpoa/91N3X7mM3r8eIlMBBJZvsz/mxKI= github.com/dlclark/regexp2 v1.11.0 h1:G/nrcoOa7ZXlpoa/91N3X7mM3r8eIlMBBJZvsz/mxKI=
@@ -47,6 +51,8 @@ github.com/eslider/go-hocr v0.2.2-0.20260827163626-8ff01582b002 h1:LOFxQG4mxvlH7
github.com/eslider/go-hocr v0.2.2-0.20260827163626-8ff01582b002/go.mod h1:fIgfH/E1j3rU8du4X4+7mxTD0GPtPQibTzytgitdJWU= github.com/eslider/go-hocr v0.2.2-0.20260827163626-8ff01582b002/go.mod h1:fIgfH/E1j3rU8du4X4+7mxTD0GPtPQibTzytgitdJWU=
github.com/eslider/go-xls/v2 v2.1.0 h1:HszWKqYQbXxACmAXXWdMsfNl1NDBfGVBnJUPtyUHQ7A= github.com/eslider/go-xls/v2 v2.1.0 h1:HszWKqYQbXxACmAXXWdMsfNl1NDBfGVBnJUPtyUHQ7A=
github.com/eslider/go-xls/v2 v2.1.0/go.mod h1:xgxO6JrfuBr9jGUB+0z5l/yDmFFZ5diGk0ATGihxlMU= github.com/eslider/go-xls/v2 v2.1.0/go.mod h1:xgxO6JrfuBr9jGUB+0z5l/yDmFFZ5diGk0ATGihxlMU=
github.com/go-sql-driver/mysql v1.10.1 h1:arlSnNLq6a5yxGxV7qg9lF4j0C+KwD6NbQyKr9QL6ME=
github.com/go-sql-driver/mysql v1.10.1/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI= github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI=
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY= github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY=
github.com/google/go-querystring v1.2.0 h1:yhqkPbu2/OH+V9BfpCVPZkNmUXhb2gBxJArfhIxNtP0= github.com/google/go-querystring v1.2.0 h1:yhqkPbu2/OH+V9BfpCVPZkNmUXhb2gBxJArfhIxNtP0=
@@ -63,8 +69,20 @@ github.com/hexops/gotextdiff v1.0.3 h1:gitA9+qJrrTCsiCl7+kh75nPqQt1cx4ZkudSTLoUq
github.com/hexops/gotextdiff v1.0.3/go.mod h1:pSWU5MAI3yDq+fZBTazCSJysOMbxWL1BSow5/V2vxeg= github.com/hexops/gotextdiff v1.0.3/go.mod h1:pSWU5MAI3yDq+fZBTazCSJysOMbxWL1BSow5/V2vxeg=
github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8= github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8=
github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw= github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw=
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM=
github.com/jackc/pgx/v5 v5.11.0 h1:IzBBtyK9AHqf98cctWFifYSci2hgQR/cd56wB4p+ogg=
github.com/jackc/pgx/v5 v5.11.0/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/joho/godotenv v1.5.1 h1:7eLL/+HRGLY0ldzfGMeQkb7vMd0as4CfYvUVzLqw0N0= github.com/joho/godotenv v1.5.1 h1:7eLL/+HRGLY0ldzfGMeQkb7vMd0as4CfYvUVzLqw0N0=
github.com/joho/godotenv v1.5.1/go.mod h1:f4LDr5Voq0i2e/R5DDNOoa2zzDfwtkZa6DnEwAbqwq4= github.com/joho/godotenv v1.5.1/go.mod h1:f4LDr5Voq0i2e/R5DDNOoa2zzDfwtkZa6DnEwAbqwq4=
github.com/kr/pretty v0.3.0 h1:WgNl7dwNpEZ6jJ9k1snq4pZsg7DOEN8hP9Xw0Tsjwk0=
github.com/kr/pretty v0.3.0/go.mod h1:640gp4NfQd8pI5XOwp5fnNeVWj67G7CFk/SaSQn7NBk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/lucasb-eyer/go-colorful v1.4.0 h1:UtrWVfLdarDgc44HcS7pYloGHJUjHV/4FwW4TvVgFr4= github.com/lucasb-eyer/go-colorful v1.4.0 h1:UtrWVfLdarDgc44HcS7pYloGHJUjHV/4FwW4TvVgFr4=
github.com/lucasb-eyer/go-colorful v1.4.0/go.mod h1:R4dSotOR9KMtayYi1e77YzuveK+i7ruzyGqttikkLy0= github.com/lucasb-eyer/go-colorful v1.4.0/go.mod h1:R4dSotOR9KMtayYi1e77YzuveK+i7ruzyGqttikkLy0=
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI= github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
@@ -98,6 +116,8 @@ github.com/rivo/uniseg v0.1.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJ
github.com/rivo/uniseg v0.2.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc= github.com/rivo/uniseg v0.2.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc=
github.com/rivo/uniseg v0.4.7 h1:WUdvkW8uEhrYfLC4ZzdpI2ztxP1I582+49Oc5Mq64VQ= github.com/rivo/uniseg v0.4.7 h1:WUdvkW8uEhrYfLC4ZzdpI2ztxP1I582+49Oc5Mq64VQ=
github.com/rivo/uniseg v0.4.7/go.mod h1:FN3SvrM+Zdj16jyLfmOkMNblXMcoc8DfTHruCPUcx88= github.com/rivo/uniseg v0.4.7/go.mod h1:FN3SvrM+Zdj16jyLfmOkMNblXMcoc8DfTHruCPUcx88=
github.com/rogpeppe/go-internal v1.16.0 h1:O9DK+vNMDVGLr2BeZqmpLeMjiMNkuXfcqntWbZV6S5g=
github.com/rogpeppe/go-internal v1.16.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs=
github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM= github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM=
github.com/sebdah/goldie/v2 v2.8.0 h1:dZb9wR8q5++oplmEiJT+U/5KyotVD+HNGCAc5gNr8rc= github.com/sebdah/goldie/v2 v2.8.0 h1:dZb9wR8q5++oplmEiJT+U/5KyotVD+HNGCAc5gNr8rc=
github.com/sebdah/goldie/v2 v2.8.0/go.mod h1:oZ9fp0+se1eapSRjfYbsV/0Hqhbuu3bJVvKI/NNtssI= github.com/sebdah/goldie/v2 v2.8.0/go.mod h1:oZ9fp0+se1eapSRjfYbsV/0Hqhbuu3bJVvKI/NNtssI=
@@ -107,6 +127,9 @@ github.com/spf13/cobra v1.10.2 h1:DMTTonx5m65Ic0GOoRY2c16WCbHxOOw6xxezuLaBpcU=
github.com/spf13/cobra v1.10.2/go.mod h1:7C1pvHqHw5A4vrJfjNwvOdzYu0Gml16OCs2GRiTUUS4= github.com/spf13/cobra v1.10.2/go.mod h1:7C1pvHqHw5A4vrJfjNwvOdzYu0Gml16OCs2GRiTUUS4=
github.com/spf13/pflag v1.0.9 h1:9exaQaMOCwffKiiiYk6/BndUBv+iRViNW+4lEMi0PvY= github.com/spf13/pflag v1.0.9 h1:9exaQaMOCwffKiiiYk6/BndUBv+iRViNW+4lEMi0PvY=
github.com/spf13/pflag v1.0.9/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg= github.com/spf13/pflag v1.0.9/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI=
github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U= github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U= github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/tiendc/go-deepcopy v1.7.2 h1:Ut2yYR7W9tWjTQitganoIue4UGxZwCcJy3orjrrIj44= github.com/tiendc/go-deepcopy v1.7.2 h1:Ut2yYR7W9tWjTQitganoIue4UGxZwCcJy3orjrrIj44=
@@ -142,8 +165,10 @@ golang.org/x/text v0.38.0 h1:sXmwo9DwP3OK9EZ7PqAdaooSGozfl/3a6/xJcbzPRhE=
golang.org/x/text v0.38.0/go.mod h1:YXZt3QhHUKYT53r2lLKFIVi6Ao1jdzrTR/KQ09qyxF4= golang.org/x/text v0.38.0/go.mod h1:YXZt3QhHUKYT53r2lLKFIVi6Ao1jdzrTR/KQ09qyxF4=
golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q= golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q=
golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA= golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA= gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI= modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI=
+104 -31
View File
@@ -111,6 +111,25 @@ func (c *Client) deleteObject(ctx context.Context, path string) (map[string]any,
return unmarshalResponseObject(raw) return unmarshalResponseObject(raw)
} }
// unmarshalResponseArray extracts the "response" field from a raw OnlyOffice
// envelope and decodes it into a list of maps. Returns (nil, nil) for a null,
// empty or scalar payload. Companion to unmarshalResponseObject for endpoints
// whose response is a list (project team, people/status, …).
func unmarshalResponseArray(raw json.RawMessage) ([]map[string]any, error) {
resp, err := responseField(raw, "response")
if err != nil {
return nil, err
}
if len(resp) == 0 || string(resp) == "null" || resp[0] != '[' {
return nil, nil
}
var list []map[string]any
if err := json.Unmarshal(resp, &list); err != nil {
return nil, err
}
return list, nil
}
// unmarshalResponseObject extracts the "response" field from a raw OnlyOffice // unmarshalResponseObject extracts the "response" field from a raw OnlyOffice
// envelope and decodes it into map[string]any. Returns (nil, nil) for a null // envelope and decodes it into map[string]any. Returns (nil, nil) for a null
// response, an empty array, or scalar payloads. When the API returns a list // response, an empty array, or scalar payloads. When the API returns a list
@@ -144,8 +163,13 @@ func unmarshalResponseObject(raw json.RawMessage) (map[string]any, error) {
} }
} }
// getJSON issues an authenticated GET and returns the raw response body. // getJSON issues an authenticated GET and returns the raw response body,
// retrying transient answers (see retryRaw).
func (c *Client) getJSON(ctx context.Context, path string) (json.RawMessage, error) { func (c *Client) getJSON(ctx context.Context, path string) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) { return c.getJSONOnce(ctx, path) })
}
func (c *Client) getJSONOnce(ctx context.Context, path string) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
@@ -166,7 +190,7 @@ func (c *Client) getJSON(ctx context.Context, path string) (json.RawMessage, err
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("GET %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "GET %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
@@ -187,6 +211,12 @@ func (c *Client) deleteForm(ctx context.Context, path string, fields url.Values)
} }
func (c *Client) formRequest(ctx context.Context, method, path string, fields url.Values) (json.RawMessage, error) { func (c *Client) formRequest(ctx context.Context, method, path string, fields url.Values) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) {
return c.formRequestOnce(ctx, method, path, fields)
})
}
func (c *Client) formRequestOnce(ctx context.Context, method, path string, fields url.Values) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
@@ -208,13 +238,17 @@ func (c *Client) formRequest(ctx context.Context, method, path string, fields ur
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("%s form %s: %d %s", method, path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "%s form %s: %d %s", method, path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
// deleteReq issues an authenticated DELETE. // deleteReq issues an authenticated DELETE.
func (c *Client) deleteReq(ctx context.Context, path string) (json.RawMessage, error) { func (c *Client) deleteReq(ctx context.Context, path string) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) { return c.deleteReqOnce(ctx, path) })
}
func (c *Client) deleteReqOnce(ctx context.Context, path string) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
@@ -235,7 +269,7 @@ func (c *Client) deleteReq(ctx context.Context, path string) (json.RawMessage, e
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("DELETE %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "DELETE %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
@@ -258,26 +292,65 @@ func (c *Client) postJSONObject(ctx context.Context, path string, body any) (map
return unmarshalResponseObject(raw) return unmarshalResponseObject(raw)
} }
// postJSON issues an authenticated POST with application/json body. // jsonBodyReader turns a request body value into an io.Reader. nil becomes
func (c *Client) postJSON(ctx context.Context, path string, body any) (json.RawMessage, error) { // "{}", []byte/string pass through, anything else is JSON-marshalled.
auth, err := c.authHeader() func jsonBodyReader(body any) (io.Reader, error) {
if err != nil {
return nil, err
}
var rdr io.Reader
switch b := body.(type) { switch b := body.(type) {
case nil: case nil:
rdr = strings.NewReader("{}") return strings.NewReader("{}"), nil
case []byte: case []byte:
rdr = bytes.NewReader(b) return bytes.NewReader(b), nil
case string: case string:
rdr = strings.NewReader(b) return strings.NewReader(b), nil
default: default:
buf, err := json.Marshal(b) buf, err := json.Marshal(b)
if err != nil { if err != nil {
return nil, err return nil, err
} }
rdr = bytes.NewReader(buf) return bytes.NewReader(buf), nil
}
}
// postJSONArray is postJSON + unmarshalResponseArray.
func (c *Client) postJSONArray(ctx context.Context, path string, body any) ([]map[string]any, error) {
raw, err := c.postJSON(ctx, path, body)
if err != nil {
return nil, err
}
return unmarshalResponseArray(raw)
}
// putJSONArray is putJSON + unmarshalResponseArray.
func (c *Client) putJSONArray(ctx context.Context, path string, body any) ([]map[string]any, error) {
raw, err := c.putJSON(ctx, path, body)
if err != nil {
return nil, err
}
return unmarshalResponseArray(raw)
}
// deleteJSONArray is deleteJSON + unmarshalResponseArray.
func (c *Client) deleteJSONArray(ctx context.Context, path string, body any) ([]map[string]any, error) {
raw, err := c.deleteJSON(ctx, path, body)
if err != nil {
return nil, err
}
return unmarshalResponseArray(raw)
}
// postJSON issues an authenticated POST with application/json body.
func (c *Client) postJSON(ctx context.Context, path string, body any) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) { return c.postJSONOnce(ctx, path, body) })
}
func (c *Client) postJSONOnce(ctx context.Context, path string, body any) (json.RawMessage, error) {
auth, err := c.authHeader()
if err != nil {
return nil, err
}
rdr, err := jsonBodyReader(body)
if err != nil {
return nil, err
} }
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL()+path, rdr) req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL()+path, rdr)
if err != nil { if err != nil {
@@ -296,32 +369,25 @@ func (c *Client) postJSON(ctx context.Context, path string, body any) (json.RawM
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("POST JSON %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "POST JSON %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
// putJSON issues an authenticated PUT with application/json body. // putJSON issues an authenticated PUT with application/json body.
func (c *Client) putJSON(ctx context.Context, path string, body any) (json.RawMessage, error) { func (c *Client) putJSON(ctx context.Context, path string, body any) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) { return c.putJSONOnce(ctx, path, body) })
}
func (c *Client) putJSONOnce(ctx context.Context, path string, body any) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
} }
var rdr io.Reader rdr, err := jsonBodyReader(body)
switch b := body.(type) {
case nil:
rdr = strings.NewReader("{}")
case []byte:
rdr = bytes.NewReader(b)
case string:
rdr = strings.NewReader(b)
default:
buf, err := json.Marshal(b)
if err != nil { if err != nil {
return nil, err return nil, err
} }
rdr = bytes.NewReader(buf)
}
req, err := http.NewRequestWithContext(ctx, http.MethodPut, c.baseURL()+path, rdr) req, err := http.NewRequestWithContext(ctx, http.MethodPut, c.baseURL()+path, rdr)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -339,7 +405,7 @@ func (c *Client) putJSON(ctx context.Context, path string, body any) (json.RawMe
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("PUT JSON %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "PUT JSON %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
@@ -352,8 +418,15 @@ func (c *Client) uploadMultipart(ctx context.Context, path, fieldName, filePath
// uploadMultipartMethod sends a single-file multipart request with the given // uploadMultipartMethod sends a single-file multipart request with the given
// HTTP method. The OnlyOffice Documents API needs PUT for /update (a new // HTTP method. The OnlyOffice Documents API needs PUT for /update (a new
// version) and POST for /upload (a new file); sending POST to /update answers // version) and POST for /upload (a new file); sending POST to /update answers
// 500 on current servers. // 500 on current servers. The file is re-opened per attempt, so transient
// answers are retried like every other request.
func (c *Client) uploadMultipartMethod(ctx context.Context, method, path, fieldName, filePath string) (json.RawMessage, error) { func (c *Client) uploadMultipartMethod(ctx context.Context, method, path, fieldName, filePath string) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) {
return c.uploadMultipartOnce(ctx, method, path, fieldName, filePath)
})
}
func (c *Client) uploadMultipartOnce(ctx context.Context, method, path, fieldName, filePath string) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
@@ -393,7 +466,7 @@ func (c *Client) uploadMultipartMethod(ctx context.Context, method, path, fieldN
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("upload %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "upload %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
+23
View File
@@ -48,3 +48,26 @@ func TestUnmarshalResponseObjectNull(t *testing.T) {
t.Fatalf("expected nil, got %#v", out) t.Fatalf("expected nil, got %#v", out)
} }
} }
func TestUnmarshalResponseArrayList(t *testing.T) {
raw := json.RawMessage(`{"response":[{"id":"a","displayName":"A"},{"id":"b","displayName":"B"}]}`)
out, err := unmarshalResponseArray(raw)
if err != nil {
t.Fatal(err)
}
if len(out) != 2 || out[0]["id"] != "a" || out[1]["displayName"] != "B" {
t.Fatalf("unexpected list: %#v", out)
}
}
func TestUnmarshalResponseArrayNullAndScalar(t *testing.T) {
for _, raw := range []string{`{"response":null}`, `{"response":{}}`, `{"response":"x"}`} {
out, err := unmarshalResponseArray(json.RawMessage(raw))
if err != nil {
t.Fatalf("%s: %v", raw, err)
}
if out != nil {
t.Fatalf("%s: expected nil, got %#v", raw, out)
}
}
}
+3
View File
@@ -5,6 +5,7 @@
// - pandoc — md↔docx // - pandoc — md↔docx
// - ocrmypdf — OCR into a searchable PDF // - ocrmypdf — OCR into a searchable PDF
// - pdftotext — extract text layer // - pdftotext — extract text layer
// - pdfdetach — list/save embedded PDF attachments
// - tesseract — OCR single images when ocrmypdf is unsuitable // - tesseract — OCR single images when ocrmypdf is unsuitable
// - ghostscript (gs) — PDF rewrite/optimize via PostScript (pdfwrite) // - ghostscript (gs) — PDF rewrite/optimize via PostScript (pdfwrite)
package docpipe package docpipe
@@ -26,6 +27,7 @@ type Tools struct {
Pandoc string Pandoc string
OCRMyPDF string OCRMyPDF string
PDFToText string PDFToText string
PDFDetach string
Tesseract string Tesseract string
Ghostscript string Ghostscript string
} }
@@ -44,6 +46,7 @@ func LookPath() Tools {
Pandoc: find("pandoc"), Pandoc: find("pandoc"),
OCRMyPDF: find("ocrmypdf"), OCRMyPDF: find("ocrmypdf"),
PDFToText: find("pdftotext"), PDFToText: find("pdftotext"),
PDFDetach: find("pdfdetach"),
Tesseract: find("tesseract"), Tesseract: find("tesseract"),
Ghostscript: find("gs", "ghostscript"), Ghostscript: find("gs", "ghostscript"),
} }
+278
View File
@@ -0,0 +1,278 @@
package docpipe
// Embedded PDF attachments (F6 #42). Digitised invoices often carry the
// original scan as a PDF attachment; the searchable body may hold only a
// summary. pdfdetach (poppler) lists/saves them; each saved attachment is run
// through the normal docpipe extraction (pdftotext/OCR).
import (
"bytes"
"encoding/xml"
"fmt"
"io"
"os"
"os/exec"
"path/filepath"
"strconv"
"strings"
"unicode"
"unicode/utf8"
)
// PDFAttachment is one embedded file in a PDF.
type PDFAttachment struct {
Index int // 1-based number as `pdfdetach -list` reports it
Name string // embedded file name
}
// AttachmentText is the extracted text of one embedded attachment.
type AttachmentText struct {
Name string
Text string
}
// parseAttachmentList parses `pdfdetach -list` output. The first line is a
// count ("N embedded files"); every following line is "<index>: <name>".
// Pure, so it is unit-tested.
func parseAttachmentList(out string) []PDFAttachment {
var atts []PDFAttachment
for _, line := range strings.Split(out, "\n") {
line = strings.TrimSpace(line)
if line == "" {
continue
}
colon := strings.Index(line, ":")
if colon <= 0 {
continue
}
n, err := strconv.Atoi(strings.TrimSpace(line[:colon]))
if err != nil {
continue
}
name := strings.TrimSpace(line[colon+1:])
if name == "" {
continue
}
atts = append(atts, PDFAttachment{Index: n, Name: name})
}
return atts
}
// safeAttachmentName strips directories and leading dots so a hostile
// attachment name cannot escape the extraction directory.
func safeAttachmentName(name string) string {
name = strings.ReplaceAll(strings.TrimSpace(name), "\\", "/")
name = filepath.Base(name)
name = strings.TrimLeft(name, ".")
if name == "" || name == "." || name == "/" {
return ""
}
return name
}
// JoinWithAttachments appends attachment text to the document body, each
// section preceded by an "[attachment: <name>]" marker so a search hit shows
// its source. Empty attachments are skipped. Pure, so it is unit-tested.
func JoinWithAttachments(body string, atts []AttachmentText) string {
var b strings.Builder
b.WriteString(strings.TrimRight(body, "\n"))
for _, a := range atts {
text := strings.TrimSpace(a.Text)
if text == "" {
continue
}
b.WriteString("\n\n[attachment: ")
b.WriteString(a.Name)
b.WriteString("]\n\n")
b.WriteString(text)
}
return b.String()
}
// ListAttachments returns the embedded files of a PDF. A PDF without
// attachments yields an empty slice and no error.
func (t Tools) ListAttachments(pdfPath string) ([]PDFAttachment, error) {
if t.PDFDetach == "" {
return nil, fmt.Errorf("pdfdetach not found on PATH")
}
cmd := exec.Command(t.PDFDetach, "-list", pdfPath)
var stderr bytes.Buffer
cmd.Stderr = &stderr
out, err := cmd.Output()
if err != nil {
return nil, fmt.Errorf("pdfdetach -list %s: %w (%s)", filepath.Base(pdfPath), err, strings.TrimSpace(stderr.String()))
}
return parseAttachmentList(string(out)), nil
}
// SaveAttachment writes the n-th embedded file (1-based) to outPath.
func (t Tools) SaveAttachment(pdfPath string, index int, outPath string) error {
if t.PDFDetach == "" {
return fmt.Errorf("pdfdetach not found on PATH")
}
if strings.TrimSpace(outPath) == "" {
return fmt.Errorf("output path required")
}
if err := EnsureDir(outPath); err != nil {
return err
}
cmd := exec.Command(t.PDFDetach, "-save", strconv.Itoa(index), "-o", outPath, pdfPath)
var stderr bytes.Buffer
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return fmt.Errorf("pdfdetach -save %d: %w (%s)", index, err, strings.TrimSpace(stderr.String()))
}
return nil
}
// ToMarkdownWithAttachments extracts the file as ToMarkdown does, then — for
// PDFs — appends the text of every embedded attachment under an
// "[attachment: <name>]" marker. Attachment failures are non-fatal: the body
// is returned unchanged.
func (t Tools) ToMarkdownWithAttachments(path, workDir, lang string, minChars int) (string, error) {
res, err := t.ToMarkdown(path, workDir, lang, minChars)
if err != nil {
return "", err
}
if Ext(path) != ".pdf" {
return res.Markdown, nil
}
atts, err := t.attachmentTexts(path, workDir, lang, minChars)
if err != nil {
return res.Markdown, nil
}
return JoinWithAttachments(res.Markdown, atts), nil
}
// attachmentTexts saves and extracts every embedded attachment, skipping the
// ones that cannot be read. It returns an error only when the attachment list
// itself cannot be obtained.
func (t Tools) attachmentTexts(pdfPath, workDir, lang string, minChars int) ([]AttachmentText, error) {
list, err := t.ListAttachments(pdfPath)
if err != nil || len(list) == 0 {
return nil, err
}
if workDir == "" {
workDir = os.TempDir()
}
dir := filepath.Join(workDir, "att-"+trimExt(filepath.Base(pdfPath)))
if err := os.MkdirAll(dir, 0o755); err != nil {
return nil, err
}
defer os.RemoveAll(dir)
out := make([]AttachmentText, 0, len(list))
for _, a := range list {
name := safeAttachmentName(a.Name)
if name == "" {
continue
}
saved := filepath.Join(dir, fmt.Sprintf("%d-%s", a.Index, name))
if err := t.SaveAttachment(pdfPath, a.Index, saved); err != nil {
continue
}
text, err := t.attachmentMarkdown(saved, dir, lang, minChars)
if err != nil {
continue
}
out = append(out, AttachmentText{Name: a.Name, Text: text})
}
return out, nil
}
// attachmentMarkdown extracts a saved attachment with the regular pipeline.
// Structured attachments that docpipe does not convert (e-invoice XML,
// CuraSoft JSON, CSV/HTML) fall back to their text content, so the embedded
// original is still searchable. Other unreadable formats return an error and
// the caller skips them.
func (t Tools) attachmentMarkdown(path, workDir, lang string, minChars int) (string, error) {
if res, err := t.ToMarkdown(path, workDir, lang, minChars); err == nil {
return res.Markdown, nil
}
switch Ext(path) {
case ".xml", ".html", ".htm":
raw, err := os.ReadFile(path)
if err != nil {
return "", err
}
return xmlToText(raw), nil
case ".json", ".csv", ".yaml", ".yml", ".toml", ".txt", ".md", ".markdown":
raw, err := os.ReadFile(path)
if err != nil {
return "", err
}
return string(raw), nil
default:
return textFallback(path)
}
}
// textFallback reads an attachment of an unknown or missing extension as plain
// text when it looks textual (valid UTF-8, mostly printable runes). Binary
// payloads (images, archives, NUL-padded blobs) are rejected with an error so
// the caller skips them instead of poisoning the index. Classified digitised
// PDFs (Scanner-*.ocr.pdf) carry .yaml/.md attachments; some exporters omit the
// extension, which this covers.
func textFallback(path string) (string, error) {
f, err := os.Open(path)
if err != nil {
return "", err
}
defer f.Close()
raw, err := io.ReadAll(io.LimitReader(f, 1<<20))
if err != nil {
return "", err
}
if !utf8.Valid(raw) {
return "", fmt.Errorf("unsupported attachment type %q", Ext(path))
}
if !mostlyPrintable(raw) {
return "", fmt.Errorf("unsupported attachment type %q", Ext(path))
}
return string(raw), nil
}
// mostlyPrintable reports whether at least 90% of the runes are printable text
// (newlines, carriage returns and tabs count as text). Pure, so it is tested.
func mostlyPrintable(b []byte) bool {
if len(b) == 0 {
return false
}
printable, total := 0, 0
for _, r := range string(b) {
if r == utf8.RuneError {
continue
}
total++
if unicode.IsPrint(r) || r == '\n' || r == '\r' || r == '\t' {
printable++
}
}
return total > 0 && printable*10 >= total*9
}
// xmlToText returns the character data of an XML/HTML document: element text
// values with decoded entities, one per line. Used for invoice XML (EN 16931
// CII / ZUGFeRD) and HTML attachments. Pure, so it is unit-tested.
func xmlToText(raw []byte) string {
dec := xml.NewDecoder(bytes.NewReader(raw))
dec.Strict = false
var b strings.Builder
for {
tok, err := dec.Token()
if err != nil {
break
}
cd, ok := tok.(xml.CharData)
if !ok {
continue
}
s := strings.TrimSpace(string(cd))
if s == "" {
continue
}
b.WriteString(s)
b.WriteByte('\n')
}
return b.String()
}
+189
View File
@@ -0,0 +1,189 @@
package docpipe
import (
"os"
"path/filepath"
"strings"
"testing"
)
func TestParseAttachmentList(t *testing.T) {
out := "2 embedded files\n1: original.pdf\n2: scan_001.png\n"
got := parseAttachmentList(out)
want := []PDFAttachment{{Index: 1, Name: "original.pdf"}, {Index: 2, Name: "scan_001.png"}}
if len(got) != len(want) {
t.Fatalf("got %+v, want %+v", got, want)
}
for i := range want {
if got[i] != want[i] {
t.Errorf("att[%d] = %+v, want %+v", i, got[i], want[i])
}
}
}
func TestParseAttachmentListEmptyAndMalformed(t *testing.T) {
for _, in := range []string{"", "0 embedded files\n", "garbage\n\n \n"} {
if got := parseAttachmentList(in); len(got) != 0 {
t.Errorf("parseAttachmentList(%q) = %+v, want empty", in, got)
}
}
}
func TestSafeAttachmentName(t *testing.T) {
cases := map[string]string{
"note.txt": "note.txt",
"../../evil.pdf": "evil.pdf",
`..\..\evil.pdf`: "evil.pdf",
"/abs/scan_001.pdf": "scan_001.pdf",
".hidden": "hidden",
" spaced name.txt ": "spaced name.txt",
"..": "",
"": "",
}
for in, want := range cases {
if got := safeAttachmentName(in); got != want {
t.Errorf("safeAttachmentName(%q) = %q, want %q", in, got, want)
}
}
}
func TestJoinWithAttachments(t *testing.T) {
body := "# scan.pdf\n\nbody token\n"
atts := []AttachmentText{
{Name: "original.pdf", Text: " original token "},
{Name: "empty.txt", Text: " "},
}
got := JoinWithAttachments(body, atts)
if !strings.Contains(got, "body token") {
t.Errorf("body text lost: %q", got)
}
if !strings.Contains(got, "[attachment: original.pdf]") {
t.Errorf("marker missing: %q", got)
}
if !strings.Contains(got, "original token") {
t.Errorf("attachment text missing: %q", got)
}
if strings.Contains(got, "empty.txt") {
t.Errorf("empty attachment must be skipped: %q", got)
}
}
func TestJoinWithAttachmentsNoAttachments(t *testing.T) {
got := JoinWithAttachments("# a.pdf\n\ntext\n\n", nil)
if got != "# a.pdf\n\ntext" {
t.Errorf("got %q, want trimmed body only", got)
}
}
func TestListAttachmentsWithoutTool(t *testing.T) {
if _, err := (Tools{}).ListAttachments("x.pdf"); err == nil || !strings.Contains(err.Error(), "pdfdetach") {
t.Fatalf("want pdfdetach error, got %v", err)
}
}
// TestToMarkdownWithAttachmentsFixture exercises the real pdfdetach + pdftotext
// pipeline on testdata/pdf-with-attachment.pdf (body token + embedded
// goo-note.txt). Skips when poppler is not installed.
func TestToMarkdownWithAttachmentsFixture(t *testing.T) {
tools := LookPath()
if tools.PDFDetach == "" || tools.PDFToText == "" {
t.Skip("pdfdetach/pdftotext not on PATH — skipping attachment extraction test")
}
fixture := filepath.Join("..", "..", "testdata", "pdf-with-attachment.pdf")
got, err := tools.ToMarkdownWithAttachments(fixture, t.TempDir(), "eng", 1)
if err != nil {
t.Fatalf("ToMarkdownWithAttachments: %v", err)
}
for _, want := range []string{"goobodytoken", "[attachment: goo-note.txt]", "gooattachmenttoken"} {
if !strings.Contains(got, want) {
t.Errorf("result missing %q:\n%s", want, got)
}
}
}
func TestXMLToText(t *testing.T) {
raw := []byte(`<?xml version="1.0" encoding="UTF-8"?>
<rsm:CrossIndustryInvoice><rsm:ExchangedDocument>
<ram:ID>S1063</ram:ID></rsm:ExchangedDocument>
<ram:Name>Acme &amp; Co</ram:Name><ram:GrandTotalAmount>42.00</ram:GrandTotalAmount>
</rsm:CrossIndustryInvoice>`)
got := xmlToText(raw)
for _, want := range []string{"S1063", "Acme & Co", "42.00"} {
if !strings.Contains(got, want) {
t.Errorf("xmlToText missing %q:\n%s", want, got)
}
}
if strings.ContainsAny(got, "<>") {
t.Errorf("xmlToText left markup: %q", got)
}
}
// TestAttachmentMarkdownFallback verifies structured attachments that docpipe
// cannot convert are still reduced to searchable text, and unknown binary
// formats error (so the caller skips them).
func TestAttachmentMarkdownFallback(t *testing.T) {
dir := t.TempDir()
xmlPath := filepath.Join(dir, "factur-x.xml")
if err := os.WriteFile(xmlPath, []byte(`<Invoice><Number>S1063</Number></Invoice>`), 0o644); err != nil {
t.Fatal(err)
}
got, err := (Tools{}).attachmentMarkdown(xmlPath, dir, "", 0)
if err != nil {
t.Fatalf("attachmentMarkdown(xml): %v", err)
}
if !strings.Contains(got, "S1063") {
t.Errorf("xml attachment text = %q, want S1063", got)
}
// Classified digitised PDFs (Scanner-*.ocr.pdf) carry .yaml metadata.
yamlPath := filepath.Join(dir, "Scanner-123-003.ocr.yaml")
if err := os.WriteFile(yamlPath, []byte("document:\n type: Rechnung\nnumber: S1063\n"), 0o644); err != nil {
t.Fatal(err)
}
gotYAML, err := (Tools{}).attachmentMarkdown(yamlPath, dir, "", 0)
if err != nil {
t.Fatalf("attachmentMarkdown(yaml): %v", err)
}
if !strings.Contains(gotYAML, "S1063") {
t.Errorf("yaml attachment text = %q, want S1063", gotYAML)
}
// Extensionless textual attachment falls back to raw text.
noExt := filepath.Join(dir, "attachment")
if err := os.WriteFile(noExt, []byte("plain attachment token goonoext"), 0o644); err != nil {
t.Fatal(err)
}
gotNoExt, err := (Tools{}).attachmentMarkdown(noExt, dir, "", 0)
if err != nil {
t.Fatalf("attachmentMarkdown(no extension): %v", err)
}
if !strings.Contains(gotNoExt, "goonoext") {
t.Errorf("extensionless attachment text = %q, want goonoext", gotNoExt)
}
binPath := filepath.Join(dir, "data.bin")
if err := os.WriteFile(binPath, []byte{0, 1, 2, 3}, 0o644); err != nil {
t.Fatal(err)
}
if _, err := (Tools{}).attachmentMarkdown(binPath, dir, "", 0); err == nil {
t.Error("unsupported attachment: want error, got nil")
}
}
// TestToMarkdownWithAttachmentsPlainPDF ensures a PDF without attachments
// returns just the body (pdfdetach prints "0 embedded files").
func TestToMarkdownWithAttachmentsPlainPDF(t *testing.T) {
tools := LookPath()
if tools.PDFDetach == "" || tools.PDFToText == "" {
t.Skip("pdfdetach/pdftotext not on PATH")
}
// The fixture itself is a PDF with one attachment; strip it by extracting
// the body only through ToMarkdown and compare JoinWithAttachments(nil).
res, err := tools.ToMarkdown(filepath.Join("..", "..", "testdata", "pdf-with-attachment.pdf"), t.TempDir(), "eng", 1)
if err != nil {
t.Fatalf("ToMarkdown: %v", err)
}
if strings.Contains(res.Markdown, "gooattachmenttoken") {
t.Fatalf("body must not contain attachment text: %q", res.Markdown)
}
}
-393
View File
@@ -1,393 +0,0 @@
package xlspipe
import (
"fmt"
"github.com/xuri/excelize/v2"
)
// Sheet names (Russian tabs) for cutover Portugal workbook.
const (
SheetInputs = "Ввод"
SheetDom6 = "Вс 6.09"
SheetJue3 = "Чт 3.09"
SheetWedHyp = "Ср гип"
SheetSummary = "Сводка"
)
// CutoverPortugalDefaults holds live FO values (2026-08-28).
type CutoverPortugalDefaults struct {
FerryDom6 float64
FerryJue3 float64
FerrySuperiorDelta float64
HousingLow float64
HousingMid float64
HousingHigh float64
DriveLow float64
DriveMid float64
DriveHigh float64
BoardLow float64
BoardMid float64
BoardHigh float64
FoodLow float64
FoodMid float64
FoodHigh float64
SimLow float64
SimMid float64
SimHigh float64
MonthCap float64
ExtraNightsDom6 float64
ExtraNightsJue3 float64
ExtraNightsWed float64
}
// DefaultCutoverPortugal returns FO live snapshot from portugal track (28.08.2026).
func DefaultCutoverPortugal() CutoverPortugalDefaults {
return CutoverPortugalDefaults{
FerryDom6: 457.89,
FerryJue3: 484.09,
FerrySuperiorDelta: 19.64,
HousingLow: 509,
HousingMid: 600,
HousingHigh: 650,
DriveLow: 70,
DriveMid: 85,
DriveHigh: 100,
BoardLow: 40,
BoardMid: 60,
BoardHigh: 80,
FoodLow: 150,
FoodMid: 220,
FoodHigh: 300,
SimLow: 50,
SimMid: 100,
SimHigh: 150,
MonthCap: 2500,
ExtraNightsDom6: 0,
ExtraNightsJue3: 0,
ExtraNightsWed: 4,
}
}
type inputField struct {
name string // defined name (ASCII, for formulas)
label string
value float64
note string
comment string
}
// BuildCutoverPortugalWorkbook creates a multi-sheet cutover budget with Russian labels,
// cell comments on non-obvious inputs, named ranges, and cross-sheet formulas.
func BuildCutoverPortugalWorkbook(d CutoverPortugalDefaults) (*excelize.File, error) {
f := excelize.NewFile()
defaultSheet := f.GetSheetName(0)
if err := f.SetSheetName(defaultSheet, SheetInputs); err != nil {
f.Close()
return nil, err
}
for _, name := range []string{SheetDom6, SheetJue3, SheetWedHyp, SheetSummary} {
if _, err := f.NewSheet(name); err != nil {
f.Close()
return nil, err
}
}
if err := writeInputsSheet(f, d); err != nil {
f.Close()
return nil, err
}
scenarios := []struct {
sheet, ferry, extra, note string
}{
{SheetDom6, "ferry_dom6", "extra_nights_dom6", "Живой слот вс 6.09 20:30; заезд в квартиру вт 8.09"},
{SheetJue3, "ferry_jue3", "extra_nights_jue3", "Живой чт 3.09; T1a 05–12 если TF-крыша кончается раньше вс"},
{SheetWedHyp, "ferry_jue3", "extra_nights_wed", "Гипотеза: ср 2.09 20:00 по тарифу Jue3; заезд пт 4.09"},
}
for _, sc := range scenarios {
if err := writeScenarioSheet(f, sc.sheet, sc.ferry, sc.extra, sc.note); err != nil {
f.Close()
return nil, err
}
}
if err := writeSummarySheet(f); err != nil {
f.Close()
return nil, err
}
f.SetActiveSheet(0)
if err := finalizeWorkbook(f); err != nil {
f.Close()
return nil, err
}
return f, nil
}
func writeInputsSheet(f *excelize.File, d CutoverPortugalDefaults) error {
if err := setHeaders(f, SheetInputs, "Параметр", "Значение €", "Кратко"); err != nil {
return err
}
fields := []inputField{
{
name: "ferry_dom6", label: "Паром вс 6.09 (Básica, без residencia)",
value: d.FerryDom6, note: "FO live: 2взр+младенец+авто",
comment: "Fred Olsen Dom 6.09 20:30 SC→Huelva, прибытие Mar 8 09:00. Butaca Normal/Básica без субсидии канарского residencia (−183€). Меняйте после нового live FO.",
},
{
name: "ferry_jue3", label: "Паром чт 3.09 (Básica, без residencia)",
value: d.FerryJue3, note: "FO live Jue 3 20:00",
comment: "Прямой рейс 35 ч. Дороже вс на ~26€. Используется также для листа «Ср гип», если своего рейса ср 2.09 нет в продаже.",
},
{
name: "ferry_superior_uplift", label: "Доплата VIP / Butaca Superior",
value: d.FerrySuperiorDelta, note: "Superior − Normal (Jue3 live)",
comment: "Разница между Superior и Normal на live Jue3 ≈19,64€. На Dom6 Superior в сессии не перевыбирали — оценка по этой дельте. VIP = salón, не каюта.",
},
{
name: "housing_7n_low", label: "Крыша 7 ночей Setúbal — минимум",
value: d.HousingLow, note: "Airbnb low band",
comment: "Короткая аренда T1a (#11): 08–15.09, 1–2BR с парковкой. Низкая граница live Airbnb Setúbal.",
},
{
name: "housing_7n_mid", label: "Крыша 7 ночей Setúbal — целевой mid",
value: d.HousingMid, note: "Цель #11: 500–650€/нед",
comment: "Рабочая оценка для брони. Основной столбец mid на листах сценариев.",
},
{
name: "housing_7n_high", label: "Крыша 7 ночей Setúbal — максимум",
value: d.HousingHigh, note: "Airbnb high band",
},
{
name: "drive_low", label: "Проезд Huelva → Setúbal — мин",
value: d.DriveLow, note: "OSRM ~3,8 ч",
comment: "≈342 км: топливо + платные дороги. Низкая/средняя/высокая оценка.",
},
{name: "drive_mid", label: "Проезд Huelva → Setúbal — mid", value: d.DriveMid},
{name: "drive_high", label: "Проезд Huelva → Setúbal — макс", value: d.DriveHigh},
{
name: "board_low", label: "Еда на пароме — мин",
value: d.BoardLow, note: "Меню FO",
comment: "Питание на борту (Fred Olsen). George 0–3 обычно бесплатно как пассажир — еда отдельно.",
},
{name: "board_mid", label: "Еда на пароме — mid", value: d.BoardMid},
{name: "board_high", label: "Еда на пароме — макс", value: d.BoardHigh},
{
name: "food_7d_low", label: "Еда 7 дней в PT — мин",
value: d.FoodLow, note: "Готовим в apt",
comment: "Первая неделя в Setúbal после парома — продукты, не рестораны.",
},
{name: "food_7d_mid", label: "Еда 7 дней в PT — mid", value: d.FoodMid},
{name: "food_7d_high", label: "Еда 7 дней в PT — макс", value: d.FoodHigh},
{
name: "sim_low", label: "SIM + документы — мин",
value: d.SimLow, note: "Разовые cutover",
comment: "eSIM, копии, мелкие госпошлины при cutover. Не включает депозит аренды.",
},
{name: "sim_mid", label: "SIM + документы — mid", value: d.SimMid},
{name: "sim_high", label: "SIM + документы — макс", value: d.SimHigh},
{
name: "month_cap", label: "Потолок бюджета на месяц (€)",
value: d.MonthCap, note: "SoT: 2500€",
comment: "Жёсткий потолок Sep из source-of-truth. «Остаток» = потолок − итого mid сценария.",
},
{
name: "extra_nights_dom6", label: "Лишние ночи в PT (вс 6.09)",
value: d.ExtraNightsDom6, note: "0 = TF до вс",
comment: "Платные ночи в PT до начала 7-дневной крыши. Для Dom6 обычно 0: остаёмся на Тенерифе до вс, заезд вт 8.09.",
},
{
name: "extra_nights_jue3", label: "Лишние ночи в PT (чт 3.09)",
value: d.ExtraNightsJue3, note: "0 если TF до вс",
comment: "Если крыша TF кончается раньше вс — нужны ночи 05–07.09 до Airbnb. По умолчанию 0.",
},
{
name: "extra_nights_wed", label: "Лишние ночи в PT (ср гип)",
value: d.ExtraNightsWed, note: "Гип: заезд пт 4.09",
comment: "Гипотетический слот ср 2.09 → заезд пт 4.09 = 4 лишних ночи до типичного 7н блока. Рейса ср в FO нет — тариф как Jue3.",
},
}
for i, fld := range fields {
row := i + 2
if err := f.SetCellStr(SheetInputs, fmt.Sprintf("A%d", row), fld.label); err != nil {
return err
}
if err := defineInput(f, SheetInputs, fld.name, row, fld.value, fld.note, fld.comment); err != nil {
return err
}
}
if err := f.SetCellStr(SheetInputs, "A25", "—"); err != nil {
return err
}
if err := f.SetCellStr(SheetInputs, "B25", "Редактируйте жёлтые ячейки"); err != nil {
return err
}
if err := f.SetCellStr(SheetInputs, "C25", "Формулы на листах сценариев и «Сводка» пересчитаются в OnlyOffice"); err != nil {
return err
}
return styleSheet(f, SheetInputs, SheetInputs)
}
func writeScenarioSheet(f *excelize.File, sheet, ferryName, extraNightsName, scenarioNote string) error {
if err := setHeaders(f, sheet, "Статья", "Мин €", "Mid €", "Макс €", "Пояснение"); err != nil {
return err
}
if err := f.SetCellStr(sheet, "A2", "Паром Básica"); err != nil {
return err
}
for col, ref := range []string{ferryName, ferryName, ferryName} {
cell, _ := excelize.CoordinatesToCellName(col+2, 2)
if err := f.SetCellFormula(sheet, cell, "="+ref); err != nil {
return err
}
}
if err := addCellComment(f, sheet, "A2", "Тариф парома для этого сценария. Берётся с листа «Ввод»."); err != nil {
return err
}
lines := []struct {
label string
low, mid, high, note string
comment string
}{
{"Крыша 7 н Setúbal", "housing_7n_low", "housing_7n_mid", "housing_7n_high", "Airbnb T1a #11", ""},
{"Huelva → Setúbal", "drive_low", "drive_mid", "drive_high", "OSRM ~3,8 ч", ""},
{"Еда на борту", "board_low", "board_mid", "board_high", "Меню FO", ""},
{"Еда 7 д в PT", "food_7d_low", "food_7d_mid", "food_7d_high", "Готовим дома", ""},
{"SIM / документы", "sim_low", "sim_mid", "sim_high", "Cutover", ""},
}
for i, ln := range lines {
row := i + 3
if err := setFormulaRow(f, sheet, row, ln.label,
"="+ln.low, "="+ln.mid, "="+ln.high, ln.note); err != nil {
return err
}
}
if err := f.SetCellStr(sheet, "A8", "Итого Básica"); err != nil {
return err
}
for col := 2; col <= 4; col++ {
cell, _ := excelize.CoordinatesToCellName(col, 8)
colL, _ := excelize.CoordinatesToCellName(col, 2)
colH, _ := excelize.CoordinatesToCellName(col, 7)
if err := f.SetCellFormula(sheet, cell, fmt.Sprintf("=SUM(%s:%s)", colL, colH)); err != nil {
return err
}
}
if err := f.SetCellStr(sheet, "E8", "SUM строк 2–7"); err != nil {
return err
}
if err := setFormulaRow(f, sheet, 9, "Итого Superior",
"=B8+ferry_superior_uplift", "=C8+ferry_superior_uplift", "=D8+ferry_superior_uplift",
"Básica + VIP"); err != nil {
return err
}
if err := addCellComment(f, sheet, "A9", "Butaca Superior / VIP salón. Доплата с листа «Ввод»."); err != nil {
return err
}
if err := f.SetCellStr(sheet, "A10", "Лишние ночи PT"); err != nil {
return err
}
for col, housing := range []string{"housing_7n_low", "housing_7n_mid", "housing_7n_high"} {
cell, _ := excelize.CoordinatesToCellName(col+2, 10)
formula := fmt.Sprintf("=%s*%s/7", extraNightsName, housing)
if err := f.SetCellFormula(sheet, cell, formula); err != nil {
return err
}
}
if err := f.SetCellStr(sheet, "E10", "ночей × (крыша/7)"); err != nil {
return err
}
if err := addCellComment(f, sheet, "A10", "Платное жильё до начала 7-дневной брони. Число ночей — на листе «Ввод» для этого сценария."); err != nil {
return err
}
if err := setFormulaRow(f, sheet, 11, "Итого с ночами",
"=B8+B10", "=C8+C10", "=D8+D10", scenarioNote); err != nil {
return err
}
if err := f.SetCellStr(sheet, "A12", "Остаток от потолка"); err != nil {
return err
}
for _, col := range []string{"B", "C", "D"} {
if err := f.SetCellFormula(sheet, col+"12", "=month_cap-C11"); err != nil {
return err
}
}
if err := f.SetCellStr(sheet, "E12", "потолок − mid итого"); err != nil {
return err
}
if err := addCellComment(f, sheet, "C12", "Сколько остаётся от месячного потолка 2500€ после cutover (mid)."); err != nil {
return err
}
if err := f.SetCellStr(sheet, "A13", "Среднее по строкам"); err != nil {
return err
}
for col := 2; col <= 4; col++ {
cell, _ := excelize.CoordinatesToCellName(col, 13)
colL, _ := excelize.CoordinatesToCellName(col, 2)
colH, _ := excelize.CoordinatesToCellName(col, 7)
if err := f.SetCellFormula(sheet, cell, fmt.Sprintf("=AVERAGE(%s:%s)", colL, colH)); err != nil {
return err
}
}
if err := f.SetCellStr(sheet, "E13", "AVG статей 2–7"); err != nil {
return err
}
return styleSheet(f, sheet, SheetInputs)
}
func writeSummarySheet(f *excelize.File) error {
if err := setHeaders(f, SheetSummary,
"Сценарий", "Básica mid", "Superior mid", "Итого mid", "Остаток", "Δ vs вс"); err != nil {
return err
}
rows := []struct {
label, sheet string
}{
{"Вс 6.09 (live)", SheetDom6},
{"Чт 3.09 (live)", SheetJue3},
{"Ср 2.09 (гипотеза)", SheetWedHyp},
}
qs := quoteSheet
for i, r := range rows {
row := i + 2
if err := f.SetCellStr(SheetSummary, fmt.Sprintf("A%d", row), r.label); err != nil {
return err
}
pairs := []struct {
col int
ref string
}{
{2, fmt.Sprintf("%s!C8", qs(r.sheet))},
{3, fmt.Sprintf("%s!C9", qs(r.sheet))},
{4, fmt.Sprintf("%s!C11", qs(r.sheet))},
{5, fmt.Sprintf("%s!C12", qs(r.sheet))},
}
for _, p := range pairs {
cell, _ := excelize.CoordinatesToCellName(p.col, row)
if err := f.SetCellFormula(SheetSummary, cell, "="+p.ref); err != nil {
return err
}
}
cell, _ := excelize.CoordinatesToCellName(6, row)
if err := f.SetCellFormula(SheetSummary, cell, fmt.Sprintf("=D%d-$D$2", row)); err != nil {
return err
}
}
if err := f.SetCellStr(SheetSummary, "A6", "Потолок месяца"); err != nil {
return err
}
if err := f.SetCellFormula(SheetSummary, "D6", "=month_cap"); err != nil {
return err
}
if err := f.SetCellStr(SheetSummary, "A7", "Лучший итого mid"); err != nil {
return err
}
if err := f.SetCellFormula(SheetSummary, "D7", "=MIN(D2:D4)"); err != nil {
return err
}
if err := f.SetCellStr(SheetSummary, "E7", "MIN по сценариям"); err != nil {
return err
}
if err := addCellComment(f, SheetSummary, "D7", "Минимальный mid «Итого с ночами» среди трёх сценариев. Сейчас обычно вс 6.09."); err != nil {
return err
}
return styleSheet(f, SheetSummary, SheetInputs)
}
-68
View File
@@ -1,68 +0,0 @@
package xlspipe
import (
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"github.com/xuri/excelize/v2"
)
func parseCalcFloat(s string) (float64, error) {
s = strings.ReplaceAll(s, ",", "")
s = strings.TrimSpace(s)
var val float64
_, err := fmt.Sscan(s, &val)
return val, err
}
func TestBuildCutoverPortugalWorkbook_Formulas(t *testing.T) {
f, err := BuildCutoverPortugalWorkbook(DefaultCutoverPortugal())
if err != nil {
t.Fatal(err)
}
defer f.Close()
dir := t.TempDir()
path := filepath.Join(dir, "cutover.xlsx")
if err := Save(f, path); err != nil {
t.Fatal(err)
}
st, err := os.Stat(path)
if err != nil || st.Size() < 4096 {
t.Fatalf("xlsx too small: %v", err)
}
opened, err := excelize.OpenFile(path)
if err != nil {
t.Fatal(err)
}
defer opened.Close()
cases := []struct {
sheet, cell string
want float64
tol float64
}{
{SheetDom6, "C8", 1522.89, 0.02},
{SheetJue3, "C8", 1549.09, 0.02},
{SheetWedHyp, "C11", 1891.95, 0.05},
{SheetSummary, "D2", 1522.89, 0.02},
{SheetSummary, "D7", 1522.89, 0.02},
}
for _, tc := range cases {
got, err := opened.CalcCellValue(tc.sheet, tc.cell)
if err != nil {
t.Fatalf("%s!%s calc: %v", tc.sheet, tc.cell, err)
}
val, err := parseCalcFloat(got)
if err != nil {
t.Fatalf("%s!%s parse %q: %v", tc.sheet, tc.cell, got, err)
}
if diff := val - tc.want; diff < -tc.tol || diff > tc.tol {
t.Fatalf("%s!%s = %v want ~%v", tc.sheet, tc.cell, val, tc.want)
}
}
}
-26
View File
@@ -1,26 +0,0 @@
package xlspipe
import (
"fmt"
"strings"
"github.com/xuri/excelize/v2"
)
// Template names for oo docs put-xlsx --template.
const (
TemplateCutoverPortugal = "cutover-portugal"
)
// BuildTemplate returns an xlsx workbook for a known template name.
func BuildTemplate(name string) (*File, error) {
switch strings.ToLower(strings.TrimSpace(name)) {
case TemplateCutoverPortugal, "portugal-cutover", "cutover":
return BuildCutoverPortugalWorkbook(DefaultCutoverPortugal())
default:
return nil, fmt.Errorf("unknown xlsx template %q (try: %s)", name, TemplateCutoverPortugal)
}
}
// File is an alias so cmd/oo can refer to excelize.File without importing excelize in every handler.
type File = excelize.File
-135
View File
@@ -1,135 +0,0 @@
package xlspipe
import (
"fmt"
"os"
"path/filepath"
"strings"
"github.com/xuri/excelize/v2"
)
const commentAuthor = "oo"
// Save writes the workbook to path (creates parent dirs).
func Save(f *excelize.File, path string) error {
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
return err
}
return f.SaveAs(path)
}
// quoteSheet returns an Excel sheet reference safe for formulas ('Name'!A1).
func quoteSheet(name string) string {
return "'" + strings.ReplaceAll(name, "'", "''") + "'"
}
// defineInput registers a named range on sheet!B{row}, sets value, note, optional comment.
func defineInput(f *excelize.File, sheet, name string, row int, value float64, note, comment string) error {
cell := fmt.Sprintf("B%d", row)
if err := f.SetCellFloat(sheet, cell, value, 2, 64); err != nil {
return err
}
if note != "" {
if err := f.SetCellStr(sheet, fmt.Sprintf("C%d", row), note); err != nil {
return err
}
}
ref := fmt.Sprintf("%s!$B$%d", quoteSheet(sheet), row)
if err := f.SetDefinedName(&excelize.DefinedName{Name: name, RefersTo: ref}); err != nil {
return err
}
if comment != "" {
return addCellComment(f, sheet, cell, comment)
}
return nil
}
func addCellComment(f *excelize.File, sheet, cell, text string) error {
return f.AddComment(sheet, excelize.Comment{
Cell: cell,
Author: commentAuthor,
Text: text,
Width: 280,
Height: 120,
})
}
func setHeaders(f *excelize.File, sheet string, headers ...string) error {
for i, h := range headers {
cell, err := excelize.CoordinatesToCellName(i+1, 1)
if err != nil {
return err
}
if err := f.SetCellStr(sheet, cell, h); err != nil {
return err
}
}
return nil
}
func setFormulaRow(f *excelize.File, sheet string, row int, label string, low, mid, high, note string) error {
if err := f.SetCellStr(sheet, fmt.Sprintf("A%d", row), label); err != nil {
return err
}
for col, formula := range []string{low, mid, high} {
cell, err := excelize.CoordinatesToCellName(col+2, row)
if err != nil {
return err
}
if err := f.SetCellFormula(sheet, cell, formula); err != nil {
return err
}
}
if note != "" {
return f.SetCellStr(sheet, fmt.Sprintf("E%d", row), note)
}
return nil
}
func styleSheet(f *excelize.File, sheet, inputsSheet string) error {
widths := map[string]float64{"A": 34, "B": 12, "C": 12, "D": 12, "E": 40}
for col, w := range widths {
if err := f.SetColWidth(sheet, col, col, w); err != nil {
return err
}
}
headerStyle, err := f.NewStyle(&excelize.Style{
Font: &excelize.Font{Bold: true},
Fill: excelize.Fill{Type: "pattern", Color: []string{"#E8F0FE"}, Pattern: 1},
Alignment: &excelize.Alignment{Horizontal: "center", WrapText: true},
})
if err != nil {
return err
}
euroStyle, err := f.NewStyle(&excelize.Style{NumFmt: 4})
if err != nil {
return err
}
inputStyle, err := f.NewStyle(&excelize.Style{
NumFmt: 4,
Fill: excelize.Fill{Type: "pattern", Color: []string{"#FFF9E6"}, Pattern: 1},
})
if err != nil {
return err
}
if err := f.SetCellStyle(sheet, "A1", "E1", headerStyle); err != nil {
return err
}
if sheet == inputsSheet {
return f.SetCellStyle(sheet, "B2", "B30", inputStyle)
}
if err := f.SetCellStyle(sheet, "B2", "D20", euroStyle); err != nil {
return err
}
return nil
}
func finalizeWorkbook(f *excelize.File) error {
if err := f.SetCalcProps(&excelize.CalcPropsOptions{FullCalcOnLoad: boolPtr(true)}); err != nil {
return err
}
return f.UpdateLinkedValue()
}
func boolPtr(v bool) *bool { return &v }
+23 -12
View File
@@ -131,7 +131,7 @@ type UpdateInvoiceParams struct {
// UpdateInvoice PUTs a full invoice body (OnlyOffice requires complete payload). // UpdateInvoice PUTs a full invoice body (OnlyOffice requires complete payload).
// //
// Linking an opportunity via EntityID on an existing invoice often returns HTTP 400 // Linking an opportunity via EntityID on an existing invoice often returns HTTP 400
// on this portal — prefer CreateInvoice with EntityID set. See docs/crm-associations.md. // on this portal — prefer CreateInvoice with EntityID set. See the CRM association rules.
func (c *Client) UpdateInvoice(ctx context.Context, id string, p UpdateInvoiceParams) (map[string]any, error) { func (c *Client) UpdateInvoice(ctx context.Context, id string, p UpdateInvoiceParams) (map[string]any, error) {
inv, err := c.GetInvoice(ctx, id) inv, err := c.GetInvoice(ctx, id)
if err != nil { if err != nil {
@@ -249,7 +249,7 @@ const (
// SetInvoiceStatus sets CRM invoice status for one or more invoice ids // SetInvoiceStatus sets CRM invoice status for one or more invoice ids
// (PUT /api/2.0/crm/invoice/status/{statusId} with invoiceids). // (PUT /api/2.0/crm/invoice/status/{statusId} with invoiceids).
// Note: Billed→Draft often does not stick; recreate Draft instead (see docs/crm-associations.md). // Note: Billed→Draft often does not stick; recreate Draft instead (see the CRM association rules).
func (c *Client) SetInvoiceStatus(ctx context.Context, statusID int, invoiceIDs ...int64) (map[string]any, error) { func (c *Client) SetInvoiceStatus(ctx context.Context, statusID int, invoiceIDs ...int64) (map[string]any, error) {
if statusID <= 0 { if statusID <= 0 {
return nil, fmt.Errorf("status id is required") return nil, fmt.Errorf("status id is required")
@@ -443,21 +443,32 @@ func (c *Client) DeleteInvoiceItem(ctx context.Context, id string) (map[string]a
return c.deleteObject(ctx, fmt.Sprintf("/api/2.0/crm/invoiceitem/%s.json", url.PathEscape(id))) return c.deleteObject(ctx, fmt.Sprintf("/api/2.0/crm/invoiceitem/%s.json", url.PathEscape(id)))
} }
// contactAddress is the JSON payload OO stores in a ContactInfo row of type
// Address (ASC.Api.CRM.Wrappers.Address).
type contactAddress struct {
Street string `json:"street"`
City string `json:"city"`
State string `json:"state"`
Zip string `json:"zip"`
Country string `json:"country"`
}
// AddContactAddress attaches a postal address to a contact. // AddContactAddress attaches a postal address to a contact.
// category: Home|Postal|Office|Billing|Other|Work (or numeric string). // category: Home|Postal|Office|Billing|Other|Work.
//
// OO stores addresses as ContactInfo rows of infoType Address whose `data` is
// the Address object as JSON (the dedicated /addressdata endpoint binds the
// model from the body and is not accepted by all builds). The generic
// /contact/{id}/data endpoint is the one `oo contacts info-add` uses.
func (c *Client) AddContactAddress(ctx context.Context, contactID, street, city, state, zip, country, category string, isPrimary bool) (map[string]any, error) { func (c *Client) AddContactAddress(ctx context.Context, contactID, street, city, state, zip, country, category string, isPrimary bool) (map[string]any, error) {
if category == "" { if category == "" {
category = "Billing" category = "Billing"
} }
fields := url.Values{} payload, err := json.Marshal(contactAddress{Street: street, City: city, State: state, Zip: zip, Country: country})
fields.Set("street", street) if err != nil {
fields.Set("city", city) return nil, err
fields.Set("state", state) }
fields.Set("zip", zip) return c.AddContactInfo(ctx, contactID, "Address", string(payload), category, isPrimary)
fields.Set("country", country)
fields.Set("category", category)
fields.Set("isPrimary", strconv.FormatBool(isPrimary))
return c.postFormObject(ctx, fmt.Sprintf("/api/2.0/crm/contact/%s/address", url.PathEscape(contactID)), fields)
} }
// UpdateCompany updates company name and optional about text. // UpdateCompany updates company name and optional about text.

Some files were not shown because too many files have changed in this diff Show More