Compare commits

...
149 Commits
Author SHA1 Message Date
eSlider 13a4b9f846 Merge pull request 'fix(dav): re-upload converted ext instead of UpdateFile (#84)' (#85) from fix/dav-updatefile-convert#84 into main
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go 1.25) (push) Successful in 1m33s
Tests / Test (Go stable) (push) Successful in 1m34s
2026-09-27 15:12:36 +01:00
eSlider aad6315731 fix(dav): re-upload converted ext instead of UpdateFile (#84)
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m30s
Tests / Test (Go 1.25) (pull_request) Successful in 1m31s
UploadToFolderReplacing matched legacy->OOXML counterparts (f.xls vs the
saved f.xlsx) and updated them in place. UpdateFile replaces the body without
re-running OnlyOffice's server-side conversion, so the raw OLE2 .xls landed
under the stored .xlsx name and the file would not open (#84).

UpdateFile now runs only when the stored and local extensions are equal.
A conversion-equivalent match is deleted (all duplicates collapsed) and
uploaded afresh, so the server converts again. AssertNoFileConflict stays
conversion-aware for --no-replace and never had an UpdateFile path.
2026-09-27 15:12:13 +01:00
eSlider 329c0b9561 Merge pull request 'fix(dav): upload --replace matches server-converted ext (xls→xlsx) (#82)' (#83) from fix/dav-replace-convert#82 into main
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 6s
Tests / Test (Go 1.25) (push) Successful in 1m5s
Tests / Test (Go stable) (push) Successful in 1m13s
2026-09-27 15:06:16 +01:00
eSlider 74e72d0e06 fix(dav): upsert upload matches server-converted ext (xls→xlsx) (#82)
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 1m3s
Tests / Test (Go stable) (pull_request) Successful in 1m11s
OnlyOffice converts legacy binary Office uploads (.xls/.doc/.ppt) into
OOXML (.xlsx/.docx/.pptx) on the server. UploadToFolderReplacing matched
by the exact stem|ext via FindFilesByDedupKey, so a repeated
`oo dav upload FOLDER f.xls --replace` never found the stored f.xlsx and
appended a second file (live: ids 3799+3887, 3800+3888).

- EquivalentUploadExt / FindFilesByStemExt: match by stem with a
  legacy↔OOXML extension equivalence, so .xls finds the saved .xlsx.
- planUploadReplacement: pick the surviving file and the redundant
  duplicate ids for a replacing upload.
- UploadToFolderReplacing updates the existing file in place (UpdateFile,
  stable id, no delete window), removes extra duplicates, and falls back
  to conversion-aware delete + upload when the portal rejects the update.
- AssertNoFileConflict (--no-replace) uses the same conversion-aware
  matching so a .xls upload conflicts with an existing .xlsx.
- Offline tests cover the matcher, the plan and the repeated-.xls
  regression; pdf/xlsx behaviour unchanged.
2026-09-27 15:05:13 +01:00
eSlider 534661dec7 Merge pull request 'feat(dav): oo dav upload (локальный файл) + ensure-path (mkdir -p) (#80)' (#81) from feat/dav-upload#80 into main
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 6s
Tests / Test (Go stable) (push) Successful in 1m10s
Tests / Test (Go 1.25) (push) Successful in 1m15s
2026-09-25 15:08:36 +01:00
eSlider 55356558a1 feat(dav): oo dav upload (local file) + ensure-path (mkdir -p) in Documents (#80)
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go stable) (pull_request) Successful in 1m9s
Tests / Test (Go 1.25) (pull_request) Successful in 1m11s
2026-09-25 15:07:57 +01:00
eSlider aa069bea56 Merge pull request 'chore(release): v0.21.0 (#78)' (#79) from chore/release-v0.21.0#78 into main
Release / GoReleaser (push) Skipped
Tests / Test (Go 1.25) (push) Successful in 1m23s
Tests / Secret scan (gitleaks) (push) Successful in 3s
Tests / Test (Go stable) (push) Successful in 1m28s
2026-09-23 18:25:08 +01:00
eSlider a8fa202496 chore(release): v0.21.0 (#78)
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 1m29s
Tests / Test (Go stable) (pull_request) Successful in 1m33s
2026-09-23 18:25:00 +01:00
eSlider 87ff317166 Merge pull request 'chore: move business tooling out of the public library; tidy filestore naming' (#77) from chore/split-business-tools into main
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go 1.25) (push) Successful in 1m25s
Tests / Test (Go stable) (push) Successful in 1m28s
2026-09-23 09:05:48 +01:00
eSlider dc62570a5c chore: move business tooling out of the public library; tidy filestore naming
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m19s
Tests / Test (Go 1.25) (pull_request) Successful in 1m24s
Keep the public tree project-generic. Business/one-off tools, deployment
and business docs move to the private oo-workspace repo.

Moved to oo-workspace:
- cmd/ooscan, cmd/pdfamount, cmd/kontoblatt, cmd/kontolink
- internal/xlspipe (cutover-portugal workbook) -> oow workbook build
  (drops the --template/--title flags from oo docs put-xlsx)
- deploy/docker-compose.rclone-webdav.yml + docs/rclone-webdav.md
- docs/crm-associations.md

Removed GitHub-era leftovers:
- .github/workflows/release-please.yml, release-please-config.json,
  .release-please-manifest.json (tags are created on Gitea per SemVer)

Naming: the FileStore subsystem is now filestore_*.go (was file_*.go) to
match files_*.go (project Documents). Docs/AGENTS/README updated.
2026-09-23 09:05:40 +01:00
eSlider a9a77e110d Merge pull request 'feat: upstream generic workspace tooling (board-sync, crm audit, catalog names)' (#76) from feat/upstream-oow-workspace into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 1m23s
Tests / Test (Go 1.25) (push) Successful in 1m23s
2026-09-22 23:13:49 +01:00
eSlider 82da45dc4d feat: upstream generic workspace tooling (board-sync, crm audit, catalog names)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 1m16s
Tests / Test (Go stable) (pull_request) Successful in 1m36s
Promote the generic, reusable parts of the private oo-workspace into the
public library/CLI, so oo-workspace can shrink to business glue.

- board.go: Board types + (c *Client) SyncBoard — upsert project
  milestones/tasks by exact title from a YAML board (dry-run when
  apply=false). CLI: oo projects board-sync.
- crm_audit.go: OpportunityAudit + (c *Client) AuditOpportunities —
  file/task/member counts with a generic class (ok|dup|empty|junk-title).
  CLI: oo crm audit [--out].
- catalog/names.go: CleanPersonNames / GuessNameFromEmail /
  FormatProjectTitle (ported from oo-workspace).
- catalog: adopt the newer apply/match logic — name repair (never encode
  company in lastName), company grouping key, contact-info type
  normalization, preserve an already-applied oo_id. Keeps the local
  Address/OOProjects fields and the config-driven classifier.

Business-only oo-workspace commands (crm clean/migrate, folders, search)
deliberately stay private.
2026-09-22 23:13:40 +01:00
eSlider 8848ee00f1 Merge pull request 'feat(oo): native conversion, deep links, sheet export (re-land)' (#75) from feat/oo-conversion-links into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go 1.25) (push) Successful in 1m29s
Tests / Test (Go stable) (push) Successful in 1m30s
2026-09-22 22:42:07 +01:00
eSlider ff46aad615 chore: keep the re-landed feature files showroom-safe
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m21s
Tests / Test (Go 1.25) (pull_request) Successful in 1m27s
The conversion / link / sheet commits predate the tree cleanup, so their
comments and README/test examples still carried internal refs. Re-apply
the genericization (no internal hosts, personal names or client file ids).
2026-09-22 22:41:58 +01:00
eSlider 5e73f6c335 feat(oo): sheet-aware spreadsheet export (docs csv/json)
- lib: WorkbookSheetNames/CSV/JSON (excelize) — sheet selection + JSON rows
- oo docs csv  SRC... [--sheet N] [--delimiter ,|;|||tab] [--out]
- oo docs json SRC... [--sheet N] [--out]
- SRC = OO file id or local path; --sheet is 1-based (default first)
- DS csv output is first-sheet-only (verified) -> local reader for selection/JSON
- test: TestWorkbookSheetExport
2026-09-22 22:41:14 +01:00
eSlider 7c089a6439 feat(oo): docs pdf --stream/--pipe (converted bytes to stdout) 2026-09-22 22:41:14 +01:00
eSlider 4515a88fe8 docs(readme): examples for links, conversion, team/users CRUD 2026-09-22 22:41:14 +01:00
eSlider 58b905cc49 feat(oo): native document conversion (docs pdf/presigned)
- lib: PresignedURI, SignJWT (HS256, stdlib), ConvertDocument (DocumentServer
  /converter; legacy /ConvertService.ashx), DownloadURLTo
- oo docs pdf FILE_ID|PATH...: OO file ids or local files (temp upload to
  --folder, convert, download, cleanup) -> PDF/other
- oo docs presigned FILE_ID
- docs base from $ONLYOFFICE_DOCS_URL else $ONLYOFFICE_URL/ds-vpath; JWT secret
  from $ONLYOFFICE_DS_SECRET (DocumentServer CoAuthoring secret)
- test: SignJWT
2026-09-22 22:41:14 +01:00
eSlider 520d6d05c8 docs(oo): comment where the new helpers are used
- link.go/links.go: third-party doc packs (fileid links), id-stability caveat
- auth.go AuthenticateAs / users check: email-vs-userName login finding
- projects files replace-in: version-history rationale
2026-09-22 22:41:14 +01:00
eSlider df38b75011 feat(oo): file deep links, login check, replace-in (fresh upload)
- link: print Products/Files/DocEditor.aspx?fileid=… deep links (title+url)
- lib: FileEditorURL/FolderURL helpers (+tests)
- users check: verify credentials (userName vs email) via authentication.json
- projects files replace-in FOLDER_ID FILE...: hard delete same stem|ext in a
  folder + fresh upload → single clean version (no version history)
2026-09-22 22:41:14 +01:00
eSlider ee8e88fb9b feat(oo): projects files update (overwrite existing file content) 2026-09-22 22:41:14 +01:00
eSlider e323a43317 feat(oo): projects milestone-delete 2026-09-22 22:41:14 +01:00
eSlider 71d421fb41 Merge pull request 'chore: keep host/client specifics out of the tree (env & config)' (#74) from chore/showroom-generic into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 55s
Tests / Test (Go stable) (push) Successful in 1m6s
2026-09-22 22:40:49 +01:00
eSlider afb93feae5 chore: keep host/client specifics out of the tree (env & config)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 59s
Tests / Test (Go stable) (pull_request) Successful in 1m7s
Showroom-safe: the tree no longer carries internal hosts, IPs, ports,
personal names, client domains or client file names. Behaviour is
unchanged and now supplied per deployment.

- catalog: hardcoded mail-org / project classifiers become a YAML-driven
  Classifier (OO_CATALOG_CONFIG or --config); neutral default classifies
  nothing as work. New catalog/classify.go + example + tests.
- storage_fallback: drop the baked-in MinIO endpoint IP; require
  MINIO_ENDPOINT (+ keys) from the env.
- kontolink: build DocEditor links from $ONLYOFFICE_URL instead of a
  hardcoded portal host; kontoblatt: no client file id in the output name.
- oo: load .env CLI-wide (bootstrap.LoadEnv in execute) so
  non-authenticating commands (catalog scans) also see config.
- genericize comments/docs/fixtures (AGENTS, README, .env.example,
  crm-associations, catalog tests, ES/pdfattach tests, mails).
2026-09-22 22:03:36 +01:00
eSlider b493472f2e Merge pull request 'feat(oo): project team CRUD and user lifecycle' (#73) from feat/oo-team-users-crud into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 1m1s
Tests / Test (Go 1.25) (push) Successful in 1m2s
Reviewed-on: #73
2026-09-22 14:29:47 +01:00
eSlider 1dc99d6997 feat(oo): project team CRUD and user lifecycle (block/unblock/password/delete)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 1m2s
Tests / Test (Go stable) (pull_request) Successful in 1m40s
- project team: oo projects team list|add|remove|set (portal users)
- users: get|create|update|delete|block|unblock|password
- lib: CreateUser, DeleteUser (auto-suspend before delete), Add/Remove/
  SetProjectTeam, ListProjectTeam; deleteJSON reused; jsonBodyReader helper
- tests: unmarshalResponseArray
2026-09-22 14:26:23 +01:00
eSlider 9b21d7e52a Merge pull request 'fix(crm): адреса контактов через ContactInfo типа Address (#288)' (#72) from feat/catalog-addresses#288 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 58s
Tests / Test (Go 1.25) (push) Successful in 59s
2026-09-22 08:32:25 +01:00
eSlider c4beba5173 fix(crm): store contact addresses via ContactInfo Address type (#288)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m1s
Tests / Test (Go 1.25) (pull_request) Successful in 1m1s
2026-09-22 07:08:01 +00:00
eSlider 2a187ae132 Merge pull request 'fix(retry): глобальный rate-limit + Retry-After + cooldown (#70)' (#71) from fix/oo-backoff#70 into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 21s
Tests / Test (Go 1.25) (push) Successful in 22s
2026-09-17 13:39:56 +01:00
eSlider c9e16c7169 fix(retry): глобальный rate-limit + Retry-After + cooldown (#70)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 25s
Tests / Test (Go stable) (pull_request) Successful in 27s
2026-09-17 12:39:25 +00:00
eSlider 194ae62720 Merge pull request 'docs: убрать потребительские детали match из index-and-search (#34)' (#69) from docs/index-and-search-scope into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 26s
Tests / Test (Go 1.25) (push) Successful in 26s
2026-09-17 08:51:45 +01:00
eSlider 0d73146ebc docs: убрать потребительские детали match из index-and-search (#34)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 3s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
Tests / Test (Go stable) (pull_request) Successful in 29s
2026-09-17 07:51:02 +00:00
eSlider 46661f51f8 Merge pull request 'docs: карта поиска и обновления индексов (#34)' (#68) from docs/index-and-search into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 25s
Tests / Test (Go 1.25) (push) Successful in 29s
2026-09-17 08:49:31 +01:00
eSlider f22d714be7 docs: карта поиска и обновления индексов (#34)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
Tests / Test (Go stable) (pull_request) Successful in 21s
2026-09-17 07:48:40 +00:00
eSlider 6f85c206cf Merge pull request 'test(files): live-тест файлового dedup (#63)' (#65) from test/w6-file-dedup#63 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Test (Go stable) (push) Failing after 3s
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 28s
2026-09-16 23:03:23 +01:00
eSlider 1c00edbd0b Merge pull request 'test(files): live CRUD папок + фасадный CRUD (#62)' (#66) from test/w5-folder-crud#62 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Test (Go stable) (push) Failing after 4s
Tests / Secret scan (gitleaks) (push) Successful in 6s
Tests / Test (Go 1.25) (push) Successful in 23s
2026-09-16 23:03:17 +01:00
eSlider 9cd157c59f test(files): live CRUD папок + фасадный CRUD (#62)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Test (Go stable) (pull_request) Failing after 4s
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
2026-09-16 21:49:10 +00:00
eSlider 932cd2895a fix(files): REST FileStore resolves folders for stat/rename/move/delete (#62) 2026-09-16 21:49:08 +00:00
eSlider 5588c2d3f7 fix(files): ретраить transient-ответы при удалении файлов (#63)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 20s
Tests / Test (Go 1.25) (pull_request) Successful in 20s
DeleteDavItems (через ApplyDedupGroups/DeleteFilesByDedupKey) ходил в
deleteJSON напрямую и падал на 429/502/503/504. Теперь каждый delete идёт
через DoRetry, как остальные bulk-пути.
2026-09-16 21:47:37 +00:00
eSlider f8783b86a8 fix(files): пропускать пустой dedup-ключ (dotfiles) (#63)
Файлы без stem (.env, .gitignore, .npmrc) давали FileDedupKey="", и
findWithinFolderDuplicates/findCrossFolderDuplicates считали их одной
группой дублей и удаляли лишние. Пустой ключ больше не образует группу.
2026-09-16 21:47:36 +00:00
eSlider d9db1c5e56 test(files): live-тест файлового dedup (#63)
TestIntegrationFileDedup: throwaway-проект, две папки + root, дубли по
stem|ext, FindProjectDuplicates/mergeProjectRootForDedupe,
ApplyDedupGroups/DeleteFilesByDedupKey, survivor и удаления подтверждены
polling-ом. Папки/аплоады ретраят временный post-create 500 портала.
2026-09-16 21:47:33 +00:00
eSlider c551634e55 Merge pull request 'docs: Testing + rclone/SQL + docs index (#53)' (#64) from docs/w1-w3#53 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 19s
Tests / Test (Go 1.25) (push) Successful in 24s
2026-09-16 22:41:12 +01:00
eSlider a88fd5b40c docs: Testing section, rclone/SQL refs, docs index (#53)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
Tests / Test (Go stable) (pull_request) Successful in 22s
2026-09-16 21:39:41 +00:00
eSlider 9b149281a8 Merge pull request 'fix(retry): починка красного, полный прогон тестов (#57)' (#61) from chore/test-green#57 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 23s
Tests / Test (Go stable) (push) Successful in 24s
2026-09-16 22:37:05 +01:00
eSlider d6b8eb7777 fix(retry): retry transient edge answers centrally, fix stale task test (#57)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 21s
Tests / Test (Go stable) (pull_request) Successful in 21s
- Route HTTP helpers (getJSON/formRequest/deleteReq/postJSON/putJSON/
  multipart upload), Query and AuthenticateContext/ensureToken through the
  deterministic 429/502/503/504 retry (retryRaw), so the integration suite
  no longer fails on the shared openresty rate limit under parallel runs.
- Fix stale cmd/office/fetch integration test: loader.TaskFields was
  replaced by loader.DetailForm (broke go vet -tags=integration).
- Add unit tests for Transient/DoRetry.
2026-09-16 21:32:14 +00:00
eSlider 8a17a32269 Merge pull request 'feat(deploy): rclone WebDAV mount (docker compose) + smoke + docs (#54)' (#59) from feat/rclone-webdav#54 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 8s
Tests / Test (Go 1.25) (push) Successful in 25s
Tests / Test (Go stable) (push) Successful in 27s
2026-09-16 22:30:16 +01:00
eSlider 43a1a47faa Merge pull request 'test(es): доказать, что oo search идёт в Elasticsearch (#56)' (#58) from feat/es-integration#56 into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 26s
Tests / Test (Go stable) (push) Successful in 28s
2026-09-16 22:30:14 +01:00
eSlider 620505cb24 Merge pull request 'feat(files): SQL backend via Client.FileStore + live MySQL facade test (#55)' (#60) from feat/db-integration#55 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 30s
Tests / Test (Go 1.25) (push) Successful in 37s
2026-09-16 22:30:05 +01:00
eSlider a6bba30438 feat(files): SQL file backend via Client.FileStore/SQLFileStore + live MySQL integration (#55)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 24s
Tests / Test (Go 1.25) (pull_request) Successful in 26s
2026-09-16 21:26:15 +00:00
SE 6549fbf7e2 feat(deploy): rclone WebDAV mount compose + smoke + docs (#54)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 20s
Tests / Test (Go stable) (pull_request) Successful in 22s
2026-09-16 21:25:38 +00:00
eSlider 1a8bf0f83c test(es): prove oo search uses the Elasticsearch backend (#56)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 3s
Tests / Test (Go stable) (pull_request) Successful in 24s
Tests / Test (Go 1.25) (pull_request) Successful in 25s
2026-09-16 21:22:43 +00:00
eSlider 046882b8ef Merge pull request 'feat(search): полный уникальный path первым + ближайшая папка в результатах (#51)' (#52) from feat/search-path#51 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 22s
Tests / Test (Go 1.25) (push) Successful in 22s
2026-09-16 22:19:46 +01:00
Yftyr a3ef83a961 feat(search): full unique path first + immediate folder in results (#51)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 21s
Tests / Test (Go stable) (pull_request) Successful in 26s
2026-09-16 21:19:32 +00:00
eSlider 54f825f8e8 Merge pull request 'feat(search): подстрока/AND, scope по поддереву, лимит 1000 (#49)' (#50) from fix/es-substring#49 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 24s
Tests / Test (Go 1.25) (push) Successful in 26s
2026-09-16 22:16:01 +01:00
Yftyr 6a4940ba82 feat(search): substring/AND terms, nested folder scope, limit 1000 (#49)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 26s
Tests / Test (Go 1.25) (pull_request) Successful in 39s
2026-09-16 21:15:44 +00:00
eSlider 1da368fd5d Merge pull request 'fix(docpipe): индекс вложений .yaml и без расширения (#47)' (#48) from fix/attachments-yaml#47 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 26s
Tests / Test (Go 1.25) (push) Successful in 27s
2026-09-16 19:54:20 +01:00
Yftyr 66a6b57401 fix(docpipe): index .yaml and extensionless PDF attachments (#47)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 21s
Tests / Test (Go 1.25) (pull_request) Successful in 22s
2026-09-16 18:54:09 +00:00
eSlider 075c66d4d7 Merge pull request 'docs: единый контракт файлового клиента (#39)' (#46) from docs/file-client#39 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 6s
Tests / Test (Go stable) (push) Successful in 21s
Tests / Test (Go 1.25) (push) Successful in 1m20s
2026-09-16 18:37:36 +01:00
eSlider c59ad645ea docs(files): add unified file client contract (#39)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 17s
Tests / Test (Go stable) (pull_request) Successful in 1m20s
2026-09-16 17:37:15 +00:00
eSlider c76ab7267d Merge pull request 'feat(search): PDF-контент в поиске — свой индекс oo_docs_text (#42)' (#44) from feat/pdf-content#42 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 1m21s
Tests / Test (Go 1.25) (push) Successful in 1m23s
2026-09-16 18:34:54 +01:00
eSlider 3cc288d281 feat(search): index embedded PDF attachment text (#42)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 25s
Tests / Test (Go 1.25) (pull_request) Successful in 28s
2026-09-16 17:34:04 +00:00
eSlider ce4778bdf1 fix(files): dedupe ProviderPG after facade/SQL merge, rename test fake 2026-09-16 17:29:01 +00:00
eSlider 3fe43ee82c feat(search): PDF content via own ES index and oo index (#42) 2026-09-16 17:28:05 +00:00
eSlider 6d93ab5b1f Merge pull request 'feat(files): read-only SQL file store (Community Server DB) (#36)' (#45) from feat/pg-store#36 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 3s
Tests / Test (Go stable) (push) Failing after 26s
Tests / Test (Go 1.25) (push) Failing after 28s
2026-09-16 18:27:21 +01:00
eSlider 95d79925ea Merge pull request 'refactor(files): единый файловый фасад + миграция CLI/TUI (#38)' (#43) from feat/file-facade#38 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Failing after 37s
Tests / Test (Go stable) (push) Failing after 38s
2026-09-16 18:27:21 +01:00
eSlider bf3ef025aa feat(files): read-only SQL file store over Community Server DB (#36)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 1m56s
Tests / Test (Go stable) (pull_request) Successful in 1m57s
2026-09-16 17:24:22 +00:00
eSlider ba29738482 refactor(files): single file client facade + CLI/TUI migration (#38)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 19s
Tests / Test (Go 1.25) (pull_request) Successful in 22s
- FileClient composes FileStore/Searcher backends; Read()/Write()/Search()
  select REST/DAV/PG/ES with transient read fallback. Client.Files() returns
  it and *FileClient implements FileStore, so existing callers keep working.
- Entry gains backend-native Updated + folder FilesCount/FoldersCount so
  dav ls output round-trips.
- cmd/oo dav/projects files and cmd/office/fetch download/preview/delete go
  through FileStore; oo search goes through the facade.
- Mark transport methods that FileStore now abstracts as deprecated.
2026-09-16 17:20:41 +00:00
eSlider 72c7cd5173 Merge pull request 'feat(search): Elasticsearch-клиент (имя + контент) + oo search (#37)' (#41) from feat/es-search#37 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 3s
Tests / Test (Go stable) (push) Successful in 17s
Tests / Test (Go 1.25) (push) Successful in 24s
2026-09-16 17:43:21 +01:00
Yftyr b3dfff1a4c refactor(search): use canonical model from file_core.go (#37) 2026-09-16 16:43:20 +00:00
eSlider b5a61ad420 feat(search): Elasticsearch searcher (name+content) and oo search (#37) 2026-09-16 16:42:57 +00:00
eSlider 1aba545e1d Merge pull request 'feat(files): канонический Entry + FileStore (REST/DAV) (#35)' (#40) from feat/file-store#35 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 3s
Tests / Test (Go stable) (push) Successful in 18s
Tests / Test (Go 1.25) (push) Successful in 18s
2026-09-16 17:42:49 +01:00
eSlider 94951fc6c5 feat(files): canonical Entry + FileStore REST/DAV adapters (#35)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 18s
Tests / Test (Go stable) (pull_request) Successful in 18s
2026-09-16 16:30:07 +00:00
eSlider d95985e6d0 Merge pull request 'fix(pdfamount): ставка НДС не сумма; итог DKV (#29)' (#33) from fix/pdfamount-tax#29 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go 1.25) (push) Successful in 19s
Tests / Test (Go stable) (push) Successful in 20s
2026-09-15 12:58:30 +01:00
eSlider 57b383d579 fix(pdfamount): не считать ставку НДС суммой; итог DKV (#29)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 20s
Tests / Test (Go stable) (pull_request) Successful in 23s
- Строки с %/MwSt/USt/Prozent/Steuer больше не дают сумму (было: Storchen
  19,00 % -> amount 19.00 и ложная отбраковка кандидатов в матчере).
- Приоритет меток: zu zahlender betrag > rechnungsbetrag > rechnungsendbetrag
  > gesamtbetrag > gesamtsumme (inkl. Steuern); low-priority не перебивает.
- DKV: секция Gesamtsummenaufstellung (значение после «»») важнее повторяющихся
  TOTAL-строк по машинам.
- Тесты: 19,00 % не сумма; fallback на usable-метку; DKV-итог.
2026-09-15 11:57:31 +00:00
eSlider a73f6c172d Merge pull request 'feat(oo): dav rm — удаление папок/файлов Documents (#30)' (#31) from feat/dav-rm#30 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 7s
Tests / Test (Go stable) (push) Successful in 29s
Tests / Test (Go 1.25) (push) Successful in 42s
2026-09-15 09:44:02 +01:00
eSlider e49693ad2c feat(oo): dav rm — удаление папок/файлов Documents (#30)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 6s
Tests / Test (Go stable) (pull_request) Successful in 22s
Tests / Test (Go 1.25) (pull_request) Successful in 25s
DeleteDavItems был только в клиенте; добавлена команда
`oo dav rm [FILE_ID...] --folders F1,F2`.
2026-09-15 08:31:47 +00:00
eSlider b86d71211f merge: GitHub release-please v0.18+ into Gitea main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 23s
Tests / Test (Go 1.25) (push) Successful in 30s
2026-09-15 09:29:13 +01:00
eSlider a2b919478e Merge pull request 'fix(pdfamount): суммы DKV/Diashop и др. (метки и форматы) (#27)' (#28) from fix/pdfamount-amounts#27 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 18s
Tests / Test (Go stable) (push) Successful in 19s
2026-09-15 08:55:40 +01:00
eSlider ccabc11153 fix(pdfamount): widen amount labels and formats (#27)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 19s
Tests / Test (Go stable) (pull_request) Successful in 18s
Match more payable-amount labels by priority and accept DE/EN/plain
number formats, normalising to a dot-separated 2-decimal string. Adds
unit tests with synthetic DKV-style and Diashop-style lines.
2026-09-15 07:51:23 +00:00
eSlider 9a6da1a4f7 Merge pull request 'fix(files): UpdateFile uses PUT (#25)' (#26) from fix/update-file-put#25 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 18s
Tests / Test (Go 1.25) (push) Successful in 23s
2026-09-15 08:22:50 +01:00
eSlider 504d13ed08 fix(files): UpdateFile uses PUT /api/2.0/files/{id}/update (#25)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 18s
Tests / Test (Go 1.25) (pull_request) Successful in 20s
2026-09-15 07:19:42 +00:00
eSlider efc0864d42 Merge pull request 'docs(oo): dav, documents files api, bulk tools, fix verbs (#22)' (#24) from docs/oo-reference#22 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 19s
Tests / Test (Go 1.25) (push) Successful in 19s
Reviewed-on: #24
2026-09-14 22:44:25 +01:00
eSlider 59caf5e560 docs(oo): dav, documents files api, bulk tools, fix verbs (#22)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 17s
Tests / Test (Go stable) (pull_request) Successful in 18s
2026-09-14 21:43:37 +00:00
eSlider e7803c5269 Merge pull request 'feat(files): Documents Dav ops + UpdateFile, fileops errors (#152)' (#20) from feat/oo-automation#152 into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 14s
Tests / Test (Go 1.25) (push) Successful in 1m6s
2026-09-14 22:35:09 +01:00
eSlider a08e7c49ad Merge branch 'main' into feat/oo-automation#152
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m49s
Tests / Test (Go 1.25) (pull_request) Successful in 2m20s
2026-09-14 22:31:44 +01:00
eSlider 4310aa7002 Merge pull request 'feat(files): MinIO fallback for stale S3 downloads (#152)' (#21) from feat/oo-minio-fallback#152 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 12s
Tests / Test (Go 1.25) (push) Successful in 23s
Tests / Test (Go stable) (push) Successful in 2m19s
Reviewed-on: #21
2026-09-14 22:31:32 +01:00
eSlider 1a7a962183 Merge pull request 'feat(oo): kontoblatt bulk tools, oo dav, update, retry (#22)' (#23) from feat/kontoblatt-tools#22 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go 1.25) (push) Successful in 16s
Tests / Test (Go stable) (push) Successful in 24s
Reviewed-on: #23
2026-09-14 22:31:22 +01:00
eSlider af8e3b053a feat(oo): kontoblatt bulk tools, oo dav, update, retry (#22)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 1m22s
Tests / Test (Go stable) (pull_request) Successful in 1m21s
2026-09-14 21:03:39 +00:00
eSlider 8ac777c031 feat(files): MinIO fallback for stale S3 downloads (#152)
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m15s
Tests / Test (Go 1.25) (pull_request) Successful in 1m22s
2026-09-14 15:41:39 +00:00
eSlider c576bfe2ee feat(files): Documents Dav ops + UpdateFile, fileops errors (#152)
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 1m17s
Tests / Test (Go 1.25) (pull_request) Successful in 1m20s
2026-09-14 12:37:57 +00:00
eSliderandGitHub 47a2c256f3 Merge pull request #48 from eSlider/release-please--branches--main--components--go-onlyoffice
chore(main): release 0.18.0
2026-09-14 09:18:11 +01:00
github-actions[bot]andGitHub c1881e1e61 chore(main): release 0.18.0 2026-09-04 14:47:27 +00:00
eSliderandGitHub 7f56332285 Merge pull request #49 from eSlider/sync/gitea-main-20260904
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 2m4s
Tests / Test (Go stable) (push) Successful in 2m4s
sync: gitea main (sortBy fix #18 + funding)
2026-09-04 15:47:06 +01:00
eSlider a5807b9c31 Merge pull request 'docs(funding): eSlider support links (reverse-import GitHub)' (#19) from docs/funding-github into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 13s
Tests / Test (Go stable) (push) Successful in 30s
Tests / Test (Go 1.25) (push) Successful in 32s
2026-09-04 12:31:47 +01:00
eSlider d650a16a36 docs(funding): eSlider support links (reverse-import GitHub e9c969a)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go stable) (pull_request) Successful in 34s
Tests / Test (Go 1.25) (pull_request) Successful in 37s
2026-09-04 12:29:45 +01:00
eSlider 3da11586f9 Merge pull request 'fix(crm): детерминированный sortBy=id в постраничных списках контактов' (#18) from fix/crm-filter-sortby into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 10s
Tests / Test (Go 1.25) (push) Successful in 34s
Tests / Test (Go stable) (push) Successful in 1m38s
2026-09-03 13:14:36 +01:00
eSlider e9c969a89c funding: eSlider support links (sponsors, ko-fi, liberapay, patreon, polar) 2026-09-03 05:20:18 +01:00
eSlider 4d91726179 fix(crm): deterministic sortBy=id in contact paged lists
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 26s
Tests / Test (Go stable) (pull_request) Successful in 1m21s
2026-09-03 04:50:15 +01:00
eSlider 4d8af7a2fe feat(crm): add UpdateContactName and CloseCRMTask helpers
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 6s
Tests / Test (Go stable) (push) Successful in 28s
Tests / Test (Go 1.25) (push) Successful in 38s
- UpdateContactName renames a company contact via PUT /crm/contact/company/{id}
- CloseCRMTask closes a CRM task via PUT /crm/task/{id}/close.json
2026-08-31 11:38:50 +01:00
eSlider 34745e349e chore(release): trigger CI (#16)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 26s
Tests / Test (Go stable) (pull_request) Successful in 28s
2026-08-31 11:38:48 +01:00
eSliderandGitHub dfea57a57b Merge pull request #46 from eSlider/release-please--branches--main--components--go-onlyoffice
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 23s
Tests / Test (Go stable) (pull_request) Successful in 26s
Release / GoReleaser (push) Skipped
chore(main): release 0.17.0
2026-08-31 11:34:41 +01:00
github-actions[bot]andGitHub 8595c17f25 chore(main): release 0.17.0 2026-08-31 10:34:21 +00:00
eSlider 61da2fb88b Merge pull request 'fix(files): upsert uploads by default + dedupe project root (#45)' (#15) from fix/docs-upsert-dedupe into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 24s
Tests / Test (Go 1.25) (push) Successful in 31s
2026-08-31 11:33:51 +01:00
eSliderandCursor 24ca144b22 fix(files): upsert uploads by default and dedupe project root
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go stable) (pull_request) Successful in 25s
Tests / Test (Go 1.25) (pull_request) Successful in 27s
OnlyOffice allows duplicate stem|ext in the same folder; agents hit this
via projects files upload and put-md without --folder. Default all upload
paths to replace-by-stem, add no-clobber via --no-replace, and scan
projectFolder in dedupe.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 11:18:54 +01:00
eSlider 93828ee19d Merge pull request 'feat(mailsync): FetchMailFolder — integration-layer mail walk for ETL consumers (2dph)' (#4) from feat/2dph-mail-integration-layer into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go 1.25) (push) Successful in 29s
Tests / Test (Go stable) (push) Successful in 35s
2026-08-31 10:08:19 +01:00
mdx-1andeSlider 35f0cb8d20 feat(mailsync): FetchMailFolder — integration-layer walk for ETL consumers
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 24s
Tests / Test (Go stable) (pull_request) Successful in 40s
Adds the high-level mail folder walk that sync pipelines need on top of
the raw mail API (list -> get -> download-attachment), so consumers stop
re-implementing it against private client copies.

  type MailSyncMessage struct { ID, Folder, Subject, From, Date, IsNew,
                                HasAttachments, Attachments }
  type MailSyncAttachment struct { ID, Name, Size, Body }
  func (c *Client) FetchMailFolder(ctx, folderID, MailSyncOptions)
                                   ([]MailSyncMessage, error)

Options: Limit / StartIndex for checkpointed walks, FetchBodies to
eagerly download attachment bytes via download.ashx (session-cookie path).

Hydration details:
- list items may omit the attachment array; when hasAttachments is set
  the full record is fetched and its attachments merged
- attachment ids accepted from id/fileId/attachmentId variants
- timestamps parsed from RFC3339 (any fractional digits) and
  second-precision forms

This is the first step of the 2dph integration layer (#1): the brain's
mail-ingest pipeline can now drop its private OOClient copy and consume
this canonical walk directly.

Tests: httptest-backed coverage for pagination, hydration with
full-record fallback, body download incl. auth-cookie requirement,
Limit/StartIndex windows, timestamp parsing.
2026-08-31 10:05:47 +01:00
eSlider ca5c85a4e2 Merge pull request 'chore(release): v0.16.0 — put-xlsx (excelize) (#13)' (#14) from chore/release-v0.16.0#13 into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 23s
Tests / Test (Go 1.25) (push) Successful in 24s
2026-08-30 13:44:47 +01:00
eSlider 220c7e8265 chore(release): v0.16.0 — put-xlsx (excelize) (#13)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 8s
Tests / Test (Go 1.25) (pull_request) Successful in 2m3s
Tests / Test (Go stable) (pull_request) Successful in 2m5s
2026-08-30 13:41:51 +01:00
eSlider ff4c0481ba Merge pull request 'feat(docs): put-xlsx с excelize — multi-sheet бюджеты (cutover-portugal)' (#12) from feat/docs-put-xlsx into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go 1.25) (push) Successful in 2m10s
Tests / Test (Go stable) (push) Successful in 2m12s
2026-08-30 13:41:25 +01:00
eSliderandCursor 68445b0dfd feat(docs): put-xlsx with excelize cutover workbook
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go 1.25) (pull_request) Successful in 1m59s
Tests / Test (Go stable) (pull_request) Successful in 2m5s
Add internal/xlspipe (github.com/xuri/excelize/v2) and oo docs put-xlsx
to generate multi-sheet budgets with named inputs, SUM/AVG/MIN formulas,
cross-sheet links, Russian labels, and cell comments for OnlyOffice.

Template cutover-portugal: sheets Ввод, Вс 6.09, Чт 3.09, Ср гип, Сводка.
Upload via UploadProjectFileReplacing (--replace default).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-30 13:38:27 +01:00
eSlider 8739d45b1a Merge pull request 'chore(release): v0.15.0 — files folder-ops + docpipe (#8)' (#9) from chore/release-v0.15.0#8 into main
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Test (Go 1.25) (push) Successful in 23s
Tests / Secret scan (gitleaks) (push) Successful in 5s
Tests / Test (Go stable) (push) Successful in 26s
2026-08-29 16:19:41 +01:00
eSlider f3150436fd chore(release): rerun CI после gitleaks fix (#8)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 23s
Tests / Test (Go stable) (pull_request) Successful in 25s
2026-08-29 16:18:42 +01:00
eSlider a30ae20f32 Merge pull request 'fix(ci): gitleaks через бинарник (Gitea runner) (#10)' (#11) from fix/gitleaks-gitea#10 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 6s
Tests / Test (Go 1.25) (push) Successful in 21s
Tests / Secret scan (gitleaks) (push) Successful in 4s
Tests / Test (Go stable) (push) Successful in 27s
Tests / Test (Go 1.25) (pull_request) Successful in 27s
Tests / Test (Go stable) (pull_request) Successful in 31s
2026-08-29 16:14:51 +01:00
eSlider 5a0b2ab412 fix(ci): gitleaks через бинарник на $GITHUB_WORKSPACE (Gitea runner) (#10)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 5s
Tests / Test (Go stable) (pull_request) Successful in 21s
Tests / Test (Go 1.25) (pull_request) Successful in 30s
2026-08-29 16:14:00 +01:00
eSlider 69844c9122 chore(release): v0.15.0 — files folder-ops + docpipe (#8)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Test (Go 1.25) (pull_request) Successful in 25s
Tests / Test (Go stable) (pull_request) Successful in 34s
Tests / Secret scan (gitleaks) (pull_request) Failing after 3s
2026-08-29 16:06:47 +01:00
eSliderandGitHub b04d610275 Merge pull request #42 from eSlider/feat/docs-optimize-gs
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 4s
Tests / Test (Go 1.25) (push) Successful in 20s
Tests / Test (Go stable) (push) Successful in 27s
feat(docs): Ghostscript PDF optimize for OO
2026-08-28 10:03:04 +01:00
eSliderandCursor 6b3f40e1a1 feat(docs): optimize PDF via Ghostscript pdfwrite
Add oo docs optimize for native-text PDFs (OO-friendly rewrite without
ocrmypdf overlay). Default ocr --force to false when text layer exists.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-28 10:02:58 +01:00
eSliderandGitHub ba8148a9e1 Merge pull request #41 from eSlider/fix/put-txt-fixed-width
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 4s
Tests / Test (Go 1.25) (push) Successful in 20s
Tests / Test (Go stable) (push) Successful in 27s
fix(docs): fixed-width txt as monospace code block
2026-08-28 09:59:59 +01:00
eSliderandCursor f1739dc9dc fix(docs): put-txt uses code block for fixed-width extracts
INE/pdftotext -layout txt (e.g. loan-ine) opens reliably in OO as
monospace; prose txt still uses hard line breaks.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-28 09:59:34 +01:00
eSliderandGitHub 5b812ce346 Merge pull request #40 from eSlider/feat/put-txt-line-breaks
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 4s
Tests / Test (Go 1.25) (push) Successful in 26s
Tests / Test (Go stable) (push) Successful in 27s
feat(docs): put-txt with preserved line breaks
2026-08-28 09:58:04 +01:00
eSliderandCursor 309ae44ee4 feat(docs): put-txt preserves line breaks in DOCX
TxtToMarkdown uses markdown hard breaks so pandoc keeps each source
line on its own row. Adds oo docs put-txt and projects files alias.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-28 09:57:55 +01:00
eSliderandGitHub ff2bc539f5 Merge pull request #39 from eSlider/fix/delete-immediately
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 5s
Tests / Test (Go 1.25) (push) Successful in 28s
Tests / Test (Go stable) (push) Successful in 29s
fix(files): permanent delete via DeleteDavItems
2026-08-28 09:56:21 +01:00
eSliderandCursor d9a7adcd94 fix(files): DeleteDavItems with Immediately true
Soft-delete (Immediately false) failed on some project files with
"You don't have enough permission to create" when moving to Trash.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-28 09:56:10 +01:00
eSliderandGitHub 3713ed6c51 Merge pull request #37 from eSlider/feat/files-dedupe
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 3s
Tests / Test (Go 1.25) (push) Successful in 26s
Tests / Test (Go stable) (push) Successful in 27s
feat(files): dedupe command + stem|ext matching
2026-08-27 23:40:16 +01:00
eSliderandCursor ac06d86363 feat(files): dedupe by stem|ext + oo projects files dedupe
- FileDedupKey groups true duplicates (same name+extension)
- put-md --replace deletes matching stem|ext only (not other formats)
- oo projects files dedupe PROJECT_ID [--apply] [--cross]
- Cross-folder mode prefers non-_trash folders, keeps newest file

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 23:40:07 +01:00
eSliderandGitHub 7d397ac680 Merge pull request #35 from eSlider/release-please--branches--main--components--go-onlyoffice
chore(main): release 0.14.0
2026-08-27 23:36:54 +01:00
github-actions[bot]andGitHub e52cb31bb7 chore(main): release 0.14.0 2026-08-27 22:36:44 +00:00
eSliderandGitHub a8cb4e9780 Merge pull request #36 from eSlider/fix/delete-files-dav
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 4s
Tests / Test (Go 1.25) (push) Successful in 1m29s
Tests / Test (Go stable) (push) Successful in 1m31s
fix(files): DeleteFiles actually removes files on produktor OO
2026-08-27 23:36:19 +01:00
eSliderandCursor 622dcdb7bf fix(files): DeleteFiles via per-file DELETE API
fileops/delete returns 200 on produktor.io OO without removing files;
route DeleteFiles through DeleteDavItems so put-md --replace and oo rm work.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 23:36:10 +01:00
eSliderandGitHub 2b267d36a8 Merge pull request #34 from eSlider/fix/put-md-upsert
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 4s
Tests / Test (Go stable) (push) Successful in 1m23s
Tests / Test (Go 1.25) (push) Successful in 1m31s
fix(docs): put-md upsert — no duplicate folder files
2026-08-27 23:29:09 +01:00
eSliderandCursor 50bd47570b fix(docs): put-md upsert by stem to avoid duplicate folder files
UploadToFolderReplacing deletes same-stem files before upload (--replace,
default true). OO may still display title+fileExst as double extension.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 23:28:11 +01:00
eSliderandGitHub 096a996726 Merge pull request #33 from eSlider/release-please--branches--main--components--go-onlyoffice
chore(main): release 0.13.0
2026-08-27 17:57:02 +01:00
github-actions[bot]andGitHub 625dd0a60a chore(main): release 0.13.0 2026-08-27 16:40:41 +00:00
eSliderandGitHub e531de280f Merge pull request #32 from eSlider/feat/docs-hocr
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 3s
Tests / Test (Go stable) (push) Successful in 1m43s
Tests / Test (Go 1.25) (push) Successful in 1m48s
feat(docs): hOCR → Markdown via go-hocr
2026-08-27 17:40:15 +01:00
eSliderandCursor 129e3a58cc chore: tidy go.sum after go-hocr bump
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 17:37:39 +01:00
eSliderandCursor 70e605b4ca docs(readme): document oo docs hocr
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 17:36:37 +01:00
eSliderandCursor a264cbd61c feat(docs): oo docs hocr — tesseract hOCR → go-hocr Markdown/YAML
Structured OCR path for agents alongside ocrmypdf searchable PDF.
as-md --hocr selects the same pipeline for OO downloads.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 17:36:20 +01:00
eSliderandGitHub a97a160c44 Merge pull request #29 from eSlider/release-please--branches--main--components--go-onlyoffice
chore(main): release 0.12.0
2026-08-27 17:05:08 +01:00
github-actions[bot]andGitHub df447f2673 chore(main): release 0.12.0 2026-08-27 15:59:02 +00:00
eSliderandGitHub db9ef12ff4 Merge pull request #31 from eSlider/feat/docs-md-ocr
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Failing after 4s
Tests / Test (Go 1.25) (push) Successful in 1m57s
Tests / Test (Go stable) (push) Successful in 2m1s
feat(docs): md↔docx, OCR→PDF, as-md/put-md
2026-08-27 16:58:36 +01:00
eSliderandCursor 2c0df0d55f fix(ci): point gitleaks at /github/workspace in docker action
The docker:// runner mounts the checkout at /github/workspace; using
${{ github.workspace }} (host path) made gitleaks fail with ENOENT.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 16:56:40 +01:00
eSliderandCursor 239c5ea672 docs(readme): document oo docs md↔docx / OCR agent workflow
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 16:49:18 +01:00
eSliderandCursor 6ea2fbadf7 feat(docs): md↔docx convert, OCR→PDF, as-md/put-md for agents
Add oo docs pipeline (pandoc/ocrmypdf) so agents keep Markdown locally while OnlyOffice stores versioned DOCX; OCR weak PDFs/images before returning MD. Also expose ListFolder/CreateFolder/MoveFiles/UploadToFolder.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-27 16:49:06 +01:00
eSlider 50d8974154 Merge pull request 'feat(security): secret-scan в CI + pre-push хуки (#142)' (#5) from feat/secret-scan#142 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Secret scan (gitleaks) (push) Successful in 3s
Tests / Test (Go 1.25) (push) Successful in 44s
Tests / Test (Go stable) (push) Successful in 45s
2026-08-23 22:46:45 +01:00
eSlider 3fd3f797dc Merge pull request 'reconcile(github): sync GitHub-only commits into Gitea main (#144)' (#6) from reconcile/github-sync#144 into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Test (Go 1.25) (push) Successful in 50s
Tests / Test (Go stable) (push) Successful in 44s
merge(github): sync GitHub-only commits into Gitea main (#144)
2026-08-23 19:55:48 +01:00
eSlider 0e299d0692 merge(github): bring GitHub-only commits into Gitea main (#144)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Test (Go 1.25) (pull_request) Successful in 48s
Tests / Test (Go stable) (pull_request) Successful in 50s
Merges GitHub main content (mail-send SendMail, files ops/WebDAV,
contact-indexes CRM, history whitelist) onto Gitea main so Gitea (source
of truth) contains every commit from both sides.
2026-08-23 19:55:39 +01:00
Andriy Oblivantsev 9ead554f5c feat(security): secret-scan via gitleaks in CI + pre-push/pre-commit hooks (#142)
Release Please / Release Please (push) Skipped
Release / GoReleaser (push) Skipped
Tests / Secret scan (gitleaks) (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Tests / Secret scan (gitleaks) (pull_request) Successful in 4s
Tests / Test (Go 1.25) (pull_request) Successful in 48s
Tests / Test (Go stable) (pull_request) Successful in 51s
2026-08-23 19:51:06 +01:00
eSlider 20a09530cd Merge pull request 'Add oo CLI support for mail attachment downloads' (#3) from feat/mails-download-attachment-cli into main
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Test (Go 1.25) (push) Skipped
Tests / Test (Go stable) (push) Skipped
Reviewed-on: #3
2026-08-19 12:17:10 +01:00
eSlider 65cd3f5c74 Merge pull request #2 from feat/mail-download-for-2dph
Release / GoReleaser (push) Skipped
Release Please / Release Please (push) Skipped
Tests / Test (Go 1.25) (push) Successful in 41s
Tests / Test (Go stable) (push) Successful in 41s
Add mail attachment download support for 2dph integration.
2026-08-19 12:08:04 +01:00
127 changed files with 16068 additions and 531 deletions
+46
View File
@@ -22,6 +22,52 @@ ONLYOFFICE_PROJECT_ID=33
# oo mails uses ONLYOFFICE_URL/USER/PASS above (Workspace Mail addon). # oo mails uses ONLYOFFICE_URL/USER/PASS above (Workspace Mail addon).
# deploy/docker-compose.rclone-webdav.yml — rclone FUSE mount of the oo-webdav
# sidecar. Reuses ONLYOFFICE_USER and accepts ONLYOFFICE_PASSWORD (alias
# ONLYOFFICE_PASS). Set ONLYOFFICE_WEBDAV_URL only if the sidecar is not on
# the default docker bridge address. See docs/rclone-webdav.md.
# ONLYOFFICE_WEBDAV_URL=http://172.17.0.1:8098/webdav
# cmd/office TUI — optional Document Server for DOCX→HTML preview: # cmd/office TUI — optional Document Server for DOCX→HTML preview:
# ONLYOFFICE_DOCS_URL=https://docs.example.com # ONLYOFFICE_DOCS_URL=https://docs.example.com
# ONLYOFFICE_DOCS_SECRET= # ONLYOFFICE_DOCS_SECRET=
# MinIO download fallback for the portal's stale AWS S3 consumer (older
# Documents folders). When the portal redirects to amazonaws.com with access
# key "minio" (403 InvalidAccessKeyId), files are fetched from the configured
# MinIO store instead. Without endpoint + key/secret the fallback is disabled.
# MINIO_ENDPOINT=http://minio.example.com:9000
# MINIO_BUCKET=office
# MINIO_ACCESS_KEY=
# MINIO_SECRET_KEY=
# Catalog scan classification rules (client mail domains, project remotes/names).
# Copy catalog/classify.example.yaml and point this at it; no rules means
# nothing is classified as work.
# OO_CATALOG_CONFIG=~/.config/oo/catalog-classify.yaml
# oo search — direct Elasticsearch access for name + content search. ES lives
# inside the OnlyOffice VM on localhost:9200; expose it with an SSH tunnel
# (see docs/elasticsearch.md). ONLYOFFICE_ES_INDEX defaults to files_file.
# ONLYOFFICE_ES_URL=http://127.0.0.1:9200
# ONLYOFFICE_ES_INDEX=files_file
# ONLYOFFICE_TENANT=
# Read-only SQL file store over the Community Server database (see
# docs/community-server-db.md). The live portal runs MySQL; ONLYOFFICE_DSN is
# `user:pass@tcp(host:port)/onlyoffice?parseTime=true`, or a `postgres://` URL.
# ONLYOFFICE_DSN=
# ONLYOFFICE_PG_DRIVER= # postgres | mysql (auto-detected from DSN)
# ONLYOFFICE_PG_TENANT=1 # falls back to ONLYOFFICE_TENANT
# Alternatively build a PostgreSQL DSN from parts:
# ONLYOFFICE_PG_HOST=
# ONLYOFFICE_PG_PORT=5432
# ONLYOFFICE_PG_USER=
# ONLYOFFICE_PG_PASSWORD=
# ONLYOFFICE_PG_DBNAME=onlyoffice
# ONLYOFFICE_PG_SSLMODE=disable
# oo index / oo search --backend own — own full-text index for PDF/scans,
# filled by `oo index` from internal/docpipe (pdftotext + OCR). Defaults to
# oo_docs_text. Uses the same ONLYOFFICE_ES_URL.
# ONLYOFFICE_ES_TEXT_INDEX=oo_docs_text
+6
View File
@@ -0,0 +1,6 @@
github: eSlider
ko_fi: eslider
liberapay: eslider
patreon: eslider
custom:
- https://polar.sh/eslider
-42
View File
@@ -1,42 +0,0 @@
name: Release Please
on:
push:
branches: [main]
workflow_dispatch:
permissions:
contents: write
pull-requests: write
issues: write
actions: write
jobs:
release-please:
name: Release Please
if: github.server_url == 'https://github.com'
runs-on: ubuntu-latest
steps:
- name: Run release-please
uses: googleapis/release-please-action@v4
id: release
with:
config-file: release-please-config.json
manifest-file: .release-please-manifest.json
token: ${{ secrets.RELEASE_PLEASE_TOKEN != '' && secrets.RELEASE_PLEASE_TOKEN || secrets.GITHUB_TOKEN }}
# Tags/releases created with GITHUB_TOKEN do not trigger other workflows
# (push:tags on Release.yml never fires). Dispatch GoReleaser explicitly.
- name: Trigger GoReleaser
if: ${{ steps.release.outputs.release_created == 'true' }}
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
TAG: ${{ steps.release.outputs.tag_name }}
run: |
set -euo pipefail
echo "Dispatching Release workflow for $TAG"
gh workflow run release.yml --repo "${{ github.repository }}" -f "tag=${TAG}"
outputs:
release_created: ${{ steps.release.outputs.release_created }}
tag_name: ${{ steps.release.outputs.tag_name }}
+2 -2
View File
@@ -20,8 +20,8 @@ jobs:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
# Always clone default branch first. workflow_dispatch often races with # Always clone default branch first. workflow_dispatch often races with
# release-please creating the tag; fetching refs/tags/X before it exists # tag creation; fetching refs/tags/X before it exists fails checkout.
# fails checkout. Retry fetch, then check out the tag. # Retry fetch, then check out the tag.
- uses: actions/checkout@v4 - uses: actions/checkout@v4
with: with:
fetch-depth: 0 fetch-depth: 0
+45
View File
@@ -10,6 +10,51 @@ permissions:
contents: read contents: read
jobs: jobs:
secret-scan:
name: Secret scan (gitleaks)
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Compute scan range (diff of new commits only)
id: range
run: |
if [ "$GITHUB_EVENT_NAME" = "pull_request" ]; then
RANGE="${{ github.event.pull_request.base.sha }}...${{ github.event.pull_request.head.sha }}"
else
BEFORE="${{ github.event.before }}"
if [ "$BEFORE" = "0000000000000000000000000000000000000000" ]; then
RANGE="$(git rev-list --max-parents=0 HEAD | tail -1)..$GITHUB_SHA"
else
RANGE="$BEFORE..$GITHUB_SHA"
fi
fi
echo "RANGE=$RANGE" >> "$GITHUB_ENV"
echo "Scanning range: $RANGE"
# Install the gitleaks binary instead of a docker action: the
# docker://zricethezav/gitleaks action hardcodes /github/workspace,
# which does not exist on the Gitea (act) runner. $GITHUB_WORKSPACE is
# the checkout dir on BOTH runners (GitHub and Gitea act). Mirrors the
# fix applied to 2dph (issue #142).
- name: Gitleaks (diff-only, fail on leak)
env:
GITLEAKS_RANGE: ${{ env.RANGE }}
run: |
set -euo pipefail
curl -fsSLo /tmp/gitleaks.tar.gz \
https://github.com/gitleaks/gitleaks/releases/download/v8.30.1/gitleaks_8.30.1_linux_x64.tar.gz
tar -xzf /tmp/gitleaks.tar.gz -C /tmp gitleaks
chmod +x /tmp/gitleaks
/tmp/gitleaks detect \
--source "$GITHUB_WORKSPACE" \
--log-opts="$GITLEAKS_RANGE" \
--redact \
--verbose
test: test:
name: Test (Go ${{ matrix.go }}) name: Test (Go ${{ matrix.go }})
runs-on: ubuntu-latest runs-on: ubuntu-latest
-3
View File
@@ -1,3 +0,0 @@
{
".": "0.11.0"
}
+16 -7
View File
@@ -9,17 +9,19 @@ Canonical Go client for OnlyOffice Workspace (Projects + Calendar + CRM) and the
- `request.go` — `Request`, `Query`, `Time`, `Token`, `MetaResponse`, `Permissions`. - `request.go` — `Request`, `Query`, `Time`, `Token`, `MetaResponse`, `Permissions`.
- `auth.go` — `Authenticate`, `AuthenticateContext`, `InvalidateToken`, `Auth`, token lifecycle. - `auth.go` — `Authenticate`, `AuthenticateContext`, `InvalidateToken`, `Auth`, token lifecycle.
- `http.go` — transport + DRY response decoders (`ResponseArray`/`ResponseObject`/`postFormObject`/`putFormObject`/`deleteObject`). - `http.go` — transport + DRY response decoders (`ResponseArray`/`ResponseObject`/`postFormObject`/`putFormObject`/`deleteObject`).
- `projects.go`, `tasks.go`, `users.go`, `calendar.go`, `crm.go`, `files.go`, `mails.go`, `invoices.go` — typed / untyped domain methods. **`files.go`** — CRM opportunity upload plus **project/task Documents**. **`mails.go`** — OnlyOffice Workspace Mail. **`invoices.go`** — CRM invoices, PDF regen/cleanup, status. Association rules: [`docs/crm-associations.md`](docs/crm-associations.md). - `projects.go`, `tasks.go`, `users.go`, `calendar.go`, `crm.go`, `files.go`, `files_webdav.go`, `files_stem.go`, `retry.go`, `mails.go`, `invoices.go` — typed / untyped domain methods. **`files.go`** — CRM opportunity upload plus **project/task Documents** (`UpdateFile`, `UploadToFolderReplacing`). **`files_webdav.go`** — Documents module by id (`ListDavFolder`, `MoveDavItems`/`CopyDavItems` with per-operation error surfacing, `ListFileOps`). **`retry.go`** — `DoRetry`: deterministic exponential backoff (no jitter) on 429/502/503/504; every bulk tool routes API calls through it, and the HTTP transport + auth (`retryRaw`, `AuthenticateContext`) retry transient answers centrally. `ratelimit.go` adds a process-wide token bucket (`OO_RATE_LIMIT`/`OO_BURST`) and a shared 429 cooldown gate, installed via `pacedTransport` in `NewClient`; `Retry-After` is parsed into `*TransientError` and honoured. See [`docs/rate-limiting.md`](docs/rate-limiting.md). **`mails.go`** — OnlyOffice Workspace Mail. **`invoices.go`** — CRM invoices, PDF regen/cleanup, status. Association rules live with the private `oo-workspace` tooling.
- **Unified file client (epic #34) — `filestore_core.go`, `filestore_rest.go`, `filestore_dav.go`, `filestore_pg.go`, `filestore_es.go`, `filestore_es_text.go`, `filestore_text_index.go`, `filestore_facade.go`.** `filestore_core.go` — model (`Entry`, `Kind`) + `FileStore`/`Searcher`; `filestore_rest.go`/`filestore_dav.go` — REST/WebDAV adapters; `filestore_pg.go` — **read-only** SQL store (PostgreSQL/MySQL, `ErrReadOnly` on writes); `filestore_es.go` — OnlyOffice Elasticsearch searcher; `filestore_es_text.go`/`filestore_text_index.go` — own PDF/scan index (`oo_docs_text`, PDF attachments via pdfdetach); `filestore_facade.go` — `FileClient` with read/write/search order and transient fallback. Use `c.Files()` (facade), `c.FileStore("rest"|"dav"|"pg"|"sql")` or `c.SQLFileStore()`; contract and how to add a backend: [`docs/unified-file-client.md`](docs/unified-file-client.md).
- Pure stdlib + `google/go-querystring`; no UI, no dotenv. - Pure stdlib + `google/go-querystring`; no UI, no dotenv.
- **CLI — `cmd/oo/` as `package main`.** Cobra wrapper that loads `.env` via `godotenv` at startup. **Subject-based command tree** mirroring [`tea`](https://gitea.com/gitea/tea): - **CLI — `cmd/oo/` as `package main`.** Cobra wrapper that loads `.env` via `godotenv` at startup. **Subject-based command tree** mirroring [`tea`](https://gitea.com/gitea/tea):
- `main.go` — entry point (docstring lists the command tree). - `main.go` — entry point (docstring lists the command tree).
- `common.go` — `rootCmd`, `newOO`, `printTable`/`printObject`, `--output table|json` flag. - `common.go` — `rootCmd`, `newOO`, `printTable`/`printObject`, `--output table|json` flag.
- `calendar.go`, `projects.go`, `projects_files.go`, `tasks.go`, `tasks_files.go`, `users.go`, `contacts.go`, `opportunities.go`, `cases.go`, `crm_tasks.go`, `mails.go`, `invoices.go` — one file per subject (or per subject facet), each registers in `init()`. - `calendar.go`, `projects.go`, `projects_files.go`, `tasks.go`, `tasks_files.go`, `users.go`, `contacts.go`, `opportunities.go`, `cases.go`, `crm.go`, `crm_tasks.go`, `catalog.go`, `docs.go`, `dav.go`, `search.go`, `index.go`, `mails.go`, `invoices.go` — one file per subject (or per subject facet), each registers in `init()`. `dav.go` exposes the Documents module by id (`oo dav ls|move|copy|mkdir|ensure-path|upload|rename-file|rename-folder|download|fileops`); `search.go` runs `oo search QUERY` (name/content, `--backend oo|own`); `index.go` fills the own full-text index (`oo index folder|files`, see [`docs/unified-file-client.md`](docs/unified-file-client.md)).
- CLI-only deps (`spf13/cobra`, `joho/godotenv`) stay out of the library. - CLI-only deps (`spf13/cobra`, `joho/godotenv`) stay out of the library.
- **TUI — `cmd/office/` as `package main`.** Bubble Tea three-pane browser (module tree, selectable list, markdown preview). Reuses `cmd/internal/bootstrap` for env/auth and the root `onlyoffice` library for all API calls. UI logic in `cmd/office/ui/`; preview/formatting in `cmd/office/preview/`; list loaders in `cmd/office/fetch/`. - **TUI — `cmd/office/` as `package main`.** Bubble Tea three-pane browser (module tree, selectable list, markdown preview). Reuses `cmd/internal/bootstrap` for env/auth and the root `onlyoffice` library for all API calls. UI logic in `cmd/office/ui/`; preview/formatting in `cmd/office/preview/`; list loaders in `cmd/office/fetch/`.
- **List table (`DataTable`)** — `cmd/office/ui/table*.go`. Column layout policies live in `cmd/office/model/table_layout.go` (`TableFlexLayoutFor`); cell rendering uses the bubbles/table inline pattern in `table_render.go` (`renderTableCell`, `padANSIWidth`). See `.cursor/skills/office-tui-table/SKILL.md` before changing center-pane tables. - **List table (`DataTable`)** — `cmd/office/ui/table*.go`. Column layout policies live in `cmd/office/model/table_layout.go` (`TableFlexLayoutFor`); cell rendering uses the bubbles/table inline pattern in `table_render.go` (`renderTableCell`, `padANSIWidth`). See `.cursor/skills/office-tui-table/SKILL.md` before changing center-pane tables.
- **Shared bootstrap — `cmd/internal/bootstrap/`.** `LoadEnv()` + `NewClient(ctx)` extracted from `oo`; both binaries import it. - **Shared bootstrap — `cmd/internal/bootstrap/`.** `LoadEnv()` + `NewClient(ctx)` extracted from `oo`; both binaries import it.
- **Personal ops tooling** (disk inventory, dossier→CRM sync, SearXNG) lives in private [`eSlider/oo-workspace`](https://git.produktor.io/eSlider/oo-workspace) (`oow`), not in this public tree. - **Bulk Documents tools live in the private `oo-workspace` repo** (`ooscan`, `pdfamount`, `kontoblatt`, `kontolink`), not in this public tree. They use the public client and its `DoRetry` pacing.
- **Personal ops tooling** (disk inventory, dossier→CRM sync, SearXNG) lives in a private companion repo `eSlider/oo-workspace` (the `oow` CLI), not in this public tree.
## Rules ## Rules
@@ -27,9 +29,15 @@ Canonical Go client for OnlyOffice Workspace (Projects + Calendar + CRM) and the
- New endpoints go into the library first; CLI commands are thin wrappers. - New endpoints go into the library first; CLI commands are thin wrappers.
- Prefer `ResponseObject` / `postFormObject` / `putFormObject` / `deleteObject` over hand-rolled `json.Unmarshal(responseField(...))` blocks — they exist for DRY, use them. - Prefer `ResponseObject` / `postFormObject` / `putFormObject` / `deleteObject` over hand-rolled `json.Unmarshal(responseField(...))` blocks — they exist for DRY, use them.
- Domain split is by file, **not** by subpackage. Don't introduce `internal/` or `pkg/*` subpackages inside the library — it flattens the `*Client` call surface for a reason. - Domain split is by file, **not** by subpackage. Don't introduce `internal/` or `pkg/*` subpackages inside the library — it flattens the `*Client` call surface for a reason.
- CLI commands follow **subject → verb** structure (`oo <subject> <verb>`), never `oo <verb>-<subject>`. Add new commands to the existing subject file if one fits; create a new `cmd/oo/<subject>.go` for a genuinely new domain. - CLI commands follow **subject → verb** structure (`oo <subject> <verb>`), never `oo <verb>-<subject>`. Add new commands to the existing subject file if one fits; create a new `cmd/oo/<subject>.go` for a genuinely new domain. The subject→verb tree in `cmd/oo/main.go` and the README table are documentation — update them with the code.
- **Documents for agents:** prefer Markdown in git; OnlyOffice UI is weak for `.md`/`.txt`. Use `oo docs put-md` (md→docx) and `oo docs put-txt` (txt→docx, preserves line breaks). All upload paths default to **upsert** by `stem` with the extension matched conversion-aware (legacy `.xls/.doc/.ppt` ↔ OOXML `.xlsx/.docx/.pptx`, since OnlyOffice converts them on upload), so a repeated `.xls` upload updates the saved `.xlsx` instead of appending a duplicate (`--replace`, default true); `--no-replace` fails on conflict; `--allow-duplicate` opts into raw OO append. `oo projects files dedupe PROJECT_ID` reports/removes duplicate stem|ext copies (`--apply`, `--cross`; includes project root folder).
- Every table output goes through `printTable(headers, rows)`; every single-object through `printObject(v)`. Do not `fmt.Println` rows ad-hoc or the `--output json` flag breaks for that command. - Every table output goes through `printTable(headers, rows)`; every single-object through `printObject(v)`. Do not `fmt.Println` rows ad-hoc or the `--output json` flag breaks for that command.
- No secrets in the repo; use `.env` (gitignored). Commit `.env.example` only. - No secrets in the repo; use `.env` (gitignored). Commit `.env.example` only.
- **No host/client specifics in the tree.** Endpoints, IPs/ports, client mail
domains, project remotes/names and personal names stay out of source and
fixtures — they come from env/config (`MINIO_*`, `OO_CATALOG_CONFIG`, see
[`catalog/classify.example.yaml`](catalog/classify.example.yaml)). This repo is
mirrored to GitHub as a public showroom, so the tree must stay project-generic.
- Follow SemVer on tags; this repo is tagged at GitHub under `git@github.com:eSlider/go-onlyoffice.git`. - Follow SemVer on tags; this repo is tagged at GitHub under `git@github.com:eSlider/go-onlyoffice.git`.
### Testing policy (2026-04-24) ### Testing policy (2026-04-24)
@@ -53,6 +61,7 @@ write `mux.HandleFunc("/api/2.0/...")` to emulate OnlyOffice, we write an
## Related ## Related
- [`eSlider/inventar`](https://git.produktor.io/eSlider/inventar) — ASR/ADR (see ASR-0008 Go library module conventions). - [`docs/README.md`](docs/README.md) — reference index (file client, ES, SQL, rclone).
- [`eSlider/inventar-sync`](https://git.produktor.io/eSlider/inventar-sync) — OnlyOffice → Gitea issue sync, consumes this library. - `eSlider/inventar` — ASR/ADR (see ASR-0008 Go library module conventions).
- [`produktor.io/vidarr`](https://git.produktor.io/produktor.io/vidarr) — legacy consumer being migrated from `pkg/onlyoffice` to this module. - `eSlider/inventar-sync` — OnlyOffice → Gitea issue sync, consumes this library.
- `vidarr` — legacy consumer being migrated from `pkg/onlyoffice` to this module.
+154 -1
View File
@@ -4,7 +4,65 @@ All notable changes to this project are documented here. The format is based on
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project
adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## Unreleased ## [0.21.0](https://git.produktor.io/eSlider/go-onlyoffice/compare/v0.20.0...v0.21.0) (2026-09-23)
### Features
* **retry:** global token-bucket rate limit (`OO_RATE_LIMIT`/`OO_BURST`), typed
`TransientError` with `Retry-After`, exponential backoff
(`OO_RETRY_ATTEMPTS`/`_BASE`/`_MAX`) and a process-wide 429 cooldown gate.
All HTTP paths are paced via `pacedTransport` (#70).
* **oo:** native document conversion (`oo docs pdf`), deep links, sheet-aware
export (`oo docs csv`/`json`), `oo docs presigned`, plus project team and
user lifecycle, milestone-delete and file update/replace (#73, #75).
* **catalog:** upstream generic workspace tooling (board-sync, CRM audit,
catalog names) (#76).
### Chores
* move business tooling (`ooscan`, `pdfamount`, `kontoblatt`, `kontolink`) out
of the public library into the private `oo-workspace` repo; tidy filestore
naming (#77).
* keep host/client specifics out of the tree (#74).
## [0.18.0](https://github.com/eSlider/go-onlyoffice/compare/v0.17.0...v0.18.0) (2026-09-04)
### Features
* **crm:** add UpdateContactName and CloseCRMTask helpers ([4d8af7a](https://github.com/eSlider/go-onlyoffice/commit/4d8af7a2fef1fcfd96ad913cefc6497e3389963b))
### Bug Fixes
* **crm:** deterministic sortBy=id in contact paged lists ([4d91726](https://github.com/eSlider/go-onlyoffice/commit/4d917261792203732ac739644c9c82124b2166eb))
### Documentation
* **funding:** eSlider support links (reverse-import GitHub e9c969a) ([d650a16](https://github.com/eSlider/go-onlyoffice/commit/d650a16a36037949505eb017c92bd3283e6e2f0f))
## [0.17.0](https://github.com/eSlider/go-onlyoffice/compare/v0.16.0...v0.17.0) (2026-08-31)
### Features
* **mailsync:** FetchMailFolder — integration-layer walk for ETL consumers ([35f0cb8](https://github.com/eSlider/go-onlyoffice/commit/35f0cb8d20076244141065c07e203e633bc3612a))
### Bug Fixes
* **files:** upsert uploads by default and dedupe project root ([24ca144](https://github.com/eSlider/go-onlyoffice/commit/24ca144b22abc5d056a5fd1ed9a1887f26a79d15))
## [0.16.0](https://github.com/eSlider/go-onlyoffice/compare/v0.15.0...v0.16.0) (2026-08-30)
### Features
* **docs:** `put-xlsx` — multi-sheet бюджеты с named inputs, SUM/AVG/MIN
формулами, cross-sheet ссылками и cell comments (`internal/xlspipe`,
excelize) ([68445b0](https://github.com/eSlider/go-onlyoffice/commit/68445b0))
### Added ### Added
@@ -16,6 +74,101 @@ adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
* Document that invoice→deal must be set at create (`update --opportunity` often HTTP 400) * Document that invoice→deal must be set at create (`update --opportunity` often HTTP 400)
## [0.15.0](https://github.com/eSlider/go-onlyoffice/compare/v0.14.0...v0.15.0) (2026-08-29)
### Features
* **files:** `ListFolder`, `CreateFolder`, `MoveFiles`, `UploadToFolder` — Documents folder helpers for OO Documents ingestion ([6ea2fba](https://github.com/eSlider/go-onlyoffice/commit/6ea2fba))
* **files:** dedupe by stem|ext (`files_stem.go`) + `oo projects files dedupe` ([ac06d86](https://github.com/eSlider/go-onlyoffice/commit/ac06d86))
* **docs:** `internal/docpipe` — md↔docx convert, OCR→PDF, hOCR→Markdown via go-hocr, as-md/put-md for agents ([6ea2fba](https://github.com/eSlider/go-onlyoffice/commit/6ea2fba), [a264cbd](https://github.com/eSlider/go-onlyoffice/commit/a264cbd))
* **docs:** optimize PDF via Ghostscript pdfwrite ([6b3f40e](https://github.com/eSlider/go-onlyoffice/commit/6b3f40e))
* **security:** gitleaks secret-scan in CI + pre-push/pre-commit hooks ([9ead554](https://github.com/eSlider/go-onlyoffice/commit/9ead554))
### Fixes
* **files:** `DeleteFiles` via per-file DELETE API; `DeleteDavItems` with `Immediately` true ([622dcdb](https://github.com/eSlider/go-onlyoffice/commit/622dcdb), [d9a7adc](https://github.com/eSlider/go-onlyoffice/commit/d9a7adc))
* **docs:** put-txt preserves line breaks in DOCX; fixed-width extracts in code block; put-md upsert by stem ([309ae44](https://github.com/eSlider/go-onlyoffice/commit/309ae44), [f1739dc](https://github.com/eSlider/go-onlyoffice/commit/f1739dc), [50bd475](https://github.com/eSlider/go-onlyoffice/commit/50bd475))
* **ci:** point gitleaks at `/github/workspace` in docker action ([2c0df0d](https://github.com/eSlider/go-onlyoffice/commit/2c0df0d))
## [0.14.0](https://github.com/eSlider/go-onlyoffice/compare/v0.13.0...v0.14.0) (2026-08-27)
### Features
* **crm:** contact email & person-opportunity indexes, history entity whitelist ([9c450a0](https://github.com/eSlider/go-onlyoffice/commit/9c450a0352c25f12ac16d91062744ab026ffe660))
* **docs:** hOCR → Markdown via go-hocr ([e531de2](https://github.com/eSlider/go-onlyoffice/commit/e531de280f2c521a7411f75212803649b177db58))
* **docs:** md↔docx convert, OCR→PDF, as-md/put-md for agents ([6ea2fba](https://github.com/eSlider/go-onlyoffice/commit/6ea2fbadf768d7d85f6c7a49a9bc1050e642bc56))
* **docs:** md↔docx, OCR→PDF, as-md/put-md ([db9ef12](https://github.com/eSlider/go-onlyoffice/commit/db9ef12ff4096f8aadb557ca128b6edcb503028c))
* **docs:** oo docs hocr — tesseract hOCR → go-hocr Markdown/YAML ([a264cbd](https://github.com/eSlider/go-onlyoffice/commit/a264cbd61c12be9098c8d16e7c1c1253eb460816))
* **mail:** oo mails send — SendMail + guard empty-by-id send ([dda2bd3](https://github.com/eSlider/go-onlyoffice/commit/dda2bd3ca5cea6ebb89e1ac02b785b9838727bda))
* **mail:** oo mails send — SendMail client + guard empty-by-id ([ab7dfbd](https://github.com/eSlider/go-onlyoffice/commit/ab7dfbd8594de5724286e2b460e789f4acb114bc))
* **security:** secret-scan via gitleaks in CI + pre-push/pre-commit hooks ([#142](https://github.com/eSlider/go-onlyoffice/issues/142)) ([9ead554](https://github.com/eSlider/go-onlyoffice/commit/9ead554f5cc229687d3a27934a023fb1dddaef15))
### Bug Fixes
* **ci:** point gitleaks at /github/workspace in docker action ([2c0df0d](https://github.com/eSlider/go-onlyoffice/commit/2c0df0d55fb03555b60481695554d75dceffce48))
* **docs:** put-md upsert — no duplicate folder files ([2b267d3](https://github.com/eSlider/go-onlyoffice/commit/2b267d36a8086b14a94706468f7e92e9179e62d4))
* **docs:** put-md upsert by stem to avoid duplicate folder files ([50bd475](https://github.com/eSlider/go-onlyoffice/commit/50bd47570b6430cebde43641a7b1d2f8d54d32f2))
* **files:** DeleteFiles actually removes files on produktor OO ([a8cb4e9](https://github.com/eSlider/go-onlyoffice/commit/a8cb4e97805226e8af694ec6a64ee3ac6a5e9fdd))
* **files:** DeleteFiles via per-file DELETE API ([622dcdb](https://github.com/eSlider/go-onlyoffice/commit/622dcdb7bffd49acbea5ef07f5d5d009f44b1358))
### Documentation
* **readme:** document oo docs hocr ([70e605b](https://github.com/eSlider/go-onlyoffice/commit/70e605b4ca8ee005e4351b2fc2d2eec081604998))
* **readme:** document oo docs md↔docx / OCR agent workflow ([239c5ea](https://github.com/eSlider/go-onlyoffice/commit/239c5ea67201e62de6ba067f0e7963c7c9bda188))
## [0.13.0](https://github.com/eSlider/go-onlyoffice/compare/v0.12.0...v0.13.0) (2026-08-27)
### Features
* **crm:** contact email & person-opportunity indexes, history entity whitelist ([9c450a0](https://github.com/eSlider/go-onlyoffice/commit/9c450a0352c25f12ac16d91062744ab026ffe660))
* **docs:** hOCR → Markdown via go-hocr ([e531de2](https://github.com/eSlider/go-onlyoffice/commit/e531de280f2c521a7411f75212803649b177db58))
* **docs:** md↔docx convert, OCR→PDF, as-md/put-md for agents ([6ea2fba](https://github.com/eSlider/go-onlyoffice/commit/6ea2fbadf768d7d85f6c7a49a9bc1050e642bc56))
* **docs:** md↔docx, OCR→PDF, as-md/put-md ([db9ef12](https://github.com/eSlider/go-onlyoffice/commit/db9ef12ff4096f8aadb557ca128b6edcb503028c))
* **docs:** oo docs hocr — tesseract hOCR → go-hocr Markdown/YAML ([a264cbd](https://github.com/eSlider/go-onlyoffice/commit/a264cbd61c12be9098c8d16e7c1c1253eb460816))
* **mail:** oo mails send — SendMail + guard empty-by-id send ([dda2bd3](https://github.com/eSlider/go-onlyoffice/commit/dda2bd3ca5cea6ebb89e1ac02b785b9838727bda))
* **mail:** oo mails send — SendMail client + guard empty-by-id ([ab7dfbd](https://github.com/eSlider/go-onlyoffice/commit/ab7dfbd8594de5724286e2b460e789f4acb114bc))
* **security:** secret-scan via gitleaks in CI + pre-push/pre-commit hooks ([#142](https://github.com/eSlider/go-onlyoffice/issues/142)) ([9ead554](https://github.com/eSlider/go-onlyoffice/commit/9ead554f5cc229687d3a27934a023fb1dddaef15))
### Bug Fixes
* **ci:** point gitleaks at /github/workspace in docker action ([2c0df0d](https://github.com/eSlider/go-onlyoffice/commit/2c0df0d55fb03555b60481695554d75dceffce48))
* **files:** rewrite viewUrl host to API base on download ([ecf34ba](https://github.com/eSlider/go-onlyoffice/commit/ecf34ba51aae77f1626a60795f7329f25d9344e5))
### Documentation
* **readme:** document oo docs hocr ([70e605b](https://github.com/eSlider/go-onlyoffice/commit/70e605b4ca8ee005e4351b2fc2d2eec081604998))
* **readme:** document oo docs md↔docx / OCR agent workflow ([239c5ea](https://github.com/eSlider/go-onlyoffice/commit/239c5ea67201e62de6ba067f0e7963c7c9bda188))
## [0.12.0](https://github.com/eSlider/go-onlyoffice/compare/v0.11.0...v0.12.0) (2026-08-27)
### Features
* **crm:** contact email & person-opportunity indexes, history entity whitelist ([9c450a0](https://github.com/eSlider/go-onlyoffice/commit/9c450a0352c25f12ac16d91062744ab026ffe660))
* **docs:** md↔docx convert, OCR→PDF, as-md/put-md for agents ([6ea2fba](https://github.com/eSlider/go-onlyoffice/commit/6ea2fbadf768d7d85f6c7a49a9bc1050e642bc56))
* **docs:** md↔docx, OCR→PDF, as-md/put-md ([db9ef12](https://github.com/eSlider/go-onlyoffice/commit/db9ef12ff4096f8aadb557ca128b6edcb503028c))
* **files:** add WebDAV-oriented Files operations ([ebbd5d5](https://github.com/eSlider/go-onlyoffice/commit/ebbd5d5373abfeca5f160312f201cbd683f55e42))
* **mail:** oo mails send — SendMail + guard empty-by-id send ([dda2bd3](https://github.com/eSlider/go-onlyoffice/commit/dda2bd3ca5cea6ebb89e1ac02b785b9838727bda))
* **mail:** oo mails send — SendMail client + guard empty-by-id ([ab7dfbd](https://github.com/eSlider/go-onlyoffice/commit/ab7dfbd8594de5724286e2b460e789f4acb114bc))
* **security:** secret-scan via gitleaks in CI + pre-push/pre-commit hooks ([#142](https://github.com/eSlider/go-onlyoffice/issues/142)) ([9ead554](https://github.com/eSlider/go-onlyoffice/commit/9ead554f5cc229687d3a27934a023fb1dddaef15))
### Bug Fixes
* **ci:** point gitleaks at /github/workspace in docker action ([2c0df0d](https://github.com/eSlider/go-onlyoffice/commit/2c0df0d55fb03555b60481695554d75dceffce48))
* **files:** rewrite viewUrl host to API base on download ([ecf34ba](https://github.com/eSlider/go-onlyoffice/commit/ecf34ba51aae77f1626a60795f7329f25d9344e5))
### Documentation
* **readme:** document oo docs md↔docx / OCR agent workflow ([239c5ea](https://github.com/eSlider/go-onlyoffice/commit/239c5ea67201e62de6ba067f0e7963c7c9bda188))
## [0.11.0](https://github.com/eSlider/go-onlyoffice/compare/v0.10.0...v0.11.0) (2026-08-18) ## [0.11.0](https://github.com/eSlider/go-onlyoffice/compare/v0.10.0...v0.11.0) (2026-08-18)
+271 -17
View File
@@ -476,6 +476,11 @@ type Task struct {
| `UpdateProject(req)` | Update project details | | `UpdateProject(req)` | Update project details |
| `DeleteProject(id)` | Delete a project | | `DeleteProject(id)` | Delete a project |
| `GetProjectMilestones(project)` | Get milestones with task counts | | `GetProjectMilestones(project)` | Get milestones with task counts |
| `DeleteMilestone(id)` | Remove a milestone |
| `ListProjectTeam(ctx, id)` | Portal users on the project team |
| `AddProjectTeamUser(ctx, id, userID)` | Add a portal user to the team |
| `RemoveProjectTeamUser(ctx, id, userID)` | Remove a portal user from the team |
| `SetProjectTeam(ctx, id, participants, notify)` | Replace the team with the given user ids |
### Tasks ### Tasks
@@ -502,6 +507,61 @@ type Task struct {
| Method | Description | | Method | Description |
|---|---| |---|---|
| `GetUsers()` | List all users with profiles | | `GetUsers()` | List all users with profiles |
| `GetUser(ctx, id)` | One user profile by id |
| `CreateUser(ctx, NewUserRequest)` | Add a portal user |
| `UpdateUser(ctx, id, body)` | Update profile fields (JSON PUT) |
| `DeleteUser(ctx, id)` | Delete permanently (auto-terminates first — OO refuses active users) |
| `BlockUser(ctx, id)` / `UnblockUser(ctx, id)` | Terminate / reactivate (login kept/denied) |
| `ChangeUserPassword(ctx, id, pw)` | Set a new password |
| `ChangeUserStatus(ctx, id, active)` | Activate / Terminate via `people/status` |
| `AuthenticateAs(ctx, login, pw)` | Verify a login without mutating the cached token |
### Documents Files
| Method | Description |
|---|---|
| `ListDavFolder(ctx, id)` | List a Documents folder (`@root` for virtual sections) |
| `ListDavSections(ctx)` | Virtual sections (Documents, Projects, …) |
| `CreateDavFolder(ctx, parentID, title)` | Create a subfolder |
| `RenameDavFolder(ctx, id, title)` / `RenameDavFile(ctx, id, title)` | Rename folder / file |
| `DownloadFile(ctx, id, dst)` / `DownloadDavFile(ctx, id, w)` | Download file bytes |
| `UploadDavFile(ctx, folderID, fileName, src)` | Upload from a reader |
| `UploadToFolder(ctx, folderID, localPath)` | Upload a local file into a folder |
| `UploadToFolderReplacing(ctx, folderID, localPath)` | Upsert by `stem\|ext`; returns replaced ids |
| `UpdateFile(ctx, fileID, localPath)` | New version of an existing file (same id, no copy) |
| `MoveDavItems(ctx, folderIDs, fileIDs, dest)` | Move (`resolveType=Skip`); per-operation errors surfaced, not silent nil |
| `CopyDavItems(ctx, folderIDs, fileIDs, dest)` | Copy (`conflictResolveType=Skip`); errors surfaced |
| `MoveFiles(ctx, destFolderID, fileIDs)` | Move with `resolveType=Skip` + `holdResult`; errors surfaced |
| `ListFileOps(ctx)` | Active file operations (move/copy status polling) |
| `FolderFiles(ctx, folderID)` | Flat file list of a folder (stem helpers) |
| `DeleteFilesByStem(ctx, folderID, stem)` | Remove `stem\|ext` copies |
| `DoRetry(ctx, policy, fn)` | Deterministic exponential backoff (`Base·2^(N-1)`, no jitter) on 429/502/503/504; honours `Retry-After` and the process-wide cooldown gate |
| `DefaultRetryPolicy()` | From env: 7 attempts, 2s base, 2m cap (`OO_RETRY_ATTEMPTS/_BASE/_MAX`) |
| `Transient(err)` | True for retriable OnlyOffice answers (`*TransientError` or HTTP 429/502/503/504 text) |
### Deep links & conversion
| Method | Description |
|---|---|
| `FileEditorURL(portalBase, id)` / `(c *Client).FileEditorURL(id)` | DocEditor deep link `/Products/Files/DocEditor.aspx?fileid=` |
| `FolderURL(portalBase, id)` / `(c *Client).FolderURL(id)` | Documents folder link |
| `PresignedURI(ctx, fileID)` | Short-lived fetchable URL of a portal file (`presigneduri`) |
| `SignJWT(secret, payload)` | HS256 JWT, stdlib only |
| `ConvertDocument(ctx, docsBase, secret, req)` | OnlyOffice DocumentServer conversion (`/converter`; legacy `/ConvertService.ashx`) |
| `DownloadURLTo(ctx, url, w)` | Stream an absolute URL into a writer |
| `SyncBoard(ctx, board, apply)` | Upsert project milestones/tasks from a YAML board (dry-run when apply=false) |
| `AuditOpportunities(ctx)` | Opportunities with file/task/member counts and a coarse class |
| `WorkbookSheetNames(data)` | Worksheet names of an XLS/XLSX/ODS workbook |
| `WorkbookSheetCSV(data, sheet, delim)` | One worksheet → CSV (sheet-aware, excelize) |
| `WorkbookSheetJSON(data, sheet)` | One worksheet → rows as objects (first row = header) |
Example — convert a portal file (or a local file) to PDF with the native engine:
```bash
export ONLYOFFICE_DS_SECRET=<DocumentServer CoAuthoring secret>
oo docs pdf 1234 --out out.pdf # OO file id → PDF
oo docs pdf ./report.docx --to pdf # local file → PDF (temp upload, auto-cleanup)
```
### Helper Types ### Helper Types
@@ -559,6 +619,41 @@ oo opportunities list
oo opportunities stages oo opportunities stages
oo cases list oo cases list
oo crm-tasks categories oo crm-tasks categories
# Deep links & native document conversion
oo link 1234 2345 # DocEditor URL for file ids (title + url)
oo docs presigned 1234 # short-lived fetchable URL of an OO file
oo docs pdf 1234 --out out.pdf # OO file → PDF (via DocumentServer)
oo docs pdf ./report.docx --to pdf # local file → PDF on the fly (temp upload+cleanup)
oo docs pdf 1234 --stream > out.pdf # pipe: bytes to stdout (alias --pipe)
# Spreadsheet export (sheet-aware; local reader — the DS csv output is first-sheet-only)
oo docs csv 1234 --sheet 2 # XLS/XLSX/ODS worksheet → CSV (--sheet N, 1-based)
oo docs csv 1234 --delimiter ';' # ; | | \t via --delimiter
oo docs json 1234 --sheet 2 # worksheet → JSON rows (first row = header)
oo docs csv ./book.xlsx --sheet 1 --out sheet1.csv
# export ONLYOFFICE_DS_SECRET=<DocumentServer CoAuthoring secret>
# docs base: $ONLYOFFICE_DOCS_URL, else $ONLYOFFICE_URL + /ds-vpath
# Project team (portal users) CRUD
oo projects team list 42
oo projects team add 42 <user_id> [<user_id>...]
oo projects team remove 42 <user_id>
oo projects team set 42 <user_id> [...] # replace team (may lag; verify with list)
oo projects milestone-delete 7
# Project documents: fresh re-upload (single clean version) / new version
oo projects files replace-in 1 ./contract.pdf # hard delete same stem|ext in folder + upload
oo projects files update 1 ./contract.docx # overwrite content, same file id
# Users lifecycle
oo users list ; oo users get <user_id>
oo users create --first Jane --last Doe --email jane.doe@example.com --password '…'
oo users check --login jane.doe@example.com # verify login (email works when userName 500s)
oo users update <user_id> --title "…" --location "…"
oo users block <user_id> ; oo users unblock <user_id>
oo users password <user_id> # reads the new password from stdin
oo users delete <user_id> [<user_id>...]
``` ```
### office (TUI) ### office (TUI)
@@ -607,30 +702,140 @@ go test -tags=integration ./cmd/office/fetch/... ./cmd/office/preview/...
```bash ```bash
# Project Documents (files module) # Project Documents (files module)
oo projects files list 33 oo projects files list 33
oo projects files upload 33 ./notes.md oo projects files upload 33 ./notes.docx
oo projects files download 12345 --to ./copy.md oo projects files download 12345 --to ./copy.docx
oo projects files rename 12345 notes-v2.md oo projects files rename 12345 notes-v2.docx
oo projects files delete 12345 oo projects files delete 12345
oo projects files dedupe 7 # dry-run duplicate report
oo projects files dedupe 7 --apply # remove older stem|ext copies per folder
oo projects files dedupe 7 --cross --apply # cross-folder; keep non-_trash
# Agent document pipeline (md in git ↔ docx in OO; OCR scans)
oo docs tools
oo docs convert ./note.md # → note.docx
oo docs convert ./note.docx # → note.md
oo docs ocr ./scan.jpg --md ./scan.md # searchable PDF + markdown
oo docs hocr ./scan.jpg --lang spa --md ./scan.hocr.md --yaml ./scan.yml
oo docs put-md 7 ./OO-HONDA-7-INDEX.md --folder 490
oo docs put-txt 7 ./notes.txt --folder 490
oo docs put-xlsx 7 ./table.xlsx --folder 490
oo docs as-md 2815 --to ./parte.md # download OO file as MD (OCR if needed)
oo docs as-md 307 --hocr --lang spa # OO download via go-hocr structure
oo projects files put-md 7 ./note.md # alias
oo projects files as-md 2815 # alias
oo tasks files list 208 oo tasks files list 208
oo tasks files upload 208 ./notes.pdf oo tasks files upload 208 ./notes.pdf
oo tasks files detach 208 12345 oo tasks files detach 208 12345
``` ```
### Documents module (`oo dav`)
Direct access to the Documents module by folder/file id — the same calls that
back `oo-webdav` and the project/task file commands. `move` sends
`resolveType=Skip` + `holdResult=true`: without those params the legacy
`fileops/move` endpoint answers 200 without moving anything, and the library
surfaces such per-operation errors instead of a silent nil
(`MoveDavItems` / `CopyDavItems` / `MoveFiles`).
```bash
oo dav ls 659
oo dav ls @root # virtual sections (Documents, Projects, …)
oo dav mkdir 659 "2026 inbox"
oo dav ensure-path "Banks/Caixa" # resolve-or-create; idempotent, prints the folder id
oo dav ensure-path "Banks/Caixa" --under 659 # start under an explicit folder id
oo dav upload 659 ./historico.xlsx # multipart upload, upsert by stem (legacy↔OOXML ext aware)
oo dav upload 659 ./a.xlsx ./b.pdf --replace=false # fail on name conflict instead of replacing
oo dav move 659 22881 22882 # DEST_FOLDER_ID FILE_ID…
oo dav move 659 22881 --folders 670 # move folders along with files
oo dav copy 659 22881
oo dav rename-file 22881 invoice-v2.pdf
oo dav rename-folder 671 o2-archive
oo dav download 22881 --to ./copy.pdf # default path: ./<server title>
oo dav fileops # active move/copy operations (status polling)
```
`ensure-path` defaults to the concrete **My documents** section id (resolved
from `@root`, `rootFolderType=5`); `--under FOLDER_ID` overrides it. `upload`
reuses `UploadToFolderReplacing` (`--replace`, default) or `UploadToFolder`
after an `AssertNoFileConflict` check (`--no-replace`), so no new HTTP paths.
### Search and index (`oo search`, `oo index`)
Full-text search over the Documents index. The REST endpoint
`/api/2.0/files/@search/{query}` only searches file names in the database, so
`oo search` talks to the OnlyOffice **Elasticsearch** directly (index
`files_file`). Name search is default; `--content` also matches extracted
document text (`document.attachment.content`, Office formats only).
See [`docs/elasticsearch.md`](docs/elasticsearch.md) for the tunnel setup.
```bash
oo search "Rechnung" # names only
oo search "Mahngebühr" --content # names + document text
oo search "Rechnung" --folder 649 --limit 50
oo search "Rechnung" --json # shorthand for -o json
```
Requires `ONLYOFFICE_ES_URL` (plus optional `ONLYOFFICE_ES_INDEX`,
`ONLYOFFICE_TENANT`).
#### PDF/scans: own index (`oo index` + `--backend own`)
The OnlyOffice index covers Office formats only, so PDFs (`S1019`-style invoice
numbers) are not searchable by content. `oo index` extracts PDF text with
`internal/docpipe` (pdftotext, OCR for scans) — including the text of embedded
PDF attachments (`pdfdetach`: `<doc>.md`, `.xml`, covers the original/scan and
ZUGFeRD e-invoice XML) — into a separate index (`ONLYOFFICE_ES_TEXT_INDEX`,
default `oo_docs_text`); the OnlyOffice server and its index are **not**
modified. Then search it with `--backend own`.
```bash
oo index folder 634 --recursive --exts pdf # populate (idempotent upsert)
oo index files 3576 3578 # specific files
oo index folder 634 --dry-run # plan only
oo search "S1021" --content --backend own # finds the PDF
oo search "Rechnung" --backend own --folder 634 --json
```
See [`docs/elasticsearch.md`](docs/elasticsearch.md) for the decision and
trade-offs.
### Unified file client
All file backends (REST, WebDAV, read-only SQL, Elasticsearch) sit behind one
facade: `c.Files()` returns a `*FileClient` that also implements `FileStore`,
so old call sites keep working. Pick a transport per call with
`c.FileStore("rest"|"dav"|"pg"|"sql")` (SQL is read-only), open the SQL store
with `c.SQLFileStore()`, or register a backend on the facade
(`RegisterStore`/`RegisterSearcher`). Contract, model (`Entry`/`Kind`),
fallback rules, env names and how to add a backend:
[`docs/unified-file-client.md`](docs/unified-file-client.md).
### Bulk tools
Business / one-off bulk tools (`ooscan`, `pdfamount`, `kontoblatt`,
`kontolink`) live in the private `oo-workspace` repo, not in this public
library. They build on the public client and the same `DoRetry` pacing.
| Subject | Verbs | | Subject | Verbs |
|---|---| |---|---|
| `calendar` | `list`, `events`, `add`, `delete` | | `calendar` | `list`, `events`, `add`, `delete` |
| `projects` | `list`, `get`, `milestones`, `create`, `update`, `delete`, **`files`** (`list`, `upload`, `download`, `rename`, `delete`) | | `projects` | `list`, `get`, `milestones`, `milestone-create`, `board-sync`, `create`, `update`, `delete`, `contacts` (`add`, `remove`), `link-authors`, `link-git`, **`files`** (`list`, `upload`, `download`, `rename`, `delete`, `dedupe`, `as-md`, `put-md`, `put-txt`, `put-xlsx`) |
| `tasks` | `list`, `get`, `create`, `update`, `delete`, `subtask add`, **`files`** (`list`, `upload`, `detach`) | | `tasks` | `list`, `get`, `create`, `update`, `delete`, `subtask add`, **`files`** (`list`, `upload`, `detach`) |
| `users` | `list`, `self` (alias: `oo whoami`) | | `users` | `list`, `self` (alias: `oo whoami`) |
| `contacts` | `list`, `get`, `delete`, `info-add`, `merge`, `dedupe-info` | | `contacts` | `list`, `get`, `delete`, `info-add`, `merge`, `dedupe-info`, `tags`, `tag-add`, `tag-create`, `tag-remove` |
| `persons` | `list`, `create`, `delete`, `dedupe` | | `persons` | `list`, `create`, `delete`, `dedupe` |
| `companies` | `list`, `create`, `delete`, `dedupe`, `dedupe-persons` | | `companies` | `list`, `create`, `delete`, `dedupe`, `dedupe-persons` |
| `opportunities` | `list`, `get`, `create`, `delete`, `stages`, `member-add`, `dedupe`, `dedupe-members`, `fix-titles` | | `opportunities` | `list`, `get`, `create`, `update`, `delete`, `stages`, `member-add`, `dedupe`, `dedupe-members`, `fix-titles` |
| `invoices` | `list`, `get`, `create`, `update`, `pdf`, `pdf-cleanup`, `status`, `delete`, `items …` | | `invoices` | `list`, `get`, `create`, `update`, `pdf`, `pdf-cleanup`, `status`, `delete`, `items …` |
| `crm` | `cleanup` | | `crm` | `audit`, `cleanup` |
| `mails` | `accounts`, `folders`, `list`, `get`, `draft`, `attach`, `draft-invoice`, `delete` | | `mails` | `accounts`, `folders`, `list`, `get`, `download-attachment`, `draft`, `attach`, `draft-invoice`, `send`, `delete` |
| `cases` | `list`, `create`, `delete`, `member-add` | | `cases` | `list`, `create`, `delete`, `member-add` |
| `crm-tasks` | `list`, `create`, `delete`, `categories` | | `crm-tasks` | `list`, `create`, `delete`, `categories`, `reassign-self` |
| `docs` | `tools`, `convert`, `pdf`, `presigned`, `csv`, `json`, `optimize`, `ocr`, `hocr`, `as-md`, `put-md`, `put-txt`, `put-xlsx` |
| `catalog` | `match`, `merge`, `apply`, `scan-contacts`, `scan-projects`, `scan-thunderbird` |
| `dav` | `ls`, `move`, `copy`, `mkdir`, `ensure-path`, `upload`, `rename-file`, `rename-folder`, `download`, `fileops` |
| `search` | `QUERY` (`--content`, `--folder ID`, `--limit N`, `--backend oo\|own`, `--json`) |
| `index` | `folder FOLDER_ID`, `files FILE_ID...` (`--recursive`, `--exts pdf`, `--limit N`, `--dry-run`) |
The CLI reads only `.env` from the current working directory (godotenv is a The CLI reads only `.env` from the current working directory (godotenv is a
CLI-only concern — the library itself never loads dotfiles). CLI-only concern — the library itself never loads dotfiles).
@@ -638,6 +843,12 @@ CLI-only concern — the library itself never loads dotfiles).
Canonical `ONLYOFFICE_*` variables win over aliases. Optional CLI-only aliases: Canonical `ONLYOFFICE_*` variables win over aliases. Optional CLI-only aliases:
`OO_URL` / `OO_USER` / `OO_PASS` → `ONLYOFFICE_URL` / `ONLYOFFICE_USER` / `ONLYOFFICE_PASS`. `OO_URL` / `OO_USER` / `OO_PASS` → `ONLYOFFICE_URL` / `ONLYOFFICE_USER` / `ONLYOFFICE_PASS`.
Catalog scanning (`oo catalog scan-projects` / `scan-thunderbird`) classifies
clients from rules in `$OO_CATALOG_CONFIG` (or `--config`); see
[`catalog/classify.example.yaml`](catalog/classify.example.yaml). Without rules
nothing is classified as work. The MinIO download fallback is off unless
`MINIO_ENDPOINT` + `MINIO_ACCESS_KEY` + `MINIO_SECRET_KEY` are set.
Run `oo --help` or `oo <subject> --help` for the full command reference. Run `oo --help` or `oo <subject> --help` for the full command reference.
> **0.5.0 migration note:** the command tree was flattened per-subject. Old > **0.5.0 migration note:** the command tree was flattened per-subject. Old
@@ -717,8 +928,8 @@ Merge two known company ids (keeps `INTO`):
oo contacts merge FROM_ID INTO_ID oo contacts merge FROM_ID INTO_ID
``` ```
Company ↔ person ↔ deal ↔ project ↔ invoice ↔ mail rules and OO quirks: Company ↔ person ↔ deal ↔ project ↔ invoice ↔ mail rules and OO quirks live
[docs/crm-associations.md](docs/crm-associations.md). with the private `oo-workspace` tooling.
### Invoices (`oo invoices`) ### Invoices (`oo invoices`)
@@ -863,22 +1074,65 @@ oo projects files list 33
| `ONLYOFFICE_CALENDAR_ID` | Default calendar id used when omitted (default `1`) | | `ONLYOFFICE_CALENDAR_ID` | Default calendar id used when omitted (default `1`) |
| `ONLYOFFICE_PROJECT_ID` | Default project id used when omitted (default `33`) | | `ONLYOFFICE_PROJECT_ID` | Default project id used when omitted (default `33`) |
| `OO_URL`, `OO_USER`, `OO_PASS` | Optional CLI-only aliases for `ONLYOFFICE_*` | | `OO_URL`, `OO_USER`, `OO_PASS` | Optional CLI-only aliases for `ONLYOFFICE_*` |
| `OO_RATE_LIMIT` | Process-wide request pacing, req/s (default `4`; `0` disables) |
| `OO_BURST` | Token-bucket burst (default `1`) |
| `OO_RETRY_ATTEMPTS` | Transient retries, total attempts (default `7`) |
| `OO_RETRY_BASE` | Exponential backoff base (default `2s`) |
| `OO_RETRY_MAX` | Backoff cap (default `2m`) |
| `ONLYOFFICE_ES_URL` | Elasticsearch base URL (`oo search`, own index); see [`docs/elasticsearch.md`](docs/elasticsearch.md) |
| `ONLYOFFICE_ES_INDEX` | OnlyOffice index (default `files_file`) |
| `ONLYOFFICE_ES_TEXT_INDEX` | Own PDF/scan index (default `oo_docs_text`) |
| `ONLYOFFICE_TENANT` | `tenantId` filter for ES/SQL (empty = all) |
| `ONLYOFFICE_DSN` | Read-only SQL DSN (MySQL or `postgres://`); see [`docs/community-server-db.md`](docs/community-server-db.md) |
| `ONLYOFFICE_PG_DRIVER`, `ONLYOFFICE_PG_TENANT`, `ONLYOFFICE_PG_HOST/_PORT/_USER/_PASSWORD/_DBNAME/_SSLMODE` | SQL store override / DSN by parts (PostgreSQL) |
| `MINIO_ENDPOINT`, `MINIO_BUCKET`, `MINIO_ACCESS_KEY`, `MINIO_SECRET_KEY` | Object-store layout for SQL `Download` |
| `ONLYOFFICE_WEBDAV_URL` | rclone WebDAV sidecar URL (default `http://172.17.0.1:8098/webdav`) |
Mail and CRM cleanup are documented in [oo CLI use cases](#oo-cli-use-cases) above. Personal disk inventory / dossier sync lives in the private `oo-workspace` (`oow`) tooling. Mail and CRM cleanup are documented in [oo CLI use cases](#oo-cli-use-cases) above. Personal disk inventory / dossier sync lives in the private `oo-workspace` (`oow`) tooling.
### CI / releases ## Testing
GitHub Actions (pattern from [`eSlider/go-config`](https://github.com/eSlider/go-config)): ```bash
go build ./... && go vet ./...
go test ./... # unit — no network, no vendor mocks
go test -race ./...
go test -tags=integration ./... # live OnlyOffice (skip without creds)
```
Unit tests are pure Go (parsers, encoders, conversions). Integration tests
(`//go:build integration`) hit a live instance and **skip** cleanly when the
env is missing, so `go test ./...` stays green offline. New endpoints ship with
an integration test before merge (policy in [`AGENTS.md`](AGENTS.md)).
Live runs need credentials (`ONLYOFFICE_URL`, `ONLYOFFICE_USER`,
`ONLYOFFICE_PASS`) and, per backend:
- **Elasticsearch** (`oo search`, own index) — ES lives on `127.0.0.1:9200`
inside the OnlyOffice VM; expose it over SSH
(`-L 9200:127.0.0.1:9200`) and set `ONLYOFFICE_ES_URL`
(see [`docs/elasticsearch.md`](docs/elasticsearch.md)).
- **SQL backend** (`FileClient`, `SQLFileStore`) — MySQL on `127.0.0.1:3306`
in the same VM; tunnel `-L 3306:127.0.0.1:3306`, then set `ONLYOFFICE_DSN`
(see [`docs/community-server-db.md`](docs/community-server-db.md)).
`ONLYOFFICE_PG_TEST_FILE_ID` / `ONLYOFFICE_PG_TEST_FOLDER_ID` select a real
file for the REST cross-check; `MINIO_*` enable the download check.
### rclone WebDAV mount
The rclone WebDAV mount (compose + `oo-webdav` sidecar) is deployment
tooling and lives in the private `oo-workspace` repo. Set
`ONLYOFFICE_WEBDAV_URL` to point the client at it.
### CI / releases
| Workflow | Trigger | Purpose | | Workflow | Trigger | Purpose |
|---|---|---| |---|---|---|
| `test.yml` | push / PR | `go vet`, unit tests, build `oo` + `office` | | `test.yml` | push / PR | `go vet`, unit tests, build `oo` + `office` |
| `release-please.yml` | push to `main` | semver PR from conventional commits |
| `release.yml` | tag `v*` | GoReleaser cross-platform `oo` + `office` binaries | | `release.yml` | tag `v*` | GoReleaser cross-platform `oo` + `office` binaries |
Repo setting required once: **Settings → Actions → General → Allow GitHub Actions to create and approve pull requests**. Gitea is canonical; tags are created there per SemVer (`fix:` → patch,
`feat:` → minor, `!` → major). GoReleaser publishes assets to
Merge the release-please PR to tag a version; GoReleaser publishes assets to [GitHub Releases](https://github.com/eSlider/go-onlyoffice/releases). [GitHub Releases](https://github.com/eSlider/go-onlyoffice/releases).
## Examples ## Examples
+61 -8
View File
@@ -43,6 +43,11 @@ func (c *Client) Authenticate() error { return c.ensureToken() }
// cached token is still valid it returns immediately; otherwise it performs // cached token is still valid it returns immediately; otherwise it performs
// a POST to /api/2.0/authentication.json that is cancellable via ctx. // a POST to /api/2.0/authentication.json that is cancellable via ctx.
// //
// Transient answers from the edge (openresty 429/502/503/504) are retried with
// the same deterministic policy as every other request (see retry.go), because
// the server rate-limits authentication and the integration suite otherwise
// fails with a raw HTML 429 page.
//
// This is the recommended entry point for long-running syncs (cron, // This is the recommended entry point for long-running syncs (cron,
// watchers) because it guarantees that a stalled auth call will not block // watchers) because it guarantees that a stalled auth call will not block
// the caller past its deadline. // the caller past its deadline.
@@ -50,6 +55,14 @@ func (c *Client) AuthenticateContext(ctx context.Context) error {
if c.tokenValid() { if c.tokenValid() {
return nil return nil
} }
return DoRetry(ctx, DefaultRetryPolicy(), func() error {
return c.authenticateOnce(ctx)
})
}
// authenticateOnce performs a single authentication POST. Callers must handle
// retries; use AuthenticateContext.
func (c *Client) authenticateOnce(ctx context.Context) error {
body, err := json.Marshal(c.credentials) body, err := json.Marshal(c.credentials)
if err != nil { if err != nil {
return fmt.Errorf("marshal credentials: %w", err) return fmt.Errorf("marshal credentials: %w", err)
@@ -69,6 +82,51 @@ func (c *Client) AuthenticateContext(ctx context.Context) error {
if err != nil { if err != nil {
return err return err
} }
if resp.StatusCode >= 400 {
return statusError(resp.StatusCode, retryAfterOf(resp), "auth: %d %s", resp.StatusCode, truncate(string(raw), 400))
}
var env struct {
Response *Token `json:"response"`
}
if err := json.Unmarshal(raw, &env); err != nil {
return fmt.Errorf("auth decode: %w", err)
}
if env.Response == nil || env.Response.Value == "" {
return fmt.Errorf("auth: empty token in response")
}
c.token = env.Response
return nil
}
// AuthenticateAs verifies a login/password pair against the portal WITHOUT
// mutating the client's cached token. It returns nil when the portal issues a
// token, and the portal error otherwise.
//
// OnlyOffice accepts either the userName or the account email as the login. On
// some portals the account userName login returns HTTP 500 "User authentication
// failed" while the account email succeeds — confirmed for a freshly created
// guest user. Use this probe before sharing credentials (see `oo users check`),
// and prefer the email as the login.
func (c *Client) AuthenticateAs(ctx context.Context, login, password string) error {
body, err := json.Marshal(Credentials{User: login, Password: password})
if err != nil {
return fmt.Errorf("marshal credentials: %w", err)
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL()+"/api/2.0/authentication.json", bytes.NewReader(body))
if err != nil {
return err
}
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Accept", "application/json")
resp, err := c.client.Do(req)
if err != nil {
return fmt.Errorf("auth request: %w", err)
}
defer resp.Body.Close()
raw, err := io.ReadAll(resp.Body)
if err != nil {
return err
}
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return fmt.Errorf("auth: %d %s", resp.StatusCode, truncate(string(raw), 400)) return fmt.Errorf("auth: %d %s", resp.StatusCode, truncate(string(raw), 400))
} }
@@ -81,7 +139,6 @@ func (c *Client) AuthenticateContext(ctx context.Context) error {
if env.Response == nil || env.Response.Value == "" { if env.Response == nil || env.Response.Value == "" {
return fmt.Errorf("auth: empty token in response") return fmt.Errorf("auth: empty token in response")
} }
c.token = env.Response
return nil return nil
} }
@@ -99,17 +156,13 @@ func (c *Client) tokenValid() bool {
// ensureToken refreshes the authentication token when missing or expired. // ensureToken refreshes the authentication token when missing or expired.
// Mirrors the logic inline in Query() but is safe to call from helpers that // Mirrors the logic inline in Query() but is safe to call from helpers that
// bypass the typed Request abstraction. // bypass the typed Request abstraction. It shares AuthenticateContext so the
// transient-retry policy applies to every code path.
func (c *Client) ensureToken() error { func (c *Client) ensureToken() error {
if c.tokenValid() { if c.tokenValid() {
return nil return nil
} }
tok, err := c.Auth(c.credentials) return c.AuthenticateContext(context.Background())
if err != nil {
return err
}
c.token = tok
return nil
} }
// authHeader returns the value for the Authorization header, ensuring a token. // authHeader returns the value for the Authorization header, ensuring a token.
+205
View File
@@ -0,0 +1,205 @@
package onlyoffice
// Project board (Gantt) upsert from a YAML board file.
//
// A board describes projects, their milestones and tasks by exact title. Sync
// creates only what is missing: existing milestones/tasks (matched by title)
// are left untouched, so the file can be the source of truth for a project
// plan and re-applied safely. Dry-run (apply=false) reports counts without
// writing.
//
// The file format is deliberately small and presentation-free:
//
// projects:
// - id: 42
// name: "Example"
// milestones:
// - title: "Kickoff"
// deadline: "2026-01-15"
// key: true
// tasks:
// - title: "Draft"
// start: "2026-01-02"
// deadline: "2026-01-10"
// description: "…"
import (
"context"
"fmt"
"os"
"strconv"
"time"
"gopkg.in/yaml.v3"
)
// Board is a YAML mapping from a project plan onto OnlyOffice milestones/tasks.
type Board struct {
Projects []BoardProject `yaml:"projects"`
}
// BoardProject is one project with its milestones.
type BoardProject struct {
ID int `yaml:"id"`
Name string `yaml:"name,omitempty"`
Milestones []BoardMilestone `yaml:"milestones"`
}
// BoardMilestone is a milestone ("key" marks it as a key milestone).
type BoardMilestone struct {
Title string `yaml:"title"`
Deadline string `yaml:"deadline"`
Key bool `yaml:"key,omitempty"`
Tasks []BoardTask `yaml:"tasks"`
}
// BoardTask is a task inside a milestone. Start falls back to Deadline.
type BoardTask struct {
Title string `yaml:"title"`
Start string `yaml:"start,omitempty"`
Deadline string `yaml:"deadline"`
Description string `yaml:"description,omitempty"`
}
// BoardSyncResult counts what SyncBoard created or skipped.
type BoardSyncResult struct {
CreatedMilestones int
SkippedMilestones int
CreatedTasks int
SkippedTasks int
DryRun bool
}
// LoadBoard reads a board YAML file.
func LoadBoard(path string) (*Board, error) {
b, err := os.ReadFile(path)
if err != nil {
return nil, err
}
return ParseBoard(b)
}
// ParseBoard decodes a board from YAML bytes.
func ParseBoard(data []byte) (*Board, error) {
var board Board
if err := yaml.Unmarshal(data, &board); err != nil {
return nil, fmt.Errorf("board: %w", err)
}
if len(board.Projects) == 0 {
return nil, fmt.Errorf("board: no projects")
}
return &board, nil
}
func boardDay(s string) (Time, error) {
t, err := time.Parse("2006-01-02", s)
if err != nil {
return Time{}, err
}
return Time(t), nil
}
func boardMilestoneIDs(ms []*Milestone) map[string]int64 {
out := map[string]int64{}
for _, m := range ms {
if m == nil || m.Title == nil || m.ID == nil {
continue
}
out[*m.Title] = *m.ID
}
return out
}
func boardTaskTitles(rows []map[string]any) map[string]struct{} {
out := map[string]struct{}{}
for _, r := range rows {
if t, _ := r["title"].(string); t != "" {
out[t] = struct{}{}
}
}
return out
}
// SyncBoard upserts milestones and tasks by exact title. With apply=false it
// only counts what would be created.
func (c *Client) SyncBoard(ctx context.Context, board *Board, apply bool) (*BoardSyncResult, error) {
if board == nil || len(board.Projects) == 0 {
return nil, fmt.Errorf("board: no projects")
}
res := &BoardSyncResult{DryRun: !apply}
for _, p := range board.Projects {
pid := p.ID
existing, err := c.GetProjectMilestones(&Project{ID: &pid})
if err != nil {
return res, fmt.Errorf("project %d milestones: %w", pid, err)
}
haveMS := boardMilestoneIDs(existing)
tasks, err := c.ListTasks(ctx, strconv.Itoa(pid), "")
if err != nil {
return res, fmt.Errorf("project %d tasks: %w", pid, err)
}
haveTask := boardTaskTitles(tasks)
for _, m := range p.Milestones {
msID, ok := haveMS[m.Title]
if !ok {
res.CreatedMilestones++
if apply {
dl, err := boardDay(m.Deadline)
if err != nil {
return res, fmt.Errorf("milestone %q deadline: %w", m.Title, err)
}
created, err := c.CreateMilestone(NewMilestoneRequest{
ProjectID: pid,
Title: m.Title,
Deadline: dl,
IsKey: m.Key,
})
if err != nil {
return res, fmt.Errorf("create milestone %q: %w", m.Title, err)
}
if created.ID != nil {
msID = *created.ID
}
haveMS[m.Title] = msID
}
} else {
res.SkippedMilestones++
}
for _, t := range m.Tasks {
if _, exists := haveTask[t.Title]; exists {
res.SkippedTasks++
continue
}
res.CreatedTasks++
if !apply {
continue
}
start := t.Start
if start == "" {
start = t.Deadline
}
st, err := boardDay(start)
if err != nil {
return res, fmt.Errorf("task %q start: %w", t.Title, err)
}
dl, err := boardDay(t.Deadline)
if err != nil {
return res, fmt.Errorf("task %q deadline: %w", t.Title, err)
}
if _, err := c.CreateProjectTask(NewProjectTaskRequest{
ProjectId: pid,
Title: t.Title,
Description: t.Description,
StartDate: st,
Deadline: dl,
MilestoneId: int(msID),
}); err != nil {
return res, fmt.Errorf("create task %q: %w", t.Title, err)
}
haveTask[t.Title] = struct{}{}
}
}
}
return res, nil
}
+70
View File
@@ -0,0 +1,70 @@
package onlyoffice
import "testing"
func TestParseBoard(t *testing.T) {
b, err := ParseBoard([]byte(`
projects:
- id: 13
name: Example
milestones:
- title: "[lq] Test"
deadline: "2026-01-15"
key: true
tasks:
- title: Draft
start: "2026-01-02"
deadline: "2026-01-10"
description: "…"
`))
if err != nil {
t.Fatal(err)
}
if len(b.Projects) != 1 || b.Projects[0].ID != 13 {
t.Fatalf("projects: %+v", b.Projects)
}
ms := b.Projects[0].Milestones[0]
if ms.Title != "[lq] Test" || !ms.Key || ms.Deadline != "2026-01-15" {
t.Fatalf("milestone: %+v", ms)
}
if len(ms.Tasks) != 1 || ms.Tasks[0].Title != "Draft" || ms.Tasks[0].Start != "2026-01-02" {
t.Fatalf("task: %+v", ms.Tasks)
}
}
func TestParseBoardEmpty(t *testing.T) {
if _, err := ParseBoard([]byte("projects: []")); err == nil {
t.Fatal("expected error for empty board")
}
}
func TestBoardMilestoneIDs(t *testing.T) {
title := "[lq] Test"
id := int64(9)
got := boardMilestoneIDs([]*Milestone{{Title: &title, ID: &id}, nil})
if got[title] != 9 {
t.Fatalf("%v", got)
}
}
func TestBoardTaskTitles(t *testing.T) {
got := boardTaskTitles([]map[string]any{{"title": "a"}, {"title": "b"}, {"nope": 1}})
if _, ok := got["a"]; !ok {
t.Fatal("missing a")
}
if _, ok := got["b"]; !ok {
t.Fatal("missing b")
}
if len(got) != 2 {
t.Fatalf("%v", got)
}
}
func TestBoardDay(t *testing.T) {
if _, err := boardDay("2026-01-15"); err != nil {
t.Fatal(err)
}
if _, err := boardDay("15.01.2026"); err == nil {
t.Fatal("expected error for non-ISO date")
}
}
+28
View File
@@ -0,0 +1,28 @@
package catalog
import "testing"
func TestMergeAddressesDedup(t *testing.T) {
dst := []Address{{Street: "Weg 1", City: "Stadt", Zip: "1"}}
got := mergeAddresses(dst, []Address{
{Street: "weg 1", City: "stadt", Zip: "1"}, // duplicate (case-insensitive)
{Street: "Weg 2", City: "Stadt", Zip: "2"}, // new
})
if len(got) != 2 {
t.Fatalf("got %d addresses, want 2: %+v", len(got), got)
}
if got[1].Street != "Weg 2" {
t.Errorf("second = %+v", got[1])
}
}
func TestEntryAddressesYAML(t *testing.T) {
doc := &Document{Entries: []Entry{{
ID: "person:max maier", Kind: "person", Name: "Max Maier",
Addresses: []Address{{Street: "Weg 1", City: "Stadt", Zip: "12345", Category: "Billing", Primary: true}},
}}}
merged := MergeDocs(doc)
if len(merged.Entries) != 1 || len(merged.Entries[0].Addresses) != 1 {
t.Fatalf("merged = %+v", merged.Entries)
}
}
+25 -15
View File
@@ -115,20 +115,13 @@ func applyCompany(ctx context.Context, client *onlyoffice.Client, e *Entry) (boo
} }
func applyPerson(ctx context.Context, client *onlyoffice.Client, e *Entry) (bool, error) { func applyPerson(ctx context.Context, client *onlyoffice.Client, e *Entry) (bool, error) {
first := strings.TrimSpace(e.First) org := strings.TrimSpace(e.Org)
last := strings.TrimSpace(e.Last) first, last := CleanPersonNames(e.First, e.Last, e.Name, org, e.Emails)
if first == "" && last == "" { e.First, e.Last = first, last
first, last = SplitDisplayName(e.Name)
}
if first == "" {
first = strings.TrimSpace(e.Name)
}
if first == "" { if first == "" {
return false, fmt.Errorf("person missing name") return false, fmt.Errorf("person missing name")
} }
if last == "" { e.Name = strings.TrimSpace(first + " " + strings.Trim(last, "-"))
last = "-"
}
var p map[string]any var p map[string]any
var err error var err error
@@ -144,21 +137,27 @@ func applyPerson(ctx context.Context, client *onlyoffice.Client, e *Entry) (bool
} }
created := false created := false
companyID := 0 companyID := 0
if e.Org != "" { if org != "" {
if co, ferr := client.FindCompany(ctx, e.Org); ferr == nil && co != nil { if co, ferr := client.FindCompany(ctx, org); ferr == nil && co != nil {
companyID, _ = strconv.Atoi(contactIDString(co)) companyID, _ = strconv.Atoi(contactIDString(co))
} }
} }
if p == nil { if p == nil {
about := "" about := ""
if e.Org != "" { if org != "" {
about = "org: " + e.Org about = "org: " + org
} }
p, err = client.CreatePerson(ctx, first, last, companyID, "", about) p, err = client.CreatePerson(ctx, first, last, companyID, "", about)
if err != nil { if err != nil {
return false, err return false, err
} }
created = true created = true
} else {
// Repair names + ensure company link (never encode company in lastName).
id := contactIDString(p)
if _, err := client.UpdatePerson(ctx, id, first, last, companyID, "", ""); err != nil {
return false, fmt.Errorf("update person %s: %w", id, err)
}
} }
id := contactIDString(p) id := contactIDString(p)
e.OOID = id e.OOID = id
@@ -195,5 +194,16 @@ func ensureContactInfos(ctx context.Context, client *onlyoffice.Client, contactI
return fmt.Errorf("add phone %s: %w", ph, err) return fmt.Errorf("add phone %s: %w", ph, err)
} }
} }
for _, a := range e.Addresses {
if strings.TrimSpace(a.Street) == "" && strings.TrimSpace(a.City) == "" && strings.TrimSpace(a.Zip) == "" {
continue
}
if onlyoffice.HasContactAddress(existing, a.Street, a.City, a.Zip, a.Category) {
continue
}
if _, err := client.AddContactAddress(ctx, contactID, a.Street, a.City, a.State, a.Zip, a.Country, a.Category, a.Primary); err != nil {
return fmt.Errorf("add address %s: %w", strings.TrimSpace(a.Street), err)
}
}
return nil return nil
} }
+17 -5
View File
@@ -100,26 +100,38 @@ END:VCARD
func TestScanProjectsRoot(t *testing.T) { func TestScanProjectsRoot(t *testing.T) {
root := t.TempDir() root := t.TempDir()
repo := filepath.Join(root, "produktor-demo") repo := filepath.Join(root, "acme-demo")
if err := os.MkdirAll(filepath.Join(repo, ".git"), 0o755); err != nil { if err := os.MkdirAll(filepath.Join(repo, ".git"), 0o755); err != nil {
t.Fatal(err) t.Fatal(err)
} }
doc, err := ScanProjectsRoot(root, 3) cl := &Classifier{WorkNames: []string{"acme"}}
doc, err := ScanProjectsRootOpts(root, 3, ScanOptions{Classifier: cl})
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
found := false found := false
for _, e := range doc.Entries { for _, e := range doc.Entries {
if e.Kind == "company" && e.Name == "produktor-demo" { if e.Kind == "company" && e.Name == "acme-demo" {
found = true found = true
if e.Role != "work" { if e.Role != "work" || e.Zone != "warm" {
t.Fatalf("role=%q", e.Role) t.Fatalf("role=%q zone=%q", e.Role, e.Zone)
} }
} }
} }
if !found { if !found {
t.Fatalf("missing company: %+v", doc.Entries) t.Fatalf("missing company: %+v", doc.Entries)
} }
// Neutral default leaves it unclassified.
doc, err = ScanProjectsRoot(root, 3)
if err != nil {
t.Fatal(err)
}
for _, e := range doc.Entries {
if e.Kind == "company" && e.Name == "acme-demo" && e.Role != "unknown" {
t.Fatalf("neutral role=%q", e.Role)
}
}
} }
func TestEntryID(t *testing.T) { func TestEntryID(t *testing.T) {
+26
View File
@@ -0,0 +1,26 @@
# Deployment classification rules for `oo catalog scan-projects` /
# `oo catalog scan-thunderbird`. Point OO_CATALOG_CONFIG (or --config) at a copy
# of this file. Keep your real rules out of the repository — they name your
# clients and hosts. The library defaults to no rules (nothing is "work").
#
# work_remotes: a git remote containing any of these substrings → work/hot.
work_remotes:
- git.internal.example
- github.com/acme
#
# work_names: a project directory name containing any of these substrings → work/warm.
work_names:
- acme
#
# mail_orgs: ordered rules for Thunderbird/mbox identities; the first match wins.
# Match by exact `domain`, `suffix` (e.g. ".example.com") and/or display `name`.
# `zone` defaults to hot, `role` to work.
mail_orgs:
- domain: acme.example
org: Acme GmbH
zone: hot
role: work
- suffix: .gov.example
org: Public Sector
zone: warm
role: work
+127
View File
@@ -0,0 +1,127 @@
package catalog
import (
"fmt"
"os"
"strings"
"gopkg.in/yaml.v3"
)
// Classifier maps project trees and mail identities to catalog org/zone/role.
//
// The library ships with neutral defaults: nothing is classified as work unless
// the deployment supplies rules. Those rules are deployment-specific, so they
// live in a YAML config file (path from OO_CATALOG_CONFIG or the --config flag),
// not in the code. See catalog/classify.example.yaml.
type Classifier struct {
// WorkRemotes: a git remote containing any of these substrings → work/hot.
WorkRemotes []string `yaml:"work_remotes,omitempty"`
// WorkNames: a project name containing any of these substrings → work/warm.
WorkNames []string `yaml:"work_names,omitempty"`
// MailOrgs: ordered mail-identity rules; the first match wins.
MailOrgs []MailRule `yaml:"mail_orgs,omitempty"`
}
// MailRule maps an email domain and/or a display-name substring to an org with
// a zone/role. At least one of Domain, Suffix or Name must be set.
type MailRule struct {
Domain string `yaml:"domain,omitempty"` // exact domain, case-insensitive
Suffix string `yaml:"suffix,omitempty"` // domain suffix, e.g. ".example.com"
Name string `yaml:"name,omitempty"` // substring of the display name
Org string `yaml:"org"`
Zone string `yaml:"zone,omitempty"` // default "hot"
Role string `yaml:"role,omitempty"` // default "work"
}
// DefaultClassifier returns the neutral classifier (no deployment rules).
func DefaultClassifier() *Classifier { return &Classifier{} }
// LoadClassifier reads a classifier config from a YAML file.
func LoadClassifier(path string) (*Classifier, error) {
b, err := os.ReadFile(path)
if err != nil {
return nil, err
}
var c Classifier
if err := yaml.Unmarshal(b, &c); err != nil {
return nil, fmt.Errorf("parse classifier config %s: %w", path, err)
}
return &c, nil
}
// LoadClassifierFromEnv loads the classifier named by OO_CATALOG_CONFIG. An
// empty variable yields the neutral classifier.
func LoadClassifierFromEnv() (*Classifier, error) {
path := strings.TrimSpace(os.Getenv("OO_CATALOG_CONFIG"))
if path == "" {
return DefaultClassifier(), nil
}
return LoadClassifier(path)
}
// ClassifyProject returns (role, zone) for a project name and git remote.
// Generic name heuristics come first; deployment rules supply the work cases.
func (c *Classifier) ClassifyProject(name, remote string) (role, zone string) {
lower := strings.ToLower(name)
remoteL := strings.ToLower(remote)
switch {
case strings.Contains(lower, "experiment") || strings.HasPrefix(lower, "test"):
return "experiment", "cold"
case lower == "mama" || lower == "personal" || strings.Contains(lower, "private"):
return "personal", "private"
}
if c != nil {
for _, r := range c.WorkRemotes {
if r != "" && strings.Contains(remoteL, strings.ToLower(r)) {
return "work", "hot"
}
}
for _, n := range c.WorkNames {
if n != "" && strings.Contains(lower, strings.ToLower(n)) {
return "work", "warm"
}
}
}
return "unknown", "warm"
}
// ClassifyMail returns (org, zone, role) for a mail identity. name is the
// display name (may be empty); email is the address.
func (c *Classifier) ClassifyMail(name, email string) (org, zone, role string) {
em := NormalizeEmail(email)
_, domain, _ := strings.Cut(em, "@")
nameL := strings.ToLower(strings.TrimSpace(name))
if c != nil {
for _, r := range c.MailOrgs {
if !mailRuleMatches(r, domain, nameL) {
continue
}
z, ro := r.Zone, r.Role
if z == "" {
z = "hot"
}
if ro == "" {
ro = "work"
}
return r.Org, z, ro
}
}
if strings.HasSuffix(domain, ".de") && looksPublicSector(domain) {
return domain, "warm", "work"
}
return "", "private", "unknown"
}
func mailRuleMatches(r MailRule, domain, nameL string) bool {
if r.Domain != "" && domain == strings.ToLower(strings.TrimSpace(r.Domain)) {
return true
}
if r.Suffix != "" && strings.HasSuffix(domain, strings.ToLower(strings.TrimSpace(r.Suffix))) {
return true
}
if r.Name != "" && nameL != "" && strings.Contains(nameL, strings.ToLower(strings.TrimSpace(r.Name))) {
return true
}
return false
}
+64
View File
@@ -0,0 +1,64 @@
package catalog
import (
"os"
"path/filepath"
"testing"
)
func TestClassifierRules(t *testing.T) {
cl := &Classifier{
WorkRemotes: []string{"git.internal.example"},
WorkNames: []string{"acme"},
MailOrgs: []MailRule{
{Domain: "acme.example", Org: "Acme", Zone: "hot", Role: "work"},
{Suffix: ".gov.example", Org: "Public", Zone: "warm", Role: "work"},
},
}
if role, zone := cl.ClassifyProject("acme-app", "git@git.internal.example:team/acme-app.git"); role != "work" || zone != "hot" {
t.Fatalf("remote: %s/%s", role, zone)
}
if role, zone := cl.ClassifyProject("acme-demo", ""); role != "work" || zone != "warm" {
t.Fatalf("name: %s/%s", role, zone)
}
if role, zone := cl.ClassifyProject("experiment-x", ""); role != "experiment" || zone != "cold" {
t.Fatalf("experiment: %s/%s", role, zone)
}
if org, zone, role := cl.ClassifyMail("", "bob@acme.example"); org != "Acme" || zone != "hot" || role != "work" {
t.Fatalf("mail: %s/%s/%s", org, zone, role)
}
if org, _, _ := cl.ClassifyMail("X", "x@team.gov.example"); org != "Public" {
t.Fatalf("suffix: %s", org)
}
if org, zone, role := cl.ClassifyMail("", "someone@unknown.example"); org != "" || zone != "private" || role != "unknown" {
t.Fatalf("neutral: %s/%s/%s", org, zone, role)
}
}
func TestLoadClassifier(t *testing.T) {
dir := t.TempDir()
path := filepath.Join(dir, "classify.yaml")
if err := os.WriteFile(path, []byte("work_names:\n - acme\nmail_orgs:\n - domain: acme.example\n org: Acme\n"), 0o644); err != nil {
t.Fatal(err)
}
cl, err := LoadClassifier(path)
if err != nil {
t.Fatal(err)
}
if role, _ := cl.ClassifyProject("acme-app", ""); role != "work" {
t.Fatalf("role=%s", role)
}
if org, zone, _ := cl.ClassifyMail("", "a@acme.example"); org != "Acme" || zone != "hot" {
t.Fatalf("org=%s zone=%s", org, zone)
}
t.Setenv("OO_CATALOG_CONFIG", path)
if envCl, err := LoadClassifierFromEnv(); err != nil || envCl == nil || len(envCl.WorkNames) == 0 {
t.Fatalf("env: %+v %v", envCl, err)
}
t.Setenv("OO_CATALOG_CONFIG", "")
neutral, err := LoadClassifierFromEnv()
if err != nil || len(neutral.WorkNames) != 0 || len(neutral.MailOrgs) != 0 {
t.Fatalf("neutral: %+v %v", neutral, err)
}
}
+9 -4
View File
@@ -27,7 +27,7 @@ func MatchAgainstOO(ctx context.Context, client *onlyoffice.Client, doc *Documen
byEmail[NormalizeEmail(em)] = c byEmail[NormalizeEmail(em)] = c
} }
if isCo { if isCo {
key := NormalizeName(fmt.Sprint(c["displayName"])) key := onlyoffice.CompanyGroupingKey(fmt.Sprint(c["displayName"]))
if key != "" { if key != "" {
byCompanyName[key] = c byCompanyName[key] = c
} }
@@ -63,7 +63,7 @@ func MatchAgainstOO(ctx context.Context, client *onlyoffice.Client, doc *Documen
} }
if !matched { if !matched {
if e.Kind == "company" { if e.Kind == "company" {
if c, ok := byCompanyName[NormalizeName(e.Name)]; ok { if c, ok := byCompanyName[onlyoffice.CompanyGroupingKey(e.Name)]; ok {
oo = c oo = c
matched = true matched = true
} }
@@ -92,6 +92,12 @@ func MatchAgainstOO(ctx context.Context, client *onlyoffice.Client, doc *Documen
e.OOID = contactIDString(oo) e.OOID = contactIDString(oo)
continue continue
} }
// Keep a previously applied oo_id (list payloads often omit emails, so
// email match can miss persons that already exist in CRM).
if strings.TrimSpace(e.OOID) != "" {
e.Status = "exists"
continue
}
e.Status = "new" e.Status = "new"
e.OOID = "" e.OOID = ""
} }
@@ -120,8 +126,7 @@ func contactEmails(c map[string]any) []string {
out = append(out, em) out = append(out, em)
} }
for _, row := range onlyoffice.ContactInfoRows(c) { for _, row := range onlyoffice.ContactInfoRows(c) {
t := strings.ToLower(fmt.Sprint(row["infoType"])) if onlyoffice.NormalizeContactInfoType(fmt.Sprint(row["infoType"])) != "email" {
if t != "email" {
continue continue
} }
data := strings.TrimSpace(fmt.Sprint(row["data"])) data := strings.TrimSpace(fmt.Sprint(row["data"]))
+11 -5
View File
@@ -16,6 +16,8 @@ type ScanOptions struct {
MboxHeaders bool MboxHeaders bool
// MboxMaxBytes skips individual mbox files larger than this (0 = 256 MiB default). // MboxMaxBytes skips individual mbox files larger than this (0 = 256 MiB default).
MboxMaxBytes int64 MboxMaxBytes int64
// Classifier supplies deployment classification rules; nil → neutral default.
Classifier *Classifier
} }
// ScanThunderbirdRoot finds Thunderbird profiles under root and emits person rows // ScanThunderbirdRoot finds Thunderbird profiles under root and emits person rows
@@ -37,6 +39,10 @@ func ScanThunderbirdRootOpts(root string, opts ScanOptions) (*Document, error) {
if opts.MboxMaxBytes <= 0 { if opts.MboxMaxBytes <= 0 {
opts.MboxMaxBytes = 256 << 20 opts.MboxMaxBytes = 256 << 20
} }
cl := opts.Classifier
if cl == nil {
cl = DefaultClassifier()
}
var entries []Entry var entries []Entry
seenDB := map[string]struct{}{} seenDB := map[string]struct{}{}
@@ -64,7 +70,7 @@ func ScanThunderbirdRootOpts(root string, opts ScanOptions) (*Document, error) {
return nil return nil
} }
seenDB[path] = struct{}{} seenDB[path] = struct{}{}
parsed, perr := parseGlodaContacts(path) parsed, perr := parseGlodaContacts(path, cl)
if perr != nil { if perr != nil {
entries = append(entries, Entry{ entries = append(entries, Entry{
ID: EntryID("person", "", filepath.Base(path)), ID: EntryID("person", "", filepath.Base(path)),
@@ -84,7 +90,7 @@ func ScanThunderbirdRootOpts(root string, opts ScanOptions) (*Document, error) {
return nil return nil
} }
seenMAB[path] = struct{}{} seenMAB[path] = struct{}{}
parsed, perr := parseMABEmails(path) parsed, perr := parseMABEmails(path, cl)
if perr != nil { if perr != nil {
return nil return nil
} }
@@ -101,7 +107,7 @@ func ScanThunderbirdRootOpts(root string, opts ScanOptions) (*Document, error) {
if info.Size() > opts.MboxMaxBytes { if info.Size() > opts.MboxMaxBytes {
return nil return nil
} }
parsed, perr := parseMboxHeaderEmails(path) parsed, perr := parseMboxHeaderEmails(path, cl)
if perr != nil { if perr != nil {
return nil return nil
} }
@@ -150,7 +156,7 @@ func isLikelyMboxFile(name, path string) bool {
} }
// parseMboxHeaderEmails extracts addresses from From/To/Cc/Reply-To headers only. // parseMboxHeaderEmails extracts addresses from From/To/Cc/Reply-To headers only.
func parseMboxHeaderEmails(path string) ([]Entry, error) { func parseMboxHeaderEmails(path string, cl *Classifier) ([]Entry, error) {
f, err := os.Open(path) f, err := os.Open(path)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -220,7 +226,7 @@ func parseMboxHeaderEmails(path string) ([]Entry, error) {
var out []Entry var out []Entry
for em := range emails { for em := range emails {
org, zone, role := classifyMailIdentity("", em) org, zone, role := cl.ClassifyMail("", em)
out = append(out, Entry{ out = append(out, Entry{
ID: EntryID("person", em, ""), ID: EntryID("person", em, ""),
Kind: "person", Kind: "person",
+7 -7
View File
@@ -10,22 +10,22 @@ func TestParseMboxHeaderEmails(t *testing.T) {
dir := t.TempDir() dir := t.TempDir()
path := filepath.Join(dir, "INBOX") path := filepath.Join(dir, "INBOX")
body := `From - Mon Jul 1 00:00:00 2016 body := `From - Mon Jul 1 00:00:00 2016
From: Axel Schaefer <axel.schaefer@wheregroup.com> From: Alice Smith <alice.smith@acme.example>
To: Andriy Oblivantsev <andriy.oblivantsev@wheregroup.com> To: Bob Jones <bob.jones@acme.example>
Cc: noreply@example.com, client@stadt-example.de Cc: noreply@example.com, client@stadt-example.de
Subject: test Subject: test
Body line ignored Body line ignored
From - Mon Jul 2 00:00:00 2016 From - Mon Jul 2 00:00:00 2016
From: Someone <paul.schmidt@wheregroup.com> From: Someone <paul.schmidt@acme.example>
To: list@wheregroup.com To: list@acme.example
more body more body
` `
if err := os.WriteFile(path, []byte(body), 0o644); err != nil { if err := os.WriteFile(path, []byte(body), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
} }
ents, err := parseMboxHeaderEmails(path) ents, err := parseMboxHeaderEmails(path, DefaultClassifier())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -35,7 +35,7 @@ more body
got[e.Emails[0]] = true got[e.Emails[0]] = true
} }
} }
if !got["axel.schaefer@wheregroup.com"] || !got["andriy.oblivantsev@wheregroup.com"] { if !got["alice.smith@acme.example"] || !got["bob.jones@acme.example"] {
t.Fatalf("%v", got) t.Fatalf("%v", got)
} }
if got["noreply@example.com"] { if got["noreply@example.com"] {
@@ -53,7 +53,7 @@ func TestScanThunderbirdRootOptsMbox(t *testing.T) {
t.Fatal(err) t.Fatal(err)
} }
if err := os.WriteFile(filepath.Join(imap, "INBOX"), []byte( if err := os.WriteFile(filepath.Join(imap, "INBOX"), []byte(
"From - x\nFrom: a@wheregroup.com\nTo: b@wheregroup.com\n\nbody\n", "From - x\nFrom: a@acme.example\nTo: b@acme.example\n\nbody\n",
), 0o644); err != nil { ), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
} }
+152
View File
@@ -0,0 +1,152 @@
package catalog
import (
"regexp"
"strings"
"unicode"
)
var (
parenSuffixRE = regexp.MustCompile(`(?i)\s*[\(\[\{][^)\]\}]*[\)\]\}]\s*$`)
dashCompanyRE = regexp.MustCompile(`(?i)\s+[-–—]\s+[A-Za-z0-9].*$`)
emailLocalRE = regexp.MustCompile(`(?i)^[a-z0-9._%+\-]+@[a-z0-9.\-]+\.[a-z]{2,}$`)
nonNameTokenRE = regexp.MustCompile(`[^a-zA-ZÀ-öø-ÿĀ-ž0-9'’.\-]+`)
)
// CleanPersonNames strips company annotations from display names and fills
// first/last from the email local-part when the source used an address as the
// name. Company affiliation belongs on Org / the CRM companyId — never in LastName.
func CleanPersonNames(first, last, display, org string, emails []string) (cleanFirst, cleanLast string) {
first = strings.TrimSpace(first)
last = strings.TrimSpace(last)
display = strings.TrimSpace(display)
org = strings.TrimSpace(org)
if looksLikeEmail(first) {
ef, el := GuessNameFromEmail(first)
first, last = ef, el
}
if looksLikeEmail(display) && first == "" && last == "" {
display = ""
}
if first == "" && last == "" && display != "" {
first, last = SplitDisplayName(display)
}
first = stripCompanyAnnotation(first, org)
last = stripCompanyAnnotation(last, org)
// "Smith - Acme" / "Jones (Acme)" landed in last.
last = stripCompanyAnnotation(last, org)
if i := strings.IndexAny(first, "(["); i > 0 {
first = strings.TrimSpace(first[:i])
}
// Entire last name is just the company (e.g. last="Acme").
if org != "" && personLastIsOrg(last, org) {
last = ""
}
if (first == "" || looksLikeEmail(first)) && len(emails) > 0 {
ef, el := GuessNameFromEmail(emails[0])
if first == "" || looksLikeEmail(first) {
first = ef
}
if last == "" || last == "-" {
last = el
}
}
first = strings.TrimSpace(first)
last = strings.TrimSpace(last)
if last == "" {
last = "-"
}
return first, last
}
func stripCompanyAnnotation(s, org string) string {
s = strings.TrimSpace(s)
if s == "" {
return ""
}
s = parenSuffixRE.ReplaceAllString(s, "")
s = strings.TrimSpace(s)
s = dashCompanyRE.ReplaceAllString(s, "")
s = strings.TrimSpace(s)
if org != "" {
for _, sep := range []string{" - ", " – ", " — ", " / "} {
if i := strings.LastIndex(strings.ToLower(s), strings.ToLower(sep+org)); i >= 0 {
s = strings.TrimSpace(s[:i])
}
}
suf := " (" + org + ")"
if strings.HasSuffix(strings.ToLower(s), strings.ToLower(suf)) {
s = strings.TrimSpace(s[:len(s)-len(suf)])
}
}
return strings.TrimSpace(s)
}
func personLastIsOrg(last, org string) bool {
last = NormalizeName(last)
org = NormalizeName(org)
if last == "" || org == "" {
return false
}
if last == org {
return true
}
// "Acme" vs "Acme GmbH & Co. KG"
return strings.HasPrefix(org, last+" ") || strings.HasPrefix(org, last+",")
}
func looksLikeEmail(s string) bool {
return emailLocalRE.MatchString(strings.TrimSpace(s))
}
// GuessNameFromEmail turns local@domain into Title-Case first/last when the
// local part looks like first.last / first_last / first-last.
func GuessNameFromEmail(email string) (first, last string) {
email = NormalizeEmail(email)
local, _, ok := strings.Cut(email, "@")
if !ok || local == "" {
return "", ""
}
local = strings.Split(local, "+")[0]
parts := strings.FieldsFunc(local, func(r rune) bool {
return r == '.' || r == '_' || r == '-'
})
if len(parts) == 0 {
return titleToken(local), ""
}
if len(parts) == 1 {
return titleToken(parts[0]), ""
}
return titleToken(parts[0]), titleToken(strings.Join(parts[1:], " "))
}
func titleToken(s string) string {
s = nonNameTokenRE.ReplaceAllString(s, " ")
s = strings.TrimSpace(s)
if s == "" {
return ""
}
runes := []rune(strings.ToLower(s))
runes[0] = unicode.ToTitle(runes[0])
return string(runes)
}
// FormatProjectTitle builds "CC | Company | Title" (spaces around |).
// Country should be a short code (DE, TF, UA, …). Empty segments are dropped.
func FormatProjectTitle(country, company, title string) string {
parts := make([]string, 0, 3)
for _, p := range []string{country, company, title} {
p = strings.TrimSpace(p)
p = strings.ReplaceAll(p, "|", "/")
if p != "" {
parts = append(parts, p)
}
}
return strings.Join(parts, " | ")
}
+41
View File
@@ -0,0 +1,41 @@
package catalog
import "testing"
func TestCleanPersonNamesStripsCompanyParen(t *testing.T) {
f, l := CleanPersonNames("John", "Smith (Acme)", "John Smith (Acme)", "Acme", nil)
if f != "John" || l != "Smith" {
t.Fatalf("got %q %q", f, l)
}
}
func TestCleanPersonNamesStripsDashCompany(t *testing.T) {
f, l := CleanPersonNames("Jens", "Meyer - Acme", "", "Acme", nil)
if f != "Jens" || l != "Meyer" {
t.Fatalf("got %q %q", f, l)
}
}
func TestCleanPersonNamesFromEmail(t *testing.T) {
f, l := CleanPersonNames("david.patzke@acme.example", "-", "", "Acme",
[]string{"david.patzke@acme.example"})
if f != "David" || l != "Patzke" {
t.Fatalf("got %q %q", f, l)
}
}
func TestCleanPersonNamesLastIsCompany(t *testing.T) {
f, l := CleanPersonNames("Thorsten", "Acme", "", "Acme GmbH & Co. KG", nil)
if f != "Thorsten" || l != "-" {
t.Fatalf("got %q %q", f, l)
}
}
func TestFormatProjectTitle(t *testing.T) {
if got := FormatProjectTitle("DE", "Acme", "Golang"); got != "DE | Acme | Golang" {
t.Fatalf("got %q", got)
}
if got := FormatProjectTitle("", "Acme", ""); got != "Acme" {
t.Fatalf("got %q", got)
}
}
+14 -24
View File
@@ -10,10 +10,19 @@ import (
// ScanProjectsRoot finds git roots under root (max depth) and emits company rows. // ScanProjectsRoot finds git roots under root (max depth) and emits company rows.
func ScanProjectsRoot(root string, maxDepth int) (*Document, error) { func ScanProjectsRoot(root string, maxDepth int) (*Document, error) {
return ScanProjectsRootOpts(root, maxDepth, ScanOptions{})
}
// ScanProjectsRootOpts is ScanProjectsRoot with deployment classification rules.
func ScanProjectsRootOpts(root string, maxDepth int, opts ScanOptions) (*Document, error) {
root = filepath.Clean(root) root = filepath.Clean(root)
if maxDepth <= 0 { if maxDepth <= 0 {
maxDepth = 4 maxDepth = 4
} }
cl := opts.Classifier
if cl == nil {
cl = DefaultClassifier()
}
st, err := os.Stat(root) st, err := os.Stat(root)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -23,7 +32,7 @@ func ScanProjectsRoot(root string, maxDepth int) (*Document, error) {
} }
var entries []Entry var entries []Entry
err = walkGitRoots(root, root, 0, maxDepth, &entries) err = walkGitRoots(root, root, 0, maxDepth, cl, &entries)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -43,7 +52,7 @@ func ScanProjectsRoot(root string, maxDepth int) (*Document, error) {
} }
path := filepath.Join(root, name) path := filepath.Join(root, name)
id := EntryID("company", "", name) id := EntryID("company", "", name)
role, zone := classifyProjectName(name, "") role, zone := cl.ClassifyProject(name, "")
entries = append(entries, Entry{ entries = append(entries, Entry{
ID: id, ID: id,
Kind: "company", Kind: "company",
@@ -59,7 +68,7 @@ func ScanProjectsRoot(root string, maxDepth int) (*Document, error) {
return MergeDocs(&Document{Entries: entries}), nil return MergeDocs(&Document{Entries: entries}), nil
} }
func walkGitRoots(root, dir string, depth, maxDepth int, out *[]Entry) error { func walkGitRoots(root, dir string, depth, maxDepth int, cl *Classifier, out *[]Entry) error {
if depth > maxDepth { if depth > maxDepth {
return nil return nil
} }
@@ -67,7 +76,7 @@ func walkGitRoots(root, dir string, depth, maxDepth int, out *[]Entry) error {
if st, err := os.Stat(filepath.Join(dir, ".git")); err == nil && (st.IsDir() || st.Mode().IsRegular()) { if st, err := os.Stat(filepath.Join(dir, ".git")); err == nil && (st.IsDir() || st.Mode().IsRegular()) {
name := filepath.Base(dir) name := filepath.Base(dir)
remote := gitRemoteOrigin(dir) remote := gitRemoteOrigin(dir)
role, zone := classifyProjectName(name, remote) role, zone := cl.ClassifyProject(name, remote)
*out = append(*out, Entry{ *out = append(*out, Entry{
ID: EntryID("company", "", name), ID: EntryID("company", "", name),
Kind: "company", Kind: "company",
@@ -93,7 +102,7 @@ func walkGitRoots(root, dir string, depth, maxDepth int, out *[]Entry) error {
if name == ".git" || name == "node_modules" || name == "vendor" || name == ".venv" || name == "dist" { if name == ".git" || name == "node_modules" || name == "vendor" || name == ".venv" || name == "dist" {
continue continue
} }
_ = walkGitRoots(root, filepath.Join(dir, name), depth+1, maxDepth, out) _ = walkGitRoots(root, filepath.Join(dir, name), depth+1, maxDepth, cl, out)
} }
return nil return nil
} }
@@ -106,22 +115,3 @@ func gitRemoteOrigin(dir string) string {
} }
return strings.TrimSpace(string(b)) return strings.TrimSpace(string(b))
} }
func classifyProjectName(name, remote string) (role, zone string) {
lower := strings.ToLower(name)
remoteL := strings.ToLower(remote)
switch {
case strings.Contains(lower, "experiment") || strings.HasPrefix(lower, "test"):
return "experiment", "cold"
case lower == "mama" || lower == "personal" || strings.Contains(lower, "private"):
return "personal", "private"
case strings.Contains(remoteL, "git.produktor.io") || strings.Contains(remoteL, "github.com/eslider"):
return "work", "hot"
case strings.Contains(lower, "produktor") || strings.Contains(lower, "eslider") ||
strings.Contains(lower, "asesoria") || strings.Contains(lower, "dyvenia") ||
strings.Contains(lower, "onlyoffice"):
return "work", "warm"
default:
return "unknown", "warm"
}
}
+4 -23
View File
@@ -64,7 +64,7 @@ func noisyEmail(email string) bool {
return false return false
} }
func parseMABEmails(path string) ([]Entry, error) { func parseMABEmails(path string, cl *Classifier) ([]Entry, error) {
b, err := os.ReadFile(path) b, err := os.ReadFile(path)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -76,7 +76,7 @@ func parseMABEmails(path string) ([]Entry, error) {
if noisyEmail(em) { if noisyEmail(em) {
continue continue
} }
org, zone, role := classifyMailIdentity("", em) org, zone, role := cl.ClassifyMail("", em)
out = append(out, Entry{ out = append(out, Entry{
ID: EntryID("person", em, ""), ID: EntryID("person", em, ""),
Kind: "person", Kind: "person",
@@ -93,7 +93,7 @@ func parseMABEmails(path string) ([]Entry, error) {
return out, nil return out, nil
} }
func parseGlodaContacts(dbPath string) ([]Entry, error) { func parseGlodaContacts(dbPath string, cl *Classifier) ([]Entry, error) {
// read-only URI; immutable=1 helps when WAL/shm are missing // read-only URI; immutable=1 helps when WAL/shm are missing
dsn := "file:" + dbPath + "?mode=ro&_pragma=query_only(1)" dsn := "file:" + dbPath + "?mode=ro&_pragma=query_only(1)"
db, err := sql.Open("sqlite", dsn) db, err := sql.Open("sqlite", dsn)
@@ -137,7 +137,7 @@ func parseGlodaContacts(dbPath string) ([]Entry, error) {
if display != "" { if display != "" {
first, last = SplitDisplayName(display) first, last = SplitDisplayName(display)
} }
org, zone, role := classifyMailIdentity(display, em) org, zone, role := cl.ClassifyMail(display, em)
out = append(out, Entry{ out = append(out, Entry{
ID: EntryID("person", em, display), ID: EntryID("person", em, display),
Kind: "person", Kind: "person",
@@ -157,25 +157,6 @@ func parseGlodaContacts(dbPath string) ([]Entry, error) {
return out, rows.Err() return out, rows.Err()
} }
func classifyMailIdentity(name, email string) (org, zone, role string) {
em := NormalizeEmail(email)
_, domain, _ := strings.Cut(em, "@")
switch {
case domain == "wheregroup.com" || strings.Contains(strings.ToLower(name), "wheregroup"):
return "WhereGroup", "warm", "work"
case domain == "produktor.io" || domain == "eslider.de" || strings.HasSuffix(domain, ".produktor.io"):
return "produktor.io", "hot", "work"
case domain == "dyvenia.com":
return "Dyvenia", "warm", "work"
case domain == "immowelt.de" || domain == "immowelt.com":
return "Immowelt", "warm", "work"
case strings.HasSuffix(domain, ".de") && looksPublicSector(domain):
return domain, "warm", "work"
default:
return "", "private", "unknown"
}
}
func looksPublicSector(domain string) bool { func looksPublicSector(domain string) bool {
d := strings.ToLower(domain) d := strings.ToLower(domain)
hints := []string{ hints := []string{
+23 -11
View File
@@ -16,8 +16,8 @@ func TestNoisyEmail(t *testing.T) {
if !noisyEmail("x@marketplace.amazon.de") { if !noisyEmail("x@marketplace.amazon.de") {
t.Fatal("amazon marketplace") t.Fatal("amazon marketplace")
} }
if noisyEmail("andriy.oblivantsev@wheregroup.com") { if noisyEmail("alice.smith@acme.example") {
t.Fatal("should keep wheregroup") t.Fatal("should keep a human work address")
} }
} }
@@ -25,14 +25,15 @@ func TestParseMABEmails(t *testing.T) {
dir := t.TempDir() dir := t.TempDir()
path := filepath.Join(dir, "abook.mab") path := filepath.Join(dir, "abook.mab")
body := `// mork junk body := `// mork junk
PrimaryEmail=andriy.oblivantsev@wheregroup.com PrimaryEmail=alice.smith@acme.example
noreply@github.com noreply@github.com
axel.schaefer@wheregroup.com bob.jones@acme.example
` `
if err := os.WriteFile(path, []byte(body), 0o644); err != nil { if err := os.WriteFile(path, []byte(body), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
} }
ents, err := parseMABEmails(path) // Neutral default: no deployment rules → unclassified.
ents, err := parseMABEmails(path, DefaultClassifier())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -40,7 +41,18 @@ func TestParseMABEmails(t *testing.T) {
t.Fatalf("got %d: %+v", len(ents), ents) t.Fatalf("got %d: %+v", len(ents), ents)
} }
for _, e := range ents { for _, e := range ents {
if e.Org != "WhereGroup" || e.Role != "work" { if e.Org != "" || e.Role != "unknown" {
t.Fatalf("%+v", e)
}
}
// A deployment rule classifies the domain as work.
cl := &Classifier{MailOrgs: []MailRule{{Domain: "acme.example", Org: "Acme", Zone: "warm", Role: "work"}}}
ents, err = parseMABEmails(path, cl)
if err != nil {
t.Fatal(err)
}
for _, e := range ents {
if e.Org != "Acme" || e.Role != "work" || e.Zone != "warm" {
t.Fatalf("%+v", e) t.Fatalf("%+v", e)
} }
} }
@@ -56,8 +68,8 @@ func TestParseGlodaContacts(t *testing.T) {
_, err = db.Exec(` _, err = db.Exec(`
CREATE TABLE contacts (id INTEGER PRIMARY KEY, name TEXT); CREATE TABLE contacts (id INTEGER PRIMARY KEY, name TEXT);
CREATE TABLE identities (id INTEGER PRIMARY KEY, contactID INTEGER, kind TEXT, value TEXT); CREATE TABLE identities (id INTEGER PRIMARY KEY, contactID INTEGER, kind TEXT, value TEXT);
INSERT INTO contacts VALUES (1, 'Axel Schaefer'); INSERT INTO contacts VALUES (1, 'Alice Smith');
INSERT INTO identities VALUES (1, 1, 'email', 'axel.schaefer@wheregroup.com'); INSERT INTO identities VALUES (1, 1, 'email', 'alice.smith@acme.example');
INSERT INTO contacts VALUES (2, 'Noise Bot'); INSERT INTO contacts VALUES (2, 'Noise Bot');
INSERT INTO identities VALUES (2, 2, 'email', 'noreply@example.com'); INSERT INTO identities VALUES (2, 2, 'email', 'noreply@example.com');
`) `)
@@ -66,14 +78,14 @@ func TestParseGlodaContacts(t *testing.T) {
} }
_ = db.Close() _ = db.Close()
ents, err := parseGlodaContacts(dbPath) ents, err := parseGlodaContacts(dbPath, DefaultClassifier())
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
if len(ents) != 1 { if len(ents) != 1 {
t.Fatalf("got %d %+v", len(ents), ents) t.Fatalf("got %d %+v", len(ents), ents)
} }
if ents[0].First != "Axel" || ents[0].Emails[0] != "axel.schaefer@wheregroup.com" { if ents[0].First != "Alice" || ents[0].Emails[0] != "alice.smith@acme.example" {
t.Fatalf("%+v", ents[0]) t.Fatalf("%+v", ents[0])
} }
} }
@@ -84,7 +96,7 @@ func TestScanThunderbirdRoot(t *testing.T) {
if err := os.MkdirAll(prof, 0o755); err != nil { if err := os.MkdirAll(prof, 0o755); err != nil {
t.Fatal(err) t.Fatal(err)
} }
if err := os.WriteFile(filepath.Join(prof, "abook.mab"), []byte("mail=paul.schmidt@wheregroup.com\n"), 0o644); err != nil { if err := os.WriteFile(filepath.Join(prof, "abook.mab"), []byte("mail=paul.schmidt@acme.example\n"), 0o644); err != nil {
t.Fatal(err) t.Fatal(err)
} }
doc, err := ScanThunderbirdRoot(root) doc, err := ScanThunderbirdRoot(root)
+33
View File
@@ -12,6 +12,18 @@ import (
"gopkg.in/yaml.v3" "gopkg.in/yaml.v3"
) )
// Address is one postal address of a catalog row. Category is the OO
// AddressCategory label (Home|Postal|Office|Billing|Other|Work).
type Address struct {
Street string `yaml:"street,omitempty"`
City string `yaml:"city,omitempty"`
State string `yaml:"state,omitempty"`
Zip string `yaml:"zip,omitempty"`
Country string `yaml:"country,omitempty"`
Category string `yaml:"category,omitempty"`
Primary bool `yaml:"primary,omitempty"`
}
// Entry is one catalog row (person or company). // Entry is one catalog row (person or company).
type Entry struct { type Entry struct {
ID string `yaml:"id"` ID string `yaml:"id"`
@@ -21,6 +33,7 @@ type Entry struct {
Last string `yaml:"last,omitempty"` Last string `yaml:"last,omitempty"`
Emails []string `yaml:"emails,omitempty"` Emails []string `yaml:"emails,omitempty"`
Phones []string `yaml:"phones,omitempty"` Phones []string `yaml:"phones,omitempty"`
Addresses []Address `yaml:"addresses,omitempty"`
Org string `yaml:"org,omitempty"` Org string `yaml:"org,omitempty"`
Sources []string `yaml:"sources,omitempty"` Sources []string `yaml:"sources,omitempty"`
Zone string `yaml:"zone"` Zone string `yaml:"zone"`
@@ -136,6 +149,7 @@ func mergeEntry(dst, src *Entry) {
dst.Sources = uniqueStrings(append(dst.Sources, src.Sources...)) dst.Sources = uniqueStrings(append(dst.Sources, src.Sources...))
dst.Emails = uniqueEmails(append(dst.Emails, src.Emails...)) dst.Emails = uniqueEmails(append(dst.Emails, src.Emails...))
dst.Phones = uniqueStrings(append(dst.Phones, src.Phones...)) dst.Phones = uniqueStrings(append(dst.Phones, src.Phones...))
dst.Addresses = mergeAddresses(dst.Addresses, src.Addresses)
if dst.First == "" { if dst.First == "" {
dst.First = src.First dst.First = src.First
} }
@@ -176,6 +190,25 @@ func mergeEntry(dst, src *Entry) {
} }
} }
// mergeAddresses appends src addresses not already present (by street/city/zip).
func mergeAddresses(dst, src []Address) []Address {
for _, a := range src {
exists := false
for _, b := range dst {
if strings.EqualFold(strings.TrimSpace(a.Street), strings.TrimSpace(b.Street)) &&
strings.EqualFold(strings.TrimSpace(a.City), strings.TrimSpace(b.City)) &&
strings.EqualFold(strings.TrimSpace(a.Zip), strings.TrimSpace(b.Zip)) {
exists = true
break
}
}
if !exists {
dst = append(dst, a)
}
}
return dst
}
func uniqueEmails(in []string) []string { func uniqueEmails(in []string) []string {
seen := map[string]struct{}{} seen := map[string]struct{}{}
var out []string var out []string
+10 -2
View File
@@ -14,6 +14,7 @@ import (
"net/http/cookiejar" "net/http/cookiejar"
"os" "os"
"strings" "strings"
"sync"
) )
// Client of OnlyOffice API uses credentials to get a token and query the API // Client of OnlyOffice API uses credentials to get a token and query the API
@@ -32,13 +33,20 @@ type Client struct {
defaults Defaults // optional fallbacks for calendar/project IDs defaults Defaults // optional fallbacks for calendar/project IDs
selfID string // cached /api/2.0/people/@self id selfID string // cached /api/2.0/people/@self id
noteCatID int // cached CRM history category id for "note" noteCatID int // cached CRM history category id for "note"
folderTitles map[string]string // cached Documents folder id -> title (F9)
folderTitlesMu sync.Mutex
} }
// NewClient returns a new Client backed by http.DefaultClient. // NewClient returns a new Client whose transport is paced by the process-wide
// rate limiter and 429 cooldown gate (OO_RATE_LIMIT/OO_BURST; see ratelimit.go).
func NewClient(c Credentials) *Client { func NewClient(c Credentials) *Client {
jar, _ := cookiejar.New(nil) jar, _ := cookiejar.New(nil)
return &Client{ return &Client{
client: &http.Client{Jar: jar}, client: &http.Client{
Jar: jar,
Transport: &pacedTransport{base: http.DefaultTransport},
},
credentials: &c, credentials: &c,
} }
} }
+14 -6
View File
@@ -17,6 +17,18 @@ const MailListPageSize = 25
// Loader fetches list items for a menu subject using the OnlyOffice client. // Loader fetches list items for a menu subject using the OnlyOffice client.
type Loader struct { type Loader struct {
Client *onlyoffice.Client Client *onlyoffice.Client
// Files is the backend-agnostic file store used for file download, preview
// and delete. When nil it falls back to Client.FileStore(ProviderREST).
Files onlyoffice.FileStore
}
// fileStore returns the configured file store, defaulting to REST.
func (l *Loader) fileStore() onlyoffice.FileStore {
if l.Files != nil {
return l.Files
}
return l.Client.FileStore(onlyoffice.ProviderREST)
} }
// List returns items for the given list spec (nav leaf). // List returns items for the given list spec (nav leaf).
@@ -171,11 +183,7 @@ func (l *Loader) executeDelete(ctx context.Context, item model.Item) (string, er
} }
return fmt.Sprintf("Deleted message %s", item.Title), nil return fmt.Sprintf("Deleted message %s", item.Title), nil
case model.KindFile: case model.KindFile:
id, err := strconv.Atoi(item.ID) if err := l.fileStore().Delete(ctx, []string{item.ID}); err != nil {
if err != nil {
return "", err
}
if err := l.Client.DeleteFiles(ctx, []int{id}); err != nil {
return "", err return "", err
} }
return fmt.Sprintf("Deleted file %s", item.Title), nil return fmt.Sprintf("Deleted file %s", item.Title), nil
@@ -199,7 +207,7 @@ func (l *Loader) executeDownload(ctx context.Context, item model.Item, destPath
return "", err return "", err
} }
defer f.Close() defer f.Close()
if _, err := l.Client.DownloadFile(ctx, item.ID, f); err != nil { if _, err := l.fileStore().Download(ctx, item.ID, f); err != nil {
return "", err return "", err
} }
return fmt.Sprintf("Downloaded to %s", destPath), nil return fmt.Sprintf("Downloaded to %s", destPath), nil
+3 -6
View File
@@ -5,7 +5,6 @@ import (
"context" "context"
"fmt" "fmt"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/eslider/go-onlyoffice/cmd/office/model" "github.com/eslider/go-onlyoffice/cmd/office/model"
"github.com/eslider/go-onlyoffice/cmd/office/preview" "github.com/eslider/go-onlyoffice/cmd/office/preview"
) )
@@ -30,13 +29,11 @@ func (l *Loader) filePreviewMarkdown(ctx context.Context, item model.Item) (stri
return "", fmt.Errorf("file id missing") return "", fmt.Errorf("file id missing")
} }
name := item.Title name := item.Title
if meta, err := l.Client.GetFile(ctx, item.ID); err == nil && meta != nil { if e, err := l.fileStore().Stat(ctx, item.ID); err == nil && e.Title != "" {
if t := onlyoffice.FileEntryTitle(meta); t != "" { name = e.Title
name = t
}
} }
var buf bytes.Buffer var buf bytes.Buffer
if _, err := l.Client.DownloadFile(ctx, item.ID, &buf); err != nil { if _, err := l.fileStore().Download(ctx, item.ID, &buf); err != nil {
return "", err return "", err
} }
return preview.FileBytesToMarkdown(name, buf.Bytes()) return preview.FileBytesToMarkdown(name, buf.Bytes())
+3 -7
View File
@@ -4,7 +4,6 @@ package fetch_test
import ( import (
"context" "context"
"os"
"testing" "testing"
"github.com/eslider/go-onlyoffice/cmd/office/model" "github.com/eslider/go-onlyoffice/cmd/office/model"
@@ -46,9 +45,6 @@ func TestIntegrationUpdateTaskTitleDescription(t *testing.T) {
} }
func TestIntegrationTaskFieldsFromLiveAPI(t *testing.T) { func TestIntegrationTaskFieldsFromLiveAPI(t *testing.T) {
if os.Getenv("ONLYOFFICE_URL") == "" && os.Getenv("ONLYOFFICE_HOST") == "" {
t.Skip("ONLYOFFICE_URL not set")
}
loader, ctx := liveLoader(t) loader, ctx := liveLoader(t)
items, err := loader.List(ctx, model.ListSpec{Subject: model.SubjectTasks}) items, err := loader.List(ctx, model.ListSpec{Subject: model.SubjectTasks})
if err != nil { if err != nil {
@@ -57,12 +53,12 @@ func TestIntegrationTaskFieldsFromLiveAPI(t *testing.T) {
if len(items) == 0 { if len(items) == 0 {
t.Skip("no tasks") t.Skip("no tasks")
} }
title, desc, err := loader.TaskFields(ctx, items[0]) fields, err := loader.DetailForm(ctx, items[0])
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
if title == "" { if fields.Primary == "" {
t.Fatal("empty title") t.Fatal("empty title")
} }
_ = desc _ = fields.Secondary
} }
+25 -3
View File
@@ -64,6 +64,7 @@ func catalogScanContactsCmd() *cobra.Command {
func catalogScanProjectsCmd() *cobra.Command { func catalogScanProjectsCmd() *cobra.Command {
var outPath string var outPath string
var maxDepth int var maxDepth int
var configPath string
cmd := &cobra.Command{ cmd := &cobra.Command{
Use: "scan-projects", Use: "scan-projects",
Short: "Git roots / remotes / top-level dirs → company rows", Short: "Git roots / remotes / top-level dirs → company rows",
@@ -72,7 +73,11 @@ func catalogScanProjectsCmd() *cobra.Command {
if root == "" { if root == "" {
return fmt.Errorf("--root is required") return fmt.Errorf("--root is required")
} }
doc, err := catalog.ScanProjectsRoot(root, maxDepth) cl, err := catalogClassifier(configPath)
if err != nil {
return err
}
doc, err := catalog.ScanProjectsRootOpts(root, maxDepth, catalog.ScanOptions{Classifier: cl})
if err != nil { if err != nil {
return err return err
} }
@@ -81,6 +86,7 @@ func catalogScanProjectsCmd() *cobra.Command {
} }
cmd.Flags().String("root", "", "projects directory (local path)") cmd.Flags().String("root", "", "projects directory (local path)")
cmd.Flags().IntVar(&maxDepth, "max-depth", 4, "max directory depth for git roots") cmd.Flags().IntVar(&maxDepth, "max-depth", 4, "max directory depth for git roots")
cmd.Flags().StringVar(&configPath, "config", "", "classification rules YAML (default $OO_CATALOG_CONFIG)")
cmd.Flags().StringVarP(&outPath, "out", "O", "", "write YAML to this path") cmd.Flags().StringVarP(&outPath, "out", "O", "", "write YAML to this path")
_ = cmd.MarkFlagRequired("root") _ = cmd.MarkFlagRequired("root")
return cmd return cmd
@@ -89,6 +95,7 @@ func catalogScanProjectsCmd() *cobra.Command {
func catalogScanThunderbirdCmd() *cobra.Command { func catalogScanThunderbirdCmd() *cobra.Command {
var outPath string var outPath string
var mboxHeaders bool var mboxHeaders bool
var configPath string
cmd := &cobra.Command{ cmd := &cobra.Command{
Use: "scan-thunderbird", Use: "scan-thunderbird",
Short: "Thunderbird profiles: abook/history.mab + Gloda SQLite contacts", Short: "Thunderbird profiles: abook/history.mab + Gloda SQLite contacts",
@@ -98,13 +105,18 @@ func catalogScanThunderbirdCmd() *cobra.Command {
- optional --mbox-headers: From/To/Cc/Reply-To from mbox folder files (no bodies) - optional --mbox-headers: From/To/Cc/Reply-To from mbox folder files (no bodies)
Noisy senders (noreply, Amazon marketplace, GitHub reply, …) are skipped. Noisy senders (noreply, Amazon marketplace, GitHub reply, …) are skipped.
Default zone is private; known work domains (e.g. wheregroup.com) get zone=warm role=work.`, Default zone is private; work domains/names are classified from the rules in
$OO_CATALOG_CONFIG (or --config) — none are hardcoded.`,
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
root, _ := cmd.Flags().GetString("root") root, _ := cmd.Flags().GetString("root")
if root == "" { if root == "" {
return fmt.Errorf("--root is required") return fmt.Errorf("--root is required")
} }
doc, err := catalog.ScanThunderbirdRootOpts(root, catalog.ScanOptions{MboxHeaders: mboxHeaders}) cl, err := catalogClassifier(configPath)
if err != nil {
return err
}
doc, err := catalog.ScanThunderbirdRootOpts(root, catalog.ScanOptions{MboxHeaders: mboxHeaders, Classifier: cl})
if err != nil { if err != nil {
return err return err
} }
@@ -113,11 +125,21 @@ Default zone is private; known work domains (e.g. wheregroup.com) get zone=warm
} }
cmd.Flags().String("root", "", "Thunderbird profile or parent directory (local path)") cmd.Flags().String("root", "", "Thunderbird profile or parent directory (local path)")
cmd.Flags().BoolVar(&mboxHeaders, "mbox-headers", false, "also extract emails from mbox From/To/Cc headers") cmd.Flags().BoolVar(&mboxHeaders, "mbox-headers", false, "also extract emails from mbox From/To/Cc headers")
cmd.Flags().StringVar(&configPath, "config", "", "classification rules YAML (default $OO_CATALOG_CONFIG)")
cmd.Flags().StringVarP(&outPath, "out", "O", "", "write YAML to this path") cmd.Flags().StringVarP(&outPath, "out", "O", "", "write YAML to this path")
_ = cmd.MarkFlagRequired("root") _ = cmd.MarkFlagRequired("root")
return cmd return cmd
} }
// catalogClassifier loads classification rules from --config, else
// $OO_CATALOG_CONFIG, else the neutral default (no rules).
func catalogClassifier(configPath string) (*catalog.Classifier, error) {
if configPath == "" {
return catalog.LoadClassifierFromEnv()
}
return catalog.LoadClassifier(configPath)
}
func catalogMergeCmd() *cobra.Command { func catalogMergeCmd() *cobra.Command {
var inputs []string var inputs []string
var outPath string var outPath string
+3 -1
View File
@@ -31,7 +31,9 @@ func init() {
} }
// execute runs the root command. Exported only to main.go in the same package. // execute runs the root command. Exported only to main.go in the same package.
func execute() error { return rootCmd.Execute() } // .env is loaded CLI-wide so non-authenticating commands (e.g. catalog scans)
// also see configuration such as OO_CATALOG_CONFIG.
func execute() error { bootstrap.LoadEnv(); return rootCmd.Execute() }
// newOO loads env (only .env in CWD) and returns an authenticated client. // newOO loads env (only .env in CWD) and returns an authenticated client.
// godotenv is a CLI-only concern; the library itself never loads dotfiles. // godotenv is a CLI-only concern; the library itself never loads dotfiles.
+1 -1
View File
@@ -7,7 +7,7 @@ import (
var crmCmd = &cobra.Command{ var crmCmd = &cobra.Command{
Use: "crm", Use: "crm",
Short: "CRM maintenance (dedupe, cleanup)", Short: "CRM maintenance (audit, dedupe, cleanup)",
} }
func init() { func init() {
+63
View File
@@ -0,0 +1,63 @@
package main
import (
"encoding/json"
"fmt"
"os"
"sort"
"github.com/spf13/cobra"
)
func init() {
crmCmd.AddCommand(crmAuditCmd())
}
func crmAuditCmd() *cobra.Command {
var outPath string
cmd := &cobra.Command{
Use: "audit",
Short: "Audit opportunities (files/tasks/members per deal) and classify",
Long: `Lists every opportunity with its file/task/member counts and a coarse class
(ok | dup | empty | junk-title). --out writes the full JSON audit for later use.`,
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
audits, err := c.AuditOpportunities(cmd.Context())
if err != nil {
return err
}
if outPath != "" {
b, merr := json.MarshalIndent(audits, "", " ")
if merr != nil {
return merr
}
if werr := os.WriteFile(outPath, b, 0o644); werr != nil {
return werr
}
}
byClass := map[string]int{}
for _, a := range audits {
byClass[a.Class]++
}
classes := make([]string, 0, len(byClass))
for k := range byClass {
classes = append(classes, k)
}
sort.Strings(classes)
rows := make([]map[string]any, 0, len(classes))
for _, k := range classes {
rows = append(rows, map[string]any{"class": k, "count": byClass[k]})
}
printTable([]string{"class", "count"}, rows)
if outPath != "" {
fmt.Fprintf(cmd.OutOrStdout(), "wrote %d audits → %s\n", len(audits), outPath)
}
return nil
},
}
cmd.Flags().StringVar(&outPath, "out", "", "write audit JSON to this file")
return cmd
}
+512
View File
@@ -0,0 +1,512 @@
package main
import (
"context"
"fmt"
"os"
"strings"
"time"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra"
)
func init() {
rootCmd.AddCommand(davCmd())
}
// davCmd exposes the Documents module through the backend-agnostic FileStore
// (DAV backend). The underlying Dav calls are the oo-webdav proven path:
// MoveDavItems sends resolveType=Skip + holdResult=true, which the legacy
// fileops/move call without those params silently ignores (200 without move).
func davCmd() *cobra.Command {
cmd := &cobra.Command{
Use: "dav",
Short: "Documents module by folder/file id (oo-webdav proven path)",
}
cmd.AddCommand(davLsCmd())
cmd.AddCommand(davMoveCmd())
cmd.AddCommand(davCopyCmd())
cmd.AddCommand(davMkdirCmd())
cmd.AddCommand(davEnsurePathCmd())
cmd.AddCommand(davUploadCmd())
cmd.AddCommand(davRemoveCmd())
cmd.AddCommand(davRenameFileCmd())
cmd.AddCommand(davRenameFolderCmd())
cmd.AddCommand(davDownloadCmd())
cmd.AddCommand(davFileOpsCmd())
return cmd
}
func davLsCmd() *cobra.Command {
cmd := &cobra.Command{
Use: "ls FOLDER_ID",
Short: "List a Documents folder (@root for virtual sections)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
ctx := cmd.Context()
if args[0] == "@root" {
sections, err := c.ListDavSections(ctx)
if err != nil {
return err
}
rows := make([]map[string]any, 0, len(sections))
for _, s := range sections {
rows = append(rows, map[string]any{
"id": s.ID,
"title": s.Title,
"filesCount": s.FilesCount,
"foldersCount": s.FoldersCount,
})
}
printTable([]string{"id", "title", "filesCount", "foldersCount"}, rows)
return nil
}
entries, err := c.FileStore(onlyoffice.ProviderDAV).List(ctx, args[0])
if err != nil {
return err
}
folders := make([]onlyoffice.Entry, 0, len(entries))
files := make([]onlyoffice.Entry, 0, len(entries))
for _, e := range entries {
if e.Kind == onlyoffice.Folder {
folders = append(folders, e)
} else {
files = append(files, e)
}
}
frows := make([]map[string]any, 0, len(folders))
for _, f := range folders {
frows = append(frows, map[string]any{
"id": f.ID,
"title": f.Title,
"filesCount": f.FilesCount,
"foldersCount": f.FoldersCount,
})
}
rows := make([]map[string]any, 0, len(files))
for _, f := range files {
rows = append(rows, map[string]any{
"id": f.ID,
"title": f.Title,
"size": f.Size,
"updated": entryUpdated(f),
})
}
if outputFormat == "json" {
printObject(map[string]any{"folders": frows, "files": rows})
return nil
}
if len(frows) > 0 {
fmt.Println("folders:")
printTable([]string{"id", "title", "filesCount", "foldersCount"}, frows)
}
fmt.Println("files:")
printTable([]string{"id", "title", "size", "updated"}, rows)
return nil
},
}
return cmd
}
// entryUpdated prefers the backend-native timestamp string so table/JSON output
// round-trips what the API returned.
func entryUpdated(e onlyoffice.Entry) string {
if e.Updated != "" {
return e.Updated
}
if e.Modified.IsZero() {
return ""
}
return e.Modified.Format(time.RFC3339)
}
func davMoveCmd() *cobra.Command {
var folderIDs []string
cmd := &cobra.Command{
Use: "move DEST_FOLDER_ID FILE_ID [FILE_ID...]",
Short: "Move file(s) into a Documents folder (resolveType=Skip, holdResult)",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
ids := append(append([]string{}, folderIDs...), args[1:]...)
if err := c.FileStore(onlyoffice.ProviderDAV).Move(cmd.Context(), ids, args[0]); err != nil {
return err
}
printObject(map[string]any{"moved_files": args[1:], "moved_folders": folderIDs, "dest": args[0]})
return nil
},
}
cmd.Flags().StringSliceVar(&folderIDs, "folders", nil, "folder ids to move along with the files")
return cmd
}
func davCopyCmd() *cobra.Command {
var folderIDs []string
cmd := &cobra.Command{
Use: "copy DEST_FOLDER_ID FILE_ID [FILE_ID...]",
Short: "Copy file(s) into a Documents folder (conflictResolveType=Skip)",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
ids := append(append([]string{}, folderIDs...), args[1:]...)
if err := c.FileStore(onlyoffice.ProviderDAV).Copy(cmd.Context(), ids, args[0]); err != nil {
return err
}
printObject(map[string]any{"copied_files": args[1:], "copied_folders": folderIDs, "dest": args[0]})
return nil
},
}
cmd.Flags().StringSliceVar(&folderIDs, "folders", nil, "folder ids to copy along with the files")
return cmd
}
func davMkdirCmd() *cobra.Command {
return &cobra.Command{
Use: "mkdir PARENT_FOLDER_ID TITLE",
Short: "Create a subfolder in Documents",
Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
f, err := c.FileStore(onlyoffice.ProviderDAV).CreateFolder(cmd.Context(), args[0], args[1])
if err != nil {
return err
}
printObject(map[string]any{"id": f.ID, "title": f.Title, "parent": args[0]})
return nil
},
}
}
func davEnsurePathCmd() *cobra.Command {
var under string
cmd := &cobra.Command{
Use: "ensure-path PATH",
Short: "Resolve or create nested Documents folders (mkdir -p), print the final folder id",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
root := under
if root == "" {
root, err = myDocumentsID(cmd.Context(), c)
if err != nil {
return err
}
}
f, err := ensurePath(cmd.Context(), c.FileStore(onlyoffice.ProviderDAV), root, args[0])
if err != nil {
return err
}
printObject(map[string]any{"id": f.ID, "title": f.Title, "under": root})
return nil
},
}
cmd.Flags().StringVar(&under, "under", "", "parent folder id (default: My documents)")
return cmd
}
func davUploadCmd() *cobra.Command {
var replace bool
cmd := &cobra.Command{
Use: "upload DEST_FOLDER_ID LOCAL_FILE [LOCAL_FILE...]",
Short: "Upload local file(s) into a Documents folder (multipart, upsert by name)",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
results, err := uploadLocal(cmd.Context(), c, args[0], args[1:], replace)
if err != nil {
return err
}
rows := make([]map[string]any, 0, len(results))
for _, r := range results {
replaced := make([]string, 0, len(r.Replaced))
for _, id := range r.Replaced {
replaced = append(replaced, fmt.Sprint(id))
}
rows = append(rows, map[string]any{
"id": r.Entry.ID,
"title": r.Entry.Title,
"size": r.Entry.Size,
"replaced": strings.Join(replaced, ", "),
})
}
printTable([]string{"id", "title", "size", "replaced"}, rows)
return nil
},
}
cmd.Flags().BoolVar(&replace, "replace", true, "replace the same logical file (stem, legacy↔OOXML ext) before upload (default); false = fail if name taken")
return cmd
}
// uploadedFile pairs an upload result with the ids deleted to make room for it.
type uploadedFile struct {
Entry onlyoffice.Entry
Replaced []int
}
// uploader is the slice of *onlyoffice.Client that dav upload needs, so the
// command logic is unit-testable with a fake.
type uploader interface {
UploadToFolderReplacing(ctx context.Context, folderID, localPath string) (*onlyoffice.FileEntry, []int, error)
AssertNoFileConflict(ctx context.Context, folderID, localPath string) error
UploadToFolder(ctx context.Context, folderID, localPath string) (*onlyoffice.FileEntry, error)
}
// uploadLocal uploads each local file into folderID. With replace the existing
// same logical file is overwritten in place (the UploadToFolderReplacing
// upsert, which matches the server-converted extension too); without replace a
// name clash fails with ErrFileExists before anything is sent.
func uploadLocal(ctx context.Context, up uploader, folderID string, paths []string, replace bool) ([]uploadedFile, error) {
results := make([]uploadedFile, 0, len(paths))
for _, path := range paths {
var (
fe *onlyoffice.FileEntry
deleted []int
err error
)
if replace {
fe, deleted, err = up.UploadToFolderReplacing(ctx, folderID, path)
} else {
if err := up.AssertNoFileConflict(ctx, folderID, path); err != nil {
return nil, err
}
fe, err = up.UploadToFolder(ctx, folderID, path)
}
if err != nil {
return nil, fmt.Errorf("dav upload: %s: %w", path, err)
}
results = append(results, uploadedFile{
Entry: onlyoffice.FileEntryToEntry(fe, onlyoffice.ProviderDAV),
Replaced: deleted,
})
}
return results, nil
}
// ensurePath resolves-or-creates every "/"-separated segment under rootID
// through store. Existing folders are reused by title (case-insensitive), so
// repeated calls return the same id without creating duplicates.
func ensurePath(ctx context.Context, store onlyoffice.FileStore, rootID, path string) (onlyoffice.Entry, error) {
segments, err := splitDavPath(path)
if err != nil {
return onlyoffice.Entry{}, err
}
parent := rootID
var current onlyoffice.Entry
for _, name := range segments {
entries, err := store.List(ctx, parent)
if err != nil {
return onlyoffice.Entry{}, fmt.Errorf("dav ensure-path: list %s: %w", parent, err)
}
found := false
for _, e := range entries {
if e.Kind == onlyoffice.Folder && strings.EqualFold(strings.TrimSpace(e.Title), name) {
current, found = e, true
break
}
}
if !found {
current, err = store.CreateFolder(ctx, parent, name)
if err != nil {
return onlyoffice.Entry{}, fmt.Errorf("dav ensure-path: mkdir %s/%s: %w", parent, name, err)
}
}
parent = current.ID
}
return current, nil
}
// splitDavPath normalizes a Documents path into non-empty segments. "." is
// ignored; ".." is rejected rather than creating a literal folder named "..".
func splitDavPath(path string) ([]string, error) {
var segments []string
for _, s := range strings.Split(path, "/") {
s = strings.TrimSpace(s)
switch s {
case "", ".":
continue
case "..":
return nil, fmt.Errorf("dav ensure-path: %q not allowed", s)
}
segments = append(segments, s)
}
if len(segments) == 0 {
return nil, fmt.Errorf("dav ensure-path: empty path")
}
return segments, nil
}
// myDocumentsID resolves the concrete id of the "My documents" section so
// ensure-path can create folders under it (the symbolic "@my" is list-only on
// some portals). rootFolderType 5 = My; the title match covers portals that
// omit rootFolderType.
func myDocumentsID(ctx context.Context, c *onlyoffice.Client) (string, error) {
sections, err := c.ListDavSections(ctx)
if err != nil {
return "", err
}
for _, s := range sections {
if s.RootType == 5 && s.ID != "" {
return s.ID, nil
}
}
for _, s := range sections {
if strings.EqualFold(strings.TrimSpace(s.Title), "My documents") && s.ID != "" {
return s.ID, nil
}
}
return "", fmt.Errorf("dav: no My documents section in @root; pass --under FOLDER_ID")
}
func davRemoveCmd() *cobra.Command {
var folderIDs []string
cmd := &cobra.Command{
Use: "rm [FILE_ID...]",
Aliases: []string{"delete"},
Short: "Permanently delete file(s) and/or folder(s) from Documents",
Args: cobra.ArbitraryArgs,
RunE: func(cmd *cobra.Command, args []string) error {
if len(args) == 0 && len(folderIDs) == 0 {
return fmt.Errorf("dav rm: give at least one FILE_ID or --folders")
}
c, err := newOO(cmd)
if err != nil {
return err
}
ids := append(append([]string{}, args...), folderIDs...)
if err := c.FileStore(onlyoffice.ProviderDAV).Delete(cmd.Context(), ids); err != nil {
return err
}
printObject(map[string]any{"deleted_files": args, "deleted_folders": folderIDs})
return nil
},
}
cmd.Flags().StringSliceVar(&folderIDs, "folders", nil, "folder ids to delete")
return cmd
}
func davRenameFileCmd() *cobra.Command {
return &cobra.Command{
Use: "rename-file FILE_ID NEW_TITLE",
Short: "Rename a Documents file (include extension in NEW_TITLE)",
Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
if err := c.FileStore(onlyoffice.ProviderDAV).Rename(cmd.Context(), args[0], args[1]); err != nil {
return err
}
printObject(map[string]any{"id": args[0], "title": args[1]})
return nil
},
}
}
func davRenameFolderCmd() *cobra.Command {
return &cobra.Command{
Use: "rename-folder FOLDER_ID NEW_TITLE",
Short: "Rename a Documents folder",
Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
if err := c.FileStore(onlyoffice.ProviderDAV).Rename(cmd.Context(), args[0], args[1]); err != nil {
return err
}
printObject(map[string]any{"id": args[0], "title": args[1]})
return nil
},
}
}
func davDownloadCmd() *cobra.Command {
var to string
cmd := &cobra.Command{
Use: "download FILE_ID",
Short: "Download Documents file bytes (default path: ./<title>)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
ctx := cmd.Context()
store := c.FileStore(onlyoffice.ProviderDAV)
e, err := store.Stat(ctx, args[0])
if err != nil {
return err
}
path := to
if path == "" {
path = onlyoffice.SafeLocalFileName(e.Title)
}
out, err := os.Create(path)
if err != nil {
return err
}
defer out.Close()
n, err := store.Download(ctx, args[0], out)
if err != nil {
_ = os.Remove(path)
return err
}
printObject(map[string]any{"path": path, "bytes": n})
return nil
},
}
cmd.Flags().StringVar(&to, "to", "", "output path (default: ./<server title>)")
return cmd
}
func davFileOpsCmd() *cobra.Command {
return &cobra.Command{
Use: "fileops",
Short: "List active file operations (move/copy status polling)",
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
ops, err := c.ListFileOps(cmd.Context())
if err != nil {
return err
}
rows := make([]map[string]any, 0, len(ops))
for _, op := range ops {
rows = append(rows, map[string]any{
"id": fmt.Sprint(op["id"]),
"operation": fmt.Sprint(op["operation"]),
"progress": fmt.Sprint(op["progress"]),
"finished": fmt.Sprint(op["finished"]),
"error": fmt.Sprint(op["error"]),
})
}
printTable([]string{"id", "operation", "progress", "finished", "error"}, rows)
return nil
},
}
}
+121
View File
@@ -0,0 +1,121 @@
//go:build integration
package main
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"time"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/eslider/go-onlyoffice/cmd/internal/bootstrap"
)
// TestIntegrationDavEnsurePathUpload exercises the dav ensure-path/upload
// helpers against a live instance: idempotent folder creation, upload, replace
// and no-replace conflict. Destructive (creates and removes a throwaway
// folder); skips cleanly without credentials. Only run against instances you
// own: go test -tags=integration ./cmd/oo -run TestIntegrationDavEnsurePathUpload
func TestIntegrationDavEnsurePathUpload(t *testing.T) {
ctx := context.Background()
c, err := bootstrap.NewClient(ctx)
if err != nil {
t.Skipf("no live OnlyOffice credentials: %v", err)
}
root, err := myDocumentsID(ctx, c)
if err != nil {
t.Skipf("My documents root unavailable: %v", err)
}
store := c.FileStore(onlyoffice.ProviderDAV)
suffix := time.Now().UTC().Format("20060102-150405")
base := "oo-it-" + suffix
first, err := ensurePath(ctx, store, root, base+"/nested")
if err != nil {
t.Fatalf("ensurePath: %v", err)
}
baseID := findFolderID(t, ctx, store, root, base)
t.Cleanup(func() { _ = store.Delete(context.Background(), []string{baseID}) })
second, err := ensurePath(ctx, store, root, base+"/nested")
if err != nil {
t.Fatalf("ensurePath (second): %v", err)
}
if second.ID != first.ID {
t.Fatalf("ensure-path not idempotent: %s != %s", second.ID, first.ID)
}
if n := countFolders(t, ctx, store, baseID, "nested"); n != 1 {
t.Fatalf("nested folder duplicated: %d children named nested, want 1", n)
}
local := filepath.Join(t.TempDir(), "oo-it-"+suffix+".xlsx")
if err := os.WriteFile(local, []byte("integration dav upload "+suffix+"\n"), 0o600); err != nil {
t.Fatal(err)
}
name := filepath.Base(local)
if _, err := uploadLocal(ctx, c, first.ID, []string{local}, true); err != nil {
t.Fatalf("uploadLocal: %v", err)
}
if n := countFiles(t, ctx, store, first.ID, name); n != 1 {
t.Fatalf("after upload: %d files named %s, want 1", n, name)
}
if _, err := uploadLocal(ctx, c, first.ID, []string{local}, true); err != nil {
t.Fatalf("uploadLocal (replace): %v", err)
}
if n := countFiles(t, ctx, store, first.ID, name); n != 1 {
t.Fatalf("after replace: %d files named %s, want 1", n, name)
}
if _, err := uploadLocal(ctx, c, first.ID, []string{local}, false); !errors.Is(err, onlyoffice.ErrFileExists) {
t.Fatalf("no-replace conflict err = %v, want ErrFileExists", err)
}
}
func findFolderID(t *testing.T, ctx context.Context, store onlyoffice.FileStore, parent, title string) string {
t.Helper()
for _, e := range mustList(t, ctx, store, parent) {
if e.Kind == onlyoffice.Folder && strings.EqualFold(e.Title, title) {
return e.ID
}
}
t.Fatalf("folder %q not found under %s", title, parent)
return ""
}
func countFolders(t *testing.T, ctx context.Context, store onlyoffice.FileStore, parent, title string) int {
t.Helper()
n := 0
for _, e := range mustList(t, ctx, store, parent) {
if e.Kind == onlyoffice.Folder && strings.EqualFold(e.Title, title) {
n++
}
}
return n
}
func countFiles(t *testing.T, ctx context.Context, store onlyoffice.FileStore, parent, title string) int {
t.Helper()
n := 0
for _, e := range mustList(t, ctx, store, parent) {
if e.Kind == onlyoffice.File && e.Title == title {
n++
}
}
return n
}
func mustList(t *testing.T, ctx context.Context, store onlyoffice.FileStore, parent string) []onlyoffice.Entry {
t.Helper()
entries, err := store.List(ctx, parent)
if err != nil {
t.Fatalf("List(%s): %v", parent, err)
}
return entries
}
+260
View File
@@ -0,0 +1,260 @@
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"strings"
"testing"
onlyoffice "github.com/eslider/go-onlyoffice"
)
// fakeFolderStore is an in-memory onlyoffice.FileStore for the ensure-path
// logic. It records created folders so idempotency can be asserted.
type fakeFolderStore struct {
nextID int
entries map[string][]onlyoffice.Entry
created []string
}
func newFakeFolderStore() *fakeFolderStore {
return &fakeFolderStore{entries: map[string][]onlyoffice.Entry{}}
}
func (f *fakeFolderStore) Name() string { return "fake" }
func (f *fakeFolderStore) List(_ context.Context, parent string) ([]onlyoffice.Entry, error) {
return append([]onlyoffice.Entry(nil), f.entries[parent]...), nil
}
func (f *fakeFolderStore) CreateFolder(_ context.Context, parent, title string) (onlyoffice.Entry, error) {
f.nextID++
e := onlyoffice.Entry{
ID: fmt.Sprintf("id-%d", f.nextID),
ParentID: parent,
Title: title,
Kind: onlyoffice.Folder,
}
f.entries[parent] = append(f.entries[parent], e)
f.created = append(f.created, parent+"/"+title)
return e, nil
}
func (f *fakeFolderStore) Stat(context.Context, string) (onlyoffice.Entry, error) {
return onlyoffice.Entry{}, errors.New("not implemented")
}
func (f *fakeFolderStore) Upload(context.Context, string, string, io.Reader) (onlyoffice.Entry, error) {
return onlyoffice.Entry{}, errors.New("not implemented")
}
func (f *fakeFolderStore) Download(context.Context, string, io.Writer) (int64, error) {
return 0, errors.New("not implemented")
}
func (f *fakeFolderStore) Move(context.Context, []string, string) error {
return errors.New("not implemented")
}
func (f *fakeFolderStore) Copy(context.Context, []string, string) error {
return errors.New("not implemented")
}
func (f *fakeFolderStore) Rename(context.Context, string, string) error {
return errors.New("not implemented")
}
func (f *fakeFolderStore) Delete(context.Context, []string) error {
return errors.New("not implemented")
}
// fakeUploader is the *onlyoffice.Client slice dav upload depends on.
type fakeUploader struct {
conflict bool
uploads []string
replaced []string
}
func (f *fakeUploader) UploadToFolderReplacing(_ context.Context, folderID, localPath string) (*onlyoffice.FileEntry, []int, error) {
f.uploads = append(f.uploads, localPath)
ids := []int(nil)
if f.conflict {
ids = []int{7}
f.replaced = append(f.replaced, folderID+"#7")
}
return fakeFileEntry(99, localPath, folderID), ids, nil
}
func (f *fakeUploader) AssertNoFileConflict(_ context.Context, folderID, localPath string) error {
if f.conflict {
return fmt.Errorf("%w: conflict in folder %s for %s", onlyoffice.ErrFileExists, folderID, localPath)
}
return nil
}
func (f *fakeUploader) UploadToFolder(_ context.Context, folderID, localPath string) (*onlyoffice.FileEntry, error) {
f.uploads = append(f.uploads, localPath)
return fakeFileEntry(99, localPath, folderID), nil
}
func fakeFileEntry(id int, localPath, folderID string) *onlyoffice.FileEntry {
num := json.Number(fmt.Sprintf("%d", id))
title := localPath[strings.LastIndex(localPath, "/")+1:]
return &onlyoffice.FileEntry{
ID: &num,
Title: &title,
FolderID: &num,
}
}
func TestEnsurePathCreatesNestedFolders(t *testing.T) {
store := newFakeFolderStore()
ctx := context.Background()
got, err := ensurePath(ctx, store, "root", "Banks/Caixa")
if err != nil {
t.Fatalf("ensurePath: %v", err)
}
if got.Kind != onlyoffice.Folder || got.ID == "" {
t.Fatalf("ensurePath returned %+v, want a folder with an id", got)
}
if want := []string{"root/Banks", "id-1/Caixa"}; !equalStrings(store.created, want) {
t.Fatalf("created %v, want %v", store.created, want)
}
}
func TestEnsurePathIsIdempotent(t *testing.T) {
store := newFakeFolderStore()
ctx := context.Background()
first, err := ensurePath(ctx, store, "root", "Banks/Caixa")
if err != nil {
t.Fatalf("ensurePath first: %v", err)
}
created := len(store.created)
second, err := ensurePath(ctx, store, "root", "Banks/Caixa")
if err != nil {
t.Fatalf("ensurePath second: %v", err)
}
if second.ID != first.ID {
t.Fatalf("second run id = %q, want %q (no duplicate)", second.ID, first.ID)
}
if len(store.created) != created {
t.Fatalf("second run created folders: %v", store.created)
}
}
func TestEnsurePathReusesExistingFolder(t *testing.T) {
store := newFakeFolderStore()
store.entries["root"] = []onlyoffice.Entry{
{ID: "banks-id", Title: "Banks", Kind: onlyoffice.Folder},
}
store.entries["banks-id"] = []onlyoffice.Entry{
{ID: "caixa-id", Title: "Caixa", Kind: onlyoffice.Folder},
}
got, err := ensurePath(context.Background(), store, "root", "Banks/Caixa")
if err != nil {
t.Fatalf("ensurePath: %v", err)
}
if got.ID != "caixa-id" {
t.Fatalf("id = %q, want caixa-id", got.ID)
}
if len(store.created) != 0 {
t.Fatalf("created %v, want none", store.created)
}
}
func TestEnsurePathRejectsEmptyAndDotDot(t *testing.T) {
store := newFakeFolderStore()
for _, path := range []string{"", "/", "Banks/../Caixa"} {
if _, err := ensurePath(context.Background(), store, "root", path); err == nil {
t.Fatalf("ensurePath(%q) = nil error, want failure", path)
}
}
}
func TestUploadLocalReplacesByDefault(t *testing.T) {
up := &fakeUploader{conflict: true}
results, err := uploadLocal(context.Background(), up, "folder-1", []string{"a/f.xlsx"}, true)
if err != nil {
t.Fatalf("uploadLocal: %v", err)
}
if len(results) != 1 || results[0].Entry.Title != "f.xlsx" {
t.Fatalf("results = %+v", results)
}
if len(results[0].Replaced) != 1 || results[0].Replaced[0] != 7 {
t.Fatalf("replaced = %v, want [7]", results[0].Replaced)
}
if len(up.replaced) != 1 {
t.Fatalf("UploadToFolderReplacing not used for replace: %v", up.replaced)
}
}
func TestUploadLocalNoReplaceFailsOnConflict(t *testing.T) {
up := &fakeUploader{conflict: true}
_, err := uploadLocal(context.Background(), up, "folder-1", []string{"a/f.xlsx"}, false)
if !errors.Is(err, onlyoffice.ErrFileExists) {
t.Fatalf("err = %v, want ErrFileExists", err)
}
if len(up.uploads) != 0 {
t.Fatalf("uploaded despite conflict: %v", up.uploads)
}
}
func TestUploadLocalNoReplaceUploadsWhenFree(t *testing.T) {
up := &fakeUploader{}
results, err := uploadLocal(context.Background(), up, "folder-1", []string{"a/f.xlsx", "a/g.pdf"}, false)
if err != nil {
t.Fatalf("uploadLocal: %v", err)
}
if len(results) != 2 {
t.Fatalf("results = %+v", results)
}
if len(up.uploads) != 2 {
t.Fatalf("uploads = %v", up.uploads)
}
}
func TestDavRegistersUploadAndEnsurePath(t *testing.T) {
for _, name := range []string{"ensure-path", "upload"} {
cmd, _, err := rootCmd.Find([]string{"dav", name})
if err != nil {
t.Fatalf("dav %s not registered: %v", name, err)
}
if cmd.Name() != name {
t.Fatalf("resolved %q, want %q", cmd.Name(), name)
}
}
upload, _, err := rootCmd.Find([]string{"dav", "upload"})
if err != nil {
t.Fatal(err)
}
flag := upload.Flags().Lookup("replace")
if flag == nil || flag.DefValue != "true" {
t.Fatalf("upload --replace flag = %+v, want default true", flag)
}
ensure, _, err := rootCmd.Find([]string{"dav", "ensure-path"})
if err != nil {
t.Fatal(err)
}
if ensure.Flags().Lookup("under") == nil {
t.Fatal("ensure-path missing --under flag")
}
}
func equalStrings(a, b []string) bool {
if len(a) != len(b) {
return false
}
for i := range a {
if a[i] != b[i] {
return false
}
}
return true
}
+890
View File
@@ -0,0 +1,890 @@
package main
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"os"
"path/filepath"
"strconv"
"strings"
"time"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/eslider/go-onlyoffice/internal/docpipe"
"github.com/spf13/cobra"
)
func init() {
rootCmd.AddCommand(docsCmd())
}
func docsCmd() *cobra.Command {
cmd := &cobra.Command{
Use: "docs",
Short: "Local document pipeline: md↔docx, OCR→PDF, extract Markdown",
Long: `Agent-friendly conversions (requires pandoc / ocrmypdf / pdftotext on PATH).
OnlyOffice Documents UI is poor for .md/.txt — keep sources in git, store .docx in OO.
Upload Markdown as DOCX: oo docs put-md PROJECT_ID file.md
Upload plain text: oo docs put-txt PROJECT_ID file.txt (preserves line breaks)
Upload XLSX: oo docs put-xlsx PROJECT_ID FILE.xlsx
Read an OO file as MD: oo docs as-md FILE_ID
OCR a scan locally: oo docs ocr scan.pdf --md out.md
Structured OCR (hOCR→MD): oo docs hocr scan.jpg --md out.md --yaml out.yml`,
}
cmd.AddCommand(docsConvertCmd())
cmd.AddCommand(docsPDFCmd())
cmd.AddCommand(docsPresignedCmd())
cmd.AddCommand(docsCSVCmd())
cmd.AddCommand(docsJSONCmd())
cmd.AddCommand(docsOptimizeCmd())
cmd.AddCommand(docsOCRCmd())
cmd.AddCommand(docsHOCRCmd())
cmd.AddCommand(docsAsMDCmd())
cmd.AddCommand(docsPutMDCmd())
cmd.AddCommand(docsPutTxtCmd())
cmd.AddCommand(docsPutXlsxCmd())
cmd.AddCommand(docsToolsCmd())
return cmd
}
func docsToolsCmd() *cobra.Command {
return &cobra.Command{
Use: "tools",
Short: "Show which converter binaries are on PATH",
RunE: func(cmd *cobra.Command, args []string) error {
t := docpipe.LookPath()
printObject(map[string]any{
"pandoc": strOrNil(t.Pandoc),
"ocrmypdf": strOrNil(t.OCRMyPDF),
"pdftotext": strOrNil(t.PDFToText),
"tesseract": strOrNil(t.Tesseract),
"ghostscript": strOrNil(t.Ghostscript),
})
return nil
},
}
}
func docsPDFCmd() *cobra.Command {
var out, docsURL, secret, output, folder string
var stream, pipe bool
cmd := &cobra.Command{
Use: "pdf FILE_ID | PATH [ARG...]",
Short: "Convert files to PDF via the DocumentServer converter (OO file ids or local paths)",
Long: `Native OnlyOffice conversion (the engine behind the portal's "Download as PDF"):
1. GET /api/2.0/files/file/{id}/presigneduri → fetchable source URL
2. POST <docs>/converter with a JWT → converted file URL
3. download the result
Arguments may be OnlyOffice file ids OR local file paths. A local path is
uploaded to the scratch folder (--folder, default 2 = "My documents"), converted,
downloaded and then removed — so any local document yields a PDF on the fly.
Docs base defaults to $ONLYOFFICE_DOCS_URL, else $ONLYOFFICE_URL + "/ds-vpath"
(/ds-vpath is the usual reverse-proxy mount for the DocumentServer). The JWT secret is
$ONLYOFFICE_DS_SECRET (DocumentServer services.CoAuthoring.secret).`,
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
base := docsBaseURL(docsURL)
if base == "" {
return fmt.Errorf("docs base url unknown; set --docs-url or ONLYOFFICE_DOCS_URL")
}
sec := secret
if sec == "" {
sec = firstEnv("ONLYOFFICE_DS_SECRET", "OO_DS_SECRET")
}
if sec == "" {
return fmt.Errorf("JWT secret required: --secret or ONLYOFFICE_DS_SECRET")
}
ot := output
if ot == "" {
ot = "pdf"
}
toStdout := stream || pipe
if toStdout && len(args) > 1 {
return fmt.Errorf("--stream/--pipe writes one file to stdout; pass a single input")
}
for _, arg := range args {
id, local := arg, false
title := ""
if fi, statErr := os.Stat(arg); statErr == nil && !fi.IsDir() {
// Local file → temporary upload into the scratch folder.
ent, uerr := c.UploadToFolder(cmd.Context(), folder, arg)
if uerr != nil {
return fmt.Errorf("upload %s: %w", arg, uerr)
}
id = strconv.FormatInt(onlyoffice.FileEntryNumericID(ent), 10)
local = true
title = filepath.Base(arg)
} else {
if f, ferr := c.GetFile(cmd.Context(), id); ferr == nil && f != nil && f.Title != nil {
title = *f.Title
}
}
src, err := c.PresignedURI(cmd.Context(), id)
if err != nil {
return fmt.Errorf("presigneduri %s: %w", id, err)
}
res, err := c.ConvertDocument(cmd.Context(), base, sec, onlyoffice.ConvertRequest{
URL: src,
OutputType: ot,
FileType: strings.TrimPrefix(filepath.Ext(title), "."),
Title: title,
Key: fmt.Sprintf("oo-%s-%d", id, time.Now().UnixNano()),
})
if err != nil {
return fmt.Errorf("convert %s: %w", id, err)
}
var w io.Writer
dst := out
if toStdout {
w = os.Stdout
} else {
if dst == "" {
stem := strings.TrimSuffix(title, filepath.Ext(title))
if stem == "" {
stem = "file-" + id
}
dst = stem + "." + ot
}
fh, err := os.Create(dst)
if err != nil {
return err
}
w = fh
defer fh.Close()
}
n, derr := c.DownloadURLTo(cmd.Context(), res.FileURL, w)
if local {
// Best-effort cleanup of the temporary upload.
if nid, e := strconv.Atoi(id); e == nil {
_ = c.DeleteFiles(cmd.Context(), []int{nid})
}
}
if derr != nil {
return fmt.Errorf("download: %w", derr)
}
if toStdout {
// Keep stdout byte-clean for pipes; status goes to stderr.
fmt.Fprintf(os.Stderr, "converted %s -> stdout (%d bytes, %s)\n", arg, n, ot)
} else {
printObject(map[string]any{"source": arg, "fileid": id, "title": title, "output": dst, "bytes": n, "type": ot})
}
}
return nil
},
}
cmd.Flags().StringVar(&out, "out", "", "output path (default: ./<title>.<format>)")
cmd.Flags().StringVar(&output, "to", "pdf", "output format (pdf, docx, xlsx, …)")
cmd.Flags().StringVar(&docsURL, "docs-url", "", "DocumentServer base (default $ONLYOFFICE_DOCS_URL or $ONLYOFFICE_URL/ds-vpath)")
cmd.Flags().StringVar(&secret, "secret", "", "JWT secret (default $ONLYOFFICE_DS_SECRET)")
cmd.Flags().StringVar(&folder, "folder", "2", "scratch folder id for local-file uploads")
cmd.Flags().BoolVar(&stream, "stream", false, "write the converted bytes to stdout (pipe-friendly)")
cmd.Flags().BoolVar(&pipe, "pipe", false, "alias for --stream")
return cmd
}
func docsPresignedCmd() *cobra.Command {
return &cobra.Command{
Use: "presigned FILE_ID",
Short: "Print a short-lived fetchable URI for a portal file",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
u, err := c.PresignedURI(cmd.Context(), args[0])
if err != nil {
return err
}
printObject(map[string]any{"fileid": args[0], "uri": u})
return nil
},
}
}
// loadWorkbookBytes reads an argument that is either an OnlyOffice file id
// (downloaded via the client) or a local path.
func loadWorkbookBytes(cmd *cobra.Command, c *onlyoffice.Client, arg string) ([]byte, error) {
if fi, err := os.Stat(arg); err == nil && !fi.IsDir() {
return os.ReadFile(arg)
}
var buf bytes.Buffer
if _, err := c.DownloadFile(cmd.Context(), arg, &buf); err != nil {
return nil, err
}
return buf.Bytes(), nil
}
// parseDelimiter maps a flag value to a CSV delimiter rune ("," default).
func parseDelimiter(s string) rune {
switch s {
case "", ",":
return ','
case "\\t", "tab", "\t":
return '\t'
case ";":
return ';'
case "|":
return '|'
default:
r := []rune(s)
return r[0]
}
}
func docsCSVCmd() *cobra.Command {
var sheet int
var delim, out string
cmd := &cobra.Command{
Use: "csv SRC [SRC...]",
Short: "Export a worksheet (XLS/XLSX/ODS) to CSV (sheet-aware)",
Long: `SRC is an OnlyOffice file id or a local path. --sheet is 1-based (default 1 = first).
Why not the DocumentServer: its csv output covers only the first worksheet and
ignores a sheet selector (verified). Sheet selection and JSON use a local reader.`,
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
d := parseDelimiter(delim)
for _, arg := range args {
data, err := loadWorkbookBytes(cmd, c, arg)
if err != nil {
return fmt.Errorf("read %s: %w", arg, err)
}
text, err := onlyoffice.WorkbookSheetCSV(data, sheet-1, d)
if err != nil {
return fmt.Errorf("%s: %w", arg, err)
}
if out != "" {
if err := os.WriteFile(out, []byte(text), 0o644); err != nil {
return err
}
printObject(map[string]any{"source": arg, "sheet": sheet, "output": out, "bytes": len(text)})
continue
}
fmt.Print(text)
}
return nil
},
}
cmd.Flags().IntVar(&sheet, "sheet", 1, "worksheet number (1-based, default first)")
cmd.Flags().StringVar(&delim, "delimiter", ",", "CSV delimiter (',', ';', '|', 'tab')")
cmd.Flags().StringVar(&out, "out", "", "write to this file instead of stdout (single input)")
return cmd
}
func docsJSONCmd() *cobra.Command {
var sheet int
var out string
cmd := &cobra.Command{
Use: "json SRC [SRC...]",
Short: "Export a worksheet (XLS/XLSX/ODS) to JSON rows (first row = header)",
Long: `SRC is an OnlyOffice file id or a local path. --sheet is 1-based (default 1 = first).
Each data row becomes an object keyed by the header cells of that sheet.`,
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, arg := range args {
data, err := loadWorkbookBytes(cmd, c, arg)
if err != nil {
return fmt.Errorf("read %s: %w", arg, err)
}
rows, err := onlyoffice.WorkbookSheetJSON(data, sheet-1)
if err != nil {
return fmt.Errorf("%s: %w", arg, err)
}
b, err := json.MarshalIndent(rows, "", " ")
if err != nil {
return err
}
b = append(b, '\n')
if out != "" {
if err := os.WriteFile(out, b, 0o644); err != nil {
return err
}
printObject(map[string]any{"source": arg, "sheet": sheet, "rows": len(rows), "output": out})
continue
}
os.Stdout.Write(b)
}
return nil
},
}
cmd.Flags().IntVar(&sheet, "sheet", 1, "worksheet number (1-based, default first)")
cmd.Flags().StringVar(&out, "out", "", "write to this file instead of stdout (single input)")
return cmd
}
// docsBaseURL resolves the DocumentServer base: --docs-url, $ONLYOFFICE_DOCS_URL,
// else the standard /ds-vpath reverse-proxy mount.
func docsBaseURL(flag string) string {
if flag != "" {
return flag
}
if v := os.Getenv("ONLYOFFICE_DOCS_URL"); v != "" {
return v
}
if v := firstEnv("ONLYOFFICE_URL", "ONLYOFFICE_HOST", "OO_URL"); v != "" {
return strings.TrimRight(v, "/") + "/ds-vpath"
}
return ""
}
func firstEnv(keys ...string) string {
for _, k := range keys {
if v := os.Getenv(k); v != "" {
return v
}
}
return ""
}
func strOrNil(s string) any {
if s == "" {
return nil
}
return s
}
func docsConvertCmd() *cobra.Command {
var to string
cmd := &cobra.Command{
Use: "convert PATH",
Short: "Convert a local file with pandoc (md↔docx by default)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
in := args[0]
t := docpipe.LookPath()
out := to
if out == "" {
switch docpipe.Ext(in) {
case ".md", ".markdown":
out = docpipe.SiblingDOCX(in)
case ".docx":
out = strings.TrimSuffix(in, docpipe.Ext(in)) + ".md"
default:
return fmt.Errorf("--to required for input type %s", docpipe.Ext(in))
}
}
if err := docpipe.EnsureDir(out); err != nil {
return err
}
if err := t.ConvertFile(in, out); err != nil {
return err
}
printObject(map[string]any{"in": in, "out": out})
return nil
},
}
cmd.Flags().StringVar(&to, "to", "", "output path (default: sibling .docx or .md)")
return cmd
}
func docsOptimizeCmd() *cobra.Command {
var out string
cmd := &cobra.Command{
Use: "optimize PDF_PATH",
Short: "Rewrite PDF via Ghostscript (PostScript pdfwrite) for OO-friendly size/text",
Long: `Use for InDesign/iText PDFs with a good text layer — avoids ocrmypdf invisible
text overlays that break OnlyOffice DocEditor. Skips OCR; rewrites via gs pdfwrite.`,
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
in := args[0]
t := docpipe.LookPath()
if out == "" {
base := strings.TrimSuffix(filepath.Base(in), filepath.Ext(in))
out = filepath.Join(filepath.Dir(in), base+".optimized.pdf")
}
if err := docpipe.EnsureDir(out); err != nil {
return err
}
chars, _ := t.PDFTextLayerChars(in)
if err := t.OptimizePDF(in, out); err != nil {
return err
}
outChars, _ := t.PDFTextLayerChars(out)
printObject(map[string]any{
"in": in,
"pdf": out,
"text_chars_in": chars,
"text_chars_out": outChars,
"note": "native text layer preserved; no OCR overlay",
})
return nil
},
}
cmd.Flags().StringVar(&out, "out", "", "output PDF (default: <name>.optimized.pdf)")
return cmd
}
func docsOCRCmd() *cobra.Command {
var out, mdOut, lang string
var force bool
var writeMD bool
cmd := &cobra.Command{
Use: "ocr PATH",
Short: "OCR image/PDF → searchable PDF (and optional Markdown)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
in := args[0]
t := docpipe.LookPath()
if out == "" {
base := strings.TrimSuffix(filepath.Base(in), filepath.Ext(in))
out = filepath.Join(filepath.Dir(in), base+".ocr.pdf")
}
if err := docpipe.EnsureDir(out); err != nil {
return err
}
if err := t.OCRToPDF(in, out, force, lang); err != nil {
return err
}
res := map[string]any{"in": in, "pdf": out}
if writeMD || mdOut != "" {
if mdOut == "" {
mdOut = strings.TrimSuffix(out, filepath.Ext(out)) + ".md"
}
text, err := t.ExtractPDFText(out)
if err != nil {
return err
}
body := "# " + filepath.Base(in) + "\n\n" + strings.TrimSpace(text) + "\n"
if err := os.WriteFile(mdOut, []byte(body), 0o644); err != nil {
return err
}
res["md"] = mdOut
}
printObject(res)
return nil
},
}
cmd.Flags().StringVar(&out, "out", "", "output searchable PDF (default: <name>.ocr.pdf)")
cmd.Flags().StringVar(&mdOut, "md", "", "write Markdown extraction to this path")
cmd.Flags().BoolVar(&writeMD, "markdown", false, "also write sibling .md next to OCR PDF")
cmd.Flags().StringVar(&lang, "lang", "eng", "OCR language(s) for tesseract/ocrmypdf")
cmd.Flags().BoolVar(&force, "force", false, "force OCR even if a text layer exists (default: skip when pdftotext finds enough text)")
return cmd
}
func docsHOCRCmd() *cobra.Command {
var mdOut, yamlOut, hocrOut, lang string
var dpi int
var minConf float32
cmd := &cobra.Command{
Use: "hocr PATH",
Short: "Tesseract hOCR → structured Markdown/YAML (via go-hocr)",
Long: `Runs tesseract with hOCR output, parses with go-hocr, writes Markdown
(and optional YAML). Better reading order / confidence than plain pdftotext.
For Spanish scans: --lang spa or spa+eng (needs tesseract-ocr-spa / TESSDATA_PREFIX).
Phone photos: --dpi 200..300.`,
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
in := args[0]
t := docpipe.LookPath()
dir, err := os.MkdirTemp("", "oo-docs-hocr-*")
if err != nil {
return err
}
defer os.RemoveAll(dir)
res, err := t.ToHOCRMarkdown(in, dir, lang, dpi, minConf)
if err != nil {
return err
}
if mdOut == "" {
base := strings.TrimSuffix(filepath.Base(in), filepath.Ext(in))
mdOut = filepath.Join(filepath.Dir(in), base+".hocr.md")
}
if err := docpipe.EnsureDir(mdOut); err != nil {
return err
}
if err := os.WriteFile(mdOut, []byte(res.Markdown), 0o644); err != nil {
return err
}
obj := map[string]any{
"in": in,
"md": mdOut,
"did_ocr": res.DidOCR,
}
if hocrOut != "" {
if err := docpipe.EnsureDir(hocrOut); err != nil {
return err
}
b, err := os.ReadFile(res.HOCRPath)
if err != nil {
return err
}
if err := os.WriteFile(hocrOut, b, 0o644); err != nil {
return err
}
obj["hocr"] = hocrOut
} else {
obj["hocr_tmp"] = res.HOCRPath
}
if yamlOut != "" {
if err := docpipe.EnsureDir(yamlOut); err != nil {
return err
}
if err := os.WriteFile(yamlOut, []byte(res.YAML), 0o644); err != nil {
return err
}
obj["yaml"] = yamlOut
}
printObject(obj)
return nil
},
}
cmd.Flags().StringVar(&mdOut, "md", "", "Markdown output (default: <name>.hocr.md)")
cmd.Flags().StringVar(&yamlOut, "yaml", "", "also write structured YAML from go-hocr")
cmd.Flags().StringVar(&hocrOut, "hocr", "", "also keep raw .hocr file at this path")
cmd.Flags().StringVar(&lang, "lang", "eng", "tesseract language(s), e.g. spa+eng")
cmd.Flags().IntVar(&dpi, "dpi", 220, "hint DPI for phone photos / scans (0 = tesseract default)")
cmd.Flags().Float32Var(&minConf, "min-conf", 0, "drop words with OCR confidence below this (0 = keep all)")
return cmd
}
func docsAsMDCmd() *cobra.Command {
var to, lang string
var minChars int
var uploadOCR bool
var useHOCR bool
var dpi int
var minConf float32
cmd := &cobra.Command{
Use: "as-md FILE_ID",
Short: "Download an OO Documents file and emit Markdown (OCR PDF/image if needed)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
ctx := cmd.Context()
meta, err := c.GetFile(ctx, args[0])
if err != nil {
return err
}
title := onlyoffice.FileEntryTitle(meta)
dir, err := os.MkdirTemp("", "oo-docs-as-md-*")
if err != nil {
return err
}
defer os.RemoveAll(dir)
local := filepath.Join(dir, onlyoffice.SafeLocalFileName(title))
f, err := os.Create(local)
if err != nil {
return err
}
if _, err := c.DownloadFile(ctx, args[0], f); err != nil {
_ = f.Close()
return err
}
_ = f.Close()
tools := docpipe.LookPath()
var res docpipe.Result
if useHOCR {
hr, err := tools.ToHOCRMarkdown(local, dir, lang, dpi, minConf)
if err != nil {
return err
}
res = docpipe.Result{Markdown: hr.Markdown, DidOCR: hr.DidOCR, Source: hr.Source}
} else {
res, err = tools.ToMarkdown(local, dir, lang, minChars)
if err != nil {
return err
}
}
outPath := to
if outPath == "" {
base := strings.TrimSuffix(onlyoffice.SafeLocalFileName(title), filepath.Ext(onlyoffice.SafeLocalFileName(title)))
if base == "" || base == "download" {
base = "file-" + args[0]
}
outPath = base + ".md"
}
if err := docpipe.EnsureDir(outPath); err != nil {
return err
}
if err := os.WriteFile(outPath, []byte(res.Markdown), 0o644); err != nil {
return err
}
obj := map[string]any{
"file_id": args[0],
"title": title,
"md": outPath,
"did_ocr": res.DidOCR,
}
if res.OCRPDFPath != "" && uploadOCR {
folderID := onlyoffice.FileFolderID(meta)
upName := strings.TrimSuffix(onlyoffice.SafeLocalFileName(title), filepath.Ext(onlyoffice.SafeLocalFileName(title))) + ".ocr.pdf"
tmpUp := filepath.Join(dir, upName)
data, err := os.ReadFile(res.OCRPDFPath)
if err != nil {
return err
}
if err := os.WriteFile(tmpUp, data, 0o644); err != nil {
return err
}
if folderID == "" {
obj["ocr_pdf_local"] = res.OCRPDFPath
obj["note"] = "file has no folderId; OCR PDF left local — pass after moving into a folder"
} else {
ent, err := c.UploadToFolder(ctx, folderID, tmpUp)
if err != nil {
return err
}
obj["ocr_pdf_file_id"] = fileIDStr(ent)
obj["ocr_pdf_title"] = onlyoffice.FileEntryTitle(ent)
}
} else if res.OCRPDFPath != "" {
// Keep OCR PDF outside temp by copying beside md if requested via env-less default:
kept := strings.TrimSuffix(outPath, filepath.Ext(outPath)) + ".ocr.pdf"
if b, err := os.ReadFile(res.OCRPDFPath); err == nil {
_ = os.WriteFile(kept, b, 0o644)
obj["ocr_pdf_local"] = kept
} else {
obj["ocr_pdf_local"] = res.OCRPDFPath
}
}
printObject(obj)
return nil
},
}
cmd.Flags().StringVar(&to, "to", "", "write Markdown to this path (default: ./<title>.md)")
cmd.Flags().StringVar(&lang, "lang", "eng", "OCR language")
cmd.Flags().IntVar(&minChars, "min-chars", docpipe.DefaultMinTextChars, "OCR PDF if text layer shorter than this")
cmd.Flags().BoolVar(&uploadOCR, "upload-ocr", false, "upload searchable OCR PDF back into the same OO folder")
cmd.Flags().BoolVar(&useHOCR, "hocr", false, "use tesseract hOCR + go-hocr instead of ocrmypdf/pdftotext")
cmd.Flags().IntVar(&dpi, "dpi", 220, "DPI hint when --hocr (phone photos)")
cmd.Flags().Float32Var(&minConf, "min-conf", 0, "drop low-confidence words when --hocr")
return cmd
}
func docsPutMDCmd() *cobra.Command {
var folderID string
var keepLocalDOCX string
var replace bool
cmd := &cobra.Command{
Use: "put-md PROJECT_ID MARKDOWN_PATH",
Short: "Convert Markdown→DOCX and upload DOCX into a project (OO-friendly)",
Long: `Agents edit .md locally; this uploads .docx so OnlyOffice can open/version it.`,
Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid, mdPath := args[0], args[1]
c, err := newOO(cmd)
if err != nil {
return err
}
tools := docpipe.LookPath()
dir, err := os.MkdirTemp("", "oo-docs-put-md-*")
if err != nil {
return err
}
defer os.RemoveAll(dir)
docxName := strings.TrimSuffix(filepath.Base(mdPath), filepath.Ext(mdPath)) + ".docx"
docxPath := filepath.Join(dir, docxName)
if err := tools.MDToDOCX(mdPath, docxPath); err != nil {
return err
}
if keepLocalDOCX != "" {
if err := docpipe.EnsureDir(keepLocalDOCX); err != nil {
return err
}
b, err := os.ReadFile(docxPath)
if err != nil {
return err
}
if err := os.WriteFile(keepLocalDOCX, b, 0o644); err != nil {
return err
}
}
ctx := cmd.Context()
ent, deleted, err := uploadProjectDoc(ctx, c, pid, docxPath, folderID, replace)
if err != nil {
return err
}
obj := map[string]any{
"project_id": pid,
"md": mdPath,
"uploaded": fileEntryToMap(ent),
}
if folderID != "" {
obj["folder_id"] = folderID
}
if len(deleted) > 0 {
obj["replaced_file_ids"] = deleted
}
printObject(obj)
return nil
},
}
cmd.Flags().StringVar(&folderID, "folder", "", "Documents folder id (default: project root)")
cmd.Flags().StringVar(&keepLocalDOCX, "keep-docx", "", "also write the generated DOCX to this local path")
cmd.Flags().BoolVar(&replace, "replace", true, "replace same stem|ext before upload (default); false = fail if name taken")
return cmd
}
func docsPutTxtCmd() *cobra.Command {
var folderID string
var keepLocalDOCX string
var replace bool
cmd := &cobra.Command{
Use: "put-txt PROJECT_ID TEXT_PATH",
Short: "Convert plain text→DOCX (preserve line breaks) and upload into a project",
Long: `OnlyOffice cannot render .txt well. This keeps each source line on its own DOCX line.`,
Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid, txtPath := args[0], args[1]
c, err := newOO(cmd)
if err != nil {
return err
}
tools := docpipe.LookPath()
dir, err := os.MkdirTemp("", "oo-docs-put-txt-*")
if err != nil {
return err
}
defer os.RemoveAll(dir)
docxName := strings.TrimSuffix(filepath.Base(txtPath), filepath.Ext(txtPath)) + ".docx"
docxPath := filepath.Join(dir, docxName)
if err := tools.TXTToDOCX(txtPath, docxPath); err != nil {
return err
}
if keepLocalDOCX != "" {
if err := docpipe.EnsureDir(keepLocalDOCX); err != nil {
return err
}
b, err := os.ReadFile(docxPath)
if err != nil {
return err
}
if err := os.WriteFile(keepLocalDOCX, b, 0o644); err != nil {
return err
}
}
ctx := cmd.Context()
ent, deleted, err := uploadProjectDoc(ctx, c, pid, docxPath, folderID, replace)
if err != nil {
return err
}
obj := map[string]any{
"project_id": pid,
"txt": txtPath,
"uploaded": fileEntryToMap(ent),
}
if folderID != "" {
obj["folder_id"] = folderID
}
if len(deleted) > 0 {
obj["replaced_file_ids"] = deleted
}
printObject(obj)
return nil
},
}
cmd.Flags().StringVar(&folderID, "folder", "", "Documents folder id (default: project root)")
cmd.Flags().StringVar(&keepLocalDOCX, "keep-docx", "", "also write the generated DOCX to this local path")
cmd.Flags().BoolVar(&replace, "replace", true, "replace same stem|ext before upload (default); false = fail if name taken")
return cmd
}
func docsPutXlsxCmd() *cobra.Command {
var folderID, keepLocal string
var replace bool
cmd := &cobra.Command{
Use: "put-xlsx PROJECT_ID LOCAL_XLSX",
Short: "Upload an XLSX workbook into a project (upsert by stem|ext)",
Long: `Spreadsheets live in OnlyOffice — not in git. Uploads an existing .xlsx into the
project Documents (or --folder), replacing the same stem|ext by default.
oo docs put-xlsx 218 ./my.xlsx`,
Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid := args[0]
c, err := newOO(cmd)
if err != nil {
return err
}
xlsxPath := args[1]
if docpipe.Ext(xlsxPath) != ".xlsx" {
return fmt.Errorf("expected .xlsx, got %s", docpipe.Ext(xlsxPath))
}
srcLabel := xlsxPath
if keepLocal != "" {
b, err := os.ReadFile(xlsxPath)
if err != nil {
return err
}
if err := docpipe.EnsureDir(keepLocal); err != nil {
return err
}
if err := os.WriteFile(keepLocal, b, 0o644); err != nil {
return err
}
}
ctx := cmd.Context()
ent, deleted, err := uploadProjectDoc(ctx, c, pid, xlsxPath, folderID, replace)
if err != nil {
return err
}
obj := map[string]any{
"project_id": pid,
"source": srcLabel,
"uploaded": fileEntryToMap(ent),
}
if folderID != "" {
obj["folder_id"] = folderID
}
if len(deleted) > 0 {
obj["replaced_file_ids"] = deleted
}
printObject(obj)
return nil
},
}
cmd.Flags().StringVar(&folderID, "folder", "", "Documents folder id (default: project root)")
cmd.Flags().StringVar(&keepLocal, "keep-xlsx", "", "also copy the uploaded bytes to this local path")
cmd.Flags().BoolVar(&replace, "replace", true, "replace same stem|ext before upload (default); false = fail if name taken")
return cmd
}
// uploadProjectDoc upserts (--replace, default) or no-clobbers into project/folder Documents.
func uploadProjectDoc(ctx context.Context, c *onlyoffice.Client, pid, localPath, folderID string, replace bool) (*onlyoffice.FileEntry, []int, error) {
if folderID != "" {
if replace {
return c.UploadToFolderReplacing(ctx, folderID, localPath)
}
if err := c.AssertNoFileConflict(ctx, folderID, localPath); err != nil {
return nil, nil, err
}
ent, err := c.UploadToFolder(ctx, folderID, localPath)
return ent, nil, err
}
if replace {
return c.UploadProjectFileReplacing(ctx, pid, localPath)
}
ent, err := c.UploadProjectFileNoClobber(ctx, pid, localPath)
return ent, nil, err
}
+158
View File
@@ -0,0 +1,158 @@
package main
import (
"strings"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra"
)
func init() {
rootCmd.AddCommand(indexCmd())
}
// indexFlags are shared by the `oo index folder` and `oo index files` verbs.
type indexFlags struct {
recursive bool
exts string
limit int
workers int
lang string
minChars int
workDir string
backend string
dryRun bool
asJSON bool
}
// indexCmd populates the own full-text index (oo_docs_text) that makes PDF and
// scanned content searchable. The OnlyOffice index is left untouched.
func indexCmd() *cobra.Command {
f := &indexFlags{}
cmd := &cobra.Command{
Use: "index",
Short: "Populate the own full-text index for PDF/scan content",
Long: "Index document text that the OnlyOffice Elasticsearch index does not\n" +
"cover (PDFs and scans) into a separate index (ONLYOFFICE_ES_TEXT_INDEX,\n" +
"default oo_docs_text). Text is extracted with docpipe (pdftotext, OCR)\n" +
"and the OnlyOffice server is never modified.\n\n" +
"Requires ONLYOFFICE_URL/USER/PASS (to download files) and ONLYOFFICE_ES_URL\n" +
"(to write the index). See docs/elasticsearch.md.",
}
cmd.PersistentFlags().BoolVar(&f.recursive, "recursive", false, "folder: descend into subfolders")
cmd.PersistentFlags().StringVar(&f.exts, "exts", "pdf", "comma-separated extensions to index")
cmd.PersistentFlags().IntVar(&f.limit, "limit", 0, "maximum number of files to index (0 = all)")
cmd.PersistentFlags().IntVar(&f.workers, "workers", 3, "parallel downloads/extractions")
cmd.PersistentFlags().StringVar(&f.lang, "lang", "deu+eng", "OCR language(s)")
cmd.PersistentFlags().IntVar(&f.minChars, "min-chars", 0, "text-layer threshold below which OCR runs")
cmd.PersistentFlags().StringVar(&f.workDir, "work-dir", "", "temp dir for downloads (default: system temp)")
cmd.PersistentFlags().StringVar(&f.backend, "backend", "rest", "file backend: rest|dav")
cmd.PersistentFlags().BoolVar(&f.dryRun, "dry-run", false, "list what would be indexed, without changes")
cmd.PersistentFlags().BoolVar(&f.asJSON, "json", false, "shorthand for --output json")
cmd.AddCommand(indexFolderCmd(f), indexFilesCmd(f))
return cmd
}
func indexFolderCmd(f *indexFlags) *cobra.Command {
return &cobra.Command{
Use: "folder FOLDER_ID",
Short: "Index every matching file in a Documents folder",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
return runIndex(cmd, f, args[0], nil)
},
}
}
func indexFilesCmd(f *indexFlags) *cobra.Command {
return &cobra.Command{
Use: "files FILE_ID...",
Short: "Index specific Documents files",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
return runIndex(cmd, f, "", args)
},
}
}
func runIndex(cmd *cobra.Command, f *indexFlags, folderID string, ids []string) error {
if f.asJSON {
outputFormat = "json"
}
c, err := newOO(cmd)
if err != nil {
return err
}
idx, err := onlyoffice.NewESTextIndex(onlyoffice.ESTextConfigFromEnv())
if err != nil {
return err
}
ti := onlyoffice.NewTextIndexer(c.FileStore(f.backend), idx)
ti.WorkDir = f.workDir
opts := onlyoffice.IndexOptions{
Recursive: f.recursive,
Extensions: splitList(f.exts),
Limit: f.limit,
Lang: f.lang,
MinChars: f.minChars,
Workers: f.workers,
}
ctx := cmd.Context()
if f.dryRun {
var entries []onlyoffice.Entry
if folderID != "" {
entries, err = ti.PlanFolder(ctx, folderID, opts)
} else {
entries, err = ti.PlanFiles(ctx, ids, opts)
}
if err != nil {
return err
}
rows := make([]map[string]any, 0, len(entries))
for _, e := range entries {
rows = append(rows, map[string]any{
"id": e.ID,
"title": e.Title,
"folder": e.ParentID,
})
}
printTable([]string{"id", "title", "folder"}, rows)
return nil
}
if err := ti.Ensure(ctx); err != nil {
return err
}
var res onlyoffice.IndexResult
if folderID != "" {
res, err = ti.IndexFolder(ctx, folderID, opts)
} else {
res, err = ti.IndexFiles(ctx, ids, opts)
}
if err != nil {
return err
}
printObject(map[string]any{
"index": idx.Index(),
"scanned": res.Scanned,
"indexed": res.Indexed,
"skipped": res.Skipped,
"failed": res.Failed,
"errors": res.Errors,
})
return nil
}
// splitList parses a comma-separated flag value, dropping blanks.
func splitList(s string) []string {
parts := strings.Split(s, ",")
out := make([]string, 0, len(parts))
for _, p := range parts {
if p = strings.TrimSpace(p); p != "" {
out = append(out, p)
}
}
return out
}
+55
View File
@@ -0,0 +1,55 @@
package main
import (
"reflect"
"strings"
"testing"
)
func TestIndexCommandRegistered(t *testing.T) {
cmd, _, err := rootCmd.Find([]string{"index"})
if err != nil {
t.Fatal(err)
}
if cmd.Name() != "index" {
t.Fatalf("index resolved to %q", cmd.Name())
}
for _, name := range []string{"exts", "limit", "workers", "lang", "min-chars", "work-dir", "backend", "dry-run", "recursive", "json"} {
if cmd.PersistentFlags().Lookup(name) == nil {
t.Errorf("index: missing --%s flag", name)
}
}
if cmd.PersistentFlags().Lookup("exts").DefValue != "pdf" {
t.Errorf("--exts default = %q, want pdf", cmd.PersistentFlags().Lookup("exts").DefValue)
}
for _, verb := range []string{"index folder", "index files"} {
sub, _, err := rootCmd.Find(strings.Fields(verb))
if err != nil {
t.Fatalf("%s: %v", verb, err)
}
if sub.Name() != strings.Fields(verb)[1] {
t.Errorf("%s resolved to %q", verb, sub.Name())
}
}
}
func TestIndexSearchBackendFlag(t *testing.T) {
cmd, _, err := rootCmd.Find([]string{"search"})
if err != nil {
t.Fatal(err)
}
if cmd.Flags().Lookup("backend") == nil {
t.Fatal("search: missing --backend flag")
}
if cmd.Flags().Lookup("backend").DefValue != "oo" {
t.Errorf("--backend default = %q, want oo", cmd.Flags().Lookup("backend").DefValue)
}
}
func TestSplitList(t *testing.T) {
got := splitList(" pdf , .PDF, docx ,, ")
want := []string{"pdf", ".PDF", "docx"}
if !reflect.DeepEqual(got, want) {
t.Errorf("splitList = %v, want %v", got, want)
}
}
+4 -4
View File
@@ -91,7 +91,7 @@ func invoiceCreateCmd() *cobra.Command {
Long: `Create a CRM invoice (Draft) with a single line. Long: `Create a CRM invoice (Draft) with a single line.
Always pass --opportunity when a deal exists (entity link at create). Updating Always pass --opportunity when a deal exists (entity link at create). Updating
--opportunity later often fails with HTTP 400 — see docs/crm-associations.md. --opportunity later often fails with HTTP 400 — see the CRM association rules.
Example: Example:
oo invoices create --number INV-2026-01 --contact CONTACT_ID --item ITEM_ID \ oo invoices create --number INV-2026-01 --contact CONTACT_ID --item ITEM_ID \
@@ -108,7 +108,7 @@ Example:
issueDate = issueDate + "T00:00:00.0000000+01:00" issueDate = issueDate + "T00:00:00.0000000+01:00"
} }
if dueDate == "" { if dueDate == "" {
dueDate = time.Now().Add(14 * 24 * time.Hour).Format("2006-01-02") + "T00:00:00.0000000+01:00" dueDate = time.Now().Add(14*24*time.Hour).Format("2006-01-02") + "T00:00:00.0000000+01:00"
} else if !strings.Contains(dueDate, "T") { } else if !strings.Contains(dueDate, "T") {
dueDate = dueDate + "T00:00:00.0000000+01:00" dueDate = dueDate + "T00:00:00.0000000+01:00"
} }
@@ -175,7 +175,7 @@ func invoiceUpdateCmd() *cobra.Command {
Long: `Update Draft invoice fields. Long: `Update Draft invoice fields.
--opportunity often returns HTTP 400 on existing invoices. Prefer --opportunity often returns HTTP 400 on existing invoices. Prefer
oo invoices create … --opportunity, or delete+recreate. See docs/crm-associations.md. oo invoices create … --opportunity, or delete+recreate. See the CRM association rules.
`, `,
Args: cobra.ExactArgs(1), Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
@@ -274,7 +274,7 @@ func invoiceStatusCmd() *cobra.Command {
Long: `PUT /api/2.0/crm/invoice/status/{id}. Long: `PUT /api/2.0/crm/invoice/status/{id}.
Billed invoices are not content-editable. Billed→Draft often does not work — Billed invoices are not content-editable. Billed→Draft often does not work —
recreate as Draft instead (docs/crm-associations.md). recreate as Draft instead (the CRM association rules).
`, `,
Args: cobra.MinimumNArgs(1), Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
+45
View File
@@ -0,0 +1,45 @@
package main
import (
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra"
)
func init() {
rootCmd.AddCommand(linkCmd())
}
// linkCmd prints deep links for file ids. Used to build third-party document
// packs whose cover embeds links to contracts and supporting files. File ids
// come from `oo projects files list` / `oo dav ls`; `oo projects files
// replace-in` keeps them clean when a document is re-uploaded.
func linkCmd() *cobra.Command {
return &cobra.Command{
Use: "link FILE_ID [FILE_ID...]",
Short: "Print OnlyOffice DocEditor deep links (Products/Files/DocEditor.aspx?fileid=…)",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, id := range args {
printObject(map[string]any{
"fileid": id,
"title": fileTitle(cmd, c, id),
"url": c.FileEditorURL(id),
})
}
return nil
},
}
}
// fileTitle best-effort resolves a file title; never fails the command.
func fileTitle(cmd *cobra.Command, c *onlyoffice.Client, id string) string {
f, err := c.GetFile(cmd.Context(), id)
if err != nil || f == nil || f.Title == nil {
return ""
}
return *f.Title
}
+14 -8
View File
@@ -3,20 +3,26 @@
// Command tree is subject-based (mirrors the library split and the `tea` CLI): // Command tree is subject-based (mirrors the library split and the `tea` CLI):
// //
// oo calendar list | events | add | delete // oo calendar list | events | add | delete
// oo projects list | get | milestones | create | update | delete | files (list|upload|download|rename|delete) // oo projects list | get | milestones | milestone-create | milestone-delete | board-sync | create | update | delete | contacts (add|remove) | team (list|add|remove|set) | link-authors | link-git | files (list|upload|replace-in|update|download|rename|delete|dedupe|as-md|put-md|put-txt|put-xlsx)
// oo tasks list | get | create | update | delete | subtask add | files (list|upload|detach) // oo tasks list | get | create | update | delete | subtask add | files (list|upload|detach)
// oo users list | self (alias: oo whoami) // oo users list | self | get | create | update | delete | block | unblock | password | check (alias: oo whoami)
// oo contacts list | get | delete | info-add | merge | dedupe-info // oo link FILE_ID [FILE_ID...] DocEditor deep links (Products/Files/DocEditor.aspx?fileid=…)
// oo contacts list | get | delete | info-add | merge | dedupe-info | tags | tag-add | tag-create | tag-remove
// oo persons list | create | delete | dedupe // oo persons list | create | delete | dedupe
// oo companies list | create | delete | dedupe | dedupe-persons // oo companies list | create | delete | dedupe | dedupe-persons
// oo opportunities list | get | create | delete | stages | member-add | dedupe | dedupe-members | fix-titles // oo opportunities list | get | create | update | delete | stages | member-add | dedupe | dedupe-members | fix-titles
// oo cases list | create | delete | member-add // oo cases list | create | delete | member-add
// oo crm-tasks list | create | delete | categories // oo crm-tasks list | create | delete | categories | reassign-self
// oo crm cleanup // oo crm audit | cleanup
// oo mails accounts | folders | list | get | download-attachment | draft | attach | draft-invoice | delete // oo mails accounts | folders | list | get | download-attachment | draft | attach | draft-invoice | send | delete
// oo invoices list | get | create | update | pdf | pdf-cleanup | status | delete | items … // oo invoices list | get | create | update | pdf | pdf-cleanup | status | delete | items …
// oo docs tools | convert | pdf | presigned | csv | json | optimize | ocr | hocr | as-md | put-md | put-txt | put-xlsx
// oo catalog match | merge | apply | scan-contacts | scan-projects | scan-thunderbird
// oo dav ls | move | copy | mkdir | ensure-path | upload | rename-file | rename-folder | download | fileops
// oo search QUERY [--content] [--folder ID] [--limit N] [--backend oo|own] [--json]
// oo index folder FOLDER_ID | files FILE_ID... [--recursive] [--exts pdf] [--dry-run]
// //
// CRM association rules: docs/crm-associations.md // CRM association rules live with the private oo-workspace tooling.
// //
// Every list supports `--output/-o json|table` (table is the default). // Every list supports `--output/-o json|table` (table is the default).
// //
+158
View File
@@ -24,10 +24,142 @@ func init() {
projectsCmd.AddCommand(prjGetCmd()) projectsCmd.AddCommand(prjGetCmd())
projectsCmd.AddCommand(prjMilestonesCmd()) projectsCmd.AddCommand(prjMilestonesCmd())
projectsCmd.AddCommand(prjMilestoneCreateCmd()) projectsCmd.AddCommand(prjMilestoneCreateCmd())
projectsCmd.AddCommand(prjMilestoneDeleteCmd())
projectsCmd.AddCommand(prjCreateCmd()) projectsCmd.AddCommand(prjCreateCmd())
projectsCmd.AddCommand(prjUpdateCmd()) projectsCmd.AddCommand(prjUpdateCmd())
projectsCmd.AddCommand(prjDeleteCmd()) projectsCmd.AddCommand(prjDeleteCmd())
projectsCmd.AddCommand(prjContactsCmd()) projectsCmd.AddCommand(prjContactsCmd())
projectsCmd.AddCommand(prjTeamCmd())
}
func prjTeamCmd() *cobra.Command {
cmd := &cobra.Command{
Use: "team",
Short: "Project team (portal users) — CRUD",
Long: `Project team members are portal users (People), not CRM contacts.
CRM companies/persons linked to a project live under 'oo projects contacts'.`,
}
cmd.AddCommand(prjTeamListCmd())
cmd.AddCommand(prjTeamAddCmd())
cmd.AddCommand(prjTeamRemoveCmd())
cmd.AddCommand(prjTeamSetCmd())
return cmd
}
func prjTeamListCmd() *cobra.Command {
return &cobra.Command{
Use: "list PROJECT_ID",
Short: "List portal users on the project team",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
pid, err := strconv.Atoi(args[0])
if err != nil {
return fmt.Errorf("project id: %w", err)
}
c, err := newOO(cmd)
if err != nil {
return err
}
list, err := c.ListProjectTeam(cmd.Context(), pid)
if err != nil {
return err
}
printTable([]string{"id", "displayName", "userName", "email", "isAdmin"}, teamRows(list))
return nil
},
}
}
func prjTeamAddCmd() *cobra.Command {
return &cobra.Command{
Use: "add PROJECT_ID USER_ID [USER_ID...]",
Short: "Add portal user(s) to the project team",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid, err := strconv.Atoi(args[0])
if err != nil {
return fmt.Errorf("project id: %w", err)
}
c, err := newOO(cmd)
if err != nil {
return err
}
for _, uid := range args[1:] {
if _, err := c.AddProjectTeamUser(cmd.Context(), pid, uid); err != nil {
return fmt.Errorf("add %s: %w", uid, err)
}
printObject(map[string]any{"project_id": pid, "user_id": uid, "added": true})
}
return nil
},
}
}
func prjTeamRemoveCmd() *cobra.Command {
return &cobra.Command{
Use: "remove PROJECT_ID USER_ID [USER_ID...]",
Short: "Remove portal user(s) from the project team",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid, err := strconv.Atoi(args[0])
if err != nil {
return fmt.Errorf("project id: %w", err)
}
c, err := newOO(cmd)
if err != nil {
return err
}
for _, uid := range args[1:] {
if _, err := c.RemoveProjectTeamUser(cmd.Context(), pid, uid); err != nil {
return fmt.Errorf("remove %s: %w", uid, err)
}
printObject(map[string]any{"project_id": pid, "user_id": uid, "removed": true})
}
return nil
},
}
}
func prjTeamSetCmd() *cobra.Command {
var notify bool
cmd := &cobra.Command{
Use: "set PROJECT_ID USER_ID [USER_ID...]",
Short: "Replace the project team with the given users (register several at once)",
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
pid, err := strconv.Atoi(args[0])
if err != nil {
return fmt.Errorf("project id: %w", err)
}
c, err := newOO(cmd)
if err != nil {
return err
}
team, err := c.SetProjectTeam(cmd.Context(), pid, args[1:], notify)
if err != nil {
return err
}
printTable([]string{"id", "displayName", "userName", "email", "isAdmin"}, teamRows(team))
return nil
},
}
cmd.Flags().BoolVar(&notify, "notify", false, "notify added members")
return cmd
}
// teamRows maps raw team member maps into table rows.
func teamRows(list []map[string]any) []map[string]any {
rows := make([]map[string]any, 0, len(list))
for _, m := range list {
rows = append(rows, map[string]any{
"id": idString(m, "id"),
"displayName": m["displayName"],
"userName": m["userName"],
"email": m["email"],
"isAdmin": m["isAdmin"],
})
}
return rows
} }
func prjListCmd() *cobra.Command { func prjListCmd() *cobra.Command {
@@ -166,6 +298,32 @@ func prjMilestoneCreateCmd() *cobra.Command {
return cmd return cmd
} }
func prjMilestoneDeleteCmd() *cobra.Command {
return &cobra.Command{
Use: "milestone-delete MILESTONE_ID [MILESTONE_ID...]",
Aliases: []string{"milestone-rm"},
Short: "Delete project milestone(s) by id",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, raw := range args {
id, err := strconv.ParseInt(raw, 10, 64)
if err != nil {
return fmt.Errorf("milestone id %q must be integer: %w", raw, err)
}
if err := c.DeleteMilestone(id); err != nil {
return fmt.Errorf("delete milestone %d: %w", id, err)
}
printObject(map[string]any{"milestone_id": id, "deleted": true})
}
return nil
},
}
}
func prjCreateCmd() *cobra.Command { func prjCreateCmd() *cobra.Command {
var desc, resp string var desc, resp string
var country, company string var country, company string
+54
View File
@@ -0,0 +1,54 @@
package main
import (
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra"
)
func init() {
projectsCmd.AddCommand(prjBoardSyncCmd())
}
// prjBoardSyncCmd upserts project milestones/tasks from a board YAML.
func prjBoardSyncCmd() *cobra.Command {
var apply bool
cmd := &cobra.Command{
Use: "board-sync BOARD.yaml",
Short: "Upsert project milestones/tasks from a board YAML (dry-run by default)",
Long: `Reads a board YAML (projects → milestones → tasks) and creates only the
milestones/tasks that are missing, matching by exact title. Existing entries are
left untouched, so the same file can be re-applied safely.
Dry-run by default; pass --apply to write to OnlyOffice.`,
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
board, err := onlyoffice.LoadBoard(args[0])
if err != nil {
return err
}
c, err := newOO(cmd)
if err != nil {
return err
}
res, err := c.SyncBoard(cmd.Context(), board, apply)
if err != nil {
return err
}
mode := "dry-run"
if apply {
mode = "apply"
}
printObject(map[string]any{
"mode": mode,
"projects": len(board.Projects),
"created_milestones": res.CreatedMilestones,
"skipped_milestones": res.SkippedMilestones,
"created_tasks": res.CreatedTasks,
"skipped_tasks": res.SkippedTasks,
})
return nil
},
}
cmd.Flags().BoolVar(&apply, "apply", false, "write to OnlyOffice (default: dry-run)")
return cmd
}
+216 -10
View File
@@ -3,6 +3,7 @@ package main
import ( import (
"fmt" "fmt"
"os" "os"
"path/filepath"
"strconv" "strconv"
"time" "time"
@@ -21,12 +22,48 @@ func projectFilesCmd() *cobra.Command {
} }
cmd.AddCommand(prjFilesListCmd()) cmd.AddCommand(prjFilesListCmd())
cmd.AddCommand(prjFilesUploadCmd()) cmd.AddCommand(prjFilesUploadCmd())
cmd.AddCommand(prjFilesReplaceInCmd())
cmd.AddCommand(prjFilesUpdateCmd())
cmd.AddCommand(prjFilesDownloadCmd()) cmd.AddCommand(prjFilesDownloadCmd())
cmd.AddCommand(prjFilesRenameCmd()) cmd.AddCommand(prjFilesRenameCmd())
cmd.AddCommand(prjFilesDeleteCmd()) cmd.AddCommand(prjFilesDeleteCmd())
cmd.AddCommand(prjFilesDedupeCmd())
// Convenience aliases into oo docs (md↔docx / OCR pipeline).
cmd.AddCommand(aliasDocsAsMD())
cmd.AddCommand(aliasDocsPutMD())
cmd.AddCommand(aliasDocsPutTxt())
cmd.AddCommand(aliasDocsPutXlsx())
return cmd return cmd
} }
func aliasDocsAsMD() *cobra.Command {
c := docsAsMDCmd()
c.Use = "as-md FILE_ID"
c.Short = "Alias of `oo docs as-md` — download OO file as Markdown (OCR if needed)"
return c
}
func aliasDocsPutMD() *cobra.Command {
c := docsPutMDCmd()
c.Use = "put-md PROJECT_ID MARKDOWN_PATH"
c.Short = "Alias of `oo docs put-md` — Markdown→DOCX upload into project"
return c
}
func aliasDocsPutTxt() *cobra.Command {
c := docsPutTxtCmd()
c.Use = "put-txt PROJECT_ID TEXT_PATH"
c.Short = "Alias of `oo docs put-txt` — plain text→DOCX upload into project"
return c
}
func aliasDocsPutXlsx() *cobra.Command {
c := docsPutXlsxCmd()
c.Use = "put-xlsx PROJECT_ID [LOCAL_XLSX]"
c.Short = "Alias of `oo docs put-xlsx` — generate/upload XLSX with formulas"
return c
}
func prjFilesListCmd() *cobra.Command { func prjFilesListCmd() *cobra.Command {
var showFolders bool var showFolders bool
cmd := &cobra.Command{ cmd := &cobra.Command{
@@ -73,9 +110,14 @@ func prjFilesListCmd() *cobra.Command {
} }
func prjFilesUploadCmd() *cobra.Command { func prjFilesUploadCmd() *cobra.Command {
return &cobra.Command{ var replace, allowDuplicate bool
cmd := &cobra.Command{
Use: "upload PROJECT_ID LOCAL_PATH [LOCAL_PATH...]", Use: "upload PROJECT_ID LOCAL_PATH [LOCAL_PATH...]",
Short: "Upload file(s) into the project's Documents folder", Short: "Upload file(s) into the project's Documents folder (upsert by stem, legacy↔OOXML ext)",
Long: `Default: replace an existing file with the same logical name (stem, treating
legacy .xls/.doc/.ppt and their OOXML .xlsx/.docx/.pptx as one file), like cp
overwrite. Pass --no-replace to fail when the name is taken; --allow-duplicate
to always create a new file id.`,
Args: cobra.MinimumNArgs(2), Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd) c, err := newOO(cmd)
@@ -84,12 +126,90 @@ func prjFilesUploadCmd() *cobra.Command {
} }
pid := args[0] pid := args[0]
for _, p := range args[1:] { for _, p := range args[1:] {
entry, err := c.UploadProjectFile(cmd.Context(), pid, p) var entry *onlyoffice.FileEntry
var deleted []int
switch {
case allowDuplicate:
entry, err = c.UploadProjectFile(cmd.Context(), pid, p)
case replace:
entry, deleted, err = c.UploadProjectFileReplacing(cmd.Context(), pid, p)
default:
entry, err = c.UploadProjectFileNoClobber(cmd.Context(), pid, p)
}
if err != nil {
return err
}
obj := fileEntryToMap(entry)
if len(deleted) > 0 {
obj["replaced_file_ids"] = deleted
}
printObject(obj)
}
return nil
},
}
cmd.Flags().BoolVar(&replace, "replace", true, "replace same stem|ext in project folder (default)")
cmd.Flags().BoolVar(&allowDuplicate, "allow-duplicate", false, "always create a new file even when the name exists")
return cmd
}
func prjFilesReplaceInCmd() *cobra.Command {
return &cobra.Command{
Use: "replace-in FOLDER_ID LOCAL_PATH [LOCAL_PATH...]",
Short: "Replace same-named file(s) in a folder: hard delete + fresh upload (no version history)",
Long: `Deletes any file in FOLDER_ID with the same stem|ext (hard delete — the CLI
delete is permanent) and uploads the local file fresh. Unlike 'update' this
leaves a single clean version.
Why it exists: repeated 'update' of a shared document accumulated a visible
version history and left a stale id. replace-in yields one clean revision; then
point links at the returned id (or keep an nginx alias for the legacy fileid).
Note: file ids are server-assigned; a fresh upload gets a new id.`,
Args: cobra.MinimumNArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
folderID := args[0]
for _, p := range args[1:] {
stem := onlyoffice.UploadStemFromLocal(p)
ext := onlyoffice.UploadExtFromLocal(p)
deleted, derr := c.DeleteFilesByDedupKey(cmd.Context(), folderID, stem, ext)
if derr != nil {
return derr
}
ent, uerr := c.UploadToFolder(cmd.Context(), folderID, p)
if uerr != nil {
return uerr
}
obj := fileEntryToMap(ent)
if len(deleted) > 0 {
obj["replaced_file_ids"] = deleted
}
printObject(obj)
}
return nil
},
}
}
func prjFilesUpdateCmd() *cobra.Command {
return &cobra.Command{
Use: "update FILE_ID LOCAL_PATH",
Short: "Overwrite an existing Documents file with new content (new version)",
Args: cobra.ExactArgs(2),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
entry, err := c.UpdateFile(cmd.Context(), args[0], args[1])
if err != nil { if err != nil {
return err return err
} }
printObject(fileEntryToMap(entry)) printObject(fileEntryToMap(entry))
}
return nil return nil
}, },
} }
@@ -106,20 +226,22 @@ func prjFilesDownloadCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
f, err := c.GetFile(cmd.Context(), args[0]) ctx := cmd.Context()
store := c.Files()
e, err := store.Stat(ctx, args[0])
if err != nil { if err != nil {
return err return err
} }
path := to path := to
if path == "" { if path == "" {
path = onlyoffice.SafeLocalFileName(onlyoffice.FileEntryTitle(f)) path = onlyoffice.SafeLocalFileName(e.Title)
} }
out, err := os.Create(path) out, err := os.Create(path)
if err != nil { if err != nil {
return err return err
} }
defer out.Close() defer out.Close()
n, err := c.DownloadFile(cmd.Context(), args[0], out) n, err := store.Download(ctx, args[0], out)
if err != nil { if err != nil {
_ = os.Remove(path) _ = os.Remove(path)
return err return err
@@ -146,11 +268,16 @@ func prjFilesRenameCmd() *cobra.Command {
if err != nil { if err != nil {
return err return err
} }
entry, err := c.RenameFile(cmd.Context(), args[0], args[1]) store := c.Files()
if err := store.Rename(cmd.Context(), args[0], args[1]); err != nil {
return err
}
entry, err := store.Stat(cmd.Context(), args[0])
if err != nil { if err != nil {
return err return err
} }
printObject(fileEntryToMap(entry)) entry.Title = args[1]
printObject(entryToMap(entry))
return nil return nil
}, },
} }
@@ -175,7 +302,7 @@ func prjFilesDeleteCmd() *cobra.Command {
} }
ids = append(ids, id) ids = append(ids, id)
} }
if err := c.DeleteFiles(cmd.Context(), ids); err != nil { if err := c.Files().Delete(cmd.Context(), args); err != nil {
return err return err
} }
printObject(map[string]any{"deleted": ids}) printObject(map[string]any{"deleted": ids})
@@ -184,6 +311,62 @@ func prjFilesDeleteCmd() *cobra.Command {
} }
} }
func prjFilesDedupeCmd() *cobra.Command {
var apply, cross bool
cmd := &cobra.Command{
Use: "dedupe PROJECT_ID",
Short: "Find (and optionally remove) duplicate files in project Documents folders",
Long: `Duplicates share the same logical name: stem|ext (OO title+fileExst).
Default: dry-run report. Pass --apply to delete older copies (keeps newest; --cross prefers non-trash folders).`,
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
groups, deleted, err := c.DedupeProject(cmd.Context(), args[0], onlyoffice.DedupOptions{
CrossFolder: cross,
}, apply)
if err != nil {
return err
}
rows := make([]map[string]any, 0, len(groups))
for _, g := range groups {
row := map[string]any{
"key": g.Key,
"folder_id": g.FolderID,
"folder_title": g.FolderTitle,
"keep_id": fileIDStr(g.Keep),
"keep_title": onlyoffice.FileEntryTitle(g.Keep),
"remove_count": len(g.Remove),
}
removeIDs := make([]string, 0, len(g.Remove))
for _, f := range g.Remove {
removeIDs = append(removeIDs, fileIDStr(f))
}
row["remove_ids"] = removeIDs
rows = append(rows, row)
}
out := map[string]any{
"project_id": args[0],
"dry_run": !apply,
"cross": cross,
"groups": len(groups),
"duplicates": rows,
}
if apply {
out["deleted_ids"] = deleted
}
printObject(out)
return nil
},
}
cmd.Flags().BoolVar(&apply, "apply", false, "delete duplicate files (default: report only)")
cmd.Flags().BoolVar(&cross, "cross", false, "also dedupe same stem|ext across folders (prefers non-_trash)")
return cmd
}
func fileEntryRows(files []*onlyoffice.FileEntry) []map[string]any { func fileEntryRows(files []*onlyoffice.FileEntry) []map[string]any {
rows := make([]map[string]any, 0, len(files)) rows := make([]map[string]any, 0, len(files))
for _, f := range files { for _, f := range files {
@@ -208,6 +391,29 @@ func fileEntryToMap(f *onlyoffice.FileEntry) map[string]any {
return m return m
} }
// entryToMap renders a canonical Entry with the same keys as fileEntryToMap.
func entryToMap(e onlyoffice.Entry) map[string]any {
m := map[string]any{
"id": e.ID,
"title": e.Title,
"fileExst": filepath.Ext(e.Title),
"contentLength": contentLengthString(e.Size),
}
if e.Updated != "" {
m["updated"] = e.Updated
} else if !e.Modified.IsZero() {
m["updated"] = e.Modified.Format(time.RFC3339)
}
return m
}
func contentLengthString(n int64) string {
if n <= 0 {
return ""
}
return strconv.FormatInt(n, 10)
}
func fileIDStr(f *onlyoffice.FileEntry) string { func fileIDStr(f *onlyoffice.FileEntry) string {
if f == nil || f.ID == nil { if f == nil || f.ID == nil {
return "" return ""
+102
View File
@@ -0,0 +1,102 @@
package main
import (
"fmt"
"strings"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/eslider/go-onlyoffice/cmd/internal/bootstrap"
"github.com/spf13/cobra"
)
func init() {
rootCmd.AddCommand(searchCmd())
}
// searchCmd queries the OnlyOffice document index through the file facade. The
// REST /api/2.0/files/@search endpoint only searches file names in the database;
// content search needs Elasticsearch (see docs/elasticsearch.md).
func searchCmd() *cobra.Command {
var (
content bool
folder string
limit int
backend string
asJSON bool
substring bool
)
cmd := &cobra.Command{
Use: "search QUERY...",
Short: "Full-text search over documents by name, optionally by content (Elasticsearch)",
Long: "Search the OnlyOffice Documents index.\n\n" +
"By default only file names are matched. With --content the query also\n" +
"matches extracted document text (document.attachment.content); this covers\n" +
"Office formats (docx/xlsx/pptx) and is slower.\n\n" +
"--backend own queries the separate index populated by `oo index`\n" +
"(ONLYOFFICE_ES_TEXT_INDEX, default oo_docs_text) instead, which also holds\n" +
"PDFs and scans (see docs/elasticsearch.md).\n\n" +
"Requires ONLYOFFICE_ES_URL (and optionally ONLYOFFICE_ES_INDEX,\n" +
"ONLYOFFICE_TENANT). See docs/elasticsearch.md for the tunnel setup.",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
if asJSON {
outputFormat = "json"
}
bootstrap.LoadEnv()
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
var (
searcher onlyoffice.Searcher
err error
)
switch strings.ToLower(strings.TrimSpace(backend)) {
case "", "oo", "elasticsearch":
searcher, err = c.Files().Search()
case "own", "es-text":
searcher, err = onlyoffice.NewESTextIndex(onlyoffice.ESTextConfigFromEnv())
default:
return fmt.Errorf("unknown search backend %q (want oo|own)", backend)
}
if err != nil {
return err
}
hits, err := searcher.Search(cmd.Context(), onlyoffice.SearchQuery{
Text: strings.Join(args, " "),
InContent: content,
FolderID: folder,
Limit: limit,
Substring: substring,
})
if err != nil {
return err
}
rows := make([]map[string]any, 0, len(hits))
for _, h := range hits {
folderPath := h.Path
if len(folderPath) == 0 && h.ParentID != "" {
folderPath = []string{h.ParentID}
}
rows = append(rows, map[string]any{
"path": c.UniquePath(cmd.Context(), folderPath, h.Title),
"id": h.ID,
"title": h.Title,
"folder": h.ParentID,
"score": h.Score,
"highlight": h.Highlight,
})
}
if outputFormat == "json" {
printJSON(rows)
return nil
}
printTable([]string{"path", "id", "title", "folder", "score", "highlight"}, rows)
return nil
},
}
cmd.Flags().BoolVar(&content, "content", false, "also match extracted document content")
cmd.Flags().BoolVar(&substring, "substring", false, "case-insensitive *term* title match; multiple QUERY args are ANDed")
cmd.Flags().StringVar(&folder, "folder", "", "limit to a Documents folder id (matches the folder subtree)")
cmd.Flags().IntVar(&limit, "limit", 20, "maximum number of results")
cmd.Flags().StringVar(&backend, "backend", "oo", "index to query: oo (OnlyOffice) | own (oo index)")
cmd.Flags().BoolVar(&asJSON, "json", false, "shorthand for --output json")
return cmd
}
+46
View File
@@ -0,0 +1,46 @@
package main
import (
"bytes"
"strings"
"testing"
)
func TestSearchCommandRegisteredWithFlags(t *testing.T) {
cmd, _, err := rootCmd.Find([]string{"search"})
if err != nil {
t.Fatal(err)
}
if cmd.Name() != "search" {
t.Fatalf("search resolved to %q", cmd.Name())
}
for _, name := range []string{"content", "folder", "limit", "json"} {
if cmd.Flags().Lookup(name) == nil {
t.Errorf("search: missing --%s flag", name)
}
}
if got := cmd.Flags().Lookup("limit").DefValue; got != "20" {
t.Errorf("--limit default = %q, want 20", got)
}
}
func TestSearchWithoutESURLIsClearError(t *testing.T) {
clearEnv(t, "ONLYOFFICE_ES_URL", "ONLYOFFICE_ES_INDEX", "ONLYOFFICE_TENANT")
errBuf := &bytes.Buffer{}
rootCmd.SetErr(errBuf)
rootCmd.SetOut(&bytes.Buffer{})
rootCmd.SetArgs([]string{"search", "Rechnung"})
t.Cleanup(func() {
rootCmd.SetArgs(nil)
rootCmd.SetOut(nil)
rootCmd.SetErr(nil)
})
err := rootCmd.Execute()
if err == nil {
t.Fatal("expected error without ONLYOFFICE_ES_URL")
}
if !strings.Contains(err.Error(), "ONLYOFFICE_ES_URL") {
t.Fatalf("error %q missing ONLYOFFICE_ES_URL", err.Error())
}
}
+308
View File
@@ -1,6 +1,12 @@
package main package main
import ( import (
"bufio"
"fmt"
"os"
"strings"
onlyoffice "github.com/eslider/go-onlyoffice"
"github.com/spf13/cobra" "github.com/spf13/cobra"
) )
@@ -13,6 +19,14 @@ func init() {
rootCmd.AddCommand(usersCmd) rootCmd.AddCommand(usersCmd)
usersCmd.AddCommand(usersListCmd()) usersCmd.AddCommand(usersListCmd())
usersCmd.AddCommand(usersSelfCmd()) usersCmd.AddCommand(usersSelfCmd())
usersCmd.AddCommand(usersGetCmd())
usersCmd.AddCommand(usersCreateCmd())
usersCmd.AddCommand(usersUpdateCmd())
usersCmd.AddCommand(usersDeleteCmd())
usersCmd.AddCommand(usersBlockCmd())
usersCmd.AddCommand(usersUnblockCmd())
usersCmd.AddCommand(usersPasswordCmd())
usersCmd.AddCommand(usersCheckCmd())
rootCmd.AddCommand(whoamiCmd()) rootCmd.AddCommand(whoamiCmd())
} }
@@ -65,6 +79,300 @@ func usersSelfCmd() *cobra.Command {
} }
} }
func usersGetCmd() *cobra.Command {
return &cobra.Command{
Use: "get USER_ID",
Short: "Show one portal user profile",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
u, err := c.GetUser(cmd.Context(), args[0])
if err != nil {
return err
}
printObject(u)
return nil
},
}
}
func usersCreateCmd() *cobra.Command {
var first, last, email, password, title, location, sex, comment string
var visitor bool
cmd := &cobra.Command{
Use: "create",
Short: "Create a portal user",
Long: `Create a portal user (POST /api/2.0/people).
Without --password the portal generates one and the account stays NotActivated
until the user follows the activation link. With --password the account is
Active immediately. Use --visitor for a guest account.`,
RunE: func(cmd *cobra.Command, args []string) error {
if email == "" || first == "" || last == "" {
return fmt.Errorf("--email, --first and --last are required")
}
c, err := newOO(cmd)
if err != nil {
return err
}
req := onlyoffice.NewUserRequest{
FirstName: first,
LastName: last,
Email: email,
Password: password,
Title: title,
Location: location,
Sex: sex,
Comment: comment,
}
if cmd.Flags().Changed("visitor") {
req.IsVisitor = &visitor
}
u, err := c.CreateUser(cmd.Context(), req)
if err != nil {
return err
}
printObject(map[string]any{
"id": idString(u, "id"),
"displayName": idString(u, "displayName"),
"email": idString(u, "email"),
"status": u["status"],
})
return nil
},
}
cmd.Flags().StringVar(&first, "first", "", "first name (required)")
cmd.Flags().StringVar(&last, "last", "", "last name (required)")
cmd.Flags().StringVar(&email, "email", "", "email (required)")
cmd.Flags().StringVar(&password, "password", "", "initial password (default: portal-generated)")
cmd.Flags().StringVar(&title, "title", "", "job title")
cmd.Flags().StringVar(&location, "location", "", "location")
cmd.Flags().StringVar(&sex, "sex", "", "sex: male|female")
cmd.Flags().StringVar(&comment, "comment", "", "comment")
cmd.Flags().BoolVar(&visitor, "visitor", false, "create as guest (isVisitor=true)")
return cmd
}
func usersUpdateCmd() *cobra.Command {
var first, last, email, title, location, sex, comment string
cmd := &cobra.Command{
Use: "update USER_ID",
Short: "Update portal user profile fields (only flags passed)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
body := map[string]any{}
if cmd.Flags().Changed("first") {
body["firstname"] = first
}
if cmd.Flags().Changed("last") {
body["lastname"] = last
}
if cmd.Flags().Changed("email") {
body["email"] = email
}
if cmd.Flags().Changed("title") {
body["title"] = title
}
if cmd.Flags().Changed("location") {
body["location"] = location
}
if cmd.Flags().Changed("sex") {
body["sex"] = sex
}
if cmd.Flags().Changed("comment") {
body["comment"] = comment
}
if len(body) == 0 {
return fmt.Errorf("nothing to update: pass at least one of --first/--last/--email/--title/--location/--sex/--comment")
}
c, err := newOO(cmd)
if err != nil {
return err
}
u, err := c.UpdateUser(cmd.Context(), args[0], body)
if err != nil {
return err
}
printObject(map[string]any{
"id": idString(u, "id"),
"displayName": idString(u, "displayName"),
})
return nil
},
}
cmd.Flags().StringVar(&first, "first", "", "first name")
cmd.Flags().StringVar(&last, "last", "", "last name")
cmd.Flags().StringVar(&email, "email", "", "email")
cmd.Flags().StringVar(&title, "title", "", "job title")
cmd.Flags().StringVar(&location, "location", "", "location")
cmd.Flags().StringVar(&sex, "sex", "", "sex: male|female")
cmd.Flags().StringVar(&comment, "comment", "", "comment")
return cmd
}
func usersDeleteCmd() *cobra.Command {
return &cobra.Command{
Use: "delete USER_ID [USER_ID...]",
Aliases: []string{"rm"},
Short: "Delete portal user(s) permanently",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, id := range args {
u, err := c.DeleteUser(cmd.Context(), id)
if err != nil {
return fmt.Errorf("delete %s: %w", id, err)
}
printObject(map[string]any{"id": id, "deleted": true, "displayName": idString(u, "displayName")})
}
return nil
},
}
}
func usersBlockCmd() *cobra.Command {
return &cobra.Command{
Use: "block USER_ID [USER_ID...]",
Aliases: []string{"disable"},
Short: "Block (terminate) user(s): login denied, profile kept",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, id := range args {
if err := c.BlockUser(cmd.Context(), id); err != nil {
return fmt.Errorf("block %s: %w", id, err)
}
printObject(map[string]any{"id": id, "blocked": true})
}
return nil
},
}
}
func usersUnblockCmd() *cobra.Command {
return &cobra.Command{
Use: "unblock USER_ID [USER_ID...]",
Aliases: []string{"enable", "activate"},
Short: "Unblock (activate) user(s)",
Args: cobra.MinimumNArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
c, err := newOO(cmd)
if err != nil {
return err
}
for _, id := range args {
if err := c.UnblockUser(cmd.Context(), id); err != nil {
return fmt.Errorf("unblock %s: %w", id, err)
}
printObject(map[string]any{"id": id, "unblocked": true})
}
return nil
},
}
}
func usersPasswordCmd() *cobra.Command {
var password string
cmd := &cobra.Command{
Use: "password USER_ID",
Short: "Set a user password (reads stdin when --password is empty)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
pwd := password
if pwd == "" {
b, err := readLine(os.Stdin)
if err != nil {
return fmt.Errorf("read password: %w", err)
}
pwd = b
}
if pwd == "" {
return fmt.Errorf("password is empty")
}
c, err := newOO(cmd)
if err != nil {
return err
}
if err := c.ChangeUserPassword(cmd.Context(), args[0], pwd); err != nil {
return err
}
printObject(map[string]any{"id": args[0], "password_changed": true})
return nil
},
}
cmd.Flags().StringVar(&password, "password", "", "new password (omit to read one line from stdin)")
return cmd
}
func usersCheckCmd() *cobra.Command {
var login, password string
cmd := &cobra.Command{
Use: "check",
Short: "Check that a login can authenticate (userName or email)",
Long: `Probes POST /api/2.0/authentication.json with the given credentials and
discards the token.
Where this is used: before handing portal credentials to an external party
(e.g. a guest given read access to a document pack), verify the login actually
works. On some portals the account email is the reliable login identifier — the
userName login fails with 500 for a freshly created user — so share the email,
not the userName.`,
RunE: func(cmd *cobra.Command, args []string) error {
if login == "" {
return fmt.Errorf("--login is required (userName or email)")
}
if password == "" {
if b, err := readLine(os.Stdin); err == nil {
password = b
}
}
c, err := newOO(cmd)
if err != nil {
return err
}
if err := c.AuthenticateAs(cmd.Context(), login, password); err != nil {
printObject(map[string]any{"login": login, "ok": false, "error": trimAuthErr(err)})
return fmt.Errorf("login failed for %s", login)
}
printObject(map[string]any{"login": login, "ok": true})
return nil
},
}
cmd.Flags().StringVar(&login, "login", "", "userName or email")
cmd.Flags().StringVar(&password, "password", "", "password (omit to read one line from stdin)")
return cmd
}
// trimAuthErr keeps the error short for table output.
func trimAuthErr(err error) string {
s := err.Error()
if len(s) > 160 {
s = s[:160] + "…"
}
return s
}
// readLine reads a single trimmed line from r.
func readLine(r *os.File) (string, error) {
sc := bufio.NewScanner(r)
if !sc.Scan() {
if err := sc.Err(); err != nil {
return "", err
}
return "", nil
}
return strings.TrimSpace(sc.Text()), nil
}
// whoamiCmd is a convenience shortcut at the root level. // whoamiCmd is a convenience shortcut at the root level.
func whoamiCmd() *cobra.Command { func whoamiCmd() *cobra.Command {
return &cobra.Command{ return &cobra.Command{
+170
View File
@@ -0,0 +1,170 @@
package onlyoffice
// Document conversion via the OnlyOffice DocumentServer converter.
//
// The DocumentServer (the same engine behind the portal's "Download as PDF")
// converts any office format. From a portal-reachable host the converter is
// exposed at "<portal>/ds-vpath/converter" (reverse proxy) or directly at
// "http://<docs-server>:8083/converter" (legacy path: /ConvertService.ashx).
//
// Flow: PresignedURI(fileId) → Convert(docsBase, secret, req) → download
// result.FileURL. The JWT is HS256 signed with the DocumentServer's
// services.CoAuthoring.secret (NOT storage.fs.secretString).
import (
"bytes"
"context"
"crypto/hmac"
"crypto/sha256"
"encoding/base64"
"encoding/json"
"fmt"
"io"
"net/http"
"net/url"
"strings"
)
// PresignedURI returns a short-lived, fetchable download URI for a portal file
// (GET /api/2.0/files/file/{fileId}/presigneduri). The DocumentServer can fetch
// it without the caller's session, so it is the input for Convert.
func (c *Client) PresignedURI(ctx context.Context, fileID string) (string, error) {
if fileID == "" {
return "", fmt.Errorf("file id is required")
}
raw, err := c.getJSON(ctx, fmt.Sprintf("/api/2.0/files/file/%s/presigneduri", url.PathEscape(fileID)))
if err != nil {
return "", err
}
resp, err := responseField(raw, "response")
if err != nil {
return "", err
}
var s string
if err := json.Unmarshal(resp, &s); err == nil && s != "" {
return s, nil
}
// Some builds return an object instead of a bare string.
var o map[string]any
if err := json.Unmarshal(resp, &o); err == nil {
for _, k := range []string{"uri", "url", "Uri", "Url"} {
if v, ok := o[k].(string); ok && v != "" {
return v, nil
}
}
}
return "", fmt.Errorf("presigneduri: unexpected response %s", truncate(string(resp), 200))
}
// ConvertRequest is the DocumentServer converter body.
type ConvertRequest struct {
URL string `json:"url"`
OutputType string `json:"outputtype"`
FileType string `json:"filetype,omitempty"`
Key string `json:"key"`
Title string `json:"title,omitempty"`
}
// ConvertResult is the DocumentServer converter reply.
type ConvertResult struct {
FileURL string `json:"fileUrl"`
FileType string `json:"fileType"`
Percent int `json:"percent"`
EndConvert bool `json:"endConvert"`
Error *int `json:"error,omitempty"`
}
// SignJWT builds an HS256 JWT with the given payload (stdlib only).
func SignJWT(secret string, payload any) (string, error) {
if secret == "" {
return "", fmt.Errorf("jwt secret is empty")
}
hb, err := json.Marshal(map[string]string{"alg": "HS256", "typ": "JWT"})
if err != nil {
return "", err
}
pb, err := json.Marshal(payload)
if err != nil {
return "", err
}
enc := base64.RawURLEncoding.EncodeToString
signing := enc(hb) + "." + enc(pb)
mac := hmac.New(sha256.New, []byte(secret))
mac.Write([]byte(signing))
return signing + "." + enc(mac.Sum(nil)), nil
}
// ConvertDocument asks a DocumentServer to convert req.URL into req.OutputType.
// docsBase is e.g. "https://portal/ds-vpath" or "http://localhost:8083";
// secret is the DocumentServer CoAuthoring JWT secret. Passes the JWT both as
// the AuthorizationJwt header and as a body token.
func (c *Client) ConvertDocument(ctx context.Context, docsBase, secret string, req ConvertRequest) (*ConvertResult, error) {
if strings.TrimSpace(docsBase) == "" {
return nil, fmt.Errorf("docs base url is required")
}
if req.URL == "" {
return nil, fmt.Errorf("source url is required")
}
if req.OutputType == "" {
return nil, fmt.Errorf("outputtype is required")
}
if req.Key == "" {
return nil, fmt.Errorf("conversion key is required")
}
jwt, err := SignJWT(secret, req)
if err != nil {
return nil, err
}
body, err := json.Marshal(req)
if err != nil {
return nil, err
}
endpoint := strings.TrimRight(docsBase, "/") + "/converter"
httpReq, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(body))
if err != nil {
return nil, err
}
httpReq.Header.Set("Content-Type", "application/json")
httpReq.Header.Set("Accept", "application/json")
httpReq.Header.Set("AuthorizationJwt", "Bearer "+jwt)
resp, err := c.client.Do(httpReq)
if err != nil {
return nil, fmt.Errorf("converter request: %w", err)
}
defer resp.Body.Close()
raw, err := io.ReadAll(resp.Body)
if err != nil {
return nil, err
}
if resp.StatusCode >= 400 {
return nil, fmt.Errorf("converter: %d %s", resp.StatusCode, truncate(string(raw), 300))
}
var out ConvertResult
if err := json.Unmarshal(raw, &out); err != nil {
return nil, fmt.Errorf("converter decode: %w (%s)", err, truncate(string(raw), 200))
}
if out.Error != nil {
return &out, fmt.Errorf("converter error %d", *out.Error)
}
if out.FileURL == "" {
return &out, fmt.Errorf("converter returned no fileUrl")
}
return &out, nil
}
// DownloadURLTo streams an absolute URL (no portal auth) into dst.
func (c *Client) DownloadURLTo(ctx context.Context, rawurl string, dst io.Writer) (int64, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, rawurl, nil)
if err != nil {
return 0, err
}
resp, err := c.client.Do(req)
if err != nil {
return 0, err
}
defer resp.Body.Close()
if resp.StatusCode >= 400 {
return 0, fmt.Errorf("download: %d", resp.StatusCode)
}
return io.Copy(dst, resp.Body)
}
+24
View File
@@ -0,0 +1,24 @@
package onlyoffice
import (
"strings"
"testing"
)
func TestSignJWT(t *testing.T) {
payload := map[string]any{"url": "u", "outputtype": "pdf"}
tok, err := SignJWT("secret", payload)
if err != nil {
t.Fatal(err)
}
if n := len(strings.Split(tok, ".")); n != 3 {
t.Fatalf("JWT must have 3 parts, got %d", n)
}
tok2, _ := SignJWT("secret", payload)
if tok != tok2 {
t.Fatal("SignJWT must be deterministic for identical input")
}
if _, err := SignJWT("", payload); err == nil {
t.Fatal("expected error for empty secret")
}
}
+29
View File
@@ -15,10 +15,15 @@ import (
) )
// ListContacts returns a page of CRM contacts and the total count. // ListContacts returns a page of CRM contacts and the total count.
// sortBy=id is always set: OnlyOffice filter.json without an explicit sort
// order is non-deterministic on large contact sets, so a paged walk
// (ListAllContacts, ListContactsByTag, FindCompany, FindPerson) can skip or
// duplicate contacts across page boundaries.
func (c *Client) ListContacts(ctx context.Context, count, startIndex int, search string) ([]map[string]any, int, error) { func (c *Client) ListContacts(ctx context.Context, count, startIndex int, search string) ([]map[string]any, int, error) {
q := url.Values{} q := url.Values{}
q.Set("count", strconv.Itoa(count)) q.Set("count", strconv.Itoa(count))
q.Set("startIndex", strconv.Itoa(startIndex)) q.Set("startIndex", strconv.Itoa(startIndex))
q.Set("sortBy", "id")
if search != "" { if search != "" {
q.Set("filterValue", search) q.Set("filterValue", search)
} }
@@ -243,6 +248,24 @@ func (c *Client) DeleteContact(ctx context.Context, contactID string) (map[strin
return c.deleteObject(ctx, fmt.Sprintf("/api/2.0/crm/contact/%s.json", url.PathEscape(contactID))) return c.deleteObject(ctx, fmt.Sprintf("/api/2.0/crm/contact/%s.json", url.PathEscape(contactID)))
} }
// UpdateContactName renames a CRM company contact displayName.
// Uses the company endpoint (person names go through /crm/contact/person/{id}).
func (c *Client) UpdateContactName(ctx context.Context, contactID, newName string) (map[string]any, error) {
body := map[string]any{
"displayName": newName,
"companyName": newName,
"isCompany": true,
}
out, err := c.putJSONObject(ctx, fmt.Sprintf("/api/2.0/crm/contact/company/%s.json", url.PathEscape(contactID)), body)
if err != nil {
return out, err
}
if fresh, gerr := c.GetContact(ctx, contactID); gerr == nil && fresh != nil {
out = fresh
}
return out, nil
}
// ListContactTags returns all CRM contact tags (title + relativeItemsCount). // ListContactTags returns all CRM contact tags (title + relativeItemsCount).
func (c *Client) ListContactTags(ctx context.Context) ([]map[string]any, error) { func (c *Client) ListContactTags(ctx context.Context) ([]map[string]any, error) {
return c.ResponseArray(ctx, "/api/2.0/crm/contact/tag.json") return c.ResponseArray(ctx, "/api/2.0/crm/contact/tag.json")
@@ -288,6 +311,7 @@ func (c *Client) ListContactsByTag(ctx context.Context, tagName string, count, s
q.Set("count", strconv.Itoa(count)) q.Set("count", strconv.Itoa(count))
q.Set("startIndex", strconv.Itoa(startIndex)) q.Set("startIndex", strconv.Itoa(startIndex))
q.Set("tags", tagName) q.Set("tags", tagName)
q.Set("sortBy", "id")
raw, err := c.getJSON(ctx, "/api/2.0/crm/contact/filter.json?"+q.Encode()) raw, err := c.getJSON(ctx, "/api/2.0/crm/contact/filter.json?"+q.Encode())
if err != nil { if err != nil {
return nil, 0, err return nil, 0, err
@@ -717,6 +741,11 @@ func (c *Client) DeleteCRMTask(ctx context.Context, id string) (map[string]any,
return c.deleteObject(ctx, fmt.Sprintf("/api/2.0/crm/task/%s.json", url.PathEscape(id))) return c.deleteObject(ctx, fmt.Sprintf("/api/2.0/crm/task/%s.json", url.PathEscape(id)))
} }
// CloseCRMTask closes (completes) a CRM task via the task close endpoint.
func (c *Client) CloseCRMTask(ctx context.Context, id string) (map[string]any, error) {
return c.putFormObject(ctx, fmt.Sprintf("/api/2.0/crm/task/%s/close.json", url.PathEscape(id)), url.Values{})
}
// ListTaskCategories returns CRM task categories. // ListTaskCategories returns CRM task categories.
func (c *Client) ListTaskCategories(ctx context.Context) ([]map[string]any, error) { func (c *Client) ListTaskCategories(ctx context.Context) ([]map[string]any, error) {
return c.ResponseArray(ctx, "/api/2.0/crm/task/category.json") return c.ResponseArray(ctx, "/api/2.0/crm/task/category.json")
+91
View File
@@ -0,0 +1,91 @@
package onlyoffice
import (
"fmt"
"strconv"
"strings"
)
// addressCategoryCodes maps ASC.CRM.Core.AddressCategory names to their numeric
// codes. The OO API expects the code, the UI/docs use the label.
var addressCategoryCodes = map[string]int{
"home": 0,
"postal": 1,
"office": 2,
"billing": 3,
"other": 4,
"work": 5,
}
// AddressCategoryCode returns the numeric code for an AddressCategory label
// (Home|Postal|Office|Billing|Other|Work) or a numeric string. Unknown/empty
// labels fall back to Billing, the category `oo companies create` used.
func AddressCategoryCode(category string) int {
s := strings.ToLower(strings.TrimSpace(category))
if s == "" {
return addressCategoryCodes["billing"]
}
if n, err := strconv.Atoi(s); err == nil {
if n >= 0 && n <= 5 {
return n
}
return addressCategoryCodes["billing"]
}
if n, ok := addressCategoryCodes[s]; ok {
return n
}
return addressCategoryCodes["billing"]
}
// ContactAddresses returns the postal address rows of a contact map.
func ContactAddresses(contact map[string]any) []map[string]any {
if rows, ok := contact["addresses"].([]any); ok {
return mapsFromAnySlice(rows)
}
if rows, ok := contact["addresses"].([]map[string]any); ok {
return rows
}
return nil
}
// HasContactAddress reports whether a contact already has the given postal
// address. street+city+zip+category identify it; comparison is normalized.
func HasContactAddress(contact map[string]any, street, city, zip, category string) bool {
wantStreet, wantCity, wantZip := normalizeAddressPart(street), normalizeAddressPart(city), normalizeAddressPart(zip)
wantCat := AddressCategoryCode(category)
for _, row := range ContactAddresses(contact) {
if normalizeAddressPart(fmt.Sprint(row["street"])) != wantStreet {
continue
}
if normalizeAddressPart(fmt.Sprint(row["city"])) != wantCity {
continue
}
if normalizeAddressPart(fmt.Sprint(row["zip"])) != wantZip {
continue
}
if int(anyToFloat(row["category"])) != wantCat {
continue
}
return true
}
return false
}
func normalizeAddressPart(s string) string {
s = strings.ToLower(strings.TrimSpace(s))
return strings.Join(strings.Fields(s), " ")
}
func anyToFloat(v any) float64 {
switch n := v.(type) {
case float64:
return n
case int:
return float64(n)
case string:
f, _ := strconv.ParseFloat(strings.TrimSpace(n), 64)
return f
default:
return 0
}
}
+42
View File
@@ -0,0 +1,42 @@
package onlyoffice
import "testing"
func TestAddressCategoryCode(t *testing.T) {
cases := map[string]int{
"Home": 0, "Postal": 1, "Office": 2, "Billing": 3, "Other": 4, "Work": 5,
"billing": 3, " work ": 5, "3": 3, "5": 5,
"": 3, "nonsense": 3, "99": 3,
}
for in, want := range cases {
if got := AddressCategoryCode(in); got != want {
t.Errorf("AddressCategoryCode(%q) = %d, want %d", in, got, want)
}
}
}
func TestHasContactAddress(t *testing.T) {
contact := map[string]any{
"addresses": []any{
map[string]any{
"street": "Lubanas st. 125a-25", "city": "Riga",
"zip": "LV-1021", "country": "Latvia", "category": float64(3),
},
},
}
if !HasContactAddress(contact, " Lubanas St. 125a-25 ", "riga", "lv-1021", "Billing") {
t.Error("want match (normalized, case-insensitive)")
}
if HasContactAddress(contact, "Lubanas st. 125a-25", "Riga", "LV-1021", "Work") {
t.Error("different category must not match")
}
if HasContactAddress(contact, "Lubanas st. 125a-25", "Riga", "00000", "Billing") {
t.Error("different zip must not match")
}
if HasContactAddress(map[string]any{}, "x", "y", "z", "Billing") {
t.Error("empty contact must not match")
}
if got := ContactAddresses(contact); len(got) != 1 {
t.Fatalf("ContactAddresses = %d rows", len(got))
}
}
+126
View File
@@ -0,0 +1,126 @@
package onlyoffice
import (
"context"
"fmt"
"strconv"
"strings"
)
// OpportunityAudit is one CRM opportunity with its resource counts and a coarse
// class, for hygiene reporting (see `oo crm audit`).
type OpportunityAudit struct {
ID int64 `json:"id"`
Title string `json:"title"`
Created string `json:"created,omitempty"`
Files int `json:"files"`
OpenTasks int `json:"open_tasks"`
ClosedTasks int `json:"closed_tasks"`
Members int `json:"members"`
GroupKey string `json:"group_key,omitempty"`
Class string `json:"class"` // ok | dup | empty | junk-title
}
// AuditOpportunities lists every opportunity with file/task/member counts and a
// coarse class. The classification is generic and rule-free:
//
// dup — another opportunity shares the same title key
// empty — no files, tasks or members
// junk-title — title is not of the "Role @ Company" shape
// ok — everything else
//
// Callers that need stricter business rules can post-process the result.
func (c *Client) AuditOpportunities(ctx context.Context) ([]OpportunityAudit, error) {
deals, err := c.ListAllOpportunities(ctx)
if err != nil {
return nil, err
}
tasks, _, err := c.ListCRMTasks(ctx, 5000, 0)
if err != nil {
return nil, err
}
open, closed := taskCountsByOpportunity(tasks)
out := make([]OpportunityAudit, 0, len(deals))
for _, row := range deals {
id := auditID(row["id"])
if id == 0 {
continue
}
title := auditStr(row["title"])
files := 0
if fl, ferr := c.ListOpportunityFiles(ctx, strconv.FormatInt(id, 10)); ferr == nil {
files = len(fl)
}
key := strconv.FormatInt(id, 10)
out = append(out, OpportunityAudit{
ID: id,
Title: title,
Created: auditStr(row["created"]),
Files: files,
OpenTasks: open[key],
ClosedTasks: closed[key],
Members: len(OpportunityMembers(row)),
GroupKey: DealTitleKey(title, false),
})
}
groupCount := map[string]int{}
for _, a := range out {
groupCount[a.GroupKey]++
}
for i := range out {
a := &out[i]
switch {
case groupCount[a.GroupKey] > 1:
a.Class = "dup"
case a.Files == 0 && a.OpenTasks == 0 && a.ClosedTasks == 0 && a.Members == 0:
a.Class = "empty"
case !strings.Contains(a.Title, "@") || strings.HasPrefix(strings.TrimSpace(a.Title), "@"):
a.Class = "junk-title"
default:
a.Class = "ok"
}
}
return out, nil
}
// taskCountsByOpportunity buckets CRM task statuses per opportunity id.
func taskCountsByOpportunity(tasks []map[string]any) (open, closed map[string]int) {
open, closed = map[string]int{}, map[string]int{}
for _, t := range tasks {
ent, ok := t["entity"].(map[string]any)
if !ok || auditStr(ent["entityType"]) != "opportunity" {
continue
}
eid := auditStr(ent["entityId"])
status := strings.ToLower(auditStr(t["status"]))
if status == "2" || status == "closed" {
closed[eid]++
} else {
open[eid]++
}
}
return open, closed
}
func auditID(v any) int64 {
switch x := v.(type) {
case float64:
return int64(x)
case int:
return int64(x)
case int64:
return x
default:
n, _ := strconv.ParseInt(strings.TrimSpace(fmt.Sprint(x)), 10, 64)
return n
}
}
func auditStr(v any) string {
if v == nil {
return ""
}
return fmt.Sprint(v)
}
+36
View File
@@ -0,0 +1,36 @@
package onlyoffice
import "testing"
func TestAuditID(t *testing.T) {
cases := []struct {
in any
want int64
}{
{float64(12), 12},
{7, 7},
{int64(9), 9},
{"42", 42},
{nil, 0},
}
for _, c := range cases {
if got := auditID(c.in); got != c.want {
t.Fatalf("auditID(%v) = %d, want %d", c.in, got, c.want)
}
}
}
func TestTaskCountsByOpportunity(t *testing.T) {
open, closed := taskCountsByOpportunity([]map[string]any{
{"id": 1, "status": 1, "entity": map[string]any{"entityType": "opportunity", "entityId": 10}},
{"id": 2, "status": "2", "entity": map[string]any{"entityType": "opportunity", "entityId": 10}},
{"id": 3, "status": 1, "entity": map[string]any{"entityType": "contact", "entityId": 10}},
{"id": 4, "entity": "not-a-map"},
})
if open["10"] != 1 || closed["10"] != 1 {
t.Fatalf("open=%v closed=%v", open, closed)
}
if len(open) != 1 || len(closed) != 1 {
t.Fatalf("unexpected buckets: open=%v closed=%v", open, closed)
}
}
+27
View File
@@ -0,0 +1,27 @@
---
type: reference
status: current
related:
- README.md
---
# go-onlyoffice — docs
Индекс справочников. Общее — [README.md](../README.md), правила — [AGENTS.md](../AGENTS.md).
## Файлы
- [unified-file-client.md](unified-file-client.md) — единый файловый клиент:
`Entry`/`FileStore`/`FileClient`, бэкенды REST/DAV/SQL/ES, env, как добавить
бэкенд.
- [community-server-db.md](community-server-db.md) — read-only SQL-бэкенд
(MySQL/PostgreSQL): схема, SSH-туннель, DSN, MinIO download.
- [elasticsearch.md](elasticsearch.md) — поиск: индекс OnlyOffice `files_file`
и свой `oo_docs_text` (PDF/сканы), туннель.
- [index-and-search.md](index-and-search.md) — карта контуров поиска и как
обновлять индексы (`oo index`, `oo search`).
- [rate-limiting.md](rate-limiting.md) — rate limit, exponential backoff,
`Retry-After`, общий cooldown против 429; env `OO_RATE_LIMIT`/`OO_BURST`/
`OO_RETRY_*`.
## Тесты
Команды и туннели — раздел Testing в [README.md](../README.md#testing).
+179
View File
@@ -0,0 +1,179 @@
---
type: reference
status: current
related:
- README.md
- filestore_pg.go
- docs/elasticsearch.md
---
# Community Server DB — прямой SQL-доступ (read-only)
## Что это
Бэкенд `pgStore` (`filestore_pg.go`) читает файлы и папки **напрямую из БД
Community Server**, без HTTP-слоя. Реализует `FileStore` (`List`/`Stat`/
`Download`) и `Searcher` по имени. Запись запрещена: все write-методы
возвращают `ErrReadOnly`.
## Что за БД (research, live)
Проверено на VM `onlyoffice-v2` (SSH `127.0.0.1:32`):
- Community Server работает на **MySQL 8.0**, не на PostgreSQL.
- Хост: `127.0.0.1:3306` внутри VM, база `onlyoffice`.
- Конфиг: `/etc/onlyoffice/communityserver/appsettings.production.json`,
`providerName: MySql.Data.MySqlClient`.
- Таблицы: `files_file`, `files_folder`, `files_folder_tree`,
`files_security`, тенанты — `tenants_tenants` (не `tenants`).
- PostgreSQL 16 в той же VM — **наш** контур (`edw_docs`, роли `edw`/`edw_ro`,
office-assistant), к OnlyOffice отношения не имеет. `files_file` в PG нет.
- Портал хранит файлы в **S3/MinIO** (DiscStorage только для мелочи).
Бакет `office`, объект — по ключу (см. ниже).
Вывод: бэкенд назван по issue «PostgreSQL», но живой источник — MySQL.
`database/sql` + драйвер по DSN: `mysql` для MySQL, `pgx` для PostgreSQL.
`Name()` возвращает фактический движок (`mysql` или `postgres`).
## Схема
`files_file` — одна строка **на версию** (PK `tenant_id, id, version`):
| поле | смысл |
|------|-------|
| `id` | id файла (тот же, что в REST/ES) |
| `version` | номер версии этой строки |
| `version_group` | номер версии |
| `current_version` | `1` = текущая версия, `0` = старая |
| `folder_id` | id родительской папки |
| `title` | имя файла с расширением |
| `content_length` | размер в байтах |
| `create_on`, `modified_on` | даты (UTC, без зоны) |
| `tenant_id` | тенант (портал) |
`files_folder`: `id`, `parent_id`, `title`, `create_on`, `modified_on`,
`tenant_id`. `files_folder_tree`: `folder_id`, `parent_id`, `level` — готовое
дерево, пока не используется.
Текущую строку файла берём по `current_version = 1`.
## Доступ (SSH-туннель)
MySQL слушает только `127.0.0.1:3306` внутри VM. Снаружи — SSH-туннель
(SSH в VM открыт как `127.0.0.1:32`):
```bash
ssh -f -N -o ControlMaster=no -o ControlPath=none \
-p 32 -i ~/.ssh/id_ed25519 \
-L 3306:127.0.0.1:3306 root@127.0.0.1
# MySQL DSN затем:
# root:<pw>@tcp(127.0.0.1:3306)/onlyoffice?parseTime=true
```
Любой свободный локальный порт подойдёт (напр. `13306`); тогда тот же порт —
в DSN. `-o ControlMaster=no -o ControlPath=none` обязательны: иначе forward
уходит в persistent master из `~/.ssh/config`.
Креды MySQL — в конфиге Community Server внутри VM:
`/etc/onlyoffice/communityserver/appsettings.production.json` →
`ConnectionStrings.connectionString` (поля `User ID`, `Password`), база
`onlyoffice`. В самом MySQL-контейнере (`onlyoffice-mysql-server`) база пустая;
рабочий сервер — host-mysqld на `127.0.0.1:3306` (207 таблиц). Не печатать
пароль.
## Переменные
| env | default | смысл |
|-----|---------|-------|
| `ONLYOFFICE_DSN` | — | DSN драйвера (MySQL `...@tcp(...)/...` или `postgres://...`) |
| `ONLYOFFICE_PG_DRIVER` | авто | `postgres` или `mysql`; иначе по форме DSN |
| `ONLYOFFICE_PG_TENANT` | `ONLYOFFICE_TENANT` | фильтр `tenant_id` (пусто = все) |
| `ONLYOFFICE_PG_HOST/PORT/USER/PASSWORD/DBNAME/SSLMODE` | — | собрать PG DSN, если `ONLYOFFICE_DSN` пуст |
Имена — в [`.env.example`](../.env.example). Секретов нет.
## Использование
Напрямую: `NewPGStore(PGConfigFromEnv())`.
Через фасад (эпик #34): SQL-стор регистрируется на `FileClient`. После этого
`Read()` и все чтения (`Stat`/`List`) идут в БД, `Write()` остаётся REST/DAV.
```go
c := onlyoffice.NewClient(onlyoffice.GetEnvironmentCredentials())
sql, err := c.SQLFileStore() // открыть из env; caller закрывает
if err != nil { /* нет DSN / нет связи */ }
if closer, ok := sql.(interface{ Close() error }); ok { defer closer.Close() }
f := c.Files()
f.RegisterStore(onlyoffice.ProviderPG, sql)
e, _ := f.Stat(ctx, "19423") // e.Provider == "mysql" — ответил SQL
```
`Client.FileStore("pg"|"sql"|"postgres"|"mysql")` тоже отдаёт SQL-стор
(открывает из env). Если DSN нет/битый — возвращается не `nil`, а заглушка,
чей метод отдаёт ошибку открытия; ошибку как таковую даёт `SQLFileStore()`.
Различить бэкенд в ответе можно по `Entry.Provider` (`mysql` у SQL, `rest` у
REST).
## Download (MinIO)
`Download` не ходит в REST. Ключ объекта собирается из строки `files_file`:
```
00/00/<tenant>/files/folder_<shard>/file_<id>/v<version>/content.<ext>
shard = (id/1000 + 1) * 1000
```
`shard` — не `folder_id`, а следующая тысяча над `id` (файл 3727 →
`folder_4000`). Проверено live по бакету `office`.
Стриминг переиспользует `downloadMinioObject` из `storage_fallback.go`
(та же подпись SigV4 и `MINIO_*`), без дублирования.
Ограничение: схема валидна только для файлов, лежащих в **MinIO/S3** (старые
папки). Файлы в **Disc**-хранилище портала (`Data/Products/Files/...`, новые
папки) по этому ключу недоступны — `Download` вернёт `404`. Если
`MINIO_ACCESS_KEY`/`MINIO_SECRET_KEY` не заданы, `Download` вернёт явную
ошибку; `Stat`/`List`/`Search` работают и без них.
## Тесты
```bash
go test ./... # unit: rebind, csObjectKey, маппинг
go test -race ./...
# integration (нужен DSN; skip без него)
ONLYOFFICE_DSN='root:<pw>@tcp(127.0.0.1:3306)/onlyoffice?parseTime=true' \
ONLYOFFICE_PG_TENANT=1 \
ONLYOFFICE_PG_TEST_FILE_ID=19423 \
ONLYOFFICE_PG_TEST_FOLDER_ID=676 \
go test -tags=integration -run 'TestIntegrationPGStore|TestIntegrationSQLFacade' -v ./...
# плюс MINIO_* для сверки Download с REST (иначе этот шаг skip)
MINIO_ENDPOINT=http://127.0.0.1:9000 MINIO_BUCKET=office \
MINIO_ACCESS_KEY=... MINIO_SECRET_KEY=... \
go test -tags=integration -run TestIntegrationPGStore -v ./...
```
- `TestIntegrationPGStore` — `Stat`/`List`/`Download` SQL против REST и
`ErrReadOnly` у write-методов.
- `TestIntegrationSQLFacade` — SQL-стор, зарегистрированный на фасаде, реально
обслуживает чтения: `Read().Name()` = SQL-бэкенд, `Entry.Provider == "mysql"`
(у REST — `"rest"`), сверка `Stat`/`List` с REST, и прямой
`Client.FileStore("pg")`.
Без `ONLYOFFICE_DSN` оба теста делают чистый `skip`.
## Грабли
- MySQL хранит `datetime` без зоны; `parseTime=true` (ставится автоматически)
читает их как UTC. REST отдаёт `+02:00` — сравнивать моменты, не строки.
- `GetFile` (REST) не отдаёт `contentLength` — размер сверять с `Stat` SQL.
- Один файл = много строк `files_file` (по версиям). Без `current_version = 1`
получите дубликаты.
- `folder_id` не входит в ключ MinIO; ключ считает `shard` от `id`.
- Searcher SQL ищет только по имени (`LIKE`). Контент — Elasticsearch
([elasticsearch.md](elasticsearch.md)).
-124
View File
@@ -1,124 +0,0 @@
# CRM associations (company ↔ person ↔ deal ↔ project ↔ invoice ↔ mail)
Operational rules for the `oo` CLI and this library. Business SSOT remains
OnlyOffice Workspace CRM + Projects.
## Canonical graph
One **legal company** owns the relationship. Do not invent a second “bill-to”
company just for PDF layout.
```text
Company
├── Person (buyer contact) oo persons create --company-id
├── Opportunity / Deal oo opportunities … ; member-add company + person
├── Project (hub) oo projects … ; contacts add company + person
│ └── Epic + subtasks
└── Invoice (Draft → …) oo invoices create --contact COMPANY --opportunity DEAL
└── PDF file oo invoices pdf ID
└── Mail draft oo mails draft-invoice --invoice ID --to …
```
| Layer | CLI | Must link |
|-------|-----|-----------|
| Company | `oo companies create` | website, email, phone, **one** Billing address |
| Person | `oo persons create --company-id` / `oo persons update ID` | job title; never encode employer in `lastName`; **update uses JSON** (form PUT ignores `companyId`/`about`) |
| Deal | `oo opportunities create` + `member-add` | company **and** person as members |
| Project | `oo projects create` + `contacts add` | same company + person |
| Invoice | `oo invoices create --contact COMPANY --opportunity DEAL` | `entityId` at **create** |
| Mail | `oo mails draft-invoice` | attach current PDF; **do not send** until confirmed |
UI checks (same company card):
- `#contacts` → person
- `#deals` → opportunity
- `#projects` → hub project
- `#invoices` on the **deal** → invoice (needs `entity`)
- `#files` → preferably **one** current invoice PDF
**Project Team ≠ Project Contacts.** Team = portal users. CRM people/companies
show under the project **Contacts** tab (`oo projects contacts list`).
## Hard rules
1. **One company per legal entity.** Duplicate “bill-to” contacts empty Deals /
Projects / Contacts tabs and break merge. Prefer
`oo contacts merge FROM INTO` (keeps `INTO`) or `oo companies dedupe`.
2. **Link invoice → deal at create.**
`POST /crm/invoice` with `entityId` + `entityType: 0` (Opportunity).
`oo invoices update … --opportunity` often returns **400**
(“Value does not fall within the expected range”). If the link is missing,
delete the Draft and recreate with `--opportunity`.
3. **Bill To = company id**, not a throwaway contact. Person stays under the
company (`companyId`). Optional `consigneeId` for Empfänger when the portal
template prints it.
4. **Stay Draft until mail is ready.** Billed (`status id=2`) is **not editable**
via content PUT. Going Billed → Draft via `…/crm/invoice/status/1` usually
**does not work** — delete + recreate Draft instead.
5. **Do not regenerate PDF in a loop** without cleanup. Each
`GET …/crm/invoice/{id}/pdf` attaches a new file to the company (and often
the deal). Keep `invoice.fileID`; delete older PDFs with
`oo invoices pdf-cleanup ID` / Documents `fileops/delete`.
## Invoice PDF quirks
| Symptom | Workaround |
|---------|------------|
| Cached / stale PDF | Touch invoice (Draft PUT that clears `fileID`), then `GET …/pdf` — `oo invoices pdf ID --force` |
| Billing address missing on **new** PDFs | Temporary multiline `companyName` (`Line1\nLine2\n…`) on the **canonical** company → force PDF → restore clean name. Cached `fileID` keeps the multiline Bill To. |
| Separate bill-to company for newlines | **Forbidden** — merge back to the real company |
| Invoice **number** won’t change on PUT | Delete Draft and recreate with the desired number |
| Notizen / Bedingungen spacing | Leading `\n` and blank lines only — no HTML (tags print literally) |
| Issuer street lines | Organisation profile address (`street` with `\n`), not only terms |
Status ids commonly used: `1` Draft, `2` Billed, `3` Rejected, `4` Paid.
## Mail quirks
| Symptom | Workaround |
|---------|------------|
| Signature / body doubles chat URL | Put chat in **one** place only. UI drafts: signature. API send: body (API **does not** append signature). |
| Signature / body cuts URL at `#` | Plain text URLs — avoid `<a href="…#…">` (or encode `#` as `%23` in href) |
| German letter spacing | Blank `<p>&nbsp;</p>` between blocks (`MailHTMLWithBlankParagraphs`) |
| Send | `PUT /api/2.0/mail/messages/send.json` with `id/from/to/subject/body`; omit empty `cc`/`bcc`. Never auto-send; draft only until the human confirms |
Prefer OnlyOffice Mail (`/addons/mail/#drafts`) for invoice delivery until confirmed.
## Project / task quirks
- Hub title: `CC | Company` (e.g. `DE | Acme GmbH`).
- Streams = epics/tasks under the hub, not a third title segment (unless the
project itself is a named delivery stream).
- Closing a **subtask**:
`PUT /api/2.0/project/task/{epicId}/{subtaskId}/status` with `status=2`.
`oo tasks update SUBTASK -s closed` returns **404** for subtasks.
- After deleting a CRM contact, `GET /project/contact/{deletedId}` may still
return projects (ghost). Official project contact list should only show live
ids; unlink may 400 if the contact is gone.
## Merge / cleanup cheat sheet
```bash
# Keep the preferred company (INTO), drop the duplicate (FROM)
oo contacts merge FROM_ID INTO_ID
# Or by normalized name (careful — whole CRM)
oo companies dedupe
# Invoice ↔ deal must exist at create
oo invoices create --number P-YYYY-NN --contact COMPANY_ID --item ITEM_ID \
--price 300 --opportunity DEAL_ID --language de-DE …
# Fresh PDF + prune older PDFs on company/deal
oo invoices pdf INVOICE_ID --force
oo invoices pdf-cleanup INVOICE_ID
# Mail draft (no send)
oo mails draft-invoice --invoice INVOICE_ID --to billing@example.com
```
## Related
- README § invoices / mail / CRM cleanup
- Personal workspace tooling (disk inventory, dossier sync): private
`git.produktor.io/eSlider/oo-workspace` (`oow` CLI)
+260
View File
@@ -0,0 +1,260 @@
---
type: reference
status: current
related:
- README.md
- filestore_es.go
---
# Elasticsearch — полнотекстовый поиск OnlyOffice
## Что это
Полнотекстовый поиск OnlyOffice Workspace работает на **Elasticsearch**.
Клиент на сервере — NEST. Индекс — имя таблицы.
Для файлов индекс `files_file`:
| поле | тип | смысл |
|------|-----|-------|
| `id` | integer | id файла (тот же, что в REST/Documents) |
| `title` | text (`whitespacecustom`) | имя файла |
| `tenantId` | integer | тенант (портал) |
| `folders` | nested | список папок: `folderId` (строка), `id`, `tenantId` |
| `document.attachment.content` | text (`document`) | извлеченный текст (ingest-attachment) |
| `document.attachment.content_type` | text | MIME |
Важно:
- Живой сервер — **Elasticsearch 7.16.3**, кластер `elasticsearch`.
- REST `GET /api/2.0/files/@search/{query}` ищет **только по имени в БД**
(`fileDao.Search`), ES не задействует. Для поиска по содержимому нужен
прямой ES — это и делает `oo search`.
- `title` analyzer `whitespacecustom` режет по пробелам и lower-case. Полное
имя файла — один токен (`Rechnung-4711.pdf`), поэтому поиск по имени ищет
слово целиком, а не подстроку.
- `document.attachment.content` заполняется **только для Office-форматов**
(docx / xlsx / pptx). У PDF/txt, залитых через API, контент не извлекается.
- Индексация асинхронная (TeamLabSvc) — файл появляется в ES не мгновенно.
## Доступ
ES слушает `127.0.0.1:9200` **внутри** VM OnlyOffice. Снаружи порт закрыт,
SSH в VM открыт на хосте как `127.0.0.1:32` (контейнер `onlyoffice-v2`,
QEMU). Схема — SSH-туннель.
```bash
# из корня go-onlyoffice (ключ и хост — как в infra-доках)
ssh -f -N -o ControlMaster=no -o ControlPath=none \
-p 32 -i ~/.ssh/id_ed25519 \
-L 9200:127.0.0.1:9200 root@127.0.0.1
curl -s http://127.0.0.1:9200/ | head # tagline + version
curl -s 'http://127.0.0.1:9200/_cat/indices?h=index,docs.count'
```
`-o ControlMaster=no -o ControlPath=none` обязательны: иначе forward уходит
в persistent master-соединение из `~/.ssh/config` и порт остаётся занят.
Проверить, что туннель жив:
```bash
curl -s http://127.0.0.1:9200/files_file/_count
```
## Переменные
| env | default | смысл |
|-----|---------|-------|
| `ONLYOFFICE_ES_URL` | — (обязателен) | `scheme://host:port` ES |
| `ONLYOFFICE_ES_INDEX` | `files_file` | индекс |
| `ONLYOFFICE_TENANT` | пусто (все) | фильтр `tenantId` |
Имена — в [`.env.example`](../.env.example). Секретов нет: ES без пароля.
## CLI
```bash
ONLYOFFICE_ES_URL=http://127.0.0.1:9200 oo search "Rechnung"
ONLYOFFICE_ES_URL=http://127.0.0.1:9200 oo search "Mahngebühr" --content
oo search "Rechnung" --folder 649 --limit 50 --json
```
Флаги: `--content` (искать и по тексту), `--folder ID` (папка
`folders.folderId`), `--limit N` (по умолчанию 20, максимум 200),
`--json` = `-o json`.
## Обновление индекса и карта поиска
Обзор всех контуров поиска и как обновлять индексы (`oo index`) —
[index-and-search.md](index-and-search.md).
## Библиотека
`filestore_es.go` — `ESSearcher` (`Name() = "elasticsearch"`), прямой ES REST на
stdlib `net/http`:
```go
es, _ := onlyoffice.NewESSearcher(onlyoffice.ESConfigFromEnv())
hits, _ := es.Search(ctx, onlyoffice.SearchQuery{
Text: "Rechnung", InContent: true, Limit: 20,
})
```
Запрос: `multi_match` по `title^2` (+ `document.attachment.content` при
`InContent`), фильтры `tenantId` и `folders.folderId`, `_source`
id/title/folders, `highlight` для фрагмента. Ответ → `[]SearchHit` (модель из
эпика #34; пока объявлена в `filestore_es.go`, переедет в `filestore_core.go` с F1 #35).
## Тесты
```bash
# unit — чистые builders/парсеры, без сети
go test ./ -run ES
# integration — нужен ONLYOFFICE_ES_URL (+ креды REST для залива)
set -a; . .env; set +a
ONLYOFFICE_ES_URL=http://127.0.0.1:9200 ONLYOFFICE_TENANT=1 \
go test -tags=integration -run TestIntegrationESSearch -v .
```
Интеграционный тест заливает временный xlsx (в имени и в ячейке — уникальные
токены), ждёт индексации, проверяет поиск по имени и по содержимому, затем
удаляет проект.
## Грабли
- `locale`/версия ES: 7.16.3, `_search` совместим с REST 7.x.
- ES без auth и слушает только localhost — туннель обязателен.
- Фильтр `tenantId` сузит выдачу; без него видны документы всех тенантов.
- Поиск по содержимому PDF в индексе OnlyOffice не работает (для PDF нет
`attachment.content`) — только Office-форматы. Решение для PDF — свой индекс
`oo_docs_text` (F6 #42), см. ниже.
# PDF и сканы — свой индекс (F6 #42)
## Проблема
`oo search --content "S1019"` не находил номер внутри PDF-счёта: в индексе
OnlyOffice PDF лежит только по имени.
## Почему PDF исключён (исходники CommunityServer)
Разобрано в `ONLYOFFICE/CommunityServer`:
- `web/core/ASC.Web.Core/Files/FileUtility.cs` — `CanIndex(fileName)` читает
серверную настройку `files.index.formats` (в `web/studio/ASC.Web.Studio/web.appsettings.config`
значение по умолчанию `".pptx|.xlsx|.docx"`).
- `web/studio/ASC.Web.Studio/Products/Files/Core/Search/FilesWrapper.cs` —
`GetDocumentStream*` возвращает `null`, если `!FileUtility.CanIndex(Title)`,
файл зашифрован или больше `MaxFileSize`.
- `module/ASC.ElasticSearch/Core/WrapperWithDoc.cs` + mapping в `Wrapper.cs` —
маппинг `document.attachment.content` и ingest-pipeline `attachments`
формат-агностичны: они распарсят любой поток.
Вывод: PDF исключён **только настройкой** `files.index.formats`; жёсткого
ограничения на формат в коде нет.
## Варианты и решение
| # | Вариант | Оценка |
|---|---------|--------|
| a | Включить `.pdf` в `files.index.formats` + reindex | Правка сервера OO; настройка может потеряться при обновлении; полный reindex 39k док-в; Tika **не OCR** — сканы без текстового слоя дадут пустой контент. Отклонён без решения PO. |
| b | Server-side ingest/attachment для PDF | По факту то же, что (a): сервер кормит поток только для `CanIndex`. |
| c | **Свой индекс** `oo_docs_text`, наполняемый `internal/docpipe` | **Выбран.** Сервер OO не трогаем; детерминированно; работает OCR для сканов; независимо от обновлений OO; любые форматы; фильтры папка/тип. |
| d | Локальный поиск без индекса | Отклонён как основной: качаем и извлекаем на каждый запрос, нет выдачи/ранжирования/highlight. |
Итог: **вариант c**. Индекс OnlyOffice (`files_file`) не изменяется; наш
индекс живёт рядом.
## Устройство
- `filestore_es_text.go` — `ESTextIndex` (`Name() = "es-text"`):
`Ensure` (создаёт индекс с явным маппингом), `Put` (bulk, `refresh`),
`Delete` (по `id`), `Search` (`multi_match` по `title^2` + `content`,
фильтры `folder`/`ext`, highlight).
- `filestore_text_index.go` — `TextIndexer`: листает папки (`FileStore.List`),
качает файлы (`FileStore.Download`), извлекает текст через
`internal/docpipe` (`pdftotext`, для сканов — `ocrmypdf`/`tesseract`),
пишет в `TextIndex`. Пул воркеров (по умолчанию 3).
- CLI: `oo index folder|files` наполняет индекс; `oo search --backend own`
ищет по нему.
### Встроенные вложения PDF
Оцифрованные PDF несут вложения (`<doc>.md` — текст/таблицы скана,
`<doc>.yaml`/`.json` — метаданные, `.xml` — EN 16931 CII eRechnung,
`factur-x.xml` у ZUGFeRD; см. `office-assistant/docs/reference/document-metadata.md`).
`TextIndexer` обходит их: `pdfdetach -list` перечисляет, `-save` сохраняет,
каждое вложение проходит штатный `docpipe.ToMarkdown` (PDF/картинки → OCR,
`.md`/`.txt` — как есть). Форматы, которые docpipe не конвертирует
(`.xml`/`.html` — снимаются теги; `.json`/`.csv` — как текст), извлекаются
текстом; нечитаемые — пропускаются.
Текст склеивается: тело, затем по секции на вложение с маркером
`[attachment: <имя>]` (функция `docpipe.JoinWithAttachments`). Индекс — тот же
`file_id`, upsert идемпотентен. Нет вложений или pdfdetach/формат нечитаем —
индексируется тело (без падения).
Поля `oo_docs_text`:
| поле | тип | смысл |
|------|-----|-------|
| `id` | keyword | id файла Documents |
| `title` | text (+`.keyword`) | имя файла |
| `folder` | keyword | id папки |
| `ext` | keyword | расширение |
| `content` | text | извлечённый текст (pdftotext/OCR) |
## CLI
```bash
set -a; . .env; set +a # ONLYOFFICE_URL/USER/PASS + ONLYOFFICE_ES_URL
oo index folder 634 # PDF в папке 634
oo index folder 634 --recursive --exts pdf,png --limit 100
oo index files 3576 3578 # точечно
oo index folder 634 --dry-run # показать план, ничего не менять
oo search "S1021" --content --backend own
oo search "S1021" --backend own --folder 634 --json
```
`--backend` у `oo search`: `oo` (по умолчанию, индекс OnlyOffice) или `own`
(наш `ONLYOFFICE_ES_TEXT_INDEX`).
## Переменные (дополнение)
| env | default | смысл |
|-----|---------|-------|
| `ONLYOFFICE_ES_TEXT_INDEX` | `oo_docs_text` | индекс своего конвейера |
`ONLYOFFICE_ES_URL` — общий для обоих индексов.
## Тесты
```bash
go test -run 'ESText|TextIndexer|Index' ./ ./cmd/oo/ # unit, без сети
ONLYOFFICE_ES_URL=http://127.0.0.1:9200 \
go test -tags=integration -run TestIntegrationESTextIndex -v .
```
Интеграционный тест создаёт временный индекс, наполняет, ищет по контенту,
проверяет фильтры и удаление, затем удаляет индекс;
`TestIntegrationESTextIndexPDFAttachment` индексирует
`testdata/pdf-with-attachment.pdf` реальным конвейером (pdfdetach + pdftotext)
и ищет токен, лежащий только во вложении. Unit-тесты используют
fake-store/fake-extractor и не требуют pdftotext/OCR (парсер списка, склейка
`JoinWithAttachments`, снятие тегов `xmlToText` — чистые).
## Грабли
- Наполнение — ручное (`oo index`); после изменения/добавления PDF повтори.
Повтор идемпотентен (upsert по id файла).
- В индексе ищется только то, что проиндексировано; `oo index` качает каждый
файл и (для сканов) гоняет OCR — это медленно, отсюда `--limit`/`--exts`.
- `folder` фильтруется как id папки, а не как путь.
- Дубликаты (напр. `S1055.pdf` и `2026-08-20-S1055-…`) дадут несколько строк —
это ожидаемо, дедуп — на стороне потребителя.
- Вложения: нужен `pdfdetach` (poppler); если его нет — индексируется только
тело. Вложенный PDF/картинка с плохим текстовым слоем проходит OCR, это
медленно. `.json`-метаданные (CuraSoft) индексируются как текст и могут
добавить шумовых токенов.
+96
View File
@@ -0,0 +1,96 @@
---
type: reference
status: current
related:
- docs/elasticsearch.md
- docs/unified-file-client.md
- docs/community-server-db.md
---
# Поиск и индексация
Четыре разных контура поиска. Не путать: у каждого свой индекс, свои входы и
свой способ обновления.
| Контур | Что ищет | Индекс | Обновление | Вход |
|--------|----------|--------|------------|------|
| REST `@search` | только имена в БД | нет | — (живой запрос) | `oo search` (по умолчанию `--backend oo`) |
| ES `files_file` | имя + текст Office | Elasticsearch портала | сервер, асинхронно | `oo search --content` |
| ES `oo_docs_text` | PDF/сканы (свой) | Elasticsearch портала | `oo index` | `oo search --backend own` |
## Карта кода
- `filestore_core.go` — интерфейсы `Searcher`, модели `SearchQuery`/`SearchHit`.
- `filestore_es.go` — `ESSearcher` (индекс OnlyOffice `files_file`).
- `filestore_es_text.go` — `ESTextIndex` (`oo_docs_text`): `Ensure`, `Put`, `Delete`,
`Search`.
- `filestore_text_index.go` — `TextIndexer`: обход папок (`FileStore.List`), download
(`FileStore.Download`), извлечение текста (`internal/docpipe`), запись в
`TextIndex`; пул воркеров.
- `filestore_facade.go` — связка бэкендов (`Files().Search()`, порядок и fallback).
- CLI: `cmd/oo/search.go`, `cmd/oo/index.go`.
- Разовые бинари для match (`ooscan`, `pdfamount`) живут в приватном
`oo-workspace`.
- `internal/docpipe` — текст из PDF (pdftotext), для сканов OCR
(ocrmypdf/tesseract), вложения PDF (pdfdetach).
## Поиск
```bash
# имя, индекс портала
oo search "Rechnung" --limit 50 --json
# имя + текст Office (docx/xlsx/pptx)
oo search "Mahngebühr" --content
# свой индекс: PDF и сканы
oo search "S1019" --backend own --folder 649
```
Флаги `oo search`: `--content`, `--folder ID`, `--limit N`, `--backend oo|own`,
`--substring`, `--json`. Требует `ONLYOFFICE_ES_URL` (см.
[elasticsearch.md](elasticsearch.md)); `--backend own` дополнительно ничего не
требует от сервера — читает `oo_docs_text`.
Почему не REST: `GET /api/2.0/files/@search/{query}` ищет только имя в БД
(`fileDao.Search`), ES не трогает. Почему PDF не в `files_file`: сервер индексит
контент только для форматов из `files.index.formats` (по умолчанию
`.pptx|.xlsx|.docx`) — отсюда свой `oo_docs_text`.
## Обновление своего индекса (`oo index`)
```bash
# одна папка
oo index folder 649 --recursive --exts pdf
# точечно по id
oo index files 3576 3578
# без записи: что было бы проиндексировано
oo index folder 649 --recursive --dry-run
```
Флаги: `--recursive`, `--exts pdf` (по умолчанию), `--limit N`,
`--workers 3`, `--lang deu+eng`, `--min-chars N` (порог текстового слоя, ниже
которого включается OCR), `--work-dir`, `--backend rest|dav`, `--dry-run`,
`--json`.
Свойства:
- Идемпотентно: upsert по `id` файла; повтор не двоит.
- Сервер OnlyOffice не меняется: индекс живёт рядом (`ONLYOFFICE_ES_TEXT_INDEX`,
по умолчанию `oo_docs_text`).
- Медленно на сканах (OCR на каждый файл). Ограничивай `--folder`/`--limit`,
не индексируй корень целиком.
- Индексация PDF в `files_file` не делается — только `oo_docs_text`.
## Bulk-инструменты
Плоские TSV-инструменты (`ooscan`, `pdfamount`) и сверка Excel живут в
приватном `oo-workspace`, не в публичной библиотеке.
## Грабли
- ES слушает только `127.0.0.1:9200` внутри VM — SSH-туннель обязателен
(`-o ControlMaster=no -o ControlPath=none`, см. [elasticsearch.md](elasticsearch.md)).
- Портальные листинги/скачивание упираются в 429; все bulk-пути идут через
`DoRetry` (линейный бэкофф), `ooscan` дополнительно спит 350 мс на папку.
- `files_file` обновляется сервером асинхронно — свежий файл виден не сразу.
- `title` analyzer `whitespacecustom`: имя — один токен, подстрока только через
`--substring` (или wildcard).
+55
View File
@@ -0,0 +1,55 @@
---
type: reference
status: current
related:
- README.md
- ../AGENTS.md
---
# Rate limit, backoff и cooldown
Устойчивость к 429 (openresty). Всё встроено в библиотеку — отдельный пакет не
нужен. Реализация: `ratelimit.go`, `retry.go`.
## Что происходит с каждым запросом
1. **Cooldown-гейт** — общий на процесс. Если недавно пришёл 429, все запросы
ждут конца окна.
2. **Rate limiter** — token bucket на процесс. Пейсит все HTTP-пути: листинг,
создание папок, загрузку, `get project`, auth.
3. Запрос уходит.
4. Ответ ≥400 → `*TransientError` (для 429/502/503/504) с `Retry-After`.
5. `DoRetry` — экспоненциальный backoff, без jitter.
6. `Retry-After` длиннее backoff → ждём его; окно уходит в общий cooldown.
Установлено в `NewClient` через `pacedTransport`; отдельный код трогать не надо.
## Env
| Переменная | Default | Смысл |
|---|---|---|
| `OO_RATE_LIMIT` | `4` | запросов/с на процесс; `0` — лимитер выключен |
| `OO_BURST` | `1` | запас токенов token bucket |
| `OO_RETRY_ATTEMPTS` | `7` | всего попыток, включая первую |
| `OO_RETRY_BASE` | `2s` | база экспоненты: ждать перед попыткой N = `Base*2^(N-1)` |
| `OO_RETRY_MAX` | `2m` | потолок ожидания |
Битые значения → default. `OO_RETRY_*` — формат `time.ParseDuration`
(`2s`, `30s`, `2m`).
## Правила
- Детерминированно, без jitter — повторный прогон ждёт столько же.
- `Retry-After` — секунды (`120`) или HTTP-date.
- Cooldown общий: параллельные и последовательные вызовы не бьют в стену.
- Backoff cap не ограничивает `Retry-After` — серверу верим больше.
- Только stdlib.
## Когда руками снять нагрузку
`OO_RATE_LIMIT` ниже (`2`), `OO_BURST=1`; при массовом apply — батчами.
## Тесты
`retry_test.go` — `Retry-After`, экспонента, cap; `ratelimit_test.go` — burst,
`OO_RATE_LIMIT=0`, cooldown. Фейковый сервер отдаёт 429 с заголовком.
+204
View File
@@ -0,0 +1,204 @@
---
type: reference
status: current
related:
- README.md
- filestore_core.go
- filestore_facade.go
- docs/elasticsearch.md
- docs/community-server-db.md
---
# Unified file client — контракт файловых бэкендов
## Что это
Один файловый клиент на все бэкенды (эпик #34). Модель и интерфейсы —
`filestore_core.go`. Фасад `FileClient` — `filestore_facade.go`. Бэкенды:
REST, WebDAV, SQL (PostgreSQL/MySQL), Elasticsearch. Правило одно:
код зовёт `c.Files()` и не знает про транспорт.
## Модель
- `Kind` — `File` (0) или `Folder` (1).
- `Entry` — бэкенд-независимая строка: `ID`, `ParentID`, `Title`, `Kind`,
`Size`, `MIME`, `Created`, `Modified`, `Updated` (сырая строка API),
`Version`, `Provider`, `FilesCount`/`FoldersCount` (папки).
Чего бэкенд не даёт — остаётся в нуле.
- `SearchQuery` — `Text`, `InContent`, `FolderID`, `Extensions`, `Limit`.
- `SearchHit` — `Entry` + `Score`, `Highlight`, `Path`.
## Интерфейсы
`FileStore` — операции с файлами:
```go
type FileStore interface {
Name() string
List(ctx, parentID) ([]Entry, error)
Stat(ctx, id) (Entry, error)
CreateFolder(ctx, parentID, title) (Entry, error)
Upload(ctx, parentID, title, r) (Entry, error)
Download(ctx, id, w) (int64, error)
Move(ctx, ids, parentID) error
Copy(ctx, ids, parentID) error
Rename(ctx, id, title) error
Delete(ctx, ids) error
}
```
`Searcher` — поиск (необязательный):
```go
type Searcher interface {
Search(ctx, q SearchQuery) ([]SearchHit, error)
Name() string
}
```
`TextIndex` (`filestore_es_text.go`) — свой индекс: `Put`, `Delete`, `Search`,
`Name`. `ESTextIndex` реализует и `Searcher`, и `TextIndex`.
## Бэкенды
| бэкенд | провайдер | файл | что умеет |
|--------|-----------|------|-----------|
| REST | `rest` | `filestore_rest.go` | read + write, Documents API |
| WebDAV | `dav` | `filestore_dav.go` | read + write, Documents fileops |
| SQL | `postgres` / `mysql` | `filestore_pg.go` | **read-only** |
| OnlyOffice ES | `elasticsearch` | `filestore_es.go` | поиск (имя + контент Office) |
| свой ES-индекс | `es-text` | `filestore_es_text.go` | поиск + запись (PDF/сканы) |
- REST: `Stat` знает только файлы; папки — через `List`.
- WebDAV: `Move`/`Copy`/`Delete` сперва `Stat`-ят id (папка/файл), потом зовут
fileops.
- SQL: `List`/`Stat`/`Download`/`Search` (по имени). Все write-методы →
`ErrReadOnly`. `Download` идёт в S3/MinIO по layout портала.
- OnlyOffice ES: индекс `files_file`, контент только для docx/xlsx/pptx.
- Свой ES: индекс `oo_docs_text`, контент из `internal/docpipe`, в т.ч.
встроенные PDF-вложения.
## Фасад `FileClient`
`c.Files()` → `*FileClient`. Он же реализует `FileStore`, старый код
компилируется.
- `Read()` — первый зарегистрированный из `readOrder`:
`postgres` → `mysql` → `rest` → `dav`.
- `Write()` — первый из `writeOrder`: `rest` → `dav`. SQL не пишет.
- `Search()` — первый из `searchOrder`: `elasticsearch`. Нет бэкенда →
ошибка (`ONLYOFFICE_ES_URL`).
- `RegisterStore(name, s)` / `RegisterSearcher(name, s)` — добавить бэкенд.
Fallback:
- `List`/`Stat` идут по `readOrder`; переходят к следующему только на
transient-ошибке (429/502/503/504). Иначе ошибка финальная.
- `Download` **без** fallback: часть байтов уже в `w`, второй бэкенд допишет.
- Запись (`CreateFolder`/`Upload`/`Move`/`Copy`/`Rename`/`Delete`) — только
`Write()`, без fallback.
`newFileClient` сам кладёт `rest` и `dav`; ES-поиск — если задан
`ONLYOFFICE_ES_URL`. SQL-стор регистрирует вызывающий: фасад создаётся на
каждый `c.Files()`, регистрируй на том же экземпляре.
```go
sql, err := c.SQLFileStore() // открыть из env (ONLYOFFICE_DSN)
if err != nil { /* нет DSN */ }
if closer, ok := sql.(interface{ Close() error }); ok { defer closer.Close() }
f := c.Files()
f.RegisterStore(onlyoffice.ProviderPG, sql) // или sql.Name() == "mysql"
e, _ := f.Stat(ctx, "19423") // e.Provider == "mysql"
entries, _ := f.List(ctx, "676") // пойдёт в SQL
```
`Client.FileStore("pg"|"sql"|"postgres"|"mysql")` — одноразовый доступ к
SQL-стору без фасада: открывает из env; при ошибке возвращает заглушку,
которая отдаёт ошибку открытия на каждом вызове (не `nil`). `SQLFileStore()`
— тот же открыватель, но с ошибкой. Отвечавший бэкенд видно по
`Entry.Provider` (`mysql` / `postgres` у SQL, `rest` у REST).
## CLI
```bash
# поиск: --backend oo (индекс OnlyOffice) | own (свой oo_docs_text)
oo search "Rechnung"
oo search "Mahngebühr" --content
oo search "S1021" --content --backend own --folder 634 --limit 50 --json
# наполнение своего индекса (PDF/сканы, idempotent upsert по file id)
oo index folder 634
oo index folder 634 --recursive --exts pdf,png --limit 100
oo index files 3576 3578
oo index folder 634 --dry-run
```
`oo index` флаги: `--recursive`, `--exts` (default `pdf`), `--limit`,
`--workers` (3), `--lang` (`deu+eng`), `--min-chars`, `--work-dir`,
`--backend rest|dav`, `--dry-run`, `--json`.
Библиотека:
```go
idx, _ := onlyoffice.NewESTextIndex(onlyoffice.ESTextConfigFromEnv())
ti := onlyoffice.NewTextIndexer(store, idx) // store = FileStore
res, _ := ti.IndexFolder(ctx, "634", onlyoffice.IndexOptions{Recursive: true})
```
## Env (только имена)
| env | default | кто читает |
|-----|---------|------------|
| `ONLYOFFICE_ES_URL` | — | ES (оба индекса), обязателен |
| `ONLYOFFICE_ES_INDEX` | `files_file` | индекс OnlyOffice |
| `ONLYOFFICE_ES_TEXT_INDEX` | `oo_docs_text` | свой индекс |
| `ONLYOFFICE_TENANT` | пусто | фильтр `tenantId` |
| `ONLYOFFICE_DSN` | — | SQL DSN (MySQL/PostgreSQL) |
| `ONLYOFFICE_PG_DRIVER` | auto | `postgres` / `mysql` |
| `ONLYOFFICE_PG_TENANT` | `ONLYOFFICE_TENANT` | SQL tenant |
| `ONLYOFFICE_PG_HOST` `_PORT` `_USER` `_PASSWORD` `_DBNAME` `_SSLMODE` | — | DSN по частям |
| `MINIO_ENDPOINT` `MINIO_BUCKET` `MINIO_ACCESS_KEY` `MINIO_SECRET_KEY` | — | download SQL-стора |
| `OO_URL` `OO_USER` `OO_PASS` | — | CLI-алиасы |
Имена — в [`.env.example`](../.env.example). Секретов в репо нет.
## Ограничения
- OnlyOffice ES: контент только Office-форматов. PDF — только по имени.
Встроенные вложения PDF сервер не индексирует.
- Свой индекс `oo_docs_text`: покрывает PDF/сканы и вложения (pdfdetach), но
наполняется вручную (`oo index`) и идемпотентен. Фильтр `folder` — id папки,
не путь. Дубли дают несколько строк — дедуп на потребителе.
- SQL: read-only. `InContent` игнорируется (только имя). Download — через
MinIO-схему, не HTTP.
- ES: без auth, слушает localhost внутри VM — нужен SSH-туннель
(см. [elasticsearch.md](elasticsearch.md)).
- `oo index` качает каждый файл и для сканов гоняет OCR — медленно; отсюда
`--limit` и `--exts`. Нужен `pdfdetach` (poppler); без него — только тело PDF.
## Как добавить бэкенд
1. Файл `file_<name>.go`. Реализуй `FileStore` (`Name` + 9 методов). Нужен
поиск — добавь `Searcher`; нужна запись своего индекса — `TextIndex`.
2. Добавь const провайдера рядом с `ProviderREST`/`ProviderDAV`.
3. Зарегистрируй: в `newFileClient` или снаружи через
`RegisterStore`/`RegisterSearcher`.
4. Внеси имя в `readOrder` / `writeOrder` / `searchOrder`.
5. Есть CLI-команда — добавь значение в `--backend`.
6. Тесты: unit (чистые builders/парсеры, без сети) + интеграционный
(`//go:build integration`, skip без кред).
## Тесты
```bash
go test ./... # unit, без сети
go test -tags=integration ./... # live (креды в .env)
go test ./ -run 'FileStore|Facade|ESText|PG'
```
## См. также
- [README.md](README.md) — индекс справочников.
- [elasticsearch.md](elasticsearch.md) — индекс OnlyOffice и свой `oo_docs_text`.
- [community-server-db.md](community-server-db.md) — SQL-стор и схема БД.
+144 -35
View File
@@ -237,7 +237,19 @@ func (c *Client) UploadProjectFile(ctx context.Context, projectID, localPath str
return decodeResponseFileEntry(raw) return decodeResponseFileEntry(raw)
} }
// UploadProjectFileReplacing upserts by stem and (server-converted) extension
// in the project Documents folder.
func (c *Client) UploadProjectFileReplacing(ctx context.Context, projectID, localPath string) (*FileEntry, []int, error) {
folderID, err := c.projectFolderID(ctx, projectID)
if err != nil {
return nil, nil, err
}
return c.UploadToFolderReplacing(ctx, folderID, localPath)
}
// GetFile returns file metadata including viewUrl for download. // GetFile returns file metadata including viewUrl for download.
//
// Deprecated: use FileStore.Stat via Client.Files()/Client.FileStore.
func (c *Client) GetFile(ctx context.Context, fileID string) (*FileEntry, error) { func (c *Client) GetFile(ctx context.Context, fileID string) (*FileEntry, error) {
if fileID == "" { if fileID == "" {
return nil, fmt.Errorf("file id is required") return nil, fmt.Errorf("file id is required")
@@ -251,6 +263,8 @@ func (c *Client) GetFile(ctx context.Context, fileID string) (*FileEntry, error)
} }
// RenameFile sets a new title (including extension) for the file. // RenameFile sets a new title (including extension) for the file.
//
// Deprecated: use FileStore.Rename via Client.Files()/Client.FileStore.
func (c *Client) RenameFile(ctx context.Context, fileID, newTitle string) (*FileEntry, error) { func (c *Client) RenameFile(ctx context.Context, fileID, newTitle string) (*FileEntry, error) {
if fileID == "" || newTitle == "" { if fileID == "" || newTitle == "" {
return nil, fmt.Errorf("file id and new title are required") return nil, fmt.Errorf("file id and new title are required")
@@ -263,55 +277,150 @@ func (c *Client) RenameFile(ctx context.Context, fileID, newTitle string) (*File
return decodeResponseFileEntry(raw) return decodeResponseFileEntry(raw)
} }
type deleteFilesBody struct {
FileIDs []int `json:"fileIds"`
FolderIDs []int `json:"folderIds"`
}
// DeleteFiles permanently deletes files by numeric id (Documents module). // DeleteFiles permanently deletes files by numeric id (Documents module).
// Uses per-file DELETE (DeleteDavItems); fileops/delete returns 200 on some
// portals without actually removing the file.
//
// Deprecated: use FileStore.Delete via Client.Files()/Client.FileStore.
func (c *Client) DeleteFiles(ctx context.Context, fileIDs []int) error { func (c *Client) DeleteFiles(ctx context.Context, fileIDs []int) error {
if len(fileIDs) == 0 { if len(fileIDs) == 0 {
return fmt.Errorf("no file ids to delete") return fmt.Errorf("no file ids to delete")
} }
body := deleteFilesBody{FileIDs: fileIDs, FolderIDs: nil} strIDs := make([]string, len(fileIDs))
_, err := c.putJSON(ctx, "/api/2.0/files/fileops/delete.json", body) for i, id := range fileIDs {
if err != nil { strIDs[i] = strconv.Itoa(id)
_, err = c.putJSON(ctx, "/api/2.0/files/fileops/delete", body)
} }
return err return c.DeleteDavItems(ctx, nil, strIDs)
}
// ListFolder returns the Documents module listing for a folder id
// (GET /api/2.0/files/{folderId}).
//
// Deprecated: use FileStore.List via Client.Files()/Client.FileStore.
func (c *Client) ListFolder(ctx context.Context, folderID string) (map[string]any, error) {
if folderID == "" {
return nil, fmt.Errorf("folder id is required")
}
out, err := c.ResponseObject(ctx, "/api/2.0/files/"+url.PathEscape(folderID)+".json")
if err != nil {
out, err = c.ResponseObject(ctx, "/api/2.0/files/"+url.PathEscape(folderID))
}
return out, err
}
// CreateFolder creates a subfolder under parentFolderID.
func (c *Client) CreateFolder(ctx context.Context, parentFolderID, title string) (map[string]any, error) {
if parentFolderID == "" || title == "" {
return nil, fmt.Errorf("parent folder id and title are required")
}
body := map[string]any{"title": title}
out, err := c.postJSONObject(ctx, "/api/2.0/files/folder/"+url.PathEscape(parentFolderID)+".json", body)
if err != nil {
out, err = c.postJSONObject(ctx, "/api/2.0/files/folder/"+url.PathEscape(parentFolderID), body)
}
return out, err
}
// MoveFiles moves file ids into destFolderID (Documents fileops/move).
//
// Deprecated: use FileStore.Move via Client.Files()/Client.FileStore.
func (c *Client) MoveFiles(ctx context.Context, destFolderID int, fileIDs []int) (map[string]any, error) {
if destFolderID == 0 || len(fileIDs) == 0 {
return nil, fmt.Errorf("dest folder and file ids are required")
}
body := map[string]any{
"folderIds": []int{},
"fileIds": fileIDs,
"destFolderId": destFolderID,
"resolveType": "Skip",
"holdResult": true,
}
// fileops/move answers an operations envelope (like MoveDavItems), not a
// single object, so parse the raw body before unwrapping and surface any
// per-operation error. Unwrapping first (putJSONObject) made fileopsError
// look for a "response" key that was already stripped.
raw, err := c.putJSON(ctx, "/api/2.0/files/fileops/move", body)
if err != nil {
raw, err = c.putJSON(ctx, "/api/2.0/files/fileops/move.json", body)
if err != nil {
return nil, err
}
}
if ferr := fileopsError(raw); ferr != nil {
return nil, ferr
}
out, _ := unmarshalResponseObject(raw)
return out, nil
}
// UploadToFolder uploads a local file into an arbitrary Documents folder id.
//
// Deprecated: use FileStore.Upload via Client.Files()/Client.FileStore.
func (c *Client) UploadToFolder(ctx context.Context, folderID, localPath string) (*FileEntry, error) {
if folderID == "" || localPath == "" {
return nil, fmt.Errorf("folder id and local path are required")
}
uploadPath := fmt.Sprintf("/api/2.0/files/%s/upload.json", url.PathEscape(folderID))
raw, err := c.uploadMultipart(ctx, uploadPath, "file", localPath)
if err != nil {
uploadPath = fmt.Sprintf("/api/2.0/files/%s/upload", url.PathEscape(folderID))
raw, err = c.uploadMultipart(ctx, uploadPath, "file", localPath)
if err != nil {
return nil, err
}
}
return decodeResponseFileEntry(raw)
}
// UpdateFile uploads a new version of an existing file (same id, name and
// folder). It does not delete and does not create a second file.
//
// The Documents API method is PUT /api/2.0/files/{id}/update; POST is kept as
// a fallback for older servers. The path is tried with and without .json.
func (c *Client) UpdateFile(ctx context.Context, fileID, localPath string) (*FileEntry, error) {
if fileID == "" || localPath == "" {
return nil, fmt.Errorf("file id and local path are required")
}
base := fmt.Sprintf("/api/2.0/files/%s/update", url.PathEscape(fileID))
attempts := []struct {
method, path string
}{
{http.MethodPut, base},
{http.MethodPut, base + ".json"},
{http.MethodPost, base},
{http.MethodPost, base + ".json"},
}
var lastErr error
for _, a := range attempts {
raw, err := c.uploadMultipartMethod(ctx, a.method, a.path, "file", localPath)
if err == nil {
return decodeResponseFileEntry(raw)
}
lastErr = err
}
return nil, lastErr
}
// FileFolderID returns the parent folder id string for a file entry, if known.
func FileFolderID(f *FileEntry) string {
if f == nil || f.FolderID == nil {
return ""
}
return f.FolderID.String()
} }
// DownloadFile streams file bytes from the file's viewUrl using the same auth // DownloadFile streams file bytes from the file's viewUrl using the same auth
// as API calls. Writes into dst. // as API calls. Writes into dst. When the portal serves the file from its stale
// AWS S3 consumer, the bytes are fetched from the local MinIO store instead
// (see storage_fallback.go).
//
// Deprecated: use FileStore.Download via Client.Files()/Client.FileStore.
func (c *Client) DownloadFile(ctx context.Context, fileID string, dst io.Writer) (int64, error) { func (c *Client) DownloadFile(ctx context.Context, fileID string, dst io.Writer) (int64, error) {
f, err := c.GetFile(ctx, fileID) f, err := c.GetFile(ctx, fileID)
if err != nil { if err != nil {
return 0, err return 0, err
} }
if f.ViewURL == nil || *f.ViewURL == "" { return c.downloadFileEntry(ctx, f, dst)
return 0, fmt.Errorf("file %s has no viewUrl", fileID)
}
downloadURL := c.resolveAPIURL(*f.ViewURL)
auth, err := c.authHeader()
if err != nil {
return 0, err
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, downloadURL, nil)
if err != nil {
return 0, err
}
req.Header.Set("Authorization", auth)
resp, err := c.client.Do(req)
if err != nil {
return 0, err
}
defer resp.Body.Close()
if resp.StatusCode >= 400 {
b, _ := io.ReadAll(io.LimitReader(resp.Body, 512))
return 0, fmt.Errorf("GET viewUrl: %d %s", resp.StatusCode, truncate(string(b), 400))
}
n, err := io.Copy(dst, resp.Body)
return n, err
} }
func (c *Client) resolveAPIURL(ref string) string { func (c *Client) resolveAPIURL(ref string) string {
+468
View File
@@ -0,0 +1,468 @@
package onlyoffice
import (
"context"
"encoding/json"
"path/filepath"
"sort"
"strings"
)
// FileEntryExt returns a normalized extension (lowercase, with leading dot).
func FileEntryExt(f *FileEntry) string {
if f == nil {
return ""
}
exst := ""
if f.FileExst != nil {
exst = strings.TrimSpace(*f.FileExst)
}
if exst != "" {
if !strings.HasPrefix(exst, ".") {
exst = "." + exst
}
return strings.ToLower(exst)
}
if f.Title != nil {
if ext := filepath.Ext(*f.Title); ext != "" {
return strings.ToLower(ext)
}
}
return ""
}
// FileDedupKey is stem|ext — two files with the same key are duplicates.
func FileDedupKey(f *FileEntry) string {
st := FileEntryStem(f)
ext := FileEntryExt(f)
if st == "" {
return ""
}
if ext == "" {
return st
}
return st + "|" + strings.TrimPrefix(ext, ".")
}
// FindFilesByDedupKey returns folder files matching stem and extension.
func FindFilesByDedupKey(files []*FileEntry, stem, ext string) []*FileEntry {
key := dedupKeyFromParts(stem, ext)
if key == "" {
return nil
}
var out []*FileEntry
for _, f := range files {
if FileDedupKey(f) == key {
out = append(out, f)
}
}
return out
}
func dedupKeyFromParts(stem, ext string) string {
stem = strings.TrimSpace(stem)
if stem == "" {
return ""
}
ext = strings.ToLower(strings.TrimSpace(ext))
if ext != "" && !strings.HasPrefix(ext, ".") {
ext = "." + ext
}
if ext == "" {
return stem
}
return stem + "|" + strings.TrimPrefix(ext, ".")
}
// UploadExtFromLocal returns the lowercase extension from a local path.
func UploadExtFromLocal(localPath string) string {
ext := filepath.Ext(localPath)
if ext == "" {
return ""
}
return strings.ToLower(ext)
}
// legacyToOOXMLExt maps the legacy binary Office extensions OnlyOffice accepts
// on upload to the OOXML extension the server converts them into.
var legacyToOOXMLExt = map[string]string{
".xls": ".xlsx",
".doc": ".docx",
".ppt": ".pptx",
}
// normalizeExt lowercases an extension and ensures a leading dot.
func normalizeExt(ext string) string {
ext = strings.ToLower(strings.TrimSpace(ext))
if ext == "" {
return ""
}
if !strings.HasPrefix(ext, ".") {
ext = "." + ext
}
return ext
}
// EquivalentUploadExt reports whether two extensions designate the same
// document once the server-side conversion is taken into account: equal
// extensions, or a legacy binary Office format and its OOXML equivalent
// (.xls/.xlsx, .doc/.docx, .ppt/.pptx). Empty extensions only match each other.
func EquivalentUploadExt(a, b string) bool {
na, nb := normalizeExt(a), normalizeExt(b)
if na == nb {
return true
}
if na == "" || nb == "" {
return false
}
return legacyToOOXMLExt[na] == nb || legacyToOOXMLExt[nb] == na
}
// FindFilesByStemExt returns folder files matching stem and a same-or-converted
// extension (see EquivalentUploadExt). Unlike FindFilesByDedupKey, foo.xls and
// foo.xlsx are one logical file, so a replacing upload after OnlyOffice's
// legacy→OOXML conversion finds the saved file instead of appending a copy.
func FindFilesByStemExt(files []*FileEntry, stem, ext string) []*FileEntry {
stem = strings.TrimSpace(stem)
if stem == "" {
return nil
}
var out []*FileEntry
for _, f := range files {
if FileEntryStem(f) != stem {
continue
}
if EquivalentUploadExt(FileEntryExt(f), ext) {
out = append(out, f)
}
}
return out
}
// DeleteFilesByStemExt removes every folder file matching stem and a
// same-or-converted extension (legacy ↔ OOXML).
func (c *Client) DeleteFilesByStemExt(ctx context.Context, folderID, stem, ext string) ([]int, error) {
files, err := c.FolderFiles(ctx, folderID)
if err != nil {
return nil, err
}
matches := FindFilesByStemExt(files, stem, ext)
ids := make([]int, 0, len(matches))
for _, f := range matches {
if n := int(FileEntryNumericID(f)); n != 0 {
ids = append(ids, n)
}
}
if len(ids) == 0 {
return nil, nil
}
if err := c.DeleteFiles(ctx, ids); err != nil {
return nil, err
}
return ids, nil
}
// IsTrashFolderTitle reports staging/trash folders (e.g. _trash-md).
func IsTrashFolderTitle(title string) bool {
t := strings.ToLower(strings.TrimSpace(title))
return strings.HasPrefix(t, "_") || strings.Contains(t, "trash")
}
// ProjectFolderFile ties a file to its project Documents subfolder.
type ProjectFolderFile struct {
FolderID string
FolderTitle string
File *FileEntry
}
// DedupGroup is one duplicate set: keep the newest (or non-trash) file.
type DedupGroup struct {
Key string
FolderID string
FolderTitle string
Keep *FileEntry
Remove []*FileEntry
}
// DedupOptions controls project-wide duplicate scans.
type DedupOptions struct {
CrossFolder bool
}
// FindProjectDuplicates scans project folders for duplicate files.
func FindProjectDuplicates(folders []*FolderEntry, filesByFolder map[string][]*FileEntry, opts DedupOptions) []DedupGroup {
var indexed []ProjectFolderFile
for _, folder := range folders {
if folder == nil || folder.ID == nil {
continue
}
fid := folder.ID.String()
title := ""
if folder.Title != nil {
title = *folder.Title
}
for _, f := range filesByFolder[fid] {
if f == nil {
continue
}
indexed = append(indexed, ProjectFolderFile{
FolderID: fid, FolderTitle: title, File: f,
})
}
}
if opts.CrossFolder {
return findCrossFolderDuplicates(indexed)
}
return findWithinFolderDuplicates(indexed)
}
func findWithinFolderDuplicates(indexed []ProjectFolderFile) []DedupGroup {
byFolder := map[string][]ProjectFolderFile{}
for _, it := range indexed {
byFolder[it.FolderID] = append(byFolder[it.FolderID], it)
}
var out []DedupGroup
for fid, items := range byFolder {
title := ""
if len(items) > 0 {
title = items[0].FolderTitle
}
byKey := map[string][]*FileEntry{}
for _, it := range items {
k := FileDedupKey(it.File)
byKey[k] = append(byKey[k], it.File)
}
for k, group := range byKey {
if k == "" {
continue // dotfiles etc. have no stem: never treat as duplicates
}
if len(group) < 2 {
continue
}
keep, remove := pickDuplicateKeeper(group, false)
if keep == nil || len(remove) == 0 {
continue
}
out = append(out, DedupGroup{
Key: k, FolderID: fid, FolderTitle: title, Keep: keep, Remove: remove,
})
}
}
sortDedupGroups(out)
return out
}
func findCrossFolderDuplicates(indexed []ProjectFolderFile) []DedupGroup {
byKey := map[string][]ProjectFolderFile{}
for _, it := range indexed {
k := FileDedupKey(it.File)
byKey[k] = append(byKey[k], it)
}
var out []DedupGroup
for k, items := range byKey {
if k == "" {
continue // dotfiles etc. have no stem: never treat as duplicates
}
if len(items) < 2 {
continue
}
files := make([]*FileEntry, len(items))
folders := make([]string, len(items))
folderTitles := make([]string, len(items))
for i, it := range items {
files[i] = it.File
folders[i] = it.FolderID
folderTitles[i] = it.FolderTitle
}
keep, remove := pickDuplicateKeeperWithFolders(files, folders, folderTitles)
if keep == nil || len(remove) == 0 {
continue
}
fid, ftitle := "", ""
for _, it := range items {
if it.File == keep {
fid, ftitle = it.FolderID, it.FolderTitle
break
}
}
out = append(out, DedupGroup{
Key: k, FolderID: fid, FolderTitle: ftitle, Keep: keep, Remove: remove,
})
}
sortDedupGroups(out)
return out
}
func pickDuplicateKeeper(files []*FileEntry, _ bool) (*FileEntry, []*FileEntry) {
return pickDuplicateKeeperWithFolders(files, nil, nil)
}
func pickDuplicateKeeperWithFolders(files []*FileEntry, folderIDs, folderTitles []string) (*FileEntry, []*FileEntry) {
if len(files) == 0 {
return nil, nil
}
type ranked struct {
file *FileEntry
trash bool
}
rankedFiles := make([]ranked, len(files))
for i, f := range files {
trash := false
if folderTitles != nil && i < len(folderTitles) {
trash = IsTrashFolderTitle(folderTitles[i])
}
rankedFiles[i] = ranked{file: f, trash: trash}
}
sort.SliceStable(rankedFiles, func(i, j int) bool {
ri, rj := rankedFiles[i], rankedFiles[j]
if ri.trash != rj.trash {
return !ri.trash // non-trash first
}
ti, tj := rankedFiles[i].file.Updated, rankedFiles[j].file.Updated
if ti == nil {
return false
}
if tj == nil {
return true
}
return ti.After(*tj) // newest first
})
keep := rankedFiles[0].file
var remove []*FileEntry
for _, r := range rankedFiles[1:] {
remove = append(remove, r.file)
}
return keep, remove
}
func sortDedupGroups(groups []DedupGroup) {
sort.Slice(groups, func(i, j int) bool {
if groups[i].FolderTitle != groups[j].FolderTitle {
return groups[i].FolderTitle < groups[j].FolderTitle
}
return groups[i].Key < groups[j].Key
})
}
// ApplyDedupGroups deletes Remove files from each group.
func (c *Client) ApplyDedupGroups(ctx context.Context, groups []DedupGroup) ([]int, error) {
seen := map[int]struct{}{}
var ids []int
for _, g := range groups {
for _, f := range g.Remove {
n := int(FileEntryNumericID(f))
if n == 0 {
continue
}
if _, ok := seen[n]; ok {
continue
}
seen[n] = struct{}{}
ids = append(ids, n)
}
}
if len(ids) == 0 {
return nil, nil
}
if err := c.DeleteFiles(ctx, ids); err != nil {
return ids, err
}
return ids, nil
}
// DeleteFilesByDedupKey removes all files in folderID matching stem+ext.
func (c *Client) DeleteFilesByDedupKey(ctx context.Context, folderID, stem, ext string) ([]int, error) {
files, err := c.FolderFiles(ctx, folderID)
if err != nil {
return nil, err
}
matches := FindFilesByDedupKey(files, stem, ext)
if len(matches) == 0 {
return nil, nil
}
ids := make([]int, 0, len(matches))
for _, f := range matches {
n := int(FileEntryNumericID(f))
if n != 0 {
ids = append(ids, n)
}
}
if len(ids) == 0 {
return nil, nil
}
if err := c.DeleteFiles(ctx, ids); err != nil {
return nil, err
}
return ids, nil
}
// mergeProjectRootForDedupe includes projectFolder files in dedupe scans. OO often lists
// root documents only in pf.Files while pf.Folders is empty.
func mergeProjectRootForDedupe(rootID string, folders []*FolderEntry, filesByFolder map[string][]*FileEntry, rootFiles []*FileEntry) ([]*FolderEntry, map[string][]*FileEntry) {
if rootID == "" {
return folders, filesByFolder
}
if filesByFolder == nil {
filesByFolder = map[string][]*FileEntry{}
}
for _, folder := range folders {
if folder != nil && folder.ID != nil && folder.ID.String() == rootID {
if len(rootFiles) > 0 {
filesByFolder[rootID] = rootFiles
}
return folders, filesByFolder
}
}
if len(rootFiles) == 0 {
return folders, filesByFolder
}
id := json.Number(rootID)
title := "(project root)"
folders = append(folders, &FolderEntry{ID: &id, Title: &title})
filesByFolder[rootID] = rootFiles
return folders, filesByFolder
}
// DedupeProject scans project folders and optionally deletes duplicates.
func (c *Client) DedupeProject(ctx context.Context, projectID string, opts DedupOptions, apply bool) ([]DedupGroup, []int, error) {
pf, err := c.GetProjectFiles(ctx, projectID)
if err != nil {
return nil, nil, err
}
rootID, err := c.projectFolderID(ctx, projectID)
if err != nil {
return nil, nil, err
}
var rootFiles []*FileEntry
if rootID != "" {
rootFiles, err = c.FolderFiles(ctx, rootID)
if err != nil {
return nil, nil, err
}
}
filesByFolder := make(map[string][]*FileEntry, len(pf.Folders)+1)
folders := make([]*FolderEntry, 0, len(pf.Folders)+1)
for _, folder := range pf.Folders {
if folder == nil || folder.ID == nil {
continue
}
fid := folder.ID.String()
if fid == rootID {
filesByFolder[fid] = rootFiles
} else {
files, err := c.FolderFiles(ctx, fid)
if err != nil {
return nil, nil, err
}
filesByFolder[fid] = files
}
folders = append(folders, folder)
}
folders, filesByFolder = mergeProjectRootForDedupe(rootID, folders, filesByFolder, rootFiles)
groups := FindProjectDuplicates(folders, filesByFolder, opts)
if !apply || len(groups) == 0 {
return groups, nil, nil
}
deleted, err := c.ApplyDedupGroups(ctx, groups)
return groups, deleted, err
}
+416
View File
@@ -0,0 +1,416 @@
//go:build integration
package onlyoffice
import (
"context"
"encoding/json"
"os"
"path/filepath"
"strconv"
"testing"
"time"
)
// TestIntegrationFileDedup proves on a live OnlyOffice portal that the file
// dedup helpers find real duplicates and delete only the redundant copies.
// It creates a throwaway "go-onlyoffice-test-" project (removed by cleanup),
// places same stem|ext files in two subfolders and in the project root, then
// exercises FindProjectDuplicates, mergeProjectRootForDedupe,
// ApplyDedupGroups and DeleteFilesByDedupKey. Destructive — run only against
// an instance you own.
func TestIntegrationFileDedup(t *testing.T) {
c := liveClient(t)
t.Cleanup(func() { cleanupTestProjects(t, c) })
ctx := context.Background()
suffix := time.Now().UTC().Format("20060102-150405")
project, err := c.CreateProject(NewProjectRequest{
Title: testProjectPrefix + "dedup-" + suffix,
Description: "go-onlyoffice file dedup integration",
})
if err != nil {
t.Fatalf("CreateProject: %v", err)
}
if project.ID == nil {
t.Fatal("created project without id")
}
pid := strconv.Itoa(*project.ID)
root := projectFolderEventually(t, ctx, c, pid)
aID := createFolderLive(t, ctx, c, root, "A-"+suffix)
bID := createFolderLive(t, ctx, c, root, "B-"+suffix)
createFolderLive(t, ctx, c, root, "_trash-"+suffix)
// Check IsTrashFolderTitle against a real live folder title.
if !IsTrashFolderTitle("_trash-" + suffix) {
t.Fatalf("IsTrashFolderTitle(%q) = false for a live _trash folder", "_trash-"+suffix)
}
stem := "dedup-" + suffix
local := writeLocalFile(t, stem+".txt", []byte("dedup integration "+suffix+"\n"))
a1 := uploadFolderLive(t, ctx, c, aID, local)
a2 := uploadFolderLive(t, ctx, c, aID, local)
b1 := uploadFolderLive(t, ctx, c, bID, local)
r1 := uploadFolderLive(t, ctx, c, root, local)
r2 := uploadFolderLive(t, ctx, c, root, local)
t.Logf("uploaded a1=%d a2=%d b1=%d r1=%d r2=%d",
FileEntryNumericID(a1), FileEntryNumericID(a2), FileEntryNumericID(b1),
FileEntryNumericID(r1), FileEntryNumericID(r2))
if n := dedupWaitCount(t, ctx, c, aID, stem, ".txt", 2, 30*time.Second); n != 2 {
t.Fatalf("folder A has %d copies after upload, want 2", n)
}
if n := dedupWaitCount(t, ctx, c, bID, stem, ".txt", 1, 30*time.Second); n != 1 {
t.Fatalf("folder B has %d copies after upload, want 1", n)
}
if n := dedupWaitCount(t, ctx, c, root, stem, ".txt", 2, 30*time.Second); n != 2 {
t.Fatalf("project root has %d copies after upload, want 2", n)
}
// #2 mergeProjectRootForDedupe on the live tree: the project root that
// carries documents must be part of the scan exactly once.
folders, byFolder, rootFiles := liveProjectIndex(t, ctx, c, pid, root)
merged, mergedBy := mergeProjectRootForDedupe(root, folders, byFolder, rootFiles)
rootCount := 0
for _, folder := range merged {
if folder != nil && folder.ID != nil && folder.ID.String() == root {
rootCount++
}
}
if rootCount != 1 {
t.Fatalf("mergeProjectRootForDedupe: project root appears %d times in live tree, want 1", rootCount)
}
if got := len(FindFilesByDedupKey(mergedBy[root], stem, ".txt")); got != 2 {
t.Fatalf("mergeProjectRootForDedupe: root carries %d matching files, want 2", got)
}
// Within-folder scan: a duplicate pair in A and in the project root.
within := FindProjectDuplicates(merged, mergedBy, DedupOptions{})
removesByFolder := map[string]int{}
for _, g := range within {
removesByFolder[g.FolderID] = len(g.Remove)
}
if len(within) != 2 || removesByFolder[aID] != 1 || removesByFolder[root] != 1 {
t.Fatalf("within-folder groups = %d (%v), want exactly A:1 root:1", len(within), removesByFolder)
}
// DeleteFilesByDedupKey removes every stem|ext copy in one folder.
removed, err := c.DeleteFilesByDedupKey(ctx, aID, stem, ".txt")
if err != nil {
t.Fatalf("DeleteFilesByDedupKey(A): %v", err)
}
if len(removed) != 2 {
t.Fatalf("DeleteFilesByDedupKey(A) removed %v, want 2 ids", removed)
}
if n := dedupWaitCount(t, ctx, c, aID, stem, ".txt", 0, 30*time.Second); n != 0 {
t.Fatalf("folder A still has %d copies after DeleteFilesByDedupKey", n)
}
// Re-create the A duplicates so the cross-folder project scan can be
// applied and a single survivor proven by polling.
uploadFolderLive(t, ctx, c, aID, local)
uploadFolderLive(t, ctx, c, aID, local)
if n := dedupWaitCount(t, ctx, c, aID, stem, ".txt", 2, 30*time.Second); n != 2 {
t.Fatalf("folder A has %d re-uploaded copies, want 2", n)
}
// Cross-folder dry-run over the live project: one key, five copies, and
// the project-root copies must be part of the group (root merge live).
groups, deleted, err := dedupeProjectEventually(t, ctx, c, pid, DedupOptions{CrossFolder: true}, false)
if err != nil {
t.Fatalf("DedupeProject: %v", err)
}
if len(deleted) != 0 {
t.Fatalf("dry-run DedupeProject deleted %v", deleted)
}
if len(groups) != 1 {
t.Fatalf("DedupeProject cross-folder groups = %d, want 1 (%+v)", len(groups), groups)
}
if len(groups[0].Remove) != 4 {
t.Fatalf("cross-folder group removes %d files, want 4", len(groups[0].Remove))
}
if !dedupGroupHasID(groups[0], FileEntryNumericID(r1)) && !dedupGroupHasID(groups[0], FileEntryNumericID(r2)) {
t.Fatalf("cross-folder group does not include a project-root copy (merge not applied)")
}
keepID := FileEntryNumericID(groups[0].Keep)
deleted, err = c.ApplyDedupGroups(ctx, groups)
if err != nil {
t.Fatalf("ApplyDedupGroups: %v", err)
}
if len(deleted) != 4 {
t.Fatalf("ApplyDedupGroups deleted %v, want 4 ids", deleted)
}
survivorID, total := dedupWaitTotal(t, ctx, c, []string{aID, bID, root}, stem, ".txt", 1, 40*time.Second)
if total != 1 {
t.Fatalf("after ApplyDedupGroups %d copies survive, want 1", total)
}
if survivorID != keepID {
t.Fatalf("remaining copy id = %d, want kept id %d", survivorID, keepID)
}
// The removed ids must really be gone from every folder.
gone := map[int64]bool{
FileEntryNumericID(a2): true,
FileEntryNumericID(b1): true,
FileEntryNumericID(r1): true,
FileEntryNumericID(r2): true,
}
for _, fid := range []string{aID, bID, root} {
files, err := c.FolderFiles(ctx, fid)
if err != nil {
t.Fatalf("FolderFiles %s: %v", fid, err)
}
for _, f := range FindFilesByDedupKey(files, stem, ".txt") {
id := FileEntryNumericID(f)
if gone[id] {
t.Fatalf("removed id %d still present in folder %s", id, fid)
}
}
}
t.Logf("survivor id=%d keep id=%d, deleted=%v", survivorID, keepID, deleted)
}
// createFolderLive creates a subfolder (retrying transient 5xx) and returns
// its Documents folder id.
func createFolderLive(t *testing.T, ctx context.Context, c *Client, parentID, title string) string {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var (
m map[string]any
err error
)
for {
m, err = c.CreateFolder(ctx, parentID, title)
if err == nil || time.Now().After(deadline) {
break
}
time.Sleep(time.Second)
}
if err != nil {
t.Fatalf("CreateFolder %q: %v", title, err)
}
if m == nil {
t.Fatalf("CreateFolder %q: empty response", title)
}
switch v := m["id"].(type) {
case float64:
return strconv.FormatInt(int64(v), 10)
case json.Number:
return v.String()
case string:
if v != "" {
return v
}
}
t.Fatalf("CreateFolder %q: no id in response %#v", title, m)
return ""
}
// writeLocalFile writes content to a temp file and returns its path.
func writeLocalFile(t *testing.T, name string, content []byte) string {
t.Helper()
p := filepath.Join(t.TempDir(), name)
if err := os.WriteFile(p, content, 0o600); err != nil {
t.Fatal(err)
}
return p
}
// uploadFolderLive uploads localPath into folderID (retrying the portal's
// transient post-create 500) and returns the file entry.
func uploadFolderLive(t *testing.T, ctx context.Context, c *Client, folderID, localPath string) *FileEntry {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var (
e *FileEntry
err error
)
for {
e, err = c.UploadToFolder(ctx, folderID, localPath)
if err == nil || time.Now().After(deadline) {
break
}
time.Sleep(time.Second)
}
if err != nil {
t.Fatalf("UploadToFolder %s: %v", folderID, err)
}
if e == nil || e.ID == nil {
t.Fatalf("UploadToFolder %s: no file entry (%+v)", folderID, e)
}
return e
}
// liveProjectIndex rebuilds the project folder/file index the same way
// DedupeProject does, for direct mergeProjectRootForDedupe assertions.
func liveProjectIndex(t *testing.T, ctx context.Context, c *Client, projectID, rootID string) ([]*FolderEntry, map[string][]*FileEntry, []*FileEntry) {
t.Helper()
pf := getProjectFilesEventually(t, ctx, c, projectID)
rootFiles := folderFilesEventually(t, ctx, c, rootID)
folders := make([]*FolderEntry, 0, len(pf.Folders)+1)
byFolder := make(map[string][]*FileEntry, len(pf.Folders)+1)
for _, folder := range pf.Folders {
if folder == nil || folder.ID == nil {
continue
}
fid := folder.ID.String()
if fid == rootID {
byFolder[fid] = rootFiles
} else {
byFolder[fid] = folderFilesEventually(t, ctx, c, fid)
}
folders = append(folders, folder)
}
return folders, byFolder, rootFiles
}
// projectFolderEventually resolves the project Documents root id, retrying on
// a transient portal answer.
func projectFolderEventually(t *testing.T, ctx context.Context, c *Client, projectID string) string {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var (
root string
err error
)
for {
root, err = c.projectFolderID(ctx, projectID)
if err == nil || time.Now().After(deadline) {
break
}
time.Sleep(time.Second)
}
if err != nil {
t.Fatalf("projectFolderID: %v", err)
}
if root == "" {
t.Fatal("projectFolderID returned empty id")
}
return root
}
// getProjectFilesEventually lists a project's files/folders, retrying on a
// transient portal answer.
func getProjectFilesEventually(t *testing.T, ctx context.Context, c *Client, projectID string) *ProjectFilesResponse {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var last error
for {
pf, err := c.GetProjectFiles(ctx, projectID)
if err == nil {
return pf
}
last = err
if time.Now().After(deadline) {
t.Fatalf("GetProjectFiles %s: %v", projectID, last)
}
time.Sleep(500 * time.Millisecond)
}
}
// folderFilesEventually lists a folder, retrying while the portal answers
// transiently (a freshly created folder can 500 until its parent map settles).
func folderFilesEventually(t *testing.T, ctx context.Context, c *Client, folderID string) []*FileEntry {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var last error
for {
files, err := c.FolderFiles(ctx, folderID)
if err == nil {
return files
}
last = err
if time.Now().After(deadline) {
t.Fatalf("FolderFiles %s: %v", folderID, last)
}
time.Sleep(500 * time.Millisecond)
}
}
// dedupeProjectEventually runs a project dedup scan, retrying the whole scan on
// a transient portal error. It is used for dry-runs only (apply must stay a
// single deliberate call).
func dedupeProjectEventually(t *testing.T, ctx context.Context, c *Client, projectID string, opts DedupOptions, apply bool) ([]DedupGroup, []int, error) {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
var (
groups []DedupGroup
deleted []int
err error
)
for {
groups, deleted, err = c.DedupeProject(ctx, projectID, opts, apply)
if err == nil || time.Now().After(deadline) {
return groups, deleted, err
}
time.Sleep(time.Second)
}
}
// dedupWaitCount polls folderID until want stem|ext copies are visible.
func dedupWaitCount(t *testing.T, ctx context.Context, c *Client, folderID, stem, ext string, want int, d time.Duration) int {
t.Helper()
deadline := time.Now().Add(d)
got := -1
for {
if files, err := c.FolderFiles(ctx, folderID); err == nil {
got = len(FindFilesByDedupKey(files, stem, ext))
if got == want {
return got
}
}
if time.Now().After(deadline) {
return got
}
time.Sleep(500 * time.Millisecond)
}
}
// dedupWaitTotal polls the given folders until the total number of stem|ext
// copies reaches want, returning the last seen file id and count.
func dedupWaitTotal(t *testing.T, ctx context.Context, c *Client, folderIDs []string, stem, ext string, want int, d time.Duration) (int64, int) {
t.Helper()
deadline := time.Now().Add(d)
var survivor int64
total := -1
for {
survivor, total = 0, 0
for _, fid := range folderIDs {
files, err := c.FolderFiles(ctx, fid)
if err != nil {
continue
}
for _, f := range FindFilesByDedupKey(files, stem, ext) {
total++
survivor = FileEntryNumericID(f)
}
}
if total == want {
return survivor, total
}
if time.Now().After(deadline) {
return survivor, total
}
time.Sleep(500 * time.Millisecond)
}
}
func dedupGroupHasID(g DedupGroup, id int64) bool {
if id == 0 {
return false
}
if FileEntryNumericID(g.Keep) == id {
return true
}
for _, f := range g.Remove {
if FileEntryNumericID(f) == id {
return true
}
}
return false
}
+127
View File
@@ -0,0 +1,127 @@
package onlyoffice
import (
"encoding/json"
"testing"
"time"
)
func TestFileDedupKey(t *testing.T) {
title := "OO-HONDA-7-INDEX.docx"
exst := ".docx"
f := &FileEntry{Title: &title, FileExst: &exst}
if got := FileDedupKey(f); got != "OO-HONDA-7-INDEX|docx" {
t.Fatalf("got %q", got)
}
}
func TestFindFilesByDedupKey(t *testing.T) {
a := &FileEntry{Title: strPtr("foo.docx"), FileExst: strPtr(".docx")}
b := &FileEntry{Title: strPtr("foo.md"), FileExst: strPtr(".md")}
files := []*FileEntry{a, b}
got := FindFilesByDedupKey(files, "foo", ".docx")
if len(got) != 1 || got[0] != a {
t.Fatalf("got %+v", got)
}
}
func TestFindWithinFolderDuplicates(t *testing.T) {
t1 := time.Date(2026, 8, 27, 16, 0, 0, 0, time.UTC)
t2 := t1.Add(time.Hour)
old := &FileEntry{ID: jsonNum("1"), Title: strPtr("idx.docx"), FileExst: strPtr(".docx"), Updated: &t1}
new := &FileEntry{ID: jsonNum("2"), Title: strPtr("idx.docx"), FileExst: strPtr(".docx"), Updated: &t2}
indexed := []ProjectFolderFile{
{FolderID: "490", FolderTitle: "00-Index", File: old},
{FolderID: "490", FolderTitle: "00-Index", File: new},
}
groups := findWithinFolderDuplicates(indexed)
if len(groups) != 1 {
t.Fatalf("groups=%d", len(groups))
}
if FileEntryNumericID(groups[0].Keep) != 2 {
t.Fatalf("keep id=%d", FileEntryNumericID(groups[0].Keep))
}
if len(groups[0].Remove) != 1 || FileEntryNumericID(groups[0].Remove[0]) != 1 {
t.Fatalf("remove=%v", groups[0].Remove)
}
}
func TestCrossFolderPrefersNonTrash(t *testing.T) {
t1 := time.Date(2026, 8, 27, 18, 0, 0, 0, time.UTC)
t2 := t1.Add(-time.Hour)
trash := &FileEntry{ID: jsonNum("10"), Title: strPtr("INDEX.md"), FileExst: strPtr(".md"), Updated: &t1}
good := &FileEntry{ID: jsonNum("20"), Title: strPtr("INDEX.md"), FileExst: strPtr(".md"), Updated: &t2}
indexed := []ProjectFolderFile{
{FolderID: "493", FolderTitle: "_trash-md", File: trash},
{FolderID: "492", FolderTitle: "OCR", File: good},
}
groups := findCrossFolderDuplicates(indexed)
if len(groups) != 1 {
t.Fatalf("groups=%d", len(groups))
}
if FileEntryNumericID(groups[0].Keep) != 20 {
t.Fatalf("keep id=%d", FileEntryNumericID(groups[0].Keep))
}
}
func TestMergeProjectRootForDedupe(t *testing.T) {
old := &FileEntry{ID: jsonNum("1"), Title: strPtr("a.docx"), FileExst: strPtr(".docx")}
newer := &FileEntry{ID: jsonNum("2"), Title: strPtr("a.docx"), FileExst: strPtr(".docx")}
rootFiles := []*FileEntry{old, newer}
folders, byFolder := mergeProjectRootForDedupe("489", nil, nil, rootFiles)
if len(folders) != 1 || folders[0].ID.String() != "489" {
t.Fatalf("folders=%+v", folders)
}
if len(byFolder["489"]) != 2 {
t.Fatalf("root files=%d", len(byFolder["489"]))
}
groups := findWithinFolderDuplicates([]ProjectFolderFile{
{FolderID: "489", FolderTitle: "(project root)", File: old},
{FolderID: "489", FolderTitle: "(project root)", File: newer},
})
if len(groups) != 1 {
t.Fatalf("groups=%d", len(groups))
}
}
func TestFindDuplicatesSkipsEmptyKey(t *testing.T) {
// Dotfiles (".env", ".gitignore", ".npmrc") normalize to an empty stem, so
// FileDedupKey is "". They are not duplicates of each other and must never
// form a dedup group that would delete one of them.
env := &FileEntry{ID: jsonNum("1"), Title: strPtr(".env")}
gitignore := &FileEntry{ID: jsonNum("2"), Title: strPtr(".gitignore")}
npmrc := &FileEntry{ID: jsonNum("3"), Title: strPtr(".npmrc")}
if FileDedupKey(env) != "" || FileDedupKey(gitignore) != "" || FileDedupKey(npmrc) != "" {
t.Fatalf("dotfiles should have empty dedup key")
}
within := []ProjectFolderFile{
{FolderID: "500", FolderTitle: "Cfg", File: env},
{FolderID: "500", FolderTitle: "Cfg", File: gitignore},
}
if groups := findWithinFolderDuplicates(within); len(groups) != 0 {
t.Fatalf("within-folder empty-key groups = %d, want 0 (%+v)", len(groups), groups)
}
cross := []ProjectFolderFile{
{FolderID: "500", FolderTitle: "Cfg", File: env},
{FolderID: "501", FolderTitle: "Other", File: npmrc},
}
if groups := findCrossFolderDuplicates(cross); len(groups) != 0 {
t.Fatalf("cross-folder empty-key groups = %d, want 0 (%+v)", len(groups), groups)
}
}
func TestIsTrashFolderTitle(t *testing.T) {
if !IsTrashFolderTitle("_trash-md") {
t.Fatal("expected trash")
}
if IsTrashFolderTitle("00-Index") {
t.Fatal("expected not trash")
}
}
func strPtr(s string) *string { return &s }
func jsonNum(s string) *json.Number {
n := json.Number(s)
return &n
}
+79
View File
@@ -0,0 +1,79 @@
package onlyoffice
// Human-readable folder paths for search results (F9). The OnlyOffice ES
// index stores only ancestor folder ids; titles live in the Documents tree, so
// resolving a path costs one GET /api/2.0/files/{id} per distinct folder,
// cached on the client. Folders that cannot be listed (e.g. a section root)
// fall back to their id, so a path is always produced.
import (
"context"
"strings"
)
// FolderTitle returns the title of a Documents folder id, cached on the client.
// An empty id yields an empty title. Unknown/unlistable ids (section roots)
// return ("", nil) so callers can fall back to the id.
func (c *Client) FolderTitle(ctx context.Context, folderID string) (string, error) {
folderID = strings.TrimSpace(folderID)
if folderID == "" {
return "", nil
}
c.folderTitlesMu.Lock()
if c.folderTitles != nil {
if t, ok := c.folderTitles[folderID]; ok {
c.folderTitlesMu.Unlock()
return t, nil
}
}
c.folderTitlesMu.Unlock()
title := ""
out, err := c.ListFolder(ctx, folderID)
if err == nil {
if cur, ok := out["current"].(map[string]any); ok {
if s, ok := cur["title"].(string); ok {
title = strings.TrimSpace(s)
}
}
}
c.folderTitlesMu.Lock()
if c.folderTitles == nil {
c.folderTitles = map[string]string{}
}
c.folderTitles[folderID] = title
c.folderTitlesMu.Unlock()
return title, nil
}
// FolderPath resolves an ancestor folder id chain (root → leaf, as the ES
// backend reports it) into folder titles, falling back to the id when a title
// cannot be read. The result never fails on a single lookup: only the whole
// call honours ctx cancellation.
func (c *Client) FolderPath(ctx context.Context, ids []string) []string {
out := make([]string, 0, len(ids))
for _, id := range ids {
if err := ctx.Err(); err != nil {
break
}
title, err := c.FolderTitle(ctx, id)
if err != nil || title == "" {
title = id
}
out = append(out, title)
}
return out
}
// UniquePath builds a stable, human-readable, unique path for a result: the
// resolved folder chain plus the file title. "." separates nothing — the
// segments are joined with "/", matching the Documents breadcrumb the web UI
// shows.
func (c *Client) UniquePath(ctx context.Context, folderPath []string, title string) string {
parts := c.FolderPath(ctx, folderPath)
if t := strings.TrimSpace(title); t != "" {
parts = append(parts, t)
}
return strings.Join(parts, "/")
}
+35
View File
@@ -0,0 +1,35 @@
//go:build integration
package onlyoffice
import (
"context"
"strings"
"testing"
)
// TestIntegrationFolderPath resolves the real Fibu EDL folder chain
// (project root 522 → Eingangsrechnungen 647 → 2025 649) to titles.
func TestIntegrationFolderPath(t *testing.T) {
creds := GetEnvironmentCredentials()
if strings.TrimSpace(creds.Url) == "" || strings.TrimSpace(creds.User) == "" {
t.Skip("no ONLYOFFICE_URL/USER credentials")
}
c := NewClient(creds)
ctx := context.Background()
path := c.FolderPath(ctx, []string{"522", "647", "649"})
if len(path) != 3 {
t.Fatalf("FolderPath returned %v, want 3 segments", path)
}
for i, seg := range path {
if strings.TrimSpace(seg) == "" {
t.Errorf("segment %d empty: %v", i, path)
}
}
full := c.UniquePath(ctx, []string{"522", "647", "649"}, "Rechnung-x.pdf")
if !strings.HasSuffix(full, "Rechnung-x.pdf") || !strings.Contains(full, "/") {
t.Errorf("UniquePath = %q, want a slash-joined path ending in the file", full)
}
t.Logf("path=%v full=%q", path, full)
}
+180
View File
@@ -0,0 +1,180 @@
package onlyoffice
import (
"testing"
"time"
)
func TestEquivalentUploadExt(t *testing.T) {
tests := []struct {
a, b string
want bool
}{
{".xls", ".xlsx", true},
{".XLS", ".xlsx", true},
{"xls", "xlsx", true},
{".doc", ".docx", true},
{".ppt", ".pptx", true},
{".pdf", ".pdf", true},
{"", "", true},
{".xls", ".docx", false},
{".xlsx", "", false},
{".csv", ".xlsx", false},
}
for _, tc := range tests {
if got := EquivalentUploadExt(tc.a, tc.b); got != tc.want {
t.Errorf("EquivalentUploadExt(%q,%q)=%v want %v", tc.a, tc.b, got, tc.want)
}
if got := EquivalentUploadExt(tc.b, tc.a); got != tc.want {
t.Errorf("EquivalentUploadExt(%q,%q)=%v want %v (symmetric)", tc.b, tc.a, got, tc.want)
}
}
}
func TestFindFilesByStemExtMatchesConvertedXLS(t *testing.T) {
xlsx := &FileEntry{ID: jsonNum("3799"), Title: strPtr("ES29-extracto.xlsx"), FileExst: strPtr(".xlsx")}
pdf := &FileEntry{ID: jsonNum("5"), Title: strPtr("ES29-extracto.pdf"), FileExst: strPtr(".pdf")}
other := &FileEntry{ID: jsonNum("3887"), Title: strPtr("ES87-extracto.xlsx"), FileExst: strPtr(".xlsx")}
files := []*FileEntry{xlsx, pdf, other}
got := FindFilesByStemExt(files, "ES29-extracto", ".xls")
if len(got) != 1 || got[0] != xlsx {
t.Fatalf("converted .xls match = %+v, want the saved .xlsx only", got)
}
}
func TestFindFilesByStemExtKeepsExactMatch(t *testing.T) {
xlsx := &FileEntry{ID: jsonNum("3799"), Title: strPtr("foo.xlsx"), FileExst: strPtr(".xlsx")}
pdf := &FileEntry{ID: jsonNum("5"), Title: strPtr("foo.pdf"), FileExst: strPtr(".pdf")}
files := []*FileEntry{xlsx, pdf}
if got := FindFilesByStemExt(files, "foo", ".xlsx"); len(got) != 1 || got[0] != xlsx {
t.Fatalf("exact .xlsx match = %+v, want only xlsx", got)
}
if got := FindFilesByStemExt(files, "foo", ".pdf"); len(got) != 1 || got[0] != pdf {
t.Fatalf("exact .pdf match = %+v, want only pdf", got)
}
if got := FindFilesByStemExt(files, "foo", ".docx"); len(got) != 0 {
t.Fatalf("unrelated ext matched %+v, want none", got)
}
}
func TestFindFilesByStemExtEmptyStem(t *testing.T) {
f := &FileEntry{ID: jsonNum("1"), Title: strPtr("foo.xlsx"), FileExst: strPtr(".xlsx")}
if got := FindFilesByStemExt([]*FileEntry{f}, "", ".xlsx"); len(got) != 0 {
t.Fatalf("empty stem matched %+v, want none", got)
}
}
func TestPlanUploadReplacementFreshUpload(t *testing.T) {
plan := planUploadReplacement(nil, "ES29-extracto", ".xls")
if plan.UpdateID != "" || len(plan.DeleteIDs) != 0 {
t.Fatalf("empty folder plan = %+v, want a fresh upload", plan)
}
}
// TestPlanUploadReplacementUpdatesExactExt covers the in-place path: only a
// stored file with the same extension is updated via UpdateFile. PDF over PDF
// and OOXML over OOXML keep the id and rewrite the body.
func TestPlanUploadReplacementUpdatesExactExt(t *testing.T) {
tests := []struct {
name string
file *FileEntry
stem string
ext string
id string
}{
{
name: "xlsx over xlsx",
file: &FileEntry{ID: jsonNum("3799"), Title: strPtr("ES29-extracto.xlsx"), FileExst: strPtr(".xlsx")},
stem: "ES29-extracto", ext: ".xlsx", id: "3799",
},
{
name: "pdf over pdf",
file: &FileEntry{ID: jsonNum("42"), Title: strPtr("extracto.pdf"), FileExst: strPtr(".pdf")},
stem: "extracto", ext: ".pdf", id: "42",
},
{
name: "legacy xls over xls",
file: &FileEntry{ID: jsonNum("11"), Title: strPtr("legacy.xls"), FileExst: strPtr(".xls")},
stem: "legacy", ext: ".xls", id: "11",
},
}
for _, tc := range tests {
t.Run(tc.name, func(t *testing.T) {
plan := planUploadReplacement([]*FileEntry{tc.file}, tc.stem, tc.ext)
if plan.UpdateID != tc.id {
t.Fatalf("UpdateID = %q, want %q (same ext updates in place)", plan.UpdateID, tc.id)
}
if len(plan.DeleteIDs) != 0 {
t.Fatalf("DeleteIDs = %v, want none", plan.DeleteIDs)
}
})
}
}
// TestPlanUploadReplacementConvertedExtDeletesThenUploads is the regression for
// #84: UpdateFile does not re-run the server-side legacy→OOXML conversion, so a
// .xls upload must not overwrite a stored .xlsx in place (raw OLE2 under an
// .xlsx name). The stale counterpart is deleted and the file uploaded afresh.
func TestPlanUploadReplacementConvertedExtDeletesThenUploads(t *testing.T) {
xlsx := &FileEntry{ID: jsonNum("3799"), Title: strPtr("ES29-extracto.xlsx"), FileExst: strPtr(".xlsx")}
plan := planUploadReplacement([]*FileEntry{xlsx}, "ES29-extracto", ".xls")
if plan.UpdateID != "" {
t.Fatalf("UpdateID = %q, want empty (do not update across conversion)", plan.UpdateID)
}
if len(plan.DeleteIDs) != 1 || plan.DeleteIDs[0] != 3799 {
t.Fatalf("DeleteIDs = %v, want [3799] (delete the stale .xlsx before upload)", plan.DeleteIDs)
}
}
func TestPlanUploadReplacementCollapsesDuplicates(t *testing.T) {
older := time.Date(2026, 9, 1, 10, 0, 0, 0, time.UTC)
newer := older.Add(time.Hour)
first := &FileEntry{ID: jsonNum("3799"), Title: strPtr("ES29-extracto.xlsx"), FileExst: strPtr(".xlsx"), Updated: &older}
second := &FileEntry{ID: jsonNum("3887"), Title: strPtr("ES29-extracto.xlsx"), FileExst: strPtr(".xlsx"), Updated: &newer}
plan := planUploadReplacement([]*FileEntry{first, second}, "ES29-extracto", ".xlsx")
if plan.UpdateID != "3887" {
t.Fatalf("UpdateID = %q, want the newest duplicate 3887", plan.UpdateID)
}
if len(plan.DeleteIDs) != 1 || plan.DeleteIDs[0] != 3799 {
t.Fatalf("DeleteIDs = %v, want [3799]", plan.DeleteIDs)
}
}
// TestPlanUploadReplacementDeletesConvertedDuplicates: with a converted match
// every stale copy is deleted (there is no keeper — the fresh upload replaces
// them all).
func TestPlanUploadReplacementDeletesConvertedDuplicates(t *testing.T) {
first := &FileEntry{ID: jsonNum("3799"), Title: strPtr("ES29-extracto.xlsx"), FileExst: strPtr(".xlsx")}
second := &FileEntry{ID: jsonNum("3887"), Title: strPtr("ES29-extracto.xlsx"), FileExst: strPtr(".xlsx")}
plan := planUploadReplacement([]*FileEntry{first, second}, "ES29-extracto", ".xls")
if plan.UpdateID != "" {
t.Fatalf("UpdateID = %q, want empty", plan.UpdateID)
}
if len(plan.DeleteIDs) != 2 {
t.Fatalf("DeleteIDs = %v, want both stale .xlsx ids", plan.DeleteIDs)
}
}
// TestPlanUploadReplacementRepeatedXLS is the regression for #84: with the
// server-converted .xlsx already present, the second .xls upload deletes it and
// uploads anew so OnlyOffice converts again — UpdateFile would corrupt it.
func TestPlanUploadReplacementRepeatedXLS(t *testing.T) {
const stem = "ES29-extracto"
// First upload: nothing in the folder.
if plan := planUploadReplacement(nil, stem, ".xls"); plan.UpdateID != "" || len(plan.DeleteIDs) != 0 {
t.Fatalf("first upload plan = %+v, want create", plan)
}
// OnlyOffice converts .xls -> .xlsx on upload; the second upload must
// delete it and upload fresh, never UpdateFile it.
saved := &FileEntry{ID: jsonNum("3799"), Title: strPtr(stem + ".xlsx"), FileExst: strPtr(".xlsx")}
plan := planUploadReplacement([]*FileEntry{saved}, stem, ".xls")
if plan.UpdateID != "" {
t.Fatalf("second upload plan = %+v, want delete+upload (not UpdateFile)", plan)
}
if len(plan.DeleteIDs) != 1 || plan.DeleteIDs[0] != 3799 {
t.Fatalf("second upload DeleteIDs = %v, want [3799]", plan.DeleteIDs)
}
}
+254
View File
@@ -0,0 +1,254 @@
package onlyoffice
import (
"context"
"encoding/json"
"errors"
"fmt"
"path/filepath"
"strconv"
"strings"
)
// ErrFileExists is returned when --no-replace / no-clobber upload hits an existing stem|ext.
var ErrFileExists = errors.New("onlyoffice: file already exists in folder (use replace or delete first)")
// FileEntryStem returns the logical basename without duplicated extensions.
// OO often stores title="foo.docx" and fileExst=".docx" (UI shows foo.docx.docx).
func FileEntryStem(f *FileEntry) string {
if f == nil || f.Title == nil {
return ""
}
exst := ""
if f.FileExst != nil {
exst = *f.FileExst
}
return NormalizeUploadStem(*f.Title, exst)
}
// NormalizeUploadStem derives a stable stem for matching uploads.
func NormalizeUploadStem(title, exst string) string {
t := strings.TrimSpace(title)
t = strings.TrimSuffix(t, ".")
if exst != "" && strings.HasSuffix(t, exst) {
t = strings.TrimSuffix(t, exst)
}
if ext := filepath.Ext(t); ext != "" {
t = strings.TrimSuffix(t, ext)
}
return strings.TrimSpace(t)
}
// UploadStemFromLocal returns the stem used to match/replace folder files.
func UploadStemFromLocal(localPath string) string {
base := filepath.Base(localPath)
ext := filepath.Ext(base)
if ext != "" {
base = strings.TrimSuffix(base, ext)
}
return base
}
// FolderFiles returns file entries in a Documents folder.
func (c *Client) FolderFiles(ctx context.Context, folderID string) ([]*FileEntry, error) {
raw, err := c.ListFolder(ctx, folderID)
if err != nil {
return nil, err
}
return ParseFolderFileEntries(raw), nil
}
// ParseFolderFileEntries extracts []*FileEntry from ListFolder JSON.
func ParseFolderFileEntries(raw map[string]any) []*FileEntry {
items, _ := raw["files"].([]any)
out := make([]*FileEntry, 0, len(items))
for _, it := range items {
m, ok := it.(map[string]any)
if !ok {
continue
}
b, err := json.Marshal(m)
if err != nil {
continue
}
var f FileEntry
if err := json.Unmarshal(b, &f); err != nil {
continue
}
out = append(out, &f)
}
return out
}
// FindFilesByStem returns folder files whose logical stem matches.
func FindFilesByStem(files []*FileEntry, stem string) []*FileEntry {
stem = strings.TrimSpace(stem)
if stem == "" {
return nil
}
var out []*FileEntry
for _, f := range files {
if FileEntryStem(f) == stem {
out = append(out, f)
}
}
return out
}
// DeleteFilesByStem removes all files in folderID matching stem (any extension).
// Prefer DeleteFilesByDedupKey when the upload extension is known.
func (c *Client) DeleteFilesByStem(ctx context.Context, folderID, stem string) ([]int, error) {
files, err := c.FolderFiles(ctx, folderID)
if err != nil {
return nil, err
}
matches := FindFilesByStem(files, stem)
if len(matches) == 0 {
return nil, nil
}
ids := make([]int, 0, len(matches))
for _, f := range matches {
n := int(FileEntryNumericID(f))
if n != 0 {
ids = append(ids, n)
}
}
if len(ids) == 0 {
return nil, nil
}
if err := c.DeleteFiles(ctx, ids); err != nil {
return nil, err
}
return ids, nil
}
// AssertNoFileConflict reports ErrFileExists when localPath logical name is
// already in folderID. Matching is conversion-aware, so a .xls upload also
// conflicts with an existing .xlsx and does not create a hidden duplicate.
func (c *Client) AssertNoFileConflict(ctx context.Context, folderID, localPath string) error {
files, err := c.FolderFiles(ctx, folderID)
if err != nil {
return err
}
stem := UploadStemFromLocal(localPath)
ext := UploadExtFromLocal(localPath)
matches := FindFilesByStemExt(files, stem, ext)
if len(matches) == 0 {
return nil
}
ids := make([]string, 0, len(matches))
for _, f := range matches {
ids = append(ids, fmt.Sprintf("%d", FileEntryNumericID(f)))
}
return fmt.Errorf("%w: %s%s in folder %s (existing file ids: %s)",
ErrFileExists, stem, ext, folderID, strings.Join(ids, ", "))
}
// UploadProjectFileNoClobber uploads only when stem|ext is not already in the project folder.
func (c *Client) UploadProjectFileNoClobber(ctx context.Context, projectID, localPath string) (*FileEntry, error) {
folderID, err := c.projectFolderID(ctx, projectID)
if err != nil {
return nil, err
}
if err := c.AssertNoFileConflict(ctx, folderID, localPath); err != nil {
return nil, err
}
return c.UploadProjectFile(ctx, projectID, localPath)
}
// uploadReplacementPlan is how a replacing upload reconciles with a folder.
type uploadReplacementPlan struct {
// UpdateID is the existing file to overwrite in place; set only when the
// stored extension equals the local one. UpdateFile keeps the stored name,
// so overwriting across extensions would leave an unconverted body under a
// mismatched name.
UpdateID string
// DeleteIDs are the ids to remove before uploading. They are the redundant
// duplicates of an in-place update, or every conversion-equivalent
// counterpart when the body must be converted again by a fresh upload.
DeleteIDs []int
}
// planUploadReplacement matches an incoming local upload (stem + ext) against
// the files already in a folder and decides between a fresh upload, an in-place
// update and delete + reupload. Matching tolerates the legacy→OOXML conversion
// OnlyOffice performs on upload, so a repeated upload of f.xls finds the saved
// f.xlsx instead of creating a second file. Because UpdateFile replaces the body
// without re-running that conversion, an equivalent-but-different extension is
// deleted and re-uploaded rather than updated in place (#84).
func planUploadReplacement(files []*FileEntry, stem, ext string) uploadReplacementPlan {
matches := FindFilesByStemExt(files, stem, ext)
if len(matches) == 0 {
return uploadReplacementPlan{}
}
keep, remove := pickDuplicateKeeper(matches, false)
plan := uploadReplacementPlan{}
if FileEntryExt(keep) == normalizeExt(ext) {
if id := FileEntryNumericID(keep); id != 0 {
plan.UpdateID = strconv.FormatInt(id, 10)
}
} else if id := int(FileEntryNumericID(keep)); id != 0 {
plan.DeleteIDs = append(plan.DeleteIDs, id)
}
for _, f := range remove {
if id := int(FileEntryNumericID(f)); id != 0 {
plan.DeleteIDs = append(plan.DeleteIDs, id)
}
}
return plan
}
// UploadToFolderReplacing upserts localPath into folderID by logical name.
// Matching is stem + extension with OnlyOffice's server-side conversion
// accounted for: a local .xls is stored as .xlsx, so a repeated upload replaces
// the saved document instead of appending a duplicate. An exact-extension
// counterpart is updated in place (stable id, no window without the file); a
// conversion-equivalent one is deleted and uploaded afresh so the server
// converts the body again. Extra duplicates are collapsed either way.
func (c *Client) UploadToFolderReplacing(ctx context.Context, folderID, localPath string) (*FileEntry, []int, error) {
stem := UploadStemFromLocal(localPath)
ext := UploadExtFromLocal(localPath)
files, err := c.FolderFiles(ctx, folderID)
if err != nil {
return nil, nil, err
}
plan := planUploadReplacement(files, stem, ext)
if plan.UpdateID != "" {
ent, err := c.UpdateFile(ctx, plan.UpdateID, localPath)
if err == nil {
deleted, derr := c.deleteReplacementStale(ctx, plan.DeleteIDs)
return ent, deleted, derr
}
// Portal rejected the in-place update: delete + fresh upload still
// leaves one file (server-converted).
deleted, derr := c.DeleteFilesByStemExt(ctx, folderID, stem, ext)
if derr != nil {
return nil, deleted, derr
}
ent, uerr := c.UploadToFolder(ctx, folderID, localPath)
return ent, deleted, uerr
}
// Conversion-equivalent (or no counterpart): remove the stale files first so
// the fresh upload is converted and exactly one file remains.
deleted, derr := c.deleteReplacementStale(ctx, plan.DeleteIDs)
if derr != nil {
return nil, deleted, derr
}
ent, uerr := c.UploadToFolder(ctx, folderID, localPath)
return ent, deleted, uerr
}
// deleteReplacementStale removes the ids collected by planUploadReplacement:
// duplicate files after an in-place update, or every stale counterpart before a
// converted reupload.
func (c *Client) deleteReplacementStale(ctx context.Context, ids []int) ([]int, error) {
if len(ids) == 0 {
return nil, nil
}
if err := c.DeleteFiles(ctx, ids); err != nil {
return ids, err
}
return ids, nil
}
+35
View File
@@ -0,0 +1,35 @@
package onlyoffice
import "testing"
func TestNormalizeUploadStem(t *testing.T) {
tests := []struct {
title, exst, want string
}{
{"OO-HONDA-7-INDEX.docx", ".docx", "OO-HONDA-7-INDEX"},
{"README.txt", ".txt", "README"},
{"car-docs-print.docx", ".docx", "car-docs-print"},
{"plain.", "", "plain"},
{"foo", ".docx", "foo"},
}
for _, tc := range tests {
if got := NormalizeUploadStem(tc.title, tc.exst); got != tc.want {
t.Fatalf("NormalizeUploadStem(%q,%q)=%q want %q", tc.title, tc.exst, got, tc.want)
}
}
}
func TestFileEntryStem(t *testing.T) {
title := "00-INDEX.docx"
exst := ".docx"
f := &FileEntry{Title: &title, FileExst: &exst}
if got := FileEntryStem(f); got != "00-INDEX" {
t.Fatalf("got %q", got)
}
}
func TestUploadStemFromLocal(t *testing.T) {
if got := UploadStemFromLocal("/tmp/OO-HONDA-7-INDEX.docx"); got != "OO-HONDA-7-INDEX" {
t.Fatalf("got %q", got)
}
}
+95 -34
View File
@@ -52,6 +52,8 @@ type DavListing struct {
// ListDavFolder returns the contents of a folder by id, which may be a // ListDavFolder returns the contents of a folder by id, which may be a
// symbolic root such as "@my". For "@root" use ListDavSections. // symbolic root such as "@my". For "@root" use ListDavSections.
//
// Deprecated: use FileStore.List via Client.Files()/Client.FileStore.
func (c *Client) ListDavFolder(ctx context.Context, id string) (*DavListing, error) { func (c *Client) ListDavFolder(ctx context.Context, id string) (*DavListing, error) {
raw, err := c.getJSON(ctx, "/api/2.0/files/"+url.PathEscape(id)) raw, err := c.getJSON(ctx, "/api/2.0/files/"+url.PathEscape(id))
if err != nil { if err != nil {
@@ -111,6 +113,8 @@ func (c *Client) ListDavSections(ctx context.Context) ([]DavFolder, error) {
} }
// CreateDavFolder creates a folder titled title inside parentID. // CreateDavFolder creates a folder titled title inside parentID.
//
// Deprecated: use FileStore.CreateFolder via Client.Files()/Client.FileStore.
func (c *Client) CreateDavFolder(ctx context.Context, parentID, title string) (*DavFolder, error) { func (c *Client) CreateDavFolder(ctx context.Context, parentID, title string) (*DavFolder, error) {
raw, err := c.postJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(parentID), raw, err := c.postJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(parentID),
map[string]string{"title": title}) map[string]string{"title": title})
@@ -130,6 +134,8 @@ func (c *Client) CreateDavFolder(ctx context.Context, parentID, title string) (*
} }
// RenameDavFolder renames a folder. // RenameDavFolder renames a folder.
//
// Deprecated: use FileStore.Rename via Client.Files()/Client.FileStore.
func (c *Client) RenameDavFolder(ctx context.Context, id, title string) error { func (c *Client) RenameDavFolder(ctx context.Context, id, title string) error {
_, err := c.putJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(id), _, err := c.putJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(id),
map[string]string{"title": title}) map[string]string{"title": title})
@@ -137,6 +143,8 @@ func (c *Client) RenameDavFolder(ctx context.Context, id, title string) error {
} }
// RenameDavFile renames a file (title includes the extension). // RenameDavFile renames a file (title includes the extension).
//
// Deprecated: use FileStore.Rename via Client.Files()/Client.FileStore.
func (c *Client) RenameDavFile(ctx context.Context, id, title string) error { func (c *Client) RenameDavFile(ctx context.Context, id, title string) error {
_, err := c.putJSON(ctx, "/api/2.0/files/file/"+url.PathEscape(id), _, err := c.putJSON(ctx, "/api/2.0/files/file/"+url.PathEscape(id),
map[string]string{"title": title}) map[string]string{"title": title})
@@ -144,46 +152,121 @@ func (c *Client) RenameDavFile(ctx context.Context, id, title string) error {
} }
// MoveDavItems moves the given folders and/or files into destFolderID. // MoveDavItems moves the given folders and/or files into destFolderID.
// The fileops API answers 200 with per-operation error strings even when
// nothing moves (e.g. missing permission), so the response is parsed and the
// first operation error is returned instead of a silent nil.
//
// Deprecated: use FileStore.Move via Client.Files()/Client.FileStore.
func (c *Client) MoveDavItems(ctx context.Context, folderIDs, fileIDs []string, destFolderID string) error { func (c *Client) MoveDavItems(ctx context.Context, folderIDs, fileIDs []string, destFolderID string) error {
_, err := c.putJSON(ctx, "/api/2.0/files/fileops/move", map[string]any{ raw, err := c.putJSON(ctx, "/api/2.0/files/fileops/move", map[string]any{
"folderIds": nums(folderIDs), "folderIds": nums(folderIDs),
"fileIds": nums(fileIDs), "fileIds": nums(fileIDs),
"destFolderId": num(destFolderID), "destFolderId": num(destFolderID),
"resolveType": "Skip", "resolveType": "Skip",
"holdResult": true, "holdResult": true,
}) })
if err != nil {
return err return err
}
return fileopsError(raw)
} }
// CopyDavItems copies the given folders and/or files into destFolderID. // CopyDavItems copies the given folders and/or files into destFolderID.
// Per-operation errors are surfaced like in MoveDavItems.
//
// Deprecated: use FileStore.Copy via Client.Files()/Client.FileStore.
func (c *Client) CopyDavItems(ctx context.Context, folderIDs, fileIDs []string, destFolderID string) error { func (c *Client) CopyDavItems(ctx context.Context, folderIDs, fileIDs []string, destFolderID string) error {
_, err := c.putJSON(ctx, "/api/2.0/files/fileops/copy", map[string]any{ raw, err := c.putJSON(ctx, "/api/2.0/files/fileops/copy", map[string]any{
"folderIds": nums(folderIDs), "folderIds": nums(folderIDs),
"fileIds": nums(fileIDs), "fileIds": nums(fileIDs),
"destFolderId": num(destFolderID), "destFolderId": num(destFolderID),
"conflictResolveType": "Skip", "conflictResolveType": "Skip",
"deleteAfter": true, "deleteAfter": true,
}) })
if err != nil {
return err return err
}
return fileopsError(raw)
}
// ListFileOps returns the currently active file operations
// (GET /api/2.0/files/fileops) for status polling.
func (c *Client) ListFileOps(ctx context.Context) ([]map[string]any, error) {
raw, err := c.getJSON(ctx, "/api/2.0/files/fileops")
if err != nil {
return nil, err
}
resp, err := responseField(raw, "response")
if err != nil {
return nil, err
}
if len(resp) == 0 || string(resp) == "null" {
return nil, nil
}
var ops []map[string]any
if err := json.Unmarshal(resp, &ops); err != nil {
return nil, err
}
return ops, nil
}
// fileopsError extracts per-operation "error" strings from a fileops/move or
// fileops/copy envelope. A 200 with error entries means nothing moved.
func fileopsError(raw json.RawMessage) error {
resp, err := responseField(raw, "response")
if err != nil {
return err
}
var ops []struct {
Error *string `json:"error"`
Finished *bool `json:"finished"`
Progress *int `json:"progress"`
}
if err := json.Unmarshal(resp, &ops); err != nil {
return nil // not an operations envelope — nothing to report
}
var errs []string
for _, op := range ops {
if op.Error != nil && *op.Error != "" {
errs = append(errs, *op.Error)
}
}
if len(errs) > 0 {
return fmt.Errorf("onlyoffice: fileops: %s", strings.Join(errs, "; "))
}
return nil
} }
// DeleteDavItems deletes the given folders and/or files. // DeleteDavItems deletes the given folders and/or files.
//
// Deprecated: use FileStore.Delete via Client.Files()/Client.FileStore.
func (c *Client) DeleteDavItems(ctx context.Context, folderIDs, fileIDs []string) error { func (c *Client) DeleteDavItems(ctx context.Context, folderIDs, fileIDs []string) error {
body := map[string]any{"DeleteAfter": true, "Immediately": false} body := map[string]any{"DeleteAfter": true, "Immediately": true}
for _, id := range folderIDs { for _, id := range folderIDs {
if _, err := c.deleteJSON(ctx, "/api/2.0/files/folder/"+url.PathEscape(id), body); err != nil { if err := c.deleteDavItem(ctx, "/api/2.0/files/folder/"+url.PathEscape(id), body); err != nil {
return err return err
} }
} }
for _, id := range fileIDs { for _, id := range fileIDs {
if _, err := c.deleteJSON(ctx, "/api/2.0/files/file/"+url.PathEscape(id), body); err != nil { if err := c.deleteDavItem(ctx, "/api/2.0/files/file/"+url.PathEscape(id), body); err != nil {
return err return err
} }
} }
return nil return nil
} }
// deleteDavItem deletes one item, retrying transient 429/502/503/504 answers
// through DoRetry like every other bulk path (deletes are idempotent).
func (c *Client) deleteDavItem(ctx context.Context, path string, body any) error {
return DoRetry(ctx, DefaultRetryPolicy(), func() error {
_, err := c.deleteJSON(ctx, path, body)
return err
})
}
// UploadDavFile uploads src (fileName) into folderID, streaming from src. // UploadDavFile uploads src (fileName) into folderID, streaming from src.
//
// Deprecated: use FileStore.Upload via Client.Files()/Client.FileStore.
func (c *Client) UploadDavFile(ctx context.Context, folderID, fileName string, src io.Reader) (*DavFile, error) { func (c *Client) UploadDavFile(ctx context.Context, folderID, fileName string, src io.Reader) (*DavFile, error) {
raw, err := c.uploadReader(ctx, "/api/2.0/files/"+url.PathEscape(folderID)+"/upload", "file", fileName, src) raw, err := c.uploadReader(ctx, "/api/2.0/files/"+url.PathEscape(folderID)+"/upload", "file", fileName, src)
if err != nil { if err != nil {
@@ -201,34 +284,12 @@ func (c *Client) UploadDavFile(ctx context.Context, folderID, fileName string, s
return env.Response, nil return env.Response, nil
} }
// DownloadDavFile streams the file identified by id to w, returning bytes copied. // DownloadDavFile streams the file identified by id to w, returning bytes
// copied. It shares the MinIO stale-S3 fallback with DownloadFile.
//
// Deprecated: use FileStore.Download via Client.Files()/Client.FileStore.
func (c *Client) DownloadDavFile(ctx context.Context, id string, w io.Writer) (int64, error) { func (c *Client) DownloadDavFile(ctx context.Context, id string, w io.Writer) (int64, error) {
file, err := c.GetFile(ctx, id) return c.DownloadFile(ctx, id, w)
if err != nil {
return 0, err
}
if file.ViewURL == nil || *file.ViewURL == "" {
return 0, fmt.Errorf("onlyoffice: file %s has no viewUrl", id)
}
u := c.resolveAPIURL(*file.ViewURL)
auth, err := c.authHeader()
if err != nil {
return 0, err
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
if err != nil {
return 0, err
}
req.Header.Set("Authorization", auth)
resp, err := c.client.Do(req)
if err != nil {
return 0, err
}
defer resp.Body.Close()
if resp.StatusCode >= 400 {
return 0, fmt.Errorf("onlyoffice: download: %d", resp.StatusCode)
}
return io.Copy(w, resp.Body)
} }
// --- internal helpers ------------------------------------------------------- // --- internal helpers -------------------------------------------------------
@@ -264,7 +325,7 @@ func (c *Client) deleteJSON(ctx context.Context, path string, body any) (json.Ra
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("DELETE %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "DELETE %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
@@ -304,7 +365,7 @@ func (c *Client) uploadReader(ctx context.Context, path, fieldName, fileName str
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("upload %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "upload %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
+285
View File
@@ -0,0 +1,285 @@
//go:build integration
package onlyoffice
import (
"bytes"
"context"
"strconv"
"testing"
"time"
)
// TestIntegrationFileStores runs the same operation set through the REST and
// DAV FileStore adapters against a throwaway project Documents folder: file
// create/upload/list/stat/download/move/copy/rename/delete and folder
// create/stat/list/rename/move/delete. Destructive — only run against
// instances you own.
//
// The Documents fileops API is asynchronous: a move/copy/delete is accepted
// immediately and becomes visible a moment later, so effects are polled.
func TestIntegrationFileStores(t *testing.T) {
c := liveClient(t)
t.Cleanup(func() { cleanupTestProjects(t, c) })
ctx := context.Background()
suffix := time.Now().UTC().Format("20060102-150405")
project, err := c.CreateProject(NewProjectRequest{
Title: testProjectPrefix + "store-" + suffix,
Description: "go-onlyoffice file store integration",
})
if err != nil {
t.Fatalf("CreateProject: %v", err)
}
if project.ID == nil {
t.Fatal("created project without id")
}
root, err := c.projectFolderID(ctx, strconv.Itoa(*project.ID))
if err != nil {
t.Fatalf("projectFolderID: %v", err)
}
for _, backend := range []string{ProviderREST, ProviderDAV} {
t.Run(backend, func(t *testing.T) {
testFileStoreOps(t, ctx, c, c.FileStore(backend), root, suffix)
})
}
}
func testFileStoreOps(t *testing.T, ctx context.Context, c *Client, store FileStore, root, suffix string) {
t.Helper()
content := []byte("file store " + store.Name() + " " + suffix + "\n")
src, err := store.CreateFolder(ctx, root, "fs-src-"+suffix)
if err != nil {
t.Fatalf("CreateFolder src: %v", err)
}
if src.Kind != Folder || src.ID == "" {
t.Fatalf("created src folder: %+v", src)
}
dst, err := store.CreateFolder(ctx, root, "fs-dst-"+suffix)
if err != nil {
t.Fatalf("CreateFolder dst: %v", err)
}
if dst.Kind != Folder || dst.ID == "" {
t.Fatalf("created dst folder: %+v", dst)
}
t.Cleanup(func() {
if err := c.DeleteDavItems(ctx, []string{src.ID, dst.ID}, nil); err != nil {
t.Logf("cleanup folders: %v", err)
}
})
up, err := store.Upload(ctx, src.ID, "doc-"+suffix+".txt", bytes.NewReader(content))
if err != nil {
t.Fatalf("Upload: %v", err)
}
if up.Kind != File || up.ID == "" {
t.Fatalf("uploaded entry: %+v", up)
}
if !waitEntry(ctx, store, src.ID, up.ID, 15*time.Second) {
t.Fatalf("uploaded %s not listed in src", up.ID)
}
st, err := store.Stat(ctx, up.ID)
if err != nil {
t.Fatalf("Stat: %v", err)
}
if st.ID != up.ID || st.Kind != File {
t.Fatalf("stat = %+v", st)
}
var buf bytes.Buffer
n, err := store.Download(ctx, up.ID, &buf)
if err != nil {
t.Fatalf("Download: %v", err)
}
if n != int64(len(content)) || !bytes.Equal(buf.Bytes(), content) {
t.Fatalf("download mismatch: got %d bytes %q want %d", n, buf.String(), len(content))
}
moveEventually(t, ctx, store, up.ID, dst.ID)
if !waitEntry(ctx, store, dst.ID, up.ID, 20*time.Second) {
t.Fatalf("moved file %s not in dst", up.ID)
}
copied := copyEventually(t, ctx, store, up.ID, src.ID, 20*time.Second)
if copied == nil {
t.Fatalf("no copy found in src after Copy")
}
newTitle := "renamed-" + suffix + ".txt"
renameEventually(t, ctx, store, up.ID, newTitle)
if err := store.Delete(ctx, []string{up.ID, copied.ID}); err != nil {
t.Fatalf("Delete: %v", err)
}
if !waitNoEntry(ctx, store, dst.ID, up.ID, 20*time.Second) {
t.Fatalf("file %s still present in dst after delete", up.ID)
}
if !waitNoEntry(ctx, store, src.ID, copied.ID, 20*time.Second) {
t.Fatalf("copy %s still present in src after delete", copied.ID)
}
// --- CRUD on the folders themselves, reusing the throwaway src/dst ---
// A child file lets us prove it survives the folder rename and move.
child, err := store.Upload(ctx, src.ID, "child-"+suffix+".txt", bytes.NewReader(content))
if err != nil {
t.Fatalf("Upload child: %v", err)
}
if !waitEntry(ctx, store, src.ID, child.ID, 15*time.Second) {
t.Fatalf("child %s not listed in src", child.ID)
}
fst, err := store.Stat(ctx, src.ID)
if err != nil {
t.Fatalf("Stat(folder): %v", err)
}
if fst.ID != src.ID || fst.Kind != Folder {
t.Fatalf("Stat(folder) = %+v", fst)
}
flist, err := store.List(ctx, src.ID)
if err != nil {
t.Fatalf("List(folder): %v", err)
}
if entryByID(flist, child.ID) == nil {
t.Fatalf("child %s not in List(src)", child.ID)
}
folderTitle := "renamed-folder-" + suffix
renameEventually(t, ctx, store, src.ID, folderTitle)
if e, err := store.Stat(ctx, src.ID); err != nil {
t.Fatalf("Stat(folder) after rename: %v", err)
} else if e.Kind != Folder || e.Title != folderTitle {
t.Fatalf("folder after rename = %+v, want title %q", e, folderTitle)
}
moveEventually(t, ctx, store, src.ID, dst.ID)
if !waitEntry(ctx, store, dst.ID, src.ID, 20*time.Second) {
t.Fatalf("moved folder %s not in dst %s", src.ID, dst.ID)
}
if !waitNoEntry(ctx, store, root, src.ID, 20*time.Second) {
t.Fatalf("folder %s still in root after move", src.ID)
}
if !waitEntry(ctx, store, src.ID, child.ID, 20*time.Second) {
t.Fatalf("child file %s lost after moving folder %s", child.ID, src.ID)
}
if err := store.Delete(ctx, []string{child.ID}); err != nil {
t.Fatalf("Delete(child): %v", err)
}
if err := store.Delete(ctx, []string{src.ID}); err != nil {
t.Fatalf("Delete(folder): %v", err)
}
if !waitNoEntry(ctx, store, dst.ID, src.ID, 20*time.Second) {
t.Fatalf("folder %s still present in dst after delete", src.ID)
}
if err := store.Delete(ctx, []string{dst.ID}); err != nil {
t.Fatalf("Delete(dst folder): %v", err)
}
if !waitNoEntry(ctx, store, root, dst.ID, 20*time.Second) {
t.Fatalf("dst folder %s still present in root after delete", dst.ID)
}
}
// moveEventually issues Move and retries while the operation is not visible yet
// (the fileops API accepts asynchronously and occasionally rejects a move that
// raced the just-finished upload).
func moveEventually(t *testing.T, ctx context.Context, store FileStore, id, dstID string) {
t.Helper()
var lastErr error
for attempt := 0; attempt < 5; attempt++ {
if lastErr = store.Move(ctx, []string{id}, dstID); lastErr == nil {
if waitEntry(ctx, store, dstID, id, 6*time.Second) {
return
}
}
time.Sleep(time.Second)
}
t.Fatalf("Move %s -> %s: %v", id, dstID, lastErr)
}
// copyEventually issues Copy and retries while the new copy is not visible yet
// (copy is accepted asynchronously, like move).
func copyEventually(t *testing.T, ctx context.Context, store FileStore, id, dstID string, d time.Duration) *Entry {
t.Helper()
var lastErr error
for attempt := 0; attempt < 5; attempt++ {
if lastErr = store.Copy(ctx, []string{id}, dstID); lastErr == nil {
if e := waitOtherFile(ctx, store, dstID, id, d); e != nil {
return e
}
}
time.Sleep(time.Second)
}
t.Fatalf("Copy %s -> %s: %v", id, dstID, lastErr)
return nil
}
func renameEventually(t *testing.T, ctx context.Context, store FileStore, id, title string) {
t.Helper()
var lastErr error
for attempt := 0; attempt < 5; attempt++ {
if lastErr = store.Rename(ctx, id, title); lastErr == nil {
if e, err := store.Stat(ctx, id); err == nil && e.Title == title {
return
}
}
time.Sleep(time.Second)
}
t.Fatalf("Rename %s -> %q: %v", id, title, lastErr)
}
func waitEntry(ctx context.Context, store FileStore, parentID, id string, d time.Duration) bool {
deadline := time.Now().Add(d)
for time.Now().Before(deadline) {
if list, err := store.List(ctx, parentID); err == nil && entryByID(list, id) != nil {
return true
}
time.Sleep(500 * time.Millisecond)
}
return false
}
func waitNoEntry(ctx context.Context, store FileStore, parentID, id string, d time.Duration) bool {
deadline := time.Now().Add(d)
for time.Now().Before(deadline) {
if list, err := store.List(ctx, parentID); err == nil && entryByID(list, id) == nil {
return true
}
time.Sleep(500 * time.Millisecond)
}
return false
}
func waitOtherFile(ctx context.Context, store FileStore, parentID, id string, d time.Duration) *Entry {
deadline := time.Now().Add(d)
for time.Now().Before(deadline) {
if list, err := store.List(ctx, parentID); err == nil {
if e := firstFileOtherThan(list, id); e != nil {
return e
}
}
time.Sleep(500 * time.Millisecond)
}
return nil
}
func entryByID(entries []Entry, id string) *Entry {
for i := range entries {
if entries[i].ID == id {
return &entries[i]
}
}
return nil
}
func firstFileOtherThan(entries []Entry, id string) *Entry {
for i := range entries {
if entries[i].Kind == File && entries[i].ID != id {
return &entries[i]
}
}
return nil
}
+262
View File
@@ -0,0 +1,262 @@
package onlyoffice
// Canonical file model and the backend-agnostic store interface. REST
// (files.go), WebDAV (files_webdav.go) and future backends (PostgreSQL,
// Elasticsearch) implement FileStore/Searcher so callers stop depending on a
// concrete transport. This file holds only types and pure conversions — no IO.
import (
"context"
"io"
"mime"
"path/filepath"
"strconv"
"strings"
"time"
)
// Kind distinguishes files from folders in the canonical model.
type Kind int
const (
File Kind = iota
Folder
)
// String renders the kind for logs and table output.
func (k Kind) String() string {
switch k {
case File:
return "file"
case Folder:
return "folder"
default:
return "unknown"
}
}
// Provider names for the FileStore adapters.
const (
ProviderREST = "rest"
ProviderDAV = "dav"
)
// Entry is the backend-independent representation of a document or folder.
// Fields that a backend cannot supply stay at their zero value.
type Entry struct {
ID string
ParentID string
Title string
Kind Kind
Size int64
MIME string
Created time.Time
Modified time.Time
// Updated is the backend-native timestamp string, when the backend exposes
// one. It lets list output round-trip the API value; Modified is the
// parsed form for logic.
Updated string
Version int
Provider string
// Folder-only counters. Zero for files and for backends that do not
// report them.
FilesCount int
FoldersCount int
}
// FileStore is the operation surface every file backend implements.
type FileStore interface {
Name() string
List(ctx context.Context, parentID string) ([]Entry, error)
Stat(ctx context.Context, id string) (Entry, error)
CreateFolder(ctx context.Context, parentID, title string) (Entry, error)
Upload(ctx context.Context, parentID, title string, r io.Reader) (Entry, error)
Download(ctx context.Context, id string, w io.Writer) (int64, error)
Move(ctx context.Context, ids []string, parentID string) error
Copy(ctx context.Context, ids []string, parentID string) error
Rename(ctx context.Context, id, title string) error
Delete(ctx context.Context, ids []string) error
}
// SearchQuery narrows a Searcher request. InContent asks the backend to match
// document bodies, not just titles. Substring switches title matching from the
// analyzer's whole-token match to a case-insensitive "*term*" wildcard and ANDs
// every whitespace-separated term (e.g. "rechnung 2025").
type SearchQuery struct {
Text string
InContent bool
FolderID string
Extensions []string
Limit int
Substring bool
}
// SearchHit is one Searcher result: the matching entry plus backend-specific
// ranking metadata.
type SearchHit struct {
Entry
Score float64
Highlight string
Path []string
}
// Searcher is the optional content/name search surface. Only some backends
// (for example Elasticsearch) provide it.
type Searcher interface {
Search(ctx context.Context, q SearchQuery) ([]SearchHit, error)
Name() string
}
// FileStore returns the adapter for a backend name: ProviderREST (default),
// ProviderDAV (alias "webdav") or the read-only SQL store (ProviderPG,
// ProviderMySQL and the aliases "pg"/"sql"). The SQL store is opened from the
// environment (ONLYOFFICE_DSN / ONLYOFFICE_PG_*); when it cannot be opened the
// returned store surfaces that error on every operation instead of returning
// nil. Use SQLFileStore when the open error itself is needed. Unknown or empty
// names select the REST backend. The composed facade (backend
// selection/fallback) lives on FileClient in filestore_facade.go.
func (c *Client) FileStore(backend string) FileStore {
switch strings.ToLower(strings.TrimSpace(backend)) {
case ProviderDAV, "webdav":
return &davStore{c: c}
case ProviderPG, ProviderMySQL, "pg", "sql":
s, err := c.SQLFileStore()
if err != nil {
return &errStore{name: strings.ToLower(strings.TrimSpace(backend)), err: err}
}
return s
default:
return &restStore{c: c}
}
}
// errStore is the FileStore placeholder returned when a backend cannot be
// opened (for example SQL without a DSN). Every operation returns the recorded
// error instead of panicking on a nil interface.
type errStore struct {
name string
err error
}
func (s *errStore) Name() string { return s.name }
func (s *errStore) List(context.Context, string) ([]Entry, error) { return nil, s.err }
func (s *errStore) Stat(context.Context, string) (Entry, error) { return Entry{}, s.err }
func (s *errStore) CreateFolder(context.Context, string, string) (Entry, error) {
return Entry{}, s.err
}
func (s *errStore) Upload(context.Context, string, string, io.Reader) (Entry, error) {
return Entry{}, s.err
}
func (s *errStore) Download(context.Context, string, io.Writer) (int64, error) {
return 0, s.err
}
func (s *errStore) Move(context.Context, []string, string) error { return s.err }
func (s *errStore) Copy(context.Context, []string, string) error { return s.err }
func (s *errStore) Rename(context.Context, string, string) error { return s.err }
func (s *errStore) Delete(context.Context, []string) error { return s.err }
// Files returns the composed file facade. The returned *FileClient implements
// FileStore, so callers that used Files() as the plain REST store keep working.
func (c *Client) Files() *FileClient { return c.newFileClient() }
// retryStoreOp runs one store operation under the shared deterministic
// transient-error policy (429/502/503/504).
func retryStoreOp(ctx context.Context, fn func() error) error {
return DoRetry(ctx, DefaultRetryPolicy(), fn)
}
// FileEntryToEntry converts a Files-module file row to the canonical model.
func FileEntryToEntry(f *FileEntry, provider string) Entry {
e := Entry{Kind: File, Provider: provider}
if f == nil {
return e
}
if f.ID != nil {
e.ID = f.ID.String()
}
e.ParentID = FileFolderID(f)
if f.Title != nil {
e.Title = *f.Title
}
if f.ContentLength != nil {
e.Size = parseContentLength(*f.ContentLength)
}
exst := ""
if f.FileExst != nil {
exst = *f.FileExst
}
e.MIME = mimeForTitle(e.Title, exst)
if f.Updated != nil {
e.Modified = *f.Updated
e.Updated = f.Updated.Format(time.RFC3339)
}
return e
}
// DavFileToEntry converts a WebDAV file row to the canonical model.
func DavFileToEntry(f DavFile, provider string) Entry {
return Entry{
ID: f.ID,
Title: f.Title,
Kind: File,
Size: f.Size,
MIME: mimeForTitle(f.Title, ""),
Modified: f.ModTime(),
Updated: f.Updated,
Provider: provider,
}
}
// DavFolderToEntry converts a WebDAV folder row to the canonical model.
func DavFolderToEntry(f DavFolder, provider string) Entry {
return Entry{
ID: f.ID,
ParentID: f.ParentID,
Title: f.Title,
Kind: Folder,
Modified: f.ModTime(),
Updated: f.Updated,
Provider: provider,
FilesCount: f.FilesCount,
FoldersCount: f.FoldersCount,
}
}
// parseContentLength reads the leading integer of an OnlyOffice contentLength
// string (the API sometimes appends a unit, e.g. "12345 b").
func parseContentLength(s string) int64 {
fields := strings.Fields(s)
if len(fields) == 0 {
return 0
}
n, err := strconv.ParseInt(fields[0], 10, 64)
if err != nil {
return 0
}
return n
}
// mimeForTitle derives a MIME type from an explicit extension or the title.
func mimeForTitle(title, exst string) string {
ext := strings.TrimSpace(exst)
if ext == "" {
ext = filepath.Ext(title)
}
if ext == "" {
return ""
}
if !strings.HasPrefix(ext, ".") {
ext = "." + ext
}
return mime.TypeByExtension(strings.ToLower(ext))
}
+193
View File
@@ -0,0 +1,193 @@
package onlyoffice
import (
"encoding/json"
"strings"
"testing"
"time"
)
func TestFileEntryToEntry(t *testing.T) {
id := json.Number("42")
title := "invoice.pdf"
exst := ".pdf"
size := "12345"
parent := json.Number("7")
updated := time.Date(2026, 1, 2, 3, 4, 5, 0, time.UTC)
f := &FileEntry{
ID: &id,
Title: &title,
FileExst: &exst,
ContentLength: &size,
FolderID: &parent,
Updated: &updated,
}
e := FileEntryToEntry(f, ProviderREST)
if e.ID != "42" {
t.Errorf("ID = %q, want 42", e.ID)
}
if e.ParentID != "7" {
t.Errorf("ParentID = %q, want 7", e.ParentID)
}
if e.Title != title {
t.Errorf("Title = %q, want %q", e.Title, title)
}
if e.Kind != File {
t.Errorf("Kind = %v, want file", e.Kind)
}
if e.Size != 12345 {
t.Errorf("Size = %d, want 12345", e.Size)
}
if e.MIME != "application/pdf" {
t.Errorf("MIME = %q, want application/pdf", e.MIME)
}
if !e.Modified.Equal(updated) {
t.Errorf("Modified = %v, want %v", e.Modified, updated)
}
if e.Provider != ProviderREST {
t.Errorf("Provider = %q, want %q", e.Provider, ProviderREST)
}
}
func TestFileEntryToEntryNil(t *testing.T) {
e := FileEntryToEntry(nil, ProviderDAV)
if e.Kind != File {
t.Errorf("Kind = %v, want file", e.Kind)
}
if e.ID != "" || e.Title != "" {
t.Errorf("nil entry should be empty: %+v", e)
}
if e.Provider != ProviderDAV {
t.Errorf("Provider = %q, want %q", e.Provider, ProviderDAV)
}
}
func TestFileEntryToEntrySizeFormats(t *testing.T) {
cases := map[string]int64{
"12345": 12345,
"12345 b": 12345,
"0": 0,
"": 0,
"notanum": 0,
}
for in, want := range cases {
got := parseContentLength(in)
if got != want {
t.Errorf("parseContentLength(%q) = %d, want %d", in, got, want)
}
}
}
func TestDavFileToEntry(t *testing.T) {
f := DavFile{
ID: "9",
Title: "note.txt",
Size: 10,
Updated: "2026-01-02T03:04:05.0000000+01:00",
}
e := DavFileToEntry(f, ProviderDAV)
if e.ID != "9" || e.Title != "note.txt" {
t.Errorf("identity mismatch: %+v", e)
}
if e.Kind != File {
t.Errorf("Kind = %v, want file", e.Kind)
}
if e.Size != 10 {
t.Errorf("Size = %d, want 10", e.Size)
}
if !strings.HasPrefix(e.MIME, "text/plain") {
t.Errorf("MIME = %q, want text/plain*", e.MIME)
}
if e.Modified.IsZero() {
t.Error("Modified not parsed")
}
if e.Provider != ProviderDAV {
t.Errorf("Provider = %q, want %q", e.Provider, ProviderDAV)
}
}
func TestDavFolderToEntry(t *testing.T) {
f := DavFolder{
ID: "5",
Title: "inbox",
ParentID: "1",
Updated: "2026-01-02T03:04:05.0000000+01:00",
}
e := DavFolderToEntry(f, ProviderDAV)
if e.ID != "5" || e.Title != "inbox" || e.ParentID != "1" {
t.Errorf("identity mismatch: %+v", e)
}
if e.Kind != Folder {
t.Errorf("Kind = %v, want folder", e.Kind)
}
if e.MIME != "" {
t.Errorf("folder MIME = %q, want empty", e.MIME)
}
if e.Modified.IsZero() {
t.Error("Modified not parsed")
}
}
func TestEntriesFromFolderMap(t *testing.T) {
m := map[string]any{
"files": []any{
map[string]any{"id": float64(42), "title": "a.pdf", "pureContentLength": float64(7)},
},
"folders": []any{
map[string]any{"id": float64(7), "title": "sub", "parentId": float64(1)},
},
}
entries, err := entriesFromFolderMap(m, ProviderREST)
if err != nil {
t.Fatalf("entriesFromFolderMap: %v", err)
}
if len(entries) != 2 {
t.Fatalf("got %d entries, want 2: %+v", len(entries), entries)
}
byID := map[string]Entry{}
for _, e := range entries {
byID[e.ID] = e
}
if got := byID["42"]; got.Kind != File || got.Size != 7 || got.Title != "a.pdf" {
t.Errorf("file entry = %+v", got)
}
if got := byID["7"]; got.Kind != Folder || got.ParentID != "1" || got.Title != "sub" {
t.Errorf("folder entry = %+v", got)
}
}
func TestEntriesFromFolderMapNil(t *testing.T) {
entries, err := entriesFromFolderMap(nil, ProviderREST)
if err != nil || entries != nil {
t.Fatalf("got %v, %v; want nil, nil", entries, err)
}
}
func TestKindString(t *testing.T) {
if File.String() != "file" || Folder.String() != "folder" {
t.Errorf("kind strings: %q %q", File.String(), Folder.String())
}
if Kind(9).String() != "unknown" {
t.Errorf("unknown kind = %q", Kind(9).String())
}
}
func TestClientFileStoreSelection(t *testing.T) {
c := NewClient(Credentials{})
if got := c.FileStore(ProviderDAV).Name(); got != ProviderDAV {
t.Errorf("FileStore(dav).Name() = %q", got)
}
if got := c.FileStore("webdav").Name(); got != ProviderDAV {
t.Errorf("FileStore(webdav).Name() = %q", got)
}
if got := c.FileStore(ProviderREST).Name(); got != ProviderREST {
t.Errorf("FileStore(rest).Name() = %q", got)
}
if got := c.FileStore("").Name(); got != ProviderREST {
t.Errorf("FileStore(\"\").Name() = %q", got)
}
if got := c.Files().Name(); got != ProviderREST {
t.Errorf("Files().Name() = %q", got)
}
}
+189
View File
@@ -0,0 +1,189 @@
package onlyoffice
// davStore implements FileStore on top of the Documents/WebDAV methods in
// files_webdav.go. The Documents fileops calls need folder and file ids
// separated, so ids are classified through Stat before move/copy/rename/delete.
import (
"bytes"
"context"
"fmt"
"io"
)
// davStore is a FileStore over the WebDAV-oriented Documents API.
type davStore struct{ c *Client }
// Name reports the backend name.
func (s *davStore) Name() string { return ProviderDAV }
// List returns the files and folders directly below parentID.
func (s *davStore) List(ctx context.Context, parentID string) ([]Entry, error) {
var out []Entry
err := retryStoreOp(ctx, func() error {
l, err := s.c.ListDavFolder(ctx, parentID)
if err != nil {
return err
}
entries := make([]Entry, 0, len(l.Folders)+len(l.Files))
for _, f := range l.Folders {
entries = append(entries, DavFolderToEntry(f, ProviderDAV))
}
for _, f := range l.Files {
entries = append(entries, DavFileToEntry(f, ProviderDAV))
}
out = entries
return nil
})
return out, err
}
// Stat resolves a folder or file entry by id. A folder answers ListDavFolder
// with its own metadata in Current; otherwise the file metadata API is used.
func (s *davStore) Stat(ctx context.Context, id string) (Entry, error) {
return s.stat(ctx, id)
}
// CreateFolder creates a subfolder under parentID.
func (s *davStore) CreateFolder(ctx context.Context, parentID, title string) (Entry, error) {
var out Entry
err := retryStoreOp(ctx, func() error {
f, err := s.c.CreateDavFolder(ctx, parentID, title)
if err != nil {
return err
}
if f == nil {
return fmt.Errorf("onlyoffice: dav store: empty create-folder response")
}
out = DavFolderToEntry(*f, ProviderDAV)
return nil
})
return out, err
}
// Upload streams r into parentID as title. The reader is buffered once so a
// retry re-sends the same bytes instead of an exhausted stream.
func (s *davStore) Upload(ctx context.Context, parentID, title string, r io.Reader) (Entry, error) {
data, err := io.ReadAll(r)
if err != nil {
return Entry{}, err
}
var out Entry
err = retryStoreOp(ctx, func() error {
f, err := s.c.UploadDavFile(ctx, parentID, title, bytes.NewReader(data))
if err != nil {
return err
}
if f == nil {
return fmt.Errorf("onlyoffice: dav store: empty upload response")
}
out = DavFileToEntry(*f, ProviderDAV)
return nil
})
return out, err
}
// Download streams the file bytes into w.
func (s *davStore) Download(ctx context.Context, id string, w io.Writer) (int64, error) {
var n int64
err := retryStoreOp(ctx, func() error {
var e error
n, e = s.c.DownloadDavFile(ctx, id, w)
return e
})
return n, err
}
// Move moves ids into parentID, splitting folders from files.
func (s *davStore) Move(ctx context.Context, ids []string, parentID string) error {
folders, files, err := s.split(ctx, ids)
if err != nil {
return err
}
if len(folders) == 0 && len(files) == 0 {
return nil
}
return retryStoreOp(ctx, func() error {
return s.c.MoveDavItems(ctx, folders, files, parentID)
})
}
// Copy copies ids into parentID, splitting folders from files.
func (s *davStore) Copy(ctx context.Context, ids []string, parentID string) error {
folders, files, err := s.split(ctx, ids)
if err != nil {
return err
}
if len(folders) == 0 && len(files) == 0 {
return nil
}
return retryStoreOp(ctx, func() error {
return s.c.CopyDavItems(ctx, folders, files, parentID)
})
}
// Rename renames a folder or file.
func (s *davStore) Rename(ctx context.Context, id, title string) error {
e, err := s.stat(ctx, id)
if err != nil {
return err
}
return retryStoreOp(ctx, func() error {
if e.Kind == Folder {
return s.c.RenameDavFolder(ctx, id, title)
}
return s.c.RenameDavFile(ctx, id, title)
})
}
// Delete removes ids, splitting folders from files.
func (s *davStore) Delete(ctx context.Context, ids []string) error {
folders, files, err := s.split(ctx, ids)
if err != nil {
return err
}
if len(folders) == 0 && len(files) == 0 {
return nil
}
return retryStoreOp(ctx, func() error {
return s.c.DeleteDavItems(ctx, folders, files)
})
}
// stat resolves a single id to a folder or file Entry.
func (s *davStore) stat(ctx context.Context, id string) (Entry, error) {
var out Entry
err := retryStoreOp(ctx, func() error {
if l, err := s.c.ListDavFolder(ctx, id); err == nil {
if l != nil && l.Current.ID != "" && l.Current.ID == id {
out = DavFolderToEntry(l.Current, ProviderDAV)
return nil
}
} else if Transient(err) {
return err
}
f, err := s.c.GetFile(ctx, id)
if err != nil {
return err
}
out = FileEntryToEntry(f, ProviderDAV)
return nil
})
return out, err
}
// split classifies ids into folder and file id lists.
func (s *davStore) split(ctx context.Context, ids []string) (folders, files []string, err error) {
for _, id := range ids {
e, err := s.stat(ctx, id)
if err != nil {
return nil, nil, err
}
if e.Kind == Folder {
folders = append(folders, id)
} else {
files = append(files, id)
}
}
return folders, files, nil
}
+325
View File
@@ -0,0 +1,325 @@
package onlyoffice
// Elasticsearch backend of the unified file client (epic #34, F3 #37).
//
// OnlyOffice full-text search runs on Elasticsearch (index `files_file`, NEST
// client on the server). The REST endpoint GET /api/2.0/files/@search/{query}
// only searches file names in the database, so content search needs a direct
// ES query. The live server is Elasticsearch 7.16.3; the request shape below
// is plain REST and stays stdlib-only, matching the repo's no-extra-deps rule.
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"regexp"
"strconv"
"strings"
"time"
)
// The canonical model (Kind, Entry, SearchQuery, SearchHit, Searcher) lives in
// filestore_core.go (F1 #35).
const (
defaultESIndex = "files_file"
defaultESLimit = 20
maxESLimit = 1000
maxESResponseSize = 8 << 20
)
// ESConfig configures the direct Elasticsearch searcher.
type ESConfig struct {
URL string // scheme://host:port of the ES HTTP endpoint
Index string // index name, default files_file
Tenant string // tenantId filter, empty means all tenants
}
// ESConfigFromEnv reads ONLYOFFICE_ES_URL, ONLYOFFICE_ES_INDEX (default
// files_file) and ONLYOFFICE_TENANT. The library never loads dotfiles — the
// CLI does that.
func ESConfigFromEnv() ESConfig {
return ESConfig{
URL: strings.TrimRight(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL")), "/"),
Index: firstNonEmpty(os.Getenv("ONLYOFFICE_ES_INDEX"), defaultESIndex),
Tenant: strings.TrimSpace(os.Getenv("ONLYOFFICE_TENANT")),
}
}
// ESSearcher queries OnlyOffice's Elasticsearch index directly for file name
// and document content.
type ESSearcher struct {
cfg ESConfig
http *http.Client
}
// NewESSearcher returns a searcher for the OnlyOffice Elasticsearch index.
// The URL is required; an empty index falls back to files_file.
func NewESSearcher(cfg ESConfig) (*ESSearcher, error) {
if strings.TrimSpace(cfg.URL) == "" {
return nil, fmt.Errorf("onlyoffice: elasticsearch URL is empty (set ONLYOFFICE_ES_URL)")
}
cfg.URL = strings.TrimRight(cfg.URL, "/")
if cfg.Index == "" {
cfg.Index = defaultESIndex
}
return &ESSearcher{cfg: cfg, http: &http.Client{Timeout: 30 * time.Second}}, nil
}
// Name implements Searcher.
func (s *ESSearcher) Name() string { return "elasticsearch" }
// Search runs a multi_match over title (and, when q.InContent is set,
// document.attachment.content), filtered by tenant and optional folder.
func (s *ESSearcher) Search(ctx context.Context, q SearchQuery) ([]SearchHit, error) {
q.Text = strings.TrimSpace(q.Text)
if q.Text == "" {
return nil, fmt.Errorf("onlyoffice: empty search query")
}
body, err := json.Marshal(esSearchRequest(q, s.cfg.Tenant))
if err != nil {
return nil, fmt.Errorf("onlyoffice: build elasticsearch query: %w", err)
}
endpoint := s.cfg.URL + "/" + s.cfg.Index + "/_search"
req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Accept", "application/json")
resp, err := s.http.Do(req)
if err != nil {
return nil, fmt.Errorf("onlyoffice: elasticsearch search: %w", err)
}
defer resp.Body.Close()
raw, err := io.ReadAll(io.LimitReader(resp.Body, maxESResponseSize))
if err != nil {
return nil, err
}
if resp.StatusCode >= 400 {
return nil, fmt.Errorf("onlyoffice: elasticsearch search: %d %s", resp.StatusCode, truncate(string(raw), 400))
}
return parseESSearchResponse(raw)
}
// esSearchRequest builds the ES query body. Pure, so it is unit-tested.
func esSearchRequest(q SearchQuery, tenant string) esRequest {
limit := q.Limit
if limit <= 0 {
limit = defaultESLimit
}
if limit > maxESLimit {
limit = maxESLimit
}
fields := []string{"title^2"}
if q.InContent {
fields = append(fields, "document.attachment.content")
}
var must []esClause
if q.Substring {
for _, term := range strings.Fields(strings.ToLower(q.Text)) {
if term = escapeWildcard(term); term != "" {
must = append(must, esClause{Wildcard: map[string]any{"title": "*" + term + "*"}})
}
}
}
if len(must) == 0 {
must = []esClause{{MultiMatch: &esMultiMatch{Query: q.Text, Fields: fields}}}
}
var filter []esClause
if t := strings.TrimSpace(tenant); t != "" {
filter = append(filter, esClause{Term: map[string]any{"tenantId": numericOrString(t)}})
}
if f := strings.TrimSpace(q.FolderID); f != "" {
// folders is an ES nested field; a plain term on folders.folderId would
// not match. The stored Folders list holds every ancestor id, so
// filtering by a project root id scopes to its whole subtree.
filter = append(filter, esClause{Nested: &esNested{
Path: "folders",
Query: esNestedTerm{Term: map[string]any{"folders.folderId": f}},
}})
}
for _, ext := range normalizeExtensions(q.Extensions) {
filter = append(filter, esClause{Wildcard: map[string]any{"title": "*." + ext}})
}
highlightFields := map[string]struct{}{"title": {}}
if q.InContent {
highlightFields["document.attachment.content"] = struct{}{}
}
return esRequest{
Size: limit,
Source: []string{"id", "title", "folders"},
Query: esQuery{Bool: esBool{Must: must, Filter: filter}},
Highlight: esHighlight{PreTags: []string{"<em>"}, PostTags: []string{"</em>"}, Fields: highlightFields},
}
}
// normalizeExtensions lowercases, trims leading dots and drops empties.
func normalizeExtensions(exts []string) []string {
out := make([]string, 0, len(exts))
seen := map[string]bool{}
for _, e := range exts {
e = strings.ToLower(strings.TrimSpace(strings.TrimPrefix(strings.TrimSpace(e), ".")))
if e == "" || seen[e] {
continue
}
seen[e] = true
out = append(out, e)
}
return out
}
// numericOrString keeps an integer-looking filter value numeric (tenantId is
// a long) and leaves anything else as a string (folderId is a text token).
func numericOrString(s string) any {
if n, err := strconv.ParseInt(s, 10, 64); err == nil {
return n
}
return s
}
// esRequest is the subset of the ES query DSL this client emits.
type esRequest struct {
Size int `json:"size"`
Source []string `json:"_source"`
Query esQuery `json:"query"`
Highlight esHighlight `json:"highlight"`
}
type esQuery struct {
Bool esBool `json:"bool"`
}
type esBool struct {
Must []esClause `json:"must,omitempty"`
Filter []esClause `json:"filter,omitempty"`
}
type esClause struct {
MultiMatch *esMultiMatch `json:"multi_match,omitempty"`
Term map[string]any `json:"term,omitempty"`
Terms map[string]any `json:"terms,omitempty"`
Wildcard map[string]any `json:"wildcard,omitempty"`
Nested *esNested `json:"nested,omitempty"`
}
type esNested struct {
Path string `json:"path"`
Query esNestedTerm `json:"query"`
}
type esNestedTerm struct {
Term map[string]any `json:"term,omitempty"`
}
// escapeWildcard strips ES wildcard metacharacters from a user term so a query
// cannot inject wildcard syntax. Pure, so it is unit-tested.
func escapeWildcard(s string) string {
return strings.NewReplacer("*", "", "?", "", `\`, "").Replace(s)
}
type esMultiMatch struct {
Query string `json:"query"`
Fields []string `json:"fields"`
}
type esHighlight struct {
PreTags []string `json:"pre_tags,omitempty"`
PostTags []string `json:"post_tags,omitempty"`
Fields map[string]struct{} `json:"fields"`
}
// esResponse is the subset of an ES search response we consume.
type esResponse struct {
Took int `json:"took"`
Hits struct {
Total struct {
Value int `json:"value"`
Relation string `json:"relation"`
} `json:"total"`
Hits []esResponseHit `json:"hits"`
} `json:"hits"`
}
type esResponseHit struct {
ID string `json:"_id"`
Score float64 `json:"_score"`
Source struct {
ID int `json:"id"`
Title string `json:"title"`
Folders []struct {
FolderID string `json:"folderId"`
ID int `json:"id"`
} `json:"folders"`
} `json:"_source"`
Highlight map[string][]string `json:"highlight"`
}
// parseESSearchResponse converts an ES search response into SearchHit values.
// Pure, so it is unit-tested.
func parseESSearchResponse(raw []byte) ([]SearchHit, error) {
var r esResponse
if err := json.Unmarshal(raw, &r); err != nil {
return nil, fmt.Errorf("onlyoffice: decode elasticsearch response: %w", err)
}
hits := make([]SearchHit, 0, len(r.Hits.Hits))
for _, h := range r.Hits.Hits {
id := strconv.Itoa(h.Source.ID)
if h.Source.ID == 0 {
id = h.ID
}
// Folders is the ancestor breadcrumb in root → leaf order, so the last
// entry is the immediate parent (the previous "first" value was the
// project root, which made every result look like it lived in #522).
var parent string
path := make([]string, 0, len(h.Source.Folders))
for _, f := range h.Source.Folders {
if strings.TrimSpace(f.FolderID) == "" {
continue
}
path = append(path, f.FolderID)
}
if len(path) > 0 {
parent = path[len(path)-1]
}
hits = append(hits, SearchHit{
Entry: Entry{
ID: id,
ParentID: parent,
Title: h.Source.Title,
Kind: File,
Provider: "elasticsearch",
},
Score: h.Score,
Highlight: esHighlightText(h.Highlight),
Path: path,
})
}
return hits, nil
}
var esHighlightTag = regexp.MustCompile(`</?em[^>]*>`)
// esHighlightText flattens a highlight map into one plain-text snippet,
// preferring the content fragment over the title. It covers both the
// OnlyOffice content field and the own-index "content" field.
func esHighlightText(hl map[string][]string) string {
for _, key := range []string{"document.attachment.content", "content", "title"} {
frags := hl[key]
if len(frags) == 0 {
continue
}
clean := make([]string, 0, len(frags))
for _, f := range frags {
clean = append(clean, esHighlightTag.ReplaceAllString(f, ""))
}
return strings.Join(clean, " … ")
}
return ""
}
+226
View File
@@ -0,0 +1,226 @@
//go:build integration
package onlyoffice
import (
"context"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
"time"
"github.com/xuri/excelize/v2"
)
// Default known fixtures for TestIntegrationESFacadeUsesES on the live index.
const (
defaultESTestQuery = "Rechnung_986-2025.pdf"
defaultESTestSubstring = "rechnung 2025"
defaultESTestFolder = "522"
)
// TestIntegrationESFacadeUsesES proves that the public search path — `oo search`
// and Client.Files().Search() — really runs against the OnlyOffice
// Elasticsearch backend and not the REST @search endpoint, which only looks at
// file names in the database and is not a Searcher at all (see
// docs/elasticsearch.md). It pins the concrete backend and checks that a known
// document comes back with a non-empty id and folder path.
//
// Requires ONLYOFFICE_ES_URL only — the query never touches the REST API, so no
// OnlyOffice credentials are needed. Skips when it is missing. The fixture is
// overridable with ONLYOFFICE_ES_TEST_QUERY, ONLYOFFICE_ES_TEST_TITLE,
// ONLYOFFICE_ES_TEST_SUBSTRING and ONLYOFFICE_ES_TEST_FOLDER.
func TestIntegrationESFacadeUsesES(t *testing.T) {
esURL := strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL"))
if esURL == "" {
t.Skip("ONLYOFFICE_ES_URL not set — skipping Elasticsearch integration test")
}
query := firstNonEmpty(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_TEST_QUERY")), defaultESTestQuery)
wantTitle := firstNonEmpty(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_TEST_TITLE")), query)
substring := firstNonEmpty(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_TEST_SUBSTRING")), defaultESTestSubstring)
folder := firstNonEmpty(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_TEST_FOLDER")), defaultESTestFolder)
c := NewClient(Credentials{})
searcher, err := c.Files().Search()
if err != nil {
t.Fatalf("Files().Search(): %v", err)
}
if got := searcher.Name(); got != ProviderES {
t.Fatalf("searcher.Name() = %q, want %q (REST @search is not a Searcher)", got, ProviderES)
}
if _, ok := searcher.(*ESSearcher); !ok {
t.Fatalf("searcher = %T, want *ESSearcher (ES backend, not REST)", searcher)
}
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
// Known file name: multi_match over title, as `oo search <file>` does.
start := time.Now()
hits, err := searcher.Search(ctx, SearchQuery{Text: query, Limit: 20})
if err != nil {
t.Fatalf("Search(%q): %v", query, err)
}
t.Logf("ES facade query %q: %d hits in %s", query, len(hits), time.Since(start))
known := findHitByTitle(hits, wantTitle)
if known == nil {
t.Fatalf("query %q returned %d hits, none titled %q", query, len(hits), wantTitle)
}
if strings.TrimSpace(known.ID) == "" {
t.Errorf("hit %q has empty id", known.Title)
}
if len(known.Path) == 0 {
t.Errorf("hit %q has empty path", known.Title)
}
if known.Provider != ProviderES {
t.Errorf("hit provider = %q, want %q", known.Provider, ProviderES)
}
// Substring + folder subtree, as `oo search <terms> --substring --folder N`
// does: wildcard terms ANDed together, scoped to the folder's subtree.
start = time.Now()
subHits, err := searcher.Search(ctx, SearchQuery{Text: substring, Substring: true, FolderID: folder, Limit: 200})
if err != nil {
t.Fatalf("substring Search(%q, folder %s): %v", substring, folder, err)
}
t.Logf("ES facade substring %q folder %s: %d hits in %s", substring, folder, len(subHits), time.Since(start))
if len(subHits) == 0 {
t.Fatalf("substring query %q in folder %s returned no hits", substring, folder)
}
if findHitByTitle(subHits, wantTitle) == nil {
t.Errorf("substring query %q in folder %s did not return %q", substring, folder, wantTitle)
}
}
// findHitByTitle returns the first hit whose title matches, case-insensitively.
func findHitByTitle(hits []SearchHit, title string) *SearchHit {
for i := range hits {
if strings.EqualFold(strings.TrimSpace(hits[i].Title), title) {
return &hits[i]
}
}
return nil
}
// TestIntegrationESSearch uploads a throwaway workbook and verifies that the
// direct Elasticsearch search finds it by file name and by content.
//
// The content index (document.attachment.content) is only populated for Office
// formats (docx/xlsx/pptx), so the fixture is an xlsx whose cell carries a
// unique token. Requires ONLYOFFICE_ES_URL (a reachable ES endpoint — in the
// current setup a tunnel to the ES inside the OnlyOffice VM, see
// docs/elasticsearch.md) plus the regular REST credentials for the upload.
// Skips when either is missing.
func TestIntegrationESSearch(t *testing.T) {
esURL := strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL"))
if esURL == "" {
t.Skip("ONLYOFFICE_ES_URL not set — skipping Elasticsearch integration test")
}
c := liveClient(t)
t.Cleanup(func() { cleanupTestProjects(t, c) })
stamp := time.Now().UTC().Format("20060102-150405")
nameToken := "goesname" + stamp
contentToken := "goescontent" + stamp
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
defer cancel()
project, err := c.CreateProject(NewProjectRequest{
Title: testProjectPrefix + "es-" + stamp,
Description: "go-onlyoffice elasticsearch integration",
})
if err != nil {
t.Fatalf("CreateProject: %v", err)
}
if project.ID == nil {
t.Fatal("created project without id")
}
pid := strconv.Itoa(*project.ID)
title := nameToken + ".xlsx"
localPath := filepath.Join(t.TempDir(), title)
book := excelize.NewFile()
if err := book.SetCellValue("Sheet1", "A1", "OnlyOffice Elasticsearch content fixture "+contentToken); err != nil {
t.Fatalf("SetCellValue: %v", err)
}
if err := book.SaveAs(localPath); err != nil {
t.Fatalf("SaveAs: %v", err)
}
entry, err := c.UploadProjectFile(ctx, pid, localPath)
if err != nil {
t.Fatalf("UploadProjectFile: %v", err)
}
fileID := strconv.Itoa(int(FileEntryNumericID(entry)))
if fileID == "0" {
t.Fatalf("upload returned no file id: %+v", entry)
}
es, err := NewESSearcher(ESConfig{
URL: esURL,
Index: os.Getenv("ONLYOFFICE_ES_INDEX"),
Tenant: os.Getenv("ONLYOFFICE_TENANT"),
})
if err != nil {
t.Fatalf("NewESSearcher: %v", err)
}
// Indexing is asynchronous on the server; poll until the file shows up.
// The server's title analyzer splits on whitespace, so the name query is
// the full file name token (including extension), as a user would type it.
nameHit := waitForHit(t, ctx, es, SearchQuery{Text: title}, fileID)
if nameHit.Title != title {
t.Errorf("name hit title = %q, want %q", nameHit.Title, title)
}
contentHit := waitForHit(t, ctx, es, SearchQuery{Text: contentToken, InContent: true}, fileID)
if contentHit.Highlight == "" {
t.Error("content hit has no highlight fragment")
}
if !strings.Contains(contentHit.Title, nameToken) {
t.Errorf("content hit title = %q, want the uploaded workbook", contentHit.Title)
}
// The content token is absent from the title, so a name-only search must
// not return the file — this proves the content field is really queried.
if hits := searchQuiet(t, es, SearchQuery{Text: contentToken}); len(hits) != 0 {
t.Errorf("name-only search for content token returned %d hits, want 0", len(hits))
}
}
// waitForHit polls ES until the file with fileID appears and returns that hit.
func waitForHit(t *testing.T, ctx context.Context, s *ESSearcher, q SearchQuery, fileID string) SearchHit {
t.Helper()
var lastErr error
for {
hits, err := s.Search(ctx, q)
if err != nil {
lastErr = err
} else {
for _, h := range hits {
if h.ID == fileID {
return h
}
}
}
select {
case <-ctx.Done():
t.Fatalf("search %q: file %s not indexed in time (last err: %v)", q.Text, fileID, lastErr)
case <-time.After(3 * time.Second):
}
}
}
func searchQuiet(t *testing.T, s *ESSearcher, q SearchQuery) []SearchHit {
t.Helper()
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
hits, err := s.Search(ctx, q)
if err != nil {
t.Fatalf("Search(%q): %v", q.Text, err)
}
return hits
}
+217
View File
@@ -0,0 +1,217 @@
package onlyoffice
import (
"encoding/json"
"reflect"
"testing"
)
func TestESSearchRequestNameOnly(t *testing.T) {
got := esSearchRequest(SearchQuery{Text: "Rechnung"}, "1")
if got.Size != defaultESLimit {
t.Errorf("size = %d, want %d", got.Size, defaultESLimit)
}
if !reflect.DeepEqual(got.Source, []string{"id", "title", "folders"}) {
t.Errorf("_source = %v", got.Source)
}
if len(got.Query.Bool.Must) != 1 || got.Query.Bool.Must[0].MultiMatch == nil {
t.Fatalf("must = %+v, want one multi_match", got.Query.Bool.Must)
}
mm := got.Query.Bool.Must[0].MultiMatch
if mm.Query != "Rechnung" {
t.Errorf("query = %q", mm.Query)
}
if !reflect.DeepEqual(mm.Fields, []string{"title^2"}) {
t.Errorf("fields = %v, want title only", mm.Fields)
}
if _, ok := got.Highlight.Fields["document.attachment.content"]; ok {
t.Error("content highlight present without InContent")
}
if _, ok := got.Highlight.Fields["title"]; !ok {
t.Error("title highlight missing")
}
if len(got.Query.Bool.Filter) != 1 || got.Query.Bool.Filter[0].Term["tenantId"] != int64(1) {
t.Errorf("tenant filter = %+v, want numeric tenantId=1", got.Query.Bool.Filter)
}
}
func TestESSearchRequestContentFields(t *testing.T) {
got := esSearchRequest(SearchQuery{Text: "Mahnung", InContent: true}, "")
mm := got.Query.Bool.Must[0].MultiMatch
want := []string{"title^2", "document.attachment.content"}
if !reflect.DeepEqual(mm.Fields, want) {
t.Errorf("fields = %v, want %v", mm.Fields, want)
}
if _, ok := got.Highlight.Fields["document.attachment.content"]; !ok {
t.Error("content highlight missing with InContent")
}
if len(got.Query.Bool.Filter) != 0 {
t.Errorf("filter = %+v, want none without tenant/folder", got.Query.Bool.Filter)
}
}
func TestESSearchRequestFiltersAndLimit(t *testing.T) {
got := esSearchRequest(SearchQuery{
Text: "Storchen",
FolderID: "649",
Extensions: []string{".PDF", "pdf", "docx"},
Limit: 5000,
}, "42")
if got.Size != maxESLimit {
t.Errorf("size = %d, want cap %d", got.Size, maxESLimit)
}
var tenant, folder, wildcards int
for _, f := range got.Query.Bool.Filter {
switch {
case f.Term != nil && f.Term["tenantId"] != nil:
tenant++
case f.Nested != nil:
folder++
if f.Nested.Path != "folders" || f.Nested.Query.Term["folders.folderId"] != "649" {
t.Errorf("folder filter = %+v, want nested folders term 649", f.Nested)
}
case f.Wildcard != nil:
wildcards++
}
}
if tenant != 1 || folder != 1 {
t.Errorf("term filters tenant=%d folder=%d, want 1 each", tenant, folder)
}
if wildcards != 2 {
t.Errorf("wildcard filters = %d, want deduped PDF+docx", wildcards)
}
}
func TestESSearchRequestSubstringAndsTerms(t *testing.T) {
got := esSearchRequest(SearchQuery{Text: "Rechnung 2025", Substring: true, FolderID: "522"}, "")
if got.Query.Bool.Must[0].MultiMatch != nil {
t.Fatalf("substring must not use multi_match: %+v", got.Query.Bool.Must)
}
if len(got.Query.Bool.Must) != 2 {
t.Fatalf("must = %+v, want two ANDed wildcard terms", got.Query.Bool.Must)
}
want := []string{"*rechnung*", "*2025*"}
for i, m := range got.Query.Bool.Must {
if m.Wildcard == nil || m.Wildcard["title"] != want[i] {
t.Errorf("must[%d] = %+v, want title wildcard %q", i, m, want[i])
}
}
if len(got.Query.Bool.Filter) != 1 || got.Query.Bool.Filter[0].Nested == nil {
t.Errorf("folder filter = %+v, want nested", got.Query.Bool.Filter)
}
}
func TestEscapeWildcard(t *testing.T) {
cases := map[string]string{"*rechnung*": "rechnung", "a?b\\c": "abc", "plain": "plain"}
for in, want := range cases {
if got := escapeWildcard(in); got != want {
t.Errorf("escapeWildcard(%q) = %q, want %q", in, got, want)
}
}
}
func TestESSearchRequestRejectsEmptyTextAtSearch(t *testing.T) {
s, err := NewESSearcher(ESConfig{URL: "http://localhost:9200"})
if err != nil {
t.Fatalf("NewESSearcher: %v", err)
}
if _, err := s.Search(t.Context(), SearchQuery{Text: " "}); err == nil {
t.Error("empty query: want error")
}
}
func TestNewESSearcherRequiresURL(t *testing.T) {
if _, err := NewESSearcher(ESConfig{}); err == nil {
t.Error("empty URL: want error")
}
s, err := NewESSearcher(ESConfig{URL: "http://es:9200/"})
if err != nil {
t.Fatalf("NewESSearcher: %v", err)
}
if s.cfg.Index != defaultESIndex {
t.Errorf("index = %q, want %q", s.cfg.Index, defaultESIndex)
}
if s.cfg.URL != "http://es:9200" {
t.Errorf("url = %q, want trimmed", s.cfg.URL)
}
if s.Name() != "elasticsearch" {
t.Errorf("Name() = %q", s.Name())
}
}
func TestNormalizeExtensions(t *testing.T) {
got := normalizeExtensions([]string{" .PDF ", "pdf", "", "xlsx"})
want := []string{"pdf", "xlsx"}
if !reflect.DeepEqual(got, want) {
t.Errorf("normalizeExtensions = %v, want %v", got, want)
}
}
func TestParseESSearchResponse(t *testing.T) {
raw := []byte(`{
"took": 12,
"hits": {
"total": {"value": 2, "relation": "eq"},
"hits": [
{
"_id": "2395",
"_score": 7.31,
"_source": {"id": 2395, "title": "Rechnung-4711.pdf",
"folders": [{"folderId": "438", "id": 0}, {"folderId": "11", "id": 0}]},
"highlight": {
"title": ["<em>Rechnung</em>-4711.pdf"],
"document.attachment.content": ["… Zahlung der <em>Rechnung</em> …"]
}
},
{
"_id": "2318",
"_score": 6.02,
"_source": {"id": 2318, "title": "Mahnung.pdf", "folders": []},
"highlight": {"title": ["<em>Mahnung</em>.pdf"]}
}
]
}
}`)
hits, err := parseESSearchResponse(raw)
if err != nil {
t.Fatalf("parseESSearchResponse: %v", err)
}
if len(hits) != 2 {
t.Fatalf("hits = %d, want 2", len(hits))
}
h0 := hits[0]
if h0.ID != "2395" || h0.Title != "Rechnung-4711.pdf" || h0.Kind != File {
t.Errorf("hit0 entry = %+v", h0.Entry)
}
// folders is root → leaf; the immediate parent is the last entry.
if h0.ParentID != "11" || !reflect.DeepEqual(h0.Path, []string{"438", "11"}) {
t.Errorf("hit0 path = %v parent = %q, want parent 11", h0.Path, h0.ParentID)
}
if h0.Score != 7.31 {
t.Errorf("hit0 score = %v", h0.Score)
}
if h0.Highlight != "… Zahlung der Rechnung …" {
t.Errorf("hit0 highlight = %q, want content fragment", h0.Highlight)
}
if hits[1].Highlight != "Mahnung.pdf" {
t.Errorf("hit1 highlight = %q, want title without tags", hits[1].Highlight)
}
if hits[1].ParentID != "" || len(hits[1].Path) != 0 {
t.Errorf("hit1 path = %v", hits[1].Path)
}
}
func TestESSearchRequestJSONShape(t *testing.T) {
got := esSearchRequest(SearchQuery{Text: "Rechnung", InContent: true}, "1")
b, err := json.Marshal(got)
if err != nil {
t.Fatalf("marshal: %v", err)
}
var back map[string]any
if err := json.Unmarshal(b, &back); err != nil {
t.Fatalf("unmarshal: %v", err)
}
if _, ok := back["query"].(map[string]any)["bool"]; !ok {
t.Errorf("query.bool missing: %s", b)
}
}
+336
View File
@@ -0,0 +1,336 @@
package onlyoffice
// Own full-text index (epic #34, F6 #42).
//
// The OnlyOffice Elasticsearch index (files_file) only holds extracted content
// for Office formats. FileUtility.CanIndex gates extraction by the server
// setting files.index.formats, whose default is ".pptx|.xlsx|.docx", so PDFs
// are indexed by name only. Instead of patching the server (risky: lost on
// upgrade, forces a full reindex) this file implements a second, independent
// index (default oo_docs_text) that our own pipeline fills from
// internal/docpipe (pdftotext + OCR). The OnlyOffice index is never touched.
//
// See docs/elasticsearch.md for the decision and the trade-offs.
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
const defaultESTextIndex = "oo_docs_text"
// ESTextConfig configures the own full-text index.
type ESTextConfig struct {
URL string // scheme://host:port of the ES HTTP endpoint
Index string // index name, default oo_docs_text
Tenant string // reserved for future multi-tenant data; unused for now
}
// ESTextConfigFromEnv reads ONLYOFFICE_ES_URL and ONLYOFFICE_ES_TEXT_INDEX
// (default oo_docs_text). The library never loads dotfiles — the CLI does that.
func ESTextConfigFromEnv() ESTextConfig {
return ESTextConfig{
URL: strings.TrimRight(strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL")), "/"),
Index: firstNonEmpty(os.Getenv("ONLYOFFICE_ES_TEXT_INDEX"), defaultESTextIndex),
Tenant: strings.TrimSpace(os.Getenv("ONLYOFFICE_TENANT")),
}
}
// TextDoc is one document in the own full-text index. It is keyed by the
// OnlyOffice file id so hits map straight back to Documents entries.
type TextDoc struct {
ID string `json:"id"`
Title string `json:"title"`
FolderID string `json:"folder,omitempty"`
Ext string `json:"ext,omitempty"`
Content string `json:"content"`
}
// TextIndex is the storage/search surface for locally extracted document text.
// It complements Searcher: ESSearcher reads OnlyOffice's index, ESTextIndex
// reads ours.
type TextIndex interface {
Put(ctx context.Context, docs []TextDoc) error
Delete(ctx context.Context, ids []string) error
Search(ctx context.Context, q SearchQuery) ([]SearchHit, error)
Name() string
}
// ESTextIndex is a TextIndex (and Searcher) over a dedicated Elasticsearch
// index filled by TextIndexer.
type ESTextIndex struct {
cfg ESTextConfig
http *http.Client
}
var (
_ TextIndex = (*ESTextIndex)(nil)
_ Searcher = (*ESTextIndex)(nil)
)
// NewESTextIndex returns a searcher/writer for the own full-text index. The URL
// is required; an empty index falls back to oo_docs_text.
func NewESTextIndex(cfg ESTextConfig) (*ESTextIndex, error) {
if strings.TrimSpace(cfg.URL) == "" {
return nil, fmt.Errorf("onlyoffice: elasticsearch URL is empty (set ONLYOFFICE_ES_URL)")
}
cfg.URL = strings.TrimRight(cfg.URL, "/")
if cfg.Index == "" {
cfg.Index = defaultESTextIndex
}
return &ESTextIndex{cfg: cfg, http: &http.Client{Timeout: 120 * time.Second}}, nil
}
// Name implements Searcher and TextIndex.
func (x *ESTextIndex) Name() string { return "es-text" }
// Index returns the configured index name.
func (x *ESTextIndex) Index() string { return x.cfg.Index }
// esTextMapping pins explicit types: content must stay a plain text field (no
// keyword subfield) and folder/ext stay exact keywords for filters.
const esTextMapping = `{
"mappings": {
"properties": {
"id": {"type": "keyword"},
"title": {"type": "text", "fields": {"keyword": {"type": "keyword", "ignore_above": 512}}},
"folder": {"type": "keyword"},
"ext": {"type": "keyword"},
"content": {"type": "text"}
}
}
}`
// Ensure creates the index with the explicit mapping. A missing index is
// created; an already existing one is left untouched.
func (x *ESTextIndex) Ensure(ctx context.Context) error {
status, raw, err := x.do(ctx, http.MethodPut, "/"+x.cfg.Index, []byte(esTextMapping), "application/json")
if err != nil {
return err
}
if status == http.StatusOK {
return nil
}
if status == http.StatusBadRequest && bytes.Contains(raw, []byte("resource_already_exists_exception")) {
return nil
}
return fmt.Errorf("onlyoffice: create text index %s: %d %s", x.cfg.Index, status, truncate(string(raw), 300))
}
// Put upserts documents via the bulk API and refreshes so they are immediately
// searchable.
func (x *ESTextIndex) Put(ctx context.Context, docs []TextDoc) error {
if len(docs) == 0 {
return nil
}
status, raw, err := x.do(ctx, http.MethodPost, "/"+x.cfg.Index+"/_bulk?refresh=true", esTextBulkBody(docs), "application/x-ndjson")
if err != nil {
return err
}
if status >= 400 {
return fmt.Errorf("onlyoffice: bulk index %s: %d %s", x.cfg.Index, status, truncate(string(raw), 400))
}
var res esBulkResponse
if err := json.Unmarshal(raw, &res); err != nil {
return fmt.Errorf("onlyoffice: decode bulk response: %w", err)
}
if !res.Errors {
return nil
}
return fmt.Errorf("onlyoffice: bulk index %s: %s", x.cfg.Index, res.firstError())
}
// Delete removes documents by OnlyOffice file id. A missing index means there
// is nothing to delete.
func (x *ESTextIndex) Delete(ctx context.Context, ids []string) error {
if len(ids) == 0 {
return nil
}
body, err := json.Marshal(map[string]any{"query": map[string]any{"terms": map[string]any{"id": ids}}})
if err != nil {
return err
}
status, raw, err := x.do(ctx, http.MethodPost, "/"+x.cfg.Index+"/_delete_by_query?refresh=true", body, "application/json")
if err != nil {
return err
}
if status == http.StatusNotFound {
return nil
}
if status >= 400 {
return fmt.Errorf("onlyoffice: delete from %s: %d %s", x.cfg.Index, status, truncate(string(raw), 400))
}
return nil
}
// Search runs a multi_match over title (boosted) and content, with optional
// folder and extension filters. A missing index yields no hits, not an error.
func (x *ESTextIndex) Search(ctx context.Context, q SearchQuery) ([]SearchHit, error) {
q.Text = strings.TrimSpace(q.Text)
if q.Text == "" {
return nil, fmt.Errorf("onlyoffice: empty search query")
}
body, err := json.Marshal(esTextSearchRequest(q))
if err != nil {
return nil, fmt.Errorf("onlyoffice: build elasticsearch query: %w", err)
}
status, raw, err := x.do(ctx, http.MethodPost, "/"+x.cfg.Index+"/_search", body, "application/json")
if err != nil {
return nil, err
}
if status == http.StatusNotFound {
return nil, nil
}
if status >= 400 {
return nil, fmt.Errorf("onlyoffice: search %s: %d %s", x.cfg.Index, status, truncate(string(raw), 400))
}
return parseESTextResponse(raw)
}
// do sends one request and returns the status and body (bounded). The caller
// decides which statuses are errors.
func (x *ESTextIndex) do(ctx context.Context, method, path string, body []byte, contentType string) (int, []byte, error) {
var r io.Reader
if body != nil {
r = bytes.NewReader(body)
}
req, err := http.NewRequestWithContext(ctx, method, x.cfg.URL+path, r)
if err != nil {
return 0, nil, err
}
req.Header.Set("Accept", "application/json")
if contentType != "" {
req.Header.Set("Content-Type", contentType)
}
resp, err := x.http.Do(req)
if err != nil {
return 0, nil, fmt.Errorf("onlyoffice: elasticsearch %s: %w", method, err)
}
defer resp.Body.Close()
raw, err := io.ReadAll(io.LimitReader(resp.Body, maxESResponseSize))
if err != nil {
return resp.StatusCode, nil, err
}
return resp.StatusCode, raw, nil
}
// esTextBulkBody renders the NDJSON bulk payload. Pure, so it is unit-tested.
func esTextBulkBody(docs []TextDoc) []byte {
var b bytes.Buffer
enc := json.NewEncoder(&b)
enc.SetEscapeHTML(false)
for _, d := range docs {
_ = enc.Encode(map[string]any{"index": map[string]any{"_id": d.ID}})
_ = enc.Encode(d)
}
return b.Bytes()
}
// esTextSearchRequest builds the own-index query. Pure, so it is unit-tested.
func esTextSearchRequest(q SearchQuery) esRequest {
limit := q.Limit
if limit <= 0 {
limit = defaultESLimit
}
if limit > maxESLimit {
limit = maxESLimit
}
fields := []string{"title^2", "content"}
must := []esClause{{MultiMatch: &esMultiMatch{Query: q.Text, Fields: fields}}}
var filter []esClause
if f := strings.TrimSpace(q.FolderID); f != "" {
filter = append(filter, esClause{Term: map[string]any{"folder": f}})
}
if exts := normalizeExtensions(q.Extensions); len(exts) > 0 {
filter = append(filter, esClause{Terms: map[string]any{"ext": exts}})
}
return esRequest{
Size: limit,
Source: []string{"id", "title", "folder", "ext"},
Query: esQuery{Bool: esBool{Must: must, Filter: filter}},
Highlight: esHighlight{PreTags: []string{"<em>"}, PostTags: []string{"</em>"}, Fields: map[string]struct{}{"title": {}, "content": {}}},
}
}
// esBulkResponse is the subset of an ES bulk response we consume.
type esBulkResponse struct {
Errors bool `json:"errors"`
Items []map[string]struct {
ID string `json:"_id"`
Status int `json:"status"`
Error *struct {
Type string `json:"type"`
Reason string `json:"reason"`
} `json:"error"`
} `json:"items"`
}
// firstError returns a compact description of the first failed bulk item.
func (r esBulkResponse) firstError() string {
for _, item := range r.Items {
for op, res := range item {
if res.Error != nil {
return fmt.Sprintf("%s %s: %s %s", op, res.ID, res.Error.Type, res.Error.Reason)
}
}
}
return "unknown bulk error"
}
// esTextResponse is the subset of an own-index search response we consume.
type esTextResponse struct {
Hits struct {
Total struct {
Value int `json:"value"`
} `json:"total"`
Hits []struct {
ID string `json:"_id"`
Score float64 `json:"_score"`
Source TextDoc `json:"_source"`
HL map[string][]string `json:"highlight"`
} `json:"hits"`
} `json:"hits"`
}
// parseESTextResponse converts an own-index search response into SearchHit
// values. Pure, so it is unit-tested.
func parseESTextResponse(raw []byte) ([]SearchHit, error) {
var r esTextResponse
if err := json.Unmarshal(raw, &r); err != nil {
return nil, fmt.Errorf("onlyoffice: decode elasticsearch response: %w", err)
}
hits := make([]SearchHit, 0, len(r.Hits.Hits))
for _, h := range r.Hits.Hits {
id := h.Source.ID
if id == "" {
id = h.ID
}
parent := h.Source.FolderID
var path []string
if parent != "" {
path = []string{parent}
}
hits = append(hits, SearchHit{
Entry: Entry{
ID: id,
ParentID: parent,
Title: h.Source.Title,
Kind: File,
Provider: "es-text",
},
Score: h.Score,
Highlight: esHighlightText(h.HL),
Path: path,
})
}
return hits, nil
}
+158
View File
@@ -0,0 +1,158 @@
//go:build integration
package onlyoffice
import (
"context"
"net/http"
"os"
"strings"
"testing"
"time"
"github.com/eslider/go-onlyoffice/internal/docpipe"
)
// TestIntegrationESTextIndex verifies the own full-text index end to end
// against a live Elasticsearch: create the index with its mapping, index a
// document, find it by content (and reject it via a folder filter and after
// deletion), then drop the throwaway index.
//
// Requires ONLYOFFICE_ES_URL (a reachable ES endpoint — in the current setup a
// tunnel to the ES inside the OnlyOffice VM, see docs/elasticsearch.md). It
// does not need OnlyOffice credentials because no file is downloaded: the
// TextIndexer write path is covered by unit tests with a fake extractor.
func TestIntegrationESTextIndex(t *testing.T) {
esURL := strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL"))
if esURL == "" {
t.Skip("ONLYOFFICE_ES_URL not set — skipping Elasticsearch integration test")
}
stamp := time.Now().UTC().Format("20060102150405")
idx, err := NewESTextIndex(ESTextConfig{URL: esURL, Index: "oo_docs_text_it_" + stamp})
if err != nil {
t.Fatalf("NewESTextIndex: %v", err)
}
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
t.Cleanup(func() {
cleanupCtx, done := context.WithTimeout(context.Background(), 30*time.Second)
defer done()
_, _, _ = idx.do(cleanupCtx, http.MethodDelete, "/"+idx.Index(), nil, "")
})
if err := idx.Ensure(ctx); err != nil {
t.Fatalf("Ensure: %v", err)
}
// Ensure is idempotent.
if err := idx.Ensure(ctx); err != nil {
t.Fatalf("Ensure (second): %v", err)
}
token := "gotes" + stamp
doc := TextDoc{
ID: "3578",
Title: "2026-07-28-S1021-acme-rechnung.pdf",
FolderID: "634",
Ext: "pdf",
Content: "Begleitzettel SGB XI — Rechnung " + token,
}
if err := idx.Put(ctx, []TextDoc{doc}); err != nil {
t.Fatalf("Put: %v", err)
}
hits, err := idx.Search(ctx, SearchQuery{Text: token})
if err != nil {
t.Fatalf("Search: %v", err)
}
if len(hits) != 1 || hits[0].ID != "3578" {
t.Fatalf("content search hits = %+v, want doc 3578", hits)
}
if !strings.Contains(hits[0].Highlight, token) {
t.Errorf("highlight = %q, want token", hits[0].Highlight)
}
if hits, err := idx.Search(ctx, SearchQuery{Text: token, FolderID: "999"}); err != nil {
t.Fatalf("Search with folder filter: %v", err)
} else if len(hits) != 0 {
t.Errorf("folder filter returned %d hits, want 0", len(hits))
}
if hits, err := idx.Search(ctx, SearchQuery{Text: token, Extensions: []string{"docx"}}); err != nil {
t.Fatalf("Search with ext filter: %v", err)
} else if len(hits) != 0 {
t.Errorf("ext filter returned %d hits, want 0", len(hits))
}
if err := idx.Delete(ctx, []string{"3578"}); err != nil {
t.Fatalf("Delete: %v", err)
}
if hits, err := idx.Search(ctx, SearchQuery{Text: token}); err != nil {
t.Fatalf("Search after delete: %v", err)
} else if len(hits) != 0 {
t.Errorf("after delete search returned %d hits, want 0", len(hits))
}
}
// TestIntegrationESTextIndexPDFAttachment indexes testdata/pdf-with-attachment.pdf
// through the real pipeline (TextIndexer + docpipe: pdfdetach + pdftotext) and
// verifies that text living only in the embedded attachment is searchable.
//
// Requires ONLYOFFICE_ES_URL plus poppler (pdfdetach/pdftotext). No OnlyOffice
// credentials are needed: a fixture FileStore serves the PDF bytes.
func TestIntegrationESTextIndexPDFAttachment(t *testing.T) {
esURL := strings.TrimSpace(os.Getenv("ONLYOFFICE_ES_URL"))
if esURL == "" {
t.Skip("ONLYOFFICE_ES_URL not set — skipping Elasticsearch integration test")
}
if docpipe.LookPath().PDFDetach == "" {
t.Skip("pdfdetach not on PATH — skipping PDF attachment integration test")
}
pdf, err := os.ReadFile("testdata/pdf-with-attachment.pdf")
if err != nil {
t.Fatalf("read fixture: %v", err)
}
stamp := time.Now().UTC().Format("20060102150405")
idx, err := NewESTextIndex(ESTextConfig{URL: esURL, Index: "oo_docs_text_it_att_" + stamp})
if err != nil {
t.Fatalf("NewESTextIndex: %v", err)
}
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
t.Cleanup(func() {
cleanupCtx, done := context.WithTimeout(context.Background(), 30*time.Second)
defer done()
_, _, _ = idx.do(cleanupCtx, http.MethodDelete, "/"+idx.Index(), nil, "")
})
if err := idx.Ensure(ctx); err != nil {
t.Fatalf("Ensure: %v", err)
}
store := &textFakeStore{files: map[string][]byte{"9001": pdf}}
ix := NewTextIndexer(store, idx)
res, err := ix.IndexEntries(ctx, []Entry{{ID: "9001", Title: "scan.pdf", ParentID: "777", Kind: File}}, IndexOptions{MinChars: 1})
if err != nil {
t.Fatalf("IndexEntries: %v", err)
}
if res.Indexed != 1 || res.Failed != 0 {
t.Fatalf("result = %+v, want one indexed doc", res)
}
// Token appears only inside the embedded goo-note.txt attachment.
hits, err := idx.Search(ctx, SearchQuery{Text: "gooattachmenttoken"})
if err != nil {
t.Fatalf("Search attachment token: %v", err)
}
if len(hits) != 1 || hits[0].ID != "9001" {
t.Fatalf("attachment-token hits = %+v, want doc 9001", hits)
}
if !strings.Contains(hits[0].Highlight, "gooattachmenttoken") {
t.Errorf("highlight = %q, want attachment token", hits[0].Highlight)
}
// Body text is indexed as before.
if hits, err := idx.Search(ctx, SearchQuery{Text: "goobodytoken"}); err != nil {
t.Fatalf("Search body token: %v", err)
} else if len(hits) != 1 {
t.Errorf("body-token hits = %d, want 1", len(hits))
}
}
+275
View File
@@ -0,0 +1,275 @@
package onlyoffice
import (
"context"
"fmt"
"io"
"os"
"reflect"
"strings"
"testing"
)
func TestESTextSearchRequestShape(t *testing.T) {
got := esTextSearchRequest(SearchQuery{
Text: "S1021",
FolderID: "634",
Extensions: []string{".PDF", "pdf"},
Limit: 5,
})
if got.Size != 5 {
t.Errorf("size = %d, want 5", got.Size)
}
mm := got.Query.Bool.Must[0].MultiMatch
if mm == nil || !reflect.DeepEqual(mm.Fields, []string{"title^2", "content"}) {
t.Fatalf("multi_match = %+v, want title^2 + content", mm)
}
if _, ok := got.Highlight.Fields["content"]; !ok {
t.Error("content highlight missing")
}
if _, ok := got.Highlight.Fields["title"]; !ok {
t.Error("title highlight missing")
}
var folder, exts int
for _, f := range got.Query.Bool.Filter {
switch {
case f.Term != nil && f.Term["folder"] != nil:
folder++
if f.Term["folder"] != "634" {
t.Errorf("folder term = %+v", f.Term)
}
case f.Terms != nil:
exts++
if !reflect.DeepEqual(f.Terms["ext"], []string{"pdf"}) {
t.Errorf("ext terms = %+v, want deduped pdf", f.Terms)
}
}
}
if folder != 1 || exts != 1 {
t.Errorf("filters folder=%d exts=%d, want 1 each", folder, exts)
}
}
func TestESTextBulkBody(t *testing.T) {
body := esTextBulkBody([]TextDoc{
{ID: "3578", Title: "S1021.pdf", FolderID: "634", Ext: "pdf", Content: "Begleitzettel <S1021> & mehr"},
{ID: "3579", Title: "S1023.pdf", Ext: "pdf", Content: "x"},
})
lines := strings.Split(strings.TrimRight(string(body), "\n"), "\n")
if len(lines) != 4 {
t.Fatalf("bulk body has %d lines, want 4:\n%s", len(lines), body)
}
if !strings.Contains(lines[0], `"index"`) || !strings.Contains(lines[0], `"_id":"3578"`) {
t.Errorf("action line = %q", lines[0])
}
if !strings.Contains(lines[1], `"content":"Begleitzettel <S1021> & mehr"`) {
t.Errorf("source line should keep HTML unescaped, got %q", lines[1])
}
if !strings.Contains(lines[2], `"_id":"3579"`) {
t.Errorf("second action line = %q", lines[2])
}
}
func TestParseESTextResponse(t *testing.T) {
raw := []byte(`{
"hits": {
"total": {"value": 1, "relation": "eq"},
"hits": [
{
"_id": "3578",
"_score": 3.21,
"_source": {"id": "3578", "title": "2026-07-28-S1021-acme-rechnung.pdf", "folder": "634", "ext": "pdf"},
"highlight": {"content": ["Begleitzettel … <em>S1021</em> …"]}
}
]
}
}`)
hits, err := parseESTextResponse(raw)
if err != nil {
t.Fatalf("parseESTextResponse: %v", err)
}
if len(hits) != 1 {
t.Fatalf("hits = %d, want 1", len(hits))
}
h := hits[0]
if h.ID != "3578" || h.Title != "2026-07-28-S1021-acme-rechnung.pdf" || h.Kind != File {
t.Errorf("entry = %+v", h.Entry)
}
if h.ParentID != "634" || !reflect.DeepEqual(h.Path, []string{"634"}) {
t.Errorf("path = %v parent = %q", h.Path, h.ParentID)
}
if h.Provider != "es-text" {
t.Errorf("provider = %q", h.Provider)
}
if h.Highlight != "Begleitzettel … S1021 …" {
t.Errorf("highlight = %q, want tags stripped", h.Highlight)
}
}
func TestNewESTextIndexDefaults(t *testing.T) {
if _, err := NewESTextIndex(ESTextConfig{}); err == nil {
t.Error("empty URL: want error")
}
x, err := NewESTextIndex(ESTextConfig{URL: "http://es:9200/"})
if err != nil {
t.Fatalf("NewESTextIndex: %v", err)
}
if x.Index() != defaultESTextIndex {
t.Errorf("index = %q, want %q", x.Index(), defaultESTextIndex)
}
if x.cfg.URL != "http://es:9200" {
t.Errorf("url = %q, want trimmed", x.cfg.URL)
}
if x.Name() != "es-text" {
t.Errorf("Name() = %q", x.Name())
}
}
func TestESTextConfigFromEnvIndexDefault(t *testing.T) {
t.Setenv("ONLYOFFICE_ES_URL", "http://es:9200/")
t.Setenv("ONLYOFFICE_ES_TEXT_INDEX", "")
cfg := ESTextConfigFromEnv()
if cfg.Index != defaultESTextIndex {
t.Errorf("index = %q, want %q", cfg.Index, defaultESTextIndex)
}
}
func TestTextIndexerIndexEntries(t *testing.T) {
store := &textFakeStore{
files: map[string][]byte{"1": []byte("PDFBYTES")},
}
idx := &textFakeIndex{}
ix := NewTextIndexer(store, idx)
ix.Extractor = textFakeExtractor{prefix: "TEXT "}
res, err := ix.IndexEntries(context.Background(), []Entry{
{ID: "1", Title: "Rechnung.PDF", ParentID: "649", Kind: File},
{ID: "2", Title: "Tabelle.xlsx", ParentID: "649", Kind: File},
{ID: "3", Title: "Unterordner", Kind: Folder},
}, IndexOptions{})
if err != nil {
t.Fatalf("IndexEntries: %v", err)
}
if res.Scanned != 3 || res.Indexed != 1 || res.Skipped != 2 || res.Failed != 0 {
t.Errorf("result = %+v, want scanned=3 indexed=1 skipped=2 failed=0", res)
}
if len(idx.docs) != 1 {
t.Fatalf("indexed docs = %d, want 1", len(idx.docs))
}
got := idx.docs[0]
want := TextDoc{ID: "1", Title: "Rechnung.PDF", FolderID: "649", Ext: "pdf", Content: "TEXT PDFBYTES"}
if !reflect.DeepEqual(got, want) {
t.Errorf("doc = %+v, want %+v", got, want)
}
}
func TestTextIndexerRecordsExtractionFailure(t *testing.T) {
store := &textFakeStore{files: map[string][]byte{"1": []byte("x")}}
idx := &textFakeIndex{}
ix := NewTextIndexer(store, idx)
ix.Extractor = textFailingExtractor{}
res, err := ix.IndexEntries(context.Background(), []Entry{{ID: "1", Title: "a.pdf", Kind: File}}, IndexOptions{})
if err != nil {
t.Fatalf("IndexEntries: %v", err)
}
if res.Indexed != 0 || res.Failed != 1 || len(res.Errors) != 1 {
t.Errorf("result = %+v, want one failure recorded", res)
}
}
func TestTextIndexerPlanFolder(t *testing.T) {
store := &textFakeStore{dirs: map[string][]Entry{
"root": {
{ID: "10", Title: "a.pdf", Kind: File},
{ID: "11", Title: "sub", Kind: Folder},
},
"11": {
{ID: "12", Title: "b.PDF", Kind: File},
{ID: "13", Title: "c.xlsx", Kind: File},
},
}}
ix := NewTextIndexer(store, &textFakeIndex{})
flat, err := ix.PlanFolder(context.Background(), "root", IndexOptions{})
if err != nil {
t.Fatalf("PlanFolder: %v", err)
}
if len(flat) != 1 || flat[0].ID != "10" {
t.Errorf("flat plan = %+v, want only a.pdf", flat)
}
deep, err := ix.PlanFolder(context.Background(), "root", IndexOptions{Recursive: true, Limit: 10})
if err != nil {
t.Fatalf("PlanFolder recursive: %v", err)
}
if len(deep) != 2 {
t.Errorf("recursive plan = %d entries, want 2", len(deep))
}
}
// --- fakes -----------------------------------------------------------------
type textFakeStore struct {
dirs map[string][]Entry
files map[string][]byte
stat map[string]Entry
}
func (f *textFakeStore) Name() string { return "fake" }
func (f *textFakeStore) List(_ context.Context, parentID string) ([]Entry, error) {
return f.dirs[parentID], nil
}
func (f *textFakeStore) Stat(_ context.Context, id string) (Entry, error) {
if e, ok := f.stat[id]; ok {
return e, nil
}
return Entry{}, fmt.Errorf("not found: %s", id)
}
func (f *textFakeStore) Download(_ context.Context, id string, w io.Writer) (int64, error) {
b, ok := f.files[id]
if !ok {
return 0, fmt.Errorf("no bytes for %s", id)
}
n, err := w.Write(b)
return int64(n), err
}
func (f *textFakeStore) CreateFolder(context.Context, string, string) (Entry, error) {
return Entry{}, nil
}
func (f *textFakeStore) Upload(context.Context, string, string, io.Reader) (Entry, error) {
return Entry{}, nil
}
func (f *textFakeStore) Move(context.Context, []string, string) error { return nil }
func (f *textFakeStore) Copy(context.Context, []string, string) error { return nil }
func (f *textFakeStore) Rename(context.Context, string, string) error { return nil }
func (f *textFakeStore) Delete(context.Context, []string) error { return nil }
type textFakeIndex struct{ docs []TextDoc }
func (f *textFakeIndex) Put(_ context.Context, docs []TextDoc) error {
f.docs = append(f.docs, docs...)
return nil
}
func (f *textFakeIndex) Delete(context.Context, []string) error { return nil }
func (f *textFakeIndex) Search(context.Context, SearchQuery) ([]SearchHit, error) { return nil, nil }
func (f *textFakeIndex) Name() string { return "fake" }
type textFakeExtractor struct{ prefix string }
func (f textFakeExtractor) Extract(path, _, _ string, _ int) (string, error) {
b, err := os.ReadFile(path)
if err != nil {
return "", err
}
return f.prefix + string(b), nil
}
type textFailingExtractor struct{}
func (textFailingExtractor) Extract(string, string, string, int) (string, error) {
return "", fmt.Errorf("boom")
}
+249
View File
@@ -0,0 +1,249 @@
package onlyoffice
// Single file client (epic #34, F4 #38). FileClient composes the registered
// FileStore and Searcher backends and picks one per operation: REST/DAV for
// writes, PostgreSQL (when registered) for fast reads, Elasticsearch for name
// and content search. Client.Files returns the facade; it also implements
// FileStore, so existing callers keep compiling.
import (
"context"
"errors"
"io"
"strings"
)
// ProviderES is the composed Elasticsearch searcher. The SQL store owns
// ProviderPG/ProviderMySQL (filestore_pg.go); the facade references ProviderPG in
// readOrder.
const ProviderES = "elasticsearch"
var (
errNoReadBackend = errors.New("onlyoffice: no file backend registered for reads")
errNoWriteBackend = errors.New("onlyoffice: no file backend registered for writes")
errNoSearcher = errors.New("onlyoffice: no search backend registered (set ONLYOFFICE_ES_URL)")
)
// FileClient is the single entry point for file operations. It holds the
// registered backends and the order in which each operation tries them.
type FileClient struct {
stores map[string]FileStore
searchers map[string]Searcher
readOrder []string
writeOrder []string
searchOrder []string
}
// newFileClient builds the facade over the built-in REST and DAV stores. The
// Elasticsearch searcher is registered when ONLYOFFICE_ES_URL is set; the
// missing-credential case is left to Search so read-only commands still work.
func (c *Client) newFileClient() *FileClient {
f := &FileClient{
stores: map[string]FileStore{
ProviderREST: &restStore{c: c},
ProviderDAV: &davStore{c: c},
},
searchers: map[string]Searcher{},
readOrder: []string{ProviderPG, ProviderMySQL, ProviderREST, ProviderDAV},
writeOrder: []string{ProviderREST, ProviderDAV},
searchOrder: []string{ProviderES},
}
if cfg := ESConfigFromEnv(); cfg.URL != "" {
if es, err := NewESSearcher(cfg); err == nil {
f.searchers[ProviderES] = es
}
}
return f
}
// RegisterStore adds or replaces a named backend (for example the PostgreSQL
// read store). The name is matched case-insensitively.
func (f *FileClient) RegisterStore(name string, s FileStore) {
if f == nil || s == nil {
return
}
name = normalizeProvider(name)
if name == "" {
return
}
if f.stores == nil {
f.stores = map[string]FileStore{}
}
f.stores[name] = s
}
// RegisterSearcher adds or replaces a named search backend.
func (f *FileClient) RegisterSearcher(name string, s Searcher) {
if f == nil || s == nil {
return
}
name = normalizeProvider(name)
if name == "" {
return
}
if f.searchers == nil {
f.searchers = map[string]Searcher{}
}
f.searchers[name] = s
}
// Read returns the preferred backend for reads: the SQL store (PostgreSQL or
// MySQL) when registered, then REST, then WebDAV.
func (f *FileClient) Read() FileStore { return f.firstStore(f.readOrder) }
// Write returns the preferred backend for writes: REST, then WebDAV.
func (f *FileClient) Write() FileStore { return f.firstStore(f.writeOrder) }
// Search returns the preferred name/content searcher (Elasticsearch), or an
// error when no search backend is configured.
func (f *FileClient) Search() (Searcher, error) {
if f == nil {
return nil, errNoSearcher
}
for _, name := range f.searchOrder {
if s := f.searchers[normalizeProvider(name)]; s != nil {
return s, nil
}
}
return nil, errNoSearcher
}
// firstStore returns the first registered store in the order.
func (f *FileClient) firstStore(order []string) FileStore {
if f == nil {
return nil
}
for _, name := range order {
if s := f.stores[normalizeProvider(name)]; s != nil {
return s
}
}
return nil
}
// orderedStores returns the registered stores in the order.
func (f *FileClient) orderedStores(order []string) []FileStore {
if f == nil {
return nil
}
out := make([]FileStore, 0, len(order))
for _, name := range order {
if s := f.stores[normalizeProvider(name)]; s != nil {
out = append(out, s)
}
}
return out
}
func normalizeProvider(name string) string {
return strings.ToLower(strings.TrimSpace(name))
}
// Name implements FileStore and reports the preferred read backend.
func (f *FileClient) Name() string {
if s := f.Read(); s != nil {
return s.Name()
}
return ""
}
// List reads from the preferred backend, falling back to the next read backend
// only on a transient error (429/502/503/504).
func (f *FileClient) List(ctx context.Context, parentID string) ([]Entry, error) {
return fallbackRead(ctx, f.orderedStores(f.readOrder), func(s FileStore) ([]Entry, error) {
return s.List(ctx, parentID)
})
}
// Stat reads from the preferred backend, with the same transient fallback.
func (f *FileClient) Stat(ctx context.Context, id string) (Entry, error) {
return fallbackRead(ctx, f.orderedStores(f.readOrder), func(s FileStore) (Entry, error) {
return s.Stat(ctx, id)
})
}
// Download streams file bytes. It does not fall back: a failed attempt may have
// already written partial bytes into w, so a second backend would append.
func (f *FileClient) Download(ctx context.Context, id string, w io.Writer) (int64, error) {
s := f.Read()
if s == nil {
return 0, errNoReadBackend
}
return s.Download(ctx, id, w)
}
// CreateFolder writes to the preferred write backend.
func (f *FileClient) CreateFolder(ctx context.Context, parentID, title string) (Entry, error) {
s := f.Write()
if s == nil {
return Entry{}, errNoWriteBackend
}
return s.CreateFolder(ctx, parentID, title)
}
// Upload writes to the preferred write backend.
func (f *FileClient) Upload(ctx context.Context, parentID, title string, r io.Reader) (Entry, error) {
s := f.Write()
if s == nil {
return Entry{}, errNoWriteBackend
}
return s.Upload(ctx, parentID, title, r)
}
// Move writes to the preferred write backend.
func (f *FileClient) Move(ctx context.Context, ids []string, parentID string) error {
s := f.Write()
if s == nil {
return errNoWriteBackend
}
return s.Move(ctx, ids, parentID)
}
// Copy writes to the preferred write backend.
func (f *FileClient) Copy(ctx context.Context, ids []string, parentID string) error {
s := f.Write()
if s == nil {
return errNoWriteBackend
}
return s.Copy(ctx, ids, parentID)
}
// Rename writes to the preferred write backend.
func (f *FileClient) Rename(ctx context.Context, id, title string) error {
s := f.Write()
if s == nil {
return errNoWriteBackend
}
return s.Rename(ctx, id, title)
}
// Delete writes to the preferred write backend.
func (f *FileClient) Delete(ctx context.Context, ids []string) error {
s := f.Write()
if s == nil {
return errNoWriteBackend
}
return s.Delete(ctx, ids)
}
// fallbackRead runs op against each store in order, moving on only when the
// error is transient. Non-transient errors (not found, forbidden) are final.
func fallbackRead[T any](ctx context.Context, stores []FileStore, op func(FileStore) (T, error)) (T, error) {
var zero T
if len(stores) == 0 {
return zero, errNoReadBackend
}
var err error
for i, s := range stores {
var v T
v, err = op(s)
if err == nil {
return v, nil
}
if i == len(stores)-1 || !Transient(err) {
return zero, err
}
}
return zero, err
}
+131
View File
@@ -0,0 +1,131 @@
//go:build integration
package onlyoffice
import (
"bytes"
"context"
"strconv"
"testing"
"time"
)
// TestIntegrationFacadeCRUD drives the whole operation set through the composed
// facade c.Files(): folder create, upload, stat, list, rename, move, copy,
// delete. Writes must go to REST (the default writeOrder), reads follow
// readOrder (REST when no SQL backend is registered) and every returned Entry
// must report its provider. Destructive — throwaway project, cleaned up.
func TestIntegrationFacadeCRUD(t *testing.T) {
c := liveClient(t)
t.Cleanup(func() { cleanupTestProjects(t, c) })
ctx := context.Background()
suffix := time.Now().UTC().Format("20060102-150405")
project, err := c.CreateProject(NewProjectRequest{
Title: testProjectPrefix + "facade-" + suffix,
Description: "go-onlyoffice facade CRUD integration",
})
if err != nil {
t.Fatalf("CreateProject: %v", err)
}
if project.ID == nil {
t.Fatal("created project without id")
}
root, err := c.projectFolderID(ctx, strconv.Itoa(*project.ID))
if err != nil {
t.Fatalf("projectFolderID: %v", err)
}
f := c.Files()
if got := f.Write().Name(); got != ProviderREST {
t.Fatalf("Write().Name() = %q, want %q", got, ProviderREST)
}
if got := f.Read().Name(); got != ProviderREST {
t.Fatalf("Read().Name() = %q, want %q (no SQL backend registered)", got, ProviderREST)
}
src, err := f.CreateFolder(ctx, root, "facade-src-"+suffix)
if err != nil {
t.Fatalf("CreateFolder src: %v", err)
}
if src.Kind != Folder || src.ID == "" {
t.Fatalf("created src folder: %+v", src)
}
if src.Provider != ProviderREST {
t.Fatalf("CreateFolder provider = %q, want %q", src.Provider, ProviderREST)
}
dst, err := f.CreateFolder(ctx, root, "facade-dst-"+suffix)
if err != nil {
t.Fatalf("CreateFolder dst: %v", err)
}
if dst.Provider != ProviderREST {
t.Fatalf("CreateFolder dst provider = %q, want %q", dst.Provider, ProviderREST)
}
t.Cleanup(func() {
if err := c.DeleteDavItems(ctx, []string{src.ID, dst.ID}, nil); err != nil {
t.Logf("cleanup folders: %v", err)
}
})
content := []byte("facade crud " + suffix + "\n")
up, err := f.Upload(ctx, src.ID, "facade-doc-"+suffix+".txt", bytes.NewReader(content))
if err != nil {
t.Fatalf("Upload: %v", err)
}
if up.Kind != File || up.ID == "" {
t.Fatalf("uploaded entry: %+v", up)
}
if up.Provider != ProviderREST {
t.Fatalf("Upload provider = %q, want %q (write order REST first)", up.Provider, ProviderREST)
}
if !waitEntry(ctx, f, src.ID, up.ID, 15*time.Second) {
t.Fatalf("uploaded %s not listed in src", up.ID)
}
st, err := f.Stat(ctx, up.ID)
if err != nil {
t.Fatalf("Stat: %v", err)
}
if st.ID != up.ID || st.Kind != File {
t.Fatalf("Stat = %+v", st)
}
if st.Provider != ProviderREST {
t.Fatalf("Stat provider = %q, want %q (read order REST)", st.Provider, ProviderREST)
}
list, err := f.List(ctx, src.ID)
if err != nil {
t.Fatalf("List: %v", err)
}
if e := entryByID(list, up.ID); e == nil {
t.Fatalf("uploaded %s not in List(src)", up.ID)
} else if e.Provider != ProviderREST {
t.Fatalf("List provider = %q, want %q", e.Provider, ProviderREST)
}
renamed := "facade-renamed-" + suffix + ".txt"
renameEventually(t, ctx, f, up.ID, renamed)
moveEventually(t, ctx, f, up.ID, dst.ID)
if !waitEntry(ctx, f, dst.ID, up.ID, 20*time.Second) {
t.Fatalf("moved file %s not in dst", up.ID)
}
copied := copyEventually(t, ctx, f, up.ID, src.ID, 20*time.Second)
if copied == nil {
t.Fatalf("no copy found in src after Copy")
}
if copied.Provider != ProviderREST {
t.Fatalf("Copy provider = %q, want %q", copied.Provider, ProviderREST)
}
if err := f.Delete(ctx, []string{up.ID, copied.ID}); err != nil {
t.Fatalf("Delete: %v", err)
}
if !waitNoEntry(ctx, f, dst.ID, up.ID, 20*time.Second) {
t.Fatalf("file %s still present in dst after delete", up.ID)
}
if !waitNoEntry(ctx, f, src.ID, copied.ID, 20*time.Second) {
t.Fatalf("copy %s still present in src after delete", copied.ID)
}
}
+298
View File
@@ -0,0 +1,298 @@
package onlyoffice
import (
"context"
"errors"
"fmt"
"io"
"strings"
"testing"
)
// fakeStore is a FileStore test double; it records which backend served a call
// and returns a canned result or error.
type fakeStore struct {
name string
entries []Entry
err error
calls *[]string
}
func (f *fakeStore) record(op string) {
if f.calls != nil {
*f.calls = append(*f.calls, op+":"+f.name)
}
}
func (f *fakeStore) Name() string { return f.name }
func (f *fakeStore) List(_ context.Context, _ string) ([]Entry, error) {
f.record("list")
if f.err != nil {
return nil, f.err
}
return f.entries, nil
}
func (f *fakeStore) Stat(_ context.Context, id string) (Entry, error) {
f.record("stat")
if f.err != nil {
return Entry{}, f.err
}
return Entry{ID: id, Title: "t-" + f.name, Provider: f.name}, nil
}
func (f *fakeStore) CreateFolder(_ context.Context, _, title string) (Entry, error) {
f.record("mkdir")
if f.err != nil {
return Entry{}, f.err
}
return Entry{ID: "new", Title: title, Provider: f.name}, nil
}
func (f *fakeStore) Upload(_ context.Context, _, title string, _ io.Reader) (Entry, error) {
f.record("upload")
return Entry{ID: "up", Title: title, Provider: f.name}, f.err
}
func (f *fakeStore) Download(_ context.Context, _ string, _ io.Writer) (int64, error) {
f.record("download")
return 0, f.err
}
func (f *fakeStore) Move(_ context.Context, _ []string, _ string) error {
f.record("move")
return f.err
}
func (f *fakeStore) Copy(_ context.Context, _ []string, _ string) error {
f.record("copy")
return f.err
}
func (f *fakeStore) Rename(_ context.Context, _, _ string) error {
f.record("rename")
return f.err
}
func (f *fakeStore) Delete(_ context.Context, _ []string) error {
f.record("delete")
return f.err
}
type fakeSearcher struct{ name string }
func (s *fakeSearcher) Name() string { return s.name }
func (s *fakeSearcher) Search(_ context.Context, _ SearchQuery) ([]SearchHit, error) {
return []SearchHit{{Entry: Entry{Title: s.name}}}, nil
}
func newFacadeTestClient(stores map[string]FileStore, read, write []string) *FileClient {
return &FileClient{
stores: stores,
searchers: map[string]Searcher{},
readOrder: read,
writeOrder: write,
}
}
// TestFileClientIsFileStore guarantees the facade can stand in for the
// interface anywhere a plain FileStore is expected.
func TestFileClientIsFileStore(t *testing.T) {
var _ FileStore = (*FileClient)(nil)
}
func TestClientFilesPrefersRESTForReadsAndWrites(t *testing.T) {
c := NewClient(Credentials{})
f := c.Files()
if got := f.Read().Name(); got != ProviderREST {
t.Errorf("Read().Name() = %q, want %q", got, ProviderREST)
}
if got := f.Write().Name(); got != ProviderREST {
t.Errorf("Write().Name() = %q, want %q", got, ProviderREST)
}
if got := f.Name(); got != ProviderREST {
t.Errorf("Name() = %q, want %q", got, ProviderREST)
}
}
func TestFileClientPostgresTakesReadPriority(t *testing.T) {
pg := &fakeStore{name: ProviderPG}
f := newFacadeTestClient(
map[string]FileStore{ProviderREST: &fakeStore{name: ProviderREST}, ProviderPG: pg},
[]string{ProviderPG, ProviderREST},
[]string{ProviderREST},
)
if got := f.Read().Name(); got != ProviderPG {
t.Errorf("Read().Name() = %q, want %q", got, ProviderPG)
}
if got := f.Write().Name(); got != ProviderREST {
t.Errorf("Write().Name() = %q, want %q (PG is read-only)", got, ProviderREST)
}
}
func TestFileClientRegisterStoreNormalizesName(t *testing.T) {
pg := &fakeStore{name: "pg"}
f := newFacadeTestClient(map[string]FileStore{}, []string{ProviderPG}, nil)
f.RegisterStore(" POSTGRES ", pg)
if got := f.Read(); got != pg {
t.Fatalf("Read() = %v, want registered postgres store", got)
}
f.RegisterStore("", pg)
f.RegisterStore("pg", nil)
}
func TestFileClientReadFallsBackOnlyOnTransient(t *testing.T) {
var calls []string
primary := &fakeStore{name: "primary", err: fmt.Errorf("onlyoffice: list: 503 unavailable"), calls: &calls}
secondary := &fakeStore{name: "secondary", entries: []Entry{{ID: "1"}}, calls: &calls}
f := newFacadeTestClient(
map[string]FileStore{"primary": primary, "secondary": secondary},
[]string{"primary", "secondary"},
nil,
)
got, err := f.List(context.Background(), "root")
if err != nil {
t.Fatalf("List: %v", err)
}
if len(got) != 1 || got[0].ID != "1" {
t.Fatalf("List() = %+v, want secondary entry", got)
}
want := []string{"list:primary", "list:secondary"}
if fmt.Sprint(calls) != fmt.Sprint(want) {
t.Fatalf("call order = %v, want %v", calls, want)
}
}
func TestFileClientReadStopsOnPermanentError(t *testing.T) {
var calls []string
primary := &fakeStore{name: "primary", err: errors.New("onlyoffice: not found"), calls: &calls}
secondary := &fakeStore{name: "secondary", entries: []Entry{{ID: "1"}}, calls: &calls}
f := newFacadeTestClient(
map[string]FileStore{"primary": primary, "secondary": secondary},
[]string{"primary", "secondary"},
nil,
)
if _, err := f.List(context.Background(), "root"); err == nil {
t.Fatal("expected permanent error to be returned")
}
if len(calls) != 1 || calls[0] != "list:primary" {
t.Fatalf("secondary backend must not run on a permanent error: %v", calls)
}
}
func TestFileClientWriteUsesWriteBackend(t *testing.T) {
var calls []string
rest := &fakeStore{name: ProviderREST, calls: &calls}
dav := &fakeStore{name: ProviderDAV, calls: &calls}
f := newFacadeTestClient(
map[string]FileStore{ProviderREST: rest, ProviderDAV: dav},
[]string{ProviderREST},
[]string{ProviderREST, ProviderDAV},
)
if _, err := f.CreateFolder(context.Background(), "p", "t"); err != nil {
t.Fatalf("CreateFolder: %v", err)
}
if _, err := f.Upload(context.Background(), "p", "t", nil); err != nil {
t.Fatalf("Upload: %v", err)
}
if len(calls) != 2 || calls[0] != "mkdir:rest" || calls[1] != "upload:rest" {
t.Fatalf("write calls = %v, want REST", calls)
}
}
func TestFileClientWriteWithoutBackend(t *testing.T) {
f := newFacadeTestClient(map[string]FileStore{}, nil, nil)
if err := f.Delete(context.Background(), []string{"1"}); !errors.Is(err, errNoWriteBackend) {
t.Fatalf("Delete err = %v, want errNoWriteBackend", err)
}
if _, err := f.List(context.Background(), "root"); !errors.Is(err, errNoReadBackend) {
t.Fatalf("List err = %v, want errNoReadBackend", err)
}
}
func TestNewFileClientReadOrderIncludesSQL(t *testing.T) {
f := NewClient(Credentials{}).newFileClient()
want := []string{ProviderPG, ProviderMySQL, ProviderREST, ProviderDAV}
if fmt.Sprint(f.readOrder) != fmt.Sprint(want) {
t.Fatalf("readOrder = %v, want %v", f.readOrder, want)
}
}
// TestClientFileStoreSQLRoutingWithoutDSN checks that the SQL backend names are
// recognised and never yield nil: without a DSN the returned store surfaces the
// open error on use.
func TestClientFileStoreSQLRoutingWithoutDSN(t *testing.T) {
t.Setenv("ONLYOFFICE_DSN", "")
t.Setenv("ONLYOFFICE_PG_HOST", "")
c := NewClient(Credentials{})
for _, name := range []string{"pg", "sql", ProviderPG, ProviderMySQL} {
s := c.FileStore(name)
if s == nil {
t.Fatalf("FileStore(%q) = nil", name)
}
if _, err := s.Stat(context.Background(), "1"); err == nil {
t.Errorf("FileStore(%q).Stat without DSN: want error", name)
}
}
}
func TestFileClientMySQLStoreIsPreferredForReads(t *testing.T) {
mysql := &fakeStore{name: ProviderMySQL}
f := newFacadeTestClient(
map[string]FileStore{ProviderREST: &fakeStore{name: ProviderREST}, ProviderMySQL: mysql},
[]string{ProviderPG, ProviderMySQL, ProviderREST},
[]string{ProviderREST},
)
if got := f.Read().Name(); got != ProviderMySQL {
t.Errorf("Read().Name() = %q, want %q", got, ProviderMySQL)
}
}
// TestFileClientWriteToReadOnlyStore guarantees the facade surfaces ErrReadOnly
// when the configured write backend is the read-only SQL store.
func TestFileClientWriteToReadOnlyStore(t *testing.T) {
pg := &pgStore{driver: ProviderPG}
f := newFacadeTestClient(
map[string]FileStore{ProviderREST: &fakeStore{name: ProviderREST}, ProviderPG: pg},
[]string{ProviderPG, ProviderREST},
[]string{ProviderPG, ProviderREST},
)
ctx := context.Background()
if _, err := f.CreateFolder(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
t.Errorf("CreateFolder err = %v", err)
}
if _, err := f.Upload(ctx, "1", "x", strings.NewReader("x")); !errors.Is(err, ErrReadOnly) {
t.Errorf("Upload err = %v", err)
}
if err := f.Move(ctx, []string{"1"}, "2"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Move err = %v", err)
}
if err := f.Copy(ctx, []string{"1"}, "2"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Copy err = %v", err)
}
if err := f.Rename(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Rename err = %v", err)
}
if err := f.Delete(ctx, []string{"1"}); !errors.Is(err, ErrReadOnly) {
t.Errorf("Delete err = %v", err)
}
}
func TestFileClientSearchSelection(t *testing.T) {
f := &FileClient{searchers: map[string]Searcher{}, searchOrder: []string{ProviderES}}
_, err := f.Search()
if err == nil || !strings.Contains(err.Error(), "ONLYOFFICE_ES_URL") {
t.Fatalf("Search without backend = %v, want ONLYOFFICE_ES_URL hint", err)
}
es := &fakeSearcher{name: "fake-es"}
f.RegisterSearcher(ProviderES, es)
got, err := f.Search()
if err != nil {
t.Fatalf("Search: %v", err)
}
if got.Name() != "fake-es" {
t.Fatalf("searcher = %q, want fake-es", got.Name())
}
}
+525
View File
@@ -0,0 +1,525 @@
package onlyoffice
// Read-only SQL backend of the unified file client (epic #34, F2 #36).
//
// The goal is to read files and folders straight from the Community Server
// database, without the REST layer. Research on the live portal (VM
// `onlyoffice-v2`) showed the server runs on **MySQL 8.0** (`files_file`,
// `files_folder`, `files_folder_tree`, tenant `tenants_tenants`), not
// PostgreSQL — see docs/community-server-db.md. The store below therefore
// speaks `database/sql` and selects its driver from the DSN, so it works
// against the live MySQL today and against PostgreSQL if the portal is ever
// migrated. Every query is a SELECT; the write methods of FileStore return
// ErrReadOnly.
//
// Downloads follow the portal's S3/MinIO object layout through the shared
// MinIO helper in storage_fallback.go — no HTTP file endpoint is used.
import (
"context"
"database/sql"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"path/filepath"
"strconv"
"strings"
"time"
"github.com/go-sql-driver/mysql"
_ "github.com/jackc/pgx/v5/stdlib"
)
// Provider names for the SQL backend. ProviderPG is the value Name reports for
// a PostgreSQL connection and ProviderMySQL for MySQL.
const (
ProviderPG = "postgres"
ProviderMySQL = "mysql"
)
// ErrReadOnly is returned by every FileStore write method of the SQL backend.
var ErrReadOnly = errors.New("onlyoffice: sql file store is read-only")
const (
pgConnectTimeout = 10 * time.Second
pgSearchLimit = 50
pgSearchMaxLimit = 500
)
// PGConfig configures the read-only SQL store. DSN is a driver DSN:
// `user:pass@tcp(host:port)/onlyoffice?parseTime=true` for MySQL or a
// `postgres://` / libpq keyword string for PostgreSQL. Driver, when set,
// forces the engine ("postgres" or "mysql"); otherwise it is detected from the
// DSN. Tenant filters rows (empty means all tenants).
type PGConfig struct {
DSN string
Driver string
Tenant string
}
// PGConfigFromEnv reads ONLYOFFICE_DSN (or the ONLYOFFICE_PG_* parts),
// ONLYOFFICE_PG_DRIVER and the tenant from ONLYOFFICE_PG_TENANT /
// ONLYOFFICE_TENANT. The library never loads dotfiles — the CLI does that.
func PGConfigFromEnv() PGConfig {
dsn := strings.TrimSpace(os.Getenv("ONLYOFFICE_DSN"))
if dsn == "" {
dsn = pgDSNFromParts()
}
return PGConfig{
DSN: dsn,
Driver: strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_DRIVER")),
Tenant: firstNonEmpty(os.Getenv("ONLYOFFICE_PG_TENANT"), os.Getenv("ONLYOFFICE_TENANT")),
}
}
// pgDSNFromParts builds a libpq keyword DSN from ONLYOFFICE_PG_* variables.
// It returns "" unless a host is set, which keeps the MySQL path (ONLYOFFICE_DSN)
// the default.
func pgDSNFromParts() string {
host := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_HOST"))
if host == "" {
return ""
}
port := firstNonEmpty(os.Getenv("ONLYOFFICE_PG_PORT"), "5432")
dbname := firstNonEmpty(os.Getenv("ONLYOFFICE_PG_DBNAME"), "onlyoffice")
sslmode := firstNonEmpty(os.Getenv("ONLYOFFICE_PG_SSLMODE"), "disable")
return fmt.Sprintf("host=%s port=%s user=%s password=%s dbname=%s sslmode=%s",
host, port, os.Getenv("ONLYOFFICE_PG_USER"), os.Getenv("ONLYOFFICE_PG_PASSWORD"), dbname, sslmode)
}
// SQLFileStore opens the read-only SQL store from the environment
// (PGConfigFromEnv: ONLYOFFICE_DSN or the ONLYOFFICE_PG_* parts). It is the
// error-aware counterpart of Client.FileStore("pg"/"sql"), which returns an
// errStore when the open fails. The caller owns the returned store and should
// close it (the concrete type has a Close method).
func (c *Client) SQLFileStore() (FileStore, error) {
return NewPGStore(PGConfigFromEnv())
}
// pgStore is a read-only FileStore/Searcher over the Community Server database.
type pgStore struct {
db *sql.DB
driver string
tenantID int64
hasTenant bool
http *http.Client
}
var (
_ FileStore = (*pgStore)(nil)
_ Searcher = (*pgStore)(nil)
)
// NewPGStore opens the database and verifies connectivity. It never writes.
func NewPGStore(cfg PGConfig) (*pgStore, error) {
dsn := strings.TrimSpace(cfg.DSN)
if dsn == "" {
return nil, fmt.Errorf("onlyoffice: sql file store: empty DSN (set ONLYOFFICE_DSN)")
}
driver := pgDriver(dsn, cfg.Driver)
dsn, err := normalizeSQLDSN(driver, dsn)
if err != nil {
return nil, err
}
db, err := sql.Open(sqlDriverName(driver), dsn)
if err != nil {
return nil, fmt.Errorf("onlyoffice: sql file store: open %s: %w", driver, err)
}
ctx, cancel := context.WithTimeout(context.Background(), pgConnectTimeout)
defer cancel()
if err := db.PingContext(ctx); err != nil {
db.Close()
return nil, fmt.Errorf("onlyoffice: sql file store: ping %s: %w", driver, err)
}
s := &pgStore{db: db, driver: driver, http: &http.Client{}}
if t := strings.TrimSpace(cfg.Tenant); t != "" {
n, err := strconv.ParseInt(t, 10, 64)
if err != nil {
db.Close()
return nil, fmt.Errorf("onlyoffice: sql file store: non-numeric tenant %q", t)
}
s.tenantID, s.hasTenant = n, true
}
return s, nil
}
// Close releases the database handle.
func (s *pgStore) Close() error { return s.db.Close() }
// Name implements FileStore and Searcher.
func (s *pgStore) Name() string { return s.driver }
// pgDriver resolves the engine: the explicit value wins, otherwise the DSN
// shape decides. A leading postgres:// scheme or a libpq keyword DSN (which
// always carries '=') selects PostgreSQL; anything else is MySQL.
func pgDriver(dsn, explicit string) string {
switch strings.ToLower(strings.TrimSpace(explicit)) {
case ProviderPG, "pg", "postgresql", "pgx":
return ProviderPG
case ProviderMySQL, "mariadb":
return ProviderMySQL
}
l := strings.ToLower(strings.TrimSpace(dsn))
switch {
case strings.HasPrefix(l, "postgres://"), strings.HasPrefix(l, "postgresql://"):
return ProviderPG
case strings.HasPrefix(l, "mysql://"), strings.Contains(l, "@tcp("), strings.Contains(l, "@unix("):
return ProviderMySQL
case strings.Contains(l, "="):
return ProviderPG
default:
return ProviderMySQL
}
}
// sqlDriverName maps the engine to its registered database/sql driver.
func sqlDriverName(driver string) string {
if driver == ProviderPG {
return "pgx"
}
return "mysql"
}
// normalizeSQLDSN converts a mysql:// URL to the go-sql-driver form and forces
// parseTime so datetime columns scan into time.Time. PostgreSQL DSNs pass
// through untouched.
func normalizeSQLDSN(driver, dsn string) (string, error) {
if driver != ProviderMySQL {
return dsn, nil
}
if strings.HasPrefix(strings.ToLower(dsn), "mysql://") {
converted, err := mysqlDSNFromURL(dsn)
if err != nil {
return "", err
}
dsn = converted
}
cfg, err := mysql.ParseDSN(dsn)
if err != nil {
return "", fmt.Errorf("onlyoffice: sql file store: parse mysql DSN: %w", err)
}
cfg.ParseTime = true
return cfg.FormatDSN(), nil
}
// mysqlDSNFromURL turns mysql://user:pass@host:port/db into the driver DSN.
func mysqlDSNFromURL(raw string) (string, error) {
u, err := url.Parse(raw)
if err != nil || u.Host == "" {
return "", fmt.Errorf("onlyoffice: sql file store: bad mysql URL %q", raw)
}
user := ""
if u.User != nil {
user = u.User.Username()
if p, ok := u.User.Password(); ok {
user += ":" + p
}
}
q := u.Query()
q.Set("parseTime", "true")
return fmt.Sprintf("%s@tcp(%s)/%s?%s", user, u.Host, strings.TrimPrefix(u.Path, "/"), q.Encode()), nil
}
// rebind rewrites '?' placeholders to PostgreSQL's $1..$n. MySQL keeps them.
func rebind(query, driver string) string {
if driver != ProviderPG {
return query
}
var b strings.Builder
b.Grow(len(query) + 8)
n := 0
for _, r := range query {
if r == '?' {
n++
b.WriteByte('$')
b.WriteString(strconv.Itoa(n))
continue
}
b.WriteRune(r)
}
return b.String()
}
// List returns the folders and files directly below parentID, folders first.
func (s *pgStore) List(ctx context.Context, parentID string) ([]Entry, error) {
pid, err := parseEntryID(parentID)
if err != nil {
return nil, err
}
folders, err := s.queryFolders(ctx, "parent_id = ?", pid)
if err != nil {
return nil, err
}
files, err := s.queryFiles(ctx, "folder_id = ? AND current_version = 1", pid)
if err != nil {
return nil, err
}
out := make([]Entry, 0, len(folders)+len(files))
for _, f := range folders {
out = append(out, folderRowToEntry(f, s.Name()))
}
for _, f := range files {
out = append(out, fileRowToEntry(f, s.Name()))
}
return out, nil
}
// Stat resolves a folder or file id to an Entry. Folders win when both id
// spaces overlap (they never do on a real portal, but the lookup is cheap).
func (s *pgStore) Stat(ctx context.Context, id string) (Entry, error) {
n, err := parseEntryID(id)
if err != nil {
return Entry{}, err
}
folders, err := s.queryFolders(ctx, "id = ?", n)
if err != nil {
return Entry{}, err
}
if len(folders) > 0 {
return folderRowToEntry(folders[0], s.Name()), nil
}
files, err := s.queryFiles(ctx, "id = ? AND current_version = 1", n)
if err != nil {
return Entry{}, err
}
if len(files) == 0 {
return Entry{}, fmt.Errorf("onlyoffice: sql file store: id %s not found", id)
}
return fileRowToEntry(files[0], s.Name()), nil
}
// Download streams the file's current version from the portal's S3/MinIO store.
// The object key is reconstructed from the file id and version; the parent
// folder id is not part of the key.
func (s *pgStore) Download(ctx context.Context, id string, w io.Writer) (int64, error) {
n, err := parseEntryID(id)
if err != nil {
return 0, err
}
files, err := s.queryFiles(ctx, "id = ? AND current_version = 1", n)
if err != nil {
return 0, err
}
if len(files) == 0 {
return 0, fmt.Errorf("onlyoffice: sql file store: file %s not found", id)
}
f := files[0]
key := csObjectKey(s.tenantID, f.id, f.version, filepath.Ext(f.title))
return downloadMinioObject(ctx, s.http, key, w)
}
// CreateFolder is unavailable: the SQL backend is read-only.
func (s *pgStore) CreateFolder(context.Context, string, string) (Entry, error) {
return Entry{}, ErrReadOnly
}
// Upload is unavailable: the SQL backend is read-only.
func (s *pgStore) Upload(context.Context, string, string, io.Reader) (Entry, error) {
return Entry{}, ErrReadOnly
}
// Move is unavailable: the SQL backend is read-only.
func (s *pgStore) Move(context.Context, []string, string) error { return ErrReadOnly }
// Copy is unavailable: the SQL backend is read-only.
func (s *pgStore) Copy(context.Context, []string, string) error { return ErrReadOnly }
// Rename is unavailable: the SQL backend is read-only.
func (s *pgStore) Rename(context.Context, string, string) error { return ErrReadOnly }
// Delete is unavailable: the SQL backend is read-only.
func (s *pgStore) Delete(context.Context, []string) error { return ErrReadOnly }
// Search matches file titles by substring. Content search lives in the
// Elasticsearch backend; q.InContent is ignored here.
func (s *pgStore) Search(ctx context.Context, q SearchQuery) ([]SearchHit, error) {
text := strings.TrimSpace(q.Text)
if text == "" {
return nil, fmt.Errorf("onlyoffice: empty search query")
}
limit := q.Limit
if limit <= 0 {
limit = pgSearchLimit
}
if limit > pgSearchMaxLimit {
limit = pgSearchMaxLimit
}
where := "title LIKE ? AND current_version = 1"
args := []any{"%" + text + "%"}
if s.hasTenant {
where += " AND tenant_id = ?"
args = append(args, s.tenantID)
}
if fid := strings.TrimSpace(q.FolderID); fid != "" {
n, err := parseEntryID(fid)
if err != nil {
return nil, err
}
where += " AND folder_id = ?"
args = append(args, n)
}
for _, ext := range normalizeExtensions(q.Extensions) {
where += " AND LOWER(title) LIKE ?"
args = append(args, "%."+ext)
}
query := rebind(`SELECT id, folder_id, title, content_length, version, create_on, modified_on
FROM files_file WHERE `+where+` ORDER BY modified_on DESC, id DESC LIMIT ?`, s.driver)
args = append(args, limit)
rows, err := s.db.QueryContext(ctx, query, args...)
if err != nil {
return nil, fmt.Errorf("onlyoffice: sql search: %w", err)
}
defer rows.Close()
var hits []SearchHit
for rows.Next() {
f, err := scanFileRow(rows)
if err != nil {
return nil, err
}
hits = append(hits, SearchHit{Entry: fileRowToEntry(f, s.Name())})
}
return hits, rows.Err()
}
// queryFolders runs a folder SELECT with the tenant filter applied.
func (s *pgStore) queryFolders(ctx context.Context, where string, arg any) ([]pgFolderRow, error) {
args := []any{arg}
if s.hasTenant {
where += " AND tenant_id = ?"
args = append(args, s.tenantID)
}
query := rebind(`SELECT id, parent_id, title, create_on, modified_on
FROM files_folder WHERE `+where+` ORDER BY title, id`, s.driver)
rows, err := s.db.QueryContext(ctx, query, args...)
if err != nil {
return nil, fmt.Errorf("onlyoffice: sql list folders: %w", err)
}
defer rows.Close()
var out []pgFolderRow
for rows.Next() {
var r pgFolderRow
if err := rows.Scan(&r.id, &r.parentID, &r.title, &r.created, &r.modified); err != nil {
return nil, fmt.Errorf("onlyoffice: sql folder row: %w", err)
}
out = append(out, r)
}
return out, rows.Err()
}
// queryFiles runs a file SELECT for the current version with the tenant filter.
func (s *pgStore) queryFiles(ctx context.Context, where string, arg any) ([]pgFileRow, error) {
args := []any{arg}
if s.hasTenant {
where += " AND tenant_id = ?"
args = append(args, s.tenantID)
}
query := rebind(`SELECT id, folder_id, title, content_length, version, create_on, modified_on
FROM files_file WHERE `+where+` ORDER BY title, id`, s.driver)
rows, err := s.db.QueryContext(ctx, query, args...)
if err != nil {
return nil, fmt.Errorf("onlyoffice: sql list files: %w", err)
}
defer rows.Close()
var out []pgFileRow
for rows.Next() {
f, err := scanFileRow(rows)
if err != nil {
return nil, err
}
out = append(out, f)
}
return out, rows.Err()
}
// pgFileRow is one current files_file row.
type pgFileRow struct {
id int64
folderID int64
title string
size int64
version int
created time.Time
modified time.Time
}
// pgFolderRow is one files_folder row.
type pgFolderRow struct {
id int64
parentID int64
title string
created time.Time
modified time.Time
}
// scanFileRow reads the canonical file column order.
func scanFileRow(rows *sql.Rows) (pgFileRow, error) {
var f pgFileRow
if err := rows.Scan(&f.id, &f.folderID, &f.title, &f.size, &f.version, &f.created, &f.modified); err != nil {
return f, fmt.Errorf("onlyoffice: sql file row: %w", err)
}
return f, nil
}
// fileRowToEntry maps a files_file row to the canonical model.
func fileRowToEntry(f pgFileRow, provider string) Entry {
return Entry{
ID: strconv.FormatInt(f.id, 10),
ParentID: strconv.FormatInt(f.folderID, 10),
Title: f.title,
Kind: File,
Size: f.size,
MIME: mimeForTitle(f.title, ""),
Created: f.created.UTC(),
Modified: f.modified.UTC(),
Version: f.version,
Provider: provider,
}
}
// folderRowToEntry maps a files_folder row to the canonical model.
func folderRowToEntry(f pgFolderRow, provider string) Entry {
return Entry{
ID: strconv.FormatInt(f.id, 10),
ParentID: strconv.FormatInt(f.parentID, 10),
Title: f.title,
Kind: Folder,
Created: f.created.UTC(),
Modified: f.modified.UTC(),
Provider: provider,
}
}
// csObjectKey reconstructs the object key the portal's S3 consumer uses:
//
// 00/00/<tenant>/files/folder_<shard>/file_<id>/v<version>/content.<ext>
//
// The shard is the next thousand above the file id (file 3727 -> folder_4000),
// NOT the parent folder id — verified live against the MinIO bucket.
func csObjectKey(tenant int64, fileID int64, version int, ext string) string {
shard := (fileID/1000 + 1) * 1000
ext = strings.TrimPrefix(strings.ToLower(strings.TrimSpace(ext)), ".")
if ext == "" {
ext = "bin"
}
if version < 1 {
version = 1
}
if tenant <= 0 {
tenant = 1
}
return fmt.Sprintf("00/00/%02d/files/folder_%d/file_%d/v%d/content.%s", tenant, shard, fileID, version, ext)
}
// parseEntryID parses a numeric OnlyOffice id or returns a store error.
func parseEntryID(id string) (int64, error) {
n, err := strconv.ParseInt(strings.TrimSpace(id), 10, 64)
if err != nil {
return 0, fmt.Errorf("onlyoffice: sql file store: non-numeric id %q", id)
}
return n, nil
}
+224
View File
@@ -0,0 +1,224 @@
//go:build integration
package onlyoffice
import (
"bytes"
"context"
"errors"
"os"
"strings"
"testing"
"time"
)
// TestIntegrationPGStore exercises the read-only SQL backend against the live
// Community Server database and cross-checks list/stat/download with the REST
// FileStore. It needs ONLYOFFICE_DSN plus the usual ONLYOFFICE_URL/USER/PASS;
// ONLYOFFICE_PG_TEST_FILE_ID / ONLYOFFICE_PG_TEST_FOLDER_ID pick a real file
// (a file reachable over REST too). Download streams from MinIO, so it also
// needs MINIO_ACCESS_KEY/MINIO_SECRET_KEY.
//
// The live Community Server runs on MySQL; PostgreSQL is supported by the same
// code path when the DSN says so.
func TestIntegrationPGStore(t *testing.T) {
cfg := PGConfigFromEnv()
if strings.TrimSpace(cfg.DSN) == "" {
t.Skip("ONLYOFFICE_DSN not set — skipping SQL store integration test")
}
store, err := NewPGStore(cfg)
if err != nil {
t.Fatalf("NewPGStore: %v", err)
}
t.Cleanup(func() { _ = store.Close() })
t.Logf("sql store backend: %s", store.Name())
ctx := context.Background()
if err := testPGStoreReadOnly(ctx, store); err != nil {
t.Fatal(err)
}
fileID := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_TEST_FILE_ID"))
folderID := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_TEST_FOLDER_ID"))
if fileID == "" || folderID == "" {
t.Skip("ONLYOFFICE_PG_TEST_FILE_ID / ONLYOFFICE_PG_TEST_FOLDER_ID not set — skipping live comparison")
}
c := liveClient(t)
rest := c.Files()
dbFile, err := store.Stat(ctx, fileID)
if err != nil {
t.Fatalf("sql Stat(%s): %v", fileID, err)
}
restFile, err := rest.Stat(ctx, fileID)
if err != nil {
t.Fatalf("rest Stat(%s): %v", fileID, err)
}
if dbFile.Kind != File {
t.Errorf("sql kind = %v, want file", dbFile.Kind)
}
if dbFile.ID != restFile.ID || dbFile.Title != restFile.Title || dbFile.ParentID != restFile.ParentID {
t.Errorf("stat mismatch sql=%+v rest=%+v", dbFile, restFile)
}
// GetFile omits contentLength, so size is only comparable when REST has it.
if restFile.Size > 0 && dbFile.Size != restFile.Size {
t.Errorf("size sql=%d rest=%d", dbFile.Size, restFile.Size)
}
if d := dbFile.Modified.Sub(restFile.Modified); d > 2*time.Minute || d < -2*time.Minute {
t.Errorf("modified sql=%v rest=%v", dbFile.Modified, restFile.Modified)
}
list, err := store.List(ctx, folderID)
if err != nil {
t.Fatalf("sql List(%s): %v", folderID, err)
}
if entryByID(list, fileID) == nil {
t.Errorf("file %s not in sql List(%s)", fileID, folderID)
}
// Every file the REST layer can see in the folder must be in the SQL list
// (the SQL store sees more, so only assert this direction).
restList, err := rest.List(ctx, folderID)
if err != nil {
t.Fatalf("rest List(%s): %v", folderID, err)
}
dbIDs := make(map[string]bool, len(list))
for _, e := range list {
dbIDs[e.ID] = true
}
for _, e := range restList {
if e.Kind == File && !dbIDs[e.ID] {
t.Errorf("rest file %s (%q) missing from sql list", e.ID, e.Title)
}
}
// Download reads the object store, not the database, so it only runs with
// the MinIO credentials configured (the portal's S3 layout). Without them
// the DSN-only assertions above still prove the SQL reads.
if os.Getenv("MINIO_ACCESS_KEY") == "" || os.Getenv("MINIO_SECRET_KEY") == "" {
t.Log("MINIO_ACCESS_KEY/MINIO_SECRET_KEY not set — SQL download cross-check skipped")
return
}
var buf bytes.Buffer
n, err := store.Download(ctx, fileID, &buf)
if err != nil {
t.Fatalf("sql Download(%s): %v", fileID, err)
}
if n == 0 || n != dbFile.Size {
t.Errorf("sql Download = %d bytes, stat says %d", n, dbFile.Size)
}
var restBuf bytes.Buffer
rn, err := rest.Download(ctx, fileID, &restBuf)
if err != nil {
t.Fatalf("rest Download(%s): %v", fileID, err)
}
if rn != n || !bytes.Equal(restBuf.Bytes(), buf.Bytes()) {
t.Errorf("download mismatch sql=%d rest=%d bytes", n, rn)
}
}
// TestIntegrationSQLFacade proves that reads are served by the SQL store when
// it is part of the file client, not by REST. It needs ONLYOFFICE_DSN plus
// ONLYOFFICE_PG_TEST_FILE_ID / ONLYOFFICE_PG_TEST_FOLDER_ID and the usual REST
// credentials (for the cross-check). Every Entry served by SQL carries
// Provider "mysql"; REST entries carry "rest", so the provider is the proof of
// which backend answered.
func TestIntegrationSQLFacade(t *testing.T) {
cfg := PGConfigFromEnv()
if strings.TrimSpace(cfg.DSN) == "" {
t.Skip("ONLYOFFICE_DSN not set — skipping SQL facade integration test")
}
fileID := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_TEST_FILE_ID"))
folderID := strings.TrimSpace(os.Getenv("ONLYOFFICE_PG_TEST_FOLDER_ID"))
if fileID == "" || folderID == "" {
t.Skip("ONLYOFFICE_PG_TEST_FILE_ID / ONLYOFFICE_PG_TEST_FOLDER_ID not set — skipping SQL facade integration test")
}
c := liveClient(t)
ctx := context.Background()
sqlStore, err := c.SQLFileStore()
if err != nil {
t.Fatalf("SQLFileStore: %v", err)
}
if closer, ok := sqlStore.(interface{ Close() error }); ok {
t.Cleanup(func() { _ = closer.Close() })
}
if sqlStore.Name() == ProviderREST {
t.Fatalf("SQLFileStore returned REST")
}
// Direct constructor: Client.FileStore("pg") must not be REST either.
direct := c.FileStore("pg")
if direct == nil || direct.Name() == ProviderREST {
t.Fatalf("FileStore(\"pg\") = %v, want SQL backend", direct)
}
if closer, ok := direct.(interface{ Close() error }); ok {
t.Cleanup(func() { _ = closer.Close() })
}
f := c.Files()
f.RegisterStore(ProviderPG, sqlStore)
if got := f.Read().Name(); got != sqlStore.Name() {
t.Fatalf("facade read backend = %q, want %q (SQL)", got, sqlStore.Name())
}
got, err := f.Stat(ctx, fileID)
if err != nil {
t.Fatalf("facade Stat(%s): %v", fileID, err)
}
if got.Provider != ProviderMySQL {
t.Errorf("facade Stat provider = %q, want %q (SQL, not REST)", got.Provider, ProviderMySQL)
}
want, err := c.FileStore(ProviderREST).Stat(ctx, fileID)
if err != nil {
t.Fatalf("rest Stat(%s): %v", fileID, err)
}
if got.ID != want.ID || got.Title != want.Title || got.ParentID != want.ParentID {
t.Errorf("facade SQL stat %+v != REST %+v", got, want)
}
list, err := f.List(ctx, folderID)
if err != nil {
t.Fatalf("facade List(%s): %v", folderID, err)
}
entry := entryByID(list, fileID)
if entry == nil {
t.Fatalf("file %s not in facade List(%s)", fileID, folderID)
}
if entry.Provider != ProviderMySQL {
t.Errorf("facade List provider = %q, want %q", entry.Provider, ProviderMySQL)
}
d, err := direct.Stat(ctx, fileID)
if err != nil {
t.Fatalf("FileStore(\"pg\").Stat(%s): %v", fileID, err)
}
if d.Provider != ProviderMySQL {
t.Errorf("FileStore(\"pg\") provider = %q, want %q", d.Provider, ProviderMySQL)
}
}
// testPGStoreReadOnly asserts that every write method returns ErrReadOnly.
func testPGStoreReadOnly(ctx context.Context, s *pgStore) error {
if _, err := s.CreateFolder(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
return errors.New("CreateFolder did not return ErrReadOnly")
}
if _, err := s.Upload(ctx, "1", "x", strings.NewReader("x")); !errors.Is(err, ErrReadOnly) {
return errors.New("Upload did not return ErrReadOnly")
}
if err := s.Move(ctx, []string{"1"}, "2"); !errors.Is(err, ErrReadOnly) {
return errors.New("Move did not return ErrReadOnly")
}
if err := s.Copy(ctx, []string{"1"}, "2"); !errors.Is(err, ErrReadOnly) {
return errors.New("Copy did not return ErrReadOnly")
}
if err := s.Rename(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
return errors.New("Rename did not return ErrReadOnly")
}
if err := s.Delete(ctx, []string{"1"}); !errors.Is(err, ErrReadOnly) {
return errors.New("Delete did not return ErrReadOnly")
}
return nil
}
+165
View File
@@ -0,0 +1,165 @@
package onlyoffice
import (
"context"
"errors"
"testing"
"time"
)
func TestRebind(t *testing.T) {
mysqlQuery := "SELECT id FROM files_file WHERE folder_id = ? AND title = ? LIMIT ?"
if got := rebind(mysqlQuery, ProviderMySQL); got != mysqlQuery {
t.Errorf("mysql query changed: %q", got)
}
want := "SELECT id FROM files_file WHERE folder_id = $1 AND title = $2 LIMIT $3"
if got := rebind(mysqlQuery, ProviderPG); got != want {
t.Errorf("rebind = %q, want %q", got, want)
}
}
func TestCSPObjectKey(t *testing.T) {
cases := []struct {
tenant int64
fileID int64
version int
ext string
want string
}{
{1, 2, 1, ".docx", "00/00/01/files/folder_1000/file_2/v1/content.docx"},
{1, 999, 1, ".pdf", "00/00/01/files/folder_1000/file_999/v1/content.pdf"},
{1, 1000, 1, ".xlsx", "00/00/01/files/folder_2000/file_1000/v1/content.xlsx"},
{1, 3727, 1, ".pdf", "00/00/01/files/folder_4000/file_3727/v1/content.pdf"},
{1, 22484, 1, ".PDF", "00/00/01/files/folder_23000/file_22484/v1/content.pdf"},
{1, 4, 6, "xlsx", "00/00/01/files/folder_1000/file_4/v6/content.xlsx"},
{0, 7, 0, "", "00/00/01/files/folder_1000/file_7/v1/content.bin"},
{2, 11, 3, ".doc", "00/00/02/files/folder_1000/file_11/v3/content.doc"},
}
for _, tc := range cases {
if got := csObjectKey(tc.tenant, tc.fileID, tc.version, tc.ext); got != tc.want {
t.Errorf("csObjectKey(%d,%d,%d,%q) = %q, want %q", tc.tenant, tc.fileID, tc.version, tc.ext, got, tc.want)
}
}
}
func TestPGDriverDetection(t *testing.T) {
cases := []struct {
dsn, explicit, want string
}{
{"postgres://u:p@h:5432/onlyoffice", "", ProviderPG},
{"postgresql://u:p@h/db", "", ProviderPG},
{"host=h user=u password=p dbname=onlyoffice sslmode=disable", "", ProviderPG},
{"root:secret@tcp(127.0.0.1:3306)/onlyoffice?parseTime=true", "", ProviderMySQL},
{"mysql://root:secret@127.0.0.1:3306/onlyoffice", "", ProviderMySQL},
{"root:secret@tcp(h:3306)/db", "postgres", ProviderPG},
{"postgres://u:p@h/db", "mysql", ProviderMySQL},
}
for _, tc := range cases {
if got := pgDriver(tc.dsn, tc.explicit); got != tc.want {
t.Errorf("pgDriver(%q, %q) = %q, want %q", tc.dsn, tc.explicit, got, tc.want)
}
}
}
func TestNormalizeSQLDSNMySQL(t *testing.T) {
got, err := normalizeSQLDSN(ProviderMySQL, "mysql://root:secret@127.0.0.1:3306/onlyoffice")
if err != nil {
t.Fatalf("normalizeSQLDSN: %v", err)
}
want := "root:secret@tcp(127.0.0.1:3306)/onlyoffice?parseTime=true"
if got != want {
t.Errorf("normalize = %q, want %q", got, want)
}
// A driver DSN keeps parseTime and gains it when missing.
got, err = normalizeSQLDSN(ProviderMySQL, "root:secret@tcp(127.0.0.1:3306)/onlyoffice")
if err != nil {
t.Fatalf("normalizeSQLDSN: %v", err)
}
if got != want {
t.Errorf("normalize = %q, want %q", got, want)
}
}
func TestFileRowToEntry(t *testing.T) {
created := time.Date(2026, 9, 12, 18, 0, 37, 0, time.UTC)
modified := time.Date(2026, 9, 13, 13, 50, 36, 0, time.UTC)
e := fileRowToEntry(pgFileRow{
id: 22484, folderID: 649, title: "Rechnung.pdf",
size: 123433, version: 2, created: created, modified: modified,
}, ProviderMySQL)
if e.ID != "22484" || e.ParentID != "649" {
t.Errorf("ids = %q/%q", e.ID, e.ParentID)
}
if e.Title != "Rechnung.pdf" || e.Kind != File {
t.Errorf("title/kind = %q/%v", e.Title, e.Kind)
}
if e.Size != 123433 || e.Version != 2 {
t.Errorf("size/version = %d/%d", e.Size, e.Version)
}
if e.MIME != "application/pdf" {
t.Errorf("mime = %q", e.MIME)
}
if !e.Created.Equal(created) || !e.Modified.Equal(modified) {
t.Errorf("times = %v/%v", e.Created, e.Modified)
}
if e.Provider != ProviderMySQL {
t.Errorf("provider = %q", e.Provider)
}
}
func TestFolderRowToEntry(t *testing.T) {
modified := time.Date(2026, 8, 1, 10, 30, 0, 0, time.UTC)
e := folderRowToEntry(pgFolderRow{id: 649, parentID: 647, title: "2025", modified: modified}, ProviderMySQL)
if e.ID != "649" || e.ParentID != "647" || e.Title != "2025" {
t.Errorf("folder = %+v", e)
}
if e.Kind != Folder {
t.Errorf("kind = %v, want folder", e.Kind)
}
if e.Size != 0 || e.MIME != "" {
t.Errorf("folder size/mime = %d/%q", e.Size, e.MIME)
}
if !e.Modified.Equal(modified) {
t.Errorf("modified = %v", e.Modified)
}
}
func TestPGStoreWriteMethodsReadOnly(t *testing.T) {
s := &pgStore{driver: ProviderPG}
ctx := context.Background()
if _, err := s.CreateFolder(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
t.Errorf("CreateFolder err = %v", err)
}
if _, err := s.Upload(ctx, "1", "x", nil); !errors.Is(err, ErrReadOnly) {
t.Errorf("Upload err = %v", err)
}
if err := s.Move(ctx, nil, "1"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Move err = %v", err)
}
if err := s.Copy(ctx, nil, "1"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Copy err = %v", err)
}
if err := s.Rename(ctx, "1", "x"); !errors.Is(err, ErrReadOnly) {
t.Errorf("Rename err = %v", err)
}
if err := s.Delete(ctx, nil); !errors.Is(err, ErrReadOnly) {
t.Errorf("Delete err = %v", err)
}
}
func TestPGStoreName(t *testing.T) {
if got := (&pgStore{driver: ProviderPG}).Name(); got != ProviderPG {
t.Errorf("Name = %q, want %q", got, ProviderPG)
}
if got := (&pgStore{driver: ProviderMySQL}).Name(); got != ProviderMySQL {
t.Errorf("Name = %q, want %q", got, ProviderMySQL)
}
}
func TestPGStoreStatRejectsNonNumeric(t *testing.T) {
s := &pgStore{driver: ProviderPG}
if _, err := s.Stat(context.Background(), "not-a-number"); err == nil {
t.Error("Stat accepted a non-numeric id")
}
}
+256
View File
@@ -0,0 +1,256 @@
package onlyoffice
// restStore implements FileStore on top of the REST Documents methods in
// files.go. It is a thin adapter: no endpoint logic lives here, and every call
// is wrapped in DoRetry.
import (
"context"
"encoding/json"
"io"
"os"
"path/filepath"
)
// restStore is a FileStore over the REST Documents API.
type restStore struct{ c *Client }
// Name reports the backend name.
func (s *restStore) Name() string { return ProviderREST }
// List returns the files and folders directly below parentID.
func (s *restStore) List(ctx context.Context, parentID string) ([]Entry, error) {
var out []Entry
err := retryStoreOp(ctx, func() error {
raw, err := s.c.ListFolder(ctx, parentID)
if err != nil {
return err
}
entries, err := entriesFromFolderMap(raw, ProviderREST)
if err != nil {
return err
}
out = entries
return nil
})
return out, err
}
// Stat returns file or folder metadata. Folders are resolved through the
// listing endpoint (their own id appears as the listing's Current); other ids
// fall back to the file metadata API.
func (s *restStore) Stat(ctx context.Context, id string) (Entry, error) {
return s.stat(ctx, id)
}
// stat resolves a single id to a folder or file Entry.
func (s *restStore) stat(ctx context.Context, id string) (Entry, error) {
var out Entry
err := retryStoreOp(ctx, func() error {
if l, err := s.c.ListDavFolder(ctx, id); err == nil {
if l != nil && l.Current.ID != "" && l.Current.ID == id {
out = DavFolderToEntry(l.Current, ProviderREST)
return nil
}
} else if Transient(err) {
return err
}
f, err := s.c.GetFile(ctx, id)
if err != nil {
return err
}
out = FileEntryToEntry(f, ProviderREST)
return nil
})
return out, err
}
// CreateFolder creates a subfolder under parentID.
func (s *restStore) CreateFolder(ctx context.Context, parentID, title string) (Entry, error) {
var out Entry
err := retryStoreOp(ctx, func() error {
m, err := s.c.CreateFolder(ctx, parentID, title)
if err != nil {
return err
}
e, err := folderEntryFromMap(m, parentID, ProviderREST)
if err != nil {
return err
}
if e.ParentID == "" {
e.ParentID = parentID
}
if e.Title == "" {
e.Title = title
}
out = e
return nil
})
return out, err
}
// Upload streams r into parentID as title. UploadToFolder is path based, so
// the reader is spooled to a temporary file first (ponytail: OnlyOffice
// multipart upload buffers the whole body anyway).
func (s *restStore) Upload(ctx context.Context, parentID, title string, r io.Reader) (Entry, error) {
dir, err := os.MkdirTemp("", "oo-rest-upload-")
if err != nil {
return Entry{}, err
}
defer os.RemoveAll(dir)
local := filepath.Join(dir, SafeLocalFileName(title))
f, err := os.Create(local)
if err != nil {
return Entry{}, err
}
if _, err := io.Copy(f, r); err != nil {
f.Close()
return Entry{}, err
}
if err := f.Close(); err != nil {
return Entry{}, err
}
var out Entry
err = retryStoreOp(ctx, func() error {
fe, err := s.c.UploadToFolder(ctx, parentID, local)
if err != nil {
return err
}
out = FileEntryToEntry(fe, ProviderREST)
return nil
})
return out, err
}
// Download streams the file bytes into w.
func (s *restStore) Download(ctx context.Context, id string, w io.Writer) (int64, error) {
var n int64
err := retryStoreOp(ctx, func() error {
var e error
n, e = s.c.DownloadFile(ctx, id, w)
return e
})
return n, err
}
// Move moves folders and/or files into parentID. Ids are classified through
// stat so folder moves use folderIds and file moves use fileIds on the shared
// fileops/move endpoint.
func (s *restStore) Move(ctx context.Context, ids []string, parentID string) error {
folders, files, err := s.split(ctx, ids)
if err != nil {
return err
}
if len(folders) == 0 && len(files) == 0 {
return nil
}
return retryStoreOp(ctx, func() error {
return s.c.MoveDavItems(ctx, folders, files, parentID)
})
}
// Copy copies folders and/or files into parentID. files.go has no copy method,
// so the shared REST fileops copy endpoint (CopyDavItems) is used.
func (s *restStore) Copy(ctx context.Context, ids []string, parentID string) error {
folders, files, err := s.split(ctx, ids)
if err != nil {
return err
}
if len(folders) == 0 && len(files) == 0 {
return nil
}
return retryStoreOp(ctx, func() error {
return s.c.CopyDavItems(ctx, folders, files, parentID)
})
}
// Rename sets a new title (including extension) for a file or folder.
func (s *restStore) Rename(ctx context.Context, id, title string) error {
e, err := s.stat(ctx, id)
if err != nil {
return err
}
if e.Kind == Folder {
return retryStoreOp(ctx, func() error {
return s.c.RenameDavFolder(ctx, id, title)
})
}
return retryStoreOp(ctx, func() error {
_, err := s.c.RenameFile(ctx, id, title)
return err
})
}
// Delete permanently deletes folders and/or files.
func (s *restStore) Delete(ctx context.Context, ids []string) error {
folders, files, err := s.split(ctx, ids)
if err != nil {
return err
}
if len(folders) == 0 && len(files) == 0 {
return nil
}
return retryStoreOp(ctx, func() error {
return s.c.DeleteDavItems(ctx, folders, files)
})
}
// split classifies ids into folder and file id lists.
func (s *restStore) split(ctx context.Context, ids []string) (folders, files []string, err error) {
for _, id := range ids {
e, err := s.stat(ctx, id)
if err != nil {
return nil, nil, err
}
if e.Kind == Folder {
folders = append(folders, id)
} else {
files = append(files, id)
}
}
return folders, files, nil
}
// entriesFromFolderMap converts a ListFolder response map into canonical
// entries, reusing the DavFile/DavFolder decoders for robust size handling.
func entriesFromFolderMap(m map[string]any, provider string) ([]Entry, error) {
if m == nil {
return nil, nil
}
b, err := json.Marshal(m)
if err != nil {
return nil, err
}
var listing DavListing
if err := json.Unmarshal(b, &listing); err != nil {
return nil, err
}
out := make([]Entry, 0, len(listing.Folders)+len(listing.Files))
for _, f := range listing.Folders {
out = append(out, DavFolderToEntry(f, provider))
}
for _, f := range listing.Files {
out = append(out, DavFileToEntry(f, provider))
}
return out, nil
}
// folderEntryFromMap converts a CreateFolder response map into a folder Entry.
func folderEntryFromMap(m map[string]any, parentID, provider string) (Entry, error) {
e := Entry{Kind: Folder, Provider: provider, ParentID: parentID}
if m == nil {
return e, nil
}
b, err := json.Marshal(m)
if err != nil {
return e, err
}
var f DavFolder
if err := json.Unmarshal(b, &f); err != nil {
return e, err
}
e = DavFolderToEntry(f, provider)
return e, nil
}
+337
View File
@@ -0,0 +1,337 @@
package onlyoffice
// Text extraction pipeline for the own full-text index (epic #34, F6 #42).
//
// TextIndexer downloads stored documents, extracts text through docpipe
// (pdftotext; OCR for scans) and writes the result to a TextIndex. For PDFs it
// also indexes the text of embedded attachments (pdfdetach), so a scan filed
// as an attachment is searchable too. It is the write side of ESTextIndex and
// never touches the OnlyOffice server's own ES index.
import (
"context"
"fmt"
"os"
"path/filepath"
"strings"
"sync"
"github.com/eslider/go-onlyoffice/internal/docpipe"
)
// defaultTextIndexExts are the formats extracted by default. The OnlyOffice
// index already covers docx/xlsx/pptx; F6 adds PDF.
var defaultTextIndexExts = []string{"pdf"}
const (
defaultTextIndexWorkers = 3
defaultTextIndexLang = "deu+eng"
maxTextIndexErrors = 20
)
// TextExtractor turns a local file into indexable plain text. The default uses
// docpipe (pdftotext + OCR); tests inject a fake to stay offline.
type TextExtractor interface {
Extract(path, workDir, lang string, minChars int) (string, error)
}
// docpipeExtractor is the production TextExtractor.
type docpipeExtractor struct{ tools docpipe.Tools }
// Extract renders the file as Markdown, OCRing PDFs/images with a weak text
// layer first and appending the text of embedded PDF attachments
// (docpipe.ToMarkdownWithAttachments).
func (d docpipeExtractor) Extract(path, workDir, lang string, minChars int) (string, error) {
text, err := d.tools.ToMarkdownWithAttachments(path, workDir, lang, minChars)
if err != nil {
return "", err
}
return strings.TrimSpace(text), nil
}
// IndexOptions controls a TextIndexer run.
type IndexOptions struct {
Recursive bool // IndexFolder: descend into subfolders
Extensions []string // empty = defaultTextIndexExts (pdf)
Limit int // max files to index, 0 = all
Lang string // OCR language(s), default deu+eng
MinChars int // OCR threshold, default docpipe.DefaultMinTextChars
Workers int // parallel downloads/extractions, default 3
}
// IndexResult summarises a run.
type IndexResult struct {
Scanned int
Indexed int
Skipped int
Failed int
Errors []string
}
// TextIndexer wires a FileStore, a TextIndex and an extractor together.
type TextIndexer struct {
Store FileStore
Index TextIndex
Extractor TextExtractor // nil = local docpipe tools
WorkDir string // temp dir for downloads/extraction
}
// NewTextIndexer returns a TextIndexer over the given store and index.
func NewTextIndexer(store FileStore, index TextIndex) *TextIndexer {
return &TextIndexer{Store: store, Index: index}
}
// textIndexEnsurer is implemented by indexes that can be created up front.
type textIndexEnsurer interface {
Ensure(ctx context.Context) error
}
// Ensure creates the backing index when the TextIndex supports it.
func (ix *TextIndexer) Ensure(ctx context.Context) error {
if e, ok := ix.Index.(textIndexEnsurer); ok {
return e.Ensure(ctx)
}
return nil
}
// IndexFiles stats the given file ids and indexes them.
func (ix *TextIndexer) IndexFiles(ctx context.Context, ids []string, opts IndexOptions) (IndexResult, error) {
entries := make([]Entry, 0, len(ids))
for _, id := range ids {
e, err := ix.Store.Stat(ctx, id)
if err != nil {
return IndexResult{}, fmt.Errorf("onlyoffice: stat %s: %w", id, err)
}
entries = append(entries, e)
}
return ix.IndexEntries(ctx, entries, opts)
}
// IndexFolder lists a folder and indexes every matching file.
func (ix *TextIndexer) IndexFolder(ctx context.Context, folderID string, opts IndexOptions) (IndexResult, error) {
entries, err := ix.collect(ctx, folderID, opts.Recursive)
if err != nil {
return IndexResult{}, err
}
return ix.IndexEntries(ctx, entries, opts)
}
// PlanFolder lists the files IndexFolder would process, without downloading or
// extracting anything.
func (ix *TextIndexer) PlanFolder(ctx context.Context, folderID string, opts IndexOptions) ([]Entry, error) {
entries, err := ix.collect(ctx, folderID, opts.Recursive)
if err != nil {
return nil, err
}
return selectEntries(entries, opts), nil
}
// PlanFiles stats the ids and returns those that would be indexed.
func (ix *TextIndexer) PlanFiles(ctx context.Context, ids []string, opts IndexOptions) ([]Entry, error) {
entries := make([]Entry, 0, len(ids))
for _, id := range ids {
e, err := ix.Store.Stat(ctx, id)
if err != nil {
return nil, fmt.Errorf("onlyoffice: stat %s: %w", id, err)
}
entries = append(entries, e)
}
return selectEntries(entries, opts), nil
}
// IndexEntries extracts and indexes the given files (folders are ignored).
func (ix *TextIndexer) IndexEntries(ctx context.Context, entries []Entry, opts IndexOptions) (IndexResult, error) {
opts = opts.withDefaults()
var res IndexResult
work := selectEntries(entries, opts)
res.Scanned = len(entries)
res.Skipped = len(entries) - len(work)
if len(work) == 0 {
return res, nil
}
workers := opts.Workers
if workers > len(work) {
workers = len(work)
}
if workers < 1 {
workers = 1
}
type outcome struct {
doc TextDoc
err error
}
jobs := make(chan Entry)
results := make(chan outcome, workers)
var wg sync.WaitGroup
for i := 0; i < workers; i++ {
wg.Add(1)
go func() {
defer wg.Done()
for e := range jobs {
if err := ctx.Err(); err != nil {
results <- outcome{err: err}
continue
}
doc, err := ix.indexOne(ctx, e, opts)
results <- outcome{doc: doc, err: err}
}
}()
}
go func() {
defer close(jobs)
for _, e := range work {
select {
case jobs <- e:
case <-ctx.Done():
return
}
}
}()
go func() {
wg.Wait()
close(results)
}()
var docs []TextDoc
for r := range results {
if r.err != nil {
res.Failed++
if len(res.Errors) < maxTextIndexErrors {
res.Errors = append(res.Errors, r.err.Error())
}
continue
}
docs = append(docs, r.doc)
}
if err := ctx.Err(); err != nil {
return res, err
}
if len(docs) > 0 {
if err := ix.Index.Put(ctx, docs); err != nil {
return res, fmt.Errorf("onlyoffice: index %d docs: %w", len(docs), err)
}
res.Indexed = len(docs)
}
return res, nil
}
// indexOne downloads and extracts a single file.
func (ix *TextIndexer) indexOne(ctx context.Context, e Entry, opts IndexOptions) (TextDoc, error) {
ext := fileExt(e.Title)
dir := ix.WorkDir
if dir == "" {
dir = os.TempDir()
}
if err := os.MkdirAll(dir, 0o755); err != nil {
return TextDoc{}, err
}
tmp, err := os.CreateTemp(dir, "ooidx-*."+ext)
if err != nil {
return TextDoc{}, err
}
tmpPath := tmp.Name()
defer os.Remove(tmpPath)
if _, err := ix.Store.Download(ctx, e.ID, tmp); err != nil {
tmp.Close()
return TextDoc{}, fmt.Errorf("download %s (%s): %w", e.ID, e.Title, err)
}
if err := tmp.Close(); err != nil {
return TextDoc{}, err
}
text, err := ix.extractor().Extract(tmpPath, dir, opts.Lang, opts.MinChars)
if err != nil {
return TextDoc{}, fmt.Errorf("extract %s: %w", e.Title, err)
}
return TextDoc{ID: e.ID, Title: e.Title, FolderID: e.ParentID, Ext: ext, Content: text}, nil
}
// collect lists files under folderID, breadth-first when recursive.
func (ix *TextIndexer) collect(ctx context.Context, folderID string, recursive bool) ([]Entry, error) {
var files []Entry
queue := []string{folderID}
for len(queue) > 0 {
if err := ctx.Err(); err != nil {
return nil, err
}
id := queue[0]
queue = queue[1:]
entries, err := ix.Store.List(ctx, id)
if err != nil {
return nil, fmt.Errorf("onlyoffice: list folder %s: %w", id, err)
}
for _, e := range entries {
if e.Kind == Folder {
if recursive {
queue = append(queue, e.ID)
}
continue
}
if e.ParentID == "" {
e.ParentID = id
}
files = append(files, e)
}
}
return files, nil
}
// selectEntries filters files by extension and applies the limit.
func selectEntries(entries []Entry, opts IndexOptions) []Entry {
allowed := extensionSet(opts.Extensions)
work := make([]Entry, 0, len(entries))
for _, e := range entries {
if e.Kind != File {
continue
}
if opts.Limit > 0 && len(work) >= opts.Limit {
break
}
if !allowed[fileExt(e.Title)] {
continue
}
work = append(work, e)
}
return work
}
// extensionSet normalises the extension allow-list (default: pdf).
func extensionSet(exts []string) map[string]bool {
if len(exts) == 0 {
exts = defaultTextIndexExts
}
set := make(map[string]bool, len(exts))
for _, e := range normalizeExtensions(exts) {
set[e] = true
}
return set
}
// fileExt returns the lower-case extension without the dot.
func fileExt(title string) string {
return strings.ToLower(strings.TrimPrefix(filepath.Ext(strings.TrimSpace(title)), "."))
}
// withDefaults fills zero-valued options.
func (o IndexOptions) withDefaults() IndexOptions {
if o.Workers <= 0 {
o.Workers = defaultTextIndexWorkers
}
if o.MinChars <= 0 {
o.MinChars = docpipe.DefaultMinTextChars
}
if strings.TrimSpace(o.Lang) == "" {
o.Lang = defaultTextIndexLang
}
return o
}
// extractor returns the configured extractor or the local docpipe default.
func (ix *TextIndexer) extractor() TextExtractor {
if ix.Extractor != nil {
return ix.Extractor
}
return docpipeExtractor{tools: docpipe.LookPath()}
}
+21 -3
View File
@@ -4,26 +4,33 @@ go 1.25.0
require ( require (
github.com/JohannesKaufmann/html-to-markdown/v2 v2.5.2 github.com/JohannesKaufmann/html-to-markdown/v2 v2.5.2
github.com/aws/aws-sdk-go-v2 v1.41.1
github.com/charmbracelet/bubbles v0.18.0 github.com/charmbracelet/bubbles v0.18.0
github.com/charmbracelet/bubbletea v0.25.0 github.com/charmbracelet/bubbletea v0.25.0
github.com/charmbracelet/glamour v0.8.0 github.com/charmbracelet/glamour v0.8.0
github.com/charmbracelet/lipgloss v0.12.1 github.com/charmbracelet/lipgloss v0.12.1
github.com/charmbracelet/x/ansi v0.1.4 github.com/charmbracelet/x/ansi v0.1.4
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3
github.com/eslider/go-hocr v0.2.2-0.20260827163626-8ff01582b002
github.com/eslider/go-xls/v2 v2.1.0 github.com/eslider/go-xls/v2 v2.1.0
github.com/go-sql-driver/mysql v1.10.1
github.com/google/go-querystring v1.2.0 github.com/google/go-querystring v1.2.0
github.com/jackc/pgx/v5 v5.11.0
github.com/joho/godotenv v1.5.1 github.com/joho/godotenv v1.5.1
github.com/mattn/go-runewidth v0.0.15 github.com/mattn/go-runewidth v0.0.15
github.com/muesli/termenv v0.16.0 github.com/muesli/termenv v0.16.0
github.com/spf13/cobra v1.10.2 github.com/spf13/cobra v1.10.2
github.com/xuri/excelize/v2 v2.11.0
gopkg.in/yaml.v3 v3.0.1 gopkg.in/yaml.v3 v3.0.1
modernc.org/sqlite v1.56.0 modernc.org/sqlite v1.56.0
) )
require ( require (
filippo.io/edwards25519 v1.2.0 // indirect
github.com/JohannesKaufmann/dom v0.3.1 // indirect github.com/JohannesKaufmann/dom v0.3.1 // indirect
github.com/alecthomas/chroma/v2 v2.14.0 // indirect github.com/alecthomas/chroma/v2 v2.14.0 // indirect
github.com/atotto/clipboard v0.1.4 // indirect github.com/atotto/clipboard v0.1.4 // indirect
github.com/aws/smithy-go v1.24.0 // indirect
github.com/aymanbagabas/go-osc52/v2 v2.0.1 // indirect github.com/aymanbagabas/go-osc52/v2 v2.0.1 // indirect
github.com/aymerick/douceur v0.2.0 // indirect github.com/aymerick/douceur v0.2.0 // indirect
github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81 // indirect github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81 // indirect
@@ -32,6 +39,10 @@ require (
github.com/google/uuid v1.6.0 // indirect github.com/google/uuid v1.6.0 // indirect
github.com/gorilla/css v1.0.1 // indirect github.com/gorilla/css v1.0.1 // indirect
github.com/inconshreveable/mousetrap v1.1.0 // indirect github.com/inconshreveable/mousetrap v1.1.0 // indirect
github.com/jackc/pgpassfile v1.0.0 // indirect
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
github.com/jackc/puddle/v2 v2.2.2 // indirect
github.com/kr/text v0.2.0 // indirect
github.com/lucasb-eyer/go-colorful v1.4.0 // indirect github.com/lucasb-eyer/go-colorful v1.4.0 // indirect
github.com/mattn/go-isatty v0.0.24 // indirect github.com/mattn/go-isatty v0.0.24 // indirect
github.com/mattn/go-localereader v0.0.1 // indirect github.com/mattn/go-localereader v0.0.1 // indirect
@@ -41,15 +52,22 @@ require (
github.com/muesli/reflow v0.3.0 // indirect github.com/muesli/reflow v0.3.0 // indirect
github.com/ncruces/go-strftime v1.0.0 // indirect github.com/ncruces/go-strftime v1.0.0 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/richardlehane/mscfb v1.0.7 // indirect
github.com/richardlehane/msoleps v1.0.6 // indirect
github.com/rivo/uniseg v0.4.7 // indirect github.com/rivo/uniseg v0.4.7 // indirect
github.com/rogpeppe/go-internal v1.16.0 // indirect
github.com/spf13/pflag v1.0.9 // indirect github.com/spf13/pflag v1.0.9 // indirect
github.com/tiendc/go-deepcopy v1.7.2 // indirect
github.com/xuri/efp v0.0.1 // indirect
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9 // indirect
github.com/yuin/goldmark v1.8.2 // indirect github.com/yuin/goldmark v1.8.2 // indirect
github.com/yuin/goldmark-emoji v1.0.3 // indirect github.com/yuin/goldmark-emoji v1.0.3 // indirect
golang.org/x/net v0.55.0 // indirect golang.org/x/crypto v0.53.0 // indirect
golang.org/x/net v0.56.0 // indirect
golang.org/x/sync v0.21.0 // indirect golang.org/x/sync v0.21.0 // indirect
golang.org/x/sys v0.47.0 // indirect golang.org/x/sys v0.47.0 // indirect
golang.org/x/term v0.43.0 // indirect golang.org/x/term v0.44.0 // indirect
golang.org/x/text v0.37.0 // indirect golang.org/x/text v0.38.0 // indirect
modernc.org/libc v1.74.4 // indirect modernc.org/libc v1.74.4 // indirect
modernc.org/mathutil v1.7.1 // indirect modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect modernc.org/memory v1.11.0 // indirect
+58 -7
View File
@@ -1,3 +1,5 @@
filippo.io/edwards25519 v1.2.0 h1:crnVqOiS4jqYleHd9vaKZ+HKtHfllngJIiOpNpoJsjo=
filippo.io/edwards25519 v1.2.0/go.mod h1:xzAOLCNug/yB62zG1bQ8uziwrIqIuxhctzJT18Q77mc=
github.com/JohannesKaufmann/dom v0.3.1 h1:J16l9JAHWgkFPR3VIPbQ1gvS0cWab6laK1q7PFL3qh0= github.com/JohannesKaufmann/dom v0.3.1 h1:J16l9JAHWgkFPR3VIPbQ1gvS0cWab6laK1q7PFL3qh0=
github.com/JohannesKaufmann/dom v0.3.1/go.mod h1:BZPkf8ZeYrBgABjwJn9iiKt8aiCtkxpHkevms+Yp2DE= github.com/JohannesKaufmann/dom v0.3.1/go.mod h1:BZPkf8ZeYrBgABjwJn9iiKt8aiCtkxpHkevms+Yp2DE=
github.com/JohannesKaufmann/html-to-markdown/v2 v2.5.2 h1:XFJZFWESIWlUEHHjzBuv8RvrtCWnSGlimEX17ysSDb8= github.com/JohannesKaufmann/html-to-markdown/v2 v2.5.2 h1:XFJZFWESIWlUEHHjzBuv8RvrtCWnSGlimEX17ysSDb8=
@@ -10,6 +12,10 @@ github.com/alecthomas/repr v0.4.0 h1:GhI2A8MACjfegCPVq9f1FLvIBS+DrQ2KQBFZP1iFzXc
github.com/alecthomas/repr v0.4.0/go.mod h1:Fr0507jx4eOXV7AlPV6AVZLYrLIuIeSOWtW57eE/O/4= github.com/alecthomas/repr v0.4.0/go.mod h1:Fr0507jx4eOXV7AlPV6AVZLYrLIuIeSOWtW57eE/O/4=
github.com/atotto/clipboard v0.1.4 h1:EH0zSVneZPSuFR11BlR9YppQTVDbh5+16AmcJi4g1z4= github.com/atotto/clipboard v0.1.4 h1:EH0zSVneZPSuFR11BlR9YppQTVDbh5+16AmcJi4g1z4=
github.com/atotto/clipboard v0.1.4/go.mod h1:ZY9tmq7sm5xIbd9bOK4onWV4S6X0u6GY7Vn0Yu86PYI= github.com/atotto/clipboard v0.1.4/go.mod h1:ZY9tmq7sm5xIbd9bOK4onWV4S6X0u6GY7Vn0Yu86PYI=
github.com/aws/aws-sdk-go-v2 v1.41.1 h1:ABlyEARCDLN034NhxlRUSZr4l71mh+T5KAeGh6cerhU=
github.com/aws/aws-sdk-go-v2 v1.41.1/go.mod h1:MayyLB8y+buD9hZqkCW3kX1AKq07Y5pXxtgB+rRFhz0=
github.com/aws/smithy-go v1.24.0 h1:LpilSUItNPFr1eY85RYgTIg5eIEPtvFbskaFcmmIUnk=
github.com/aws/smithy-go v1.24.0/go.mod h1:LEj2LM3rBRQJxPZTB4KuzZkaZYnZPnvgIhb4pu07mx0=
github.com/aymanbagabas/go-osc52/v2 v2.0.1 h1:HwpRHbFMcZLEVr42D4p7XBqjyuxQH5SMiErDT4WkJ2k= github.com/aymanbagabas/go-osc52/v2 v2.0.1 h1:HwpRHbFMcZLEVr42D4p7XBqjyuxQH5SMiErDT4WkJ2k=
github.com/aymanbagabas/go-osc52/v2 v2.0.1/go.mod h1:uYgXzlJ7ZpABp8OJ+exZzJJhRNQ2ASbcXHWsFqH8hp8= github.com/aymanbagabas/go-osc52/v2 v2.0.1/go.mod h1:uYgXzlJ7ZpABp8OJ+exZzJJhRNQ2ASbcXHWsFqH8hp8=
github.com/aymanbagabas/go-udiff v0.2.0 h1:TK0fH4MteXUDspT88n8CKzvK0X9O2xu9yQjWpi6yML8= github.com/aymanbagabas/go-udiff v0.2.0 h1:TK0fH4MteXUDspT88n8CKzvK0X9O2xu9yQjWpi6yML8=
@@ -31,14 +37,22 @@ github.com/charmbracelet/x/exp/golden v0.0.0-20240715153702-9ba8adf781c4/go.mod
github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81 h1:q2hJAaP1k2wIvVRd/hEHD7lacgqrCPS+k8g1MndzfWY= github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81 h1:q2hJAaP1k2wIvVRd/hEHD7lacgqrCPS+k8g1MndzfWY=
github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81/go.mod h1:YynlIjWYF8myEu6sdkwKIvGQq+cOckRm6So2avqoYAk= github.com/containerd/console v1.0.4-0.20230313162750-1ae8d489ac81/go.mod h1:YynlIjWYF8myEu6sdkwKIvGQq+cOckRm6So2avqoYAk=
github.com/cpuguy83/go-md2man/v2 v2.0.6/go.mod h1:oOW0eioCTA6cOiMLiUPZOpcVxMig6NIQQ7OS05n1F4g= github.com/cpuguy83/go-md2man/v2 v2.0.6/go.mod h1:oOW0eioCTA6cOiMLiUPZOpcVxMig6NIQQ7OS05n1F4g=
github.com/creack/pty v1.1.9/go.mod h1:oKZEueFk5CKHvIhNR5MUki03XCEU+Q6VDXinZuGJ33E=
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dlclark/regexp2 v1.11.0 h1:G/nrcoOa7ZXlpoa/91N3X7mM3r8eIlMBBJZvsz/mxKI= github.com/dlclark/regexp2 v1.11.0 h1:G/nrcoOa7ZXlpoa/91N3X7mM3r8eIlMBBJZvsz/mxKI=
github.com/dlclark/regexp2 v1.11.0/go.mod h1:DHkYz0B9wPfa6wondMfaivmHpzrQ3v9q8cnmRbL6yW8= github.com/dlclark/regexp2 v1.11.0/go.mod h1:DHkYz0B9wPfa6wondMfaivmHpzrQ3v9q8cnmRbL6yW8=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY= github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto= github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 h1:B9YK+Tck5mTccyDhtxBzWyqGYcFxLyB6+noMNW4/VgI= github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 h1:B9YK+Tck5mTccyDhtxBzWyqGYcFxLyB6+noMNW4/VgI=
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3/go.mod h1:HMJKR5wlh/ziNp+sHEDV2ltblO4JD2+IdDOWtGcQBTM= github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3/go.mod h1:HMJKR5wlh/ziNp+sHEDV2ltblO4JD2+IdDOWtGcQBTM=
github.com/eslider/go-hocr v0.2.2-0.20260827163626-8ff01582b002 h1:LOFxQG4mxvlH7+lffbO2SIn2ThlClJlrHHfs1OqMNXs=
github.com/eslider/go-hocr v0.2.2-0.20260827163626-8ff01582b002/go.mod h1:fIgfH/E1j3rU8du4X4+7mxTD0GPtPQibTzytgitdJWU=
github.com/eslider/go-xls/v2 v2.1.0 h1:HszWKqYQbXxACmAXXWdMsfNl1NDBfGVBnJUPtyUHQ7A= github.com/eslider/go-xls/v2 v2.1.0 h1:HszWKqYQbXxACmAXXWdMsfNl1NDBfGVBnJUPtyUHQ7A=
github.com/eslider/go-xls/v2 v2.1.0/go.mod h1:xgxO6JrfuBr9jGUB+0z5l/yDmFFZ5diGk0ATGihxlMU= github.com/eslider/go-xls/v2 v2.1.0/go.mod h1:xgxO6JrfuBr9jGUB+0z5l/yDmFFZ5diGk0ATGihxlMU=
github.com/go-sql-driver/mysql v1.10.1 h1:arlSnNLq6a5yxGxV7qg9lF4j0C+KwD6NbQyKr9QL6ME=
github.com/go-sql-driver/mysql v1.10.1/go.mod h1:M+cqaI7+xxXGG9swrdeUIoPG3Y3KCkF0pZej+SK+nWk=
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI= github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI=
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY= github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY=
github.com/google/go-querystring v1.2.0 h1:yhqkPbu2/OH+V9BfpCVPZkNmUXhb2gBxJArfhIxNtP0= github.com/google/go-querystring v1.2.0 h1:yhqkPbu2/OH+V9BfpCVPZkNmUXhb2gBxJArfhIxNtP0=
@@ -55,8 +69,20 @@ github.com/hexops/gotextdiff v1.0.3 h1:gitA9+qJrrTCsiCl7+kh75nPqQt1cx4ZkudSTLoUq
github.com/hexops/gotextdiff v1.0.3/go.mod h1:pSWU5MAI3yDq+fZBTazCSJysOMbxWL1BSow5/V2vxeg= github.com/hexops/gotextdiff v1.0.3/go.mod h1:pSWU5MAI3yDq+fZBTazCSJysOMbxWL1BSow5/V2vxeg=
github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8= github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8=
github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw= github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw=
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
github.com/jackc/pgpassfile v1.0.0/go.mod h1:CEx0iS5ambNFdcRtxPj5JhEz+xB6uRky5eyVu/W2HEg=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 h1:iCEnooe7UlwOQYpKFhBabPMi4aNAfoODPEFNiAnClxo=
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761/go.mod h1:5TJZWKEWniPve33vlWYSoGYefn3gLQRzjfDlhSJ9ZKM=
github.com/jackc/pgx/v5 v5.11.0 h1:IzBBtyK9AHqf98cctWFifYSci2hgQR/cd56wB4p+ogg=
github.com/jackc/pgx/v5 v5.11.0/go.mod h1:mal1tBGAFfLHvZzaYh77YS/eC6IX9OWbRV1QIIM0Jn4=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/joho/godotenv v1.5.1 h1:7eLL/+HRGLY0ldzfGMeQkb7vMd0as4CfYvUVzLqw0N0= github.com/joho/godotenv v1.5.1 h1:7eLL/+HRGLY0ldzfGMeQkb7vMd0as4CfYvUVzLqw0N0=
github.com/joho/godotenv v1.5.1/go.mod h1:f4LDr5Voq0i2e/R5DDNOoa2zzDfwtkZa6DnEwAbqwq4= github.com/joho/godotenv v1.5.1/go.mod h1:f4LDr5Voq0i2e/R5DDNOoa2zzDfwtkZa6DnEwAbqwq4=
github.com/kr/pretty v0.3.0 h1:WgNl7dwNpEZ6jJ9k1snq4pZsg7DOEN8hP9Xw0Tsjwk0=
github.com/kr/pretty v0.3.0/go.mod h1:640gp4NfQd8pI5XOwp5fnNeVWj67G7CFk/SaSQn7NBk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/lucasb-eyer/go-colorful v1.4.0 h1:UtrWVfLdarDgc44HcS7pYloGHJUjHV/4FwW4TvVgFr4= github.com/lucasb-eyer/go-colorful v1.4.0 h1:UtrWVfLdarDgc44HcS7pYloGHJUjHV/4FwW4TvVgFr4=
github.com/lucasb-eyer/go-colorful v1.4.0/go.mod h1:R4dSotOR9KMtayYi1e77YzuveK+i7ruzyGqttikkLy0= github.com/lucasb-eyer/go-colorful v1.4.0/go.mod h1:R4dSotOR9KMtayYi1e77YzuveK+i7ruzyGqttikkLy0=
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI= github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
@@ -82,10 +108,16 @@ github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZb
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4= github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE= github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo= github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo=
github.com/richardlehane/mscfb v1.0.7 h1:oeoiM0WE79vHwE8RpIYYvIAc8ajTH2mb6UZm55/+EB0=
github.com/richardlehane/mscfb v1.0.7/go.mod h1:pe0+IUIc0AHh0+teNzBlJCtSyZdFOGgV4ZK9bsoV+Jo=
github.com/richardlehane/msoleps v1.0.6 h1:9BvkpjvD+iUBalUY4esMwv6uBkfOip/Lzvd93jvR9gg=
github.com/richardlehane/msoleps v1.0.6/go.mod h1:BWev5JBpU9Ko2WAgmZEuiz4/u3ZYTKbjLycmwiWUfWg=
github.com/rivo/uniseg v0.1.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc= github.com/rivo/uniseg v0.1.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc=
github.com/rivo/uniseg v0.2.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc= github.com/rivo/uniseg v0.2.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc=
github.com/rivo/uniseg v0.4.7 h1:WUdvkW8uEhrYfLC4ZzdpI2ztxP1I582+49Oc5Mq64VQ= github.com/rivo/uniseg v0.4.7 h1:WUdvkW8uEhrYfLC4ZzdpI2ztxP1I582+49Oc5Mq64VQ=
github.com/rivo/uniseg v0.4.7/go.mod h1:FN3SvrM+Zdj16jyLfmOkMNblXMcoc8DfTHruCPUcx88= github.com/rivo/uniseg v0.4.7/go.mod h1:FN3SvrM+Zdj16jyLfmOkMNblXMcoc8DfTHruCPUcx88=
github.com/rogpeppe/go-internal v1.16.0 h1:O9DK+vNMDVGLr2BeZqmpLeMjiMNkuXfcqntWbZV6S5g=
github.com/rogpeppe/go-internal v1.16.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs=
github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM= github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM=
github.com/sebdah/goldie/v2 v2.8.0 h1:dZb9wR8q5++oplmEiJT+U/5KyotVD+HNGCAc5gNr8rc= github.com/sebdah/goldie/v2 v2.8.0 h1:dZb9wR8q5++oplmEiJT+U/5KyotVD+HNGCAc5gNr8rc=
github.com/sebdah/goldie/v2 v2.8.0/go.mod h1:oZ9fp0+se1eapSRjfYbsV/0Hqhbuu3bJVvKI/NNtssI= github.com/sebdah/goldie/v2 v2.8.0/go.mod h1:oZ9fp0+se1eapSRjfYbsV/0Hqhbuu3bJVvKI/NNtssI=
@@ -95,29 +127,48 @@ github.com/spf13/cobra v1.10.2 h1:DMTTonx5m65Ic0GOoRY2c16WCbHxOOw6xxezuLaBpcU=
github.com/spf13/cobra v1.10.2/go.mod h1:7C1pvHqHw5A4vrJfjNwvOdzYu0Gml16OCs2GRiTUUS4= github.com/spf13/cobra v1.10.2/go.mod h1:7C1pvHqHw5A4vrJfjNwvOdzYu0Gml16OCs2GRiTUUS4=
github.com/spf13/pflag v1.0.9 h1:9exaQaMOCwffKiiiYk6/BndUBv+iRViNW+4lEMi0PvY= github.com/spf13/pflag v1.0.9 h1:9exaQaMOCwffKiiiYk6/BndUBv+iRViNW+4lEMi0PvY=
github.com/spf13/pflag v1.0.9/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg= github.com/spf13/pflag v1.0.9/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI=
github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
github.com/tiendc/go-deepcopy v1.7.2 h1:Ut2yYR7W9tWjTQitganoIue4UGxZwCcJy3orjrrIj44=
github.com/tiendc/go-deepcopy v1.7.2/go.mod h1:4bKjNC2r7boYOkD2IOuZpYjmlDdzjbpTRyCx+goBCJQ=
github.com/xuri/efp v0.0.1 h1:fws5Rv3myXyYni8uwj2qKjVaRP30PdjeYe2Y6FDsCL8=
github.com/xuri/efp v0.0.1/go.mod h1:ybY/Jr0T0GTCnYjKqmdwxyxn2BQf2RcQIIvex5QldPI=
github.com/xuri/excelize/v2 v2.11.0 h1:HxaEFl6sRN2+8J5a8HaKq+0M4FsjBGMnWWtjOCPSG88=
github.com/xuri/excelize/v2 v2.11.0/go.mod h1:jxFLbzaIwGQ5ufFNvYfUOHqXhfPaNmP14KWfmNz2Uak=
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9 h1:+C0TIdyyYmzadGaL/HBLbf3WdLgC29pgyhTjAT/0nuE=
github.com/xuri/nfp v0.0.2-0.20250530014748-2ddeb826f9a9/go.mod h1:WwHg+CVyzlv/TX9xqBFXEZAuxOPxn2k1GNHwG41IIUQ=
github.com/yuin/goldmark v1.7.1/go.mod h1:uzxRWxtg69N339t3louHJ7+O03ezfj6PlliRlaOzY1E= github.com/yuin/goldmark v1.7.1/go.mod h1:uzxRWxtg69N339t3louHJ7+O03ezfj6PlliRlaOzY1E=
github.com/yuin/goldmark v1.8.2 h1:kEGpgqJXdgbkhcOgBxkC0X0PmoPG1ZyoZ117rDVp4zE= github.com/yuin/goldmark v1.8.2 h1:kEGpgqJXdgbkhcOgBxkC0X0PmoPG1ZyoZ117rDVp4zE=
github.com/yuin/goldmark v1.8.2/go.mod h1:ip/1k0VRfGynBgxOz0yCqHrbZXhcjxyuS66Brc7iBKg= github.com/yuin/goldmark v1.8.2/go.mod h1:ip/1k0VRfGynBgxOz0yCqHrbZXhcjxyuS66Brc7iBKg=
github.com/yuin/goldmark-emoji v1.0.3 h1:aLRkLHOuBR2czCY4R8olwMjID+tENfhyFDMCRhbIQY4= github.com/yuin/goldmark-emoji v1.0.3 h1:aLRkLHOuBR2czCY4R8olwMjID+tENfhyFDMCRhbIQY4=
github.com/yuin/goldmark-emoji v1.0.3/go.mod h1:tTkZEbwu5wkPmgTcitqddVxY9osFZiavD+r4AzQrh1U= github.com/yuin/goldmark-emoji v1.0.3/go.mod h1:tTkZEbwu5wkPmgTcitqddVxY9osFZiavD+r4AzQrh1U=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg= go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
golang.org/x/crypto v0.53.0 h1:QZ4Muo8THX6CizN2vPPd5fBGHyogrdK9fG4wLPFUsto=
golang.org/x/crypto v0.53.0/go.mod h1:DNLU434OwVakk9PzuwV8w62mAJpRJL3vsgcfp4Qnsio=
golang.org/x/image v0.38.0 h1:5l+q+Y9JDC7mBOMjo4/aPhMDcxEptsX+Tt3GgRQRPuE=
golang.org/x/image v0.38.0/go.mod h1:/3f6vaXC+6CEanU4KJxbcUZyEePbyKbaLoDOe4ehFYY=
golang.org/x/mod v0.37.0 h1:vF1DjpVEshcIqoEaauuHebaLk1O1forxjxBaVn884JQ= golang.org/x/mod v0.37.0 h1:vF1DjpVEshcIqoEaauuHebaLk1O1forxjxBaVn884JQ=
golang.org/x/mod v0.37.0/go.mod h1:m8S8VeM9r4dzDwjrKO0a1sZP3YjeMamRRlD+fmR2Q/0= golang.org/x/mod v0.37.0/go.mod h1:m8S8VeM9r4dzDwjrKO0a1sZP3YjeMamRRlD+fmR2Q/0=
golang.org/x/net v0.55.0 h1:bcvxaJn3e1U6InsFWt1JUq1aSjnRxLzT2rtD2KfkDF8= golang.org/x/net v0.56.0 h1:Rw8j/hFzGvJUZwNBXnAtf5sVDVt+65SK2C7IxCxZt5o=
golang.org/x/net v0.55.0/go.mod h1:L5U2KuzuOe1lY7Z+aWVIKK6qEeJXnXV9yzGA+WCHJww= golang.org/x/net v0.56.0/go.mod h1:D3Ku6r+V6JROoZK144D2XfMHFcMq/0zSfLelVTCFKec=
golang.org/x/sync v0.21.0 h1:HLII4xRRTtCRkxYp4HNFF0Js/Og6q2i++KXbg0gHCwM= golang.org/x/sync v0.21.0 h1:HLII4xRRTtCRkxYp4HNFF0Js/Og6q2i++KXbg0gHCwM=
golang.org/x/sync v0.21.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0= golang.org/x/sync v0.21.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.1.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg= golang.org/x/sys v0.1.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs= golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw= golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/term v0.43.0 h1:S4RLU2sB31O/NCl+zFN9Aru9A/Cq2aqKpTZJ6B+DwT4= golang.org/x/term v0.44.0 h1:0rLvDRCtNj0gZkyIXhCyOb2OAzEhLVqc4B+hrsBhrmc=
golang.org/x/term v0.43.0/go.mod h1:lrhlHNdQJHO+1qVYiHfFKVuVioJIheAc3fBSMFYEIsk= golang.org/x/term v0.44.0/go.mod h1:7ze4MdzUzLXpSAoFP1H0bOI9aXDqveSvatT5vKcFh2Y=
golang.org/x/text v0.37.0 h1:Cqjiwd9eSg8e0QAkyCaQTNHFIIzWtidPahFWR83rTrc= golang.org/x/text v0.38.0 h1:sXmwo9DwP3OK9EZ7PqAdaooSGozfl/3a6/xJcbzPRhE=
golang.org/x/text v0.37.0/go.mod h1:a5sjxXGs9hsn/AJVwuElvCAo9v8QYLzvavO5z2PiM38= golang.org/x/text v0.38.0/go.mod h1:YXZt3QhHUKYT53r2lLKFIVi6Ao1jdzrTR/KQ09qyxF4=
golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q= golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q=
golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA= golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405 h1:yhCVgyC4o1eVCa2tZl7eS0r+SDo693bJlVdllGtEeKM=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk=
gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q=
gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA= gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI= modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI=
+112 -31
View File
@@ -111,6 +111,25 @@ func (c *Client) deleteObject(ctx context.Context, path string) (map[string]any,
return unmarshalResponseObject(raw) return unmarshalResponseObject(raw)
} }
// unmarshalResponseArray extracts the "response" field from a raw OnlyOffice
// envelope and decodes it into a list of maps. Returns (nil, nil) for a null,
// empty or scalar payload. Companion to unmarshalResponseObject for endpoints
// whose response is a list (project team, people/status, …).
func unmarshalResponseArray(raw json.RawMessage) ([]map[string]any, error) {
resp, err := responseField(raw, "response")
if err != nil {
return nil, err
}
if len(resp) == 0 || string(resp) == "null" || resp[0] != '[' {
return nil, nil
}
var list []map[string]any
if err := json.Unmarshal(resp, &list); err != nil {
return nil, err
}
return list, nil
}
// unmarshalResponseObject extracts the "response" field from a raw OnlyOffice // unmarshalResponseObject extracts the "response" field from a raw OnlyOffice
// envelope and decodes it into map[string]any. Returns (nil, nil) for a null // envelope and decodes it into map[string]any. Returns (nil, nil) for a null
// response, an empty array, or scalar payloads. When the API returns a list // response, an empty array, or scalar payloads. When the API returns a list
@@ -144,8 +163,13 @@ func unmarshalResponseObject(raw json.RawMessage) (map[string]any, error) {
} }
} }
// getJSON issues an authenticated GET and returns the raw response body. // getJSON issues an authenticated GET and returns the raw response body,
// retrying transient answers (see retryRaw).
func (c *Client) getJSON(ctx context.Context, path string) (json.RawMessage, error) { func (c *Client) getJSON(ctx context.Context, path string) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) { return c.getJSONOnce(ctx, path) })
}
func (c *Client) getJSONOnce(ctx context.Context, path string) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
@@ -166,7 +190,7 @@ func (c *Client) getJSON(ctx context.Context, path string) (json.RawMessage, err
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("GET %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "GET %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
@@ -187,6 +211,12 @@ func (c *Client) deleteForm(ctx context.Context, path string, fields url.Values)
} }
func (c *Client) formRequest(ctx context.Context, method, path string, fields url.Values) (json.RawMessage, error) { func (c *Client) formRequest(ctx context.Context, method, path string, fields url.Values) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) {
return c.formRequestOnce(ctx, method, path, fields)
})
}
func (c *Client) formRequestOnce(ctx context.Context, method, path string, fields url.Values) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
@@ -208,13 +238,17 @@ func (c *Client) formRequest(ctx context.Context, method, path string, fields ur
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("%s form %s: %d %s", method, path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "%s form %s: %d %s", method, path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
// deleteReq issues an authenticated DELETE. // deleteReq issues an authenticated DELETE.
func (c *Client) deleteReq(ctx context.Context, path string) (json.RawMessage, error) { func (c *Client) deleteReq(ctx context.Context, path string) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) { return c.deleteReqOnce(ctx, path) })
}
func (c *Client) deleteReqOnce(ctx context.Context, path string) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
@@ -235,7 +269,7 @@ func (c *Client) deleteReq(ctx context.Context, path string) (json.RawMessage, e
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("DELETE %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "DELETE %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
@@ -258,26 +292,65 @@ func (c *Client) postJSONObject(ctx context.Context, path string, body any) (map
return unmarshalResponseObject(raw) return unmarshalResponseObject(raw)
} }
// postJSON issues an authenticated POST with application/json body. // jsonBodyReader turns a request body value into an io.Reader. nil becomes
func (c *Client) postJSON(ctx context.Context, path string, body any) (json.RawMessage, error) { // "{}", []byte/string pass through, anything else is JSON-marshalled.
auth, err := c.authHeader() func jsonBodyReader(body any) (io.Reader, error) {
if err != nil {
return nil, err
}
var rdr io.Reader
switch b := body.(type) { switch b := body.(type) {
case nil: case nil:
rdr = strings.NewReader("{}") return strings.NewReader("{}"), nil
case []byte: case []byte:
rdr = bytes.NewReader(b) return bytes.NewReader(b), nil
case string: case string:
rdr = strings.NewReader(b) return strings.NewReader(b), nil
default: default:
buf, err := json.Marshal(b) buf, err := json.Marshal(b)
if err != nil { if err != nil {
return nil, err return nil, err
} }
rdr = bytes.NewReader(buf) return bytes.NewReader(buf), nil
}
}
// postJSONArray is postJSON + unmarshalResponseArray.
func (c *Client) postJSONArray(ctx context.Context, path string, body any) ([]map[string]any, error) {
raw, err := c.postJSON(ctx, path, body)
if err != nil {
return nil, err
}
return unmarshalResponseArray(raw)
}
// putJSONArray is putJSON + unmarshalResponseArray.
func (c *Client) putJSONArray(ctx context.Context, path string, body any) ([]map[string]any, error) {
raw, err := c.putJSON(ctx, path, body)
if err != nil {
return nil, err
}
return unmarshalResponseArray(raw)
}
// deleteJSONArray is deleteJSON + unmarshalResponseArray.
func (c *Client) deleteJSONArray(ctx context.Context, path string, body any) ([]map[string]any, error) {
raw, err := c.deleteJSON(ctx, path, body)
if err != nil {
return nil, err
}
return unmarshalResponseArray(raw)
}
// postJSON issues an authenticated POST with application/json body.
func (c *Client) postJSON(ctx context.Context, path string, body any) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) { return c.postJSONOnce(ctx, path, body) })
}
func (c *Client) postJSONOnce(ctx context.Context, path string, body any) (json.RawMessage, error) {
auth, err := c.authHeader()
if err != nil {
return nil, err
}
rdr, err := jsonBodyReader(body)
if err != nil {
return nil, err
} }
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL()+path, rdr) req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL()+path, rdr)
if err != nil { if err != nil {
@@ -296,32 +369,25 @@ func (c *Client) postJSON(ctx context.Context, path string, body any) (json.RawM
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("POST JSON %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "POST JSON %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
// putJSON issues an authenticated PUT with application/json body. // putJSON issues an authenticated PUT with application/json body.
func (c *Client) putJSON(ctx context.Context, path string, body any) (json.RawMessage, error) { func (c *Client) putJSON(ctx context.Context, path string, body any) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) { return c.putJSONOnce(ctx, path, body) })
}
func (c *Client) putJSONOnce(ctx context.Context, path string, body any) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
} }
var rdr io.Reader rdr, err := jsonBodyReader(body)
switch b := body.(type) {
case nil:
rdr = strings.NewReader("{}")
case []byte:
rdr = bytes.NewReader(b)
case string:
rdr = strings.NewReader(b)
default:
buf, err := json.Marshal(b)
if err != nil { if err != nil {
return nil, err return nil, err
} }
rdr = bytes.NewReader(buf)
}
req, err := http.NewRequestWithContext(ctx, http.MethodPut, c.baseURL()+path, rdr) req, err := http.NewRequestWithContext(ctx, http.MethodPut, c.baseURL()+path, rdr)
if err != nil { if err != nil {
return nil, err return nil, err
@@ -339,13 +405,28 @@ func (c *Client) putJSON(ctx context.Context, path string, body any) (json.RawMe
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("PUT JSON %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "PUT JSON %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
// uploadMultipart posts a single file to path under the given form field name. // uploadMultipart posts a single file to path under the given form field name.
func (c *Client) uploadMultipart(ctx context.Context, path, fieldName, filePath string) (json.RawMessage, error) { func (c *Client) uploadMultipart(ctx context.Context, path, fieldName, filePath string) (json.RawMessage, error) {
return c.uploadMultipartMethod(ctx, http.MethodPost, path, fieldName, filePath)
}
// uploadMultipartMethod sends a single-file multipart request with the given
// HTTP method. The OnlyOffice Documents API needs PUT for /update (a new
// version) and POST for /upload (a new file); sending POST to /update answers
// 500 on current servers. The file is re-opened per attempt, so transient
// answers are retried like every other request.
func (c *Client) uploadMultipartMethod(ctx context.Context, method, path, fieldName, filePath string) (json.RawMessage, error) {
return retryRaw(ctx, func() (json.RawMessage, error) {
return c.uploadMultipartOnce(ctx, method, path, fieldName, filePath)
})
}
func (c *Client) uploadMultipartOnce(ctx context.Context, method, path, fieldName, filePath string) (json.RawMessage, error) {
auth, err := c.authHeader() auth, err := c.authHeader()
if err != nil { if err != nil {
return nil, err return nil, err
@@ -368,7 +449,7 @@ func (c *Client) uploadMultipart(ctx context.Context, path, fieldName, filePath
if err := mw.Close(); err != nil { if err := mw.Close(); err != nil {
return nil, err return nil, err
} }
req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL()+path, &buf) req, err := http.NewRequestWithContext(ctx, method, c.baseURL()+path, &buf)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -385,7 +466,7 @@ func (c *Client) uploadMultipart(ctx context.Context, path, fieldName, filePath
return nil, err return nil, err
} }
if resp.StatusCode >= 400 { if resp.StatusCode >= 400 {
return nil, fmt.Errorf("upload %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400)) return nil, statusError(resp.StatusCode, retryAfterOf(resp), "upload %s: %d %s", path, resp.StatusCode, truncate(string(raw), 400))
} }
return raw, nil return raw, nil
} }
+23
View File
@@ -48,3 +48,26 @@ func TestUnmarshalResponseObjectNull(t *testing.T) {
t.Fatalf("expected nil, got %#v", out) t.Fatalf("expected nil, got %#v", out)
} }
} }
func TestUnmarshalResponseArrayList(t *testing.T) {
raw := json.RawMessage(`{"response":[{"id":"a","displayName":"A"},{"id":"b","displayName":"B"}]}`)
out, err := unmarshalResponseArray(raw)
if err != nil {
t.Fatal(err)
}
if len(out) != 2 || out[0]["id"] != "a" || out[1]["displayName"] != "B" {
t.Fatalf("unexpected list: %#v", out)
}
}
func TestUnmarshalResponseArrayNullAndScalar(t *testing.T) {
for _, raw := range []string{`{"response":null}`, `{"response":{}}`, `{"response":"x"}`} {
out, err := unmarshalResponseArray(json.RawMessage(raw))
if err != nil {
t.Fatalf("%s: %v", raw, err)
}
if out != nil {
t.Fatalf("%s: expected nil, got %#v", raw, out)
}
}
}
+422
View File
@@ -0,0 +1,422 @@
// Package docpipe converts documents for OnlyOffice agent workflows:
// Markdown ↔ DOCX (pandoc) and image/PDF OCR → searchable PDF + Markdown text.
//
// External tools (optional at runtime; helpers skip/error clearly when missing):
// - pandoc — md↔docx
// - ocrmypdf — OCR into a searchable PDF
// - pdftotext — extract text layer
// - pdfdetach — list/save embedded PDF attachments
// - tesseract — OCR single images when ocrmypdf is unsuitable
// - ghostscript (gs) — PDF rewrite/optimize via PostScript (pdfwrite)
package docpipe
import (
"bytes"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
)
// DefaultMinTextChars: below this, a PDF is treated as needing OCR.
const DefaultMinTextChars = 200
// Tools reports which converters are available on PATH.
type Tools struct {
Pandoc string
OCRMyPDF string
PDFToText string
PDFDetach string
Tesseract string
Ghostscript string
}
// LookPath resolves converter binaries (empty string if missing).
func LookPath() Tools {
find := func(names ...string) string {
for _, n := range names {
if p, err := exec.LookPath(n); err == nil {
return p
}
}
return ""
}
return Tools{
Pandoc: find("pandoc"),
OCRMyPDF: find("ocrmypdf"),
PDFToText: find("pdftotext"),
PDFDetach: find("pdfdetach"),
Tesseract: find("tesseract"),
Ghostscript: find("gs", "ghostscript"),
}
}
func (t Tools) requirePandoc() error {
if t.Pandoc == "" {
return fmt.Errorf("pandoc not found on PATH (needed for md↔docx)")
}
return nil
}
// Ext returns lower-case extension including dot (".pdf").
func Ext(path string) string {
return strings.ToLower(filepath.Ext(path))
}
// ConvertFile converts between md and docx (and other pandoc formats) via pandoc.
// outExt may be ".md", ".docx", or a full output path.
func (t Tools) ConvertFile(inPath, outPath string) error {
if err := t.requirePandoc(); err != nil {
return err
}
if strings.TrimSpace(outPath) == "" {
return fmt.Errorf("output path required")
}
cmd := exec.Command(t.Pandoc, inPath, "-o", outPath)
var stderr bytes.Buffer
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return fmt.Errorf("pandoc %s → %s: %w (%s)", inPath, outPath, err, strings.TrimSpace(stderr.String()))
}
return nil
}
// TXTToDOCX converts plain text to DOCX preserving line breaks (via markdown hard breaks).
func (t Tools) TXTToDOCX(txtPath, docxPath string) error {
if Ext(txtPath) != ".txt" {
return fmt.Errorf("expected .txt input, got %q", txtPath)
}
if docxPath == "" {
docxPath = strings.TrimSuffix(txtPath, Ext(txtPath)) + ".docx"
}
b, err := os.ReadFile(txtPath)
if err != nil {
return err
}
dir := filepath.Dir(docxPath)
if dir == "" || dir == "." {
dir = os.TempDir()
}
tmpMD := filepath.Join(dir, trimExt(filepath.Base(txtPath))+".txt2docx.md")
md := TxtToMarkdown(string(b))
if err := os.WriteFile(tmpMD, []byte(md), 0o644); err != nil {
return err
}
defer os.Remove(tmpMD)
return t.MDToDOCX(tmpMD, docxPath)
}
// TxtToMarkdown converts plain text to Markdown for DOCX output.
// Prose text: each line is a hard break. Fixed-width extracts (INE, pdftotext -layout):
// wrapped in a fenced code block (monospace, columns preserved).
func TxtToMarkdown(content string) string {
content = normalizeTxtNewlines(content)
if isFixedWidthTxt(content) {
return "```\n" + content + "\n```\n"
}
var b strings.Builder
for _, line := range strings.Split(content, "\n") {
if strings.TrimSpace(line) == "" {
b.WriteByte('\n')
continue
}
b.WriteString(line)
b.WriteString(" \n")
}
return b.String()
}
func normalizeTxtNewlines(content string) string {
content = strings.ReplaceAll(content, "\r\n", "\n")
return strings.ReplaceAll(content, "\r", "\n")
}
// isFixedWidthTxt detects pdftotext -layout style extracts (many indented/spaced columns).
func isFixedWidthTxt(content string) bool {
lines := strings.Split(content, "\n")
if len(lines) < 8 {
return false
}
indented, long := 0, 0
for _, line := range lines {
if strings.TrimSpace(line) == "" {
continue
}
if len(line) >= 72 {
long++
}
if len(line) > 0 && (line[0] == ' ' || line[0] == '\t') {
indented++
}
}
n := len(lines)
return indented*100/n >= 20 || (long >= 5 && indented*100/n >= 10)
}
// MDToDOCX writes a DOCX next to or at outPath from a Markdown file.
func (t Tools) MDToDOCX(mdPath, docxPath string) error {
if Ext(mdPath) != ".md" && Ext(mdPath) != ".markdown" {
return fmt.Errorf("expected markdown input, got %q", mdPath)
}
if docxPath == "" {
docxPath = strings.TrimSuffix(mdPath, Ext(mdPath)) + ".docx"
}
return t.ConvertFile(mdPath, docxPath)
}
// DOCXToMD writes Markdown from a DOCX file.
func (t Tools) DOCXToMD(docxPath, mdPath string) error {
if Ext(docxPath) != ".docx" {
return fmt.Errorf("expected .docx input, got %q", docxPath)
}
if mdPath == "" {
mdPath = strings.TrimSuffix(docxPath, Ext(docxPath)) + ".md"
}
return t.ConvertFile(docxPath, mdPath)
}
// OptimizePDF rewrites a PDF through Ghostscript (PostScript pdfwrite).
// Preserves native text layers; strips broken OCR overlays; shrinks for OO preview.
// Use instead of ocrmypdf when pdftotext already extracts enough text.
func (t Tools) OptimizePDF(inPath, outPath string) error {
if t.Ghostscript == "" {
return fmt.Errorf("ghostscript (gs) not found on PATH")
}
if outPath == "" {
return fmt.Errorf("output PDF path required")
}
args := []string{
"-sDEVICE=pdfwrite",
"-dCompatibilityLevel=1.5",
"-dNOPAUSE", "-dQUIET", "-dBATCH",
"-dPDFSETTINGS=/ebook",
"-dEmbedAllFonts=true",
"-dSubsetFonts=true",
"-dCompressFonts=true",
"-dCompressPages=true",
"-dDetectDuplicateImages=true",
"-dAutoRotatePages=/None",
"-sOutputFile=" + outPath,
inPath,
}
cmd := exec.Command(t.Ghostscript, args...)
var stderr bytes.Buffer
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return fmt.Errorf("ghostscript pdfwrite: %w (%s)", err, strings.TrimSpace(stderr.String()))
}
return nil
}
// PDFTextLayerChars returns approximate extracted character count (0 if unavailable).
func (t Tools) PDFTextLayerChars(pdfPath string) (int, error) {
if t.PDFToText == "" {
return 0, fmt.Errorf("pdftotext not found on PATH")
}
cmd := exec.Command(t.PDFToText, "-layout", pdfPath, "-")
out, err := cmd.Output()
if err != nil {
return 0, err
}
return len(bytes.TrimSpace(out)), nil
}
// NeedsOCR reports whether path likely needs OCR before text extraction.
func (t Tools) NeedsOCR(path string, minChars int) (bool, error) {
if minChars <= 0 {
minChars = DefaultMinTextChars
}
switch Ext(path) {
case ".jpg", ".jpeg", ".png", ".tif", ".tiff", ".webp", ".gif", ".bmp":
return true, nil
case ".pdf":
n, err := t.PDFTextLayerChars(path)
if err != nil {
// If we cannot measure, prefer OCR.
return true, nil
}
return n < minChars, nil
default:
return false, nil
}
}
// OCRToPDF runs ocrmypdf into outPDF (searchable). Forces OCR when force is true.
func (t Tools) OCRToPDF(inPath, outPDF string, force bool, lang string) error {
if t.OCRMyPDF == "" {
return fmt.Errorf("ocrmypdf not found on PATH")
}
if outPDF == "" {
return fmt.Errorf("output PDF path required")
}
if lang == "" {
lang = "eng"
}
args := []string{"-l", lang, "--skip-big", "100"}
if force {
args = append(args, "--force-ocr")
} else {
args = append(args, "--skip-text")
}
args = append(args, inPath, outPDF)
cmd := exec.Command(t.OCRMyPDF, args...)
var stderr bytes.Buffer
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
// Retry with force if skip-text refused.
if !force && strings.Contains(stderr.String(), "PriorOcrFoundError") == false {
args2 := []string{"-l", lang, "--force-ocr", inPath, outPDF}
cmd2 := exec.Command(t.OCRMyPDF, args2...)
var stderr2 bytes.Buffer
cmd2.Stderr = &stderr2
if err2 := cmd2.Run(); err2 == nil {
return nil
}
}
return fmt.Errorf("ocrmypdf: %w (%s)", err, strings.TrimSpace(stderr.String()))
}
return nil
}
// ImageToText OCRs a raster image with tesseract (stdout text).
func (t Tools) ImageToText(imgPath, lang string) (string, error) {
if t.Tesseract == "" {
return "", fmt.Errorf("tesseract not found on PATH")
}
if lang == "" {
lang = "eng"
}
cmd := exec.Command(t.Tesseract, imgPath, "stdout", "-l", lang)
var stderr bytes.Buffer
cmd.Stderr = &stderr
out, err := cmd.Output()
if err != nil {
return "", fmt.Errorf("tesseract: %w (%s)", err, strings.TrimSpace(stderr.String()))
}
return string(out), nil
}
// ExtractPDFText returns layout text from a PDF via pdftotext.
func (t Tools) ExtractPDFText(pdfPath string) (string, error) {
if t.PDFToText == "" {
return "", fmt.Errorf("pdftotext not found on PATH")
}
cmd := exec.Command(t.PDFToText, "-layout", pdfPath, "-")
out, err := cmd.Output()
if err != nil {
return "", err
}
return string(out), nil
}
// Result of ToMarkdown.
type Result struct {
Markdown string
OCRPDFPath string // set when a searchable PDF was produced
DidOCR bool
Source string
}
// ToMarkdown turns a local file into Markdown text.
// PDFs/images with weak/no text layer are OCR'd to a searchable PDF first (when tools exist).
func (t Tools) ToMarkdown(path string, workDir string, lang string, minChars int) (Result, error) {
res := Result{Source: path}
ext := Ext(path)
switch ext {
case ".md", ".markdown", ".txt":
b, err := os.ReadFile(path)
if err != nil {
return res, err
}
res.Markdown = string(b)
return res, nil
case ".docx", ".odt", ".rtf", ".html", ".htm":
if err := t.requirePandoc(); err != nil {
return res, err
}
tmp := filepath.Join(workDir, "out.md")
if err := t.ConvertFile(path, tmp); err != nil {
return res, err
}
b, err := os.ReadFile(tmp)
if err != nil {
return res, err
}
res.Markdown = string(b)
return res, nil
case ".pdf":
need, _ := t.NeedsOCR(path, minChars)
pdf := path
if need {
if workDir == "" {
workDir = os.TempDir()
}
outPDF := filepath.Join(workDir, trimExt(filepath.Base(path))+".ocr.pdf")
if err := t.OCRToPDF(path, outPDF, true, lang); err != nil {
return res, err
}
res.DidOCR = true
res.OCRPDFPath = outPDF
pdf = outPDF
}
text, err := t.ExtractPDFText(pdf)
if err != nil {
return res, err
}
res.Markdown = wrapMD(filepath.Base(path), text)
return res, nil
case ".jpg", ".jpeg", ".png", ".tif", ".tiff", ".webp", ".gif", ".bmp":
if workDir == "" {
workDir = os.TempDir()
}
outPDF := filepath.Join(workDir, trimExt(filepath.Base(path))+".ocr.pdf")
if t.OCRMyPDF != "" {
if err := t.OCRToPDF(path, outPDF, true, lang); err == nil {
res.DidOCR = true
res.OCRPDFPath = outPDF
text, err := t.ExtractPDFText(outPDF)
if err != nil {
return res, err
}
res.Markdown = wrapMD(filepath.Base(path), text)
return res, nil
}
}
text, err := t.ImageToText(path, lang)
if err != nil {
return res, err
}
res.DidOCR = true
res.Markdown = wrapMD(filepath.Base(path), text)
return res, nil
default:
return res, fmt.Errorf("unsupported type %q for markdown extraction", ext)
}
}
func wrapMD(title, body string) string {
body = strings.TrimSpace(body)
if body == "" {
return "# " + title + "\n\n_(empty text layer)_\n"
}
return "# " + title + "\n\n" + body + "\n"
}
func trimExt(name string) string {
return strings.TrimSuffix(name, filepath.Ext(name))
}
// SiblingDOCX returns path with .docx extension replacing the original ext.
func SiblingDOCX(mdPath string) string {
return strings.TrimSuffix(mdPath, Ext(mdPath)) + ".docx"
}
// EnsureDir creates parent directories for path.
func EnsureDir(path string) error {
dir := filepath.Dir(path)
if dir == "" || dir == "." {
return nil
}
return os.MkdirAll(dir, 0o755)
}
+167
View File
@@ -0,0 +1,167 @@
package docpipe
import (
"os"
"path/filepath"
"strings"
"testing"
)
func TestExt(t *testing.T) {
if Ext("Foo.PDF") != ".pdf" {
t.Fatalf("Ext: %q", Ext("Foo.PDF"))
}
}
func TestSiblingDOCX(t *testing.T) {
if got := SiblingDOCX("notes.md"); got != "notes.docx" {
t.Fatalf("got %q", got)
}
}
func TestWrapMD(t *testing.T) {
s := wrapMD("a.pdf", " hello ")
if !strings.HasPrefix(s, "# a.pdf\n") || !strings.Contains(s, "hello") {
t.Fatalf("wrap: %q", s)
}
}
func TestNeedsOCR_Image(t *testing.T) {
tools := LookPath()
need, err := tools.NeedsOCR("x.jpg", 0)
if err != nil || !need {
t.Fatalf("jpg should need OCR: need=%v err=%v", need, err)
}
}
func TestTxtToMarkdown_FixedWidthUsesCodeBlock(t *testing.T) {
var lines []string
for i := 0; i < 12; i++ {
lines = append(lines, " column layout line "+strings.Repeat("x", 40))
}
in := strings.Join(lines, "\n")
md := TxtToMarkdown(in)
if !strings.HasPrefix(md, "```\n") || !strings.Contains(md, "```") {
t.Fatalf("expected code block: %q", md[:min(80, len(md))])
}
}
func min(a, b int) int {
if a < b {
return a
}
return b
}
func TestTxtToMarkdownPreservesLines(t *testing.T) {
in := "line1\nline2\n\nline4"
md := TxtToMarkdown(in)
if !strings.Contains(md, "line1 \n") || !strings.Contains(md, "line2 \n") {
t.Fatalf("hard breaks missing: %q", md)
}
if !strings.Contains(md, "line4 \n") {
t.Fatalf("last line: %q", md)
}
}
func TestTXTToDOCXPreservesLines(t *testing.T) {
tools := LookPath()
if tools.Pandoc == "" {
t.Skip("pandoc not installed")
}
dir := t.TempDir()
txt := filepath.Join(dir, "sample.txt")
docx := filepath.Join(dir, "sample.docx")
body := "MyBox Auto — resumen\nNº contrato: 123\n\nEstado: Vigente\n"
if err := os.WriteFile(txt, []byte(body), 0o644); err != nil {
t.Fatal(err)
}
if err := tools.TXTToDOCX(txt, docx); err != nil {
t.Fatal(err)
}
mdOut := filepath.Join(dir, "out.md")
if err := tools.DOCXToMD(docx, mdOut); err != nil {
t.Fatal(err)
}
got, err := os.ReadFile(mdOut)
if err != nil {
t.Fatal(err)
}
s := string(got)
for _, want := range []string{"MyBox Auto", "Nº contrato", "Estado: Vigente"} {
if !strings.Contains(s, want) {
t.Fatalf("missing %q in %q", want, s)
}
}
if strings.Contains(s, "MyBox Auto — resumen Nº") {
t.Fatalf("lines collapsed: %q", s)
}
}
func TestMDDocxRoundTrip(t *testing.T) {
tools := LookPath()
if tools.Pandoc == "" {
t.Skip("pandoc not installed")
}
dir := t.TempDir()
md := filepath.Join(dir, "n.md")
docx := filepath.Join(dir, "n.docx")
md2 := filepath.Join(dir, "n2.md")
if err := os.WriteFile(md, []byte("# Title\n\nHello **world**.\n"), 0o644); err != nil {
t.Fatal(err)
}
if err := tools.MDToDOCX(md, docx); err != nil {
t.Fatal(err)
}
if _, err := os.Stat(docx); err != nil {
t.Fatal(err)
}
if err := tools.DOCXToMD(docx, md2); err != nil {
t.Fatal(err)
}
b, err := os.ReadFile(md2)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(string(b), "Hello") {
t.Fatalf("round-trip missing Hello: %s", b)
}
}
func TestOptimizePDF(t *testing.T) {
tools := LookPath()
if tools.Ghostscript == "" || tools.PDFToText == "" {
t.Skip("ghostscript/pdftotext not installed")
}
in := "/tmp/ccgg-original.pdf"
if _, err := os.Stat(in); err != nil {
t.Skip("local fixture not present")
}
dir := t.TempDir()
out := filepath.Join(dir, "out.pdf")
charsIn, err := tools.PDFTextLayerChars(in)
if err != nil || charsIn < 1000 {
t.Skip("fixture has no text layer")
}
if err := tools.OptimizePDF(in, out); err != nil {
t.Fatal(err)
}
charsOut, err := tools.PDFTextLayerChars(out)
if err != nil {
t.Fatal(err)
}
if charsOut < charsIn/2 {
t.Fatalf("text layer lost: in=%d out=%d", charsIn, charsOut)
}
}
func TestToMarkdown_PlainMD(t *testing.T) {
tools := LookPath()
dir := t.TempDir()
p := filepath.Join(dir, "a.md")
_ = os.WriteFile(p, []byte("hi"), 0o644)
res, err := tools.ToMarkdown(p, dir, "eng", 0)
if err != nil || res.Markdown != "hi" {
t.Fatalf("got %+v err=%v", res, err)
}
}
+153
View File
@@ -0,0 +1,153 @@
package docpipe
import (
"bytes"
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
hocr "github.com/eslider/go-hocr"
)
// HOCRResult is structured OCR output for agents.
type HOCRResult struct {
HOCRPath string
Markdown string
YAML string
DidOCR bool
Source string
}
// ImageToHOCR runs tesseract hOCR into outBase+".hocr" (tesseract adds the extension).
// outBase must not include ".hocr". dpi 0 uses tesseract default; phone photos often need 200–300.
func (t Tools) ImageToHOCR(imgPath, outBase, lang string, dpi int) (string, error) {
if t.Tesseract == "" {
return "", fmt.Errorf("tesseract not found on PATH")
}
if lang == "" {
lang = "eng"
}
if outBase == "" {
return "", fmt.Errorf("hOCR output base path required")
}
args := []string{imgPath, outBase, "-l", lang}
if dpi > 0 {
args = append(args, "--dpi", fmt.Sprintf("%d", dpi))
}
args = append(args, resolveHOCRConfig())
cmd := exec.Command(t.Tesseract, args...)
var stderr bytes.Buffer
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return "", fmt.Errorf("tesseract hocr: %w (%s)", err, strings.TrimSpace(stderr.String()))
}
out := outBase + ".hocr"
if _, err := os.Stat(out); err != nil {
// Some builds write .html
alt := outBase + ".html"
if _, err2 := os.Stat(alt); err2 == nil {
return alt, nil
}
return "", fmt.Errorf("tesseract hocr: missing output %s (%s)", out, strings.TrimSpace(stderr.String()))
}
return out, nil
}
// resolveHOCRConfig returns a tesseract config path that works when TESSDATA_PREFIX
// points at a custom traineddata dir without relative "hocr" configs.
func resolveHOCRConfig() string {
var candidates []string
if p := strings.TrimSpace(os.Getenv("TESSDATA_PREFIX")); p != "" {
candidates = append(candidates,
filepath.Join(p, "configs", "hocr"),
filepath.Join(p, "tessdata", "configs", "hocr"),
)
}
candidates = append(candidates,
"/usr/share/tesseract-ocr/5/tessdata/configs/hocr",
"/usr/share/tesseract-ocr/4.00/tessdata/configs/hocr",
"/usr/share/tessdata/configs/hocr",
)
for _, c := range candidates {
if _, err := os.Stat(c); err == nil {
return c
}
}
return "hocr"
}
// HOCRToMarkdown parses an hOCR file via go-hocr and returns Markdown (+ optional YAML).
// Words with confidence in (0, minConf) are dropped; minConf 0 keeps all.
func HOCRToMarkdown(hocrPath string, minConf float32) (md string, yml string, err error) {
doc, err := hocr.ReadFile(hocrPath)
if err != nil {
return "", "", fmt.Errorf("go-hocr: %w", err)
}
md = doc.ToMarkdown(minConf)
yml, err = doc.ToYaml()
if err != nil {
return md, "", err
}
return md, yml, nil
}
// ToHOCRMarkdown OCRs an image (or rasterizes first PDF page) to hOCR → Markdown/YAML.
func (t Tools) ToHOCRMarkdown(path, workDir, lang string, dpi int, minConf float32) (HOCRResult, error) {
res := HOCRResult{Source: path}
if workDir == "" {
workDir = os.TempDir()
}
ext := Ext(path)
img := path
switch ext {
case ".jpg", ".jpeg", ".png", ".tif", ".tiff", ".webp", ".gif", ".bmp":
// ok
case ".pdf":
raster, err := t.pdfFirstPagePNG(path, workDir)
if err != nil {
return res, err
}
img = raster
if dpi == 0 {
dpi = 300
}
default:
return res, fmt.Errorf("hOCR path expects image or PDF, got %q", ext)
}
base := filepath.Join(workDir, trimExt(filepath.Base(path))+".hocr-out")
hocrPath, err := t.ImageToHOCR(img, base, lang, dpi)
if err != nil {
return res, err
}
res.DidOCR = true
res.HOCRPath = hocrPath
md, yml, err := HOCRToMarkdown(hocrPath, minConf)
if err != nil {
return res, err
}
res.Markdown = wrapMD(filepath.Base(path), strings.TrimSpace(md))
res.YAML = yml
return res, nil
}
// pdfFirstPagePNG uses pdftoppm when available.
func (t Tools) pdfFirstPagePNG(pdfPath, workDir string) (string, error) {
pdftoppm, err := exec.LookPath("pdftoppm")
if err != nil {
return "", fmt.Errorf("pdftoppm not found (needed to rasterize PDF for hOCR)")
}
outBase := filepath.Join(workDir, trimExt(filepath.Base(pdfPath))+".page")
cmd := exec.Command(pdftoppm, "-png", "-f", "1", "-singlefile", "-r", "200", pdfPath, outBase)
var stderr bytes.Buffer
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return "", fmt.Errorf("pdftoppm: %w (%s)", err, strings.TrimSpace(stderr.String()))
}
png := outBase + ".png"
if _, err := os.Stat(png); err != nil {
return "", fmt.Errorf("pdftoppm: missing %s", png)
}
return png, nil
}
+49
View File
@@ -0,0 +1,49 @@
package docpipe
import (
"os"
"path/filepath"
"strings"
"testing"
)
func TestHOCRToMarkdown_Fixture(t *testing.T) {
// Minimal hOCR 1.2 snippet
hocrXML := `<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en">
<head>
<title>tesseract</title>
<meta http-equiv="Content-Type" content="text/html;charset=utf-8"/>
<meta name="ocr-system" content="tesseract 5"/>
<meta name="ocr-capabilities" content="ocr_page ocr_carea ocr_par ocr_line ocrx_word"/>
</head>
<body>
<div class="ocr_page" id="page_1" title="image &quot;x.png&quot;; bbox 0 0 100 50; ppageno 0">
<div class="ocr_carea" id="block_1_1" title="bbox 0 0 100 50">
<p class="ocr_par" id="par_1_1" lang="eng" title="bbox 0 0 100 50">
<span class="ocr_line" id="line_1_1" title="bbox 0 0 100 20; baseline 0 0; x_size 20">
<span class="ocrx_word" id="word_1_1" title="bbox 0 0 40 20; x_wconf 96">Hello</span>
<span class="ocrx_word" id="word_1_2" title="bbox 45 0 100 20; x_wconf 92">world</span>
</span>
</p>
</div>
</div>
</body>
</html>`
dir := t.TempDir()
p := filepath.Join(dir, "sample.hocr")
if err := os.WriteFile(p, []byte(hocrXML), 0o644); err != nil {
t.Fatal(err)
}
md, yml, err := HOCRToMarkdown(p, 0)
if err != nil {
t.Fatal(err)
}
if !strings.Contains(md, "Hello") || !strings.Contains(md, "world") {
t.Fatalf("md=%q", md)
}
if !strings.Contains(yml, "Hello") {
t.Fatalf("yaml missing word: %s", yml)
}
}

Some files were not shown because too many files have changed in this diff Show More