CI & Release / Verify documentation (push) Skipped
CI & Release / Verify simulator (push) Skipped
CI & Release / Trivy scan (push) Skipped
CI & Release / Semantic Release (push) Skipped
CI & Release / Verify documentation (pull_request) Successful in 6s
CI & Release / Verify simulator (pull_request) Failing after 10s
CI & Release / Trivy scan (pull_request) Skipped
CI & Release / Semantic Release (pull_request) Skipped
Restore docs/adr to main; add check_adrs.py CI gate so existing ADR-*.md files cannot be modified — supersede with a new numbered ADR instead.
33 lines
1.3 KiB
Markdown
33 lines
1.3 KiB
Markdown
# ADR-0002: One storage format on every tier — Parquet + DuckDB, Hive layout
|
|
|
|
## Status
|
|
|
|
Accepted (2026-07-08)
|
|
|
|
## Context
|
|
|
|
Data lives in three places: on the drone's NVMe in flight (T1), in the ground
|
|
warehouse after landing (T3), and in the dev/simulation environment. If each
|
|
tier used a different format or layout, offload would need transformation code,
|
|
schemas would drift, and a query proven on the bench would not run unchanged on
|
|
flight data.
|
|
|
|
## Decision
|
|
|
|
Use **one storage format on every floor**: columnar **Parquet** files in an
|
|
identical **Hive-partitioned** layout (`dataset=…/flight=…/drone=…/…`), queried
|
|
with **DuckDB**. Flight offload T1 → T3 is therefore a plain **mirror**
|
|
(copy/rsync of partition directories), not an ETL step.
|
|
|
|
## Consequences
|
|
|
|
- Offload is a file operation, testable with a checksum; no format converters
|
|
to maintain or version.
|
|
- The same SQL runs against T1, T3, and simulated data — the simulator
|
|
([`../../simulator/`](../../simulator/)) generates the *real* layout, so tests
|
|
exercise the production format.
|
|
- Partitioning choices are load-bearing; changing them is a schema migration.
|
|
Detailed in [`../03-data-platform.md`](../03-data-platform.md).
|
|
- DuckDB's single-file, no-server model fits an air-gapped drone with no room
|
|
for a database daemon.
|