docs: record architecture decisions (ADR-0001..0009) and operations knowledge
Add docs/adr/ with a chronological ADR set reconstructed from the project history (no orchestrator, one storage format, derived-state sync, SQL-over-SSH contract, WireGuard beneath SSH, IaC boundaries, non-root image, runtime data paths, bounded lake scans), an index README, and a template. Distinguish repo-local ADRs from ecosystem-wide ASRs. Expand AGENTS.md with a decision-records section and a runtime/operations guide (service URLs, rebuild workflow, CPU budget knobs). Reference the ADRs from README, CONTRIBUTING, and the affected design docs, and add an open question for T1 partition pruning.
This commit is contained in:
@@ -2,6 +2,8 @@
|
||||
|
||||
Zero trust, fully air-gapped, everything provisioned before wheels-up.
|
||||
|
||||
> **Decision:** [ADR-0005 — WireGuard beneath SSH](adr/ADR-0005-wireguard-beneath-ssh.md).
|
||||
|
||||
## Trust model
|
||||
|
||||
- **Nothing joins the swarm at runtime.** Every device receives its identity (key pair + certificate) during ground provisioning, signed by the fleet's offline CA. There is no trust-on-first-use, no runtime enrollment endpoint, no exception path.
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
Three infrastructures, one data contract. The layout, schemas, and pipelines are identical everywhere; only scale and lifetime differ.
|
||||
|
||||
> **Decisions:** [ADR-0006 — IaC boundaries](adr/ADR-0006-iac-boundaries.md), [ADR-0008 — runtime data paths](adr/ADR-0008-runtime-data-paths.md).
|
||||
|
||||
```mermaid
|
||||
graph LR
|
||||
subgraph dev [3 — Dev and simulation]
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
The platform watches itself with the same discipline it applies to sensor data — and largely through the same pipeline.
|
||||
|
||||
> **Decision:** [ADR-0009 — bound observability lake scans to a flight window and cache](adr/ADR-0009-bounded-lake-scans.md).
|
||||
|
||||
## Two kinds of signals, one storage
|
||||
|
||||
| Signal | Examples | Where it goes |
|
||||
|
||||
@@ -61,3 +61,9 @@ Who owns the device registry (drone ids, key issuance, revocation) organizationa
|
||||
Sub-250 g class units cannot run the full stack (no GPU, minimal CPU/storage).
|
||||
|
||||
*Proposal:* define a minimal profile early — state publisher + UDP-multicast pose broadcast, no local Parquet store, no video — so the swarm protocol never assumes full-stack peers.
|
||||
|
||||
## 11 — Lake retention and partition pruning
|
||||
|
||||
The observability path now scans only the most recent flights ([ADR-0009](adr/ADR-0009-bounded-lake-scans.md)), but old `flight=` partitions still accumulate on T1 disk after offload to T3.
|
||||
|
||||
*Proposal:* a scheduled prune of T1 partitions whose flights are confirmed present in the T3 warehouse (offload as the retention gate), configurable by age and free-space watermark. On the drone the same policy is bounded by the NVMe quota ladder ([03](03-data-platform.md)); in the sim environment it is a ground CronJob alongside the offload job.
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
How software, models, and configuration reach the fleet — reproducibly, scanned, versioned, and atomically rollback-able. Everything lives inside the air gap.
|
||||
|
||||
> **Decisions:** [ADR-0006 — IaC boundaries & fleet manifest](adr/ADR-0006-iac-boundaries.md), [ADR-0007 — non-root image](adr/ADR-0007-non-root-image.md).
|
||||
|
||||
## Source and pipelines: GitLab
|
||||
|
||||
GitLab (self-managed, on-prem) is the backbone: repositories, CI/CD, and the container registry in one system.
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
# ADR-0001: No swarm-wide orchestrator; coordinate through data
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08)
|
||||
|
||||
## Context
|
||||
|
||||
The fleet flies fully autonomously with no internet uplink and only
|
||||
intermittent mesh connectivity between drones. A conventional control plane
|
||||
(Kubernetes across the swarm, a leader election, a central scheduler) assumes
|
||||
stable membership and low-latency links — exactly what a swarm does not have.
|
||||
Partitions are the normal case, not the exception.
|
||||
|
||||
## Options
|
||||
|
||||
| Option | Pros | Cons |
|
||||
| --- | --- | --- |
|
||||
| A — Cluster orchestrator spanning the swarm | Familiar tooling | Assumes stable membership; a partition stalls coordination; single point of failure in the air |
|
||||
| B — Each drone autonomous, coordinating through exchanged data | Survives partitions; no air-side control plane to fail | No global view; behaviour must be derivable from local state |
|
||||
|
||||
## Decision
|
||||
|
||||
Every drone is an autonomous node. There is **no orchestrator spanning the
|
||||
swarm**. Coordination happens by exchanging derived state (see
|
||||
[ADR-0003](ADR-0003-sync-derived-state-only.md)), and each drone makes its own
|
||||
in-flight decisions against its local data window. Orchestration tooling
|
||||
(Kubernetes, Flux) is confined to the **ground** segment
|
||||
([ADR-0006](ADR-0006-iac-boundaries.md)).
|
||||
|
||||
## Consequences
|
||||
|
||||
- The platform degrades gracefully under partition: a drone that loses peers
|
||||
keeps flying and recording.
|
||||
- There is no single "fleet state" to query in the air; health questions are
|
||||
answered from each drone's own store and reconciled on the ground.
|
||||
- Design principle 1 in [`../../README.md`](../../README.md) and the on-board
|
||||
model in [`../02-architecture.md`](../02-architecture.md) follow from this.
|
||||
@@ -0,0 +1,32 @@
|
||||
# ADR-0002: One storage format on every tier — Parquet + DuckDB, Hive layout
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08)
|
||||
|
||||
## Context
|
||||
|
||||
Data lives in three places: on the drone's NVMe in flight (T1), in the ground
|
||||
warehouse after landing (T3), and in the dev/simulation environment. If each
|
||||
tier used a different format or layout, offload would need transformation code,
|
||||
schemas would drift, and a query proven on the bench would not run unchanged on
|
||||
flight data.
|
||||
|
||||
## Decision
|
||||
|
||||
Use **one storage format on every floor**: columnar **Parquet** files in an
|
||||
identical **Hive-partitioned** layout (`dataset=…/flight=…/drone=…/…`), queried
|
||||
with **DuckDB**. Flight offload T1 → T3 is therefore a plain **mirror**
|
||||
(copy/rsync of partition directories), not an ETL step.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Offload is a file operation, testable with a checksum; no format converters
|
||||
to maintain or version.
|
||||
- The same SQL runs against T1, T3, and simulated data — the simulator
|
||||
([`../../simulator/`](../../simulator/)) generates the *real* layout, so tests
|
||||
exercise the production format.
|
||||
- Partitioning choices are load-bearing; changing them is a schema migration.
|
||||
Detailed in [`../03-data-platform.md`](../03-data-platform.md).
|
||||
- DuckDB's single-file, no-server model fits an air-gapped drone with no room
|
||||
for a database daemon.
|
||||
@@ -0,0 +1,33 @@
|
||||
# ADR-0003: Sync derived state only; raw telemetry stays local
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08)
|
||||
|
||||
## Context
|
||||
|
||||
Each drone produces high-rate sensor telemetry and on-board video detections —
|
||||
far more than the mesh can carry. But peers only need enough to coordinate
|
||||
flight: where everyone is, where they are heading, what was detected. Pushing
|
||||
raw feeds across the air would saturate the radio and drain batteries for data
|
||||
nobody reads in flight.
|
||||
|
||||
## Decision
|
||||
|
||||
Only **derived state** crosses the air — position, attitude, detections — as
|
||||
small pub/sub broadcasts (~5 Hz). **Raw telemetry stays on local NVMe** until
|
||||
the drone lands, then travels to the warehouse during offload
|
||||
([ADR-0002](ADR-0002-one-storage-format.md)). The design is **event-driven**:
|
||||
new derived data triggers downstream action through hooks; nothing polls asking
|
||||
"anything new yet?".
|
||||
|
||||
## Consequences
|
||||
|
||||
- Air bandwidth scales with the number of peers and the broadcast rate, not
|
||||
with sensor resolution.
|
||||
- The warehouse is the only place with the full picture; in-flight decisions use
|
||||
the local window plus peer broadcasts.
|
||||
- Broadcast staleness per peer becomes a first-class health signal
|
||||
([`../07-observability.md`](../07-observability.md)).
|
||||
- What syncs and what does not is specified in
|
||||
[`../04-swarm-sync.md`](../04-swarm-sync.md).
|
||||
@@ -0,0 +1,41 @@
|
||||
# ADR-0004: Read-only SQL-over-SSH as the peer query contract; no debug API
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08). Supersedes the earlier sketch of a bespoke query service.
|
||||
|
||||
## Context
|
||||
|
||||
Beyond the 5 Hz broadcast, a peer (or an engineer on the bench) sometimes needs
|
||||
to *ask* a drone a richer question: "show me your detections in the last minute
|
||||
near this position". Inventing a service for this means a new port, a new
|
||||
protocol, an auth layer, and a second code path that only runs during
|
||||
debugging — the classic "works on the bench, untested in flight" trap.
|
||||
|
||||
## Options
|
||||
|
||||
| Option | Pros | Cons |
|
||||
| --- | --- | --- |
|
||||
| A — Custom query microservice + REST API | Flexible | New port/protocol/auth to secure; a debug-only path that never flies |
|
||||
| B — Read-only DuckDB SQL over an SSH forced command | Reuses SSH trust + the query language already on both ends; identical path on bench and in flight | SQL must be gated to read-only |
|
||||
|
||||
## Decision
|
||||
|
||||
Peer data access is **read-only DuckDB SQL executed through an SSH forced
|
||||
command**. SQL is already the query language on both ends
|
||||
([ADR-0002](ADR-0002-one-storage-format.md)), and SSH already carries the trust
|
||||
model ([ADR-0005](ADR-0005-wireguard-beneath-ssh.md)), so no new service, port,
|
||||
or protocol is invented. A **SQL gate** rejects anything that writes
|
||||
(`COPY`, `INSERT`, `ATTACH`, pragmas). There is **no separate debug API**: an
|
||||
engineer debugging on the ground runs the identical query through the identical
|
||||
wrapper, permissions, and output format a peer drone would use.
|
||||
|
||||
## Consequences
|
||||
|
||||
- One access path is built, secured, and tested — "what you test is what
|
||||
flies".
|
||||
- The SQL gate is safety-critical and is covered by unit tests
|
||||
(`simulator/tests/test_sql_gate.py`).
|
||||
- The explorer and the prototype's live mode are *just another read-only
|
||||
consumer* of this same contract ([`../04-swarm-sync.md`](../04-swarm-sync.md)).
|
||||
- Design principles 7 and 8 in [`../../README.md`](../../README.md) restate this.
|
||||
@@ -0,0 +1,40 @@
|
||||
# ADR-0005: WireGuard beneath SSH for the mesh transport
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08)
|
||||
|
||||
## Context
|
||||
|
||||
Peer bulk sync and SQL-over-SSH ([ADR-0004](ADR-0004-sql-over-ssh-contract.md))
|
||||
need an encrypted, authenticated channel across an ad-hoc radio mesh. SSH alone
|
||||
would work at the application layer, but it leaves the drones' listening
|
||||
surface (sshd, any other service port) exposed on the raw mesh network, and it
|
||||
gives no cheap, uniform way to authenticate *the peer* before the application
|
||||
handshake.
|
||||
|
||||
## Options
|
||||
|
||||
| Option | Pros | Cons |
|
||||
| --- | --- | --- |
|
||||
| A — SSH only | One daemon, familiar | Service ports exposed on the raw mesh; per-service auth; no network-layer peer identity |
|
||||
| B — WireGuard tunnel, SSH inside it | Network-layer peer auth via static keys; only the WG port is exposed; SSH speaks over a private, encrypted subnet | Two layers to provision |
|
||||
|
||||
## Decision
|
||||
|
||||
Run **WireGuard as the network layer** and **SSH on top of it**. WireGuard
|
||||
authenticates peers by pre-shared static public keys and presents a private
|
||||
encrypted subnet; SSH forced commands then carry the SQL and bulk-sync contract
|
||||
over that subnet. Keys are issued **per device before deployment** — nothing
|
||||
joins the mesh at runtime (design principle 5).
|
||||
|
||||
## Consequences
|
||||
|
||||
- The only thing exposed on the raw radio is the WireGuard port; sshd and the
|
||||
data services listen only on the WG interface.
|
||||
- Peer identity is established once, at the network layer, and reused by every
|
||||
application above it.
|
||||
- Provisioning must place both WG and SSH keys during host prep — handled by
|
||||
`infra/ansible/drone-provision.yml` ([ADR-0006](ADR-0006-iac-boundaries.md)).
|
||||
- Rationale and the "why not SSH alone" argument live in
|
||||
[`../05-network-security.md`](../05-network-security.md).
|
||||
@@ -0,0 +1,41 @@
|
||||
# ADR-0006: IaC split — Ansible hosts, Terraform ground, Flux ground-only, fleet manifest for drones
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08)
|
||||
|
||||
## Context
|
||||
|
||||
Three very different things need provisioning: (1) drone hosts, which are
|
||||
air-gapped and never reachable by a controller in flight; (2) the ground
|
||||
segment (warehouse, observability, offload), a normal k3s cluster; (3) the
|
||||
simulation environment that mirrors the ground segment locally. Using one tool
|
||||
for all three would force a GitOps controller or a Terraform apply loop onto the
|
||||
drone — impossible for an autonomous, disconnected node
|
||||
([ADR-0001](ADR-0001-no-swarm-orchestrator.md)).
|
||||
|
||||
## Decision
|
||||
|
||||
Split infrastructure ownership by lifecycle:
|
||||
|
||||
| Layer | Tool | Scope |
|
||||
| --- | --- | --- |
|
||||
| Host preparation | **Ansible** | Drone bench prep (keys, forced commands, WireGuard, Compose bundle) and the local k3d sim cluster |
|
||||
| Ground workloads | **Terraform** | k3s/k3d resources: virtual fleet + observability (`sim-env`), warehouse + T1→T3 offload (`ground`) |
|
||||
| Ground continuous delivery | **Flux (GitOps)** | Ground segment **only** — policy labels, dashboard bundles |
|
||||
| Drone delivery | **Fleet release manifest** | A versioned, digest-pinned artifact list; no controller pulls to the drone |
|
||||
|
||||
Drones are delivered by the **fleet release manifest** (one version for the
|
||||
whole fleet, atomic rollback), never by a controller reaching into the air.
|
||||
|
||||
## Consequences
|
||||
|
||||
- No GitOps agent or `terraform apply` ever targets a drone; the air side stays
|
||||
controller-free.
|
||||
- Terraform state is environment-local and kept out of Git
|
||||
([ADR-0008](ADR-0008-runtime-data-paths.md)).
|
||||
- The sim cluster is a faithful miniature: the same Terraform describes it and
|
||||
the ground segment.
|
||||
- Boundaries are documented in
|
||||
[`../06-environments.md`](../06-environments.md#iac-boundaries); delivery in
|
||||
[`../11-cicd-delivery.md`](../11-cicd-delivery.md).
|
||||
@@ -0,0 +1,27 @@
|
||||
# ADR-0007: Run the simulator image as a non-root user
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08)
|
||||
|
||||
## Context
|
||||
|
||||
The Trivy scan in CI flagged **DS-0002**: the simulator image ran as `root`
|
||||
because no `USER` was set. A container writing the data lake as root also
|
||||
produces root-owned Parquet on bind mounts, which then blocks a later non-root
|
||||
process from the same files.
|
||||
|
||||
## Decision
|
||||
|
||||
Create a dedicated unprivileged user **`swarm` (uid/gid 10001)** in the
|
||||
`simulator/Dockerfile` and switch to it with `USER swarm` before the entrypoint.
|
||||
The uid is fixed (not auto-assigned) so host-side ownership of the shared data
|
||||
paths ([ADR-0008](ADR-0008-runtime-data-paths.md)) is deterministic.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Trivy DS-0002 clears; the image follows least-privilege.
|
||||
- Bind-mounted data is owned by `10001:10001`; host prep (`sim-cluster.yml`) and
|
||||
Ansible `chown` the `var/t1` / `var/t3` paths to this uid so k3d pods can write.
|
||||
- Any future service image reusing this data must run as the same uid or share
|
||||
the group.
|
||||
@@ -0,0 +1,40 @@
|
||||
# ADR-0008: Shared `var/t1` / `var/t3` runtime paths; state out of Git
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08)
|
||||
|
||||
## Context
|
||||
|
||||
Compose, the k3d fleet, and the ground offload job all need to read and write
|
||||
the same data lake, so that "go live" in the prototype shows the data the
|
||||
simulator just generated. Early on, Terraform state and the provider cache were
|
||||
accidentally committed (~6 MB of `*.tfstate` and a vendored provider binary),
|
||||
polluting history and risking stale/secret data in Git.
|
||||
|
||||
## Decision
|
||||
|
||||
Standardise two **gitignored** runtime directories at the repo root:
|
||||
|
||||
| Path | Tier | Contents |
|
||||
| --- | --- | --- |
|
||||
| `var/t1/` | T1 live lake | In-flight Parquet from Compose or the k3d fleet |
|
||||
| `var/t3/` | T3 warehouse | Offloaded historical partitions |
|
||||
|
||||
They are created automatically and overridable via
|
||||
`SWARM_T1_DIR` / `SWARM_WAREHOUSE_DATA` / `SWARM_SIM_DATA`. Treat `var/` as
|
||||
**runtime-only** per the ecosystem canonical layout. `.gitignore` excludes
|
||||
`*.tfstate`, `*.tfstate.backup`, and `.terraform/` while **keeping**
|
||||
`.terraform.lock.hcl` (the lock is source, the state is not).
|
||||
|
||||
## Consequences
|
||||
|
||||
- One lake feeds every component; the live prototype reflects real generated
|
||||
data.
|
||||
- Git history carries no machine state or large binaries; `terraform` treats
|
||||
its state as local/remote-backend concern, not versioned source.
|
||||
- Ownership of these paths is fixed to uid 10001
|
||||
([ADR-0007](ADR-0007-non-root-image.md)).
|
||||
- Recovery from earlier state drift is documented in
|
||||
[`../../infra/terraform/README.md`](../../infra/terraform/README.md)
|
||||
(`terraform import`).
|
||||
@@ -0,0 +1,52 @@
|
||||
# ADR-0009: Bound observability lake scans to a flight window and cache
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-07-08)
|
||||
|
||||
## Context
|
||||
|
||||
The k3d sim node was running at ~245% CPU. Three causes compounded:
|
||||
|
||||
1. **Explorer** rebuilt a DuckDB view over the *entire* lake on every HTTP
|
||||
request (~2000 Parquet files, 200 MB+) — roughly 1.5 cores.
|
||||
2. **Exporter** scanned the full lake every 5 s — roughly 0.7 cores.
|
||||
3. **Drones** exited after sealing a flight; Kubernetes restarted each pod,
|
||||
and every restart appended a fresh `flight=` partition tree, so the lake grew
|
||||
without bound and every scan got more expensive.
|
||||
|
||||
The cost is inherent to scanning an append-only lake whose size grows with pod
|
||||
churn, not to any single slow query. Metrics and the live view only ever need
|
||||
recent flights.
|
||||
|
||||
## Decision
|
||||
|
||||
Bound and cache the scans instead of reading the whole lake:
|
||||
|
||||
- **Flight window.** The exporter and explorer scan only the most recent
|
||||
`METRIC_FLIGHT_WINDOW` flights (default 5) via a shared helper
|
||||
(`simulator/monitoring/lake.py`). Older partitions stay on disk but out of the
|
||||
hot path.
|
||||
- **Caching.** Explorer caches DuckDB views (`VIEW_REFRESH_S=30`) and the
|
||||
partition tree (`TREE_CACHE_S=15`); the exporter scans every
|
||||
`SCAN_INTERVAL_S=30` instead of 5 s.
|
||||
- **Stop restart churn.** Drones set `KEEP_ALIVE=1` and idle after sealing
|
||||
instead of exiting, so no new flight tree is created per restart. In k3d,
|
||||
`IMU_HZ=20` and resource limits cap per-drone load; the live prototype polls
|
||||
every 2 s.
|
||||
- **Self-observability.** Add node-exporter, a
|
||||
`swarm_exporter_scan_duration_seconds` metric, and a **Swarm Platform**
|
||||
Grafana dashboard (node CPU/memory + scan health) alongside the existing fleet
|
||||
dashboard.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Steady state dropped to explorer ~29 m, exporter ~16 m CPU; the k3d node to
|
||||
~56% (from ~245%).
|
||||
- Metrics reflect recent flights only; long-horizon analysis is a warehouse
|
||||
(T3) query, not an exporter concern — consistent with
|
||||
[ADR-0003](ADR-0003-sync-derived-state-only.md).
|
||||
- Old flight partitions accumulate on disk; pruning them is a follow-up
|
||||
(tracked in [`../09-open-questions.md`](../09-open-questions.md)).
|
||||
- Observability design is documented in
|
||||
[`../07-observability.md`](../07-observability.md).
|
||||
@@ -0,0 +1,40 @@
|
||||
# Architecture Decision Records
|
||||
|
||||
This directory records the significant decisions behind Swarm House, in the
|
||||
order they were made. Each record is immutable once **Accepted** — a later
|
||||
decision that changes course gets a new number and supersedes the old one
|
||||
rather than editing it.
|
||||
|
||||
## ADR vs ASR
|
||||
|
||||
| Kind | What it captures | Where it lives |
|
||||
| --- | --- | --- |
|
||||
| **ADR** (Architecture Decision Record) | A point-in-time decision for *this* repository, with the options considered and their consequences | Here, in `docs/adr/` |
|
||||
| **ASR** (Architecture Standard Record) | A standing, cross-repository policy every project must obey (Go-first, secrets, layout) | The ecosystem `inventar` repo, not here |
|
||||
|
||||
The rules in [`../../AGENTS.md`](../../AGENTS.md) are this repo's local standing
|
||||
constraints; anything ecosystem-wide belongs in `inventar` as an ASR. Design
|
||||
decisions specific to the platform are ADRs and belong here.
|
||||
|
||||
## Index
|
||||
|
||||
| ADR | Decision | Status |
|
||||
| --- | --- | --- |
|
||||
| [ADR-0001](ADR-0001-no-swarm-orchestrator.md) | No swarm-wide orchestrator; coordinate through data | Accepted |
|
||||
| [ADR-0002](ADR-0002-one-storage-format.md) | One storage format on every tier — Parquet + DuckDB, Hive layout | Accepted |
|
||||
| [ADR-0003](ADR-0003-sync-derived-state-only.md) | Sync derived state only; raw telemetry stays local | Accepted |
|
||||
| [ADR-0004](ADR-0004-sql-over-ssh-contract.md) | Read-only SQL-over-SSH as the peer query contract; no debug API | Accepted |
|
||||
| [ADR-0005](ADR-0005-wireguard-beneath-ssh.md) | WireGuard beneath SSH for the mesh transport | Accepted |
|
||||
| [ADR-0006](ADR-0006-iac-boundaries.md) | IaC split: Ansible hosts, Terraform ground, Flux ground-only, fleet manifest for drones | Accepted |
|
||||
| [ADR-0007](ADR-0007-non-root-image.md) | Run the simulator image as a non-root user | Accepted |
|
||||
| [ADR-0008](ADR-0008-runtime-data-paths.md) | Shared `var/t1` / `var/t3` runtime paths; state out of Git | Accepted |
|
||||
| [ADR-0009](ADR-0009-bounded-lake-scans.md) | Bound observability lake scans to a flight window and cache | Accepted |
|
||||
|
||||
## Writing a new ADR
|
||||
|
||||
1. Copy [`TEMPLATE.md`](TEMPLATE.md) to `ADR-NNNN-short-slug.md` (next free number).
|
||||
2. Fill in Context, Options (if more than one was real), Decision, Consequences.
|
||||
3. Add a row to the index above.
|
||||
4. Cross-link the ADR from the design doc it affects (and vice versa).
|
||||
|
||||
Keep the tone laconic: state the decision and why, skip the narrative.
|
||||
@@ -0,0 +1,29 @@
|
||||
# ADR-NNNN: Short decision title
|
||||
|
||||
## Status
|
||||
|
||||
Proposed | Accepted | Superseded by ADR-XXXX
|
||||
|
||||
## Context
|
||||
|
||||
What forces are in play? What problem or constraint triggered the decision?
|
||||
Keep it to the facts that actually shaped the choice.
|
||||
|
||||
## Options
|
||||
|
||||
| Option | Pros | Cons |
|
||||
| --- | --- | --- |
|
||||
| A — … | … | … |
|
||||
| B — … | … | … |
|
||||
|
||||
(Drop this section when only one option was ever real.)
|
||||
|
||||
## Decision
|
||||
|
||||
The choice, stated plainly. One paragraph.
|
||||
|
||||
## Consequences
|
||||
|
||||
- What becomes easier.
|
||||
- What becomes harder or is now forbidden.
|
||||
- Follow-ups, links to affected docs.
|
||||
Reference in New Issue
Block a user