eSlider dfb0556eda
CI & Release / Verify simulator (push) Skipped
CI & Release / Trivy scan (push) Skipped
CI & Release / Semantic Release (push) Skipped
docs: add design journey, a narrative walk-through linking into the code
A chronological companion to the ADRs: how the design came together, from
the first mental model and 2D prototype through the three data tiers, the
relative-pose broadcast, the layered WireGuard/SSH trust model, SQL as the
peer contract, offload and replication, and the end-to-end proof of concept.
Each chapter links straight into the relevant source and decision records.
Flagged as the recommended first read in the README.
2026-07-09 12:30:40 +01:00

Swarm House

An on-prem data platform for autonomous drone swarms — a design proposal from a DevOps perspective.

A fleet of drones flies fully autonomously: no internet uplink, no external access, all data stays inside the swarm. Each drone collects high-rate sensor telemetry and runs on-board video object detection. The swarm must exchange just enough state to coordinate flight, store everything else locally, and hand its data over to a ground warehouse after landing — reproducibly, securely, and at scale.

This repository describes how to build, deliver, test, and operate that platform: the storage layout, the sync strategy, the network trust model, the CI/CD chain, and the simulation environment that makes it all testable without touching real hardware.

Design principles

  1. Every drone is autonomous. No cluster orchestrator spans the swarm; intermittent mesh connectivity makes that an anti-pattern. Coordination happens through data exchange, not through a control plane.
  2. One storage format on every floor. Parquet + DuckDB with an identical Hive-partitioned layout on the drone, in the ground warehouse, and in the dev environment. Flight offload is a plain mirror operation.
  3. Sync results, not raw data. Only derived state (position, attitude, detections) crosses the air. Raw telemetry stays on local NVMe until the drone lands.
  4. Event-driven, not polled. New derived data triggers downstream actions through hooks; nothing burns cycles asking "anything new yet?".
  5. Zero trust, provisioned ahead. Keys are issued per device before deployment; nothing joins the swarm at runtime. Dev channels are physically absent from production builds.
  6. The whole fleet has one version. A release manifest pins every image digest, model weight, schema, and config. Rollback is atomic.
  7. SQL is the contract. Peer data access is read-only DuckDB SQL over SSH forced commands — the query language already lives on both ends, so no service, port, or protocol is invented for it.
  8. No debug API. The bench and the flight use the same channel: an engineer debugging on the ground runs the identical query through the identical wrapper, permissions, and output format a peer drone would use. What you test is what flies.

Documentation

Document Contents
00 — Glossary Terms and abbreviations used throughout
01 — Problem statement Goal, constraints, knowns vs assumptions
02 — Architecture On-board layers, service composition, communication planes
03 — Data platform Parquet/DuckDB storage tiers, partitioning, hot-write path
04 — Swarm sync What syncs, what does not, pub/sub transport, query standard
05 — Network & security Zero-trust mesh, key provisioning, encrypted transport
06 — Environments Drone / ground warehouse / dev-simulation infrastructures
07 — Observability Logs, service metrics, hardware telemetry
08 — Roadmap Dependency graph, simple to complex
09 — Open questions Known unknowns and proposed answers
10 — Domain context Swarm autonomy principles this design builds on
11 — CI/CD & delivery Pipelines, registry, dev containers, fleet releases
12 — Design journey Start here — a narrative walk-through of how the design came together, linking into the code
ADRs Architecture Decision Records — the decisions behind the above, in the order they were made

New here? Read the design journey first: it tells the story chronologically and links straight into the code and decisions. The design principles below are the what; the ADRs are the why and when — each principle traces to a dated, immutable decision record.

Runnable parts

Component Purpose Stack
simulator/ Virtual drone fleet: generates realistic telemetry through the same Parquet/DuckDB pipeline a real drone would use; N drones via Docker Compose Python, pyarrow, DuckDB
prototype/ 2D top-down swarm visualization: drones, obstacles, inter-drone links with live transfer metrics; live mode renders real positions from the k3d fleet React, TypeScript, Vite, SVG
infra/ansible/ Host preparation: drone bench provisioning (keys, forced commands, WireGuard, Compose bundle) and local k3d simulation cluster Ansible
infra/terraform/ Declared workloads on the sim cluster: virtual fleet + observability (sim-env), warehouse + T1→T3 offload (ground) Terraform, kubernetes provider
infra/gitops/ Flux overlays for the ground segment — policy labels, dashboard bundles; drones stay on the fleet manifest Flux, Kustomize

Shared runtime data (gitignored): var/t1/ (live lake), var/t3/ (warehouse). See CONTRIBUTING.md.

Quick start

# Spin up a virtual swarm of 10 drones (Compose, no cluster needed)
cd simulator && DRONE_COUNT=10 docker compose up

# Run the visualization
cd prototype && npm install && npm run dev

Or run the full miniature of the ground segment — k3d cluster, fleet as StatefulSet, Grafana, explorer, warehouse with post-flight offload:

ansible-playbook infra/ansible/sim-cluster.yml
cd infra/terraform/sim-env && terraform init && terraform apply
cd ../ground && terraform init && terraform apply
# then press "go live" in the prototype header

See 06 — Environments → IaC boundaries for why Ansible owns hosts, Terraform owns the ground segment, and neither flies on the drone.

Contributing (TDD-first workflow): CONTRIBUTING.md.

CI & releases

Every push runs the pipeline in .github/workflows/release.yaml: pytest unit tests, a compile check, a 10-second virtual smoke flight verified with DuckDB, and a Trivy vulnerability/misconfiguration scan. The simulator Docker build also runs pytest — a failing test blocks the image. Pushes to main then cut a semantic version from the conventional-commit type (feat: → minor, !/BREAKING CHANGE → major, else patch), tag the commit, and publish a release with a fleet release manifest (versioned artifact list with checksums — the delivery unit described in 11 — CI/CD & delivery) plus a source bundle.

S
Description
No description provided
Readme
20 MiB
2026-07-17 12:18:13 +01:00
Languages
Python 38.6%
TypeScript 26.8%
HCL 25.4%
HTML 7.3%
Dockerfile 1.1%
Other 0.8%