# 04 — Swarm sync What crosses the air, what never does, and how it moves. ## The golden rule > **Sync results, not raw data — and only when they are actually needed.** | Data | Synced in flight? | Why | | --- | --- | --- | | `state` (pose) | **Yes — broadcast** | Every drone needs every peer's position to coordinate and avoid collisions | | `detections` | **Yes — published, pulled on demand** | Peers re-task around detections; but not every peer needs every frame of context | | `telemetry` (raw) | **No** | Nobody in the air needs a peer's raw IMU; it offloads at the base ([03](03-data-platform.md)) | Bandwidth is the scarcest resource in the system. Every message class gets an explicit budget; anything unbudgeted stays local. ## Broadcast payload: small on the wire, precise at rest The `state` broadcast is a fixed compact frame: | Field | Type | Notes | | --- | --- | --- | | `drone_id` | uint16 | Fleet-scoped registry | | `ts_ns` | int64 | Epoch nanoseconds, same clock domain as storage | | `pos_x/y/z` | int32 | **Millimeters** in the mission frame — quantized only here, storage keeps full float precision | | `att_roll/pitch/yaw` | int16 | Centi-degrees | | `vel_x/y/z` | int16 | cm/s | | `frame_ref` | uint8 | Frame of reference id (GPS-denied: local/visual-odometry frames must be explicit) | | `flags` | uint8 | Battery-low, returning, degraded-sensors, … | ~40 bytes per frame → a 50-drone swarm at 5 Hz is ~10 KB/s of pose traffic before transport overhead. Trivial even on a congested mesh. `detections` events are slightly larger (class, confidence, bounding volume, ego-pose) but event-shaped and rare by comparison. ## Transport: event-driven pub/sub over an ad-hoc mesh Requirements: peer discovery without infrastructure, pub/sub with late-joiner catch-up, graceful behavior under partitions, tiny footprint, ARM64 support. ```mermaid graph LR subgraph droneA [Drone A] HA["event hook"] --> PA["publisher"] end subgraph droneB [Drone B] SB["subscriber"] --> LB["local state cache + DuckDB"] end subgraph droneC [Drone C — out of direct range] SC["subscriber"] end PA -->|"pose @5Hz, detections on event"| SB SB -->|"store-and-forward relay"| SC ``` ### Recommended: Zenoh - Designed exactly for constrained, dynamic networks: built-in peer discovery, brokerless peer-to-peer mode, store-and-forward, and a query layer on top of pub/sub. - First-class robotics citizenship: an official ROS 2 RMW implementation exists, so the ingestion side and the sync side can share one middleware. - Tiny footprint, ARM64-native. ### Alternatives considered | Option | Verdict | | --- | --- | | **DDS multicast** (ROS 2 default) | Works, battle-tested; but discovery storms and tuning pain on lossy wireless meshes are well documented. Keep as fallback since ROS 2 speaks it natively | | **MQTT** | Needs a broker — a per-drone broker bridge is possible but adds moving parts for no gain over Zenoh | | **Raw UDP multicast** | Perfect as a last-resort minimal profile for the pose broadcast alone (fixed frame, no discovery); no query layer, no reliability — documented as the degraded mode | | **MinIO bucket replication** | Wrong tool for the 5 Hz pose path, right tool for bulk derived datasets — see below | ## Two sync mechanisms, deliberately separate 1. **Fast path — pub/sub (Zenoh):** pose frames and detection events. Fire-and-forget with bounded staleness; consumers keep a peer-state cache. 2. **Bulk path — rsync over persistent SSH:** sealed `detections`/`state` Parquet partitions are pulled opportunistically between drones when links allow. This is how a drone that was out of range catches up on mission history without anyone re-sending events. Why SSH-based bulk sync over object-store replication: - **Identity is already there.** Every drone holds pre-provisioned ed25519 keys and a fixed `known_hosts`/`authorized_keys` set from ground provisioning ([05](05-network-security.md)) — the trust model needs no new machinery. - **One persistent multiplexed session** (`ControlMaster`) per peer costs almost nothing at idle and survives as a single TCP stream; every transfer rides it without new handshakes. - **rsync delta transfer is resumable** across link drops — exactly the failure mode of an ad-hoc mesh — and the shared partition layout makes it trivially incremental: same paths, same files, pull only what is missing. - **Zero extra services** on the flight-critical node. MinIO remains available where an S3 API is genuinely wanted (ground warehouse, and optionally on board), but the in-flight bulk path does not depend on it. Partition healing is automatic: replication is pull-based, addressed by partition path, and idempotent (each drone only ever writes its own `drone=` subtree — **no write conflicts by construction**). Two deliberate non-features of the SSH channel: - **No remote filesystem mounts in flight.** FUSE/sshfs over a lossy mesh inherits NFS hang semantics — processes block uninterruptibly when the link drops. Mounts are fine on the bench and at the dock; in flight, data moves by pull, never by mount. - **No free-form remote execution.** Arbitrary drone-to-drone exec would mean one compromised unit owns the fleet. Instead, every permitted operation is an **SSH forced command** (`command="…"` in `authorized_keys`, one key pair per operation): the key *is* the API. Flexible like a CLI, auditable like an RPC, least-privilege by construction. ## Peer query standard For everything that is not broadcast, drones (and mission services) query each other explicitly. The contract: - **Schema-first:** all datasets share the versioned schemas from [03](03-data-platform.md); a query addresses `dataset + partition predicates + column projection + time range`. - **In flight — query over forced command:** a peer query is a DuckDB statement (from an allow-listed template set) executed via its dedicated SSH forced command, streaming the result back as Parquet/Arrow over the same multiplexed session. No server process, no new port, no new protocol — and the receiving side lands data straight into its own store. - **On the ground — GraphQL as the integration standard:** the warehouse exposes the same schemas through GraphQL for the multi-team surface, where flexibility and introspection matter more than grams and milliwatts. **Arrow Flight** is the upgrade path for heavy columnar serving from T3 if GraphQL result sizes become the bottleneck. - Responses are always columnar batches (Parquet/Arrow), never JSON blobs. ## Event-driven end to end No component polls. The chain from sensor to swarm reaction: ```mermaid sequenceDiagram participant V as video-analytics participant W as parquet-writer participant H as event hook participant Z as pub/sub mesh participant P as peer drone participant M as peer mission logic V->>W: detection row W->>H: derived-data event H->>Z: publish detection Z->>P: deliver (direct or relayed) P->>M: local hook fires M->>M: re-task decision (out of scope) ``` The same hook mechanism that feeds the mesh also feeds local mission logic — one eventing model on board and across the swarm.