aboutsummaryrefslogtreecommitdiffstats
path: root/docs/MESHBAY_DESIGN.md
diff options
context:
space:
mode:
Diffstat (limited to 'docs/MESHBAY_DESIGN.md')
-rw-r--r--docs/MESHBAY_DESIGN.md160
1 files changed, 149 insertions, 11 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md
index 18b23cf..06ccaca 100644
--- a/docs/MESHBAY_DESIGN.md
+++ b/docs/MESHBAY_DESIGN.md
@@ -16,7 +16,7 @@
> them — it names the invariant that holds today, not the incident that produced
> it. §13 is the register of those labels.
>
-> Wire versions at the time of writing: **MNP 6.0** (oldest peer accepted 4.0),
+> Wire versions at the time of writing: **MNP 6.1** (oldest peer accepted 4.0),
> **MHP 0.1**, packages **0.19.0**. The normative source for the wire format is
> `MESHBAY_NODE_PROTOCOL.md`; this document states the design the protocol
> serves, not its byte layout.
@@ -1693,6 +1693,17 @@ Five protections, and they are the substance:
disk one capped file at a time;
- the target root must be **writable and available**, enforced by the node.
+**A full disk is a stated refusal, `disk_full`.** At chunk 0 the node compares the
+announced size (`total_chunks` × the chunk's length, an upper bound within one
+chunk) with the free space of the destination's filesystem, and refuses when the
+upload would leave less than `DISK_RESERVE_BYTES` (1 GiB) — a disk filled to its
+last byte breaks the node's own databases and the operator's system too. Two
+uploads can pass that check together, so a write that fails with `ENOSPC` or
+`EDQUOT` is refused the same way, and its `.part` and state are dropped: they
+cannot be finished until the operator makes room, and keeping them holds the
+space that ran out. The client carries the code on the error and translates it,
+so a caller can stop rather than retry.
+
**There is no quarantine subdirectory.** A folder appearing beside the operator's
library because somebody sent a file is the node deciding how their disk is
arranged. What made a quarantine worth having was never the subdirectory — it is
@@ -1838,7 +1849,10 @@ chat opens at the newest page. A forwards pager is not what a chat opens with.
`meshbay-node chat prune <days>` deletes **messages only, never an epoch key**. An
epoch with no messages is harmless; an epoch key deleted while messages still need
-it is an unreadable archive.
+it is an unreadable archive. The operator's purge (the signed `chat_purge`, from
+the Chat settings) is the same rule with no age: every message goes, the epoch
+keys and the attachments stay. The replay index goes with the rows, which costs
+nothing, since a sealed message is taken only from the device that signed it.
**A message is bounded in size and in rate, like every other member-supplied
write.** Sending one costs the operator a row that nothing expires, every other
@@ -1987,8 +2001,9 @@ by accident (§2.4).
**Stores:** accounts (username, encrypted email, status, role), the group registry
and membership, IP logs (one year, legal retention), node registrations, refresh
-tokens, notifications, the moderation blocklist, instance policy, and per-account
-device keys for hub login.
+tokens, notifications, the moderation blocklist, instance policy, per-account
+device keys for hub login, and the phones that asked to be notified — a poll
+secret's hash and, for a phone with a push distributor, its endpoint (§11.3).
**Does not store:** file content, file names, private-group indexes, message
content, private keys, group keys, keypair bundles, user identity keys, node IPs
@@ -3308,6 +3323,88 @@ are two albums.
---
+### 9.12 Photo backup (Android)
+
+The Android application sends the photos taken on the phone to **one folder of
+one group**, chosen by the member, once a day. It is a client feature over the
+upload path that exists: a backed-up photo is a `file_upload` into a writable
+root under a slot the node granted, with every rule of §6.4, and its owner is
+recorded as for any upload. **No message, no node state and no hub state was
+added for it**; the node cannot tell a backed-up photo from one sent by hand,
+and does not need to.
+
+**Split.** The phone (`photos/` in the Android package, reached as
+`platform.photoSync`) lists the photos through MediaStore, keeps the ledger and
+hands the bytes over. The page (`photo-sync.js`) decides when a run is due and
+does the sending, because the transport and the group key are there. Bytes
+cross by an opaque token the page fetches from its own origin,
+`/photosync/<token>`, served by the shell's request interception: valid for the
+run that issued it and for the one photo it names. The page never receives a
+`content://` URI. A browser and the desktop have no camera roll and no such
+object — absent, not refusing (§11.3).
+
+**When.** A run is due 24 hours after the last run that *finished*; an
+interrupted one is retried at the next chance, at most every fifteen minutes.
+Due-ness is checked at start, on return to the foreground, when the network
+turns unmetered, and by one timer at the due time — nothing polls. A scheduled
+run needs an **unmetered** network (`NET_CAPABILITY_NOT_METERED`, not "Wi-Fi":
+a phone on another phone's hotspot is on Wi-Fi and spending that phone's data).
+"Back up now" runs whenever; on a metered network it first states the count and
+size and asks — the one confirmation in this feature that is not about where
+the photos go. Turning metered mid-run stops after the file in flight; the
+node's partial upload resumes it (§8.5).
+
+**One photo at a time, under a slot.** Each upload asks for a transfer slot
+like a member sending by hand and gives it back, so a family's daily backups
+never occupy the node's upload pool ahead of a person.
+
+**Where.** The member chooses the group (one per phone) and a folder among
+those that are writable, available and shown by the group's Photos tab — a
+folder outside those would take the photos and show them nowhere. Inside it,
+each photo goes under `YYYY/MM` from when it was taken, because an album is a
+directory (§9.9) and one folder of twenty thousand is a slow album. Albums on
+the phone are chosen too; the camera alone is the default, because screenshots
+and saved images are where photos nobody meant to share live. **By default the
+photos already on the phone are sent, newest first**, then each new one; "from
+now on" is an option. A confirmation is drawn **when the destination or the
+starting point changes, never per run and never for a change of albums alone**:
+it names the group, its owner, its member count, the folder, the count and size
+about to go, and says that members can download them and that what they
+downloaded cannot be taken back.
+
+**Additive by construction.** The backup never deletes, renames or replaces
+anything on the node: `photo-sync.js` reaches no such operation
+(`test_photo_sync.py` reads it for them). **The ledger is the memory of what was
+sent, never a mirror of the phone**: per (account, group, folder), each photo's
+MediaStore id, modification date, size, the SHA-256 of the bytes sent and the
+name the node's ack gave. A run sends what the ledger lacks and never compares
+the other way, so a photo deleted on the phone stays in the group, and one
+deleted on the node is not sent again. An empty ledger (a reinstall) is
+reconciled against the folder by name and size before anything is sent.
+
+**Edits.** A photo edited in place keeps its id and changes its modification
+date; when the size also changed, or the bytes no longer hash to what was sent,
+the edit is sent **beside the original**, as `<stem>-edited-<YYYYMMDD-HHMMSS>.<ext>`
+from the edit's date. A touch that left the bytes alone sends nothing. "Save as
+copy" makes a new photo and needs no rule.
+
+**Location.** The manifest does not ask for `ACCESS_MEDIA_LOCATION`, so a
+photo read through MediaStore has its EXIF location **redacted by the
+platform**: a camera roll going to a group does not say where its owner lives,
+and nobody had to do anything for it.
+
+**Refusals.** A refusal that will hold tomorrow — `disk_full` (§6.4), a folder
+no longer writable, gone or on a drive that is not plugged, the member removed
+from the group, the photo permission withdrawn — stops the run, is said once in
+a notification and in Settings, and is retried a day later rather than at every
+opening. Anything else (the node offline, the network gone) is an interruption
+and is retried.
+
+**Screen off.** The sending lives in the page, so a run holds a `dataSync`
+foreground service (`BackupService`) and the WebView reported visible, like a
+cast (§11.3). Android 15 limits that type to six hours a day; the service stops
+when told and the run carries on at the next opening.
+
## 10. Filesystem portability
**exFAT and NTFS on Windows are the common case, not an edge case.** Most users are
@@ -3442,11 +3539,48 @@ narrow bridge — and Android has all three:
locks and the renderer's priority policy do not prevent it. A cast's pipeline
lives in the page (WebRTC → decrypt → relay), so while one runs the shell
keeps the WebView reported visible and holds a media-playback foreground
- service; nowhere else, because a page never hidden is never throttled.
+ service. A photo backup does the same with a `dataSync` service of its own
+ (§9.12); nowhere else, because a page never hidden is never throttled.
+- **Notifications reach a closed application with nothing to install, and
+ never through a vendor push service.** Android lets nothing hold a connection
+ for an application that is not running, so there are three ways to wake a
+ phone: a vendor push service (Google sees who is notified and when; not
+ F-Droid), a push distributor the owner installs, or the phone asking. **The
+ phone asks by default**, and uses a distributor when one is already there.
+ Turned on, the page registers the phone (`POST /v1/push/subscriptions`) and
+ gets a row id and a **poll secret**; a system job (`JobScheduler`, no library)
+ then fetches what is new with `POST /v1/push/poll` every fifteen minutes,
+ which Android stretches under Doze. The secret is the only credential the
+ background holds and it reads notification lines and nothing else — not a
+ session, so nothing renewing in the background can collide with the page's
+ rotating refresh token, and a sign-out, which deletes the row, ends it. **When
+ a UnifiedPush distributor is already installed** (ntfy, or an application
+ carrying one), the row also gets its endpoint and P-256 key and each
+ notification is sent there at once as one RFC 8291 record, so the push server
+ relays bytes it cannot read and learns only *when* — the metadata §7.1 already
+ concedes to the hub; the fetch then runs every four hours as a net under it. A
+ distributor that refuses, disappears or answers 404/410 is not an error: the
+ row loses its endpoint and the phone fetches again. Both paths carry the same
+ payload — the hub's own row (kind, title, group, a `#/` route, its date),
+ never a message, which the hub does not hold — and a line pushed then fetched
+ is drawn once. **Nothing reaches a phone that was not created**, and
+ `create_notification` creates nothing for a muted group *or for an account
+ that turned every notification off* — the second switch used to be read only
+ by the interface, which hid rows the hub went on writing, and a switch only a
+ renderer honours is no switch once a phone is told about every row. A push
+ endpoint is a URL a member chose and the hub fetches it, so a send resolves
+ it, refuses any non-public address and connects to the address it checked
+ (SNI and Host carry the name); no redirect is followed. The shell draws a
+ pushed message only if the connector decrypted it with the phone's key, and a
+ notification's link is a route in the page, applied as `location.hash`, never
+ loaded. No VAPID yet: a distributor that requires it refuses, and the phone
+ fetches instead.
-Release builds are signed with the development key until the release key exists
-(Stage D12); §2.3's sentence about who holds a signing key applies to whichever
-store distributes them.
+Release builds are signed with the release key, distributed as a direct APK;
+the key stays outside the repository and a release build without it fails
+rather than falling back to the development key. §2.3's sentence about who
+holds a signing key applies to whichever store distributes them, if one ever
+does.
### 11.4 Casting
@@ -3804,6 +3938,7 @@ had already been asked.
| **AV30** | **What one member's offers cost a node is bounded per account and per node, and the bound admits the heaviest ordinary account** (§7.2). Each offer makes the node allocate a peer connection. A budget of 120 per node refilled at two a second bounds a member there without touching their other nodes, and it is counted by account because a mobile carrier shares one IPv4 address among many subscribers. Pending offers are capped at 32 per account. Both refusals carry `Retry-After` and the client retries them, because a refused offer otherwise reads as a node that is down. An offer carries at most 64 ICE candidates (32 KiB), and its IP-log row — kept a year — is written only once it goes to a node, so naming nodes that do not exist costs the hub nothing. A node's `update_groups`, a database read each, is budgeted like `chat_notify` (ten a minute) and claims at most 1000 groups |
| **AV32** | **A node hosts a group because its owner said so, not because its account belongs to it** (§7.2). Every member holds the key, so a member's node passes the handshake like the real host and could be the one a client keeps. A node may claim the groups its account owns and those whose owner approved it; any other claim is a pending request the owner sees |
| **AV33** | **Nobody is made a member without saying yes** (§7.3). A membership makes the account's client list the group, name it in its tokens and dial its nodes, so an owner's addition is an invitation until the invitee accepts it. The MNP token names only the group it is minted for |
+| **AV34** | **What a member's chat costs other members' phones is bounded twice, and what a phone's fetching costs the hub once** (§11.3). A pushed notification is one outbound request per subscription of the recipient, so the fan-out of one chat line is members × phones. It is bounded where it starts (`chat_notify`, ten a minute per node), a conversation reaches each phone at most once per 30 s (the phone shows one line per group, so the pushes in between would only replace it), and an account holds ten subscriptions at most. A subscription the push server reports gone (404/410) loses its endpoint rather than being retried for ever. A phone fetching instead costs one indexed read per poll, refused below a minute per row, and returns at most twenty lines |
| **AV31** | **What waits for a signature is bounded** (§5.4). Any authenticated member can ask for an admin challenge, since the signature is checked afterwards, and a pending challenge kept its whole request until answered — measured, 200 requests of 1 MiB held 400 MiB for the life of one connection. At most eight pending per connection, 64 KiB each, expired ones dropped |
### 13.6 Chat design findings
@@ -3996,10 +4131,11 @@ seeking, audio-language and subtitle selection, transfer leases with queueing,
pause and resume, the group-application framework with Chat, Files, Videos,
Music and Photos, cross-group search with source merging, per-account playlists,
casting to a Chromecast with subtitles rebased onto the relay's clock, the
+phone's photo backup (§9.12), the
operator CLI and loopback control API, the desktop client through its identity
and download stages, account recovery, the Windows port through packaging, and
-the Android client — the shell, native keys, downloads and uploads, and casting
-(§11.3).
+the Android client — the shell, native keys, downloads and uploads, casting,
+and notifications while it is closed, fetched or pushed (§11.3).
The packages install: a machine has been taken from the built artefacts to a
running hub and node on **Ubuntu 26.04 (`.deb`), Fedora 44 (`.rpm`) and
@@ -4027,7 +4163,9 @@ process runs it — `systemctl --user` on Linux, Task Scheduler on Windows.
| — | **Bitmap subtitles** (PGS, VOBSUB — about a fifth of the embedded streams). No WebVTT without OCR; they are not listed rather than listed and blank. Burn-in covers them and costs `-c:v copy`, which is what the eight-slot sizing assumes never happens |
| — | Delegation (§3.4) |
| — | Tier 3 roster attestation (§3.3) |
-| — | **Android: phone behaviour and release** — the back button driving the page, recovery from a network handover, keeping a download alive with the screen off, lock-screen media controls, and a release key (§11.3) |
+| — | **HEIC/HEIF in Photos.** The indexer does not classify `.heic`/`.heif` as images and no browser but Safari draws them, so a phone that shoots HEIC has its backed-up photos stored and listed in Files but absent from Photos. Needs `pillow-heif` in every package and a JPEG rendition for the viewer (§9.12) |
+| — | **Photo backup with the application closed**, and **"remove what this phone sent"** (§9.12). A closed application has no page and so no transport; it runs at the next opening |
+| — | **Android: phone behaviour and release** — the back button driving the page, recovery from a network handover, keeping a download alive with the screen off, lock-screen media controls, and an update channel (§11.3) |
| — | **Federation between two hubs.** The protocol is written and switched off in the code (§7.6); what is not built is one run between two machines |
### 15.3 Open, and why each is where it is