From 7662484cae8e74b7d9aa383bd6cd0dad4690aadc Mon Sep 17 00:00:00 2001 From: Christophe Besson Date: Sun, 27 Sep 2026 22:20:35 +0200 Subject: fix(node): a Windows daemon that stops properly, starts honestly and runs once Found by installing the builds and driving every startup mode live: - Stop through the node's own control API first (POST /api/shutdown, loopback and per-run token): the one channel that reaches a daemon in any session without elevation -- a service node runs in session 0 -- and the one that runs its shutdown. Then Task Scheduler, then a forced stop. Nine stops in a row used to log no shutdown at all: each was a TerminateProcess. - The forced stop spares the command running it. The frozen meshbay-node.exe is the daemon and every CLI verb, so `taskkill /IM meshbay-node.exe` killed `autostart stop` and `restart-daemon` themselves: exit 1, no output, and no node after a restart. It excludes its own pid and its parent's, and /T takes a venv launcher's python child and a daemon's ffmpeg children with it. - Start and restart report the version that answered, never "started" about a node nobody asked; `service start` says so when no node answered, and where the log is. - A second instance fails before it touches anything. The daemon wrote ui-token, then failed to bind inside uvicorn's task and exited with the reason on a hidden console; the node still running then refused every stop and status, its token file naming a dead process. The control port is now bound first (exclusively on Windows, where SO_REUSEADDR would share it), and a refusal is logged and exits 2. Linux had the same order. - The daemon logs to %LOCALAPPDATA%\meshbay\state\node.log: Task Scheduler discards its stderr. Only the daemon run opens it, never a CLI verb. - Hub sign-in waits are interruptible, a stop requested before the node is up is honoured, and a hub that answers 429 or restarts leaves the node in waiting_for_hub rather than looking dead. - operator_paired is null until the roster is read, instead of a false that showed "No operator paired" about a node whose pairing was intact. The node test conftest also points HOME, USERPROFILE, LOCALAPPDATA and APPDATA at a throwaway directory for every test, and keeps log_file() away from the developer's own node: redirecting HOME alone isolates nothing on Windows, and the CLI tests had been writing invite and pairing codes into the real profile. Co-Authored-By: Claude Opus 5.5 --- packages/meshbay-node/src/meshbay_node/ops/groups.py | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) (limited to 'packages/meshbay-node/src/meshbay_node/ops') diff --git a/packages/meshbay-node/src/meshbay_node/ops/groups.py b/packages/meshbay-node/src/meshbay_node/ops/groups.py index 905aa44..1d8003c 100644 --- a/packages/meshbay-node/src/meshbay_node/ops/groups.py +++ b/packages/meshbay-node/src/meshbay_node/ops/groups.py @@ -122,7 +122,10 @@ async def list_groups(state: dict) -> dict: "peers": sum(1 for p in peers.values() if p.get("group_id") == gid), }) roster = state.get("roster") - has_operator = False + # None, not False, until the roster is published (after the hub login): a + # node still waiting for its account has not lost its operator, and saying + # "no operator paired" sent the reader after a pairing that was intact. + has_operator = None if roster: members = await roster.list_members() has_operator = any(m["role"] == "operator" and m["status"] == "active" -- cgit v1.2.3