summaryrefslogtreecommitdiffstats
path: root/packages/meshbay-node/tests/test_login_retry_is_resilient.py
Commit message (Collapse)AuthorAgeFilesLines
* fix(node): a Windows daemon that stops properly, starts honestly and runs onceChristophe Besson40 hours1-2/+18
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Found by installing the builds and driving every startup mode live: - Stop through the node's own control API first (POST /api/shutdown, loopback and per-run token): the one channel that reaches a daemon in any session without elevation -- a service node runs in session 0 -- and the one that runs its shutdown. Then Task Scheduler, then a forced stop. Nine stops in a row used to log no shutdown at all: each was a TerminateProcess. - The forced stop spares the command running it. The frozen meshbay-node.exe is the daemon and every CLI verb, so `taskkill /IM meshbay-node.exe` killed `autostart stop` and `restart-daemon` themselves: exit 1, no output, and no node after a restart. It excludes its own pid and its parent's, and /T takes a venv launcher's python child and a daemon's ffmpeg children with it. - Start and restart report the version that answered, never "started" about a node nobody asked; `service start` says so when no node answered, and where the log is. - A second instance fails before it touches anything. The daemon wrote ui-token, then failed to bind inside uvicorn's task and exited with the reason on a hidden console; the node still running then refused every stop and status, its token file naming a dead process. The control port is now bound first (exclusively on Windows, where SO_REUSEADDR would share it), and a refusal is logged and exits 2. Linux had the same order. - The daemon logs to %LOCALAPPDATA%\meshbay\state\node.log: Task Scheduler discards its stderr. Only the daemon run opens it, never a CLI verb. - Hub sign-in waits are interruptible, a stop requested before the node is up is honoured, and a hub that answers 429 or restarts leaves the node in waiting_for_hub rather than looking dead. - operator_paired is null until the roster is read, instead of a false that showed "No operator paired" about a node whose pairing was intact. The node test conftest also points HOME, USERPROFILE, LOCALAPPDATA and APPDATA at a throwaway directory for every test, and keeps log_file() away from the developer's own node: redirecting HOME alone isolates nothing on Windows, and the CLI tests had been writing invite and pairing codes into the real profile. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(node,client): survive a transient hub state on login, and surface a ↵Christophe Besson4 days1-0/+109
failed node-key link Two defensive gaps turned a routine reset-and-reonboard into "impossible de démarrer le node": 1. daemon._login_with_retry retried a 401 (node key not linked yet) but `raise`d on every other status, so a 429 — the daemon's own 5s retries hitting the sign-in rate limit — or a 502/503 while the hub restarts during a deploy killed the process, and systemd crash-looped it. Those statuses (429, 5xx) are now retried with a back-off that respects Retry-After, so a freshly reset node stays alive (the operator needs it up to read its key) instead of dying. A genuine 4xx (400/422) still raises. 2. create-group's linkNodeKey swallowed every error as "already linked or same key" — but PUT /me/node_key is idempotent and returns 200 on a re-link, so there was no benign error to hide: the catch only ever hid a real failure (a rejected session, a bad key), letting the wizard proceed against a node that looked linked but was not, which then could not authenticate. The link failure now surfaces (detectNode shows it). test_login_retry_is_resilient.py holds the retry behaviour (429/5xx retried, Retry-After honoured, 401 stays alive, 400 still raises); red before, green after. common/node/hub suites green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>