| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Found by installing the builds and driving every startup mode live:
- Stop through the node's own control API first (POST /api/shutdown, loopback
and per-run token): the one channel that reaches a daemon in any session
without elevation -- a service node runs in session 0 -- and the one that
runs its shutdown. Then Task Scheduler, then a forced stop. Nine stops in a
row used to log no shutdown at all: each was a TerminateProcess.
- The forced stop spares the command running it. The frozen meshbay-node.exe
is the daemon and every CLI verb, so `taskkill /IM meshbay-node.exe` killed
`autostart stop` and `restart-daemon` themselves: exit 1, no output, and no
node after a restart. It excludes its own pid and its parent's, and /T takes
a venv launcher's python child and a daemon's ffmpeg children with it.
- Start and restart report the version that answered, never "started" about a
node nobody asked; `service start` says so when no node answered, and where
the log is.
- A second instance fails before it touches anything. The daemon wrote
ui-token, then failed to bind inside uvicorn's task and exited with the
reason on a hidden console; the node still running then refused every stop
and status, its token file naming a dead process. The control port is now
bound first (exclusively on Windows, where SO_REUSEADDR would share it), and
a refusal is logged and exits 2. Linux had the same order.
- The daemon logs to %LOCALAPPDATA%\meshbay\state\node.log: Task Scheduler
discards its stderr. Only the daemon run opens it, never a CLI verb.
- Hub sign-in waits are interruptible, a stop requested before the node is up
is honoured, and a hub that answers 429 or restarts leaves the node in
waiting_for_hub rather than looking dead.
- operator_paired is null until the roster is read, instead of a false that
showed "No operator paired" about a node whose pairing was intact.
The node test conftest also points HOME, USERPROFILE, LOCALAPPDATA and APPDATA
at a throwaway directory for every test, and keeps log_file() away from the
developer's own node: redirecting HOME alone isolates nothing on Windows, and
the CLI tests had been writing invite and pairing codes into the real profile.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
|
|
failed node-key link
Two defensive gaps turned a routine reset-and-reonboard into "impossible de
démarrer le node":
1. daemon._login_with_retry retried a 401 (node key not linked yet) but `raise`d
on every other status, so a 429 — the daemon's own 5s retries hitting the
sign-in rate limit — or a 502/503 while the hub restarts during a deploy
killed the process, and systemd crash-looped it. Those statuses (429, 5xx)
are now retried with a back-off that respects Retry-After, so a freshly
reset node stays alive (the operator needs it up to read its key) instead of
dying. A genuine 4xx (400/422) still raises.
2. create-group's linkNodeKey swallowed every error as "already linked or same
key" — but PUT /me/node_key is idempotent and returns 200 on a re-link, so
there was no benign error to hide: the catch only ever hid a real failure
(a rejected session, a bad key), letting the wizard proceed against a node
that looked linked but was not, which then could not authenticate. The link
failure now surfaces (detectNode shows it).
test_login_retry_is_resilient.py holds the retry behaviour (429/5xx retried,
Retry-After honoured, 401 stays alive, 400 still raises); red before, green
after. common/node/hub suites green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|