Summary
Three robustness problems in the daemon transport, each reproduced on Linux with rush-client/rushd from main @ 60007c9:
- Socket paths longer than
sun_path (108 bytes) are silently truncated, so different workspaces share one daemon. resolveDaemonPaths (libraries/rush-daemon-transport/src/DaemonPaths.ts:61-67) builds <base>/rushd-<uid>/rushd-<32hex>.sock (the base plus 55 bytes) and never checks the result against the limit. Node/libuv then truncate the path on listen()/connect(). The part that gets cut off is the workspace key: with a 77-character base the socket is named rushd-44c88aeea54e6, and with 90 characters it is just rushd-. So with a long XDG_RUNTIME_DIR or TMPDIR (deep CI, sandbox, Nix or Bazel-style directories), workspace B talks to workspace A's daemon: daemon status in ws2 reports ws1's PID, build fails with Cannot attest the restarting daemon ownership, and daemon stop in ws1 can stop the wrong daemon. Expected: detect overlong paths and switch to a shorter derived location (for example a hashed path under a short base), or fail with a clear error. Never truncate silently.
- The reclaim mutex's steal path is not exclusive. On EEXIST with a dead
mutexPid, tryAcquireReclaimLock (DaemonReclaimLock.ts:73-86) calls rewriteMutexEntry (:47-64), which does unlinkSync(lock) and then writeFileSync(lock, { flag: 'wx' }). Two stealers can both see "dead holder"; B then unlinks A's fresh file and creates its own, so both acquire. In 10 trials with 16 concurrent stealers, every trial had more than one holder (2-6). In addition, readMutexPid treats a just-created, still-empty file as dead, and the finally in DaemonReclaim.ts:46-52 unlinks .reclaim even when another process owns it. Expected: steal atomically (for example rename a uniquely named file over the lock and re-verify ownership after writing), and release only a lock you still own.
- Deleting the runtime dir or socket leaves an unreachable live daemon running alongside a new one. systemd-logind removes
/run/user/<uid> at logout, and tmp cleaners can remove it too. The old daemon keeps running, unreachable, until its idle timeout (900 s by default); the next client auto-starts a second daemon for the same workspace; and daemon stop only reaches the new one. Expected: after binding, the host periodically re-validates that lstat(socketPath) still matches the bound inode and that the lockfile still names its own PID, and shuts down (or re-publishes its endpoint) when it doesn't.
Repro steps
- Start the daemon with
XDG_RUNTIME_DIR=<90-char dir> in two different workspaces, then run daemon status in the second: it reports the first workspace's PID.
- Seed
<key>.pid.json.reclaim with {"mutexPid": <dead pid>} and have 16 forked processes call tryAcquireReclaimLock at the same moment: more than one acquires.
rush-client daemon start && rush-client build, then rm -rf $XDG_RUNTIME_DIR/rushd-<uid>, then rush-client build: two daemon processes now exist for the workspace.
Expected result: As described for each item above.
Actual result: As described for each item above.
Details
This was found during an automated performance/behavior analysis of rush-client/rushd on Linux. Each item was confirmed independently (item 1 twice).
Standard questions
| Question |
Answer |
@microsoft/rush globally installed version? |
built from main @ 60007c9 (5.179.0) |
rushVersion from rush.json? |
5.179.0 |
pnpmVersion, npmVersion, or yarnVersion from rush.json? |
pnpm@10.27.0 |
(if pnpm) useWorkspaces from pnpm-config.json? |
true |
| Operating system? |
Linux (WSL2 Ubuntu 24.04) |
| Would you consider contributing a PR? |
Yes |
Node.js version (node -v)? |
22.23.2 |
Summary
Three robustness problems in the daemon transport, each reproduced on Linux with
rush-client/rushdfrommain@ 60007c9:sun_path(108 bytes) are silently truncated, so different workspaces share one daemon.resolveDaemonPaths(libraries/rush-daemon-transport/src/DaemonPaths.ts:61-67) builds<base>/rushd-<uid>/rushd-<32hex>.sock(the base plus 55 bytes) and never checks the result against the limit. Node/libuv then truncate the path onlisten()/connect(). The part that gets cut off is the workspace key: with a 77-character base the socket is namedrushd-44c88aeea54e6, and with 90 characters it is justrushd-. So with a longXDG_RUNTIME_DIRorTMPDIR(deep CI, sandbox, Nix or Bazel-style directories), workspace B talks to workspace A's daemon:daemon statusin ws2 reports ws1's PID,buildfails withCannot attest the restarting daemon ownership, anddaemon stopin ws1 can stop the wrong daemon. Expected: detect overlong paths and switch to a shorter derived location (for example a hashed path under a short base), or fail with a clear error. Never truncate silently.mutexPid,tryAcquireReclaimLock(DaemonReclaimLock.ts:73-86) callsrewriteMutexEntry(:47-64), which doesunlinkSync(lock)and thenwriteFileSync(lock, { flag: 'wx' }). Two stealers can both see "dead holder"; B then unlinks A's fresh file and creates its own, so both acquire. In 10 trials with 16 concurrent stealers, every trial had more than one holder (2-6). In addition,readMutexPidtreats a just-created, still-empty file as dead, and thefinallyinDaemonReclaim.ts:46-52unlinks.reclaimeven when another process owns it. Expected: steal atomically (for example rename a uniquely named file over the lock and re-verify ownership after writing), and release only a lock you still own./run/user/<uid>at logout, and tmp cleaners can remove it too. The old daemon keeps running, unreachable, until its idle timeout (900 s by default); the next client auto-starts a second daemon for the same workspace; anddaemon stoponly reaches the new one. Expected: after binding, the host periodically re-validates thatlstat(socketPath)still matches the bound inode and that the lockfile still names its own PID, and shuts down (or re-publishes its endpoint) when it doesn't.Repro steps
XDG_RUNTIME_DIR=<90-char dir>in two different workspaces, then rundaemon statusin the second: it reports the first workspace's PID.<key>.pid.json.reclaimwith{"mutexPid": <dead pid>}and have 16 forked processes calltryAcquireReclaimLockat the same moment: more than one acquires.rush-client daemon start && rush-client build, thenrm -rf $XDG_RUNTIME_DIR/rushd-<uid>, thenrush-client build: two daemon processes now exist for the workspace.Expected result: As described for each item above.
Actual result: As described for each item above.
Details
This was found during an automated performance/behavior analysis of
rush-client/rushdon Linux. Each item was confirmed independently (item 1 twice).Standard questions
@microsoft/rushglobally installed version?main@ 60007c9 (5.179.0)rushVersionfrom rush.json?pnpmVersion,npmVersion, oryarnVersionfrom rush.json?useWorkspacesfrom pnpm-config.json?node -v)?