Summary
Most of the daemon-side time for every warm request (tier-0 reuse path) goes to re-fingerprinting the whole workspace from scratch, before anything is scheduled. The cost grows with repo size, and part of the work is synchronous, so it blocks the event loop that serves all clients. After client startup (#6054) is fixed, this becomes the dominant fixed cost of a warm daemon request.
| workspace |
captureWorkspaceInputFingerprintAsync (warm median) |
captureProjectConfigurationFingerprintAsync (median) |
sync runtime _hashPaths |
| 30-project synthetic |
61 ms |
22 ms |
22-27 ms |
| rushstack (196 projects) |
263 ms |
372 ms |
23 ms |
Fully warm no-op latency vs graph size (synthetic): 1.26 s at 1 project, 2.84 s at 50, 8.67 s at 300. That is about 25 ms per project, bound by re-validation. A daemon CPU profile over 12 warm requests in a 30-project workspace shows:
- lstat 466 ms and realpath 263 ms (from
_hashPaths / hashFilesAsync);
- ajv schema compilation 530 ms (
JsonSchema.ensureCompiled via heft-config-file: uncached project config loads recompile the schemas);
LockFile.tryAcquire -> getProcessStartTime -> spawnSync('ps') 353 ms, about 30-54 ms per call. It runs on every request (acquireExecutionLeaseAsync), in warm-set maintenance, and in #prepareAsync.
As a result, the daemon event loop stalls for up to 190-260 ms (p99) while preparing warm requests, delaying every other client.
Repro steps
Call the same rush-lib functions that WorkspaceRequestLifecycle calls on the tier-0 path (#captureAsync at WorkspaceRequestLifecycle.ts:373 and captureProjectConfigurationFingerprintAsync at :384), with a persistent WorkspaceRuntimeFingerprintCache, 10 times against a 30-project and a 196-project workspace. Or profile the daemon (--cpu-prof) during repeated warm no-op rush-client build requests.
Expected result: Warm requests do incremental work proportional to what changed: stat-identity memoization of definition files, cached compiled schemas, and cached runtime fingerprints. Nothing synchronous or subprocess-based sits on the daemon event loop per request.
Actual result: Full recomputation on every request, with a synchronous file walk, schema recompiles, and a synchronous ps spawn.
Details
Root cause (main @ 60007c9):
libraries/rush-lib/src/api/WorkspaceInputFingerprint.ts:62-99 (_hashPaths): a synchronous recursive readdir plus statSync/realpathSync of the ~500+ runtime files on every capture.
:221-241 (hashFilesAsync): re-reads and re-hashes every definition file (rush.json, common/config/**, 4 files per project) with concurrency 3, with no stat memo.
:199-215 (captureProjectConfigurationFingerprintAsync): uncached project configuration loads recompile ajv schemas.
LockFile getProcessStartTime spawns ps synchronously. On Linux, read /proc/<pid>/stat instead.
Suggested fix: memoize by file identity (dev/ino/size/mtime/ctime), making hashes a cheap stat check; cache compiled schemas process-wide; compute the runtime fingerprint once per process, since runtime files cannot change without a restart; replace the synchronous ps with procfs; and move the remaining synchronous I/O off the event loop.
This was found during an automated performance/behavior analysis of rush-client/rushd on Linux and independently confirmed.
Standard questions
| Question |
Answer |
@microsoft/rush globally installed version? |
built from main @ 60007c9 (5.179.0) |
rushVersion from rush.json? |
5.179.0 |
pnpmVersion, npmVersion, or yarnVersion from rush.json? |
pnpm@10.27.0 |
(if pnpm) useWorkspaces from pnpm-config.json? |
true |
| Operating system? |
Linux (WSL2 Ubuntu 24.04) |
| Would you consider contributing a PR? |
Yes |
Node.js version (node -v)? |
22.23.2 |
Summary
Most of the daemon-side time for every warm request (tier-0 reuse path) goes to re-fingerprinting the whole workspace from scratch, before anything is scheduled. The cost grows with repo size, and part of the work is synchronous, so it blocks the event loop that serves all clients. After client startup (#6054) is fixed, this becomes the dominant fixed cost of a warm daemon request.
captureWorkspaceInputFingerprintAsync(warm median)captureProjectConfigurationFingerprintAsync(median)_hashPathsFully warm no-op latency vs graph size (synthetic): 1.26 s at 1 project, 2.84 s at 50, 8.67 s at 300. That is about 25 ms per project, bound by re-validation. A daemon CPU profile over 12 warm requests in a 30-project workspace shows:
_hashPaths/hashFilesAsync);JsonSchema.ensureCompiledvia heft-config-file: uncached project config loads recompile the schemas);LockFile.tryAcquire -> getProcessStartTime -> spawnSync('ps')353 ms, about 30-54 ms per call. It runs on every request (acquireExecutionLeaseAsync), in warm-set maintenance, and in#prepareAsync.As a result, the daemon event loop stalls for up to 190-260 ms (p99) while preparing warm requests, delaying every other client.
Repro steps
Call the same rush-lib functions that
WorkspaceRequestLifecyclecalls on the tier-0 path (#captureAsyncatWorkspaceRequestLifecycle.ts:373andcaptureProjectConfigurationFingerprintAsyncat:384), with a persistentWorkspaceRuntimeFingerprintCache, 10 times against a 30-project and a 196-project workspace. Or profile the daemon (--cpu-prof) during repeated warm no-oprush-client buildrequests.Expected result: Warm requests do incremental work proportional to what changed: stat-identity memoization of definition files, cached compiled schemas, and cached runtime fingerprints. Nothing synchronous or subprocess-based sits on the daemon event loop per request.
Actual result: Full recomputation on every request, with a synchronous file walk, schema recompiles, and a synchronous
psspawn.Details
Root cause (main @ 60007c9):
libraries/rush-lib/src/api/WorkspaceInputFingerprint.ts:62-99(_hashPaths): a synchronous recursive readdir plusstatSync/realpathSyncof the ~500+ runtime files on every capture.:221-241(hashFilesAsync): re-reads and re-hashes every definition file (rush.json, common/config/**, 4 files per project) with concurrency 3, with no stat memo.:199-215(captureProjectConfigurationFingerprintAsync): uncached project configuration loads recompile ajv schemas.LockFilegetProcessStartTimespawnspssynchronously. On Linux, read/proc/<pid>/statinstead.Suggested fix: memoize by file identity (dev/ino/size/mtime/ctime), making hashes a cheap stat check; cache compiled schemas process-wide; compute the runtime fingerprint once per process, since runtime files cannot change without a restart; replace the synchronous
pswith procfs; and move the remaining synchronous I/O off the event loop.This was found during an automated performance/behavior analysis of
rush-client/rushdon Linux and independently confirmed.Standard questions
@microsoft/rushglobally installed version?main@ 60007c9 (5.179.0)rushVersionfrom rush.json?pnpmVersion,npmVersion, oryarnVersionfrom rush.json?useWorkspacesfrom pnpm-config.json?node -v)?