Skip to content

feat(render-farm): chunked GPU Studio export service - #2356

Merged
richiemcilroy merged 15 commits into
mainfrom
building/render-farm-83a16bf3
Sep 26, 2026
Merged

richiemcilroy merged 15 commits into
mainfrom
building/render-farm-83a16bf3

Conversation

@richiemcilroy

@richiemcilroy richiemcilroy commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

ELI5

Before: a Studio export renders on one machine from start to finish, so a long or 4K export takes about as long to produce as the render itself, and nobody can watch it until the whole file exists.

Now: the export is cut into short chunks that GPU workers render at the same time. Each chunk streams playable video segments while it renders, so the export can be watched a few seconds after it is requested, and the finished MP4 is stitched together from the chunks without copying them again.

Why it matters: on three NVIDIA L4 GPUs a 2 hour 1080p export plays within about 4 seconds and finishes in about 2 minutes 15 seconds; a 20 minute 1080p export finishes in about 23 seconds. Workers keep every frame on the GPU (decode, compositing and encode), and the service recovers on its own from crashed or hung renders, restarted workers and a restarted coordinator.

What's in this PR

  • apps/render-farm (new, Bun): coordinator and worker services.
    • The coordinator plans chunks by render work, with a short lead-in chunk so the first segment plays quickly. It schedules them fairly across exports: playback-gating chunks first, then the export holding the fewest slots.
    • Workers fetch only the byte ranges each chunk needs and upload their bytes straight into one S3 multipart upload; the coordinator writes the faststart moov as part 1.
    • HLS output: 2 s fMP4 segments, published as each GOP finishes, under an EVENT playlist.
    • A versioned S3 journal lets a restarted coordinator resume exports: request receipt, plan, part-range reservations for retries and hedges, and accepted results. Details are under "Recovery" below.
    • Includes a Dockerfile and a README with the configuration reference.
  • crates/render-farm (new, Rust): the engine process workers run.
    • Operations: probe, video chunk, audio section, warm.
    • Keeps the wgpu device, the project and the rendering layers alive between chunks.
    • Streams GOP and extradata events back to the worker.
    • Fetches both source clips for chunks that fall inside a transition.
  • crates/rendering: Linux GPU-only frame path.
    • NVDEC CUDA frames are imported into Vulkan through external memory, composited in wgpu, and handed to NVENC as CUDA frames, with no CPU round trip.
    • Partial Vulkan/CUDA allocations unwind on failure.
    • Decoder read-ahead (CAP_DECODER_READAHEAD).
    • Pooled rendering layers, and a blur cache for static backgrounds.
    • The GPU path and layer pooling are gated on Linux plus CAP_LINUX_GPU_FRAMES / CAP_LINUX_HW_DECODE=cuda. Read-ahead is off unless CAP_DECODER_READAHEAD is set.
    • The blur result cache is off unless a process enables it, and only the render farm engine does. On the farm, a static image, colour or gradient background with blur is blurred once, then reused until the background, blur amount or output size changes. The desktop editor and exports keep blurring every frame.
  • crates/ffmpeg-hw-device, crates/audio, crates/export: small supporting changes (CUDA primary context, sample-buffer constructor, an added field default).
  • CI: render farm TypeScript tests and typecheck in the Typecheck job; engine unit tests in sync-tests.

Benchmarks

Measured on 3 workers (1× NVIDIA L4, 8 vCPU each, NVENC p1) with the production image and default settings. The full matrix ran at 1042c6e2a. "Playable" is time from request until the HLS playlist can start playback; "File" is time until the final MP4 is complete.

Export Playable File
1080p, 1 min 2.9 s 3.7 s
1080p, 5 min 2.6 s 7.6 s
1080p, 20 min 2.8 s 23.3 s
1080p, 1 h 2.8 s 67.9 s
1080p, 2 h 3.5 s 135.2 s
4K, 1 min 2.9 s 5.5 s
4K, 5 min 3.2 s 15.2 s
4K, 20 min 3.0 s 50.7 s
4K, 1 h 3.2 s 136.3 s
4K, 2 h 4.1 s 265.7 s
  • Later commits: they change validation and hedging only. Spot checks at 5c25675b2 and 5d7e5bd3d match the table: 1080p 20 min 2.6 s / 23.6 s, 1080p 1 h 2.6 s / 66.8 s, 1080p 2 h 3.1 s / 134.1 s, and 4K 20 min 2.9 s / 48.9 s. None of those runs dispatched a duplicate chunk.
  • Durable dispatch writes: writing a reservation before every dispatch had cost 0.6-1.2 s of time to first playable segment on long exports. Here the journaled plan reserves each chunk's first range, so only retries and hedges wait on a write.
  • HLS coverage: the playlist reached every video frame and audio packet of the output. After worker and coordinator faults, the playlists ended with ENDLIST, covered the full 1200 s, and every listed segment was fetchable.
  • Where the time goes: 4K is NVENC-bound (encoder at 100%). 1080p runs at about 79% SM and 75% NVDEC.

Reliability

Every fault below was injected into a 1080p 20 min export that takes about 23 s without faults. Each export completed with a verified file:

Fault Completed in
Engine killed (SIGKILL) mid-chunk 26.0-26.7 s
Engine hung (SIGSTOP), 4 runs 24.9-26.1 s
Worker app restarted 29.9-32.3 s
Worker container stopped (drains, then exits) 27.1-31.5 s
Coordinator restarted mid-export 31.6-32.3 s
Coordinator restarted during planning 26.6-29.3 s

The mechanisms behind this:

  • Retries and hedging:
    • Failed chunks retry with backoff.
    • Slow chunks are hedged, and so is a copy whose frame count stops advancing for 5 s. Before that rule, an engine hung near the end of its chunk waited for the 30 s watchdog.
    • The worker watchdog and hedge-loser cancellation use SIGKILL, so a stopped or wedged engine releases its slot.
  • Stale work:
    • Every dispatch has its own attempt id and multipart part range, so late or stale results can't collide.
    • Heartbeats list every held task; tasks the coordinator can't account for are requeued.
  • Shutdown: SIGTERM drains a worker.
  • Restart and GPU health:
    • Journal resume re-attaches chunks that are still rendering.
    • RF_REQUIRE_GPU=1 rejects any chunk that fell back to software decode or rendering.
    • Engines are recycled when VRAM runs high.

Recovery

  • Journal writes:
    • A request receipt is persisted before a job is acknowledged, and the plan before scheduling.
    • Retries and hedges persist a part-range reservation before dispatch.
    • Attempts, hedge ownership and range high-water marks are restored after a restart. Six ranges cover five attempts plus one hedge; exhaustion fails the job instead of reusing live parts.
  • Accepted results:
    • Video metadata and audio packets are persisted before they are acknowledged or consumed downstream.
    • Competing reports coalesce, and completed work stays out of requeue and scheduling.
  • Multipart completion: an already completed MP4 is recognised by its size and header hash, including when the completion response was lost.
  • HLS playlist: failed playlist writes retry without advancing the cursor, and recovery state is kept until ENDLIST lands.
  • Probe engine: a crashed or hung probe engine is recreated, with a bounded retry.

Hardening

  • Auth and requests:
    • A bearer token is required (constant-time compare).
    • Request validation, and a cap on active jobs.
  • Recording manifests (bounded by RF_MAX_SOURCE_FILES, RF_MAX_SOURCE_BYTES, RF_MAX_EXPORT_SECONDS):
    • Manifest and metadata files are bounded when read.
    • File count, total bytes, sidecar sizes, moov size and export duration are all capped.
    • Keys are restricted to RF_SOURCE_KEY_PREFIXES, with path containment on materialised sources.
  • Source MP4 indexing: sample tables must fit inside their boxes, the sample count is capped, and sample-to-chunk runs are clamped to the chunk table.
  • HLS segments:
    • A segment report is listed only when its key is one the reporting dispatch wrote for that chunk.
    • Each dispatch writes its own segment objects, so a late hedge never replaces listed bytes.
  • Presigned URLs:
    • Lifetimes are bounded by the credential lifetime.
    • Job reads return a freshly signed playlist URL.

Verification

  • Unit tests (91 Bun tests, at c01767878): scheduler, planning, HLS segmenting and report validation, request and manifest validation, MP4/fMP4 box builders and source indexing, S3 XML decoding, and coordinator recovery. Typecheck and Biome also pass.
  • Rust: cargo fmt --check and clippy with -D warnings pass for the touched crates. The Linux build and cargo test -p cap-render-farm pass in CI (sync-tests on Ubuntu) and in the production image build.
  • Hardware: all benchmarks and faults above ran on AWS g6.2xlarge (NVIDIA L4).

Demo

Not applicable: server-side infrastructure with no UI. It was verified by the benchmarks and fault injection above instead of a screen recording.

Platform coverage and limits

  • Verified on: Linux with NVIDIA L4 (driver via the NVIDIA container toolkit).
  • Other GPUs: other NVIDIA GPUs should work; non-NVIDIA Linux GPUs aren't supported by the GPU-only path.
  • Desktop apps:
    • The shared rendering changes are gated off. The desktop editor and exports run the same code paths as before.
    • The Linux shared-buffer types declare Send + Sync explicitly, so the GPUI editor build no longer overflows the trait-recursion limit.
    • The desktop build jobs fail only on assets::tests::every_embedded_icon_is_referenced in cap-gpui, which fails the same way on other open PRs.
  • Windows clippy: the job fails on existing dead code in cap-utils (crates/utils/src/export_resources.rs), which this PR doesn't touch.
  • Presigned URLs: segment URLs inside a playlist are signed for 6 h, which covers the 4 h default export cap.

Deployment notes

  • Coordinator: needs a bucket-scoped role or keys, a shared RF_TOKEN, and RF_SOURCE_KEY_PREFIXES set to the prefixes recordings live under.
  • Bucket: add a lifecycle rule that aborts incomplete multipart uploads and expires hls/ and jobs/.
  • Workers: run with --gpus all --init. The image defaults to the GPU-only path.
  • Journal upgrade: this coordinator refuses the earlier unversioned journal and aborts unfinished legacy uploads. Drain active exports before replacing an older coordinator. Nothing runs the older coordinator today.
  • Not wired in yet: this PR adds the service only. Nothing in the web app or desktop app calls it yet.

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on this re-review; no new blocking issue was established.

Summary

This PR adds a GPU render-farm service for chunked Studio exports, progressive HLS playback, multipart MP4 assembly, and recovery after worker or coordinator failures. Since the previous review, it binds each segment report to the reporting dispatch’s part range and removes redundant field comments.

  • The latest segment-validation change addresses the earlier cross-copy key finding.
  • All previous Greptile threads are resolved; no new actionable finding was established.

Reviews (3) · Last reviewed commit: "chore(render-farm): drop comments that r..."

Adds a chunked Studio export service: a Bun coordinator plans GOP-aligned
chunks and audio sections, fair-schedules them across exports and assembles a
faststart MP4 from parts that GPU workers upload directly into one S3
multipart upload. Workers render with frames kept on the GPU (NVDEC, CUDA to
Vulkan interop, NVENC), stream 2 s fMP4 segments per finished GOP so an HLS
playlist is playable within seconds, and render Studio Sound audio on spare
cores.

Reliability: retries with backoff, progress-based hedging, engine and job
watchdogs, graceful SIGTERM drain, worker restart detection, an S3 journal
that resumes in-flight jobs after a coordinator restart, and fail-fast
refusal of chunks whose engine fell back to software.

Renderer changes (Linux render hosts only unless noted): CUDA-backed decoded
frames and NV12 output, decoder read-ahead (CAP_DECODER_READAHEAD), a
per-project layer pool with prebuilt spares, per-stage loop timings
(CAP_RENDER_LOOP_STATS), and a blurred-background cache that applies to all
renders.
Comment thread apps/render-farm/src/s3.ts Fixed
…s safe

- Every dispatch of a chunk uploads into its own part range and carries an
  attempt number; stale failures are ignored and requeues are idempotent, so
  two live copies can never write the same parts.
- Workers report every held task (including reserved and upload phases);
  the coordinator re-attaches them after a restart and reclaims tasks a live
  worker stops reporting. Resumed jobs are not failed by the stall watchdog,
  and their HLS playlists are rebuilt from finished chunks.
- Only silent predecessors are superseded, so workers sharing a hostname
  coexist; the first audio result for a section wins.
- Worker promise handling cannot crash the process; draining aborts open
  polls. Manifests cannot escape the project directory or recording prefix,
  presigned URLs are capped at credential lifetime, and job requests reject
  inherited compression names.
- Engine: always joins the encoder thread, fails chunks whose CUDA frames
  did not reach the compositor, destroys per-thread CUDA streams, pools
  shared buffers by power-of-two size class, clamps decoder read-ahead to
  the cache, resets reused layers' frame state and keeps held frames' storage
  identity so repeated frames are not re-uploaded.
@richiemcilroy richiemcilroy changed the title feat(render-farm): parallel GPU export farm with progressive HLS output feat(render-farm): chunked GPU Studio export service Sep 25, 2026
@richiemcilroy
richiemcilroy marked this pull request as ready for review September 25, 2026 23:11

@superagent-security superagent-security Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Superagent found 1 security concern(s).

Comment thread apps/render-farm/src/coordinator.ts
Comment thread apps/render-farm/src/coordinator.ts
Comment thread apps/render-farm/src/mp4.ts
Comment thread apps/render-farm/src/worker.ts Outdated
Comment thread apps/render-farm/src/coordinator.ts
Comment thread apps/render-farm/src/validate.ts Outdated

@superagent-security superagent-security Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Superagent found 2 security concern(s).

Comment thread apps/render-farm/src/mp4.ts
@richiemcilroy

Copy link
Copy Markdown
Member Author

hey @greptileai, please re-review the PR

Comment thread apps/render-farm/src/hls.ts Outdated
Comment thread apps/render-farm/src/validate.ts Outdated
@richiemcilroy

Copy link
Copy Markdown
Member Author

hey @greptileai, please re-review the PR

@richiemcilroy
richiemcilroy merged commit 8f4a360 into main Sep 26, 2026
29 of 33 checks passed

This branch was successfully deployed

1 active deployment
Preview — c0176787 Deployed Sep 26, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants