feat(render-farm): chunked GPU Studio export service - #2356
Merged
Merged
Conversation
Adds a chunked Studio export service: a Bun coordinator plans GOP-aligned chunks and audio sections, fair-schedules them across exports and assembles a faststart MP4 from parts that GPU workers upload directly into one S3 multipart upload. Workers render with frames kept on the GPU (NVDEC, CUDA to Vulkan interop, NVENC), stream 2 s fMP4 segments per finished GOP so an HLS playlist is playable within seconds, and render Studio Sound audio on spare cores. Reliability: retries with backoff, progress-based hedging, engine and job watchdogs, graceful SIGTERM drain, worker restart detection, an S3 journal that resumes in-flight jobs after a coordinator restart, and fail-fast refusal of chunks whose engine fell back to software. Renderer changes (Linux render hosts only unless noted): CUDA-backed decoded frames and NV12 output, decoder read-ahead (CAP_DECODER_READAHEAD), a per-project layer pool with prebuilt spares, per-stage loop timings (CAP_RENDER_LOOP_STATS), and a blurred-background cache that applies to all renders.
…s safe - Every dispatch of a chunk uploads into its own part range and carries an attempt number; stale failures are ignored and requeues are idempotent, so two live copies can never write the same parts. - Workers report every held task (including reserved and upload phases); the coordinator re-attaches them after a restart and reclaims tasks a live worker stops reporting. Resumed jobs are not failed by the stall watchdog, and their HLS playlists are rebuilt from finished chunks. - Only silent predecessors are superseded, so workers sharing a hostname coexist; the first audio result for a section wins. - Worker promise handling cannot crash the process; draining aborts open polls. Manifests cannot escape the project directory or recording prefix, presigned URLs are capped at credential lifetime, and job requests reject inherited compression names. - Engine: always joins the encoder thread, fails chunks whose CUDA frames did not reach the compositor, destroys per-thread CUDA streams, pools shared buffers by power-of-two size class, clamps decoder read-ahead to the cache, resets reused layers' frame state and keeps held frames' storage identity so repeated frames are not re-uploaded.
richiemcilroy
marked this pull request as ready for review
September 25, 2026 23:11
Member
Author
|
hey @greptileai, please re-review the PR |
Member
Author
|
hey @greptileai, please re-review the PR |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ELI5
Before: a Studio export renders on one machine from start to finish, so a long or 4K export takes about as long to produce as the render itself, and nobody can watch it until the whole file exists.
Now: the export is cut into short chunks that GPU workers render at the same time. Each chunk streams playable video segments while it renders, so the export can be watched a few seconds after it is requested, and the finished MP4 is stitched together from the chunks without copying them again.
Why it matters: on three NVIDIA L4 GPUs a 2 hour 1080p export plays within about 4 seconds and finishes in about 2 minutes 15 seconds; a 20 minute 1080p export finishes in about 23 seconds. Workers keep every frame on the GPU (decode, compositing and encode), and the service recovers on its own from crashed or hung renders, restarted workers and a restarted coordinator.
What's in this PR
apps/render-farm(new, Bun): coordinator and worker services.moovas part 1.crates/render-farm(new, Rust): the engine process workers run.crates/rendering: Linux GPU-only frame path.CAP_DECODER_READAHEAD).CAP_LINUX_GPU_FRAMES/CAP_LINUX_HW_DECODE=cuda. Read-ahead is off unlessCAP_DECODER_READAHEADis set.crates/ffmpeg-hw-device,crates/audio,crates/export: small supporting changes (CUDA primary context, sample-buffer constructor, an added field default).Benchmarks
Measured on 3 workers (1× NVIDIA L4, 8 vCPU each, NVENC p1) with the production image and default settings. The full matrix ran at
1042c6e2a. "Playable" is time from request until the HLS playlist can start playback; "File" is time until the final MP4 is complete.5c25675b2and5d7e5bd3dmatch the table: 1080p 20 min 2.6 s / 23.6 s, 1080p 1 h 2.6 s / 66.8 s, 1080p 2 h 3.1 s / 134.1 s, and 4K 20 min 2.9 s / 48.9 s. None of those runs dispatched a duplicate chunk.ENDLIST, covered the full 1200 s, and every listed segment was fetchable.Reliability
Every fault below was injected into a 1080p 20 min export that takes about 23 s without faults. Each export completed with a verified file:
SIGKILL) mid-chunkSIGSTOP), 4 runsThe mechanisms behind this:
SIGKILL, so a stopped or wedged engine releases its slot.SIGTERMdrains a worker.RF_REQUIRE_GPU=1rejects any chunk that fell back to software decode or rendering.Recovery
ENDLISTlands.Hardening
RF_MAX_SOURCE_FILES,RF_MAX_SOURCE_BYTES,RF_MAX_EXPORT_SECONDS):moovsize and export duration are all capped.RF_SOURCE_KEY_PREFIXES, with path containment on materialised sources.Verification
c01767878): scheduler, planning, HLS segmenting and report validation, request and manifest validation, MP4/fMP4 box builders and source indexing, S3 XML decoding, and coordinator recovery. Typecheck and Biome also pass.cargo fmt --checkand clippy with-D warningspass for the touched crates. The Linux build andcargo test -p cap-render-farmpass in CI (sync-tests on Ubuntu) and in the production image build.Demo
Not applicable: server-side infrastructure with no UI. It was verified by the benchmarks and fault injection above instead of a screen recording.
Platform coverage and limits
Send + Syncexplicitly, so the GPUI editor build no longer overflows the trait-recursion limit.assets::tests::every_embedded_icon_is_referencedincap-gpui, which fails the same way on other open PRs.cap-utils(crates/utils/src/export_resources.rs), which this PR doesn't touch.Deployment notes
RF_TOKEN, andRF_SOURCE_KEY_PREFIXESset to the prefixes recordings live under.hls/andjobs/.--gpus all --init. The image defaults to the GPU-only path.The PR appears safe to merge based on this re-review; no new blocking issue was established.
Summary
This PR adds a GPU render-farm service for chunked Studio exports, progressive HLS playback, multipart MP4 assembly, and recovery after worker or coordinator failures. Since the previous review, it binds each segment report to the reporting dispatch’s part range and removes redundant field comments.
Reviews (3) · Last reviewed commit: "chore(render-farm): drop comments that r..."