Skip to content

2.5 (a): worker mode in the core — process() on a thread of its own - #29

Merged
tap merged 5 commits into
development/v2from
v2/worker-core
Oct 1, 2026
Merged

tap merged 5 commits into
development/v2from
v2/worker-core

Conversation

@tap

@tap tap commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Stacked on #28 (the design). Only the top commit, 65a0694, is new here. Merge #28 first.

Step (a) of plan 2.5, following the decided design.

What changes

core/include/tap/python/worker.h runs a processor's process() on a worker thread, a fixed number of vectors behind the audio thread.

  • The audio thread only copies vectors in and out of a lock-free ring. It never takes the GIL, allocates, locks or prints. It wakes the worker with a C++20 atomic notify, which never blocks.
  • The worker takes the slots in order and runs the existing processor::process(). Per-sample and per-vector classes work unchanged, with 2.4's channel matching.
  • Underruns: if a vector's outputs aren't ready when they're due, that vector is output as silence. The worker still processes it and catches up, so the latency stays fixed and the class still sees every vector in order.
  • A whole ring behind (the latency plus 0.25 s, at least 16 vectors): new inputs are dropped until the worker catches up, so the class sees a gap. This corrects the design text, which said "oldest". Dropping new ones is what a single-producer, single-consumer ring can do without racing the worker for a slot it's using; the plan is updated.
  • Reports: both kinds are reported once per load, from the main thread. The processor gains load_count() so the worker can tell when a load has happened.
  • Starting and stopping: the audio thread borrows the ring for each vector by taking an atomic pointer. start() and stop() therefore wait at most one vector, even while a host's old signal chain is still running, and a vector that arrives meanwhile is silence.
  • Threads: the core's CMake target now links Threads.

Tests (test_worker.cpp, fixture stalls.py)

  • Latency: worker output equals direct output L vectors later, per sample and per vector, L = 1, 2, 3, with two channels.
  • Channel matching: a missing host input reads as silence, an extra output is silent, and short vectors work.
  • Stalls: a 0.2 s stall goes silent and is reported once; after it, the output is back at the same latency. A second stall is not reported again; after a reload, a stall is reported again.
  • Drops: a stall that runs past the ring drops exactly the 12 vectors beyond it. Their outputs are silent and aren't counted as late.
  • Lifecycle: restart with a new vector size and latency, stop, and stop during a stall (it waits for the vector in progress).
  • Reloads: ten reloads while an audio thread runs at real-time pace; the last reload is what plays.

Checked locally (macOS)

  • Core battery 83/83.
  • The worker scenarios passed 15 runs in a row.
  • ThreadSanitizer: clean on the worker scenarios and the existing threading and reload ones.
  • clang-format and clang-tidy-18 pass.

Step (b), the Max object (@mode, @latency, @latencysamples, and a runtime test in Max), is next.

🤖 Generated with Claude Code

https://claude.ai/code/session_019rVV9SkjkWmiBXFMT4whkz

tap and others added 5 commits September 30, 2026 17:46
Python on a thread of its own, a fixed latency of whole vectors behind the audio thread, which
never takes the GIL: a lock-free ring of slots in the core, the existing processor::process() run
on the worker, underruns output as silence with the latency kept, @mode and @Latency applied at
the next chain compile, and a read-only latency attribute (Max documents no call to report
latency to the host).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019rVV9SkjkWmiBXFMT4whkz
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019rVV9SkjkWmiBXFMT4whkz
@Latency sets it in vectors; the read-only @latencysamples reports it in samples.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019rVV9SkjkWmiBXFMT4whkz
worker.h runs a processor's process() on a worker thread a fixed number of vectors behind the
audio thread, which only copies vectors through a lock-free ring and never takes the GIL. A vector
whose outputs are not ready in time is output as silence; the worker still processes it and
catches up, so the latency stays fixed and the class sees every vector in order. A worker a whole
ring behind drops new inputs until it catches up. Both are reported once per load from the main
thread (the processor gains load_count() for that). The audio thread borrows the ring by taking an
atomic pointer, so start() and stop() can run while a host's old signal chain still does.

Core battery: the output is the direct output L vectors later, per sample and per vector; channel
matching; underruns, drops and their reports; restart and stop; reloads under an audio thread.
Clean under ThreadSanitizer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019rVV9SkjkWmiBXFMT4whkz
#25–#28 landed on development/v2 (#28 as a squash of this branch's first three commits), so the
plan conflicts: resolved as development/v2's plan with 2.5 (a)'s own change applied. Every file
now equals development/v2 plus exactly this PR's commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EaBE7c7xYRveqLrVTQTpQ5
@tap
tap merged commit 49fb308 into development/v2 Oct 1, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants