Repository navigation
Conversation
Dispatching a set of jobs with concurrency controls acquires each job's semaphore one by one, inside the transaction that dispatches the whole set, and in whatever order the jobs come in. Two dispatchers whose batches share concurrency keys can then acquire them in opposite order and deadlock, and so can bulk enqueues. The batch that's aborted stays scheduled, so no job is lost and the limits hold, but the dispatcher that hits the error is delayed. Go through the jobs by concurrency key instead, so everyone acquires the locks they share in the same order. Within a key, jobs keep their id order, so the same job as before gets the lock first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Acq8BRNz9v7w9mKYLYRdCi
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
dispatch_next_batchdispatches a whole batch in one transaction and, for jobs with concurrency controls, acquires each semaphore in job ID order. Two dispatchers whose batches share keys can take them in opposite order and deadlock onsolid_queue_semaphores.perform_all_latergoes through the same path.However, no job is lost and limits hold, the batch rolls back and stays scheduled, and the dispatcher is restarted.
It only shows up under load (3 dispatchers, 200 limited jobs/s over 50 keys: ~20 deadlocks in 20 s on PostgreSQL 17).
This issue is different from #609 and #162.
There is a script to reproduce the issue with a rails single file https://gist.github.com/ceritium/82d13702de27d36831204bd126370568
Solution
The fix sorts by
[concurrency_key, id]so everyone takes shared locks in the same order. The batch commits atomically and workers claim by priority and id, so the order isn't visible outside.Let me know what you think.
Many thanks!