Repository navigation
Retry batches answered with 429 or 5xx, and report rejected batches once - #55
Merged
Merged
Conversation
The HTTP outlet treats every response as delivered: 429 and 5xx are never retried, Retry-After is ignored, and 401 or 413 drop the batch without a word. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under `ruby -w`, warnings from other code reach stderr while the outlet runs and broke the exact stderr matches. Expect the warning through `warn` instead, and match the 413 one as a line of stderr. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
deliver_requests treated every response as a delivery. Now: - 2xx: delivered, the reconnect wait starts over. - 429 and 5xx: the request goes back on the queue like after a failed connection, with the same backoff and limit of 3 attempts, without starting the wait over. Retry-After (seconds or an HTTP date) sets the minimum wait, capped at 60 s. - Any other status: the batch is dropped, and the first rejection with each status in the process prints a warning through Kernel#warn, e.g. "Logtail: Better Stack rejected 37 log lines with HTTP 401 Unauthorized - check your source token. Further rejections with this status won't be reported." RequestAttempt takes the batch's line count for the warning, and the retry-or-drop logic moved into a helper that the exception path uses too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Better Stack's ingesting proxy answers 408 Request Time-out, and closes the connection, when a new connection stays unused for more than about 10 to 15 seconds. The outlet opens a new connection after every requests_per_conn requests and then waits for lines, so after a quiet spell the next batch gets that 408. It is dropped with a warning that Better Stack rejected it, although the server never read the request. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A 408 means the server didn't read the request, so the batch is kept and retried with the same backoff and attempt limit, waiting as long as a Retry-After header asks, and no rejection is reported. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PetrHeinz
marked this pull request as ready for review
October 2, 2026 09:07
Conflicts in lib/logtail/log_devices/http.rb, both sides kept: - require block: #54's require "set" and this branch's require "time" - constants after MAX_RECONNECT_WAIT: this branch's MAX_RETRY_AFTER, REPORTED_REJECTIONS and REPORTED_REJECTIONS_LOCK, then SYNCHRONOUS_DELIVERY_TIMEOUT from #57/#58 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ivery With #57 and #58 merged, a delivery in the calling thread (flush without an outlet thread, lines written after close or during shutdown) counts any HTTP response as delivered. So a batch rejected with 401 or 403 isn't reported, and lines written after close keep going out after a 429 or 5xx. Three of the four new examples fail for that. The fourth, a batch answered with 408, 429 or 5xx is neither retried nor reported, passes already and keeps the fix from reporting those statuses as rejected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…very deliver_synchronously now counts only 2xx as delivered. 408, 429 and 5xx are logged at debug level and not retried, since nothing would deliver the retry; any other status is dropped and reported once per status through report_rejected_batch. It takes the RequestAttempts, so the warning can count the lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The HTTP outlet treats any response as a delivery (
deliver_requests). A 429 or 5xx is never retried and itsRetry-Afteris ignored, and a 401, 403 or 413 drops the batch without a word. Only exceptions are retried, and since 0.1.20 any response, a 503 included, also resets the reconnect backoff. The red-team found that a wrong source token loses 100% of the logs with no output at all.Retry-Afterheader (seconds, or an HTTP date) sets the minimum wait, capped at 60 s. These responses don't reset the backoff.Kernel#warn, never through a Logtail logger. For example:Logtail: Better Stack rejected 37 log lines with HTTP 401 Unauthorized - check your source token. Further rejections with this status won't be reported.Logtail::Configdebug output stays. It now calls any 2xx a success, not just 202.Behaviour and compatibility:
flush/closewait for it the same way. Against a local stub answering 503 at exit,closetook 2 s (three attempts); withRetry-After: 60it waited its full 20 s limit, as for an unreachable host. The shutdown PR (claude/fast-quiet-shutdown) shortens that wait.ruby -W0silences it, like anywarn.RequestAttempttakes an optional line count for the warning.Follow-up after the E2E run of the release candidate (two more commits, the first one only a test, red on CI): Better Stack's ingesting proxy answers
408 Request Time-out, and closes the connection, when a new connection stays unused for more than about 10 to 15 seconds. The outlet opens a new connection after everyrequests_per_connrequests and then waits for lines, so after a quiet spell the next batch got that 408 and was dropped with a misleading "rejected … with HTTP 408" warning; 0.1.20 lost it silently. A 408 means the server didn't read the request, so it's now retried like a 429 or 5xx, with the same backoff and attempt limit, honouringRetry-After, and without a warning.Targets the logtail 0.1.21 patch release. No dependencies on the other open PRs; the fork-safety and shutdown PRs change other methods of
http.rb.The first commit only adds the tests and is expected to fail on CI; the fix follows in the next commit.
🤖 Generated with Claude Code