Conversation
Event delivery now follows the shared rules (flagsmith-private#316, #318): - Retry 408, 429, 502, 503, 504, network errors and timeouts: 3 attempts in total, full-jitter backoff from 1s doubling to a 10s cap. Every other status, 500 included, drops the batch. - A batch that still fails is put back at the head of the buffer and waits for the next timed flush or flushEvents(). On close() it is dropped. - The buffer is bounded by maxBufferItems, dropping the oldest events, instead of max(maxBufferItems, 1000). A batch never exceeds it. - A 401 or 403 stops the processor: the timer stops, the buffer is dropped, later events are dropped, and one error is logged. - An exposure stays deduplicated until its event is delivered or dropped. - rejected[] entries on a 2xx are logged and counted as dropped. - FlagsmithClient.getDroppedEventCount() returns the monotonic drop count. - The close() bound covers the three attempts and their backoffs. - README and Javadoc cover flushEvents() and close() for short-lived processes.
|
@themis-blindfold review |
⚖️ Themis review: ✅ Ship itThe event queue now follows the stated delivery rules: bounded buffering, selective retries with jitter, requeueing, authentication shutdown, and monotonic drop accounting. CI is still running.
📝 Walkthrough
🧪 How to verify
Product take: A solid reliability improvement for experimentation telemetry, especially for intermittent events API failures and short-lived clients. The bounded buffer makes its delivery trade-off explicit and observable. 🧭 Assumptions & unverified claims
The event queue now has a clear exit strategy when the network gets dramatic. · reviewed at b087b19 |
|
@themis-blindfold review |
⚖️ Themis review: ✅ Ship itThe shared event-delivery rules are implemented consistently across batching, retries, shutdown, counters, tests, and user-facing guidance. The CI test matrix is still in progress; Maven is not installed in this environment, so I could not rerun the suite locally.
📝 Walkthrough
🧪 How to verify
Product take: Solid reliability and operability improvement for experimentation telemetry, especially for transient outages and short-lived workloads. The new drop counter makes the intentional bounded-buffer trade-off observable. 🧭 Assumptions & unverified claims
The events now fail with receipts instead of disappearing into the night · reviewed at 5896682 |
Closes Flagsmith/flagsmith-private#316. Follow-up to #226.
Applies the shared event delivery rules (flagsmith-private#318). The Go SDK already implements them.
Changes
flushEvents(). It is dropped if it fails duringclose().maxBufferItems, dropping the oldest events first. feat: experimentation support #226 usedmax(maxBufferItems, 1000). A batch is never larger thanmaxBufferItems.rejected[]: entries on a 2xx are logged, never resent, and counted as dropped.FlagsmithClient.getDroppedEventCount(), monotonic.close()bound: one batch's worst case, which is now 3 attempts at the client's timeouts plus the 1s and 2s backoff ceilings.flushEvents(),close()andwithEventsMaxBufferItems(), including the requirement that short-lived processes callclose().Behaviour change
With a very small
maxBufferItems, a burst of events can now drop events even when the events API is healthy, because the buffer is capped atmaxBufferItemswhile two batches are in flight. This is what the shared rule specifies, and Go behaves the same. The default is 1,000.Testing
mvn verifyandmvn verify -P test-okhttp4both pass (502 tests).