Skip to content

fix(primitives): stop stale DaemonSets and ReplicaSets reporting healthy - #223

Merged
sourcehawk merged 1 commit into
mainfrom
fix/rollout-health--daemonset-replicaset
Oct 3, 2026
Merged

sourcehawk merged 1 commit into
mainfrom
fix/rollout-health--daemonset-replicaset

Conversation

@sourcehawk

Copy link
Copy Markdown
Owner

Description

A DaemonSet or ReplicaSet whose controller had not observed the current spec could report Healthy after the grace
period, and a DaemonSet rollout whose new pods never become ready reported Healthy. This PR brings the DaemonSet and
ReplicaSet status handlers in line with the rule from #220: the converging and grace handlers agree, and Healthy
needs the current generation observed, the ready count matched, and the rollout or scale-down finished. A DaemonSet with
the OnDelete strategy still reports Healthy once its pods are ready, because the controller does not replace them.

Changes

  • DaemonSet: DefaultConvergingStatusHandler and DefaultGraceStatusHandler require numberReady to equal
    desiredNumberScheduled, and updatedNumberScheduled to reach it. A ready but unfinished rollout reports Updating
    (Creating right after create) with Waiting for rollout: N/M pods updated. OnDelete skips the rollout check. A
    desired count of zero is still Healthy once the generation is observed.
  • ReplicaSet: both handlers require status.replicas not to be more than the desired count, and report
    Waiting for scale-down: N/M replicas while it is higher.
  • Both grace handlers report Down when pods are desired and none are ready, Healthy only when the converging handler
    reports Healthy, and Degraded for all other states, including a stale observedGeneration.
  • The DaemonSet converging handler now requires numberReady to equal the desired count, not to be at least it. The
    grace handler already reported Degraded for more ready pods than desired, so the two handlers disagreed there. The
    upstream controller never reports that state.
  • The DaemonSet grace reason for a stale generation is now Waiting for DaemonSet controller to observe latest spec, the
    same text as the converging handler.
  • "Status Handlers" sections in docs/primitives/daemonset.md (rewritten) and docs/primitives/replicaset.md (new),
    synced to the plugin references. GoDoc on the handlers, builders and resources matches.

Challenges

The DaemonSet controller counts only the oldest pod on each node when it computes numberReady and
updatedNumberScheduled (updateDaemonSetStatus in Kubernetes v1.34.1). With maxSurge, the old pod stays until the
new pod is ready, so numberReady can equal the desired count while updatedNumberScheduled stays below it. This is
why the rollout check is needed. A DaemonSet has no scale-down analogue: the controller counts at most one pod for each
node that must run one, and kubectl rollout status does not wait for numberMisscheduled. kubectl rollout status
also waits for numberAvailable. The handlers use numberReady, like the other workload kinds.

Related

Testing

TestDefaultHandlers_UnfinishedRollout (DaemonSet) and TestDefaultHandlers_NotConverged (ReplicaSet) call the
converging handler (what the component reports before the grace period expires) and the grace handler (what it reports
after) on the same object. Both tables cover a stale generation with all pods ready and with none ready. The DaemonSet table
adds a surge rollout with all old pods ready and an unfinished rollout just after create. The ReplicaSet table adds a
scale-down with one pod above the desired count. TestDefaultHandlers_OnDelete checks that an OnDelete DaemonSet with
old pods stays Healthy. Every new case failed before the fix except two that pin existing results: the stale
generation with none ready (Down) and OnDelete (Healthy). The healthy
DaemonSet fixtures now set updatedNumberScheduled, and two existing DaemonSet tests changed with the contract (more
ready than desired, and the grace reason for a stale generation).

make all passes, and make e2e-primitives passes for PRIMITIVE=daemonset and PRIMITIVE=replicaset.

🤖 Generated with Claude Code

…caSet

The DaemonSet and ReplicaSet grace handlers did not check
observedGeneration when pods were desired, so a stale generation with
all pods ready gave Healthy for grace while convergence reported
Updating. The DaemonSet handlers also did not read
updatedNumberScheduled. With maxSurge the controller keeps the old pod
on a node until the new pod is ready and counts only the oldest pod, so
a rollout whose new pods never become ready reported Healthy.

Both kinds now use the rule of the Deployment and StatefulSet handlers.
The controller must observe the current generation and the ready count
must equal the desired count. For a DaemonSet, updatedNumberScheduled
must reach desiredNumberScheduled, except with the OnDelete strategy,
because the controller does not replace those pods. A DaemonSet has no
scale-down to wait for. For a ReplicaSet, status.replicas must not be
more than the desired count. The grace handlers report Down when pods
are desired and none are ready, Healthy only when convergence is
Healthy, and Degraded otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 3, 2026 09:06
@sourcehawk sourcehawk changed the title fix(primitives): stop reporting stuck DaemonSets and stale ReplicaSets as healthy fix(primitives): stop stale DaemonSets and ReplicaSets reporting healthy Oct 3, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The health predicates, strategy exceptions, tests, and synchronized documentation are consistent and complete.

Review effort: Balanced
Findings: None

What changed in this PR

Aligns DaemonSet and ReplicaSet health reporting with other workload primitives, preventing stale or incomplete states from reporting healthy.

Changes:

  • Adds generation, readiness, rollout, and scale-down health checks.
  • Adds regression coverage for stale status and incomplete convergence.
  • Updates GoDoc and synchronized user/plugin documentation.
File Description
pkg/​primitives/​daemonset/​handlers.go Strengthens DaemonSet health checks.
pkg/​primitives/​daemonset/​handlers_test.go Tests stale and unfinished rollouts.
pkg/​primitives/​daemonset/​builder.go Updates builder GoDoc.
pkg/​primitives/​daemonset/​resource.go Updates resource GoDoc.
pkg/​primitives/​replicaset/​handlers.go Adds scale-down and generation checks.
pkg/​primitives/​replicaset/​handlers_test.go Tests stale and scaling states.
pkg/​primitives/​replicaset/​builder.go Updates builder GoDoc.
pkg/​primitives/​replicaset/​resource.go Updates resource GoDoc.
docs/​primitives/​daemonset.md Documents DaemonSet status handling.
docs/​primitives/​replicaset.md Documents ReplicaSet status handling.
plugin/​skills/​using-primitives/​references/​primitives/​daemonset.md Synchronizes DaemonSet plugin reference.
plugin/​skills/​using-primitives/​references/​primitives/​replicaset.md Synchronizes ReplicaSet plugin reference.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@sourcehawk
sourcehawk merged commit 613fcd3 into main Oct 3, 2026
8 checks passed
@sourcehawk
sourcehawk deleted the fix/rollout-health--daemonset-replicaset branch October 3, 2026 22:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants