Skip to content

fix(linux): one logging-service alert per host - #2789

Merged
kryonsx merged 2 commits into
utmstack:v11from
kryonsx:codex/linux-alert-noise-20260929
Oct 1, 2026
Merged

kryonsx merged 2 commits into
utmstack:v11from
kryonsx:codex/linux-alert-noise-20260929

Conversation

@kryonsx

@kryonsx kryonsx commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Why

"Audit or Logging Service Disabled" grouped by origin.host and origin.user. An alert never carries origin.* (the alert plugin stores the event's origin side as adversary), and grouping keys that do not resolve are skipped, so every matching record opened a new top-level alert. On one v11 deployment the rule stored 335 alerts in 30 days, up to 88 in a day and 10 in one hour, and the rule flood guard can switch it off.

Noise is reduced here only by de-duplication. The condition is unchanged.

What changes

Before After
Alerts groupBy origin.host, origin.user (never resolve on an alert) deduplicateBy dataSource: one alert per host, repeats dropped for seven days

Expected volume

Every match was stored as its own alert, so the rule's alerts over the last 30 days on the v11 deployments are its matches (341 on three deployments). Replayed with the new key and seven-day de-duplication: 9 alerts in 30 days (6, 2 and 1), worst hour 1, worst day 1.

What the condition matched is worth a separate review: kernel out-of-memory messages that name the journald or rsyslog cgroup (oom-kill: … cpuset=systemd-journald.service), rsyslog reloads after log rotation (systemctl kill -s HUP rsyslog.service), package scripts that unmask rsyslog (unmask contains mask), journald watchdog restarts and a Filebeat status line that lists process capabilities.

Tests

  • plugins/alerts/linux_alert_volume_test.go (new) runs fabricated journald records through the Linux raw model and the pinned go-sdk v1.1.36 CEL. It pins the de-duplication key, checks that it resolves, and checks that stop, disable and mask of a logging service match while a restart and an unrelated message do not.
  • go test ./... in plugins/alerts passes on this branch, which is based on current v11 (89cd26c4).
  • Local EventProcessor lab (engine 8a3ade7, go-sdk v1.1.36, the production events and alerts plugins, OpenSearch 2.19.1, this branch's Linux filter from current v11) with fabricated journald records in two runs: systemctl stop auditd, a restart and an unrelated message on one host, then systemctl disable rsyslog.service on that host and systemctl stop auditd on a second host. Both runs processed all of their events.
    • With this branch: 3 alerts produced and 2 stored (one per host); the second match on the first host was dropped as a duplicate. The restart and the unrelated message produced none.
    • With the rule as shipped: 3 alerts produced and 3 top-level alerts stored; origin.host and origin.user resolved on none of them.

🤖 Generated with Claude Code

kryonsx and others added 2 commits September 29, 2026 19:07
"Audit or Logging Service Disabled" grouped by origin.host and
origin.user, which an alert never carries (the alert holds the event's
origin side as adversary). Grouping keys that do not resolve are
skipped, so every matching record opened a new top-level alert.

The rule now raises one alert per host (dataSource) and drops repeats
for seven days. Its condition is unchanged.

linux_alert_volume_test.go pins the key and checks that it resolves on
fabricated journald records.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kryonsx
kryonsx marked this pull request as ready for review October 1, 2026 02:10
@kryonsx
kryonsx requested a review from a team October 1, 2026 02:10
@kryonsx
kryonsx merged commit 1fed817 into utmstack:v11 Oct 1, 2026
2 of 5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant