Skip to content

v12 dashboard: UTMStack platform #2730

Description

@kryonsx

Part of: #2696 · Needs first: #2697 (click-through to the Log Explorer and the value/count table)
Related: #1739 (parsers and rules for this integration, owned by the detection team)

Goal

Ship a built-in UTMStack platform dashboard that shows, at a glance, how much UTMStack platform is sending, whether it is still sending, and what kinds of events they are. The heart of it is the list of UTMStack platform logs grouped by messages (widget W7): click one and the Log Explorer opens on exactly those logs.

Where the data comes from

Integration (catalog name) UTMSTACK
Data type utmstack
How the logs arrive The UTMStack collector, installed by hand on the UTMStack server from the integration's setup guide, follows the Docker logs of the platform's own containers and sends every line as a utmstack log.
Parser definitions/filters/utmstack/utmstack.yaml
What dataSource holds The host name of the machine that runs the collector, which is the UTMStack server itself. Proof: collectors/utmstack/collector/docker.go line 356 sets DataSource: osInfo.Hostname, which utils/os.go GetOsInfo fills from os.Hostname() (lines 38-42); log-input/ingest/server.go applyDefaults only fills 'unknown' when it is empty. The collector reads only the local Docker daemon, so each swarm node needs its own collector.
Grouped by log.msg (messages)

Why log.msg: Every Go service in the stack logs through go-sdk catcher. Its JSON line holds a fixed message text in msg, with the changing details in args and cause (go-sdk catcher/log.go SdkLog and errors.go SdkError; code is just the MD5 of msg, so msg is meant to name a kind of event). The json step (utmstack.yaml lines 8-10) keeps it as log.msg. The values are not too free-form: they are fixed strings written by developers, readable, and moderate in number (backend examples below). A few embed values, such as 'CORS rejected origin "..." (devMode=false) — request will be aborted with 403' or 'database not ready (attempt 3/10); retrying in 5s', but those are rare. The real gap is coverage: log.msg exists only when the line is clean JSON and has an args key. Today that means backend lines after start-up that pass arguments (every catcher.Warn, and Error or Info calls with an args map). Event processor, agent manager and log input lines start with an icon and are not parsed at all, and ClickHouse, NATS and Redis print plain text into log.message (see caveats). No other field is better: log.message holds whole lines (every row unique), raw cannot be grouped, and log.containerName is nearly always the backend.

Typical values: http, sso: sign-in failed, alerts: notification failed, incidents: failed to write history, soar dispatch: claim failed, compliance: scheduled evaluation failed, integrations: failed to register datasource, session purge failed, audit entries past retention removed, dashboards: seeding a system dashboard failed

Widgets

Standard layout from the parent issue; rows W2, W3 and W9 onward are specific to this integration.

# Title Shown as Query Why
W1 Total logs number logs: count All platform log lines, from every collected container.
W2 Errors number logs: count; filter raw contains "severity":"ERROR" Error lines from every Go service, parsed or not; searching raw is the only way to count event processor, agent manager and log input errors today.
W3 Failed API requests number logs: count; filter log.msg is one of [http] Backend API answers with status 400 or more (failed sign-ins, missing permissions, server errors).
W4 Alerts number alerts: count Alerts from rules written for dataType utmstack; none ship today, so it stays at zero until one is added.
W5 Log volume over time area chart logs: count over time Gaps mean the collector stopped; a steep climb can mean an error storm.
W6 Logs by UTMStack host bar chart logs: top 10 values of dataSource dataSource is the host name of the server running the collector; one bar per node that has one.
W7 Top messages value and count table logs: top 25 values of log.msg; filter log.msg exists The main list: backend log lines grouped by message; click one to open its lines.
W8 Messages over time (top 5) line chart logs: count over time, one line per value of log.msg (top 5); filter log.msg exists Spots a message that suddenly repeats, such as a burst of failed API requests.
W9 Top error messages value and count table logs: top 10 values of log.msg; filter log.severity is one of [ERROR, CRITICAL, ALERT]; log.msg exists The backend failures behind the errors, by message (backend only, so its total is below W2).
W10 Failed API requests by path value and count table logs: top 10 values of log.args.path; filter log.msg is one of [http]; log.args.path exists Which API endpoints fail most; paths are long, so a value and count table.
W11 Failed API requests by status code bar chart logs: top 10 values of log.args.status; filter log.msg is one of [http]; log.args.status exists 401 and 403 point to sign-in or permission problems, 5xx to backend faults.
W12 Failed API requests by client IP bar chart logs: top 10 values of log.args.client_ip; filter log.msg is one of [http]; log.args.client_ip exists Addresses behind the failures, for example someone guessing passwords on the console.
W13 Alerts by rule bar chart alerts: top 10 values of name Kept for the standard layout; empty until a rule for dataType utmstack exists.
W14 Alerts by severity bar chart alerts: top 10 values of severity Kept for the standard layout; empty until a rule for dataType utmstack exists.
W15 Latest logs table of latest logs logs: latest 20 records; columns @timestamp, dataSource, log.msg, log.severity, log.cause, raw raw is included because most lines (event processor, agent manager, log input, ClickHouse, NATS, Redis) have no parsed field.

Fields used and where they come from

  • log.msg: Catcher message: one fixed text per log call in the code. Examples: http, sso: sign-in failed, soar dispatch: claim failed. Source: json step lines 8-10 (top-level key msg of go-sdk catcher SdkLog/SdkError). Call sites: backend/server.go line 89 ('http'), backend/modules/iam/handler/idp.go line 278, backend/modules/soar/usecase/dispatch.go line 136.
  • log.severity: Catcher level, upper case. INFO by default; when args carry a status: 1xx DEBUG, 2xx INFO, 3xx NOTICE, 4xx WARNING, 500-501 ERROR, 502-508 CRITICAL, 509-510 ALERT; catcher.Error without a status is ERROR and catcher.Warn always adds status 400 (WARNING). Examples: INFO, WARNING, ERROR. Source: json step lines 8-10; go-sdk catcher/errors.go calculateSeverity, catcher/log.go Warn and Log. The parser maps nothing to the top-level severity column.
  • log.args.path / log.args.status / log.args.method / log.args.client_ip: Details of a failed API request, written when the backend answers with status 400 or more (msg 'http'). status is a number; client_ip is gin's ClientIP(). Examples: /api/v1/alerts, 401, GET, 203.0.113.9. Source: backend/server.go lines 86-95 (requestLogger; /api/v1/health is skipped). The json step keeps nested keys as they are (EventProcessor plugins/json/main.go sanitizes top-level keys only).
  • log.cause: Error text given to catcher.Error. For msg 'http' it is 'status '. Examples: status 404, context deadline exceeded. Source: json step lines 8-10; go-sdk catcher/errors.go Cause field.
  • log.containerName: Swarm service name the collector adds to clean JSON lines only; in practice almost always the backend. Examples: utmstack_backend. Source: collectors/utmstack/models/docker.go EnrichLogWithContainer lines 84-103 (only when isValidJSON, lines 105-109) and CleanContainerName lines 111-121.
  • log.message: The whole line when it lacks "msg": or "args": (plain-text containers, catcher calls without arguments). Not usable for grouping. Examples: [1] 2026/09/24 10:15:02.123456 [INF] Server is ready. Source: grok lines 13-18.
  • raw: The line as the container printed it, after the collector removes Docker's 8-byte header. The only field every utmstack log has. Examples: <icon> {"timestamp":"...","code":"...","msg":"failed to parse log during pipeline execution",...,"severity":"ERROR"}. Source: collectors/utmstack/collector/docker.go lines 334-357 (Raw: enrichedLog).

Watch out for

  • A dashboard makes sense, but only part of it can be broken down. log.msg, and with it W3 and W7 to W12, exists only on backend lines that are clean JSON with arguments: every catcher.Warn, plus Error or Info calls that pass an args map. The event processor (manager and worker), agent manager and log input print an icon before their JSON, so their lines keep no parsed field (see parser issues). ClickHouse, NATS and Redis print plain text, which lands in log.message. The collector skips postgres and frontend (collectors/utmstack/config/const.go lines 36-77; the other names in that list are v11 containers). So W1, W2, W5, W6 and W15 cover every component, while W3 and W7 to W12 cover the backend only.
  • W2 counts every line whose catcher JSON says "severity":"ERROR", parsed or not, with a case-insensitive contains on raw (positionCaseInsensitive in go-sdk store clickhouse filter.go). The text comes from go-sdk catcher, which marshals the JSON without spaces. CRITICAL and ALERT lines are not counted, since a contains filter takes one value. The click-through must pass the full filter list, because contains cannot travel as a URL parameter (foundation task, item 1).
  • W9 lists backend errors only, so its total is smaller than W2.
  • No shipped correlation rule uses dataType utmstack (nothing under definitions/rules), so W4, W13 and W14 stay at zero until someone writes one. They are kept so the layout matches the other dashboards.
  • On a single-server install W6 shows one bar. The collector reads only the local Docker daemon, so a multi-node swarm needs a collector on each node to see every container.
  • log.containerName is added only to lines that are clean JSON (collectors/utmstack/models/docker.go EnrichLogWithContainer), so today it is almost always utmstack_backend; for global swarm services (event-processor-worker, log-input) it would also keep the node and task IDs, because CleanContainerName only strips '..'. That is why there is no 'Logs by container' widget. Add one once the icon problem is fixed.
  • log.args.status is stored as a number. The 'in' filters are safe (the store compares as text); check once on a live instance that clicking a status bar (eq '404' on a number) matches.
  • log.args.client_ip is gin's ClientIP(). The backend trusts every proxy (gin's default; backend/server.go uses gin.New() with no SetTrustedProxies), so it is the left-most X-Forwarded-For address, which a client can set itself. Good for spotting a noisy address, not proof of who sent a request.
  • The top-level severity column is empty for utmstack logs; the level is only in log.severity, upper case (INFO, WARNING, ERROR, CRITICAL, ALERT, NOTICE, DEBUG).
  • Field names do not depend on the underscore change: catcher's top-level keys (timestamp, code, msg, cause, args, severity, containerName) have no underscores, and nested args keys such as client_ip and latency_ms were never stripped.
  • If the feedback loop described in the parser issues happens, W1, W2 and W5 will be dominated by the event processor's own parse errors about platform logs.
  • The collector drops the first 8 bytes of every line as Docker's stream header and splits the stream on newline bytes (collectors/utmstack/collector/docker.go lines 260-285 and 334-337). When a header's length bytes contain a newline byte (roughly one line in 256), the line is split and the next piece loses 8 real characters, so a small share of lines arrive broken.
  • Checked by replaying utmstack.yaml over sample catcher lines (clean JSON with and without args, icon-prefixed, plain text) with a local copy of the step logic that follows the EventProcessor code. The icon prefix and the feedback loop come from reading the code; neither was seen on a live system.

Parser problems found while designing this

These are not dashboard work, but they limit what the dashboard can show. They belong to #1739; raise them there rather than working around them in the dashboard.

  • Lines that start with an icon are lost. Every Go service in the stack logs through go-sdk catcher, which prints an icon and a space before the JSON unless CATCHER_BEAUTY=false (go-sdk catcher/catcher.go init; log.go and errors.go print the icon, a space, then the JSON). Only the backend turns it off (backend/modules/audit/module.go line 22); installer/docker/compose.go sets no CATCHER_BEAUTY for the event processor, agent manager or log input. For their lines the json step (lines 8-10) runs because raw contains "msg": and "args":, fails on the icon, and the grok fallback (lines 13-18) is skipped because its condition is the exact opposite, so the log keeps no field at all, not even log.message. Fix: set CATCHER_BEAUTY=false for every service, and have the parser drop anything before the first '{' (or fall back to log.message when the json step fails).
  • Likely feedback loop (from code reading, not seen on a live system). Each utmstack line that fails the json step makes the event processor print two new error lines: 'failed to decode JSON' from the json plugin (EventProcessor plugins/json/main.go lines 42-48; plugin output goes to the container's stdout, pkg/tools/plugins.go lines 36-40) and 'failed to parse log during pipeline execution' (pkg/parsing/parsing.go lines 469-479). Both start with an icon and carry msg and args, so when the collector ships them they fail the same way and each makes two more. Volume can keep growing until the collector's queue starts dropping logs (collectors/utmstack/collector/docker.go lines 365-382). Fixing the icon problem stops it; so would excluding the event processor's containers from the collector.
  • Clean catcher lines without arguments are not parsed: args is left out of the JSON when empty (json tag omitempty), so the condition at lines 8-10 fails and the whole line goes to log.message. At least 123 of the backend's 214 catcher call sites pass no arguments to Error or Info. Testing for "msg": alone would be enough.
  • The parser maps nothing to the standard columns: the catcher level stays in log.severity (the severity column is empty) and the service name only in log.containerName, so the platform's logs cannot be filtered like other sources by severity.

How to build it

Before you start: parser field names are changing while the parsers are updated for the engine's new underscore handling (see "Field names are about to move" in #2696). Check every field in this issue against the v12 parser at that moment and against real logs, and build against what you find.

  1. Add definitions/dashboards/integration-utmstack.yaml. Keep it in the top folder: the test that checks shipped dashboards (TestEveryShippedDashboardDefinitionIsValid) only reads the top folder.
  2. Start from the file below; it follows the table above and passes the same checks as the backend (domain.Spec.Validate).
  3. W7 is a value/count table: a category query shown with chart type table. It already renders; #2697 makes its rows clickable and its headers readable.
  4. Load real logs from this technology on a v12 test server (or replay samples) and check every widget before opening the pull request.
Starting dashboard file
# Dashboard version v1.0.0
#
# System-owned default dashboard for the UTMStack platform integration, seeded by
# backend/modules/dashboards/repository/dashboard_bootstrap.go.
# Field names come from definitions/filters/utmstack/utmstack.yaml.
# Verify every widget against real logs before shipping.
name: "UTMStack platform"
description: "What UTMStack platform is sending: volume, event types, and the activity worth a look."
widgets:
  - layout: { x: 0, y: 0, w: 3, h: 2 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: metric
      metric:
        agg: count
    config:
      __builder:
        chartType: metric
        title: "Total logs"

  - layout: { x: 3, y: 0, w: 3, h: 2 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: metric
      metric:
        agg: count
      filters:
        - field: "raw"
          op: contains
          value: "\"severity\":\"ERROR\""
    config:
      __builder:
        chartType: metric
        title: "Errors"

  - layout: { x: 6, y: 0, w: 3, h: 2 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: metric
      metric:
        agg: count
      filters:
        - field: "log.msg"
          op: in
          value: ["http"]
    config:
      __builder:
        chartType: metric
        title: "Failed API requests"

  - layout: { x: 9, y: 0, w: 3, h: 2 }
    spec:
      dataset: alerts
      dataType: "utmstack"
      chart: metric
      metric:
        agg: count
    config:
      __builder:
        chartType: metric
        title: "Alerts"

  - layout: { x: 0, y: 2, w: 8, h: 4 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: time
      metric:
        agg: count
    config:
      __builder:
        chartType: area
        title: "Log volume over time"

  - layout: { x: 8, y: 2, w: 4, h: 4 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: category
      metric:
        agg: count
      dimension: "dataSource"
      limit: 10
    config:
      __builder:
        chartType: bar
        title: "Logs by UTMStack host"

  - layout: { x: 0, y: 6, w: 6, h: 6 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: category
      metric:
        agg: count
      dimension: "log.msg"
      filters:
        - field: "log.msg"
          op: exists
      limit: 25
    config:
      __builder:
        chartType: table
        title: "Top messages"

  - layout: { x: 6, y: 6, w: 6, h: 6 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: time
      metric:
        agg: count
      dimension: "log.msg"
      filters:
        - field: "log.msg"
          op: exists
      limit: 5
    config:
      __builder:
        chartType: line
        title: "Messages over time (top 5)"

  - layout: { x: 0, y: 12, w: 6, h: 4 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: category
      metric:
        agg: count
      dimension: "log.msg"
      filters:
        - field: "log.severity"
          op: in
          value: ["ERROR", "CRITICAL", "ALERT"]
        - field: "log.msg"
          op: exists
      limit: 10
    config:
      __builder:
        chartType: table
        title: "Top error messages"

  - layout: { x: 6, y: 12, w: 6, h: 4 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: category
      metric:
        agg: count
      dimension: "log.args.path"
      filters:
        - field: "log.msg"
          op: in
          value: ["http"]
        - field: "log.args.path"
          op: exists
      limit: 10
    config:
      __builder:
        chartType: table
        title: "Failed API requests by path"

  - layout: { x: 0, y: 16, w: 6, h: 4 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: category
      metric:
        agg: count
      dimension: "log.args.status"
      filters:
        - field: "log.msg"
          op: in
          value: ["http"]
        - field: "log.args.status"
          op: exists
      limit: 10
    config:
      __builder:
        chartType: bar
        title: "Failed API requests by status code"

  - layout: { x: 6, y: 16, w: 6, h: 4 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: category
      metric:
        agg: count
      dimension: "log.args.client_ip"
      filters:
        - field: "log.msg"
          op: in
          value: ["http"]
        - field: "log.args.client_ip"
          op: exists
      limit: 10
    config:
      __builder:
        chartType: bar
        title: "Failed API requests by client IP"

  - layout: { x: 0, y: 20, w: 6, h: 4 }
    spec:
      dataset: alerts
      dataType: "utmstack"
      chart: category
      metric:
        agg: count
      dimension: "name"
      limit: 10
    config:
      __builder:
        chartType: bar
        title: "Alerts by rule"

  - layout: { x: 6, y: 20, w: 6, h: 4 }
    spec:
      dataset: alerts
      dataType: "utmstack"
      chart: category
      metric:
        agg: count
      dimension: "severity"
      limit: 10
    config:
      __builder:
        chartType: bar
        title: "Alerts by severity"

  - layout: { x: 0, y: 24, w: 12, h: 6 }
    spec:
      dataset: logs
      dataType: "utmstack"
      chart: table
      metric:
        agg: count
      limit: 20
      columns: ["@timestamp", "dataSource", "log.msg", "log.severity", "log.cause", "raw"]
    config:
      __builder:
        chartType: table
        title: "Latest logs"

Done when

  • definitions/dashboards/integration-utmstack.yaml is merged to release/v12.0.0 and go test ./modules/dashboards/... passes in backend/.
  • With UTMStack platform logs flowing on a v12 test server, every widget shows data. A widget that stays empty while logs arrive means a wrong field or value: fix it, don't ship it.
  • Clicking a row, bar or line point opens the Log Explorer on the same logs, and the Log Explorer count matches the widget.
  • A time range with no logs shows empty states, not errors.
  • A screenshot of the finished dashboard is attached to this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions