Skip to main content

Alerts

Alerts are emitted by the monitoring engine when a Rule trigger condition is met for a monitored resource. The engine evaluates every active rule once per minute and opens at most one alert per (rule, resource) pair — repeated triggers update the existing alert rather than creating a new one.

Accessing Alerts​

Navigate to Monitoring > Alerts for the full list. The Overview page summarizes alert activity with summary widgets and the most recent activity.

Alerts List​

The alerts table presents every alert across the tenant.

Columns:

ColumnDescription
Alert NameAlert identifier combining rule type and affected resource (for example, Disk utilization for web-01)
CreatedTimestamp the alert first opened
Last UpdatedTimestamp of the most recent trigger, severity change, or status change
SeverityCritical / High / Medium / Low — current level after any overrides
Alert Rule NameRule that produced the alert
Resource GroupDirectors / Devices / Targets
StatusOpen / Acknowledged / Resolved

Controls:

  • Search alerts — free-text search across rule name and resource name
  • Severity filter — All, Critical, High, Medium, Low
  • Status filter — All, Open, Acknowledged, Resolved
  • Alert Type filter — All, Director, Device, Target
  • Row checkboxes — select alerts to Acknowledge or Resolve them together (see Acting on several alerts)
  • Pagination at the foot of the table

Severity and Status​

Severity levels rank the urgency of an alert. A multi-severity rule sets the level from its configured thresholds; a single-severity rule applies the same level to every trigger.

SeverityBadge
Criticalred
Highred (lighter)
Mediumorange
Lowyellow

Status values describe where the alert sits in its lifecycle.

StatusDescription
OpenAlert was opened by a trigger and has not been acknowledged or resolved
AcknowledgedAn operator has taken ownership; the alert is still active but is being worked
ResolvedThe alert is closed, either manually or automatically by the engine

Alert Detail​

Clicking a row opens the alert detail page. The main panel surfaces:

  • Current status badge and severity badge
  • An alert summary line naming the rule type and resource (for example, "Disk utilization alert for web-01"), followed by a state-dependent description
  • Rule Name and Rule Type
  • A See alert rule details link that opens the rule's detail page
  • Acknowledged By and Acknowledgement Note (shown once the alert is acknowledged)
  • Resolved By and Resolution Note (shown once the alert is resolved)

A side info panel lists the Created and Last updated timestamps, plus Acknowledged and Resolved timestamps when those states have been reached.

Below the panel, a timeline section records every state change — the original trigger, severity upgrades, acknowledgements, and the resolution. Its columns are Event Time, Event, Severity, and Trigger Condition.

Acknowledging and Resolving​

Both actions are available from the alerts list row menu, from the Actions menu on the alert detail page, and for several alerts at once from the list's batch bar (see Acting on several alerts). The menu items are state-conditional: Acknowledge alert appears while the alert is Open, Remove acknowledgement while it is Acknowledged, and Resolve alert as long as it is not Resolved. These actions are gated by the alert-edit permission; opening the detail page itself requires the alert-read permission.

Acknowledge​

Acknowledge alert moves the alert from Open to Acknowledged and records the operator and timestamp. The acknowledgement note is optional. Remove acknowledgement returns the alert to Open.

Resolve​

Resolve alert moves the alert to Resolved and closes the lifecycle. The resolution note is required, captured in the timeline alongside the operator and timestamp. Resolution is final — a resolved alert cannot be reopened, but a new trigger against the same (rule, resource) pair will open a fresh alert.

Acting on several alerts​

The alerts list has a checkbox on every row. Selecting one or more alerts replaces the table toolbar with a batch bar, so a backlog can be acknowledged or resolved in one step instead of one alert at a time.

  1. Tick the alerts on the list. The selection is kept across pages, up to 500 alerts per action.
  2. The batch bar shows N alerts selected with Acknowledge, Resolve, and Cancel.
  3. Acknowledge opens Acknowledge N alerts. The note is optional, as it is for a single alert.
  4. Resolve opens Resolve N alerts. The note is required, as it is for a single alert.
  5. The note is entered once and saved on every selected alert. The result is reported in one notification:
    • N alerts acknowledged (or resolved) when every selected alert moved
    • X of Y alerts acknowledged (or resolved) when some were skipped, followed by how many had already changed state

Permissions. The checkboxes appear only for users with the alert-edit permission, the same permission as the single-alert actions. Without it the list has no checkboxes.

Resolved alerts cannot be selected. Their checkbox is shown but disabled.

Alerts the action no longer applies to are skipped, not failed. If another operator moved one of the selected alerts first, to a state the chosen action does not apply to, the action leaves that alert as it is and counts it in the notification; the rest of the selection is still actioned. Bulk Acknowledge skips an alert that is already Acknowledged, and bulk Resolve skips one that is already Resolved. An alert another operator acknowledged is still resolved by bulk Resolve.

Every actioned alert gets its own timeline entry and audit log entry, exactly as if it had been acted on alone.

There is no bulk Remove acknowledgement and no bulk Stop monitoring log source. Those actions stay per alert.

Alert Lifecycle​

Evaluation tick​

The monitoring engine re-evaluates every active rule once a minute. Threshold and metric checks run against the most recent observation window; status checks look at the current connection state.

Severity override​

When a multi-severity rule fires at a higher level against a resource that already has an Open alert, the existing alert's severity is upgraded in place — the timeline gains an entry recording the change, but no new alert is created. Triggers at the same or a lower severity update the last-triggered timestamp and increment the occurrence count without changing the severity.

Auto-resolve​

Auto-resolve is time-based only, and only for the rule types that accept a Resolve after period. When the period elapses since the alert first triggered, the alert closes on its own.

Rule typeAuto-resolve
Crash detection, Backpressure, Director status, Device status, Target status, No data received, Queue usage, Bus unhealthy, Bus restart, Consumer lag, JetStream storage utilizationCloses automatically once the configured Resolve after period has elapsed
Processor / memory / disk utilization, High and low data volume, High and low event volume, Total ingest amountNever closes on its own. Resolve it by hand once the condition has passed

Auto-resolved alerts carry the system-generated resolution note This alert was automatically resolved by the system after the defined resolution period expired. and are flagged in the timeline.

warning

The utilization and volume rules do not close when the triggering condition stops being met. A CPU spike that has long since passed leaves its alert Open until someone resolves it, so a busy list of these alerts is a backlog of past events, not a picture of the present.

Multi-resource fan-out​

A rule scoped to All directors, All devices, or All targets — with or without exceptions — fans out to every resource it matches. If five Directors breach the rule simultaneously, five independent alerts open — one per resource — and each follows its own acknowledge / resolve lifecycle.