Alerts

Create alert rules and route firing alerts to Slack, email, and incident-sync targets like Jira and ServiceNow.

The Alerts page shows alerts and lets you manage the rules that produce them. An alert is either firing or resolved. When an alert is firing you can Resolve it from the list, or click Declare incident to escalate it into Incident Management (Enterprise) for team response and a postmortem.

Alert rules

Open the Rules tab and click + Add rule. A rule has:

  • Name
  • Description
  • Metric, grouped by area
    • Resource Usage, Container CPU %, Container Memory %, Container Memory MB, Pod CPU %, Pod Memory %, Pod Restart Count, Node Not Ready
    • Fleet Health, Agent Silent (minutes)
    • CI/CD, Pipeline Failed, Pipeline Stuck, Pipeline Flaky, Deploy Regression, VCS Connector Unhealthy
    • Backups, Cluster/Host Backup Failed, Cluster/Host Backup Overdue, Disaster Recovery Backup Failed. Covers both Cluster & Host Backups and Disaster Recovery, which back up different things (see each page for the distinction)
    • Ingress & Gateway API, NGINX Ingress Reload Failed, NGINX Ingress Config Reload Errors, NGINX Gateway Fabric Reload Errors
    • Autoscaling, Cluster Autoscaler Unschedulable Pods, Cluster Autoscaler Not Safe to Scale
  • Operator, greater than, greater-or-equal, less than, or less-or-equal
  • Threshold
  • Severity

You can enable or disable, edit, and delete rules at any time. The alert engine checks each enabled rule every cycle against the latest sampled value. It fires as soon as the condition is true and resolves as soon as it isn't, since there's no sustained-duration window to wait out.

Some metrics are absolute values rather than percentages (for example a raw MB figure instead of a 0-100% one). A threshold that made sense as a percentage almost never makes sense as the same number against an absolute value. A rule like "greater than 70" against a percent-based metric fires around 70% usage as intended, but the identical "greater than 70" against a raw MB metric fires the moment usage passes 70 MB, which most workloads exceed immediately and never drop back below, so the alert never auto-resolves and has to be resolved by hand every time. If an alert never auto-resolves, double check the rule is using the percent-based version of its metric, not the absolute one, before assuming it's a bug.

Notifications

Configure notification channels under Alerts → Notifications (or Settings → Notifications). Six channel types are supported, Slack, Email, Jira, ServiceNow, Statuspage, and Salesforce. Every channel has Test (sends a real test message so you can confirm delivery actually works, independent of whether any alert has fired recently), Edit (change its config without deleting and recreating it, since the channel type itself can't be changed after creation), and Delete.

Slack and Email fire on every alert, the moment it fires or resolves. Jira, ServiceNow, Statuspage, and Salesforce are **incident-sync** channels instead. They fire on an [incident's](/dashboard/incident-management) declare, status-change, and resolve events, not on a plain alert firing, and they update the same external ticket across that whole lifecycle rather than creating a new one each time.

Slack

Paste an incoming webhook URL from the Slack app you want to post as.

Email

Set a From address and one or more To recipients.

Jira

Base URL (e.g. https://yourteam.atlassian.net), Email (the Atlassian account the API token belongs to), API Token (create one at id.atlassian.com/manage-profile/security/api-tokens), Project Key (the short uppercase key shown in your Jira project's settings, e.g. OPS), and Issue Type (defaults to Bug). KubeWatch creates one issue per incident, adds a comment on every status change, and best-effort transitions it toward Done/Closed/Resolved when the incident resolves (a transition that doesn't exist under those names in your workflow is silently skipped rather than failing the sync).

ServiceNow

Instance URL (e.g. https://yourcompany.service-now.com), Username, Password, and Table (defaults to incident). KubeWatch creates a Table API record on declare and updates its state field (New / In Progress / Resolved) as the incident's status changes.

Statuspage

API Key (Statuspage account settings → API) and Page ID (found in the page's own settings). KubeWatch creates a Statuspage incident on declare and updates it on every status change, mapping KubeWatch's status onto Statuspage's own vocabulary (investigating / identified / monitoring / resolved).

Salesforce

Instance URL (your org's My Domain URL, e.g. https://yourcompany.my.salesforce.com), Access Token, and API Version (defaults to v58.0). KubeWatch creates a standard Salesforce Case on declare and updates its Status field (New / Working / Closed) as the incident changes.

Alerting evaluates the metrics your agents report, so a rule only fires once the relevant metric is flowing in.