AI Diagnostics
Every org gets a free, local model for log-based root-cause diagnosis out of the box. Bring your own key if you'd rather use a hosted provider.
Every other AI-Ops feature on this platform is deliberately rule-based, not LLM-backed (see the AI-Ops overview for why). AI Diagnostics is the first feature that actually calls an LLM. Every org gets a free, local model connection the moment they first open the feature, Embedded (free, local), a local model (Qwen2.5-7B-Instruct) running on KubeWatch's own infrastructure, never a third-party API, at no cost and with no key to configure. If you'd rather use your own OpenAI, Anthropic, Azure OpenAI, or OpenAI-compatible connection instead (or alongside the embedded one), you can add one under LLM connections, and every diagnosis that runs against it spends your own budget, not a shared platform key.
What you get
- Automatic investigation means that when one of your alert rules fires, AI Diagnostics pulls the last 15 minutes of logs for the affected container or pod, and asks your configured LLM (the free embedded model by default, or your own connection if you've added one) for a root-cause summary and, when it's confident enough, a suggested fix. To stop a flapping alert from re-diagnosing the same resource over and over, an alert-triggered diagnosis is skipped if that same resource was already diagnosed in the last 15 minutes (configurable); a manual click never gets skipped this way.
- On-demand diagnosis means clicking Diagnose with AI on a container or pod's Logs tab, on any firing alert, or on any incident linked to one, without waiting for a second occurrence.
- A structured diagnosis, not a wall of text, made up of a one-line summary, a confidence score, supporting detail, and (when applicable) a specific suggested action, all shown on the Root Cause AI page alongside the exact log excerpt that was sent to the model. A diagnosis's status is one of pending, completed, failed, or rate_limited (skipped because the org's hourly cap was already reached, not an in-progress state).
- The suggested fix is never trusted blindly, since it must resolve to one of the same four actions Autonomous Remediation already supports (restart a Docker container, rolling-restart a Kubernetes workload, roll back a Kubernetes Deployment, or bump a container's memory limit for a repeated OOMKilled pattern) and must actually apply to the diagnosed resource. Anything else, including a malformed or invented action type, is silently downgraded to a read-only suggestion, never shown as executable.
How it connects to Autonomous Remediation
A resolvable, safe suggestion doesn't execute anything by itself, and it only ever becomes a draft in the first place when the model's own confidence score clears an org-configurable threshold (default 60%); below that, the suggestion stays visible on the diagnosis as advisory only. When it does clear the bar, it becomes a draft remediation playbook, visible on the Remediation page, tagged as AI-suggested. It starts in dry-run exactly like a playbook you'd write by hand, and only reaches unattended execution after a human reviews a dry run and explicitly promotes it, only on the Enterprise plan.
On every other plan, AI Diagnostics still gives you the full diagnosis and suggested fix (captioned with an upgrade prompt), it just never drafts a playbook for it. One org-wide setting controls unattended execution platform-wide, the same Autonomous Remediation has always used, rather than a second gate that could drift out of sync with it.
Setting up a connection
You don't need to do anything to get started, since the first time your org uses this feature, KubeWatch automatically provisions the Embedded (free, local) connection and sets it as your default, at no cost. To use a hosted provider instead, add a connection under Root Cause AI → LLM connections. Supported providers:
- OpenAI, the standard
api.openai.comendpoint. - Anthropic, the standard
api.anthropic.comendpoint. - Azure OpenAI / OpenAI-compatible, any endpoint that speaks the OpenAI Chat Completions wire format, with a base URL you supply (a self-hosted gateway or proxy, for example).
Your key is encrypted at rest and is never returned to the browser after you save it, only a Test connection button that reports whether it's reachable. Whichever connection is marked default is the one every new diagnosis uses.
What it doesn't do
There's no continuous background scanning of your logs, since diagnosis only runs when an alert actually fires, or when you explicitly ask for one. A defensive per-hour cap (configurable, default 20) also protects your own LLM budget from a flapping alert triggering a diagnosis storm; once an org hits it, further diagnoses for that hour are recorded as rate_limited rather than run.
It also doesn't create or update incidents on its own. The link between an incident and a diagnosis is a single button in the dashboard; there's no automatic incident creation or resolution tied to a diagnosis result.