AI-Ops
Seven features that bring AI agents, GPU-backed inference, self-hosted model stacks, and edge AI fleets into the same platform already watching your containers, pods, and nodes.
Your AI workloads are still workloads, since a LangChain agent burns real tokens, a vLLM server holds a real GPU, and an edge inference box can still go dark on a bad network. AI-Ops is that same monitoring, alerting, and (where it makes sense) automated response, applied to those specifics, on the one dashboard you already use for everything else, not a separate tool with its own login and its own alert queue.

The seven features
| Feature | Plan |
|---|---|
| Agent Cost-and-Reliability Ops | Pro |
| GPU and Inference Cost Control | Pro |
| Docker-vs-Kubernetes Workload Advisor | Pro |
| Edge AI Fleet Observability | Pro |
| Self-Hosted AI Stack Observability | Enterprise |
| Autonomous Remediation | Enterprise |
| AI Diagnostics | Every plan (free embedded model included; bring your own key optional) |
AI Diagnostics is the odd one out, deliberately
Every feature above is rule-based, meaning a threshold, a policy, a scoring heuristic you can read and predict. AI Diagnostics is the one exception, since it actually calls an LLM, a free local model included for every org by default, or your own key with your own choice of provider if you add one, to review the logs behind a firing alert, explain a likely root cause, and, only for a resolvable, confident, safe suggestion and only on the Enterprise plan, draft a dry-run playbook for Autonomous Remediation to evaluate. Everything else on this page works the same way whether or not you ever touch an LLM connection at all.
Why one platform instead of six point tools
An AI agent's cost problem and a Node's CPU problem are the same shape of problem, since something is consuming more than it should, and someone needs to know before the bill or the outage does. Splitting AI workloads into a separate observability stack means a separate dashboard, a separate alert queue, and a separate on-call handoff for exactly the teams who are already using KubeWatch for everything around the AI workload. These seven features exist so that boundary never has to happen.
Shared platform investments
Three pieces of infrastructure are shared across more than one feature:
- Loki, a log-oriented sibling to Mimir, holding AI agent spans/traces and other high-cardinality, text-heavy records that don't fit a time-series store.
- Extended KubeWatch Agent, new pluggable collectors (OTel span receiver, GPU metrics, inference-engine scraping, a reduced-footprint "Lite" build for edge hardware), so a deployment only runs what it needs.
- The Remediation Engine, where Agent Cost-and-Reliability Ops and GPU and Inference Cost Control each own a self-contained detect-decide-act loop of their own, since neither publishes an event for a separate engine to consume. Autonomous Remediation instead plugs into KubeWatch's original, general-purpose Alert Rules (CPU/memory/restarts/node-readiness/agent-silence), which is the real shared alert stream every other alert in the platform already flows through.
Why this order
Agent Cost-and-Reliability Ops and the Workload Advisor came first because they stand up the shared infrastructure (Loki, the extended Agent) everything after them leans on. Autonomous Remediation came last among the AI-specific features on purpose, since its value is a direct function of how many alert types and command actions already exist for it to act on, and that list kept growing as the others shipped.