Agent Cost-and-Reliability Ops

Bring your AI agents into KubeWatch as a first-class monitored entity, with cost and reliability visibility and policy rules that evaluate continuously.

**Available now**, on the Pro plan and above. See the **AI Agents** page in your dashboard. Cost/reliability monitoring and policy evaluation are fully live; automated enforcement (throttle/kill/rollback actually acting) is landing next, see the note below.

Customers running AI agents, whether that's LangChain, CrewAI, or a custom system built on the Claude or OpenAI SDKs, previously had no infrastructure-level visibility into them inside KubeWatch. They showed up as a generic container process, indistinguishable from any other workload. This feature brings AI agents into KubeWatch as a first-class monitored entity, alongside pods and containers, with the same alerting machinery the platform's autoscaler already uses, and the same policy-evaluation groundwork its automated response plugs into.

**Terminology**, "KubeWatch Agent" is KubeWatch's own telemetry daemon, the same one that reports on your containers and nodes. "AI agent" is your LLM-based autonomous system, the thing actually being monitored. They're never the same object, even though the word "agent" applies to both.

What you get

  • AI Agents page lists every registered AI agent process, with cost/hour, error rate, and status at a glance, the same visual density as the existing Pods list.
  • Session drill-down is a waterfall view of a single session's LLM calls, tool calls, and retrievals.
  • Policies define automated responses to cost, error-rate, or runaway-loop conditions, either throttle (cap concurrent sessions), kill (terminate in-flight sessions), or rollback (revert to a previous agent version). Every policy evaluation and decision is logged regardless of arming state, so you get a full audit trail of what the engine would have done even before enabling enforcement.
**Enforcement is not yet wired up in this release.** A policy's condition-matching, dry-run evaluation, and decision logging all run for real. What doesn't yet happen is the KubeWatch Agent actually throttling, killing, or rolling back the AI agent process. An armed policy that matches records the action it *would* take and returns a clear "not implemented" failure rather than a false success, since it will never silently pretend to have acted. Until real enforcement ships, treat this feature as monitoring and alerting on AI agent cost/reliability, not as automated remediation.

Safety

  • Every new policy starts in forced dry-run for at least 30 minutes before it can be armed. The engine evaluates and logs what it would do, applying nothing until you explicitly arm it. This is enforced server-side, not just hidden in the UI.
  • The UI requires typing the agent's name to confirm before arming a kill policy, in anticipation of enforcement landing, since killing a session will be irreversible once it's real.

Getting data in

Point your AI agent framework's OpenTelemetry exporter at the KubeWatch Agent running alongside it (same pod, or same host for non-containerized deployments). Framework coverage starts with LangChain and raw OTel, the broadest existing instrumentation base, with more frameworks added based on demand.

Cost accuracy depends on your model pricing staying current. A stale price list produces confidently wrong cost dashboards, which is worse than no cost dashboard at all. Review your pricing periodically.