Agent Cost-and-Reliability Ops
Bring your AI agents into KubeWatch as a first-class monitored entity, with cost and reliability visibility and policy rules that evaluate continuously.
Customers running AI agents, whether that's LangChain, CrewAI, or a custom system built on the Claude or OpenAI SDKs, previously had no infrastructure-level visibility into them inside KubeWatch. They showed up as a generic container process, indistinguishable from any other workload. This feature brings AI agents into KubeWatch as a first-class monitored entity, alongside pods and containers, with the same alerting machinery the platform's autoscaler already uses, and the same policy-evaluation groundwork its automated response plugs into.
What you get
- AI Agents page lists every registered AI agent process, with cost/hour, error rate, and status at a glance, the same visual density as the existing Pods list.
- Session drill-down is a waterfall view of a single session's LLM calls, tool calls, and retrievals.
- Policies define automated responses to cost, error-rate, or runaway-loop conditions, either throttle (cap concurrent sessions), kill (terminate in-flight sessions), or rollback (revert to a previous agent version). Every policy evaluation and decision is logged regardless of arming state, so you get a full audit trail of what the engine would have done even before enabling enforcement.
Safety
- Every new policy starts in forced dry-run for at least 30 minutes before it can be armed. The engine evaluates and logs what it would do, applying nothing until you explicitly arm it. This is enforced server-side, not just hidden in the UI.
- The UI requires typing the agent's name to confirm before arming a kill policy, in anticipation of enforcement landing, since killing a session will be irreversible once it's real.
Getting data in
Point your AI agent framework's OpenTelemetry exporter at the KubeWatch Agent running alongside it (same pod, or same host for non-containerized deployments). Framework coverage starts with LangChain and raw OTel, the broadest existing instrumentation base, with more frameworks added based on demand.