Self-Hosted AI Stack Observability
Monitor local inference engines, model routers, and sovereign fallback readiness.
Running your own inference stack instead of calling out to a hosted API means the usual observability gaps of self-hosting apply to your models too, since no one's dashboard shows you vLLM's queue depth or whether your fallback model actually loads if the primary one goes down. This feature extends KubeWatch's existing self-hosted deployment option into the AI stack itself, covering local inference engines, model routers, and the fallback path an organization keeps ready for exactly this kind of disruption.
What you get
- Self-hosted inference engines treated as a first-class monitored workload type (vLLM, Ollama, SGLang, TGI), scraped by the KubeWatch Agent the same way container-level metrics are collected today, visible on the AI Observability page.
- Model Registry shows which models are deployed where, at what version, behind which router, on the Model Fleet page. Register an endpoint's engine, role (primary or fallback), and version hash.
- Sovereign Fallback Health Checker is a synthetic probe that runs from the KubeWatch Agent itself (not the backend, which generally can't reach a cluster-internal endpoint) and periodically sends a minimal real request (an OpenAI-compatible chat completion, or Ollama's generate endpoint) to each configured model, turning "we have a fallback" into "we know our fallback works, checked five minutes ago." Every check's result, latency, and any error is kept in an append-only history per model.
What it doesn't do
This feature observes and health-checks routing infrastructure. It does not replace a model router, and it does not evaluate model quality (accuracy, hallucination rate), which is an eval-tooling concern.