Edge AI Fleet Observability

Fleet visibility and OTA model updates for edge AI devices, over the same long-poll channel built for the autoscaler.

**Available now**, on the Pro plan and above. See the **Edge Fleet** page in your dashboard.

As inference moves off centralized cloud infrastructure and onto cameras, sensors, and industrial controllers, fleet management is becoming its own category, currently led by hardware and IoT vendors rather than developer-facing observability tools. This feature is a strong architectural fit specifically because KubeWatch's Command Channel (agent-initiated long-poll, built because agents have no inbound command-receive path) solves the exact problem edge devices behind NAT and firewalls present, only more so than the original Kubernetes/Docker case it was designed for.

What you get

  • An edge mode for the existing KubeWatch Agent (KUBEWATCH_MODE=edge), no separate binary, just a leaner collection path with no container/pod/node discovery, plus store-and-forward buffering, where telemetry queues locally through a connectivity gap and flushes on reconnect, timestamped by when it actually happened rather than when the device finally reported it.
  • Fleet visibility as a normal agent registration, extended with location, hardware profile, and deployed model version. No separate device registry to keep in sync.
  • Three-state connectivity, not two, meaning online, offline (expected), and offline (unexpected) read differently on purpose. A factory-floor camera that goes quiet every night during a maintenance window reads calmly (gray), while a device that's actually gone dark reads as an alert (red). Configure expected offline windows per device (recurring, by day of week and UTC time range) so a real outage never gets lost in nightly noise, and a nightly window never trains you to ignore a real one.
  • OTA model updates, via the same Command Channel every other action on the platform uses, with a 24-hour TTL instead of the usual couple of minutes, since a device might only poll once a day.

Why model updates call out to your own script

KubeWatch has no universal way to know how your inference stack is actually deployed on a given piece of edge hardware. There's no Kubernetes API or Docker socket to reuse the way a cluster or container-host restart does, so an update command invokes a script you provide (KUBEWATCH_MODEL_UPDATE_HOOK), passing the target version as its argument, and reports back whatever that script reports. This is a real, working trigger for teams who wire up their own apply logic, not KubeWatch guessing at a model-swap mechanism it's never seen.

What it doesn't do

This feature assumes a device has already been enrolled by some existing process. It observes and commands enrolled devices, registering with a normal agent API key just like any Docker host or Kubernetes node, but it doesn't replace an MDM/provisioning system. It's also not designed for real-time (sub-second) control loops on-device. The existing push/long-poll cadence fits fleet observability and configuration, not latency-sensitive on-device control logic.