GPU and Inference Cost Control

Real-time GPU utilization, cost-per-token attribution, and GPU-aware autoscaling.

**Available now**, on the Pro plan and above. See the **GPU cost control** section on the **AI Observability** page, and GPU-aware policies on the **Auto Scaling** page.

A GPU sitting at 30% utilization still bills at 100%, and unlike a CPU box you can't just right-size it after the fact without re-provisioning hardware or renegotiating a cloud commitment. This feature extends KubeWatch's existing autoscaler policy engine to be GPU- and cost-aware, and gives GPU spend the same live visibility as CPU/memory cost today, on the same dashboard, not a separate GPU-cost tool with its own login.

What you get

  • Real-time GPU utilization, memory, temperature, and power visibility, at the same fidelity as existing CPU/memory metrics, read via nvidia-smi on the KubeWatch Agent, no separate collector to deploy.
  • GPU pricing and cost attribution lets you enter an hourly rate per GPU type (cloud rate card or an amortized estimate for owned hardware); KubeWatch then computes live fleet cost and a cost-per-1k-tokens figure, joining GPU-hour cost to actual token volume from Agent Cost-and-Reliability Ops.
  • GPU-aware autoscaling policies, on the same Auto Scaling page as CPU/memory policies, using the same dry-run/rollback conventions:
    • Scale to zero on idle stops paying for GPUs behind an inference workload with no traffic for a configurable idle window. Works for both Kubernetes and KubeWatch-managed Docker.
    • Rightsize adjusts a Deployment's nvidia.com/gpu resource request/limit toward a target utilization (Kubernetes only).
    • Spot shift renders a Karpenter NodePool pair (spot-preferred, on-demand fallback) sized to your chosen max spot fraction (Kubernetes only).
GPU utilization is a fleet-wide signal, not a per-pod one. `nvidia-smi` reports per physical device, with no native pod attribution. Policies key on GPU type as an org-wide signal and act on the policy's own target workload when triggered. That's a deliberate v1 simplification, not a bug.

Out of scope

Model-level optimization (quantization, distillation), multi-cloud GPU price arbitrage, and automatic provider-pricing-API sync (pricing is manual entry) are explicitly not part of this feature.