GPU and Inference Cost Control
Real-time GPU utilization, cost-per-token attribution, and GPU-aware autoscaling.
**Available now**, on the Pro plan and above. See the **GPU cost control** section on the **AI Observability**
page, and GPU-aware policies on the **Auto Scaling** page.
A GPU sitting at 30% utilization still bills at 100%, and unlike a CPU box you can't just right-size it after the fact without re-provisioning hardware or renegotiating a cloud commitment. This feature extends KubeWatch's existing autoscaler policy engine to be GPU- and cost-aware, and gives GPU spend the same live visibility as CPU/memory cost today, on the same dashboard, not a separate GPU-cost tool with its own login.
What you get
- Real-time GPU utilization, memory, temperature, and power visibility, at the same fidelity as existing
CPU/memory metrics, read via
nvidia-smion the KubeWatch Agent, no separate collector to deploy. - GPU pricing and cost attribution lets you enter an hourly rate per GPU type (cloud rate card or an amortized estimate for owned hardware); KubeWatch then computes live fleet cost and a cost-per-1k-tokens figure, joining GPU-hour cost to actual token volume from Agent Cost-and-Reliability Ops.
- GPU-aware autoscaling policies, on the same Auto Scaling page as CPU/memory policies, using the same
dry-run/rollback conventions:
- Scale to zero on idle stops paying for GPUs behind an inference workload with no traffic for a configurable idle window. Works for both Kubernetes and KubeWatch-managed Docker.
- Rightsize adjusts a Deployment's
nvidia.com/gpuresource request/limit toward a target utilization (Kubernetes only). - Spot shift renders a Karpenter NodePool pair (spot-preferred, on-demand fallback) sized to your chosen max spot fraction (Kubernetes only).
GPU utilization is a fleet-wide signal, not a per-pod one. `nvidia-smi` reports per physical device, with no
native pod attribution. Policies key on GPU type as an org-wide signal and act on the policy's own target
workload when triggered. That's a deliberate v1 simplification, not a bug.
Out of scope
Model-level optimization (quantization, distillation), multi-cloud GPU price arbitrage, and automatic provider-pricing-API sync (pricing is manual entry) are explicitly not part of this feature.