Autonomous Remediation
Playbook-driven auto-response to your existing alerts: restart what's broken, safely.
Most teams already know exactly what they'd do the third time a pod crash-loops at 3am, which is restart it. Most of the delay between the alert firing and the fix landing is a human relaying that same obvious action. This feature builds the loop that skips the relay (detect, dry-run, execute with a clear audit trail) on top of KubeWatch's existing Alert Rules, as a shared substrate rather than one-off automation per feature.
What you get
- Playbooks let you attach an automated action to any existing alert rule (CPU %, memory %, pod restart count, node readiness, agent silence, VCS & CI/CD Observability's deploy-regression rule, or a custom rule). No separate condition taxonomy to learn, since a playbook just points at a rule you've already configured.
- Four actions are supported, meaning restart a single Docker container in place, trigger a Kubernetes rolling restart of the
owning Deployment, StatefulSet, or DaemonSet (the same mechanism as
kubectl rollout restart, a pod-template annotation bump applied via server-side apply), roll a Kubernetes Deployment back to its previous ReplicaSet revision (the same mechanism askubectl rollout undo), which is built specifically to close the loop on a deploy-regression alert, or bump a Kubernetes pod's memory limit for a repeated OOMKilled pattern (a single, bounded, automatic increase, applied at most once per resource). Restarting the workload instead of deleting the pod directly is what actually clears a stuck process, since otherwise Kubernetes would just recreate it identically. - Non-negotiable guardrails mean every new playbook starts in dry-run and can only reach live execution after
you explicitly promote it. A per-playbook circuit breaker also caps automated executions per hour, and
requires_human_approvaldefaults on for both action types until you turn it off. - A complete, append-only audit trail of every match (executed, deferred to a human, blocked by the circuit breaker, or dry-run) with the reasoning behind each one.
A real bug fix that came with this
Building this surfaced (and fixed) a real gap in Alert Rules that predates this feature, since CPU %/Memory % rules previously matched against metric names nothing ever wrote, so they silently never fired, and no alert ever carried enough identity to know which container or pod triggered it. Both are now fixed. Alerts carry real resource identity, and a rule now fires once per breaching resource instead of only the first one ever seen.
Why the guardrails aren't configurable away
Every new playbook is dry-run-only until its first match has been reviewed by a human, with no exceptions, including playbooks you'd consider urgent. The scenario this feature is most likely to get wrong is many alerts firing at once during a real incident, which triggers many playbooks simultaneously. The circuit breaker's rolling-window budget exists specifically for that case.
What it doesn't do
Restart actions have no rollback, since there's no meaningful "undo" for a restart, only a repeat of it. This is called out explicitly in the audit trail rather than implying an undo path that doesn't exist.
The Deployment rollback action's automatic trigger currently only resolves for an Argo CD Application managing exactly one Deployment, the only source that reliably names a specific namespace and Deployment today. Flux, Tekton, and CI/CD-provider deploys still fire the deploy-regression alert normally, but a playbook attached to that rule just shows "does not apply" for those until a resolvable target exists. The rollback itself also has an honest limit, since it reverts the Deployment's pod template via Kubernetes' own revision history, not a database migration or any other side effect the bad deploy may have caused.
Playbooks also don't invent responses to alert types they aren't attached to. An unmatched alert simply does nothing differently than it does today.