active
KubeRescue
Kubernetes failure detection and auto-remediation
Autonomous Kubernetes failure detection and policy-driven auto-remediation engine for SRE teams.
- Go
- Kubernetes
- SRE
- Incident response
01 problem
Most Kubernetes auto-remediation tools act first and explain later, if at all. Restarting a crash-looping pod can destroy the evidence needed to understand it, restart a bare pod nobody can recreate, or loop forever. On-call engineers need a tool that understands a failure before touching anything, and that is honest about what it did.
02 approach
- 01 Safety over automation: pods without a controller are never deleted, every action is budgeted with
--max-restarts, and--dry-runis first-class. - 02 Explainability over magic: every action carries its evidence (restart count, exit code, termination reason, owner), and every skip carries its reason.
- 03 Truthful reporting: a dry run is never counted as a remediation, and report counters reflect what actually happened.
- 04 Detectors are pure functions over pod state with no API calls, so they are trivially testable. A shared
Findingtype is the contract between detection, remediation and reporting. - 05 Resilience: a transient API error degrades one scan, never the process; retries use capped exponential backoff.
03 architecture
-
01 · observe
-
02 · detect
-
03 · act
-
04 · report
trigger · monitor · --interval 30s
Engine.Scan
The engine lists pods in the target namespace (optionally filtered by a label selector) on each interval, or once with --once. It works with in-cluster config or a local kubeconfig.
All steps as text
-
1. observe
- Engine.Scan: The engine lists pods in the target namespace (optionally filtered by a label selector) on each interval, or once with --once. It works with in-cluster config or a local kubeconfig.
-
2. detect
- Detector.Detect(pod): Detectors are pure functions over pod state with no API calls. They emit a Finding carrying the evidence (container, restart count, last termination reason, exit code and owning controller).
- diagnose (read-only): kuberescue diagnose reuses the same detectors in a read-only path that never calls remediate. It adds Kubernetes events and a plain-language explanation for each finding.
-
3. act
- remediate.RestartPod: Only controller-managed pods are restarted, since bare pods are never deleted. Actions are capped by --max-restarts per scan, and with --dry-run nothing changes. Every outcome is recorded as restarted, dry-run, skipped or failed.
-
4. report
- Report: Human-readable text, or versioned JSON for automation. Reports go to stdout and structured logs to stderr. Exit code 2 signals findings, so it can gate a CI job.
In practice
Preview what KubeRescue would do. This changes nothing:
kuberescue monitor -n default --once --dry-run
Scan once and remediate, restarting at most three pods:
kuberescue monitor -n default --once --max-restarts 3
Every finding carries its evidence, so the output explains the action:
CrashLoopBackOff default/api-7c8f9f6d9b-x2q4m
container=api restarts=7 lastReason=OOMKilled exitCode=137 owner=ReplicaSet/api-7c8f9f6d9b
action: restarted
Summary: detected=1 restarted=1 skipped=0 failed=0
diagnose only reads pods, Deployments and events, and explains what it finds:
$ kuberescue diagnose -n default
OOMKilled default/api-7c8f9f6d9b-x2q4m container=api owner=ReplicaSet/api-7c8f9f6d9b
container "api" was killed for exceeding its memory limit; raise the limit or
investigate a possible memory leak.
event: BackOff x3 — Back-off restarting failed container api in pod ...
Summary: detected=1
For automation, the JSON report is versioned and its counters are truthful. A dry run counts as dry-run, never as restarted:
{
"schemaVersion": "v1alpha1",
"namespace": "default",
"dryRun": true,
"detected": 1,
"restarted": 0,
"findings": [
{
"pod": "api-7c8f9f6d9b-x2q4m",
"reason": "CrashLoopBackOff",
"restartCount": 7,
"lastTerminationReason": "OOMKilled",
"lastExitCode": 137,
"ownerKind": "ReplicaSet"
}
],
"actions": [{ "pod": "api-7c8f9f6d9b-x2q4m", "outcome": "dry-run" }]
}
05 outcome
-
kuberescue diagnoseexplains five failure classes (CrashLoopBackOff, OOMKilled, ImagePullBackOff, Pending/FailedScheduling and stuck rollouts) with the events behind each, entirely read-only. - Versioned JSON reports (
schemaVersion: v1alpha1) and CI-friendly exit codes:0clean,1error,2findings. - Runs in-cluster with namespaced RBAC (a Role, no ClusterRole), a non-root distroless image and
--dry-runon by default. - Released as prebuilt binaries with checksums and a container image on GHCR. Next milestones: a policy gate, an operator with CRDs, and Prometheus metrics.