Skip to main content

Kubernetes incidents, explained

See what broke.Understand why.Know what to do next.

kwatch turns Kubernetes failures into clear alerts with the likely cause, useful evidence, and a practical next step.

Open sourceRuns in your clusterNo hosted account
Example incidentHigh severity

OOMKilled

production / orders-api

Pod: orders-api-7ffc9d4f9-x9p4t · Node: worker-3

Likely cause

The container exceeded its 512Mi memory limit.

Next step

Increase limits.memory or reduce memory usage.

Recent logs and Kubernetes events add context to the alert.

From signal to action

Understand the incident without piecing it together yourself

Kubernetes shows symptoms. kwatch connects the story so your team can decide what needs attention.

01

Detect the problem

Watch for crashes, stuck workloads, unhealthy nodes, and other cluster signals.

02

Connect the clues

Bring together status, recent logs, Kubernetes events, and affected resources.

03

Send a useful alert

Give responders a likely cause and a next step in the channel they already use.

Coverage

Start with common failures. Add checks as you grow.

Safe defaults cover everyday incidents. Heartbeat, Metrics Server usage, TLS checks, and active probes are available when you need them.

Pods and scheduling

Crashes, OOM kills, restarts, readiness, and pending Pods.

Workloads

Rollouts, Jobs, CronJobs, autoscaling, and availability.

Infrastructure and storage

Node pressure, persistent storage, and platform health.

Networking and security

Services, Ingress, webhooks, TLS, RBAC, and policy findings.

Get started

From command to useful alerts.

The interactive kwatch.sh manager guides installation, channel setup, and verification. You need Bash, curl, kubectl, and cluster install permissions.

Install kwatchRecommended

Run the manager on a machine with access to your cluster.

/bin/bash -c "$(curl -fsSL https://kwatch.dev/kwatch.sh)"

Run it again to change settings, upgrade, check status, or uninstall.

01

Choose your cluster

The manager shows the current kubectl context before installing.

02

Connect a channel

It stores credentials in a Kubernetes Secret.

03

Verify the install

It checks the deployment before you start monitoring.

Read the installation guide
Built for production

Two replicas by default: one active leader and one standby. A single-replica option is available without kwatch self-failover.

How failover works

Notifications

Send alerts where your team works.

Connect a familiar destination first. Choose from 56 integrations when your team needs more routes.

See all channels