Skip to main content

5 posts tagged with "alerting"

View All Tags

Reduce Kubernetes alert noise without hiding incidents

· 3 min read
kwatch contributors
Project documentation

An alert is useful when it changes what the on-call person does. Repeated notifications for one failure, or dozens of symptoms from one shared problem, make that decision harder. Muting everything is risky too: a real incident can disappear with the noise.

kwatch has several controls for different kinds of noise. Choose the one that matches the problem you are trying to solve.

Send Kubernetes incident alerts to PagerDuty with kwatch

· 2 min read
Amgad Ramses
Maintainer of kwatch

PagerDuty is a good destination when a Kubernetes failure needs an on-call response. kwatch adds the Kubernetes context before it sends the incident: the affected workload, likely cause, recent evidence, and what to investigate next.

The screenshots are from the original guide. Check the current PagerDuty channel guide before configuring a new integration.

Detect Kubernetes crashes with kwatch and Slack

· 3 min read
Andrew Attallah
Maintainer of kwatch

When a Kubernetes workload restarts, a Slack notification is useful only if it helps you decide what to do next. kwatch turns crash signals, events, and recent logs into a focused incident message.

This guide shows how to connect Slack and install kwatch with the supported interactive manager. The manager keeps the webhook in a Kubernetes Secret; do not paste it directly into config.yaml.

The screenshots show the original Slack setup flow. Use the current Slack guide for supported settings and delivery options.

What is kwatch? Kubernetes incidents, explained

· 2 min read
Abdelrahman Ahmed
Owner of kwatch

See what broke. Understand why. Know what to do next. 👀🧠⚡

Kubernetes gives teams a powerful way to run applications, but a failed pod or stalled rollout can leave you with a short status such as CrashLoopBackOff and not enough context to act quickly.

kwatch watches Kubernetes resources, events, and recent logs. When something needs attention, it groups related symptoms and sends an alert with the likely cause, impact, evidence, and a practical next step.