Troubleshoot missing or delayed alerts
Follow the path from observation to delivery. Change one setting at a time so you can tell which step fixed the problem.
1. Check the active leader
The default installation runs one active leader and one standby. A standby can
be healthy without being ready to monitor. Do not use its /readyz response
to decide that the whole installation is broken.
Start with the managed workload:
kubectl get pods -n kwatch -l app=kwatch
kubectl logs -n kwatch deployment/kwatch
If you need endpoint details, port-forward the Deployment and inspect health:
kubectl port-forward -n kwatch deployment/kwatch 8060:8060
curl http://localhost:8060/healthz
curl http://localhost:8060/readyz
curl http://localhost:8060/health
/healthz reports process liveness. /readyz succeeds only on an active
leader whose required state and informer sources are ready. /health exposes
safe leadership and degradation reason codes. See
production readiness for the full
endpoint contract.
2. Confirm that kwatch can see the source
Check that the monitor is enabled and that the affected namespace is in scope. A missing or unavailable lister causes kwatch to skip detection; it does not create a synthetic incident. Source and permission problems appear in health or logs.
If the failure is a Pod crash, inspect the same Pod directly:
kubectl describe pod -n production <pod-name>
kubectl logs -n production <pod-name> --previous
For node, network, or storage findings, inspect the corresponding Kubernetes resource and confirm the monitor's coverage. If access is denied, review the RBAC guide before changing permissions.
3. Check scope and suppression
If kwatch sees the condition, look for a deliberate decision to stay quiet:
- Namespace or reason filters may exclude the resource.
- A silence or maintenance marker may suppress the finding.
- Startup baseline can summarize pre-existing problems instead of sending individual alerts.
- Cooldown, node inhibition, mass-failure suppression, cascading suppression, or grouping may fold related symptoms into another incident.
Review the configuration guide for these
controls. The stable audit skip reasons include baseline,
node_inhibition, mass_failure, cascading_suppression, and cooldown.
Check the corresponding audit record before removing a silence or changing a
threshold.
4. Separate incident creation from delivery
If an incident exists but no message arrived, check the selected provider, its Secret-backed credentials, routes, rate limits, retries, and any fallback. Provider errors may be permanent, retryable, or rate limited. Increasing retry counts will not repair an invalid credential or destination.
Use kwatch lint after a configuration change. kwatch lint --check can
test credentials for providers that support checks. Protected diagnostics such
as /incidents, /deadletters, and /test-alert are disabled unless you
enable them and supply a bearer token. See production readiness
before exposing diagnostics.
5. Verify recovery
After fixing the source, configuration, or provider, wait for the active leader to become ready and check a real incident or a supported test notification. Confirm that the expected destination received it. If the leader changed, review replication and failover for state restore and monitoring-gap behavior.