Skip to main content

Feature reference

This page is generated from the Kwatch feature catalog. Feature IDs are stable product vocabulary and should not be renamed without a migration.

Feature IDLifecycleDescriptionDependencies
core.detection.podsruntimeDetect pod and container failuresnone
core.pods.schedulingruntimeDetect pod scheduling failurescore.detection.pods
core.pods.pendingruntimeDetect pods stuck Pendingcore.detection.pods
core.pods.oomruntimeDetect repeated out-of-memory failurescore.detection.pods
core.pods.readinessruntimeDetect sustained pod readiness failurescore.detection.pods
core.pods.restartsruntimeDetect excessive container restartscore.detection.pods
core.detection.workloadsruntimeDetect workload rollout and execution failuresnone
core.workloads.deployment-rolloutruntimeDetect stuck Deployment rolloutscore.detection.workloads
core.workloads.statefulset-rolloutruntimeDetect stuck StatefulSet rolloutscore.detection.workloads
core.workloads.daemonset-rolloutruntimeDetect stuck DaemonSet rolloutscore.detection.workloads
core.workloads.job-failuresruntimeDetect failed and suspended Jobscore.detection.workloads
core.workloads.cronjob-failuresruntimeDetect failed and missed CronJobscore.detection.workloads
core.workloads.pdb-violationsruntimeDetect PodDisruptionBudget violationscore.detection.workloads
core.workloads.hpa-diagnosticsruntimeDiagnose HorizontalPodAutoscaler failurescore.detection.workloads
core.detection.nodesruntimeDetect node readiness and resource failuresnone
core.nodes.conditionsruntimeDetect node conditions and lifecycle failurescore.detection.nodes
core.nodes.resourcesruntimeDetect node resource pressurecore.detection.nodes
core.detection.storageruntimeDetect persistent storage failuresnone
core.storage.pvc-usageruntimeDetect PVC usage and volume failurescore.detection.storage
core.detection.networkruntimeDetect service and network failuresnone
core.network.service-endpointsruntimeDetect Service and EndpointSlice failurescore.detection.network
core.network.ingress-backendsruntimeDetect Ingress backend failurescore.detection.network
core.network.policiesruntimeDetect NetworkPolicy failurescore.detection.network
core.detection.securityruntimeDetect security and admission failuresnone
core.security.admission-webhooksruntimeDetect admission webhook failurescore.detection.security
core.cluster-resources.statusruntimeDetect cluster resource status failuresnone
core.security.tlsruntimeDetect TLS certificate expirycore.detection.security
intelligence.diagnosis.directruntimeExplain the most likely direct causenone
intelligence.diagnosis.dependency-graphruntimeTrace related Kubernetes dependenciesintelligence.diagnosis.direct
intelligence.diagnosis.impactruntimeEstimate affected resources and blast radiusintelligence.diagnosis.dependency-graph
intelligence.diagnosis.change-diffruntimeRelate incidents to recent changesintelligence.diagnosis.direct
intelligence.diagnosis.timelineruntimeKeep a compact incident timelinenone
intelligence.diagnosis.confidenceruntimeShow confidence and supporting evidenceintelligence.diagnosis.direct
intelligence.diagnosis.feedbackruntimePersist operator feedback for RCA improvementintelligence.diagnosis.direct
incidents.lifecycle.cooldownruntimeSuppress repeated notifications during cooldownnone
incidents.lifecycle.groupingruntimeGroup related incidents into one narrativenone
incidents.lifecycle.mass-failureruntimeReduce noise during broad failuresincidents.lifecycle.grouping
incidents.lifecycle.cascade-suppressionruntimeSuppress symptoms after a root cause is knownintelligence.diagnosis.dependency-graph
incidents.persistence.activestartupRestore active incident lifecycle after restartnone
incidents.persistence.baselinestartupPersist startup baseline statenone
incidents.persistence.change-historyruntimePersist recent change historynone
telemetry.kubelet.summaryruntimeRead built-in kubelet summary telemetrynone
telemetry.cpu.usageruntimeDetect CPU usage pressuretelemetry.kubelet.summary
telemetry.cpu.throttlingruntimeDetect container CPU throttlingtelemetry.kubelet.summary
telemetry.memory.usageruntimeDetect memory pressure and overusetelemetry.kubelet.summary
telemetry.storage.usageruntimeDetect ephemeral storage and inode pressuretelemetry.kubelet.summary
telemetry.pressureruntimeDetect cgroup pressure signalstelemetry.kubelet.summary
telemetry.network.errorsruntimeDetect kubelet-observed network errorstelemetry.kubelet.summary
telemetry.runtime.errorsruntimeDetect container runtime error ratestelemetry.kubelet.summary
telemetry.metrics-apiruntimeRead the optional Kubernetes metrics APInone
telemetry.adaptive-baselineruntimeAdapt bounded thresholds to observed usagenone
probes.httpruntimeRun configured HTTP checksnone
probes.tcpruntimeRun configured TCP checksnone
probes.dnsruntimeRun configured DNS checksnone
probes.services.automaticruntimeDerive safe probe targets from servicesnone
probes.latencyruntimeDetect probe latency regressionsnone
control-plane.podsruntimeObserve control-plane component podsnone
control-plane.api.healthruntimeCheck Kubernetes API health endpointsnone
control-plane.api.latencyruntimeMeasure Kubernetes API latencycontrol-plane.api.health
control-plane.schedulerruntimeObserve scheduler healthcontrol-plane.pods
control-plane.controller-managerruntimeObserve controller-manager healthcontrol-plane.pods
control-plane.etcdruntimeObserve etcd health signalscontrol-plane.api.health
cluster-resources.statusruntimeObserve status conditions on cluster resourcescore.cluster-resources.status
cluster-resources.crd-discoverystartupDiscover supported custom resources dynamicallynone
security.rbac.auditruntimeReport missing permissions and RBAC driftnone
security.tlsruntimeMonitor configured Kubernetes TLS secretsnone
security.audit-logruntimeWrite structured incident audit recordsnone
delivery.escalationruntimeEscalate incidents through alert tiersnone
delivery.templatesruntimeRender operator-selected alert templatesnone
delivery.runbooksruntimeAttach reason-aware runbook linksnone