You can monitor the fencing health of your two-node OKD cluster to prevent failures that could impact your workloads. The cluster provides metrics and alerts that tell you when critical issues occur, such as nodes going offline or fence devices becoming unreachable.
A two-node OKD cluster with fencing (TNF) uses Pacemaker to manage cluster membership, fencing, and the etcd resource. The cluster exposes this Pacemaker state as Prometheus metrics and ships a set of alerts that fire on TNF-specific failure conditions.
You can monitor the fencing layer of your two-node OKD cluster by using the same observability tools you use for the rest of the cluster.
Integrated Prometheus metrics and alerts for Pacemaker help you detect node failures, fence device issues, and resource problems before they compromise high availability.
TNF monitoring reflects the state of the Pacemaker cluster:
Pacemaker cluster membership and expected node count: This is not the etcd raft membership, which fluctuates by design during fencing and recovery.
Node availability, maintenance, and standby state: This is distinct from the state of OKD nodes.
Fence device health for each node.
The state of Pacemaker-managed resources, such as etcd, as reported by the resource agent.
Every TNF metric name is prefixed with tnf_ and is grouped by level: tnf_cluster_*, tnf_node_*, and tnf_resource_*. Each metric is a gauge that mirrors a Pacemaker condition, where 1 means the condition is true (the healthy state) and 0 means the condition is not true.
TNF alerts are delivered as part of the cluster Etcd Operator alerting rules, in a tnf-pacemaker.rules group that is separate from the standard etcd alerts. The alerts fire only on TNF. Each alert targets a specific leaf condition, such as a node being offline or a fence device being unavailable, so that the firing alert names the actionable problem rather than a generic cluster unhealthy rollup.
You do not need to configure anything to enable TNF metrics and alerts. They are active automatically on a cluster that uses the TNF topology.
You can build queries and dashboards to monitor health of your two-node OKD cluster with fencing (TNF) by using Prometheus metrics. These gauge metrics track cluster health, node status, and Pacemaker resource state to help you identify and respond to availability issues.
The cluster-level metrics describe the cluster as a whole and have no labels.
| Metric | Description |
|---|---|
|
Overall cluster health. |
|
Whether the cluster is in service. |
|
Whether the cluster has the expected node count. |
The node-level metrics describe an individual node and carry a node label set to the node name:
| Metric | Description |
|---|---|
|
Overall node health. |
|
Whether the node is online. |
|
Whether the node is in service. |
|
Whether the node is active. |
|
Whether the node is ready. |
|
Whether the node is in a clean state. |
|
Whether the node is a cluster member. |
|
Whether at least one fence device for the node is healthy. |
|
Whether all fence devices for the node are healthy. |
The resource-level metrics describe a Pacemaker-managed resource, such as etcd, on a specific node. They carry a node label and a resource label.
| Metric | Description |
|---|---|
|
Overall resource health. |
|
Whether the resource is in service. |
|
Whether the resource is managed by Pacemaker. |
|
Whether the resource is enabled. |
|
Whether the resource is operational. |
|
Whether the resource is active. |
|
Whether the resource is started. |
|
Whether the resource can be scheduled. |
A two-node OKD cluster with fencing (TNF) ships several alerts in the tnf-pacemaker.rules group.
critical alertsIndicate a condition that breaks cluster functionality and requires immediate action.
warning alertsIndicate a degraded state that needs attention while the cluster remains functional.
Each alert includes a runbook_url annotation that links to detailed investigation and remediation steps.
The cluster-level alerts describe the cluster as a whole and have no labels.
| Alert | Severity | Fires when | Runbook |
|---|---|---|---|
|
|
The cluster does not have exactly two nodes, for example after a node is added or removed. Restore the cluster to two control plane nodes. |
|
|
|
The cluster is in maintenance mode. Expected during planned maintenance; investigate if unexpected. |
The node-level alerts carry a node label that identifies the affected node.
| Alert | Severity | Fires when | Runbook |
|---|---|---|---|
|
|
The node is offline because of a node failure, reboot, or network partition. The cluster has lost high-availability redundancy until the node recovers. |
|
|
|
All fence devices for the node are unhealthy, so the node cannot be fenced. Common causes are an unreachable BMC or invalid credentials. Restore fencing before relying on the cluster for high availability. |
|
|
|
Fence device redundancy is lost, but at least one device remains healthy. The node can still be fenced. Repair the failed device. |
|
|
|
Pacemaker cannot confirm the node state, for example after a fencing, communication, or configuration problem. |
|
|
|
The node is in maintenance mode. Expected during planned maintenance; investigate if unexpected. |
|
|
|
The node is in standby mode and is not running resources. Expected during planned maintenance; investigate if unexpected. |
The resource-level alerts carry a node label and a resource label, for example resource="etcd".
| Alert | Severity | Fires when | Runbook |
|---|---|---|---|
|
|
A Pacemaker-managed resource, such as |
|
|
|
A resource has failed and cannot start, for example because of a resource agent failure or a configuration error. |
|
|
|
A resource is not managed by Pacemaker, typically after a manual unmanage operation. Pacemaker does not recover an unmanaged resource. |
|
|
|
A resource is disabled, typically after a manual disable operation. |