Cluster checkup framework - Monitoring | Virtualization

Running predefined latency checkups
- Running a latency checkup by using the web console
- Running a latency checkup by using the CLI
Running predefined storage checkups
Additional resources

A checkup is an automated test workload that allows you to verify if a specific cluster functionality works as expected. The cluster checkup framework uses native Kubernetes resources to configure and execute the checkup.

The OKD Virtualization cluster checkup framework is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.

For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.

As a developer or cluster administrator, you can use predefined checkups to improve cluster maintainability, troubleshoot unexpected behavior, minimize errors, and save time. You can review the results of the checkup and share them with experts for further analysis. Vendors can write and publish checkups for features or services that they provide and verify that their customer environments are configured correctly.

Running predefined latency checkups

You can use a latency checkup to verify network connectivity and measure latency between two virtual machines (VMs) that are attached to a secondary network interface. The predefined latency checkup uses the ping utility.

Before you run a latency checkup, you must first create a bridge interface on the cluster nodes to connect the VM’s secondary interface to any interface on the node. If you do not create a bridge interface, the VMs do not start and the job fails.

Running a predefined checkup in an existing namespace involves setting up a service account for the checkup, creating the Role and RoleBinding objects for the service account, enabling permissions for the checkup, and creating the input config map and the checkup job. You can run a checkup multiple times.

You must always:

Verify that the checkup image is from a trustworthy source before applying it.
Review the checkup permissions before creating the Role and RoleBinding objects.

Running a latency checkup by using the web console

Run a latency checkup to verify network connectivity and measure the latency between two virtual machines attached to a secondary network interface.

Prerequisites

You must add a NetworkAttachmentDefinition to the namespace.

Procedure

Navigate to Virtualization → Checkups in the web console.
Click the Network latency tab.
Click Install permissions.
Click Run checkup.
Enter a name for the checkup in the Name field.
Select a NetworkAttachmentDefinition from the drop-down menu.
Optional: Set a duration for the latency sample in the Sample duration (seconds) field.
Optional: Define a maximum latency time interval by enabling Set maximum desired latency (milliseconds) and defining the time interval.
Optional: Target specific nodes by enabling Select nodes and specifying the Source node and Target node.
Click Run.

Verification

To view the status of the latency checkup, go to the Checkups list on the Latency checkup tab. Click on the name of the checkup for more details.

Running a latency checkup by using the CLI

You run a latency checkup using the CLI by performing the following steps:

Create a service account, roles, and rolebindings to provide cluster access permissions to the latency checkup.
Create a config map to provide the input to run the checkup and to store the results.
Create a job to run the checkup.
Review the results in the config map.
Optional: To rerun the checkup, delete the existing config map and job and then create a new config map and job.
When you are finished, delete the latency checkup resources.

Prerequisites

You installed the OpenShift CLI (oc).
The cluster has at least two worker nodes.
You configured a network attachment definition for a namespace.

Procedure

Create a ServiceAccount, Role, and RoleBinding manifest for the latency checkup:

Example role manifest file

---
apiVersion: v1
kind: ServiceAccount
metadata:
  name: vm-latency-checkup-sa
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: kubevirt-vm-latency-checker
rules:
- apiGroups: ["kubevirt.io"]
  resources: ["virtualmachineinstances"]
  verbs: ["get", "create", "delete"]
- apiGroups: ["subresources.kubevirt.io"]
  resources: ["virtualmachineinstances/console"]
  verbs: ["get"]
- apiGroups: ["k8s.cni.cncf.io"]
  resources: ["network-attachment-definitions"]
  verbs: ["get"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: kubevirt-vm-latency-checker
subjects:
- kind: ServiceAccount
  name: vm-latency-checkup-sa
roleRef:
  kind: Role
  name: kubevirt-vm-latency-checker
  apiGroup: rbac.authorization.k8s.io
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: kiagnose-configmap-access
rules:
- apiGroups: [ "" ]
  resources: [ "configmaps" ]
  verbs: ["get", "update"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: kiagnose-configmap-access
subjects:
- kind: ServiceAccount
  name: vm-latency-checkup-sa
roleRef:
  kind: Role
  name: kiagnose-configmap-access
  apiGroup: rbac.authorization.k8s.io

Apply the ServiceAccount, Role, and RoleBinding manifest:
```
$ oc apply -n <target_namespace> -f <latency_sa_roles_rolebinding>.yaml (1)
```
1 <target_namespace> is the namespace where the checkup is to be run. This must be an existing namespace where the NetworkAttachmentDefinition object resides.

Create a ConfigMap manifest that contains the input parameters for the checkup:

Example input config map

apiVersion: v1
kind: ConfigMap
metadata:
  name: kubevirt-vm-latency-checkup-config
  labels:
    kiagnose/checkup-type: kubevirt-vm-latency
data:
  spec.timeout: 5m
  spec.param.networkAttachmentDefinitionNamespace: <target_namespace>
  spec.param.networkAttachmentDefinitionName: "blue-network" (1)
  spec.param.maxDesiredLatencyMilliseconds: "10" (2)
  spec.param.sampleDurationSeconds: "5" (3)
  spec.param.sourceNode: "worker1" (4)
  spec.param.targetNode: "worker2" (5)

1	The name of the `NetworkAttachmentDefinition` object.
2	Optional: The maximum desired latency, in milliseconds, between the virtual machines. If the measured latency exceeds this value, the checkup fails.
3	Optional: The duration of the latency check, in seconds.
4	Optional: When specified, latency is measured from this node to the target node. If the source node is specified, the `spec.param.targetNode` field cannot be empty.
5	Optional: When specified, latency is measured from the source node to this node.

Apply the config map manifest in the target namespace:

$ oc apply -n <target_namespace> -f <latency_config_map>.yaml

Create a Job manifest to run the checkup:

Example job manifest

apiVersion: batch/v1
kind: Job
metadata:
  name: kubevirt-vm-latency-checkup
  labels:
    kiagnose/checkup-type: kubevirt-vm-latency
spec:
  backoffLimit: 0
  template:
    spec:
      serviceAccountName: vm-latency-checkup-sa
      restartPolicy: Never
      containers:
        - name: vm-latency-checkup
          image: registry.redhat.io/container-native-virtualization/vm-network-latency-checkup-rhel9:v4.0
          securityContext:
            allowPrivilegeEscalation: false
            capabilities:
              drop: ["ALL"]
            runAsNonRoot: true
            seccompProfile:
              type: "RuntimeDefault"
          env:
            - name: CONFIGMAP_NAMESPACE
              value: <target_namespace>
            - name: CONFIGMAP_NAME
              value: kubevirt-vm-latency-checkup-config
            - name: POD_UID
              valueFrom:
                fieldRef:
                  fieldPath: metadata.uid

Apply the Job manifest:

$ oc apply -n <target_namespace> -f <latency_job>.yaml

Wait for the job to complete:

$ oc wait job kubevirt-vm-latency-checkup -n <target_namespace> --for condition=complete --timeout 6m

Review the results of the latency checkup by running the following command. If the maximum measured latency is greater than the value of the spec.param.maxDesiredLatencyMilliseconds attribute, the checkup fails and returns an error.

$ oc get configmap kubevirt-vm-latency-checkup-config -n <target_namespace> -o yaml

Example output config map (success)

apiVersion: v1
kind: ConfigMap
metadata:
  name: kubevirt-vm-latency-checkup-config
  namespace: <target_namespace>
  labels:
    kiagnose/checkup-type: kubevirt-vm-latency
data:
  spec.timeout: 5m
  spec.param.networkAttachmentDefinitionNamespace: <target_namespace>
  spec.param.networkAttachmentDefinitionName: "blue-network"
  spec.param.maxDesiredLatencyMilliseconds: "10"
  spec.param.sampleDurationSeconds: "5"
  spec.param.sourceNode: "worker1"
  spec.param.targetNode: "worker2"
  status.succeeded: "true"
  status.failureReason: ""
  status.completionTimestamp: "2022-01-01T09:00:00Z"
  status.startTimestamp: "2022-01-01T09:00:07Z"
  status.result.avgLatencyNanoSec: "177000"
  status.result.maxLatencyNanoSec: "244000" (1)
  status.result.measurementDurationSec: "5"
  status.result.minLatencyNanoSec: "135000"
  status.result.sourceNode: "worker1"
  status.result.targetNode: "worker2"

1	The maximum measured latency in nanoseconds.

Optional: To view the detailed job log in case of checkup failure, use the following command:
```
$ oc logs job.batch/kubevirt-vm-latency-checkup -n <target_namespace>
```

Delete the job and config map that you previously created by running the following commands:

$ oc delete job -n <target_namespace> kubevirt-vm-latency-checkup

$ oc delete config-map -n <target_namespace> kubevirt-vm-latency-checkup-config

Optional: If you do not plan to run another checkup, delete the roles manifest:
```
$ oc delete -f <latency_sa_roles_rolebinding>.yaml
```

Running predefined storage checkups

You can use a storage checkup to verify that the cluster storage is optimally configured for OKD Virtualization.

You must always:

Verify that the checkup image is from a trustworthy source before applying it.
Review the checkup permissions before creating the Role and RoleBinding objects.

Retaining resources for troubleshooting storage checkups

The predefined storage checkup includes skipTeardown configuration options, which control resource clean up after a storage checkup runs. By default, the skipTeardown field value is Never, which means that the checkup always performs teardown steps and deletes all resources after the checkup runs.

You can retain resources for further inspection in case a failure occurs by setting the skipTeardown field to onfailure.

Prerequisites

You have installed the OpenShift CLI (oc).

Procedure

Run the following command to edit the storage-checkup-config config map:
```
$ oc edit configmap storage-checkup-config -n <checkup_namespace>
```

Configure the skipTeardown field to use the onfailure value. You can do this by modifying the storage-checkup-config config map, stored in the storage_checkup.yaml file:

apiVersion: v1
kind: ConfigMap
metadata:
  name: storage-checkup-config
  namespace: <checkup_namespace>
data:
  spec.param.skipTeardown: onfailure
# ...

Reapply the storage-checkup-config config map by running the following command:
```
$ oc apply -f storage_checkup.yaml -n <checkup_namespace>
```

Running a storage checkup by using the web console

Run a storage checkup to validate that storage is working correctly for virtual machines.

Procedure

Navigate to Virtualization → Checkups in the web console.
Click the Storage tab.
Click Install permissions.
Click Run checkup.
Enter a name for the checkup in the Name field.
Enter a timeout value for the checkup in the Timeout (minutes) fields.
Click Run.

You can view the status of the storage checkup in the Checkups list on the Storage tab. Click on the name of the checkup for more details.

Running a storage checkup by using the CLI

Use a predefined checkup to verify that the OKD cluster storage is configured optimally to run OKD Virtualization workloads.

Prerequisites

You have installed the OpenShift CLI (oc).

The cluster administrator has created the required cluster-reader permissions for the storage checkup service account and namespace, such as in the following example:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: kubevirt-storage-checkup-clustereader
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: cluster-reader
subjects:
- kind: ServiceAccount
  name: storage-checkup-sa
  namespace: <target_namespace> (1)

1	The namespace where the checkup is to be run.

Procedure

Create a ServiceAccount, Role, and RoleBinding manifest file for the storage checkup:

Example service account, role, and rolebinding manifest

---
apiVersion: v1
kind: ServiceAccount
metadata:
  name: storage-checkup-sa
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: storage-checkup-role
rules:
  - apiGroups: [ "" ]
    resources: [ "configmaps" ]
    verbs: ["get", "update"]
  - apiGroups: [ "kubevirt.io" ]
    resources: [ "virtualmachines" ]
    verbs: [ "create", "delete" ]
  - apiGroups: [ "kubevirt.io" ]
    resources: [ "virtualmachineinstances" ]
    verbs: [ "get" ]
  - apiGroups: [ "subresources.kubevirt.io" ]
    resources: [ "virtualmachineinstances/addvolume", "virtualmachineinstances/removevolume" ]
    verbs: [ "update" ]
  - apiGroups: [ "kubevirt.io" ]
    resources: [ "virtualmachineinstancemigrations" ]
    verbs: [ "create" ]
  - apiGroups: [ "cdi.kubevirt.io" ]
    resources: [ "datavolumes" ]
    verbs: [ "create", "delete" ]
  - apiGroups: [ "" ]
    resources: [ "persistentvolumeclaims" ]
    verbs: [ "delete" ]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: storage-checkup-role
subjects:
  - kind: ServiceAccount
    name: storage-checkup-sa
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: storage-checkup-role

Apply the ServiceAccount, Role, and RoleBinding manifest in the target namespace:
```
$ oc apply -n <target_namespace> -f <storage_sa_roles_rolebinding>.yaml
```

Create a ConfigMap and Job manifest file. The config map contains the input parameters for the checkup job.

Example input config map and job manifest

---
apiVersion: v1
kind: ConfigMap
metadata:
  name: storage-checkup-config
  namespace: $CHECKUP_NAMESPACE
data:
  spec.timeout: 10m
  spec.param.storageClass: ocs-storagecluster-ceph-rbd-virtualization
  spec.param.vmiTimeout: 3m
---
apiVersion: batch/v1
kind: Job
metadata:
  name: storage-checkup
  namespace: $CHECKUP_NAMESPACE
spec:
  backoffLimit: 0
  template:
    spec:
      serviceAccount: storage-checkup-sa
      restartPolicy: Never
      containers:
        - name: storage-checkup
          image: quay.io/kiagnose/kubevirt-storage-checkup:main
          imagePullPolicy: Always
          env:
            - name: CONFIGMAP_NAMESPACE
              value: $CHECKUP_NAMESPACE
            - name: CONFIGMAP_NAME
              value: storage-checkup-config

Apply the ConfigMap and Job manifest file in the target namespace to run the checkup:
```
$ oc apply -n <target_namespace> -f <storage_configmap_job>.yaml
```

Wait for the job to complete:

$ oc wait job storage-checkup -n <target_namespace> --for condition=complete --timeout 10m

Review the results of the checkup by running the following command:

$ oc get configmap storage-checkup-config -n <target_namespace> -o yaml

Example output config map (success)

apiVersion: v1
kind: ConfigMap
metadata:
  name: storage-checkup-config
  labels:
    kiagnose/checkup-type: kubevirt-storage
data:
  spec.timeout: 10m
  status.succeeded: "true" (1)
  status.failureReason: "" (2)
  status.startTimestamp: "2023-07-31T13:14:38Z" (3)
  status.completionTimestamp: "2023-07-31T13:19:41Z" (4)
  status.result.cnvVersion: 4.20.2 (5)
  status.result.defaultStorageClass: trident-nfs (6)
  status.result.goldenImagesNoDataSource: <data_import_cron_list> (7)
  status.result.goldenImagesNotUpToDate: <data_import_cron_list> (8)
  status.result.ocpVersion: 4.20.0 (9)
  status.result.pvcBound: "true" (10)
  status.result.storageProfileMissingVolumeSnapshotClass: <storage_class_list> (11)
  status.result.storageProfilesWithEmptyClaimPropertySets: <storage_profile_list> (12)
  status.result.storageProfilesWithSmartClone: <storage_profile_list> (13)
  status.result.storageProfilesWithSpecClaimPropertySets: <storage_profile_list> (14)
  status.result.storageProfilesWithRWX: |-
    ocs-storagecluster-ceph-rbd
    ocs-storagecluster-ceph-rbd-virtualization
    ocs-storagecluster-cephfs
    trident-iscsi
    trident-minio
    trident-nfs
    windows-vms
  status.result.vmBootFromGoldenImage: VMI "vmi-under-test-dhkb8" successfully booted
  status.result.vmHotplugVolume: |-
    VMI "vmi-under-test-dhkb8" hotplug volume ready
    VMI "vmi-under-test-dhkb8" hotplug volume removed
  status.result.vmLiveMigration: VMI "vmi-under-test-dhkb8" migration completed
  status.result.vmVolumeClone: 'DV cloneType: "csi-clone"'
  status.result.vmsWithNonVirtRbdStorageClass: <vm_list> (15)
  status.result.vmsWithUnsetEfsStorageClass: <vm_list> (16)

1	Specifies if the checkup is successful (`true`) or not (`false`).
2	The reason for failure if the checkup fails.
3	The time when the checkup started, in RFC 3339 time format.
4	The time when the checkup has completed, in RFC 3339 time format.
5	The OKD Virtualization version.
6	Specifies if there is a default storage class.
7	The list of golden images whose data source is not ready.
8	The list of golden images whose data import cron is not up-to-date.
9	The OKD version.
10	Specifies if a PVC of 10Mi has been created and bound by the provisioner.
11	The list of storage profiles using snapshot-based clone but missing VolumeSnapshotClass.
12	The list of storage profiles with unknown provisioners.
13	The list of storage profiles with smart clone support (CSI/snapshot).
14	The list of storage profiles spec-overriden claimPropertySets.
15	The list of virtual machines that use the Ceph RBD storage class when the virtualization storage class exists.
16	The list of virtual machines that use an Elastic File Store (EFS) storage class where the GID and UID are not set in the storage class.

Delete the job and config map that you previously created by running the following commands:

$ oc delete job -n <target_namespace> storage-checkup

$ oc delete config-map -n <target_namespace> storage-checkup-config

Optional: If you do not plan to run another checkup, delete the ServiceAccount, Role, and RoleBinding manifest:
```
$ oc delete -f <storage_sa_roles_rolebinding>.yaml
```

Troubleshooting a failed storage checkup

If a storage checkup fails, there are steps that you can take to identify the reason for failure.

Prerequisites

You have installed the OpenShift CLI (oc).
You have downloaded the directory provided by the must-gather tool.

Procedure

Review the status.failureReason field in the storage-checkup-config config map by running the following command and observing the output:

$ oc get configmap storage-checkup-config -n <namespace> -o yaml

Example output config map

apiVersion: v1
kind: ConfigMap
metadata:
  name: storage-checkup-config
  labels:
    kiagnose/checkup-type: kubevirt-storage
data:
  spec.timeout: 10m
  status.succeeded: "false" (1)
  status.failureReason: "ErrNoDefaultStorageClass" (2)
# ...

1	If the checkup has failed, the `status.succeeded` value is `false`.
2	If the checkup has failed, the `status.failureReason` field contains an error message. In this example output, the `ErrNoDefaultStorageClass` error message means that no default storage class is configured.

Search the directory provided by the must-gather tool for logs, events, or terms related to the error in the data.status.failureReason field value.

Additional resources

Storage checkup error codes

The following error codes might appear in the storage-checkup-config config map after a storage checkup fails.

Error code Meaning

Error code	Meaning
`ErrNoDefaultStorageClass`	No default storage class is configured.
`ErrPvcNotBound`	One or more persistent volume claims (PVCs) failed to bind.
`ErrMultipleDefaultStorageClasses`	Multiple default storage classes are configured.
`ErrEmptyClaimPropertySets`	There are `StorageProfile` objects containing empty `ClaimPropertySets` specs.
`ErrVMsWithUnsetEfsStorageClass`	There are VMs using elastic file system (EFS) storage classes, where the GID and UID are not set in the `StorageClass` object.
`ErrGoldenImagesNotUpToDate`	One or more golden images has a `DataImportCron` object that is either not up to date or has a `DataSource` object which is not ready.
`ErrGoldenImageNoDataSource`	The `DataSource` object of the golden image has either no PVC or no snapshot source configured.
`ErrBootFailedOnSomeVMs`	Some VMs failed to boot within the expected time.

ErrNoDefaultStorageClass

No default storage class is configured.

ErrPvcNotBound

One or more persistent volume claims (PVCs) failed to bind.

ErrMultipleDefaultStorageClasses

Multiple default storage classes are configured.

ErrEmptyClaimPropertySets

There are StorageProfile objects containing empty ClaimPropertySets specs.

ErrVMsWithUnsetEfsStorageClass

There are VMs using elastic file system (EFS) storage classes, where the GID and UID are not set in the StorageClass object.

ErrGoldenImagesNotUpToDate

One or more golden images has a DataImportCron object that is either not up to date or has a DataSource object which is not ready.

ErrGoldenImageNoDataSource

The DataSource object of the golden image has either no PVC or no snapshot source configured.

ErrBootFailedOnSomeVMs

Some VMs failed to boot within the expected time.

Additional resources

Connecting a virtual machine to a Linux bridge network