interfaces:
- ipv4:
dhcp: true
enabled: true
mac-address: <xx:xx:xx:xx:xx:xx>
mtu: 1500 # Set to 1500 for standard MTU or 9000 for jumbo frames
name: <interface-name>
state: up
type: ethernet
Before installing the NVIDIA DPF Operator, you must set up the management cluster, configure worker nodes, and install and configure the required Operators.
The management cluster is a standard OKD 4 cluster installed by using the Assisted Installer. This cluster hosts the DPF Operators and the hosted control planes for managing the hosted cluster on DPUs.
You have access to the Red Hat Hybrid Cloud Console.
You have the OpenShift CLI (oc) installed.
Go to the Red Hat Hybrid Cloud Console cluster creation page and create a cluster with control-plane nodes only. Select Data center → Assisted Installer.
Optional: Configure jumbo MTU for each control plane node.
Under Hosts' network configuration in the Assisted Installer wizard, select Static IP, bridges, and bonds.
Set the Static network configurations section per node according to the following template, using the relevant MAC address and interface name for each node:
interfaces:
- ipv4:
dhcp: true
enabled: true
mac-address: <xx:xx:xx:xx:xx:xx>
mtu: 1500 # Set to 1500 for standard MTU or 9000 for jumbo frames
name: <interface-name>
state: up
type: ethernet
|
Select the following operators to install with the cluster:
Storage → Logical Volume Manager Storage
Platform Operations & Lifecycle → MultiCluster Engine
Scheduling → Node Feature Discovery
Click Add hosts to add hosts to the cluster. Only control plane nodes are required at this stage.
After the installation completes, download the KUBECONFIG file and save it as mgmt-kubeconfig.
Set the KUBECONFIG environment variable:
$ export KUBECONFIG="$(pwd)/mgmt-kubeconfig"
Verify that all nodes are in a Ready state:
$ oc get nodes
You must deploy the dpu-worker-config Helm chart to configure worker nodes with DPUs before adding those nodes to the management cluster.
The dpu-worker-config Helm chart creates the MachineConfigPool, dpu-worker-configuration MachineConfig, and other required resources that configure the bridge, OVS services, and IP routing on worker nodes.
The MachineConfigPool groups DPU-equipped worker nodes so that the Machine Config Operator can apply DPU-specific configurations to them.
The MachineConfig resource performs several configuration tasks required by DPF:
Bridge Configuration: Creates a br-ex bridge interface that enables communication between the DPU and the hosted cluster control plane running on the management cluster. This name must match the dpuNodeOOBBridgeName value in the DPFOperatorConfig resource, or DPU provisioning fails. For more information refer to DPF Operator Prerequisites.
OVS Service Management: Disables OpenShift’s default OVS services on x86 worker nodes. This is required for OVN-Kubernetes DPU Host mode operation, where networking functions are offloaded to the DPU rather than running on the host CPU.
IP Rules Configuration: Sets routing rules required for pod-to-host control-plane traffic.
You have access to the cluster as a user with the cluster-admin role.
You have installed the OpenShift CLI (oc).
The Helm CLI (helm) is installed on your workstation.
You have a pull secret file that includes credentials for registry.redhat.io. Helm reads its registry credentials from this file, which is separate from the container runtime configuration.
Set the OPENSHIFT_PULL_SECRET environment variable to the path of your pull secret file:
$ export OPENSHIFT_PULL_SECRET="/root/pull-secret.txt"
Deploy the dpu-worker-config Helm chart:
$ helm upgrade --install dpu-worker-config \
oci://registry.redhat.io/dpu-kit-for-nvidia/dpu-worker-config-chart \
--version 4.22.0 \
--registry-config "${OPENSHIFT_PULL_SECRET}" \
--namespace dpf-hcp-provisioner-system \
--create-namespace \
--disable-openapi-validation
Verify that the dpu-worker-config Helm release is deployed:
$ helm list -n dpf-hcp-provisioner-system
Verify that the MachineConfigPool was created automatically:
$ oc get mcp worker-dpu
NAME CONFIG UPDATED UPDATING DEGRADED MACHINECOUNT READYMACHINECOUNT UPDATEDMACHINECOUNT DEGRADEDMACHINECOUNT AGE
worker-dpu rendered-worker-dpu-b4a44cc606c2ffcf07067d1e943a8758 True False False 0 0 0 0 2m
Verify that the dpu-worker-configuration MachineConfig is created:
$ oc get machineconfig dpu-worker-configuration
NAME GENERATEDBYCONTROLLER IGNITIONVERSION AGE
dpu-worker-configuration 3.2.0 2m
|
The Machine Config Operator automatically reboots worker nodes to apply DPU-specific configurations after nodes with the |
You must create a dedicated namespace for the DPF Operator and its components before installing the Operator.
You have access to the cluster as a user with the cluster-admin role.
You have installed the OpenShift CLI (oc).
Create the dpf-operator-system namespace:
$ oc create namespace dpf-operator-system
Verify that the namespace was created:
$ oc get namespace dpf-operator-system
Before you install the DPF Operator, you must install the cert-manager Operator for Red Hat OpenShift, MetalLB Operator, Red Hat OpenShift GitOps, and NVIDIA Maintenance Operator.
The multicluster engine Operator and the Node Feature Discovery Operator can be installed during management cluster creation by using the Assisted Installer. If you did not install the multicluster engine Operator, follow "Install the multicluster engine Operator" before configuring the required Operators.
If you did not install the multicluster engine Operator by using the Assisted Installer, install it from the OpenShift CLI. You do not need to install Red Hat Advanced Cluster Management or create a MultiClusterHub resource to use the multicluster engine Operator. If the Assisted Installer already created a MultiClusterEngine resource, skip this procedure.
You have access to the management cluster as a user with the cluster-admin role.
You have installed the OpenShift CLI (oc).
No MultiClusterEngine resource exists on the cluster.
If you already have an existing MultiClusterEngine resource with a different name, note that the following verification commands use mce. Substitute your resource’s name where needed.
Check which multicluster engine channels are available from the redhat-operators catalog:
$ oc get packagemanifests -n openshift-marketplace -l catalog=redhat-operators \
-o jsonpath='{.items[?(@.metadata.name=="multicluster-engine")].status.channels[*].name}{"\n"}'
The following example uses stable-2.17, which is the channel used by this deployment. If your catalog does not offer this channel, select a supported channel for your OpenShift version before continuing.
Create a file named mce-operator.yaml with the following content:
apiVersion: v1
kind: Namespace
metadata:
name: multicluster-engine
---
apiVersion: operators.coreos.com/v1
kind: OperatorGroup
metadata:
name: multicluster-engine
namespace: multicluster-engine
spec:
targetNamespaces:
- multicluster-engine
---
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
name: multicluster-engine
namespace: multicluster-engine
spec:
channel: stable-2.17
name: multicluster-engine
source: redhat-operators
sourceNamespace: openshift-marketplace
installPlanApproval: Automatic
If you selected a different supported channel, replace stable-2.17 with that channel in the file.
Apply the file:
$ oc apply -f mce-operator.yaml
Wait until the multicluster engine Operator CSV reports Succeeded and the MultiClusterEngine CRD is established:
$ oc get csv -n multicluster-engine
$ oc wait crd/multiclusterengines.multicluster.openshift.io --for=create --timeout=10m
$ oc wait crd/multiclusterengines.multicluster.openshift.io --for=condition=Established --timeout=5m
Repeat the CSV check until its phase is Succeeded before creating the custom resource.
Create a file named mce.yaml with the following content:
apiVersion: multicluster.openshift.io/v1
kind: MultiClusterEngine
metadata:
name: mce
spec:
overrides:
components:
- name: hypershift
enabled: true
Apply the file:
$ oc apply -f mce.yaml
Wait for the MultiClusterEngine resource to become available:
$ oc wait multiclusterengine/mce --for=jsonpath='{.status.phase}'=Available --timeout=15m
Verify that the hosted control plane component is enabled:
$ oc get multiclusterengine mce -o jsonpath='{.spec.overrides.components[?(@.name=="hypershift")].enabled}{"\n"}'
true
The cert-manager Operator for Red Hat OpenShift manages TLS certificates for DPF components. You install this operator by using the OpenShift CLI.
You have access to the cluster as a user with the cluster-admin role.
You have installed the OpenShift CLI (oc).
Create a file named cert-manager-operator.yaml with the following content:
apiVersion: v1
kind: Namespace
metadata:
name: cert-manager
---
apiVersion: operators.coreos.com/v1
kind: OperatorGroup
metadata:
name: openshift-cert-manager-operator
namespace: cert-manager
spec:
targetNamespaces:
- cert-manager
---
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
name: openshift-cert-manager-operator
namespace: cert-manager
spec:
channel: stable-v1
name: openshift-cert-manager-operator
source: redhat-operators
sourceNamespace: openshift-marketplace
Apply the file:
$ oc apply -f cert-manager-operator.yaml
Verify that the Operator is installed:
$ oc get pods -n cert-manager
The MetalLB Operator provides load balancing services for DPF components on the management cluster. You install this operator by using the OpenShift CLI.
You have access to the cluster as a user with the cluster-admin role.
You have installed the OpenShift CLI (oc).
Create a file named metallb-operator.yaml with the following content:
apiVersion: v1
kind: Namespace
metadata:
name: metallb-system
---
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
name: metallb-operator
namespace: openshift-operators
spec:
channel: "stable"
name: metallb-operator
source: redhat-operators
sourceNamespace: openshift-marketplace
installPlanApproval: Automatic
config:
# Tolerate the taint on the master nodes
tolerations:
- key: "node-role.kubernetes.io/control-plane"
operator: "Exists"
effect: "NoSchedule"
# Force scheduling only on nodes with the control-plane label
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: "node-role.kubernetes.io/control-plane"
operator: "Exists"
Apply the file:
$ oc apply -f metallb-operator.yaml
Wait for the MetalLB custom resource definition to be created and established before configuring MetalLB:
$ oc wait crd/metallbs.metallb.io --for=create --timeout=10m
$ oc wait crd/metallbs.metallb.io --for=condition=Established --timeout=5m
The Red Hat OpenShift GitOps manages DPF service deployments and configurations using GitOps principles. You install this operator by using the OpenShift CLI.
You have access to the cluster as a user with the cluster-admin role.
You have installed the OpenShift CLI (oc).
Create a file named gitops-operator.yaml with the following content:
apiVersion: v1
kind: Namespace
metadata:
name: openshift-gitops-operator
labels:
openshift.io/cluster-monitoring: "true"
---
apiVersion: operators.coreos.com/v1
kind: OperatorGroup
metadata:
name: openshift-gitops-operator
namespace: openshift-gitops-operator
spec:
upgradeStrategy: Default
---
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
name: openshift-gitops-operator
namespace: openshift-gitops-operator
spec:
channel: gitops-1.21
config:
env:
- name: ARGOCD_CLUSTER_CONFIG_NAMESPACES
value: "openshift-gitops,dpf-operator-system"
- name: CONTROLLER_CLUSTER_ROLE
value: "cluster-admin"
- name: SERVER_CLUSTER_ROLE
value: "cluster-admin"
installPlanApproval: Automatic
name: openshift-gitops-operator
source: redhat-operators
sourceNamespace: openshift-marketplace
Apply the file:
$ oc apply -f gitops-operator.yaml
Wait for the Argo CD custom resource definition to be created and established before creating an ArgoCD resource:
$ oc wait crd/argocds.argoproj.io --for=create --timeout=10m
$ oc wait crd/argocds.argoproj.io --for=condition=Established --timeout=5m
The NVIDIA Maintenance Operator assists in performing maintenance tasks and gracefully draining DPU worker nodes. You install this operator by using Helm.
You have access to the cluster as a user with the cluster-admin role.
You have installed the OpenShift CLI (oc).
You have installed the Helm CLI (helm).
Create a Helm values file named maintenance-operator-values.yaml with the following content:
operatorConfig:
maxParallelOperations: 60%
operator:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: "node-role.kubernetes.io/master"
operator: Exists
- matchExpressions:
- key: "node-role.kubernetes.io/control-plane"
operator: Exists
tolerations:
- key: node-role.kubernetes.io/master
operator: Exists
effect: NoSchedule
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
Install the Operator by using Helm:
$ helm upgrade --install maintenance-operator oci://ghcr.io/mellanox/maintenance-operator-chart \
--namespace dpf-operator-system \
--create-namespace \
--disable-openapi-validation \
--version 0.3.0 \
--values maintenance-operator-values.yaml \
--wait
Verify that the Operator pod is running:
$ oc get pods -n dpf-operator-system
After the required Operators are installed, configure Node Feature Discovery, MetalLB, GitOps, and Cluster Network Operator for the DPF environment. This procedure also verifies that the multicluster engine and hosted control plane components are ready.
You have access to the cluster as a user with the cluster-admin role.
You have installed the OpenShift CLI (oc).
You have installed the cert-manager Operator for Red Hat OpenShift, MetalLB Operator, Red Hat OpenShift GitOps, and NVIDIA Maintenance Operator.
You have installed the Logical Volume Manager Storage Operator, multicluster engine Operator, and the Node Feature Discovery Operator. You can install them by using the Assisted Installer during cluster creation. If you did not install the multicluster engine Operator, follow "Install the multicluster engine Operator" before continuing.
Define the cluster variables:
$ export CLUSTER_NAME="doca-mgmt"
$ export BASE_DOMAIN="example.com"
$ export HOST_CLUSTER_API="api.${CLUSTER_NAME}.${BASE_DOMAIN}"
where:
CLUSTER_NAMESpecifies the management cluster name.
BASE_DOMAINSpecifies the management cluster base domain.
HOST_CLUSTER_APISpecifies the management cluster API endpoint.
Create a file named nfd-instance.yaml with the following NodeFeatureDiscovery resource definition:
apiVersion: nfd.openshift.io/v1
kind: NodeFeatureDiscovery
metadata:
name: nfd-instance
namespace: openshift-nfd
spec:
operand:
workerEnvs:
- name: KUBERNETES_SERVICE_HOST
value: $HOST_CLUSTER_API
- name: KUBERNETES_SERVICE_PORT
value: "6443"
workerConfig:
configData: |
sources:
pci:
deviceClassWhitelist:
- "0200"
- "03"
- "12"
- "0207"
deviceLabelFields:
- "vendor"
- "device"
- "class"
Apply the file by using envsubst to substitute the environment variables:
$ envsubst < nfd-instance.yaml | oc apply -f -
Create a file named nfd-rule.yaml with the following NodeFeatureRule resource definition:
apiVersion: nfd.openshift.io/v1alpha1
kind: NodeFeatureRule
metadata:
name: dpu-detection-rule
namespace: openshift-nfd
spec:
rules:
- labels:
dpu-enabled: ""
matchFeatures:
- feature: pci.device
matchExpressions:
device:
op: In
value:
- a2d6
- a2dc
vendor:
op: In
value:
- 15b3
name: DPU-detection-rule
Apply the file:
$ oc apply -f nfd-rule.yaml
Create a file named metallb-config.yaml with the following MetalLB resource definition:
apiVersion: metallb.io/v1beta1
kind: MetalLB
metadata:
name: metallb
namespace: openshift-operators
spec:
nodeSelector:
node-role.kubernetes.io/control-plane: ""
speakerTolerations:
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
Apply the MetalLB resource file:
$ oc apply -f metallb-config.yaml
Create a file named argocd-instance.yaml with the following ArgoCD resource definition:
apiVersion: argoproj.io/v1beta1
kind: ArgoCD
metadata:
name: argocd
namespace: dpf-operator-system
spec:
nodePlacement:
nodeSelector:
node-role.kubernetes.io/control-plane: ""
tolerations:
- key: node-role.kubernetes.io/master
operator: Exists
effect: NoSchedule
- key: node-role.kubernetes.io/control-plane
operator: Exists
effect: NoSchedule
server:
route:
enabled: true
controller: {}
repo: {}
applicationSet:
enabled: false
resourceExclusions: |
- apiGroups:
- packages.operators.coreos.com
kinds:
- PackageManifest
sso:
provider: dex
dex:
openShiftOAuth: true
notifications:
enabled: false
Apply the Argo CD file:
$ oc apply -f argocd-instance.yaml
Wait for the ArgoCD Redis deployment to be ready:
$ oc wait deployment argocd-redis -n dpf-operator-system \
--for=condition=Available --timeout=120s
Enable global IP forwarding on the OVN-Kubernetes configuration:
This command enables IP packet forwarding between different networks managed by OVN-Kubernetes.
$ oc patch network.operator.openshift.io cluster --type=merge -p \
'{"spec":{"defaultNetwork":{"ovnKubernetesConfig":{"gatewayConfig":{"ipForwarding":"Global"}}}}}'
Verify that the MultiClusterEngine instance is created:
$ oc get multiclusterengine mce
NAME STATUS AGE CURRENTVERSION DESIREDVERSION MESSAGE
mce Available 4m58s 2.17.2 2.17.2 All components available
Verify that the hosted control plane component is enabled:
$ oc get multiclusterengine mce -o jsonpath='{.spec.overrides.components[?(@.name=="hypershift")].enabled}{"\n"}'
|
The manual installation procedure creates |
To enable the hypershift component on the MultiClusterEngine resource, use the following steps.
In current multicluster engine Operator versions the component is named hypershift. Earlier versions use hypershift-preview.
If the result is empty, check whether the components list exists:
$ oc get multiclusterengine mce -o jsonpath='{.spec.overrides.components}{"\n"}'
If the list exists but has no hypershift entry, append it without replacing the other components:
$ oc patch multiclusterengine mce --type=json \
-p='[{"op":"add","path":"/spec/overrides/components/-","value":{"name":"hypershift","enabled":true}}]'
If the list is absent, edit the resource instead and create spec.overrides.components with an entry named hypershift set to enabled: true.
If the result is false, an entry exists but is disabled. Edit the resource and set the hypershift component to enabled: true:
$ oc edit multiclusterengine mce
|
Do not use |
Verify that the NodeFeatureDiscovery instance and NodeFeatureRule are configured:
$ oc get nodefeaturediscovery,nodefeaturerule -n openshift-nfd
Verify that the MetalLB instance was created:
$ oc get metallb -n openshift-operators
Verify that the Argo CD pods are running:
$ oc get pods -n dpf-operator-system | grep argocd
Verify that IP forwarding is set to Global:
$ oc get network.operator.openshift.io cluster -o jsonpath='{.spec.defaultNetwork.ovnKubernetesConfig.gatewayConfig.ipForwarding}'
Global