Storage ports remain down after node reboot when NVIDIA ConnectX NICs are connected to Cisco Nexus 9336C-FX2 switches
Applies to
- AFF A900
- Cisco Nexus 9336C-FX2 switches
- NVIDIA ConnectX-based Ethernet adapters, including:
- CX6: X91153A
- CX6-DX: X50130A
- CX7: X50133A
- Storage ports connected using copper cables
Issue
- A storage port remains down or offline after a node reboot, power cycle, takeover/giveback, automated nondisruptive upgrade (ANDU), link toggle, switch-port reset, or cable reconnect
- ONTAP reports the following health alert:
Health Monitor process nchm:NoPathToNSMA_Alert…
- From
sysconfig -a,System Storage Configuration changed fromQuad-Path HAtoSingle-Path HAand storage port isauto-unknown-fd-down.
System Rev: C0
System Storage Configuration: Single-Path HA
System ACP Connectivity: Inband Active
・
slot 2: Dual 40G/100G/200G Ethernet Controller CX6
e2a MAC Address: XX:00:XX:11:XX:XX (auto-unknown-fd-down)
- EMS logs show link down on storage ports upon reboot.
[?] Thu May 09 10:12:45 +0900 [node-1: kernel: netif.linkDown:info]: Ethernet e2a: Link down, check cable.
[?] Thu May 09 10:23:07 +0900 [node-1: kernel: netif.linkDown:info]: Ethernet e10b: Link down, check cable.
[?] Thu May 09 10:12:45 +0900 [node-2: kernel: netif.linkDown:info]: Ethernet e2a: Link down, check cable.
[?] Thu May 09 10:23:07 +0900 [node-2: kernel: netif.linkDown:info]: Ethernet e10b: Link down, check cable.
storage port showshows that one of the storage ports goes down randomly after reboot.
Speed VLAN
Node Port Type Mode (Gb/s) State Status ID
------------------ ---- ----- ------- ------ -------- --------- ----
node-1
e10a ENET network - - - -
e10b ENET storage 100 enabled online 30
e11a ENET network - - - -
e11b ENET network - - - -
e2a ENET storage 0 enabled offline 30
e2b ENET network - - - -
- No errors/discards seen in ifstat but we see
Up to downindicating several times flapping occured.
- The affected storage port remain in a link-down or offline state, or require an extended time to recover.
- Switch logs show an error
UNSUPPORTED_TRANSCEIVERandunsupported vendorfrom the connected port, but Hardware Universe confirms that it is supported.
[2024-05-09 19:15:36.280] 2024 May 9 01:27:10 switch02 %ETHPORT-5-IF_HARDWARE: Interface Ethernet1/11, hardware type changed to 100G
[2024-05-09 19:15:36.280] 2024 May 9 01:27:10 switch02 %ETHPORT-3-IF_UNSUPPORTED_TRANSCEIVER: Transceiver on interface Ethernet1/11 is not supported
[2024-05-09 19:15:36.311] 2024 May 9 01:27:11 switch02 %ETHPORT-5-IF_HARDWARE: Interface Ethernet1/12, hardware type changed to 100G
[2024-05-09 19:15:36.311] 2024 May 9 01:27:11 switch02 %ETHPORT-3-IF_UNSUPPORTED_TRANSCEIVER: Transceiver on interface Ethernet1/12 is not supported
[2024-05-09 19:21:17.984] 2024-05-09T09:02:43.200667000+00:00 [M 1] [ethpm] E_DEBUG Ifindex (Ethernet1/12)0x1a000000, SFP security check: unsupported vendor id 0x54
[2024-05-09 19:21:18.001] 2024-05-09T09:02:37.697065000+00:00 [M 1] [ethpm] E_DEBUG Ifindex (Ethernet1/11)0x1a000000, SFP security check: unsupported vendor id 0x54
- Cisco switch logs contain messages similar to:
kr_tune_success=0Took too long to tune links. Restarting AN
