StorageGRID appliance intermediately reboots unexpectedly due to HICs issue
Applies to
NetApp StorageGRID SG5712
Issue
StorageGRID node rebooted unexpectedly with the alert "
NODE_DOWN-MAJOR"In
/var/log/kern.log :2025-09-01T16:31:17.098398+00:00 SG kernel: [365419.665123] NETDEV WATCHDOG: hic4 (qede): transmit queue 3 timed outMar 15 00:45:35 localhost kernel: [9675498.842739] bond0: (slave hic2): link status down for interface, disabling it in 200 ms
Mar 15 00:45:35 localhost wdogd[2842]: NTMN: /usr/bin/storagegrid-network-monitor failed to read SFP data for hic2; suspect unplugged
Mar 15 00:45:35 localhost kernel: [9675498.842783] qede 0000:42:00.1: Ending qede_remove successfully
Mar 15 00:45:35 localhost kernel: [9675498.843130] [qed_load_mcp_offsets:269()]Assume that no MFW is running [MCP_REG_CACHE_PAGING_ENABLE 0x10040]
Mar 15 00:45:35 localhost kernel: [9675498.843138] [qed_hw_info_port_num:7524()]Unknown port mode 0x00010040
Mar 15 00:45:35 localhost kernel: [9675498.843149] [qed_hw_get_nvm_info:7121()]Unknown port mode in 0x00010040
Mar 15 00:45:35 localhost kernel: [9675498.843165] [qed_mcp_fill_shmem_func_info:3376()]MAC is 0 in shmem
Mar 15 00:45:35 localhost kernel: [9675498.843426] [qed_hw_prepare_single:7899()]Failed to get HW information
Mar 15 00:45:35 localhost kernel: [9675498.843466] qede 0000:42:00.1 hic2: Recovery handling has failed. Power cycle is needed.