Errors on All HA Mailbox Disks After Adding Disks to NS224
Applies to
- NetApp AFF-A400
- ONTAP 9.16.1P1
- NS224
Issue
After adding new disks to an existing NS224 shelf during a maintenance window, both nodes in the HA pair became unresponsive. One node panicked and the partner rebooted. Both entered a reboot loop, reporting no disks present. The following panic string was observed:
PANIC: Permanent errors on all HA mailbox disks (while marshalling header) in SK process fmmbx_instanceWorker on release 9.16.1P1(C)
Relevant log excerpts:
mbx_inst_header_marshal: Error writing to all mailbox disks. mbx_sequenceNo=23670722[PA400EUPOLIS02:fmmbx_instanceWorker:cf.multidisk.fatalProblem:error]: Node encountered a multidisk error or other fatal error while waiting to be taken over. Permanent errors on all HA mailbox disks (while marshalling header).[PA400EUPOLIS01:scsi_cmdblk_strthr_admin:disk.readReservationFailed:error]: Disk read reservation failed on e5b.00.3.16CDB0x5e:01-SCSI:nosense(000)[PA400EUPOLIS01:sanown_io:diskown.errorDuringIO:error]: error 25 (no valid path to disk) on disk e0c.00.0.16 while reading reservation state
Symptoms:
- Both controllers become unresponsive after disk insertion.
- One node panics, partner reboots.
- Both nodes enter reboot loop, reporting no disks.
- Recovery only after entire NS224 shelf power-cycled.
