Node panic due to persistent DIMM ECC errors
Applies to
- AFF and FAS systems
- ONTAP 9
Issue
- Node experience multiple panics and automatic HA (High Availability) takeover events.
CLTFLT:HAGroupNotification from DCNETAFFBKCEUSPR1-02 (CONTROLLER TAKEOVER COMPLETE AUTOMATIC-Communication Error) - Repeated uncorrectable ECC (Error-Correcting Code) memory errors reported on DIMM (Dual In-line Memory Module) slot, even after prior DIMM replacement.
Mon Aug 3 19:43:20 2026 SRAM record type (CPU) from DataONTAP: socket (1) core (12) bank (7)Mon Aug 3 19:43:20 2026 SRAM record type (LOG) from DataONTAP: UECC Addr 0x6xxx9b80Mon Aug 3 19:43:20 2026 SRAM record type (DIMM) from DataONTAP: slot (13)- After replacement of DIMM, errors recurred with a new DIMM at a different address.
Thu Mar 1 01:35:23 2026 SRAM record type (CPU) from DataONTAP: socket (0) core (0) bank (7)Thu Mar 1 01:35:23 2026 SRAM record type (LOG) from DataONTAP: UECC Addr 0x8x0Thu Mar 1 01:35:23 2026 SRAM record type (DIMM) from DataONTAP: slot (2)