PANIC : ECC error at DIMM-XX, Uncorrectable Machine Check Error Resolution guide
Applies to
- ONTAP 9
- Data ONTAP 8 7-Mode
- AFF / FAS Platforms
- Memory, UECC (Uncorrectable Error Correction Code)
- DIMM
- NVRAM
Description
This procedure provides a guide to the correct remediation actions when a controller experiences a panic and reboot due to an uncorrectable memory error on a DIMM.
Examples:
-
PANIC: ECC error at DIMM-18: 2C-0F-2007-2664E6BE,ADDR 0x180a048b40,(Node(1), Memory controller(1), CH(3), DIMM(0), Rank(0), Bank Group(1), Bank(0x0), Row(0xb8b1), Col(0x2f8), Uncorrectable Machine Check Error at CPU21.
-
NVRAM in slot 6: uncorrectable memory error at address 0x99f60628 DIMM(1), Rank(0), Bank Group(0), Bank(0x0), Row(0x99f6), Col(0x1d) in process idle
-
PANIC: Uncorrectable Machine Check Error at CPU14. ECC error at DIMM-13: CE-01-1941-03A203B8,ADDR 0x15f09e1f40,(Node(1), Memory controller(0), CH(1), DIMM(0), Rank(0), Bank Group(3), Bank(0x0), Row(0x15e0f), Col(0x70)) SKL_IMC0 Error: STATUS<0xfe0000c001010091>(VALID,OVERFLOW,UC,EN,MISCV,ADDRV,PCC,CORR_ERR_STATUS(0),CORR_ERR_CNT(0x3),OTHER_INFO(0),MscodDdrType(0x1),MscodDataRdErr,MCACOD(0x91))MISC<0x200400c00fc02086>(DataErrorChunk(0x2),McCmdChnl(0x1),McCmdMemRegion(0),McCmdOpcode(0),McCmdVld,SmiAD,SmiMsgClass(0),SmiOpcode(0),TrkId(0x7e),Error_Type(0x4),ADDRMODE(0x2),ADDRLSB(0x6))ADDR<0x00000015f09e1f40>(HIPHYADDR(0x15),LOPHYADDR(0x3c2787d))(Node(1), Memory controller(0), CH(1), DIMM(0), Rank(0), Bank Group(3), Bank(0x0), Row(0x15e0f), Col(0x70)
-
cf_hwassist: cf.hwassist.takeoverTrapRecv:debug]: hw_assist: Received takeover hw_assist alert from partner(node02), system_down because dimm_uecc_error.