AFF-A300 node enters unknown state during ONTAP upgrade
Applies to
- FAS/AFF
- ONTAP 9
- Automated Non-Disruptive Upgrade (ANDU)
Issue
- During a planned ONTAP cluster upgrade on the AFF-A300 system, node 2 entered an 'Unknown' state.
Cluster::> storage failover showTakeoverNode Partner Possible State Description-------------- -------------- -------- -------------------------------------Clusterc1aClusterc1b false In takeoverClusterc1bClusterc1a - Unknown2 entries were displayed. - The node became unresponsive and its Service Processor (SP)/Baseboard Management Controller (BMC) was inaccessible remotely.
- The node had an uptime of 699 days prior to the event.
- SP or BMC reports heartbeat stop and raise "
Power Reset" event:
[IPMI.notice]: 9200 | c0 | OEM: ffff7000ff00 | ManufId: 150300 | SP Power Reset
[IPMI.notice]: 9300 | c0 | OEM: fcff70560000 | ManufId: 150300 | POS Register: Power on Reset(Normal Power Cycle)
or
[IPMI.notice]: 00d4 | c0 | OEM: ffff7000ff00 | ManufId: 150300 | BMC Power Reset
[IPMI.notice]: 00d5 | c0 | OEM: fcff70560000 | ManufId: 150300 | POS Register: Power on Reset(Normal Power Cycle)
- Partner takeover due to missed heartbeat or lost communication
[node_name: cf_main: cf.fsm.takeover.noHeartbeat:alert]: Failover monitor: Takeover initiated after no heartbeat was detected from the partner node.
or
cf_fastTimeout: cf.ic.heartBeatFailed:error]: HA interconnect: Heartbeat failed
or
cf_takeover: callhome.sfo.takeover:alert]: Call home for CONTROLLER TAKEOVER COMPLETE AUTOMATIC - Communication Error
- No other suspicious error messages.
