Unexpected Node Reboot and HA Interconnect Loss on ASA-A90
Applies to
- NetApp ASA-A90 systems
- ONTAP 9.17.1P3
Issue
A node in an ASA-A90 cluster (example: Node b) unexpectedly rebooted, resulting in an automatic takeover by its HA partner (Node-a). This triggered the following alert:
CLTFLT: HAGroupNotification from Node-a (CONTROLLER TAKEOVER COMPLETE AUTOMATIC - Communication Error)
Additional log excerpts:
[Node-a:vifmgr.clus.linkdown:EMERGENCY]: The cluster port e1a on Node-a has gone down unexpectedly. [Node-a:vifmgr.clus.linkdown:EMERGENCY]: The cluster port e7a on Node-a has gone down unexpectedly. [Node-a:callhome.clam.node.ooq:EMERGENCY]: Callhome for NODE(S) OUT OF CLUSTER QUORUM. [Node-a:monitor.globalStatus.critical:EMERGENCY]: This node has taken over Node-b.
The affected node was unavailable for approximately 125 seconds before returning to service through a fresh boot sequence. No controller panic, watchdog timeout, power supply failure, thermal shutdown, or explicit hardware fault was logged.
