Disk missing on CVO causing system to panic
Applies to
- Cloud Volumes ONTAP (CVO)
- NetApp Console ( Blue XP / Cloud Manager)
- Microsoft Azure
- Amazon Web Services (AWS)
- Google Cloud Platform (GCP)
- Single Node or HA Pair
Issue
- One or more disks becomes unreachable due to an issue in the underline infrastructure and causes a panic:
[Cluster-01: pha_remove000: mlm.array.lun.removed:notice]: Array LUN '0b.29' (00000000i3g268fHE60S) is no longer being presented to this node.
[Cluster-01: dmgr_thread: raid.disk.missing:info]: Disk /aggr04/plex0/rg0/0b.29 S/N [00000000i3g268fHE60S] UID [00000000i3g268fHE60S] is missing from the system
[Cluster-01: config_thread: sk.panic:alert]: Panic String: aggr aggr04: raid volfsm, fatal disk error in RAID group with no parity disk.. Raid type - raid0 Group name plex0/rg0 state NORMAL. 1 disk failed in the group. Disk 0b.29 S/N [00000000i3g268fHE60S] UID [00000000i3g268fHE60S] error: disk does not exist. in SK process config_thread on release 9.7P7 (C)
[Cluster-01: config_thread: sk.panic:alert]: params: {'reason': 'aggr aggr04: raid volfsm, fatal disk error in RAID group with no parity disk.. Raid type - raid0 Group name plex0/rg0 state NORMAL. 1 disk failed in the group. Disk 0b.29 S/N [00000000i3g268fHE60S] UID [00000000i3g268fHE60S] error: adapter error prevents command from being sent to device. in SK process config_thread on release 9.7P7 (C)'}
- Under some circumstances, the system may instead panic with a
WAFL Hungpanic:
Panic String: WAFL hung for aggr1. in SK process wafl_exempt02 on release 9.9.0 (C) - In AWS/GCP it may result in plex failure and Node may come back as "unknow" status.
SYMPFA:HA Group Notification from Node-02 (SYNCMIRROR PLEX FAILED) ALERT
- On Azure it may result in panic if disks (page blobs in case of Azure HA root/data aggregate) are not reachable.
Thu Nov 20 22:06:40 -0500 [Cluster-01: rc: sk.panic:alert]: Panic String: DIAGNOSTIC PANIC Disk deleted or missing on cloud shared HA in SK process rc on release 9.16.1P8 (C)
- Support Case might get created automatically due to
HA Group Notification (PARTNER DOWN, TAKEOVER IMPOSSIBLE ) EMERGENCYalert
Cause
In a RAID group (RAID 0) configuration with no parity disk, any single disk error is unrecoverable and the storage controller will reboot in order to attempt to clear the issue. Hence system panic is an expected behavior.
Solution
- Stop and restart OR reboot the impacted CVO instance from the cloud provider console. Post that the node impacted should be ready for Giveback. Proceed with the Give back of the impacted node.
- For the full RCA, please engage the relevant Cloud Provider (AWS/Azure/GCP) for a root cause analysis.
Internal Notes
- PLEASE DO NOT UPLOAD A CORE FILE unless it is specifically requested by an EE/Engineering.
- If losing the EBS volume results in a WAFL Hang, it could be due to bug 1366793.
- With single node if reboot does not bring back the failed disk it may need to detach and reattach the EBS volume.
- Under certain occasions if the EBS volume does not detach you may have to halt the node to detach the EBS volume. Forced detach is not recommended.
