AWS fatal disk error in RAID group
- Views:
- 127
- Visibility:
- Public
- Votes:
- 0
- Category:
- cloud-volumes-ontap-cvo
- Specialty:
- amazon_web_service
- Last Updated:
Applies to
- Cloud Volumes ONTAP (CVO)
- AWS
- Single node
Issue
AWS disk failed on single node cluster
Fri Dec 08 08:02:25 +0800 [Cluster01: config_thread: sk.panic:alert]: Panic String: aggr aggr1: raid volfsm, fatal disk error in RAID group with no parity disk.. Raid type - raid0 Group name plex0/rg0 state NORMAL. 1 disk failed in the group. Disk 0b.11 S/N [0000000007C65fh8/lNP] UID [0000000007C65fh8/lNP] error: aborted error disk event. in SK process config_thread on release 9.12.1P3 (C)Cause
An intermittent network error between the EC2 instance and EBS storage
Solution
- Use ASUPs to figure out which disk failed
- Use the AWS vol name to search for the disk in the AWS console
- If the disk is healthy, reattach the disk to the virtual machine
- If the disk is not healthy a new disk will need to be created and added to the virtual machine
Note: If the disk is not healthy and needs to be replaced, there is no way to recover the data for that aggregate unless there is a backup of that data available
Partner Notes
partnerNotes_text
Additional Information
- Sometimes the disk will fail in ONTAP but will still be attached to the instance in AWS and needs manual intervention to resolve
Internal Notes
internalNotes_text
