ONTAP Select Cluster Unresponsive Due to Severe Backend Disk Latency
Applies to
- ONTAP Select High Avalability cluster
- Cluster hosted on VMware ESXi
Issue
ONTAP Select cluster becomes unavailable. Nodes and cluster management LIF are not reachable by SSH or System Manager, virtual machines (VMs) are slow or unavailable.
- Symptoms can include:
- Node fails to boot and panic due to latency.
- Partner node still serves data, but with severe latency.
- System Manager becomes unavailable when the failed node tries to come online.
- Example EMS messages:
[kernel:device.latency.threshold:notice]: The device latency in microseconds exceeded the threshold value for device "isp18" and the associated target device "/dev/da4". The threshold is "25000", execution latency is "5555955"