NFSv4.1 Latency During Cluster Reboot on CVO
Applies to
- Cloud Volumes ONTAP (CVO)
- NFSv4.1 protocol environments
Issue
CVO planned cluster reboots with NFSv4.1 clients experience latency or timeouts (up to 2 minutes in rare cases, but typically around 45 seconds). This is observed as increased delay in protocol takeover and LIF movement, impacting business workloads such as SAP.
Cause
- The latency is caused by the NFSv4.1 grace period and protocol negotiation during cluster or node reboot/failover.
- Delays may be exacerbated by high workload, aggregate fill, or ongoing root/boot snapshots.
- This is expected behavior for NFSv4.1 clients during failover, as per protocol standards and ONTAP design.
- A one-time 2-minute delay may occur during instance type changes or specific environmental triggers (e.g., GCP outage or plex resync).
Solution
- Upgrade ONTAP to the latest recommended version to benefit from protocol improvements and bug fixes for NFSv4.1 latency.
- Confirm that the observed NFSv4.1 delay during planned takeover/giveback is within protocol and ONTAP tolerance.
- Monitor latency using
qos statistics volume latency showand cluster metrics before and after reboot. - For critical workloads, schedule reboots during off-peak hours and inform application owners of expected NFSv4.1 grace period delays.
Partner Notes
partnerNotes_text
Additional Information
additionalInformation_text
Internal Notes
internalNotes_text
