CONTAP-766563: MetroCluster switchover fails because an SVM stays degraded after a FlexGroup create job is rolled back
Issue
- On a MetroCluster IP system, a FlexGroup volume create job can fail after some constituents have already been created (for example, when a node reaches the maximum number of volumes).
- ONTAP then rolls back those constituents by deleting them and placing them in the volume recovery queue.
- Configuration replication of those deleted volumes can fail.
- The SVM configuration state is set to degraded and does not recover on its own.
- The "metrocluster vserver resync" command does not clear the condition. Negotiated MetroCluster switchover then fails with a non-overridable veto similar to:
Failed to validate the node and cluster components before the switchover operation.
The Vserver configuration state was set to "degraded" because configuration replication failed.
Reason: MetroCluster Vserver stream operation failed.- The "metrocluster vserver show" command reports a corrective action to run "metrocluster vserver resync". That command queues the same failed replication again and does not return the SVM to a healthy configuration state.
- You might also see EMS event callhome.dr.apply.failed with a reason similar to:
File system analytics cannot be enabled when a volume is being created with the state being "offline".