Skip to main content
NetApp Knowledge Base

SwitchIfOutErrorsWarn_Alert - AutoSupport Message

Views:
6
Visibility:
Public
Votes:
0
Category:
fabric-interconnect-and-management-switches
Specialty:
hw
Last Updated:

Applies to

  • ONTAP 9
  • Cluster Network Switches
  • Call home for Health Monitor process cshm: SwitchIfOutErrorsWarn_Alert

Event Summary

This message occurs when an error is detected during the periodic health monitoring.

  • System health monitors create alerts for potential problems detected while monitoring the subsystem.
  • The alerts contain information about probable cause along with recommended actions to rectify the problem.
  • The percentage of outbound packet errors of switch interface "Switch Name/Slot: 0 Port: 4 10G - Level" is above the warning threshold.
  • Outbound packet errors indicate that the switch interface is experiencing errors while transmitting traffic through the cluster interconnect.
  • Degradations in the cluster interconnect can result in cluster instability or potentially outages.

Validate

AutoSupport Message

HA Group Notification from Node Name (Health Monitor process cshm: SwitchIfOutErrorsWarn_Alert[Node Name/Slot: 0 Port: 4 10G - Level]) ALERT

Event Log

event log show -severity * -message-name callhome*

[Node Name Name: mgwd: callhome.hm.alert.major:alert]: Call home for Health Monitor process cshm: SwitchIfOutErrorsWarn_Alert[Node Name/Slot: 0 Port: 4 10G - Level].
Command Line

system health alert show -node <node name> -monitor cluster-switch -alert-id SwitchIfOutErrorsWarn_Alert

                  Node: NetApp-a
                   Monitor: cluster-switch
            Class of Alert: SwitchIfOutErrorsWarn_Alert
         Severity of Alert: Major
            Probable Cause: Threshold_crossed
Probable Cause Description: The percentage of outbound packet errors of switch interface "$(cluster_switch_analytics.unique-name)" is above the warning threshold.
           Possible Effect: Communication between nodes in the cluster might be degraded.
        Corrective Actions: 1) Migrate any cluster LIF that uses this connection to another port connected to a cluster switch.
For example, if cluster LIF "clus1" is on port e0a and the other LIF is on e0b,
run the following command to move "clus1" to e0b:
"network interface migrate -vserver vs1 -lif clus1 -sourcenode node1 -destnode node1 -dest-port e0b"
2) Replace the network cable with a known-good cable.
If errors are corrected, stop. No further action is required.
Otherwise, continue to Step 3.
3) Move the network cable to another port on the node (if available).
Migrate the cluster LIF to the new port.
If errors are corrected, contact technical support to troubleshoot the original node port.
Otherwise, continue to Step 4.
4) Move the network cable to another available cluster switch port.
Migrate the cluster LIF back to the original port.

Resolution

  1. Perform link troubleshooting on the port reporting SwitchIfOutErrorsWarn_Alert.
  2. Attempt reseat of cable and SFP/transceiver.
  3. Check for patch panels or intermediate connections and bypass if possible.
  4. Replace the cable and/or SFP with known-good components.
  5. Node Side - Migrate the affected cluster LIF to a healthy cluster port and move the connection to an alternate node cluster port (if available).
    • If errors stop, investigate the original node port. (Contact NetApp Technical Support or log into the NetApp Support Site to create a case. Reference this article for further assistance.)
  6. Switch Side - Move the connection to a known-good switch port.
    • If errors stop, investigate the original switch port, collect switch logs and engage Switch support 
      • For a Broadcom switch, contact Broadcom for assistance
      • For a Cisco switch, contact Cisco for assistance
      • For a NVIDIA switch, contact NVIDIA for assistance

    Additional Information

     

    NetApp provides no representations or warranties regarding the accuracy or reliability or serviceability of any information or recommendations provided in this publication or with respect to any results that may be obtained by the use of the information or observance of any recommendations provided herein. The information in this document is distributed AS IS and the use of this information or the implementation of any recommendations or techniques herein is a customer's responsibility and depends on the customer's ability to evaluate and integrate them into the customer's operational environment. This document and the information contained herein may be used solely in connection with the NetApp products discussed in this document.